Setting device, method, and program

By employing an open-ear device and acoustic signal analysis, the technique addresses the inaccuracies of conventional HRTF optimization methods, providing a precise and user-specific HRTF selection for enhanced sound localization.

WO2025169295A1PCT designated stage Publication Date: 2025-08-14NT T INC
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/003892
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-02-06
Publication Date
2025-08-14

AI Technical Summary

Technical Problem

Conventional methods for personal optimization of head-related transfer functions (HRTFs) are cumbersome and prone to deviations when sound is produced, as they rely on external head shape measurements, which may not accurately reflect individual ear characteristics.

Method used

A technique that uses an open-ear acoustic signal output device and a sound collecting unit positioned away from the ear canal to measure and select an optimized HRTF based on the similarity of acoustic characteristics, utilizing spectral distances and principal components to match the user's head-related transfer function.

Benefits of technology

Enables easy and accurate personal optimization of HRTFs, improving the accuracy of sound localization and direction perception by aligning the selected HRTF with the user's individual ear shape and characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024003892_14082025_PF_FP_ABST
    Figure JP2024003892_14082025_PF_FP_ABST
Patent Text Reader

Abstract

According to the present invention, an acoustic signal is emitted from a sound emission unit that does not seal the external auditory canal of a user, and an observation signal based on the acoustic signal is collected by a sound collection unit disposed at the position of a head that is neither the external auditory canal nor an open end of the external auditory canal. The observation signal is used for setting a head-related transfer function of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Setting device, method and program

[0001] The present invention relates to a technique for setting an optimum head-related transfer function for a user.

[0002] To achieve stereophonic sound, the head-related transfer function (HRTF), which is the characteristic of how sound travels from a sound source to both ears, is essential. The head-related transfer function is highly dependent on individual factors such as ear size and shape, and in recent years, personalized optimization services for head-related transfer functions have been launched. Previous research suggesting a relationship between the HeSTF (Hearable Speaker Transfer Function), which is the transfer characteristic from an audio signal output device to the ear, and pinna shape is known from the study described in Non-Patent Document 1.

[0003] Yuki Watanabe and five others, "Individual Analysis of Transfer Characteristics Based on Pinna Shape in Open-Ear Earphones," 150th Annual Meeting of the Acoustical Society of Japan (Autumn 2023)

[0004] Conventional personal optimization of head-related transfer functions is based on the external shape of the head obtained by image capture, LiDAR (Light Detection and Ranging) measurements, etc. Therefore, measurements for personal optimization can be cumbersome, and the personally optimized head-related transfer function may deviate from the characteristics when sound is actually produced.

[0005] The present invention provides a technique for easily and accurately optimizing a head-related transfer function for individual use.

[0006] An acoustic signal is emitted from a sound emitting unit that does not seal the ear canal of the user, and an observation signal based on the acoustic signal is collected by a sound collecting unit located at a position on the head that is neither the ear canal nor the open end of the ear canal. Multiple pairs of head-related transfer function candidates and acoustic characteristics of assumed observation signals assumed for the head-related transfer function candidates are associated with each other, and a head-related transfer function of the user is selected from the head-related transfer function candidates based on the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the assumed observation signal. The similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the expected observed signal is a value based on at least one of: (1) the logarithmic spectral distance between the spectral envelope of the observed signal and the spectral envelope of the expected observed signal; (2) the distance between a vector representing the principal components of the acoustic characteristics of the observed signal and a vector representing the principal components of the acoustic characteristics of the expected observed signal; and (3) a distance measure whose value increases as the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the expected observed signal increases, and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the expected observed signal increases. N11 is the notch in the spectral envelope of the observed signal that is located at the lowest frequency in the specified frequency range, N12 is the notch in the spectral envelope of the observed signal that is located next to N11 at the lowest frequency in the specified frequency range, N21 is the notch in the spectral envelope of the assumed observed signal that is located at the lowest frequency in the specified frequency range, and N22 is the notch in the spectral envelope of the assumed observed signal that is located next to N21 at the lowest frequency in the specified frequency range.

[0007] This allows for easy and accurate personal optimization of head-related transfer functions.

[0008] FIG. 1 is a diagram illustrating the configuration of a setting device according to an embodiment. FIG. 2A is a rear view of the acoustic signal output device of FIG. 1. FIG. 2B is a right side view of the acoustic signal output device of FIG. 2A. FIG. 3 is a diagram illustrating the configuration and usage of a setting device according to an embodiment. FIG. 4 is a diagram illustrating the configuration and usage of an evaluation device according to an embodiment. FIG. 5 is a diagram illustrating the configuration and usage of an evaluation device according to an embodiment. FIGS. 6A and 6B are diagrams for explaining experimental conditions for subjective evaluation. FIG. 7A is a diagram modeling the directional perception of normal sound. FIG. 7B is a diagram modeling the directional perception of sound reproduced by the device. FIG. 7C is a diagram modeling the directional perception when the component corresponding to HeSTF is corrected from the sound reproduced by the device. FIG. 8 is a graph illustrating the relationship between the acoustic characteristics of an observation signal observed by a sound collector placed at the open end of the ear canal and the acoustic characteristics of an observation signal observed by a sound collector placed behind the pinna. 9A is a graph illustrating the relationship between the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal and the acoustic characteristics of an assumed observation signal corresponding to a head-related transfer function selected from candidate head-related transfer functions based on the Euclidean distance between the acoustic characteristics of the observation signal and the acoustic characteristics of an assumed observation signal corresponding to the candidate head-related transfer function. 9B is a graph illustrating the relationship between the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal and the acoustic characteristics of an assumed observation signal corresponding to a head-related transfer function selected from candidate head-related transfer functions based on the cosine similarity between the acoustic characteristics of the observation signal and the acoustic characteristics of an assumed observation signal corresponding to the candidate head-related transfer function. 9C is a graph illustrating the relationship between the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal and the acoustic characteristics of an assumed observation signal corresponding to a head-related transfer function selected from candidate head-related transfer functions based on the spectral distortion of the acoustic characteristics of the assumed observation signal corresponding to the candidate head-related transfer function relative to the acoustic characteristics of the observation signal.Fig. 10A is a bubble chart illustrating subjective evaluation results when the azimuth angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics highly similar to the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal. Fig. 10B is a bubble chart illustrating subjective evaluation results when the azimuth angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics less similar to the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal. Fig. 10C is a bubble chart illustrating subjective evaluation results when the elevation angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics highly similar to the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal. Fig. 10D is a bubble chart illustrating subjective evaluation results when the elevation angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics that are low in similarity to the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal. Fig. 11A is a bubble chart illustrating subjective evaluation results when the azimuth angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics that are high in similarity to the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal. Fig. 11B is a bubble chart illustrating subjective evaluation results when the azimuth angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics that are low in similarity to the acoustic characteristics of an observation signal observed by a sound collector located at a position on the head that is neither the ear canal nor the open end of the ear canal. Figure 11C is a bubble chart illustrating the subjective evaluation results when the elevation angle of the presentation angle is randomly changed based on the head-related transfer function corresponding to an assumed observed signal having acoustic characteristics that are highly similar to the acoustic characteristics of an observed signal observed with a sound collection unit placed at a position on the head that is neither the ear canal nor the open end of the ear canal.Fig. 11D is a bubble chart illustrating subjective evaluation results when the elevation angle of the presentation angle is randomly changed based on a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics that are low in similarity to the acoustic characteristics of an observation signal observed by a sound collection unit placed at a position on the head that is neither the ear canal nor the open end of the ear canal. Fig. 12 is a diagram illustrating the configuration and usage state of a setting device of an embodiment. Figs. 13A and 13B are diagrams for explaining a modified example of an acoustic signal output device. Fig. 14 is a diagram for explaining a modified example of an acoustic signal output device. Fig. 15 is a block diagram illustrating the hardware configuration of a head-related transfer function setting device and an evaluation processing device.

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings. [First Embodiment] First, a first embodiment of the present invention will be described. <Configuration of Setting Device 1> As illustrated in Figures 1 to 3, the setting device 1 of this embodiment has an acoustic signal output device 11-1 and a head-related transfer function setting device 12.

[0010] The acoustic signal output device 11-1 of this embodiment includes a sound emitting unit 111-1 configured to emit an acoustic signal without sealing the ear canal 1011-1 of the user 1000, a sound collecting unit 113-1 that is placed at a position on the head that is neither the ear canal 1011-1 nor the open end (entrance) of the ear canal 1011-1 and is configured to collect an observation signal based on the acoustic signal, and a wearing unit 112-1 to which the sound emitting unit 111-1 and the sound collecting unit 113-1 are fixed and that is worn on the auricle 1012-1 of one ear 1010-1 (for example, the right ear) of the user 1000. A specific example of the sound emitting unit 111 is an open-ear earphone, and a specific example of the sound collecting unit 113-1 is a microphone.

[0011] In this embodiment, the sound emitting unit 111-1 houses a driver unit (not shown) that emits an acoustic signal based on an input playback signal, and a sound hole 111a-1 that emits this acoustic signal to the outside is provided on one surface 111b-1 of the sound emitting unit 111-1. Meanwhile, no sound hole is provided on the other (opposite) surface 111c-1 of the sound emitting unit 111-1. However, this does not limit the present invention, and a sound hole may be provided in a location other than the surface 111b-1 of the sound emitting unit 111-1. The attachment unit 112-1 in this embodiment is substantially C-shaped, with the sound emitting unit 111-1 attached to one end. Furthermore, a sound collecting unit 113-1 is attached between the one end of the attachment unit 112-1 to which the sound emitting unit 111-1 is attached and the other end to which the sound emitting unit 111-1 is not attached.

[0012] 3, the acoustic signal output device 11-1 is worn by the user 1000 by hooking the attachment unit 112-1 onto the auricle 112-1. As a result, the sound collection unit 113-1 and the sound emitting unit 111-1 are worn on the ear 1010-1 of the user 1000. That is, the sound collection unit 113-1 is disposed on the back side of the auricle 112-1 (the side where the ear canal 1011-1 does not exist), and the sound emitting unit 111-1 is disposed on the front side of the auricle 1012-1 (the side where the ear canal 1011-1 exists). In this case, for example, the surface 111b of the sound emitting unit 111-1 faces the user 1000, and the sound emitting unit 111-1 is disposed without sealing the ear canal 1011-1. For example, the sound hole 111a-1 is disposed at a position away from the open end of the ear canal 1011-1. The sound collection unit 113-1 may be located anywhere behind the auricle 112-1. However, it is preferable that the sound collection unit 113-1 be located above the ear canal 1011-1 on the rear side of the auricle 112-1. For example, it is preferable that the sound collection unit 113-1 be located in a region above the ear canal 1011-1 on the rear side of the auricle 112-1. This is because if the sound collection unit 113-1 is covered too much by the auricle 112-1, it may be impossible to measure the characteristics of the observation signal based on the acoustic signal emitted from the sound emission unit 111-1. The sound collection unit 113-1 may be located on the rear side of the auricle 112-1 (for example, behind the helix), or may be located at a position on the head behind the auricle 112-1. For example, it is preferable that the sound collection unit 113-1 be located in any position within the region 1001 illustrated in FIG. 3.

[0013] The head-related transfer function setting device 12 has a storage unit 121, a playback unit 122, and a setting processing unit 123. The storage unit 121 stores acoustic signal information A for identifying an acoustic signal to be emitted from the sound emitting unit 111-1, candidate head-related transfer functions HRTF-1, ..., HRTF-N, and a correspondence table T that associates the candidate head-related transfer functions HRTF-1, ..., HRTF-N with acoustic characteristics AOS-1, ..., AOS-N of an assumed observation signal, respectively. Here, the acoustic characteristic AOS-n (where n = 1, ..., N) of the assumed observation signal is the acoustic characteristic of the observation signal (the observation signal based on the acoustic signal) observed by the sound collecting unit 113-1 when the head-related transfer function of the user 1000 is HRTF-n and an acoustic signal identified by the acoustic signal information A is emitted from the sound emitting unit 111-1 of the acoustic signal output device 11-1 worn by the user 1000 as described above, or an acoustic characteristic similar thereto. For example, when the head-related transfer function of the user 1000 is HRTF-n, the acoustic characteristic AOS-n is the acoustic characteristic of the observed signal observed at the position of the sound collection unit 113-1 or a position approximate thereto when an acoustic signal identified by the acoustic signal information A is emitted from the position of the sound emitting unit 111-1 of the acoustic signal output device 11-1 worn by the user 1000 as described above or a position approximate thereto, or an acoustic characteristic approximate thereto. The acoustic characteristic AOS-n may be obtained by actually measuring the observed signal or may be obtained by simulation. The acoustic signal information A is, for example, information for perceiving an acoustic signal emitted from a sound source (virtual sound source) localized at a specific position (virtual position). The virtual position is, for example, a position at an elevation angle of 10 to 20 degrees upward relative to the horizontal direction in front of the user 1000 wearing the acoustic signal output device 11-1, but may be another position. The acoustic signal may be, for example, a Time Stretched Pulse (TSP) signal, but may also be a signal representing other sounds such as speech or music. The acoustic characteristic may be, for example, a frequency spectrum, but may also be other acoustic characteristics. Here, N is an integer of 2 or greater.That is, in correspondence table T, multiple pairs of candidate head-related transfer functions HRTF-1, ..., HRTF-N are associated with acoustic characteristics AOS-1, ..., AOS-N of assumed observation signals assumed for the candidate head-related transfer functions HRTF-1, ..., HRTF-N. An example of correspondence table T is shown below. For the sake of simplicity, this embodiment illustrates an example in which one piece of acoustic signal information A, corresponding candidate head-related transfer functions HRTF-1, ..., HRTF-N, and acoustic characteristics AOS-1, ..., AOS-N of the assumed observation signal are associated in the correspondence table T. However, this does not limit the present invention. The correspondence table T may also associate multiple types of acoustic signal information A, corresponding candidate head-related transfer functions HRTF-1, ..., HRTF-N, and acoustic characteristics AOS-1, ..., AOS-N of the assumed observation signal. The multiple types of acoustic signal information A may, for example, have different virtual positions or different sound types. Furthermore, a specific piece of acoustic signal information A may be selected from multiple types of acoustic signal information A depending on the environment.

[0014] The playback unit 122 is electrically connected to the storage unit 121 and the sound emission unit 111-1. The setting processing unit 123 is electrically connected to the storage unit 121 and the sound collection unit 113-1. However, this does not limit the present invention, and at least some of these may be configured to be able to communicate wirelessly or via a network.

[0015] <Configuration of Evaluation Device 100> As illustrated in FIGS. 4 and 5, the evaluation device 100 of this embodiment has sound signal output devices 11-1 and 11-2 and an evaluation processing device 13.

[0016] The acoustic signal output device 11-i (where i = 1, 2) of this embodiment includes a sound emitting unit 111-i configured to emit an acoustic signal without sealing the ear canal 1011-i of the user 1000, and a mounting unit 112-i to which the sound emitting unit 111-i is fixed and which is mounted on the auricle 1012-i of the ear 1010-i of the user 1000. The sound emitting unit 111-i of this embodiment houses a driver unit (not shown) that emits an acoustic signal based on an input playback signal, and a sound hole 111a-i is provided on one surface 111b-i of the sound emitting unit 111-i to emit this acoustic signal to the outside. Meanwhile, no sound hole is provided on the other (opposite) surface 111c-i of the sound emitting unit 111-i. However, this does not limit the present invention, and a sound hole may be provided in a location other than the surface 111b-i of the sound emitting unit 111-i. The wearing unit 112-i in this embodiment is substantially C-shaped, and the sound emitting unit 111-i is attached to one end thereof. A specific example of the sound emitting unit 111-i is an open-ear earphone. The acoustic signal output device 11-1 may be the same as or different from the acoustic signal output device 11-1 of the setting device 1. Furthermore, the acoustic signal output device 11-2 may or may not be configured symmetrically with the acoustic signal output device 11-1 of the setting device 1.

[0017] As illustrated in Figures 4 and 5, the acoustic signal output device 11-i of this embodiment is worn by the user 1000 by hooking the wearing unit 112-i onto the auricle 112-i. As a result, the sound collection unit 113-1 is worn on one ear 1010-1 (e.g., the right ear) of the user 1000, and the sound collection unit 113-2 is worn on the other ear 1010-2 (e.g., the left ear) of the user 1000. The sound emitting unit 111-i of this embodiment is disposed on the front side of the auricle 1012-i. In this case, for example, the surface 111b-i of the sound emitting unit 111-i faces the user 1000, and the sound emitting unit 111-i is disposed without sealing the ear canal 1011-i. For example, the sound hole 111a-i is disposed at a position away from the open end of the ear canal 1011-i.

[0018] 4 and 5, the evaluation processing device 13 of this embodiment has a storage unit 131, a playback unit 132, and an input unit 133. The playback unit 132 is electrically connected to the storage unit 131 and the sound emitting units 111-1 and 111-2. The input unit 133 is electrically connected to the storage unit 131. However, this does not limit the present invention, and at least some of these may be configured to be able to communicate wirelessly or via a network.

[0019] <Individual Optimization of Head-Related Transfer Functions> Next, we will explain the individual optimization of head-related transfer functions for the user 1000 in this embodiment. The individual optimization of head-related transfer functions is performed using the setting device 1. First, the user 1000 wears the acoustic signal output device 11-1 of the setting device 1 on the pinna 1012-1 of one ear 1010-1 as described above (FIG. 3).

[0020] Next, the playback unit 122 of the head-related transfer function setting device 12 (FIGS. 1 to 3) extracts acoustic signal information A from the correspondence table T stored in the storage unit 121 and sends it to the playback unit 122. The playback unit 122 outputs a playback signal for outputting an acoustic signal identified by the acoustic signal information A. The playback signal is input to the driver unit of the sound emitting unit 111-1, and the driver unit outputs an acoustic signal based on the input playback signal. This acoustic signal is emitted to the outside from the sound hole 111a-1. As described above, the sound emitting unit 111-1 does not seal the ear canal 1011-1 of the user 1000, and emits this acoustic signal to the outside. A portion of the acoustic signal emitted from the sound emitting unit 111-1 reaches the sound collecting unit 113-1. The sound collecting unit 113-1 collects an observation signal based on this acoustic signal and sends information representing the observation signal to the setting processing unit 123. The setting processing unit 123 sets the head-related transfer function F of the user 1000 based on this observation signal.

[0021] The setting processing unit 123 of this embodiment refers to the correspondence table T stored in the storage unit 121 and selects a head-related transfer function F∈{HRTF-1, ..., HRTF-N} of the user 1000 from the candidate head-related transfer functions HRTF-1, ..., HRTF-N based on the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics AOS-1, ..., AOS-N of the assumed observed signal. For example, the setting processing unit 123 sets the candidate head-related transfer function HRTF-n associated with the acoustic characteristic AOS-n (here, n∈{1, ..., N}) of the assumed observed signal that has the highest similarity to the acoustic characteristics of the observed signal as the head-related transfer function F=HRTF-n of the user 1000. If a larger similarity value indicates higher similarity, the similarity is highest when the similarity value is greatest. On the other hand, if a smaller similarity value indicates higher similarity, the similarity is highest when the similarity value is smallest. Alternatively, the setting processing unit 123 may set a candidate head-related transfer function HRTF-n associated with the acoustic characteristics AOS-n of the assumed observation signal whose similarity to the acoustic characteristics of the observed signal exceeds a threshold as the head-related transfer function F=HRTF-n of the user 1000. Alternatively, the setting processing unit 123 may set one of the candidate head-related transfer functions HRTF-n associated with the acoustic characteristics AOS-n of the assumed observation signal whose similarity to the acoustic characteristics of the observed signal exceeds a threshold, which satisfies other conditions, as the head-related transfer function F=HRTF-n of the user 1000. Alternatively, the setting processing unit 123 may set a function value (e.g., an average value) of multiple candidate head-related transfer functions HRTF-n associated with the acoustic characteristics AOS-n of the assumed observation signal whose similarity to the acoustic characteristics of the observed signal exceeds a threshold as the head-related transfer function F=HRTF-n of the user 1000.

[0022] The similarity used to set the head-related transfer function F may be, for example, the smallness of the Euclidean distance between the features of the acoustic characteristics of the observed signal and the features of the acoustic characteristics AOS-n of the assumed observed signal, the magnitude of the cosine similarity between the features of the acoustic characteristics of the observed signal and the features of the acoustic characteristics AOS-n of the assumed observed signal, the smallness of the spectral distortion of the acoustic characteristics AOS-n of the assumed observed signal relative to the acoustic characteristics of the observed signal, or some other measure.

[0023] For example, it is desirable to select the head-related transfer function F based on the similarity between feature amounts corresponding to the notches and peaks in the spectral envelope of the observed signal and feature amounts corresponding to the notches and peaks in the spectral envelope of the assumed observed signal. This is because the notches and peaks in the spectral envelope tend to reflect features of the shape of the head, including the ears 1010-1 of the user 1000, and using these features is expected to enable appropriate individual optimization of the head-related transfer function F. For example, the feature amounts corresponding to the notches and peaks in the spectral envelope are pairs of the frequencies and magnitudes (e.g., absolute values ​​of power and amplitude) of the notches and peaks in the spectral envelope. For example, the setting processing unit 123 may determine the similarity between a vector whose elements are the frequencies and magnitudes of the notches and peaks in the spectral envelope of the observed signal and a vector whose elements are the frequencies and magnitudes of the notches and peaks in the spectral envelope of the assumed observed signal, and select the head-related transfer function F based on the similarity.

[0024] It is also preferable that the notch and peak of the spectral envelope of the observed signal are [N11, P11], and the notch and peak of the spectral envelope of the assumed observed signal are [N21, P21]. Similarly, it is preferable that the notch and peak of the spectral envelope of the observed signal are [N11, P11, N12], and the notch and peak of the spectral envelope of the assumed observed signal are [N21, P21, N22]. For example, it is preferable that the setting processing unit 123 determines the similarity between a vector having the frequency and magnitude of [N11, P11] of the spectral envelope of the observed signal as elements and a vector having the frequency and magnitude of [N21, P21] of the spectral envelope of the assumed observed signal as elements, and selects the head-related transfer function F based on the similarity. For example, it is desirable that the setting processing unit 123 determines the similarity between a vector whose elements are the frequency and magnitude of [N11, P11, N12] of the spectral envelope of the observed signal and a vector whose elements are the frequency and magnitude of [N21, P21, N22] of the spectral envelope of the assumed observed signal, and selects the head-related transfer function F based on the similarity. Here, N11 is a frequency within a predetermined frequency range F range is the notch in the spectrum envelope of the observed signal that exists on the lowest frequency side in the predetermined frequency range Frange P11 is a notch in the spectrum envelope of the observed signal that exists next to N11 on the low frequency side in the predetermined frequency range F range is the peak of the spectrum envelope of the observed signal that exists at the lowest frequency side in the predetermined frequency range F range is the notch in the spectrum envelope of the assumed observed signal that exists at the lowest frequency in the predetermined frequency range F range P21 is a notch in the spectral envelope of the assumed observed signal that exists next to N21 on the low frequency side in the predetermined frequency range F range In particular, these peaks and notches are likely to represent the characteristics of the shape of the head of the user 1000, including the ear 1010-1, and by using these, the head-related transfer function F can be appropriately optimized individually.

[0025] Furthermore, in the frequency band of 5 to 13 kHz of the spectral envelope of the observed signal, characteristics of the shape of the head, including the ear 1010-1 of the user 1000, are particularly likely to appear, and by using information of this frequency band, it is possible to appropriately individually optimize the head-related transfer function F. Therefore, it is preferable that the setting processing unit 123 selects the head-related transfer function F of the user 1000 from the candidate head-related transfer functions HRTF-1, ..., HRTF-N based on the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics AOS-1, ..., AOS-N of the assumed observed signal in the frequency band including 5 to 13 kHz. For example, it is preferable that the setting processing unit 123 sets the candidate head-related transfer function HRTF-n associated with the acoustic characteristic AOS-n of the assumed observed signal that has the highest similarity to the acoustic characteristics of the observed signal in the frequency band including 5 to 13 kHz as the head-related transfer function F=HRTF-n of the user 1000. An example of a frequency band including 5 to 13 kHz is the frequency band of 5 to 13 kHz. For example, in the above-mentioned frequency range F range is preferably 5 to 13 kHz. rangeFor the same reason, if the smallness of spectral distortion is used as the degree of similarity, it is desirable to select the head-related transfer function F by using the smallness of spectral distortion when the band is limited to 5 to 13 kHz as the degree of similarity.

[0026] The similarity used to set the head-related transfer function F, in other words, the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the assumed observed signal, may be (1) the logarithmic spectral distance between the spectral envelope of the observed signal and the spectral envelope of the assumed observed signal, or (2) the distance between a vector representing the principal components of the acoustic characteristics of the observed signal and a vector representing the principal components of the acoustic characteristics of the assumed observed signal, or (3) a distance measure whose value increases as the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal increases.

[0027] An example of the log-spectral distance between the spectral envelope of the observed signal and the spectral envelope of the assumed observed signal is the value of LSD, defined by the following equation: ^p b,ch,l is the value of the spectral envelope of the assumed observed signal frequency bin l∈{1,…,L} corresponding to ch∈{left, right} when an acoustic signal is output from sound source position b∈{1,…,B}. p b,ch,l,n is the value of the spectral envelope of frequency bin l∈{1,…,L} of the n∈{1,…,N}th observed signal corresponding to ch∈{left, right} when an acoustic signal is output from sound source position b∈{1,…,B}. B is the number of sound source positions. L is the number of frequency bins. N is the number of observed signals.

[0028] Note that the assumed observed signal corresponding to left is the assumed observed signal corresponding to the left ear. The assumed observed signal corresponding to right is the assumed observed signal corresponding to the right ear. The observed signal corresponding to left is the observed signal corresponding to the left ear. The observed signal corresponding to right is the observed signal corresponding to the right ear.

[0029] As in this example, the acoustic characteristics of a plurality of assumed observed signals are stored as the acoustic characteristics AOS-n of the assumed observed signal, in other words, as the acoustic characteristics of the assumed observed signals of the same person, and the similarity may be calculated using the acoustic characteristics of these plurality of assumed observed signals.Furthermore, as in this example, the similarity may be calculated using the feature quantities of the acoustic characteristics of the plurality of observed signals.

[0030] A "vector representing the principal components of the acoustic characteristics of the observed signal" is, for example, a vector of a predetermined dimension or less obtained from a predetermined range of the spectrum of the observed signal and having a predetermined contribution rate or more. Similarly, a "vector representing the principal components of the acoustic characteristics of the assumed observed signal" is, for example, a vector of a predetermined dimension or less obtained from a predetermined range of the spectrum of the assumed observed signal and having a predetermined contribution rate or more. An example of the predetermined range is 188 Hz to 16 kHz. An example of the predetermined contribution rate is 90%. The "vector representing the principal components of the acoustic characteristics of the observed signal" and the "vector representing the principal components of the acoustic characteristics of the assumed observed signal" can be obtained using, for example, a PCA (Principal Component Analysis) technique. Any distance measure that satisfies the distance axioms can be used as the distance between the vector representing the principal components of the acoustic characteristics of the observed signal and the vector representing the principal components of the acoustic characteristics of the assumed observed signal. Examples of such distance measures include Euclidean distance and Manhattan distance.

[0031] An example of a distance measure that has a larger value as the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal are larger is the NFD defined by the following equation. α and β are predetermined positive real numbers. For example, α=β=1. NFD1 and NFD2 are values ​​defined by the following equations, for example. Let Y be an arbitrary number, and |Y| means the absolute value of Y. f N1 (HeSTF j ) denotes the frequency of the notch N11 in the spectral envelope of the observed signal. N1 (HeSTF k) denotes the frequency of the notch N21 in the spectral envelope of the assumed observed signal. N2 (HeSTF j ) denotes the frequency of the notch N12 in the spectral envelope of the observed signal. N2 (HeSTF k ) denotes the frequency of the notch N22 in the spectral envelope of the assumed observed signal.

[0032] Note that NFD1 and NFD2 may be values ​​defined by the following formulas, for example. f N1 (HeSTF j,b ) denotes the frequency of the notch N11 in the spectral envelope of the observed signal when an acoustic signal is output from a sound source position b∈{1,...,B}. N1 (HeSTF k,b ) denotes the frequency of the notch N21 in the spectral envelope of the assumed observed signal when an acoustic signal is output from the sound source position b∈{1,...,B}. N2 (HeSTF j,b ) denotes the frequency of the notch N12 in the spectral envelope of the observed signal when an acoustic signal is output from the sound source position b∈{1,...,B}. N2 (HeSTF k,b ) denotes the frequency of the notch N22 in the spectral envelope of the assumed observed signal when an acoustic signal is output from a sound source position b∈{1,...,B}. B is the number of sound source positions.

[0033] As in this example, the acoustic characteristics of a plurality of assumed observed signals are stored as the acoustic characteristics AOS-n of the assumed observed signal, in other words, as the acoustic characteristics of the assumed observed signals of the same person, and the similarity may be calculated using the acoustic characteristics of these plurality of assumed observed signals.Furthermore, as in this example, the similarity may be calculated using the feature quantities of the acoustic characteristics of the plurality of observed signals.

[0034] The difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal is defined as the "first difference," the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal is defined as the "second difference," and the value of the distance measure that has a larger value the greater the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal, and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal, is defined as the "value of distance measure (3)."

[0035] In this case, as in this example, the value of the distance measure (3) may be calculated from the average value of multiple first differences corresponding to multiple different sound source positions and the average value of multiple second differences corresponding to multiple different sound source positions.

[0036] In the following explanation, (1) the logarithmic spectral distance between the spectral envelope of the observed signal and the spectral envelope of the assumed observed signal is referred to as the “distance (1),” (2) the distance between the vector representing the principal components of the acoustic characteristics of the observed signal and the vector representing the principal components of the acoustic characteristics of the assumed observed signal is referred to as the “distance (2),” and (3) the value of the distance measure that increases as the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal are increased is referred to as the “distance (3).”

[0037] The similarity used to set the head-related transfer function F, in other words, the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the assumed observed signal, may be a value based on at least one of the distances (1), (2), and (3).

[0038] An example of a value based on at least one of the distances (1), (2), and (3) is an output value when at least one of the distances (1), (2), and (3) is input to a non-decreasing function. When X is an arbitrary number, a "value based on X" may be X itself or a value calculated based on X.

[0039] The distance (1) above may be the sum of the logarithmic spectral distance between the spectral envelope of the observed signal corresponding to the left ear and the spectral envelope of the assumed observed signal corresponding to the left ear, and the logarithmic spectral distance between the spectral envelope of the observed signal corresponding to the right ear and the spectral envelope of the assumed observed signal corresponding to the right ear.

[0040] Furthermore, the distance (2) may be the sum of the distance between the vector representing the principal component of the acoustic characteristics of the observed signal corresponding to the left ear and the vector representing the principal component of the acoustic characteristics of the assumed observed signal corresponding to the left ear, and the distance between the vector representing the principal component of the acoustic characteristics of the observed signal corresponding to the right ear and the vector representing the principal component of the acoustic characteristics of the assumed observed signal corresponding to the right ear.

[0041] Alternatively, the logarithmic spectral distance between the spectral envelope of the observed signal and the spectral envelope of the assumed observed signal may be calculated based on the spectral envelopes of k different observed signals, and the average of the k calculated logarithmic spectral distances may be used as the distance in (1) above. k is a predetermined positive integer. For example, k=5.

[0042] Alternatively, the distance between the vector representing the principal components of the acoustic characteristics of the observed signal and the vector representing the principal components of the acoustic characteristics of the assumed observed signal may be calculated based on the spectral envelopes of each of k different observed signals, and the average of the k calculated distances may be used as the distance in (2) above.

[0043] Furthermore, when there are multiple observed signals, the frequency of notch N11 in the average spectrum of the spectral envelopes of the multiple observed signals may be defined as "notch N11 in the spectral envelope of the observed signals," and the frequency of notch N12 in the average spectrum of the spectral envelopes of the multiple observed signals may be defined as "notch N12 in the spectral envelope of the observed signals," and the distance in (3) above may be calculated. When there are multiple assumed observed signals, the frequency of notch N21 in the average spectrum of the spectral envelopes of the multiple assumed observed signals may be defined as "the frequency of notch N21 in the spectral envelope of the assumed observed signals," and the frequency of notch N22 in the spectral envelopes of the multiple assumed observed signals may be defined as "the frequency of notch N22 in the spectral envelope of the assumed observed signals," and the distance in (3) above may be calculated.

[0044] <Subjective Evaluation of Head-Related Transfer Function> Next, the subjective evaluation of the head-related transfer function F set for the user 1000 will be described. The subjective evaluation of the head-related transfer function F in this embodiment is performed using the evaluation device 100. First, the head-related transfer function F set for the user 1000, who is the subject, is stored in the storage unit 131 of the evaluation device 100. Next, the user 1000 wears the acoustic signal output device 11-1 of the evaluation device 100 on the pinna 1012-1 of one ear 1010-1 (for example, the right ear) ( FIG. 4 ), and wears the acoustic signal output device 11-2 of the evaluation device 100 on the pinna 1012-2 of the other ear 1010-2 (for example, the left ear) ( FIG. 5 ).

[0045] Next, the reproduction unit 132 of the evaluation processing device 13 of the evaluation device 100 (FIGS. 4 and 5) extracts the head-related transfer function F set for the user 1000, who is the subject, from the storage unit 131, and generates a reproduction signal X for emitting an acoustic signal at an arbitrary (for example, random) localization angle (direction of a localized virtual sound source, hereinafter referred to as "presentation angle") θ based on the head-related transfer function F using a known stereophonic technique. F -1, X F -2 is generated and output. Here, the playback signal X F -1 is supplied to the acoustic signal output device 11-1, and the playback signal X F The audio signal θ-2 is supplied to the audio signal output device 11-2. This audio signal is, for example, a TSP signal, but may also be a signal representing other sounds such as speech or music. As illustrated in FIG. 6A, the elevation angle of this presentation angle θ is 0° in the horizontal direction in front of the user 1000, 90° above the user 1000, and 180° in the horizontal direction behind the user 1000. As illustrated in FIG. 6B, the azimuth angle of this presentation angle θ is 0° in the horizontal direction in front of the user 1000, 90° to the left, 180° to the back, and 270° to the right. There is no limitation on the presentation angle θ, but for example, the azimuth angle of the presentation angle θ is selected from eight directions: 0°, 40°, 90°, 140°, 180°, 220°, 270°, and 320°, and the elevation angle of the presentation angle θ is selected from five directions: 0°, 50°, 90°, 130°, and 180°.

[0046] Playback signal X F -1 is input to the acoustic signal output device 11-1, and the reproduced signal XF The sound output unit 111-1 of the sound signal output device 11-1 outputs the reproduced signal X F -1, the sound output unit 111-2 of the sound signal output device 11-2 emits a sound signal at a presentation angle θ, and the sound output unit 111-2 outputs a reproduction signal X F -2, an acoustic signal having a presentation angle θ is emitted. As a result, the subject user 1000 perceives an acoustic signal arriving from a certain direction. Here, the more appropriate the set head-related transfer function F is for the user 1000, the closer the angle of the arrival direction of the acoustic signal perceived by the user 1000 is to the presentation angle θ. By evaluating the degree of similarity between the presentation angle θ and the angle of the arrival direction of the acoustic signal perceived by the user 1000, the appropriateness of the set head-related transfer function F for the user 1000 can be evaluated. To evaluate this, the angle of the arrival direction of the acoustic signal perceived by the user 1000 (hereinafter referred to as the "answer angle") D (answer regarding the localization angle of the acoustic signal) is input to the input unit 133 of the evaluation processing device 13. Information specifying the input arrival direction D is stored in the memory unit 131 in association with information specifying the presentation angle θ of the presented acoustic signal. Such subjective evaluation of the head-related transfer function F may be performed for the user 1000 only once or multiple times. As a result, one or more pairs of the presentation angle θ and the answer angle D are stored in the storage unit 131. For the pairs of the presentation angle θ and the answer angle D obtained in this manner, the similarity between the presentation angle θ and the answer angle D is evaluated, thereby enabling the appropriateness of the head-related transfer function F set for the user 1000 to be evaluated. That is, the higher the similarity between the presentation angle θ and the answer angle D, the more appropriate the head-related transfer function F set for the user 1000 is, and the lower the similarity between the presentation angle θ and the answer angle D, the more inappropriate the head-related transfer function F set for the user 1000 is. The evaluation of the similarity between the presentation angle θ and the answer angle D may be performed locally by the evaluation processing device 13 or online. In the latter case, the pair of the presentation angle θ and the answer angle D may be transmitted to an external device via a network, and the external device may evaluate the similarity between the presentation angle θ and the answer angle D and evaluate the appropriateness of the head-related transfer function F.

[0047] <Features of the Present Embodiment> As described above, the setting device 1 (FIGS. 1 to 3) of the present embodiment emits an acoustic signal from the sound emitting unit 111-1 that does not seal the ear canal 1011-1 of the user 1000, and collects an observation signal based on the acoustic signal with the sound collecting unit 113-1 that is arranged at a position on the head that is neither the ear canal 1011-1 nor the open end of the ear canal 1011-1, and the collected observation signal is used to set the head-related transfer function F of the user 1000. In the present embodiment, an acoustic signal is actually emitted from the sound emitting unit 111-1, and the head-related transfer function F of the user 1000 is set based on the observation signal collected by the sound collecting unit 113-1, so that personal optimization of the head-related transfer function can be performed more easily and accurately than before. Here, strictly speaking, the head-related transfer function is a transfer function from a sound source to the ear canal 1011-1 or the open end of the ear canal 1011-1, but even when using an observation signal collected by a sound collection unit 113-1 arranged at a position on the head that is neither the ear canal 1011-1 nor the open end of the ear canal 1011-1, it is possible to perform personal optimization of the head-related transfer function with sufficient accuracy. In particular, in this embodiment, the sound collection unit 113-1 is attached to the auricle 1012-1 of the ear 1010-1 of the user 1000, so it is possible to perform personal optimization of the head-related transfer function with sufficient accuracy. The results of this experiment will be described later.

[0048] Preferably, there are multiple pairs of correspondences between the head-related transfer function candidates HRTF-1, ..., HRTF-N and the acoustic characteristics AOS-1, ..., AOS-N of the assumed observation signal assumed for the head-related transfer function candidates HRTF-1, ..., HRTF-N, and the setting processing unit 123 desirably selects the head-related transfer function F of the user 1000 from the head-related transfer function candidates HRTF-1, ..., HRTF-N based on the similarity between the acoustic characteristics of the observation signal and the acoustic characteristics AOS-1, ..., AOS-N of the assumed observation signal. This allows for personal optimization of the head-related transfer function with sufficient accuracy based on the similarity, even if the sound collection unit 113-1 is placed at a position on the head that is neither the ear canal 1011-1 nor the open end of the ear canal 1011-1, and the observation signal collected by the sound collection unit 113-1 does not strictly correspond to the head-related transfer function of the user 1000.

[0049] Preferably, the head-related transfer function F of the user 1000 is selected based on the similarity between the feature amounts corresponding to the notches and peaks in the spectral envelope of the observed signal and the feature amounts corresponding to the notches and peaks in the spectral envelope of the assumed observed signal. This is because the notches and peaks in the spectral envelope tend to reflect the characteristics of the shape of the head of the user 1000, including the ear 1010-1, and by using these, the head-related transfer function F can be appropriately optimized individually.

[0050] Preferably, the notches and peaks of the spectral envelope of the observed signal used in the individual optimization of the head-related transfer function F are [N11, P11], and the notches and peaks of the spectral envelope of the assumed observation signal are [N21, P21]. Similarly, preferably, the notches and peaks of the spectral envelope of the observed signal used in the individual optimization of the head-related transfer function F are [N11, P11, N12], and the notches and peaks of the spectral envelope of the assumed observation signal are [N21, P21, N22]. This is because these peaks and notches tend to reflect the characteristics of the shape of the head, including ear 1010-1 of user 1000, and by using these, the head-related transfer function F can be appropriately optimized individually.

[0051] Preferably, the setting processing unit 123 selects the head-related transfer function F of the user 1000 from the candidate head-related transfer functions HRTF-1, ..., HRTF-N based on the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics AOS-1, ..., AOS-N of the assumed observed signal in a frequency band including 5 to 13 kHz. This is because, in the frequency band of 5 to 13 kHz of the spectral envelope of the observed signal, the characteristics of the shape of the head of the user 1000, including the ear 1010-1, are particularly likely to appear, and by using information in this frequency band, it is possible to appropriately individually optimize the head-related transfer function F.

[0052] Furthermore, the evaluation processing device 13 of the evaluation device 100 (FIGS. 4 and 5) of this embodiment generates a reproduction signal X for emitting an acoustic signal at an arbitrary presentation angle θ based on the head-related transfer function F set based on the observation signal collected by the sound collection unit 113-1 of the setting device 1 described above. F -1, X FThe sound output unit 111-1 of the acoustic signal output device 11-1 outputs the reproduction signal X F -1, the sound output unit 111-2 of the sound signal output device 11-2 emits a sound signal at a presentation angle θ, and the sound output unit 111-2 outputs a reproduction signal X F -2, an acoustic signal having a presentation angle θ is emitted. As a result, the user 1000, who is the subject, perceives an acoustic signal arriving from a certain direction. The user 1000 inputs an answer angle D, which is the direction from which the perceived acoustic signal arrives (an answer related to the localization angle of the acoustic signal), to the evaluation processing device 13. Based on the similarity between the presentation angle θ and the answer angle D thus obtained, the appropriateness of the head-related transfer function F set for the user 1000 can be evaluated.

[0053] [Second Embodiment] The second embodiment is a modification of the first embodiment, and performs subjective evaluation of the head-related transfer function taking into account the HeSTF (Hearable Speaker Transfer Function), which is the transfer characteristic from the acoustic signal output device to the ear. The following description will focus on differences from the first embodiment, and the same reference numerals will be used to simplify the description of parts that are common to the first embodiment.

[0054] <HeSTF> HeSTF refers to the transfer characteristics from the device that outputs the acoustic signal to the ear, and indicates the characteristics of the device itself and the characteristics from the device to the ear. The characteristics from the device to the ear vary from person to person depending on the shape of the ear in the frequency band of 4 to 10 kHz. However, the individual differences in HeSTF are smaller than the individual differences in head-related transfer functions, and the impact due to individual differences in HeSTF is sufficiently smaller than the impact due to individual differences in head-related transfer functions.

[0055] <Direction Perception Model> Figure 7A shows a model of the direction perception of normal sounds. In the frequency domain, if the sound emitted from a normal sound source is S and the head-related transfer function from this sound source to the user's ear canal is F, the sound Y observed in the ear canal can be expressed as Y = FS.

[0056] Figure 7B shows a model of the directional perception of sound reproduced by a device that outputs an acoustic signal. In the frequency domain, if the sound emitted from a localized virtual sound source is S, the head-related transfer function from this virtual sound source to the user's ear canal is F, and the HeSTF from the device to the ear is H, then the sound Y observed in the ear canal can be expressed as Y = HFS. Therefore, the sound Y = HFS perceived by the user when sound is reproduced from a device that outputs an acoustic signal differs from the sound Y = FS perceived by the user when sound is emitted from a normal sound source.

[0057] 7C shows a model of directional perception when the component corresponding to HeSTF is corrected from the sound reproduced by a device that outputs an acoustic signal. In the frequency domain, if the sound emitted from a localized virtual sound source is S, the head-related transfer function from this virtual sound source to the user's ear canal is F, the HeSTF from the device to the ear is H, and the inverse filter that cancels H is C, then the sound Y observed in the ear canal can be expressed as Y = CHFS = FS. In other words, ideally, by using such an inverse filter C, the influence of HeSTF can be eliminated, and a sound equivalent to the sound Y = FS perceived when sound is emitted from a normal sound source can be perceived. Such an inverse filter C can be expressed, for example, as C in the following equation (1): conv (f) (See, for example, Reference 1). conv (f)=B(f)D * (f) / (D(f)D * (f)+βH(f)) (1) where f represents frequency, and the inverse filter C depends on the frequency f. D(f) is the transfer function from the device to the dummy head or user, which is measured in advance, and D * (f) is the complex conjugate of D(f), H(f) is a high-pass filter, and H * (f) is the complex conjugate of H(f), β is a real number greater than or equal to 0 that represents the weight of the high-pass filter, and B(f) represents the band-pass filter. In addition, to improve the accuracy of upward localization, C in the following equation (2) prop(f) may be used as the inverse filter C (see, for example, Reference 1). As mentioned above, there are individual differences in HeSTF depending on the shape of the ear, and there are also individual differences in D(f). However, since these individual differences are smaller than the individual differences in HeSTF, D(f) is common to all users. C prop (f)=C conv (f)+αL * (f)L(f) (2) where L(f) is a high-pass filter and L * (f) is the complex conjugate of L(f), and α is a real number greater than or equal to 0 that represents the weight of the high-pass filter. [Reference 1] Zora Scharer and Alexander Lindau, “Evaluation of Equalization Methods for Binaural Signals”, May 2009, Physics, Journal of the Audio Engineering Society.

[0058] <Configuration of Setting Device 1> The setting device 1 is the same as that in the first embodiment.

[0059] <Configuration of Evaluation Device 200> As illustrated in FIGS. 4 and 5, the evaluation device 200 of this embodiment has sound signal output devices 11-1 and 11-2 and an evaluation processing device 23.

[0060] The acoustic signal output devices 11-1 and 11-2 are the same as those in the first embodiment.

[0061] 4 and 5, the evaluation processing device 23 of this embodiment has a storage unit 131, a playback unit 232, and an input unit 133. The above-mentioned inverse filter C is input and set in the playback unit 232. The playback unit 232 is electrically connected to the storage unit 131 and the sound emitting units 111-1 and 111-2. The input unit 133 is electrically connected to the storage unit 131. However, this does not limit the present invention, and at least some of these units may be configured to be able to communicate wirelessly or via a network.

[0062] <Individual Optimization of Head-Related Transfer Functions> In this embodiment, the individual optimization of the head-related transfer functions of the user 1000 is the same as in the first embodiment.

[0063] <Subjective Evaluation of Head-Related Transfer Function> Next, the subjective evaluation of the head-related transfer function F set for the user 1000 will be described. The subjective evaluation of the head-related transfer function F in this embodiment is performed using the evaluation device 200. First, the head-related transfer function F set for the user 1000, who is the subject, is stored in the storage unit 131 of the evaluation device 200. Next, the user 1000 wears the acoustic signal output device 11-1 of the evaluation device 100 on the pinna 1012-1 of one ear 1010-1 (for example, the right ear) ( FIG. 4 ), and wears the acoustic signal output device 11-2 of the evaluation device 100 on the pinna 1012-2 of the other ear 1010-2 (for example, the left ear) ( FIG. 5 ).

[0064] Next, the reproduction unit 232 of the evaluation processing device 23 of the evaluation device 100 (FIGS. 4 and 5) extracts the head-related transfer function F set for the user 1000, who is the subject, from the storage unit 131, and generates a reproduction signal X for emitting an acoustic signal at an arbitrary presentation angle θ based on the head-related transfer function F and the inverse filter C using a known stereophonic technology. F,C -1, X F,C The difference from the first embodiment is that the inverse filter C is used to remove the influence of the HeSTF to generate and output the reproduced signal X F,C -1, X F,C That is, the reproduction signal X −2 in the first embodiment is generated. F -1, X F The playback signal X-2 of this embodiment allows the user 1000 to perceive the sound Y=HFS as shown in FIG. F,C -1, X F,C The playback signal X -2 allows the user 1000 to perceive the sound Y=CHFS=FS obtained by applying the inverse filter C to the Y=HFS illustrated in FIG. F,C -1 is supplied to the acoustic signal output device 11-1, and the playback signal X F,C -2 is supplied to the audio signal output device 11-2. This audio signal is also, for example, a TSP signal, but may also be a signal representing other sounds such as voice or music.

[0065] Playback signal X F,C -1 is input to the acoustic signal output device 11-1, and the reproduced signal X F,CThe sound output unit 111-1 of the sound signal output device 11-1 outputs the reproduced signal X F,C -1, the sound output unit 111-2 of the sound signal output device 11-2 emits a sound signal at a presentation angle θ, and the sound output unit 111-2 outputs a reproduction signal X F,C -2, an acoustic signal having a presentation angle θ is emitted. As a result, the user 1000, who is the subject, perceives an acoustic signal arriving from a certain direction. As in the first embodiment, the user 1000 inputs the response angle D of the acoustic signal perceived by the user 1000 into the input unit 133 of the evaluation processing device 23. Information specifying the input response angle D is stored in the storage unit 131 in association with information specifying the presentation angle θ of the presented acoustic signal. Such subjective evaluation of the head-related transfer function F may be performed for the user 1000 only once or multiple times. As a result, one or more pairs of the presentation angle θ and the response angle D are stored in the storage unit 131. For the pairs of the presentation angle θ and the response angle D obtained in this manner, the similarity between the presentation angle θ and the response angle D can be evaluated to evaluate the appropriateness of the head-related transfer function F set for the user 1000. The rest is the same as in the first embodiment.

[0066] <Characteristics of this embodiment> In this embodiment, the same effects as in embodiment 1 can be obtained. Furthermore, in this embodiment, the component corresponding to the HeSTF is corrected by the inverse filter C, so that the appropriateness of the head-related transfer function F can be evaluated more accurately than in embodiment 1.

[0067] [Third Embodiment] In the first and second embodiments, an acoustic signal is emitted from the sound emitting unit 111-1 of the acoustic signal output device 11-1 worn on one ear 1010-1, an observation signal based on the acoustic signal is collected by the sound collecting unit 113-1, and the collected observation signal is used to set the head-related transfer function F of the user 1000. However, it is also possible to emit an acoustic signal from the sound emitting units 111-1 of the acoustic signal output devices 11-1 and 11-2 worn on both ears 1010-1 and 1010-2, collect an observation signal based on the acoustic signal by the sound collecting units 113-1 and 113-2, and set the head-related transfer functions of both ears 1010-1 and 1010-2 of the user 1000, respectively, using the collected observation signal. That is, an acoustic signal is emitted from the sound emitting unit 111-1 of the acoustic signal output device 11-1 worn on one ear 1010-1 (for example, the right ear), an observation signal based on the acoustic signal is collected by the sound collecting unit 113-1, and the collected observation signal is used to set the head-related transfer function of one ear 1010-1 of the user 1000. Furthermore, an acoustic signal is emitted from the sound emitting unit 111-2 of the acoustic signal output device 11-2 worn on the other ear 1010-2, an observation signal based on the acoustic signal is collected by the sound collecting unit 113-2, and the collected observation signal is used to set the head-related transfer function of the other ear 1010-2 of the user 1000. This allows the head-related transfer functions of the left and right ears to be set independently. This allows for more accurate personal optimization of the head-related transfer function.

[0068] When the head-related transfer functions for the left and right ears are set independently, the head-related transfer functions set independently for the left and right ears may be used to subjectively evaluate these head-related transfer functions. In this case, the playback signal X F -1, X F -2 and playback signal X F,C -1, X F,C -2 is generated. The rest is the same as described in the first and second embodiments.

[0069] [Experimental Results] Next, the experimental results are shown. Figure 8 illustrates the relationship between the acoustic characteristics of an observation signal observed by a sound collection unit placed at the opening end of the ear canal 1011-1 and the acoustic characteristics of an observation signal observed by a sound collection unit 113-1 placed at a position on the head that is not the ear canal or the opening end of the ear canal. Here, "HATS Front," "HATS 10°," "HATS 20°," and "HATS 30°" illustrate the spectral envelope of an observation signal obtained by observing an acoustic signal emitted from a sound source localized at elevation angles of 0°, 10°, 20°, and 30° with a sound collection unit placed at the opening end of the ear canal 1011-1 of the user 1000. Furthermore, "DOWN" illustrates the spectral envelope of an observation signal obtained by observing an acoustic signal emitted from a sound emission unit 111-1 placed on the front side of the pinna 1012-1 with a sound collection unit 113-1 placed on the back side of the pinna 1012-1. "UP" illustrates the spectral envelope of an observed signal obtained by observing an acoustic signal emitted from sound emitting unit 111-1 located on the front side of pinna 1012-1 with a sound collecting unit located at a higher position (near the helix) than sound collecting unit 113-1 on the back side of pinna 1012-1. As illustrated in Figure 8, in the frequency range of 5 to 13 kHz, the spectral envelopes of "HATS10°" and "HATS20°" are similar to the spectral envelopes of "UP" and "DOWN," and the positions of the notches and peaks in the spectral envelopes of "HATS10°" and "HATS20°" show a tendency similar to the positions of the notches and peaks in the spectral envelopes of "UP" and "DOWN." Therefore, it is understood that it is possible to emit an acoustic signal from sound emitting unit 111-1 that does not seal ear canal 1011-1 of user 1000, collect an observation signal based on the acoustic signal with sound collecting unit 113-1 that is arranged at a position on the head that is neither ear canal 1011-1 nor the open end of ear canal 1011-1, and use the collected observation signal to set a head-related transfer function F for user 1000. In particular, it is understood that when the acoustic signal emitted from sound emitting unit 111-1 is an acoustic signal emitted from a sound source localized at a virtual position in a direction at an elevation angle of 10 to 20 degrees with respect to user 1000, it is possible to set a head-related transfer function suitable for user 1000 with higher accuracy.

[0070] 9A is a graph illustrating the relationship between the acoustic characteristics of an observation signal observed by a sound collecting unit 113-1 arranged behind the pinna 1012-1 and the acoustic characteristic AOS-n of an assumed observation signal corresponding to a head-related transfer function F∈{HRTF-1, ..., HRTF-N} selected from candidate head-related transfer functions HRTF-1, ..., HRTF-N based on the Euclidean distance between the acoustic characteristics of the observation signal and the acoustic characteristic AOS-n of an assumed observation signal corresponding to the candidate head-related transfer function HRTF-n. FIG. 9B is a graph illustrating the relationship between the acoustic characteristics of an observation signal observed by a sound collecting unit 113-1 arranged behind the pinna 1012-1 and the acoustic characteristic AOS-n of an assumed observation signal corresponding to a head-related transfer function F∈{HRTF-1, ..., HRTF-N} selected from candidate head-related transfer functions HRTF-1, ..., HRTF-N based on the cosine similarity between the acoustic characteristics of the observation signal and the acoustic characteristic AOS-n of an assumed observation signal corresponding to the candidate head-related transfer function HRTF-n. 9C is a graph illustrating the relationship between the acoustic characteristics of an observed signal observed by a sound collection unit 113-1 disposed behind the pinna 1012-1 and the acoustic characteristic AOS-n of an assumed observed signal corresponding to a head-related transfer function Fε{HRTF-1, ..., HRTF-N} selected from candidate head-related transfer functions HRTF-1, ..., HRTF-N based on the spectral distortion of the acoustic characteristic AOS-n of an assumed observed signal corresponding to the candidate head-related transfer function HRTF-n for the acoustic characteristics of the observed signal. It can be seen that, regardless of whether the Euclidean distance, cosine similarity, or spectral distortion measure is used, an assumed observed signal having an acoustic characteristic AOS-n similar to the acoustic characteristics of the observed signal can be selected, and the candidate head-related transfer function HRTF-n corresponding to this acoustic characteristic AOS-n can be set as the head-related transfer function F of the user 1000. In particular, subjective evaluation of head-related transfer functions showed that using Euclidean distance enabled the setting of a more accurate head-related transfer function F.

[0071] Next, an example experiment for subjective evaluation of head-related transfer functions according to the second embodiment will be shown. <Experimental conditions for subjective evaluation> The experimental conditions for subjective evaluation are as follows. (1) Features used to select head-related transfer functions: A four-dimensional vector whose elements are the frequency and power of [N11, P11] of the spectral envelope of the observed signal, and a four-dimensional vector whose elements are the frequency and power of [N21, P21] of the spectral envelope of the assumed observed signal. A six-dimensional vector whose elements are the frequency and power of [N11, P11, N12] of the spectral envelope of the observed signal, and a six-dimensional vector whose elements are the frequency and power of [N21, P21, N22] of the spectral envelope of the assumed observed signal. (2) Similarity measure: Closeness of Euclidean distance. (3) Presentation angle θ: Elevation angle: Five directions (FIG. 6A) of 0°, 50°, 90°, 130°, and 180° were selected randomly three times each. Azimuth angle: Eight directions (Fig. 6B) of 0°, 40°, 90°, 140°, 180°, 220°, 270°, and 320° are randomly selected three times each. (4) Inverse filter C: C calculated using the transfer function D(f) from the sound emitting unit 111-1 to the entrance of the ear canal 1011-1 of the user 1000. prop (f) (5) Evaluating the appropriateness of the head-related transfer function: This was carried out online. The evaluation processing device 23 was installed on each subject's PC.

[0072] <Subjective Evaluation Experimental Results> Figures 10A to 10D show the experimental results for subject 1, and Figures 11A to 11D show the experimental results for subject 2. Figures 10A and 11A are bubble charts illustrating the subjective evaluation results when the azimuth of the presentation angle θ was randomly changed based on head-related transfer functions corresponding to assumed observation signals having acoustic characteristics that are first to third most similar to the acoustic characteristics of the observed signal. Figures 10B and 11B are bubble charts illustrating the subjective evaluation results when the azimuth of the presentation angle θ was randomly changed based on head-related transfer functions corresponding to assumed observation signals having acoustic characteristics that are least similar to the acoustic characteristics of the observed signal. Figures 10C and 11C are bubble charts illustrating the subjective evaluation results when the elevation angle of the presentation angle θ was randomly changed based on head-related transfer functions corresponding to assumed observation signals having acoustic characteristics that are first to third most similar to the acoustic characteristics of the observed signal. 10D and 10D are bubble charts illustrating subjective evaluation results when the elevation angle of the presentation angle θ is randomly changed based on the head-related transfer function corresponding to the assumed observation signal having acoustic characteristics least similar to those of the observed signal. The horizontal axis of FIGS. 10A to 10D and 11A to 11D represents the presentation angle θ, and the vertical axis represents the answer angle D. As can be seen from a comparison between FIGS. 10C and 10D and between FIGS. 11C and 11D in particular, the answer angle D is closer to the presentation angle θ when a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics highly similar to those of the observed signal is used than when a head-related transfer function corresponding to an assumed observation signal having acoustic characteristics less similar to those of the observed signal is used. This shows that personal optimization of a head-related transfer function is possible by selecting a head-related transfer function from candidate head-related transfer functions based on the similarity between the acoustic characteristics of the observed signal and those of the assumed observation signal.

[0073] [Modifications, etc.] The present invention is not limited to the above-described embodiments. For example, the positions and numbers of the sound emitting units and sound collecting units are not limited to those exemplified in the embodiments. Furthermore, the sound emitting units and sound collecting units may be devices other than open-ear earphones. For example, the sound emitting units and sound collecting units may be open-ear headphones or neck speakers, or may be devices fixed to the seat.

[0074] In the second embodiment, a transfer function D(f) measured in advance from the device to the dummy head or the user is used to calculate the HeSTF from the device to the ear. As described above, individual differences in HeSTF are smaller than individual differences in HRTF. However, using an HeSTF optimized for each user 1000 improves the estimation accuracy of the HRTF. Therefore, when each user 1000 first individually optimizes their HRTF, an observation signal based on the acoustic signal emitted from the sound emitting unit 111-1 may be collected by a sound collecting unit arranged in the ear canal 1011-1 of each user 1000 or the open end of the ear canal 1011-1, and the HeSTF of each user 1000 may be obtained using the transfer function D(f) based on this observation signal.

[0075] 12, such a setting device 4 includes an acoustic signal output device 41-1 and a head-related transfer function setting device 42. In addition to a sound emitting unit 111-1, a sound collecting unit 113-1, and a wearing unit 112-1, the acoustic signal output device 41-1 further includes a sound collecting unit 413-1 that is disposed in the ear canal 1011-1 of the user 1000 or at the open end of the ear canal 1011-1 and is configured to collect an observation signal based on the acoustic signal emitted from the sound emitting unit 111-1. A specific example of the sound collecting unit 413-1 is a microphone.

[0076] The head-related transfer function setting device 42 has a storage unit 121, a playback unit 122, and a setting processing unit 423. The setting processing unit 423 has the functions of the setting processing unit 123 described above, and further has the function of obtaining and outputting an inverse filter C based on an observation signal collected by the sound collection unit 413-1. Here, only the inverse filter setting process of obtaining and outputting the inverse filter C, which is a difference between the setting processing unit 423 and the setting processing unit 123, will be described. The inverse filter setting process may be performed when each user 1000 first performs individual optimization of the head-related transfer function, but may also be performed multiple times for at least one user 1000. Furthermore, the inverse filter setting process may be performed before the individual optimization of the head-related transfer function, may be performed after the individual optimization of the head-related transfer function, or may be performed at a timing unrelated to the individual optimization of the head-related transfer function.

[0077] When performing the inverse filter setting process, the playback unit 122 of the head-related transfer function setting device 42 ( FIG. 12 ) extracts acoustic signal information A from the correspondence table T stored in the storage unit 121 and sends it to the playback unit 122. The playback unit 122 outputs a playback signal for outputting an acoustic signal specified by the acoustic signal information A. The playback signal is input to the driver unit of the sound emitting unit 111-1, and the driver unit outputs an acoustic signal based on the input playback signal, and this acoustic signal is emitted to the outside from the sound hole 111a-1. The sound collecting unit 413-1 collects an observation signal based on this acoustic signal and sends information representing the observation signal to the setting processing unit 423. The acoustic signal information A is also sent to the setting processing unit 423. The setting processing unit 423 obtains and outputs an inverse filter C based on the input acoustic signal information A and the information representing the observation signal. For example, the setting processing unit 423 calculates the transfer function D(f) of the user 1000 based on the input acoustic signal information A and information representing the observed signal, and obtains and outputs an inverse filter C based on the transfer function D(f). Note that if the acoustic signal information A is set in the setting processing unit 423 in advance, input of the acoustic signal information A to the setting processing unit 423 may be omitted. A specific example of the inverse filter C is C in the above-mentioned formula (1): conv (f) or C in Eq. (2) propThe output inverse filter C may be input to and set by the reproduction unit 232 of the evaluation processing device 23 of the second embodiment. The rest is as described in the first and second embodiments.

[0078] The setting processing unit 423 may use the HeSTF of the user 1000 obtained as described above to calculate the head-related transfer function F of this user 1000. For example, a model that uses the HeSTF as input and outputs the head-related transfer function F or a probability distribution of the head-related transfer function F may be trained by machine learning, and this model may be set in the setting processing unit 423. Any model may be used, for example, a model based on deep learning may be exemplified. In this case, the setting processing unit 423 may input the HeSTF generated as described above into the model, obtain and output the head-related transfer function F or a probability distribution of the head-related transfer function F. Alternatively, an observation signal based on an acoustic signal emitted from the sound emitting unit 111-1 may be collected by the sound collecting unit 113-1, and the setting processing unit 423 may calculate the head-related transfer function F using this observation signal. For example, a model that uses the HeSTF and the observation signal as input and outputs the head-related transfer function F or a probability distribution of the head-related transfer function F may be trained by machine learning, and this model may be set in the setting processing unit 423. In this case, the setting processing unit 423 may input the above-mentioned HeSTF and the observation signal into a model, and obtain and output the head-related transfer function F or the probability distribution of the head-related transfer function F. For example, a model that takes the observation signal as input and outputs the head-related transfer function F or the probability distribution of the head-related transfer function F may be learned by machine learning, and this model may be set in the setting processing unit 423. In this case, the setting processing unit 423 may input the above-mentioned observation signal into the model, and obtain and output the probability distribution of the head-related transfer function F or the head-related transfer function F.

[0079] An acoustic signal output device having a shape other than the acoustic signal output devices 11-1 and 11-2 described above (e.g., open-ear earphones) may be used as long as it has a sound emitting unit configured to emit an acoustic signal without sealing the ear canal 1011-1 of the user 1000. For example, the acoustic signal output device 51-i (where i = 1 or 2) illustrated in FIG. 13A may be used. The acoustic signal output device 51-i illustrated in FIG. 13A includes the driver unit (not shown), a substantially spherical sound emitting unit 511-i that houses the driver unit therein, a substantially spherical mounting unit 512-i that is placed on the auricle 1010-i when worn, a curved portion 514-i that is an elastic body connecting the sound emitting unit 511-i and the mounting unit 512-i, and a sound collecting unit 513-i provided in the mounting unit 512-i. When the acoustic signal output device 51-i is worn, the sound emitting unit 511-i is positioned on the front side of the auricle 1010-i (the side of the ear canal 1011-i), and the attachment unit 512-i is positioned on the back side of the auricle 1010-i (the side where the ear canal 1011-i is not present), with the auricle 1010-i sandwiched between the sound emitting unit 511-i and the attachment unit 512-i. As a result, when the acoustic signal output device 51-i is worn, the sound emitting unit 511-i is positioned on the front side of the auricle 1010-i, and the sound collecting unit 513-i is positioned on the back side of the auricle 1010-i. Alternatively, the sound collecting unit 513-i may be configured to be positioned on the back side of the auricle 1010-i in this manner. For example, the acoustic signal output device 51-i' (where i = 1, 2) illustrated in FIG. 13B may be used. The acoustic signal output device 51-i' differs from the acoustic signal output device 51-i in that the sound collection unit 513-i is disposed inside the curved portion 514-i on the sound emission unit 511-i side, a sound hole 516-i is provided in the wearing portion 512-i, and a sound conduit 515-i is provided inside the wearing portion 512-i, connecting the sound collection unit 513-i and the sound hole 516-i and guiding the acoustic signal arriving at the sound hole 516-i to the sound collection unit 513-i. As a result, when the acoustic signal output device 51-i' is worn, the sound emission unit 511-i and the sound collection unit 513-i are disposed on the front side of the auricle 1010-i, and the sound guide hole 516-i is disposed on the back side of the auricle 1010-i. With such a configuration, it is necessary to estimate the head-related transfer function F taking into account the transfer characteristics of the sound conduit 515-i and the sound guide hole 516-i.

[0080] A glasses-type (eyeglasses-type) acoustic signal output device 61 as shown in FIG. 14 may be used. The acoustic signal output device 61 includes a glasses portion 612 having a tip cell portion 6122, a temple portion 6121, and a front portion 6123; a sound emitting portion 611 configured to emit an acoustic signal without sealing the ear canal 1011-1 of the user 1000; and a sound collecting portion 613 arranged at a position on the head that is neither the ear canal 1011-1 nor the open end of the ear canal 1011-1 and configured to collect an observation signal based on the acoustic signal. The sound emitting portion 611 and the sound collecting portion 613 are fixed to the glasses portion 612. In this example, the sound emitting portion 611 is attached to the temple portion 6121. As a result, when the user 1000 wears the acoustic signal output device 61, the sound emitting portion 611 is configured to be arranged on the back side of the user's 1000's auricle 112-1 (the side where the ear canal 1011-1 is not present). Furthermore, in this configuration, the sound emitting unit 611 is positioned above the ear canal 1011-1 at a position behind the auricle 112-1. For example, the sound collection unit 613 is positioned in a region above the ear canal 1011-1 at a position behind the auricle 112-1. However, this is merely an example and does not limit the present invention. The sound emitting unit 611 may be provided in a portion of the tip cell unit 6122 other than the position illustrated in FIG. 14 , or the sound emitting unit 611 may be provided at a position on the tip cell unit 6122 side of the temple unit 6121. Furthermore, the sound emitting unit 611 in this example is provided in the temple unit 6121. Preferably, the sound emitting unit 611 is provided at a position on the front unit 6123 side of the temple unit 6121. The sound emitting unit 611 may also be provided in the front unit 6123. This makes it possible to increase the distance between the sound emitting unit 611 and the sound collecting unit 613, making it easier to observe the reflected component of the acoustic signal emitted from the sound emitting unit 611 at the head of the user 1000 at the sound collecting unit 613. As a result, in the individual optimization of the head-related transfer function, it becomes easier to select (match) the head-related transfer function F from directly in front of the user 1000 to the ear canal 1011-1, which corresponds to the observed signal at the sound collecting unit 613. Furthermore, the directional sound collecting unit 613 may have directionality (for example, the sound collecting unit 613 is a directional microphone), and may be arranged so that the sound collection direction of the sound collecting unit 613 faces the ear canal 1012-1 when the user 1000 wears the acoustic signal output device 61.This also makes it easier to observe the reflected components of the acoustic signal emitted from the sound emission unit 611 at the head of the user 1000 at the sound collection unit 613. Alternatively, the sound collection unit 613 may be a microphone array. This may suppress ambient noise contained in the observed signal at the sound collection unit 613 while emphasizing the reflection at the pinna 1011-1. Furthermore, the number of times the acoustic signal (for example, the TSP signal) emitted from the sound emission unit 611 is played may be increased, and the noise components may be reduced by synchronously adding the observed signal at the sound collection unit 613. This may improve the accuracy of individual optimization of the head-related transfer function.

[0081] Furthermore, the various processes described above may not only be executed in chronological order as described, but may also be executed in parallel or individually depending on the processing capacity of the device executing the processes or as necessary. Needless to say, other modifications are possible within the scope of the present invention.

[0082] [Hardware Configuration] The head-related transfer function setting device 12 and the evaluation processing device 13, 23 in each embodiment are devices configured by a general-purpose or dedicated computer having a processor (hardware processor) such as a CPU (central processing unit) and memories such as RAM (random-access memory) and ROM (read-only memory) executing a predetermined program. That is, the head-related transfer function setting device 12 and the evaluation processing device 13, 23 in each embodiment have, for example, a processing circuit configured to implement each unit possessed by each of them. This computer may have one processor and memory, or may have multiple processors and memories. This program may be installed on the computer or may be pre-recorded in a ROM or the like. Furthermore, some or all of the processing units may be configured using electronic circuits that independently realize processing functions, rather than electronic circuits that realize functional configurations by loading programs like a CPU. Furthermore, the electronic circuits constituting one device may include multiple CPUs.

[0083] FIG. 15 is a block diagram illustrating the hardware configuration of the head-related transfer function setting device 12, 42 and the evaluation processing device 13, 23 in each embodiment. As illustrated in FIG. 15, the head-related transfer function setting device 12, 42 and the evaluation processing device 13, 23 in this example include a CPU (Central Processing Unit) 10a, an input unit 10b, an output unit 10c, a RAM (Random Access Memory) 10d, a ROM (Read Only Memory) 10e, an auxiliary storage device 10f, a communication unit 10h, and a bus 10g. The CPU 10a in this example includes a control unit 10aa, a calculation unit 10ab, and a register 10ac, and executes various calculation processes according to various programs loaded into the register 10ac. The input unit 10b is an input terminal, keyboard, mouse, touch panel, etc., through which data is input. The output unit 10c is an output terminal, display, etc., through which data is output. The communication unit 10h is a LAN card, etc., controlled by the CPU 10a that has loaded a predetermined program. The RAM 10d is a static random access memory (SRAM), dynamic random access memory (DRAM), or the like, and has a program area 10da where predetermined programs are stored and a data area 10db where various data are stored. The auxiliary storage device 10f is a hard disk, magneto-optical disc (MO), semiconductor memory, or the like, and has a program area 10fa where predetermined programs are stored and a data area 10fb where various data are stored. The bus 10g connects the CPU 10a, input unit 10b, output unit 10c, RAM 10d, ROM 10e, communication unit 10h, and auxiliary storage device 10f so that information can be exchanged. The CPU 10a writes the program stored in the program area 10fa of the auxiliary storage device 10f to the program area 10da of RAM 10d in accordance with the loaded OS (Operating System) program. Similarly, the CPU 10a writes various data stored in the data area 10fb of the auxiliary storage device 10f to the data area 10db of the RAM 10d.The addresses on the RAM 10d where these programs and data are written are stored in the register 10ac of the CPU 10a. The control unit 10aa of the CPU 10a sequentially reads out these addresses stored in the register 10ac, reads out the programs and data from the areas on the RAM 10d indicated by the read addresses, causes the calculation unit 10ab to sequentially execute the calculations indicated by the programs, and stores the calculation results in the register 10ac. With this configuration, the functional configurations of the head-related transfer function setting devices 12, 42 and the evaluation processing devices 13, 23 are realized.

[0084] The above-mentioned program can be recorded on a computer-readable recording medium. Examples of computer-readable recording media include non-transitory recording media. Examples of such recording media include magnetic recording devices, optical disks, magneto-optical recording media, and semiconductor memories.

[0085] This program may be distributed, for example, by selling, transferring, or lending a portable recording medium, such as a DVD or CD-ROM, on which the program is recorded. Furthermore, the program may be distributed by storing the program in a storage device of a server computer and transferring the program from the server computer to other computers via a network. As described above, a computer that executes such a program may first temporarily store the program recorded on a portable recording medium or transferred from the server computer in its own storage device. Then, when executing a process, the computer reads the program stored in its own storage device and executes processing in accordance with the read program. Alternatively, the program may be executed by a computer that reads the program directly from a portable recording medium and executes processing in accordance with the program. Furthermore, the computer may execute processing in accordance with the received program each time a program is transferred from the server computer to the computer. Alternatively, the server computer may not transfer the program to the computer, but may instead execute the processing function simply by issuing an execution instruction and obtaining the results, thereby executing the processing described above through a so-called ASP (Application Service Provider) service. In this embodiment, the program includes information used for processing by an electronic computer that is equivalent to a program (such as data that is not a direct instruction to a computer but has properties that dictate computer processing).

[0086] In each embodiment, the device is configured by executing a predetermined program on a computer, but at least a part of the processing contents may be realized by hardware.

[0087] 1 Setting device 11, 41, 51, 51', 61 Acoustic signal output device 12, 42 Head-related transfer function setting device 13, 23 Evaluation processing device 100, 200 Evaluation device

Claims

1. A sound emitting unit configured to emit an acoustic signal without sealing the ear canal of a user; a sound collecting unit arranged at a position on the head that is neither the ear canal nor the open end of the ear canal and configured to collect an observation signal based on the acoustic signal; and a setting processing unit in which a plurality of pairs of head-related transfer function candidates and acoustic characteristics of assumed observation signals assumed for the head-related transfer function candidates are associated with each other, and which selects a head-related transfer function of the user from the head-related transfer function candidates based on the similarity between the acoustic characteristics of the observation signal and the acoustic characteristics of the assumed observation signal, The similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the assumed observed signal is a value based on at least one of: (1) the logarithmic spectral distance between the spectral envelope of the observed signal and the spectral envelope of the assumed observed signal; (2) the distance between a vector representing the principal components of the acoustic characteristics of the observed signal and a vector representing the principal components of the acoustic characteristics of the assumed observed signal; and (3) a distance measure value that increases as the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal increase. A setting device in which N11 is a notch in the spectral envelope of the observed signal that is located at the lowest frequency side within a specified frequency range, N12 is a notch in the spectral envelope of the observed signal that is located next to N11 at the lowest frequency side within the specified frequency range, N21 is a notch in the spectral envelope of the assumed observed signal that is located at the lowest frequency side within the specified frequency range, and N22 is a notch in the spectral envelope of the assumed observed signal that is located next to N21 at the lowest frequency side within the specified frequency range.

2. A setting device according to claim 1, wherein the sound collection unit is configured to be worn on the ear of the user.

3. The setting device according to claim 1, wherein the lower limit of the predetermined frequency range is 5 kHz.

4. A setting device according to claim 1, wherein the setting processing unit selects the user's head-related transfer function from the candidate head-related transfer functions based on the similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the assumed observed signal in a frequency band including 5 to 13 kHz.

5. A method for selecting a head-related transfer function of the user from the candidate head-related transfer functions, the method including the steps of: emitting an acoustic signal from a sound emitting unit that does not seal the ear canal of the user; collecting an observation signal based on the acoustic signal with a sound collecting unit that is placed at a position on the head that is neither the ear canal nor the open end of the ear canal; and associating a plurality of pairs of head-related transfer function candidates with acoustic characteristics of assumed observation signals assumed for the head-related transfer function candidates, based on the similarity between the acoustic characteristics of the observation signal and the acoustic characteristics of the assumed observation signal, The similarity between the acoustic characteristics of the observed signal and the acoustic characteristics of the assumed observed signal is a value based on at least one of: (1) the logarithmic spectral distance between the spectral envelope of the observed signal and the spectral envelope of the assumed observed signal; (2) the distance between a vector representing the principal components of the acoustic characteristics of the observed signal and a vector representing the principal components of the acoustic characteristics of the assumed observed signal; and (3) a distance measure value that increases as the difference between the frequency of notch N11 in the spectral envelope of the observed signal and the frequency of notch N21 in the spectral envelope of the assumed observed signal and the difference between the frequency of notch N12 in the spectral envelope of the observed signal and the frequency of notch N22 in the spectral envelope of the assumed observed signal increase. A setting method in which N11 is a notch in the spectral envelope of the observed signal that is located at the lowest frequency side within a specified frequency range, N12 is a notch in the spectral envelope of the observed signal that is located next to N11 at the lowest frequency side within the specified frequency range, N21 is a notch in the spectral envelope of the assumed observed signal that is located at the lowest frequency side within the specified frequency range, and N22 is a notch in the spectral envelope of the assumed observed signal that is located next to N21 at the lowest frequency side within the specified frequency range.

6. A program for causing a computer to function as the setting device of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Sound reproducing device

    JP2000324590A

  • Audio metrics for head-related transfer function (HRTF) selection or adaptation

    US20120328107A1

  • Audio personalisation method and system

    US20220124448A1