Acoustic signal processing device, and program
The acoustic signal processing device uses spherical harmonic spectra to render sound in the frequency domain, addressing the limitations of traditional binaural playback by allowing smooth movement of sound images and maintaining accurate sound localization.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON HOSO KYOKAI
- Filing Date
- 2022-02-22
- Publication Date
- 2026-04-17
AI Technical Summary
Existing binaural playback technologies struggle to smoothly render a sound image that moves along an arbitrary trajectory due to the limited number of sound source and listening point pairs in HRIR measurement, leading to disruptions in frequency characteristics when mixing output signals from different HRTFs.
An acoustic signal processing device that calculates conversion coefficients using spherical harmonic spectra to render sound in the frequency domain, allowing the sound source or listening point to move arbitrarily by linearly combining basis functions, with the head position or each ear's position as the listening point, thereby maintaining frequency characteristics for accurate sound localization.
Enables continuous and accurate sound image localization by interpolating sound in the frequency domain, preserving spectral cues for sound image localization without interference-induced distortion.
Smart Images

Figure 0007847447000020 
Figure 0007847447000021 
Figure 0007847447000022
Abstract
Description
[Technical Field]
[0001] This invention relates to an acoustic signal processing device and a program, and more particularly to a technology for realizing binaural playback. [Background technology]
[0002] Augmented Reality (AR) and Virtual Reality (VR) content is sometimes created or presented by rendering sound in accordance with the movements of video objects and the viewer. Sound rendering is used to create an effective sense of presence. AR / VR content is sometimes created with the assumption that it will be presented while wearing headphones or a head-mounted display (HMD) with built-in headphones. Binaural playback technology is sometimes employed as a technique to virtually recreate a three-dimensional sound space using headphones.
[0003] Binaural playback is achieved by reproducing the sound pressure obtained when sound waves emitted from a sound source placed at an arbitrary position reach the entrance of the ear canal of both ears of the listener. In binaural playback, sound is presented directly using playback sound sources close to each ear, such as headphones or earphones. The transmission characteristics of sound from the sound source to each ear include the effects of reflection, diffraction, and attenuation of sound waves in the listener's head, auricle, torso, etc. (hereinafter collectively referred to as "head, etc."). By presenting sound with these transmission characteristics added, placed virtually at a set position, it is possible to provide the listener with a high sense of presence.
[0004] In binaural playback, the head-related transfer function (HRTF) or head-related impulse response (HRIR) is used as a feature that indicates the transfer characteristics from the sound source position in space to the listening position (listening point). HRTF is expressed in the frequency domain, while HRIR is the time-domain representation of HRTF. Generally, HRTF or HRIR is measured in a special acoustic environment such as an anechoic chamber. In the measurement, sound based on the measurement signal is emitted from speakers placed around the listener, and the sound is captured using a microphone placed at the entrance of the listener's ear canal. The HRTF or HRIR (hereinafter collectively referred to as "HRIR, etc.") is obtained using the captured signal and a known measurement signal.
[0005] A binaural signal is a two-channel acoustic signal that includes a left-ear acoustic signal (hereinafter referred to as the "left-ear signal") obtained by convolving the left ear's HRIR into the input acoustic signal (hereinafter referred to as the "input signal"), and a right-ear signal obtained by convolving the right ear's HRIR into the input signal. Each ear's acoustic signal (hereinafter referred to as the "each ear's signal") y is obtained by performing a convolving operation of the input signal with HRIR, as illustrated in equation (1). In equation (1), y(t), x(t), and h(t) represent the sample values of each ear's signal, the input signal, and HRIR at time t, respectively.
[0006]
number
[0007] Each ear signal y(t) can also be obtained by calculating the product Y(ω) of the Fourier transform X(ω) of the input signal and the HRTF H(ω), as exemplified in equation (2), and then performing an inverse Fourier transform on the product Y(ω). In equation (2), ω represents frequency.
[0008]
number
[0009] As described above, binaural playback is achieved by using headphones to present sounds based on signals for each ear to the corresponding ear. The sound pressure of the sound presented to the external auditory canals of both ears is equivalent to the sound pressure from sound waves arriving from the sound source location used in measurements such as HRIR. Therefore, the listener can perceive a virtual sound image at that sound source location.
[0010] In recent years, the notation of sound fields based on spatial Fourier series expansion has attracted attention as a technique for representing acoustics in three-dimensional space. A representative example is a method based on spherical harmonic expansion (for example, Non-Patent Document 1). This method is based on the fact that the sound pressure distribution p(r,θ,φ,ω) in three-dimensional space can be expressed as a linear combination of basis functions of the spherical harmonic expansion, as shown in equation (3). In other words, the spherical harmonic expansion describes the sound pressure distribution in three-dimensional space in a form in which the radial component and angular component are separated by variables. Equation (3) corresponds to the general solution of the three-dimensional wave equation given in polar coordinates by transforming the variables of the three-dimensional wave equation given in Cartesian coordinates. In equation (3), (r,θ,φ) represents the three-dimensional coordinates expressed in polar coordinates. n (2) This shows the nth order spherical Hankel function of the second kind. The spherical Hankel function of the second kind gives an orthogonal basis for the radial r component. Y n m This shows the nth-th order m-th spherical harmonics. Spherical harmonics provide an orthogonal basis in the angular direction. A n m This shows a spherical harmonic spectrum. In a spherical harmonic spectrum, the product of the second kind spherical Hankel function and spherical harmonics corresponds to the weighting coefficients for the basis functions of the spherical harmonic expansion, and it can represent the spatial distribution of the sound field, such as the directivity of the sound source. Applications of the spherical harmonic spectrum to sound field control, such as rotation of a virtual sound source in an arbitrary direction and interpolation of sound pressure at an arbitrary position, have been proposed.
[0011]
number
[0012] Ambisonics (e.g., Non-Patent Document 2) is a well-known technology that utilizes this principle. This technology has also been adopted and standardized in MPEG-H 3DA (Non-Patent Document 3), a next-generation speech encoding scheme. Furthermore, a method has been proposed to encode binaural signals using Ambisonics with HRTF measured in three-dimensional space (Patent Document 1). [Prior art documents] [Patent Documents]
[0013] [Patent Document 1] Patent No. 6067934 [Non-patent literature]
[0014] [Non-Patent Document 1] Yoichi Haneda, "Signal Processing in the Wavenumber Domain of Sound," Fundamentals Review, IEICE Fundamentals and Boundary Society, Vol. 11, No. 4, pp. 243-255, 2017. [Non-Patent Document 2] DH Cooper, T. Shiga, Discrete-matrix multichannel stereo. Journal of Audio Engineering Society 20(5), pp.346-360, 1972. [Non-Patent Document 3] ISO / IEC 23008-3:2019 “Information technology - High efficiency coding and media delivery in heterogeneous environments - Part 3: 3D audio, Second edition” (2019) [Non-Patent Document 4] AV Oppenheim, RW Schafer, Digital signal processing, Englewood Cliffs, NJ: Prentice-Hall, 1975. [Non-Patent Document 5] PA Martin, Multiple Scattering: Interaction of Time-Harmonic Waves with N obstacles, Cambridge university press, 2006. [Overview of the project] [Problems that the invention aims to solve]
[0015] HRIRs are determined for pairs of sound source positions and listening points within an acoustic space. In binaural playback using a single pair of HRIRs, a sound image is perceived as being stationary at the sound source position, assuming the listening point is located there. To represent the movement of the sound image or the listener, it is common to switch HRIRs according to changes in the positional relationship between the sound source and the listening point. However, since the number of sound source and listening point pairs involved in HRIR measurement is finite, even if the sound source position is fixed, the positions of the sound image that can be represented are limited. Simply switching HRIRs makes it difficult to move the listening point along an arbitrary trajectory and make the listener perceive a smoothly moving sound image.
[0016] In binaural playback using HRTFs, the sound source signal is Fourier transformed buffer by buffer, and rendering is performed by multiplying the buffer with the HRTF in the frequency domain. Different HRTFs are used for each buffer, each representing a different sound source position or listening point. By presenting a sound that is a mix of the output signals from at least two buffers, the sound source position or listening point of the sound image is linearly approximated between the buffers. To perceive a smoothly moving sound image, it is desirable to perform rendering in time samples. Furthermore, mixing output signals based on different HRTFs in the time domain may disrupt the frequency characteristics of each individual HRTF, which serve as cues for sound image localization.
[0017] One of the objectives of this invention is to render sound in the frequency domain when the sound source or listening point moves at will. [Means for solving the problem]
[0018] [1] One aspect of the present invention is an acoustic signal processing device comprising: a spherical harmonic spectrum calculation unit that calculates conversion coefficients from basis functions of a spherical harmonic expansion at a sound source position with respect to the listening point to basis functions at a sound source position with respect to the origin, and for each ear calculates a spherical harmonic spectrum of a sound field based on the spherical harmonic spectrum of the head-related transfer function, the conversion coefficients, and the sound source signal; and a binaural signal generation unit that calculates a sound pressure spectrum for each ear by linearly combining the basis functions using the spherical harmonic spectrum of the sound field, and converts the sound pressure spectrum into an acoustic signal in the time domain. According to the configuration in [1], for each ear, a spherical harmonic spectrum representing the sound field at the listening point is calculated without recalculating the spherical harmonic spectrum of the head-related transfer function in the frequency domain. By linearly combining the basis functions in the spherical harmonic expansion using the calculated spherical harmonic spectrum, it is possible to render in the frequency domain a sound in which the sound source or listening point moves arbitrarily. Through rendering, the frequency characteristics of the sound continuously change in accordance with the smooth fluctuation of the sound source position or listening point.
[0019] [2] One aspect of the present invention is the above-described acoustic signal processing device, wherein the listening point is the head position, and the reference point for acquiring the head-related transfer function is the head position. According to the configuration in [2], using the head position as the listening point simplifies the rendering operations based on the listening point compared to using the position of each ear. Furthermore, using the head position as the reference point for acquiring the head-related transfer function makes it easier to acquire and manage the head-related transfer functions for each ear simultaneously.
[0020] [3] One aspect of the present invention is the above-described acoustic signal processing device, wherein the listening point is the position of each ear, and the reference point for acquiring the head-related transfer function is the position of each ear. According to the configuration in [3], the position of each ear is used as a listening point, and that position is used as the expansion center for the spherical harmonic expansion of the head-related transfer function. Therefore, the estimation accuracy of the calculated sound pressure spectrum can be improved compared to when the head position is used as the listening point.
[0021] [4] One aspect of the present invention is a program that causes a computer to function as the above-described acoustic signal processing device. According to the configuration in [4], the position of each ear is used as the listening point, and this position is used as the expansion center for the spherical harmonic expansion of the head-related transfer function. Therefore, the estimation accuracy of the calculated sound pressure spectrum can be improved compared to when the center of the head is used as the listening point. [Effects of the Invention]
[0022] According to the present invention, sound in which the sound source or listening point moves arbitrarily can be rendered in the frequency domain. [Brief explanation of the drawing]
[0023] [Figure 1] This is a schematic block diagram showing an example of the functional configuration of the acoustic signal processing device according to the first embodiment. [Figure 2] This figure illustrates the global coordinate system and local coordinate system according to the first embodiment. [Figure 3] This is a flowchart illustrating the acoustic signal processing according to the first embodiment. [Figure 4] This figure illustrates the global coordinate system and local coordinate system according to the second embodiment. [Figure 5] This is a flowchart illustrating the acoustic signal processing according to the second embodiment. [Modes for carrying out the invention]
[0024] <First Embodiment> Embodiments of the present invention will be described below with reference to the drawings. First, an example of the functional configuration of the acoustic signal processing device 10 according to the first embodiment will be described. Figure 1 is a schematic block diagram showing an example of the functional configuration of the acoustic signal processing device 10 according to this embodiment. The acoustic signal processing device 10 receives the coordinates of the listening point and the sound source signal for each ear (left and right). The acoustic signal processing device 10 calculates conversion coefficients from the basis functions of the spherical harmonic expansion at the sound source position relative to the listening point to the basis functions at the sound source position relative to the origin. For each ear, the acoustic signal processing device 10 calculates the spherical harmonic spectrum of the sound field based on the spherical harmonic spectrum of the head-related transfer function, the conversion coefficients, and the sound source signal. For each ear, the acoustic signal processing device 10 calculates the sound pressure spectrum by linearly combining the basis functions using the spherical harmonic spectrum of the sound field, and converts the sound pressure spectrum into a time-domain acoustic signal.
[0025] The acoustic signal processing device 10 comprises an input unit 110, a control unit 120, a storage unit 130, and an output unit 140. The input unit 110 receives listening point information indicating the coordinates of the listening point and a sound source signal. The input unit 110 outputs the listening point information and the sound source signal to the control unit 120. The listening point is, for example, the position of the listener's head in three-dimensional space (hereinafter referred to as "head position"). For example, the position of the center of the head is indicated as the head position. The time series of the listening point for each time corresponds to the movement trajectory. The input unit 110 may input and output listening point information for each time, or it may input and output listening point information indicating the movement trajectory over a certain period. The input unit 110 may acquire the sound source signal for each time, or it may acquire the sound source signal for that period all at once. The input unit 110 is, for example, an input interface.
[0026] The control unit 120 performs processing to realize the functions of the acoustic signal processing device 10. The control unit 120 includes a spherical harmonic spectrum calculation unit 122, a binaural signal generation unit 124, and a spherical harmonic expansion unit 126.
[0027] The spherical harmonic spectrum calculation unit 122 calculates, for each time τ, the conversion coefficient S h from the basis function of the n-th order and m-th degree in the spherical harmonic expansion at the head position [r n m ν μ (τ)] as the listening point to the basis function of the ν-th order and μ-th degree at the sound source position r, or S’<μ The product of (θ,φ) or the ν-th order sphere Bessel function j ν (kr) and the ν-th order μ-th spherical harmonic Y ν μ This corresponds to the weighting coefficients determined such that the order of the ν-th order μ-th basis function of the spherical harmonic expansion, which is the product with (θ,φ), and the weighted sum between the orders are equal. The spherical harmonic spectrum calculation unit 122 calculates r <r s When (τ), the transformation coefficient S n m ν μ Calculate r>r s (τ) Transformation coefficient S' n m ν μ Calculate r, r s These represent the radial movement of the sound source position in the global coordinate system and the radial movement of the sound source position in the local coordinate system relative to the listening point, respectively. That is, the transformation coefficient S n m ν μ ([r h (τ)]), S' n m ν μ ([r h (τ)]) is the listening point [r h (τ)] relative to the sound source position [r s This shows the degree of contribution of each basis function in the transformation from the nth-order m-th basis function of the spherical harmonic expansion at [r] to a weighted sum of ν-order μ-th basis functions at the sound source position [r] relative to the origin.
[0028] The spherical harmonic spectrum calculation unit 122 converts the sound source signal s(τ) input from the input unit 110 using the calculated conversion coefficient S n m ν μ ([r h (τ)]) or S' n m ν μ ([r h The converted sound source signal z(t) (described later), obtained by multiplying by (τ)), is subjected to a Fourier transform to form the converted sound source spectrum Z in the frequency domain. nm ν μ Convert it to (ω). The spherical harmonic spectrum calculation unit 122 reads the spherical harmonic spectrum α n m (ω) of the HRTF for each of the left and right ears from the storage unit 130, and multiplies the read spherical harmonic spectrum α n m (ω) of the HRTF by the converted sound source spectrum to calculate the spherical harmonic spectrum P ν μ (ω) of the sound field. The spherical harmonic spectrum P ν μ (ω) of the sound field is calculated using Equation (13) (described later). The spherical harmonic spectrum calculation unit 122 outputs the spherical harmonic spectrum P ν μ (ω) of the sound field calculated for each ear to the binaural signal generation unit 124.[[ID=??]] [[ID=??]]
[0029] [[ID=??]] The binaural signal generation unit 124 uses the spherical harmonic spectrum P ν μ (ω) of the sound field input from the spherical harmonic spectrum calculation unit 122 as a weighting coefficient, and calculates a weighted sum with the product of the spherical Bessel function j ν (kr) and the spherical harmonic function Y ν μ (θ, φ) or the product of the second kind spherical Hankel function h ν (2) (kr) and the spherical harmonic function Y ν μ (θ, φ) as the sound pressure spectrum P(r, ω). The sound pressure spectrum P(r, ω) is calculated using Equation (12) (described later). The binaural signal generation unit 124 performs an inverse Fourier transform on the sound pressure spectrum P(r, ω) in the frequency domain calculated for each ear to generate an acoustic signal in the time domain as a signal for each ear. The binaural signal generation unit 124 outputs a binaural signal composed of signals for each ear as an output signal via the output unit 140.
[0030] It seems there are some tags with "??" in the translated text which might be due to the unclear nature of the original tags in that part. If there is more context or specific instructions regarding those tags, it would be possible to provide a more accurate translation.The spherical harmonic expansion unit 126 acquires HRTF for each pair of listening point and sound source position for each left and right ear, performs spherical harmonic expansion on the acquired HRTF, and generates a spherical harmonic spectrum α common to the listening point and sound source position pair. n m (ω) is calculated in advance. Each HRTF is expressed as a linear combination of the product of the nth-order second-kind spherical Hankel function and the nth-order m-th spherical harmonic function by spherical harmonic expansion as shown in equation (6) (described later). The spherical harmonic expansion section 126 uses a common weighting coefficient between the listening point and sound source position so that the dependence in the radial direction is explained by the nth-order second-kind spherical Hankel function and the dependence in the angular direction is explained by the nth-order m-th spherical harmonic function in the spherical harmonic spectrum α n m It can be calculated as (ω). The spherical harmonic expansion section 126 calculates the spherical harmonic spectrum α n m The HRTF data representing (ω) is stored in the storage unit 130.
[0031] The memory unit 130 stores data used for processing in the control unit 120 and data acquired by the control unit 120. The memory unit 130 is configured to include a storage medium such as RAM (Random Access memory) or ROM (Read Only Memory). The output unit 140 outputs the binaural signal input from the binaural signal generation unit 124 to the outside. The output unit 140 is, for example, an output interface. The output unit 140 may be integrated with the input unit 110 and configured as an input / output interface.
[0032] According to the method described above, the spherical harmonic spectrum of the HRTF based on the listening point, which is the origin of the local coordinate system, is transformed into the spherical harmonic spectrum of the HRTF based on the origin of the global coordinate system, and separated into radial and angular components. Then, the transformed source spectrum is obtained from the transformation coefficients of the basis functions from the local coordinate system to the global coordinate system and the sound source signal. Furthermore, it is transformed into a sound pressure spectrum using the transformed source spectrum and the spherical harmonic spectrum of the HRTF common to the sound source position and the listening point. As a result, the sound source signal with the HRTF convolved is interpolated in the frequency domain for any listening point and sound source position. Since the frequency characteristics of the sound (spectral cue), which are clues for sound image localization, are not disturbed by the calculation, more reliable sound image localization to the sound source position can be expected for the listener. In contrast, in the overlapping summation method described in Non-Patent Literature 4, binaural signals relating to different target directions are added in the time domain. Interference occurs between the two due to the phase difference of the HRTFs contained in each binaural signal. Due to interference-induced distortion of frequency characteristics, it was sometimes impossible to achieve sound image localization in the target direction. This embodiment can serve as a solution to this problem.
[0033] Next, the acoustic signal processing according to this embodiment will be described in more detail. This embodiment is based on the spherical harmonic expansion of the HRTF in three-dimensional space. Here, the global coordinate system, which covers the entire sound field including the playback sound source and the listener, and the local coordinate system, which has the listener's head position as the origin, will be explained using Figure 2. [r] is a vector representing an arbitrary position S from the origin O in the global coordinate system, and [r] is a vector representing the position H of the head at time t (hereinafter sometimes referred to as "head position") moving along an arbitrary trajectory. h (t)], a vector representing an arbitrary position S in a local coordinate system with head position H as the origin [r s (t)] is expressed as vector [r], [r h (t)]..r s (t) represents the three-dimensional spherical coordinates [r(t),θ(t),φ(t)],[r h (t), θ h (t), φ h (t)]..r s(t), θ s (t), φ s It is expressed as (t)).
[0034] In HRTF measurement, assuming that position S is the sound source position where the measurement speaker is installed, the transfer function showing the frequency domain transfer characteristics of sound from sound source position S to head position H corresponds to HRTF, and the impulse response that expresses these transfer characteristics in the time domain corresponds to HRIR. In this embodiment, the movement of the head position H as the listening point is described as "the movement of a head with directivity represented by a spherical harmonic spectrum" obtained by expanding the HRTF with spherical harmonics. Note that since the HRTF is determined by the relative positional relationship between the sound source and the head, the movement of the target point may be considered instead of the listening point. The target point refers to the position where the sound source is virtually installed and which serves as the target for sound image localization.
[0035] Assume that in the global coordinate system, the listening point H moves along an arbitrary trajectory, and sound is emitted from the sound source position S. In this case, the listening point H (coordinates [r]) at time τ is relative to the sound source position S (coordinates [r]). h (τ)]) Impulse response g([r]‐[r h (τ)]) is a time-varying impulse response. As shown in equation (4), the binaural signal p([r],t) at time t is an impulse response g([r]‐[r h It is obtained by performing a convolution operation using (τ)]). The spectrum of the binaural signal in the frequency domain is obtained by performing a Fourier transform on the binaural signal p([r],t) in the time domain, as shown in equation (5).
[0036]
number
number
[0037] Since HRTF is measured using a microphone placed on the head, it is represented as a distribution around the head by performing a spherical harmonic expansion in a local coordinate system with the head position as the origin. Sound source position S in the local coordinate system (coordinate [r s ]) Head position H (coordinates [r h The transfer function in the frequency domain at ]), i.e., the HRTF, is given by G([r s ]‐[r h It is expressed as (τ)],ω). According to the spherical harmonic expansion, as shown in equation (6), the HRTF is transformed into a linear combination of basis functions that are the product of the nth-order second kind spherical Hankel function and the nth-order m-th spherical harmonic function, that is, a weighted sum that crosses order and degree. Therefore, the HRTF is the spherical harmonic spectrum α, which consists of weight coefficients multiplied by each basis function. n m It is represented by (ω). Since HRTF is acquired separately for the left and right ears, the spherical harmonic spectrum α n m The shape of (ω) is also different in the left and right ears.
[0038]
number
[0039] According to the spherical harmonic expansion shown in equation (6), the head position H, which is the center of the expansion, changes over time. Therefore, the orthogonal basis functions, the second kind spherical Hankel function and the spherical harmonics, also need to be recalculated each time the head position H moves. In this embodiment, the center of the expansion is shifted to the origin of the global coordinate system using the addition theorem of spherical Bessel functions, and the sound field is represented using orthogonal basis functions related to a fixed sound source position, regardless of the head position H. The addition theorem of spherical Bessel functions states that when [r2] = [r1] + [b] (where b is the coordinate between the three-dimensional coordinates [r1] and [r2]) between the three-dimensional coordinates [r1] (= [r1, θ1, φ1]) and [r2] (= [r2, θ2, φ2]), the relationship shown in equation (7) holds. In equation (7), the transformation coefficient S n m ν μ ([b]), S' n mν μ ([b]) is given by equations (8) and (9), respectively.
[0040]
number
[0041]
number
[0042]
number
[0043] In equations (8) and (9), Y * q μ-m (θ b ,φ b ) is a spherical harmonic function Y q μ-m (θ b ,φ b The complex conjugate of ) is shown. W1 and W2 represent the Wigner 3j notation in equation (10).
[0044]
number
[0045] By applying the addition theorem for spherical Bessel functions to the HRTF shown in equation (6) and shifting the expansion center of the spherical harmonics to the origin O of the global coordinate system, the HRTF can be expressed as a linear combination of basis functions relating to the sound source position in the global coordinate system. Specifically, the three-dimensional coordinates [r1] and [b] in equation (7) are replaced with the coordinates of the sound source position [r] and the head position [r], respectively, with respect to the origin O in the global coordinate system. h By substituting (τ) into equation (6) and applying it, the HRTF is transformed into equation (11). Then, substituting the HRTF expressed in equation (11) into equation (5), the sound pressure spectrum P([r],ω) is given as shown in equation (12).
[0046]
number
[0047]
number
[0048] In equation (12), the spherical Bessel function j ν (kr) and spherical harmonics Y ν μ The spherical harmonic spectrum P of the sound field multiplied by a basis function that is the product of (θ,φ) ν μ (ω) is given by equation (13). The above converted sound source signal z(τ) is given by the sound source signal s(τ) in equation (13) and the conversion coefficient S n m ν μ (r h This corresponds to the product with (τ). The time-domain converted sound source signal z(τ) is the frequency-domain converted sound source spectrum Z n m ν μ It is converted to (ω). The converted sound source spectrum Z is shown in equation (13). n m ν μ (ω) is further represented by the spherical harmonic spectrum α of the HRTF. n m It can be superimposed on (ω).
[0049] Spherical harmonic spectrum P of the sound field ν μ (ω) is given by the origin O of the global coordinate system as the expansion center. Therefore, the spherical harmonic spectrum calculation unit 122 uses equation (13) to calculate the coordinates [r] of the sound source position given in the global coordinate system and the coordinates [r] of the listening point. h Based on (τ), the spherical harmonic spectra of the HRTF are sequentially determined α n m Without calculating (ω), the spherical harmonic spectrum P of the sound field ν μ(ω) can be calculated. The binaural signal generation unit 124 can then generate a binaural signal as an output signal by performing an inverse Fourier transform on the spectrum of the binaural signal given by equation (12).
[0050]
number
[0051] Next, an example of acoustic signal processing according to this embodiment will be described. Figure 3 is a flowchart illustrating acoustic signal processing according to this embodiment. (Step S102) The input unit 110 receives the coordinates of the listening point and the sound source signal. (Step S104) The spherical harmonic spectrum calculation unit 122 calculates, for each frequency, the conversion coefficients from the basis functions of the spherical harmonic expansion at the sound source position with the center of the head as the reference point as the listening point to the basis functions at the sound source position with the origin as the expansion center, at each time step. (Step S106) The spherical harmonic spectrum calculation unit 122 converts the converted sound source signal obtained by multiplying the sound source signal by a conversion coefficient for each ear into a converted sound source spectrum in the frequency domain. The spherical harmonic spectrum calculation unit 122 calculates the spherical harmonic spectrum of the sound field by multiplying the converted sound source spectrum by the spherical harmonic spectrum of the HRTF according to equation (13).
[0052] (Step S108) The binaural signal generation unit 124 calculates the sound pressure spectrum for each ear by linearly combining the basis functions of the spherical harmonic expansion using the spherical harmonic spectrum of the sound field according to equation (12). As a linear combination, the weighted sum of the basis functions of the spherical harmonic expansion, with the spherical harmonic spectrum of the sound field as the weight coefficient, is obtained as the sound pressure spectrum. (Step S110) The binaural signal generation unit 124 converts the frequency domain sound pressure spectrum for each ear into a time domain output signal and outputs the converted output signal to the output unit 140. After that, the process shown in Figure 3 is completed.
[0053] <Second Embodiment> Next, a second embodiment will be described. The following description will mainly focus on the differences from the first embodiment, and for similarities, refer to the description in the first embodiment. According to spherical harmonic expansion, the closer a location is to the expansion center, the more accurately the sound field spectrum can be estimated. Typically, in binaural playback, the listener's head position is used as the listening point. However, the entrances to the external auditory canals of each ear are used as measurement points for HRTF. The entrances to the external auditory canals are located about 7-10 cm away from the head position. This deviation of the measurement points from the head position can cause a decrease in the accuracy of the sound field spectrum. Therefore, in this embodiment, the expansion center of the spherical harmonic expansion is set to the position of each ear. This is expected to improve the accuracy of sound field spectrum estimation.
[0054] Figure 4 shows the relationship between the global coordinate system and the local coordinate system according to this embodiment. However, the left ear is used as an example. The spherical harmonic spectrum calculation unit 122 receives listening point information indicating the head position via the input unit 110. The spherical harmonic spectrum calculation unit 122 can, for example, determine the position of the left ear as a position located a predetermined distance from the center of the head and to the left in a predetermined direction of the head.
[0055] Here, the vector indicating the position of the left ear E relative to the head position H is [r e ](=[r e ,θ e ,φ e ]), the vector of the sound source position S with respect to the left ear E is [r' s (τ)](=[r' s (τ),θ' s (τ), φ' s (τ)]), the vector of the position of ear E relative to the origin O in the global coordinate system is [r' h (τ)](=[r' h (τ),θ' h (τ), φ' h This is expressed as (τ). Using the addition theorem for spherical Bessel functions, in the spherical harmonic expansion of the HRTF expressed by equation (6), if the expansion center is shifted from the head position H to the left ear E, the HRTF G from the sound source to the left ear is obtained. L ([r's ]‐[r' h (τ)],ω) are given as shown in equation (14).
[0056]
number
[0057] In equation (14), the second type of sphere Hankel function j ν (2) (kr' s (τ) and spherical harmonics Y ν μ (θ' s (τ), φ' s The spherical harmonic spectrum of the left ear β multiplied by the basis function that is the product of (τ)) ν μ (ω) is given by equation (15). Coordinates of left ear E [r e ] may be predetermined by the distance and direction from the center of the head. In that case, the spherical harmonic development section 126 is the coordinates of the left ear E with respect to the head position H [r e ] Conversion coefficient S n m ν μ ([r e The conversion coefficient S calculated using equation (8) is obtained by calculating the ]) n m ν μ ([r e ]) and the spherical harmonic spectrum α of the HRTF in the left ear n m (ω) is the spherical harmonic spectrum β n m It may be corrected to (ω). The spherical harmonic expansion section 126 uses the same method for the right ear to obtain the spherical harmonic spectrum β of the HRTF of the right ear. n m (ω) can be corrected. The spherical harmonic expansion section 126 calculates the corrected spherical harmonic spectrum β for each ear. n m HRTF data representing (ω) is stored in the storage unit 130 beforehand.
[0058]
number
[0059] Equation (14) and the spherical harmonic spectrum β of Equation (15) ν μ Substituting (ω), the HRTF G related to the left ear is L ([r' s ]‐[r' h (τ)],ω) is transformed as shown in equation (16). The transformed HRTF is the spherical harmonic spectrum α of the HRTF in equation (6). n m (ω) is replaced with the corrected spherical harmonic spectrum β ν μ Except for the use of (ω), it has the same form as the HRTF shown in equation (6). This correction of the spherical harmonic spectrum can also be considered as a shift of the reference point for obtaining the HRTF from the center of the head to the position of the left ear.
[0060] Next, using the addition formula for spherical harmonics, we shift the expansion center in the spherical harmonic expansion of the HRTF in equation (16) to the origin O in the global coordinate system. In equation (7), the three-dimensional coordinates [r1] and [b] are replaced with the coordinates of the sound source position [r] and the position of the left ear [r'], respectively, with respect to the origin O in the global coordinate system. h By substituting (τ) into equation (16) and applying it, the HRTF is transformed into equation (17). Then, substituting the HRTF expressed in equation (17) into equation (5), the sound pressure spectrum P of the left ear signal is obtained as shown in equation (18). L ([r],ω) is obtained.
[0061]
number
[0062]
number
[0063]
number
[0064] In equation (18), the ηth-order sphere Bessel function j η (kr) and the ηth-order ξ-th spherical harmonic Y η ξ An η-th order ξ-th basis function that is the product of (θ,φ), or an η-th order second kind sphere Hankel function h η (2) (kr) and the ηth-order ξ-th spherical harmonic Y η ξ The spherical harmonic spectrum P of the sound field is obtained by multiplying it by an η-th order ξ-th basis function which is the product of (θ,φ). Lη ξ (ω) is expressed by equation (19).
[0065]
number
[0066] Therefore, the spherical harmonic spectrum calculation unit 122 calculates the coordinates [r] of the sound source position S given in the global coordinate system and the coordinates [r'] of the left ear E as the listening point. h (τ)], and the spherical harmonic spectrum β of the left ear E read from the memory unit 130. ν μ Using (ω), the spherical harmonic spectrum P of the sound field. Lη ξ (ω) can be calculated. The binaural signal generation unit 124 can generate the left ear signal by performing an inverse Fourier transform on the left ear spectrum given by equation (18).
[0067] The spherical harmonic spectrum calculation unit 122 calculates the spherical harmonic spectrum P for the right ear using the same method as for the left ear. Rη ξ (ω) can be calculated. The binaural signal generation unit 124 also uses the same method for the right ear as for the left ear to calculate the spherical harmonic spectrum P Rη ξBy calculating the sound pressure spectrum for the right ear from (ω) and the sound source signal s(τ), and then performing an inverse Fourier transform on the calculated sound pressure spectrum, a signal for the right ear can be generated.
[0068] Next, an example of acoustic signal processing according to this embodiment will be described. Figure 5 is a flowchart illustrating acoustic signal processing according to this embodiment. (Step S122) The input unit 110 receives the coordinates of the center of the head and the sound source signal as the listening point. (Step S124) The spherical harmonic spectrum calculation unit 122 calculates, for each frequency, the conversion coefficients from the basis functions of the spherical harmonic expansion at the sound source position, with the position of each ear as the listening point as the reference, to the basis functions at the sound source position, with the origin as the expansion center as the reference, at each time step. (Step S126) The spherical harmonic spectrum calculation unit 122 converts the converted sound source signal obtained by multiplying the sound source signal by a conversion coefficient for each ear into a converted sound source spectrum in the frequency domain. The spherical harmonic spectrum calculation unit 122 calculates the spherical harmonic spectrum of the sound field by multiplying the converted sound source spectrum obtained using equation (19) by the corrected spherical harmonic spectrum.
[0069] (Step S128) The binaural signal generation unit 124 calculates the sound pressure spectrum for each ear by linearly combining the basis functions of the spherical harmonic expansion using the spherical harmonic spectrum of the sound field according to equation (18). (Step S130) The binaural signal generation unit 124 converts the frequency domain sound pressure spectrum for each ear into a time domain output signal and outputs the converted output signal to the output unit 140. After that, the process shown in Figure 5 is completed.
[0070] In the above explanation, it is assumed that the head position is used as the listening point for measuring HRTF, and that a pair of HRTFs (left and right) are associated with the central position of each head. When the position of each ear and the HRTF measured for that ear are associated as the reference for measuring HRTF, the spherical harmonic expansion unit 126 performs a spherical harmonic expansion on the HRTF measured for each sound source position for each ear, and calculates the spherical harmonic spectrum α of the HRTF related to that ear. n m You may also calculate (ω). In that case, the spherical harmonic spectrum α n m (ω) is the corrected spherical harmonic spectrum β for that ear. ν μ This corresponds to (ω). Therefore, in steps S124 and S126, the spherical harmonic spectrum calculation unit 122 calculates the spherical harmonic spectrum α. n m (ω) Corrected spherical harmonic spectrum β ν μ (ω) can be used in place of (ω). Also, the binaural signal generation unit 124 also calculates the spherical harmonic spectrum α in step S128. n m (ω) Corrected spherical harmonic spectrum β ν μ You can use it by substituting (ω).
[0071] Generally, the HRTF (Head-Reaching Frequency Response) is determined by the relative positional relationship between the sound source and the listening point. While the above description uses the case where the listening point moves and the sound source is stationary as an example, it is not limited to this case. This embodiment can also be applied, for example, when the listening point is stationary and the sound source is moving.
[0072] The above description uses the example of the spherical harmonic spectrum calculation unit 122 and the binaural signal generation unit 124 generating one binaural signal for one sound source, but is not limited to this. The input unit 110 may receive input for each of multiple sound sources, associated with listening point information and sound source signals. The input unit 110 may also receive input associated with sound source position information indicating the sound source position of each sound source. The control unit 120 may then have the sound source position set for each individual sound source. The spherical harmonic spectrum calculation unit 122 and the binaural signal generation unit 124 may generate binaural signals for each sound source given a sound source signal and sound source position, and the generated binaural signals may be mixed between the sound sources to obtain an acoustic signal which is output as an output signal via the output unit 140.
[0073] When the control unit 120 acquires a sound source signal and listening point information from an external device, it may read the sound source signal and listening point information previously stored in the storage unit 130 instead of using the input unit 110. Also, instead of outputting the output signal to an external device using the output unit 140, the control unit 120 may store it in the storage unit 130, or output it to a playback unit (speaker) installed in or connected to the sound signal processing device 10 and emit sound. In the acoustic signal processing device 10, the spherical harmonic expansion unit 126 may be omitted. The storage unit 130 may have HRTF data acquired from an external device pre-stored in it.
[0074] As described above, the acoustic signal processing device 10 according to this embodiment comprises a spherical harmonic spectrum calculation unit 122 and a binaural signal generation unit 124. The spherical harmonic spectrum calculation unit 122 calculates conversion coefficients from the basis functions of the spherical harmonic expansion at the sound source position with respect to the listening point to the basis functions at the sound source position with respect to the origin, and calculates the spherical harmonic spectrum of the sound field for each ear based on the spherical harmonic spectrum of the head-related transfer function, the conversion coefficients, and the sound source signal. The binaural signal generation unit 124 calculates the sound pressure spectrum of the output signal for each ear by linearly combining the basis functions using the spherical harmonic spectrum of the sound field, and converts the sound pressure spectrum into a time-domain output signal. With this configuration, for each ear, the spherical harmonic spectrum representing the sound field at the listening point is calculated without recalculating the spherical harmonic spectrum of the head-related transfer function in the frequency domain. By linearly combining the basis functions in the spherical harmonic expansion using the calculated spherical harmonic spectrum, it is possible to render sound in the frequency domain with the sound source or listening point moving arbitrarily. Through rendering, the frequency characteristics of the sound continuously change in accordance with the smooth fluctuations of the sound source position or listening point. Furthermore, since it does not involve the addition of the head-related transfer function or the sound source signal convolved with the head-related transfer function in the time domain, the sound pressure spectrum of the sound field at the listening point is estimated without disturbing the frequency characteristics that provide clues for sound image localization. Therefore, sound image localization to that sound source position is not hindered.
[0075] Furthermore, the listening point and the reference point for acquiring the head-related transfer function may both be head positions. By using the head position as the listening point, the rendering operations related to the listening point can be simplified compared to when the position of each ear is used. Furthermore, by using the head position as the reference point for acquiring head-related transfer functions, it becomes easier to acquire and manage head-related transfer functions for each ear simultaneously.
[0076] Furthermore, the listening point and the reference point for acquiring the head-related transfer function may be the position of each ear, respectively. By using the position of each ear as the listening point, that position is used as the expansion center for the spherical harmonic expansion of the head-related transfer function. Therefore, the estimation accuracy of the calculated sound pressure spectrum can be improved compared to when the center of the head is used as the listening point.
[0077] The acoustic signal processing device 10 may be implemented as a dedicated acoustic signal processing device, or it may be implemented as a device whose primary function is not the processing of acoustic signals, such as a personal computer or tablet device. The acoustic signal processing device 10 may also be implemented as part of equipment (for example, a mixing console) related to the production, editing, and distribution (including broadcasting) of various types of content.
[0078] Furthermore, some or all of the above-described acoustic signal processing device 10 may be configured using dedicated components (such as integrated circuits) or implemented using a computer. For example, either the spherical harmonic spectrum calculation unit 122 or the binaural signal generation unit 124, or a combination thereof, may be implemented by a general-purpose arithmetic processing unit such as a CPU (Central Processing Unit) executing processes instructed by commands written in a predetermined program read from a storage medium such as ROM (Read Only Memory).
[0079] Although one embodiment of this invention has been described in detail above with reference to the drawings, the specific configuration is not limited to that described above, and various design changes can be made without departing from the spirit of this invention. [Explanation of Symbols]
[0080] 10...Acoustic signal processing unit, 110...Input unit, 120...Control unit, 122...Spherical harmonic spectrum calculation unit, 124...Binaural signal generation unit, 126...Spherical harmonic expansion unit, 130...Storage unit, 140...Output unit
Claims
1. The conversion coefficients from the basis functions of the spherical harmonic expansion at the sound source position relative to the listening point to the basis functions at the sound source position relative to the origin are calculated. For each ear, a spherical harmonic spectrum calculation unit calculates the spherical harmonic spectrum of the sound field based on the spherical harmonic spectrum of the head-related transfer function, the conversion coefficient, and the sound source signal. For each ear, the sound pressure spectrum is calculated by linearly combining the basis functions using the spherical harmonic spectrum of the sound field. The system includes a binaural signal generation unit that converts the sound pressure spectrum into a time-domain acoustic signal. Acoustic signal processing device.
2. The receiving point is the head position, The reference point for obtaining the head transfer function is the head position. The acoustic signal processing device according to claim 1.
3. The aforementioned listening points are the positions of each ear. The reference points for obtaining the aforementioned head-related transfer function are the positions of each ear. The acoustic signal processing device according to claim 1.
4. A program for causing a computer to function as an acoustic signal processing device according to any one of claims 1 to 3.
Citation Information
Patent Citations
IEC23008-3
Silver halide photosensitive material
JP1985067934A
Signal processing device and method, and program
WO2020255810A1