Acoustic signal processing method, acoustic signal processing apparatus, and acoustic signal processing program
The method optimizes HRIRs for binaural rendering by learning representative directions, reducing bias and degradation in synthesized sound quality in VR and AR devices.
Patent Information
- Application Number
- JP2024100096
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-30
- Filing Date
- 2024-06-21
- Publication Date
- 2025-08-12
AI Technical Summary
Existing acoustic signal processing methods using head-related impulse responses (HRIRs) for binaural rendering in VR and AR devices suffer from biased distortion due to artificial selection of representative directions, leading to degradation in synthesized sound quality.
An acoustic signal processing method that acquires sound source direction, pans the signal by time shifting and gain adjusting, and allocates it to learned representative HRIRs, minimizing error and distortion through a cost function that optimizes HRIRs for binaural rendering.
Reduces bias and degradation in synthesized sound by optimizing representative HRIRs, improving sound quality and reducing computational load in VR and AR devices.
Smart Images

Figure 2025117508000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention particularly relates to an acoustic signal processing method, an acoustic signal processing device, and an acoustic signal processing program. [Background technology]
[0002] VR headphones and HMDs (Head Mounted Displays) capable of playing content such as movies, VR (Virtual Reality), and AR (Augmented Reality) have been available for some time. In these VR headphones and HMDs, a head-related transfer function (HRTF) is used to generate out-of-head stereophonic sound (binaural rendering) in order to create a wider sound field. This function takes into account the direction from the listener to the sound source. When playing binaurally rendered audio using HRTFs on headphones, etc., the calculation of the actual acoustic signal often uses a head-related impulse response (hereinafter referred to as "HRIR"), which represents the head-related transfer function on the time axis.
[0003] As a typical device using HRIR, Patent Document 1 describes a device for generating 3D sound using HRIR that reduces the computational load even when there are a large number of sound sources (hereinafter referred to as "prior art"). In the prior art, multiple sound sources (target signals) are grouped into a smaller number of representative directions, and sound images are synthesized using only the HRIRs of the representative directions, thereby reducing the amount of computation required to generate signals at the ears. In this conventional technology, the representative directions for grouping HRIRs are artificially selected at predetermined intervals from a predetermined set of HRIRs covering the entire celestial sphere, such as six directions at 60-degree intervals in the horizontal plane, the zenith, the nadir, etc. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 2023-164284 Summary of the Invention [Problem to be solved by the invention]
[0005] In the prior art, the representative direction is selected artificially, and the distortion of the synthesized playback signal, which is dependent on the selection, can be biased depending on the direction from which the sound source signal arrives. Therefore, there has been a technical demand to reduce this bias and further reduce the degradation caused by synthesis by panning.
[0006] The present invention has been made in view of the above circumstances, and aims to solve the above-mentioned problems. [Means for solving the problem]
[0007] The acoustic signal processing method of the present invention is an acoustic signal processing method executed by an acoustic signal processing device, characterized in that it acquires the sound source direction of a sound source signal, pans the sound source signal by time shifting and gain adjusting the sound source signal based on the acquired sound source direction and allocating it to a plurality of representative head-impulse responses (HRIRs), and performs binaural rendering by convolving the plurality of representative HRIRs with the allocated signals, each of which is calculated by learning. The acoustic signal processing method of the present invention is characterized in that each of the plurality of representative HRIRs is trained based on a cost function that minimizes the expected value of the error between a synthetic HRIR obtained by panning all around the celestial sphere or a horizontal plane and a true HRIR. The acoustic signal processing method of the present invention is characterized in that each of the plurality of representative HRIRs is learned based on a cost function that includes a term that minimizes the expected value of the error between a synthetic HRIR obtained by panning all around the celestial sphere or the horizontal plane and a true HRIR, and a term that minimizes the expected value of the change in error between adjacent HRIRs in the azimuth and / or elevation directions. The acoustic signal processing method of the present invention is characterized in that the expected value of the error is an average value of the error. The acoustic signal processing method of the present invention is characterized in that the expected value of the error uses an average value of the ratio of the power of the error to the power of the signal. The acoustic signal processing method of the present invention is characterized in that each of the plurality of representative HRIRs is different for the left ear signal and the right ear signal of the binaural rendering. In the acoustic signal processing method of the present invention, the cost function includes a distortion measure that adds, as a penalty, a change in the gain by which the HRIR of each representative direction is multiplied in the learning. The acoustic signal processing method of the present invention is characterized in that the distortion measure is used during the learning and / or reproduction. In the acoustic signal processing method of the present invention, the gain is calculated by the following equation (13):
number
number
number
[0008] According to the present invention, when performing binaural rendering by panning based on the acquired sound source direction, an acoustic signal processing method can be provided that uses HRIRs calculated by learning as the representative direction of a specific representative direction and optimizes the representative direction used for panning itself, thereby reducing bias due to HRIR distortion and reducing the degradation caused by synthesis due to panning more than conventional techniques. [Brief explanation of the drawings]
[0009] [Figure 1]1 is a control configuration diagram of an acoustic signal processing device according to a first embodiment of the present invention. [Figure 2] FIG. 2 is a conceptual diagram showing the concept of synthesis of reproduced sounds by panning shown in FIG. [Figure 3] 4 is a flowchart of a learning process according to the first embodiment of the present invention. [Figure 4] FIG. 5 is a conceptual diagram of an initial value of a representative direction in the learning process shown in FIG. [Figure 5] 3 is a flowchart of a reproduction process according to the first embodiment of the present invention. [Figure 6] FIG. 10 is a control configuration diagram of an acoustic signal processing device according to another embodiment of the present invention. [Figure 7] 1 is a graph of SNR according to Example 1 of the present invention. [Figure 8] 1 is a graph showing a mapping of HRIRs in a representative direction (elevation angle 0°) according to Example 1 of the present invention. [Figure 9] 1 is a graph showing a mapping of HRIRs in a representative direction (elevation angle 46°) according to Example 1 of the present invention. [Figure 10] 1 is a graph showing a mapping of HRIR in a representative direction (elevation angle -46°) according to Example 1 of the present invention. [Figure 11] FIG. 2 is a conceptual diagram illustrating generation of a moving sound source according to the first embodiment of the present invention. [Figure 12] 1 is a graph of a moving sound source waveform (elevation angle 0°) according to Example 1 of the present invention. [Figure 13] 1 is a graph of a moving sound source waveform (elevation angle 46°) according to Example 1 of the present invention. [Figure 14] 1 is a graph of a moving sound source waveform (elevation angle -46°) according to Example 1 of the present invention. [Figure 15] FIG. 2 is a plan view in which HRIRs in representative directions according to Example 1 of the present invention are plotted. [Figure 16] FIG. 2 is a rear view in which HRIRs in representative directions according to Example 1 of the present invention are plotted. [Figure 17] 10 is a graph of the gain of the left ear (elevation angle 0) of a synthetic HRIR (α=1, β=0) according to Example 2 of the present invention. [Figure 18] 10 is a graph of a moving sound source waveform (elevation angle 0°) of a synthetic HRIR (α=1, β=0) according to Example 2 of the present invention. [Figure 19] 10 is a graph of the gain of the left ear (elevation angle 0) of a synthetic HRIR (α=0.8, β=0.2) according to Example 2 of the present invention. [Figure 20] 10 is a graph of a moving sound source waveform (elevation angle 0°) of a synthetic HRIR (α=0.8, β=0.2) according to Example 2 of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0010] First Embodiment [Control configuration of the acoustic signal reproducing device 1] First, with reference to FIG. 1, the control configuration of the acoustic signal reproducing device 1 according to the first embodiment of the present invention will be described.
[0011] The acoustic signal reproducing device 1 is a device worn by a listener and capable of reproducing audio, such as reproducing the acoustic signal (sound signal) of content such as video, audio, and text data, or for making calls between remote locations. Specifically, the audio signal reproducing device 1 may be, for example, a stereophonic reproduction device using a PC (Personal Computer) or smartphone with headphones connected, a dedicated game console, a content reproducing device that reproduces content stored on an optical medium or a flash memory card, equipment at a movie theater or public viewing venue, headphones equipped with a dedicated decoder and head tracking sensor, an HMD (Head-Mounted Display) for VR (Virtual Reality), AR (Augmented Reality), or MR (Mixed Reality), a headphone-type smartphone, a television (video) conferencing system, remote conferencing equipment, an audio hearing aid, a hearing aid, or other home appliances.
[0012] The acoustic signal reproducing device 1 according to this embodiment includes a direction acquisition unit 10, a panning unit 20, a learning unit 30, an output unit 40, and a reproducing unit 50 as a control configuration. In this embodiment, the direction acquisition unit 10, the panning unit 20, and the learning unit 30 are configured as an acoustic signal processing device 2 that generates an acoustic signal.
[0013] In this embodiment, stereophonic sound is generated from sound source signals S-1 to Sn, which are multiple acoustic signals (sound source signals, target signals). Any of the multiple sound source signals S-1 to Sn will be simply referred to as a "sound source signal S" below. As the sound source signal S according to this embodiment, an audio signal of the content, an audio signal of a remote call participant, or the like can be used.
[0014] This content may be various types of content, such as games, movies, VR, AR, MR, etc. The movies also include musical instrument performances, lectures, etc. In this case, the sound source signal S may be an audio signal (hereinafter referred to as an "audio object") derived from an object such as an instrument, vehicle, game character, etc. (hereinafter simply referred to as an "object, etc."), or an audio signal from a human being who is the source of sound, such as an actor, narrator, storyteller, storyteller, or other speaker. A spatial positional relationship is set for these audio signals within the content.
[0015] Alternatively, when the sound source signal S is an acoustic signal of a remote call participant, it may be an acoustic signal uttered by a user (participant) of various messengers or video conferencing application software (hereinafter simply referred to as "application") on a personal computer (PC) or smartphone. This acoustic signal may be acquired by a microphone such as a headset, or may be acquired by a device fixed to a desk or the like. Directional information may include the orientation of the participant's head within a camera, or the orientation of an avatar placed in a virtual space. Furthermore, the sound source signal S may be an acoustic signal of a remote conference participant in a one-to-one, one-to-many, or many-to-many video conferencing system between locations. In this case, the orientation of each call participant relative to the camera may also be set as direction information.
[0016] In either case, an audio signal recorded by a network or directly connected microphone can also be used as the sound source signal S. In this case, directional information may also be added to the audio signal. Alternatively, any combination of the audio signals of the above-mentioned contents and remote participants may be used. Furthermore, in this embodiment, the acoustic signal of the sound source signal S also serves as a "target signal" for reproducing the direction of the stereophonic sound.
[0017] The direction acquisition unit 10 acquires the sound source direction of the sound source signal S. In this embodiment, the direction acquisition unit 10 acquires the direction of the sound source signal S relative to the front direction of the listener. Furthermore, the direction acquisition unit 10 may acquire the direction of the listener relative to the radiation direction of the sound source signal S. Specifically, the direction obtaining unit 10 obtains the direction of the sound source signal S as seen from the listener. Alternatively, the direction obtaining unit 10 may obtain the direction of the listener as seen from the sound source signal S.
[0018] Here, for the sound source signal S according to this embodiment, directional information when uttering a voice is calculated or set. Therefore, the direction acquisition unit 10 acquires the emission direction of the sound from the sound source signal S. In this embodiment, for example, the direction acquisition unit 10 can acquire the direction of the head of the participant, which is the sound source signal S. Furthermore, the direction acquisition unit 10 can also acquire the head direction of the listener from head tracking using a gyro sensor or the like of an HMD or smartphone, directional information such as the orientation of an avatar in a virtual space, etc. Based on this directional information, the direction acquisition unit 10 can calculate the directions of the sound source signal S and the listener relative to each other in a spatial arrangement including the virtual space.
[0019] The panning unit 20 pans each of the multiple sound source signals S (target signals) acquired by the direction acquisition unit 10 by allocating the sound source signals S to multiple HRIRs of representative directions (hereinafter simply referred to as "representative HRIRs" or "specific representative HRIRs") through time shifting and gain adjustment based on the sound source directions acquired by the direction acquisition unit 10. That is, the panning unit 20 generates a playback signal (signal) that combines the sound source signals S (target signals) by panning using a representative direction that approximates the sound source direction of the sound source signals S. The panning unit 20 performs binaural rendering by convolving the representative HRIRs with the allocated signals. In this way, the panning unit 20 can group the sound source signals S into a number of representative directions that is fewer than the number of sound source signals S, and synthesize a sound image using only the representative HRIRs for these representative directions. This reduces the amount of calculation required to generate signals at the ears.
[0020] In this embodiment, "equivalent" and "equivalently" refer to signals with an error below a specific level and that are substantially similar. Specifically, the panning unit 20 pans the sound source signal S to synthesize HRIRs from several directions that are closest to the sound source direction of the sound source signal S or that are most similar to the HRIR for the sound source direction, thereby generating a synthesized HRIR for that direction equivalently. In this embodiment, this direction will be referred to as a "specific representative direction" (hereinafter simply referred to as a "representative direction"), as described below. Furthermore, in this embodiment, the representative HRIR is calculated through learning by the learning unit 30. In this case, for example, the representative HRIR can use two horizontal directions or three directions including two horizontal directions and an elevation angle as the representative direction, as will be described in detail later.
[0021] In this embodiment, the panning unit 20 calculates, for each sound source direction of the sound source signal S, a gain value to be multiplied when allocating the sound source signal to the HRIR of the representative direction, and a time shift value corresponding to the time shift set in the sound source signal S to be allocated, and can store these in an HRIR table 200, which will be described later. Then, the panning unit 20 time-shifts each sound source signal S by a time shift value and gain value corresponding to the sound source direction of each sound source signal S, multiplies the result by a gain, and adds up the sound source signals S assigned to each representative HRIR to generate a sum signal. The panning unit 20 treats this sum signal as being present at a position corresponding to the representative HRIR. The panning unit 20 can convolve the representative HRIR into this sum signal to generate a signal at the listener's ear.
[0022] To generate a specific representative HRIR, the learning unit 30 iteratively learns a specific representative direction from an initial value and stores the learned direction in the representative direction information 210. In the learning unit 30 according to this embodiment, the representative HRIR is learned based on a cost function that minimizes the expected value of the error between the true HRIR and a synthesized HRIR obtained by panning all around the celestial sphere or the horizontal plane. The expected value of this error may be the mean value of the error. This error is expressed as an error vector ei It may be expressed as: Furthermore, the particular representative HRIR may be different for the left and right ear signals of a binaural rendering.
[0023] Here, when the sound source is an impulse, minimizing the difference (distortion) between the original playback signal (true playback signal) generated by convolving the HRIR of the sound source direction with the sound source signal S and the playback signal generated using panning according to this embodiment is mathematically equivalent to minimizing the error (distortion) between the true HRIR in that direction and the synthesized HRIR obtained by adding together the representative HRIR of this embodiment that has been time-shifted and gain-adjusted. Therefore, in this embodiment, it is possible to optimize the representative HRIR by minimizing the latter distortion.
[0024] Here, the learning according to this embodiment may be learning in a broad sense performed by a computer (control calculation unit), such as various types of machine learning, statistical methods including clustering, heuristic methods, etc. In this embodiment, the learning unit 30 may update the representative direction using an LBG algorithm (Linde-Buzo-Gray algorithm) based on alternating learning.
[0025] The output unit 40 outputs the acoustic signal generated by the acoustic signal processing device 2. In this embodiment, the output unit 40 includes, for example, a D / A converter, a headphone amplifier, etc., and outputs a binaurally rendered acoustic signal as a playback acoustic signal for the playback unit 50, which is a headphone. Here, the playback acoustic signal may be, for example, an acoustic signal that can be heard by a listener by decoding digital data based on information included in the content and playing it back in the playback unit 50. Alternatively, the output unit 40 may play back the acoustic signal by encoding it and outputting it as an audio file or streaming audio.
[0026] The reproduction unit 50 reproduces the reproduction sound signal output by the output unit 40. The reproduction unit 50 may also include a speaker (hereinafter referred to as a "speaker, etc.") equipped with an electromagnetic driver and diaphragm of headphones or earphones, or earmuffs or earpieces worn by a listener. Alternatively, the reproduction unit 50 may be capable of outputting the digital reproduction sound signal as a digital signal or converting it into an analog sound signal using a D / A converter from a speaker or the like so that the listener can hear it. Alternatively, the reproduction unit 50 may output the sound signal separately to headphones or earphones of an HMD worn by the listener.
[0027] The HRIR table 200 contains data of representative HRIRs used by the panning unit 20. Furthermore, the HRIR table 200 includes values for generating composite HRIRs calculated by the panning unit 20, which will be described later. Specifically, the HRIR table 200 includes, for example, gain values calculated for each representative direction for each sound source direction at 2° intervals over a full 360°. For example, when panning in two directions (left and right) is performed for two representative directions, two gain values (A value, B value) may be used for each sound source direction, and when panning in three directions including the elevation angle direction is performed, three gain values (A value, B value, C value) may be used. Furthermore, the HRIR table 200 may also include a time shift value for time-shifting the sound source signal S. The HRIR table 200 can store this time shift value in association with the gain value. These gain values and time shift values may be calculated offline in advance by the learning unit 30 or the like. Alternatively, these gain values and time shift values may be calculated when the acoustic signal reproducing device 1 according to this embodiment is turned on or initialized.
[0028] In this embodiment, as described above, the HRIR obtained by time-shifting the representative HRIR through panning, multiplying it by a gain, and adding it together (synthesizing) is called the "synthetic HRIR." The degree to which this synthetic HRIR approximates the true HRIR for that direction is expressed by the degree to which the target signal synthesized by panning approximates the original sound source signal S convolved with the HRIR for that direction.
[0029] Representative direction information 210 is information on representative HRIRs calculated by learning by learning unit 30. This representative direction information 210 may be set, for example, according to two or three representative HRIRs of the celestial sphere that have been learned. That is, when panning in two directions (left and right) is performed and the number of representative HRIRs is two, two representative HRIRs are set as representative direction information 210. Alternatively, when panning in three directions including the elevation angle direction is performed, three representative HRIRs are set as representative direction information 210.
[0030] [Hardware configuration of the acoustic signal reproduction device 1] The audio signal reproducing device 1 includes, for example, various circuits such as control means (control units) such as an ASIC (Application Specific Processor), a DSP (Digital Signal Processor), a CPU (Central Processing Unit), an MPU (Micro Processing Unit), and a GPU (Graphics Processing Unit).
[0031] Furthermore, the acoustic signal reproducing device 1 may include, as a storage means (storage unit), a storage unit such as a semiconductor memory such as a ROM (Read Only Memory) or a RAM (Random Access Memory), a magnetic recording medium such as a HDD (Hard Disk Drive), or an optical recording medium. The ROM may include a flash memory or other writable or recordable recording medium. Furthermore, an SSD (Solid State Drive) may be provided instead of an HDD. This storage unit may store the control program and various contents according to this embodiment. Of these, the control program is a program for realizing each functional configuration and each method, including the acoustic signal processing program of this embodiment. This control program includes an embedded program such as firmware, an OS (Operating System), and an application.
[0032] The various types of content may be, for example, movie or music data, games, audiobooks, e-book data capable of voice synthesis, television or radio broadcast data, various audio data related to operating instructions for car navigation systems or various home appliances, entertainment content including VR, AR, MR, etc., or other data capable of audio output. Alternatively, the content may include background music or sound effects from games, MIDI files, voice call data for mobile phones or walkie-talkies, or synthesized voice data for text in messengers. These types of content may be acquired by downloading files or data chunks transmitted via wire or wirelessly, or may be acquired in stages via streaming or the like. Furthermore, the application according to this embodiment may be an application such as a media player that plays back content, or an application for a messenger or video conference.
[0033] The acoustic signal reproducing device 1 may also be equipped with a direction calculation means including a GNSS (Global Navigation Satellite System) receiver that calculates the direction in which the listener is facing, an in-room position and direction detector, an acceleration sensor, a gyro sensor, a geomagnetic sensor, etc. that are capable of head tracking, and a circuit that converts the outputs of these sensors into direction information.
[0034] Furthermore, the acoustic signal reproducing device 1 may include a display unit such as a liquid crystal display or an organic EL display, an input unit such as buttons, a keyboard, a pointing device such as a mouse or a touch panel, and an interface unit for connecting to various devices wirelessly or via a wired connection. Of these, the interface unit may include an interface for a flash memory medium such as a microSD (registered trademark) card or a USB (Universal Serial Bus) memory, a LAN board, a wireless LAN board, a serial interface, a parallel interface, etc.
[0035] Furthermore, the sound signal reproducing device 1 can realize each method according to the present embodiment using hardware resources by the control means executing various programs mainly stored in the storage means. Note that a part or any combination of the above configurations may be configured in terms of hardware or circuits using ICs, programmable logic, FPGAs (Field-Programmable Gate Arrays), or the like.
[0036] [Learning process by the acoustic signal reproducing device 1] Next, the learning process performed by the sound signal reproducing device 1 according to the first embodiment of the present invention will be described with reference to FIGS.
[0037] First, an overview of panning and learning according to this embodiment will be explained with reference to FIG. In the audio signal processing according to this embodiment, the HRIR from each sound source signal S to the ear is not directly convolved with each sound source signal S, but rather each sound source signal S is synthesized and expressed by panning from representative points r1 to rn (hereinafter, when referring to one of these representative points, it will simply be referred to as "representative point r" or "r"), and the HRIR from representative point r to the ear (synthesized HRIR) is convolved to perform binaural rendering. In this case, in the learning process according to this embodiment, a specific representative HRIR is generated by repeatedly learning from initial values using HRIRs of the entire celestial sphere.
[0038] The learning process of this embodiment is mainly carried out in the audio signal reproduction device 1, in which the control means cooperates with each part to control and execute the control program stored in the storage means using hardware resources, or executes it directly in each circuit. The learning process will be described in detail below for each step with reference to the flowchart in FIG.
[0039] (Step S100) First, the learning unit 30 performs an initial cross-correlation calculation process. In the learning according to this embodiment, for example, HRIRs (a set of true HRIRs, or original HRIRs) from all around the celestial sphere that can be convoluted for sound source directions (hereinafter referred to as "target directions") that differ every 2° across the 360° celestial sphere are used with their phases aligned. Specifically, the learning unit 30 calculates the cross-correlation of the HRIRs for each target direction of the true HRIRs after applying a low-pass filter (LPF) with a cutoff frequency of, for example, 3000 Hz to 6000 Hz and approximately 6 db / oct (octave) to 12 db / oct. This calculation of cross-correlation can be performed in the same way as with conventional techniques.
[0040] (Step S101) Next, the learning unit 30 performs a shift amount application process. Before learning, the learning unit 30 aligns the phases of the HRIRs in each target direction by time shifting them so that the cross-correlation between them and the HRIR in the 0° direction of the horizontal plane is maximized. For this reason, the time shift amount that maximizes the cross-correlation is applied to the HRIRs in the target directions of the entire celestial sphere. The learning unit 30 stores the HRIRs to which these time shift amounts have been applied in the HRIR table 200.
[0041] (Step S102) Next, the learning unit 30 performs a representative direction initial value setting process. The learning unit 30 stores the initial values of various data related to the representative HRIR in the representative direction information 210. According to Fig. 4, in this embodiment, a total of eight points, six directions at 60° intervals on the horizontal plane and elevation angles of ±90°, are set as initial values for specific representative HRIRs. Fig. 4(a) is a plan view, and Fig. 4(b) is a right side view. Here, in this embodiment, an example will be described in which the initial values are set as the initial values for the representative HRIRs from the true HRIRs at 30° from representative point r1 to listener L in a counterclockwise direction to 330° from representative point r6 to listener L, at an elevation angle of 90° from upper r7 to listener L, and at an elevation angle of -90° from lower r8 to listener L. That is, in this embodiment, the HRIR table 200 and the representative direction information 210 store initial values of representative HRIRs from 30° to 330° as X1 to X6, an elevation angle of 90° as X7, and an elevation angle of −90° as X8.
[0042] (Step S103) Here, the learning unit 30 performs a representative direction candidate selection process. The learning unit 30 references the representative direction information 210 and selects three HRIRs to be used as representative HRIRs for each target direction HRIR. Of these three, two representative HRIRs are selected that have a small cosine distance to the HRIR for each target direction, and one representative HRIR where an HRIR for that target direction exists. For the representative HRIR for the elevation angle direction, the learning unit 30 selects X7, which is the zenith if the elevation angle is 0° or greater. On the other hand, the learning unit 30 selects X8, which is the nadir if the elevation angle is less than 0°. After making this selection for the HRIR of each target direction, the learning unit 30 calculates the gains A, B, and C using these.
[0043] The calculation of the gains A, B, and C will now be described. First, the vector of the HRIR time waveform for each target direction is expressed as x i Then, the three representative HRIRs are X 1(i) , X 2(i) , X 3(i) That is, X 1(i) is the HRIR in the target direction x i This is the first representative HRIR selected for X. 2(i) is x i This is the second representative HRIR selected for X. 3(i) is x i This is the third representative HRIR selected for A i , B i , C i is the weight (gain) applied to each representative HRIR.
[0044] Considering the characteristics of human hearing and the frequency distribution of sound sources, it is believed that auditory performance can be improved by prioritizing errors in the low frequency range. For this reason, in this embodiment, a convolution matrix W is created using an LPF that convolves an impulse response with a cutoff frequency of 3 kHz (6 dB / oct) to perform low frequency weighting in distortion calculation. This matrix is multiplied by the difference (error) between the HRIR of the target direction and the synthesized HRIR to obtain an error vector e i Let's say. Specifically, the error vector e i can be expressed by the following formula (1).
[0045]
number
[0046] In Equation (1), i is the index of the HRIR in the target direction, and the subscript i of the representative HRIR is the HRIR in the target direction. xiThis means that the In this embodiment, this equation (1) is used as an example of a cost function that minimizes the expected value of the error between the synthetic HRIR obtained by panning all around the celestial sphere and the true HRIR. This error vector e i From the magnitude of , solve the following equation (2).
[0047]
number
[0048] Then, the following equation (3) is obtained.
[0049]
number
[0050] Here, in formula (3), [ ] -1 represents the inverse matrix. Using this equation (3), the learning unit 30 determines the optimal gain A i , B i , C i Calculate. The learning unit 30 can store the gains A, B, and C and the time shift values calculated for sound source directions (target directions) that differ every 2° in the 360° celestial sphere in the HRIR table 200 and use them in the playback process described later.
[0051] (Step S104) Next, the learning unit 30 performs the representative direction update process. The learning unit 30 updates X1 using the gain, the HRIR of the target direction, and the representative HRIR. Specifically, ||e in the following equation (4) i || 2 Minimize the expectation of .
[0052]
number
[0053] This, X 1(i) When partially differentiated with respect to
[0054]
number
[0055] The following equation (6) is the HRIR update equation:
[0056]
number
[0057] Although the convolution matrix W is lost during the expansion process of equation (6), a gain that emphasizes low-frequency error is used. As a result, the HRIRs of the representative direction that are generated can also emphasize low-frequency error.
[0058] In this embodiment, after updating the representative direction, which is one centroid, using the LBG algorithm as described above, nearest neighbor search and encoding (clustering) using the k-means method are performed again. In this nearest neighbor search, the learning unit 30 sequentially selects representative directions from those closest in cosine distance to the target direction. The learning unit 30 sets this selected representative direction as the representative direction information 210.
[0059] (Step S105) Next, the learning unit 30 determines whether all representative HRIRs have been updated. If the HRIRs for all representative directions have been updated, the learning unit 30 determines Yes. In this embodiment, the learning unit 30 determines Yes when the updating of X1 to X6, excluding the zenith and nadir, is complete. If all representative HRIRs have not yet been updated, the learning unit 30 determines No.
[0060] If the answer is Yes, the learning unit 30 advances the process to step S106. If the answer is No, the learning unit 30 returns the process to step S103 and continues updating the other representative directions. After that, the same procedure is performed again to update X2, which is the HRIR of the next representative direction. The learning unit 30 continues these processes until X6 is updated. Note that the learning unit 30 does not update X7 and X8, as described above. As a result, ||e i || 2 Minimize the expected value of X1, X2, ... X N In this embodiment, N=8 representative HRIRs are generated by learning.
[0061] (Step S106) When all the representative HRIRs have been updated, the learning unit 30 performs a convergence rate calculation process. The learning unit 30 calculates the error vector e i The rate of change of the sum of these is calculated as the convergence rate.
[0062] (Step S107) Next, the learning unit 30 determines whether or not convergence has occurred. If the convergence rate is equal to or less than a specific convergence threshold, the learning unit 30 determines that convergence has occurred and judges it as Yes. Alternatively, the learning unit 30 may also judge it as Yes if the number of steps is equal to or greater than a specific value. If the convergence rate is greater than the specific threshold, the learning unit 30 judges it as No. If the answer is Yes, the learning unit 30 ends the learning process. If the answer is No, the learning unit 30 returns the process to step S103 and continues updating X1 to X6. Learning is performed by repeating these steps until convergence is reached. This makes it possible to repeatedly learn the representative HRIR from its initial value. Note that in the learning process according to this embodiment, the two elevation angles X7 and X8 are not subject to updating.
[0063] The learning unit 30 can perform these learning processes separately for the left ear signal and the right ear signal of binaural rendering, and store the results in the HRIR table 200 and the representative direction information 210. This allows a specific representative direction to converge differently for the left ear signal and the right ear signal. This completes the learning process according to this embodiment.
[0064] [Reproduction processing by the audio signal reproduction device 1] Next, with reference to FIG. 5, a description will be given of the reproduction process of an acoustic signal by the acoustic signal reproduction device 1 according to the first embodiment of the present invention. In the playback process according to this embodiment, when panning a sound source signal S (target signal), first, a time shift is applied to each of the sound source signals S to be assigned to the three learned representative HRIRs, using the time shift value calculated in the learning process described above so as to maximize the cross-correlation between the HRIR in the target direction and the representative HRIR. These signals are then multiplied by the gain calculated in the learning process described above so as to minimize the error energy between the HRIR in the target direction and the synthesized HRIR. A sum signal of the representative point signals for the number of sound source signals S to be grouped into the representative point is then calculated, and the representative HRIR at the representative point position is convolved with this sum signal to generate a signal at the ear of the listener L.
[0065] The playback process of this embodiment is mainly carried out in the audio signal playback device 1, with the control means working in cooperation with each part to control and execute the control program stored in the storage means using hardware resources, or the control means directly executing the control program in each circuit. The playback process will be described in detail below for each step with reference to the flowchart in FIG.
[0066] (Step S110) First, the direction acquisition unit 10 of the sound signal reproducing device 1 performs sound source and direction acquisition processing. The direction acquisition unit 10 acquires the direction of the sound source signal S as seen by the listener L. Specifically, the direction acquisition unit 10 acquires an acoustic signal (target signal) of the sound source signal S. This acoustic signal may have any sampling frequency and any number of quantization bits. Furthermore, the direction acquisition unit 10 acquires directional information of the sound source signal S that is added to the acoustic signal of the content or the acoustic signals of the participants in a remote call. Then, the direction acquisition unit 10 grasps the spatial arrangement of the sound source signal S and the listener L. As described above, this arrangement may be an arrangement within a space including a virtual space set in content or the like. Then, the direction acquisition unit 10 calculates the direction of the sound source signal S as seen by the listener L, i.e., the sound source direction, according to the grasped spatial arrangement. Similarly, for the sound signal of the content, the direction acquisition unit 10 can also calculate the sound source direction based on the arrangement of the listener L by referring to the direction information of the sound source signal S. The direction acquisition unit 10 may also calculate the direction of the listener L as seen from the sound source signal S.
[0067] (Step S111) Next, the panning unit 20 performs panning processing. Here, the panning unit 20 performs panning of the sound source signal S by referring to the direction information. Specifically, the panning unit 20 obtains a time shift value corresponding to the (target direction) from the HRIR table 200. This time shift value is calculated in the learning process described above so as to maximize the cross-correlation between the HRIR for the target direction multiplied by the LPF and the HRIR for the representative direction multiplied by the LPF. The panning unit 20 performs a time shift on the sound source signal S corresponding to the representative HRIR in order to generate a signal to be assigned to the representative HRIR. The panning unit 20 may perform this time shift, for example, by adding "0" to the HRIR, which is the number of samples for the time. The panning unit 20 also calculates the gain A i , B i , C i is obtained and applied to each of the time-shifted signals to be assigned to each representative HRIR.
[0068] Then, the panning unit 20 takes the sum of the representative point signals of the sound source signals S that are grouped together at the representative point r, and generates a sum signal. Then, the panning unit 20 performs binaural rendering by convolving the representative HRIR with this sum signal, thereby generating a signal at the listener L's ear.
[0069] (Step S112) Next, the panning unit 20 and the output unit 40 perform audio output processing. The output unit 40 reproduces the signals at the ears generated by the panning unit 20 by outputting them to the reproduction unit 50. This output may be, for example, a two-channel analog acoustic signal corresponding to the left and right ears of the listener. This allows the reproduction unit 50 to reproduce a stereophonic signal corresponding to a virtual sound field as a two-channel acoustic signal through headphones. This completes the audio reproduction process according to the first embodiment of the present invention.
[0070] The above configuration can provide the following effects. Conventionally, in headphone playback of 3D sound for AR or VR using conventional technology such as that described in Patent Document 1, HRIRs for the corresponding sound source directions are individually convolved with multiple sound source signals to generate binaural signals. In this case, representative points for grouping the HRIRs are artificially selected at appropriate intervals from a set of true HRIRs of the entire celestial sphere, such as six directions at 60° intervals on the horizontal plane, the zenith, and the nadir. In this way, in conventional technology where an appropriate HRIR is artificially selected from a set of HRIRs covering the entire celestial sphere and used as the representative direction, a bias occurs in the distortion of the reproduced signal depending on the selection, making it difficult to make an "optimal" selection.
[0071] In contrast, (A) an acoustic signal processing method according to a first embodiment of the present invention is an acoustic signal processing method executed by an acoustic signal processing device 2, which acquires the sound source direction (target direction) of a sound source signal S, time-shifts and gain-adjusts the sound source signal S based on the acquired sound source direction (target direction), and pans it by allocating it to a representative HRIR (each of a plurality of representative HRIRs), and performs binaural rendering by convolving the representative HRIR (a plurality of representative HRIRs) with the allocated signal, and is characterized in that the specific representative HRIR is calculated by learning and is stored in the HRIR table 200 and the representative direction information 210.
[0072] This configuration minimizes the expected value of the degradation caused by panning-based rendering. In other words, it minimizes the expected value of the difference (distortion) between the synthesized HRIR and the true HRIR in the celestial sphere. In this way, by setting a distortion evaluation criterion and using "learning" to arrive at an "optimal" solution, it is possible to minimize the bias caused by distortion in the synthesized HRIR. This makes it possible to reproduce 3D sound of higher quality than conventional technology.
[0073] Therefore, the acoustic signal processing method according to this embodiment can be applied to VR / AR applications such as games and movies as a 3D sound field playback system. Furthermore, by applying it to smartphones and home appliances, the amount of calculation required to generate 3D sound can be reduced, leading to cost savings. Furthermore, as a method with even lower calculation amounts, it can be applied to international standardization, etc. In addition, when playing content, high-quality audio can be generated while reducing the load for multiple audio sources such as one-to-one connection, one-to-multipoint connection, and multipoint-to-multipoint connection messengers, remote conferences, etc.
[0074] Furthermore, (B) is an acoustic signal processing method according to a first embodiment of the present invention, characterized in that the representative HRIR (each of a plurality of representative HRIRs) is learned based on a cost function that minimizes the expected value of the error between the true HRIR and a synthesized HRIR obtained by panning all around the celestial sphere or the horizontal plane.
[0075] This configuration allows us to optimize the representative HRIR based on an appropriate cost function, which makes it possible to make the sound synthesized at the ear by panning closer to the sound at the ear generated by convolving multiple sound sources with the original HRIR than with conventional techniques.
[0076] Furthermore, (C) the acoustic signal processing method according to the first embodiment of the present invention is characterized in that the acoustic signal processing method is the acoustic signal processing method according to (B) or (C), in which the expected value of the error is the average value of the error.
[0077] This configuration allows the signal-to-noise ratio (hereinafter referred to as "SNR") to be increased through learning. In other words, the SNR can be improved, enabling the reproduction of higher quality 3D sound. This saves computing resources while enabling audio reproduction suitable for listening via headphones, such as AR / VR.
[0078] Furthermore, (D) the acoustic signal processing method according to the first embodiment of the present invention is characterized in that the representative HRIR (each of the multiple representative HRIRs) is different for the left ear signal and the right ear signal of the binaural rendering, being an acoustic signal processing method described in any one of (A) to (C).
[0079] This configuration makes it possible to generate optimal representative HRIRs for the right and left ears. Furthermore, if the HRIRs for the right and left ears are measured using a dummy head as the true HRIRs, optimization can be performed taking into account errors in the right and left ears at the time of measurement.
[0080] (E) An acoustic signal processing device 2 according to a first embodiment of the present invention includes a direction acquisition unit 10 that acquires the sound source direction of a sound source signal S, and a panning unit 20 that performs panning by time-shifting and gain-adjusting the sound source signal based on the sound source direction acquired by the direction acquisition unit and allocating it to a plurality of representative head impulse responses (HRIRs), and performs binaural rendering by convolving the representative HRIRs (a plurality of representative HRIRs) with the allocated signals, each of which is calculated by learning.
[0081] (F) The acoustic signal processing device 2 according to the first embodiment of the present invention is characterized by including a learning unit 30 that iteratively learns, from initial values, representative HRIRs (each of multiple representative HRIRs) for binaural rendering by panning the sound source signal by time shifting and gain adjusting it and allocating it to representative HRIRs (multiple representative head HRIRs), and convolving the representative HRIRs (multiple representative HRIRs) with the allocated signals.
[0082] By configuring in this way, it is possible to minimize bias in the distortion of the synthesized HRIR, and to reproduce stereophonic sound of higher quality than with conventional techniques.
[0083] (G) The acoustic signal reproduction device 1 according to the first embodiment of the present invention is characterized by comprising an acoustic signal processing device 2 that executes the acoustic signal processing method described in (A) to (F) above, and an audio output unit 40 that outputs the acoustic signal generated by the acoustic signal processing device 2. With this configuration, the generated audio can be output through headphones, an HMD, or the like, allowing you to experience realistic audio.
[0084] Note that in the above-described embodiment, an example has been described in which the representative HRIR (each of multiple representative HRIRs) is learned based on a cost function that minimizes the expected value of the error between the true HRIR and a synthesized HRIR obtained by panning all around the celestial sphere or the horizontal plane. In contrast, the representative HRIR (each of the multiple representative HRIRs) may be learned based on a cost function that includes a term that minimizes the expected value of the error between the true HRIR and a synthetic HRIR obtained by panning all around the celestial sphere or the horizontal plane, and a term that minimizes the expected value of the change in error between adjacent HRIRs in the azimuth and / or elevation directions.
[0085] In this case, the cost function can be expressed as follows:
[0086]
number
[0087] Here, w is a weighting matrix, and for example, a convolution matrix that convolves an impulse response of an LPF of about 3 kHz is used.
[0088] On this basis, α|e i | 2 + β|e i+1 -e i | 2 minimize the expected value of X1~X N The representative HRIR may be generated by learning. Specifically, as in the above-described embodiment, X1 to X N It is possible to train the system to learn representative directions (N=8).
[0089] Furthermore, α|e i | 2 + β|e i+1 -e i | 2 + γ|e i -e i-1 | 2 minimize the expected value of X1~X N The representative HRIR may be generated by learning. In this case, x i+1 , x i-1 is x i The HRIR in the target direction may be adjacent to the target direction. This adjacentness may be adjacent in the azimuth direction, adjacent in the elevation direction, or a combination of these. Also in this case, as in the above-described embodiment, X1 to X N It is possible to train the system to learn representative directions (N=8).
[0090] In the above embodiment, the expected value of the error in the cost function is the mean value of the error, which is the error vector e i An example expressed by the following has been described. However, the expected value of the error in the cost function may be the average value of the ratio of the error power to the signal power, which may be, for example, the SNR when the difference between the target direction HRIR and the synthesized HRIR according to this embodiment is taken as the log-scaled noise (error power).
[0091] Furthermore, in the above embodiment, an example has been described in which the direction from the representative point r, similar to the prior art, is used as the initial value of the representative HRIR. However, it is also possible to use more appropriate initial values in accordance with the characteristics of the HRIR of the celestial sphere and the characteristics of human hearing. By configuring in this way, it is possible to converge the representative direction more quickly, and it is expected that the SNR will be increased.
[0092] Furthermore, although the above embodiment describes convolving the HRIR into the acoustic signal of the sound source signal S, similar processing can also be performed by converting the acoustic signal of the sound source signal S into the frequency domain and applying the HRTF. In this case, it is possible to apply different HRTFs to each frequency range. Specifically, as in the above-described embodiment, by using HRTFs for low and high frequencies based on frequencies near or slightly higher than the frequency band where human hearing sensitivity is high, it becomes possible to synthesize with higher accuracy.
[0093] Second Embodiment In the first embodiment described above, a small change in angle can cause a large change in gain. For this reason, in the acoustic signal processing according to the second embodiment of the present invention, an example will be described in which a distortion measure is used to add (impose) a change in the gain multiplied by the HRIR of each representative direction as a penalty in order to reproduce moving sounds more smoothly.
[0094] (distortion measure) Specifically, the gains in three directions (A i , B i , C i) is written as the following equation (10):
[0095]
number
[0096] To this, a term (β) that takes into account the change from the adjacent gain can be added as in the following equation (11), which can be used as the distortion measure according to this embodiment:
[0097] E=α||{w(x i -(A i X 1(i) +B i X 2(i) +C i X 3(i) )}|| 2 +β{(A i -A i-1 ) 2 +(B i -B i-1 ) 2 +(C i -C i-1 ) 2} …… Formula (11) where x i where is the HRIR in the target direction, w is a weighting matrix, and α and β are constants. α is a coefficient for performing calculations equivalent to the gain equation according to the first embodiment, and β is a coefficient used when imposing a penalty on the distortion measure for a change in gain. For example, values such as α=0.5 to 1.0 and β=0 to 0.5 can be used. In Example 2 described below, values of α=0.8 and β=0.2 are used as an example, but these values are not limiting and any optimum value can be used.
[0098] To calculate the gain, the following equation (12) is solved.
[0099]
number
[0100] As a result, the following equation (13) is obtained.
[0101]
number
[0102] In equation (13), [] -1 represents the inverse matrix. From this equation (13), the learning unit 30 determines the optimal gain A i , B i , C i It is possible to calculate
[0103] Here, in the same way as in the first embodiment described above, when selecting three directions out of five directions, the three directions selected and used change for each i, so the distortion measure is expressed as the following equation (14):
[0104]
number
[0105] In this equation (14), X 1(i) The coefficient (gain) applied to A is A i , X 2(i) The coefficient applied to B i , X 3(i) The coefficient of C i , X 4(i) The coefficient of D i , X 5(i) The coefficient applied to E i As, X i Only the coefficients in the top three directions that are closest in cosine distance to can be made non-zero.
[0106] Table 5 below shows examples of coefficients calculated using equation (14):
[0107] [Table 5]
[0108] In the example of Table 5, the change indicated by the arrow can be added to the distortion scale as a penalty.
[0109] (5x5 gain calculation) The error vector e using the distortion scale is given in the following equation (15). i Here's an example:
[0110]
number
[0111] This is partially differentiated using the following equation (16) to obtain the gain A i , B i , C i , D i , E i Solve:
[0112]
number
[0113] By summarizing this equation, the following equation (17) is obtained.
[0114]
number
[0115] [] -1 represents the inverse matrix. This is an example of a 5x5 gain formula using distortion measures.
[0116] An example of the 5x5 gain transition calculated by the above equation (17) is shown in Table 6 below:
[0117] [Table 6]
[0118] As shown in the bracketed areas in Table 6, there may be cases where the previous gain is not available, and conversely, there may be cases where the previous gain is available.
[0119] (Update formula for the representative direction) Therefore, the HRIR update equation for the representative direction update process during the learning process can be set to the following equation (18):
[0120]
number
[0121] (gain difference term) In the above update formula, there may be no adjacent gains in the first target direction in the calculation. For example, when calculating 0°, there is no gain at 358°. Therefore, it is preferable to repeatedly perform gain calculations in the HRIR representative direction update process until convergence is achieved. Specifically, gain calculations can be performed sequentially for each elevation angle, and the calculations can be repeated until the rate of change becomes a specific value, for example, 0.001 or less.
[0122] Additionally, in this embodiment, the distortion measure may be used only during playback, or may be used during playback and learning. When the distortion measure is used during learning, the representative direction update process of the learning process of the first embodiment described above can be executed using the above-described formula (18). Also, Gain A i , B i , C i , D i , E i During actual rendering (reproduction), it can be used as a gain to be multiplied by each signal distributed to the HRIR of each representative direction.
[0123] The above configuration can provide the following effects. (H) An acoustic signal processing method according to a second embodiment of the present invention is characterized in that it is an acoustic signal processing method according to any one of (A) to (G), in which the cost function includes a distortion measure that adds (imposes as a penalty) a change in gain multiplied by the HRIR of each representative direction during learning.
[0124] By configuring it this way and using a distortion measure that imposes penalties based on gain changes, high-quality stereophonic sound can be reproduced. As will be shown in Example 2 below, sudden changes in gain due to the direction of the sound source signal are reduced, making it possible to reduce distortion, especially of moving sounds.
[0125] (I) The acoustic signal processing method according to the second embodiment of the present invention is characterized in that the distortion measure is used during learning and / or panning in the acoustic signal processing method described in (H).
[0126] This configuration allows the system to use the distortion measure during training to select the optimal representative direction of the HRIR, taking into account the connection between the gain of the synthesized HRIR when the direction of the sound source changes, and to calculate an appropriate gain. Furthermore, during playback, moving sounds can be reproduced smoothly.
[0127] (J) The acoustic signal processing method according to the second embodiment of the present invention is characterized in that the gain is calculated by equation (13) in the acoustic signal processing method described in (H) or (I).
[0128] This configuration allows high-quality stereophonic sound to be reproduced.
[0129] (K) In the acoustic signal processing method according to the second embodiment of the present invention, i , B i , C i is an acoustic signal processing method described in any one of (H) to (J), in which, during playback, the gain is used as a gain to be multiplied by each signal distributed to each of the plurality of representative HRIRs.
[0130] By configuring it this way, during playback, A i , B i , C i By using this as the gain to be multiplied by each signal distributed to the HRIR of each representative direction, high-quality stereophonic sound can be reproduced.
[0131] (L) The acoustic signal processing method according to the second embodiment of the present invention is characterized in that the distortion measure is an acoustic signal processing method according to any one of (H) to (K) that is used during learning and / or panning.
[0132] This configuration allows us to use the distortion measure during training to select the optimal representative direction of the HRIR, taking into account the relationship between the gain of the synthesized HRIR and the angle, and calculate the appropriate gain. Furthermore, during playback, moving sounds can be reproduced smoothly.
[0133] (M) The acoustic signal processing method according to the second embodiment of the present invention is characterized in that the gain is calculated by equation (17) in the acoustic signal processing method described in (L).
[0134] This configuration allows high-quality stereophonic sound to be reproduced.
[0135] (N) In the acoustic signal processing method according to the second embodiment of the present invention, i , B i , C i , D i , E i is an acoustic signal processing method according to (L) or (M), wherein, during playback, the signal is used as a gain to be multiplied by each signal distributed to each of the plurality of representative HRIRs.
[0136] By configuring it this way, during playback, A i , B i , C i , D i , E i By using this as the gain to be multiplied by each signal distributed to the HRIR of each representative direction, high-quality stereophonic sound can be reproduced.
[0137] (O) The acoustic signal processing method according to the second embodiment of the present invention is characterized in that the acoustic signal processing method is any one of (L) to (N), in which the plurality of representative HRIRs are updated by equation (18) during learning.
[0138] This configuration allows for updating and learning to minimize changes in the gain multiplied by the HRIR for each representative direction, resulting in higher quality stereophonic sound reproduction.
[0139] In addition to the above example, it is also possible to use a distortion measure that adds a change in the synthetic HRIR to the penalty. As such a distortion measure, for example, the error vector e i can be used:
[0140]
number
[0141] When this is partially differentiated with respect to each gain, the following equation (20) is obtained:
[0142]
number
[0143] The following equation (21) obtained by solving this may be used as the gain equation:
[0144]
number
[0145] When α=1 and β=0, this gain equation (21) becomes the same as equation (3) that does not take into account gain changes but is extended to five directions.
[0146] Furthermore, as an update formula for the HRIR in the representative direction, it is also possible to perform partial differentiation of the above formula (19) with respect to the HRIR in each representative direction, as in the following formula (22):
[0147]
number
[0148] Then, the following equation (23) is obtained:
[0149]
number
[0150] This can also be used as the update formula for the HRIR in the representative direction.
[0151] Thus, (P) in the acoustic signal processing method according to the second embodiment of the present invention, the cost function is characterized in that, in learning, in order to smoothly reproduce moving sounds, the cost function is an acoustic signal processing method described in any one of (A) to (O), which includes a distortion measure that adds (imposes) the magnitude of the energy as a penalty by adding frequency weighting to the change in adjacent synthetic HRIRs to calculate energy.
[0152] By configuring it in this way, frequency weighting is applied to the changes (differences) in adjacent synthesized HRIRs, their energy is calculated, and the calculated magnitude is used as a distortion measure to impose a penalty, making it possible to reproduce moving sounds smoothly.
[0153] In addition, in the above-described first embodiment, in the representative direction candidate selection process in step S103, an example is described in which the learning unit 30 refers to the representative direction information 210 and selects three HRIRs to be used as representative HRIRs for the HRIRs of each target direction. In the example of this first embodiment, two representative HRIRs with small cosine distances to the HRIR of each target direction, one representative HRIR in which the HRIR of the target direction exists, and either the zenith or the nadir are selected.
[0154] However, in the representative direction candidate selection process according to the second embodiment, it is possible to select three representative HRIRs that have small cosine distances to the HRIRs of each target direction, regardless of whether they are the zenith or nadir. When selecting these three HRIRs, the cosine distance may be weighted so that the representative HRIR selected one frame earlier, i.e., during the panning process of the adjacent target direction signal, is more likely to be selected continuously, before the comparison is performed.
[0155] Specifically, (Q) the acoustic signal processing method according to the second embodiment of the present invention is characterized in that, when selecting a representative direction, weighting is performed so as to make it easier to continue selecting a representative direction that has been selected in the past.
[0156] By configuring in this way, sudden changes in the representative direction are suppressed, and the moving sound can be changed smoothly.
[0157] Other Embodiments In the above embodiment, an example in which the playback unit 50 plays back on two channels, left and right, has been described. Regarding this, it is also possible to play it using headphones or the like that are capable of playing back multiple channels.
[0158] Furthermore, in the above embodiment, the sound signal reproducing device 1 has been described as being integrally configured. However, the acoustic signal reproducing device 1 may be configured as a reproduction system in which an information processing device such as a smartphone, PC, or home appliance is connected to a terminal such as a headset, headphones, or separated earphones. In such a configuration, the direction acquisition unit 10 and the reproduction unit 50 may be provided in the terminal, and the functions of the direction acquisition unit 10, panning unit 20, and learning unit 30 may be executed by either the information processing device or the terminal. In addition, data may be transmitted between the information processing device and the terminal via, for example, Bluetooth (registered trademark), HDMI (registered trademark), Wi-Fi (registered trademark), USB (Universal Serial Bus), or other wired or wireless information transmission means. In this case, the functions of the information processing device may also be executed by a server on an intranet or the Internet.
[0159] In the above-described first to third embodiments, the acoustic signal reproducing device 1 includes the learning unit 30, the output unit 40, and the reproducing unit 50. However, in the above-described first to third embodiments, the acoustic signal reproducing device 1 includes the learning unit 30, the output unit 40, and the reproducing unit 50. However, a configuration that does not include the learning unit 30, the output unit 40, and the reproduction unit 50 is also possible. An example of the configuration of an audio signal processing device 2b that only generates such audio signals is shown in Figure 6. In this audio signal processing device 2b, for example, binaural rendering is performed using an HRIR table 200 that has been learned in advance by a learning unit 30 and learned representative direction information 210, and the generated audio signal data can be stored on a recording medium M.
[0160] Furthermore, the acoustic signal processing device 2b according to such another embodiment can be incorporated into various devices such as content playback devices such as PCs, smartphones, game devices, media players, VR, AR, MR, videophones, video conference systems, remote conference systems, game devices, and other home appliances. In other words, the acoustic signal processing device 2b can be applied to all devices that can acquire the direction of a sound source signal S in a virtual space, such as devices equipped with a television or display, videophones through a display, videoconferencing, telepresence, and the like.
[0161] The acoustic signal processing program according to this embodiment can also be executed by these devices. Furthermore, when creating or distributing content, these acoustic signal processing programs can also be executed by a PC or server of a production company or distributor. Furthermore, this acoustic signal processing program can also be executed by the acoustic signal reproduction device 1 according to the above-described embodiment.
[0162] In this way, processing by the above-described audio signal processing devices 2, 2b and / or audio signal processing programs enables playback of movies, games, VR, AR, MR, and the like with headphones and / or HMDs that provide a more immersive and realistic experience. Realism can also be enhanced in remote conferences, etc. Furthermore, application to movie theaters, field games, 3D sound field capture, transmission, and playback systems, and application to AR and VR applications, etc., is also possible.
[0163] Contrary to the above-described acoustic signal processing device 2b, it is also possible to use an acoustic signal processing device that mainly includes only the learning unit 30 as its functional configuration. In this case, an acoustic signal processing program for performing the above-described learning process may be executed to function as the acoustic signal processing device. This acoustic signal processing device may include a GPU or the like for accelerating vector calculations. In this case, the acoustic signal processing program according to this embodiment may be executed by a separate server or the like that has a built-in GPU, with respect to the learning process of the learning unit 30.
[0164] In the above embodiment, an example in which directional information is added to the acoustic signal of the sound source signal S has been described. In this regard, in situations where a conversation is taking place in which the speaker and listener change roles at any time, such as the above-mentioned remote conference, direction information does not need to be added to the acoustic signal of the sound source signal S. In other words, when the current listener is the speaker, the direction of the speaker (current listener) can be estimated using the acoustic signal of the speaker, and this can be used as the direction of the listener as seen from the current speaker.
[0165] In this case, the direction acquisition unit 10 calculates the direction of arrival of, for example, the L (left) channel signal (hereinafter referred to as the "L signal") and the R (right) channel signal (hereinafter referred to as the "R signal") of the acoustic signal as seen from the listener. In this case, the direction acquisition unit 10 may calculate the ratio of the intensities of the L channel and the R channel. It is also possible to estimate the direction of arrival of the signal of each frequency component from this intensity ratio.
[0166] Alternatively, the direction acquisition unit 10 may estimate the direction of arrival of the acoustic signal from the relationship between the ITD (Interaural Time Difference) of the signal at each frequency in the HRTF (Head-Related Transfer Function) and the direction of arrival. The direction acquisition unit 10 may refer to a database stored in the storage unit for the relationship between the ITD and the direction of arrival.
[0167] Alternatively, it is possible to estimate the direction of a speaker or listener by performing face recognition from facial image data of a person, such as a speaker or listener in a content or video conference. That is, it is possible to estimate the direction even in a configuration without head tracking. Similarly, it may be possible to grasp the position of a speaker or listener in a space. This configuration makes it possible to accommodate a variety of flexible configurations. In applications such as VR and Social VR, the position of the sound source is known in advance, so the direction of the sound source signal S can be obtained from the positional relationship between the sound source signal S and the listener L without estimating the sound source direction.
[0168] The present invention will now be further described by way of examples with reference to the drawings, but the following specific examples are not intended to limit the present invention. [Example]
[0169] [Objective evaluation] In this embodiment, to perform an objective evaluation, the objective evaluation was performed using three items: the SNR when the HRIR in the target direction was considered as the signal and the difference between the HRIR in the target direction and the synthesized HRIR obtained by this embodiment was considered as the noise; a mapping diagram showing which representative HRIR each true HRIR belongs to; and the waveform of a moving sound source obtained by applying the synthesized HRIR to a sine wave. The original HRIRs (a set of true HRIRs) used are from FABIAN (<URL=”https: / / depositonce.tu-berlin.de / handle / 11303 / 6153”> This FABIAN contains data at 2° intervals across the entire celestial sphere.
[0170] (SNR) First, as a comparative example, similar to the prior art, no learning was performed, and HRIRs for each direction shown in Figure 3 were artificially selected from the FABIAN dataset and used as representative HRIRs to generate synthetic HRIRs. In this comparative example, the HRIR for the target direction was generated using nearby representative HRIRs. For example, HRIRs within the range of 30° to 90° in the horizontal plane were synthesized using the HRIRs at 30° and 90° and the HRIR at the zenith. In contrast, the representative HRIR data were subjected to iterative learning using the above-mentioned learning process to provide examples.
[0171] Figure 7 shows a comparison of the SNRs of the composite HRIRs of the example and comparative example, mapped together. Figure 7(a) shows the SNR at an elevation angle of 0°, Figure 7(b) shows the SNR at 46°, and Figure 7(c) shows the SNR at -46°. In each case, the solid line represents the SNR of the example, and the dotted line represents the SNR of the comparative example. The horizontal axis of the graph represents the azimuth, and the vertical axis represents the SNR value.
[0172] The average SNR values calculated over the entire circumference for this example and comparative example are shown in Table 1 below. In Table 1, the SNR is infinite for the direction where the representative HRIR for the comparative example and the HRIR for the target direction match. For this reason, these directions were excluded when calculating the average SNR values and outputting the graph.
[0173] [Table 1]
[0174] As a result, the learning results in an HRIR that is different from the initial representative HRIR in the Example. Also, although the value is lower than that of the conventional technology at an elevation angle of 0°, the SNR of the Example is improved over the Comparative Example at both elevation angles of ±46°.
[0175] (Mapping diagram) Next, Figures 8, 9, and 10 show mapping diagrams indicating which representative HRIR is used for the HRIR of each target direction. Figure 8 shows an elevation angle of 0°, Figure 9 shows 46°, and Figure 10 shows -46°. In each mapping diagram, the first vertical axis is the index of the representative HRIR, the second axis is the magnitude of the error vector, and the horizontal axis is the azimuth of the target direction HRIR. Here, the representative HRIR indices are numbered 1 to 6 counterclockwise from the front of the horizontal plane, using the angles shown in Figure 3 above, with an elevation angle of 90° being number 7 and an elevation angle of -90° being number 8.
[0176] In each mapping diagram, graph (a) shows the results at the initial value, and graph (b) shows the results at the time of convergence through learning. The points on the graphs indicate which representative HRIR is used for each target direction HRIR. The circle (○), cross (×), and triangle (△) at each point indicate the order (1st, 2nd, 3rd) in which the representative HRIR for that point is used.
[0177] To give a specific example, when looking at the graph in Figure 8(b), the HRIR for a target direction of 90° is synthesized using representative HRIRs No. 2, No. 4, and No. 7. Also, the HRIR for a target direction of 120° is synthesized using representative HRIRs No. 4, No. 2, and No. 7.
[0178] Table 2 below shows the sum of all error vectors in the entire celestial sphere and the sum of error vectors for each elevation angle.
[0179] [Table 2]
[0180] These results show that learning reduced errors across the entire celestial sphere and at elevation angles of ±46°, demonstrating the effectiveness of learning.
[0181] (moving sound source) It is possible to generate a moving sound source that circles around the head by sequentially convolving 360° HRIRs with a sound signal of a predetermined sufficient length.To examine the effectiveness of this embodiment, a moving sound source was created by convolving a synthetic HRIR using representative HRIRs generated by learning with a sound signal, and the waveform was confirmed.
[0182] The generation of this moving sound source will be explained with reference to FIG. The sound signal used was a 12-second sine wave with a sampling frequency of 48 kHz and an amplitude of 1 kHz. To create the moving sound source, the 12-second signal was extracted in 256-sample increments, as shown in Figure 11, and a Hanning window was applied to the extracted signals. To account for signal overlap, the signals were extracted with a 128-sample offset, and each extracted signal was convolved with an HRIR and added together. This was repeated for the entire circumference, generating a signal that circles the head in 12 seconds.
[0183] Fig. 12 shows the results of generating a moving sound source waveform with an elevation angle of 0°, Fig. 13 shows it with an elevation angle of 46°, and Fig. 14 shows it with an elevation angle of -46°. (a) shows the true HRIR, (b) shows the synthesized HRIR from the comparative example, and (c) shows the results of convolving the synthesized HRIR from this example. In each graph, the vertical axis represents the amplitude of the waveform, and the horizontal axis represents the azimuth (degrees). In all cases, the HRIR from the left ear is used.
[0184] At an elevation angle of 0°, the comparative example and the example of the prior art show waveforms similar to those of a true HRIR, and a smooth transition can be seen. Furthermore, at elevation angles of ±46°, the example exhibits a waveform more similar to the true HRIR and transitions more smoothly than the comparative example. This is thought to be because each representative HRIR changes from its initial value toward elevation angles of ±46°, resulting in a waveform closer to the true HRIR for the example at elevation angles of ±46°.
[0185] (Position of representative HRIR after learning) In this embodiment, the representative HRIR is updated using equation (6). As a result, by repeating learning, the representative HRIR changes to an HRIR with a different direction from the initial value. As a result, as described above, the HRIR changes to a more similar HRIR with an elevation angle of ±46° than the initial value, which is thought to have an impact on improving the SNR and smoothing the waveform of a moving sound source compared to conventional technology. Therefore, to confirm the change in the representative HRIR, the cosine similarity between each representative HRIR and each target direction HRIR was calculated, and the direction for which this value was maximum was plotted.
[0186] 15 and 16 are plots of representative HRIRs. FIG. 15 is a rotated plan view of the head of listener L viewed from directly above for clarity. FIG. 16 is a rear view of listener L's head viewed from behind. In both figures, (a) shows the initial orientation, and (b) shows the orientation after convergence. r1 to r8 correspond to the representative HRIRs X1 to X8 described above, respectively. Both figures show the HRIR results for the left ear. In this example, the representative HRIRs (r7, r8) at elevation angles of ±90° are not subject to updating, and therefore do not change from their initial values.
[0187] Tables 3 and 4 below show the changes in azimuth and elevation angle of representative HRIRs.
[0188] [Table 3]
[0189] [Table 4]
[0190] These results show that many representative HRIRs are distributed across the celestial sphere, not just horizontally, but are generally biased to the left. The representative HRIRs are updated using celestial sphere HRIRs, and because these are results for the left ear, they are thought to be biased to the left, where HRIR power is greater. Here, increasing the number of forward target directions and training is expected to cause the representative HRIRs to move forward. This is thought to enable training that places more emphasis on directions that are more important to auditory perception.
[0191] (amount of calculation) The computational complexity of binaural rendering according to this embodiment is evaluated. Assuming that sound sources exist in the entire celestial sphere, the total number of sound source objects is M, the number of HRIR samples is L, and the total number of representative HRIRs is P. When the method according to this embodiment is used, signals at the ears are generated by the following procedure: a. Gain adjustment of the sound source signal b. Addition to representative points c. Convolution of representative HRIRs
[0192] The number of multiplications, sums, and sum-of-products steps required per sample is as follows: a. 3M b. 3M-P c. PL
[0193] Therefore, the amount of calculation is calculated by the following equation (8): 3M+3M-P+PL=6M-P+PL …… Formula (8)
[0194] In contrast, if the method of this embodiment is not used and the HRIR is directly convolved, the amount of calculation required will be as follows: M×L=ML …… Formula (9)
[0195] Here, as an example, we calculate the amount of calculation when the number of sound source objects is 100, the number of HRIR samples is 256, and the number of representative HRIRs is 8. When direct convolution is performed, the amount is 25,600, whereas when panning using this method is performed, the amount is 2,640. In this way, even when a large number of sound source objects exist in the celestial sphere and binaural rendering is performed, the use of the method of this embodiment is expected to have the effect of significantly reducing the amount of calculation.
[0196] (summary) In this example, a representative HRIR was generated by learning using spherical HRIRs to minimize the expected degradation caused by panning during rendering. Objective evaluation results confirmed many advantages over conventional techniques. While the SNR was inferior to conventional techniques on the horizontal plane, it was superior at elevation angles of 46° and -46°. This demonstrates that the simulation performance of the HRIR in this example is improved. Furthermore, repeated learning reduced the error in the simulated HRIR. In a comparison using a moving sound source in which the synthetic HRIR was convolved with the sound source, it was confirmed that a shape more similar to the true HRIR could be created than with conventional technology, as shown in Figures 11 and 12.
[0197] These objective evaluation results suggest that the learning effect is evident. In addition, it was found that repeated learning biases the representative HRIR toward the more powerful HRIR. This suggests that learning can be performed with an emphasis on directions that are important to auditory perception. [Example]
[0198] As Example 2 of the present invention, a moving sound source waveform was generated using the distortion measure according to the second embodiment and equations (11) to (18) that take into account changes in gain for the moving sound source. There are five representative directions.
[0199] FIG. 17 is a graph of the gain (elevation angle 0°) of the left ear for the composite HRIR (α=1, β=0). Since α=1 and β=0, the situation is the same as in the first embodiment. Here, the gain A i , B i , C i The horizontal axis shows azimuth (deg) and the vertical axis shows gain. There are some jagged edges depending on the direction. Fig. 18 is a graph of a moving sound source waveform using this gain (elevation angle 0°). The vertical axis represents amplitude, and the horizontal axis represents azimuth (deg). Since no distortion scale is used, the results are the same as in Example 1.
[0200] Figure 19 shows the left ear gain (elevation angle 0) graph for the synthetic HRIR (α=0.8, β=0.2). As in Figure 17, the gain A i , B i , C i The horizontal axis shows the direction in degrees, and the vertical axis shows the gain. As a distortion measure, β, which takes into account the change from the adjacent gain, is used to calculate the gain, which shows that the connection between gains is smoother. Figure 20 is a graph of the moving sound source waveform (elevation angle 0°) using this gain. The vertical axis represents amplitude, and the horizontal axis represents azimuth (°). It can be seen that the change is smoother than in Figure 117.
[0201] It goes without saying that the configurations and operations of the above-described embodiments are merely examples, and can be modified as appropriate within the scope of the present invention. [Industrial Applicability]
[0202] The audio signal processing method of the present invention can provide an audio signal processing device that improves playback quality while reducing the amount of calculation required to generate stereophonic sound, and can be used industrially. [Explanation of symbols]
[0203] 1. Audio signal reproduction device 2, 2b Acoustic signal processing device 10 Direction acquisition part 20 Panning Section 30 Learning Department 40 Output section 50 Playback Department 200HRIR table 210 Representative direction information L listener S, S-1~Sn sound source M Recording medium
Claims
1. An acoustic signal processing method executed by an acoustic signal processing device, comprising: Obtaining the source direction of the sound source signal; panning the sound source signal by time-shifting and gain-adjusting the sound source signal based on the acquired sound source direction and dividing it into a plurality of representative head impulse responses (HRIRs); performing binaural rendering by convolving the plurality of representative HRIRs with the distributed signals; Each of the plurality of representative HRIRs is calculated by learning.
1. An acoustic signal processing method comprising:
2. Each of the plurality of representative HRIRs is The learning is based on a cost function that minimizes the expected value of the error between the synthetic HRIR and the true HRIR by panning all around the celestial sphere or the horizontal plane.
2. The acoustic signal processing method according to claim 1.
3. Each of the plurality of representative HRIRs is The learning is based on a cost function including a term that minimizes the expected value of the error between the synthetic HRIR and the true HRIR due to panning in the entire celestial sphere or the entire circumference in the horizontal plane, and a term that minimizes the expected value of the change in error between adjacent HRIRs in the azimuth and / or elevation angle directions.
2. The acoustic signal processing method according to claim 1.
4. The expected value of the error is the average value of the error.
4. The acoustic signal processing method according to claim 2 or 3.
5. The expected value of the error is the average value of the ratio of the error power to the signal power.
4. The acoustic signal processing method according to claim 2 or 3.
6. Each of the plurality of representative HRIRs is different for the left ear signal and the right ear signal of the binaural rendering.
2. The acoustic signal processing method according to claim 1.
7. The cost function includes a distortion measure that adds a change in the gain multiplied by the HRIR of each representative direction as a penalty in the learning.
4. The acoustic signal processing method according to claim 2 or 3.
8. The distortion measure is used during the training and / or the reproduction.
8. The acoustic signal processing method according to claim 7.
9. The gain is calculated by the following equation (13): [Equation 1] 8. The acoustic signal processing method according to claim 7.
10. The above A i , B i , C i is used as a gain by which each signal distributed to each of the plurality of representative HRIRs is multiplied during reproduction.
10. The acoustic signal processing method according to claim 9.
11. The cost function includes a distortion measure that calculates energy by adding frequency weighting to changes in adjacent synthesized HRIRs in order to reproduce moving sounds smoothly during the training, and adds the magnitude of the energy as a penalty.
4. The acoustic signal processing method according to claim 2 or 3.
12. The distortion measure is used during the training and / or the reproduction.
12. The acoustic signal processing method according to claim 11.
13. The gain is calculated by the following equation (17): [Equation 2] 8. The acoustic signal processing method according to claim 7.
14. The above A i , B i , C i , D i , E i is used as a gain by which each signal distributed to each of the plurality of representative HRIRs is multiplied during reproduction.
14. The acoustic signal processing method according to claim 13.
15. During the learning, the plurality of representative HRIRs are updated using the following equation (18): [Equation 3] 2. The acoustic signal processing method according to claim 1.
16. When selecting a representative direction, weighting is performed so that a representative direction selected in the past is likely to continue to be selected.
2. The acoustic signal processing method according to claim 1.
17. An acoustic signal processing method executed by an acoustic signal processing device, comprising: Panning is performed by time-shifting and gain-adjusting a sound source signal and allocating it to a plurality of representative head impulse responses (HRIRs), and each of the plurality of representative HRIRs for binaural rendering is iteratively learned from an initial value by convolving the plurality of representative HRIRs with the allocated signals.
1. An acoustic signal processing method comprising:
18. a direction acquisition unit for acquiring a sound source direction of a sound source signal; a panning unit that performs panning by time-shifting and gain-adjusting the sound source signal based on the sound source direction acquired by the direction acquisition unit and allocating it to a plurality of representative head impulse responses (HRIRs), and performs binaural rendering by convolving the plurality of representative HRIRs with the allotted signals; Each of the plurality of representative HRIRs is calculated by learning.
1. An acoustic signal processing device comprising:
19. An audio signal processing device that processes an audio signal, The binaural rendering system includes a learning unit that iteratively learns each of a plurality of representative head impulse responses (HRIRs) from an initial value, by panning a sound source signal by time shifting and adjusting its gain and allocating it to a plurality of representative HRIRs, and convolving the plurality of representative HRIRs with the allocated signals.
1. An acoustic signal processing device comprising:
20. An audio signal processing program executed by an audio signal processing device, the audio signal processing device Acquire the sound source direction of the sound source signal, panning the sound source signal by time-shifting and gain-adjusting the sound source signal based on the acquired sound source direction and dividing it into a plurality of representative head impulse responses (HRIRs); performing binaural rendering by convolving the plurality of representative HRIRs with the assigned signals; Each of the plurality of representative HRIRs is calculated by learning. An acoustic signal processing program comprising:
21. An audio signal processing program executed by an audio signal processing device, Panning is performed by time-shifting and gain-adjusting a sound source signal and allocating it to a plurality of representative head impulse responses (HRIRs), and each of the plurality of representative HRIRs for binaural rendering is iteratively learned from an initial value by convolving the plurality of representative HRIRs with the allocated signals. An acoustic signal processing program comprising:
Citation Information
Patent Citations
Sound generation apparatus, sound reproducing apparatus, sound generation method, and sound signal processing program
JP2023164284A