Signal generation device, vehicle, and signal generation method
The signal generation device enhances sound image localization in closed spaces by adjusting frequency characteristics and levels of output signals using HRTFs and panning processing, addressing blurring issues in DBAP processing.
Patent Information
- Application Number
- JP2021114159
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-07-09
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-07-09
AI Technical Summary
Existing DBAP processing in closed spaces can lead to blurring of sound image localization.
A signal generation device that adjusts frequency characteristics of sound signals using HRTFs and executes panning processing to individually adjust levels of output signals based on the target position of virtual sound sources, enhancing sound image localization in closed spaces.
Suppresses blurring of sound image localization in closed spaces by accurately positioning sound images using HRTFs and panning processing, improving clarity and localization accuracy.
Smart Images

Figure 0007707704000001 
Figure 0007707704000002 
Figure 0007707704000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to a signal generation device, a vehicle, and a signal generation method.
Background Art
[0002] Non-Patent Document 1 discloses DBAP (Distance Based Amplitude Panning) processing. The DBAP processing is a process of controlling sound image localization by adjusting the volume of sound emitted from a speaker according to the distance between the position of a virtual sound source and the position of the speaker.
Prior Art Documents
Non-Patent Documents
[0003]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0004] In the DBAP processing described in Non-Patent Document 1, there is a possibility that sound image localization may be blurred in a closed space.
[0005] One aspect of the present disclosure aims to provide a technology capable of suppressing blurring of sound image localization in a closed space.
Means for Solving the Problems
[0006] A signal generation device according to one aspect of the present disclosure includes a first generation unit that generates a processed signal by adjusting the frequency characteristics of a sound signal indicating the sound of a virtual sound source based on a HRTF (Head-Related Transfer Function) corresponding to the target position of the virtual sound source, and a second generation unit that generates a plurality of output signals corresponding one-to-one to a plurality of speakers based on the processed signal generated by the first generation unit and executes panning processing for individually adjusting the levels of the plurality of output signals based on the target position.
[0007] A signal generation device according to another aspect of the present disclosure includes a signal processing unit that generates a plurality of processed signals by executing panning processing for generating a plurality of signals corresponding one-to-one to a plurality of speakers based on a sound signal indicating the sound of a virtual sound source and individually adjusting the levels of the plurality of signals based on the target position of the virtual sound source, and a generation unit that generates a plurality of output signals by adjusting the frequency characteristics of the plurality of processed signals generated by the signal processing unit based on a HRTF (Head-Related Transfer Function) corresponding to the target position.
[0008] A signal generation method according to another aspect of the present disclosure is a signal generation method realized by a computer, which includes generating a processed signal by adjusting the frequency characteristics of a sound signal indicating the sound of a virtual sound source based on a HRTF (Head-Related Transfer Function) corresponding to the target position of the virtual sound source, generating a plurality of output signals corresponding one-to-one to a plurality of speakers based on the generated processed signal, and executing panning processing for individually adjusting the levels of the plurality of output signals based on the target position.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Mode for Carrying Out the Invention
[0010] A: First Embodiment A1: Signal Generation Device 1 FIG. 1 is a diagram showing an example of the signal generation device 1 according to the first embodiment. The signal generation device 1 is mounted on the vehicle 100. The vehicle 100 includes the signal generation device 1, wheels 2a to 2d, an operation unit 3, a sound source 4, a notification generation unit 4A, and speakers 51 to 54.
[0011] The signal generation device 1 generates output signals h1 to h4 that correspond one-to-one to the speakers 51 to 54. The output signal h1 is provided to the speaker 51. The output signal h2 is provided to the speaker 52. The output signal h3 is provided to the speaker 53. The output signal h4 is provided to the speaker 54. The signal generation device 1 controls the sound image localization of the sound emitted from the speakers 51 to 54 by using the output signals h1 to h4. The sound image is the perceived sound source of the sound emitted from the speakers 51 to 54. The sound image is an example of a virtual sound source. The sound image localization means the position of the sound image.
[0012] The signal generation device 1 controls sound localization with only the driver sitting in the driver's seat of the vehicle 100 as the target of sound localization by causing speakers 51 to 54 to emit sound using output signals h1 to h4. The signal generation device 1 may target passengers other than the driver in the vehicle 100 or all passengers in the vehicle 100 for sound localization.
[0013] Each of the wheels 2a and 2b is a front wheel of the vehicle 100. Each of the wheels 2c and 2d is a rear wheel of the vehicle 100. The vehicle 100 may have additional wheels in addition to the wheels 2a to 2d.
[0014] The operation unit 3 is a touch panel. The operation unit 3 may be not limited to a touch panel but an operation panel having various operation buttons. The operation unit 3 receives operations performed by the passengers of the vehicle 100. Hereinafter, the "passengers of the vehicle 100" will be referred to as "users".
[0015] The sound source 4 generates a sound signal a1. The sound signal a1 indicates sound by its waveform. The sound signal a1 indicates a piece of music. The sound signal a1 may indicate a sound different from a piece of music, for example, a natural sound such as a sound of a wave, or a virtual engine sound. The sound signal a1 is a 1ch (channel) signal.
[0016] The notification generation unit 4A generates an alarm (alert) and various information. When the notification generation unit 4A determines that an alert or information should be issued based on the information received from the devices in the vehicle 100, it instructs the sound source 4 to generate the sound signal a1 and generates target position information b1 described later. The devices in the vehicle 100 are, for example, a measurement device that measures the speed of the vehicle 100 or a detection device that detects a human body located around the vehicle 100.
[0017] FIG. 2 is a diagram showing an example of the vehicle 100. FIG. 2 shows the x-axis 10a, the y-axis 10b, and the z-axis 10c in addition to the vehicle 100. The x-axis 10a is an axis along the left-right direction of the vehicle 100. The y-axis 10b is an axis along the front-rear direction of the vehicle 100. The z-axis 10c is an axis along the up-down direction of the vehicle 100. The x-axis 10a, the y-axis 10b, and the z-axis 10c constitute a three-dimensional coordinate system 10d.
[0018] The vehicle 100 includes an FL door 61, an FR door 62, an RL door 63, an RR door 64, a front glass 71, a rear glass 72, a roof panel 73, a floor panel 74, and a passenger compartment 100a.
[0019] The FL door 61 is a front left door. The FR door 62 is a front right door. The RL door 63 is a rear left door. The RR door 64 is a rear right door.
[0020] The passenger compartment 100a is an enclosed space. The passenger compartment 100a is defined by, for example, the FL door 61, the FR door 62, the RL door 63, the RR door 64, the front glass 71, the rear glass 72, the roof panel 73, and the floor panel 74. The passenger compartment 100a has speakers 51 to 54, a dashboard 75, and seats 81 to 84.
[0021] The speakers 51 to 54 are an example of a plurality of speakers. The plurality of speakers is not limited to four speakers, and may be, for example, two, three, or five or more speakers. Each of the speakers 51 to 54 emits sound into the passenger compartment 100a. The speaker 51 is located in the left part 75a of the dashboard 75. The speaker 52 is located in the right part 75b of the dashboard 75. The speaker 53 is located in the RL door 63. The speaker 54 is located in the RR door 64. The sound emitted from each of the speakers 51 to 54 is reflected in the passenger compartment 100a. For example, the sound emitted from each of the speakers 51 and 52 is reflected by at least the front glass 71. The positions of the speakers 51 to 54 are not limited to the positions shown in FIG. 2 and can be changed as appropriate.
[0022] Seat 81 is the driver's seat. Seat 82 is the passenger seat. Seat 83 is the rear right seat. Seat 84 is the rear left seat.
[0023] In FIG. 1, the signal generation device 1 includes a storage device 11 and a processing device 12. The storage device 11 may be an external element of the signal generation device 1.
[0024] The storage device 11 is a computer-readable recording medium (for example, a non-transitory computer-readable recording medium). The storage device 11 includes a non-volatile memory and a volatile memory. The non-volatile memory is, for example, ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), and EEPROM (Electrically Erasable Programmable Read Only Memory). The volatile memory is, for example, RAM (Random Access Memory).
[0025] The storage device 11 stores HRTF (Head-Related Transfer Function) information i1, position information i2, and a program p1.
[0026] The HRTF information i1 is information indicating the HRTF. The HRTF is a transfer function representing the change in sound reaching both ears of the human body from a sound source. The HRTF reflects the change in sound caused by parts of the human body including the auricle of the human body, the head of the human body, and the shoulders of the human body.
[0027] FIG. 3 is a diagram showing an example of HRTF information i1. The HRTF information i1 indicates a set c of HRTFs for each position t of a sound source. The set c of HRTFs includes an R-HRTF61 and an L-HRTF62. The R-HRTF61 and the L-HRTF62 are HRTFs for the right ear and the left ear, respectively, according to the position t. Further, the R-HRTF61 and the L-HRTF62 are transfer functions representing changes in sound reaching the right ear and the left ear of a human body from a sound source existing at the position t, respectively. The R-HRTF61 and the L-HRTF62 are generated based on sound signals captured from respective microphones attached to the right ear and the left ear of a dummy head simulating a human body for a sound (impulse) generated from a predetermined position. That is, for any sound, by guiding the respective sounds added with the R-HRFT61 and the L-HRTF62 to the right ear and the left ear of a listener, any sound can be localized at a predetermined position.
[0028] FIG. 4 is a diagram showing an example of a set c of HRTFs (R-HRTF61 and L-HRTF62). The set c of HRTFs represents the relationship between frequency and sound pressure. The R-HRTF61 and the L-HRTF62 define filter coefficients in a FIR (Finite Impulse Response) filter, respectively. For example, the R-HRTF61 and the L-HRTF62 define coefficients (filter coefficients) in a plurality of taps included in the FIR filter, respectively. The plurality of taps is, for example, 512 taps. The plurality of taps is not limited to 512 taps and may be, for example, 1024 taps.
[0029] FIG. 5 is a diagram showing an example of the position t of a sound source. The position t of the sound source is an arbitrary position on the circumference k2 of the circle k1. The circle k1 is located in the plane m1. The plane m1 is parallel to both the x-axis 10a and the y-axis 10b and includes the point 81a included in the seat 81 (driver's seat). The point 81a is the center point of the seat 81. The point 81a is not limited to the center point of the seat 81 and may be, for example, an end point of the seat 81. The point 81a is located at the center of the circle k1. The radius of the circle k1 is 1.5 m. The radius of the circle k1 is not limited to 1.5 m and may be smaller than 1.5 m or larger than 1.5 m.
[0030] FIG. 5 shows, in addition to the position t of the sound source, a straight line n1 and a straight line n2. The straight line n1 is a straight line parallel to the y-axis 10b and passing through the point 81a. The straight line n2 is a straight line passing through the point 81a and the position t of the sound source.
[0031] The position t of the sound source is indicated by an angle q1. The angle q1 is the inclination angle of the straight line n2 with respect to the straight line n1. The angle q1 in the counterclockwise direction with respect to the straight line n1 is indicated by a positive (+) value. The angle q1 in the clockwise direction with respect to the straight line n1 is indicated by a negative (-) value.
[0032] FIG. 5 further shows a target position t1 of a virtual sound source and a straight line n3. The target position t1 is located within a region having vertices at the positions of each of the speakers 51 to 54. The target position t1 may be located on the circumference k2 or may not be located on the circumference k2. The straight line n3 is a straight line passing through the point 81a and the target position t1.
[0033] The target position t1 is indicated by an angle q2 and the distance between the target position t1 and the point 81a. The angle q2 is the inclination angle of the straight line n3 with respect to the straight line n1. The angle q2 in the counterclockwise direction with respect to the straight line n1 is indicated by a positive (+) value. The angle q2 in the clockwise direction with respect to the straight line n1 is indicated by a negative (-) value.
[0034] The HRTF information i1 shown in FIG. 3 indicates the position t (angle q1) of the sound source every 5 degrees in the range from -180 degrees to 180 degrees. The HRTF information i1 may indicate the position t (angle q1) of the sound source at angles different from 5 degrees every 5 degrees in the range from -180 degrees to 180 degrees.
[0035] In FIG. 1, the position information i2 includes speaker position information and position conversion information. The speaker position information is information indicating the position of each of speakers 51 to 54. The speaker position information indicates the position of each of speakers 51 to 54 in the coordinates of the three-dimensional coordinate system 10d. The position conversion information indicates the correspondence between the target position t1 of the virtual sound source indicated by the angle q2 and the distance (the distance between the target position t1 and point 81a) and the coordinates of the three-dimensional coordinate system 10d.
[0036] The program p1 defines the operation of the signal generation device 1. The storage device 11 may store the program p1 read from a storage device in a server (not shown). In this case, the storage device in the server is an example of a computer-readable recording medium.
[0037] The processing device 12 includes one or more CPUs (Central Processing Units). The one or more CPUs are an example of one or more processors. Each of the processing device, the processor, and the CPU is an example of a computer.
[0038] The processing device 12 reads the program p1 from the storage device 11. The processing device 12 functions as an instruction unit 13, a determination unit 14, an addition unit 15, a generation unit 16, and a panning processing unit 17 by executing the program p1.
[0039] The instruction unit 13 receives the target position information b1 from the operation unit 3 or the notification generation unit 4A. The target position information b1 is information indicating the target position t1 (angle q2 and distance) of the virtual sound source.
[0040] The indicating unit 13 determines the coordinates in the three-dimensional coordinate system 10d of the target position t1 (angle q2 and distance) of the virtual sound source by using the position conversion information included in the position information i2. The indicating unit 13 generates position-related information j1 including the target position t1 of the virtual sound source indicated by the coordinates of the three-dimensional coordinate system 10d and the speaker position information included in the position information i2.
[0041] The indicating unit 13 provides the panning processing unit 17 with the position-related information j1. Further, the indicating unit 13 provides the determination unit 14 with the target position information b1.
[0042] The determination unit 14 determines an HRTF 9a that is an HRTF corresponding to the target position t1 of the virtual sound source based on the target position information b1. For example, the determination unit 14 determines the HRTF 9a by using the target position information b1 and the HRTF information i1. An example of a method for determining the HRTF 9a will be described later. The HRTF 9a corresponding to the target position t1 mainly determines the position in the front-rear direction of the sheet 81 in the sound image localization of the sound output from the speakers 51 to 54 based on the output signals h1 to h4. The front-rear direction of the sheet 81 means the front-rear direction of the vehicle 100.
[0043] The determination unit 14 provides the generation unit 16 with the HRTF 9a. The HRTF 9a is a 2ch signal having an R-HRTF 9r and an L-HRTF 9l. The R-HRTF 9r and the L-HRTF 9l are HRTFs for the right ear and the left ear corresponding to the target position t1 of the virtual sound source.
[0044] The additional unit 15 generates a sound signal f1 by expanding the frequency band of the sound signal a1. For example, the additional unit 15 generates the sound signal f1 by applying a distortion process to the sound signal a1. The distortion process is a process of expanding the frequency band of the sound signal a1 by distorting the waveform of the sound signal a1 (performing processes such as non-linear conversion). The sound signal f1 includes, in addition to the sound signal a1, a sound signal indicating the higher harmonics of the sound indicated by the sound signal a1. The sound signal f1 is a 1ch signal. The additional unit 15 provides the sound signal f1 to the generation unit 16. The sound signal f1 is an example of a sound signal indicating the sound of a virtual sound source. The additional unit 15 is an example of a third generation unit.
[0045] The generation unit 16 generates a processed signal g1 by adjusting the frequency characteristics of the sound signal f1 based on the HRTF9a corresponding to the target position t1 of the virtual sound source. For example, the generation unit 16 generates the processed signal g1 by adjusting the frequency characteristics of the sound signal f1 with the HRTF9a corresponding to the target position t1 of the virtual sound source. The generation unit 16 may generate the processed signal g1 by adjusting the frequency characteristics of the sound signal f1 with the result obtained by multiplying the HRTF9a by a constant w. The processed signal g1 is a 1ch signal. The generation unit 16 is an example of a first generation unit. The generation unit 16 includes a synthesis unit 161 and a signal generation unit 162.
[0046] The synthesis unit 161 generates an HRTF9b by synthesizing the R-HRTF9r and L-HRTF9l included in the HRTF9a. The HRTF9b is a 1ch signal.
[0047] The signal generation unit 162 generates a processed signal g1 by adjusting the frequency characteristics of the sound signal f1 based on the HRTF9b. The signal generation unit 162 includes a FIR filter 163. The FIR filter 163 has a plurality of taps. The filter coefficients of the FIR filter 163 are determined by the HRTF9b. The filter coefficients of the FIR filter 163 may be determined by the result obtained by multiplying the HRTF9b by a constant w. The FIR filter 163 generates the processed signal g1 by performing a convolution process on the sound signal f1.
[0048] Since the R-HRTF9r and L-HRTF9l included in HRTF9a originally represent the positions of virtual sound sources in all directions surrounding the user (including not only the front-back direction but also the left-right direction), the position information in the left-right direction of the virtual sound source is lost by synthesizing the R-HRTF9r and L-HRTF9l. However, in the present invention, since the HRTF process is adopted to recover the weakness of the DBAP process described later (the front-back direction localization in a specific environment is unclear), the above disappearance in the left-right direction does not pose any problem. On the contrary, it brings the merit that the filter processing amount is reduced by half.
[0049] The panning processing unit 17 executes panning processing. In the panning processing, the panning processing unit 17 generates output signals h1 to h4 based on the processing signal g1. The output signals h1 to h4 are sound signals of the FL (front left) ch, FR (front right) ch, RL (rear left) ch, and RR (rear right) ch, respectively. Subsequently, in the panning processing, the panning processing unit 17 individually adjusts the levels of the output signals h1 to h4 based on the position-related information j1.
[0050] The panning processing determines at least the position in the left-right direction of the sheet 81 in the sound image localization of the sound output from the speakers 51 to 54 based on the output signals h1 to h4. The left-right direction of the sheet 81 means the left-right direction of the vehicle 100.
[0051] The panning processing unit 17 executes the DBAP process as the panning processing. The DBAP process is a process for controlling the sound image localization by adjusting the volume of the sound emitted from the speaker according to the distance between the position of the virtual sound source and the position of the speaker.
[0052] A2: Description of the operation FIG. 6 is a diagram showing an example of the operation of the signal generation device 1. Hereinafter, the FIR filter 163 has 512 taps. R-HRTF61 and L-HRTF62 in the HRTF information i1 respectively indicate the coefficients at 512 taps in the FIR filter 163. The adder 15 generates the sound signal f1.
[0053] When the operation unit 3 receives an instruction indicating the target position t1 of the virtual sound source from the user, it provides the target position information b1 to the instruction unit 13. Alternatively, when the notification generation unit 4A determines that an alert or information should be generated based on the information received from the device in the vehicle 100, it provides the target position information b1 corresponding to the alert or information to the instruction unit 13. The target position information b1 is information indicating the target position t1 of the virtual sound source by the angle q2 and the distance (described above).
[0054] The angle q2 satisfies the relationship of "-180 degrees ≤ q2 ≤ 180 degrees". The location of the target position t1 is specified by the angle q2 and the distance. When the instruction unit 13 receives the target position information b1, the operation shown in FIG. 6 starts.
[0055] In step S101, the instruction unit 13 determines the coordinates in the three-dimensional coordinate system 10d of the target position t1 (angle q2 and distance) of the virtual sound source indicated by the target position information b1 by using the position conversion information included in the position information i2. The position conversion information indicates the correspondence between the target position t1 (angle q2 and distance) of the virtual sound source and the coordinates of the three-dimensional coordinate system 10d.
[0056] Subsequently, in step S102, the instruction unit 13 generates position-related information j1. The position-related information j1 includes the target position t1 of the virtual sound source indicated by the coordinates of the three-dimensional coordinate system 10d and the speaker position information included in the position information i2. The speaker position information indicates the positions of each of the speakers 51 to 54 in the coordinates of the three-dimensional coordinate system 10d. Therefore, if the position-related information j1 is used, the distances between the target position t1 of the virtual sound source and each of the speakers 51 to 54 are specified. The distances between the target position t1 of the virtual sound source and each of the speakers 51 to 54 are information necessary for the DBAP process.
[0057] Subsequently, the instruction unit 13 provides the position-related information j1 to the panning processing unit 17. Subsequently, the instruction unit 13 provides the target position information b1 to the determination unit 14. The provision of the target position information b1 may be executed before the position-related information j1 is provided.
[0058] Subsequently, in step S103, the determination unit 14 determines the HRTF 9a corresponding to the target position t1 of the virtual sound source based on the target position information b1.
[0059] In step S103, the determination unit 14 reads out a set c (in 5-degree increments) of two HRTFs (before and after) sandwiching the angle based on the angle information (for example, in 1-degree increments) included in the target position information b1, and determines the HRTF 9a by performing an interpolation operation. The determination unit 14 uses a linear interpolation operation as the interpolation operation. The interpolation operation is not limited to a linear interpolation operation. For example, the interpolation operation may be a spline interpolation operation.
[0060] Subsequently, the determination unit 14 provides the HRTF 9a corresponding to the target position t1 of the virtual sound source to the synthesis unit 161.
[0061] Subsequently, in step S104, the synthesis unit 161 generates the HRTF 9b by synthesizing the R-HRTF 9r and the L-HRTF 9l included in the HRTF 9a.
[0062] In step S104, the synthesizing unit 161 generates HRTF9b by adding R-HRTF9r to L-HRTF9l. The synthesizing unit 161 may generate the arithmetic mean of R-HRTF9r and L-HRTF9l as HRTF9b. The synthesizing unit 161 may generate, as HRTF9b, the sum of the result obtained by multiplying R-HRTF9r by a first constant and the result obtained by multiplying L-HRTF9l by a second constant. The first constant may be equal to the second constant or different from the second constant.
[0063] Subsequently, in step S105, the synthesizing unit 161 determines the filter coefficients of the FIR filter 163 using HRTF9b. For example, the synthesizing unit 161 sets the coefficients indicated by HRTF9b to 512 taps in the FIR filter 163.
[0064] Subsequently, in step S106, the FIR filter 163 generates a processed signal g1 by performing a convolution process on the sound signal f1. Subsequently, the FIR filter 163 provides the processed signal g1 to the panning processing unit 17.
[0065] Subsequently, in step S107, the panning processing unit 17 performs a panning process on the processed signal g1 based on the position-related information j1.
[0066] In step S107, the panning processing unit 17 executes the DBAP process as the panning process. Hereinafter, the DBAP process will be described. First, the panning processing unit 17 identifies the distances between the target position t1 of the virtual sound source and each of the speakers 51 to 54 based on the position-related information j1. Subsequently, the panning processing unit 17 divides the processing signal g1 into output signals h1 to h4. Subsequently, the panning processing unit 17 individually adjusts the levels of the output signals h1 to h4 based on the distances between the target position t1 of the virtual sound source and each of the speakers 51 to 54. For example, the panning processing unit 17 may individually adjust the levels of the output signals h1 to h4 based on the distances in the left-right direction of the sheet 82 between the target position t1 of the virtual sound source and each of the speakers 51 to 54. Since the DBAP process is a known technique, detailed description thereof is omitted.
[0067] The panning processing unit 17 provides the speaker 51 with the output signal h1 (the sound signal of the FLch) having the adjusted level, the speaker 52 with the output signal h2 (the sound signal of the FRch), the speaker 53 with the output signal h3 (the sound signal of the RLch), and the speaker 54 with the output signal h4 (the sound signal of the RRch).
[0068] The speakers 51 to 54 emit sounds based on the output signals h1 to h4 having the adjusted levels.
[0069] The sounds emitted from the speakers 51 to 54 are affected by both the influence of the process based on the HRTF9b and the influence of the panning process. Therefore, the user sitting on the seat 81 can recognize the sounds emitted from the speakers 51 to 54 as the sounds emitted from the target position t1 of the virtual sound source. That is, for the user sitting on the seat 81, sound image localization occurs at the target position t1 of the virtual sound source.
[0070] FIG. 7 is a diagram for explaining target positions (assumed sound image localizations) d1 to d4 of virtual sound sources in a situation where only the DBAP process is executed without executing the process based on HRTF9b in the passenger compartment 100a (hereinafter referred to as "the situation of only DBAP"). FIG. 8 is a diagram for explaining actual positions (actual sound image localizations) e1 to e4 of virtual sound sources in the situation of only DBAP. In the situation of only DBAP, the DBAP process is executed on the sound signal a1 output from the sound source 4.
[0071] In the situation of only DBAP, when the position d1 is set as the target position t1 of the virtual sound source, the actual position of the virtual sound source (sound image) occurs at the position e1. When the position d2 is set as the target position t1 of the virtual sound source, the actual position of the virtual sound source (sound image) occurs at the position e2. When the position d3 is set as the target position t1 of the virtual sound source, the actual position of the virtual sound source (sound image) occurs at the position e3. When the position d4 is set as the target position t1 of the virtual sound source, the actual position of the virtual sound source (sound image) occurs at the position e4.
[0072] In the situation of only DBAP, the following problems occur. When the speaker is panned from left to right in front of the seat 81, a user sitting on the seat 81 hears a sound like being trapped due to the influence of the reflected sound in the passenger compartment 100a. For this reason, some people do not necessarily feel that the sound image is in the front. In particular, in the region to the right of the center in the left - right direction of the vehicle 100 in front of the seat 81, the sound image is localized inside the head, and it is difficult to feel that the sound image exists in the front. Also, in the region to the right of the seat 81, the speaker is too close to the user sitting on the seat 81, and the sounds of the FRch and RRch do not mix with each other, and the localization of the sound image is ambiguous.
[0073] In this embodiment (a situation where both the process based on HRTF9a and the DBAP process are executed in the passenger compartment 100a), the actual position (actual sound image localization) of the virtual sound source is almost the same as the target position (assumed sound image localization) of the virtual sound source.
[0074] This embodiment has the following advantages compared to the situation with only DBAP. In the region to the right of the center in the left - right direction of the vehicle 100, among the front of the seat 81 and in front of the seat 81, it is likely to give the feeling that the sound image exists in the front. In the region to the right of the seat 81, the sound - image localization is improved. In other directions as well, the direction from the seat 81 to the sound image becomes clear.
[0075] Also, in this embodiment, for the sound signal f1 generated by expanding the frequency band of the sound signal a1, processing based on HRTF9a is executed. Therefore, compared to the configuration in which processing based on HRTF9a is executed for the sound signal a1, the frequency band affected by HRTF9a increases. For this reason, the sound image is sharper compared to the configuration in which processing based on HRTF9a is executed for the sound signal a1.
[0076] A3: Summary of the First Embodiment The generation unit 16 generates the processing signal g1 by adjusting the frequency characteristics of the sound signal f1 based on HRTF9a corresponding to the target position t1 of the virtual sound source. The panning processing unit 17 executes panning processing. In the panning processing, output signals h1 - h4 are generated based on the processing signal g1, and the levels of the output signals h1 - h4 are adjusted based on the target position t1 of the virtual sound source.
[0077] Therefore, compared to the configuration that performs panning processing without performing adjustment based on HRTF9a, it is possible to suppress the blurring of sound - image localization in a closed space.
[0078] B: Variation The modes of variation in the first embodiment are shown below. Two or more modes arbitrarily selected from the following modes may be appropriately combined within a range where they do not conflict with each other.
[0079] B1: The First Variation Example In the first embodiment, the generation unit 16 may use the R-HRTF9r or the L-HRTF9l instead of the HRTF9b. In the first modification, the generation unit 16 includes a setting unit instead of the synthesis unit 161. The setting unit determines the filter coefficients of the FIR filter 163 using the R-HRTF9r or the L-HRTF9l. For example, the synthesis unit 161 sets the coefficients indicated by the R-HRTF9r or the L-HRTF9l to a plurality of taps in the FIR filter 163. In this case, among the R-HRTF9r and the L-HRTF9l, the HRTF used to determine the filter coefficients of the FIR filter 163 is an example of the HRTF corresponding to the target position.
[0080] According to the first modification, compared with the first embodiment in which the HRTF9b is generated by synthesizing the R-HRTF9r and the L-HRTF9l, the synthesis process can be made unnecessary.
[0081] In the first embodiment, the HRTF9b is generated by synthesizing the R-HRTF9r and the L-HRTF9l. Therefore, the relationship between the frequency and the sound pressure of the HRTF9b tends to be more complicated than that of either the R-HRTF9r or the L-HRTF9l. The more complicated the relationship between the frequency and the sound pressure in the HRTF used to determine the filter coefficients of the FIR filter 163 is, the easier the sound corresponding to the signal generated by the FIR filter 163 is to be perceived by humans and the easier it is to affect the sound image localization. For this reason, compared with the first modification, the first embodiment makes it easier to localize the sound image at the target position t1 of the virtual sound source.
[0082] B2: Second modification The frequency band of the sound that humans can perceive is limited. For example, a man in his forties tends to have difficulty hearing sounds having frequencies higher than 12 kHz. For this reason, in a situation where the highest frequency among the frequencies of the sound signal a1 is higher than a threshold value (for example, 12 kHz), even if the addition unit 15 expands the frequency band of the sound signal a1, the user may not be able to hear the sound having the expanded frequency.
[0083] Therefore, in the first embodiment and the first modification example, the addition unit 15 may expand the frequency band of the sound signal a1 only when the highest frequency among the frequencies of the sound signal a1 is lower than a threshold value (for example, 12 kHz). Note that the threshold value is not limited to 12 kHz and can be changed as appropriate.
[0084] According to the second modification example, it is possible to limit the addition unit 15 from performing an operation with low necessity (an operation that hardly affects sound image localization).
[0085] B3: Third modification example In the first embodiment and the first modification example, the addition unit 15 may be omitted. In this case, the sound signal a1 is provided to the generation unit 16 instead of the sound signal f1.
[0086] According to the third modification example, compared with the configuration having the addition unit 15, it is possible to reduce the processing load and simplify the configuration.
[0087] B4: Fourth modification example In the first embodiment and the first to third modification examples, the panning processing unit 17 may execute VBAP (Vector Based Amplitude Panning) processing as panning processing instead of DBPA processing.
[0088] According to the fourth modification example, even when VBAP processing is used as panning processing, it is possible to suppress the blurring of sound image localization in a closed space compared with a configuration that performs panning processing without performing adjustment based on HRTF9a.
[0089] B5: Fifth modification example In the first embodiment and the first to fourth modification examples, panning processing is executed after the processing based on HRTF is executed. In the first embodiment and the first to fourth modification examples, the processing based on HRTF may be executed after the panning processing is executed.
[0090] FIG. 9 is a diagram showing an example of a fifth modification. The panning processing unit 17 in the fifth modification generates processing signals g11 to g14 by performing panning processing on the sound signal f1. The processing signals g11 to g14 are an example of a plurality of processing signals. The number of the plurality of processing signals is not limited to four and may be the same as the number of the plurality of speakers 51 to 54. The panning processing in the fifth modification is, for example, DBAP processing or VBAP processing.
[0091] In the panning processing in the fifth modification, based on the sound signal f1, four signals corresponding one-to-one to the speakers 51 to 54 are generated, and the levels of the four signals are individually adjusted based on the target position t1 of the virtual sound source. The four signals are an example of a plurality of signals. The number of the plurality of signals is not limited to four and may be the same as the number of the plurality of speakers 51 to 54. The plurality of signals (four signals) are generated by dividing the sound signal f1. The processing signals g11 to g14 are four signals having levels individually adjusted based on the target position t1 of the virtual sound source.
[0092] In the fifth modification, the generation unit 16 generates output signals h1 to h4 by adjusting the frequency characteristics of the processing signals g11 to g14 based on the HRTF9a corresponding to the target position t1.
[0093] The generation unit 16 in the fifth modification includes a synthesis unit 161 and four FIR filters 163. The four FIR filters 163 correspond one-to-one to the processing signals g11 to g14. The four FIR filters 163 correspond one-to-one to the output signals h1 to h4. The synthesis unit 161 sets the filter coefficients of each of the four FIR filters 163 based on the HRTF9a. Each of the four FIR filters 163 generates a corresponding output signal by performing a convolution process on the corresponding processing signal.
[0094] According to the fifth modification, similarly to the first embodiment, it is possible to suppress blurring of sound image localization in a closed space as compared with a configuration that performs panning processing without performing adjustment based on the HRTF9a.
[0095] In the fifth modification example, after the panning process is executed, the process based on HRTF is executed. On the other hand, in the first embodiment and the first to fourth modification examples, after the process based on HRTF is executed, the panning process is executed. For this reason, the number of FIR filters 163 used in the first embodiment and the first to fourth modification examples is smaller than the number of FIR filters 163 used in the fifth modification example. Therefore, the first embodiment and the first to fourth modification examples can reduce the processing load and simplify the configuration as compared with the fifth modification example.
[0096] B6: Sixth modification example In the first embodiment and the first to fifth modification examples, the closed space is not limited to the vehicle compartment 100a, and may be, for example, an indoor space.
[0097] C: Aspects grasped from the above-described embodiments and modification examples The following aspects are grasped from at least one of the above-described embodiments and modification examples.
[0098] C1: First aspect The signal generation device according to the aspect (first aspect) of the present disclosure includes: a first generation unit that generates a processing signal by adjusting the frequency characteristics of a sound signal indicating the sound of a virtual sound source based on an HRTF (Head-Related Transfer Function) corresponding to the target position of the virtual sound source; and a second generation unit that generates a plurality of output signals corresponding one-to-one to a plurality of speakers based on the processing signal generated by the first generation unit and executes a panning process for individually adjusting the levels of the plurality of output signals based on the target position.
[0099] According to this aspect, compared with a configuration that performs panning processing without performing adjustment based on HRTF, it is possible to suppress blurring of sound image localization in a closed space. Further, in a configuration that performs adjustment based on HRTF after performing panning processing, it is necessary to perform adjustment based on HRTF for each of a plurality of signals generated by the panning processing. On the other hand, according to this aspect, it is not necessary to perform adjustment based on HRTF for each of the plurality of signals generated by the panning processing, and the processing load can be reduced.
[0100] C2: Second aspect In the example of the first aspect (second aspect), the HRTF corresponding to the target position is the R-HRTF corresponding to the target position or the L-HRTF corresponding to the target position. According to this aspect, compared with a configuration that generates HRTF by synthesizing R-HRTF and L-HRTF, the synthesis process can be made unnecessary, and the processing load can be reduced.
[0101] C3: Third aspect In the example of the first aspect (third aspect), the HRTF corresponding to the target position includes an R-HRTF that is an HRTF for the right ear corresponding to the target position and an L-HRTF that is an HRTF for the left ear corresponding to the target position, and the first generation unit includes a synthesis unit that generates HRTF by synthesizing the R-HRTF and the L-HRTF, and a signal generation unit that generates the processed signal by adjusting the frequency characteristics of the sound signal based on the HRTF generated by the synthesis unit.
[0102] The HRTF generated by the synthesis unit is more likely to include a gap that affects sound image localization than each of the R-HRTF and the L-HRTF. Therefore, according to this aspect, the accuracy of sound image localization is improved compared with a configuration that performs adjustment based on the R-HRTF or the L-HRTF. Further, the processing amount of the FIR filter is reduced by half by synthesis.
[0103] C4: Fourth aspect In the example of the first to third aspects (the fourth aspect), the HRTF corresponding to the target position determines the position in the front-rear direction of the seat in the sound image localization of the sound output from the plurality of speakers based on the plurality of output signals, and the panning process determines the position in the left-right direction of the seat in the sound image localization. According to this aspect, since the position of the sound image localization in the front-rear direction of the seat, which is difficult to determine by the panning process, is determined using the HRTF, the difference between the position of the sound image localization and the target position can be reduced compared to a configuration that uses only the panning process without using the HRTF.
[0104] C5: The fifth aspect In the example of the first to fourth aspects (the fifth aspect), the vehicle further includes a third generation unit that generates the sound signal by expanding the frequency band of the signal indicating the sound, and the first generation unit generates the processing signal by adjusting the frequency characteristics of the sound signal generated by the third generation unit based on the HRTF corresponding to the target position. According to this aspect, the frequency band of the signal affected by the HRTF increases. Therefore, the sound image localization caused by the HRTF is likely to occur.
[0105] C6: The sixth aspect The vehicle according to the aspect of the present disclosure (the sixth aspect) includes the plurality of speakers, the seat, and the signal generation device according to claim 4. According to this aspect, it is possible to suppress the blurring of the sound image localization in both.
[0106] C7: The seventh aspect A signal generation device according to an aspect (seventh aspect) of the present disclosure generates a plurality of signals corresponding one-to-one to a plurality of speakers based on a sound signal indicating the sound of a virtual sound source, and executes panning processing for individually adjusting the levels of the plurality of signals based on the target position of the virtual sound source to generate a plurality of processed signals, and a generation unit that generates a plurality of output signals by adjusting the frequency characteristics of the plurality of processed signals generated by the signal processing unit based on an HRTF (Head-Related Transfer Function) corresponding to the target position. According to this aspect, it is possible to suppress blurring of sound image localization in a closed space as compared with a configuration that performs panning processing without performing adjustment based on HRTF.
[0107] C8: Eighth aspect A signal generation method according to an aspect (eighth aspect) of the present disclosure is a signal generation method realized by a computer, which generates a processed signal by adjusting the frequency characteristics of a sound signal indicating the sound of a virtual sound source based on an HRTF (Head-Related Transfer Function) corresponding to the target position of the virtual sound source, generates a plurality of output signals corresponding one-to-one to a plurality of speakers based on the generated processed signal, and executes panning processing for individually adjusting the levels of the plurality of output signals based on the target position. According to this aspect, it is possible to suppress blurring of sound image localization in a closed space as compared with a configuration that performs panning processing without performing adjustment based on HRTF.
Description of reference numerals
[0108] 1... Signal generation device, 3... Operation unit, 4... Sound source, 11... Storage device, 12... Processing device, 13... Instruction unit, 14... Decision unit, 15... Addition unit, 16... Generation unit, 161... Synthesis unit, 162... Signal generation unit, 163... FIR filter, 17... Panning processing unit, 51 to 54... Speakers, 81 to 84 sheets, 100... Vehicle.
Claims
A third generation unit that generates a sound signal indicating the sound of a virtual sound source by expanding the frequency band of the signal indicating the sound only when the highest frequency among the frequencies of the signal indicating the sound is lower than a threshold value; A first generation unit that generates a processed signal by adjusting the frequency characteristics of the sound signal based on an HRTF (Head-Related Transfer Function) corresponding to the target position of the virtual sound source; A second generation unit that generates a plurality of output signals corresponding one-to-one to a plurality of speakers based on the processed signal generated by the first generation unit, and executes panning processing for individually adjusting the levels of the plurality of output signals based on the target position; A signal generation device including the above.
2. The HRTF corresponding to the target position is an R-HRTF corresponding to the target position or an L-HRTF corresponding to the target position. The signal generation device according to claim 1.
3. The HRTF corresponding to the target position has an R-HRTF that is an HRTF for the right ear corresponding to the target position and an L-HRTF that is an HRTF for the left ear corresponding to the target position. The first generation unit includes: A synthesis unit that generates an HRTF by synthesizing the R-HRTF and the L-HRTF; A signal generation unit that generates the processed signal by adjusting the frequency characteristics of the sound signal based on the HRTF generated by the synthesis unit. Including The signal generation device according to claim 1.
4. The HRTF corresponding to the target position determines the position in the front-rear direction of the seat in the sound image localization of the sound output from the plurality of speakers based on the plurality of output signals. The panning processing determines the position in the left-right direction of the seat in the sound image localization. The signal generation device according to any one of claims 1 to 3.
5. The threshold value is 12 kHz. The signal generation device according to any one of claims 1 to 4.
6. The plurality of speakers; The seat; The signal generation device according to claim 4; A vehicle including the above. A third generation unit that generates a sound signal indicating the sound of a virtual sound source by expanding the frequency band of the signal indicating the sound only when the highest frequency among the frequencies of the signal indicating the sound is lower than a threshold value; Based on the sound signal, a signal processing unit generates a plurality of signals corresponding one-to-one to a plurality of speakers, and generates a plurality of processed signals by executing a panning process of individually adjusting the levels of the plurality of signals based on the target position of the virtual sound source; a generating unit generates a plurality of output signals by adjusting the frequency characteristics of the plurality of processed signals generated by the signal processing unit based on an HRTF (Head-Related Transfer Function) corresponding to the target position; A signal generation device including the above.
8. A signal generation method realized by a computer, comprising: Only when the highest frequency among the frequencies of the signal indicating sound is lower than a threshold value, generating a sound signal indicating the sound of the virtual sound source by expanding the frequency band of the signal indicating the sound; generating a processed signal by adjusting the frequency characteristics of the sound signal based on an HRTF (Head-Related Transfer Function) corresponding to the target position of the virtual sound source; generating a plurality of output signals corresponding one-to-one to a plurality of speakers based on the generated processed signal, and executing a panning process of individually adjusting the levels of the plurality of output signals based on the target position; A signal generation method.
Citation Information
Patent Citations
Sound source localization method, sound source localization apparatus, and program
JP2012004816A
Acoustic reproduction device, acoustic reproduction method, and acoustic reproduction program
JP2015163909A