Signal processing device, signal processing method, signal processing program, and acoustic system
The signal processing device corrects the head transfer function based on environmental disturbances to maintain accurate sound localization in vehicles, addressing issues with head position fluctuations and noise, enhancing acoustic reproduction.
Patent Information
- Application Number
- JP2024005643
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-01-17
- Publication Date
- 2025-07-30
AI Technical Summary
Existing acoustic systems in vehicles face challenges in maintaining acoustic effects due to disturbances such as fluctuations in the user's head position and vehicle noise, which can impair the localization of sound images.
A signal processing device that utilizes a head-related transfer function (HRTF) to correct the head transfer function based on the ambient acoustic environment's disturbance state, emphasizing specific feature amounts of the HRTF to suppress the influence of disturbances and enhance acoustic reproduction.
The device effectively localizes sound images at the intended position by correcting the HRTF, providing a rich sense of presence and acoustic effect despite disturbances, ensuring accurate sound localization for each user.
Smart Images

Figure 2025111302000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a signal processing apparatus, a signal processing method, a signal processing program, and an acoustic system.
Background Art
[0002] In an acoustic system mounted on a vehicle, a speaker incorporated in a headrest of a seat is known. When a speaker is incorporated in the headrest, since the speaker is disposed behind the head of a user (listener), when an acoustic signal from a general sound source is reproduced and sound is output from such a headrest speaker, a sound image is formed on the rear side of the user. Therefore, when outputting sound from the headrest speaker, instead of reproducing the acoustic signal as it is, a technique of using a head-related transfer function (HRTF) to control so that a sound image is formed on the front side of the user may be adopted.
[0003] In Patent Document 1, the characteristics of a first band extracted from a first head-related transfer function of a user and the characteristics of a second band other than the first band extracted from a second head-related transfer function measured in a second measurement environment different from the first measurement environment in which the first head-related transfer function is measured are synthesized to generate a third head-related transfer function. Thereby, the signal processing apparatus of Patent Document 1 can realize personalization of the head-related transfer function in the entire band and perform acoustic control suitable for each individual.
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] Even if a head transfer function specialized for an individual is generated as in Patent Document 1, disturbances such as fluctuations in the user's head position and vehicle driving noise occur. When the influence of the disturbance becomes large, the acoustic effect using the head transfer function may not be sufficiently obtained.
[0006] An object of the present disclosure is to provide a technology that suppresses the influence of disturbances such as fluctuations in the user's head position and vehicle noise, appropriately obtains an acoustic effect using a head transfer function, and enables acoustic reproduction with a rich sense of presence.
Means for Solving the Problem
[0007] To solve the above problems, the signal processing device of the present disclosure is a signal processing device that performs signal processing of a sound source signal using a head transfer function of a user who listens to an acoustic reproduction sound, acquires the head transfer function of the user, acquires a disturbance state in the ambient acoustic environment of the user, and corrects the head transfer function based on the disturbance state.
Effect of the Invention
[0008] According to the present disclosure, by emphasizing the feature amount of the head transfer function according to the magnitude of the disturbance, when the disturbance is large, the feature amount of the head transfer function is greatly emphasized, and the ratio of the influence by the disturbance is suppressed to be small with respect to the feature amount of the head transfer function. For this reason, the present disclosure can provide a technology that suppresses the influence of disturbances such as fluctuations in the user's head position and noise, appropriately obtains an acoustic effect using a head transfer function, and enables acoustic reproduction with a rich sense of presence.
Brief Description of the Drawings
[0009]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
MODE FOR CARRYING OUT THE INVENTION
[0010] Hereinafter, embodiments of a signal processing apparatus, a signal processing method, a signal processing program, and an acoustic system disclosed in the present application will be described with reference to the drawings. Note that the present invention is not limited only to the embodiments shown below.
[0011] <First Embodiment> FIG. 1 is a diagram showing a schematic configuration of an acoustic system 100 mounted on a vehicle 1. As shown in FIG. 1, the acoustic system 100 includes a head unit 10, speaker units 40 (H1 to H8), and a sensor 30.
[0012] The head unit 10 uses a storage medium such as a CD, DVD, or semiconductor memory as a sound source, reads the content stored in the storage medium, reproduces an acoustic signal, and outputs sound from a plurality of speaker units 40. In the acoustic system of this embodiment, a speaker unit (hereinafter also referred to as a headrest speaker) 40 is disposed on the headrest 20H of each seat 20 in the vehicle 1, and sound is mainly output from the headrest speaker 40 to the user sitting on each seat 20. In this case, the headrest speaker 40 is disposed behind the user, but acoustic control is performed using the head related transfer function (hereinafter also referred to as HRTF) so that the sound image is localized in front of the user. The specific control method will be described later.
[0013] The sound source that supplies content to the head unit 10 is not limited to a storage medium, and may be a tuner or a network receiver. For example, the head unit 10 may receive radio or television broadcasts and generate an acoustic signal of this broadcast (content). Further, the head unit 10 may receive content from the user's smartphone, music player, or a server on the network, etc., and generate an acoustic signal based on this content. Also, when the content includes an image signal, the image signal may be reproduced together with the acoustic signal and displayed on the display devices 61, 62. The head unit 10 of this embodiment may be an audio-visual-navigation integrated electronic device (in-vehicle device) having, in addition to the audio function, a visual function such as video playback and television broadcast display, and a navigation function for setting a destination and a route via point according to the operation of the occupant and performing route guidance (navigation) to the destination.
[0014] FIG. 2 is a diagram schematically showing the arrangement of the speaker units 40 provided in the vehicle. In FIG. 2, the header side of the figure is the front of the vehicle, i.e., the direction in which the vehicle travels forward, and the footer side of the figure is the rear of the vehicle and the left side of the figure is the left side of the vehicle and the right side of the figure is the right side of the vehicle as shown.
[0015] In the example of FIG. 2, at least eight speaker units 40 (H1 to H8) are installed in the vehicle. In FIG. 2, only the headrest speakers 40 (H1 to H8) arranged on the headrests 20H (21H to 24H) of each seat 20 (21 to 24) are shown, but it may also include speaker units 40 arranged other than the headrests 20H (21H to 24H). The headrest speaker H1 is arranged on the right side of the headrest 21H in the right front seat (driver's seat) 21, and the headrest speaker H2 is arranged on the left side of the headrest 21H.
[0016] The headrest speaker H3 is arranged on the right side of the headrest 22H in the left front seat (passenger seat) 22, and the headrest speaker H4 is arranged on the left side of the headrest 22H. The headrest speaker H5 is arranged on the right side of the headrest 23H in the right rear seat 23, and the headrest speaker H6 is arranged on the left side of the headrest 23H. The headrest speaker H7 is arranged on the right side of the headrest 24H in the left rear seat 24, and the headrest speaker H8 is arranged on the left side of the headrest 24H.
[0017] The speaker unit 40 may be configured to be connected to the head unit 10 by wire, and the diaphragm is driven by an audio signal (electrical signal) supplied from the head unit 10 to physically output sound (vibration of air). Further, the speaker unit 40 may include a receiving unit, a driving unit, and a speaker. The receiving unit wirelessly receives an audio signal from the head unit 10, and the driving unit converts the audio signal into an electrical signal for driving the speaker and supplies it to the speaker, and the speaker outputs sound.
[0018] When the head unit 10 performs acoustic control such as outputting sound from the headrest speakers 40 (H1 to H8) to localize the sound image in front of the user, the sensor 30 detects the magnitude of a factor (disturbance state) that disrupts this acoustic control. Here, the disturbance state may be, for example, a variation in the user's head position. Note that the position of the user's ear is also referred to as the listening position (sound reception position), and a variation in the listening position may be detected as a variation in the head position. The variation in the head position is, for example, assumed when the user is sitting on each of the seats 21 to 24 and looking forward, and the difference between the reference position defined as the positions of both ears of the user and the actual head position of the user may be detected as the variation in the head position. The variation in the head position may be, for example, a variation in position within a horizontal plane or a variation (rotation) in the orientation of the head. Examples of the sensor (hereinafter, the head sensor) 31 that detects this variation in the head position include a camera, a ToF sensor, a three-dimensional scanner, and the like. The head sensor 31 is arranged facing the head of the user sitting on the corresponding seat in front of each of the seats 21 to 24. The head sensors 31 for the right front seat 21 and the left front seat 22 are provided, for example, on the instrument panel. The head sensors 31 for the right rear seat 23 and the left rear seat 24 are provided, for example, at the rear of the headrests 21H and 22H of the seats 21 and 22 located in front of them, respectively. Note that the head sensor 31 may also serve as a sensor of another device, such as the camera of a drive recorder.
[0019] Further, the disturbance state may be noise in the acoustic environment (ambient acoustic environment) of the listening space (inside the vehicle) where the user is present, such as the running sound or wind noise of the vehicle 1, the sound emitted from the surroundings of the vehicle 1, the sound of air blown out by the air conditioner, etc. As the sensor (hereinafter, also referred to as noise sensor) 32 for detecting noise, for example, a microphone (hereinafter, also referred to as a mic) can be mentioned. The noise sensor 32 is provided, for example, on the headrests 20H (21H to 24H) of each seat 21 to 24. Further, the sensor 30 may be a temperature sensor for detecting a disturbing temperature, a vibration sensor (acceleration sensor) for detecting vibration, a vehicle speed sensor for detecting the running speed of the vehicle 1, etc. Note that it is also possible to obtain the signal of the speed meter of the vehicle 1, the control state (operation state) signal of the air conditioner, etc. from the vehicle communication line (vehicle LAN, so-called CAN) and use them as sensors. Possible.
[0020] FIG. 3 is a functional block diagram of the head unit 10. The head unit 10 is a form of a signal processing device that performs signal processing of an acoustic signal using the head transfer function of a user who listens to the acoustic reproduction sound. As shown in FIG. 3 for this signal processing, the head unit 10 includes each processing unit such as an audio signal acquisition unit 11, an HRTF acquisition unit 12, a disturbance acquisition unit 13, an HRTF enhancement unit 14, an audio signal generation unit 15, an inverse filter unit 16, and an output control unit 17. The audio signal acquisition unit 11 reads and acquires the acoustic signal s(f) from a sound source device such as a CD, DVD, USB memory, memory card, etc. Further, the audio signal acquisition unit 11 can also acquire the acoustic signal s(f) from an external sound source device such as a content server or NAS (Network Attached Storage) via a network.
[0021] The HRTF acquisition unit 12 acquires the HRTF of the user. For example, the HRTF acquisition unit stores a plurality of patterns of HRTFs, and based on each pattern, plays a test sound source while performing control to localize the sound image to a specified position such as the front side → right side → rear side → left side of the user, and allows the user to select the pattern in which the sound image localization can be perceived most as specified, thereby acquiring the HRTF suitable for the user. Not limited to this, the HRTF acquisition unit 12 may read an image of the user's ear taken by a camera or an MRI image, and determine the HRTF from the shape of the user's ear. Further, the HRTF acquisition unit 12, with the user wearing a measurement microphone in the external auditory canal, acquires impulse sounds from a plurality of surrounding coordinates (positions of each speaker) using the microphone, measures the impulse response (such as the difference in timing and sound pressure level in the output and acquisition of the impulse sound), and may obtain the HRTF based on this measurement result (for example, by performing a Fourier transform). The head-related transfer function (HRTF) is a function having, for example, a frequency f, an azimuth angle azi, and an elevation angle ele as elements as shown in FIG. 3. Note that instead of the HRTF acquisition unit 12 calculating the HRTF, HRTFs calculated in advance for various values of the above-described various measurement parameters may be stored in a storage device, and the HRTF acquisition unit 12 may acquire the HRTF corresponding to the measurement result of the various measurement parameters from the storage device by reading it out.
[0022] The disturbance acquisition unit 13 acquires the disturbance state of the acoustic environment around the user via the sensor 30. In the present embodiment, the disturbance acquisition unit 13 acquires, for example, the magnitude (variation amount) of the variation in the head position of the user sitting on each of the seats 21 to 24 by the head sensor 31. For example, the disturbance acquisition unit 13 acquires the variation amounts for the azimuth angle Tazi, elevation angle Tele, and distance Tdis of the head. Then, the disturbance acquisition unit 13 obtains the head variation error w l,r (f) as follows. N1 l,r (f) = (Tazi × Tele × Tdis) x f w l,r (f) = a × N1 l,r (f) However, Ni l,r : External disturbance (l: left ear, r: right ear) a: Load
[0023] Further, the external disturbance acquisition unit 13 uses the noise sensor 32 to obtain the magnitude (volume) N2 of noise such as the running sound of the vehicle or the sound of air blowing out by the air conditioner l,r (f). Then, the external disturbance acquisition unit 13 determines the spectrum E of the sound input to both ears of the user as follows l,r (f) and the magnitude N2 of the noise l,r (f) to obtain the noise error w l,r (f feat , azi, ele). SNR l,r (f) = E l,r (f) / N2 l,r (f) w l,r (f feat , azi, ele) = b / SNR l,r (f) In addition, the external disturbance acquisition unit 13 may acquire external disturbances other than the amount of change in the head position and the volume of the noise
[0024] The HRTF enhancement unit 14 enhances the characteristic amount of the HRTF according to the external disturbance state acquired by the external disturbance acquisition unit 13 quantity. Fig. 4 is an explanatory diagram of the process of enhancing the characteristic amount of the HRTF. In the graph of Fig. 4, the horizontal axis represents the frequency and the vertical axis represents the level, and the solid line 51 represents the HRTF. In Fig. 4, state A represents the HRTF before enhancement. The HRTF in state A is, for example, the HRTF for each user acquired by the HRTF acquisition unit 12. Hereinafter, state A is also referred to as the reference state. The HRTF enhancement unit 14, for example, increases the amplitude of the HRTF in the reference state to perform enhancement as in state B. At this time, in the waveform of the HRTF, the mountain-shaped portions protruding upward are defined as peaks P1 to P3, and the valley-shaped portions depressed downward are defined as notches N1 to N2. Specific characteristic amounts may be enhanced using the levels of these peaks P1 to P3 and notches N1 to N2 as characteristic amounts
[0025] For example, the HRTF emphasis unit 14 may emphasize by reducing the levels of the notches N1 to N2 while keeping the peaks P1 to P3 in the reference state, or by increasing the levels of the peaks P1 to P3 while keeping the notches N1 to N2 in the reference state, etc.
[0026] Also, when the HRTF emphasis unit 14 identifies in order from the lower frequency side to the higher frequency side of the HRTF as the first peak P1, the first notch N1, the second peak P2, the second notch N2, etc., it may emphasize the peaks P1 to P3 and the notches N1 to N2 within a predetermined frequency range (for example, a frequency band where the difference in HRTF is easily felt) (the peaks become higher and the notches become lower). For example, it may be configured to emphasize only the first notch N1, the second peak P2, and the second notch N2 and not emphasize other peaks and notches. In addition, the feature amount to be emphasized may be specified according to the traveling speed of the vehicle 1 and the frequency of the traveling noise. For example, when the traveling speed of the vehicle 1 is low, since there are many low-frequency components in the traveling noise, the HRTF emphasis unit 14 acquires the traveling speed of the vehicle 1, and when the traveling speed is equal to or lower than a predetermined value, it may emphasize a predetermined number of notches (for example, the first notch N1 with the lowest frequency) from the lower frequency side of the HRTF.
[0027] Also, since the wavelength is shorter at higher frequencies and the error due to head displacement or movement becomes larger, the peaks P1 to P3 and the notches N1 to N2 may be emphasized as the wavelength of the HRTF increases. Further, the HRTF emphasis unit 14 may emphasize by lowering the notch N1 of a specific frequency when the sound blown out by the air conditioner is loud, or by raising the peak P1 of a specific frequency when the wind noise is loud. Also, the feature amount of the HRTF may be, for example, the amount of variation of the curve (spectrum profile) connecting the peak vertices as the feature amount, and this spectrum profile may be emphasized in the same manner as the above-mentioned peaks and notches. Note that when the HRTF emphasis unit 14 emphasizes the feature amount of the HRTF as described above, the peak portion in the head transfer function may be emphasized more strongly as the level of the disturbance state is larger. Thereby, the HRTF emphasis unit 14 appropriately corrects the HRTF according to the magnitude of the disturbance, enabling acoustic reproduction that effectively and efficiently suppresses the influence of the disturbance.
[0028] The sound signal generating unit 15 processes the sound signal acquired from the sound source and generates a sound signal to be output to the speaker unit 40. For example, the sound signal generating unit 15 separates a center component and left and right sound components from the sound signal s(f) acquired from the sound source, and generates an enhanced HRTF:H' for each component S(f). l,r (f,azi,ele) are convolved to generate the acoustic signal for each headrest speaker. When a performance is held in a concert hall or live venue, the position of the main vocals (center component) and the guitar and drums (left and right components) can be recognized by the way the reverberation propagates (phase difference), etc. However, in the case of the headrest speakers H1 to H8, the way in which the sound propagates is simulated by convolving the HRTF. As a result, the sound signal generator 15 controls the sound so that when the user hears the sound output from the headrest speakers H1 to H8, the sound image of the sound components of each sound source is perceived at a predetermined position (for example, the position of each sound source when the content was recorded (the sound source position intended by the content creator)). In other words, the sound signal generator 15 of this embodiment localizes the sound image of the center component directly in front of the user.
[0029] The inverse filter unit 16 applies an inverse filter G to the acoustic signal from the sound source. ll,lr,rl,rr -1 (f, azi, ele) is used to cancel out crosstalk, and the received The inverse filter cancels the spatial characteristics up to the listening position. Note that the inverse filter can be obtained by, for example, finding the inverse matrix of the transfer function from the headrest speakers H1 to H8 to the listening position.
[0030] The output control unit 17 supplies the acoustic signal to the speaker units 40 via the amplifier 104, causing each speaker unit 40 to output sound.
[0031] FIG. 5 is a configuration diagram of an acoustic system 100 including a head unit 10. In FIG. 5, mainly the components necessary for explaining the features of the present embodiment are shown, and the description of general components is omitted. In other words, each component illustrated in FIG. 5 is a functional concept, and it is not necessarily physically configured as shown in the figure. For example, the specific form of the dispersion / integration of each functional block is not limited to that shown in the figure, and all or part of it can be functionally or physically dispersed / integrated in any unit according to various loads, usage situations, etc.
[0032] As illustrated in FIG. 5, the head unit 10 is an information processing device (computer) having a control unit 101, a memory 102, an input / output interface (IF) 103, and an amplifier 104 that are interconnected by a connection bus 110. In FIG. 5, the head unit 10 is configured to include the amplifier 104, but the head unit (signal processing device) and the amplifier 104 may be separate entities.
[0033] The control unit 101 controls the entire head unit and is configured by, for example, a CPU (Central Processing Unit), an MPU (Micro Processing Unit), a main storage device, etc. The control unit 101 is also referred to as a controller or a processor. The control unit 101 is not limited to a configuration including a single processor and may have a multiprocessor configuration. Also, a single control unit 101 connected by a single socket may have a multi-core configuration. The main storage device is used, for example, as a work area of the control unit 101, a storage area for programs and data, and a buffer area for communication data. The main storage device includes, for example, a Random Access Memory (RAM), or a combination of a RAM and a Read Only Memory (ROM). Note that, when the control unit 101 is configured to include a DSP (digital signal processor), which is a processor dedicated to audio signal processing, the audio signal may be processed by the DSP, and the CPU may control the operation of the DSP. For example, the CPU may output various parameter values required for the audio signal to the DSP, and the DSP may perform arithmetic processing on the audio signal using the various parameter values to generate a desired signal.
[0034] The memory 102 is an auxiliary storage device that stores programs executed by the control unit 101, operation setting information, and the like. The memory 102 is not limited to an internal storage device built into the head unit 10, and may be an external storage device such as an external attached storage device or a NAS (Network Attached Storage). The memory 102 is, for example, an HDD (Hard-disk Drive), an SSD (Solid State Drive), an EPROM (Erasable Programmable ROM), a flash memory, a USB memory, a memory card, or the like.
[0035] The input / output IF 103 is an interface for inputting and outputting data between other devices such as a content server, a speaker unit 40, an amplifier 104, and an ECU. The input / output IF 103 performs input and output of data, for example, with a disk drive that reads data from a storage medium such as a CD or a DVD, an operation unit that receives operations by a user, display devices 61 and 62 that perform displays for the user, a communication module, and other devices. Also, the input / output IF 103 performs input and output of data, for example, with a tuner that receives radio or TV broadcast waves, a reader / writer that reads and writes data to a storage medium such as a memory card, a camera (head sensor) 31, a microphone (noise sensor) 32, and other sensors 30. The operation unit It is an input means that receives an operation by and inputs operation information indicating this operation to the control unit 101. The operation unit may be, for example, various switches such as push button switches, a volume operated by a dial (rotary knob), or a touch panel provided on the display surface of the display device 61. The display devices 61 and 62 are output means for displaying information related to music playback and the like to the user, and are realized by a liquid crystal display panel or the like. The communication module is an interface for communicating with other devices such as a content server and a speaker unit via a communication line. Note that a plurality of each of the above components may be provided, or some of the components may not be provided.
[0036] In the head unit 10, the control unit 101 functions as each processing unit such as the sound signal acquisition unit 11, the HRTF acquisition unit 12, the disturbance acquisition unit 13, the HRTF enhancement unit 14, the sound signal generation unit 15, the inverse filter unit 16, and the output control unit 17 shown in FIG. 3 by executing an application program. However, at least part of the processing of each of the above processing units may be provided by a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), or the like. Also, at least part of each of the above processing units may be a dedicated LSI (large scale integration) such as an FPGA (Field-Programmable Gate Array) or other digital circuits. Also, a configuration including an analog circuit in at least part of each of the above processing units may be used.
[0037] [Signal Processing Method] FIG. 6 is a flowchart when the control unit executes the signal processing method according to the present embodiment based on a signal processing program. For example, when the accessory power supply of the vehicle 1 is turned on and power is supplied to the head unit 10, or when content playback is instructed by a user operation, the control unit 101 starts the processing of FIG. 6. The processing of FIG. 6 is repeatedly executed until the accessory power supply of the vehicle 1 is turned off and the operation of the head unit 10 stops, or until termination is instructed by a user operation. The cycle of repeating FIG. 6 only needs to be such that music can be played back approximately in real time. For example, when an acoustic signal is stored in a buffer for several milliseconds to several seconds, the cycle may be such that it can be repeated several times within the period of this buffer.
[0038] In step S10, the control unit 101 acquires an acoustic signal from a sound source device such as a CD, DVD, or semiconductor memory.
[0039] In step S20, the control unit 101 acquires the HRTF according to the user. For example, the control unit 101 causes the user to select a pattern suitable for the user from a plurality of patterns in advance by the operation of the HRTF acquisition unit 12 and stores it in the memory 102, and acquires the HRTF of this stored pattern from the memory 102. Note that the control unit 101 acquires the HRTF of the user for each of the seats 21 to 24 so that the subsequent processing can be performed with the HRTF for each of the seats 21 to 24.
[0040] In step S30, the control unit 101 acquires the magnitude of the disturbance via the sensor 30. For example, the control unit 101 acquires the amount of change in the head position of each user by the head sensor 31 and acquires the volume of the noise by the noise sensor 32.
[0041] In step S40, the control unit 101 emphasizes the feature amount of the HRTF according to the magnitude of the disturbance acquired in step S30 and generates an emphasized HRTF.
[0042] In step S50, the control unit 101 processes the acoustic signal so that when the acoustic signal is convolved with the emphasized HRTF and the sound based on the acoustic signal is output from the headrest speakers H1 to H8, the sound image is localized at a predetermined position in front of each user. That is, the sound image of the center component is localized in front of the user, and the sound images of the left and right components are localized to the left and right of the user respectively to generate the acoustic signal.
[0043] In step S60, the control unit 101 processes the acoustic signal using an inverse filter determined based on the inverse matrix of the transfer function from the headrest speakers H1 to H8 to the listening position, etc., and cancels the spatial characteristics from the headrest speakers H1 to H8 to the listening positions of the users sitting on each of the seats 21 to 24.
[0044] In step S70, the control unit 101 supplies the acoustic signal processed in step S60 to each of the headrest speakers H1 to H8 to output sound.
[0045] FIG. 7 is a diagram for explaining the effect of emphasizing the HRTF. As a comparative example, the head unit 10 performs convolution processing of the acoustic signal acquired from the sound source 70 with the HRTF 72 acquired by the HRTF acquisition unit 12, and processes the acoustic signal so as to cancel the spatial characteristics with the inverse filter 73 of the inverse filter unit 16, and supplies the acoustic signal to the headrest speakers H1, H2 to output sound. In this case, when the output sound reaches the listening position, the characteristics 74 based on the HRTF of the sound are superimposed with errors due to disturbances. Also, when the inverse filter 73 is created using a dummy head, since it is different from the characteristics of an actual user, this difference is also superimposed as an error. Then, disturbances such as fluctuations in the user's head position and noise increase, and as shown in FIG. 7, errors 91, 92 are superimposed on the characteristics 74, and when the characteristics of the characteristics 74 are impaired, the sound image cannot be localized at the target position.
[0046] In contrast, in this embodiment, the head unit 10 emphasizes the HRTF 72 acquired by the HRTF acquisition unit 12, convolves the emphasized HRTF 75 with the acoustic signal acquired from the sound source 70, cancels the spatial characteristics using the inverse filter 73, and supplies it to the headrest speakers H1 and H2 to output sound. In this case, even if an error due to the same disturbance as in the comparative example is superimposed on the sound reaching the listening position, since the feature amount of the HRTF 75 is emphasized, the errors 91 and 92 are relatively small. As a result, the manner in which the characteristics 74 based on the HRTF 75 are impaired is reduced, and the sound image is localized at the target position.
[0047] [Effects of the Embodiment] (1) The head unit (signal processing device) 10 of this embodiment corrects the HRTF based on the disturbance state in the user's ambient acoustic environment. Thereby, the head unit 10 can suppress the influence of disturbances such as fluctuations in the user's head position and noise, appropriately obtain the acoustic effect using the head-related transfer function, and perform acoustic reproduction with a rich sense of presence.
[0048] (2) The head unit 10 of this embodiment processes the acoustic signal processed with the corrected HRTF using the inverse function of the spatial transfer function in the space from the speaker unit 40 where the signal-processed acoustic signal is output to the user's head position. For example, when the user listens to the sound output from the speaker unit 40 in front of the speaker unit 40, an inverse filter based on the spatial characteristics is used to process the acoustic signal from the sound source and cancel the spatial characteristics at the user's listening position. Also, when the head unit 10 outputs sound based on the acoustic signal from a plurality of speaker units 40, the head unit 10 processes the acoustic signal using the emphasized HRTF so as to localize the sound image in front of the listening position. Thereby, even when the speaker unit 40 is provided in the headrest of the seat and outputs sound from behind the user, the head unit 10 can localize the sound in front of the user. Also, even when the effect of processing the acoustic signal by the HRTF or the like is inhibited by a disturbance, the head unit 10 can localize the sound image at the target position by emphasizing the feature amount of the HRTF.
[0049] (3) The head unit 10 acquires the amount of change in the user's head position as a disturbance state. By this, even when the positional relationship between the listening position and the speaker unit 40 fluctuates and the effect by the processing of the acoustic signal on the spatial characteristics is inhibited, the head unit 10 can localize the sound image at the target position by emphasizing the characteristic amount of the HRTF. (4) The head unit 10 acquires the noise in the user's ambient acoustic environment as a disturbance state. By this, even when the sound output from the speaker unit 40 is masked by the noise and the effect by the processing of the acoustic signal is inhibited, the head unit 10 can localize the sound image at the target position by emphasizing the characteristic amount of the HRTF.
[0050] (5) The head unit 10 emphasizes the characteristic amount more strongly as the high-frequency component of the acoustic signal becomes higher. The higher the frequency of the voice, the greater the influence on the user's sense of localization. Conversely, the lower the frequency of the voice, the smaller the influence on the user's sense of localization. The head unit 10 of the present embodiment emphasizes the characteristic amount of the HRTF more strongly as the frequency becomes higher, so as to effectively and efficiently make the user feel the sense of localization of the sound image.
[0051] (6) The characteristic amount of the head unit 10 is at least one of the peak level, notch level, and spectral shape in the HRTF. In this way, the head unit 10 can effectively and efficiently make the user feel the sense of localization of the sound image by emphasizing the peak level, notch level, and spectral shape in the HRTF.
[0052] (7) The head unit 10 emphasizes the characteristic amount more strongly as the high-frequency component of the acoustic signal becomes higher. The higher the frequency of the voice, the greater the influence on the user's sense of localization. Conversely, the lower the frequency of the voice, the smaller the influence on the user's sense of localization. The head unit 10 of the present embodiment emphasizes the characteristic amount of the HRTF more strongly as the frequency becomes higher, so as to effectively and efficiently make the user feel the sense of localization of the sound image.
[0053] (7) In this embodiment, the head unit 10 is for a user who is a vehicle occupant, and a speaker is incorporated into the headrest of the seat on which the user sits. Further, the head unit 10 processes an acoustic signal using an emphasized HRTF so as to localize the sound image when the sound based on the acoustic signal from a sound source is output from the speaker unit 40 in front of the user. As a result, the head unit 10 can individually provide appropriate sound to each user sitting on each seat, and can appropriately perform acoustic control such as localizing the sound image in front of the user for each seat.
[0054] <Second Embodiment> FIG. 8 is a diagram showing the configuration of the acoustic system 200 according to the second embodiment, and FIG. 9 is a diagram showing the arrangement of the speaker unit 40 of the acoustic system 200 according to the second embodiment. This embodiment is different from the above-described first embodiment in that the speaker unit 40 is arranged in addition to the headrest, and the other configurations are the same. Therefore, in this embodiment, the same reference numerals are assigned to the same elements as those in the first embodiment described above, and the description thereof will be omitted again.
[0055] As shown in FIGS. 8 and 9, the acoustic system 200 of the present embodiment has, in addition to the headrest speakers H1 to H8, ten speaker units 40 (CTR, FR, WFR, ROR, RR, WF, FL, WFL, ROL, RL) installed in the vehicle. Among the plurality of speaker units 40, the speaker unit CTR is a so-called center speaker arranged at the front center inside the vehicle. The speaker unit FR is a speaker arranged on the front right side inside the vehicle. The speaker unit WFR is a woofer arranged on the front right side inside the vehicle and below the right front seat (driver's seat) 21. The speaker unit ROR is installed on the right side of the ceiling portion substantially in the center in the front-rear direction of the passenger compartment and is a speaker for suppressing reflected sound and environmental sound. The speaker unit RR is a speaker arranged on the rear right side inside the vehicle. The speaker unit WF is a woofer arranged at the rear center inside the vehicle. The speaker unit FL is a speaker arranged on the front left side inside the vehicle. The speaker unit WFL is a woofer arranged on the front left side inside the vehicle and below the left front seat (passenger seat) 22. The speaker unit ROL is installed on the left side of the ceiling portion substantially in the center in the front-rear direction of the passenger compartment and is a speaker for suppressing reflected sound and environmental sound. The speaker unit RL is a speaker arranged on the rear left side inside the vehicle.
[0056] As described above, the acoustic system 200 of the present embodiment arranges 18 speaker units 40 so as to surround each of the seats 21 to 24 of the vehicle 1, and outputs sound from around the users sitting on each of the seats 21 to 24, thereby enabling surround reproduction. The acoustic system 200 may perform acoustic control by outputting sound from each speaker unit 40 with the interior of the vehicle 1 as one listening space. Further, the acoustic system 200 may set a listening space for each of the seats 21 to 24, and perform acoustic control for each of the seats 21 to 24 by outputting sound from the speaker units 40 surrounding each of the seats 21 to 24. For example, in the case of the right front seat 21, acoustic control may be performed with the speaker units FR and FL as front speakers and the headrest speakers H1 and H2 as rear speakers. Similarly, in the case of the left front seat 22, acoustic control may be performed with the speaker units FR and FL as front speakers and the headrest speakers H3 and H4 as rear speakers. Further, in the case of the right rear seat 23, acoustic control may be performed with the speaker units RR and RL as front speakers and the headrest speakers H5 and H6 as rear speakers. Also, in the case of the left rear seat 24, acoustic control may be performed with the speaker units RR and RL as front speakers and the headrest speakers H7 and H8 as rear speakers. Even in the case where the front speakers and the rear speakers (headrest speakers H1 to H8) are arranged in combination like this, when outputting sound from the headrest speakers H1 to H8, similar to the first embodiment described above, the HRTF is emphasized based on the disturbance state, and acoustic control is performed using the emphasized HRTF, so that the influence of the disturbance can be suppressed and the sound image can be localized at the target position.
[0057] <Third Embodiment> Compared with the first embodiment described above, the present embodiment has a different configuration in which the sound image of the center component is localized at a predetermined direction away from the front of the user, and the other configurations are the same. Therefore, in the present embodiment, the same reference numerals are given to the same elements as those in the first embodiment described above, and the repeated description is omitted.
[0058] Figure 10 is an explanatory diagram of the position for localizing the sound image in the acoustic system 100 of the third embodiment. In the aforementioned first embodiment, the control unit 101 localizes the sound image 81 of the center component among the acoustic signals in the front direction D1 of the user. In this case, the user tends to feel that the sound image 81 is close, and it is difficult to perceive the distance to the sound image 81 (sound image distance). Therefore, the control unit 101 of this embodiment localizes the sound image 82 of the center component at a position separated from the front of the user in a predetermined direction using the emphasized HRTF. For example, as shown in Figure 10, the control unit 101 localizes the sound image 82 in the left front direction D2 rotated 30 degrees to the left from the front direction D1 around the vertical axis X1 passing through the center of the user's head. When the sound image is localized at a position deviated from the front of the user in this way, the sound image distance is relatively long. Therefore, the user can feel a spatial spread for the main sound components such as the main vocal, rather than feeling as if it is trapped in the head. The control unit 101 sets the localization position of the sound image as described above, and controls the level, phase (delay amount), etc. of the acoustic signals output from each speaker so that the sound image is localized at the localization position.
[0059] The control unit 101 may also localize the right and left components based on the acoustic signals at positions separated from (rotated from) their original positions in the same direction as the center component. In the example of Figure 10, the sound image 84 of the right component is localized at a position rotated 45 degrees from the position of the sound image 83 based on the acoustic signal. Also, the sound image 85 of the left component is localized at the position based on the acoustic signal without being rotated.
[0060] Note that the direction in which the sound image 82 of the center component is separated is the direction from each seat 21 - 24 toward the center in the vehicle width direction of the vehicle 1. For example, it is the left direction for the right front seat 21 and the right rear seat 23, and the right direction for the left front seat 22 and the left rear seat 24. In this case, as shown in Figures 1, 2, and 9, a display device 6 for the front seats is provided at the center in the vehicle width direction of the instrument panel or the center console. 1 is provided, and a display device 62 for the rear seats is provided on the ceiling portion at the center of the vehicle. As a result, when the control unit 101 causes these display devices 61, 62 to display the image of the content and outputs the sound of the content from the speaker units 40 of each seat 21 to 24, the user can perceive the image and the sound image of the center component in the same direction, can enjoy the content without discomfort, and can feel the depth of the sound more than when the sound image 81 is formed in the front.
[0061] In addition, the control unit 101 may perform head tracking according to the variation of the head position detected by the sensor 30, and localize the sound image 82 of the center component absolutely in the vehicle width direction center direction (for example, the direction where the display device is provided) with respect to each seat 21 to 24. Further, the control unit 101 may perform head tracking and localize the sound image 82 of the center component obliquely forward relative to the head of each user.
Explanation of Signs
[0062] 1: Vehicle 10: Head unit 11: Sound signal acquisition unit 12: HRTF acquisition unit 13: Disturbance acquisition unit 14: HRTF enhancement unit 15: Sound signal generation unit 16: Inverse filter unit 17: Output control unit 20: Seat 20H: Headrest 30: Sensor 40: Speaker unit 100: Audio system 101: Control unit 102: Memory 104: Amplifier 110: Connection bus 200: Audio system
Claims
1. A signal processing apparatus that performs signal processing on an acoustic signal using the head-related transfer function of a user who listens to the reproduced acoustic sound, acquiring the head-related transfer function of the user, acquiring a disturbance state in the ambient acoustic environment of the user, correcting the head-related transfer function based on the disturbance state, a signal processing apparatus.
2. Processing the acoustic signal processed by the corrected head-related transfer function using the inverse function of the spatial transfer function in the space from the speaker where the signal-processed acoustic signal is output to the head position of the user, The signal processing apparatus according to claim 1.
3. The signal processing apparatus according to claim 1 or 2, wherein correction is performed to strongly emphasize the peak portion in the head-related transfer function as the level of the disturbance state increases.
4. The signal processing apparatus according to claim 1 or 2, wherein the disturbance state is the amount of variation in the head position of the user.
5. The signal processing apparatus according to claim 1, wherein the disturbance state is noise in the ambient acoustic environment of the user.
6. The signal processing apparatus according to claim 3, wherein correction is performed to strongly emphasize the feature amount as the high-frequency component of the acoustic signal increases.
7. A signal processing method for performing signal processing on an acoustic signal using the head-related transfer function of a user who listens to the reproduced acoustic sound, acquiring the head-related transfer function of the user, acquiring a disturbance state of the ambient acoustic environment of the user, correcting the head-related transfer function based on the disturbance state a signal processing method.
8. A signal processing program for performing signal processing on an acoustic signal using the head-related transfer function of a user who listens to the reproduced acoustic sound, acquiring the head-related transfer function of the user, acquiring a disturbance state of the ambient acoustic environment of the user, correcting the head-related transfer function based on the disturbance state a signal processing program for causing a computer to execute the processing.
9. An acoustic system including a speaker that outputs reproduced acoustic sound in response to an input acoustic signal, and a signal processing apparatus that performs signal processing on the acoustic signal from a sound source, wherein the signal processing apparatus is a signal processing apparatus that performs signal processing on an acoustic signal using the head-related transfer function of a user who listens to the reproduced acoustic sound, acquiring the head-related transfer function of the user, acquiring a disturbance state of the ambient acoustic environment of the user, correcting the head-related transfer function based on the disturbance state, performing signal processing on the acoustic signal using the corrected head-related transfer function, and outputting the signal-processed acoustic signal to the speaker an acoustic system.
10. A speaker that outputs an acoustic reproduction sound in response to an input acoustic signal, and A signal processing device that performs signal processing on the acoustic signal from the sound source, and is an acoustic system mounted on a vehicle, wherein The speaker is A headrest speaker mounted on the headrest of the vehicle, and The signal processing device is A signal processing device that performs signal processing on the acoustic signal using the head-related transfer function of the user who listens to the acoustic reproduction sound, and Obtains the head-related transfer function of the user, Obtains the disturbance state of the ambient acoustic environment of the user, Corrects the head-related transfer function based on the disturbance state, Performs signal processing on the acoustic signal using the corrected head-related transfer function, and Outputs the signal-processed acoustic signal to the speaker Acoustic system.
Citation Information
Patent Citations
Signal processing device, signal processing method, and program
WO2020036077A1