Audio reproduction system
The acoustic reproduction system addresses challenges in sound quality and localization by using a signal processing unit to generate optimized signal components for speakers with specific angles, ensuring front localization and presence even in non-ideal arrangements and during head rotations.
Patent Information
- Application Number
- PCT/JP2024/040682
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-21
- Filing Date
- 2024-11-15
- Publication Date
- 2025-05-30
AI Technical Summary
Existing audio reproduction systems face challenges in achieving ideal sound quality and stable sound image localization, especially in environments with obstacles and during head rotations, due to limitations in speaker arrangement and virtual sound source processing.
The proposed acoustic reproduction system includes a signal processing unit that generates specific signal components for left and right speakers, as well as a front speaker, to ensure front localization and presence, even with non-ideal speaker arrangements. This system sets the opening and elevation angles of the left and right speakers within specific ranges and uses virtual sound source processing to maintain stable front localization during head rotations.
The system effectively provides front localization and a sense of presence even when a basic speaker arrangement is not feasible, and it maintains stable front localization during head rotations, enhancing sound quality and user experience.
Smart Images

Figure JP2024040682_30052025_PF_FP_ABST
Abstract
Description
Sound reproduction system
[0001] The present disclosure relates to sound reproduction systems.
[0002] Conventionally, there is a technique for acoustic correction of a speaker.
[0003] For example, there is a technology for audio devices that corrects deviations in the center sound image within a playback sound field, asymmetric sound field spread, etc. (See Patent Document 1.) This technology calculates transfer functions based on impulse response sequences of the left and right speakers, and configures the audio device to include a correction circuit made up of transfer functions obtained by an inverse matrix with the transfer functions as elements.
[0004] There is also technology relating to an audio processing device that allows a listener to hear sound that is always in an appropriate state even if the position of the listener's head relative to the speaker is not constant (see Patent Document 2), and technology relating to a headphone device that minimizes the deviation between the image position and the sound image position (see Patent Document 3).
[0005] Furthermore, there is a technique relating to a central signal extraction process that generates a signal component with high correlation and a signal component with low correlation from two inputs, one on the left and one on the right (see Patent Documents 4 to 6).
[0006] There are also other techniques related to virtual sound source generation (see Non-Patent Documents 1 and 2).
[0007] Japanese Patent Application Laid-Open No. 2001-057699 Japanese Patent Application Laid-Open No. 2003-092799 Japanese Patent Application Laid-Open No. 08-009489 Japanese Patent Application Laid-Open No. 2009-025500 Japanese Patent Application Laid-Open No. 2009-027388 Japanese Patent Application Laid-Open No. 2011-023961
[0008] Virtual Sound Source Generation Technology for Stereo Reproduction, Yoshitaka Murayama and Haruo Hamada, JAS Journal 2011 Vol.51 No.1 (January issue) Sound Field Reproduction by Transaural System Using Two Speakers, ASJ 2011
[0009] However, if it is difficult to achieve ideal speaker placement inside a vehicle, etc., problems may arise such as deterioration of sound quality due to primary reflections caused by obstructions near the speakers, and it may be difficult to stably position the speakers forward.
[0010] Furthermore, when using virtual sound source processing to improve localization, there is a problem that the transfer function of the playback system changes when the listener's head moves during listening, resulting in sound image localization, sound field impression, and sound quality that are not as intended.
[0011] Patent Documents 1 and 2 propose a method of detecting head movement with a sensor and reflecting it in virtual sound source processing, but do not describe how to deal with head rotation, which frequently occurs when listening to an in-car system, for example.
[0012] Furthermore, the headphone playback in Patent Document 3 involves a method that uses a gyroscope to accommodate head rotation, which may be applicable, but there are problems with accuracy and the need to handle huge amounts of data, which increases the cost of the system.
[0013] The objective of the present disclosure is to provide an audio reproduction system that can provide a sound source environment in which forward localization and a sense of realism can be obtained even when basic speaker placement is not possible, and in which stable forward localization can be obtained even when head rotation occurs when virtual sound source processing is used.
[0014] The sound reproduction system of the present disclosure includes a signal processing unit that performs signal processing on multiple channels including at least a left channel signal L and a right channel signal R, a left speaker and a right speaker that are arranged on the left and right of the listener at a predetermined angle of spread or less and a predetermined angle of elevation or more, and a front speaker that is positioned further forward from the listener than the left speaker and the right speaker, and the signal processing unit generates a front speaker signal SF, a left speaker signal SLt, and a right speaker signal SRt from the left channel signal L and the right channel signal R, respectively, and the front speaker signal SF includes at least 200 to 9 kHz of the components that are highly correlated between L and R, and the left speaker signal SLt mainly comprises components of the left channel signal L that are low in correlation with the right channel signal R, and the right speaker signal SRt mainly comprises components of the right channel signal R that are low in correlation with the left channel signal L.
[0015] In addition, in the sound reproduction system, the left speaker and the right speaker may be arranged such that the opening angle is an angle within a predetermined range of 110° or less and not exceeding the listener's line of sight in front of them, and the elevation angle is an angle within a predetermined range of 30° or more and not exceeding the apex on the ear side of the listener.
[0016] The sound reproduction system may also be configured such that the front speakers are composed of a front left speaker and a front right speaker positioned in a plane parallel to the coronal plane to the left and right of the intended listener and facing the direction of the listener.
[0017] Furthermore, in the sound reproduction system, the signal processing unit may be configured to reproduce signal components in at least the band of a lower frequency of 0 to 200 Hz and an upper frequency of 400 to 600 Hz through the front speakers, regardless of the left-right correlation between the left speaker signal SLt and the right speaker signal SRt.
[0018] In addition, in the sound reproduction system, the listener may be a plurality of listeners, and the left speaker and the right speaker may be placed on the left and right sides of each listener's head, respectively, or the left speaker may be placed at the left end of each listener lined up horizontally in the left-right direction, and the right speaker may be placed at the right end.
[0019] Furthermore, in the sound reproduction system, the opening angle of the left speaker and the right speaker may be set to a range of 70 to 110°, and the elevation angle may be set to a predetermined range of 40° or more but not exceeding the apex on the ear side of the listener, and the signal processing unit may perform virtual sound source processing on the left speaker signal SLt and the right speaker signal SRt to reproduce the signals.
[0020] Furthermore, in the case of an audio reproduction system that is installed as an in-vehicle speaker system, the left speaker and the right speaker may be placed above the ears of the listeners in each seat, the front speakers may be placed in positions related to the dashboard or the bottom of the left and right doors, and the signal processing unit may be incorporated into the in-vehicle audio system.
[0021] The sound reproduction system of the present disclosure has the advantage of being able to provide a sound source environment in which forward localization and a sense of realism can be obtained even when basic speaker placement is not possible, and in which stable forward localization can be obtained even when head rotation occurs when virtual sound source processing is used.
[0022] FIG. 1 is a diagram showing a schematic configuration of an audio reproduction system. FIG. 2A is a diagram showing an example of a speaker arrangement. FIG. 2B is a diagram showing an example of a speaker arrangement. FIG. 2C is a diagram showing an example of a speaker arrangement. FIG. 3 is an example of a signal processing mode of a signal processing unit. FIG. 4 is an example of a graph explaining the basis for frequencies to ensure a sense of realism. FIG. 5A is an example of the time response and frequency response characteristics of an HRTF. FIG. 5B is an example of the time response and frequency response characteristics of an HRTF. FIG. 5C is an example of the time response and frequency response characteristics of an HRTF. FIG. 6 is the frequency response characteristics of an HRTF in front of a sound source. FIG. 7 is an example of a configuration having a front left speaker and a front right speaker. FIG. 8 is a diagram showing the overall signal processing flow of a second embodiment. FIG. 9 is an example for two listeners arranged side by side. FIG. 10 is an example in which an LSP and an RSP are installed for each listener. FIG. 11 is an example of a signal generation unit of a third embodiment. FIG. 12 is an example of a signal generation unit of the third embodiment. FIG. 13 is a diagram showing the overall signal processing flow of the third embodiment. FIG. 14 is an example of an arrangement in an experimental example where the opening angle is 90° and the elevation angle is 30°. FIG. 15 is an example of an arrangement where the opening angle is 150° and the elevation angle is 0°. FIG. 16 is a diagram showing changes during head rotation. FIG. 17 is a graph showing changes in HRTF_L at each elevation angle during head rotation. FIG. 18 is a graph showing changes in HRTF_L at each elevation angle during head rotation. FIG. 19 is a graph showing sound pressure differences and phase differences for each opening angle and elevation angle. FIG. 20 is a graph showing sound pressure differences and phase differences for each opening angle and elevation angle. FIG. 21 is a graph showing sound pressure differences and phase differences for each opening angle and elevation angle. FIG. 22 is an example of speaker arrangement when applied to a listening chair, etc. FIG. 23 is an example of center signal extraction processing when applied to a listening chair, etc. Fig. 24 is an example of center signal extraction processing when applied to a listening chair or the like. Fig. 25A is an example of speaker placement when applied to an in-vehicle system. Fig. 25B is an example of speaker placement when applied to an in-vehicle system. Fig. 26 is an example of center signal extraction processing when applied to an in-vehicle system. Fig. 27 is an example of center signal extraction processing when applied to an in-vehicle system. Fig. 28 is an example of a signal generation unit when the front speakers are capable of high-quality sound reproduction across the entire frequency range.FIG. 29 shows an example of center signal extraction processing when the front speakers have high bass reproduction capabilities and the LSP and RSP speakers are installed suitable for mid- and high-range reproduction.
[0023] An example of an embodiment of the disclosed technology will be described below with reference to the drawings. Note that the same or equivalent components and parts in each drawing are given the same reference numerals. Also, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.
[0024] The premise and overview of the embodiments of the present disclosure will be described.
[0025] Most common 2-channel stereo sources are designed to be played back on a system with basic speaker placement (usually 30 degrees to the left and right of the front) to create a sound image in front of the listener.
[0026] Many in-car audio systems have at least one speaker on each side of the listener, and they are located in various locations, such as near the door feet (below), on the roof liner (ceiling), or in the rear tray (rear). When speakers are located below the doors, the sound image is localized forward, but the diaphragm is not facing the listener, which means high frequencies are easily attenuated. Furthermore, the seats and other sound-insulating and reflective objects can cause problems, particularly with high-frequency attenuation and sound quality degradation. When speakers are located on the roof liner, the distance between the listener and the speaker is short, making them less susceptible to the effects of reflected sound compared to other placements, resulting in high-quality signal reproduction. However, the sound image is located at an upward angle, which creates a sense of incongruity. When speakers are located on the rear tray, both the sound image being located at the rear creates a sense of incongruity and the sound quality is significantly degraded by the effects of reflected sound.
[0027] To address the problem of sound image localization, the transfer function between the loudspeaker and the ear (HRTF: Head related transfer function) is used. The term "between the loudspeaker and the ear" generally refers to the distance from the sound source to the entrance of the ear canal or above the eardrum. In the method using the HRTF, a method is often applied in which a reproduced sound field in a virtual loudspeaker arrangement is obtained by performing a process (virtual sound source processing) in which the HRTF in a virtual intended loudspeaker arrangement and the inverse transfer function of the transfer function of the reproduction system including crosstalk are input and reproduced.
[0028] As mentioned above, there are problems with sound quality and stability depending on the speaker arrangement. Therefore, the method according to this embodiment provides a sound reproduction system that can obtain forward localization and a sense of realism even when basic speaker arrangement is not possible, and can obtain stable forward localization even when head rotation occurs when using virtual sound source processing.
[0029] [First Embodiment] A first embodiment will be described. Fig. 1 is a diagram showing a schematic configuration of an audio reproduction system. As shown in Fig. 1, the audio reproduction system 1 includes a signal processing unit 10, a left speaker 12L and a right speaker 12R arranged on the left and right sides of the head of a listener (UL), and a front speaker 14 located forward and farther from the listener than the left and right speakers. The signal processing unit 10 processes signals on multiple channels including at least a left channel signal L and a right channel signal R. The signal processing unit 10 may be realized by the hardware configuration of a computer having a CPU, RAM, storage, etc. (not shown).
[0030] The left speaker 12L and the right speaker 12R are disposed on the left and right of the listener at a predetermined angle of opening or less and a predetermined angle of elevation or more. In the following description, for ease of explanation, when the speakers are referred to by their abbreviations, the left speaker 12L will be referred to as (LSP), the right speaker 12R will be referred to as (RSP), and the front speaker 14 will be referred to as (FSP). In some embodiments described below, there may be a plurality of left speakers 12L, right speakers 12R, and front speakers 14.
[0031] <Combination of Speaker Arrangement and Signal Processing> Here, the combination of speaker arrangement and signal processing in the method according to this embodiment will be described. The signal processing unit 10 extracts signal band components (SF) of center-localized sound signals (signal components with high internal correlation between the L and R signals) in a frequency band of at least 200 Hz to 9 kHz (described later in A1) from the L and R signals. The signal processing unit 10 also reproduces L signal components SLt and R signal components SRt in a band of at least 600 Hz or higher (described later in A2) of low-correlation signals. The L signal components SLt and R signal components SRt are reproduced from LSPs and RSPs, respectively, positioned closer to the listener than the FSPs, with an elevation angle of 30° or more and an opening angle of 110° or less (described later in A3). The signal band component SF is reproduced from the FSPs. Note that a position closer to the listener than the FSPs may be any position that is relatively close. Each signal is corrected for time differences and level differences between at least the FSP and LSP (or RSP) caused by differences in speaker distance and efficiency.
[0032] 2A to 2C are diagrams showing examples of speaker arrangements. Fig. 2A is a diagram (2A) showing the listener (UL) as seen from above, Fig. 2B is a diagram (2B) showing the listener (UL) as seen from behind, and Fig. 2C is a diagram (2C) showing the listener (UL) as seen from the left side. The speaker arrangement examples state that "opening angle of LSP and RSP: 110° or less (90° in the example)" and "elevation angle of LSP and RSP: 40° or more (60° in the example)."
[0033] FIG. 3 shows an example of the signal processing mode of the signal processing unit. (3A) is a diagram showing the overall signal processing flow of the signal processing unit 10 of the first embodiment. (3B) is an example of the signal generation unit 10A, showing an example in which the 200 to 9 kHz band of center-localized sound is output from the FSP. (3C) is an example of a circuit mode including a center signal extraction unit 10C that performs center signal extraction processing. Note that the center signal extraction processing method can be applied with reference to the techniques of Patent Documents 4 to 6. For example, a signal processing method using an adaptive filter may be used, and other configurations may also be used as long as they are capable of separating the signal of the center-localized sound from other signals. An example of the method is described below.
[0034] As shown in (3A), the signal processing unit 10 receives a left channel signal L (Lch signal) from the LSP and a right channel signal R (Rch signal) from the RSP, and generates an L signal component SLt, an R signal component SRt, and a signal band component SF as signals in the signal generation unit 10A. Dly is a delay corresponding to the difference in distance between the LSP and RSP and the FSP. Part 10B performs gain control to correct the difference in efficiency between the LSP and RSP and the FSP and the level difference due to the difference in distance, and performs power amplification to drive the speakers.
[0035] The signal generation unit 10A (3B) extracts signals in a frequency band of 200 Hz to 9 kHz from the L signal component SLt and the R signal component SRt via a bandstop filter (BSF: L1, R2) and a bandpass filter (BPF: L2, R1). Two bandpass filters are provided on each side. Dly is a delay corresponding to the center signal extraction unit 10C. The center signal extraction unit 10C extracts low-correlation signals SLt', SRt', and high-correlation signals SF from the L-side input Sil and the R-side input Sir. SLt' and SRt' are added to a delay signal in a calculator to output the signal component SLt and the R signal component SRt.
[0036] In the center signal extraction unit 10C of (3C), the filter coefficients for extracting highly correlated signals are updated so as to minimize the error between the left and right signals, and the filter coefficients are convolved with the L-side input Sil and the R-side input Sir via an adaptive filter (ADF) to generate SLt' and SRt'. Alternatively, any other configuration may be used as long as it can separate the signal of the sound localized in the center from other signals.
[0037] In this way, the signal processing unit 10 generates a front speaker signal SF, a left speaker signal SLt, and a right speaker signal SRt from the signal L and the signal R. The front speaker signal SF includes at least 200 to 9 kHz of the components highly correlated between L and R. The left speaker signal SLt mainly comprises components of the left channel signal L that have a low correlation with the right channel signal R, and the right speaker signal SRt mainly comprises components of the right channel signal R that have a low correlation with the left channel signal L.
[0038] Next, the principles (A1) to (A3) that characterize the method of this embodiment will be described.
[0039] The frequency bands listed in (A1) are frequency bands that are particularly effective for forward localization of centrally localized sound. If sound with a frequency of 9 kHz or less is reproduced from the FSP, it can be perceived as being sufficiently forward as the head rotates. Conversely, frequencies above 9 kHz have short wavelengths, making it difficult to perceive direction based on interaural phase differences. Furthermore, the sensitivity of the ears may be low, and the contribution to horizontal localization is low, so sound may be reproduced from upper SPs. In this case, high-frequency sounds, which are strongly affected by obstructions, are reproduced from nearby LSPs and RSPs, thereby enabling high-quality sound (in this case, there is no need to place dedicated high-frequency speakers in the front). Furthermore, since the 7-8 kHz sound source, which contributes to the sense of upward localization within centrally localized sound, is not reproduced from the LSPs and RSPs located above, a strong sense of upward localization can be avoided (the 7-8 kHz range, which contributes to the sense of upward localization, will be explained later in "Perception of Sound Source Direction").
[0040] The frequencies in the 600 Hz band or higher listed in (A2) are frequencies for ensuring a sense of realism. In order to ensure a sense of realism, it is necessary that signals with low correlation between the left and right channels can be heard in a state where there is a large interaural difference (difference in sound pressure between the left and right ears). In the case of an example of speaker placement with an elevation angle of 30° or higher and an opening angle of 90°, the interaural difference (HRTF_R-L) for the output from the left ear is particularly clear above 600 Hz. Therefore, a sense of realism can be ensured by reproducing signals with low correlation, at least in the frequency band above 600 Hz, from speakers placed on the left and right ears. Note that signals with low correlation contribute to a sense of realism as signals output from one speaker.
[0041] 4 is an example of a graph explaining the basis for the frequency required to ensure a sense of realism. The graph shows the interaural difference (the value obtained by subtracting the sound pressure HRTF-R at the right ear from the sound pressure HRTF-L at the left ear) for each elevation angle with an LSP having an opening angle of 90°. The greater the interaural difference between signals with low correlation, the greater the sense of realism felt. In the graph, the interaural difference is clear (approximately 3 dB or more) at 600 Hz or higher at any elevation angle, so that a sense of realism can be ensured by reproducing at least 600 Hz or higher from the LSP and RSP.
[0042] Here, we will explain HRTF. HRTF, which serves as a clue for identifying the direction of a sound source through hearing, mainly includes the following elements: (1) Interaural difference: The difference in sound pressure level, time difference, and phase difference that occurs between the left and right ears contributes greatly to the perception of left and right direction. If there is no interaural difference, the sound source can be recognized as being on the median plane, but it is difficult to identify the direction. (2) Sound pressure frequency characteristics: These are characterized by differences depending on the direction of the sound source due to the influence of reflection and resonance by the pinna. They vary greatly not only in the left and right (horizontal) direction but also in the up and down direction, so they also affect the perception of up and down direction.
[0043] Figures 5A to 5C show examples of the time response and frequency response characteristics of HRTFs. Figure 5A shows an example (5A) where the sound source direction is 30 degrees to the front left (elevation angle 0°). The time response (5B) of Figure 5B and the frequency response characteristics (5C) of Figure 5C are obtained as HRTFs from a sound source in this direction to the left and right ears (eardrums). Meanwhile, Figure 6 shows the frequency response characteristics of the HRTF for a sound source in front of the ear when the elevation angles are 0°, 20°, and 40°, respectively.
[0044] Here, we will explain "sound source direction perception." The sound pressure characteristics of HRTFs differ in all directions of the sound source, and can therefore serve as a clue for identifying the sound source direction in all directions. However, in past experiments on the direction perception of stationary sound sources, the accuracy rate was low for the sound source direction in the median plane and in the front-to-back direction, suggesting that the interaural difference is more important than the sound pressure characteristics of HRTFs. Furthermore, for stationary sound sources, it is usually difficult to determine whether the characteristics of the generated sound pressure characteristics are due to the HRTF or the characteristics of the sound source. However, changes in the sound pressure characteristics caused by the movement of the sound source can be perceived as changes in the transfer function, making identification easier. When a person tries to clearly identify the location of a sound source, rotating or tilting their head has the same effect as moving the sound source, and it is thought that the direction is clearly identified by sensing "changes" in the sound pressure characteristics and "changes" in the interaural difference. It has also been confirmed that when the level of 7 to 8 kHz is raised for a sound source reproduced by a speaker, as in the case of HRTF changes due to elevation angle, the sound image appears to rise.
[0045] The arrangement listed in (A3) is an arrangement that does not make the listener feel as if the sound source is located at the LSP and RSP when the listener's head rotates, i.e., an arrangement that minimizes changes in HRTF when the listener's head rotates. The LSP and RSP, which mainly output signals with low LR correlation, are arranged to the left and right of the listener, respectively. Here, the arrangement is such that a strong sense of localization does not occur in the direction of the LSP and RSP so that the sense of forward localization is not impaired when the listener's head rotates. Based on the results of the "Listening Experiment" and "Confirmation of HRTF Related to Listening Experiment" described below, the elevation angle is set to 40° or more and the opening angle is set to 110° or less. The opening angle ranges from 110° or less to an angle that does not exceed the line of sight when the listener is facing forward, and the elevation angle ranges from 30° or more to an angle that does not exceed the vertex on the listener's ear side (or the vertex of the head).
[0046] As described above, the sound reproduction system of the first embodiment can provide forward localization and a sense of realism even when basic speaker placement is not possible, and can provide a sound source environment in which stable forward localization can be obtained even when head rotation occurs when virtual sound source processing is used.
[0047] Furthermore, many music sources, TV, and radio content are created to center the main instruments and voices, including the basic instruments, vocals, and narration. By reproducing these components with front speakers, the overall sound is centered, even when the LSPs and RSPs are positioned above. Furthermore, by not reproducing the 6-9 kHz centrally positioned sound from the upper LSPs and RSPs, the sound image of the centrally positioned sound can be prevented from shifting in the height direction of the upper LSPs and RSPs. Furthermore, when the head rotates, the frontal positioning provided by the FSPs dominates the upper positioning provided by the upper speakers, maintaining the sense of frontal positioning. Because the LSPs and RSPs are positioned close to the listener, the influence of primary reflections is reduced, resulting in higher sound quality. Furthermore, drive power can be reduced.
[0048] Furthermore, by setting the elevation angle of the LSP and RSP to 30° or more, (1) it is possible to reduce the change in HRTF when the head rotates, (2) it is possible to reduce the change in the distance between the SP and the ear when the head rotates, and (3) it is possible to reduce the peaks and dips in the sound pressure characteristics. These effects make it possible to suppress the strong perception of the real sound source position (SP position) due to the sound emitted from the LSP and RSP when the head rotates. Furthermore, when virtual sound source processing is used, it is possible to increase robustness against head rotation (making it less likely that localization in an unintended direction or echo sensation will occur), and it is possible to maintain forward localization even when the head moves more than expected.
[0049] Furthermore, the left speaker 12L (LSP) and the right speaker 12R (RSP) may have an opening angle in the range of 70 to 110°, and an elevation angle in a predetermined range of 40° or more but not exceeding the apex on the listener's ear side. Furthermore, the signal processing unit 10 may perform virtual sound source processing on the left speaker signal SLt and the right speaker signal SRt before playback. This minimizes changes in interaural differences when the head rotates, enabling more stable playback with virtual sound source processing. This can prevent the phenomenon of unclear sound localization and increased echo when the head rotates.
[0050] Second Embodiment In the second embodiment, two front speakers are arranged. Note that the same parts as those in the first embodiment are denoted by the same reference numerals and the description thereof will be omitted.
[0051] In the second embodiment, the front speakers 14 include a front left speaker 14L (FLSP) and a front right speaker 14R (FRSP). FIG. 7 shows an example of a configuration including a front left speaker and a front right speaker. The front left speaker 14L and the front right speaker 14R are arranged in a plane parallel to the coronal plane on the left and right sides of the target listener (UL), and facing the direction of the listener (UL). By arranging the two front speakers in this manner, a centrally localized sound (phantom center) in front of the listener is defined as the FSP sound image SFr. The front left speaker 14L and the front right speaker 14R are examples of left speakers and left speakers of the present disclosure.
[0052] Here, we will explain the characteristics of the phantom center. When speakers are placed on the left and right of the listener (UL) and the same signal (monaural signal) is played back, the sound image is localized directly in front of the listener. Furthermore, even if the listener moves to the left or right between the left and right speakers, the sound image is generally localized directly in front of the listener after the movement.
[0053] Fig. 8 is a diagram showing the overall flow of signal processing in the second embodiment. As shown in Fig. 8, the signal band component SF output from the signal generating unit 10A is split and reproduced in FLSP and FRSP.
[0054] [Modification of the Second Embodiment] Figure 9 shows an example in which two listeners are arranged side by side. When multiple listeners are arranged at different positions in the left-right direction, the FLSP is arranged further left and forward of the leftmost listener (UL1), and the FRSP is arranged further right and forward of the rightmost listener (UL2), in a plane parallel to the coronal plane of each listener. Each of the dotted front speakers is located in a center position relative to the listeners (UL1) and (UL2). In this way, a left speaker (front left speaker 14L (FLSP)) is arranged at the left end of each listener arranged horizontally in the left-right direction, and a right speaker (front right speaker 14R (FRSP)) is arranged at the right end.
[0055] 10 shows an example in which an LSP and an RSP are installed for each listener. In this way, the left speaker 12L and the right speaker 12R can be placed on the left and right sides of each listener's head.
[0056] According to the second embodiment, even if an FSP cannot be installed directly in front of the listener, a sound image of a centrally localized sound can be localized directly in front of the listener. For multiple listeners at different horizontal positions, a sound image of a centrally localized sound can be localized directly in front of each listener. Furthermore, a similar playback space can be provided to each listener at different horizontal positions.
[0057] In the third embodiment, the signal processing unit 10 reproduces signal components (midlow components) between the lower limit frequency of 0 to 200 Hz and the upper limit frequency of 400 to 600 Hz for the signals in the first and second embodiments from the front speakers 14, even if there is no correlation between the left and right. As a result, all midlow components of the left and right signals are reproduced in FSP (or FLSP and FRSP).
[0058] Since most musical instruments include the 200-400 Hz band, even sounds from instruments recorded on only one channel can be perceived as being closer to the front. Furthermore, if low-correlation signals above 600 Hz are reproduced from the front speakers, the reproduced sounds from the nearby LSP and RSP are reduced, which increases the influence of primary reflections, leading to a deterioration in sound quality and a loss of realism. A suitable method is to separate the mid-low components including 200-400 Hz from the L and R signals in advance, mix them with highly correlated signals extracted from other components, and then send them to the FSP (or FLSP and FRSP).
[0059] 11 and 12 show examples of a signal generation unit 10A according to the third embodiment. The example in FIG. 11 shows a case where all mid-low (200 to 400 Hz) signal components are reproduced by the front speakers. Bandpass filters (BPF: L1, L3, R1, R3) and bandstop filters (BSF: L2, R2) are provided on the left and right. This can be applied to both the case where only one FSP is used as in the first embodiment (FIG. 3), and the case where there are two FSPs (FLSP and FRSP) on the left and right, and the left and right signals are the same as in the second embodiment (FIG. 7).
[0060] The example in Figure 12 shows a case where all midlow signal components (200 to 400 Hz) are reproduced by the left and right front speakers. In this case, the front speakers are located on the left and right (FLSP and FRSP), and the left and right midlow components are reproduced by the respective front speakers. The highly correlated signal SF' is calculated at 0.5 and output as SFl and SFr to FLSP and FRSP.
[0061] Fig. 13 is a diagram showing the overall flow of signal processing in the third embodiment. When playing back from the front left and right speakers, as shown in Fig. 13, the SLt and SRt signals are delayed and played back with LSP and RSP, and the SFl and SFr signals are played back with FLSP and FRSP.
[0062] As described above, the signal processing unit 10 of the third embodiment reproduces signal components in at least the lower frequency band of 0 to 200 Hz and the upper frequency band of 400 to 600 Hz from the front speakers, regardless of the left-right correlation between the left speaker signal SLt and the right speaker signal SRt. According to the third embodiment, by reproducing the 200 to 400 Hz band, which is included in most musical instruments and human voices, from the front speakers, it is possible to obtain a sense of front localization in addition to the reproduced sound from the LSP and RSP, even in the case of a source that contains almost no center-localized sound. Furthermore, by setting the upper frequency band to 600 Hz or less, a sense of realism can be maintained.
[0063] [Fourth Embodiment] In the fourth embodiment, in contrast to the first to third embodiments, in the sound reproduction system 1, the opening angle of the LSP and RSP is set to 70 to 110°, and the elevation angle is set to 40° or more, and virtual sound source processing is performed on the L signal component SLt and the R signal component SRt.
[0064] When playing back signals that undergo virtual sound source processing, it is desirable that not only the HRTF but also the interaural difference change little when the head rotates. To achieve this, it is better to have little change in the distance between the left and right speakers (LSP / RSP) and the eardrum, and the greater the elevation angle of the LSP and RSP, the less change there is. Furthermore, by setting the opening angle to around 90° (coronal plane), the change in the distance between the left and right speakers and the eardrum becomes equal when the head rotates, and the change in the interaural difference can also be reduced.
[0065] A supplementary note on virtual reproduction (such as surround) of virtual sound source processing is provided. Virtual sound source processing attempts to reproduce the intended sound field and localization by reproducing the sound pressures above the left and right eardrums. Processing of the sound source signal (Sound Source) includes processing for setting a virtual sound source at a specific position and convolving the HRTF in that direction, and processing for canceling crosstalk components (also using HRTF) to control the sound pressures at the left and right ears when playing back on a 2-channel speaker, for example. Therefore, it is assumed that the HRTF assumed when processing the signal is the same as the HRTF during playback. Potential factors that lead to differences in HRTF include individual differences in HRTF, changes in the position of the head during playback, and changes in the position of the ears due to rotation, and addressing these factors is important in building a practical system.
[0066] Generally, with playback methods using HRTF, (1) virtual playback is effective for sources with movement (such as movies). (2) For sources with no movement (such as music), localization, especially in the front-to-back direction, is often unclear. (3) If the listener's head moves during playback, localization can become unclear or can be localized in an unintended direction. Furthermore, there can be a strong echo that sounds unnatural.
[0067] [Experimental Example] Here, an experimental example of a listening experiment and HRTF analysis will be described with respect to the first embodiment. The listening experiment was a listening experiment to evaluate the sense of forward localization (with FSP) when the head was rotated depending on the arrangement of the LSP and RSP. The elevation angle is the angle between the horizontal plane and the line connecting the center surface of the diaphragm of the LSP (or RSP) and the center of the head. The opening angle is the angle between the center surface of the diaphragm of the LSP (or RSP) and the front direction of the listener's reference position. The speakers were oriented in the radial direction, with the LSP and RSP positioned so that the central axis of the diaphragm faces the entrance of the ear canal of the left ear, and the RSP was positioned so that the central axis of the diaphragm faces the entrance of the ear canal of the right ear. The distance between the FSP and the center of the head was 1.4 m. The distance between the LSP and the RSP and the center of the head was 0.4 m.
[0068] Figure 14 shows an example of an arrangement in an experimental example with an opening angle of 90° and an elevation angle of 30°. Figure 15 shows an example of an arrangement with an opening angle of 150° and an elevation angle of 0°. Table 1 is a comparison table for opening angles and elevation angles. The symbols in the comparison table are ○, △, and × to indicate the state of the sound source as follows: ○: Stable forward localization, △: For some sources, the presence of the sound source is felt in the LSP and RSP directions when the head is rotated (somewhat unstable forward localization), ×: The presence of the sound source is clearly felt in the LSP and RSP directions when the head is rotated (forward localization is impaired).
[0069] (Confirmation of HRTF data under listening conditions) Fig. 16 is a diagram showing changes during head rotation. The changes in HRTF_L during head rotation were compared at each elevation angle, with the changes in HRTF-L when the head was rotated ±20° at an opening angle of 90° to the left (HRTF_L at directional angles of 70, 90, and 110° to the left). Figs. 17 and 18 are graphs showing the changes in HRTF_L at each elevation angle during head rotation. It can be seen that with speaker placement at an opening angle of 90°, the larger the elevation angle, the smaller the change in HRTF during head rotation. It can be seen that the changes are particularly large at 0° and 20°.
[0070] The situation where there is no (or little) change in sound pressure at both ears when the head rotates is similar to when a sound source is located above or when headphones are being used. The sound reproduced from the LSP and RSP in such an arrangement, which results in such a change in characteristics, will be localized above or inside the head. However, this effect is weaker than the effect of perceiving a front sound source (sound reproduced from the FSP) when the head rotates, and the sense of front localization is maintained as long as the content (sound source signal) contains front-localized sound.
[0071] Next, the difference in the change in the interaural difference (HRTF_R-HRTF_L) during head rotation due to the speaker arrangement will be compared.
[0072] 19 to 21 are graphs showing the sound pressure difference and phase difference for each opening angle and elevation angle. (19A) in FIG. 19 shows the sound pressure difference and phase difference for a typical speaker arrangement with an opening angle of 30° and an elevation angle of 0°. In this case, the influence of head rotation and the change in distance between the speaker and the ear are large, resulting in large changes in both sound pressure and phase. (19B) shows the case with an opening angle of 90° (in the coronal plane) and an elevation angle of 0°. In this case, the change in distance is small, so the change in phase is small, but the change in sound pressure characteristics is large due to the influence of the pinna. (20A) in FIG. 20 shows the case with an opening angle of 90° and an elevation angle of 40°, and (20B) shows the case with an opening angle of 77° and an elevation angle of 40°. FIG. 21 shows the case with an opening angle of 103° and an elevation angle of 40°.
[0073] When the elevation angle is 0, the change in interaural difference when the head rotates is large, which can cause unwanted echoes and prevent sound image localization from being reproduced, especially when virtual sound source processing is used.When the elevation angle is 40° or more, and the opening angle is in the range of 77 to 103°, the change in interaural difference when the head rotates is relatively small in both sound pressure level and phase.
[0074] (Application Examples of the Embodiments) Next, application examples of the above-described embodiments will be illustrated. First, an application example to a listening chair or the like will be given. FIG. 22 shows an example of speaker placement when applied to a listening chair or the like. The FSP is installed facing forward using a fixture connected to the armrest, and the signal processing unit 10 can be mounted on PC software, a game console, an audio amplifier, a playback player, or the like. The fixture may be movable for seating, or may be replaced with an FRSP and an FLSP. The LSP and RSP are located above the left and right ears, respectively (elevation angle 70 to 80°, opening angle 90°).
[0075] Figures 23 and 24 show an example of center signal extraction processing in an application example to a listening chair, etc. Note that the reference numerals are omitted because they are the same as those in the above example. When using identical speakers for all speakers, assigning the low-frequency range to the LSP and RSP side, which are located close together, is advantageous in terms of bass perception and input resistance, so the low frequencies are reproduced by the LSP and RSP side. In addition, since the FSP also has a small diameter and high high-frequency reproduction capability, all highly correlated signals above 200 Hz are reproduced by the FSP. Furthermore, when playing back a virtual sound source, as shown in Figure 24, performing this on SLt and SRt allows for stable forward localization to be maintained.
[0076] In addition, even in the case of multi-channel sources such as 5.1ch, the effect can be obtained by applying this processing to at least the front left and right channels. It is also possible to handle the other rear left and right channels by performing virtual playback using LSP and RSP. In this case, the effect is not obtained for these channels, but you can enjoy stable positioning and high sound quality for the entire multi-channel playback.
[0077] Second, an example of application to an in-vehicle system will be described. Figures 25A and 25B show examples of speaker placement when applied to an in-vehicle system. Figures 26 and 27 show an example of center signal extraction processing when applied to an in-vehicle system. In the in-vehicle systems of Figures 25A and 25B, the LSP and RSP are placed overhead (roof lining) near the ears of the listeners in the passenger seat and driver's seat, respectively. The FSPs are also placed on the left and right sides of the dashboard. If placement on the dashboard is not possible, they may be placed below the left and right doors or in a location related to the door bottoms. The signal processing unit 10 is incorporated into the audio function built into the navigation system. Figure 26 shows the center signal extraction processing for the case of (25A). Furthermore, since the LSP and RSP are those with sufficient bass and treble reproduction capabilities (e.g., Patent Document 6), only a minimum signal (highly correlated signals of 200 to 9 kHz) is reproduced from the front SPs.
[0078] Also, (25B) in Figure 25 is an example in which speakers are placed in the rear. Figure 27 shows the center signal extraction process for (25B). In this way, a circuit is provided that inputs the rear L channel and rear R channel separately from the front. In the case of a multi-channel source (e.g., 5ch), rear speakers are installed to obtain stable rear localization for the rear L channel and rear R channel. By combining the rear LSP (14Lb) and rear RSP (14Rb) and performing the same processing as for the main channels (front L channel, front R channel), the same effect can be obtained for the rear channel localization as for the front channels.
[0079] As described above, when the sound reproduction system 1 is installed as an in-vehicle speaker system, the left speaker 12L and the right speaker 12R are placed above the ears of the listeners in each seat. The front speakers 14 are placed in positions related to the dashboard or the bottom of the left and right doors. The signal processing unit 10 is incorporated into the in-vehicle audio system.
[0080] FIG. 28 shows an example of a signal generation unit in which the front speakers are capable of high-quality sound reproduction across the entire frequency range. The signals sent to the front speakers (FSP or FLSP and FRSP) may be signals with high LR correlation, including at least 200 to 9 kHz. Other highly correlated frequency components (components below 200 Hz and above 9 kHz) may be reproduced by the front speakers or by the LSP and RSP. Components below 200 Hz may be distributed between the front SP, LSP, and RSP. The appropriate speaker is selected based on the reproduction capabilities and installation conditions of each speaker. To avoid primary reflections, it may be possible to reproduce signals from the LSP and RSP, which are closer to the front speakers.
[0081] 29 shows an example of center signal extraction processing when the front speakers have high bass reproduction capabilities and the LSP and RSP speakers are set up for mid- and high-range reproduction. In this case, for example, all frequencies below 500 Hz are reproduced by the front speakers.
[0082] In addition, in an example where an LSP and an RSP are placed for each of multiple listeners, if the distance between the listeners is close and the SP playback sound from one listener that is close to another listener affects the other listener, processing may be performed to cancel this.
[0083] The present invention is not limited to the above-described embodiment, and various modifications and applications are possible without departing from the spirit and scope of the present invention.
[0084] The disclosure of Japanese Patent Application No. 2023-197575, filed on November 21, 2023, is incorporated herein by reference in its entirety.
[0085] 1 Sound reproduction system 10 Signal processing unit 12L Left speaker 12R Right speaker 14 Front speaker
Claims
1. An audio reproduction system comprising: a signal processing unit that processes signals on multiple channels including at least a left channel signal L and a right channel signal R; a left speaker and a right speaker that are positioned to the left and right of a listener at a predetermined angle or less and a predetermined angle of elevation or more, respectively; and a front speaker that is positioned further forward from the listener than the left speaker and the right speaker, wherein the signal processing unit generates a front speaker signal SF, a left speaker signal SLt, and a right speaker signal SRt from the left channel signal L and the right channel signal R, respectively, wherein the front speaker signal SF includes at least 200 to 9 kHz of the components highly correlated between L and R, the left speaker signal SLt is mainly composed of components of the left channel signal L that are low correlated with the right channel signal R, and the right speaker signal SRt is mainly composed of components of the right channel signal R that are low correlated with the left channel signal L.
2. The sound reproduction system of claim 1, wherein the left speaker and the right speaker are arranged such that the opening angle is within a predetermined range of 110° or less and not exceeding the listener's line of sight in front of the listener, and the elevation angle is within a predetermined range of 30° or more and not exceeding the apex on the ear side of the listener.
3. The sound reproduction system of claim 1, wherein the front speakers are comprised of a front left speaker and a front right speaker positioned in a plane parallel to the coronal plane to the left and right of the intended listener and facing in the direction of the listener.
4. The sound reproduction system of claim 1, wherein the signal processing unit reproduces signal components in at least a band with a lower frequency of 0 to 200 Hz and an upper frequency of 400 to 600 Hz through the front speakers, regardless of the left-right correlation of the left speaker signal SLt and the right speaker signal SRt.
5. The sound reproduction system of claim 1, wherein the listener is a plurality of listeners, and the left speaker and the right speaker are arranged on the left and right sides of each listener's head, respectively, or the left speaker is arranged at the left end of each listener lined up horizontally in the left-right direction, and the right speaker is arranged at the right end.
6. The sound reproduction system of claim 1, wherein the opening angle of the left speaker and the right speaker is in the range of 70 to 110° and the elevation angle is in a predetermined range of from 40° or more to not exceeding the apex on the ear side of the listener, and the signal processing unit applies virtual sound source processing to the left speaker signal SLt and the right speaker signal SRt to reproduce the signals.
7. The sound reproducing system according to claim 1, when installed as an in-vehicle speaker system, wherein the left speaker and the right speaker are positioned above the ears of the listener in each seat, the front speakers are positioned in positions related to the dashboard or the bottom of the left and right doors, and the signal processing unit is incorporated into the in-vehicle audio.
Citation Information
Patent Citations
Audio system
JP1999113097A
Audio device and audio processing method
JP2010103768A
Acoustic processing device, acoustic processing method, and program
JP2018170563A
3D acoustic system and vehicle
JP2020120290A