Audio signal processing apparatus
The audio signal processing device addresses sound image displacement and unclear localization by adjusting head-related transfer function contributions based on speaker positions, ensuring accurate sound localization and improved stereo output.
Patent Information
- Application Number
- JP2024106767
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-07-02
- Publication Date
- 2026-01-16
AI Technical Summary
Existing surround audio signal processing techniques result in displacement of the sound image in front of the frontal plane or unclear localization and presence when downmixing multi-channel signals to 2-channel stereo, leading to unsatisfactory sound image impressions.
An audio signal processing device that combines surround audio signals with head-related transfer functions based on speaker positions relative to the listener, adjusting ratios to minimize the contribution of these functions for speakers in front of the frontal plane, and generating downmix signals for stereo output, using a pseudo surround audio signal generator and downmix unit to maintain accurate sound localization.
Prevents displacement of the sound image in front of the frontal plane and maintains clear localization and presence by reducing the impact of head-related transfer functions for speakers in front, resulting in improved sound image quality for stereo headphones.
Smart Images

Figure 2026007181000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a technique for processing surround audio signals. [Background technology]
[0002] Known surround audio signal processing techniques include upmixing a 2-channel stereo signal to generate 5-channel surround audio signals of C, L, R, LS, and RS (see, for example, Patent Documents 1 and 2). Here, C is the audio signal for the speaker in the center in front of the listener, L is the audio signal for the speaker in front of the listener to the left, R is the audio signal for the speaker in front of the listener to the right, LS is the audio signal for the speaker to the left or left rear of the listener, and RS is the audio signal for the speaker to the right or right rear of the listener. Furthermore, as a technique for processing surround audio signals, a technique for downmixing a multi-channel surround audio signal such as 5.1ch to a 2ch stereo signal is known (for example, Patent Document 3). This technology generates an audio signal by convolving the audio signal of each channel with a head-related transfer function, which is the transfer function of sound from the corresponding speaker to each of the listener's left and right ears, and then synthesizes the audio signals of each channel convolved with the head-related transfer function for each ear, downmixing them into a 2-channel audio signal for stereo headphones. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2013-126116 [Patent Document 2] Japanese Patent Application Laid-Open No. 2010-103768 [Patent Document 3] Japanese Patent Publication No. 2023-92962 Summary of the Invention [Problem to be solved by the invention]
[0004] As described above, when a stereo 2-channel audio signal for stereo headphones is generated by downmixing the audio signals of each channel obtained by convolving a head-related transfer function into a multi-channel surround audio signal, the sound image in front of the frontal plane (coronal plane) that divides the body of the listener into front and back may be displaced upward, or the sense of positioning and presence of the sound image may become unclear, resulting in an unsatisfactory sound image impression.
[0005] Therefore, an object of the present invention is to downmix a multi-channel surround audio signal into a 2-channel audio signal so as to obtain a good sound image. [Means for solving the problem]
[0006] In order to achieve the above object, the present invention provides an audio signal processing device that converts a surround audio signal including n-channel audio signals for output from speakers installed at corresponding positions, each of which corresponds to n (n≧4) positions of speakers defined relatively to the assumed position and orientation of a listener, into an output stereo audio signal, which is a stereo audio signal consisting of two-channel audio signals output from the audio signal output device, the audio signal processing device comprising: a signal processing means that, for each of the n channels of the surround audio signal, combines the audio signal of that channel with an audio signal obtained by convolving the audio signal of that channel with a head-related transfer function from the speaker installed at the position corresponding to that channel to the listener, at a ratio according to the position corresponding to that channel, and outputs the combined signal as a downmix audio signal; and a downmixing means that generates a two-channel audio signal of the output stereo audio signal by combining the downmix audio signals generated by the downmix audio signal generating means.
[0007] In this audio signal processing device, the n placement positions may include a plurality of placement positions that are forward of the frontal plane of the listener and that have different magnitudes of directional difference from a front direction of the listener. In this case, for placement positions that are forward of the frontal plane of the listener, the ratio according to the placement position may be set so that the closer the direction of the placement position is to the front direction of the listener, the greater the ratio of the audio signal of the channel corresponding to the placement position and the smaller the ratio of the audio signal obtained by convolving the audio signal of the channel corresponding to the placement position with the head-related transfer function.
[0008] In this audio signal processing device, the n placement positions may include placement positions that are in front of the frontal plane of the listener and placement positions that are behind the frontal plane of the listener. In this case, the ratios according to the placement positions may be set so that a placement position that is in front of the frontal plane of the listener has a larger ratio of the audio signal of the channel corresponding to that placement position and a smaller ratio of the audio signal obtained by convolving the head-related transfer function with the audio signal of the channel corresponding to that placement position than a placement position that is behind the frontal plane of the listener.
[0009] Furthermore, if the n placement positions include a placement position that is behind the frontal plane of the listener, for a placement position that is behind the frontal plane of the listener, the ratio according to the placement position may be set to 0 for the ratio of the audio signal of the channel corresponding to that placement position, and 1 for the ratio of the audio signal obtained by convolving the audio signal of the channel corresponding to that placement position with the head-related transfer function.
[0010] In the audio signal processing device described above, the downmix audio signals output by the signal processing means for the n channels of the surround audio signal may include a left-channel downmix audio signal and a right-channel downmix audio signal. In this case, the signal processing means may combine, for the n channels of the surround audio signal, an audio signal of that channel and an audio signal obtained by convolving the audio signal of that channel with a head-related transfer function from a speaker installed at a placement position corresponding to that channel to the left ear of the listener, at a ratio according to the placement position corresponding to that channel, to output the result as the left-channel downmix audio signal, and may also combine, for the n channels of the surround audio signal, an audio signal of that channel and an audio signal obtained by convolving the audio signal of that channel with a head-related transfer function from a speaker installed at a placement position corresponding to that channel to the right ear of the listener, at a ratio according to the placement position corresponding to that channel, to output the result as the right-channel downmix audio signal. Furthermore, the downmixing means may generate a left channel audio signal of the output stereo audio signal by synthesis including each left channel downmix audio signal generated by the downmix audio signal generating means, and may generate a right channel audio signal of the output stereo audio signal by synthesis including each right channel downmix audio signal generated by the downmix audio signal generating means.
[0011] In this case, the surround audio signal may include an audio signal for bass reproduction of a low-frequency effect channel, and the downmixing means may combine each left-channel downmix audio signal generated by the downmix audio signal generating means with the audio signal of the low-frequency effect channel to generate a left-channel audio signal of the output stereo audio signal, and may combine each right-channel downmix audio signal generated by the downmix audio signal generating means with the audio signal of the low-frequency effect channel to generate a right-channel audio signal of the output stereo audio signal.
[0012] The above audio signal processing device may also be provided with a surround audio signal generating means for generating the surround audio signal from an input stereo audio signal, which is a stereo audio signal consisting of two-channel audio signals input to the audio signal output device.
[0013] According to such an audio signal processing device, for audio signals to be output from speakers installed at positions assumed to be in front of the listener from the frontal plane, the contribution of the head-related transfer function can be reduced compared to the audio signals to be output from such speakers, thereby generating downmix audio signals that are the source of downmixing of output stereo audio signals. This makes it possible to prevent the localization position of a sound image in front of the frontal plane (coronal plane) from being displaced from its original position due to application of the head-related transfer function. This also prevents the sense of localization and presence of the sound image from becoming unclear. [Effects of the Invention]
[0014] As described above, according to the present invention, it is possible to downmix a multi-channel surround audio signal into a 2-channel audio signal so as to obtain a good sound image. [Brief explanation of the drawings]
[0015] [Figure 1] 1 is a diagram illustrating a configuration of an audio signal processing device according to an embodiment of the present invention. [Figure 2] FIG. 1 is a diagram showing a speaker arrangement assumed in an embodiment of the present invention. [Figure 3] 1 is a diagram showing the configuration of a component separation unit that can be used in a pseudo surround audio signal generation unit according to an embodiment of the present invention. [Figure 4] FIG. 2 is a diagram illustrating a configuration of a downmix signal generation unit according to the embodiment of the present invention. [Figure 5] 10A and 10B are diagrams illustrating examples of setting head-related transfer function contribution ratios according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0016] Hereinafter, an embodiment of the present invention will be described. FIG. 1a shows the configuration of an audio signal processing device according to this embodiment. As shown in the figure, the audio signal processing device includes a pseudo surround audio signal generator 1 that upmixes an input 2-channel stereo signal SigIN of Lin and Rin to generate a 7.1.4-channel surround audio signal SigA consisting of audio signals of C, L, R, LS, RS, LB, RB, LFE, LH, RH, LBH, and RBH, a signal processor that processes the surround audio signal SigA generated by the pseudo surround audio signal generator 1 to generate C(L), C(R), L(L), L(R), R(L), R(R), LS(L), The system is equipped with a signal processing unit 2 that outputs the signals LS(R), RS(L), RS(R), LB(L), LB(R), RB(L), RB(R), LH(L), LH(R), RH(L), RH(R), LBH(L), LBH(R), RBH(L), RBH(R), and LFE as a downmix signal SigB, a downmix unit 3 that combines the downmix signals SigB and outputs a 2-channel stereo signal SigOUT of Lout and Rout, a control unit 4, and a user interface 5 such as a switch or touch panel that accepts user operations. Here, the downmix unit 3 applies appropriate gains and delays to C(L), L(L), R(L), LS(L), RS(L), LB(L), RB(L), LH(L), RH(L), LBH(L), RBH(L), and LFE, respectively, and combines them to generate Lout, and applies appropriate gains and delays to C(R), L(R), R(R), LS(R), RS(R), LB(R), RB(R), LH(R), RH(R), LBH(R), RBH(R), and LFE, respectively, and combines them to generate Rout.
[0017] The stereo signal SigOUT output by the downmix unit 3 is a 2-channel stereo signal for output from an audio output device equipped with a pair of left and right speakers, such as stereo headphones or stereo earphones, where Lout is the audio signal for the L (left) channel and Rout is the audio signal for the R (right) channel.
[0018] Each audio signal of the surround audio signal SigA is a surround audio signal that is assumed to be output using speakers arranged at 13 positions, examples of which are shown in FIGS. 2a and 2b. However, the audio signal LFE is an audio signal of a low-frequency effect channel, and since the sound image of the low-frequency effect channel's bass audio signal is not localized, the placement position of the bass subwoofer, which is the speaker corresponding to the signal LFE, is arbitrary. Next, in the illustrated example, as shown in Figure 2a, the direction directly in front of the listener as seen from the listener's listening point LP is taken as a horizontal angle of 0°, and the horizontal angle is measured counterclockwise when viewed vertically downward.The speaker corresponding to C is placed at a horizontal angle of 0°, the speaker corresponding to L at a horizontal angle of 30°, the speaker corresponding to R at a horizontal angle of 330°, the speaker corresponding to LS at a horizontal angle of 90°, the speaker corresponding to RS at a horizontal angle of 270°, the speaker corresponding to LB at a horizontal angle of 150°, the speaker corresponding to RB at a horizontal angle of 210°, the speaker corresponding to LH at a horizontal angle of 45°, the speaker corresponding to RH at a horizontal angle of 315°, the speaker corresponding to LBH at a horizontal angle of 135°, and the speaker corresponding to RBH at a horizontal angle of 225°.
[0019] In addition, in the illustrated example, as shown in Figure 2b, the elevation angles of the placement positions of the corner speakers, the speakers corresponding to C, L, R, LS, RS, LB, and RB are placed at a position with an elevation angle of 0° when viewed from the listener's listening point LP, and the speakers corresponding to LH, RH, LBH, and RBH are placed at a position with an elevation angle of 45° when viewed from the listener's listening point LP.
[0020] Now, the specific configuration of the pseudo surround audio signal generating unit 1 for generating the surround audio signal SigA can vary depending on the application, etc., but basically, each audio signal of the surround audio signal SigA can be generated by separating a component correlated with Rin of Lin, a component uncorrelated with Rin of Lin, a component correlated with Lin of Rin, and a component uncorrelated with Rin of Rin from Lin and combining each separated component or each component with Lin and Rin.
[0021] For example, the LFE of the surround audio signal SigA generated by the pseudo surround audio signal generating unit 1 can be generated as the sum of the low frequency components of Lin and Rin. Furthermore, LS is the component of Lin that is uncorrelated with Rin, RS is the component of Rin that is uncorrelated with Lin, C is the sum of the components of Lin and Rin that are correlated with each other, L can be generated as the sum of LS and C, and R can be generated as the sum of RS and C. In addition, LB can be generated by combining delayed Lin with L, LS, and C in an appropriate mixing ratio, and RB can be generated by combining delayed Rin with R, RS, and C in an appropriate mixing ratio. Also, LH can be generated by applying an appropriate delay and reverb effect to L, RH can be generated by applying an appropriate delay and reverb effect to R, LBH can be generated by applying an appropriate delay and reverb effect to LB, and RBH can be generated by applying an appropriate delay and reverb effect to RB.
[0022] The separation of correlated and uncorrelated components can be performed using a component separation unit having the configuration shown in FIG. The component separator shown in FIG. 3 separates, from signals A and B, a component CA of signal B that is correlated with signal A and a component SB of signal B that is uncorrelated with signal A. As shown in the figure, the component separation unit includes a variable filter 101, an update unit 102 that updates the transfer function (filter coefficient) W of the variable filter 101 using an adaptive algorithm such as an LMS algorithm, and an adder 103, and the variable filter 101, update unit 102, and adder 103 form an adaptive filter.
[0023] Variable filter 101 receives signal A as input, and adder 103 subtracts the output of variable filter 101 from signal B and outputs the result. Update unit 102 executes an adaptive algorithm using the output of adder 103 as an error, and updates transfer function W of variable filter 101 so that the power of the error is minimized.
[0024] The power of the output of adder 103 is minimum when the output of variable filter 101 matches component CA of signal B that is correlated with signal A, and at this time, the output of adder 103, which is obtained by subtracting the output of variable filter 101 from signal B, represents component SB of signal B that is uncorrelated with signal A.
[0025] Therefore, a component CA of signal B that is correlated with signal A can be separated as the output of variable filter 101, and a component SB of signal B that is uncorrelated with signal A can be separated as the output of adder 103. Similarly, by swapping signal A of signal B, a component CB of signal A that is correlated with signal B and a component SA of signal A that is uncorrelated with signal B can be separated.
[0026] Returning to FIG. 1, the signal processing unit 2 includes downmix signal generation units 21 provided corresponding to each of the audio signals C, L, R, LS, RS, LB, RB, LH, RH, LBH, and RBH of the surround audio signal SigA, to which the corresponding audio signal is input, and a delay unit 22 to which the audio signal LFE is input.
[0027] The downmix signal generation unit 21 receives an audio signal X (X is any one of C, L, R, LS, RS, LB, RB, LH, RH, LBH, and RBH), processes the audio signal X, generates audio signals X(L) and X(R) of the downmix signal SigB, and outputs the generated audio signals to the downmix unit 3.
[0028] The delay unit 22 delays the audio signal LFE by the amount of delay caused by the processing in the downmix signal generation unit 21 and outputs the delayed audio signal LFE to the downmix unit 3 . Here, each downmix signal generator 21 in the signal processor 2 will be described. Since each downmix signal generator 21 in the signal processing unit 2 has the same configuration, the configuration of the downmix signal generator 21 corresponding to the audio signal LH will be described below as a representative example. FIG. 4 shows the configuration of the downmix signal generator 21 corresponding to the audio signal LH, and the downmix signal generator 21 includes an Lch downmix signal generator 211 and an Rch downmix signal generator 212. The Lch downmix signal generation unit 211 and the Rch downmix signal generation unit 212 each include a head-related transfer function filter 2101 to which an audio signal LH is input, a K-times multiplier 2102 that multiplies the output of the head-related transfer function filter 2101 by K, a delay unit 2103 that delays the audio signal LH to align the delay with the output of the head-related transfer function filter 2101, a (1-K)-times multiplier 2104 that multiplies the audio signal LH delayed by the delay unit 2103 by (1-K), and an adder 2105 that adds the output of the K-times multiplier 2102 and the output of the (1-K)-times multiplier 2104 and outputs the result.
[0029] The head-related transfer function filter 2101 of the Lch downmix signal generation unit 211 has a head-related transfer function set as a transfer function of sound from the speaker corresponding to the audio signal LH to the left ear of the listener, and convolves the set head-related transfer function with the input audio signal LH and outputs the result. Also, the adder 2105 of the Lch downmix signal generation unit 211 outputs the downmix signal SigB LH(L) to the downmix unit 3.
[0030] Furthermore, the head-related transfer function filter 2101 of the Rch downmix signal generation unit 212 has the transfer function of the sound from the speaker corresponding to the audio signal LH to the right ear of the listener set as the head-related transfer function, and convolves the set head-related transfer function with the input audio signal LH and outputs the result. Furthermore, the adder 2105 of the Rch downmix signal generation unit 212 outputs the output to the downmix unit 3 as LH(R) of the downmix signal SigB.
[0031] In the same downmix signal generator 21, the multipliers K of the K-times multipliers 2102 of the Lch downmix signal generator 211 and the Rch downmix signal generator 212 are equal, and the multipliers (1-K) of the (1-K)-times multipliers 2104 of the Lch downmix signal generator 211 and the Rch downmix signal generator 212 are equal. These multipliers are values set by the controller 4 and are in the range 0≦K≦1.
[0032] The outputs of the adders 2105 of the Lch downmix signal generator 211 and the Rch downmix signal generator 212 are audio signals obtained by convolving a head-related transfer function into the audio signal LH and mixing the audio signal LH at a ratio of K:(1-K). Therefore, K represents the contribution rate of the head-related transfer function in LH(L) and LH(R) output to the downmixer 3. Therefore, in the following description, K will be referred to as the head-related transfer function contribution rate K.
[0033] The downmix signal generators 21 corresponding to the audio signals C, L, R, LS, RS, LB, RB, RH, LBH, and RBH also have the same configuration as the downmix signal generator 21 corresponding to LH, and the description thereof will be obtained by replacing LH in the configuration of the downmix signal generator 21 corresponding to LH in FIG. 4 with LH in the above description of the downmix signal generator 21 corresponding to LH, with the corresponding audio signals.
[0034] However, the head-related transfer function contribution ratios K that the control unit 4 sets to each downmix signal generation unit 21 are not the same, but are set to values that correspond to the downmix signal generation unit 21. The control unit 4 sets the head-related transfer function contribution ratio K to be set in each downmix signal generation unit 21 according to the horizontal angle shown in FIG. 2a of the speaker to which the audio signal input for downmix signal generation corresponds. Figure 5a shows the relationship between the horizontal angle of the speaker and the set head-related transfer function contribution rate K. The head-related transfer function contribution rate K is set so that at horizontal angles (0°-90°, 270°-360°) that are in front of the listener from the frontal plane (a plane whose normal direction is a horizontal angle of 0° and an elevation angle of 0°), the head-related transfer function contribution rate K approaches 0 as the horizontal angle approaches the front of the listener (0° / 360°), and so that the head-related transfer function contribution rate K is 1 at horizontal angles that are not in front of the listener from the frontal plane.
[0035] FIG. 5b shows the head-related transfer function contribution ratios K that are set in the downmix signal generator 21 corresponding to each of the audio signals C, L, R, LS, RS, LB, RB, LH, RH, LBH, and RBH of the surround audio signal SigA, in accordance with the relationship in FIG. 5a. The head-related transfer function contribution ratio K that is set in the downmix signal generator 21 corresponding to C is 0, the head-related transfer function contribution ratios K that are set in the downmix signal generator 21 corresponding to L and R are 0.33, the head-related transfer function contribution ratio K that is set in the downmix signal generator 21 corresponding to LH and RH is 0.5, and the head-related transfer function contribution ratios K that are set in the downmix signal generator 21 corresponding to the remaining LS, RS, LB, RB, LBH, and RBH are all 1.
[0036] As a result, the audio signals of the downmix signal SigB that are input to the downmix unit 3 and combined into the 2-channel stereo signals Lout and Rout, excluding LFE, are obtained by convolving the audio signal for that speaker of the surround audio signal SigA with a head-related transfer function for the audio signal that corresponds to that speaker, while the audio signals for that speaker that corresponds to a speaker that is in front of the listener, excluding LFE, are obtained by mixing the audio signal for that speaker of the surround audio signal SigA with an audio signal obtained by convolving the audio signal with a head-related transfer function, and the ratio of the mixed audio signal convolved with a head-related transfer function decreases for audio signals that correspond to speakers that are closer to the front of the listener.
[0037] This makes it possible to prevent the localization position of the sound image in front of the frontal plane from being displaced upward from its original position due to the application of the head-related transfer function, and also prevents the sense of localization and presence of the sound image from becoming unclear. Here, the control unit 4 also has a function of individually changing the head-related transfer function contribution ratio K set in each downmix signal generation unit 21 in response to a user operation received via the user interface 5. By making the head-related transfer function contribution ratio K changeable in response to a user operation in this way, each user can adjust the characteristics of sound image localization in accordance with their own preferences and sensibilities.
[0038] The embodiments of the present invention have been described above. In the above embodiment, when the horizontal angle of the speakers corresponding to the audio signals input to the downmix signal generation unit 21 is a horizontal angle that is in front of the listener from the frontal plane, the head-related transfer function contribution rate K approaches 0 as the horizontal angle approaches the front of the listener (0° / 360°). However, this may be such that the same value is used as the head-related transfer function contribution rate K in all downmix signal generation units 21 to which audio signals corresponding to speakers at horizontal angles that are in front of the listener from the frontal plane are input. However, this same value is a value that is smaller than the head-related transfer function contribution rate K of downmix signal generation units 21 to which audio signals corresponding to speakers that are not at a horizontal angle that is in front of the listener from the frontal plane are input.
[0039] Furthermore, in the above embodiment, the pseudo surround audio signal generating unit 1 generates a 7.1.4 channel surround audio signal, but L may be any number equal to or greater than 4, M may be any number equal to or greater than 2, and the pseudo surround audio signal generating unit 1 may generate a surround audio signal SigA of L.0.0 channel, L.1.0 channel, L.0.M channel, or L.1.M channel, and the signal processing unit 2 and downmixing unit 3 may perform the above-mentioned processing on each audio signal of the surround audio signal SigA in a form adapted to these audio signals.
[0040] Alternatively, without providing the pseudo surround audio signal generating unit 1, each of the surround audio signals SigA of the L.0.0 channel, L.1.0 channel, L.0.M channel, and L.1.M channel may be input to the signal processing unit 2 as an audio source, and the signal processing unit 2 and the downmixing unit 3 may perform the above-mentioned processing in a form adapted to these audio signals. [Explanation of symbols]
[0041] 1...pseudo surround audio signal generation unit, 2...signal processing unit, 3...downmix unit, 4...control unit, 5...user interface, 21...downmix signal generation unit, 22...delay unit, 101...variable filter, 102...update unit, 103...adder, 211...Lch downmix signal generation unit, 212...Rch downmix signal generation unit, 2101...head-related transfer function filter, 2102...K-times multiplier, 2103...delay unit, 2104...(1-K)-times multiplier, 2105...adder.
Claims
1. An audio signal processing device that converts a surround audio signal including n-channel audio signals for output from speakers installed at corresponding placement positions, the n-channel audio signals corresponding to placement positions of n speakers (n≧4) defined relative to an assumed position and orientation of a listener, into an output stereo audio signal that is a stereo audio signal consisting of two-channel audio signals output from the audio signal output device, a signal processing means for synthesizing, for n channels of the surround audio signal, an audio signal of the channel and an audio signal obtained by convolving the audio signal of the channel with a head-related transfer function from a speaker installed at a placement position corresponding to the channel to the listener, at a ratio according to the placement position corresponding to the channel, and outputting the result as a downmix audio signal; downmixing means for generating two-channel audio signals of the output stereo audio signal by combining the downmix audio signals generated by the downmix audio signal generating means.
2. 2. The audio signal processing device according to claim 1, the n placement positions include a plurality of placement positions in front of the frontal plane of the listener, the placement positions having different magnitudes of difference in direction from a front direction of the listener, and a ratio according to the placement position, which is set so that, for a placement position that is forward of the frontal plane of the listener, the closer the placement position is to a front direction of the listener, the greater the ratio of the audio signal of the channel corresponding to the placement position and the smaller the ratio of the audio signal obtained by convolving the audio signal of the channel corresponding to the placement position with the head-related transfer function.
3. 2. The audio signal processing device according to claim 1, the n placement positions include a placement position that is in front of the frontal plane of the listener and a placement position that is behind the frontal plane of the listener, the ratio according to the placement position is set so that a placement position that is forward of the frontal plane of the listener has a larger ratio of the audio signal of the channel corresponding to that placement position, and a smaller ratio of the audio signal obtained by convolving the head-related transfer function with the audio signal of the channel corresponding to that placement position, than a placement position that is rearward of the frontal plane of the listener.
4. 3. An audio signal processing apparatus according to claim 2, the n placement positions include a placement position that is rearward of the frontal plane of the listener, For a placement position that is behind the frontal plane of the listener, the ratio according to the placement position is 0 for the audio signal of the channel corresponding to the placement position, and 1 for the audio signal obtained by convolving the audio signal of the channel corresponding to the placement position with the head-related transfer function.
5. 5. An audio signal processing apparatus according to claim 1, 2, 3 or 4, the downmix audio signals output by the signal processing means for the n channels of the surround audio signal include a left-channel downmix audio signal and a right-channel downmix audio signal, the signal processing means combines, for n channels of the surround audio signal, an audio signal of the channel and an audio signal obtained by convolving the audio signal of the channel with a head-related transfer function from a speaker installed at a placement position corresponding to the channel to the left ear of the listener, at a ratio according to the placement position corresponding to the channel, and outputs the result as the left channel downmix audio signal; and combines, for n channels of the surround audio signal, an audio signal of the channel and an audio signal obtained by convolving the audio signal of the channel with a head-related transfer function from a speaker installed at a placement position corresponding to the channel to the right ear of the listener, at a ratio according to the placement position corresponding to the channel, and outputs the result as the right channel downmix audio signal; the downmixing means generates a left-channel audio signal of the output stereo audio signal by synthesis including each left-channel downmix audio signal generated by the downmix audio signal generating means, and generates a right-channel audio signal of the output stereo audio signal by synthesis including each right-channel downmix audio signal generated by the downmix audio signal generating means.
6. 6. An audio signal processing apparatus according to claim 5, the surround audio signal includes an audio signal for bass reproduction in a low frequency effect channel; the downmixing means combines each left-channel downmix audio signal generated by the downmix audio signal generating means with the audio signal of the low-frequency effect channel to generate a left-channel audio signal of the output stereo audio signal, and combines each right-channel downmix audio signal generated by the downmix audio signal generating means with the audio signal of the low-frequency effect channel to generate a right-channel audio signal of the output stereo audio signal.
7. 5. An audio signal processing apparatus according to claim 1, 2, 3 or 4, an audio signal processing device comprising: a surround audio signal generating means for generating the surround audio signal from an input stereo audio signal, which is a stereo audio signal consisting of two-channel audio signals input to the audio signal output device;
Citation Information
Patent Citations
Audio device and audio processing method
JP2010103768A
Audio device
JP2013126116A
Audio signal output method, audio signal output device, and audio system
JP2023092962A