Audio signal processing unit

The audio signal processing device enhances sound localization for narrow frequency bands by separating and frequency-multiplying target sounds, addressing the inadequacies of existing techniques.

JP2026059986APending Publication Date: 2026-04-08ALPS ALPINE CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2024-09-27
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

Existing audio signal processing techniques fail to provide a sufficient sense of sound localization for sounds with a narrow frequency band.

Method used

An audio signal processing device with signal processing units for each channel that separates target sounds, generates localization enhancement signals by multiplying their frequencies, and synthesizes these with the original signals to expand the frequency band and enhance localization.

Benefits of technology

The device achieves a better sense of localization for sounds with narrow frequency bands by adding harmonics and subharmonics, improving perception of such sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026059986000001_ABST
    Figure 2026059986000001_ABST
Patent Text Reader

Abstract

We provide an "audio signal processing device" that achieves good localization for sounds with a narrow frequency band. [Solution] The system has a signal processing unit corresponding to each channel of a multi-channel audio signal. Each signal processing unit includes a target sound separation unit 411 that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a localization enhancement sound signal by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by k (where k is an integer of 2 or more); and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an audio signal processing apparatus.

Background Art

[0002] There is known an audio signal processing apparatus that provides a sense of sound localization at a sound image localization target position by convolution of the acoustic transfer characteristics (head transfer function) from the sound image localization target position to each of the listener's left and right ears with an audio signal (for example, Patent Document 1).

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] Humans have a low ability to perceive the direction of sounds with a narrow frequency band, and there has been a problem that the above-described technique of convolution of the acoustic transfer characteristics (head transfer function) with an audio signal cannot provide a sufficient sense of localization for sounds with a narrow frequency band. Therefore, an object of the present invention is to provide an audio signal processing apparatus that can realize a better sense of localization for sounds with a narrow frequency band.

Means for Solving the Problems

[0005] To achieve the above objectives, the present invention provides an audio signal processing device having a signal processing unit provided for each channel of a plurality of channels of audio signals, each signal processing unit having: a target sound separation unit that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a localization enhancement sound signal by multiplying the frequency of the target sound signal separated by the target sound separation unit by k (where k is an integer of 2 or more); and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0006] Furthermore, the present invention provides an audio signal processing device having a signal processing unit provided for each channel of a plurality of audio signals, each signal processing unit comprising: a target sound separation unit that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a localization enhancement sound signal by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / L (where L is an integer of 2 or more); and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0007] Furthermore, the present invention provides an audio signal processing device having a signal processing unit provided for each channel of a plurality of audio signals, each signal processing unit having: a target sound separation unit that separates a target sound signal, which is a sound to be targeted for localization enhancement, from a source sound signal, which is an audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a first signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by k (where k is an integer of 2 or more), and a second signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / L (where L is an integer of 2 or more), and also synthesizes the first signal and the second signal to generate a localization enhancement sound signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0008] Furthermore, the present invention provides an audio signal processing device having a signal processing unit provided for each channel of a plurality of audio signals, each signal processing unit comprising: a target sound separation unit that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a signal for each integer i from 2 to n (where n>2) by multiplying the frequency of the target sound signal separated by the target sound separation unit by i, and synthesizes the generated signals to generate a localization enhancement sound signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0009] Furthermore, the present invention provides an audio signal processing device having a signal processing unit provided for each channel of a plurality of audio signals, each signal processing unit comprising: a target sound separation unit that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that, for each j, takes each integer from 2 to m (where m>2) as i, generates a signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / j, and synthesizes the generated signals to generate a localization enhancement sound signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0010] Furthermore, the present invention provides an audio signal processing device having a signal processing unit provided for each channel of a plurality of audio signals, each signal processing unit having: a target sound separation unit that separates a target sound signal, which is a sound to be targeted for localization enhancement, from a source sound signal, which is an audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a first signal for each i, where i is an integer from 2 to n (where n>2), and for each j, the frequency of the target sound signal separated by the target sound separation unit is multiplied by i; and a localization enhancement sound generation unit that generates a localization enhancement sound signal by combining the generated first signal and each second signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs it.

[0011] Herein, the audio signal processing device described above may be provided with time-stretching means in each of the signal processing units for extending the duration of the localization enhancement sound signal synthesized by the synthesis unit when the duration of the target sound signal separated by the target sound separation unit is shorter than a predetermined duration.

[0012] Furthermore, the above audio signal processing device may be provided with a direction estimation means for estimating the direction in which the target sound should be localized, and each of the signal processing units may be provided with a localization correction unit that applies a head-level transfer function to the localization enhancement sound signal synthesized by the synthesis unit, which is compatible with the direction estimated by the direction estimation means, and replaces the target sound signal in the source sound signal synthesized by the synthesis unit with a signal to which a head-level transfer function compatible with the direction estimated by the direction estimation means has been applied to the target sound signal.

[0013] In this case, the direction estimation means may estimate the direction in which the target sound should be localized based on the relationship between the target sound signals separated by the target sound separation unit of each signal processing unit. Alternatively, the direction estimation means may estimate the direction in which the target sound should be localized based on information regarding the position of the sound source of the target sound output from the device that outputs the audio signals of the multiple channels. With the audio signal processing device described above, for target sounds that have poor localization due to a narrow frequency band, the frequency band of sounds associated with the target sound can be expanded by adding harmonics and subharmonics, thereby providing a better sense of localization. [Effects of the Invention]

[0014] As described above, the present invention provides an audio signal processing device that can achieve a better sense of localization for sounds with a narrow frequency band. [Brief explanation of the drawing]

[0015] [Figure 1] This figure shows the configuration of an AV system according to an embodiment of the present invention. [Figure 2] This figure shows the configuration of an audio signal processing device according to an embodiment of the present invention. [Figure 3] This diagram shows the configuration of the spatial awareness enhancement component generating unit according to an embodiment of the present invention. [Figure 4] This figure shows another example of the configuration of the spatial awareness enhancement component generation unit according to the present invention. [Figure 5] This figure shows another example of the configuration of the spatial awareness enhancement component generation unit according to the present invention. [Figure 6] This figure shows the configuration of the target sound separation unit according to an embodiment of the present invention. [Figure 7] This figure shows another example configuration of an audio signal processing device according to an embodiment of the invention. [Figure 8] This figure shows another example configuration of an audio signal processing device according to an embodiment of the invention. [Figure 9] This figure shows another example configuration of an audio signal processing device according to an embodiment of the invention. [Modes for carrying out the invention]

[0016] Hereinafter, embodiments of the present invention will be described. FIG. 1 shows the configuration of the AV system according to this embodiment. As shown in the figure, the AV system includes an AV device 1 that outputs video signals and audio signals of the same AV content, an input device 2 that receives operations on the AV device 1, a display 3 that displays the video signal output by the AV device 1, and an audio signal processing device 4 that performs signal processing to enhance the sound localization feeling, which is a target for enhancing the sound localization feeling for the audio signal output by the AV device 1, and outputs the processed signal, an amplifier 5 that amplifies the audio signal output by the sound localization enhancement signal, and an acoustic output device 6 that emits the sound represented by the audio signal output by the amplifier 5.

[0017] The AV device 1 is, for example, a PC or a game machine, the input device 2 is, for example, a keyboard or a game pad, and the acoustic output device 6 is, for example, a speaker or headphones. Next, FIG. 2 shows the configuration of the audio signal processing device 4. Here, as shown in the figure, the AV device 1 outputs two-channel stereo audio signals of the L channel and the R channel as source sounds. Also, the audio signal processing device 4 outputs two-channel stereo audio signals of the L channel and the R channel as output sounds to the amplifier 5.

[0018] The audio signal processing device 4 includes an L-channel processing unit 41 that performs signal processing on the audio signal of the L channel of the source sound and outputs it to the amplifier 5 as the audio signal of the L channel of the output sound, and an R-channel processing unit 42 that performs signal processing on the audio signal of the R channel of the source sound and outputs it to the amplifier 5 as the audio signal of the R channel of the output sound.

[0019] The L-channel processing unit 41 and the R-channel processing unit 42 have the same configuration. As shown in Figure 2 for the L-channel processing unit 41, the L-channel is the channel corresponding to the L-channel processing unit 41 and the R-channel is the channel corresponding to the R-channel processing unit 42. The L-channel processing unit 411 separates the target sound, which is the sound targeted for localization enhancement, from the audio signal of the corresponding channel of the source sound. The R-channel processing unit 42 generates a localization enhancement component from the target sound input from the target sound of the target sound, which enhances the localization of the target sound. The R-channel processing unit 413 adds the localization enhancement component generated by the localization enhancement component 412 to the audio signal of the corresponding channel of the source sound and outputs it as the audio signal of the corresponding channel of the output sound.

[0020] Next, Figure 3 shows the configuration of the localization enhancement component generation unit 412 of the audio signal processing device 4. The configuration shown in Figure 3 is the configuration of the localization enhancement component generation unit 412 when the target sound is concentrated in the low-frequency range (for example, the range below 100 Hz), that is, when the frequency range of the target sound is roughly within the low-frequency range. In this context, an example of a sound concentrated in this low-frequency range is the sound of enemy footsteps when playing FPS (First Person Shooter) or TPS (Third Person Shooter) games on AV equipment 1. It is preferable to enhance the sense of localization of such footsteps by treating them as target sounds so that the position and movement of the enemy can be grasped by hearing.

[0021] As shown in the figure, the localization enhancement component generation unit 412 is equipped with an FFT4121 that converts the input target sound into a frequency domain signal using a Fast Fourier Transform. Furthermore, the system is equipped with N-1 i-th frequency component generators 4122, where i is an integer from 2 to N (N≧2). The i-th frequency component generators 4122 convert each frequency component of the target sound from the output of the FFT 4121 into a frequency component that is i times that frequency, with a predetermined gain. Therefore, the i-th frequency component generators 4122 generate a frequency-domain signal of the (i-1)th harmonic for each frequency component of the target sound.

[0022] Furthermore, the system includes a frequency domain summer 4123 that generates a frequency domain signal by adding together the components of each frequency of the frequency domain signal generated by N-1 i-th frequency component generation units 4122, and an IFFT 4124 that converts (returns) the frequency domain signal generated by the frequency domain summer 4123 into a time domain audio signal using an inverse fast Fourier transform. The time domain audio signal output by the IFFT 4124 becomes the localization enhancement component output by the localization enhancement component generation unit 412.

[0023] Here, the predetermined gain used by the i-th frequency component generation unit 4122 when generating a frequency domain signal by converting each frequency component of the target sound into a component of i times the frequency is set so that the sound component corresponding to the signal in that frequency domain does not sound unnatural in the output sound from the audio output device 6. Here, the above-described positional sense enhancement component generation unit 412 may be configured as an i-th frequency component generation unit 4122, where i is a single integer of 2 or more, and comprises one i-th frequency component generation unit 4122.

[0024] The configuration of the localization enhancement component generation unit 412 has been described above. However, if the target sound is concentrated in a narrow band within the mid-range (800Hz-2kHz), the localization enhancement component generation unit 412 may be configured as shown in Figure 4. As shown in the figure, the localization enhancement component generation unit 412 adds M-1 1 / j frequency component generation units 4125 to the configuration shown in Figure 3, where j is an integer from 2 to M (M≧2). The 1 / j frequency component generation unit 4125 generates a frequency domain signal by converting the components of each frequency of the target sound from the output of the FFT 4121 to components of 1 / j times their frequency with a predetermined gain. Therefore, the 1 / j frequency component generation unit 4125 generates a frequency domain signal of the (j-1)th harmonic for each frequency component of the target sound.

[0025] Furthermore, in this configuration, the frequency domain summer 4123 generates a frequency domain signal by adding together the components of each frequency of the frequency domain signal generated by N-1 i-times frequency component generators 4122 and the components of each frequency of the frequency domain signal generated by M-1 j-times frequency component generators 4122 for each frequency, and outputs this signal to the IFFT 4124.

[0026] The IFFT4124 then converts the frequency domain signal generated by the frequency domain summer 4123 into a time domain audio signal and outputs it as a localization enhancement component. Here, the predetermined gain used by the 1 / j frequency component generation unit 4125 when generating a frequency domain signal by converting each frequency component of the target sound into a 1 / j frequency component is set so that the sound component corresponding to the signal in that frequency domain does not sound unnatural in the output sound from the audio output device 6. Here, the above-described positional sense enhancement component generation unit 412 may be configured as an i-th frequency component generation unit 4122, where i is an integer of 2 or more, and it comprises one i-th frequency component generation unit 4122. Alternatively, it may be configured as a 1 / j-th frequency component generation unit 4125, where j is an integer of 2 or more, and it comprises one 1 / j-th frequency component generation unit 4125.

[0027] With the above configuration, for target sounds that have poor localization due to their narrow frequency band, the addition of harmonics and subharmonics expands the frequency band of sounds associated with the target sound, thereby providing a better sense of localization. Next, in applications where the target sound may be short in duration (for example, 20ms or less), a configuration that time-stretches the localization enhancement component may be added to the localization enhancement component generation unit 412 shown in Figures 3 and 4. Here, time stretching is a signal processing that extends the duration of the sound without changing its pitch.

[0028] In other words, in this case, as shown in Figure 5, when a configuration for time-stretching the localization enhancement component is added to the localization enhancement component generation unit 412 in Figure 4, the localization enhancement component generation unit 412 is provided with a target sound length detection unit 4126 for detecting the duration of the target sound.

[0029] Here, the target sound length detection unit 4126 detects the duration of the target sound from the target sound. However, the target sound length detection unit 4126 may also detect the duration of the target sound from the frequency domain signal output by the FFT 4121. Furthermore, a time stretching unit 4127 is provided corresponding to the i-time component generation unit 4122 and the 1 / j-time component generation unit 4125, respectively. This time stretching unit 4127 extends the duration of the frequency domain signal of the target sound, which is obtained by converting the frequencies output by the i-time component generation unit 4122 and the 1 / j-time component generation unit 4125 to i-time or 1 / j-time, to a predetermined time (for example, 20ms) or more without changing the pitch (frequency), and outputs it to the frequency domain summing unit 4123.

[0030] Then, when the target sound length detection unit 4126 detects that the target sound is shorter than a predetermined time length (for example, 20ms or less), each time stretching unit 4127 stretches the duration of the frequency domain signal of the target sound, which has been converted to a frequency of i or 1 / j, to a time longer than the predetermined time (for example, a predetermined time of 20ms or more) without changing the pitch.

[0031] In applications where the duration of the target sound is always shorter than a predetermined duration, the target sound duration detection unit 4126 may detect the target sound, and in response to this detection, each time stretching unit 4127 may extend the duration of the frequency domain signal of the target sound, which has been converted to a multiple of i or 1 / j, to a duration longer than the predetermined duration without changing the pitch.

[0032] In this way, by performing time stretching, it becomes possible to provide a sense of localization of the target sound even when the duration of the target sound is too short to be reliably perceived. Next, Figure 6a shows the configuration of the target sound separation unit 411 of the audio signal processing device 4. As shown in the figure, the target sound separation unit 411 can be a DNN4111 (Deep Neural Network) that has been pre-trained to extract target sounds from audio signals. Furthermore, if the frequency bands of the target sound and other sounds do not substantially overlap in the source sound, as shown in Figure 6b, a Band-Pass Filter (BPF4112) is used as the target sound separation unit 411 to extract the sound within the frequency band of the target sound. ) can also be used.

[0033] Embodiments of the present invention have been described above. In this embodiment, the audio signal processing device 4 may further incorporate a head-related transfer function. Figure 7 shows the configuration of the audio signal processing device 4 in this case. As shown in the figure, this audio signal processing device 4 includes an L channel processing unit 41, an R channel processing unit 42, and a direction estimation unit 701. The L channel processing unit 41 and the R channel processing unit 42 have the same configuration and each includes a target sound separation unit 411 that separates the target sound from the audio signal of the corresponding channel of the source sound as described above, a localization enhancement component generation unit 412 that generates a localization enhancement component from the target sound output by the target sound separation unit 411 as described above, a target sound subtractor 711 that subtracts the target sound output by the target sound separation unit 411 from the audio signal of the corresponding channel of the source sound, a target sound adder 712 that synthesizes the output of the localization enhancement component generation unit 412 and the target sound output by the target sound separation unit 411 by addition, a localization correction unit 713 that convolves a head transfer function into the output of the target sound adder 712, and an output adder 714 that synthesizes the output of the localization correction unit 713 and the output of the target sound subtractor 711 by addition and outputs it as an audio signal of the corresponding channel of the output sound.

[0034] Furthermore, the direction estimation unit 701 estimates the direction of the sound source of the target sound, represented by the source sound, relative to the listener, as the localization position direction of the sound image of the target sound, based on the ratio of volume levels and time delay between the target sound output by the target sound separation unit 411 of the L channel processing unit 41 and the target sound output by the target sound separation unit 411 of the R channel processing unit 42.

[0035] The head-related transfer function (HRF) convolved by the localization correction unit 713 of the L channel processing unit 41 is calculated from the target sound source to the listener's left ear, with a predetermined distance to the target sound source and the estimated localization position direction of the target sound source. The head-related transfer function (HRF) convolved by the localization correction unit 713 of the R channel processing unit 42 is calculated from the target sound source to the listener's right ear, with a predetermined distance to the target sound source and the estimated localization position direction of the target sound source.

[0036] Therefore, the outputs of the L channel processing unit 41 and the R channel processing unit 42 are a composite obtained by adding together the components of the audio signal of the corresponding channel of the source sound other than the target sound, and the component obtained by convolving the head transfer function onto the target sound and the localization enhancement component generated from the target sound.

[0037] Therefore, the configuration shown in Figure 7 is equivalent to the configuration of the audio signal processing device 4 shown in Figure 2, in which the localization enhancement component added by the summing unit 413 is replaced with a component obtained by convolving the localization enhancement component with a head-related transfer function, and the target sound in the source sound added by the summing unit 413 is replaced with a sound obtained by convolving the said target sound with a head-related transfer function.

[0038] Furthermore, if the target sound included in the source sound already has a head-related transfer function (HTF) convolved into it, the HTF convolved by the localization correction unit 713 of the L channel processing unit 41 and the R channel processing unit 42 may be set to take this already convolved HTF into consideration, so that the HTF convolved into the output of the localization correction unit 713 is an appropriate HTF for the localization position direction estimated by the direction estimation unit 701.

[0039] Furthermore, if the target sound is the footsteps or gunshots of an enemy in an FPS (First Person Shooter) or TPS (Third Person Shooter) game, and the AV equipment 1 running the game program can obtain the coordinate information of the enemy relative to the player in the game space, then instead of the direction estimation unit 701 of the audio signal processing device 4 shown in Figure 7, a direction estimation unit 701 may be provided as shown in Figure 8, which estimates the direction of the enemy relative to the player in the game space as the localization position direction of the sound image of the target sound, based on the coordinate information of the enemy relative to the player in the game space obtained from the AV equipment 1.

[0040] Furthermore, if the target sound is the footsteps or gunshots of an enemy in an FPS (First Person Shooter) or TPS (Third Person Shooter) game, and the AV device 1 running the game program displays a map image M representing the player's and enemy's positions on a map of the game space on the display 3, as shown in Figure 9a, then instead of the direction estimation unit 701 of the audio signal processing device 4 shown in Figure 7, a direction estimation unit 701 may be provided, as shown in Figure 9b, which analyzes the video output by the AV device 1 to determine the direction of the enemy relative to the player in the game space, and estimates the determined direction as the localization position direction of the sound image of the target sound.

[0041] In this way, by estimating the localization position direction of the target sound image and convolving an appropriate head-related transfer function with the localization enhancement component and the target sound based on the localization position direction estimated by the direction estimation unit 701, it is expected that a better sense of localization can be provided. In the above description, we have shown the case where AV equipment 1 outputs two stereo audio signals, the L channel and the R channel, as source sound. However, this embodiment can also be applied to cases where AV equipment 1 outputs three or more audio signals as source sound by providing a processing unit similar to the L channel processing unit 41 and the R channel processing unit 42 for each channel.

[0042] In summary, according to the embodiments described above, the following audio signal processing device is provided. 1. An audio signal processing device having signal processing units provided for each channel of a multi-channel audio signal, each signal processing unit comprising: a target sound separation unit 411 that separates a target sound signal representing a target sound that is the sound to be targeted for localization enhancement from a source sound signal which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a localization enhancement sound signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by k (where k is an integer of 2 or more); and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result. 2. An audio signal processing device having a signal processing unit provided for each channel of a multi-channel audio signal, each signal processing unit comprising: a target sound separation unit 411 that separates a target sound signal representing a target sound that is the sound to be targeted for localization enhancement from a source sound signal which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a localization enhancement sound signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by 1 / L (where L is an integer of 2 or more); and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0043] 3. An audio signal processing device having signal processing units provided for each channel of a multi-channel audio signal, each signal processing unit comprising: a target sound separation unit 411 that separates a target sound signal representing a target sound that is the sound to be enhanced for localization from a source sound signal which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a first signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by k (where k is an integer of 2 or more), and a second signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by 1 / L (where L is an integer of 2 or more), and also synthesizes the first signal and the second signal to generate a localization enhancement sound signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0044] 4. An audio signal processing device having a signal processing unit provided for each channel of a multi-channel audio signal, each signal processing unit comprising: a target sound separation unit 411 that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a signal for each integer i from 2 to n (where n>2) by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by i, and synthesizes the generated signals to generate a localization enhancement sound signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0045] 5. An audio signal processing device having a signal processing unit provided for each channel of a multi-channel audio signal, each signal processing unit having: a target sound separation unit 411 that separates a target sound signal, which is the sound to be targeted for localization enhancement, from a source sound signal, which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that, for each j, takes integers i from 2 to m (where m > 2), generates a signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit 411 by 1 / j, and synthesizes the generated signals to generate a localization enhancement sound signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0046] 6. Audio signal processing device having signal processing units provided for each channel of a multi-channel audio signal, each signal processing unit having: a target sound separation unit 411 that separates a target sound signal representing a target sound that is the sound to be enhanced for localization from a source sound signal which is the audio signal of the corresponding channel; a localization enhancement sound generation unit that generates a first signal for each i, where i is an integer from 2 to n (where n>2), and for each j, the frequency of the target sound signal separated by the target sound separation unit 411 is multiplied by i; and a localization enhancement sound generation unit that generates a localization enhancement sound signal by combining the generated first signal and each second signal; and a synthesis unit that synthesizes the localization enhancement sound signal generated by the localization enhancement sound generation unit and the source sound signal and outputs the result.

[0047] 7, an audio signal processing device in which each of the signal processing units is provided with a time-stretching means for extending the duration of the localization enhancement sound signal synthesized by the synthesis unit when the duration of the target sound signal separated by the target sound separation unit 411 is shorter than a predetermined duration.

[0048] An audio signal processing device comprising, in 8, 1 to 7, a direction estimation means for estimating the direction in which the target sound should be localized, and a localization correction unit 713 in each of the signal processing units, which applies a head-related transfer function that conforms to the direction estimated by the direction estimation means to the localization enhancement sound signal synthesized by the synthesis unit, and replaces the target sound signal in the source sound signal synthesized by the synthesis unit with a signal in which a head-related transfer function that conforms to the direction estimated by the direction estimation means has been applied to the target sound signal.

[0049] In 9 and 8, the direction estimation means is an audio signal processing device that estimates the direction in which the target sound should be localized based on the relationship between the target sound signals separated by the target sound separation unit 411 of each signal processing unit. In 10 and 8, the direction estimation means is an audio signal processing device that estimates the direction in which the target sound should be localized based on information regarding the position of the sound source of the target sound output from the device that outputs the audio signals of the multiple channels. [Explanation of Symbols]

[0050] 1...AV equipment, 2...Input device, 3...Display, 4...Audio signal processing device, 5...Amplifier, 6...Audio output device, 41...L channel processing unit, 42...R channel processing unit, 411...Target sound separation unit, 412...Localization enhancement component generation unit, 413...Addition unit, 701...Direction estimation unit, 711...Target sound subtractor, 712...Target sound adder, 713...Localization correction unit, 714...Output adder, 4111...DNN, 4112...BPF, 4121...FFT, 4122...i-thread frequency component generation unit, 4123...Frequency domain addition unit, 4124...IFFT, 4125...1 / j-thread frequency component generation unit, 4126...Target sound length detection unit, 4127...Time stretching unit.

Claims

1. It has a signal processing unit provided for each channel of a multi-channel audio signal, Each signal processing unit: A target sound separation unit separates the target sound signal, which represents the sound targeted for localization enhancement, from the source sound signal, which is the audio signal of the corresponding channel. The localization enhancement sound generation unit generates a localization enhancement sound signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by k (where k is an integer of 2 or more), An audio signal processing device characterized by having a combining unit that combines the localization enhancement sound signal generated by the localization enhancement sound generation unit with the source sound signal and outputs the result.

2. It has a signal processing unit provided for each channel of a multi-channel audio signal, Each signal processing unit: A target sound separation unit separates the target sound signal, which represents the sound targeted for localization enhancement, from the source sound signal, which is the audio signal of the corresponding channel. The localization enhancement sound generation unit generates a localization enhancement sound signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / L (where L is an integer of 2 or more), An audio signal processing device characterized by having a combining unit that combines the localization enhancement sound signal generated by the localization enhancement sound generation unit with the source sound signal and outputs the result.

3. It has a signal processing unit provided for each channel of a multi-channel audio signal, Each signal processing unit: A target sound separation unit separates the target sound signal, which represents the sound targeted for localization enhancement, from the source sound signal, which is the audio signal of the corresponding channel. A localization enhancement sound generation unit generates a first signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by k (where k is an integer of 2 or more), and a second signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / L (where L is an integer of 2 or more), and combines the first signal and the second signal to generate a localization enhancement sound signal. An audio signal processing device characterized by having a combining unit that combines the localization enhancement sound signal generated by the localization enhancement sound generation unit with the source sound signal and outputs the result.

4. It has a signal processing unit provided for each channel of a multi-channel audio signal, Each signal processing unit: A target sound separation unit separates the target sound signal, which represents the sound targeted for localization enhancement, from the source sound signal, which is the audio signal of the corresponding channel. A localization enhancement sound generation unit generates a signal for each integer i from 2 to n (where n > 2) by multiplying the frequency of the target sound signal separated by the target sound separation unit by i, and synthesizes the generated signals to generate a localization enhancement sound signal. An audio signal processing device characterized by having a combining unit that combines the localization enhancement sound signal generated by the localization enhancement sound generation unit with the source sound signal and outputs the result.

5. It has a signal processing unit provided for each channel of a multi-channel audio signal, Each signal processing unit: A target sound separation unit separates the target sound signal, which represents the sound targeted for localization enhancement, from the source sound signal, which is the audio signal of the corresponding channel. Let i be an integer from 2 to m (where m > 2), and for each j, the localization enhancement sound generation unit generates a signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / j, and synthesizes the generated signals to generate a localization enhancement sound signal. An audio signal processing device characterized by having a combining unit that combines the localization enhancement sound signal generated by the localization enhancement sound generation unit with the source sound signal and outputs the result.

6. It has a signal processing unit provided for each channel of a multi-channel audio signal, Each signal processing unit: A target sound separation unit separates the target sound signal, which represents the sound targeted for localization enhancement, from the source sound signal, which is the audio signal of the corresponding channel. A localization enhancement sound generation unit generates a localization enhancement sound signal by combining the generated first and second signals, with each i being an integer from 2 to n (where n > 2), and for each i, it generates a first signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by i, and with each j being an integer from 2 to m (where m > 2), it generates a second signal obtained by multiplying the frequency of the target sound signal separated by the target sound separation unit by 1 / j, and combines the generated first and second signals. An audio signal processing device characterized by having a combining unit that combines the localization enhancement sound signal generated by the localization enhancement sound generation unit with the source sound signal and outputs the result.

7. An audio signal processing apparatus according to claim 1, 2, 3, 4, 5, or 6, The audio signal processing apparatus is characterized in that each of the signal processing units has a time-stretching means for extending the duration of the localization enhancement sound signal synthesized by the synthesis unit when the duration of the target sound signal separated by the target sound separation unit is shorter than a predetermined duration.

8. An audio signal processing apparatus according to claim 1, 2, 3, 4, 5, or 6, The system includes a direction estimation means for estimating the direction in which the target sound should be localized, Each of the aforementioned signal processing units is: An audio signal processing device comprising: a synthesis unit that applies a head-level transfer function to the localization-enhancing sound signal synthesized by the synthesis unit, which is adapted to the direction estimated by the direction estimation means; and a localization correction unit that replaces the target sound signal in the source sound signal synthesized by the synthesis unit with a signal to which a head-level transfer function adapted to the direction estimated by the direction estimation means has been applied to the target sound signal.

9. An audio signal processing device according to claim 8, The direction estimation means is an audio signal processing device characterized in that it estimates the direction in which the target sound should be localized based on the relationship between the target sound signals separated by the target sound separation unit of each signal processing unit.

10. An audio signal processing device according to claim 8, The direction estimation means is characterized by estimating the direction in which the target sound should be localized based on information regarding the position of the sound source of the target sound output from the device that outputs the audio signals of the multiple channels.

Citation Information

Patent Citations

  • Sound image localization processing device

    JP3395809B2