Audio signal processing device

By separating and frequency-transforming the target sound in the audio signal processing device, a sound signal with enhanced positioning is generated, which solves the problem of insufficient sound positioning perception in narrow frequency bands and achieves better spatial positioning perception effect.

CN121751049APending Publication Date: 2026-03-27ALPS ALPINE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing technologies cannot effectively improve the localization and perception capabilities of narrow-band sounds.

Method used

By setting up a signal processing unit with multiple channels in the audio signal processing device, the target sound is separated and frequency transformed to generate a sound signal with enhanced positioning sense, and then synthesized with the source sound for output. This includes frequency multiplication, reduction or harmonic processing of the target sound, combined with time stretching and head transfer function correction.

Benefits of technology

It achieves better localization perception of narrow frequency band sounds and enhances the spatial localization perception effect of sounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121751049A_ABST
    Figure CN121751049A_ABST
Patent Text Reader

Abstract

The present invention addresses the problem of providing an "audio signal processing device" that achieves a good sense of localization with respect to a sound having a narrow frequency band. [Solution] The present invention is provided with signal processing units provided so as to correspond to respective channels of audio signals of a plurality of channels, and each signal processing unit has: a target sound separation unit (411) that separates a target sound signal indicating a target sound, which is a sound to be subjected to localization enhancement, from a source sound signal, which is an audio signal of the corresponding channel; a sense-of-localization-enhanced sound generation unit that generates a sense-of-localization-enhanced sound signal in which the frequency of the target sound signal separated by the target sound separation unit (411) is multiplied by k (k is an integer of 2 or more); and a synthesis unit that synthesizes and outputs the localization-enhanced sound signal generated by the localization-enhanced sound generation unit and the source sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to an audio signal processing apparatus. Background Technology

[0002] An audio signal processing device is known to provide a sense of localization of a sound image to a sound image location target by convolving the acoustic transmission characteristics (head transfer function) from the sound image location target location to the listener's left and right ears into an audio signal and outputting it (e.g., Patent Document 1).

[0003] Existing technical documents:

[0004] Patent documents:

[0005] Patent Document 1: Japanese Patent No. 3395809 Summary of the Invention

[0006] The problem that the invention aims to solve:

[0007] Humans have a low ability to perceive the direction of narrow-band sounds. Furthermore, the technology that convolves the aforementioned sound transmission characteristics (head transfer function) into the audio signal and outputs it has the problem of not being able to provide sufficient localization for narrow-band sounds.

[0008] Therefore, the objective of this invention is to provide an audio signal processing apparatus that can achieve better localization for narrow-band sounds.

[0009] Methods used to solve problems:

[0010] To achieve the above-mentioned problem, the present invention provides an audio signal processing apparatus having signal processing units respectively provided corresponding to each channel of an audio signal with multiple channels. Each signal processing unit includes: a target sound separation unit that separates a target sound signal representing the sound of the object to be enhanced for location perception from the audio signal of the corresponding channel, i.e., the source sound signal; a location perception enhancement sound generation unit that generates a location perception enhancement sound signal such that the frequency of the target sound signal separated by the target sound separation unit is k times, where k is an integer greater than or equal to 2; and a synthesis unit that synthesizes the location perception enhancement sound signal generated by the location perception enhancement sound generation unit and the source sound signal and outputs it.

[0011] Furthermore, the present invention provides an audio signal processing apparatus having signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels. Each signal processing unit includes: a target sound separation unit that separates a target sound signal representing the sound of the object to be enhanced for location perception from the audio signal of the corresponding channel, i.e., the source sound signal; a location perception enhancement sound generation unit that generates a location perception enhancement sound signal such that the frequency of the target sound signal separated by the target sound separation unit is 1 / L times, wherein L is an integer greater than or equal to 2; and a synthesis unit that synthesizes the location perception enhancement sound signal generated by the location perception enhancement sound generation unit and the source sound signal and outputs it.

[0012] Furthermore, the present invention provides an audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels. Each signal processing unit includes: a target sound separation unit, which separates a target sound signal representing the sound of the object to be enhanced for location perception from the audio signal of the corresponding channel, i.e., the source sound signal; a location perception enhancement sound generation unit, which generates a first signal whose frequency is k times that of the target sound signal separated by the target sound separation unit and a second signal whose frequency is 1 / L times that of the target sound signal separated by the target sound separation unit, and combines the first signal and the second signal to generate a location perception enhancement sound signal, wherein k is an integer of 2 or more and L is an integer of 2 or more; and a synthesis unit, which synthesizes the location perception enhancement sound signal generated by the location perception enhancement sound generation unit and the source sound signal and outputs it.

[0013] Furthermore, the present invention provides an audio signal processing apparatus having signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels. Each signal processing unit includes: a target sound separation unit that separates a target sound signal representing the sound of the object to be enhanced for location perception from the audio signal of the corresponding channel, i.e., the source sound signal; a location perception enhancement sound generation unit that sets each integer from 2 to n as i, generates a signal for each i such that the frequency of the target sound signal separated by the target sound separation unit is i times, and synthesizes the generated signals to generate a location perception enhancement sound signal, wherein n>2; and a synthesis unit that synthesizes the location perception enhancement sound signal generated by the location perception enhancement sound generation unit and the source sound signal and outputs it.

[0014] Furthermore, the present invention provides an audio signal processing apparatus having signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels. Each signal processing unit includes: a target sound separation unit that separates a target sound signal representing the sound of the object to be enhanced for location perception from the audio signal of the corresponding channel, i.e., the source sound signal; a location perception enhancement sound generation unit that sets each integer from 2 to m as j, generates a signal for each j such that the frequency of the target sound signal separated by the target sound separation unit is 1 / j times, and synthesizes the generated signals to generate a location perception enhancement sound signal, wherein m>2; and a synthesis unit that synthesizes the location perception enhancement sound signal generated by the location perception enhancement sound generation unit and the source sound signal and outputs it.

[0015] Furthermore, the present invention provides an audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels. Each signal processing unit includes: a target sound separation unit, which separates a target sound signal representing the sound of the object to be enhanced for location perception, i.e., the target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location perception enhancement sound generation unit, which sets each integer from 2 to n as i, and generates a first signal for each i such that the frequency of the target sound signal separated by the target sound separation unit is i times, and sets each integer from 2 to m as j, and generates a second signal for each j such that the frequency of the target sound signal separated by the target sound separation unit is 1 / j times, and synthesizes each of the generated first signals and each of the second signals to generate a location perception enhancement sound signal, wherein n>2, m>2; and a synthesis unit, which synthesizes the location perception enhancement sound signal generated by the location perception enhancement sound generation unit with the source sound signal and outputs it.

[0016] Here, the above-described audio signal processing apparatus may also include a time stretching unit in each of the signal processing units, which stretches the duration of the location enhancement sound signal synthesized by the synthesis unit at least when the duration of the target sound signal separated by the target sound separation unit is shorter than a predetermined time length.

[0017] Alternatively, in the above-described audio signal processing apparatus, a direction estimation unit may be provided, which estimates the direction in which the target sound should be located. Furthermore, a positioning correction unit may be provided in each of the signal processing units. This positioning correction unit applies a head transfer function suitable for the direction estimated by the direction estimation unit to the positioning-enhanced sound signal synthesized by the synthesis unit, and replaces the target sound signal in the source sound signal synthesized by the synthesis unit with a signal obtained by applying a head transfer function suitable for the direction estimated by the direction estimation unit to the target sound signal.

[0018] Alternatively, in this case, the direction estimation unit may estimate the direction in which the target sound should be located based on the relationship between the target sound signals separated by the target sound separation section of each signal processing unit.

[0019] Alternatively, the direction estimation unit may estimate the direction in which the target sound should be located based on information related to the location of the sound source of the target sound, output by the device from the output source of the audio signals of the plurality of channels.

[0020] According to the audio signal processing device described above, for target sounds that lack localization due to narrow frequency bands, by adding harmonics and subharmonics, the frequency band of the sound associated with the target sound can be expanded, thus giving it a better localization.

[0021] Invention effects:

[0022] As described above, according to the present invention, an audio signal processing apparatus can be provided that can achieve better localization for narrow-band sounds. Attached Figure Description

[0023] Figure 1 This is a diagram illustrating the configuration of an AV system according to an embodiment of the present invention.

[0024] Figure 2 This is a diagram illustrating the configuration of an audio signal processing apparatus according to an embodiment of the present invention.

[0025] Figure 3 This diagram illustrates the configuration of the positioning enhancement component generation unit according to an embodiment of the present invention.

[0026] Figure 4 This is a diagram illustrating another configuration example of the positioning enhancement component generation unit according to an embodiment of the present invention.

[0027] Figure 5 This is a diagram illustrating another configuration example of the positioning enhancement component generation unit according to an embodiment of the present invention.

[0028] Figure 6 This is a diagram illustrating the configuration of the target sound separation unit according to an embodiment of the present invention.

[0029] Figure 7 This is a diagram illustrating another configuration example of an audio signal processing apparatus according to an embodiment of the invention.

[0030] Figure 8 This is a diagram illustrating another configuration example of an audio signal processing apparatus according to an embodiment of the invention.

[0031] Figure 9This is a diagram illustrating another configuration example of an audio signal processing apparatus according to an embodiment of the invention.

[0032] Explanation of reference numerals in the attached figures:

[0033] 1…AV equipment, 2…input device, 3…display, 4…audio signal processing device, 5…amplifier, 6…audio output device, 41…L-channel processing unit, 42…R-channel processing unit, 411…target sound separation unit, 412…positioning enhancement component generation unit, 413…addition unit, 701…direction estimation unit, 711…target sound subtractor, 712…target sound adder, 713…positioning correction unit, 714…output adder, 4111…DNN, 4112…BPF, 4121…FFT, 4122…i-fold frequency component generation unit, 4123…frequency domain addition unit, 4124…IFFT, 4125…1 / j-fold frequency component generation unit, 4126…target sound length detection unit, 4127…time stretching unit. Detailed Implementation

[0034] The embodiments of the present invention will be described below.

[0035] Figure 1 This describes the configuration of the AV system in this embodiment.

[0036] As shown in the figure, the AV system includes: AV device 1, which outputs video and audio signals of the same AV content; input device 2, which accepts operations on AV device 1; display 3, which displays the video signal output by AV device 1; audio signal processing device 4, which performs localization signal processing on the audio signal output by AV device 1 to enhance the sound as a target for localization enhancement, and then outputs it; amplifier 5, which amplifies the audio signal output for localization enhancement; and sound output device 6, which emits the sound represented by the audio signal output by amplifier 5.

[0037] AV device 1 is, for example, a PC or a game console; input device 2 is, for example, a keyboard or a game controller; and audio output device 6 is, for example, a speaker or headphones.

[0038] then, Figure 2 This shows the configuration of the audio signal processing device 4.

[0039] Here, as shown in the figure, AV device 1 outputs two-channel stereo audio signals (L and R channels) as the source sound. Furthermore, audio signal processing device 4 outputs the two-channel stereo audio signals (L and R channels) as output sound to amplifier 5.

[0040] Furthermore, the audio signal processing device 4 includes: an L-channel processing unit 41, which processes the L-channel audio signal of the source sound and outputs the L-channel audio signal as the output sound to the amplifier 5; and an R-channel processing unit 42, which processes the R-channel audio signal of the source sound and outputs the R-channel audio signal as the output sound to the amplifier 5.

[0041] The L-channel processing unit 41 and the R-channel processing unit 42 have the same configuration, such as Figure 2 The L-channel processing unit 41 in the image includes: a target sound separation unit 411, which uses the L-channel as the channel corresponding to the L-channel processing unit 41 and the R-channel as the channel corresponding to the R-channel processing unit 42, and separates the target sound, i.e., the target sound, from the audio signal of the corresponding channel of the source sound; a location enhancement component generation unit 412, which generates an audio signal component that enhances the location of the target sound based on the target sound input from the target sound separation unit 411; and an addition unit 413, which adds the location enhancement component generated by the location enhancement component generation unit 412 to the audio signal of the corresponding channel of the source sound through addition, and outputs it as the audio signal of the corresponding channel of the output sound.

[0042] then, Figure 3 This describes the configuration of the positioning enhancement component generation unit 412 in the audio signal processing device 4.

[0043] Figure 3 The configuration shown is that the target sound is concentrated in the low-frequency band (e.g., the frequency band below 100Hz), that is, the frequency band of the target sound is roughly converged in the low-frequency band, which is the configuration of the localization enhancement component generation unit 412.

[0044] Here, as a sound concentrated in the low-frequency band, such as the footsteps of an enemy in an FPS (first-person shooter) game or a TPS (third-person shooter) game played on AV device 1, such footsteps are preferred as target sounds to enhance the sense of positioning by allowing the player to determine the enemy's location and activities through hearing.

[0045] As shown in the figure, the positioning enhancement component generation unit 412 is equipped with FFT4121, which transforms the input target sound into a frequency domain signal through Fast Fourier Transform.

[0046] Furthermore, i is set to integers from 2 to N (N≧2), and N-1 frequency component generation units 4122 are provided. Each frequency component generation unit 4122 generates a frequency domain signal by transforming the frequency components of the target sound to their i-th multiples with a predetermined gain, based on the output of the FFT 4121. Therefore, the frequency component generation unit 4122 generates a frequency domain signal of the (i-1)th harmonic for each frequency component of the target sound.

[0047] In addition, it includes: a frequency domain adder 4123 that generates a frequency domain signal synthesized by adding the frequency components of the frequency domain signal generated by the N-1 i-fold frequency component generation unit 4122 at each frequency; and an IFFT 4124 that transforms (returns) the frequency domain signal generated by the frequency domain adder 4123 into a time domain audio signal through an inverse fast Fourier transform, and the time domain audio signal output by the IFFT 4124 becomes the positioning enhancement component output by the positioning enhancement component generation unit 412.

[0048] Here, the predetermined gain used when the i-fold frequency component generation unit 4122 generates a frequency domain signal after transforming the components of each frequency of the target sound into components of i-fold frequency is set such that the components of the sound corresponding to the signal in that frequency domain do not become unnatural sounds in the output sound from the audio output device 6.

[0049] Here, the above-mentioned positioning enhancement component generation unit 412 may also include an i-fold frequency component generation unit 4122 as the i-fold frequency component generation unit 4122, where i is set to an integer greater than or equal to 2.

[0050] The above describes the configuration of the localization enhancement component generation unit 412. However, when the target sound is concentrated in a narrow frequency band within the mid-range (800Hz-2kHz), the localization enhancement component generation unit 412 can also be configured as follows: Figure 4 It is constructed as shown.

[0051] As shown in the figure, the positioning enhancement component generation unit 412 is in Figure 3 In the configuration shown, j is set to integers from 2 to M (M ≥ 2), and M-1 1 / j-fold frequency component generation units 4125 are added. The 1 / j-fold frequency component generation units 4125 generate frequency domain signals by transforming the frequency components of the target sound into their 1 / j-fold frequencies with a predetermined gain, based on the output of the FFT 4121. Therefore, the 1 / j-fold frequency component generation units 4125 generate the frequency domain signal of the (j-1)th harmonic for each frequency component of the target sound.

[0052] Furthermore, in this configuration, the frequency domain adder 4123 generates a frequency domain signal obtained by adding the frequency components of the frequency domain signal generated by the N-1 i-fold frequency component generation unit 4122 and the frequency components of the frequency domain signal generated by the M-1 j-fold frequency component generation unit 4122 at each frequency, and outputs it to the IFFT 4124.

[0053] Furthermore, IFFT4124 transforms the frequency domain signal generated by the frequency domain adder 4123 into a time domain audio signal, which is then output as a positioning enhancement component.

[0054] Here, the gain used by the 1 / j multiple frequency component generation unit 4125 when generating a frequency domain signal that transforms each frequency component of the target sound into a 1 / j multiple frequency component is set so that the sound component corresponding to the signal in that frequency domain does not become an unnatural sound in the output sound from the audio output device 6.

[0055] Here, the positioning enhancement component generation unit 412 described above may also include an i-fold frequency component generation unit 4122, where i is set to an integer greater than or equal to 2. Alternatively, it may include a 1 / j-fold frequency component generation unit 4125, where j is set to an integer greater than or equal to 2.

[0056] Based on the above configuration, for target sounds that lack localization due to narrow frequency bands, adding harmonics and subharmonics can expand the frequency band of sounds associated with the target sound, thus giving them a better localization.

[0057] Next, in applications where the target sound might be a short-duration sound (e.g., less than 20ms), it is also possible to... Figure 3 or Figure 4 The positioning enhancement component generation unit 412 shown includes a configuration that performs time stretching on the positioning enhancement component. Here, time stretching is a signal processing technique that stretches the duration without changing the pitch of the sound.

[0058] That is, in this case, such as Figure 5 As shown, in Figure 4 When the positioning enhancement component generation unit 412 is configured to perform time stretching on the positioning enhancement component, the positioning enhancement component generation unit 412 includes a target sound length detection unit 4126 that detects the time length of the target sound.

[0059] Here, the target sound length detection unit 4126 detects the duration of the target sound from the target sound. However, the target sound length detection unit 4126 can also detect the duration of the target sound based on the frequency domain signal output by the FFT 4121.

[0060] In addition, a time stretching unit 4127 is provided corresponding to the i-fold frequency component generation unit 4122 and the 1 / j-fold frequency component generation unit 4125, respectively. The time stretching unit 4127 stretches the duration of the signal of the target sound in the frequency domain after the frequency output by the i-fold frequency component generation unit 4122 and the 1 / j-fold frequency component generation unit 4125 is transformed to i-fold or 1 / j-fold without changing the pitch (frequency), and outputs it to the frequency domain addition unit 4123.

[0061] Furthermore, when the target sound length detection unit 4126 detects that the target sound is shorter than the specified time length (e.g., less than 20ms), in each time stretching unit 4127, the duration of the signal in the frequency domain of the target sound after the frequency is changed to i times or 1 / j times is stretched to a time longer than the specified time (e.g., a specified time of more than 20ms) without changing the pitch.

[0062] In addition, in applications where the duration of the target sound is always shorter than the specified duration, the target sound can be detected in the target sound length detection unit 4126. In response to this detection, in each time stretching unit 4127, the duration of the signal in the frequency domain of the target sound after the frequency is changed to i times or 1 / j times is stretched to a duration longer than the specified duration without changing the pitch.

[0063] In this way, by performing time stretching, the localization of the target sound can be provided even when the duration of the target sound is so short that it cannot be perceived stably.

[0064] then, Figure 6 The letter 'a' indicates the configuration of the target sound separation unit 411 of the audio signal processing device 4.

[0065] As shown in the figure, the target sound separation unit 411 can use a DNN4111 (Deep Neural Network) that has been deeply learned by pre-extracting the target sound from the audio signal.

[0066] Furthermore, in the source sound, if the frequency band of the target sound does not substantially overlap with the frequency bands of other sounds, such as Figure 6 As shown in b, the target sound separation unit 411 can also use a BPF4112 (Band-Pass Filter) to extract the frequency band of the target sound.

[0067] The embodiments of the present invention have been described above.

[0068] Here, in the above embodiment, the head transfer function may also be further implemented in the audio signal processing device 4.

[0069] Figure 7 This indicates the configuration of the audio signal processing device 4 under this condition.

[0070] As shown in the figure, the audio signal processing device 4 includes an L-channel processing unit 41, an R-channel processing unit 42, and a direction estimation unit 701.

[0071] The L-channel processing unit 41 and the R-channel processing unit 42 have the same configuration, each including: a target sound separation unit 411, which separates the target sound from the audio signal of the corresponding channel of the source sound as described above; a positioning enhancement component generation unit 412, which generates a positioning enhancement component from the target sound output by the target sound separation unit 411 as described above; a target sound subtractor 711, which subtracts the target sound output by the target sound separation unit 411 from the audio signal of the corresponding channel of the source sound; a target sound adder 712, which synthesizes the output of the positioning enhancement component generation unit 412 and the target sound output by the target sound separation unit 411 by addition; a positioning correction unit 713, which convolves the head transfer function to the output of the target sound adder 712; and an output adder 714, which synthesizes the output of the positioning correction unit 713 and the output of the target sound subtractor 711 by addition and outputs it as the audio signal of the corresponding channel of the output sound.

[0072] In addition, the direction estimation unit 701 estimates the direction of the sound source of the target sound represented by the source sound relative to the listener as the positioning position direction of the sound image of the target sound based on the ratio of the volume level between the target sound output by the target sound separation unit 411 of the L channel processing unit 41 and the target sound output by the target sound separation unit 411 of the R channel processing unit 42 or the time delay.

[0073] Furthermore, the head transfer function convolved by the positioning correction unit 713 of the L-channel processing unit 41 is a head transfer function from the source of the target sound to the listener's left ear, calculated by setting the distance to the source of the target sound to a predetermined distance and setting the direction of the source of the target sound to the estimated positioning position direction. The head transfer function convolved by the positioning correction unit 713 of the R-channel processing unit 42 is a head transfer function from the source of the target sound to the listener's right ear, calculated by setting the distance to the source of the target sound to a predetermined distance and setting the direction of the source of the target sound to the estimated positioning position direction.

[0074] Therefore, the outputs of the L-channel processing unit 41 and the R-channel processing unit 42 are synthesized by adding components other than the target sound of the audio signal of the corresponding channel of the source sound with components obtained by convolving the head transfer function into the target sound and the localization enhancement components generated based on the target sound.

[0075] Therefore, such Figure 7 The composition is equivalent to the following composition: In Figure 2 In the configuration of the audio signal processing device 4 shown, the positioning enhancement component after addition by the adder 413 is replaced with a component obtained by convolving the head transfer function into the positioning enhancement component, and the target sound in the source sound after addition by the adder 413 is replaced with the sound obtained by convolving the head transfer function into the target sound.

[0076] Furthermore, if the target sound contained in the source sound is a sound that has already been convolved with a head transfer function, the head transfer function to be convolved by the positioning correction unit 713 of the L-channel processing unit 41 or the R-channel processing unit 42 can be set in consideration of the already convolved head transfer function, so that the head transfer function convolved into the output of the positioning correction unit 713 becomes a head transfer function suitable for the positioning position direction estimated by the direction estimation unit 701.

[0077] Furthermore, when the target sound is the footsteps or gunshots of an enemy in an FPS (First-Person Shooter) or TPS (Third-Person Shooter) game, and the coordinate information of the enemy relative to the player in the game space can be obtained from the AV device 1 executing the game program, such as... Figure 8 As shown, a direction estimation unit 701 can also be provided instead. Figure 7 The direction estimation unit 701 of the audio signal processing device 4 shown estimates the direction of the enemy relative to the player in the game space as the location direction of the sound image of the target sound based on the coordinate information of the enemy relative to the player in the game space obtained from the AV device 1.

[0078] Furthermore, in situations where the target sound is the footsteps or gunshots of enemies in an FPS (First-Person Shooter) or TPS (Third-Person Shooter) game, and the AV device 1 executing the game program is such... Figure 9 As shown in Figure a, in the case where a map image M representing the positions of players and enemies on a map indicating the game space is displayed on monitor 3, such as... Figure 9 As shown in b, a direction estimation unit 701 can also be provided instead. Figure 7The direction estimation unit 701 of the audio signal processing device 4 shown performs image analysis on the video output by the AV device 1 to obtain the direction of the enemy relative to the player in the game space, and estimates the obtained direction as the location direction of the sound image of the target sound.

[0079] In this way, by estimating the location orientation of the target sound image and convolving the head transfer function, which is suitable for the location orientation estimated by the direction estimation unit 701, into the location enhancement component or the target sound, it is expected to provide a better sense of location.

[0080] The above describes the case where the AV device 1 outputs stereo audio signals from two channels, L and R, as the source sound. However, in this embodiment, even when the AV device 1 outputs audio signals from three or more channels as the source sound, the same application can be performed by setting the same processing unit as the L channel processing unit 41 or the R channel processing unit 42 for each channel.

[0081] In summary, based on the above embodiments, the following audio signal processing apparatus is provided.

[0082] 1. An audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels, each signal processing unit comprising: a target sound separation unit 411, which separates a target sound signal representing the sound of an object for location enhancement, i.e., a target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location enhancement sound generation unit, which generates a location enhancement sound signal such that the frequency of the target sound signal separated by the target sound separation unit 411 is k times, wherein k is an integer greater than or equal to 2; and a synthesis unit, which synthesizes and outputs the location enhancement sound signal generated by the location enhancement sound generation unit and the source sound signal.

[0083] 2. An audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels, each signal processing unit comprising: a target sound separation unit 411, which separates a target sound signal representing the sound of an object for location enhancement, i.e., a target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location enhancement sound generation unit, which generates a location enhancement sound signal such that the frequency of the target sound signal separated by the target sound separation unit 411 is 1 / L times, wherein L is an integer greater than or equal to 2; and a synthesis unit, which synthesizes the location enhancement sound signal generated by the location enhancement sound generation unit and the source sound signal and outputs it.

[0084] 3. An audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels, each signal processing unit comprising: a target sound separation unit 411, which separates a target sound signal representing the sound of an object for location enhancement, i.e., a target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location enhancement sound generation unit, which generates a first signal whose frequency is k times that of the target sound signal separated by the target sound separation unit 411 and a second signal whose frequency is 1 / L times that of the target sound signal separated by the target sound separation unit 411, and combines the first signal and the second signal to generate a location enhancement sound signal, wherein k is an integer of 2 or more and L is an integer of 2 or more; and a synthesis unit, which synthesizes the location enhancement sound signal generated by the location enhancement sound generation unit and the source sound signal and outputs it.

[0085] 4. An audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels, each signal processing unit comprising: a target sound separation unit 411, which separates a target sound signal representing the sound of an object for location enhancement, i.e., a target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location enhancement sound generation unit, which sets each integer from 2 to n as i, generates a signal such that the frequency of the target sound signal separated by the target sound separation unit 411 is i times the value of i, and synthesizes the generated signals to generate a location enhancement sound signal, wherein n>2; and a synthesis unit, which synthesizes the location enhancement sound signal generated by the location enhancement sound generation unit and the source sound signal and outputs it.

[0086] 5. An audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels, each signal processing unit comprising: a target sound separation unit 411, which separates a target sound signal representing the sound of an object for location enhancement, i.e., a target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location enhancement sound generation unit, which sets each integer from 2 to m as j, generates a signal for each j such that the frequency of the target sound signal separated by the target sound separation unit 411 is 1 / j times, and synthesizes the generated signals to generate a location enhancement sound signal, wherein m>2; and a synthesis unit, which synthesizes the location enhancement sound signal generated by the location enhancement sound generation unit and the source sound signal and outputs it.

[0087] 6. An audio signal processing apparatus comprising signal processing units respectively provided corresponding to each channel of an audio signal of multiple channels, each signal processing unit comprising: a target sound separation unit 411, which separates a target sound signal representing the sound of an object for location enhancement, i.e., a target sound, from the audio signal of the corresponding channel, i.e., the source sound signal; a location enhancement sound generation unit, which sets each integer from 2 to n as i, and generates a first signal for each i such that the frequency of the target sound signal separated by the target sound separation unit 411 is i times, and sets each integer from 2 to m as j, and generates a second signal for each j such that the frequency of the target sound signal separated by the target sound separation unit 411 is 1 / j times, and synthesizes the generated first signal and the generated second signal to generate a location enhancement sound signal, wherein n>2, m>2; and a synthesis unit, which synthesizes the location enhancement sound signal generated by the location enhancement sound generation unit with the source sound signal and outputs it.

[0088] 7. In the audio signal processing apparatus of 1 to 6, a time stretching unit is provided in each of the signal processing units. The time stretching unit stretches the duration of the location enhancement sound signal synthesized by the synthesis unit at least when the duration of the target sound signal separated by the target sound separation unit 411 is shorter than a predetermined time length.

[0089] 8. In the audio signal processing apparatus of 1 to 7, a direction estimation unit is provided, which estimates the direction in which the target sound should be located, and a positioning correction unit 713 is provided in each of the signal processing units. The positioning correction unit 713 applies a head transfer function suitable for the direction estimated by the direction estimation unit to the positioning enhancement sound signal synthesized by the synthesis unit, and replaces the target sound signal in the source sound signal synthesized by the synthesis unit with a signal obtained by applying a head transfer function suitable for the direction estimated by the direction estimation unit to the target sound signal.

[0090] 9. In the audio signal processing apparatus of 8, the direction estimation unit estimates the direction in which the target sound should be located based on the relationship between the target sound signals separated by the target sound separation unit 411 of each signal processing unit.

[0091] 10. In the audio signal processing apparatus of 8, the direction estimation unit estimates the direction in which the target sound should be located based on information related to the location of the sound source of the target sound, output from the device of the output source of the audio signals of the plurality of channels.

Claims

1. An audio signal processing device, characterized in that, It has a signal processing unit that is respectively provided for each channel of the audio signal of multiple channels. Each signal processing unit has: The target sound separation unit separates the target sound signal, which represents the sound of the object used for enhancing the sense of localization, from the audio signal of the corresponding channel, i.e., the source sound signal. A localization-enhancing sound generation unit generates a localization-enhancing sound signal such that the frequency of the target sound signal separated by the target sound separation unit is k times that of the target sound signal, where k is an integer greater than or equal to 2; and The synthesis unit combines and outputs the location-enhanced sound signal generated by the location-enhanced sound generation unit and the source sound signal.

2. An audio signal processing device, characterized in that, It has a signal processing unit that is respectively provided for each channel of the audio signal of multiple channels. Each signal processing unit has: The target sound separation unit separates the target sound signal, which represents the sound of the object used for enhancing the sense of localization, from the audio signal of the corresponding channel, i.e., the source sound signal. A localization-enhancing sound generation unit generates a localization-enhancing sound signal such that the frequency of the target sound signal separated by the target sound separation unit is 1 / L, where L is an integer greater than or equal to 2; and The synthesis unit combines and outputs the location-enhanced sound signal generated by the location-enhanced sound generation unit and the source sound signal.

3. An audio signal processing device, characterized in that, It has a signal processing unit that is respectively provided for each channel of the audio signal of multiple channels. Each signal processing unit has: The target sound separation unit separates the target sound signal, which represents the sound of the object used for enhancing the sense of localization, from the audio signal of the corresponding channel, i.e., the source sound signal. A location-enhancing sound generation unit generates a first signal whose frequency is k times that of the target sound signal separated by the target sound separation part, and a second signal whose frequency is 1 / L times that of the target sound signal separated by the target sound separation part. The first signal and the second signal are then combined to generate a location-enhancing sound signal, wherein k is an integer of 2 or more, and L is an integer of 2 or more. The synthesis unit combines and outputs the location-enhanced sound signal generated by the location-enhanced sound generation unit and the source sound signal.

4. An audio signal processing device, characterized in that, It has a signal processing unit that is respectively provided for each channel of the audio signal of multiple channels. Each signal processing unit has: The target sound separation unit separates the target sound signal, which represents the sound of the object used for enhancing the sense of localization, from the audio signal of the corresponding channel, i.e., the source sound signal. The localization-enhancing sound generation unit sets each integer from 2 to n as i, and for each i, generates a signal such that the frequency of the target sound signal separated by the target sound separation part is i times the original frequency. The generated signals are then combined to generate a localization-enhancing sound signal, where n > 2. The synthesis unit combines and outputs the location-enhanced sound signal generated by the location-enhanced sound generation unit and the source sound signal.

5. An audio signal processing device, characterized in that, It has a signal processing unit that is respectively provided for each channel of the audio signal of multiple channels. Each signal processing unit has: The target sound separation unit separates the target sound signal, which represents the sound of the object used for enhancing the sense of localization, from the audio signal of the corresponding channel, i.e., the source sound signal. The localization-enhancing sound generation unit sets each integer from 2 to m as j, and for each j, generates a signal such that the frequency of the target sound signal separated by the target sound separation part is 1 / j times. The generated signals are then combined to generate a localization-enhancing sound signal, where m > 2. The synthesis unit combines and outputs the location-enhanced sound signal generated by the location-enhanced sound generation unit and the source sound signal.

6. An audio signal processing device, characterized in that, It has a signal processing unit that is respectively provided for each channel of the audio signal of multiple channels. Each signal processing unit has: The target sound separation unit separates the target sound signal, which represents the sound of the object used for location enhancement, from the audio signal of the corresponding channel, i.e., the source sound signal. The location-enhancing sound generation unit assigns integers from 2 to n as i, and for each i, generates a first signal such that the frequency of the target sound signal separated by the target sound separation part is i times the original frequency. It assigns integers from 2 to m as j, and for each j, generates a second signal such that the frequency of the target sound signal separated by the target sound separation part is 1 / j times the original frequency. The generated first and second signals are then combined to generate a location-enhancing sound signal, where n>2 and m>2. The synthesis unit synthesizes the location-enhanced sound signal generated by the location-enhanced sound generation unit with the source sound signal and outputs it.

7. The audio signal processing apparatus according to any one of claims 1 to 6, characterized in that, Each of the signal processing units has a time stretching unit, which stretches the duration of the location-enhancing sound signal synthesized by the synthesis unit at least when the duration of the target sound signal separated by the target sound separation unit is shorter than a predetermined time length.

8. The audio signal processing apparatus according to any one of claims 1 to 6, characterized in that, It has a direction estimation unit that estimates the direction in which the target sound should be located. Each of the signal processing units has a positioning correction unit that applies a head transfer function suitable for the direction estimated by the direction estimation unit to the positioning enhancement sound signal synthesized by the synthesis unit, and replaces the target sound signal in the source sound signal synthesized by the synthesis unit with a signal obtained by applying a head transfer function suitable for the direction estimated by the direction estimation unit to the target sound signal.

9. The audio signal processing apparatus according to claim 8, characterized in that, The direction estimation unit estimates the direction in which the target sound should be located based on the relationship between the target sound signals separated by the target sound separation part of each signal processing unit.

10. The audio signal processing apparatus according to claim 8, characterized in that, The direction estimation unit estimates the direction in which the target sound should be located based on information related to the location of the sound source of the target sound, output by the device from the output source of the audio signals of the plurality of channels.