Impulse Response Generation Device and Impulse Response Generation Program

The impulse response generation device automatically separates early reflection and late reverberation sounds to generate a large number of impulse responses with similar timbres and low mutual correlations, addressing the inefficiencies of manual parameter adjustment in existing methods.

JP7705327B2Active Publication Date: 2025-07-09NIPPON HOSO KYOKAI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2021156154
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-09-24
Publication Date
2025-07-09
Estimated Expiration
2041-09-24

AI Technical Summary

Technical Problem

Existing methods for generating a large number of impulse responses with similar timbres and low mutual correlations require manual adjustment of processing parameters based on the acoustic characteristics of the sound field, making the process cumbersome and time-consuming.

Method used

An impulse response generation device that automatically separates an impulse response into early reflection and late reverberation sounds using a signal separation unit, time window multiplication units, random number generation units, and time shift units to generate a large number of impulse responses with similar timbres and low mutual correlations.

Benefits of technology

The device can efficiently and automatically produce a large number of impulse responses with similar timbres and low mutual correlations, reducing the need for manual parameter adjustment and shortening the measurement time required for generating multiple impulse responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007705327000013
    Figure 0007705327000013
  • Figure 0007705327000014
    Figure 0007705327000014
  • Figure 0007705327000015
    Figure 0007705327000015
Patent Text Reader

Abstract

To automatically discriminate boundary between initial reflection and rear reverberation of an impulse response, and to generate from a single impulse response a number of impulse responses that have similar timbre and low correlation with each other.SOLUTION: An impulse response generation device includes: a signal separation section that separates an impulse response into an initial reflection sound and a rear reverberation sound; a time window multiplication section that multiplies the initial reflection sound and the rear reverberation sound by a window function obtained by shifting a time window of a predetermined time length several times; a random number generation section that generates a time shift amount by a random number; a time-shifting section that time-shifts the initial reflection sound and the rear reverberation sound, for each time window based on the time shift amount; and a signal adding section that adds the time-shifted initial reflection sound and the rear reverberation sound. The signal separation section separates the impulse response into the initial reflection sound and the rear reverberation sound, based on an evaluation function obtained by applying an ε filter to the impulse response.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an impulse response generation device and an impulse response generation program, and more particularly to an impulse response generation technology for generating a large number of impulse responses having a timbre similar to a given impulse response and having a low correlation with each other from a given impulse response.

Background Art

[0002] In the sound systems of movies and televisions, in addition to the conventional two-channel (ch) stereo system and 5.1ch surround, in recent years, multi-channel sound systems with a large number of channels such as 7.1ch, 10.2ch, and 22.2ch have been adopted. In the production of multi-channel sound content, in the same way as the conventional stereo system and 5.1ch sound system, a reverberation addition device is used to give a sense of spatial expansion and rich presence.

[0003] There are various types of reverberation addition devices, such as those of the IIR (Infinite impulse response) filter type that simulate the reverberation of space by simulation and those of the FIR (Finite impulse response) filter type that generate reverberation using the impulse response measured in the actual space. However, when adding reverberation by convolving the impulse response with the sound source like the FIR filter type, the number of impulse responses corresponding to the number of channels of the playback system is required. For example, in the case of a 22.2 multi-channel sound system, 22ch of impulse responses are required. Furthermore, considering measuring these impulse responses in a diffuse sound field (a state where sounds with equal timbre and uncorrelated with each other arrive from each arrival direction) such as a concert hall, it is required to have a similar timbre and a low correlation with each other.

[0004] Although such a large number of impulse responses can be obtained by actual measurement, it is not easy because it is necessary to measure the impulse responses at points that are sufficiently separated from each other, which requires a large-scale measurement system and a long measurement time.

[0005] Therefore, a device has been proposed that generates a large number of impulse responses having similar timbres and low mutual correlation from a given single impulse response (see, for example, Patent Documents 1 and 2, Non-Patent Documents 1 and 2). The devices described in Patent Documents 1 and 2 separate the impulse response into an early reflection part (a part composed of low-order reverberations from the floor, ceiling, walls, etc., which arrive within a relatively short time after the direct sound arrives) and a late reverberation part (a part that follows the early reflection and repeats multiple reverberations, making it impossible to distinguish individual sounds). Then, predetermined processing is applied to each of the early reflection part and the late reverberation part of the impulse response, and the processed early reflection part and late reverberation part of the impulse response are synthesized and output to generate a large number of impulse responses.

Prior Art Documents

Patent Documents

[0006]

Patent Document 1

Patent Document 2

Non-Patent Documents

[0007]

Non-Patent Document 1

Non-Patent Document 2

Non-Patent Document 3

Non-Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0008] In the embodiments described in Patent Documents 1 and 2, in order to generate a large number of impulse responses with similar timbres and low mutual correlations from a given single impulse response, it is necessary to properly use different processing parameters for the early reflection sound part and the late reverberation sound part. Therefore, a signal separation unit 10 is provided to separate the early reflection sound and the late reverberation sound and process them separately. In the signal separation unit 10 of Patent Documents 1 and 2, the sound in the range where the amplitude of the input impulse response exceeds a predetermined threshold is treated as the early reflection sound, or the early reflection sound is treated as a predetermined time from the direct sound to separate the early reflection sound and the late reverberation sound. However, since the predetermined threshold and the predetermined time vary depending on the acoustic characteristics of the actual sound field where the impulse response is obtained, the user has to check the waveform of the impulse response that is the source of generation to determine the boundary between the early reflection sound and the late reverberation sound, and perform processing based on that time.

[0009] Therefore, in view of the above problems, an object of the present invention is to automatically discriminate the boundary (the length of the early reflection sound) between the early reflection sound and the late reverberation sound of an impulse response, and to provide an impulse response generation device and an impulse response generation program that can generate a large number of impulse responses with similar timbres and low mutual correlations from a given single impulse response.

Means for Solving the Problems

[0010] To solve the above problems, the impulse response generation device according to the present invention includes a signal separation unit that separates an impulse response into an early reflection sound and a late reverberation sound, a first time window multiplication unit that multiplies a first window function obtained by shifting a plurality of times a first time window having a predetermined time length suitable for the early reflection sound by the early reflection sound, and a second time window multiplication unit that multiplies a second window function obtained by shifting a plurality of times a second time window having a predetermined time length suitable for the late reverberation sound by the late reverberation sound, a first random number generation unit that generates a first time shift amount by a random number based on a distribution width suitable for the early reflection sound, and a second random number generation unit that generates a second time shift amount by a random number based on a distribution width suitable for the late reverberation sound, a first time shift unit that time-shifts, within the time range of the early reflection sound, the output of the first time window multiplication unit regarding the early reflection sound for each first time window based on the first time shift amount, and a second time shift unit that time-shifts, within the time range of the late reverberation sound, the output of the second time window multiplication unit regarding the late reverberation sound for each second time window based on the second time shift amount, and a signal addition unit that adds the output of the first time shift unit regarding the early reflection sound and the output of the second time shift unit regarding the late reverberation sound. The signal separation unit separates the impulse response into an early reflection sound and a late reverberation sound based on an evaluation function obtained by applying an ε filter to the impulse response.

[0011] Also, the impulse response generation device further includes a time setting unit that selects one or more blocks in units of time windows for at least one of the outputs of the first time window multiplication unit regarding the early reflection sound and the second time window multiplication unit regarding the late reverberation sound to delete them or selects one or more blocks in units of time windows and duplicates each of them to change the time length. The first time shift unit time-shifts, within the time range of the early reflection sound, the output of the time setting unit regarding the early reflection sound for each first time window based on the first time shift amount, and the second time shift unit time-shifts, within the time range of the late reverberation sound, the output of the time setting unit regarding the late reverberation sound for each second time window based on the second time shift amount, which is desirable.

[0012] Further, it is desirable that the impulse response generation device determines a time range in which the signal separation unit separates the early reflection sound and the late reverberation sound according to the reverberation time of the impulse response, and performs the separation at a time when the evaluation function is maximized within the time range.

[0013] Further, it is desirable that the impulse response generation device determines a time range in which the signal separation unit separates the early reflection sound and the late reverberation sound according to the reverberation time of the impulse response, determines a threshold value ε based on the maximum value of the amplitude of the impulse response within the time range, and performs the separation at the last time when the amplitude of the evaluation function is equal to or greater than the threshold value ε within the time range.

[0014] Further, it is desirable that the impulse response generation device determines a time range in which the signal separation unit separates the early reflection sound and the late reverberation sound according to the reverberation time of the impulse response, further obtains an average value of the absolute values of the amplitudes within a predetermined time window of the impulse response, generates a new impulse response by normalizing the amplitude of the impulse response with the average value while moving the time window, determines a threshold value ε based on the maximum value of the amplitude of the new impulse response within the time range, and performs the separation at the last time when the amplitude of the evaluation function obtained by applying an ε filter to the new impulse response is equal to or greater than the threshold value ε within the time range.

[0015] In order to solve the above problems, an impulse response generation program according to the present invention causes a computer to function as the impulse response generation device.

Effects of the Invention

[0016] According to the impulse response generation device and the impulse response generation program of the present invention, the boundary (length of the early reflection sound) between the early reflection sound and the late reverberation sound of the impulse response can be automatically discriminated, and a large number of impulse responses having similar timbres and low mutual correlation can be generated from one given impulse response.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Best Mode for Carrying Out the Invention

[0018] Hereinafter, embodiments of the present invention will be described in detail with reference to the accompanying drawings.

[0019] (First Embodiment) FIG. 1 is a diagram showing an example of the configuration of an impulse response generation device according to a first embodiment of the present invention. The impulse response generation device according to the present embodiment includes a signal separation unit 10 that separates an impulse response into an initial reflected sound and a late reverberation sound, a time window multiplication unit 20 that multiplies a window function obtained by shifting a time window of a predetermined time length a plurality of times by the initial reflected sound and the late reverberation sound, a random number generation unit 30 that generates a time shift amount by a random number, a time shift unit 40 that time-shifts the output of the time window multiplication unit 20 based on the time shift amount, and a signal addition unit 50 that adds the outputs of the time shift unit 40 regarding the initial reflected sound and the late reverberation sound. Each functional unit 10 to 50 of the impulse response generation device is configured by a suitable processor such as a CPU or a suitable electric circuit.

[0020] The time window multiplication unit 20 includes a first time window multiplication unit 21 and a second time window multiplication unit 22, the random number generation unit 30 includes a first random number generation unit 31 and a second random number generation unit 32, and the time shift unit 40 includes a first time shift unit 41 and a second time shift unit 42. As will be described later, the first time window multiplication unit 21, the first random number generation unit 31, and the first time shift unit 41 perform processing on the initial reflected sound of the impulse response, and the second time window multiplication unit 22, the second random number generation unit 32, and the second time shift unit 42 perform processing on the late reverberation sound of the impulse response.

[0021] The signal separation unit 10 automatically separates the input impulse response h(t) into an initial reflected sound e(t) and a late reverberation sound r(t). The impulse response h(t) serves as a reference for a newly generated impulse response h'(t).

[0022] First, with reference to FIG. 2 schematically showing the reverberant sound (impulse response) of a diffuse sound field, the initial reflected sound and the late reverberant sound of the reverberant sound will be outlined. The sound emitted from a sound source such as a musical instrument or a speaker first arrives as a direct sound. Next, the initial reflected sound reflected by the floor, ceiling, and walls arrives. The initial reflected sound is composed of low-order reflections that can be heard separately from other sounds depending on the conditions. Further, the multiple reflected sounds after repeated reflections gradually arrive as late reverberant sounds that are generally delayed by several milliseconds to about 150 milliseconds or more from the direct sound. When it becomes a late reverberant sound, the density of the reflected sound suddenly increases, and the reflected sounds arriving from all directions overlap with each other to become a diffuse sound having no directivity as a whole. Thus, the reverberant sound is composed of the initial reflected sound and the late reverberant sound having different properties.

[0023] However, particularly in the impulse response measured in an actual sound field, each initial reflected sound also has a temporal spread, and the late reverberant sound and the ambient noise of the sound field also overlap. Therefore, the boundary between the initial reflected sound and the late reverberant sound does not become clear like the waveform of the schematic diagram in FIG. 2. Therefore, in order to automatically separate the initial reflected sound and the late reverberant sound in the signal separation unit 10, it is necessary to automatically determine the time that becomes the boundary between the two.

[0024] Although details will be described later, in the present embodiment, a group of reflected sounds obtained by cutting out the impulse response with a window function (time window) having a predetermined shape and time length is randomly shifted (rearranged) within a predetermined time length to generate a large number of impulse responses having low correlation with each other. For the generated large number of impulse responses to have low correlation with each other and be reverberant sounds having similar timbres, the time length of the time window and the time shift amount of the reflected sound group are important parameters, and it has been confirmed by experiments that the optimum values of these parameters are different in the processing of the initial reflected sound and the late reverberant sound (Non-Patent Documents 1 and 2). In the initial reflected sound portion, it is necessary to divide the individual peaks due to the low-order reflections with a time window and shift them within the time from the arrival of the direct sound to the vicinity of the last arrival time of the main low-order reflected sound having a larger amplitude than the late reverberant sound. Therefore, the signal separation unit 10 of the present application is characterized in that it can automatically discriminate the length of the initial reflected sound from the arrival of the direct sound to the vicinity of the last arrival time of the main low-order reflected sound having a larger amplitude than the late reverberant sound.

[0025] The operation of the signal separation unit 10 will be described for each embodiment below.

[0026] (Embodiment 1 of the signal separation unit 10) First, from the input impulse response h(t), the attenuation curve E(t) of the shredder is obtained by the following equation (1) (Non-Patent Document 3). It is assumed that the direct sound arrives at t = 0.

[0027] [Equation]

[0028] Since the actually measured impulse response is discrete-time data by digital processing, the attenuation curve of the discrete-time system corresponding to Equation (1) is obtained as the following equation (2) (N is the maximum range of measurement).

[0029] [Equation]

[0030] Next, the reverberation time T RT is obtained from the attenuation curve of the impulse response. The reverberation time T RT is, for example, as in T 30 (the reverberation time obtained by the evaluation in the 30 dB range). From the initial level of the attenuation curve E(t), the evaluation interval from -5 dB to -35 dB is linearly approximated by the least squares method to obtain the slope, and it can be obtained from the time corresponding to 60 dB attenuation with that slope.

[0031] To simply and surely perform the discrimination process of t0 with the length of the early reflection sound as t0, the range of t0 is determined in advance. The length of the early reflection sound (the boundary time between the early reflection sound and the late reverberation sound) is such that in a space with a large room volume, since the average traveling distance for the sound wave from the sound source to reach the sound receiving point after reflection becomes long, the reverberation time and the boundary time of that space also become large values. Generally, this boundary time is from several ms to 150 ms, but in a large space such as an arena cave, considering that the early reflection sound arrives even later, when the reverberation time is T RT is used, TRT Determine the time range of t0 to be discriminated as the length of the initial reflected sound according to the following. As an example, according to the reverberation time, the following condition (3) is set as the time range where the boundary between the initial reflected sound and the late reverberation sound exists, and the maximum value of the evaluation function is searched within the time range that satisfies this condition to determine t0. The evaluation function will be described in detail later.

[0032]

Number

[0033] Find the maximum value max(|h(t0)|) of the absolute value of the impulse response h(t0) in the time range of t0 determined as condition (3), and let ε be the value obtained by multiplying this value by a predetermined constant ε0.

[0034]

Number

[0035] Here, define the non-linear function F(x) as in the following equation (5).

[0036]

Number

[0037] Figure 3 is an example of the non-linear function F(x), and F(x) is set as in the following equation (6). Note that the non-linear function F(x) may be a function other than equation (6) as long as it satisfies equation (5).

[0038]

Number

[0039] The evaluation function used in the present invention utilizes an ε-filter. The ε-filter is a non-linear digital filter having a function of removing additive small-amplitude high-frequency noise superimposed on a signal including a sudden large-amplitude change. By introducing the above-mentioned non-linear function F(x) into a non-recursive linear digital filter with x(n) as the input and y(n) as the output, the following equation (7) can be derived. Here, the coefficient a(k) is the coefficient of a non-recursive linear digital low-pass filter whose sum is 1. The filter represented by this equation (7) is called an ε-filter (Non-Patent Document 4).

[0040]

Number

[0041] Next, an evaluation function |G(t)| using an ε-filter is introduced. If the sampling frequency is f s then t = n / f s Therefore, when expressed in a discrete-time system, the evaluation function becomes the following equation (8).

[0042]

Number

[0043] As described above, the coefficient a(k) is the coefficient of a non-recursive linear digital low-pass filter (hereinafter referred to as an LP filter) whose sum is 1. In this embodiment, |G(n)| represented by equation (8) is used as the evaluation function, and based on this evaluation function, the impulse response is separated into an initial reflected sound and a late reverberation sound.

[0044] In addition, when setting the evaluation function, the value of a constant ε0 is set in consideration of the discrimination sensitivity of the initial reflected sound. Further, the cut-off frequency of the LP filter is set based on the frequency characteristics of the impulse response, and the order of the coefficient a(k) is determined in consideration of calculation accuracy, calculation processing load, etc. Here, ε0 = 0.5, and the coefficient a(k) is the coefficient of a 256th-order LP filter having the frequency characteristics shown in FIG. 4.

[0045] In Example 1, the time t0 when the evaluation function reaches its maximum is set as the boundary between the early reflected sound and the late reverberant sound. In a general space such as a hall, the early reflected sound is intermittently observed as a reflected sound with a larger amplitude compared to the late reverberant sound. Empirically, even near the boundary between the early reflected sound and the late reverberant sound within the range of condition (3), the early reflected sound is considered to be observed as a reflected sound with a relatively large amplitude. Example 1 uses this largest reflected sound that has passed through the ε filter as the boundary, which is a simple boundary detection method. In a discrete-time system, n = n0 that maximizes the evaluation function is obtained. Let the sampling frequency of the measured impulse response be f s Then the length t0 of the early reflected sound is t0 = n0 / f s That is, by obtaining n0 that satisfies condition (3) and maximizes the evaluation function |G(n)| in Equation (8), the length t0 of the early reflected sound is t0 = n0 / f s This method shows an example of discriminating the length of the early reflected sound from the impulse response measured indoors such as in a concert hall or a small booth in a studio.

[0046] Figure 5 shows an example of separating the early reflected sound and the late reverberant sound in a small space (small booth) with a reverberation time T RT = 350 ms (when T RT < 1 s). The graph in Figure 5 has the time waveform of the impulse response in the upper part and the evaluation function |G(t)| in the lower part. The time on the horizontal axis is common up to a maximum of 0.25 seconds. The dashed line in the graph is the boundary (length of the early reflected sound) t0 between the early reflected sound and the late reverberant sound discriminated by the method of Example 1 described above. In this example, t0 = 10.4 ms.

[0047] Similarly, Figure 6 shows an example of separating the early reflected sound and the late reverberant sound in a relatively large space (concert hall) with a reverberation time T RT = 1.7 s (when T RT ≥ 1 s). The graph in Figure 6 has the time waveform of the impulse response in the upper part and the evaluation function |G(t)| in the lower part. The time on the horizontal axis is common up to a maximum of 1 second. The dashed line in the graph is the boundary t0 between the early reflected sound and the late reverberant sound discriminated by the method of Example 1. In this example, t0 = 77.7 ms.

[0048] Since the ε filter has the characteristic of suppressing small-amplitude high-frequency noise while maintaining sudden changes in the input signal, when an impulse response h(t) is input to the ε filter, the initial reflected sound of the sudden change passes through the filter, and the multi-reflected sound (late reverberation sound) that can be regarded as small-amplitude noise is suppressed by the LP filter. Therefore, the initial reflected sound can be emphasized in the evaluation function.

[0049] As shown by the dashed line on the impulse response waveforms of FIGS. 5 and 6, according to the method of Example 1, the boundary between the initial reflected sound and the late reverberation sound can be discriminated as the maximum peak portion of the main initial reflected sound. In addition, when there are multiple maximum values of the evaluation function within the time range of condition (3) that are the same, it is desirable to adopt the later time as t0.

[0050] (Example 2 of the signal separation unit 10) In a sound field with an extremely large room volume or a sound field with a structure in which a plurality of rooms with greatly different volumes are connected, the initial reflection may arrive with a large time delay. In Example 1, there may be a case where this delayed initial reflected sound cannot be discriminated. The example shown in FIG. 7 is an example of separating the initial reflected sound and the late reverberation sound in a large space (T RT = 8.6 s). The upper part is the time waveform of the impulse response, and the lower part is the evaluation function |G(t)|. The dashed line in the graph is the boundary (length of the initial reflected sound) t0 between the initial reflected sound and the late reverberation sound discriminated by the method of Example 1 described above. In this method, t0 is discriminated as 80 ms, but the initial reflected sound arriving around 100 ms later should be used as the boundary. Example 2 is a method that can handle such cases.

[0051] When ε is sufficiently small with respect to the amplitude of the impulse response to which ε is input, the evaluation function, which is the output of the ε filter, reflects large-amplitude sudden changes such as initial reflected sounds, while reflected sounds with an impulse response amplitude within ±ε / 2 have the characteristic of being smoothed by a linear low-pass filter. Therefore, in Example 2, within the range of searching for the length t0 of the initial reflected sound in condition (3), with ε as the threshold, the last time when the amplitude of the evaluation function becomes ε or more is defined as the length t0 of the initial reflected sound. Examples of the separation process of the initial reflected sound and the late reverberation sound by the method of Example 2 are shown in FIGS. 8 to 10. The constant ε0 and the coefficients a(k) of the LP filter are the same as those in Example 1.

[0052] FIG. 8 is an example of processing the impulse response of FIG. 7, which was discriminated as having a short initial reflected sound length in Example 1, by the method of Example 2. The upper part is the time waveform of the impulse response, and the lower part is the evaluation function |G(t)|, which is the same as FIG. 7 as a waveform diagram. In the method of this Example 2, by obtaining the last time t0 when the amplitude of the evaluation function becomes ε or more, the initial reflected sound part can be discriminated as t0 = 99.2 ms, and it can be seen that this is an improvement from Example 1 in this regard.

[0053] FIGS. 9 and 10 show the results of obtaining the length t0 of the initial reflected sound by the method of this Example 2 for the sound fields corresponding to FIGS. 5 and 6 of Example 1 (the same as FIGS. 5 and 6 as waveform diagrams), respectively. That is, FIG. 9 is the case of applying to a small space (small booth) with a reverberation time T RT = 350 ms (T RT <1 s), and the last time when the amplitude of the evaluation function becomes ε or more is t0 = 10.8 ms. Similarly, FIG. 10 is the case of applying to a relatively large space (concert hall) with a reverberation time T RT = 1.7 s (T RT ≧ 1 s), and the last time when the amplitude of the evaluation function becomes ε or more is t0 = 80.2 ms. In these two examples, the length of the initial reflected sound can be discriminated with almost the same values as in Example 1. Therefore, the method of Example 2 can also be applied to the separation of the initial reflected sound and the late reverberation sound in a sound field with a short reverberation time.

[0054] (Example 3 of the signal separation unit 10) In Example 3 of the signal separation unit 10, in the measured impulse response h(n), the function h WE obtained by normalizing the amplitude with the average value of the absolute value of the amplitude within T nr seconds is newly obtained. That is, within the samples of n from 0 to N WE (=T WE ·fs), the average value of the absolute value |h(n)| of h(n) is obtained, and the amplitude of each h(n) from n = 0 to N is divided by the average value to normalize the amplitude of h(n). Subsequently, while moving the time window (while shifting by one sample), h(n) is repeatedly normalized, and a new impulse response is defined by the following formula (9). WE up to N WE +1, and the amplitude is divided by the average value of |h(n)| and normalized. A new impulse response obtained by normalizing h(n) is defined by the following formula (9).

[0055]

Equation

[0056] Instead of h(n) in formula (8), an evaluation function |G'(n)| obtained by applying an ε filter to h nr (n) is defined as follows in formula (10).

[0057]

Equation

[0058] In Example 3, within the range of searching for the length t0 of the initial reflected sound under condition (3) in the same manner as in Example 2, the last time when the value |G'(n)| of the above evaluation function becomes equal to or greater than the threshold value ε is defined as the length t0 of the initial reflected sound. Note that ε is set based on the newly normalized impulse response h nr (n).

[0059] Here, the features of Example 3 will be described in comparison with the method of Example 2. FIG. 11 shows the reverberation time T by the method of Example 2 RTThis is an example where, in a large space of 6.8 s, the length t0 of the initial reflected sound was discriminated as 85 ms. From the waveform of the impulse response h(t) in the upper part of the figure, it seems that the initial reflected sound arrives overlapping the late reverberant sound also around 140 ms. However, the value of the evaluation function |G(t)| around 140 ms in the lower part is very small, making it difficult to distinguish from the late reflected sound. It is difficult to lower the discrimination threshold of the initial reflected sound by reducing the value of the constant ε0 in Equation (4) for discrimination.

[0060] Therefore, Fig. 12 shows an example where the method of Example 3 is applied to the same impulse response h(t) as in Fig. 11. The upper part is the time waveform h(t) of the impulse response (the same as h(t) in the upper part of Fig. 11), the middle part is the impulse response h nr (t) normalized by the average value of a predetermined time window, and the lower part is the evaluation function |G'(t)| obtained by applying the ε filter to the normalized h nr (n). From the evaluation function |G'(t)|, the initial reflected sound around 140 ms can be discriminated (t0 = 139 ms). In this example, as shown by the waveform of h nr (t) in the middle part of Fig. 12, since the initial reflected sound overlapping or lurking in the late reverberant sound part is emphasized by the effect of Equation (9) and then the ε filter is applied, the initial reflected sound clearly appears around 140 ms in the evaluation function |G'(n)| shown in the lower part of Fig. 12. With the method of Example 3, it is possible to discriminate and separate the initial reflected sound overlapping or lurking in the late reverberant sound part in this way. In this data example, fs = 48000 Hz and T WE was set to 0.01 s.

[0061] In Examples 1 to 3, the discrimination sensitivity of the initial reflected sound can be appropriately adjusted by the value of the aforementioned constant ε0. That is, when ε0 is increased, only the initial reflected sound with a larger amplitude is discriminated, and when ε0 is set to a smaller value, the initial reflected sound with a smaller amplitude can also be discriminated. However, it can be applied with the same ε0 (ε0 = 0.5 in this example) from a small space with a reverberation time of 350 ms to an extremely large space with a reverberation time of 8.6 s as in the above-described processing examples.

[0062] In each embodiment of the signal separation unit 10, in condition (3), the reverberation time is divided into two cases: less than 1 s and 1 s or more, and the range for searching the boundary (the length of the initial reflected sound) t0 between the desired initial reflected sound and the late reverberation sound is set. However, in order to cope with various acoustic spaces, this classification may be increased.

[0063] Returning to FIG. 1, the configuration of the impulse response generation device will be described.

[0064] As described above, based on the waveform characteristics of the impulse response h(t), the signal separation unit 10 automatically separates the sound from the direct sound to the length t0 of the initial reflected sound as the initial reflected sound e(t), and the sound after t0 as the late reverberation sound r(t). Then, the signal separation unit 10 outputs the initial reflected sound e(t) to the first time window multiplication unit 21 and outputs the late reverberation sound r(t) to the second time window multiplication unit 22.

[0065] Hereinafter, the processing of the first time window multiplication unit 21, the first random number generation unit 31, and the first time shift unit 41 for the initial reflected sound e(t) will be described in detail.

[0066] The first time window multiplication unit 21 multiplies the initial reflected sound e(t) by a window function We n (t) (n = 1, 2,..., k) obtained by shifting a time window We(t) of a predetermined time length Tw a plurality of times (k times) to output e n (t)=We n (t)×e(t) (n = 1, 2,..., k) is generated. Note that as the time window, a Hanning window with little change in amplitude-frequency characteristics due to window multiplication can be used. Also, a Hamming window, a Blackman window, etc. may be used as the window. When the window function We n (t) is obtained by shifting the time window We(t) at equal intervals Ted, the output e n (t) is represented by Equation (11). The first time window multiplication unit 21 outputs the output e n (t) to the first time shift unit 41.

[0067]

Equation

[0068] The first random number generator 31 generates a time shift amount by a random number. The first time shift unit 41, based on the time shift amount, shifts the output e n (t) of the first time window multiplication unit 21 for the initial reflected sound e(t) within the time range of the initial reflected sound e(t) at each time window to generate an output e´(t). Here, if the time shift amount N * (n) generated by the first random number generator 31 is a random number of a uniform distribution taking an integer value with a distribution width N, the output e n (t) shifted in time, the output e´(t), is represented by the following equation (12).

[0069] [Number]

[0070] The first time shift unit 41 shifts the output e n (t) for the initial reflected sound e(t) within the time range of the initial reflected sound. Thus, it controls such that n+N * (n) satisfies 1≦n+N * (n)≦k. When e´(t) goes outside the range of the initial reflected sound e(t), the first time shift unit 41 causes the first random number generator 31 to generate a time shift amount by a random number again, and shifts the output e n (t) of the first time window multiplication unit 21 within the time range of the initial reflected sound e(t). The first time shift unit 41 outputs the generated e´(t) to the signal addition unit 50.

[0071] FIG. 13 is a diagram showing an outline of generating a new impulse response by time shifting. FIG. 13(a) shows an impulse response serving as a reference for generating a new impulse response, and a window function We n (t) with a window width Tw is multiplied by the initial reflected sound e(t). Here, the output e of the first time window multiplication unit 21 corresponding to the nth window function indicated by hatching nFor (t), the first random number generator 31 generates, for example, a time shift amount (-2, -1, 0, 1, 2) with a width N = 5 by random numbers. As shown in Fig. 13(b), the first time shift unit 41 shifts the output e n (t) corresponding to the n-th window function based on the time shift amount (for example, -1). In this way, by shifting the output of the first time window multiplication unit 21 for each time window based on random numbers, an impulse response h'(t) is generated that has a timbre similar to the early reflection sound e(t) of a given impulse response h(t) and has low mutual correlation early reflection sounds e'(t).

[0072] The processing of the second time window multiplication unit 22, the second random number generator 32, and the second time shift unit 42 for the late reverberation sound r(t) is the same as the processing of the first time window multiplication unit 21, the first random number generator 31, and the first time shift unit 41 for the early reflection sound e(t), respectively. Note that parameter values such as the length of the window function (window width) and the range of random numbers of the time shift amount (distribution width) can be individually set to values suitable for each of the early reflection sound e(t) and the late reverberation sound r(t). That is, in the first time window multiplication unit 21, a first time window with a predetermined time length suitable for the early reflection sound is used, in the second time window multiplication unit 22, a second time window with a predetermined time length suitable for the late reverberation sound is used, in the first random number generator 31, the first time shift amount is generated based on the distribution width suitable for the early reflection sound, and in the second random number generator 32, it is desirable to generate the second time shift amount based on the distribution width suitable for the late reverberation sound.

[0073] The second time window multiplication unit 22 multiplies the late reverberation sound r(t) by window functions Wr n (t) (n = 1, 2,... m) obtained by shifting a plurality of time windows Wr(t) with a predetermined time length Tr each, and outputs r n (t) (n = 1, 2,... m) is generated. The second random number generator 32 generates a time shift amount by random numbers. The second time shift unit 42, based on the time shift amount, for each time window, the output r of the second time window multiplication unit 22 regarding the late reverberation sound r(t) nOutput r'(t) is generated by time-shifting (t) within the time range of the late reverberation sound r(t). When r'(t) goes outside the range of the late reverberation sound r(t), the second time-shifting unit 42 causes the second random number generation unit 32 to generate again a time-shifting amount that falls within the range of the late reverberation sound r(t). The second time-shifting unit 42 outputs the generated r'(t) to the signal addition unit 50.

[0074] The signal addition unit 50 adds the output e´(t) of the first time-shifting unit 41 regarding the early reflection sound e(t) and the output r'(t) of the second time-shifting unit 42 regarding the late reverberation sound r(t) to generate a new impulse response h'(t).

[0075] Thus, according to this embodiment, the signal separation unit 10 automatically separates the impulse response into the early reflection sound and the late reverberation sound. The time window multiplication unit 20 multiplies the window functions obtained by shifting a time window of a predetermined time length a plurality of times by the early reflection sound and the late reverberation sound. The random number generation unit 30 generates the time-shifting amount by a random number. The time-shifting unit 40 time-shifts, for each time window, the output of the first time window multiplication unit 21 regarding the early reflection sound within the time range of the early reflection sound and the output of the second time window multiplication unit 22 regarding the late reverberation sound within the time range of the late reverberation sound based on the time-shifting amount. The signal addition unit 50 adds the outputs of the time-shifting unit 40 regarding the early reflection sound and the late reverberation sound to generate a new impulse response. As a result, it becomes possible to generate a large number of impulse responses having a timbre similar to a given impulse response in advance and having a low correlation with each other, and to provide an impulse response of sufficient quality for use in reverberation addition in a multi-channel acoustic system. In particular, according to this embodiment, it becomes possible to accurately and automatically separate the early reflection sound and the late reverberation sound and to generate a large number of impulse responses having a low correlation between the early reflection sound and the late reverberation sound, and it is possible to obtain a sense of spread that cannot be obtained from the same impulse response when the impulse response is convolved.

[0076] (Second Embodiment) FIG. 14 is a diagram showing an example of the configuration of an impulse response generation device according to the second embodiment of the present invention. The impulse response generation device according to the present embodiment includes a signal separation unit 10 that separates an impulse response into an initial reflection sound and a late reverberation sound, a time window multiplication unit 20 that multiplies a window function obtained by shifting a time window of a predetermined time length a plurality of times by the initial reflection sound and the late reverberation sound, a time setting unit 60 that changes the impulse response to an arbitrary time length, a random number generation unit 30 that generates a time shift amount by a random number, a time shift unit 40 that time-shifts the output of the time setting unit 60 based on the time shift amount, and a signal addition unit 50 that adds the outputs of the time shift unit 40 regarding the initial reflection sound and the late reverberation sound. Each functional unit 10 to 60 of the impulse response generation device is configured by a suitable processor such as a CPU or a suitable electric circuit.

[0077] The time window multiplication unit 20 includes a first time window multiplication unit 21 and a second time window multiplication unit 22, the time setting unit 60 includes a first time setting unit 61 and a second time setting unit 62, the random number generation unit 30 includes a first random number generation unit 31 and a second random number generation unit 32, and the time shift unit 40 includes a first time shift unit 41 and a second time shift unit 42. As will be described later, the first time window multiplication unit 21, the first time setting unit 61, the first random number generation unit 31, and the first time shift unit 41 perform processing on the initial reflection sound of the impulse response, and the second time window multiplication unit 22, the second time setting unit 62, the second random number generation unit 32, and the second time shift unit 42 perform processing on the late reverberation sound of the impulse response.

[0078] The signal separation unit 10 automatically separates the sound from the direct sound to the length t0 of the initial reflection sound as the initial reflection sound e(t) and the sound after t0 as the late reverberation sound r(t). The separation method is as described in the first embodiment above. Then, the initial reflection sound e(t) is output to the first time window multiplication unit 21, and the late reverberation sound r(t) is output to the second time window multiplication unit 22.

[0079] Next, the processing of the first time window multiplication unit 21, the first time setting unit 61, the first random number generation unit 31, and the first time shift unit 41 for the initial reflection sound e(t) will be described in detail.

[0080] The first time window multiplication unit 21 is the same as the first time window multiplication unit 21 of the first embodiment, and a window function We n (t) (n = 1, 2, … k) obtained by shifting a time window We(t) of a predetermined time length Tw a plurality of times (k times) is multiplied by the initial reflected sound e(t) and the output e n (t)=We n (t)×e(t) (n = 1, 2, … k) is generated. When the window function We n (t) is obtained by shifting the time window We(t) at equal intervals Ted, the output e n (t) is represented by Equation (11). The first time window multiplication unit 21 outputs the output e n (t) to the first time setting unit 61.

[0081] The first time setting unit 61 deletes or duplicates the output of the first time window multiplication unit 21 for each time window with respect to the initial reflected sound e n (t) (length t0) of the input impulse response h(t) to generate a new initial reflected sound e'(t) of length t e . When the reverberation time is shortened (t e <t0), one or more blocks after the window function multiplication of the original initial reflected sound e(t) are selected and deleted. When the reverberation time is lengthened (t e >t0), one or more blocks after the window function multiplication of the original initial reflected sound e(t) are selected and appropriately duplicated to generate a new initial reflected sound e'(t). Regarding the method of deletion or duplication, for example, it may be performed on equally spaced blocks, or on randomly selected blocks. Also, blocks in any continuous section may be deleted or duplicated. As will be described later, an impulse response h'(t) with a changed reverberation time is generated from the new initial reflected sound e'(t) and the new late reverberation sound r'(t).

[0082] FIG. 15 is a diagram showing an outline of the process for shortening the reverberation time. When the initial reflected sound e(t) is e1 to e12, the new initial reflected sound e'(t) is represented by, for example, e1, e3, e4, e6, e7, e9, e10, e12 in which the blocks e2, e5, e8, e11 in time window units are deleted. Thereby, the first time setting unit 61 can generate a new initial reflected sound e'(t) having a length of 2 / 3 of the original length.

[0083] FIG. 16 is a diagram showing an outline of the process for lengthening the reverberation time. When the initial reflected sound e(t) is e1 to e12, the new initial reflected sound e'(t) is represented by, for example, e1, e2, e3, e3, e4, e5, e6, e6, e7, e8, e9, e9, e10, e11, e12, e12 in which the blocks e3, e6, e9, e12 in time window units are duplicated. Thereby, the first time setting unit 61 can generate a new initial reflected sound e'(t) having a length of 4 / 3 of the original length.

[0084] The first random number generator 31 is the same as in the first embodiment, and generates a time shift amount by a random number. Here, the time shift amount N * (n) is a random number of a uniform distribution taking an integer value of the distribution width N.

[0085] The first time shift unit 41, based on the time shift amount, time-shifts each component e' n (t) of the initial reflected sound e'(t) after the time length change, which is the output of the first time setting unit 61, within the time range of the initial reflected sound e'(t) to generate an output e''(t).

[0086] The outline of generating a new impulse response by time shift is the same as in FIG. 13. That is, the output e' n (t) [corresponding to the component indicated by the hatching in FIG. 13(a)] corresponding to the nth window function is shifted in time by the time shift amount N from the first random number generator 31 *Perform time shifting based on (n) [Fig. 13(b)]. In this way, by performing time shifting based on random numbers for each time window, an initial reflected sound e''(t) is generated that has a timbre similar to the initial reflected sound e'(t) of the impulse response h'(t) and has a low correlation with each other.

[0087] Note that the first time shift unit 41 time-shifts each output e' n (t) of the initial reflected sound e'(t) after the time length change within the time range of the initial reflected sound, so n + N * (n) satisfies 1 ≤ n + N * (n) ≤ (t e -Tw + Ted) / Ted and controls accordingly. When e''(t) is outside the range of the initial reflected sound e'(t), the first random number generation unit 31 generates a time shift amount by random numbers again, and time-shifts the output e' n (t) of the first time window multiplication unit 21 within the time range of the initial reflected sound e'(t). The first time shift unit 41 outputs the generated e''(t) to the signal addition unit 50.

[0088] The processing of the second time window multiplication unit 22, the second time setting unit 62, the second random number generation unit 32, and the second time shift unit 42 for the late reverberation sound r(t) is the same as the processing of the first time window multiplication unit 21, the first random number generation unit 31, and the first time shift unit 41 for the initial reflected sound e(t), respectively. Note that parameter values such as the length of the window function (window width) and the range of random numbers of the time shift amount (distribution width) can be individually set to values suitable for the initial reflected sound e(t) and the late reverberation sound r(t) respectively. That is, the first time window multiplication unit 21 uses a first time window with a predetermined time length suitable for the initial reflected sound, the second time window multiplication unit 22 uses a second time window with a predetermined time length suitable for the late reverberation sound, the first random number generation unit 31 generates a first time shift amount based on a distribution width suitable for the initial reflected sound, and the second random number generation unit 32 generates a second time shift amount based on a distribution width suitable for the late reverberation sound.

[0089] The second time window multiplication unit 22 shifts a plurality of window functions Wr obtained by shifting a time window Wr(t) with a predetermined time length Tr respectively. nMultiply (t) (n = 1, 2, … m) by the late reverberation sound r(t) and output r n (t) (n = 1, 2, … m) is generated. The second time setting unit 62 deletes or duplicates, for each time window, the output of the second time window multiplication unit 22 regarding the late reverberation sound r(t) (length t1) of the input impulse response h(t) to generate a new late reverberation sound r'(t) of length t r When shortening the reverberation time (t r <t1), the block of the original late reverberation sound r(t) is appropriately deleted, and when lengthening the reverberation time (tr > t1), the block of the original late reverberation sound r(t) is appropriately duplicated to generate a new late reverberation sound r'(t). The second random number generation unit 32 generates the time shift amount by a random number. The second time shift unit 42 shifts each component r' n (t) of the late reverberation sound r'(t) after the time length change within the time range of the late reverberation sound r'(t) to generate the output r''(t). When r''(t) goes out of the range of the late reverberation sound r'(t) after the time length change, the second time shift unit 42 causes the second random number generation unit 32 to generate a time shift amount that is within the range of the late reverberation sound r'(t) again. The second time shift unit 42 outputs the generated r''(t) to the signal addition unit 50.

[0090] The signal addition unit 50 adds the output e''(t) of the first time shift unit 41 regarding the early reflection sound e(t) and the output r''(t) of the second time shift unit 42 regarding the late reverberation sound r(t) to generate a new impulse response h''(t).

[0091] In this embodiment, a configuration is provided in which a time shift unit 40 that shifts the reverberation component for each time window is provided after the time setting unit 60 that changes the length of the reverberation time. However, the time shift unit 40 and the related random number generation unit 30 may be omitted, and a configuration may be adopted in which a new impulse response h'(t) is generated directly from the output of the time setting unit 60. In this case, the signal addition unit 50 adds the output e'(t) of the first time setting unit 61 regarding the early reflection sound e(t) and the output r'(t) of the second time setting unit 62 regarding the late reverberation sound r(t) to generate a new impulse response h'(t).

[0092] Thus, according to this embodiment, the signal separation unit 10 automatically separates the impulse response into an initial reflection sound and a late reverberation sound. The time window multiplication unit 20 multiplies the window functions obtained by shifting a time window of a predetermined time length a plurality of times by the initial reflection sound and the late reverberation sound. The time setting unit 60 deletes or duplicates the output of the time window multiplication unit 20 for at least one of the initial reflection sound or the late reverberation sound for each time window to change the time length. By adding the initial reflection sound and the late reverberation sound with the changed time lengths, it becomes possible to generate a large number of impulse responses having a timbre similar to a given impulse response and having an arbitrarily set reverberation time with low correlation with each other, and to provide impulse responses of sufficient quality and various types for use in reverberation addition in a multi-channel acoustic system.

[0093] Furthermore, the random number generation unit 30 generates a time shift amount by a random number, and the time shift unit 40 time-shifts the output of the first time setting unit 61 regarding the initial reflection sound within the time range of the initial reflection sound for each time window based on the time shift amount. Thereby, it becomes possible to generate a large number of impulse responses with low correlation of the initial reflection sound, and it is possible to obtain a sense of spread that cannot be obtained from the impulse response of the same initial reflection sound when the impulse response is convolved. The same applies to the late reverberation sound.

[0094] In the above embodiment, the configuration and operation of the impulse response generation device have been described, but the present invention is not limited to this, and it may be configured as a method for generating a large number of impulse responses. That is, it may be configured as an impulse response generation method in an impulse response generation device that sequentially includes the processing steps in each part of the impulse response generation device according to the data flow of FIG. 1 or FIG. 14.

[0095] Note that a computer can be suitably used to function as the impulse response generation device described above. Such a computer can be realized by storing a program describing the processing content for realizing each function of the impulse response generation device in the storage unit of the computer, and reading and executing this program by the CPU of the computer. Note that this program (impulse response generation program) can be recorded on a computer-readable recording medium.

[0096] Although the above-described embodiments have been described as representative examples, it is obvious to those skilled in the art that many changes and substitutions can be made within the spirit and scope of the present invention. Therefore, the present invention should not be construed as being limited by the above-described embodiments, and various modifications or changes are possible without departing from the scope of the claims. For example, the functions included in each block, each step, etc. described in the embodiments can be rearranged so as not to be logically contradictory, and a plurality of constituent blocks, steps, etc. can be combined into one or divided.

Explanation of Reference Numerals

[0097] 10 Signal separation unit 20 Time window multiplication unit 21 First time window multiplication unit 22 Second time window multiplication unit 30 Random number generation unit 31 First random number generation unit 32 Second random number generation unit 40 Time shift unit 41 First time shift unit 42 Second time shift unit 50 Signal addition unit 60 Time setting unit 61 First time setting unit 62 Second time setting unit

Claims

1. a signal separation unit that separates an impulse response into an initial reflected sound and a late reverberation sound; a first time window multiplication unit that multiplies the initial reflected sound by a first window function obtained by shifting a plurality of times a first time window having a predetermined time length suitable for the initial reflected sound; and a second time window multiplication unit that multiplies the late reverberation sound by a second window function obtained by shifting a plurality of times a second time window having a predetermined time length suitable for the late reverberation sound; a first random number generation unit that generates a first time shift amount by a random number based on a distribution width suitable for the initial reflected sound, and a second random number generation unit that generates a second time shift amount by a random number based on a distribution width suitable for the late reverberation sound; a first time shift unit that time-shifts, within the time range of the initial reflected sound, the output of the first time window multiplication unit regarding the initial reflected sound for each first time window based on the first time shift amount, and a second time shift unit that time-shifts, within the time range of the late reverberation sound, the output of the second time window multiplication unit regarding the late reverberation sound for each second time window based on the second time shift amount; a signal addition unit that adds the output of the first time shift unit regarding the initial reflected sound and the output of the second time shift unit regarding the late reverberation sound; An impulse response generation device comprising: the signal separation unit separates the impulse response into an initial reflected sound and a late reverberation sound based on an evaluation function obtained by applying an ε filter to the impulse response.

2. The impulse response generation device according to claim 1, further comprising a time setting unit that selects one or more blocks in units of time windows for at least one of the outputs of the first time window multiplication unit regarding the initial reflected sound and the second time window multiplication unit regarding the late reverberation sound, deletes them, or selects one or more blocks in units of time windows, duplicates each of them, and changes the time length; the first time shift unit time-shifts, within the time range of the initial reflected sound, the output of the time setting unit regarding the initial reflected sound for each first time window based on the first time shift amount, and the second time shift unit time-shifts, within the time range of the late reverberation sound, the output of the time setting unit regarding the late reverberation sound for each second time window based on the second time shift amount. An impulse response generation device characterized by this.

3. An impulse response generation device according to claim 1 or 2, wherein the signal separation unit determines a time range for separating the early reflected sound and the late reverberant sound according to the reverberation time of the impulse response, and performs the separation at a time when the evaluation function is maximized within the time range.

4. An impulse response generation device according to claim 1 or 2, wherein the signal separation unit determines a time range for separating the early reflected sound and the late reverberant sound according to the reverberation time of the impulse response, determines a threshold value ε based on the maximum value of the amplitude of the impulse response within the time range, and performs the separation at the last time when the amplitude of the evaluation function is equal to or greater than the threshold value ε within the time range.

5. An impulse response generation device according to claim 1 or 2, wherein the signal separation unit determines a time range for separating the early reflected sound and the late reverberant sound according to the reverberation time of the impulse response, further obtains an average value of the absolute values of the amplitudes within a predetermined time window of the impulse response, generates a new impulse response in which the amplitude of the impulse response is normalized by the average value while moving the time window, determines a threshold value ε based on the maximum value of the amplitude of the new impulse response within the time range, and performs the separation at the last time when the amplitude of the evaluation function obtained by applying an ε filter to the new impulse response is equal to or greater than the threshold value ε within the time range.

6. An impulse response generation program for causing a computer to function as the impulse response generation device according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Rectifier for magnet type ac generator

    JP1987012336A

  • Reformer of fuel cell power generating system

    JP1988048773A

  • Sound correction device, sound output device and sound correction method

    JP2012015792A

  • Device, method, and program for generating impulse response

    JP2015219413A

  • Method, signal processor, audio encoder, audio decoder and binaural renderer for processing audio signals according to room impulse responses

    JP2016532149A