Focused sound source local reproduction device and focused sound source local reproduction method, and program

JP2026144937APending Publication Date: 2026-09-09株式会社ラダ·プロダクション +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
JP2025140677
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-02-28
Filing Date
2025-08-26
Publication Date
2026-09-09

AI Technical Summary

Benefits of technology

【0013】 本開示の一側面によれば、特定のエリア内で音が聞こえるように生成した焦点音源を3次元空間上でリアルタイムに移動させることができる。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026144937000001_ABST
    Figure 2026144937000001_ABST
Patent Text Reader

Abstract

A focused sound source, generated to be audible within a specific area, is moved in real time within a 3D space. [Solution] The target sound field calculation unit calculates a time-domain spatial domain signal on the reference line representing the target sound field, in which the focus sound source exists within the spatial range of the spatial window, by applying a spatial window having a specified spatial range to a signal representing the focus sound source sound field, which is a sound field on a preset reference line of the listening area, in which a focus sound source is generated at a specified focus sound source position. The speaker drive signal generation unit generates a time-domain speaker drive signal for reproducing the focus sound source sound field by performing a two-dimensional convolution operation between a time-domain spatial inverse filter, which is an inverse filter of the transfer function between multiple points on the reference line of the listening area and multiple speaker elements for reproducing the sound field, and the time-domain spatial domain signal. This technology can be applied, for example, to a real-time sound field generation system equipped with a linear speaker array.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure relates to a focal sound source local reproduction device and a focal sound source local reproduction method and program, and more particularly to a focal sound source local reproduction device and a focal sound source local reproduction method and program that enable the real-time movement of a focal sound source, which has been generated so that sound can be heard in a specific area, in a three-dimensional space. [Background technology]

[0002] Conventionally, a speaker array, which consists of multiple speakers, can be used to generate a focal sound source that moves in a three-dimensional space in front of or behind the speaker array. For example, methods for generating a focal sound source that moves in three-dimensional space include the focal sound source method based on wavefront synthesis, the PM (Pressure Matching) method, and the SD (Spectral Division) method. The PM method is a technique for determining the speaker drive signal of the reproduced sound field so as to reproduce the sound pressure at multiple discretized positions in the target sound field. The SD method is a technique for setting a boundary line in the target sound field and determining the speaker drive signal in the spatial Fourier transform domain so that the spatial Fourier transform (wavenumber domain) coefficients of the continuous sound pressure on that boundary line match the spatial Fourier transform (wavenumber domain) coefficients of the reproduced sound pressure on the boundary line set in the reproduced sound field.

[0003] Patent Document 1 discloses a technique for reproducing a sound field formed by a moving point source without introducing temporal aliasing when generating a drive signal for a speaker array using the SD method. [Prior art documents] [Patent Documents]

[0004] [Patent Document 1] Japanese Patent Publication No. 2022-122414 [Overview of the project] [Problems that the invention aims to solve]

[0005] By applying the local playback method, which reproduces sound locally so that it can be heard within a specific area, to the method of generating a focal sound source that moves in three-dimensional space as described above, it is thought that the focal sound source, which has been generated so that the sound can be heard within a specific area, can be moved in three-dimensional space in real time.

[0006] However, the wavefront synthesis method does not define the reproduced sound field; rather, it automatically synthesizes wavefronts as it reproduces the sound pressure gradient on the boundary line where the speaker array is placed. Therefore, even if the local reproduction method is applied to the focal sound source method based on wavefront synthesis, there was a problem in that the original sound field had to be a locally reproduced focal sound source.

[0007] Furthermore, when applying the local regeneration method to the PM method, the computational load required to move the focal sound source, which is generated so that sound is audible within a specific area, becomes extremely large, posing a challenge in achieving real-time movement of such a focal sound source. Similarly, when applying the local regeneration method to the SD method, the calculations are performed in the spatial Fourier transform domain, resulting in a large computational load required for convolution in the wavenumber domain. Consequently, the computations are not fast enough to achieve real-time movement of the focal sound source, which is generated so that sound is audible within a specific area.

[0008] Thus, even when applying the local reproduction method to conventional methods for generating focal sound sources that move in three-dimensional space, it is considered difficult to move the focal sound source, which has been generated so that the sound can be heard within a specific area, in three-dimensional space in real time.

[0009] This disclosure is made in light of these circumstances and enables the real-time movement of a focused sound source, generated so that sound is audible within a specific area, in three-dimensional space. [Means for solving the problem]

[0010] A focal sound source local reproduction device according to one aspect of the present disclosure comprises: a target sound field calculation unit that calculates a time-domain spatial domain signal on the reference line representing a target sound field in which the focal sound source exists within the spatial range of the spatial window, by applying a spatial window having a specified spatial range to a signal representing a focal sound source sound field, which is a sound field on a preset reference line of a listening area in which a focal sound source is generated at a specified focal sound source position; a time-domain spatial inverse filter which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field; and a speaker drive signal generation unit that generates a time-domain speaker drive signal for reproducing the focal sound source sound field by performing a two-dimensional convolution operation with the time-domain spatial domain signal.

[0011] A focal sound source local reproduction method or program according to one aspect of the present disclosure includes: calculating a time-domain spatial domain signal on the reference line representing a target sound field in which the focal sound source exists within the spatial range of the spatial window, by applying a spatial window having a specified spatial range to a signal representing a focal sound source sound field, which is a sound field on a preset reference line of a listening area in which a focal sound source is generated at a specified focal sound source position; and generating a time-domain speaker drive signal for reproducing the focal sound source sound field by performing a two-dimensional convolution operation between a time-domain spatial inverse filter, which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field, and the time-domain spatial domain signal.

[0012] In one aspect of this disclosure, a sound field in which a focal sound source is generated at a specified focal sound source position is used, and a spatial window having a specified spatial range is applied to a signal representing the focal sound source sound field, which is a sound field on a preset reference line of a listening area. By doing so, a time-domain spatial signal on the reference line representing the target sound field in which the focal sound source sound field exists within the spatial range of the spatial window is calculated. A time-domain speaker drive signal for reproducing the focal sound source sound field is generated by performing a two-dimensional convolution operation between the time-domain spatial inverse filter, which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field, and the time-domain spatial signal. [Effects of the Invention]

[0013] According to one aspect of this disclosure, a focused sound source, generated so that sound can be heard within a specific area, can be moved in real time in a three-dimensional space.

[0014] The effects described herein are not necessarily limited to those described herein and may include any of the effects described herein. [Brief explanation of the drawing]

[0015] [Figure 1] This block diagram shows an example configuration of a real-time sound field generation system that applies this technology. [Figure 2] This diagram illustrates the target sound field generated by the real-time sound field generation system. [Figure 3] This diagram illustrates an example configuration of the target sound field calculation unit. [Figure 4] This diagram illustrates an example of the configuration of the pre-calculation unit for the inverse transfer function filter. [Figure 5] This diagram illustrates an example configuration of the speaker drive signal generation unit. [Figure 6] This is a flowchart explaining the focus source control process. [Figure 7] This diagram illustrates a second example configuration of the target sound field calculation unit. [Figure 8] This diagram illustrates a second example configuration of the speaker drive signal generation unit. [Figure 9] This diagram illustrates a third configuration example of a focus sound source control device. [Figure 10] This is a block diagram showing an example configuration of one embodiment of a computer to which this technology is applied. [Modes for carrying out the invention]

[0016] The following describes in detail a specific embodiment of this technology, with reference to the drawings.

[0017] <Example of a real-time sound field generation system configuration> Figure 1 is a block diagram showing an example configuration of one embodiment of a real-time sound field generation system to which this technology is applied.

[0018] As shown in Figure 1, the real-time sound field generation system 11 comprises a sound source 21, a focus sound source control device 22, a linear speaker array 23, and a focus sound source position designation device 24. The linear speaker array 23 is composed of a plurality of speaker elements 31 (s speaker elements 31-1 to 31-s in the illustrated example) arranged in a straight line.

[0019] The real-time sound field generation system 11 supplies the original signal output from the sound source 21 to the linear speaker array 23 via the focal sound source control device 22, thereby generating a sound field (hereinafter referred to as the target sound field) that moves a focal sound source, generated so that sound can be heard in a specific area, in real time in three-dimensional space. Hereinafter, the predetermined area in which sound transmitted from the focal sound source generated by the real-time sound field generation system 11 can be heard will be referred to as the listening area.

[0020] For example, the focus sound source control device 22 generates the position of the focus sound source (x) using the linear speaker array 23. FS ,y FS) can be interactively controlled in accordance with the listener's movement so as to follow the tracking position obtained by the focal sound source position specifying device 24 tracking the listener. In the configuration example shown in Fig. 1, as the focal sound source position specifying device 24, the movement of the listener is sequentially tracked, and the position of the focal sound source (x FS ,y FS ) is calculated by a tracking device. Note that the focal sound source position specifying device 24 is not limited to a tracking device, and may be, for example, a device that holds focal sound source position information for each preset time and transmits the focal sound source position information to the focal sound source control device 22 as time elapses.

[0021] Furthermore, the focal sound source control device 22 can generate a focal sound source that transmits sound within a fan-shaped listening area defined by two reproduction area boundary lines (the dot-dash straight lines shown in Fig. 1) intersecting at the position (x FS ,y FS ) of the focal sound source. For example, with the direction along the linear speaker array 23 defined as the x direction and the direction orthogonal to the x direction defined as the y direction, the two reproduction area boundary lines are provided so as to be inclined at a predetermined angle (about 45 degrees in the example shown in Fig. 1) with respect to the x direction. The fan-shaped listening area extends from the position (x FS ,y FS ) of the focal sound source toward the positive side and negative side in the y direction (note that in an ideal sound field, the negative side in the y direction is also a listening area, whereas in a reproduced sound field, time is reversed on the negative side in the y direction, so the negative side in the y direction does not actually serve as a listening area), and is defined as being surrounded by an arc centered on the position (x FS ,y FS ) of the focal sound source and the two reproduction area boundary lines.

[0022] Therefore, as shown in Figure 1, the real-time sound field generation system 11 tracks the position of person A, who will be the listener, using the focal sound source position designation device 24, and generates a target sound field in which the focal sound source moves in accordance with the movement of person A, and in which the sound transmitted from the focal sound source is heard only by person A within the listening area. In other words, even if person A moves, the real-time sound field generation system 11 can make person A hear the sound reproduced locally within the listening area by the focal sound source, while generating a target sound field in which person B, who is located far away from person A, cannot hear the sound.

[0023] Referring to Figure 2, the target sound field generated by the real-time sound field generation system 11 will be described.

[0024] For example, conventionally, as shown on the left side of Figure 2, the sound field P of the focal sound source. FS The sound is generated so that it propagates in all directions from the focal sound source position (the inverted triangle mark shown on the left side of Figure 2).

[0025] Furthermore, conventionally, as shown in the center of Figure 2, the sound field P of local reproduction PSZ The sound is generated so that it travels within a listening area that follows a predetermined width (the width of the rectangular mark shown in the center of Figure 2).

[0026] Then, as shown on the right side of Figure 2, the target sound field P generated by the real-time sound field generation system 11 is generated. PSZ-FS The sound field P of the focal sound source FS and the sound field P of local reproduction PSZ These are synthesized and generated so that the sound travels within the listening area of ​​a sector-shaped region defined by two reproduction area boundaries that intersect at the focal sound source position (the inverted triangle mark shown on the right side of Figure 2). However, in the example shown on the right side of Figure 2, the listening area appears to be narrowed because only the speaker element 31 near the center of the multiple speaker elements 31 that make up the linear speaker array 23 are vibrating. However, simply vibrating only the speaker element 31 near the center would degrade the reproduction accuracy of the focal sound source.

[0027] In the following, the SD method is used to generate a focal sound source moving in three-dimensional space, and the target sound field P is described below. PSZ-FS This section describes a method for generating [the specified value].

[0028] First, in order to synthesize a sound field in which the sound reproduced by the focal sound source is heard only in the listening region, we consider generating a focal sound source such that the wavefront of the sound reproduced by the focal sound source is zero outside the listening region. For example, in a two-dimensional plane, the wavefront of the sound reproduced by the focal sound source is the position of the focal sound source (x FS ,y FS The sound will spread out in a circular pattern centered on (x). Therefore, to limit this spread, the position of the focal sound source (x FS ,y FS A focused sound source is generated such that the sound wavefront spreads only to the sector-shaped area demarcated by two intersecting reproduction region boundaries (see Figure 1).

[0029] In the area reproduction method using the SD method, only the sound pressure along a reference line parallel to the x-axis is observed. Therefore, the observed sound pressure is the sound field P generated by the focal sound source. FS (x,y ref It is represented by multiplying (ω) by a window function. Therefore, the target sound field P PSZ-FS (x,y ref ,ω) is the position of the focal sound source (x FS ,y FS Taking ) as the origin and using the rectangular window W(x-x0) as the window function, it can be expressed as shown in equation (1) below.

[0030]

number

[0031] However, in equation (1), x is the x-axis position of each speaker element 31, and the x-axis coordinate on the reference line of the target sound field is treated as the same. Also, x0 is the center position in the x-axis direction of the listening area, and y ref ω is the position in the y-axis direction of the reference line of the target sound field, ω is the angular frequency of the original signal supplied from the sound source 21, and H0(2) (·) is the Hankel function of the second kind, and L is the width of the listening region.

[0032] Then, the spatial domain signal represented by position x is transformed by spatial Fourier transform, resulting in wavenumber k x It is also possible to perform various calculations on a wavenumber spectral signal that has as a variable. Therefore, in the SD method, the target sound field P, represented by equation (1) above, PSZ-FS (x,y ref The wavenumber spectrum P of the target sound field is the spatial Fourier coefficient of ω. ~ PSZ-FS (k x ,y,ω) is the position of the focal sound source (x FS ,y FS Using ), and based on the convolution theorem and the shift theorem, it can be obtained as shown in equation (2) below. However, in equation (2), the wavenumber spectrum W of the window function ~ (k x The wavenumber spectrum P of the sound field is as shown in the following equation (3). ~ (k x ,y ref ω) is as shown in equation (4) below.

[0033]

number

number

number

[0034] To obtain the target sound field P(ω), it is actually necessary to multiply it by the Fourier coefficients S(ω) of the original signal, but this is omitted to avoid a complicated equation. This is equivalent to describing the case where the Fourier coefficients S(ω)=1 of the original signal, and generality is not lost when the Fourier coefficients S(ω)≠1 of the original signal by interpreting the target sound field P(ω) as a speaker drive filter. Similarly, when obtaining the target sound field p(t), the Fourier coefficients S(t) of the original signal are omitted, but by first convolving with the Fourier coefficients S(t) of the original signal, the speaker drive signal d(t) corresponding to the Fourier coefficients S(t) of the original signal can be obtained.

[0035] Here, in the SD method, the wavenumber spectrum D of the speaker drive signal obtained by spatially transforming the speaker drive signal D(x,ω) at position (x,0) is... ~ (k x ,ω) is the sound field P of the focus sound source. FS Wavenumber spectrum P ~ (k x ,y ref ω), and the wavenumber spectrum G of the transfer function in 3-dimensional free space. ~ (k x ,y ref Using ω, it can be expressed as shown in equation (5) below.

[0036]

number

[0037] Then, by substituting the above-mentioned equation (2) into equation (5), the target sound field P can be calculated. PSZ-FS (x,y ref Wavenumber spectrum D of speaker drive signal for generating ω ~ (k x ω) can be found as shown in equation (6) below.

[0038]

number

[0039] Therefore, by performing the inverse space Fourier transform on equation (6), the target sound field P PSZ-FS (x,y ref A speaker drive signal D(x,ω) that can generate ω can be obtained.

[0040] By the way, when using the SD method to generate a focal sound source moving in 3D space, it is necessary to perform an inverse space and inverse time Fourier transform each time the focal sound source moves, therefore, the target sound field P PSZ-FS (x,y ref It is difficult to keep up with the calculation process to determine the speaker drive signal D(x,ω) for generating the target sound field P, for example, in real time to track the movement of the listener. PSZ-FS (x,y ref By dividing the speaker drive signal D(x,ω) used to generate (x,ω) into a time-varying part and a time-invariant part, and efficiently performing calculations, it is possible to achieve real-time movement of the focused sound source, which is generated so that the sound is heard only within the listening area.

[0041] For example, the wave number spectrum G of a transfer function in 3-dimensional free space. ~ (k x ,y ref ω) is the Hankel function of the second kind H0 (2) Using (·), it can be expressed as shown in equation (7) below.

[0042]

number

[0043] Then, by substituting equation (7) into equation (5) above, the target sound field P PSZ-FS (x,y ref Wavenumber spectrum D of speaker drive signal for generating ω ~ (k x ω) can be found as shown in equation (8) below.

[0044]

number

[0045] Therefore, by performing the inverse space-inverse time Fourier transform on equation (8), we can obtain a time-domain-space domain speaker drive signal d(x,t) as shown in equation (9). In equation (9), ** represents a two-dimensional convolution, and IFT2 represents the inverse space-inverse time Fourier transform.

[0046]

number

[0047] Next, in order to efficiently calculate equation (9), we will consider an analytical solution. The target sound field p(x,y) used in equation (9) ref ,t) is the target sound field P represented by equation (1) above. PSZ-FS (x,y ref This is the inverse time Fourier transform of (ω). Also, since the window function W(x) is frequency independent, the sound field P of the focal source is independent. FS (x,y ref Let's consider ω.

[0048] First, approximating the Hankel function when the argument is sufficiently large, we get the sound field P of the focal source. FS (x,y ref ω) can be expressed as shown in the following equation (10).

[0049]

number

[0050] For this equation (10), the position of the focal sound source (x FS ,y FS The time variable depends on the position of the focal sound source (x FS ,y FS By separating the time-invariant terms (which do not depend on the time domain) and applying the inverse time Fourier transform, the sound field p of the focal sound source in the time domain and spatial domain can be determined. FS (x,y reft) can be found as shown in equation (11) below, where ITFT represents the inverse time Fourier transform, * represents the one-dimensional convolution, and δ(·) represents the Dirac delta function.

[0051]

number

[0052] Note that r used in equation (11) is calculated as shown in equation (12) below, from the position x in the x-axis direction on the reference line to the position (x) of the focal sound source. FS ,y FS This is the distance to ).

[0053]

number

[0054] Then, by multiplying equation (11) by the window function W(x), we obtain the target sound field p in the time domain and space domain. PSZ-FS (x,y ref Q(t) can be expressed using a time-invariant filter Q(t) as shown in equation (13) below.

[0055]

number

[0056] Here, the time-invariant filter Q(t) does not depend on the position of the speaker elements 31 that make up the linear speaker array 23, so it can be numerically pre-calculated and convolved into the original signal output from the sound source 21.

[0057] Furthermore, the delta function δ in equation (13) represents the delay of the signal. The delay amount is the position of the focal sound source (x FS ,y FSvaries depending on ]) and is generally a non-integer, so calculation can be performed using a first-order Thiran all-pass filter represented by the following equation (14). However, the coefficient a used in equation (14) is expressed using the delay amount D as shown in the following equation (15).

[0058]

Math

Math

[0059] Then, the second IFT term in the above-mentioned equation (9) is the wavenumber spectrum G of the transfer function ~ (k x ,y ref , ω) represents the result of inverse spatial-inverse time Fourier transform of the reciprocal of . The second IFT term requires solving an integral over an infinite interval, making it difficult to obtain an analytical solution. On the other hand, the transfer function G(x,y,ω) is a time-invariant transfer function between each speaker element 31 of the linear speaker array 23 and discrete points on a preset reference line, so the inverse spatial-inverse time Fourier transform can be pre-calculated numerically.

[0060] According to the above method, the target sound field P that can move in real time in three-dimensional space a focused sound source generated to allow sound to be heard within the listening area PSZ-FS (x,y ref , ω) can be generated.

[0061] To implement such a method, a first configuration example of the focused sound source control device 22 can be configured to include a target sound field calculation unit 41, a transfer function inverse filter pre-calculation unit 42, and a speaker drive signal generation unit 43, as shown in FIG. 1.

[0062] The target sound field calculation unit 41, as will be described later with reference to Figure 3, obtains a time-domain spatial domain signal representing the focal sound source field, which is the sound field in which a focal sound source is generated at the focal sound source position specified by the focal sound source position designation device 24, and then applies a spatial window having a specified spatial range to that focal sound source field to obtain a time-domain spatial domain signal representing the target sound field in which the focal sound source field exists within the spatial range of that spatial window.

[0063] The transfer function inverse filter pre-calculation unit 42, as will be described later with reference to Figure 4, pre-calculates the coefficients of the two-dimensional Fourier transform transfer function inverse filter, which is the inverse filter of the transfer function between discrete points on a pre-set reference line of the listening area and each speaker element 31 of the linear speaker array 23 for reproducing the sound field.

[0064] The speaker drive signal generation unit 43, as will be described later with reference to Figure 5, combines a time-domain spatial domain signal representing the target sound field of the previous frame with a time-domain spatial domain signal representing the target sound field of the current frame, applies a two-dimensional Fourier transform to the signal obtained by this combination to obtain two-dimensional Fourier transform sound field coefficients, and then applies a two-dimensional inverse Fourier transform to the result of taking the product of these two-dimensional Fourier transform sound field coefficients and the two-dimensional Fourier transform inverse filter coefficients element by element to generate a time-domain speaker drive signal.

[0065] Referring to Figure 3, an example of the configuration of the target sound field calculation unit 41 will be described.

[0066] As shown in Figure 3, the target sound field calculation unit 41 is configured to include a signal pre-calculation unit 51, a delay processing unit 52, and a time-space signal generation unit 53.

[0067] The signal pre-calculation unit 51 numerically pre-calculates the time-invariant filter Q(t) as described above with reference to equation (13), and convolves the time-invariant filter Q(t) onto the original signal output from the sound source 21. The signal pre-calculation unit 51 then supplies the signal obtained by convolving the time-invariant filter Q(t) onto the original signal to the delay processing unit 52.

[0068] The delay processing unit 52 uses a signal obtained by convolving the time-invariant filter Q(t) supplied from the signal pre-calculation unit 51 with the original signal to obtain the position (x FS ,y FS ), the sound field p of the focal sound source in the time domain and spatial domain is obtained by calculating the above-mentioned formula (11) FS (x,y ref ,t) is obtained and supplied to the time-space signal generation unit 53. At this time, the delay processing unit 52 obtains the sound field p of the focal sound source in the time domain and spatial domain FS (x,y ref ,t), according to the above-mentioned formula (12), as shown in FIG. 3, the distance r x (x=1 to 3S), 1 / √r shown in the following formula (16) x δ(t-r x / c) is calculated for each distance. This δ(t-r x / c) is a time-varying term that depends on the position (x FS ,y FS ) of the focal sound source, and can be obtained by calculating the above-mentioned formula (14) and formula (15).

[0069]

Math

[0070] Note that the position (x FS ,y FS ) of the focal sound source is a position specified by the focal sound source position specifying device 24. For example, focal sound source position information that is stored in advance as position information that moves over time can be read out and used. Alternatively, as described with reference to FIG. 1, by performing motion capture on a specific part of the listener's body (e.g., the head), the position coordinates of the specific part may be read in real time and used.

[0071] The time-space signal generation unit 53 receives the sound field p of the focal sound source in the time domain and spatial domain supplied from the delay processing unit 52 FS (x,y refFor t), three times or more spatial points are prepared on the reference line equal to the number S of speaker elements 31 constituting the linear speaker array 23, and the sound fields p of the time-domain and spatial-domain focal sound sources are set at the S central points. FS (x,y ref By inputting the calculated value of ,t) and filling the S points on the left and S points on the right with 0, a signal for 3S can be created. Alternatively, the time-space signal generation unit 53 generates the sound field p of the focal sound source in the time-domain and space-domain. FS (x,y ref For t), the signal corresponding to position x over 3S may be calculated.

[0072] At this time, the time-space signal generation unit 53 processes the window function W(x-x0) within the listening area of ​​the reference line based on equation (1) above. Here, x0, which is the center of the window function, is the position of the focal sound source specified by the focal sound source position designation device 24 (x FS ,y FS By setting it to ), that is, x0 = x FS This allows the value to change moment by moment, or to be a fixed value set in advance. Furthermore, the shape of the window function itself and the window width may also change over time. In this way, the time-space signal generation unit 53 generates the sound field p of the focal sound source in the time-domain and spatial domains. FS (x,y ref By applying a window function W(x-x0) with a specified spatial range to (t), a time-domain and spatial-domain signal of a predetermined frame size T (i.e., the target sound field in the time-domain and spatial domains where the focal sound source sound field exists within the spatial range of the spatial window) is obtained. PSZ-FS (x,y ref It is possible to generate a time-domain and spatial-domain signal (t) that represents t.

[0073] Referring to Figure 4, an example of the configuration of the transfer function inverse filter pre-calculation unit 42 will be described.

[0074] As shown in Figure 4, the transfer function inverse filter pre-calculation unit 42 is configured to include a transfer function calculation unit 61 and a two-dimensional Fourier transform unit 62.

[0075] The transfer function calculation unit 61 performs the calculation in IFT2 in equation (9) above using angular frequency ω and wavenumber k. x Each step is performed to obtain the frequency-wavenumber spectrum of the transfer function filter based on the result of the calculation, and this spectrum is supplied to the 2D Fourier transform unit 62.

[0076] Here, the angular frequency ω to be calculated is the sampling frequency F. s Using this, the range is given by equation (17) below. Alternatively, it is possible to calculate only the angular frequencies ω within a predetermined range and set the values ​​for angular frequencies ω outside that range to 0.

[0077]

number

[0078] Also, the wave number k in the x-axis direction. x This is given by the following equation (18), using the spacing d of the speaker elements 31 that constitute the linear speaker array 23.

[0079]

number

[0080] The two-dimensional Fourier transform unit 62 calculates the inverse transfer function filter in the frequency-wavenumber domain based on the frequency-wavenumber spectrum of the transfer function filter supplied from the transfer function calculation unit 61. First, the two-dimensional Fourier transform unit 62 performs an inverse space Fourier transform (ISFT) on the frequency-wavenumber spectrum of the transfer function filter to calculate the inverse transfer function filter in the wavenumber domain (k x ) is transformed into the spatial domain (x). At this time, the angular frequency remains unchanged, so a new matrix is ​​created containing the values ​​obtained by inverting the angular frequency in the ω direction and applying the complex conjugate, and this matrix is ​​subjected to the inverse time frequency Fourier transform (ITFT).

[0081] Furthermore, the 2D Fourier transform unit 62 swaps the first and second halves of the result of this inverse time-frequency Fourier transform, thereby creating an inverse transfer function filter G(x,y) in the space-time domain. refThe signal obtained is t). After further padding the obtained signal with T zeros, the 2D Fourier transform unit 62 performs a 2D spatial Fourier transform (FT2). At this time, the 2D Fourier transform unit 62 performs the 2D Fourier transform so that the result of the product operation in the speaker drive signal generation unit 43 (i.e., the result of taking the product of the 2D Fourier transform sound field coefficients and the 2D Fourier transform transfer function inverse filter coefficients element by element) matches the result of the linear convolution operation.

[0082] As a result, the transfer function inverse filter pre-calculation unit 42 calculates the transfer function inverse filter IG in the frequency-wavenumber domain. ~ (k x ,y ref The values ​​of ω are calculated in advance and output as the inverse filter coefficients of the 2D Fourier transform transfer function. The reason why the values ​​already obtained in the transfer function calculation unit 61 are not used as is is that if the values ​​obtained in the frequency and wavenumber domain are used directly in the multiplication unit 73 (Figure 5), the convolution operation becomes a cyclic convolution, causing aliasing and degrading sound quality and sound field reproduction accuracy.

[0083] Referring to Figure 5, an example of the configuration of the speaker drive signal generation unit 43 will be described.

[0084] As shown in Figure 5A, the speaker drive signal generation unit 43 is configured to include a previous frame time-space signal holding unit 71, a two-dimensional Fourier transform unit 72, a multiplication unit 73, a two-dimensional inverse Fourier transform unit 74, and a speaker drive signal extraction unit 75.

[0085] The previous frame time-space signal holding unit 71 holds the time-space signal supplied from the time-space signal generation unit 53 for one frame and supplies the time-space signal from the previous frame to the two-dimensional Fourier transform unit 72.

[0086] The two-dimensional Fourier transform unit 72 combines the time-space signal of the current frame supplied from the time-space signal generation unit 53 with the time-space signal of the previous frame supplied from the previous frame time-space signal holding unit 71. As a result, as shown in Figure 5B, the two-dimensional Fourier transform unit 72 obtains a matrix signal of size 3S × 2T as the time-space domain signal.

[0087] Then, the 2D Fourier transform unit 72 performs a 2D Fourier transform on this 3S × 2T matrix signal to obtain the sound field P of the focal sound source in the time domain and spatial domain. PSZ-FS (k x The values ​​,ω) are supplied to the multiplication unit 73 as two-dimensional Fourier transform sound field coefficients.

[0088] The multiplication unit 73 performs a multiplication operation by the two-dimensional Fourier transform sound field coefficients supplied from the two-dimensional Fourier transform unit 72 and the two-dimensional Fourier transform inverse transfer function filter coefficients supplied from the inverse transfer function filter pre-calculation unit 42, for each element (frequency and wavenumber). The multiplication unit 73 then supplies the result of this calculation to the two-dimensional inverse Fourier transform unit 74.

[0089] The 2D inverse Fourier transform unit 74 performs a 2D inverse Fourier transform on the calculation result supplied from the multiplication unit 73 to obtain a speaker drive signal in the form of a matrix of size 3S × 2T, and supplies it to the speaker drive signal extraction unit 75.

[0090] The speaker drive signal extraction unit 75 extracts speaker drive signals from the 3S × 2T matrix of speaker drive signals supplied from the 2D inverse Fourier transform unit 74, as shown in Figure 5C, for the latter half of the time domain (T time points) and for S times the number of actual speaker elements 31, starting from the center in the spatial domain. The speaker drive signal extraction unit 75 then supplies these extracted speaker drive signals, i.e., the time-domain and spatial-domain speaker drive signals d(x,t), to each speaker element 31.

[0091] In the focus sound source control device 22 configured as described above, the time-space signal generation unit 53 can perform multiplication in the spatial domain to avoid convolution operations in the wavenumber domain, as a method of limiting the listening area using the window function W(x). Furthermore, in the focus sound source control device 22, the speaker drive signal generation unit 43 can also perform calculations on the speaker drive signal d(x,t) as a time-domain / spatial domain signal.

[0092] For example, to generate a sequentially moving focal sound source, a form like the time-domain representation of wavefront synthesis is suitable. Also, when applying a spatial window to a spherical wave, the operation becomes a convolution in the wavenumber domain, so it is better to perform it in space. In other words, the calculation of the transfer function in the field becomes a convolution in the spatial domain, so the amount of computation can be reduced by performing it in the wavenumber domain.

[0093] Therefore, the real-time sound field generation system 11 can reduce the amount of computation by applying a window function in a different region than the conventional SD method, and can achieve a computation speed that enables real-time movement of the focal sound source generated so that the sound is audible within the listening area.

[0094] <Example of focus source control processing> Referring to the flowchart shown in Figure 6, an example of the focus source control process performed by the focus source control device 22 will be described.

[0095] In step S11, the transfer function inverse filter pre-calculation unit 42 calculates the coefficients of the two-dimensional Fourier transform transfer function inverse filter, which is the inverse filter of the transfer function between discrete points on a pre-set reference line of the listening area and each speaker element 31 of the linear speaker array 23 for reproducing the sound field, prior to the processing in steps S12 and S13. The transfer function inverse filter pre-calculation unit 42 then supplies the coefficients of the two-dimensional Fourier transform transfer function inverse filter to the speaker drive signal generation unit 43.

[0096] In step S12, the target sound field calculation unit 41 obtains a time-domain spatial domain signal representing the focal sound source field, which is the sound field in which a focal sound source is generated at the focal sound source position specified by the focal sound source position designation device 24. Furthermore, the target sound field calculation unit 41 applies a spatial window having a specified spatial range to the focal sound source field, and obtains a time-domain spatial domain signal representing the target sound field in which the focal sound source field exists within the spatial range of that spatial window. The target sound field calculation unit 41 then supplies the time-domain spatial domain signal representing the target sound field to the speaker drive signal generation unit 43.

[0097] In step S13, the speaker drive signal generation unit 43 combines the time-domain spatial domain signal representing the target sound field supplied from the speaker drive signal generation unit 43 in step S12 of the previous frame with the time-domain spatial domain signal representing the target sound field supplied from the speaker drive signal generation unit 43 in step S12 of the current frame. The speaker drive signal generation unit 43 then applies a two-dimensional Fourier transform to the signal obtained by this combination to obtain the two-dimensional Fourier transform sound field coefficients. Furthermore, the speaker drive signal generation unit 43 generates a time-domain speaker drive signal by applying a two-dimensional inverse Fourier transform to the result of taking the product of these two-dimensional Fourier transform sound field coefficients and the two-dimensional Fourier transform inverse filter coefficients supplied from the transfer function inverse filter pre-calculation unit 42 in step S11, element by element. This signal is then supplied to the linear speaker array 23.

[0098] Subsequently, the process returns to step S12, and the next focal sound source position specified by the focal sound source position designation device 24 is used as the processing target, and the same process is repeated thereafter.

[0099] By having the focus sound source control device 22 perform the focus sound source control processing described above, the real-time sound field generation system 11 can move the focus sound source, which has been generated so that the sound can be heard within the listening area, in three-dimensional space in real time.

[0100] For example, the real-time sound field generation system 11 can be used in use cases such as event venues to provide a user experience that instills fear only in specific listeners by using a ghost's voice or similar as the focus sound source.

[0101] <Second Configuration Example of Focus Sound Source Control Device> A second configuration example of the focus sound source control device 22 will be described with reference to Figures 7 and 8.

[0102] For example, the reciprocal space-time Fourier transform IFT2 used in equation (9) above can be calculated independently by first performing the reciprocal space Fourier transform ISFT and then the reciprocal time Fourier transform ITFT (i.e., IFT2[]=ITFT[ISFT[]]), and the reciprocal space Fourier transform ISFT can be solved as shown in equation (19).

[0103]

number

[0104] However, in equation (19), the first transformation is an approximation for the case where the argument of the Hankel function is large, and the second transformation is a stationary phase approximation. Therefore, based on equation (19), the above-mentioned equation (9) becomes as shown in the following equation (20).

[0105]

number

[0106] Then, by truncating the integral in equation (20) to a finite length and discretizing it, equation (20) becomes as shown in equation (21). However, the distance r used in equation (20) ref,m This refers to the x-coordinate of the m-th (m=1~M) reference point, which is placed at equal intervals on the reference line that serves as the listening position. ref,mUsing this, it is expressed as shown in the following equation (22). Note that, unlike the focus sound source control device 22 in the first configuration example described above, the multiple reference points arranged at equal intervals on the reference line can be set independently of the position of the speaker element 31.

[0107]

number

number

[0108] Furthermore, the inverse time Fourier transform ITFT[k] used in equation (21), together with the time-invariant filter Q(t) in equation (11) above, can be convolved into the original signal as a new time-invariant filter Q'(t), as shown in equation (23).

[0109]

number

[0110] Furthermore, in order to realize a method for acquiring the time-domain and spatial-domain speaker drive signal d(x,t) determined according to equation (22), the second configuration example, the focus sound source control device 22A, can be configured to include a target sound field calculation unit 41A and a speaker drive signal generation unit 43A, as shown in Figure 7.

[0111] As shown in Figure 7, the target sound field calculation unit 41A is configured to include a signal pre-calculation unit 51A, a delay processing unit 52A, and a target sound field signal generation unit 54.

[0112] The signal pre-calculation unit 51A numerically pre-calculates a time-invariant filter Q'(t) as shown in equation (23) above, and convolves the time-invariant filter Q'(t) onto the original signal output from the sound source 21. For example, this time-invariant filter Q'(t) has characteristics that are independent of time, similar to the time-invariant filter Q(t) in equation (11), as well as characteristics that are independent of the position of the speaker element 31 extracted in advance from the target sound field calculation unit 41A and the speaker drive signal generation unit 43A. The signal pre-calculation unit 51 then supplies the signal with the time-invariant filter Q'(t) convolved onto the original signal to the delay processing unit 52A.

[0113] The delay processing unit 52A calculates the distance r from the m-th reference point to the focal sound source for the signal obtained by convolving the original signal with the time-invariant filter Q'(t) supplied from the signal pre-calculation unit 51A. FS,m Delay δ(tr) FS,m / c) is given, and the target sound field p of the focal sound source. FS (x ref,m ,y ref The distance r from the m-th reference point to the focal sound source is calculated and supplied to the target sound field signal generation unit 54. FS,m The position of the reference point (x ref,m ,y ref ) and the position of the focal sound source (x FS ,y FS Using ), it can be calculated as shown in equation (24) below.

[0114]

number

[0115] The target sound field signal generation unit 54 generates the target sound field p of the focal sound source supplied from the signal pre-calculation unit 51A. FS (x ref,m ,y ref For t), the x-coordinate position of the m-th reference point, which is placed at equal intervals on the baseline, is x ref,m The window function W(x) follows ref,m By applying this process, the target sound field of the locally reproduced focal sound source is calculated. For example, the position x coordinate of the m-th reference point ref,mThis is the number of reference points M (=α × L) and the interval d between reference points. ref (=d ls Using / β), it can be calculated as shown in equation (25) below.

[0116]

number

[0117] As a result, the target sound field signal generation unit 54 generates a window function W(x ref,m The target sound field p in the time-domain and spatial domains exists within the spatial range of the spatial window represented by ). PSZ-FS (x ref,m ,y ref It can generate a time-domain and spatial-domain signal representing t). Furthermore, the target sound field signal generation unit 54 generates the target sound field p in the time-domain and spatial domain over time t to T-1. PSZ-FS (x ref,m ,y ref By calculating t), a time-domain and spatial-domain signal is generated such that the frame size for one frame is T × M, and this signal is supplied to the speaker drive signal generation unit 43A.

[0118] As shown in Figure 8, the speaker drive signal generation unit 43A is configured to include a previous frame time-space signal holding unit 81, a concatenation processing unit 82, a delay processing unit 83, a gain processing unit 84, an addition unit 85, and a correction processing unit 86. In the following description, the process of sequentially calculating the speaker drive signal for the l-th channel (l=1 to L) will be explained, and this process will be repeated for L channels.

[0119] The previous frame time-space signal holding unit 81 holds the time-domain and spatial domain signals supplied from the target sound field signal generation unit 54 for one frame (frame size: T × M) and supplies the time-domain and spatial domain signals from the previous frame to the concatenation processing unit 82.

[0120] The concatenation processing unit 82 concatenates the time-domain and spatial domain signals of the current frame supplied from the target sound field signal generation unit 54 with the time-domain and spatial domain signals of the previous frame supplied from the previous frame time-domain and spatial domain signal holding unit 71 to obtain a time-domain and spatial domain signal with a frame size of 2T × M, and supplies it to the delay processing unit 83.

[0121] The delay processing unit 83 calculates the time-domain and spatial-domain signals supplied from the coupling processing unit 82 by calculating the distance r from the l-th speaker element 31 to the m-th reference point for each of the multiple reference points that are equally spaced on the reference line. l,m The corresponding delay δ(t+r l,m A delay process is applied that gives / c). As a result, the delay processing unit 83 acquires the time-domain and spatial-domain signals for each reference point that have undergone the delay process and supplies them to the gain processing unit 84. Here, the distance r from the l-th speaker element 31 to the m-th reference point is... l,m The position of the reference point (x ref,m ,y ref ) and the position of the focal sound source (x FS ,y FS Using ), it can be calculated as shown in equation (26) below.

[0122]

number

[0123] The gain processing unit 84 applies a gain A to each of the multiple reference points, which are equally spaced on the reference line, with respect to the time-domain and spatial-domain signals for each reference point that have been delayed and supplied by the delay processing unit 83, as shown in equation (27) below. l,m That is, the distance r from the l-th speaker element 31 to the m-th reference point. l,m Gain A can be calculated using l,m Gain processing is applied to give the desired result. As a result, the gain processing unit 84 acquires the time-domain and spatial-domain signals for each reference point that have undergone delay processing and gain processing, and supplies them to the adder 85.

[0124]

number

[0125] The summing unit 85 calculates the sum of the time-domain and spatial-domain signals for each reference point that have undergone delay processing and gain processing, supplied by the gain processing unit 84. As a result, the summing unit 85 obtains a signal representing the sum of the time-domain and spatial-domain signals for each reference point that have undergone delay processing and gain processing, and supplies it to the correction processing unit 86.

[0126] The correction processing unit 86 applies a correction term g shown in the following equation (28) to the signal representing the sum of the time-domain and spatial-domain signals for each reference point that have been subjected to delay processing and gain processing, supplied from the adder 85. l That is, the number of reference points M and the y coordinates y of the reference points ref Correction term g is obtained using l A correction process is applied by multiplying by x. As a result, the correction processing unit 86 determines the x-coordinate x ls、l The time-domain and spatial-domain speaker drive signal d(x) of the l-th speaker element 31 located at the speaker position. ls、l The signal (t) is obtained and supplied to the l-th speaker element 31.

[0127]

number

[0128] For example, the x-coordinate of the l-th speaker element 31 is x ls、l This refers to the number L of speaker elements 31 and the spacing d between speaker elements 31. ls Using this, it can be obtained as shown in equation (29) below.

[0129]

number

[0130] By configuring the focus sound source control device 22A as described above, the real-time sound field generation system 11 can move the focus sound source, which has been generated so that the sound can be heard within the listening area, in three-dimensional space in real time.

[0131] <Third Configuration Example of Focus Sound Source Control Device> Referring to Figure 9, a third configuration example of the focus sound source control device 22 will be described.

[0132] For example, the target sound field p in the time domain and space domain obtained by equation (13) above. PSZ-FS (x,y ref Substituting ,t) into equation (21) above, we obtain the following equation (30). However, the distance r from the m-th reference point to the focal sound source used in equation (30) is... FS,m This can be calculated as shown in equation (24) above.

[0133]

number

[0134] Furthermore, in order to realize a method for acquiring a speaker drive signal d(x,t) in the time domain and spatial domain as shown in equation (30), the third configuration example, the focus sound source control device 22B, can be configured as shown in Figure 9.

[0135] As shown in Figure 9, the focus sound source control device 22B is configured to include a signal pre-calculation unit 51B, a window function processing unit 55, a delay processing unit 83B, a gain processing unit 84B, an adder 85, and a correction processing unit 86. In the following description, the process of sequentially calculating the speaker drive signal for the l-th channel (l=1 to L) will be explained, and this process will be repeated for L channels.

[0136] The signal pre-calculation unit 51B numerically pre-calculates a time-invariant filter Q'(t) as shown in equation (23) above, and convolves the original signal output from the sound source 21 with the time-invariant filter Q'(t). The signal pre-calculation unit 51 then supplies the signal obtained by convolving the time-invariant filter Q'(t) onto the original signal to the window function processing unit 55.

[0137] The window function processing unit 55 calculates the x-coordinate position of the m-th reference point, which is placed at equal intervals on the reference line, for the signal obtained by convolving the original signal with the time-invariant filter Q'(t) supplied from the signal pre-calculation unit 51B. ref,m The window function W(x) follows ref,m The window function processing unit 55 applies the window function W(x) to the signal obtained by convolving the original signal with the time-invariant filter Q'(t) for each of the multiple reference points that are evenly spaced on the reference line. ref,m The window function processing unit 55 acquires the signal that has been subjected to the window function W(x) and supplies it to the delay processing unit 83B. ref,m If the spatial range is outside the spatial window represented by ), the process is skipped.

[0138] The delay processing unit 83B applies a window function W(x) to the signal obtained by convolving the original signal with the time-invariant filter Q'(t) supplied from the window function processing unit 55. ref,m For a signal that has been subjected to ), the distance r from the l-th speaker element 31 to the m-th reference point is calculated for each of the multiple reference points that are equally spaced on the reference line. l,m , and the distance r from the m-th reference point to the focal sound source. FS,m The corresponding delay δ(t+r l,m / cr FS,m A delay process is applied that gives / c). As a result, the delay processing unit 83B acquires the time-domain and spatial-domain signals for each reference point that have been delayed and supplies them to the gain processing unit 84B. Here, the distance r from the m-th reference point to the focal sound source is FS,m This is calculated as shown in equation (24) above, and is the distance r from the l-th speaker element 31 to the m-th reference point. l,m This can be calculated as shown in equation (26) above.

[0139] The gain processing unit 84B applies a gain B to each of the multiple reference points, which are equally spaced on the reference line, for each time-domain and spatial-domain signal supplied by the delay processing unit 83B, as shown in equation (31) below. l,m That is, the distance r from the l-th speaker element 31 to the m-th reference point. l,m , and the distance r from the m-th reference point to the focal sound source. FS,mGain B can be calculated using l,m Gain processing is applied to give the desired result. As a result, the gain processing unit 84B acquires the time-domain and spatial-domain signals for each reference point that have undergone delay processing and gain processing, and supplies them to the adder 85.

[0140]

number

[0141] The addition unit 85 and the correction processing unit 86 are configured in the same way as shown in Figure 8 above, and a detailed explanation thereof will be omitted.

[0142] By configuring the focus sound source control device 22B as described above, the real-time sound field generation system 11 can move the focus sound source, which has been generated so that the sound can be heard within the listening area, in three-dimensional space in real time.

[0143] Furthermore, the delay processing units 52A, 83, and 83B perform real-time processing that calculates for each sample (more precisely, for each sample within a frame), and can therefore be implemented, for example, using a Thiran filter. In a Thiran filter, the nth sample y(n) can be obtained using equation (14) above, and the coefficient α can be calculated according to equation (15) above.

[0144] At this time, in order to calculate the nth sample y(n), the previous sample y(n-1) is needed, so the delay processing unit 83 saves sample y(n-1) before gain is applied in the subsequent processing until sample y(n) is calculated. Also, when the sampling frequency when the delay processing unit 52A, delay processing unit 83, and delay processing unit 83B perform sampling is F, the Nyquist frequency is F / 2 and the Nyquist angular frequency is πF.

[0145] Furthermore, the target sound field calculation unit 41 and the target sound field calculation unit 41A are configured to calculate a time-domain spatial domain signal on the reference line representing the target sound field, in which the focus sound source is generated at the focus sound source position specified by the focus sound source position designation device 24, and the focus sound source is a sound field on the reference line of a preset listening area, by applying a spatial window having a specified spatial range to the signal representing the focus sound field, in which the focus sound source is located within the spatial range of the spatial window. In addition, the speaker drive signal generation unit 43 and the speaker drive signal generation unit 43A are configured to generate a time-domain speaker drive signal d for reproducing the focus sound source sound field by performing a two-dimensional convolution operation between a time-domain spatial inverse filter, which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements 31 for reproducing the sound field, and a time-domain spatial domain signal.

[0146] <Example of computer configuration> Next, the series of processes described above (the localized sound source reproduction method) can be performed by hardware or by software. When the series of processes are performed by software, the programs that make up the software are installed on a general-purpose computer or the like.

[0147] Figure 10 is a block diagram showing an example configuration of one embodiment of a computer on which the program that performs the series of processes described above is installed.

[0148] The program can be pre-recorded on the hard disk 105 or ROM 103, which are recording media built into the computer.

[0149] Alternatively, the program can be stored (recorded) on a removable recording medium 111 driven by drive 109. Such a removable recording medium 111 can be provided as so-called packaged software. Examples of removable recording media 111 include flexible disks, CD-ROMs (Compact Disc Read Only Memory), MO (Magneto Optical) disks, DVDs (Digital Versatile Discs), magnetic disks, semiconductor memory, etc.

[0150] In addition to installing the program from the removable storage medium 111 as described above, the program can also be downloaded to the computer via a communication network or broadcasting network and installed on the built-in hard disk 105. That is, the program can be transferred wirelessly to the computer from a download site via a satellite for digital satellite broadcasting, or transferred via a wired connection to the computer via a network such as a LAN (Local Area Network) or the Internet.

[0151] The computer has a built-in CPU (Central Processing Unit) 102, and an input / output interface 110 is connected to the CPU 102 via a bus 101.

[0152] When the CPU 102 receives a command from the user via the input / output interface 110, such as by operating the input unit 107, it executes a program stored in the ROM (Read Only Memory) 103 accordingly. Alternatively, the CPU 102 loads a program stored in the hard disk 105 into the RAM (Random Access Memory) 104 and executes it.

[0153] As a result, the CPU 102 performs processing according to the flowchart described above, or processing according to the configuration of the block diagram described above. The CPU 102 then outputs the processing results as needed, for example, via the input / output interface 110, from the output unit 106, or transmitted from the communication unit 108, or recorded on the hard disk 105.

[0154] The input section 107 consists of a keyboard, mouse, microphone, etc. The output section 106 consists of an LCD (Liquid Crystal Display), speakers, etc.

[0155] In this specification, the processes performed by a computer according to a program do not necessarily have to be performed chronologically in the order described in the flowchart. That is, the processes performed by a computer according to a program include processes that are executed in parallel or individually (e.g., parallel processing or object-based processing).

[0156] Furthermore, the program may be processed by a single computer (processor), or it may be processed in a distributed manner by multiple computers. Moreover, the program may be transferred to a remote computer for execution.

[0157] Furthermore, in this specification, a system means a collection of multiple components (devices, modules (parts), etc.), regardless of whether all components are located in the same enclosure or not. Therefore, multiple devices housed in separate enclosures and connected via a network, and a single device in which multiple modules are housed in one enclosure, are both considered systems.

[0158] Furthermore, for example, the configuration described as a single device (or processing unit) may be divided and configured as multiple devices (or processing units). Conversely, the configurations described above as multiple devices (or processing units) may be combined and configured as a single device (or processing unit). It is also possible to add configurations other than those described above to the configuration of each device (or each processing unit). Moreover, if the overall system configuration and operation are substantially the same, a part of the configuration of one device (or processing unit) may be included in the configuration of another device (or other processing unit).

[0159] Furthermore, for example, this technology can be configured as cloud computing, where a single function is shared and processed collaboratively by multiple devices via a network.

[0160] Furthermore, for example, the program described above can be executed on any device. In that case, the device should have the necessary functions (such as functional blocks) and be able to obtain the necessary information.

[0161] Furthermore, each step described in the flowchart above can be executed by a single device or shared among multiple devices. Additionally, if a single step includes multiple processes, these processes can be executed by a single device or shared among multiple devices. In other words, multiple processes within a single step can be executed as multiple steps. Conversely, processes described as multiple steps can be combined and executed as a single step.

[0162] Furthermore, the program executed by the computer may be executed in a chronological order according to the sequence of steps described herein, or it may be executed in parallel or individually at necessary times, such as when a call is made. In other words, as long as no inconsistencies arise, the processing of each step may be executed in an order different from the sequence described above. Moreover, the processing of the steps of this program may be executed in parallel with the processing of other programs, or it may be executed in combination with the processing of other programs.

[0163] Furthermore, the technologies described in this specification can be implemented independently, as long as they do not create a contradiction. Of course, any multiple technologies can also be implemented in combination. For example, some or all of the technologies described in one embodiment can be combined with some or all of the technologies described in another embodiment. In addition, some or all of the above-mentioned technologies can be implemented in combination with other technologies not mentioned above.

[0164] It should be noted that this embodiment is not limited to the embodiment described above, and various modifications are possible without departing from the spirit of this disclosure. Furthermore, the effects described herein are merely illustrative and not limiting, and other effects may also exist. [Explanation of symbols]

[0165] 11 Real-time sound field generation system, 21 Sound source, 22 Focus sound source control device, 23 Linear speaker array, 24 Focus sound source position designation device, 31 Speaker element, 41 Target sound field calculation unit, 42 Transfer function inverse filter pre-calculation unit, 43 Speaker drive signal generation unit, 51 Signal pre-calculation unit, 52 Delay processing unit, 53 Time-space signal generation unit, 61 Transfer function calculation unit, 62 2D Fourier transform unit, 71 Previous frame time-space signal retention unit, 72 2D Fourier transform unit, 73 Multiplication unit, 74 2D inverse Fourier transform unit, 75 Speaker drive signal extraction unit

Claims

1. A target sound field calculation unit calculates a time-domain spatial domain signal on the reference line that represents the target sound field in which the focal sound source exists within the spatial range of the spatial window, by applying a spatial window having a specified spatial range to a signal representing the focal sound source sound field, which is a sound field in which a focal sound source is generated at a specified focal sound source position and is a sound field on the reference line of a preset listening area. A time-domain spatial inverse filter, which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field, and a speaker drive signal generation unit that generates a time-domain speaker drive signal for reproducing the focused sound source sound field by performing a two-dimensional convolution operation with the time-domain spatial signal. A localized sound source reproduction device equipped with a focal sound source.

2. Transfer function inverse filter pre-calculation unit calculates the coefficients of the two-dimensional Fourier transform transfer function inverse filter in advance, as an inverse filter for the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field. The focal sound source local reproduction device according to claim 1, further comprising:

3. The transfer function inverse filter pre-calculation unit, when calculating the coefficients of the two-dimensional Fourier transform transfer function inverse filter, first converts the coefficients obtained in the frequency domain and wavenumber domain to the time domain and spatial domain, and then performs a two-dimensional Fourier transform so that the result of the product operation in the speaker drive signal generation unit matches that of the linear convolution operation. The speaker drive signal generation unit combines the time-domain and spatial-domain signal representing the target sound field from the previous frame with the time-domain and spatial-domain signal representing the target sound field of the current frame. It then applies a two-dimensional Fourier transform to the signal obtained by this combination to obtain two-dimensional Fourier transform sound field coefficients. Finally, it applies a two-dimensional inverse Fourier transform to the result of taking the product of these two-dimensional Fourier transform sound field coefficients and the two-dimensional Fourier transform inverse filter coefficients for each element, thereby generating the time-domain speaker drive signal. The focal sound source local reproduction device according to claim 2.

4. The aforementioned target sound field calculation unit is: A signal pre-calculation unit numerically pre-calculates a time-invariant time-domain filter and convolves the original signal with the said time-invariant time-domain filter, A processing unit that uses the signal obtained by convolving the original signal with the aforementioned time-invariant time-domain filter to obtain a time-domain and spatial-domain signal representing the focal sound source field according to the specified focal sound source position, and has The focal sound source local reproduction device according to claim 1.

5. The aforementioned target sound field calculation unit is: A signal pre-calculation unit numerically pre-calculates a time-invariant time-domain filter and convolves the original signal with the said time-invariant time-domain filter, A delay processing unit calculates the target sound field of the focal sound source by applying a delay to the signal obtained by convolving the original signal with the aforementioned time-invariant time-domain filter, based on the distance from a plurality of reference points on the reference line to the focal sound source. A target sound field signal generation unit calculates the target sound field of the locally reproduced focal sound source by applying a window function to the target sound field of the focal sound source according to the x-coordinate of a predetermined reference point, and generates a time-domain and spatial-domain signal representing the target sound field in the time-domain and spatial-domain where the focal sound source sound field exists within the spatial range of the spatial window represented by the window function. It has, The aforementioned time-invariant time-domain filter includes the position of the speaker element, which is extracted in advance from the target sound field calculation unit and the speaker drive signal generation unit, and time-independent characteristics. The focal sound source local reproduction device according to claim 1.

6. The target sound field signal generation unit generates one frame's worth of time-domain and spatial-domain signals by calculating the time-domain and spatial-domain signals for a predetermined time for all the reference points. The speaker drive signal generation unit is, A delay processing unit performs a delay on a time-domain spatial domain signal of a frame size obtained by concatenating the time-domain spatial domain signal of the current frame and the time-domain spatial domain signal of the previous frame, by applying a delay to each of the multiple reference points on the reference line, according to the distance from the speaker element to be processed to a predetermined reference point. A gain processing unit performs gain processing on the time-domain and spatial-domain signals for each of the reference points that have undergone the aforementioned delay processing, by applying a gain to each of the plurality of reference points on the reference line, which is determined using the distance from the speaker element to be processed to a predetermined reference point. An adder that calculates the sum of the time-domain and spatial-domain signals for each reference point that have undergone the aforementioned delay processing and gain processing, A correction processing unit obtains the speaker drive signal to be supplied to the speaker element to be processed by applying a correction process to a signal representing the sum of the time-domain and spatial-domain signals for each reference point, which has been subjected to the aforementioned delay processing and gain processing, by multiplying it by a correction term determined using the number of reference points and the y-coordinate of each reference point. It has, Multiple speaker elements are processed sequentially, and the processing is repeated for the number of speaker elements. The focal sound source local reproduction device according to claim 5.

7. A signal pre-calculation unit numerically pre-calculates a time-invariant time-domain filter and convolves the original signal with the said time-invariant time-domain filter, A window function processing unit applies a window function to the signal obtained by convolving the original signal with the aforementioned time-invariant time-domain filter, for each of the multiple reference points on the reference line, according to the x-coordinate positions of the multiple reference points. A delay processing unit obtains a time-domain and spatial domain signal for each of the reference points to which the window function has been applied to the signal obtained by convolving the original signal with the time-invariant time-domain filter, by applying a delay to each of the multiple reference points on the reference line, which is adjusted according to the distance from the speaker element to be processed to a predetermined reference point and the distance from the predetermined reference point to the focal sound source, and then applying the delay processing to the signal obtained by each of the reference points to which the delay processing has been applied. A gain processing unit performs gain processing on the time-domain and spatial-domain signals for each reference point that have undergone the aforementioned delay processing, by applying a gain to each of the plurality of reference points on the reference line, which is determined using the distance from the speaker element to be processed to a predetermined reference point and the distance from the predetermined reference point to the focal sound source. An adder that calculates the sum of the time-domain and spatial-domain signals for each reference point that have undergone the aforementioned delay processing and gain processing, A correction processing unit obtains a speaker drive signal to be supplied to the speaker element being processed by applying a correction process to a signal representing the sum of the time-domain and spatial-domain signals for each reference point, which has been subjected to the aforementioned delay processing and gain processing, by multiplying it by a correction term determined using the number of reference points and the y-coordinate of each reference point. Equipped with, Multiple speaker elements are processed sequentially, and the processing is repeated for the number of speaker elements. Focal sound source local reproduction device.

8. The aforementioned designated focal sound source position is a position acquired in real time by motion capturing a specific part of the listener's body. The focal sound source local reproduction device according to claim 1.

9. The specified spatial range is the spatial range centered on the specified focal sound source position. The focal sound source local reproduction device according to claim 1.

10. The focal sound source local reproduction device, A sound field in which a focal sound source is generated at a specified focal sound source position, and a signal representing the focal sound source sound field, which is a sound field on a predetermined reference line of the listening area, is subjected to a spatial window having a specified spatial range, thereby calculating a time-domain spatial-domain signal on the reference line that represents the target sound field in which the focal sound source sound field exists within the spatial range of the spatial window. A time-domain speaker drive signal for reproducing the focused sound source sound field is generated by performing a two-dimensional convolution operation between a time-domain spatial inverse filter, which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field, and the time-domain spatial signal. A method for localized playback of a focal sound source, including the method described above.

11. In the computer of the focused sound source local playback device, A sound field in which a focal sound source is generated at a specified focal sound source position, and a signal representing the focal sound source sound field, which is a sound field on a predetermined reference line of the listening area, is subjected to a spatial window having a specified spatial range, thereby calculating a time-domain spatial-domain signal on the reference line that represents the target sound field in which the focal sound source sound field exists within the spatial range of the spatial window. A time-domain speaker drive signal for reproducing the focused sound source sound field is generated by performing a two-dimensional convolution operation between a time-domain spatial inverse filter, which is an inverse filter of the transfer function between a plurality of points on the reference line of the listening area and a plurality of speaker elements for reproducing the sound field, and the time-domain spatial signal. A program that performs a process that includes the following.

Citation Information

Patent Citations

  • Sound field reproduction device and program

    JP2022122414A