Sound field control device, sound field control method, and program
The sound field control system uses a speaker array with Fourier transform processors to achieve accurate multi-spot reproduction with reduced computational demands, addressing the high calculation burden of existing methods.
Patent Information
- Application Number
- JP2024021120
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-02-15
- Publication Date
- 2025-08-27
AI Technical Summary
Existing sound field control technologies require a large amount of calculation for multi-spot reproduction, which is impractical for mobile devices and other applications where computational resources are limited.
A sound field control system utilizing a speaker array with a short-time Fourier transform processor, time-frequency filter processing unit, and inverse Fourier transform processor to reduce the computational load while maintaining accurate multi-spot reproduction.
Enables highly accurate multi-spot sound reproduction with reduced computational requirements, suitable for mobile devices and other applications.
Smart Images

Figure 2025125209000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a multi-sound field control technology that simultaneously realizes different sound fields in multiple different areas. [Background technology]
[0002] Among sound field control technologies using multiple speakers, local reproduction technology, in which sound is heard only in a certain area, and multi-spot reproduction technology, which can present different sounds in different areas by overlapping multiple local reproductions, are expected to have a wide range of applications in personal audio for teleworking, guide systems at museums and world expositions, and translation systems. Unlike ultrasonic speakers, such local reproduction and multi-spot reproduction technologies control sounds in the audible range, resulting in high sound quality and low sound pressure levels, so there are no health concerns. Therefore, such local reproduction and multi-spot reproduction technologies are expected to be new methods to replace ultrasonic speakers.
[0003] Multi-spot reproduction technology can be realized by calculating a local reproduction filter set in advance so that sound reaches only each area, and then filtering the input sound signal with the local reproduction filter (see, for example, Non-Patent Document 1). For example, when using a circular speaker array to reproduce audio in four languages in four areas in a multi-spot manner, the audio in four languages can be reproduced in four areas in a multi-spot manner by overlapping the local reproduction of each area. [Prior art documents] [Non-patent literature]
[0004] [Non-Patent Document 1] 1. T. Okamoto and A. Sakaguchi, "Experimental validation of spatial Fourier transform-based multiple sound zone generation with a linear loudspeaker array," J. Acoust. Soc. Am., vol. 141, no. 3, pp. 1769-1780, Mar. 2017. Summary of the Invention [Problem to be solved by the invention]
[0005] However, the above-mentioned method (for example, the method disclosed in Non-Patent Document 1) is based on the superposition of local reproduction, so that, for example, to achieve multi-spot reproduction of N areas using L channel (L: natural number equal to or greater than 2) speakers, L×N filtering calculations are required, resulting in a problem of a large amount of calculation. When considering control on mobile terminals, etc., it is desirable to reduce the amount of calculation as much as possible, so this becomes a problem.
[0006] In view of the above, an object of the present invention is to provide a sound field control system, a sound field control device, a sound field control method, and a program that perform highly accurate multi-spot reproduction while reducing the amount of calculation. [Means for solving the problem]
[0007] In order to solve the above problem, a representative example (one aspect) of the invention disclosed in this application is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number of 2 or more), and is a sound field control device for performing multi-spot playback in which a plurality of audio signals are each played back in different divided areas included in the entire area using a speaker array consisting of a plurality of speakers, and is equipped with a short-time Fourier transform processing unit, a time-frequency filter processing unit for the entire area sound field, and a short-time Fourier inverse transform processing unit.
[0008] The short-time Fourier transform processor performs short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0009] The time-frequency filter processing unit for the full-area sound field acquires a harmonic spectrum for the full-area sound field to reproduce the full-area sound field based on the harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filter processing using the acquired harmonic spectrum for the full-area sound field to acquire a time-frequency domain drive signal for driving the speaker array.
[0010] The inverse short-time Fourier transform processor performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array. [Effects of the Invention]
[0011] According to the present invention, it is possible to realize a sound field control system, a sound field control device, a sound field control method, and a program that perform highly accurate multi-spot reproduction while reducing the amount of calculation. [Brief explanation of the drawings]
[0012] [Figure 1] 1 is a schematic configuration diagram of a sound field control system 1000 according to a first embodiment. [Figure 2] FIG. 2 is a schematic configuration diagram of a speaker array Spk_arry of the sound field control system 1000 according to the first embodiment. [Figure 3] 1 is a flowchart of a process executed by the sound field control system 1000. [Figure 4] 1 is a flowchart of a process executed by the sound field control system 1000. [Figure 5] 10A and 10B are diagrams for explaining a method for modeling sound pressure distribution in a reproduction area and a silent area using a function. [Figure 6] FIG. 1 is a schematic configuration diagram of a sound field control system 1000 according to a first modified example of the first embodiment. [Figure 7]10 is a flowchart of processing executed in a sound field control system 1000A according to a first modified example of the first embodiment. [Figure 8] 10 is a flowchart of processing executed in a sound field control system 1000A according to a first modified example of the first embodiment. [Figure 9] FIG. 10 is a schematic diagram of a sound field control system 2000 according to a second embodiment (during learning processing). [Figure 10] FIG. 10 is a schematic configuration diagram of a time-frequency filter processing unit 4B for a full-area sound field (during learning processing) of the sound field control system 2000 according to the second embodiment. [Figure 11] 10 is a flowchart of a learning process executed by a sound field control system 2000 according to a second embodiment. [Figure 12] FIG. 10 is a schematic diagram of a sound field control system 2000 according to a second embodiment (during inference processing). [Figure 13] FIG. 10 is a schematic configuration diagram of a time-frequency filter processing unit 4B for a full-area sound field (during inference processing) of the sound field control system 2000 according to the second embodiment. [Figure 14] A diagram showing the CPU bus configuration. DETAILED DESCRIPTION OF THE INVENTION
[0013] [First embodiment] The first embodiment will be described below with reference to the drawings.
[0014] <1.1: Sound field control system configuration> FIG. 1 is a schematic diagram of a sound field control system 1000 according to the first embodiment.
[0015] Fig. 2 is a schematic diagram of the speaker array Spk_arry of the sound field control system 1000 according to the first embodiment. It is assumed that the x-axis, y-axis, and z-axis are set as shown in Fig. 2. The origin of the xyz space (three-dimensional space) is set to point O in Fig. 2.
[0016] As shown in Fig. 1, the sound field control system 1000 includes a sound field control device 100 and a speaker array Spk_arry. For ease of explanation, the sound field control system 1000 will be described below as controlling sound fields in multiple zones (for example, in the case of Fig. 2, four zones: a first zone Zone1 (this zone is referred to as Q1), a second zone Zone2 (this zone is referred to as Q2), a third zone Zone3 (this zone is referred to as Q3), and a fourth zone Zone4 (this zone is referred to as Q4)) by driving a circular speaker array (a speaker array consisting of multiple speakers installed at equal intervals on the same circumference) as shown in Fig. 2. Furthermore, the speaker array Spk_arry is assumed to be composed of L (L: a natural number greater than or equal to 2) speakers Spk1 to SpkL (in the case of Figure 2, L = 16) that are equally spaced on the same circumference of a circle with a radius r0 from a center point O, as shown in Figure 2 in a plan view from above.
[0017] As shown in FIG. 1, the sound field control device 100 includes a reproduction area setting unit 1, a transfer function acquisition processing unit 2, a short-time Fourier transform processing unit 3, a time-frequency filter processing unit for the entire sound field 4, and a short-time inverse Fourier transform processing unit 5.
[0018] The reproduction area setting unit 1 sets information about a plurality of areas (reproduction areas) to be subjected to sound field control. For ease of explanation, the following description will be given assuming that four reproduction areas (four reproduction areas Zone1, Zone2, Zone3, and Zone4 in FIG. 2) are set as a plurality of areas (reproduction areas) to be subjected to sound field control. The reproduction area setting unit 1 outputs data including information about the set areas (reproduction areas) to be subjected to sound field control to the full-area sound field time-frequency filter processing unit 4 as data D_settings_zone(Q) (Q={Q1, Q2, Q3, Q4}).
[0019] The transfer function acquisition processing unit 2 acquires a transfer function of each of the L speakers constituting the speaker array Spk_arry and a radius r refThe processing unit 100 performs a process to obtain a two-dimensional cylindrical harmonic spectrum matrix of the transfer function on the circumference of the circle, and outputs data including the two-dimensional cylindrical harmonic spectrum matrix C obtained by the process as data D_C to the time-frequency filter processing unit 4 for the full-range sound field.
[0020] The short-time Fourier transform processor 3 converts the audio signal s (Qi) (t) (audio signal (time domain signal) to be played (output) in the playback domain Qi (i: natural number), t: time) is input, and the audio signal s (Qi) For the sake of convenience, in this embodiment, four audio signals s (Qi) (t) is input to the short-time Fourier transform processing unit 3, and the audio signal to be reproduced in the reproduction region Q1 is the audio signal s (Q1) (t), and the audio signal to be played in the playback region Q2 is the audio signal s (Q2) (t), and the audio signal to be played in the playback region Q3 is the audio signal s (Q3) (t), and the audio signal to be played in the playback region Q4 is the audio signal s (Q4) (t) will be explained below.
[0021] As shown in FIG. 1, the short-time Fourier transform processing unit 3 includes time domain window function processing units 31-1 to 32-4 and discrete Fourier transform processing units 32-1 to 32-4.
[0022] The time domain window function processing unit 31-1 is a time domain window function processing unit for processing an audio signal s (Q1) (t) is input, and the audio signal s (Q1) The time domain window function processing unit 31-1 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-1 converts the signal acquired by the time domain window function processing into a signal s1 (Q1) (t) is output to the discrete Fourier transform processing unit 32-1.
[0023] The time domain window function processing unit 31-2 is a time domain window function processing unit for processing an audio signal s (Q2) (t) is input, and the audio signal s (Q2) The time domain window function processing unit 31-2 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-2 converts the signal acquired by the time domain window function processing into a signal s1 (Q2) (t) is output to the discrete Fourier transform processing unit 32-2.
[0024] The time domain window function processing unit 31-3 processes the audio signal s to be reproduced in the reproduction domain Q3 input to the sound field control device 100. (Q3) (t) is input, and the audio signal s (Q3) The time domain window function processing unit 31-3 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-3 converts the signal acquired by the time domain window function processing into a signal s1 (Q3) (t) and outputs it to the discrete Fourier transform processing unit 32-3.
[0025] The time domain window function processing unit 31-4 processes the audio signal s to be reproduced in the reproduction domain Q4 input to the sound field control device 100. (Q4) (t) is input, and the audio signal s (Q4) The time domain window function processing unit 31-4 performs processing using a time domain window function on (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-4 converts the signal acquired by the time domain window function processing into a signal s1 (Q4) (t) is output to the discrete Fourier transform processing unit 32-4.
[0026] The discrete Fourier transform processing unit 32-1 converts the signal s1 output from the time domain window function processing unit 31-1 into (Q1) (t) is input, and the signal s1 (Q1) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q1)(ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-1 then converts the acquired frequency domain data S (Q1) (ω) is output to the time-frequency filter processing unit 4 for the full range sound field.
[0027] The discrete Fourier transform processing unit 32-2 processes the signal s1 output from the time domain window function processing unit 31-2. (Q2) (t) is input, and the signal s1 (Q2) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q2) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-2 then converts the acquired frequency domain data S (Q2) (ω) is output to the time-frequency filter processing unit 4 for the full-range sound field.
[0028] The discrete Fourier transform processing unit 32-3 processes the signal s1 output from the time domain window function processing unit 31-3. (Q3) (t) is input, and the signal s1 (Q3) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q3) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-3 then converts the acquired frequency domain data S (Q3) (ω) is output to the time-frequency filter processing unit 4 for the full-range sound field.
[0029] The discrete Fourier transform processing unit 32-4 processes the signal s1 output from the time domain window function processing unit 31-4. (Q4) (t) is input, and the signal s1 (Q4) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q4) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-4 then converts the acquired frequency domain data S (Q4)(ω) is output to the time-frequency filter processing unit 4 for the full-range sound field.
[0030] The time-frequency filter processing unit 4 for the whole area sound field receives the data D_settings_zone(Q) output from the reproduction area setting unit 1, the data D_C output from the transfer function acquisition processing unit 2, and the data S output from the discrete Fourier transform processing unit 32-1. (Q1) (ω) and the data S output from the discrete Fourier transform processing unit 32-2. (Q2) (ω) and the data S output from the discrete Fourier transform processing unit 32-3. (Q3) (ω) and the data S output from the discrete Fourier transform processing unit 32-4. (Q4) The time-frequency filter processing unit 4 for the full-range sound field receives the data D_settings_zone(Q), S (Q1) (ω), Data S (Q2) (ω), Data S (Q3) (ω), and data S (Q4) (ω), time-frequency filtering for the entire sound field is performed, and the data (time-frequency domain data) D1(ω) to D2(ω) of the signals (driving signals) for driving each speaker of the speaker array Spk_arry are obtained. L The speaker array Spk_arry is assumed to consist of L speakers (L is a natural number equal to or greater than 2) that are installed at equal intervals on the same circumference, and the signal data (time-frequency domain data) for driving the i-th speaker is obtained as D i Let (ω) (i: natural number, 1≦i≦L).
[0031] Then, the time-frequency filter processing unit 4 for the full-range sound field converts the data (data in the time-frequency domain) D1(ω) to D L (ω) is output to the short-time inverse Fourier transform processing unit 5.
[0032] As shown in FIG. 1, the short-time inverse Fourier transform processing unit 5 includes inverse discrete Fourier transform processing units 51-1 to 51-L and time domain window function processing units 52-1 to 52-L.
[0033] The inverse discrete Fourier transform processing unit 51-1 converts the time-frequency domain data D1 of the drive signal output from the time-frequency filter processing unit 4 for the full-range sound field into (all) (ω) is input, and the corresponding data D1 (all) The inverse discrete Fourier transform processing unit 51-1 performs an inverse discrete Fourier transform on (ω) to obtain a post-inverse discrete Fourier transform signal (time domain signal) d1(t). The inverse discrete Fourier transform processing unit 51-1 then outputs the obtained post-inverse discrete Fourier transform signal (time domain signal) d1(t) to the time domain window function processing unit 52-1.
[0034] The inverse discrete Fourier transform processing unit 51-i (i: natural number, 2≦i≦L) has the same configuration and function as the inverse discrete Fourier transform processing unit 51-1. That is, the inverse discrete Fourier transform processing unit 51-i (i: natural number, 2≦i≦L) converts the time-frequency domain data D of the drive signal output from the full-range sound field time-frequency filter processing unit 4 into i (all) (ω) and input the data D i (all) (ω) is subjected to an inverse discrete Fourier transform, and the signal after the inverse discrete Fourier transform (time domain signal) d i Then, the inverse discrete Fourier transform processing unit 51-i obtains the signal (time domain signal) d after the inverse discrete Fourier transform. i (t) is output to the time domain window function processing unit 52-i.
[0035] The time domain window function processing unit 52-1 inputs the signal (time domain signal) d1(t) after the inverse discrete Fourier transform output from the inverse discrete Fourier transform processing unit 51-1, processes the signal d1(t) using a time domain window function (time domain window function processing), and outputs the signal obtained by this processing as signal sig_d1 to the first speaker Spk1 of the speaker array Spk_arry.
[0036] The time domain window function processor 52-i (i: natural number, 2≦i≦L) has the same configuration and function as the time domain window function processor 52-1. That is, the time domain window function processor 52-i processes the signal (time domain signal) d after the inverse discrete Fourier transform output from the inverse discrete Fourier transform processor 51-i. i (t) is input, and the signal d i (t) is processed using a time domain window function (time domain window function processing), and the signal obtained by this processing is expressed as signal sig_d i and outputs it to the i-th speaker Spki of the speaker array Spk_arry.
[0037] As shown in Figure 2, the speaker array Spk_arry is composed of L (L: a natural number greater than or equal to 2) speakers Spk1 to SpkL (in the case of Figure 2, L = 16) that are equally spaced on the same circumference of a circle with a radius r0 from a center point O when viewed from above.
[0038] The speaker Spki (i: natural number, 1≦i≦L) receives the signal sig_d output from the short-time inverse Fourier transform processor 5. i and the signal sig_d i It outputs (radiates) a sound equivalent to
[0039] <1.2: Operation of the sound field control system> The operation of the sound field control system 1000 configured as above will be described below.
[0040] 3 and 4 are flowcharts of the processing executed by the sound field control system 1000.
[0041] 5 is a diagram for explaining a method for modeling the sound pressure distribution in the reproduction area and the silent area using a function. As shown in FIG. 5, the x-axis, y-axis, and z-axis are set.
[0042] The sound field control system 1000 performs 2.5-dimensional sound field control. Although the sound field emitted from each speaker is three-dimensional, the 2.5-dimensional sound field control method controls the sound field only on the same plane as the speakers in three-dimensional space. Therefore, the speakers and the combined sound pressure are all on the same plane, and the height direction is ignored.
[0043] In the following, for the sake of convenience, as shown in FIG. 2, the area Zone1 (area Q1), the area Zone2 (area Q2), the area Zone3 (area Q3), and the area Zone4 (area Q4) are assumed to be playback areas, and (1) the audio signal to be played in the area Zone1 (area Q1) is designated as audio signal s (Q1) (t) (for example, a Japanese speech signal), and (2) the audio signal to be reproduced in the region Zone2 (region Q2) is the audio signal s (Q2) (t) (for example, an English speech signal), and (3) the audio signal to be reproduced in the region Zone3 (region Q3) is the audio signal s (Q3) (t) (for example, a Chinese speech signal), and (4) the audio signal to be reproduced in the area Zone4 (area Q4) is the audio signal s (Q4) An example of the case where (t) (for example, a Korean voice signal) is used will be described with reference to a flowchart.
[0044] (Step S1): In step S1, a playback area setting process is executed. Specifically, the following process is executed.
[0045] The playback area setting unit 1 sets information (information for identifying area Q1, area Q2, area Q3, area Q4) about areas Q1, Q2, Q3, and Q4, which are areas (playback areas) to be subject to sound field control, and outputs data including the information about the set areas (playback areas) to be subject to sound field control to the time-frequency filter processing unit 4 for the full-area sound field as data D_settings_zone(Q) (Q = {Q1, Q2, Q3, Q4}).
[0046] (Step S2): In step S2, a transfer function acquisition process is executed. Specifically, the following process is executed.
[0047] The transfer function acquisition processing unit 2 acquires a transfer function of each of the L speakers constituting the speaker array Spk_arry and a radius r ref The process of obtaining the two-dimensional cylindrical harmonic spectrum matrix of the transfer function on the circumference of the circle is performed.
[0048] Specifically, the position of the i-th speaker (i: natural number, 1≦i≦L) of the L speakers that make up the speaker array Spk_arry is defined as r i Let D be the drive signal for each speaker (D in the formula below), and let O be the center point synthesized by the drive signal D, and let r ref The two-dimensional cylindrical harmonic spectrum of the sound field on the circumference of a(r ref ) (a(r ref )).
number
number
number
number
[0049] Then, the transfer function acquisition processing unit 2 outputs data including the two-dimensional cylindrical harmonic spectral matrix C acquired by this processing to the full-range sound field time-frequency filter processing unit 4 as data D_C.
[0050] (Steps S3 to S5): In steps S3 to S5, each audio signal s input to the sound field control device 100 (Qi) (t) (In this embodiment, the audio signal s (Q1) (t), s (Q2) (t), s (Q3) (t), and,s (Q4) A short-time Fourier transform process is performed on (t). Specifically, the following process is performed:
[0051] In step S3, all audio signals s input to the sound field control device 100 are (Qi) (t) (In this embodiment, the audio signal s (Q1) (t)~s (Q4) The loop processing (loop 1) is started (executed) for (t).
[0052] In step S4, the audio signal s (Qi) Specifically, a time domain window function processor 31-i (i: integer, 1≦i≦4) performs a short-time Fourier transform process on an audio signal s to be reproduced in a reproduction domain Qi (i: integer, 1≦i≦4) input to the sound field control device 100. (Qi) (t) is input, and the audio signal s (Qi) (t) is processed by a time domain window function g(t) (time domain window function processing) (for example, the window function g(t) of the following formula is applied to the audio signal s (Qi) (t) is convolved with the signal s1 (Qi) Get (t).
number
[0053] The discrete Fourier transform processing unit 32-i (i: integer, 1≦i≦4) converts the signal s1 output from the time domain window function processing unit 31-i. (Qi) (t) is input, and the signal s1 (Qi) A discrete Fourier transform is performed on (t) (the signal in the time domain) (for example, the process corresponding to the following formula is performed), and the data after the discrete Fourier transform (data in the frequency domain) is called data S (Qi) (ω) (ω: angular frequency, ω=2πf, f: frequency).
number
[0054] In step S5, all audio signals s input to the sound field control device 100 are (Qi) (t) (In this embodiment, the audio signal s (Q1) (t)~s (Q4) (t)), it is determined whether or not the loop process (loop 1) has been executed, and (1) all audio signals s input to the sound field control device 100 are (Qi) If it is determined that the loop process (loop 1) has been executed for (t), the process proceeds to step S6, and (2) all audio signals s input to the sound field control device 100 are (Qi) If the loop process (loop 1) has not been executed for (t), the process returns to step S3, and the loop process (loop 1) is repeatedly executed.
[0055] (Step S6): In step S6 (steps S61 to S67), a time-frequency filtering process for the entire sound field is executed. Specifically, the following process is executed.
[0056] (Step S61): In step S61, a loop process (loop 2 process) of the time-frequency filtering process for the entire sound field is started (executed). (Qi) (ω) (in this embodiment, S (Q1) (ω)~S (Q4) Since the sound field control system 1000 performs processing using discrete values, the angular frequency ω is a discrete value ω of the angular frequency. k (ω k : integer, N: natural number, -N≦ω k ≦N, n=ω k ) In other words, the loop 2 processing is executed for integer n (-N≦n≦N, N: natural number) (the loop 2 processing is executed by incrementing n by +1 from n=-N until n=N is reached).
[0057] (Step S62): In step S62, loop processing (loop 3 processing) is started (executed). Note that loop 3 processing is executed for each region Qi (Q1 to Q4 in this embodiment).
[0058] (Step S63): In step S63, a process for acquiring a two-dimensional cylindrical harmonic spectrum of the region Qi is executed. Specifically, the following process is executed.
[0059] For ease of explanation, the following description will be given assuming that the window function is set to a rectangle in all regions (regions Q1 to Q4). In region Qi, the two-dimensional cylindrical harmonic spectrum when the window function is set to a rectangle is expressed as a n,rect (Φ,φ S ) and the two-dimensional cylindrical harmonic spectrum a n,rect (Φ,φ S ) will be explained.
[0060] If the sound pressure in the direction where the sound can be heard is "1" and the sound pressure in the direction where the sound cannot be heard is "0", then when the window function is set to a rectangular window, the sound pressure from the center point O to the radius r ref Sound pressure (sound pressure distribution) P on the circumference of rect (φ) (see FIG. 5) can be expressed by the following equation (which can be modeled as a square wave):
number
number
[0061] (Step S64): In step S64, the audio signal s to be reproduced in the area Qi is (Qi) (t) angular frequency n(=ω k ) frequency component data S n (Qi) Specifically, the time-frequency filter processing unit 4 for the full-range sound field performs the acquisition process of the data S output from the short-time Fourier transform processing unit 3A. (Qi) (ω) and input the data S (Qi) From (ω), data S (Qi) S contained in (ω) n (Qi) By extracting (taking out) the audio signal s to be played in the region Qi, (Qi) (t) angular frequency n(=ω k ) frequency component data S n (Qi) Get.
[0062] (Step S65): In step S65, it is determined whether the termination condition for loop 3 processing is satisfied, and if the result of this determination is that the termination condition for loop 3 processing is not satisfied (if it is determined that loop 3 processing has not been executed for all areas Qi), the process returns to step S62 and loop 3 processing (steps S62 to S65) is executed. On the other hand, if the result of the above determination is that the termination condition for loop 3 processing is satisfied (if it is determined that loop 3 processing has been executed for all areas Qi), the process proceeds to step S66.
[0063] (Step S66): In step S66, the two-dimensional cylindrical harmonic spectrum a for the whole sound field is calculated. n,rect (all) (r ref Specifically, the time-frequency filter processing unit 4 for the entire sound field performs the acquisition process of the two-dimensional cylindrical harmonic spectrum a of the region Qi acquired in steps S62 to S65. n,rect (Φ (Qi) ,φ S (Qi) ) and the frequency component data S of the angular frequency n of the audio signal to be reproduced in the region Qi. n (Qi) By using the formula below, we can obtain the 2D cylindrical harmonic spectrum a for the entire sound field. n,rect (all) (r ref ) to get the
number
[0064] In steps S62 to S66, the processing equivalent to (Equation 9) is performed by processing as described above, but the present invention is not limited to this. For n, the time frequency components of the audio signal to be reproduced in the divided regions of the entire region (regions Q1 to Q4 in this embodiment) are multiplied by the spatial frequency components of the reproduced sound field of that region (corresponding to the two-dimensional cylindrical harmonic spectrum), and the obtained values are added each time, and when all n have been processed, the results of the above two product-sum operations can be obtained.
[0065] (Step S67): In step S67, it is determined whether the termination condition for loop 2 processing is satisfied, and if the result of this determination shows that the termination condition for loop 2 processing is not satisfied (if it is determined that loop 2 processing has not been executed for all n (n: integer, -N≦n≦N)), the process returns to step S61, and loop 2 processing (steps S61 to S67) is executed. On the other hand, if the result of the above determination shows that the termination condition for loop 2 processing is satisfied (if it is determined that loop 2 processing has been executed for all n (n: integer, -N≦n≦N)), the process proceeds to step S67.
[0066] (Step S68): In step S68, a process of acquiring an excitation signal in the time-frequency domain is executed. Specifically, the following process is executed.
[0067] The time-frequency filter processing unit 4 for the full-range sound field is a 2D cylindrical harmonic spectrum a for the full-range sound field obtained in steps S61 to S67. n,rect (Φ (Qi) ,φ S (Qi) ) (n: integer, -N≦n≦N) (2D cylindrical harmonic spectrum for full range sound field a (all) (r ref )) and perform processing equivalent to the following formula to obtain the excitation signal D^ in the time-frequency domain. (all) (Drive signals (signals in the time-frequency domain) for driving L speakers of the speaker array Spk_arry) are acquired.
number
number
[0068] (Step S7): In step S7, short-time inverse Fourier transform processing is performed. Specifically, the following processing is performed.
[0069] The inverse discrete Fourier transform processing unit 51-i (i: natural number, 1≦i≦L) converts the time-frequency domain data D of the driving signal output from the full-range sound field time-frequency filter processing unit 4. i (all) (ω)(=D i (all) (r i )=D i (all) (ω,r i ))(D i (all) (ω) is the 1×L matrix D^ (all) Enter the i-th (i-th column) element of the data D i (all) That is, the inverse discrete Fourier transform processing unit 51-i performs processing corresponding to the following equation to generate a signal d i (t) (time domain driving signal d i (t)).
number
[0070] The time domain window function processing unit 52-i (i: natural number, 1≦i≦L) processes the signal (time domain signal) d after the inverse discrete Fourier transform output from the inverse discrete Fourier transform processing unit 51-i. i (t) is input, and the signal d i (t) by a time domain window function (time domain window function processing) (for example, by applying the window function g(t) in (Equation 5) to the signal d i (t) is convolved with the signal sig_d i (t)(=g(t)*d i (t) (*: operator indicating convolution processing)) and output to the i-th speaker Spki of the speaker array Spk_arry.
[0071] (Step S8): In step S8, the i-th speaker Spki of the speaker array Spk_arry receives the drive signal sig_d obtained in step S7. i (1) By being driven by (t), an audio signal s is transmitted to the reproduction area Zone1 (Q1). (Q1) (t) is played back, and (2) the audio signal s is played back in the playback area Zone2 (Q2). (Q2) (t) is played back, and (3) the audio signal s is played back in the playback area Zone3 (Q3). (Q3) (t) is played back, and (4) the audio signal s is played back in the playback area Zone4 (Q4). (Q4) (t) is played.
[0072] <Summary> As described above, in the sound field control system 1000, the total number of regions is M (M is a natural number equal to or greater than 2, in this embodiment, M=4) (the audio signals to be reproduced in each of the M regions are s (Q1) (t)~s (QM)For the sound field divided into M divided regions (t), a two-dimensional cylindrical harmonic spectrum for the entire sound field is obtained to reproduce the sound field (the entire sound field having M divided regions), and the obtained two-dimensional cylindrical harmonic spectrum for the entire sound field is used to perform time-frequency filtering to obtain the time-frequency domain excitation signal D^ (all) (Drive signals (time-frequency domain signals) for driving L speakers of the speaker array Spk_arry) are acquired. Then, the sound field control system 1000 calculates the time-frequency domain drive signals D^ (all) By performing short-time Fourier transform processing on the L speaker arrays Spk_arry, signals (time-domain drive signals) for driving each of the L speaker arrays Spk_arry are obtained. In other words, the sound field control system 1000 adopts a method of setting a sound field in which the entire area is divided into M areas and reproducing the sound field, so the amount of calculation is dramatically reduced compared to a method (local sound field superposition method) in which a divided sound field (local sound field) is set for each of the M divided areas and drive signals acquired for each local sound field are superposed. In the local sound field superposition method, a divided sound field (local sound field) is set for each of the M divided areas and time-frequency filter processing is performed using the two-dimensional cylindrical harmonic spectrum acquired for each local sound field to generate a time-frequency domain drive signal D^ (Qi) By superimposing the drive signals of the local sound fields, the drive signal D^ for the entire area is obtained. (all) Therefore, the number of calculations for the filtering process in the time frequency filtering process is M×L (L: number of speakers, L is a natural number equal to or greater than 2). In contrast, the sound field control system 1000 sets a sound field in which the entire area is divided into M areas, and employs a method for reproducing the sound field. Therefore, the number of calculations for the filtering process in the time frequency filtering process is only L (in Equation 10, the L×2N+1 matrix ((C(r ref ) H C(r ref )+λI) -1 C(r ref ) H ) and a 2N+1×1 matrix (a (all) (rref ) (the number of operations required to obtain the product of the two matrices above is only L for 2N+1 elements).
[0073] Furthermore, in the sound field control system 1000, regardless of the type of signals that the M number of input audio signals are, the sound pressure distribution of the M number of audio signals can be acquired (calculated) at high speed by short-time Fourier transform processing, enabling high-speed processing.
[0074] As described above, the sound field control system 1000 can perform highly accurate multi-spot reproduction while reducing the number of calculations.
[0075] <First Modification> Next, a first modified example of the first embodiment will be described. Note that the same parts as those in the above embodiment are given the same reference numerals, and detailed description thereof will be omitted.
[0076] FIG. 6 is a schematic configuration diagram of a sound field control system 1000 according to a first modified example of the first embodiment.
[0077] The sound field control system 1000A of this modification has a configuration in which the sound field control device 100 in the sound field control system 1000 of the first embodiment is replaced with a sound field control device 100A.
[0078] The sound field control device 100A has a configuration in which, in the sound field control device 100 of the first embodiment, the short-time Fourier transform processing unit 3 is replaced with a short-time Fourier transform processing unit 3A, the time-frequency filter processing unit for full-area sound field 4 is replaced with a time-frequency filter processing unit for full-area sound field 4A, and a silent area setting unit 6 is further added.
[0079] The silent region setting unit 6 receives, for example, externally input data Info_dark_Q specifying a silent region, and generates data D_dark_Q instructing that silent region processing be performed on the region specified by the data Info_dark_Q.The silent region setting unit 6 then outputs the data D_dark_Q to the short-time Fourier transform processing unit 3A and the full-range sound field time-frequency filter processing unit 4A.
[0080] The short-time Fourier transform processing unit 3A has the same functions as the short-time Fourier transform processing unit 3 of the first embodiment, and further has a function of inputting data D_dark_Q output from the silent region setting unit 6 and performing the following processing. When the data D_dark_Q is input, the short-time Fourier transform processing unit 3A identifies a region designated as a silent region by the data D_dark_Q (this region is called a silent designated region Q_dark), and extracts the audio signal s corresponding to the silent designated region Q_dark. (Q_dark) (t) is forced to "0", or the audio signal s (Q_dark) (t) discrete Fourier transform data (frequency domain data) S (Q_dark) The process forces (ω) to "0".
[0081] The time-frequency filter processing unit 4A for the entire sound field has the same functions as the time-frequency filter processing unit 4 for the entire sound field in the first embodiment, and further has a function of inputting the data D_dark_Q output from the silent region setting unit 6 and performing the following processing. When the data D_dark_Q is input, the time-frequency filter processing unit 4A for the entire sound field calculates the two-dimensional cylindrical harmonic spectrum a of the region Qi. n,rect (Φ (Qi) ,φ S (Qi) )(Qi=Q_dark) is forcibly set to "0".
[0082] The operation of the sound field control system 1000A configured as above will be described below. Note that a description of the same parts as in the first embodiment will be omitted.
[0083] 7 and 8 are flowcharts of the processing executed in the sound field control system 1000A of the first modified example of the first embodiment.
[0084] The processing executed by the sound field control system 1000A will be described below with reference to the flowcharts of FIGS.
[0085] (Steps S1 to S2): In steps S1 and S2, the same processes as in steps S1 and S2 in the first embodiment are executed.
[0086] (Steps S3A to S5A): In step S3A, the same process as in step S3 is executed.
[0087] In step S4A, the same process as in step S4 is executed. However, when data D_dark_Q is input from the silent region setting unit 6 to the short-time Fourier transform processing unit 3A, the area designated as a silent region by the data D_dark_Q (this area is called a silent designated region Q_dark) is identified, and the audio signal s corresponding to the silent designated region Q_dark is extracted. (Q_dark) (t) is forcibly set to "0".
[0088] In step S5A, the same processing as in step S5 is executed.
[0089] (Step S6A): In step S6A, a time-frequency filtering process for the entire sound field is performed. Specifically, the following process is performed.
[0090] (Steps S6A1 to S6A2): In steps S6A1 and S6A2, the same processing as in steps S61 and S62 is executed.
[0091] (Steps S6A1 to S6A2): In steps S6A1 and S6A2, the same processing as in steps S61 and S62 is executed.
[0092] (Step S6A3): In step S6A3, a determination process is executed to determine whether the data to be processed is data in the region Qi that is set as a silent region. If the result of the determination process indicates that the data to be processed is data in the region Qi that is set as a silent region, the process proceeds to step S6A5. If the data to be processed is not data in the region Qi that is set as a silent region, the process proceeds to step S6A4.
[0093] (Steps S6A4 and S6A6): In steps S6A4 and S6A6, the same processing as in steps S63 and S64 is executed.
[0094] (Step S6A5): In step S6A5, silent area setting processing is executed for the area Qi (=area Q_dark set as a silent area). Specifically, the following processing is executed.
[0095] When data D_dark_Q is input from the silent region setting unit 6, the time-frequency filter processing unit 4A for the entire sound field calculates the two-dimensional cylindrical harmonic spectrum a n,rect (Φ (Qi) ,φ S (Qi) )(Qi=Q_dark) is forcibly set to "0".
[0096] Alternatively, when data D_dark_Q is input from the silent region setting unit 6, the short-time Fourier transform processing unit 3A calculates frequency component data S of angular frequency n of the audio signal to be reproduced in the region Qi (=region Q_dark set as the silent region). n (Qi) (Qi=Q_dark) is forcibly set to "0".
[0097] (Steps S6A7 to S6A10): In steps S6A7 to S6A10, the same processes as in steps S65 to S68 are executed, respectively.
[0098] By performing the above processing, the sound field control system 1000A can set the area Qi specified by the data Info_dark_Q as a silent area, while reproducing the sound of a predetermined audio signal in areas other than the silent area. Since the sound field control system 1000A can set the area Qi specified by the data Info_dark_Q as a silent area, when the sound field control system 1000A is used, for example, in a multilingual simultaneous interpretation conference, by setting the area where the speaker is located (the direction of the speaker) as a silent area, the sound does not reach the speaker, dramatically improving the ease with which the speaker can speak. For example, when the playback areas are set as shown in Figure 2, if a speaker (e.g., a Chinese speaker) is present in the playback area Q3, (1) the audio signal s is generated by translating the Chinese of the speaker into Japanese. (Q1) (t), and (2) the Chinese speech of the speaker translated into English is converted into an audio signal s (Q2) (t), and (3) the Chinese speech of the speaker translated into Korean is converted into an audio signal s (Q4) In the case of (t), sound field control system 1000A sets playback area Q3 where the speaker (e.g., a Chinese speaker) is located as a silent area Q_dark, and by performing the above processing, playback area Q3 where the speaker (e.g., a Chinese speaker) is located becomes a silent area (an area where sound does not reach), while the translated audio can be heard in other areas (Japanese audio can be played in area Q1, English audio in area Q2, and Korean audio in area Q4). In this way, when sound field control system 1000A is used in a multilingual simultaneous interpretation conference, by setting the area where the speaker is located (the direction of the speaker) as a silent area, audio does not reach the speaker, dramatically improving the ease with which the speaker can speak.
[0099] In the process of step S6A, the two-dimensional cylindrical harmonic spectrum a of the region Qi is n,rect (Φ (Qi) ,φ S (Qi)) (Qi=Q_dark) is forcibly set to "0", and the frequency component data S of the angular frequency n of the audio signal played in the region Qi (= the region Q_dark set to the silent region) n (Qi) It is also possible to perform both processes of forcibly setting (Qi=Q_dark) to "0".
[0100] The process of step S4A may be the same as the process of step S4 (the audio signal s corresponding to the silent designated region Q_dark). (Q_dark) (It is also possible to forcibly set (t) to "0" without processing.)
[0101] [Second embodiment] Next, a second embodiment will be described. Note that the same parts as those in the above embodiment (including the modified examples) are given the same reference numerals, and detailed description thereof will be omitted.
[0102] In the sound field control system of the second embodiment, processing using a trainable model (for example, a neural network model) is executed in the time-frequency filtering process for the full-area sound field.
[0103] FIG. 9 is a schematic diagram of a sound field control system 2000 according to the second embodiment (during learning processing).
[0104] FIG. 10 is a schematic diagram of the configuration of the all-area sound field time-frequency filter processing unit 4B (during learning processing) of the sound field control system 2000 according to the second embodiment.
[0105] FIG. 11 is a flowchart of the learning process executed by the sound field control system 2000 according to the second embodiment.
[0106] FIG. 12 is a schematic diagram of a sound field control system 2000 (during inference processing) according to the second embodiment.
[0107] FIG. 13 is a schematic diagram of the configuration of the all-area sound field time-frequency filter processing unit 4B (during inference processing) of the sound field control system 2000 according to the second embodiment.
[0108] <2.1: Sound field control system configuration> As shown in FIG. 9, the sound field control system 2000 of the second embodiment has a configuration in which the time frequency filter processing unit 4 for full-area sound field in the sound field control system 1000 of the first embodiment is replaced with a time frequency filter processing unit 4B for full-area sound field.
[0109] As shown in FIG. 10, the time-frequency filter processing unit 4B for the full-area sound field includes a reproduction sound pressure distribution setting processing unit 4B1, a learnable model 4B2, a transfer function processing unit 4B3, a reproduction sound pressure distribution acquisition processing unit 4B4, and a loss evaluation unit 4B5.
[0110] The reproduction sound pressure distribution setting processing unit 4B1 calculates the time frequency component data S of the audio signal output from the short time Fourier transform processing unit 3 by using the data D_settings_zone(Q) output from the reproduction area setting unit 1. (Q1) (ω)~S (QM) (ω) (in this embodiment, M=4). The reproduction sound pressure distribution setting processing unit 4B1 receives the data D_settings_zone(Q) and the time frequency component data S of the audio signal. (Q1) (ω)~S (QM) Based on (ω), the correct data of the reproduced sound pressure distribution P (correct) (Spatial frequency domain sound pressure distribution data P (correct) ) is generated. The reproduction sound pressure distribution setting processing unit 4B1 then generates correct data P (correct) The data including the above is output to the loss evaluation unit 4B5 as data D_P_correct.
[0111] The trainable model 4B2 includes a first linear processing unit 4B21, a second linear processing unit 4B22, and a drive signal acquisition processing unit 4B23.
[0112] The first linear processing unit 4B21 has M input channels (corresponding to the number of input nodes) (M=4 in this embodiment) and N output channels (the number of output nodes). f (N f The first linear processing unit 4B21 includes a linear layer (complex linear layer) (for example, a layer of a neural network model consisting of one or more layers (for example, a fully connected layer)) in which M (M: a natural number of 2 or more, M is the number of divided regions (in this embodiment, M=4)) frequency domain data S of the audio signal output from the short-time Fourier transform processing unit 3. (Q1) (ω)~S (QM) (ω) (corresponding to data of an M × 1 matrix) is input, and the input data is processed by a linear layer (complex linear layer) (the weight matrix (parameter) of the linear layer (complex linear layer) of the first linear processing unit 4B21 is set to W (1st) Then, the first linear processing unit 4B21 performs the above-described processing (N output to the output node) (which may include a process of adding a bias and a process using an activation function). f pieces of data) are output as data D4a to the second linear processing unit 4B22.
[0113] The second linear processing unit 4B22 is configured to process a linearly interpolating input signal when the number of input channels (corresponding to the number of input nodes) is N. f and includes a linear layer (complex linear layer) (for example, a layer of a neural network model consisting of one or more layers (for example, a fully connected layer)) whose output channels (number of output nodes) are L (the number of speakers of the speaker array Spk_arry). The second linear processing unit 4B22 receives the data D4a output from the first linear processing unit 4B21, and processes the input data by the linear layer (complex linear layer) (by converting the weight matrix (parameter) of the linear layer (complex linear layer) of the second linear processing unit 4B22 into W (2nd) (This may include processing to add a bias and processing using an activation function.) The second linear processing unit 4B22 then outputs the processed data (L pieces of data output to the output node) to the drive signal acquisition processing unit 4B23 as data D4b.
[0114] The drive signal acquisition processing unit 4B23 inputs data D4b output from the second linear processing unit 4B22 and acquires this data D4b as frequency component data D of the drive signal (L pieces of data (corresponding to an L×1 matrix of data) acquired by the second linear processing unit 4B22). The drive signal acquisition processing unit 4B23 then outputs data including the frequency component data D of the drive signal as data D4c to the transfer function processing unit 4B3.
[0115] The transfer function processing unit 4B3 receives data D_C output from the transfer function acquisition processing unit 2 and data D4c output from the drive signal acquisition processing unit 4B23. The transfer function processing unit 4B3 acquires a transfer function G (transfer function G in the spatial frequency domain) from the data D_C, and also acquires frequency component data D (corresponding to data in an L×1 matrix) (time frequency component data D) of the drive signal from the data D4c. Then, the transfer function processing unit 4B3 P=G×D Then, the transfer function processing unit 4B3 outputs data including the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) as data D4d to the reproduction sound pressure distribution acquisition processing unit 4B4.
[0116] The reproduction sound pressure distribution acquisition processing unit 4B4 inputs data D4d output from the transfer function processing unit 4B3 and acquires sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) included in the data D4d. Then, the reproduction sound pressure distribution acquisition processing unit 4B4 outputs the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) as data D4e to the loss evaluation unit 4B5.
[0117] The loss evaluation unit 4B5 receives as input the data D_P_correct output from the reproduction sound pressure distribution setting processing unit 4B1, the data D4e output from the reproduction sound pressure distribution acquisition processing unit 4B4, and the frequency component data D of the drive signal corresponding to the audio signal output from the drive signal acquisition processing unit 4B23. The loss evaluation unit 4B5 receives from the data D_P_correct the correct data P of the reproduction sound pressure distribution included in the data.(correct) (Spatial frequency domain sound pressure distribution data P (correct) ) and also obtains sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) included in the data D4e from the data D4e. Then, the loss evaluation unit 4B4 evaluates the correct data P of the reproduced sound pressure distribution using a predetermined loss function. (correct) (Spatial frequency domain sound pressure distribution data P (correct) ) and the sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain). The loss evaluation unit 4B5 then backpropagates the acquired loss (error) by, for example, the error backpropagation method, and executes a process of updating the parameters of the trainable model 4B2.
[0118] <2.2: Operation of the sound field control system> The operation of the sound field control system 2000 configured as above will now be described.
[0119] The operation of the sound field control system 2000 will be explained below by dividing it into (1) learning processing and (2) inference processing.
[0120] (2.2.1: Learning process) First, the learning process executed by the sound field control system 2000 will be described with reference to the flowchart of FIG.
[0121] (Step SB1): In step SB1, a transfer function setting process is executed. Specifically, the following process is executed.
[0122] The transfer function acquisition processing unit 2 acquires a transfer function of each of the L speakers constituting the speaker array Spk_arry and a radius r ref Then, data including the two-dimensional cylindrical harmonic spectrum matrix C acquired by this process is output as data D_C to the transfer function processing unit 4B3 of the time-frequency filter processing unit for full-range sound field 4B.
[0123] The transfer function processing unit 4B3 calculates a two-dimensional cylindrical harmonic spectrum matrix C(=C(r ref )) into the transfer function G(=C(r ref )) (see Equation 2).
[0124] (Step SB2): In step SB2, an initial value setting process is executed in which variable j is set to 1. Note that variable j corresponds to a frame (audio frame) number.
[0125] (Step SB3): In step SB3, the process of acquiring (inputting) time frequency component data of the audio signal of the j-th frame is executed. Specifically, the following process is executed.
[0126] The audio signal s of the jth frame (jth frame) to be played in the regions Q1 to QM (Q1) (t)~s (QM) (t) is input to the short-time Fourier transform processor 3, and processed by the short-time Fourier transform processor 3 to obtain the audio signal s (Q1) (t)~s (QM) Data (time frequency component data) S after short-time Fourier transform processing of (t) (Q1) (ω)~S (QM) (ω) and obtain the audio signal s (Q1) (t)~s (QM) Data (time frequency component data) S after short-time Fourier transform processing of (t) (Q1) (ω)~S (QM) (ω) is input to the first linear processing unit 4B21 and the reproduction sound pressure distribution setting processing unit 4B1 of the trainable model 4B2 of the time-frequency filter processing unit for full-range sound field 4B.
[0127] (Step SB4): In step SB4, a reproduction sound pressure distribution setting process is executed. Specifically, the following process is executed.
[0128] The reproduction area setting unit outputs data D_settings_zone(Q) including information about the reproduction area to the reproduction sound pressure distribution setting processing unit 4B1 of the full-area sound field time-frequency filter processing unit 4B. In this embodiment, the reproduction area Q is described as being set to Q1, Q2, Q3, and Q4 as shown in FIG.
[0129] The reproduction sound pressure distribution setting processing unit 4B1 receives the data D_settings_zone(Q) and the time frequency component data S of the audio signal. (Q1) (ω)~S (QM) Based on (ω), the correct data of the reproduced sound pressure distribution P (correct) (Spatial frequency domain sound pressure distribution data P (correct) ) (the correct data P (correct) P j (correct) In this embodiment, local reproduction is performed in the reproduction region (each of the divided regions Q1 to Q4) using a spatial Hann window (spatial Hann function).
[0130] As explained in the first embodiment, if the sound pressure in the direction where sound can be heard is "1" and the sound pressure in the direction where sound cannot be heard is "0", then the sound pressure from the center point O to the radius r when the window function is set to the Hann window (spatial Hann window) is ref Sound pressure (sound pressure distribution) P on the circumference of hann (φ) can be expressed by the following formula (can be modeled by the following formula).
number
number
number
[0131] (Step SB5): In step SB5, the drive signal acquisition process is performed using the trainable model 4B2. Specifically, the following process is performed.
[0132] The first linear processing unit 4B21 calculates the frequency component data S of the audio signal of the j-th frame to be reproduced in the regions Q1 to Q4. (Q1) (ω)~S (Q4) (ω) is input to each of M nodes of the input node, and processed by a linear layer (complex linear layer) (weight matrix (parameter) of the linear layer (complex linear layer) of the first linear processing unit 4B21): W (1st) ) (which may include a process of adding a bias and a process using an activation function). Then, the first linear processing unit 4B21 performs the above-described process on the data (N output to the output node). f pieces of data) are output as data D4a to the second linear processing unit 4B22.
[0133] The second linear processing unit 4B22 receives the data D4a output from the first linear processing unit 4B21, and performs processing on the received data using a linear layer (complex linear layer) (weight matrix (parameter): W (2nd) ) (which may include processing to add a bias and processing using an activation function). Then, the second linear processing unit 4B22 outputs the processed data (L pieces of data output to the output node) to the drive signal acquisition processing unit 4B23 as data D4b.
[0134] The excitation signal acquisition processing unit 4B23 receives the data D4b output from the second linear processing unit 4B22 and converts the data D4b into frequency component data D of the excitation signal (L pieces of data (corresponding to data in an L×1 matrix) acquired by the second linear processing unit 4B22) (the frequency component data D of the excitation signal corresponding to the audio signal of the j-th frame is D j Then, the drive signal acquisition processing unit 4B23 outputs data including the frequency component data D of the drive signal as data D4c to the transfer function processing unit 4B3.
[0135] (Step SB6): In step SB6, a reproduction sound pressure distribution acquisition process is executed. Specifically, the following process is executed.
[0136] The transfer function processing unit 4B3 uses the transfer function G (transfer function G in the spatial frequency domain) acquired in step SB1 and the frequency component data D (corresponding to data of an L×1 matrix) (time frequency component data D) of the drive signal included in the data D4c to calculate: P=G×D (P j =G×D j ) By performing a process equivalent to the above, sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) (sound pressure distribution data P corresponding to the audio signal of the j-th frame) is converted into P j Then, the transfer function processing unit 4B3 acquires the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) (=P j ) is output as data D4d to the reproduction sound pressure distribution acquisition processing unit 4B4.
[0137] The reproduction sound pressure distribution acquisition processing unit 4B4 inputs data D4d output from the transfer function processing unit 4B3 and acquires sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) included in the data D4d. Then, the reproduction sound pressure distribution acquisition processing unit 4B4 outputs the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) as data D4e to the loss evaluation unit 4B5.
[0138] (Step SB7): In step SB7, the loss evaluation process is executed. Specifically, the following process is executed.
[0139] The loss evaluation unit 4B5 receives the data D_P_correct output from the reproduction sound pressure distribution setting processing unit 4B1 and the data D4e output from the reproduction sound pressure distribution acquisition processing unit 4B4. The loss evaluation unit 4B5 receives the data D_P_correct and calculates the correct data P (correct) (Spatial frequency domain sound pressure distribution data P (correct) )(=P j (correct) ) is acquired, and from the data D4e, the sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) (=P j ) of the excitation signal corresponding to the audio signal of the j-th frame output from the excitation signal acquisition processing unit 4B23. j ) is acquired. The loss evaluation unit 4B5 acquires the sound pressure distribution data P j And the correct sound pressure distribution data P j (correct) The following information shall be stored and maintained.
[0140] Then, the loss evaluation unit 4B5 obtains the loss (error) using the following loss evaluation function:
number
[0141] The second term in the above equation is a regularization term, and the loss evaluation unit 4B5 calculates the frequency component data D (=D j ) can be used to perform processing equivalent to the second term of the above equation, or the frequency component data D (=D j ) to perform processing equivalent to the second term in the above equation.
[0142] (Step SB8): In step SB8, the variable j is set to the number N of frames that have been subjected to the learning process. F As a result of the determination, j≧N is determined. F If j≧N, the process returns to step SB2. F If not, the process proceeds to step SB9.
[0143] (Step SB9): In step SB9, the variable j is incremented by +1 (j←j+1).
[0144] (Step SB10): In step SB10, a convergence determination process is performed based on the loss (Loss) acquired in step SB7. Specifically, if the loss (Loss) acquired in step SB7 becomes equal to or less than a predetermined value, or if the amount of fluctuation in the loss (Loss) acquired in step SB7 becomes equal to or less than a predetermined value, it is determined that the learning process has converged. If it is determined that the learning process has converged, the process proceeds to step SB12, and if it is determined that the learning process has not converged, the process proceeds to step SB11.
[0145] (Step SB11): In step SB11, the error (loss) acquired in step SB7 is backpropagated by, for example, the error backpropagation method (backpropagated to the reproduction sound pressure distribution acquisition processing unit 4B4, the transfer function processing unit 4B3, the drive signal acquisition processing unit 4B23, the second linear processing unit 4B22, and the first linear processing unit 4B21), and the parameters of the trainable model 4B2 (the parameters of the second linear processing unit 4B22 (weight matrix W (2nd) ), and the parameters of the first linear processing unit 4B21 (weighting matrix W (1st) (e.g., parameter update processing is performed using the optimization method Adam).
[0146] (Step SB12): In step SB12, a trained model acquisition process is executed. Specifically, for example, when it is determined in step SB10 that the learning process has converged, the parameters set in the trainable model 4B2 are acquired as optimal parameters. Then, by setting the parameters of the trainable model 4B2 as optimal parameters, the trainable model 4B2 becomes a trained model (a model whose parameters are fixed to the optimal parameters).
[0147] As described above, the learning process is executed in the sound field control system 2000.
[0148] (2.2.2: Inference Processing) Next, the inference processing executed in the sound field control system 2000 will be described.
[0149] In the inference process, as shown in FIG. 12, the trained model acquired in the above learning process (the trainable model 4B2 in which the optimal parameters acquired in the above learning process are set) is installed in the time-frequency filter processing unit 4B for the full-range sound field of the sound field control system 2000. That is, in the inference process, as shown in FIG. 13, the first linear processing unit 4B21 and the second linear processing unit 4B22 of the trainable model 4B2 are set in the time-frequency filter processing unit 4B for the full-range sound field to set the optimal parameters (W opt(1st) , W opt (2nd) ) is installed and executed with the trained model 4B2 set.
[0150] As in the first embodiment, the audio signal s to be reproduced in the regions Q1 to Q4 is (Q1) (t)~s (Q4) (t) is input to the short-time Fourier transform processor 3. The short-time Fourier transform processor 3 executes the same processing as in the first embodiment, and outputs the audio signal s (Q1) (t)~s (Q4) (t) time frequency component data S (Q1) (ω)~S (Q4) (ω) is output to the full-range sound field time-frequency filter processing unit 4B.
[0151] The first linear processing unit 4B21 converts the frequency component data S of the audio signal to be reproduced in the regions Q1 to Q4 output from the short-time Fourier transform processing unit 3 into (Q1) (ω)~S (Q4) (ω) is input to each of M (M=4) nodes of the input node, and processed by the linear layer (complex linear layer) (weight matrix (optimal parameter) of the linear layer (complex linear layer) of the first linear processing unit 4B21): W opt (1st) ) (which may include a process of adding a bias and a process using an activation function). Then, the first linear processing unit 4B21 performs the above-described process on the data (N output to the output node). f pieces of data) are output as data D4a to the second linear processing unit 4B22.
[0152] The second linear processing unit 4B22 receives the data D4a output from the first linear processing unit 4B21, and performs processing on the received data using a linear layer (complex linear layer) (weight matrix (optimal parameters): W opt (2nd)) (which may include processing to add a bias and processing using an activation function). Then, the second linear processing unit 4B22 outputs the processed data (L pieces of data output to the output node) to the drive signal acquisition processing unit 4B23 as data D4b.
[0153] The drive signal acquisition processing unit 4B23 inputs the data D4b output from the second linear processing unit 4B22 and acquires the data D4b as frequency component data D of the drive signal (L pieces of data (corresponding to data in an L×1 matrix) acquired by the second linear processing unit 4B22). Then, the drive signal acquisition processing unit 4B23 calculates the frequency component data D of the drive signal (=[D1 (all) (ω),D2 (all) (ω),···,D L (all) (ω)] T ) to data D1 (all) (ω)~D L (all) (ω) and output to the short-time inverse Fourier transform processing unit 5.
[0154] In the processing after the short-time inverse Fourier transform processing unit 5, the same processing as in the first embodiment is executed.
[0155] As described above, in the sound field control system 2000, the time-frequency filter processing unit 4B for the full-area sound field performs processing using a trained model (e.g., a neural network model), thereby enabling highly accurate sound field control (multi-spot playback).
[0156] That is, in the sound field control system 2000, the total number of regions is M (M is a natural number equal to or greater than 2, in this embodiment, M=4) (the audio signals to be reproduced in each of the M regions are s (Q1) (t)~s (QM)For a sound field divided into M divided regions (t) (a sound field of the entire area having M divided regions), a sound pressure distribution for reproducing the sound field (a sound field of the entire area having M divided regions) is obtained as a correct sound pressure distribution, and a learnable model is trained so that frequency component data of a drive signal for outputting the correct sound pressure distribution is output. Then, the trained model of the trainable model is loaded into the time-frequency filter processing unit 4B for the entire area sound field, and the trained model is used to obtain frequency component data of the drive signal. The frequency component data of the obtained drive signal is subjected to inverse short-time Fourier transform processing to obtain a time-domain drive signal, and the time-domain drive signal is used to drive L speakers of the speaker array Spk_arry, thereby enabling desired sound field control (desired multi-spot reproduction).
[0157] In the sound field control system 2000, the time-frequency filter processing unit for full-area sound field 4B simply performs processing using a trained model, so that highly accurate multi-spot reproduction can be performed while reducing the number of calculations.
[0158] In the above, the case where the learning process is executed by the sound field control device 100B of the sound field control system 2000 has been described, but the present invention is not limited to this. For example, the above learning process (the learning process described as being executed by the sound field control device 100B of the sound field control system 2000) may be executed by another device (e.g., a learning processing device), and optimal parameters acquired by the learning process executed by the other device (e.g., a learning processing device) may be set in the first linear processing unit 4B21 and the second linear processing unit 4B22 of the full-area sound field time-frequency filter processing unit 4B (thereby acquiring the trained model 4B2). In this case, the sound field control device 100B may be configured without including the reproduction area setting unit 1, the transfer function acquisition processing unit 2, the reproduction sound pressure distribution setting processing unit 4B1, the transfer function processing unit 4B3, the reproduction sound pressure distribution acquisition processing unit 4B4, and the loss evaluation unit 4B5.
[0159] In the above, the spatial window for setting the sound pressure distribution of each divided region (regions Q1 to QM) is a Hann window, but this is not limited to this, and the spatial window for setting the sound pressure distribution of each divided region (regions Q1 to QM) may be a rectangular window. Furthermore, the sound pressure distribution of each divided region (regions Q1 to QM) may be any distribution (sound pressure distribution).
[0160] Furthermore, although the above describes a case where the error is backpropagated for each frame in the learning process, the present invention is not limited to this. For example, the loss (error) (e.g., average error) obtained after executing the process (forward propagation process) for a predetermined number of frames may be backpropagated to perform the parameter update process (the error may be backpropagated by batch processing or mini-batch processing).
[0161] The loss function described above is an example, and other loss functions may be used (for example, P j -P j (correct) Alternatively, the L1 norm or L2 norm of a vector w having elements of all updatable parameters (weights) may be used in the regularization term.
[0162] In the second embodiment, a neural network model is used as the learning model, but the learning process may be performed without using an activation function.
[0163] Furthermore, if the parameters of each linear processing unit of the trainable model are updated in real time according to the input audio signal, real-time adaptive sound field control (multi-spot reproduction) becomes possible.
[0164] [Other embodiments] The above-described embodiments and / or modifications may be combined as appropriate to realize a sound field control system and / or a sound field control device.
[0165] In addition, in the sound field control systems 1000 and 1000A described in the above embodiments (including modifications), the total area has four areas Q1 to Q4, but the present invention is not limited to this, and the total area may have M areas (M: a natural number of 2 or more). In this case, the audio signal input to the sound field control devices 100, 100A, and 100B is s (Q1) (t)~s (QM) (t).
[0166] Furthermore, in the embodiments (including modified examples), the case where the sound pressure distribution of the region Qi included in the entire region is modeled as a rectangular wave (corresponding to the window function being a rectangular window) has been described, but this is not limited to this, and the sound pressure distribution of the region Qi included in the entire region may be another sound pressure distribution (for example, the sound pressure distribution of the region Qi included in the entire region may be a sound pressure distribution obtained by modeling using a model corresponding to the window function being a Hann function).
[0167] Furthermore, in the sound field control systems 1000, 1000A and sound field control devices 100, 100A, 100B described in the above embodiments (including modified examples), each block may be individually integrated into a single chip using a semiconductor device such as an LSI, or may be integrated into a single chip to include some or all of the blocks.
[0168] Although we refer to it as an LSI here, it may also be called an IC, system LSI, super LSI, or ultra LSI depending on the level of integration.
[0169] Furthermore, the method of integration is not limited to LSI, but may be realized by dedicated circuits or general-purpose processors. FPGAs (Field Programmable Gate Arrays), which can be programmed after LSI manufacturing, or reconfigurable processors, which allow the connections and settings of circuit cells within LSI to be reconfigured, may also be used.
[0170] Furthermore, part or all of the processing of each functional block in each of the above embodiments (including modified examples) may be realized by a program. And part or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Furthermore, the programs for performing each processing are stored in a storage device such as a hard disk or ROM, and are executed in the ROM or by being read into the RAM.
[0171] Furthermore, each process in the above-described embodiment (including modifications) may be realized by hardware, software (including cases where it is realized together with an OS (operating system), middleware, or a predetermined library), or may be realized by a combination of software and hardware.
[0172] For example, when each functional unit of the above embodiment (including the modified examples) is realized by software, each functional unit may be realized by software processing using the hardware configuration shown in FIG. 14 (for example, a hardware configuration in which a CPU, GPU, ROM, RAM, input unit, output unit, communication unit, memory unit (for example, a memory unit realized by an HDD, SSD, etc.), an external media drive, etc. are connected via a bus).
[0173] Furthermore, when each functional unit of the above embodiment (including modified examples) is realized by software, the software may be realized using a single computer having the hardware configuration shown in Figure 14, or may be realized by distributed processing using multiple computers.
[0174] Furthermore, the order of execution of the processing methods in the above embodiments (including modified examples) is not necessarily limited to that described in the above embodiments, and the order of execution can be changed within the scope of the gist of the invention.
[0175] The scope of the present invention includes a computer program for causing a computer to execute the above-described method, and a computer-readable recording medium having the program recorded thereon. Examples of computer-readable recording media include flexible disks, hard disks, CD-ROMs, MOs, DVDs, DVD-ROMs, DVD-RAMs, large-capacity DVDs, next-generation DVDs, and semiconductor memories.
[0176] The computer program is not limited to one recorded on the recording medium, but may be one transmitted via a telecommunications line, a wireless or wired communication line, a network such as the Internet, or the like.
[0177] The specific configuration of the present invention is not limited to the above-described embodiment (including modified examples), and various changes and modifications are possible without departing from the gist of the invention.
[0178] [Note] The present invention can also be expressed as follows: A first invention is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number of 2 or more), and that performs multi-spot reproduction by using a speaker array consisting of a plurality of speakers to reproduce a plurality of audio signals in different divided areas included in the entire area, and that includes a short-time Fourier transform processing unit, a time-frequency filter processing unit for the entire area sound field, and an inverse short-time Fourier transform processing unit.
[0179] The short-time Fourier transform processor performs short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0180] The time-frequency filter processing unit for the full-area sound field acquires a harmonic spectrum for the full-area sound field to reproduce the full-area sound field based on the harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filter processing using the acquired harmonic spectrum for the full-area sound field to acquire a time-frequency domain drive signal for driving the speaker array.
[0181] The inverse short-time Fourier transform processor performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
[0182] In this sound field control device, for a sound field (a whole-area sound field having M divided areas) in which the whole area is divided into M areas (M: a natural number of 2 or more), a cylindrical harmonic spectrum for the whole-area sound field is acquired to reproduce the sound field (a whole-area sound field having M divided areas), and time-frequency filtering is performed using the acquired cylindrical harmonic spectrum for the whole-area sound field to acquire time-frequency domain drive signals (for example, drive signals (time-frequency domain signals) for driving L speakers of a speaker array Spk_arry).The sound field control device then performs short-time Fourier transform processing on the acquired time-frequency domain drive signals to acquire, for example, signals (time-domain drive signals) for driving each of the L speaker arrays Spk_arry. In other words, this sound field control device sets a sound field in which the entire area is divided into M areas, and then adopts a method of reproducing the sound field.Therefore, the amount of calculation is dramatically reduced compared to, for example, a method in which a divided sound field (local sound field) is set for each of the M divided areas, and the drive signals obtained for each local sound field are superimposed (local sound field superimposition method).
[0183] A second aspect of the present invention is the first aspect of the present invention, further comprising a silent area setting section that sets at least one of the divided areas as a silent area.
[0184] The time-frequency filter processing unit for the entire sound field sets the harmonic spectrum derived from the sound pressure distribution of the divided area for the divided area set in the silent area to 0, or sets the short-time Fourier transform data corresponding to the silent area to 0, and obtains the harmonic spectrum for the entire sound field to reproduce the sound field of the entire area.
[0185] This allows the sound field control device to set at least one of the divided areas as a silent area.
[0186] The third invention is a learning processing method for a parameter-configurable learnable model that receives time-frequency component data of an audio signal as input and outputs frequency component data of a drive signal for reproducing a specified sound field in a specified area using multiple speakers, and includes a transfer function setting processing step, an audio signal time-frequency component data acquisition step, a playback sound pressure distribution setting processing step, a playback sound pressure distribution acquisition processing step, and a loss evaluation processing step.
[0187] The transfer function setting process step sets a transfer function between the speaker position and the control point.
[0188] The audio signal time frequency component data acquisition step performs frequency conversion on the time domain audio signal to acquire frequency component data of the audio signal.
[0189] The reproduction sound pressure distribution setting processing step acquires correct data for the sound pressure distribution to be reproduced in the specified area based on the setting status of the specified area and the frequency component data of the audio signal acquired in the audio signal time frequency component data acquisition step.
[0190] The excitation signal acquisition processing step acquires frequency component data of the excitation signal by processing the frequency component data of the audio signal using a trainable model whose parameters can be set.
[0191] The reproduction sound pressure distribution acquisition processing step acquires reproduction sound pressure distribution data based on the transfer function acquired in the transfer function setting processing step and the frequency component data of the audio signal acquired in the drive signal acquisition processing step.
[0192] The loss evaluation processing step acquires a loss based on the ground truth data of the sound pressure distribution acquired by the reproduction sound pressure distribution setting processing step and the reproduction sound pressure distribution, and executes a process of updating the parameters of the learnable model based on the acquired loss.
[0193] As a result, in this learning processing method, for example, for a sound field (a sound field of an entire area having M divided areas) in which the entire area is divided into M areas (M: a natural number equal to or greater than 2), a sound pressure distribution for reproducing the sound field (a sound field of an entire area having M divided areas) can be acquired (set) as a correct sound pressure distribution, and a trainable model can be trained so that frequency component data of a drive signal for outputting the correct sound pressure distribution is output. Then, in this learning processing method, by processing using the trained model of the trainable model, for example, frequency component data of a drive signal can be acquired from frequency component data of an audio signal to be reproduced in the entire area having M divided areas, and the acquired frequency component data of the drive signal can be subjected to inverse short-time Fourier transform to acquire a time-domain drive signal, and multiple speakers (for example, L speakers of a speaker array Spk_arry) can be driven by the time-domain drive signal, thereby enabling desired sound field control (desired multi-spot reproduction).
[0194] The learning method may be implemented using, for example, a processor and a memory accessible from the processor, and each step of the learning method may be performed by the processor.
[0195] Furthermore, "frequency transformation" refers to a process (transformation process) for obtaining frequency component data of a time-domain signal from the signal. Examples of frequency transformation include Fourier transform (including discrete Fourier transform, fast Fourier transform, etc.), wavelet transform (including discrete wavelet transform, etc.), and discrete cosine transform.
[0196] The fourth invention is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number of 2 or more), and is a sound field control device for performing multi-spot playback in which a speaker array consisting of a plurality of speakers is used to play a plurality of audio signals in different divided areas included in the entire area, and is equipped with a short-time Fourier transform processing unit, a time-frequency filter processing unit for the entire area sound field, and a short-time Fourier inverse transform processing unit.
[0197] The short-time Fourier transform processor performs short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0198] The time-frequency filter processing unit for full-range sound field performs processing using a trained model, which is a trainable model in which the acquired optimal parameters are set, by executing a learning process using the learning processing method of the third invention.The time-frequency filter processing unit for full-range sound field then inputs the short-time Fourier transform data to the trained model and acquires data output from the trained model as a time-frequency domain driving signal for driving a speaker array.
[0199] The inverse short-time Fourier transform processor performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
[0200] As a result, in this sound field control device, for example, frequency component data of an audio signal to be reproduced in an entire area having M divided areas can be input to a trained model of a time-frequency filter processing unit for an entire area sound field, and frequency component data of a drive signal can be acquired from the trained model. Then, in this sound field control device, the frequency component data of the acquired drive signal is subjected to inverse short-time Fourier transform processing to acquire a time-domain drive signal, and by driving multiple speakers (for example, L speakers of a speaker array Spk_arry) with the time-domain drive signal, desired sound field control (desired multi-spot reproduction) can be performed.
[0201] The fifth invention is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number of 2 or more), and is a sound field control method for performing multi-spot playback using a speaker array consisting of multiple speakers to play multiple audio signals in different divided areas included in the entire area, and includes a short-time Fourier transform processing step, a time-frequency filter processing step for the entire area sound field, and a short-time inverse Fourier transform processing step.
[0202] The short-time Fourier transform processing step performs a short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0203] The time-frequency filtering process for the entire sound field acquires a harmonic spectrum for the entire sound field to reproduce the sound field of the entire area based on the harmonic spectrum derived from the sound pressure distribution of the divided area and the short-time Fourier transform data, and performs time-frequency filtering process using the acquired harmonic spectrum for the entire sound field to acquire a time-frequency domain driving signal for driving the speaker array.
[0204] The inverse short-time Fourier transform processing step performs an inverse Fourier transform on the time-frequency domain driving signal to obtain a driving signal for driving each speaker of the speaker array.
[0205] This makes it possible to realize a sound field control method that has the same effects as the first aspect of the invention.
[0206] A sixth aspect of the present invention is a program for causing a computer to execute the method (learning processing method, sound field control method) of the third or fifth aspect of the present invention.
[0207] This makes it possible to realize a program for causing a computer to execute a method (learning processing method, sound field control method) that has the same effects as the third or fifth invention. [Explanation of symbols]
[0208] 1000, 1000A, 2000A Sound Field Control System 100, 100A, 100B Sound field control device 3. 3A Short-time Fourier transform processing section 4, 4A Time-frequency filter processing unit for full-range sound field 5 Short-time inverse Fourier transform processing section 6 Silence area setting section Spk_arry Speaker Array Spk1~Spk16 speakers
Claims
1. A sound field control device for controlling a sound field of an entire area having M divided areas (M: a natural number of 2 or more), the sound field control device being for performing multi-spot reproduction in which a plurality of audio signals are reproduced in different divided areas included in the entire area using a speaker array consisting of a plurality of speakers, a short-time Fourier transform processing unit that performs a short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data; a time-frequency filtering unit for a full-area sound field that obtains a harmonic spectrum for reproducing the sound field of the full area based on the harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filtering using the obtained harmonic spectrum for the full-area sound field to obtain a time-frequency domain driving signal for driving the speaker array; a short-time inverse Fourier transform processor that performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array; A sound field control device comprising:
2. a silent area setting unit that sets at least one of the divided areas as a silent area; The full-range sound field time-frequency filter processing unit a harmonic spectrum derived from the sound pressure distribution of the divided region set in the silent region is set to 0, or the short-time Fourier transform data corresponding to the silent region is set to 0, and the harmonic spectrum for the whole region sound field for reproducing the sound field of the whole region is acquired. The sound field control device according to claim 1 .
3. A learning processing method for a parameter-configurable trainable model that receives time-frequency component data of an audio signal as input and outputs frequency component data of a drive signal for reproducing a predetermined sound field in a predetermined area using a plurality of speakers, comprising: a transfer function setting process step for setting a transfer function between the speaker position and the control point; an audio signal time frequency component data acquisition step of performing frequency conversion on a time domain audio signal to acquire frequency component data of the audio signal; a reproduction sound pressure distribution setting processing step of acquiring correct data of a sound pressure distribution to be reproduced in the predetermined region based on the setting status of the predetermined region and the frequency component data of the audio signal acquired in the audio signal time-frequency component data acquisition step; a drive signal acquisition processing step of acquiring frequency component data of a drive signal by processing the frequency component data of the audio signal using a trainable model whose parameters can be set; a reproduction sound pressure distribution acquisition processing step of acquiring reproduction sound pressure distribution data based on the transfer function acquired in the transfer function setting processing step and frequency component data of the audio signal acquired in the drive signal acquisition processing step; a loss evaluation processing step of acquiring a loss based on the ground truth data of the sound pressure distribution acquired by the reproduction sound pressure distribution setting processing step and the reproduction sound pressure distribution, and executing a process of updating parameters of the trainable model based on the acquired loss; A learning processing method comprising:
4. A sound field control device for controlling a sound field of an entire area having M divided areas (M: a natural number of 2 or more), the sound field control device being for performing multi-spot reproduction in which a plurality of audio signals are reproduced in different divided areas included in the entire area using a speaker array consisting of a plurality of speakers, a short-time Fourier transform processing unit that performs a short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data; a time-frequency filter processing unit for a full-range sound field that performs processing using a trained model, which is a trainable model in which the acquired optimal parameters are set by executing a learning process according to the learning processing method of claim 3, and inputs the short-time Fourier transform data to the trained model, and acquires data output from the trained model as a time-frequency domain driving signal for driving the speaker array; a short-time inverse Fourier transform processor that performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array; A sound field control device comprising:
5. A sound field control device controls a sound field in an entire area having M divided areas (M: a natural number of 2 or more), and a sound field control method for performing multi-spot reproduction in which a plurality of audio signals are reproduced in different divided areas included in the entire area using a speaker array consisting of a plurality of speakers, the method comprising: a short-time Fourier transform processing step of performing a short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data; a time-frequency filtering step for a full-area sound field, which obtains a harmonic spectrum for reproducing the full-area sound field based on the harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filtering using the obtained harmonic spectrum for the full-area sound field to obtain a time-frequency domain driving signal for driving the speaker array; a short-time inverse Fourier transform processing step of performing an inverse Fourier transform on the time-frequency domain drive signal to obtain drive signals for driving each speaker of the speaker array; A sound field control method comprising:
6. A program for causing a computer to execute the method according to claim 3 or 5.