Sound field control device, sound field control method, and program
The sound field control system optimizes computational efficiency for multi-spot reproduction by using Fourier transform processors, enabling accurate and resource-efficient creation of multiple sound zones.
Patent Information
- Application Number
- PCT/JP2025/000350
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-02-15
- Filing Date
- 2025-01-08
- Publication Date
- 2025-08-21
AI Technical Summary
Existing sound field control technologies require excessive computational resources for multi-spot reproduction, making them inefficient for mobile applications.
A sound field control system utilizing a speaker array with a short-time Fourier transform processor, time-frequency filter processor, and inverse Fourier transform processor to minimize calculations while achieving accurate multi-spot reproduction.
The system reduces computational load while maintaining high accuracy in creating distinct sound zones, suitable for mobile devices and applications requiring multiple sound fields.
Smart Images

Figure JP2025000350_21082025_PF_FP_ABST
Abstract
Description
Sound field control device, sound field control method and program
[0001] The present invention relates to a multi-sound field control technology that simultaneously realizes different sound fields in multiple different areas.
[0002] Among sound field control technologies using multiple speakers, local reproduction technology, in which sound is heard only in a certain area, and multi-spot reproduction technology, which can present different sounds in different areas by overlapping multiple local reproductions, are expected to have a wide range of applications in personal audio for teleworking, guide systems at museums and world expositions, and translation systems. Unlike ultrasonic speakers, such local reproduction and multi-spot reproduction technologies control sounds in the audible range, resulting in high sound quality and low sound pressure levels, so there are no health concerns. Therefore, such local reproduction and multi-spot reproduction technologies are expected to be new methods to replace ultrasonic speakers.
[0003] Multi-spot reproduction technology can be realized by calculating a local reproduction filter set in advance so that sound reaches only each region, and filtering an input sound signal with the local reproduction filter (see, for example, Non-Patent Document 1). For example, when using a circular speaker array to reproduce audio in four languages in four regions in a multi-spot manner, the audio in four languages can be reproduced in four regions in a multi-spot manner by overlapping the local reproduction of each region.
[0004] 1. T. Okamoto and A. Sakaguchi, "Experimental validation of spatial Fourier transform-based multiple sound zone generation with a linear loudspeaker array," J. Acoust. Soc. Am., vol. 141, no. 3, pp. 1769-1780, Mar. 2017.
[0005] However, the above-mentioned method (for example, the method disclosed in Non-Patent Document 1) is based on the superposition of local reproduction, and therefore, for example, to achieve multi-spot reproduction of N areas using L channel speakers (L: a natural number equal to or greater than 2), filtering calculations are required L×N times, resulting in a problem of a large amount of calculation. When taking into consideration control on mobile terminals, etc., it is desirable to reduce the amount of calculation as much as possible, which poses a problem.
[0006] In view of the above, an object of the present invention is to provide a sound field control system, a sound field control device, a sound field control method, and a program that perform highly accurate multi-spot reproduction while minimizing the amount of calculation.
[0007] In order to solve the above problem, a representative example (one aspect) of the invention disclosed in this application is a sound field control device that controls a sound field of an entire area having M (M: a natural number greater than or equal to 2) divided areas, and is a sound field control device for performing multi-spot playback in which multiple audio signals are each played back in different divided areas included in the entire area using a speaker array consisting of multiple speakers, and is equipped with a short-time Fourier transform processing unit, a time-frequency filter processing unit for the entire area sound field, and a short-time Fourier inverse transform processing unit.
[0008] The short-time Fourier transform processor performs short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0009] The time-frequency filter processing unit for the full-area sound field acquires a harmonic spectrum for the full-area sound field to reproduce the full-area sound field based on the harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filter processing using the acquired harmonic spectrum for the full-area sound field to acquire a time-frequency domain drive signal for driving the speaker array.
[0010] The inverse short-time Fourier transform processor performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
[0011] According to the present invention, it is possible to realize a sound field control system, a sound field control device, a sound field control method, and a program that perform highly accurate multi-spot reproduction while reducing the amount of calculation.
[0012] FIG. 1 is a schematic configuration diagram of a sound field control system 1000 according to the first embodiment. FIG. 2 is a schematic configuration diagram of a speaker array Spk_arry of the sound field control system 1000 according to the first embodiment. FIG. 3 is a flowchart of processing executed by the sound field control system 1000. FIG. 4 is a flowchart of processing executed by the sound field control system 1000. FIG. 5 is a diagram for explaining a method of modeling sound pressure distributions in a reproduction area and a silent area using a function. FIG. 6 is a schematic configuration diagram of a sound field control system 1000 according to a first modified example of the first embodiment. FIG. 7 is a flowchart of processing executed by a sound field control system 1000A according to a first modified example of the first embodiment. FIG. 8 is a flowchart of processing executed by a sound field control system 1000A according to a first modified example of the first embodiment. FIG. 9 is a schematic configuration diagram of a sound field control system 2000 according to a second embodiment (during learning processing). FIG. 10 is a schematic configuration diagram of a time-frequency filter processing unit 4B for a full-area sound field (during learning processing) of the sound field control system 2000 according to the second embodiment. FIG. 11 is a flowchart of learning processing executed by the sound field control system 2000 according to the second embodiment. FIG. 10 is a schematic configuration diagram of a sound field control system 2000 according to a second embodiment (during inference processing). FIG. 11 is a schematic configuration diagram of a time-frequency filter processing unit 4B for a full-area sound field (during inference processing) of the sound field control system 2000 according to the second embodiment. FIG. 12 is a diagram showing a CPU bus configuration.
[0013] First Embodiment A first embodiment will be described below with reference to the drawings.
[0014] 1.1: Configuration of Sound Field Control System FIG. 1 is a schematic diagram of a sound field control system 1000 according to the first embodiment.
[0015] Fig. 2 is a schematic diagram of the speaker array Spk_arry of the sound field control system 1000 according to the first embodiment. Note that the x-axis, y-axis, and z-axis are set as shown in Fig. 2. The origin of the xyz space (three-dimensional space) is set to point O in Fig. 2.
[0016] As shown in Fig. 1, the sound field control system 1000 includes a sound field control device 100 and a speaker array Spk_arry. For ease of explanation, the sound field control system 1000 will be described below as controlling a sound field in a plurality of zones (for example, in the case of Fig. 2, four zones: a first zone Zone1 (this zone is designated as Q1), a second zone Zone2 (this zone is designated as Q2), a third zone Zone3 (this zone is designated as Q3), and a fourth zone Zone4 (this zone is designated as Q4)) by driving a circular speaker array (a speaker array consisting of a plurality of speakers installed at equal intervals on the same circumference) as shown in Fig. 2. Furthermore, as shown in Fig. 2, the speaker array Spk_arry is arranged in a plane view from above, with a radius r from a center point O. 0 It is assumed that the system is configured with L (L: natural number of 2 or more) speakers Spk1 to SpkL (in the case of FIG. 2, L=16) that are installed at equal intervals on the same circumference.
[0017] As shown in Figure 1, the sound field control device 100 includes a playback area setting unit 1, a transfer function acquisition processing unit 2, a short-time Fourier transform processing unit 3, a time-frequency filter processing unit for the entire sound field 4, and a short-time inverse Fourier transform processing unit 5.
[0018] The reproduction area setting unit 1 sets information about multiple areas (reproduction areas) to be subject to sound field control. For ease of explanation, the following description will be given assuming that four reproduction areas (four reproduction areas Zone 1, Zone 2, Zone 3, and Zone 4 in FIG. 2) are set as multiple areas (reproduction areas) to be subject to sound field control. The reproduction area setting unit 1 outputs data including information about the set areas (reproduction areas) to be subject to sound field control to the full-area sound field time-frequency filter processing unit 4 as data D_settings_zone(Q) (Q = {Q1, Q2, Q3, Q4}).
[0019] The transfer function acquisition processing unit 2 acquires a transfer function of each of the L speakers constituting the speaker array Spk_arry and a radius r refThe processing unit 100 performs a process to obtain a two-dimensional cylindrical harmonic spectral matrix of the transfer function on the circumference of the circle, and outputs data including the two-dimensional cylindrical harmonic spectral matrix C obtained by this process as data D_C to the time-frequency filter processing unit 4 for the full-range sound field.
[0020] The short-time Fourier transform processing unit 3 converts the audio signal s (Qi) (t) (audio signal (time domain signal) to be reproduced (output) in reproduction domain Qi (i: natural number), t: time) is input, and the audio signal s (Qi) For the sake of convenience, in this embodiment, four audio signals s (Qi) (t) is input to the short-time Fourier transform processing unit 3, and the audio signal to be reproduced in the reproduction region Q1 is the audio signal s (Q1) (t), and the audio signal to be reproduced in the reproduction area Q2 is the audio signal s (Q2) (t), and the audio signal to be reproduced in the reproduction area Q3 is the audio signal s (Q3) (t), and the audio signal to be reproduced in the reproduction area Q4 is the audio signal s (Q4) (t) will be explained below.
[0021] As shown in FIG. 1, the short-time Fourier transform processing unit 3 includes time domain window function processing units 31-1 to 32-4 and discrete Fourier transform processing units 32-1 to 32-4.
[0022] The time domain window function processing unit 31-1 is a time domain window function processing unit for processing an audio signal s (Q1) (t) and the audio signal s (Q1) The time domain window function processing unit 31-1 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-1 converts the signal acquired by the time domain window function processing into a signal s1 (Q1) (t) and outputs it to the discrete Fourier transform processing unit 32-1.
[0023] The time domain window function processing unit 31-2 receives the audio signal s (Q2) (t) and the audio signal s(Q2) The time domain window function processing unit 31-2 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-2 converts the signal acquired by the time domain window function processing into a signal s1 (Q2) (t) and outputs it to the discrete Fourier transform processing unit 32-2.
[0024] The time domain window function processing unit 31-3 is a time domain window function processing unit for processing an audio signal s (Q3) (t) and the audio signal s (Q3) The time domain window function processing unit 31-3 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-3 converts the signal acquired by the time domain window function processing into a signal s1 (Q3) (t) and outputs it to the discrete Fourier transform processing unit 32-3.
[0025] The time domain window function processing unit 31-4 processes the audio signal s to be reproduced in the reproduction domain Q4 input to the sound field control device 100. (Q4) (t) and the audio signal s (Q4) The time domain window function processing unit 31-4 performs processing using a time domain window function at (t) to acquire a signal for a predetermined period (unit time). Then, the time domain window function processing unit 31-4 converts the signal acquired by the time domain window function processing into a signal s1 (Q4) (t) and outputs it to the discrete Fourier transform processing unit 32-4.
[0026] The discrete Fourier transform processing unit 32-1 converts the signal s1 output from the time domain window function processing unit 31-1 into (Q1) (t) is input, and the signal s1 (Q1) (t) (time domain signal) is subjected to a discrete Fourier transform, and the data after the discrete Fourier transform (frequency domain data) is referred to as data S (Q1) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-1 then converts the acquired frequency domain data S (Q1) (ω) is output to the time-frequency filter processing unit 4 for the full-range sound field.
[0027] The discrete Fourier transform processing unit 32-2 processes the signal s1 output from the time domain window function processing unit 31-2. (Q2) (t) is input, and the signal s1 (Q2) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q2) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-2 then converts the acquired frequency domain data S (Q2) (ω) is output to the time-frequency filter processing unit 4 for the entire sound field.
[0028] The discrete Fourier transform processing unit 32-3 processes the signal s1 output from the time domain window function processing unit 31-3. (Q3) (t) is input, and the signal s1 (Q3) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q3) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-3 then converts the acquired frequency domain data S (Q3) (ω) is output to the time-frequency filter processing unit 4 for the entire sound field.
[0029] The discrete Fourier transform processing unit 32-4 processes the signal s1 output from the time domain window function processing unit 31-4. (Q4) (t) is input, and the signal s1 (Q4) A discrete Fourier transform is performed on (t) (time domain signal), and the data after the discrete Fourier transform (frequency domain data) is called data S (Q4) (ω) (ω: angular frequency, ω=2πf, f: frequency). The discrete Fourier transform processing unit 32-4 then converts the acquired frequency domain data S (Q4) (ω) is output to the time-frequency filter processing unit 4 for the entire sound field.
[0030] The time-frequency filter processing unit 4 for the entire sound field receives the data D_settings_zone(Q) output from the reproduction area setting unit 1, the data D_C output from the transfer function acquisition processing unit 2, and the data S output from the discrete Fourier transform processing unit 32-1.(Q1) (ω) and the data S output from the discrete Fourier transform processing unit 32-2. (Q2) (ω) and the data S output from the discrete Fourier transform processing unit 32-3. (Q3) (ω) and the data S output from the discrete Fourier transform processing unit 32-4. (Q4) The time-frequency filter processing unit 4 for the full-range sound field receives the data D_settings_zone(Q), S (Q1) (ω), data S (Q2) (ω), data S (Q3) (ω), and data S (Q4) (ω) is used to perform time-frequency filtering for the entire sound field, and data (time-frequency domain data) D of signals (driving signals) for driving each speaker of the speaker array Spk_arry is obtained. 1 (ω) ~D L The speaker array Spk_arry is assumed to be composed of L speakers (L: a natural number of 2 or more) that are installed at equal intervals on the same circumference, and the signal data (time-frequency domain data) for driving the i-th speaker is obtained as D i (ω) (i: natural number, 1≦i≦L).
[0031] Then, the time-frequency filter processing unit 4 for the full-range sound field uses the data (data in the time-frequency domain) D 1 (ω) ~D L (ω) is output to the short-time inverse Fourier transform processing unit 5.
[0032] As shown in FIG. 1, the short-time inverse Fourier transform processing unit 5 includes inverse discrete Fourier transform processing units 51-1 to 51-L and time domain window function processing units 52-1 to 52-L.
[0033] The inverse discrete Fourier transform processing unit 51-1 converts the time-frequency domain data D of the drive signal output from the full-range sound field time-frequency filter processing unit 4 into 1 (all) (ω) is input, and the data D 1 (all) (ω) is subjected to an inverse discrete Fourier transform, and the signal after the inverse discrete Fourier transform (time domain signal) d 1Then, the inverse discrete Fourier transform processing unit 51-1 obtains the signal (time domain signal) d after the inverse discrete Fourier transform. 1 (t) is output to the time domain window function processing unit 52-1.
[0034] The inverse discrete Fourier transform processing unit 51-i (i: natural number, 2≦i≦L) has the same configuration and function as the inverse discrete Fourier transform processing unit 51-1. That is, the inverse discrete Fourier transform processing unit 51-i (i: natural number, 2≦i≦L) converts the time-frequency domain data D of the drive signal output from the full-range sound field time-frequency filter processing unit 4 into i (all) (ω) is input, and the data D i (all) (ω) is subjected to an inverse discrete Fourier transform, and the signal after the inverse discrete Fourier transform (time domain signal) d i Then, the inverse discrete Fourier transform processing unit 51-i obtains the signal (time domain signal) d after the inverse discrete Fourier transform. i (t) is output to the time domain window function processing unit 52-i.
[0035] The time domain window function processing unit 52-1 receives the signal (time domain signal) d after the inverse discrete Fourier transform output from the inverse discrete Fourier transform processing unit 51-1. 1 (t) is input, and the signal d 1 (t) is processed using a time domain window function (time domain window function processing), and the signal obtained by this processing is used as the signal sig_d 1 and outputs it to the first speaker Spk1 of the speaker array Spk_arry.
[0036] The time domain window function processor 52-i (i: natural number, 2≦i≦L) has the same configuration and function as the time domain window function processor 52-1. That is, the time domain window function processor 52-i processes the signal (time domain signal) d after the inverse discrete Fourier transform output from the inverse discrete Fourier transform processor 51-i. i (t) is input, and the signal d i (t) is processed using a time domain window function (time domain window function processing), and the signal obtained by this processing is used as the signal sig_d iand outputs it to the i-th speaker Spki of the speaker array Spk_arry.
[0037] As shown in FIG. 2, the speaker array Spk_arry is a speaker array having a radius r 0 The system is composed of L (L: natural number of 2 or more) speakers Spk1 to SpkL (in the case of FIG. 2, L=16) that are installed at equal intervals on the same circumference.
[0038] The speaker Spki (i: natural number, 1≦i≦L) receives the signal sig_d output from the short-time inverse Fourier transform processor 5. i and the signal sig_d i It outputs (radiates) a sound equivalent to
[0039] <1.2: Operation of Sound Field Control System> The operation of the sound field control system 1000 configured as above will be described below.
[0040] 3 and 4 are flowcharts of the processing executed by the sound field control system 1000.
[0041] 5 is a diagram for explaining a method for modeling the sound pressure distribution in the reproduction area and the silent area using a function. As shown in FIG. 5, the x-axis, y-axis, and z-axis are set.
[0042] The sound field control system 1000 performs 2.5-dimensional sound field control. Although the sound field emitted from each speaker is three-dimensional, the 2.5-dimensional sound field control method controls the sound field only on the same plane as the speakers in three-dimensional space. Therefore, the speakers and the combined sound pressure are all on the same plane, and the height direction is ignored.
[0043] In the following, for the sake of convenience, as shown in FIG. 2, the area Zone 1 (area Q1), the area Zone 2 (area Q2), the area Zone 3 (area Q3), and the area Zone 4 (area Q4) are set as reproduction areas, and (1) the audio signal to be reproduced in the area Zone 1 (area Q1) is set as the audio signal s (Q1) (t) (for example, a Japanese speech signal), and (2) an audio signal to be reproduced in the area Zone2 (area Q2) is an audio signal s (Q2)(t) (for example, an English audio signal), and (3) the audio signal to be reproduced in the area Zone 3 (area Q3) is the audio signal s (Q3) (t) (for example, a Chinese speech signal), and (4) the audio signal to be reproduced in the area Zone4 (area Q4) is the audio signal s (Q4) An example of the case where (t) (for example, a Korean voice signal) is used will be described with reference to a flowchart.
[0044] (Step S1): In step S1, a playback area setting process is executed. Specifically, the following process is executed.
[0045] The playback area setting unit 1 sets information (information for identifying area Q1, area Q2, area Q3, area Q4) about areas Q1, Q2, Q3, and Q4, which are areas to be controlled in the sound field (playback areas), and outputs data including the information about the areas to be controlled in the sound field (playback areas) that have been set as data D_settings_zone(Q) (Q = {Q1, Q2, Q3, Q4}) to the time-frequency filter processing unit 4 for the full-area sound field.
[0046] (Step S2): In step S2, a transfer function acquisition process is executed. Specifically, the following process is executed.
[0047] The transfer function acquisition processing unit 2 acquires a transfer function of each of the L speakers constituting the speaker array Spk_arry and a radius r ref A process is performed to obtain a two-dimensional cylindrical harmonic spectrum matrix of the transfer function on the circumference of the circle.
[0048] Specifically, the position of the i-th speaker (i: natural number, 1≦i≦L) among the L speakers constituting the speaker array Spk_arry is defined as r i The drive signal of each speaker is D (D in the following formula), and the center point of the drive signal D is O. ref The two-dimensional cylindrical harmonic spectrum of the sound field on the circumference of a (r ref ) (a(r ref )). At this time, each of the L speakers of the speaker array Spk_arry and theref The two-dimensional cylindrical harmonic spectral matrix C(r ref ) is expressed by the following formula. At this time, the i-th (i: natural number, 1≦i≦L) speaker and its position (r ref , φ) (φ: angle with the x-axis in the xy plane) is defined as H i (r ref , φ), then (H i (): Hankel function of the first kind), the two-dimensional cylindrical harmonic spectral matrix C(r ref Element c of n,i (r ref ) is obtained by the following formula: j: imaginary unit However, in reality, it is difficult to measure φ continuously, so the radius r from the center point O ref M microphones (M: natural number of 2 or more) are placed at equal intervals on the circumference of a circle, and the transfer function H i (r ref , φ) is measured. In this case, n,i (r ref ) is approximately obtained by the following formula: The transfer function acquisition processing unit 2 executes the process corresponding to the above formula to obtain the transfer function of each of the L speakers constituting the speaker array Spk_arry and the radius r ref Then, the transfer function acquisition processing unit 2 acquires each element of the two-dimensional cylindrical harmonic spectrum matrix C(r ref ) to obtain the
[0049] Then, the transfer function acquisition processing unit 2 outputs data including the two-dimensional cylindrical harmonic spectrum matrix C acquired by this processing to the full-range sound field time-frequency filter processing unit 4 as data D_C.
[0050] (Steps S3 to S5): In steps S3 to S5, each audio signal s input to the sound field control device 100 is (Qi) (t) (in this embodiment, audio signal s (Q1) (t), s (Q2) (t), s (Q3) (t), and s (Q4)A short-time Fourier transform process is performed on (t). Specifically, the following process is performed:
[0051] In step S3, all audio signals s input to the sound field control device 100 are (Qi) (t) (in this embodiment, audio signal s (Q1) (t) ~ s (Q4) A loop process (loop 1) is started (executed) with (t)) as the target.
[0052] In step S4, the audio signal s (Qi) Specifically, a time domain window function processor 31-i (i: integer, 1≦i≦4) performs a short-time Fourier transform process on an audio signal s to be reproduced in a reproduction domain Qi (i: integer, 1≦i≦4) input to the sound field control device 100. (Qi) (t) and the audio signal s (Qi) (t) is processed by a time domain window function g(t) (time domain window function processing) (for example, the window function g(t) of the following formula is applied to the audio signal s (Qi) (t) is convoluted into the signal s1 (Qi) (t) is obtained. t: each time (sample unit) T: number of frames to be subjected to short-time Fourier transform s1 (Qi) (t)=g(t)*s (Qi) (t) *: operator indicating convolution processing Then, the time domain window function processing unit 31-i converts the signal acquired by the time domain window function processing into a signal s1 (Qi) (t) and outputs it to the discrete Fourier transform processing unit 32-i.
[0053] The discrete Fourier transform processing unit 32-i (i: integer, 1≦i≦4) processes the signal s1 output from the time domain window function processing unit 31-i. (Qi) (t) is input, and the signal s1 (Qi) A discrete Fourier transform is performed on (t) (a signal in the time domain) (for example, a process corresponding to the following formula is performed), and the data after the discrete Fourier transform (data in the frequency domain) is referred to as data S (Qi) (ω) (ω: angular frequency, ω=2πf, f: frequency). Then, the discrete Fourier transform processing unit 32-i converts the acquired frequency domain data S (Qi) (ω) is output to the time-frequency filter processing unit 4 for the full-range sound field.
[0054] In step S5, all audio signals s input to the sound field control device 100 are (Qi) (t) (in this embodiment, audio signal s (Q1) (t) ~ s (Q4) (t)), it is determined whether or not the loop process (loop 1) has been executed, and (1) all audio signals s input to the sound field control device 100 are (Qi) If it is determined that the loop process (loop 1) has been executed for (t), the process proceeds to step S6, and (2) all audio signals s input to the sound field control device 100 are (Qi) If the loop process (loop 1) has not been executed for (t), the process returns to step S3, and the loop process (loop 1) is repeatedly executed.
[0055] (Step S6): In step S6 (steps S61 to S67), a time-frequency filtering process for the entire sound field is executed. Specifically, the following process is executed.
[0056] (Step S61): In step S61, the loop process (loop 2 process) of the time-frequency filtering process for the full-range sound field is started (executed). (Qi) (ω) (in this embodiment, S (Q1) (ω) ~ S (Q4) Since the sound field control system 1000 performs processing using discrete values, the angular frequency ω is a discrete value ω of the angular frequency. k (ω k : integer, N: natural number, -N≦ω k ≦N, n=ω k ) In other words, the loop 2 processing is executed for integer n (-N≦n≦N, N: natural number) (the loop 2 processing is executed by incrementing n by +1 from n=-N until n=N).
[0057] (Step S62): In step S62, loop processing (loop 3 processing) is started (executed). Note that loop 3 processing is executed for each region Qi (Q1 to Q4 in this embodiment).
[0058] (Step S63): In step S63, a process for acquiring a two-dimensional cylindrical harmonic spectrum of the region Qi is executed. Specifically, the following process is executed.
[0059] For ease of explanation, the following description will be given assuming that the window function is set to a rectangle in all regions (regions Q1 to Q4). In region Qi, the two-dimensional cylindrical harmonic spectrum when the window function is set to a rectangle is expressed as a n,rect (Φ, φ S ) and the two-dimensional cylindrical harmonic spectrum a n,rect (Φ, φ S This article explains how to obtain this information.
[0060] If the sound pressure in the direction where sound can be heard is "1" and the sound pressure in the direction where sound cannot be heard is "0", then when the window function is set to a rectangular window, the sound pressure from the center point O to the radius r ref Sound pressure (sound pressure distribution) on the circumference of P rect (φ) (see FIG. 5) can be expressed mathematically as follows (and can be modeled as a square wave): φ S : Azimuth angle of the center of the square wave Φ: Width of the square wave And the sound pressure P expressed by the above formula rect By performing a spatial Fourier transform on (φ), the two-dimensional cylindrical harmonic spectrum a n,rect (Φ, φ S ) (2D cylindrical harmonic spectrum a when the window function is set to a rectangular window n,rect (Φ, φ S ) is obtained by the following formula: φ S : azimuth angle of the center of the rectangular wave (corresponding to the rectangular portion of the rectangular window) Φ: length of the rectangular wave (corresponding to the rectangular portion of the rectangular window) n: value indicating the spatial frequency (n: integer, -N≦n≦N) The time frequency filter processing unit 4 for the full-area sound field determines the azimuth angle of the center of the rectangular wave (corresponding to the rectangular portion of the rectangular window) of the area Qi as φ S(Qi) and the length of the rectangular wave in the region Qi (corresponding to the rectangular portion of the rectangular window) is Φ (Qi) Then, Φ = Φ (Qi) , φ S =φ S (Qi) and perform processing corresponding to the above formula to obtain the two-dimensional cylindrical harmonic spectrum a n,rect (Φ (Qi) , φ S (Qi) ) to obtain the
[0061] (Step S64): In step S64, the audio signal s to be reproduced in the area Qi is (Qi) Angular frequency n (= ω k ) frequency component data S n (Qi) Specifically, the time-frequency filter processing unit 4 for the full-range sound field performs the acquisition process of the data S output from the short-time Fourier transform processing unit 3A. (Qi) (ω) is input, and the data S (Qi) (ω), data S (Qi) S contained in (ω) n (Qi) By extracting (taking out) the audio signal s to be reproduced in the region Qi, (Qi) Angular frequency n (= ω k ) frequency component data S n (Qi) Get.
[0062] (Step S65): In step S65, it is determined whether the termination condition for loop 3 processing is satisfied. If the result of this determination shows that the termination condition for loop 3 processing is not satisfied (if it is determined that loop 3 processing has not been executed for all regions Qi), the process returns to step S62, and loop 3 processing (steps S62 to S65) is executed. On the other hand, if the result of the above determination shows that the termination condition for loop 3 processing is satisfied (if it is determined that loop 3 processing has been executed for all regions Qi), the process proceeds to step S66.
[0063] (Step S66): In step S66, the two-dimensional cylindrical harmonic spectrum a for the full range sound field is calculated. n,rect(all) (r ref Specifically, the time-frequency filter processing unit 4 for the entire sound field performs the acquisition process of the two-dimensional cylindrical harmonic spectrum a n,rect (Φ (Qi) , φ S (Qi) ) and frequency component data S of angular frequency n of the audio signal to be reproduced in the region Qi. n (Qi) By using the above formula, a two-dimensional cylindrical harmonic spectrum a for a full range sound field is obtained by performing a process corresponding to the following formula: n,rect (all) (r ref ) to obtain the Q: A set of regions Qi, Q = {Q1, Q2, Q3, Q4} S n (Qi) Angular frequency n (=ω) of the audio signal to be reproduced in the reproduction area Qi k ) frequency component data (data Q (Qi1) (ω) element) (-N≦n≦N) Φ (Qi) φ: Length of the rectangular wave in region Qi (corresponding to the rectangular portion of the rectangular window) S (Qi) : Azimuth angle of the center of the rectangular wave of the region Qi (corresponding to the rectangular portion of the rectangular window) In other words, the time-frequency filter processing unit 4 for the whole region sound field generates a two-dimensional cylindrical harmonic spectrum a for the sound field (reproduced sound field) of the whole region (regions Q1 to Q4 in this embodiment) by the above processing. n,rect (all) (r ref That is, the time frequency filter processing unit 4 for the whole area sound field obtains the two-dimensional cylindrical harmonic spectrum a for the whole area sound field for reproducing the sound field of the whole area (areas Q1 to Q4 in this embodiment) by performing a product-sum operation between the time frequency components of the audio signal to be reproduced in the divided areas of the whole area (areas Q1 to Q4 in this embodiment) and the spatial frequency components of the reproduced sound field of the respective areas (corresponding to the two-dimensional cylindrical harmonic spectrum). n,rect (all) (r refThat is, for a sound field in which the entire region is divided into M regions (M: a natural number equal to or greater than 2, in this embodiment, M=4), the whole region sound field time-frequency filter processing unit 4 obtains a whole region sound field two-dimensional cylindrical harmonic spectrum a n,rect (all) (r ref ) to obtain the
[0064] In steps S62 to S66, the processing equivalent to (Equation 9) is performed by performing the processing as described above, but the present invention is not limited to this. For n, the time frequency components of the audio signal to be reproduced in the divided regions of the entire region (regions Q1 to Q4 in this embodiment) are multiplied by the spatial frequency components of the reproduced sound field of that region (corresponding to the two-dimensional cylindrical harmonic spectrum), and the obtained values are added up each time, and when all n have been processed, the results of the above two product-sum operations are obtained.
[0065] (Step S67): In step S67, it is determined whether the termination condition for loop 2 processing is satisfied. If the result of this determination shows that the termination condition for loop 2 processing is not satisfied (if it is determined that loop 2 processing has not been executed for all n (n: integer, -N≦n≦N)), the process returns to step S61, and loop 2 processing (steps S61 to S67) is executed. On the other hand, if the result of the above determination shows that the termination condition for loop 2 processing is satisfied (if it is determined that loop 2 processing has been executed for all n (n: integer, -N≦n≦N)), the process proceeds to step S67.
[0066] (Step S68): In step S68, a process of acquiring an excitation signal in the time-frequency domain is executed. Specifically, the following process is executed.
[0067] The time-frequency filter processing unit 4 for the full-range sound field uses the two-dimensional cylindrical harmonic spectrum a for the full-range sound field acquired in steps S61 to S67. n,rect (Φ (Qi) , φ S (Qi)) (n: integer, −N≦n≦N) (2D cylindrical harmonic spectrum for full range sound field a (all) (r ref )) to perform processing corresponding to the following formula, the excitation signal D^ in the time-frequency domain (all) (Drive signals (signals in the time-frequency domain) for driving L speakers of the speaker array Spk_arry) are acquired. X H : Hermitian transpose matrix of matrix X (adjoint matrix) I: unit matrix (unit matrix of L × L) λ: coefficient Note that the two-dimensional cylindrical harmonic spectrum a for the full range sound field (all) (r ref ) is as follows: Then, the time-frequency filter processing unit 4 for the full-range sound field generates a driving signal (a signal in the time-frequency domain) D^ for driving the L speakers of the speaker array Spk_arry acquired by the above processing. (all) (= [D 1 (all) (r 1 ), ..., D L (all) (r L )] T = [D 1 (all) (ω, r 1 ), ..., D L (all) (ω, r L )] T ) is output to the short-time inverse Fourier transform processing unit 5.
[0068] (Step S7): In step S7, a short-time inverse Fourier transform process is performed. Specifically, the following process is performed.
[0069] The inverse discrete Fourier transform processing unit 51-i (i: natural number, 1≦i≦L) converts the time-frequency domain data D of the drive signal output from the full-range sound field time-frequency filter processing unit 4. i (all) (ω) (=D i (all) (r i ) = D i (all) (ω, r i )) (D i(all) (ω) is a 1×L matrix D^ (all) The i-th element (i-th column) of the data D i (all) That is, the inverse discrete Fourier transform processing unit 51-i performs processing corresponding to the following equation to generate a signal d i (t) (time domain drive signal d i (t)) is obtained. Then, the inverse discrete Fourier transform processing unit 51-1 converts the signal (time domain signal) d obtained by the inverse discrete Fourier transform as described above into i (t) is output to the time domain window function processing unit 52-i.
[0070] The time domain window function processor 52-i (i: natural number, 1≦i≦L) receives the signal (time domain signal) d after the inverse discrete Fourier transform output from the inverse discrete Fourier transform processor 51-i. i (t) is input, and the signal d i (t) by a time domain window function (time domain window function processing) (for example, by applying the window function g(t) of (Equation 5) to the signal d i (t)), and the signal obtained by this process is converted into the signal sig_d i (t)(=g(t)*d i (t) (*: operator indicating convolution processing)) and output to the i-th speaker Spki of the speaker array Spk_arry.
[0071] (Step S8): In step S8, the i-th speaker Spki of the speaker array Spk_arry receives the drive signal sig_d obtained in step S7. i (1) By being driven by (t), an audio signal s is input to the reproduction area Zone1 (Q1). (Q1) (t) is reproduced, and (2) an audio signal s is transmitted to the reproduction area Zone2 (Q2). (Q2) (t) is reproduced, and (3) the audio signal s is transmitted to the reproduction area Zone 3 (Q3). (Q3) (t) is reproduced, and (4) the audio signal s is transmitted to the reproduction area Zone4 (Q4). (Q4)(t) is reproduced.
[0072] As described above, in the sound field control system 1000, the total number of regions is M (M is a natural number equal to or greater than 2, in this embodiment, M=4) (the audio signals to be reproduced in each of the M regions are s (Q1) (t) ~ s (QM) For a sound field divided into M divided regions (a sound field of a whole region having M divided regions), a two-dimensional cylindrical harmonic spectrum for the whole region sound field is obtained to reproduce the sound field (a sound field of a whole region having M divided regions), and a time-frequency filter process is performed using the obtained two-dimensional cylindrical harmonic spectrum for the whole region sound field, thereby obtaining a driving signal D^ in the time-frequency domain. (all) (drive signals (time-frequency domain signals) for driving L speakers of the speaker array Spk_arry) are acquired. Then, the sound field control system 1000 acquires the acquired time-frequency domain drive signals D^ (all) By performing a short-time Fourier transform process on the L speaker arrays Spk_arry, signals (time-domain drive signals) for driving each of the L speaker arrays Spk_arry are obtained. In other words, the sound field control system 1000 adopts a method of setting a sound field in which the entire area is divided into M areas and reproducing the sound field, so that the amount of calculation is dramatically reduced compared to a method (local sound field superposition method) in which a divided sound field (local sound field) is set for each of the M divided areas and drive signals acquired for each local sound field are superposed. In the local sound field superposition method, a divided sound field (local sound field) is set for each of the M divided areas and time-frequency filter processing is performed using the two-dimensional cylindrical harmonic spectrum acquired for each local sound field, thereby generating a time-frequency domain drive signal D^ (Qi) and superimposing the drive signals of the acquired local sound fields to obtain the drive signal D^ for the entire area. (all)Therefore, the number of calculations for the filtering process in the time frequency filtering process is M×L (L: number of speakers, L is a natural number equal to or greater than 2). In contrast, in the sound field control system 1000, a sound field is set in which the entire area is divided into M areas, and a method is adopted to reproduce the sound field. Therefore, the number of calculations for the filtering process in the time frequency filtering process is only L (in Equation 10, the L×2N+1 matrix ((C(r ref ) H C(r ref ) + λI) -1 C(r ref ) H ) and a 2N+1×1 matrix (a (all) (r ref ) (the number of operations required to obtain the product of the two matrices is only L for 2N+1 elements).
[0073] Furthermore, in the sound field control system 1000, regardless of the type of signals that the M audio signals are input to, the sound pressure distribution of the M audio signals can be quickly acquired (calculated) using short-time Fourier transform processing, enabling high-speed processing.
[0074] As described above, the sound field control system 1000 can perform highly accurate multi-spot reproduction while reducing the number of calculations.
[0075] <<First Modification>> Next, a first modification of the first embodiment will be described. Note that the same parts as those in the above embodiment are denoted by the same reference numerals, and detailed description thereof will be omitted.
[0076] FIG. 6 is a schematic diagram of a sound field control system 1000 according to a first modified example of the first embodiment.
[0077] The sound field control system 1000A of this modification has a configuration in which the sound field control device 100 in the sound field control system 1000 of the first embodiment is replaced with a sound field control device 100A.
[0078] The sound field control device 100A has a configuration in which, in the sound field control device 100 of the first embodiment, the short-time Fourier transform processing unit 3 is replaced with a short-time Fourier transform processing unit 3A, the full-area sound field time frequency filter processing unit 4 is replaced with a full-area sound field time frequency filter processing unit 4A, and a silent area setting unit 6 is further added.
[0079] The silent region setting unit 6 receives, for example, externally input data Info_dark_Q specifying a silent region, and generates data D_dark_Q instructing that silent region processing be performed on the region specified by the data Info_dark_Q. The silent region setting unit 6 then outputs the data D_dark_Q to the short-time Fourier transform processing unit 3A and the full-range sound field time-frequency filter processing unit 4A.
[0080] The short-time Fourier transform processing unit 3A has the same functions as the short-time Fourier transform processing unit 3 of the first embodiment, and further has a function of inputting data D_dark_Q output from the silent region setting unit 6 and performing the following processing. When the data D_dark_Q is input, the short-time Fourier transform processing unit 3A identifies a region designated as a silent region by the data D_dark_Q (this region is called a silent designated region Q_dark), and generates an audio signal s corresponding to the silent designated region Q_dark. (Q_dark) (t) is forced to "0", or the audio signal s (Q_dark) Data after discrete Fourier transform of (t) (frequency domain data) S (Q_dark) A process is performed to forcibly set (ω) to "0".
[0081] The time-frequency filter processing unit 4A for the entire sound field has the same functions as the time-frequency filter processing unit 4 for the entire sound field in the first embodiment, and further has a function of inputting the data D_dark_Q output from the silent region setting unit 6 and performing the following processing. When the data D_dark_Q is input, the time-frequency filter processing unit 4A for the entire sound field calculates the two-dimensional cylindrical harmonic spectrum a n,rect (Φ (Qi) , φ S (Qi)) (Qi=Q_dark) is forcibly set to "0".
[0082] The operation of the sound field control system 1000A configured as above will be described below. Note that the description of the same parts as those in the first embodiment will be omitted.
[0083] 7 and 8 are flowcharts of the processing executed in the sound field control system 1000A of the first modified example of the first embodiment.
[0084] The processing executed by the sound field control system 1000A will be described below with reference to the flowcharts of FIGS.
[0085] (Steps S1 to S2): In steps S1 to S2, the same processes as in steps S1 to S2 in the first embodiment are executed.
[0086] (Steps S3A to S5A): In step S3A, the same process as in step S3 is executed.
[0087] In step S4A, the same process as in step S4 is executed. However, when data D_dark_Q is input from the silent region setting unit 6 to the short-time Fourier transform processing unit 3A, a region designated as a silent region by the data D_dark_Q (this region is called a silent designated region Q_dark) is identified, and the audio signal s corresponding to the silent designated region Q_dark is generated. (Q_dark) A process is performed to forcibly set (t) to "0".
[0088] In step S5A, the same process as in step S5 is executed.
[0089] (Step S6A): In step S6A, a time-frequency filtering process for the entire sound field is executed. Specifically, the following process is executed.
[0090] (Steps S6A1 to S6A2): In steps S6A1 to S6A2, the same processes as in steps S61 to S62 are executed.
[0091] (Steps S6A1 to S6A2): In steps S6A1 to S6A2, the same processes as in steps S61 to S62 are executed.
[0092] (Step S6A3): In step S6A3, a determination process is executed to determine whether the data to be processed is data in region Qi that is set as a silent region. If the result of the determination process indicates that the data to be processed is data in region Qi that is set as a silent region, the process proceeds to step S6A5. If the data to be processed is not data in region Qi that is set as a silent region, the process proceeds to step S6A4.
[0093] (Steps S6A4, S6A6): In steps S6A4 and S6A6, the same processes as in steps S63 and S64 are executed.
[0094] (Step S6A5): In step S6A5, silent area setting processing is executed for the area Qi (=area Q_dark set as a silent area). Specifically, the following processing is executed.
[0095] When the data D_dark_Q is input from the silent region setting unit 6, the time-frequency filter processing unit 4A for the entire sound field calculates the two-dimensional cylindrical harmonic spectrum a n,rect (Φ (Qi) , φ S (Qi) ) (Qi=Q_dark) is forcibly set to "0".
[0096] Alternatively, when data D_dark_Q is input from the silent region setting unit 6, the short-time Fourier transform processing unit 3A outputs frequency component data S of angular frequency n of the audio signal reproduced in the region Qi (=region Q_dark set as the silent region). n (Qi) A process is performed to forcibly set (Qi=Q_dark) to "0".
[0097] (Steps S6A7 to S6A10): In steps S6A7 to S6A10, the same processes as in steps S65 to S68 are executed, respectively.
[0098] By performing the above processing, the sound field control system 1000A can set the area Qi specified by the data Info_dark_Q as a silent area, while reproducing the sound of a predetermined audio signal in areas other than the silent area. Since the sound field control system 1000A can set the area Qi specified by the data Info_dark_Q as a silent area, when the sound field control system 1000A is used, for example, in a multilingual simultaneous interpretation conference, by setting the area where the speaker is located (the direction of the speaker) as a silent area, the sound does not reach the speaker, dramatically improving the ease with which the speaker can speak. For example, when the playback areas are set as shown in Figure 2, a speaker (e.g., a Chinese speaker) is present in the playback area Q3, and (1) the audio signal s is generated by translating the Chinese of the speaker into Japanese. (Q1) (t), and (2) the Chinese speech of the speaker translated into English is taken as an audio signal s (Q2) (t), and (3) the Chinese speech of the speaker translated into Korean is used as the audio signal s (Q4) In the case where (t) is assumed, in the sound field control system 1000A, the playback area Q3 where the speaker (e.g., a Chinese speaker) is located is set to a silent area Q_dark, and by performing the above processing, the playback area Q3 where the speaker (e.g., a Chinese speaker) is located is made a silent area (an area where sound does not reach), while the translated speech can be heard in the other areas (Japanese speech can be reproduced in area Q1, English speech in area Q2, and Korean speech in area Q4). In this way, when the sound field control system 1000A is used in a multilingual simultaneous interpretation conference, by setting the area where the speaker is located (the direction of the speaker) to a silent area, the speech does not reach the speaker, dramatically improving the ease with which the speaker can speak.
[0099] In the process of step S6A, the two-dimensional cylindrical harmonic spectrum a n,rect (Φ (Qi) , φ S (Qi) ) (Qi=Q_dark) is forcibly set to "0", and frequency component data S of angular frequency n of the audio signal reproduced in region Qi (=region Q_dark set as a silent region) n (Qi)It is also possible to perform both processes of forcibly setting (Qi=Q_dark) to "0".
[0100] The process of step S4A may be the same as the process of step S4 (the audio signal s corresponding to the silent designated region Q_dark). (Q_dark) (It is also possible to avoid the process of forcibly setting (t) to "0").
[0101] Second Embodiment Next, a second embodiment will be described. Note that the same parts as those in the above embodiment (including the modified example) are given the same reference numerals, and detailed description thereof will be omitted.
[0102] In the sound field control system of the second embodiment, processing using a trainable model (for example, a neural network model) is executed in the time-frequency filtering process for the full-area sound field.
[0103] FIG. 9 is a schematic diagram of a sound field control system 2000 according to the second embodiment (during learning processing).
[0104] FIG. 10 is a schematic diagram of the configuration of the all-area sound field time-frequency filter processing unit 4B (during learning processing) of the sound field control system 2000 according to the second embodiment.
[0105] FIG. 11 is a flowchart of the learning process executed by the sound field control system 2000 according to the second embodiment.
[0106] FIG. 12 is a schematic diagram of a sound field control system 2000 (during inference processing) according to the second embodiment.
[0107] FIG. 13 is a schematic diagram of the configuration of the all-area sound field time-frequency filter processing unit 4B (at the time of inference processing) of the sound field control system 2000 according to the second embodiment.
[0108] <2.1: Configuration of sound field control system> As shown in Figure 9, the sound field control system 2000 of the second embodiment has a configuration in which the full-area sound field time frequency filter processing unit 4 in the sound field control system 1000 of the first embodiment is replaced with a full-area sound field time frequency filter processing unit 4B.
[0109] As shown in Figure 10, the time-frequency filter processing unit 4B for the full-area sound field includes a reproduction sound pressure distribution setting processing unit 4B1, a learnable model 4B2, a transfer function processing unit 4B3, a reproduction sound pressure distribution acquisition processing unit 4B4, and a loss evaluation unit 4B5.
[0110] The reproduction sound pressure distribution setting processing unit 4B1 receives the data D_settings_zone(Q) output from the reproduction area setting unit 1 and the time frequency component data S of the audio signal output from the short time Fourier transform processing unit 3. (Q1) (ω) ~ S (QM) (ω) (in this embodiment, M=4). The reproduction sound pressure distribution setting processing unit 4B1 receives the data D_settings_zone(Q) and the time frequency component data S of the audio signal. (Q1) (ω) ~ S (QM) Based on (ω), the correct data P of the reproduced sound pressure distribution (correct) (The correct sound pressure distribution data P (correct) Then, the reproduction sound pressure distribution setting processing unit 4B1 generates the correct data P (correct) The data including the above is output to the loss evaluation unit 4B5 as data D_P_correct.
[0111] The trainable model 4B2 includes a first linear processing unit 4B21, a second linear processing unit 4B22, and a drive signal acquisition processing unit 4B23.
[0112] The first linear processing unit 4B21 has M input channels (corresponding to the number of input nodes) (M=4 in this embodiment) and N output channels (the number of output nodes). f (N f The first linear processing unit 4B21 includes a linear layer (complex linear layer) (for example, a layer of a neural network model consisting of one or more layers (for example, a fully connected layer)) in which M (M: a natural number of 2 or more, M is the number of divided regions (in this embodiment, M=4)) frequency domain data S of the audio signal output from the short-time Fourier transform processing unit 3. (Q1) (ω) ~ S (QM)(ω) (corresponding to data of an M×1 matrix) is input, and the input data is processed by a linear layer (complex linear layer) (the weight matrix (parameter) of the linear layer (complex linear layer) of the first linear processing unit 4B21 is set to W (1st) Then, the first linear processing unit 4B21 performs the above-described processing (N output to the output node) (which may include a process of adding a bias and a process using an activation function). f pieces of data) are output as data D4a to the second linear processing unit 4B22.
[0113] The second linear processing unit 4B22 is configured to process a linearly interpolating input signal when the number of input channels (corresponding to the number of input nodes) is N. f and includes a linear layer (complex linear layer) (for example, a layer of a neural network model consisting of one or more layers (for example, a fully connected layer)) whose output channels (number of output nodes) are L (the number of speakers of the speaker array Spk_arry). The second linear processing unit 4B22 receives the data D4a output from the first linear processing unit 4B21, and processes the input data by the linear layer (complex linear layer) (by setting the weight matrix (parameter) of the linear layer (complex linear layer) of the second linear processing unit 4B22 to W (2nd) Then, the second linear processing unit 4B22 outputs the processed data (L pieces of data output to the output node) as data D4b to the drive signal acquisition processing unit 4B23.
[0114] The drive signal acquisition processing unit 4B23 inputs data D4b output from the second linear processing unit 4B22 and acquires this data D4b as frequency component data D of the drive signal (L pieces of data (corresponding to an L×1 matrix of data) acquired by the second linear processing unit 4B22). The drive signal acquisition processing unit 4B23 then outputs data including the frequency component data D of the drive signal as data D4c to the transfer function processing unit 4B3.
[0115] The transfer function processing unit 4B3 receives data D_C output from the transfer function acquisition processing unit 2 and data D4c output from the drive signal acquisition processing unit 4B23. The transfer function processing unit 4B3 acquires transfer function G (transfer function G in the spatial frequency domain) from data D_C, and acquires frequency component data D (corresponding to data in an L×1 matrix) (time frequency component data D) of the drive signal from data D4c. The transfer function processing unit 4B3 then acquires sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) by performing processing equivalent to P = G × D. The transfer function processing unit 4B3 then outputs data including the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) as data D4d to the reproduction sound pressure distribution acquisition processing unit 4B4.
[0116] The reproduction sound pressure distribution acquisition processing unit 4B4 inputs the data D4d output from the transfer function processing unit 4B3 and acquires the sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) included in the data D4d. The reproduction sound pressure distribution acquisition processing unit 4B4 then outputs the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) to the loss evaluation unit 4B5 as data D4e.
[0117] The loss evaluation unit 4B5 receives as input the data D_P_correct output from the reproduction sound pressure distribution setting processing unit 4B1, the data D4e output from the reproduction sound pressure distribution acquisition processing unit 4B4, and the frequency component data D of the drive signal corresponding to the audio signal output from the drive signal acquisition processing unit 4B23. The loss evaluation unit 4B5 receives the data D_P_correct from the data D_P_correct and calculates the correct data P of the reproduction sound pressure distribution included in the data. (correct) (The correct sound pressure distribution data P (correct) ) and also obtains sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) included in the data D4e from the data D4e. Then, the loss evaluation unit 4B4 calculates the correct data P of the reproduced sound pressure distribution using a predetermined loss function. (correct) (The correct sound pressure distribution data P (correct)) and the sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain). The loss evaluation unit 4B5 then backpropagates the acquired loss (error) by, for example, the error backpropagation method, and executes a process of updating the parameters of the trainable model 4B2.
[0118] <2.2: Operation of Sound Field Control System> The operation of the sound field control system 2000 configured as above will be described.
[0119] The operation of the sound field control system 2000 will be explained below by dividing it into (1) learning processing and (2) inference processing.
[0120] (2.2.1: Learning Process) First, the learning process executed by the sound field control system 2000 will be described with reference to the flowchart of FIG.
[0121] (Step SB1): In step SB1, a transfer function setting process is executed. Specifically, the following process is executed.
[0122] The transfer function acquisition processing unit 2 acquires a transfer function of each of the L speakers constituting the speaker array Spk_arry and a radius r ref Then, data including the two-dimensional cylindrical harmonic spectral matrix C obtained by this process is output as data D_C to the transfer function processing unit 4B3 of the time-frequency filter processing unit 4B for the full-range sound field.
[0123] The transfer function processing unit 4B3 calculates a two-dimensional cylindrical harmonic spectral matrix C (=C(r ref )) into the transfer function G(=C(r ref )) (see Equation 2).
[0124] (Step SB2): In step SB2, an initial value setting process is executed to set a variable j to 1. Note that the variable j corresponds to a frame (audio frame) number.
[0125] (Step SB3): In step SB3, the time-frequency component data of the audio signal of the j-th frame is acquired (input). Specifically, the following process is performed.
[0126] The audio signal s of the jth frame (jth frame) to be reproduced in the regions Q1 to QM (Q1) (t) ~ s (QM) (t) is input to the short-time Fourier transform processing unit 3, and processed by the short-time Fourier transform processing unit 3 to obtain the audio signal s (Q1) (t) ~ s (QM) Data (time frequency component data) S after short-time Fourier transform processing of (t) (Q1) (ω) ~ S (QM) (ω) is obtained. Then, the obtained audio signal s (Q1) (t) ~ s (QM) Data (time frequency component data) S after short-time Fourier transform processing of (t) (Q1) (ω) ~ S (QM) (ω) is input to the first linear processing section 4B21 and the reproduction sound pressure distribution setting processing section 4B1 of the trainable model 4B2 of the time-frequency filter processing section 4B for the full-range sound field.
[0127] (Step SB4): In step SB4, a reproduction sound pressure distribution setting process is executed. Specifically, the following process is executed.
[0128] The reproduction area setting unit outputs data D_settings_zone(Q) including information about the reproduction area to the reproduction sound pressure distribution setting processing unit 4B1 of the full-area sound field time-frequency filter processing unit 4B. Note that in this embodiment, the reproduction area Q will be described as being set to Q1, Q2, Q3, and Q4 as shown in FIG.
[0129] The reproduction sound pressure distribution setting processing unit 4B1 receives the data D_settings_zone(Q) and the time frequency component data S of the audio signal. (Q1) (ω) ~ S (QM) Based on (ω), the correct data P of the reproduced sound pressure distribution (correct) (The correct sound pressure distribution data P (correct)) (correct data P (correct) P j (correct) In this embodiment, local reproduction is performed in the reproduction region (each of the divided regions Q1 to Q4) using a spatial Hann window (spatial Hann function).
[0130] As explained in the first embodiment, if the sound pressure in the direction where sound can be heard is "1" and the sound pressure in the direction where sound cannot be heard is "0", then the sound pressure from the center point O to the radius r ref Sound pressure (sound pressure distribution) on the circumference of P hann (φ) can be expressed by the following formula (can be modeled by the following formula): The sound pressure P hann By performing a spatial Fourier transform on (φ), the two-dimensional cylindrical harmonic spectrum a n,hann (Φ, φ S ) (2D cylindrical harmonic spectrum a when the window function is set to Hann window n,hann (Φ, φ S )) is obtained by the following formula: φ S : azimuth angle of the center of the waveform part (part where the value is not 0) of the Hann window Φ: length of the waveform part (part where the value is not 0) of the Hann window n: value indicating the spatial frequency (n: integer, -N≦n≦N) The reproduction sound pressure distribution setting processing unit 4B1 calculates (1) time frequency component data S of the audio signal of the region Qi. (Qi) (ω) (i: integer, 1≦i≦M, M=4), and (2) the two-dimensional cylindrical harmonic spectrum a when the window function of the sound pressure in the region Qi is set to the Hann window. n,hann (Φ (Qi) , φ S (Qi) ) and the sum of the whole area Q (= {Q1, Q2, Q3, Q4}) is added together to obtain the correct data P (correct) (Sound pressure distribution data in the spatial frequency domain P (correct) That is, the reproduction sound pressure distribution setting processing unit 4B1 performs processing corresponding to the following formula to obtain the obtained a(all) (r ref ) is used as the correct data P (correct) (Sound pressure distribution data in the spatial frequency domain P (correct) ) is obtained. Q: A set of regions Qi, Q = {Q1, Q2, Q3, Q4} S n (Qi) Angular frequency n (=ω) of the audio signal to be reproduced in the reproduction area Qi k ) frequency component data (data Q (Qi1) (ω) element) (-N≦n≦N) Φ (Qi) φ: Length of the waveform portion (non-zero portion) of the Hann window in region Qi S (Qi) : azimuth angle of the center of the waveform part (part where the value is not 0) of the Hann window of the region Qi; n: value indicating the spatial frequency (n: integer, −N≦n≦N); and the reproduction sound pressure distribution setting processing unit 4B1 sets the correct data P of the reproduction sound pressure distribution data acquired by the above processing. (correct) (=P j (correct) ) is output to the loss evaluation unit 4B5 as data D_P_correct.
[0131] (Step SB5): In step SB5, a drive signal acquisition process is executed using the trainable model 4B2. Specifically, the following process is executed.
[0132] The first linear processing unit 4B21 calculates the frequency component data S of the audio signal of the j-th frame to be reproduced in the regions Q1 to Q4. (Q1) (ω) ~ S (Q4) (ω) is input to each of M nodes of the input node, and processed by the linear layer (complex linear layer) (weight matrix (parameter) of the linear layer (complex linear layer) of the first linear processing unit 4B21): W (1st) ) (which may include a process of adding a bias and a process using an activation function). Then, the first linear processing unit 4B21 performs the above-described process on the data (N output to the output node). f pieces of data) are output as data D4a to the second linear processing unit 4B22.
[0133] The second linear processing unit 4B22 receives the data D4a output from the first linear processing unit 4B21, and performs processing on the received data using a linear layer (complex linear layer) (weight matrix (parameter): W (2nd) ) (which may include processing to add a bias and processing using an activation function). The second linear processing unit 4B22 then outputs the processed data (L pieces of data output to the output node) as data D4b to the drive signal acquisition processing unit 4B23.
[0134] The drive signal acquisition processing unit 4B23 receives the data D4b output from the second linear processing unit 4B22 and converts the data D4b into frequency component data D of the drive signal (L pieces of data (corresponding to data in an L×1 matrix) acquired by the second linear processing unit 4B22) (the frequency component data D of the drive signal corresponding to the audio signal of the j-th frame is D j The drive signal acquisition processing unit 4B23 then outputs data including the frequency component data D of the drive signal as data D4c to the transfer function processing unit 4B3.
[0135] (Step SB6): In step SB6, a reproduction sound pressure distribution acquisition process is executed. Specifically, the following process is executed.
[0136] The transfer function processing unit 4B3 calculates P=G×D (P using the transfer function G (transfer function G in the spatial frequency domain) acquired in step SB1 and the frequency component data D (corresponding to data of an L×1 matrix) (time frequency component data D) of the drive signal included in the data D4c. j = G x D j ) is executed to obtain the sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) (sound pressure distribution data P corresponding to the audio signal of the j-th frame) j Then, the transfer function processing unit 4B3 acquires the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) (=P j ) is output as data D4d to the reproduction sound pressure distribution acquisition processing section 4B4.
[0137] The reproduction sound pressure distribution acquisition processing unit 4B4 inputs the data D4d output from the transfer function processing unit 4B3 and acquires the sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) included in the data D4d. The reproduction sound pressure distribution acquisition processing unit 4B4 then outputs the acquired sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) to the loss evaluation unit 4B5 as data D4e.
[0138] (Step SB7): In step SB7, a loss evaluation process is executed. Specifically, the following process is executed.
[0139] The loss evaluation unit 4B5 receives the data D_P_correct output from the reproduction sound pressure distribution setting processing unit 4B1 and the data D4e output from the reproduction sound pressure distribution acquisition processing unit 4B4. The loss evaluation unit 4B5 receives the data D_P_correct and the correct data P (correct) (The correct sound pressure distribution data P (correct) ) (=P j (correct) ) is acquired, and sound pressure distribution data P (sound pressure distribution data P in the spatial frequency domain) (=P j ) of the drive signal corresponding to the audio signal of the j-th frame output from the drive signal acquisition processing unit 4B23. j ) is acquired. The loss evaluation unit 4B5 acquires the sound pressure distribution data P j And the correct sound pressure distribution data P j (correct) The following information shall be stored and maintained.
[0140] Then, the loss evaluation unit 4B5 obtains the loss (error) using the following loss evaluation function: mean |x|: operator for obtaining the mean value of |x| (mean value of the norm (for example, absolute value) of x) ||x||: norm of x λ: coefficient The loss evaluation unit 4B5 performs processing corresponding to the first term of the above equation on the sound pressure distribution data P j And the correct sound pressure distribution data P j (correct)and the sound pressure distribution data P obtained for the j-th frame audio signal. j And the correct sound pressure distribution data P j (correct) Using the above, the sound pressure distribution data P j and the correct sound pressure distribution data P j (correct) This is done by taking the average of the norms of the differences between
[0141] The second term in the above equation is a regularization term, and the loss evaluation unit 4B5 calculates the frequency component data D (=D j ) to perform processing equivalent to the second term of the above equation, or by obtaining the L1 norm of the frequency component data D (=D j ) to perform processing equivalent to the second term in the above equation.
[0142] (Step SB8): In step SB8, the variable j is set to the number of frames N that have been subjected to the learning process. F As a result of the determination, j≧N is determined. F If j≧N, the process returns to step SB2. F If not, the process proceeds to step SB9.
[0143] (Step SB9): In step SB9, the variable j is incremented by +1 (j←j+1).
[0144] (Step SB10): In step SB10, a convergence determination process is performed based on the loss (Loss) acquired in step SB7. Specifically, if the loss (Loss) acquired in step SB7 is equal to or less than a predetermined value, or if the amount of fluctuation in the loss (Loss) acquired in step SB7 is equal to or less than a predetermined value, it is determined that the learning process has converged. If it is determined that the learning process has converged, the process proceeds to step SB12, and if it is determined that the learning process has not converged, the process proceeds to step SB11.
[0145] (Step SB11): In step SB11, the error (loss) acquired in step SB7 is backpropagated by, for example, an error backpropagation method (backpropagated to the reproduction sound pressure distribution acquisition processing unit 4B4, the transfer function processing unit 4B3, the drive signal acquisition processing unit 4B23, the second linear processing unit 4B22, and the first linear processing unit 4B21), and the parameters of the trainable model 4B2 (parameters of the second linear processing unit 4B22 (weight matrix W (2nd) ), and the parameters of the first linear processing unit 4B21 (weighting matrix W (1st) (for example, a parameter update process is performed using the Adam optimization method).
[0146] (Step SB12): In step SB12, a trained model acquisition process is executed. Specifically, for example, when it is determined in step SB10 that the learning process has converged, the parameters set in the trainable model 4B2 are acquired as optimal parameters. Then, by setting the parameters of the trainable model 4B2 as the optimal parameters, the trainable model 4B2 becomes a trained model (a model whose parameters are fixed to the optimal parameters).
[0147] As described above, the learning process is executed in the sound field control system 2000.
[0148] (2.2.2: Inference Processing) Next, the inference processing executed in the sound field control system 2000 will be described.
[0149] In the inference process, as shown in Fig. 12, the trained model acquired in the above learning process (the trainable model 4B2 in which the optimal parameters acquired in the above learning process are set) is loaded into the time-frequency filter processing unit 4B for the full-range sound field of the sound field control system 2000. That is, in the inference process, as shown in Fig. 13, the first linear processing unit 4B21 and the second linear processing unit 4B22 of the trainable model 4B2 are loaded into the time-frequency filter processing unit 4B for the full-range sound field, and the optimal parameters (W opt (1st) , W opt (2nd) ) is set, the trained model 4B2 is installed and executed.
[0150] As in the first embodiment, the audio signal s to be reproduced in the regions Q1 to Q4 (Q1) (t) ~ s (Q4) (t) is input to the short-time Fourier transform processing unit 3. The short-time Fourier transform processing unit 3 executes the same processing as in the first embodiment, and outputs the audio signal s (Q1) (t) ~ s (Q4) Time frequency component data S of (t) (Q1) (ω) ~ S (Q4) (ω) is output to the time-frequency filter processing unit for the full range sound field 4B.
[0151] The first linear processing unit 4B21 converts the frequency component data S of the audio signal to be reproduced in the regions Q1 to Q4 output from the short-time Fourier transform processing unit 3. (Q1) (ω) ~ S (Q4) (ω) is input to each of M (M=4) nodes of the input node, and processed by the linear layer (complex linear layer) (weight matrix (optimal parameter) of the linear layer (complex linear layer) of the first linear processing unit 4B21): W opt (1st) ) (which may include a process of adding a bias and a process using an activation function). Then, the first linear processing unit 4B21 performs the above-described process on the data (N output to the output node). f pieces of data) are output as data D4a to the second linear processing unit 4B22.
[0152] The second linear processing unit 4B22 receives the data D4a output from the first linear processing unit 4B21, and performs processing on the received data using a linear layer (complex linear layer) (weight matrix (optimal parameters): W opt (2nd) ) (which may include processing to add a bias and processing using an activation function). The second linear processing unit 4B22 then outputs the processed data (L pieces of data output to the output node) as data D4b to the drive signal acquisition processing unit 4B23.
[0153] The drive signal acquisition processing unit 4B23 inputs the data D4b output from the second linear processing unit 4B22 and acquires the data D4b as frequency component data D of the drive signal (L pieces of data (corresponding to data in an L×1 matrix) acquired by the second linear processing unit 4B22). Then, the drive signal acquisition processing unit 4B23 calculates the frequency component data D of the drive signal (=[D 1 (all) (ω), D 2 (all) (ω), ..., D L (all) (ω) ] T ) into data D 1 (all) (ω) ~D L (all) (ω) is output to the short-time inverse Fourier transform processing unit 5.
[0154] In the processing after the short-time inverse Fourier transform processing unit 5, the same processing as in the first embodiment is executed.
[0155] As described above, in the sound field control system 2000, the time-frequency filter processing unit 4B for the full-area sound field performs processing using a trained model (e.g., a neural network model), thereby enabling highly accurate sound field control (multi-spot playback).
[0156] That is, in the sound field control system 2000, the total number of regions is M (M is a natural number equal to or greater than 2, in this embodiment, M=4) (the audio signals to be reproduced in each of the M regions are s (Q1) (t) ~ s (QM)For a sound field divided into M divided regions (t) (a sound field of the entire area having M divided regions), a sound pressure distribution for reproducing the sound field (a sound field of the entire area having M divided regions) is obtained as a correct sound pressure distribution, and a learnable model is trained so that frequency component data of a drive signal for outputting the correct sound pressure distribution is output. The trained model of the trainable model is then loaded into the time-frequency filter processing unit 4B for the entire area sound field, and the trained model is used to obtain frequency component data of the drive signal. The frequency component data of the obtained drive signal is subjected to an inverse short-time Fourier transform to obtain a time-domain drive signal, and the time-domain drive signal is used to drive L speakers of the speaker array Spk_arry, thereby enabling desired sound field control (desired multi-spot reproduction).
[0157] In the sound field control system 2000, the time-frequency filter processing unit 4B for the full-area sound field simply performs processing using a trained model, thereby enabling highly accurate multi-spot reproduction while reducing the number of calculations.
[0158] While the above description is directed to a case where the learning process is executed by the sound field control device 100B of the sound field control system 2000, the present invention is not limited to this. For example, the above learning process (the learning process described as being executed by the sound field control device 100B of the sound field control system 2000) may be executed by another device (e.g., a learning processing device), and optimal parameters acquired by the learning process executed by the other device (e.g., a learning processing device) may be set in the first linear processing unit 4B21 and the second linear processing unit 4B22 of the full-area sound field time-frequency filter processing unit 4B (thereby acquiring the trained model 4B2). In this case, the sound field control device 100B may be configured without including the reproduction area setting unit 1, the transfer function acquisition processing unit 2, the reproduction sound pressure distribution setting processing unit 4B1, the transfer function processing unit 4B3, the reproduction sound pressure distribution acquisition processing unit 4B4, and the loss evaluation unit 4B5.
[0159] In the above, the spatial window for setting the sound pressure distribution of each divided region (regions Q1 to QM) is a Hann window, but this is not limited to this, and the spatial window for setting the sound pressure distribution of each divided region (regions Q1 to QM) may be a rectangular window. Furthermore, the sound pressure distribution of each divided region (regions Q1 to QM) may be any distribution (sound pressure distribution).
[0160] Furthermore, although the above describes a case where the error is backpropagated for each frame in the learning process, this is not limiting. For example, the loss (error) (e.g., average error) obtained after executing the process (forward propagation process) for a predetermined number of frames may be backpropagated to perform the parameter update process (the error may be backpropagated using batch processing or mini-batch processing).
[0161] The loss function described above is an example, and other loss functions may be used (for example, P j -P j (correct) Alternatively, the L1 norm or L2 norm of a vector w having elements of all updatable parameters (weights) may be used in the regularization term.
[0162] In the second embodiment, a neural network model is used as the learning model, but the learning process may be performed without using an activation function.
[0163] Furthermore, if the parameters of each linear processing unit of the trainable model are updated in real time according to the input audio signal, real-time adaptive sound field control (multi-spot reproduction) becomes possible.
[0164] Other Embodiments The above-described embodiments and / or modifications may be combined as appropriate to realize a sound field control system and / or a sound field control device.
[0165] In addition, in the sound field control systems 1000 and 1000A described in the above embodiments (including modifications), the total area has four areas Q1 to Q4, but the present invention is not limited to this, and the total area may have M areas (M: a natural number of 2 or more). In this case, the audio signal input to the sound field control devices 100, 100A, and 100B is (Q1) (t) ~ s (QM) (t).
[0166] Furthermore, in the embodiments (including modified examples), the case where the sound pressure distribution of the region Qi included in the entire region is modeled as a rectangular wave (corresponding to the window function being a rectangular window) has been described, but this is not limited to this, and the sound pressure distribution of the region Qi included in the entire region may be another sound pressure distribution (for example, the sound pressure distribution of the region Qi included in the entire region may be a sound pressure distribution obtained by modeling using a model corresponding to the window function being a Hann function).
[0167] Furthermore, in the sound field control systems 1000, 1000A and sound field control devices 100, 100A, 100B described in the above embodiments (including modified examples), each block may be individually integrated into a single chip using a semiconductor device such as an LSI, or may be integrated into a single chip to include some or all of the blocks.
[0168] Although the term "LSI" is used here, it may also be called an IC, system LSI, super LSI, or ultra LSI depending on the degree of integration.
[0169] The method of integration is not limited to LSI, but may be realized by a dedicated circuit or a general-purpose processor. It is also possible to use an FPGA (Field Programmable Gate Array) that can be programmed after LSI manufacturing, or a reconfigurable processor that can reconfigure the connections and settings of circuit cells inside the LSI.
[0170] Furthermore, some or all of the processing of each functional block in each of the above embodiments (including modified examples) may be realized by a program. And, some or all of the processing of each functional block in each of the above embodiments is performed by a central processing unit (CPU) in a computer. Furthermore, the program for performing each processing is stored in a storage device such as a hard disk or ROM, and is read out and executed in the ROM or RAM.
[0171] The processes in the above-described embodiments (including modifications) may be implemented by hardware, software (including the case where they are implemented together with an OS (operating system), middleware, or a predetermined library), or may be implemented by a combination of software and hardware.
[0172] For example, when each functional unit of the above embodiment (including modified examples) is realized by software, each functional unit may be realized by software processing using the hardware configuration shown in FIG. 14 (e.g., a hardware configuration in which a CPU, GPU, ROM, RAM, input unit, output unit, communication unit, memory unit (e.g., a memory unit realized by an HDD, SSD, etc.), an external media drive, etc. are connected via a bus).
[0173] Furthermore, when each functional unit of the above embodiment (including modified examples) is realized by software, the software may be realized using a single computer having the hardware configuration shown in Figure 14, or may be realized by distributed processing using multiple computers.
[0174] Furthermore, the order of execution of the processing methods in the above embodiments (including modified examples) is not necessarily limited to that described in the above embodiments, and the order of execution can be changed within the scope of the gist of the invention.
[0175] The scope of the present invention includes a computer program for causing a computer to execute the above-described method, and a computer-readable recording medium on which the program is recorded. Examples of computer-readable recording media include flexible disks, hard disks, CD-ROMs, MOs, DVDs, DVD-ROMs, DVD-RAMs, large-capacity DVDs, next-generation DVDs, and semiconductor memories.
[0176] The computer program is not limited to one recorded on the recording medium, but may be one transmitted via a telecommunications line, a wireless or wired communication line, a network such as the Internet, or the like.
[0177] The specific configuration of the present invention is not limited to the above-described embodiment (including modified examples), and various changes and modifications are possible without departing from the gist of the invention.
[0178] [Note] The present invention can also be expressed as follows: A first invention is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number of 2 or more), and that performs multi-spot reproduction using a speaker array consisting of a plurality of speakers to reproduce a plurality of audio signals in different divided areas included in the entire area, and that includes a short-time Fourier transform processing unit, a time-frequency filter processing unit for the entire area sound field, and an inverse short-time Fourier transform processing unit.
[0179] The short-time Fourier transform processor performs short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0180] The time-frequency filter processing unit for the full-area sound field acquires a harmonic spectrum for the full-area sound field to reproduce the full-area sound field based on the harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filter processing using the acquired harmonic spectrum for the full-area sound field to acquire a time-frequency domain drive signal for driving the speaker array.
[0181] The inverse short-time Fourier transform processor performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
[0182] In this sound field control device, for a sound field in which the entire area is divided into M areas (M: a natural number of 2 or more) (a sound field of the entire area having M divided areas), a cylindrical harmonic spectrum for the entire area sound field is acquired to reproduce the sound field (a sound field of the entire area having M divided areas), and time-frequency filtering is performed using the acquired cylindrical harmonic spectrum for the entire area sound field to acquire time-frequency domain drive signals (e.g., drive signals (time-frequency domain signals) for driving L speakers of the speaker array Spk_arry).The sound field control device then performs short-time Fourier transform processing on the acquired time-frequency domain drive signals to acquire, for example, signals for driving each of the L speaker arrays Spk_arry (time-domain drive signals). In other words, this sound field control device sets a sound field in which the entire area is divided into M areas, and adopts a method of reproducing the sound field. Therefore, the amount of calculation is dramatically reduced compared to, for example, a method in which a divided sound field (local sound field) is set for each of the M divided areas, and the drive signals obtained for each local sound field are superimposed (local sound field superimposition method).
[0183] A second aspect of the present invention is the first aspect of the present invention, further comprising a silent area setting section that sets at least one of the divided areas as a silent area.
[0184] The time-frequency filter processing unit for the entire sound field sets the harmonic spectrum derived from the sound pressure distribution of the divided area for the divided area set in the silent area to 0, or sets the short-time Fourier transform data corresponding to the silent area to 0, and obtains the harmonic spectrum for the entire sound field to reproduce the sound field of the entire area.
[0185] This allows the sound field control device to set at least one of the divided areas as a silent area.
[0186] A third invention is a learning processing method for a parameter-configurable learnable model that receives time-frequency component data of an audio signal as input and outputs frequency component data of a drive signal for reproducing a specified sound field in a specified area using multiple speakers, and includes a transfer function setting processing step, an audio signal time-frequency component data acquisition step, a reproduction sound pressure distribution setting processing step, a reproduction sound pressure distribution acquisition processing step, and a loss evaluation processing step.
[0187] The transfer function setting process step sets a transfer function between the speaker position and the control point.
[0188] The audio signal time frequency component data acquisition step performs frequency conversion on the time domain audio signal to acquire frequency component data of the audio signal.
[0189] The reproduction sound pressure distribution setting processing step acquires correct data for the sound pressure distribution to be reproduced in the specified area based on the setting status of the specified area and the frequency component data of the audio signal acquired in the audio signal time frequency component data acquisition step.
[0190] The excitation signal acquisition processing step acquires frequency component data of the excitation signal by processing the frequency component data of the audio signal using a trainable model whose parameters can be set.
[0191] The reproduction sound pressure distribution acquisition processing step acquires reproduction sound pressure distribution data based on the transfer function acquired in the transfer function setting processing step and the frequency component data of the audio signal acquired in the drive signal acquisition processing step.
[0192] The loss evaluation processing step acquires a loss based on the ground truth data of the sound pressure distribution acquired by the reproduction sound pressure distribution setting processing step and the reproduction sound pressure distribution, and executes a process of updating the parameters of the learnable model based on the acquired loss.
[0193] As a result, in this learning processing method, for example, for a sound field in which the entire area is divided into M areas (M: a natural number equal to or greater than 2) (a sound field of the entire area having M divided areas), a sound pressure distribution for reproducing the sound field (a sound field of the entire area having M divided areas) can be acquired (set) as a correct sound pressure distribution, and the learnable model can be trained so that frequency component data of a drive signal for outputting the correct sound pressure distribution is output. Then, in this learning processing method, by processing using the trained model of the learnable model, for example, frequency component data of a drive signal can be acquired from frequency component data of an audio signal to be reproduced in the entire area having M divided areas, and the acquired frequency component data of the drive signal can be subjected to inverse short-time Fourier transform to acquire a time-domain drive signal, and multiple speakers (e.g., L speakers of a speaker array Spk_arry) can be driven by the time-domain drive signal, thereby performing desired sound field control (desired multi-spot reproduction).
[0194] The learning method may be implemented using, for example, a processor and a memory accessible from the processor, and each step of the learning method may be performed by the processor.
[0195] Furthermore, "frequency transformation" refers to a process (transformation process) for obtaining frequency component data of a signal from a time domain signal. Examples of frequency transformation include Fourier transform (including discrete Fourier transform, fast Fourier transform, etc.), wavelet transform (including discrete wavelet transform, etc.), and discrete cosine transform.
[0196] The fourth invention is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number greater than or equal to 2), and is a sound field control device for performing multi-spot playback in which a plurality of audio signals are each played back in different divided areas included in the entire area using a speaker array consisting of a plurality of speakers, and is equipped with a short-time Fourier transform processing unit, a time-frequency filter processing unit for the entire area sound field, and a short-time Fourier inverse transform processing unit.
[0197] The short-time Fourier transform processor performs short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0198] The time-frequency filter processing unit for full-range sound fields performs processing using a trained model, which is a trainable model in which the acquired optimal parameters are set, by executing a learning process using the learning processing method of the third invention. The time-frequency filter processing unit for full-range sound fields then inputs the short-time Fourier transform data to the trained model and acquires data output from the trained model as a time-frequency domain driving signal for driving a speaker array.
[0199] The inverse short-time Fourier transform processor performs an inverse Fourier transform on the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
[0200] As a result, in this sound field control device, for example, frequency component data of an audio signal to be reproduced in an entire area having M divided areas can be input to a trained model of a time-frequency filter processing unit for an entire area sound field, and frequency component data of a drive signal can be acquired from the trained model. Then, in this sound field control device, the frequency component data of the acquired drive signal is subjected to inverse short-time Fourier transform processing to acquire a time-domain drive signal, and by driving multiple speakers (for example, L speakers of a speaker array Spk_arry) using the time-domain drive signal, desired sound field control (desired multi-spot reproduction) can be performed.
[0201] The fifth invention is a sound field control device that controls a sound field of an entire area having M divided areas (M: a natural number greater than or equal to 2), and is a sound field control method for performing multi-spot playback using a speaker array consisting of multiple speakers to play multiple audio signals in different divided areas included in the entire area, and includes a short-time Fourier transform processing step, a time-frequency filter processing step for the entire area sound field, and a short-time inverse Fourier transform processing step.
[0202] The short-time Fourier transform processing step performs a short-time Fourier transform on the plurality of audio signals to obtain short-time Fourier transform data.
[0203] The time-frequency filtering process for the entire sound field acquires a harmonic spectrum for the entire sound field to reproduce the sound field of the entire area based on the harmonic spectrum derived from the sound pressure distribution of the divided area and the short-time Fourier transform data, and performs time-frequency filtering process using the acquired harmonic spectrum for the entire sound field to acquire a time-frequency domain driving signal for driving the speaker array.
[0204] The inverse short-time Fourier transform processing step performs an inverse Fourier transform on the time-frequency domain driving signal to obtain a driving signal for driving each speaker of the speaker array.
[0205] This makes it possible to realize a sound field control method that has the same effects as the first aspect of the invention.
[0206] A sixth aspect of the present invention is a program for causing a computer to execute the method (learning processing method, sound field control method) of the third or fifth aspect of the present invention.
[0207] This makes it possible to realize a program for causing a computer to execute a method (learning processing method, sound field control method) that has the same effects as the third or fifth invention.
[0208] 1000, 1000A, 2000A Sound field control system 100, 100A, 100B Sound field control device 3, 3A Short-time Fourier transform processing unit 4, 4A Time frequency filter processing unit for full-range sound field 5 Inverse short-time Fourier transform processing unit 6 Silence region setting unit Spk_arry Speaker array Spk1 to Spk16 Speakers
Claims
1. A sound field control device that controls a sound field of an entire area having M (M: natural number 2 or greater) divided areas, and that performs multi-spot playback by using a speaker array consisting of a plurality of speakers to play a plurality of audio signals in different divided areas included in the entire area, comprising: a short-time Fourier transform processing unit that performs a short-time Fourier transform of the plurality of audio signals to obtain short-time Fourier transform data; a whole-area sound field time-frequency filter processing unit that obtains a whole-area sound field harmonic spectrum for reproducing the whole-area sound field based on a harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performs time-frequency filtering using the obtained whole-area sound field harmonic spectrum to obtain a time-frequency domain drive signal for driving the speaker array; and an inverse short-time Fourier transform processing unit that performs an inverse Fourier transform of the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
2. The sound field control device according to claim 1, further comprising a silent area setting unit that sets at least one of the divided areas to a silent area, wherein the time-frequency filter processing unit for the whole area sound field sets the harmonic spectrum derived from the sound pressure distribution of the divided area for the divided area set to the silent area to 0, or sets the short-time Fourier transform data corresponding to the silent area to 0, and acquires the harmonic spectrum for the whole area sound field for reproducing the sound field of the whole area.
3. A learning processing method for a parameter-configurable learnable model that receives input time frequency component data of an audio signal and outputs frequency component data of a drive signal for reproducing a predetermined sound field in a predetermined region using multiple speakers, comprising: a transfer function setting processing step that sets a transfer function between speaker positions and control points; an audio signal time frequency component data acquisition step that performs frequency conversion on the time domain audio signal to acquire frequency component data of the audio signal; a reproduction sound pressure distribution setting processing step that acquires correct data of a sound pressure distribution to be reproduced in the predetermined region based on the setting status of the predetermined region and the frequency component data of the audio signal acquired in the audio signal time frequency component data acquisition step; a drive signal acquisition processing step that acquires frequency component data of a drive signal by processing the frequency component data of the audio signal using a parameter-configurable learnable model; and a reproduction sound pressure distribution acquisition processing step that acquires reproduction sound pressure distribution data based on the transfer function acquired in the transfer function setting processing step and the frequency component data of the audio signal acquired in the drive signal acquisition processing step. a loss evaluation processing step of acquiring a loss based on the ground truth data of the sound pressure distribution acquired by the reproduction sound pressure distribution setting processing step and the reproduction sound pressure distribution, and executing a process of updating the parameters of the trainable model based on the acquired loss.
4. A sound field control device that controls a sound field of an entire area having M divided areas (M: natural number of 2 or more), and that performs multi-spot playback by using a speaker array consisting of a plurality of speakers to play a plurality of audio signals in different divided areas included in the entire area, comprising: a short-time Fourier transform processing unit that performs a short-time Fourier transform of the plurality of audio signals to obtain short-time Fourier transform data; a time-frequency filter processing unit for the entire area sound field that performs processing using a trained model, which is a trainable model in which optimal parameters obtained by executing a learning process using the learning processing method described in claim 3, and that inputs the short-time Fourier transform data to the trained model and obtains data output from the trained model as time-frequency domain drive signals for driving the speaker array; and an inverse short-time Fourier transform processing unit that performs an inverse Fourier transform of the time-frequency domain drive signals to obtain drive signals for driving each speaker of the speaker array.
5. A sound field control device for controlling a sound field of an entire area having M (M: natural number of 2 or more) divided areas, and a sound field control method for performing multi-spot reproduction in which a speaker array consisting of a plurality of speakers is used to reproduce a plurality of audio signals in different divided areas included in the entire area, the sound field control method comprising: a short-time Fourier transform processing step of performing a short-time Fourier transform of the plurality of audio signals to obtain short-time Fourier transform data; a whole-area sound field time-frequency filtering step of obtaining a whole-area sound field harmonic spectrum for reproducing the whole-area sound field based on a harmonic spectrum derived from the sound pressure distribution of the divided areas and the short-time Fourier transform data, and performing time-frequency filtering using the obtained whole-area sound field harmonic spectrum to obtain a time-frequency domain drive signal for driving the speaker array; and an inverse short-time Fourier transform processing step of performing an inverse Fourier transform of the time-frequency domain drive signal to obtain a drive signal for driving each speaker of the speaker array.
6. A program for causing a computer to execute the method according to claim 3 or 5.
Citation Information
Patent Citations
Acoustic signal processing apparatus, acoustic signal processing method, and program
JP2013102389A
Acoustic processing device
JP2013201564A
Signal processing device, signal processing system, signal processing method, and program
JP2018074437A