Sound signal processing method and sound signal processing device

The sound signal processing method enhances sound image localization in virtual spaces by classifying and positioning virtual sound sources, simulating reflections and reverberations, resulting in clear and spatially rich sound reproduction with efficient operation.

JP2026053748APending Publication Date: 2026-03-25YAMAHA CORP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2026-01-08
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Conventional sound systems struggle to clearly reproduce sound image localization in virtual spaces where speakers are installed.

Method used

The sound signal processing method classifies virtual sound sources into first and second sources, adjusts their positions, and generates control signals to simulate initial reflections and reverberations based on the geometric shape of the virtual space, using a device with units for region setting, grouping, reflection and reverberation control, and output adjustment.

Benefits of technology

This method achieves clear sound image localization and a rich sense of spatial expansion, smoothing sound transitions, and allows for easier operation to achieve desired sound fields by adjusting gain and delay values simultaneously.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053748000001_ABST
    Figure 2026053748000001_ABST
Patent Text Reader

Abstract

It clearly reproduces sound localization in a virtual space. [Solution] The sound signal processing method classifies virtual sound sources representing the reflected sound of the target acoustic space into a first virtual sound source located between the speaker position and the receiving point position, representing the first sound source of the reflected sound of the target acoustic space, and a second virtual sound source located outside the speaker, representing the second sound source of the reflected sound of the target acoustic space. Only in the case of the first virtual sound source, the position of the first virtual sound source is moved to a position where it can be reproduced using the position of a speaker in the vicinity of the first virtual sound source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of this invention relates to a sound signal processing method and a sound signal processing apparatus that perform predetermined processing on sound input from a sound source.

Background Art

[0002] In an acoustic system for a space such as a hall, sound image localization of a sound source is performed by speakers arranged in the space.

[0003] For example, the audio processing apparatus described in Patent Document 1 outputs the sound of an audio object (sound source) using two or more speakers near the audio object (sound source). At this time, the audio processing apparatus described in Patent Document 1 calculates the gain of the audio signal output to each speaker using the position information and sound image information of the audio object (sound source).

Prior Art Documents

Patent Documents

[0004]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0005] However, with the above-described conventional configuration, it is not possible to clearly reproduce sound image localization in a virtual space in the space where the speakers are installed.

[0006] Therefore, one embodiment of this invention aims to clearly reproduce sound image localization in a virtual space.

Means for Solving the Problems

[0007] The sound signal processing method classifies virtual sound sources representing the reflected sound of the target acoustic space into a first virtual sound source located between the speaker position and the receiving point position, representing the first sound source of the reflected sound of the target acoustic space, and a second virtual sound source located outside the speaker, representing the second sound source of the reflected sound of the target acoustic space. Only in the case of the first virtual sound source, the position of the first virtual sound source is moved to a position where it can be reproduced using the position of a speaker in the vicinity of the first virtual sound source. [Effects of the Invention]

[0008] The sound signal processing method can clearly reproduce the localization of sound images in a virtual space. [Brief explanation of the drawing]

[0009] [Figure 1] Figure 1 is a functional block diagram showing the configuration of an acoustic system including an audio signal processing device according to an embodiment of the present invention. [Figure 2] Figure 2 is a flowchart of an audio signal processing method according to an embodiment of the present invention. [Figure 3] Figure 3 shows discrete waveforms of sound, including typical direct sound, early reflections, and reverberation (after-reverberation). [Figure 4] Figures 4(A) and 4(B) illustrate the concept of setting up a virtual sound source. [Figure 5] Figure 5 is a functional block diagram showing an example of the configuration of the grouping unit 40. [Figure 6] Figure 6 is a flowchart showing the method for grouping sound sources. [Figure 7] Figure 7 illustrates the concept of grouping multiple sound sources into multiple regions. [Figure 8] Figure 8(A) is a flowchart showing a method for grouping sound sources using representative points, and Figure 8(B) is a flowchart showing a method for grouping sound sources using region boundaries. [Figure 9] Figure 9 is a flowchart showing an example of a grouping method by moving the sound source. [Figure 10]FIG. 10 is a functional block diagram showing an example of the configuration of the initial reflected sound control signal generation unit 50. [Figure 11] FIG. 11 is a diagram showing an example of the GUI. [Figure 12] FIG. 12 is a flowchart showing an example of the virtual sound source setting process. [Figure 13] FIGS. 13(A) and 13(B) are diagrams showing setting examples of respective virtual sound sources when the geometric shapes are different. [Figure 14] FIGS. 14(A), 14(B), and 14(C) are diagrams showing setting examples of virtual sound sources. [Figure 15] FIGS. 15(A), 15(B), and 15(C) are diagrams showing setting examples of virtual sound sources. [Figure 16] FIG. 16 is a flowchart showing the process of assigning virtual sound sources to speakers. [Figure 17] FIGS. 17(A) and 17(B) are diagrams showing the concept of assigning virtual sound sources to speakers. [Figure 18] FIG. 18 is a flowchart showing the coefficient setting process of LDtap. [Figure 19] FIGS. 19(A) and 19(B) are diagrams for explaining the concept of coefficient setting. [Figure 20] FIG. 20(A) shows an example of the LDtap coefficient when the virtual space shape is large, and FIG. 20(B) shows an example of the LDtap coefficient when the virtual space shape is small. [Figure 21] FIG. 21 is a diagram showing the waveform of the initial reflected sound control signal generated by the initial reflected sound control signal generation unit 50. [Figure 22] FIG. 22 is a functional block diagram showing an example of the configuration of the reverberation sound control signal generation unit 70. [Figure 23] FIG. 23 is a flowchart showing an example of the generation process of the reverberation sound control signal. [Figure 24] FIG. 24 is a graph showing waveform examples of the direct sound, the initial reflected sound control signal, and the reverberation sound control signal. [Figure 25] FIG. 25 is a diagram showing an example of the area setting for the reverberation sound. [Figure 26] Figure 26 is a functional block diagram showing an example of the configuration of the output adjustment unit 90. [Figure 27] Figure 27 is a flowchart showing an example of the output adjustment process. [Figure 28] Figure 28 shows an example of a GUI for output adjustment. [Figure 29] Figures 29(A) and 29(B) show examples of settings for creating sound localization and spread at the rear of the playback space. [Figure 30] Figures 30(A) and 30(B) show examples of settings for creating sound localization and spread in the lateral direction of the playback space. [Figure 31] Figure 31 is a diagram illustrating the image of sound spread when the sound has a vertical spread. [Figure 32] Figure 32 is a functional block diagram showing the configuration of an audio signal processing device with binaural playback functionality. [Modes for carrying out the invention]

[0010] The sound signal processing method and sound signal processing device according to embodiments of the present invention will be described with reference to the figures. In the following embodiments, first, an overview of the sound signal processing method and sound signal processing device will be described, and then the specific details of each process and each configuration will be described.

[0011] In this embodiment, the playback space is the space in which the user (listener) listens to sound from a sound source (direct sound, early reflections, reverberation) using a speaker or the like. The virtual space is a space with a different sound field (acoustics) from the playback space, and the early reflections and reverberation from this sound field are reproduced (simulated) in the playback space.

[0012] [Outline configuration of the sound signal processing device] Figure 1 is a functional block diagram showing the configuration of an acoustic system including an audio signal processing device according to an embodiment of the present invention.

[0013] As shown in Figure 1, the sound signal processing device 10 includes a region setting unit 30, a grouping unit 40, an initial reflection sound control signal generation unit 50, a mixer 60, a reverberation sound control signal generation unit 70, an adder 80, and an output adjustment unit 90. The sound signal processing device 10 is implemented, for example, by electronic circuits that implement the region setting unit 30, the grouping unit 40, the initial reflection sound control signal generation unit 50, the mixer 60, the reverberation sound control signal generation unit 70, the adder 80, and the output adjustment unit 90, or by a computing device such as a computer. The part consisting of the adder 80 and the output adjustment unit 90 corresponds to the "output signal generation unit" of the present invention.

[0014] The sound signal processing device 10 is connected to multiple speakers SP1-SP64. Note that Figure 1 shows a configuration using 64 speakers, but the number of speakers is not limited to this.

[0015] The sound signal processing device 10 receives sound signals S1-S96 from multiple sound sources OBJ1-OBJ96. Note that Figure 1 shows a configuration using 96 sound sources, but the number of sound sources is not limited to this.

[0016] The region setting unit 30 divides the playback space into multiple regions and sets information about the divided regions (region information). The region information includes the position coordinates that determine the boundaries of the regions and the position coordinates of the representative points set in the regions.

[0017] The area setting unit 30 outputs area information for the set multiple areas Area1-Area8 to the grouping unit 40. Although Figure 1 shows a configuration in which eight areas are set, the number of areas is not limited to this.

[0018] The grouping unit 40 groups the sound sources OBJ1-OBJ96 into multiple regions Area1-Area8. Based on the grouping results, the grouping unit 40 generates region-specific sound signals SA1-SA8 for each of the regions Area1-Area8 using the sound signals S1-S96 from the sound sources OBJ1-OBJ96. For example, the grouping unit 40 generates region-specific sound signal SA1 by mixing the sound signals of the multiple sound sources grouped in region Area1.

[0019] The grouping unit 40 outputs multiple region-specific sound signals SA1-SA8 to the initial reflection sound control signal generation unit 50. The grouping unit 40 also outputs sound signals S1-S96 from sound sources OBJ1-OBJ96 to the mixer 60.

[0020] The initial reflection sound control signal generation unit 50 generates initial reflection sound control signals ER1-ER64 for each of the multiple speakers SP1-SP64 from multiple region-specific sound signals SA1-SA8. The initial reflection sound control signals ER1-ER64 are signals output to each of the speakers SP1-SP64 in order to simulate the initial reflection sound in the virtual space in the playback space. The initial reflection sound control signal generation unit 50 outputs the generated initial reflection sound control signals ER1-ER64 to the adder 80.

[0021] In general terms (detailed configuration and processing will be described later), the initial reflection sound control signal generation unit 50 sets up a virtual sound source in the playback space using the positions of speakers SP1-SP64 placed in the playback space and the geometric shape of the virtual space. The specific setting of the virtual sound source will be described later. Using the virtual sound source, the initial reflection sound control signal generation unit 50 generates initial reflection sound control signals ER1-ER64 that simulate the initial reflection sound in the virtual space. At this time, the initial reflection sound control signal generation unit 50 performs desired timbre adjustments on the initial reflection sound control signals ER1-ER64.

[0022] Mixer 60 is a summing mixer. Mixer 60 mixes the sound signals S1-S96 from sound sources OBJ1-OBJ96 to generate a reverberation signal Sr. Mixer 60 outputs the reverberation signal Sr to the reverberation control signal generation unit 70.

[0023] The reverberation control signal generation unit 70 generates reverberation control signals REV1-REV64 for each of the multiple speakers SP1-SP64 from the reverberation generation signal Sr. The reverberation control signals REV1-REV64 are signals output to each of the speakers SP1-SP64 in order to simulate the reverberation (rear reverberation) of the virtual space in the playback space. The reverberation control signal generation unit 70 outputs the generated reverberation control signals REV1-REV64 to the adder 80.

[0024] In general terms (detailed configuration and processing will be described later), the reverberation control signal generation unit 70 divides the playback space into multiple reverberation setting areas and generates a reverberation control signal for each of the multiple reverberation setting areas. The reverberation control signal generation unit 70 assigns multiple speakers SP1-SP64 to the multiple reverberation setting areas. Based on this assignment, the reverberation control signal generation unit 70 sets reverberation control signals for each of the multiple speakers SP1-SP64 for each reverberation setting area.

[0025] In this process, the reverberation control signal generation unit 70 sets the connection timing between the initial reflection and the reverberation based on the geometric shape of the playback space. The reverberation control signal generation unit 70 gradually increases the level (amplitude) of the reverberation control signal during the period before the connection timing, and gradually decreases the level (amplitude) of the reverberation control signal during the period after the connection timing.

[0026] The adder 80 adds the initial reflection control signal and the reverberation control signal generated for each of the multiple speakers SP1-SP64 to generate multiple speaker signals Sat1-Sat64. For example, the adder 80 adds the initial reflection control signal and the reverberation control signal for speaker SP1 to generate speaker signal Sat1. The adder 80 outputs the multiple speaker signals Sat1-Sat64 to the output adjustment unit 90.

[0027] The output adjustment unit 90 generates output signals So1-So64 by performing gain control and delay control on multiple speaker signals Sat1-Sat64. The output adjustment unit 90 outputs output signals So1-So64 to multiple speakers SP1-SP64. For example, the output adjustment unit 90 generates output signal So1 by performing gain control and delay control for speaker SP1 on speaker signal Sat1. The output adjustment unit 90 outputs output signal So1 to speaker SP1.

[0028] In general terms (detailed configuration and processing will be described later), the output adjustment unit 90 accepts input of acoustic parameters in the playback space. Acoustic parameters are parameters that set, for example, the width of the sound space, the width of the space behind the sound receiving point, and the ceiling of the sound space. Based on the position coordinates of the multiple speakers SP1-SP64 and the acoustic parameters, the output adjustment unit 90 sets the gain values ​​and delay amounts (delay amounts) of the multiple speaker signals Sat1-Sat64 all at once. Setting them all at once means that instead of setting them individually for each speaker, the gain value and delay amount for each speaker are set simply by inputting the position coordinates of each speaker into a specific calculation formula common to all speakers. The output adjustment unit 90 performs gain control and delay control on the multiple speaker signals Sat1-Sat64 using the set gain values ​​and delay values.

[0029] [Outline of the audio signal processing method] Figure 2 is a flowchart of an audio signal processing method according to an embodiment of the present invention. Figure 2 shows the audio signal processing method implemented by the audio signal processing device 10 of Figure 1. The details of each process shown in Figure 2 are explained in the description of Figure 1 above, so they are described briefly here.

[0030] (Grouping of sound sources OBJ1-OBJ96) The grouping unit 40 groups the multiple sound sources OBJ1-OBJ96 into multiple regions Area1-Area8 (S11).

[0031] (Generation of early reflection control signals) The initial reflection sound control signal generation unit 50 sets a timbre for the initial reflection sound for each group (SS12). The initial reflection sound control signal generation unit 50 sets a virtual sound source for each group (S13). The initial reflection sound control signal generation unit 50 uses the timbre and virtual sound source to generate initial reflection sound control signals for each of the multiple speakers SP1-SP64 (S14).

[0032] (Generation of reverberation control signals) Mixer 60 sums the sound signals S1-S96 from multiple sound sources OBJ1-OBJ96 (S21). Reverberation control signal generation unit 70 sets the connection timing between the initial reflection and the reverberation based on the geometric shape of the playback space (S22). Reverberation control signal generation unit 70 generates a reverberation control signal using the set connection timing (S23). Reverberation control signal generation unit 70 assigns the generated reverberation control signal to the multiple speakers SP1-SP64 based on the position coordinates of the multiple speakers SP1-SP64 in the playback space (S24).

[0033] (Output processing to multiple speakers) The adder 80 adds the initial reflection control signal and the reverberation control signal for each of the multiple speakers SP1-SP64 to generate speaker signals Sat1-Sat64 (S31).

[0034] The output adjustment unit 90 generates output signals So1-So64 from speaker signals Sat1-Sat64 using acoustic parameters that realize the localization of reverberation and spatial breadth in the playback space (S32). The output adjustment unit 90 outputs the output signals So1-So64 to a plurality of speakers SP1-SP64 (S33).

[0035] By using the above configuration and processing, the sound signal processing device 10 (sound signal processing method) can obtain the following various effects.

[0036] (1) The sound signal processing device 10 (sound signal processing method) generates early reflections by grouping sound sources for each region into which the playback space is divided, thereby achieving clear sound image localization and a rich sense of spatial expansion. In this case, the reverberation remains constant throughout the entire playback space, while only the early reflections change depending on the position of the sound source. Therefore, for example, if the position of the sound source moves, the movement of the sound from this sound source becomes smoother.

[0037] (2) The sound signal processing device 10 (sound signal processing method) generates an early reflection control signal using a virtual sound source, thereby more faithfully simulating the early reflection sound based on the geometric shape of the virtual space in the playback space.

[0038] (3) The sound signal processing device 10 (sound signal processing method) can eliminate the unnaturalness of the timbre of the initial reflection sound, which is simulated by only a virtual sound source, by adjusting the timbre of the initial reflection sound control signal.

[0039] (4) The sound signal processing device 10 (sound signal processing method) can make the transition from the initial reflection sound to the reverberation sound smoother and more natural by setting the connection timing between the initial reflection sound control signal and the reverberation sound control signal based on the geometric shape of the playback space.

[0040] (5) The sound signal processing device 10 (sound signal processing method) adjusts the gain values ​​and delay amounts of speaker signals Sat1-Sat64, which include the initial reflection control signal and the reverberation control signal, all at once, thereby enabling the user to realize the desired sound field in the playback space with easier operation input.

[0041] [Detailed explanation of each signal processing unit and each process] The following provides a detailed explanation of each signal processing unit and each process described above. First, the initial reflections, reverberations, and virtual sound sources necessary for understanding the invention will be explained with reference to the diagram.

[0042] [Early reflections and reverberation] Figure 3 shows the discrete waveforms of sound, including typical direct sound, early reflections, and reverberation (reverberation). For example, a hall used for performances or content playback has a closed space surrounded by walls. When sound is generated in this closed space, the receiving point receives the direct sound, early reflections, and reverberation (reverberation).

[0043] Direct sound is sound that travels directly from the point of origin to the point of reception.

[0044] Early reflections are sounds that arrive at the receiving point relatively quickly after being reflected off walls, floors, and ceilings at the point of origin. Therefore, early reflections arrive at the receiving point immediately after the direct sound. Furthermore, the volume (level) of early reflections is lower than that of the direct sound. If the number of reflections is one, it is a first-order reflection; if it is n reflections, it is an nth-order reflection. The direction of arrival and volume of early reflections at the receiving point are greatly influenced by the sound's origin.

[0045] Reverberation arrives at the receiving point following the initial reflection. Reverberation is the sound that arrives at the receiving point after multiple reflections of the sound generated at the point of origin. In other words, reverberation is the sound that arrives at the receiving point after being reflected and attenuated many more times. Therefore, the volume (level) of reverberation is lower than the volume (level) of the initial reflection. Furthermore, the direction of arrival and volume of reverberation are less affected by the sound generation location than the initial reflection.

[0046] [imaginary sound source] Figures 4(A) and 4(B) illustrate the concept of setting up a virtual sound source. Note that while Figures 4(A) and 4(B) show the concept of setting up a virtual sound source in two dimensions for ease of explanation, the same concept can be applied to setting up a virtual sound source in three dimensions. That is, in an actual playback space, if sound sources are not aligned on a single plane but are spatially arranged, and the virtual space is set up in three dimensions, then the virtual sound source will be set up in three dimensions.

[0047] The playback space contains a sound source SS and a sound receiving point RP. Note that the sound source SS shown in Figures 4(A) and 4(B) has a different meaning from the sound source OBJ described above, and refers to something that produces general sound. In addition, a virtual wall IWL is set up in the playback space to realize the sound field of the virtual space. The virtual wall IWL is obtained from the geometric shape of the virtual space.

[0048] The sound source SS and the receiving point RP exist within the space enclosed by the virtual wall IWL. The virtual wall IWL comprises virtual wall IWL1, virtual wall IWL2, virtual wall IWL3, and virtual wall IWL4. Virtual wall IWL1 and virtual wall IWL4 are positioned in the first direction of the playback space (vertical direction in Figures 4(A) and 4(B)) so as to sandwich the sound source SS and the receiving point RP between them. Virtual wall IWL1 is positioned closer to the sound source SS than the receiving point RP, and virtual wall IWL4 is positioned closer to the receiving point RP than the sound source SS. Virtual wall IWL2 and virtual wall IWL3 are positioned in the second direction of the playback space (horizontal direction in Figures 4(A) and 4(B)) so as to sandwich the sound source SS and the receiving point RP between them. Virtual wall IWL2 is positioned closer to the sound source SS than the receiving point RP, and virtual wall IWL3 is positioned closer to the receiving point RP than the sound source SS.

[0049] If virtual walls IWL1, IWL2, IWL3, and IWL4 are real walls that reflect sound, then, as shown in Figure 4(B), the sound emitted from the sound source SS will be reflected by virtual walls IWL1, IWL2, and IWL3 and reach the receiving point RP. Although Figure 4(B) does not show the reflection from virtual wall IWL4, reflection occurs from virtual wall IWL4 in the same way as from virtual walls IWL1, IWL2, and IWL3.

[0050] However, virtual walls IWL1, IWL2, IWL3, and IWL4 do not actually exist in the playback space. Therefore, as shown in Figure 4(A), the sound signal processing device 10 sets up virtual sound sources IS1, IS2, and IS3 by treating the sound reflections on the wall surfaces as specular reflections.

[0051] Specifically, the sound signal processing device 10 sets a virtual sound source IS1 at a position symmetrical to the sound source SS, using the virtual wall IWL1 as the reference line. The sound signal processing device 10 sets a virtual sound source IS2 at a position symmetrical to the sound source SS, using the virtual wall IWL2 as the reference line. The virtual sound source IS3 at a position symmetrical to the sound source SS, using the virtual wall IWL3 as the reference line. By adjusting the acoustic power of each virtual sound source IS, it is possible to simulate energy loss in reflection at the virtual wall IWL.

[0052] By making these settings, the sound generated by the virtual sound source IS1 will be the same as the sound generated by the sound source SS and reflected by the virtual wall IW1. The sound generated by the virtual sound source IS2 will be the same as the sound generated by the sound source SS and reflected by the virtual wall IW2. The sound generated by the virtual sound source IS3 will be the same as the sound generated by the sound source SS and reflected by the virtual wall IW3. Although the virtual sound source for virtual wall IWL4 is not shown in Figures 4(A) and 4(B), a virtual sound source can be set for virtual wall IWL4 in the same way as for virtual walls IWL1, IWL2, and IWL3.

[0053] By setting up a virtual sound source in this way, the sound signal processing device 10 can simulate the initial reflections of sound in a virtual space in a playback space where there are no real walls in the virtual space.

[0054] [Configuration and processing of the grouping unit 40] Figure 5 is a functional block diagram showing an example of the configuration of the grouping unit 40. Figure 6 is a flowchart showing the method of grouping sound sources.

[0055] As shown in Figure 5, the grouping unit 40 includes a sound source position detection unit 41, a region determination unit 42, and a matrix mixer 400.

[0056] The sound source position detection unit 41 detects the position coordinates of multiple sound sources OBJ1-OBJ96 in the playback space (Figure 6: S111). For example, the sound source position detection unit 41 detects the position coordinates of sound sources OBJ1-OBJ96 based on user input. Alternatively, the sound source position detection unit 41 is equipped with position detection sensors for detecting sound sources OBJ1-OBJ96, and detects the position coordinates of sound sources OBJ1-OBJ96 based on the position detected by the position detection sensors.

[0057] The sound source position detection unit 41 outputs the position coordinates of sound sources OBJ1-OBJ96 to the region determination unit 42.

[0058] The region determination unit 42 uses region information of multiple regions Area1-Area8 from the region setting unit 30 and the position coordinates of sound sources OBJ1-OBJ96 from the sound source position detection unit 41 to group sound sources OBJ1-OBJ96 into multiple regions Area1-Area8 (Figure 6: S112). More specifically, the region determination unit 42 performs the grouping as follows.

[0059] Figure 7 illustrates the concept of grouping multiple sound sources into multiple regions. In Figure 7, the top of the figure represents the front of the hall, which is the playback space, and the bottom of the figure represents the back of the hall.

[0060] The region setting unit 30 sets a reference point Pso for region division in relation to the playback space. For example, as shown in Figure 7, the region setting unit 30 sets the center position of the hall that realizes the playback space as the reference point Pso. The region setting unit 30 can also use a point (position) set by the user as the reference point. For example, the region setting unit 30 can use a sound receiving point set by the user as the reference point.

[0061] The area setting unit 30 sets up eight areas, Area1-Area8, so as to divide the entire circumference on the plane into eight parts, centered on the reference point Pso for area division. For example, in Figure 7, the area setting unit 30 sets up multiple areas, Area1, Area2, and Area3, in front of the reference point Pso in the hall (playback space). The area setting unit 30 also sets up area4 to the left of the reference point Pso facing forward of the hall, and area5 to the right of the reference point Pso facing forward of the hall. Furthermore, the area setting unit 30 sets up multiple areas, Area6, Area7, and Area8, behind the reference point Pso in the hall (playback space).

[0062] Note that this area setting is just an example; other settings are acceptable as long as the entire playback space can be covered by multiple defined areas. Also, while this explanation shows the setting of planar areas, spatial areas can be set in a similar manner. For example, the vertical range of area Area1 is also included in area Area1.

[0063] The region setting unit 30 sets representative points RP1-RP8 for each of the multiple regions Area1-Area8. For example, the region setting unit 30 sets the multiple representative points RP1-RP8 at the center positions of the multiple regions Area1-Area8. Alternatively, in the case of radially spreading regions as shown in Figure 7, for example, the region setting unit 30 sets a representative point at a predetermined distance from the reference point Pso on a straight line passing through the center of the radially spreading corner. Note that these methods for setting representative points are just examples, and other methods may be used as long as they allow for setting one representative point per region and ensure reliable grouping of sound sources.

[0064] The region setting unit 30 outputs region information for multiple regions Area1-Area8 to the region determination unit 42 and matrix mixer 400 of the grouping unit 40. The region information for multiple regions Area1-Area8 includes the position coordinates of representative points RP1-RP8 of regions Area1-Area8, coordinate information representing the boundary lines that form the shape of regions Area1-Area8, etc.

[0065] (A method for grouping sound sources into regions using representative points) Figure 8(A) is a flowchart showing a method for grouping sound sources using representative points.

[0066] The region determination unit 42 obtains the position coordinates of representative points RP1-RP8 from the region information of multiple regions Area1-Area8 (S131). The region determination unit 42 calculates the distance between the position coordinates of the sound sources to be grouped and the position coordinates of representative points RP1-RP8 (S132). The region determination unit 42 groups the sound sources into regions that include the representative points with the shortest distance (S133).

[0067] For example, in the case of sound source OBJ1 in the example in Figure 7, the area determination unit 42 detects the position coordinates of sound source OBJ1 and obtains the position coordinates of multiple representative points RP1-RP8. The area determination unit 42 calculates the distance between sound source OBJ1 and the multiple representative points RP1-RP8 from the position coordinates of sound source OBJ1 and the position coordinates of the multiple representative points RP1-RP8. The area determination unit 42 detects that the distance between sound source OBJ1 and representative point RP1 is shorter than the distance between sound source OBJ1 and the other representative points RP2-RP8. In other words, the area determination unit 42 detects that the distance between sound source OBJ1 and representative point RP1 is the shortest distance. The area determination unit 42 groups sound source OBJ1 into the area Area1 associated with representative point RP1.

[0068] (A method of grouping sound sources into regions using region boundaries) Figure 8(B) is a flowchart showing a method for grouping sound sources using region boundaries.

[0069] The region determination unit 42 obtains coordinate information (boundary coordinates) representing the boundary lines of each region Area1-Area8 from the region information of multiple regions Area1-Area8 (S136). The region determination unit 42 determines whether the position coordinates of the sound source to be grouped are inside each region Area1-Area8 (S137). For example, the region determination unit 42 uses the Crossing Number Algorithm to determine whether the sound source is inside or outside a region. If the region determination unit 42 finds that the sound source is inside a region (S137: YES), it groups the sound source into this region (S138).

[0070] For example, in the case of sound source OBJ1 in the example in Figure 7, the area determination unit 42 detects the position coordinates of sound source OBJ1 and obtains coordinate information (boundary coordinates) representing the boundary lines of multiple areas Area1-Area8. The area determination unit 42 determines whether sound source OBJ1 is inside or outside of multiple areas Area1-Area8 based on the position coordinates of sound source OBJ1 and the boundary coordinates of multiple areas Area1-Area8. The area determination unit 42 detects that sound source OBJ1 is inside area Area1. The area determination unit 42 groups sound source OBJ1 into area Area1.

[0071] The region determination unit 42 groups the multiple input sound sources OBJ1-OBJ96 into multiple regions Area1-Area8. For example, in the example shown in Figure 7, the region determination unit 42 groups sound sources OBJ1 and OBJ4 into region Area1, sound source OBJ2 into region Area2, and sound source OBJ3 into region Area5.

[0072] The region determination unit 42 outputs grouping information to the matrix mixer 400. The grouping information is information indicating which sound source has been grouped into which region, as described above.

[0073] The matrix mixer 400 generates region-specific sound signals SA1-SA8 for each of the multiple regions Area1-Area8 using the sound signals S1-S96 from multiple sound sources OBJ1-OBJ96 based on grouping information. For example, if multiple sound sources are grouped in a region, the matrix mixer 400 mixes the sound signals from these multiple sound sources to generate region-specific sound signals for that region. The matrix mixer 400 outputs the region-specific sound signals for each region to the initial reflection sound control signal generation unit 50. Note that if even one sound source is grouped in a region, the matrix mixer 400 outputs the sound signal from this sound source as the region-specific sound signal for that region to the initial reflection sound control signal generation unit 50.

[0074] In the example shown in Figure 7, Area 1 is a grouping of sound sources OBJ1 and OBJ4. The matrix mixer 400 mixes the sound signal S1 from sound source OBJ1 and the sound signal S4 from sound source OBJ4 to generate and output the area-specific sound signal SA1 for Area 1. Similarly, Area 2 is a grouping of sound source OBJ2. The matrix mixer 400 outputs the sound signal S2 from sound source OBJ2 as the area-specific sound signal SA2 for Area 2. Furthermore, Area 5 is a grouping of sound source OBJ3. The matrix mixer 400 outputs the sound signal S3 from sound source OBJ3 as the area-specific sound signal SA5 for Area 5.

[0075] By implementing this configuration and processing, the sound signal processing device 10 can group multiple sound sources for each of the multiple regions that divide the sound space and generate an early reflection control signal. As a result, the sound signal processing device 10 can reproduce early reflections according to the position of the sound source, achieving clear sound image localization and a rich sense of spatial expansion.

[0076] Although the above explanation does not detail the case when the sound source moves, if the sound source moves, the grouping unit 40 performs the processing shown in Figure 9. Figure 9 is a flowchart showing an example of a grouping method due to the movement of the sound source.

[0077] The sound source position detection unit 41 detects the movement of the sound source (S104). The sound source position detection unit 41 detects the movement of the sound source, for example, by user input. Alternatively, the sound source position detection unit 41 detects the movement of the sound source by continuously detecting the sound source position using a position detection sensor. The sound source position detection unit 41 detects the position coordinates of the sound source after the movement and outputs them to the region determination unit 42.

[0078] The region determination unit 42 uses the position coordinates of the sound source after movement to group it into multiple regions Area1-Area8, as described above.

[0079] By performing this processing, the sound signal processing device 10 can generate an initial reflection control signal corresponding to the position of the sound source after it has moved, even if the sound source has moved. As a result, the sound signal processing device 10 can reproduce the changes in the initial reflections that occur in response to the movement of the sound source, and even if the sound source moves, it can achieve clear sound image localization and a rich sense of spatial expansion that corresponds to the movement.

[0080] Furthermore, when such movement of a sound source occurs, the sound signal processing device 10 can apply a crossfade process to the initial reflection sound control signal before the movement and the initial reflection sound control signal after the movement. For example, when a sound source moves, the sound signal processing device 10 gradually lowers the component of the sound signal of the sound source in the region-specific sound signal that includes the sound source before the movement. On the other hand, the sound signal processing device 10 gradually raises the component of the sound signal of the sound source in the region-specific sound signal that includes the sound source after the movement.

[0081] By performing this processing, the sound signal processing device 10 can suppress discontinuous changes in the initial reflections when the sound source moves. As a result, the sound signal processing device 10 can change the initial reflections more smoothly in accordance with the movement of the sound source.

[0082] Furthermore, the matrix mixer 400 outputs the sound signals S1-S96 from multiple sound sources OBJ1-OBJ96 to the mixer 60. As described above, the mixer 60 sums the sound signals S1-S96 to generate a reverberation generation signal Sr and outputs it to the reverberation control signal generation unit 70. The reverberation control signal generation unit 70 uses the reverberation generation signal Sr to generate reverberation control signals REV1-REV64.

[0083] Through this processing, the reverberation is not affected by the position or movement of the sound source. Therefore, the sound signal processing device 10 can maintain a constant reverberation in the playback space even when the sound source moves, while more clearly reproducing the movement of the sound source by changing the initial reflections.

[0084] [Generation of early reflection control signals] Figure 10 is a functional block diagram showing an example of the configuration of the initial reflection sound control signal generation unit 50. Figure 11 is a diagram showing an example of the GUI.

[0085] As shown in Figure 10, the initial reflection sound control signal generation unit 50 includes an FIR filter circuit 51, an LDtap circuit 52, an addition processing unit 53, a timbre setting unit 501, a virtual sound source setting unit 502, and an operation unit 500. The FIR filter circuit 51 includes a plurality of FIR filters 511-518. The LDtap circuit 52 includes a plurality of LDtaps 521-528, an output speaker setting unit 5201, and a coefficient setting unit 5202. Note that the order of the FIR filter circuit 51 and the LDtap circuit 52 may be reversed.

[0086] [Adjusting the timbre of the initial reflections] The control unit 500 receives timbre specification information from the user to be added to the initial reflections and outputs it to the timbre setting unit 501. The timbre specification information is information that specifies, for example, emphasis on low frequencies, emphasis on high frequencies, volume of the initial reflections, and attenuation characteristics of the initial reflections (information that represents the filter characteristics).

[0087] As a specific example, the operation unit 500 accepts operations via a GUI 100 (Graphical User Interface) as shown in Figure 11.

[0088] GUI110 includes a settings display window 111, multiple controls 112, a knob 1131, and an adjustment value display window 1132.

[0089] The settings display window 111 displays the shape of the virtual wall IWL of the virtual space, which is set by multiple operators 112 and knobs 1131. At this time, the settings display window 111 can also display the separately set positions of the sound source SS, speaker SP, receiving point RP, and the coordinate axes of the playback space, along with the virtual wall IWL.

[0090] Multiple operators 112 are associated with pre-configured virtual space samples (various halls, rooms, etc.). Although not shown in the diagram, each operator 112 displays an index (e.g., hall name) that clearly indicates the virtual space sample associated with that operator 112.

[0091] Knob 1131 is for setting the room size. The adjustment value display window 1132 displays the set value for the room size.

[0092] GUI100 accepts various operations for adjusting the tone. For example, GUI100 is equipped with multiple control elements 112, controls for the low frequency range, controls for the high frequency range, controls for volume adjustment, controls for attenuation characteristic adjustment, etc., and accepts operations through these controls.

[0093] When a user operates a desired control using the GUI 100, the operation unit 500 detects this operation and sets the timbre specification information according to these operations.

[0094] For example, when the control unit 500 receives the selection of multiple operators 112, it obtains the timbre specification information pre-set in the virtual space associated with the operator 112. Also, when the control unit 500 receives operations from operators for the low frequency range, high frequency range, volume control, attenuation characteristic adjustment, etc., it obtains the timbre specification information set by these operators.

[0095] Although not shown in the diagram, GUI100 can also display timbre specification information using, for example, the filter coefficients of the FIR filters 511-518 described later, a rough waveform, etc. In this case, GUI100 can change the display according to the adjustment of the timbre specification information it receives. For example, GUI100 can change the waveform display according to the adjustment.

[0096] The timbre setting unit 501 sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 based on the timbre specification information. For example, if the timbre setting unit 501 receives a specification information emphasizing the low frequency range, it sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 to boost the low frequencies. Also, if the timbre setting unit 501 receives a specification information emphasizing the high frequency range, it sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 to boost the high frequencies. The timbre setting unit 501 outputs the set filter coefficients to the FIR filter circuit 51. In addition to filter coefficients, the timbre setting unit 501 can also set and adjust the sampling frequency and filter length as filter characteristics.

[0097] Furthermore, the tone setting unit 501 sets the gain value of each tap of the FIR filters 511-518 of the FIR filter circuit 51 based on the specified tone information. The tone setting unit 501 outputs the set gain value to the FIR filter circuit 51.

[0098] The multiple FIR filters 511-518 are filters corresponding to the region-specific sound signals SA1-SA8, respectively. The region-specific sound signals SA1-SA8 are input to the FIR filters 511-518. For example, as shown in Figure 10, region-specific sound signal SA1 is input to FIR filter 511, region-specific sound signal SA2 is input to FIR filter 512, region-specific sound signal SA3 is input to FIR filter 513, region-specific sound signal SA4 is input to FIR filter 514, region-specific sound signal SA5 is input to FIR filter 515, region-specific sound signal SA6 is input to FIR filter 516, region-specific sound signal SA7 is input to FIR filter 517, and region-specific sound signal SA8 is input to FIR filter 518.

[0099] Multiple FIR filters 511-518 have the same number of taps. For example, multiple FIR filters 511-518 have 16,000 taps. Note that this number of taps is just an example and should be set based on the resource conditions of the sound signal processing device 10, the accuracy of the timbre of the initial reflections to be reproduced, etc.

[0100] Multiple FIR filters 511-518 perform filtering (convolution) on multiple region-specific sound signals SA1-SA8 using the filter coefficients and gain values ​​set in the timbre setting unit 501. This allows the multiple FIR filters 511-518 to generate filtered region-specific sound signals SA1f-SA8f. For example, FIR filter 511 performs filtering (convolution) on region-specific sound signal SA1 using the filter coefficients and gain values ​​set in the timbre setting unit 501, generating filtered region-specific sound signal SA1f. Similarly, multiple FIR filters 512-518 individually generate filtered region-specific sound signals SA2f-SA8f from region-specific sound signals SA2-SA8.

[0101] Multiple FIR filters 511-518 output the filtered region-specific sound signals SA1f-SA8f to multiple LDtap 521-528. For example, FIR filter 511 outputs the filtered region-specific sound signal SA1f to LDtap 521. Similarly, multiple FIR filters 512-518 output the filtered region-specific sound signals SA2f-SA8f to multiple LDtap 522-528.

[0102] Furthermore, the timbre specification information is not limited to information emphasizing frequency range, but also includes information that sets the waveform of the initial reflection sound to the user's desired characteristics. By using such timbre specification information, the sound signal processing device 10 can realize initial reflection sounds with a wider variety of timbres that suit the user's preferences.

[0103] [Void sound source settings and LDtap settings] The virtual sound source setting unit 502 sets the virtual sound source based on the position coordinates of the receiving point in the playback space and the geometric shape of the virtual space.

[0104] Figure 12 is a flowchart showing an example of the virtual sound source setting process. The virtual sound source setting unit 502 acquires the position coordinates of the sound receiving point in the playback space (S131). For example, the virtual sound source setting unit 502 acquires the position coordinates of the sound receiving point in the playback space through user input, position detection by a position detection sensor, etc.

[0105] The virtual sound source setting unit 502 acquires the geometric shape of the virtual space (S132). For example, the virtual sound source setting unit 502 acquires the geometric shape of the virtual space based on user input or the like. The geometric shape of the virtual space includes a set of coordinates representing the shape of walls placed in the virtual space.

[0106] The virtual sound source setting unit 502 is connected to the GUI 100. When the user selects a desired operator 112 from among several operators 112, the GUI 100 reads and retrieves the geometric shape of the virtual space associated with this operator 112. Also, when the user adjusts the room size using the knob 1131, the GUI 100 retrieves the adjusted value of the room size.

[0107] The virtual sound source setting unit 502 acquires the position coordinates of the geometric shape of the virtual space with the room size set, based on the settings acquired by the GUI 100 in this manner. The virtual sound source setting unit 502 also acquires the position coordinates of the sound source SS and the position coordinates of the receiving point RP (room center). Using this acquired information, the virtual sound source setting unit 502 sets the virtual sound source as shown below. The virtual sound source setting unit 502 matches the coordinate system of the playback space with the coordinate system of the virtual space. Using the position coordinates of the receiving point in the playback space and the geometric shape of the virtual space, the virtual sound source setting unit 502 sets the position coordinates of the virtual sound source in the playback space using the concepts described in Figures 4(A) and 4(B) above (S133).

[0108] Figures 13(A) and 13(B) show examples of virtual sound source settings for different geometric shapes. In Figure 13(A), the virtual wall IWL is a rectangle, and in Figure 13(B), the virtual wall IWLh is a hexagon.

[0109] Thus, if the geometric shape of the virtual space is different, even if the position coordinates of the sound source SSa and the receiving point RP do not change, the positional relationship between the sound source SSa and the receiving point RP and the virtual wall IWL will be different from the positional relationship between the sound source SSa and the receiving point RP and the virtual wall IWLh. As a result, the positions of the imaginary sound sources IS1a, IS2a, and IS3a set in Figure 13(A) will be different from the positions of the imaginary sound sources IS1ah, IS2ah, and IS3ah set in Figure 13(B).

[0110] Figures 14(A), 14(B), and 14(C) show examples of virtual sound source settings. Figures 14(A), 14(B), and 14(C) show planar changes. Figure 14(B) shows the case where the position of the sound source SSa relative to the reference point (receiving point RP) is the same as in Figure 14(A), but the size of the virtual space is different. Figure 14(C) shows the case where the size of the virtual space is the same as in Figure 14(A), but the positional relationship between the reference point of the virtual space and the reference point (receiving point) of the playback space changes (the room center of the playback space changes).

[0111] As can be seen from the comparison between Figure 14(A) and Figure 14(B), the size of the virtual space in the playback space (denoted as virtual wall IWL in Figure 14(A) and virtual wall IWLc in Figure 14(B)) is different, resulting in different distances and positional relationships between the sound source that forms the basis of the virtual sound source and the virtual wall. Consequently, the positions of the virtual sound sources IS1a, IS2a, and IS3a set in Figure 14(A) are different from the positions of the virtual sound sources IS1c, IS2c, and IS3c set in Figure 14(B).

[0112] Furthermore, as can be seen from the comparison between Figure 14(A) and Figure 14(C), the positional relationship between the reference point in the virtual space and the receiving point RP changes, causing the position of the virtual sound source in the playback space (the position of the virtual sound source relative to the receiving point RP and the speaker) to shift. As a result, the positions of the virtual sound sources IS1a, IS2a, and IS3a set in Figure 14(A) are different from the positions of the virtual sound sources IS1as, IS2as, and IS3as set in Figure 14(C).

[0113] Figures 15(A), 15(B), and 15(C) show examples of setting up a virtual sound source. Figure 15 shows the change in height. Figures 15(A), 15(B), and 15(C) show the change in height.

[0114] The ceiling heights differ between Figure 15(A) and Figure 15(B). Specifically, the distance (height) from the virtual floor wall IWFL to the virtual ceiling wall IWCL in the virtual floor wall IWL shown in Figure 15(A) is different from the distance (height) from the virtual floor wall IWFL to the virtual ceiling wall IWCLL in the virtual floor wall IWLL shown in Figure 15(B).

[0115] As can be seen from the comparison between Figure 15(A) and Figure 15(B), the difference in ceiling height results in different distances and positional relationships between the sound source that forms the basis of the virtual sound source and the virtual walls IWCL and IWCLL on the ceiling. Consequently, the position of the virtual sound source IS1Ca set in Figure 15(A) is different from the position of the virtual sound source IS1CaL set in Figure 15(B).

[0116] The ceiling shapes differ between Figure 15(A) and Figure 15(C). Specifically, the shape of the virtual wall IWCL on the ceiling in the virtual wall IWL shown in Figure 15(A) is different from the shape of the virtual wall IWCLx on the ceiling in the virtual wall IWLx shown in Figure 15(C).

[0117] As can be seen from the comparison between Figure 15(A) and Figure 15(C), the different ceiling shapes result in different positional relationships between the sound source that forms the basis of the virtual sound source and the virtual walls IWCL and IWCLx on the ceiling. Consequently, the position of the virtual sound source IS1Ca set in Figure 15(A) is different from the position of the virtual sound source IS1Cax set in Figure 15(C).

[0118] In this way, the virtual sound source setting unit 502 can optimally set the position of the virtual sound source in the playback space, corresponding to the geometric shape of the virtual space and the positional relationship between the playback space and the virtual space. As a result, the sound signal processing device 10 can clearly localize the sound image of the early reflections, corresponding to the speaker's position coordinates in the playback space, the geometric shape of the virtual space, and the positional relationship between the playback space and the virtual space.

[0119] The virtual sound source setting unit 502 outputs the position coordinates of the virtual sound sources set for each of the multiple regions Area1-Area8 to the output speaker setting unit 5201 of the LDtap circuit 52.

[0120] The output speaker setting unit 5201 sets the virtual sound source IS to be assigned to each speaker based on the position coordinates of the virtual sound source IS, the position coordinates of the receiving point RP, and the position coordinates of the multiple speakers SP1-SP64. Figure 16 is a flowchart showing the process of assigning virtual sound sources to speakers.

[0121] The output speaker setting unit 5201 obtains the position coordinates of the virtual sound source from the virtual sound source setting unit 502 (S141). The output speaker setting unit 5201 obtains the position coordinates of the sound receiving point in the playback space, for example, through user input (S142). The output speaker setting unit 5201 obtains the position coordinates of multiple speakers SP1-SP64, for example, through user input (S143).

[0122] The output speaker setting unit 5201 sets the area of ​​responsibility for the virtual sound source for each speaker based on the positional relationship between the sound receiving point RP and the multiple speakers SP1-SP64 in the playback space (S144).

[0123] More specifically, the output speaker setting unit 5201 sets the area of ​​responsibility for each speaker's virtual sound source as follows. Figures 17(A) and 17(B) illustrate the concept of assigning virtual sound sources to speakers. Figure 17(A) shows the concept of assignment using the azimuth angle φ, and Figure 17(B) shows the concept of assignment using the elevation angle θ. Furthermore, although speaker SP1 will be used as an example below, the output speaker setting unit 5201 sets the area of ​​responsibility for the other speakers SP2-SP64 in the same manner.

[0124] The output speaker setting unit 5201 uses the position coordinates of the receiving point RP and the speaker SP1 to set a straight line (dashed line in Figure 17(A)) passing through the receiving point RP and the speaker SP1. As shown in Figure 17(A), the output speaker setting unit 5201 sets an azimuth angle φ that spreads toward the speaker SP1 on a plane with respect to this straight line (dashed line in Figure 17(A)) with respect to the receiving point RP as the reference point. The azimuth angle φ is the angle in the horizontal direction with respect to the straight line passing through the receiving point RP and the speaker SP1. Also, as shown in Figure 17(B), the output speaker setting unit 5201 sets an elevation angle θ that spreads in the vertical direction perpendicular to the plane with respect to the aforementioned straight line (dashed line in Figure 17(B)). The elevation angle θ is the angle in the vertical direction (direction perpendicular to the horizontal direction) with respect to the straight line passing through the receiving point RP and the speaker SP1.

[0125] The output speaker setting unit 5201 sets the space on the speaker SP1 side of the plane determined by the azimuth angle φ and elevation angle θ as the area RGSP1 responsible for speaker SP1.

[0126] The output speaker setting unit 5201 acquires the position coordinates of multiple virtual sound sources IS (in the case of Figure 17, multiple virtual sound sources ISa-ISg).

[0127] The output speaker setting unit 5201 uses the position coordinates of the multiple virtual sound sources ISa-ISg and the coordinates representing the assigned region RGSP1 to determine whether the multiple virtual sound sources ISa-ISg are located within the assigned region RGSP1. This determination can be achieved using the same method as the grouping of sound sources into regions described above.

[0128] The output speaker setting unit 5201 performs this determination process and, for example, in the case shown in Figure 14, determines that multiple virtual sound sources ISa, ISb, ISC, and ISd are within the assigned region RGSP1, and that multiple virtual sound sources ISe, ISf, and ISg are outside the assigned region RGSP1.

[0129] The output speaker setting unit 5201 assigns the multiple virtual sound sources ISa, ISb, ISC, and ISd, which it has determined to be within its assigned area RGSP1, to speaker SP1.

[0130] The output speaker setting unit 5201 outputs assignment information for multiple virtual sound sources to multiple speakers SP1-SP64 to the coefficient setting unit 5202. At this time, the output speaker setting unit 5201 outputs the position coordinates of the sound receiving point RP, the position coordinates of the multiple speakers SP1-SP64, and the position coordinates of the multiple virtual sound sources to the coefficient setting unit 5202 along with the assignment information.

[0131] The azimuth angle φ is, for example, 60°, and the elevation angle θ is, for example, 45°. These angles of azimuth angle φ and elevation angle θ are examples and can be set and adjusted, for example, by user input.

[0132] The coefficient setting unit 5202 sets the tap coefficients to be applied to LDtap 521-528 using the distances between the receiving point RP and the multiple speakers SP1-SP64, and the distance between the receiving point RP and the virtual sound source IS. The tap coefficients applied to LDtap 521-528 are the gain value and delay amount of LDtap 521-528.

[0133] Figure 18 is a flowchart of the coefficient setting process in LDtap. Figures 19(A) and 19(B) are diagrams illustrating the concept of coefficient setting.

[0134] The coefficient setting unit 5202 uses the position coordinates of the sound receiving point RP and the position coordinates of the multiple speakers SP1-SP64 to calculate the distance between the sound receiving point PR and the multiple speakers SP1-SP64 (speaker distance) (S151).

[0135] The coefficient setting unit 5202 calculates the distance (virtual sound source distance) between the receiving point PR and the multiple virtual sound sources IS (S152).

[0136] The coefficient setting unit 5202 compares the speaker distance and the virtual sound source distance for multiple speakers SP1-SP64 and multiple virtual sound sources IS assigned to each of these speakers SP1-SP64 (S153). For example, in the example shown in Figure 17(A), the speaker distance and the virtual sound source distance are compared for speaker SP1 and multiple virtual sound sources ISa, ISb, ISc, and ISd.

[0137] The coefficient setting unit 5202 sets the tap coefficient using the imaginary sound source distance as is if the speaker distance is less than or equal to the imaginary sound source distance (S153: YES) (S154).

[0138] For example, in the case shown in Figure 19(A), the virtual sound source ISa is farther from the receiving point RP than the speaker SP1, and the virtual sound source distance Lia between the receiving point RP and the virtual sound source ISa is greater than the speaker distance Ls1 between the receiving point RP and the speaker SP1.

[0139] In this case, the coefficient setting unit 5202 sets the tap coefficient using the distance Da1 between the virtual sound source ISa and the speaker SP1. Specifically, the coefficient setting unit 5202 sets the gain value and delay amount to be set for the virtual sound source ISa based on the distance Da1. The coefficient setting unit 5202 sets the gain value to be smaller as the distance Da1 increases, and sets the delay amount to be larger as the distance Da1 increases.

[0140] The coefficient setting unit 5202 determines whether to reproduce the imaginary sound source if the speaker distance is greater than the imaginary sound source distance (S153: NO). In other words, the coefficient setting unit 5202 determines whether to reproduce the imaginary sound source on the receiving point side of the speaker (S155).

[0141] If the coefficient setting unit 5202 is to reproduce a virtual sound source that is closer to the receiving point than the speaker (S155: YES), it moves the position of this virtual sound source (S156). More specifically, the coefficient setting unit 5202 moves the position of the virtual sound source from the receiving point side of the speaker to a position further away from the receiving point than the speaker. In this case, the coefficient setting unit 5202 moves the position of the virtual sound source using the distance difference between the virtual sound source and the speaker. The coefficient setting unit 5202 sets the tap coefficient using the position coordinates of the virtual sound source after the move (S157).

[0142] For example, in the case shown in Figure 19(B), the virtual sound source ISd is closer to the receiving point RP than to the speaker SP1, and the virtual sound source distance Lid between the receiving point RP and the virtual sound source ISd is smaller than the speaker distance Ls1 between the receiving point RP and the speaker SP1.

[0143] In this case, the coefficient setting unit 5202 moves the virtual sound source ISd using the distance difference Dd between the virtual sound source distance Lid and the speaker distance Ls1. More specifically, the coefficient setting unit 5202 moves the virtual sound source ISd to a position on a straight line passing through the receiving point RP and the speaker SP1, and on the opposite side of the receiving point RP from the speaker SP1, by a distance difference Dd. Then, the coefficient setting unit 5202 sets the tap coefficient using this distance difference Dd. Specifically, the coefficient setting unit 5202 sets the gain value and delay amount to be set for the virtual sound source ISd based on the distance difference Dd. The coefficient setting unit 5202 sets the gain value to be smaller the larger the distance difference Dd, and sets the delay amount to be larger the larger the distance difference Dd. Conceptually, the virtual sound source is moved as described above, but in terms of setting the tap coefficient, the coefficient setting unit 5202 only needs to set the tap coefficient based on the distance between the speaker distance and the virtual sound source distance.

[0144] In other words, the coefficient setting unit 5202 moves only the sound receiving point located between the sound receiving point and the speaker. Preferably, the virtual sound source outside the speaker relative to the sound receiving point does not move, but this also includes cases where this outer virtual sound source moves within a predetermined range. For example, even if this outer virtual sound source moves, it is sufficient that the distance between the outer virtual sound source and the speaker is within a predetermined range, which is a range in which the change in the initial reflected sound control signal due to the movement does not cause discomfort to the listener. If the coefficient setting unit 5202 does not reproduce a virtual sound source closer to the sound receiving point than the speaker (S155: NO), it does not set a tap coefficient for this virtual sound source.

[0145] The coefficient setting unit 5202 sets the tap coefficients set for each speaker SP1-SP64 into multiple LDtap units. More specifically, the coefficient setting unit 5202 sets the tap coefficients for each speaker SP1-SP64 into LDtap 521 based on the virtual sound source positions set in area 1. Similarly, the coefficient setting unit 5202 sets the tap coefficients of the virtual sound sources assigned to each speaker SP1-SP64 into LDtap 522-528 based on the virtual sound source positions set in multiple areas 2-8, respectively.

[0146] Multiple LDtap 521-528 units apply gain processing and delay processing to the filtered region-specific sound signals SA1f-SA8f according to the set tap coefficients, and output them to the summing unit 53. More specifically, the tap coefficients are set according to the combination of the virtual sound source positions in multiple regions and each speaker, as described above. Therefore, multiple LDtap 521-528 units set a tap coefficient for each speaker based on the virtual sound source assigned to that speaker. Multiple LDtap 521-528 units apply gain processing and delay processing to the filtered region-specific sound signals SA1f-SA8f for each speaker. Multiple LDtap 521-528 units output the signal with gain processing and delay processing for each speaker.

[0147] For example, if the virtual sound sources ISa, ISb, ISc, and ISd are assigned to speaker SP1, LDtap521 applies gain processing and delay processing to the filtered region-specific sound signal SA1f based on tap coefficients (gain value and delay amount) derived from the virtual sound sources ISa, ISb, ISc, and ISd. Then, LDtap521 outputs this signal to the summing unit 53 for speaker SP1. Multiple LDtap521-528 perform this processing on virtual sound sources for which tap coefficients have been set.

[0148] The addition processing unit 53 adds the LDtap-processed signals from each of the multiple speakers SP1-SP64 output from the multiple LDtap 521-528 for each of the multiple speakers SP1-SP64. The addition processing unit 53 outputs these added signals to the adder 80 as initial reflection sound control signals ER1-ER64 for each of the multiple speakers SP1-SP64.

[0149] By performing this processing, the initial reflection sound control signal generation unit 50 can generate an initial reflection sound control signal having the following characteristics.

[0150] Figures 20(A) and 20(B) are waveform diagrams showing an example of the relationship between the shape of the virtual space and the components of the initial reflection control signal implemented by LDtap. Figure 20(A) shows the case where the virtual space shape is large, and Figure 20(B) shows the case where the virtual space shape is small. Figures 20(A) and 20(B) also show an example of the components of the initial reflection control signal when multiple virtual sound sources are set for a single speaker.

[0151] If the positional relationship between the playback space and the virtual space remains unchanged, and the positions of the receiving point and the speaker also remain unchanged, a larger virtual space shape will result in a wider distribution of the virtual sound source compared to a smaller virtual space shape. Therefore, as shown in Figures 20(A) and 20(B), a larger virtual space shape tends to result in smaller individual components set by LDtap521-528, and a wider distribution range on the time axis.

[0152] In this way, by performing the above-described process, the initial reflected sound control signal generation unit 50 can set the optimal tap coefficient according to the shape of the virtual space.

[0153] Furthermore, even if the positional relationship between the virtual space and the playback space changes, the speaker position changes, or the sound receiving point changes, the initial reflection sound control signal generation unit 50 can set the optimal tap coefficient in response to these changes, just as it does when the shape of the virtual space changes.

[0154] In this process, multiple sound sources OBJ1-OBJ96 are optimally assigned to multiple speakers SP1-SP64 through grouping by multiple regions Area1-Area8. Furthermore, multiple virtual sound sources are optimally configured for these multiple speakers SP1-SP64. Therefore, even with changes in the relationship between the virtual space and the playback space, changes in the receiving point RP, changes in the positions of the multiple speakers SP1-SP64, and changes in the positions of the sound sources OBJ1-OBJ96, the sound signal processing device 10 can clearly localize the sound image using early reflections in response to these changes.

[0155] Furthermore, in the above configuration, the initial reflection sound control signal generation unit 50 can simulate the components of the initial reflection sound control signal from the virtual sound source IS even if the virtual sound source IS is located on the receiving point RP side of the speaker SP. Therefore, for example, when the number of virtual sound sources set for the initial reflection sound control signal is small, the initial reflection sound control signal generation unit 50 can use a virtual sound source that is closer to the receiving point RP than to the speaker SP. In this case, the initial reflection sound control signal generation unit 50 repositions the virtual sound source outside the speaker using the distance difference between the virtual sound source IS and the speaker SP as described above. This allows the initial reflection sound control signal generation unit 50 to suppress the unnaturalness of the initial reflection sound caused by moving the position of the virtual sound source.

[0156] In the above configuration, if the initial reflection sound control signal generation unit 50 is located closer to the receiving point RP than to the speaker SP, the initial reflection sound control signal generation unit 50 may set the virtual sound source IS to the position of the speaker SP. This reduces the processing load on the initial reflection sound control signal generation unit 50 for moving the virtual sound source IS.

[0157] Furthermore, in the above configuration, if the virtual sound source IS is located closer to the receiving point RP than the speaker SP, the initial reflection sound control signal generation unit 50 does not need to use the virtual sound source IS for generating the initial reflection sound control signal. As a result, the initial reflection sound control signal generation unit 50 does not need to perform the processing load of moving the virtual sound source IS, and the processing load of generating the initial reflection sound control signal can be reduced.

[0158] Furthermore, in the above configuration, the initial reflection control signal generation unit 50 sets the components of the initial reflection control signal generated by the virtual sound source, and also performs timbre adjustment using the FIR filters 511-518. The FIR filters 511-518 have the above-mentioned number of taps (e.g., 16,000 taps), which is more than the LDtap 521-528. Also, the time interval between taps of the FIR filters 511-518 (depending on the sampling frequency) is shorter than the time interval between taps of the LDtap 521-528 (depending on the arrangement of the virtual sound source). Therefore, the components of the initial reflection control signal generated by the FIR filters 511-518 are arranged more precisely on the time axis than the components of the initial reflection control signal generated by the LDtap 521-528. In other words, the time axis resolution (temporal resolution) of the FIR filters 511-518 is higher than that of the LDtap 521-528, and the number of components per unit time is greater.

[0159] The initial reflection sound control signal generation unit 50 combines the processing of the FIR filters 511-518 with that of the LDtap 521-528. Therefore, the initial reflection sound control signal generation unit 50 has high resolution in the time domain and can generate initial reflection sound control signals ER1-ER64 with a wider variety of timbres. Figure 21 shows an image of the waveform of the initial reflection sound control signal generated by the initial reflection sound control signal generation unit 50.

[0160] As shown in Figure 21, the initial reflection control signal generation unit 50 can generate an initial reflection control signal with higher resolution and capable of handling a wider variety of timbres, while retaining the initial reflection component from the virtual sound source. In other words, the sound signal processing device 10 can realize initial reflections with timbres that suit the user's preferences while maintaining clear sound image localization using initial reflections from a virtual sound source.

[0161] Furthermore, due to the high resolution of the FIR filter, in the case of short sounds such as the pulse sound of a sound source, the initial reflection control signal may become coarse if only the initial reflection sound component from LDtap is used, resulting in an unnatural timbre. However, with the above configuration and processing, the sound signal processing device 10 can suppress such coarseness of the initial reflection sound and unnatural timbre.

[0162] Furthermore, in the above configuration, the initial reflection sound control signal generation unit 50 sets an allocation region for the virtual sound source IS for each speaker SP, and does not assign virtual sound sources IS outside this region to the speaker SP. As a result, the initial reflection sound control signal generation unit 50 can suppress the generation of excessive initial reflection sound components. Therefore, the sound signal processing device 10 can suppress excessive initial reflection sound and realize a more natural initial reflection sound that corresponds to the virtual space.

[0163] [Generating reverberation control signals] Figure 22 is a functional block diagram showing an example of the configuration of the reverberation control signal generation unit 70. Figure 23 is a flowchart showing an example of the reverberation control signal generation process.

[0164] As shown in Figure 22, the reverberation control signal generation unit 70 includes a PEQ 71, an FIR filter circuit 72, a router 73, a reverberation region setting unit 701, a filter coefficient setting unit 702, a reverberation playback speaker setting unit 703, and an operation unit 700. The FIR filter circuit 72 includes a plurality of FIR filters 721-728.

[0165] The reverberation area setting unit 701 sets multiple reverberation areas Arr1-Arr8 for the playback space. More specifically, the reverberation area setting unit 701 sets the playback space to be divided into multiple reverberation areas Arr1-Arr8 around the entire circumference of the plane, with reference to the center point Psr of the playback space (see Figure 25 below).

[0166] The reverberation region setting unit 701 outputs coordinate information indicating multiple reverberation regions Arr1-Arr8 to the filter coefficient setting unit 702 and the reverberation playback speaker setting unit 703.

[0167] The filter coefficient setting unit 702 sets the filter coefficients for reverberation based on user operations, etc. The filter coefficients for reverberation are set, for example, by the measured results of the impulse response in the real space of the virtual space (another space reproduced in the playback space). Alternatively, the filter coefficients for reverberation may be set virtually using the geometric shape of the virtual space, the material of the walls, etc. In this case, the filter coefficient setting unit 702 sets the filter coefficients for each of the reverberation regions Arr1-Arr8 using coordinate information for each reverberation region Arr1-Arr8.

[0168] The filter coefficient setting unit 702 accepts inputs such as the volume of the virtual space and the surface area of ​​the virtual space through user operations. The filter coefficient setting unit 702 sets a fade-in function for the filter coefficients based on parameters such as the volume of the virtual space and the surface area of ​​the virtual space.

[0169] More specifically, the filter coefficient setting unit 702 calculates the mean free path ρ using the volume V of the virtual space and the surface area S of the virtual space. The formula for calculating the mean free path ρ is ρ = 4V / S. The mean free path is the average propagation distance of sound in a closed space from one reflection off a wall to the next reflection. By dividing the mean free path by the speed of sound c0, the average time required for sound to reflect off a wall to the next reflection can be calculated.

[0170] The filter coefficient setting unit 702 sets the connection timing tc from the mean free path ρ (Figure 23: S231). Specifically, the filter coefficient setting unit 702 sets the connection timing tc using the mean free path ρ, the speed of sound c0, and the order of reflection n. The formula for calculating the connection timing tc is tc = ρ × n / c0.

[0171] As can be seen from this calculation formula, the connection timing tc corresponds to the average time required for n reflections in the virtual space, and when reproducing the nth-order initial reflection sound, it corresponds to the time when the transition to reverberation begins. In other words, the connection timing tc corresponds to the timing when the component of the initial reflection sound control signal generated by the initial reflection sound control signal generation unit 50 disappears.

[0172] By performing this process, the filter coefficient setting unit 702 can optimally set the connection timing tc between the initial reflection and the reverberation according to the geometric shape of the virtual space.

[0173] The filter coefficient setting unit 702 sets the fade-in function using the connection timing tc from the following equation (Figure 23: S232).

[0174]

number

[0175] In this equation, t is the elapsed time since the direct sound was produced, and K is determined from the following equation.

[0176]

number

[0177] Note that in this formula, G REV G is the reverberation gain value at time t=0, and is user-configurable, but for example, the reverberation time is generally the time required for decay to -60dB, so G REV It's best to set it to something like -60dB.

[0178] The filter coefficient setting unit 702 sets the reverberation filter coefficients from the filter coefficients and the fade-in function fin (Figure 23: S233), and outputs them to multiple FIR filters 721-728.

[0179] The reverberation generation signal Sr output from mixer 60 is input to PEQ71. PEQ71 performs predetermined signal processing on the reverberation generation signal Sr and outputs it to multiple FIR filters 721-728.

[0180] By performing signal processing with the PEQ71, the level (signal magnitude) and timbre of the reverberation generation signal Sr can be adjusted. For example, the PEQ71 can refer to the volume of the initial reflection control signal and adjust the level (signal magnitude) of the reverberation generation signal Sr so that the volume of the initial reflection and the volume of the reverberation are approximately the same at the connection timing tc mentioned above. In addition, the timbre and other properties of the PEQ71 can be adjusted according to user settings.

[0181] Multiple FIR filters 721-728 filter the reverberation signal Sr using reverberation filter coefficients to generate region-specific reverberation control signals REVr1-REVr8. For example, FIR filter 721 generates region-specific reverberation control signal REVr1 for region Arr1 by performing a convolution operation on the reverberation signal Sr using the reverberation filter coefficients set for reverberation region Arr1. Similarly, FIR filters 722-728 generate region-specific reverberation control signals REVr2-REVr8 for regions Arr2-Arr8 by performing a convolution operation on the reverberation signal Sr using the reverberation filter coefficients set for each region (Figure 23: S234). Multiple FIR filters 721-728 output region-specific reverberation control signals REVr1-REVr8 to router 73.

[0182] By setting the fade-in function described above, the reverberation control signal will have the waveform shown in Figure 21. Figure 24 is a graph showing example waveforms of the direct sound, early reflection control signal, and reverberation control signal. For convenience, in Figure 24, the reverberation control signal is illustrated by the envelope of each time component. The vertical axis of Figure 24 is in dB.

[0183] As shown in Figure 24(A), the reverberation control signal gradually increases in level according to a fade-in function from the direct sound output timing to the connection timing tc. More specifically, the signal level of the reverberation control signal is -60 dBFs at the direct sound output timing, gradually increases until the connection timing tc, at which point it becomes 0 dBFs. This level is set based on the signal level of the initial reflection control signal at the connection timing tc.

[0184] In the example shown in Figure 24, the fade-in function described above is used to exponentially increase the signal level as it approaches the connection timing tc; in other words, it has characteristics opposite to the decay curve of the reverberation control signal without fade-in processing. Note that the characteristics of the level change of the reverberation control signal due to fade-in processing are not limited to this example; the user can set the desired characteristics by appropriately configuring the fade-in function.

[0185] By performing this processing, the reverberation control signal generation unit 70 can generate a reverberation control signal that accurately reproduces the reverberation of the virtual space using FIR filters 721-728. Furthermore, the signal level of the reverberation control signal gradually increases in the section where the initial reflection control signal exists, reaches a peak value corresponding to the signal level of the initial reflection control signal at the connection timing tc, and then decays.

[0186] As a result, the sound signal processing device 10 can smoothly transition between the initial reflection control signal and the reverberation control signal, which are multiple LD taps that reproduce the distribution of virtual sound sources at multiple sound source locations in the virtual space using reverberation caused by the reverberation control signal. Therefore, the sound output from the sound signal processing device 10 and heard by the user is a sound in which the unnatural feeling when transitioning from the initial reflection to the reverberation is suppressed.

[0187] The reverberation sound playback speaker setting unit 703 groups multiple speakers SP1-SP64 into reverberation sound regions Arr1-Arr8.

[0188] More specifically, the reverberation speaker setting unit 703 configures the playback space to be divided into multiple reverberation regions Arr1-Arr8 around the entire circumference of the plane, for example, with respect to the center point Psr of the playback space. The reverberation speaker setting unit 703 uses the position coordinates of the multiple speakers SP1-SP64 and coordinate information indicating the multiple reverberation regions Arr1-Arr8 to group the multiple speakers SP1-SP64 with respect to the multiple reverberation regions Arr1-Arr8. This grouping can be achieved in the same way as the method for grouping the sound source OBJ described above.

[0189] Figure 25 shows an example of setting up a region for reverberation. In Figure 25, multiple speakers SP1-SP14 are shown for simplicity and clarity. For example, as shown in Figure 25, the reverberation playback speaker setting unit 703 detects that speakers SP6 and SP7 are within the reverberation region Arr1 and groups speakers SP6 and SP7 into the reverberation region Arr1. Similarly, the reverberation playback speaker setting unit 703 also groups the other speakers SP1-SP5 and SP8-SP14 into multiple reverberation regions Arr2-Arr8, respectively.

[0190] The reverberation sound playback speaker setting unit 703 outputs grouping information of multiple speakers SP1-SP64 for multiple reverberation sound regions Arr2-Arr8 to the router 73.

[0191] Router 73 uses grouping information from the reverberation speaker setting unit 703 to assign region-specific reverberation control signals REVr1-REVr8 to multiple speakers SP1-SP64. Based on the assignment, Router 73 outputs region-specific reverberation control signals REVr1-REVr8 as reverberation control signals REV1-REV48 for each of the multiple speakers SP1-SP64.

[0192] For example, router 73 extracts from the grouping information that speakers SP6 and SP7 are grouped in region Arr1. Router 73 assigns region-specific reverberation control signals REVr1 for region Arr1 to speakers SP6 and speaker SP7. Router 73 outputs the region-specific reverberation control signal REVr1 to speaker SP6 as reverberation control signal REV6 for speaker SP6. Router 73 also outputs the region-specific reverberation control signal REVr1 to speaker SP7 as reverberation control signal REV6 for speaker SP7.

[0193] By performing area-specific reverberation control signals REVr1-REVr8 using the router 73, the reverberation control signal generation unit 70 can output the optimal reverberation control signal for each of the multiple speakers SP1-SP64 according to the arrangement of the multiple speakers SP1-SP64.

[0194] [Output adjustment] Figure 26 is a functional block diagram showing an example of the configuration of the output adjustment unit 90. Figure 27 is a flowchart showing an example of the output adjustment process.

[0195] As shown in Figure 26, the output adjustment unit 90 includes a gain control unit 91, a delay control unit 92, a gain delay setting unit 901, an operation unit 900, and a display unit 909. The gain control unit 91 includes multiple gain control units 9101-9168 corresponding to multiple speakers SP1-SP64. The delay control unit 92 includes multiple delay control units 9201-9264 corresponding to multiple speakers SP1-SP64.

[0196] The control unit 900 accepts user input to set acoustic parameters for the playback space (Figure 27: S321). Acoustic parameters for the playback space are parameters for reproducing a desired sound field in the playback space.

[0197] In this case, the acoustic parameters of the playback space are not the individual gain values ​​or delay amounts of the multiple speakers SP1-SP64, but rather weight values ​​that represent the weighting of sound in a predetermined direction within the playback space, and shape values ​​that represent the spread of sound in a predetermined direction within the playback space.

[0198] The weight value is composed of the gain value and the delay amount, and includes the weight values ​​for the front and rear of the playback space, the left and right of the playback space, and the up and down of the playback space. The shape value is composed of the gain value and the delay amount, and includes the shape value in the horizontal direction.

[0199] The display unit 909 is equipped with a GUI. Figure 28 shows an example of a GUI for output adjustment.

[0200] As shown in Figure 28, the GUI 100A includes a settings display window 111, an output status display window 115, and a plurality of controls 116. The plurality of controls 116 include a knob 1161 and an adjustment value display window 1162.

[0201] The multiple operators 116 are operators for setting weight volumes for setting weight values, shape volumes for setting shape values, and so on. The operators 116 for weight volumes each have operators 116 for setting left / right weights, front / back weights, and up / down weights, and each has an operator for setting a gain value and an operator for setting a delay amount. The operators 116 for shape volumes have an operator for setting spread, and also have an operator for setting a gain value and an operator for setting a delay amount.

[0202] The output status display window 115 graphically and schematically displays the sound spread and localization achieved by the weight and shape values ​​set by the multiple operators 116. This allows the user to easily recognize the sound spread and localization set by the multiple operators 116 as an image.

[0203] The user uses the GUI 100A on the display unit 909 to set the acoustic parameters (weight values ​​and delay amounts) they wish to reproduce. The control unit 900 accepts the settings made using the GUI 100A. The control unit 900 outputs these settings (weight values ​​and delay amounts for each acoustic parameter) to the gain delay setting unit 901.

[0204] The gain delay setting unit 901 sets the gain value and delay amount for multiple speakers SP1-SP64 based on the weight value and delay amount of each acoustic parameter. More specifically, the gain delay setting unit 901 performs the following processing.

[0205] The gain delay setting unit 901 acquires the position coordinates of multiple speakers SP1-SP64 arranged in the playback space (S322). The position coordinates are represented in a coordinate system in which, for example, the x-axis is set in the left-right direction of the playback space, the y-axis is set in the front-back direction of the playback space, and the z-axis is set in the up-down direction.

[0206] The gain delay setting unit 901 extracts the maximum and minimum position coordinates of the multiple speakers SP1-SP64 in each axial direction (S323).

[0207] The gain delay setting unit 901 stores coefficient setting formulas. These coefficient setting formulas include, for example, a weight coefficient setting formula for setting weighting in a predetermined direction in the playback space, and a shape coefficient setting formula for setting weighting in a predetermined direction in the playback space.

[0208] The coefficient setting formula for weights includes a formula for setting the gain value for weights and a formula for setting the delay amount for weights. The coefficient setting formula for shapes includes a formula for setting the gain value for shapes and a formula for setting the delay amount for shapes.

[0209] The weight coefficient setting formulas include a coefficient setting formula for the front-to-back direction to set the weighting in the front-to-back direction of the playback space, a coefficient setting formula for the left-to-right direction to set the weighting in the left-to-right direction of the playback space, and a coefficient setting formula for the up-to-down direction to set the weighting in the up-to-down direction of the playback space.

[0210] The coefficient setting formula for shape includes the coefficient setting formula for the left-right direction of the playback space.

[0211] The formula for setting the gain coefficient for the weight is, for example, a linear function that combines the set weight gain value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the gain value is set. The gain value is determined proportionally to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.

[0212] The formula for setting the delay coefficient for the weight is, for example, a linear function that combines the delay amount of the set weight value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the delay amount is set. This formula determines the delay amount in proportion to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.

[0213] The formula for setting the gain coefficient for a shape is, for example, a linear function that combines the gain value of the set shape value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the gain value is set. The gain value is determined proportionally to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.

[0214] The formula for setting the delay coefficient for shapes is, for example, a linear function that combines the delay amount of the set shape value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the delay amount is set. This formula determines the delay amount in proportion to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.

[0215] The gain delay setting unit 901 calculates the gain value and delay amount for each speaker to be set using the set gain value and delay amount (acoustic parameter), the maximum and minimum values ​​of the extracted position coordinates, and the coefficient setting formula (S324).

[0216] By using this process, the gain delay setting unit 901 can automatically calculate and set the gain values ​​and delay amounts of multiple speakers SP1-SP64 arranged in the playback space using a coefficient setting formula, without having to manually set each one individually.

[0217] The gain delay setting unit 901 outputs the gain value set for each of the multiple speakers SP1-SP64 to the multiple gain control units 9101-9164. The gain delay setting unit 901 also outputs the delay amount set for each of the multiple speakers SP1-SP64 to the multiple delay control units 9201-9264.

[0218] Each of the gain control units 9101-9164 receives speaker signals Sat1-Sat64 corresponding to the multiple speakers SP1-SP64 from the adder 80.

[0219] Multiple gain control units 9101-9164 control the signal levels of speaker signals Sat1-Sat64 using the gain values ​​set in each unit, and output them to multiple delay control units 9201-9264. For example, gain control unit 9101 controls the signal level of speaker signal Sat1 using the gain value set in gain control unit 9101, and outputs it to delay control unit 9201. Similarly, gain control units 9102-9164 control the signal levels of speaker signals Sat2-Sat64 using the gain values ​​set in gain control units 9102-9164, and output them to delay control units 9202-9164, respectively.

[0220] Multiple delay control units 9201-9164 control the signal level of signals input from multiple gain control units 9101-9164 using the delay amount set for each unit, and output them to multiple speakers SP1-SP64. For example, delay control unit 9201 controls the signal level of signals input from gain control unit 9101 using the delay amount set for delay control unit 9201, and outputs it to speaker SP1. Similarly, delay control units 9202-9164 control the signal level of signals input from gain control units 9102-9164 using the delay amounts set for delay control units 9202-9164, and output them to speakers SP1-SP64, respectively.

[0221] With this configuration, the sound signal processing device 10 can easily realize a desired sound field corresponding to the set acoustic parameters using an initial reflection control signal and a reverberation control signal, without forcing the user to individually and cumbersomely configure multiple speakers. As a result, for example, the sound signal processing device 10 can easily realize a sound field that produces a Haas effect at a predetermined position in the playback space.

[0222] (Example of sound field realization through output control) Figures 29(A) and 29(B) show examples of settings for weighting sound at the rear of the playback space. Figure 29(A) shows an example of setting the gain value and delay amount, and Figure 29(B) shows an image of sound weighting based on the settings in Figure 25(A). Note that in Figures 29(A) and 29(B), the arrangement of 14 speakers SP1-SP14 is shown for the sake of simplicity and clarity.

[0223] In the embodiments shown in Figures 29(A) and 29(B), acoustic parameters such as the gain value and delay amount at the rear end are set. The gain delay setting unit 901 sets the gain value and delay amount at the front end to values ​​with the opposite sign of the gain value and delay amount at the rear end. The gain delay setting unit 901 calculates the maximum and minimum position coordinates of the 14 speakers SP1-SP14.

[0224] The gain delay setting unit 901 calculates the gain values ​​of the 14 speakers SP1-SP14 using the gain values ​​of the rear and front ends, the maximum and minimum position coordinates of the 14 speakers SP1-SP14, and a coefficient setting formula for the front-to-back direction (for setting gain values) which sets the weighting of the playback space in the front-to-back direction.

[0225] Furthermore, the gain delay setting unit 901 calculates the delay amounts for the 14 speakers SP1-SP14 using the delay amounts at the rear and front ends, the maximum and minimum position coordinates of the 14 speakers SP1-SP14, and a coefficient setting formula for the front-to-back direction (for setting the delay amount) which sets the weighting of the playback space in the front-to-back direction.

[0226] This process allows the sound signal processing device 10 to automatically and easily set acoustic parameters such that the gain and delay values ​​are higher for speakers further back in the playback space and lower for speakers further forward, as shown in Figure 29(A). As a result, the sound signal processing device 10 can easily realize a sound field that extends to the rear of the playback space and where reverberation is localized (see Figure 29(B)).

[0227] Although this explanation uses an example in the front-to-back direction, the sound signal processing device 10 can similarly realize a weighted sound field in the left-to-right direction and the height direction (up and down direction).

[0228] Figures 30(A) and 30(B) show examples of settings for creating a lateral sound spread in the playback space. Figure 30(A) shows an example of setting the gain value and delay amount, and Figure 30(B) shows an image of the sound spread resulting from the settings in Figure 30(A). Note that in Figures 30(A) and 30(B), the arrangement of 14 speakers SP1-SP14 is shown for the sake of simplicity and clarity.

[0229] In the embodiments shown in Figures 30(A) and 30(B), an acoustic parameter such as a numerical value representing the sound spread (spread setting value) is set. The gain delay setting unit 901 calculates the maximum and minimum position coordinates of the 14 speakers SP1-SP14.

[0230] The gain delay setting unit 901 calculates the gain values ​​of the 14 speakers SP1-SP14 using the spread setting value, the maximum and minimum position coordinates of the 14 speakers SP1-SP14, and the coefficient setting formula for shape (for setting gain values).

[0231] Furthermore, the gain delay setting unit 901 calculates the delay amount for the 14 speakers SP1-SP14 using the delay amounts at the rear and front ends, the maximum and minimum position coordinates of the 14 speakers SP1-SP14, and a coefficient setting formula for shape (for delay amount setting).

[0232] This process allows the sound signal processing device 10 to automatically and easily set acoustic parameters such that the gain and delay values ​​are larger for speakers closer to the lateral ends of the playback space, and smaller for speakers closer to the lateral center, as shown in Figure 30(A). As a result, the sound signal processing device 10 can easily realize a sound field with a wide lateral spread in the playback space and localized reverberation (see Figure 30(B)).

[0233] Furthermore, by setting the acoustic parameters as described above, the sound signal processing device 10 can achieve not only weighting in the front-to-back direction, left-to-right direction, and lateral spread of the playback space, but also weighting and spread in the height direction (up and down direction) of the playback space. For example, Figure 31 is a diagram that shows an image of sound spread when the height direction is extended.

[0234] The sound signal processing device 10 increases the gain value and delay amount of the speaker SPU on the ceiling side compared to the speaker SPL and SPR closer to the floor. As a result, the sound signal processing device 10 can easily realize a sound field (see Figure 31) that has more breadth in the ceiling direction of the playback space and where reverberation is localized.

[0235] Furthermore, in the above configuration, the output adjustment unit 90 outputs the output signals So1-So64 to multiple speakers SP1-SP64. However, the sound signal processing device may also output the output signals So1-So64 after binaural processing.

[0236] Figure 32 is a functional block diagram showing the configuration of a sound signal processing device with binaural playback function. As shown in Figure 32, the sound signal processing device with binaural playback function 10A differs from the sound signal processing device 10 described above in that it includes an output adjustment unit 90A, a reverberation processing unit 97, a selection unit 98, and a binaural processing unit 99.

[0237] The output adjustment unit 90A generates multiple output signals So1-So64 from multiple speaker signals Sat1-Sat64 output from the adder 80 using the same processing as the output adjustment unit 90 described above.

[0238] The output adjustment unit 90A can select the output target. The selection of the output target is performed, for example, by user input using the GUI described above. More specifically, the GUI displays an operator that allows selection between speaker output and binaural output, and the output target is selected by operating this operator.

[0239] If speaker output is selected, the output adjustment unit 90A outputs multiple output signals So1-So64 to multiple speakers SP1-AP64 respectively (same processing as output adjustment unit 90). If binaural output is selected, the output adjustment unit 90A outputs multiple output signals So1-So64 to the selection unit 98.

[0240] The reverberation processing unit 97 receives sound signals S1-S96 from multiple sound sources OBJ1-OBJ96. The reverberation processing unit 97 adds an initial reflection control signal and a reverberation control signal to the multiple sound signals S1-S96 and outputs them to the selection unit 98. The initial reflection control signals for the multiple sound signals S1-S96 are set based on the position coordinates of the multiple sound sources OBJ1-OBJ96. The reverberation processing unit 97 outputs the multiple reverberation-processed sound signals S1'-S96' to the selection unit 98.

[0241] Multiple output signals So1-So64 and multiple reverberation-processed sound signals S1'-S96' are input to the selection unit 98. The selection unit 98 selects the multiple output signals So1-So64 and the reverberation-processed sound signals S1'-S96' based on user input using the GUI described above. More specifically, the GUI displays an operator that allows the user to select between the sound processed by the sound signal processing device 10A and the sound processed virtually based on the position coordinates of the sound sources OBJ1-OBJ96. The output target is selected by operating this operator.

[0242] If a sound processed by the sound signal processing device 10A is selected, the selection unit 98 selects multiple output signals So1-So64 and outputs them to the binaural processing unit 99. If a sound processed by virtual sound processing based on the position coordinates of sound sources OBJ1-OBJ96 is selected, the selection unit 98 selects multiple reverberation-processed sound signals S1'-S96' and outputs them to the binaural processing unit 99.

[0243] The binaural processing unit 99 applies binaural processing to the input sound signal. More specifically, if multiple output signals So1-So64 are input, the binaural processing unit 99 applies binaural processing to the multiple output signals So1-So64. If multiple sound signals S1'-S96' after reverberation processing are input, the binaural processing unit 99 applies binaural processing to the multiple sound signals S1'-S96' after reverberation processing.

[0244] Furthermore, binaural processing utilizes head-related transfer functions, and its detailed nature is already known; therefore, a detailed explanation of binaural processing will be omitted.

[0245] The binaural processing unit 99 outputs a binaurally processed 2-channel audio signal.

[0246] This allows the user to hear the sound generated by the sound signal processing device 10A and the sound with virtual reverberation processing applied based on the position coordinates of the sound sources OBJ1-OBJ96, through binaural playback. Therefore, the user can easily verify, using headphones or the like, whether the acoustic processing applied by the sound signal processing device 10A is reproducing the acoustics of the virtual space without physically constructing a playback space. The acoustic processing applied by the sound signal processing device 10A includes, for example, the grouping of sound sources, setting of the initial reflection sound control signal, setting of the reverberation sound control signal, and setting of the output control signal. By comparing the sounds in this way, the user can adjust the settings of the acoustic processing described above to reproduce the acoustics of the virtual space more faithfully.

[0247] Note that binaural playback is not limited to headphones; it can also be performed using stereo speakers or other devices.

[0248] The description of this embodiment is illustrative in all respects and not restrictive. The scope of the invention is indicated by the claims, rather than by the embodiments described above. Furthermore, the scope of the invention is intended to include all modifications within the meaning and scope equivalent to the claims. [Explanation of symbols]

[0249] 10, 10A: Audio signal processing device 30: Region setting unit 40: Grouping unit 41: Sound source position detection unit 42: Region determination unit 50: Initial reflection sound control signal generation unit 51: FIR filter circuit 52: LDtap circuit 53: Addition processing unit 60: Mixer 70: Reverberation sound control signal generation unit 71: PEQ 72: FIR filter circuit 73: Router 80: Adder 90, 90A: Output adjustment unit 91: Gain control unit 92: Delay control unit 97: Reverberation processing unit 98: Selection unit 99: Binaural processing unit 100, 100A: GUI 400: Matrix mixer 500: Operation unit 501: Tone color setting unit 502: Virtual sound source setting unit 511 - 518: FIR filter 521 - 528: LDtap 700: Operation unit 701: Region setting unit for reverberation sound 702: Filter coefficient setting unit 703: Reproduction speaker setting unit for reverberation sound 721 - 728: FIR filter 900: Operation unit 901: Gain delay setting unit 909: Display unit 5201: Output speaker setting unit 5202: Coefficient setting unit 9101 - 9164: Gain control unit 9201 - 9264: Delay control unit

Claims

1. It is determined whether the virtual sound source representing the reflected sound of the target acoustic space is within a region consisting of a predetermined azimuth angle and elevation angle from the center position of the acoustic space and the position of the speaker. Only when within the aforementioned region, the speaker controls the output of the initial reflection sound control signal of the virtual sound source. Audio signal processing method.

2. The virtual sound source is defined by the shape of the acoustic space, the position of the speaker in the acoustic space, and the position of the sound receiving point. The sound signal processing method according to claim 1.

3. The gain value and delay amount of the initial reflected sound control signal are set based on the positional relationship between the virtual sound source, the receiving point, and the speaker. The sound signal processing method according to claim 1 or claim 2.

4. The virtual sound sources are classified into a first virtual sound source located between the speaker and the receiving point and representing a first sound source of reflected sound in the target acoustic space, and a second virtual sound source located outside the speaker and representing a second sound source of reflected sound in the target acoustic space. The method for setting the position of the first virtual sound source and the second virtual sound source is made different. The sound signal processing method according to any one of claims 1 to 3.

5. If the virtual sound source representing the reflected sound of the target acoustic space is the first virtual sound source, then the position of the first virtual sound source is moved to a position where it can be reproduced using the position of a speaker in the vicinity of the first virtual sound source. The sound signal processing method according to claim 4.

6. A speaker that reproduces the reflected sound of the target acoustic space, An initial reflection sound control unit determines whether a virtual sound source representing the reflected sound of the target acoustic space is within a region consisting of a predetermined azimuth angle and elevation angle from the center position of the acoustic space and the position of the speaker, and only if it is within the region, it controls the output of an initial reflection sound control signal from the virtual sound source to the speaker. Equipped with, Sound signal processing device.

7. The initial reflected sound control unit is The virtual sound source is defined by the shape of the acoustic space, the position of the speaker in the acoustic space, and the position of the sound receiving point. The sound signal processing device according to claim 6.

8. The aforementioned initial reflection sound control unit is The gain value and delay amount of the initial reflection control signal are set based on the positional relationship between the virtual sound source and the speaker. The sound signal processing device according to claim 6 or claim 7.

9. The aforementioned initial reflection sound control unit is The virtual sound sources are classified into a first virtual sound source located between the speaker and the receiving point and representing a first sound source of reflected sound in the target acoustic space, and a second virtual sound source located outside the speaker and representing a second sound source of reflected sound in the target acoustic space. The method for setting the gain value and delay amount of the initial reflection control signal is made different for the first virtual sound source and the second virtual sound source. The sound signal processing device according to any one of claims 6 to 8.

10. The aforementioned initial reflection sound control unit is If the virtual sound source is the first virtual sound source, move the position of the first virtual sound source to a position where it can be reproduced using the position of a speaker in the vicinity of the first virtual sound source. The sound signal processing device according to claim 9.

Citation Information

Patent Citations

  • Device, method, and program for processing sound

    WO2016208406A1