Sound signal processing method, sound signal processing apparatus, and recording medium

By dividing the audio system into multiple components and grouping and setting virtual sound sources, the initial reflected sound and reverberation control signals are generated and adjusted, solving the problem of unclear virtual spatial sound image positioning in the prior art, and realizing realistic sound field reproduction and smooth sound source movement processing.

CN115119103BActive Publication Date: 2025-10-21YAMAHA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210248168.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-03-19
Filing Date
2022-03-14
Publication Date
2025-10-21
Estimated Expiration
2042-03-14

Smart Images

  • Figure CN115119103B_ABST
    Figure CN115119103B_ABST
Patent Text Reader

Abstract

The sound signal processing method and the sound signal processing apparatus can clearly reproduce sound image localization in a virtual space. The sound signal processing method classifies a virtual sound source representing a reflected sound of a target sound space into a first virtual sound source located between a position of a speaker and a position of a sound receiving point, representing a first sound source of the reflected sound of the target sound space, and a second virtual sound source located outside the speaker, representing a second sound source of the reflected sound of the target sound space. The position of the first virtual sound source is moved to a position at which the first virtual sound source can be played using a position of a speaker near the first virtual sound source only in the case of the first virtual sound source.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing device for performing predetermined processing on sound input from a sound source. Background Art

[0002] In an audio system in a space such as a hall, sound image localization relative to a sound source is performed by speakers arranged in the space.

[0003] For example, the speech processing device described in Patent Document 1 outputs the sound of an audio object (sound source) through two or more speakers near the audio object (sound source). In this case, the speech processing device described in Patent Document 1 calculates the gain of the speech signal output to each speaker using the position information and sound image information of the audio object (sound source).

[0004] Patent Document 1: International Publication No. 2013 / 208406

[0005] However, in the above-mentioned conventional configuration, it is impossible to clearly reproduce the sound image localization in the virtual space in the space where the speakers are installed. Summary of the Invention

[0006] Therefore, an object of one embodiment of the present invention is to clearly reproduce the sound image localization in a virtual space.

[0007] The sound signal processing method classifies a virtual sound source representing reflected sound in a target acoustic space into a first virtual sound source located between a speaker and a sound receiving point and representing the first sound source of reflected sound in the target acoustic space, and a second virtual sound source located outside the speaker and representing the second sound source of reflected sound in the target acoustic space. In the case of only the first virtual sound source, the position of the first virtual sound source is moved to a position where the sound can be reproduced using a speaker located near the first virtual sound source.

[0008] Effects of the Invention

[0009] The sound signal processing method can clearly reproduce the sound image localization in the virtual space. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] Figure 1 This is a functional block diagram showing the configuration of an acoustic system including the audio signal processing device according to an embodiment of the present invention.

[0011] Figure 2 This is a flowchart of a sound signal processing method according to an embodiment of the present invention.

[0012] Figure 3This is a diagram showing discrete waveforms of sounds including normal direct sound, initial reflected sound, and reverberant sound (rear reverberant sound).

[0013] Figure 4 (A) Figure 4 (B) is a diagram showing the concept of setting a virtual sound source.

[0014] Figure 5 4 is a functional block diagram showing an example of the configuration of the grouping unit 40 .

[0015] Figure 6 This is a flowchart showing a method for grouping sound sources.

[0016] Figure 7 This diagram shows the concept of grouping multiple sound sources into multiple areas.

[0017] Figure 8 (A) is a flowchart showing a method of grouping sound sources using representative points. Figure 8

[0018] (B) is a flowchart showing a method of grouping sound sources using area boundaries.

[0019] Figure 9 This is a flowchart showing an example of a method of grouping based on the movement of sound sources.

[0020] Figure 10 3 is a functional block diagram showing an example of the configuration of the initial reflected sound control signal generating unit 50 .

[0021] Figure 11 This is a diagram showing an example of a GUI.

[0022] Figure 12 This is a flowchart showing an example of a virtual sound source setting process.

[0023] Figure 13 (A) Figure 13 (B) is a diagram showing an example of setting each virtual sound source when the geometric shapes are different.

[0024] Figure 14 (A) Figure 14 (B) and Figure 14 (C) is a diagram showing an example of setting a virtual sound source.

[0025] Figure 15 (A) Figure 15 (B) Figure 15 (C) is a diagram showing an example of setting a virtual sound source.

[0026] Figure 16 This is a flowchart showing the process of allocating virtual sound sources to speakers.

[0027] Figure 17 (A) Figure 17 (B) is a diagram showing the concept of allocating virtual sound sources to speakers.

[0028] Figure 18 This is a flowchart showing the coefficient setting process of LDtap.

[0029] Figure 19 (A) Figure 19 (B) is a diagram for explaining the concept of coefficient setting.

[0030] Figure 20 (A) shows an example of the LDtap coefficient when the virtual space shape is large. Figure 20 (B) shows an example of the LDtap coefficient when the virtual space shape is small.

[0031] Figure 21 3 is a diagram showing the waveform of the initial reflected sound control signal generated by the initial reflected sound control signal generating unit 50 .

[0032] Figure 22 1 is a functional block diagram showing an example of the configuration of the reverberant sound control signal generating unit 70 .

[0033] Figure 23 This is a flowchart showing an example of a process for generating a reverberant sound control signal.

[0034] Figure 24 This is a graph showing waveform examples of the direct sound, the initial reflected sound control signal, and the reverberant sound control signal.

[0035] Figure 25 This is a diagram showing an example of area settings for reverberant sound.

[0036] Figure 26 1 is a functional block diagram showing an example of the configuration of the output adjustment unit 90 .

[0037] Figure 27 This is a flowchart showing an example of output adjustment processing.

[0038] Figure 28 This is a diagram showing an example of a GUI for output adjustment.

[0039] Figure 29 (A) Figure 29 (B) is a diagram showing a setting example in which sound is localized and spread at the rear side of the reproduction space.

[0040] Figure 30 (A) Figure 30 (B) is a diagram showing a setting example in which sound is localized and spread in the lateral direction of the reproduction space.

[0041] Figure 31 This is a diagram schematically showing the spread of sound in a case where the sound spreads in the height direction.

[0042] Figure 32 This is a functional block diagram showing the structure of an audio signal processing device with a binaural playback function. DETAILED DESCRIPTION

[0043] The sound signal processing method and the sound signal processing device according to the embodiments of the present invention will be described with reference to the accompanying drawings. In the following embodiments, the sound signal processing method and the sound signal processing device will be described in general terms first, followed by a detailed description of each process and each structure.

[0044] In this embodiment, the playback space is a space where users (listeners) listen to sounds (direct sound, initial reflected sound, and reverberant sound) from a sound source using speakers, etc. The virtual space is a space with a sound field (acoustics) different from that of the playback space, and the initial reflected sound and reverberant sound generated by this sound field are reproduced (simulated) in the playback space.

[0045] [Schematic Structure of Sound Signal Processing Device]

[0046] Figure 1 This is a functional block diagram showing the configuration of an acoustic system including the audio signal processing device according to an embodiment of the present invention.

[0047] like Figure 1 As shown, the sound signal processing device 10 includes a region setting unit 30, a grouping unit 40, an initial reflected sound control signal generating unit 50, a mixer 60, a reverberant sound control signal generating unit 70, an adder 80, and an output adjustment unit 90. The sound signal processing device 10 is implemented, for example, by an electronic circuit or a processing unit such as a computer that implements the region setting unit 30, the grouping unit 40, the initial reflected sound control signal generating unit 50, the mixer 60, the reverberant sound control signal generating unit 70, the adder 80, and the output adjustment unit 90. The portion consisting of the adder 80 and the output adjustment unit 90 corresponds to the "output signal generating unit" of the present invention.

[0048] The audio signal processing device 10 is connected to a plurality of speakers SP1 to SP64. Figure 1 Although a method using 64 speakers is shown, the number of speakers is not limited thereto.

[0049] The sound signal processing device 10 receives input of sound signals S1 to S96 from a plurality of sound sources OBJ1 to OBJ96. Figure 1 Although a method using 96 sound sources is shown, the number of sound sources is not limited thereto.

[0050] The region setting unit 30 divides the broadcast space into a plurality of regions and sets information (region information) related to the divided regions. The region information is the position coordinates that determine the boundaries of the regions and the position coordinates of representative points set in the regions.

[0051] The area setting unit 30 outputs the area information of the set plurality of areas Area1 to Area8 to the grouping unit 40. Figure 1 , eight regions are set, but the number of regions is not limited thereto.

[0052] The grouping unit 40 groups sound sources OBJ1-OBJ96 into a plurality of areas Area1-Area8. Based on the grouping results, the grouping unit 40 uses the sound signals S1-S96 from the sound sources OBJ1-OBJ96 to generate area-specific sound signals SA1-SA8 for each of the areas Area1-Area8. For example, the grouping unit 40 mixes the sound signals from the plurality of sound sources grouped into area Area1 to generate the area-specific sound signal SA1.

[0053] The grouping unit 40 outputs the plurality of area-specific sound signals SA1 to SA8 to the initial reflected sound control signal generator 50 . The grouping unit 40 also outputs the sound signals S1 to S96 from the sound sources OBJ1 to OBJ96 to the mixer 60 .

[0054] The initial reflected sound control signal generator 50 generates initial reflected sound control signals ER1-ER64 for each of the multiple speakers SP1-SP64 based on the multiple area-specific sound signals SA1-SA8. The initial reflected sound control signals ER1-ER64 are signals output to the speakers SP1-SP64 to simulate the initial reflected sound in the virtual space within the playback space. The initial reflected sound control signal generator 50 outputs the generated initial reflected sound control signals ER1-ER64 to the adder 80.

[0055] In brief (the detailed structure and processing will be described later), the initial reflected sound control signal generator 50 uses the positions of the speakers SP1-SP64 arranged in the playback space and the geometry of the virtual space to set virtual sound sources in the playback space. The specific settings of the virtual sound sources will be described later. The initial reflected sound control signal generator 50 generates initial reflected sound control signals ER1-ER64 that simulate initial reflected sound in the virtual space using the virtual sound sources. At this point, the initial reflected sound control signal generator 50 adjusts the desired timbre of the initial reflected sound control signals ER1-ER64.

[0056] The mixer 60 is an adding mixer and mixes the sound signals S1 - S96 from the sound sources OBJ1 - OBJ96 to generate a reverberant sound generating signal Sr. The mixer 60 outputs the reverberant sound generating signal Sr to the reverberant sound control signal generating unit 70 .

[0057] The reverberant sound control signal generator 70 generates reverberant sound control signals REV1-REV64 for each of the multiple speakers SP1-SP64 based on the reverberant sound generation signal Sr. These reverberant sound control signals REV1-REV64 are signals output to the speakers SP1-SP64 to simulate the reverberant sound (rear reverberant sound) in the virtual space within the playback space. The reverberant sound control signal generator 70 outputs the generated reverberant sound control signals REV1-REV64 to the adder 80.

[0058] In summary (the detailed structure and processing will be described later), the reverberant sound control signal generator 70 divides the playback space into multiple reverberant sound setting areas and generates reverberant sound control signals for each of the multiple reverberant sound setting areas. The reverberant sound control signal generator 70 allocates the multiple speakers SP1-SP64 to the multiple reverberant sound setting areas. Based on this allocation, the reverberant sound control signal generator 70 sets a reverberant sound control signal for each reverberant sound setting area for the multiple speakers SP1-SP64.

[0059] At this time, the reverberant sound control signal generator 70 sets the timing for connecting the initial reflected sound and the reverberant sound based on the geometry of the playback space. The reverberant sound control signal generator 70 gradually increases the level (amplitude) of the reverberant sound control signal before the connection timing, and gradually decreases the level (amplitude) of the reverberant sound control signal during and after the connection timing.

[0060] Adder 80 adds the initial reflected sound control signal and the reverberant sound control signal generated for each of the multiple speakers SP1 to SP64 to generate multiple speaker signals Sat1 to Sat64. For example, adder 80 adds the initial reflected sound control signal for speaker SP1 and the reverberant sound control signal for speaker SP1 to generate speaker signal Sat1. Adder 80 outputs the multiple speaker signals Sat1 to Sat64 to the output adjustment unit 90.

[0061] The output adjustment unit 90 performs gain and hysteresis control on the plurality of speaker signals Sat1-Sat64 to generate output signals So1-So64. The output adjustment unit 90 outputs the output signals So1-So64 to the plurality of speakers SP1-SP64. For example, the output adjustment unit 90 performs gain and hysteresis control on the speaker signal Sat1 for speaker SP1 to generate the output signal So1. The output adjustment unit 90 outputs the output signal So1 to the speaker SP1.

[0062] Roughly speaking (the detailed structure and processing will be described later), the output adjustment unit 90 receives the input of the sound parameters of the playback space. The sound parameters are, for example, parameters for setting the expansion of the space in the width direction of the sound space, the expansion of the space behind the sound receiving point, and the expansion of the space in the ceiling direction of the sound space. The output adjustment unit 90 centrally sets the gain values ​​and hysteresis (delay amount) of the multiple speaker signals Sat1-Sat64 based on the position coordinates and sound parameters of the multiple speakers SP1-SP64. Centralized setting means that the gain value and hysteresis of each speaker are not set separately for each speaker. For example, the gain value and hysteresis of each speaker are set by inputting the position coordinates of each speaker into a specific calculation formula shared by all speakers. The output adjustment unit 90 uses the set gain value and hysteresis value to perform gain control and hysteresis control on the multiple speaker signals Sat1-Sat64.

[0063] [Overview of the Sound Signal Processing Method]

[0064] Figure 2 This is a flowchart of a sound signal processing method according to an embodiment of the present invention. Figure 2 Shown by Figure 1 The sound signal processing method implemented by the sound signal processing device 10. Figure 2 The contents of each process shown in the above Figure 1 This has been explained in the description, so it is briefly recorded.

[0065] (Grouping of sound sources OBJ1-OBJ96)

[0066] The grouping unit 40 groups the plurality of sound sources OBJ1 to OBJ96 into groups for each of the plurality of areas Area1 to Area8 ( S11 ).

[0067] (Generation of Initial Reflected Sound Control Signal)

[0068] The initial reflected sound control signal generator 50 sets the timbre for the initial reflected sound for each group (S12). The initial reflected sound control signal generator 50 sets a virtual sound source for each group (S13). The initial reflected sound control signal generator 50 uses the timbre and the virtual sound source to generate an initial reflected sound control signal for each of the multiple speakers SP1-SP64 (S14).

[0069] (Generation of reverberation control signal)

[0070] The mixer 60 adds the sound signals S1-S96 of the multiple sound sources OBJ1-OBJ96 (S21). The reverberant sound control signal generator 70 sets the connection timing of the initial reflected sound and the reverberant sound based on the geometric shape of the playback space (S22). The reverberant sound control signal generator 70 generates a reverberant sound control signal using the set connection timing (S23). The reverberant sound control signal generator 70 distributes the generated reverberant sound control signal to the multiple speakers SP1-SP64 based on the position coordinates of the multiple speakers SP1-SP64 in the playback space (S24).

[0071] (Output processing to multiple speakers)

[0072] The adder 80 adds the initial reflected sound control signal and the reverberant sound control signal for each of the plurality of speakers SP1 to SP64 to generate speaker signals Sat1 to Sat64 ( S31 ).

[0073] The output adjustment unit 90 generates output signals So1-So64 from the speaker signals Sat1-Sat64 using acoustic parameters that realize localization of reverberation and spatial expansion in the reproduction space (S32). The output adjustment unit 90 outputs the output signals So1-So64 to the plurality of speakers SP1-SP64 (S33).

[0074] By using the above-described configuration and processing, the audio signal processing device 10 (audio signal processing method) achieves the following various effects.

[0075] (1) The sound signal processing device 10 (sound signal processing method) groups sound sources for each area obtained by dividing the playback space and generates initial reflections, thereby achieving clear sound image localization and rich spatial expansion. In this case, the reverberant sound is constant throughout the entire playback space, and only the initial reflections vary depending on the position of the sound source. Therefore, for example, if the position of the sound source moves, the sound of that sound source moves more smoothly.

[0076] (2) The sound signal processing device 10 (sound signal processing method) generates the initial reflected sound control signal using a virtual sound source, thereby being able to simulate the initial reflected sound based on the geometric shape of the virtual space more realistically in the reproduction space.

[0077] (3) The sound signal processing device 10 (sound signal processing method) adjusts the timbre of the initial reflected sound control signal, thereby being able to eliminate unnaturalness in the timbre of the initial reflected sound simulated only by the virtual sound source, for example.

[0078] (4) The sound signal processing device 10 (sound signal processing method) sets the connection timing of the initial reflected sound control signal and the reverberant sound control signal according to the geometric shape of the playback space, thereby enabling a smoother and more natural connection from the initial reflected sound to the reverberant sound.

[0079] (5) The sound signal processing device 10 (sound signal processing method) centrally adjusts the gain value and lag amount of the speaker signal Sat1-Sat64 containing the initial reflected sound control signal and the reverberant sound control signal, thereby realizing the sound field desired by the user in the playback space through easier operation input.

[0080] [Details of Each Signal Processing Unit and Each Process]

[0081] The following describes in detail the signal processing units and the processing. First, referring to the drawings, the initial reflected sound, reverberant sound, and virtual sound source necessary for understanding the present invention will be described.

[0082] [Initial reflection and reverberation sound]

[0083] Figure 3 This diagram shows the discrete waveforms of a sound, including typical direct sound, initial reflected sound, and reverberant sound (reverberant sound). For example, a hall where performances or content are played is a closed space surrounded by walls. When sound is generated in this enclosed space, the direct sound, initial reflected sound, and reverberant sound (reverberant sound) reach the sound collection point.

[0084] Direct sound is the sound that reaches the receiving point directly from the place where the sound is generated.

[0085] Initial reflected sound is the sound that arrives at the sound receiving point earlier after being reflected by the walls, floor, and ceiling at the sound source. Therefore, the initial reflected sound arrives at the sound receiving point after the direct sound. In addition, the volume (level) of the initial reflected sound is lower than the volume (level) of the direct sound. If the number of reflections is 1, it is a single reflected sound, and if it is n times, it is n reflected sounds. The direction of arrival and the volume of the initial reflected sound at the sound receiving point are greatly affected by the location where the sound originates.

[0086] The reverberant sound reaches the sound receiving point after the initial reflected sound. Reverberant sound is the sound that reaches the sound receiving point after multiple reflections from the sound originating at the source. In other words, reverberant sound is the sound that reaches the sound receiving point after multiple reflections and attenuation. Therefore, the volume (level) of the reverberant sound is lower than that of the initial reflected sound. Furthermore, the direction of arrival and volume of the reverberant sound are less affected by the source location than the initial reflected sound.

[0087] [Virtual sound source]

[0088] Figure 4 (A) Figure 4 (B) is a diagram showing the concept of setting a virtual sound source. Figure 4 (A) Figure 4 While (B) illustrates the concept of setting virtual sound sources in two dimensions for ease of explanation, virtual sound sources can also be set in three dimensions using the same concept. Specifically, in the actual playback space, if sound sources are spatially arranged rather than aligned on a single plane, and the virtual space is set in three dimensions, virtual sound sources are set in three dimensions.

[0089] There are sound source SS and receiving point RP in the playback space. Figure 4 (A) Figure 4 The sound source SS shown in (B) has a different meaning from the sound source OBJ described above and refers to a sound source that generates normal sounds. Furthermore, a virtual wall IWL is set in the playback space to realize the sound field of the virtual space. The virtual wall IWL is derived from the geometric shape of the virtual space.

[0090] The sound source SS and the sound receiving point RP exist in a space surrounded by virtual walls IWL. The virtual walls IWL include virtual walls IWL1, IWL2, IWL3, and IWL4. The virtual walls IWL1 and IWL4 are arranged in a first direction ( Figure 4 (A) Figure 4 The virtual wall IWL1 is placed closer to the sound source SS than the sound receiving point RP, and the virtual wall IWL4 is placed closer to the sound receiving point RP than the sound source SS. The virtual walls IWL2 and IWL3 are placed in the second direction ( Figure 4 (A) Figure 4 (B) The virtual wall IWL2 is arranged closer to the sound source SS than the sound receiving point RP, and the virtual wall IWL3 is arranged closer to the sound source SS than the sound receiving point RP.

[0091] If the virtual walls IWL1, IWL2, IWL3, and IWL4 are walls that reflect sound in reality, then Figure 4 As shown in (B), the sound emitted from the sound source SS is reflected by the virtual wall IWL1, the virtual wall IWL2 and the virtual wall IWL3 and reaches the sound receiving point RP. Figure 4 In (B), although reflection from the virtual wall IWL4 is not described, reflection also occurs on the virtual wall IWL4 in the same manner as on the virtual walls IWL1 , IWL2 , and IWL3 .

[0092] However, the virtual walls IWL1, IWL2, IWL3 and IWL4 do not actually exist in the playback space. Figure 4 As shown in (A), the sound signal processing device 10 regards the reflection of the sound on the wall surface as a mirror reflection and sets a virtual sound source IS1, a virtual sound source IS2, and a virtual sound source IS3.

[0093] Specifically, the sound signal processing device 10 sets the virtual sound source IS1 at a position line-symmetrical with respect to the sound source SS, with the virtual wall IWL1 as the reference line. The sound signal processing device 10 sets the virtual sound source IS2 at a position line-symmetrical with respect to the sound source SS, with the virtual wall IWL2 as the reference line. The virtual sound source IS3 is set at a position line-symmetrical with respect to the sound source SS, with the virtual wall IWL3 as the reference line. Furthermore, by adjusting the sound power of each virtual sound source IS, it is possible to simulate the energy loss caused by reflection at the virtual wall IWL.

[0094] By making the above settings, the sound generated by the virtual sound source IS1 is the same as the sound generated by the sound source SS and reflected by the virtual wall IW1. The sound generated by the virtual sound source IS2 is the same as the sound generated by the sound source SS and reflected by the virtual wall IW2. The sound generated by the virtual sound source IS3 is the same as the sound generated by the sound source SS and reflected by the virtual wall IW3. Figure 4 (A) Figure 4 In (B), a virtual sound source for the virtual wall IWL4 is not described, but a virtual sound source can be set for the virtual wall IWL4 in the same manner as for the virtual walls IWL1 , IWL2 , and IWL3 .

[0095] By setting the virtual sound source in the above-described manner, the sound signal processing device 10 can simulate the initial reflected sound of the virtual space in the reproduction space without real walls in the virtual space.

[0096] [Structure and Processing of Grouping Unit 40]

[0097] Figure 54 is a functional block diagram showing an example of the configuration of the grouping unit 40 . Figure 6 This is a flowchart showing a method for grouping sound sources.

[0098] like Figure 5 As shown, the grouping unit 40 includes a sound source position detection unit 41 , an area determination unit 42 , and a matrix mixer 400 .

[0099] The sound source position detection unit 41 detects the position coordinates of a plurality of sound sources OBJ1-OBJ96 in the reproduction space ( Figure 6 : S111). For example, the sound source position detection unit 41 detects the position coordinates of the sound sources OBJ1-OBJ96 based on user input. Alternatively, the sound source position detection unit 41 includes a position detection sensor for detecting the sound sources OBJ1-OBJ96, and detects the position coordinates of the sound sources OBJ1-OBJ96 based on the positions detected by the position detection sensor.

[0100] The sound source position detection unit 41 outputs the position coordinates of the sound sources OBJ1 -OBJ96 to the area determination unit 42 .

[0101] The area determination unit 42 uses the area information of the plurality of areas Area1 to Area8 from the area setting unit 30 and the position coordinates of the sound sources OBJ1 to OBJ96 from the sound source position detection unit 41 to group the sound sources OBJ1 to OBJ96 into the plurality of areas Area1 to Area8 ( Figure 6 : S112). More specifically, the area determination unit 42 performs grouping as follows.

[0102] Figure 7 is a diagram showing the concept of grouping multiple sound sources into multiple areas. Figure 7 In the figure, the upper side is the front side of the hall serving as the playback space, and the lower side is the back side of the hall.

[0103] The area setting unit 30 sets a reference point Pso for area division for the playback space. Figure 7 As shown, the area setting unit 30 sets the center of the hall that implements the playback space as the reference point Pso. Alternatively, the area setting unit 30 may use a point (position) set by the user as the reference point. For example, the area setting unit 30 may use a sound pickup point set by the user as the reference point.

[0104] The area setting unit 30 sets eight areas Area1 to Area8 so as to divide the entire circumference on the plane into eight areas with the reference point Pso for area division as the center. Figure 7In the case of [ ], the area setting unit 30 sets the plurality of areas Area 1, Area 2, and Area 3 in the lobby (playback space) in front of the reference point Pso. Furthermore, the area setting unit 30 sets Area 4 to the left of the reference point Pso toward the front of the lobby, and Area 5 to the right of the reference point Pso toward the front of the lobby. Furthermore, the area setting unit 30 sets the plurality of areas Area 6, Area 7, and Area 8 in the lobby (playback space) in the rear of the reference point Pso.

[0105] The area settings are just an example; other settings are possible as long as the entire playback space can be covered by the multiple areas set. Furthermore, while this description shows the setting of planar areas, the same settings can be applied to spatial areas. For example, the vertical extent of Area 1 is also included in Area 1.

[0106] The area setting unit 30 sets representative points RP1 to RP8 for each of the plurality of areas Area1 to Area8. For example, the area setting unit 30 sets the plurality of representative points RP1 to RP8 at the center positions of the plurality of areas Area1 to Area8. Alternatively, Figure 7 In the case of such a radially expanding region, for example, the region setting unit 30 sets a representative point at a predetermined distance from the reference point Pso on a straight line passing through the center of the radially expanding angle. The method of setting these representative points is merely an example; for example, one representative point may be set for each region, and other methods are also acceptable as long as they reliably perform sound source grouping processing.

[0107] The area setting unit 30 outputs area information for the plurality of areas Area 1 to Area 8 to the area determination unit 42 of the grouping unit 40 and the matrix mixer 400. The area information for the plurality of areas Area 1 to Area 8 includes the position coordinates of representative points RP1 to RP8 of the areas Area 1 to Area 8 and the coordinate information of the boundary lines that define the shapes of the areas Area 1 to Area 8.

[0108] (Method of grouping sound sources into areas using representative points)

[0109] Figure 8 (A) is a flowchart showing a method of grouping sound sources using representative points.

[0110] The area determination unit 42 obtains the position coordinates of representative points RP1-RP8 based on the area information of the multiple areas Area1-Area8 (S1121). The area determination unit 42 calculates the distance between the position coordinates of the sound source to be grouped and the position coordinates of the representative points RP1-RP8 (S1122). The area determination unit 42 groups the sound source into the area containing the representative point with the shortest distance (S1123).

[0111] For example, in Figure 7 In the example of sound source OBJ1, the area determination unit 42 detects the position coordinates of sound source OBJ1 and obtains the position coordinates of multiple representative points RP1-RP8. Based on the position coordinates of sound source OBJ1 and the position coordinates of multiple representative points RP1-RP8, the area determination unit 42 calculates the distance between sound source OBJ1 and each of the multiple representative points RP1-RP8. The area determination unit 42 detects when the distance between sound source OBJ1 and representative point RP1 is shorter than the distance between sound source OBJ1 and the other representative points RP2-RP8. In other words, the area determination unit 42 detects when the distance between sound source OBJ1 and representative point RP1 is the shortest distance. The area determination unit 42 groups the sound source OBJ1 into area Area1 associated with the representative point RP1.

[0112] (Method of grouping sound sources into regions using region boundaries)

[0113] Figure 8 (B) is a flowchart showing a method of grouping sound sources using area boundaries.

[0114] Based on the area information for the plurality of areas Area 1 to Area 8, the area determination unit 42 obtains coordinate information (boundary coordinates) indicating the boundaries of each area Area 1 to Area 8 (S1124). The area determination unit 42 then determines whether the position coordinates of the sound source to be grouped are inside each area Area 1 to Area 8 (S1125). For example, the area determination unit 42 uses a crossing number algorithm to determine whether the sound source is inside or outside the area. If the sound source is within the area (S1125: YES), the area determination unit 42 groups the sound source into that area (S1126).

[0115] For example, in Figure 7In the example of sound source OBJ1, the area determination unit 42 detects the position coordinates of sound source OBJ1 and obtains coordinate information (boundary coordinates) representing the boundaries of multiple areas Area1-Area8. Based on the position coordinates of sound source OBJ1 and the boundary coordinates of multiple areas Area1-Area8, the area determination unit 42 determines whether sound source OBJ1 is inside or outside the multiple areas Area1-Area8. The area determination unit 42 detects whether sound source OBJ1 is inside area Area1. The area determination unit 42 groups sound source OBJ1 into area Area1.

[0116] The area determination unit 42 groups the inputted plurality of sound sources OBJ1 to OBJ96 into a plurality of areas Area1 to Area8. For example, if Figure 7 For example, the area determination unit 42 groups the sound sources OBJ1 and OBJ4 into the area Area1, groups the sound source OBJ2 into the area Area2, and groups the sound source OBJ3 into the area Area5.

[0117] The region determination unit 42 outputs the grouping information to the matrix mixer 400. The grouping information is information indicating which sound source is grouped into which region.

[0118] Based on the grouping information, the matrix mixer 400 uses the sound signals S1-S96 from multiple sound sources OBJ1-OBJ96 to generate area-specific sound signals SA1-SA8 for multiple areas Area1-Area8. For example, if multiple sound sources are grouped in an area, the matrix mixer 400 mixes the sound signals from these multiple sound sources to generate area-specific sound signals for that area. The matrix mixer 400 outputs the area-specific sound signals for each area to the initial reflected sound control signal generator 50. Furthermore, even if only one sound source is grouped in an area, the matrix mixer 400 outputs the sound signal from that source as the area-specific sound signal for that area to the initial reflected sound control signal generator 50.

[0119] in the case of Figure 7 For example, area Area 1 is grouped with sound sources OBJ1 and OBJ4. Matrix mixer 400 mixes sound signal S1 from sound source OBJ1 and sound signal S4 from sound source OBJ4 to generate and output area-specific sound signal SA1 for area Area 1. Furthermore, area Area 2 is grouped with sound source OBJ2. Matrix mixer 400 outputs sound signal S2 from sound source OBJ2 as area-specific sound signal SA2 for area Area 2. Furthermore, area Area 5 is grouped with sound source OBJ3. Matrix mixer 400 outputs sound signal S3 from sound source OBJ3 as area-specific sound signal SA5 for area Area 5.

[0120] By implementing the above-described structure and processing, the sound signal processing device 10 can group multiple sound sources into multiple areas that divide the sound space and generate initial reflected sound control signals. As described above, the sound signal processing device 10 can reproduce initial reflected sounds corresponding to the positions of the sound sources, achieving clear sound image localization and rich spatial expansion.

[0121] In addition, in the above description, the case where the sound source moves is not shown in detail, but when the sound source moves, the grouping unit 40 performs Figure 9 The processing shown. Figure 9 This is a flowchart showing an example of a method of grouping based on the movement of sound sources.

[0122] The sound source position detection unit 41 detects the movement of the sound source (S104). The sound source position detection unit 41 detects the movement of the sound source, for example, through user input. Alternatively, the sound source position detection unit 41 detects the movement of the sound source by continuously detecting the sound source position using a position detection sensor. The region determination unit 42 then regroups the moved sound source (S105). The sound source position detection unit 41 detects the position coordinates of the moved sound source and outputs them to the region determination unit 42.

[0123] The area determination unit 42 uses the position coordinates of the moved sound source to group the sound source into the plurality of areas Area 1 to Area 8 as described above ( S105 ).

[0124] By performing the above-described processing, even if the sound source moves, the sound signal processing device 10 can generate an initial reflected sound control signal corresponding to the position of the moved sound source. As described above, the sound signal processing device 10 can reproduce the changes in the initial reflected sound corresponding to the movement of the sound source, achieving clear sound image localization and rich spatial expansion corresponding to the movement of the sound source.

[0125] Furthermore, when a sound source moves as described above, the sound signal processing device 10 can perform crossfade processing on the initial reflected sound control signal before and after the movement. For example, when a sound source moves, the sound signal processing device 10 gradually decreases the sound signal component of the sound source in the sound signal segmented by region that includes the sound source before the movement. On the other hand, the sound signal processing device 10 gradually increases the sound signal component of the sound source in the sound signal segmented by region that includes the sound source after the movement.

[0126] By performing the above-described processing, the sound signal processing device 10 can suppress discontinuous changes in the initial reflected sound when the sound source moves. As described above, the sound signal processing device 10 can make the initial reflected sound change more smoothly in accordance with the movement of the sound source when the sound source moves.

[0127] Matrix mixer 400 also outputs audio signals S1-S96 from multiple sound sources OBJ1-OBJ96 to mixer 60. As described above, mixer 60 adds audio signals S1-S96 to generate reverberant sound generation signal Sr, which is then output to reverberant sound control signal generator 70. Reverberant sound control signal generator 70 uses reverberant sound generation signal Sr to generate reverberant sound control signals REV1-REV64.

[0128] Through the above-described processing, the reverberant sound is not affected by the position or movement of the sound source. Therefore, even if the sound source moves, the sound signal processing device 10 can maintain the reverberant sound in the playback space constant and more clearly reproduce the movement of the sound source through changes in the initial reflected sound.

[0129] [Generation of Initial Reflected Sound Control Signal]

[0130] Figure 10 3 is a functional block diagram showing an example of the configuration of the initial reflected sound control signal generating unit 50 . Figure 11 This is a diagram showing an example of a GUI.

[0131] like Figure 10 As shown, the initial reflected sound control signal generator 50 includes an FIR filter circuit 51, an LDtap circuit 52, an addition processing unit 53, a timbre setting unit 501, a virtual sound source setting unit 502, and an operation unit 500. The LDtap circuit 52 amplifies and delays the input signal and outputs it. The FIR filter circuit 51 includes multiple FIR filters 511-518. The LDtap circuit 52 includes multiple LDtap circuits 521-528, an output speaker setting unit 5201, and a coefficient setting unit 5202. The order of the FIR filter circuit 51 and the LDtap circuit 52 may be reversed.

[0132] [Adjusting the tone of the initial reflected sound]

[0133] The operation unit 500 receives information from the user specifying the timbre to be added to the initial reflected sound and outputs it to the timbre setting unit 501. The timbre specifying information includes, for example, information specifying whether to prioritize low-pitched sounds or high-pitched sounds, the volume of the initial reflected sound, and the attenuation characteristics of the initial reflected sound (information indicating filter characteristics).

[0134] As a specific example, the operating unit 500 can Figure 11The GUI 100 (Graphical User Interface) shown in the figure receives the operation.

[0135] The GUI 100 includes a setting display window 111 , a plurality of operating elements 112 , a knob 1131 , and an adjustment value display window 1132 .

[0136] The setting display window 111 displays the shape of the virtual wall IWL of the virtual space, which is set using the plurality of operating elements 112 and the knob 1131. In this case, the setting display window 111 can display the separately set positions of the sound source SS, the speaker SP, the sound pickup point RP, and the coordinate axes of the playback space, along with the virtual wall IWL.

[0137] The plurality of operating elements 112 are associated with pre-set virtual space samples (various halls, rooms, etc.). Although not shown in the figure, the plurality of operating elements 112 are displayed with an index (e.g., hall name, etc.) that clearly indicates the virtual space sample associated with each operating element 112.

[0138] The knob 1131 is used to set the room size of the virtual space. The adjustment value display window 1132 displays the set value of the room size of the virtual space.

[0139] The GUI 100 receives various operations for adjusting the timbre. For example, the GUI 100 includes a plurality of operating elements 112, including an operating element for low-range sound, an operating element for high-range sound, an operating element for adjusting volume, and an operating element for adjusting attenuation characteristics, and receives operations through these operating elements.

[0140] When the user operates a desired operating element using the GUI 100 , the operating unit 500 detects the operation and sets the tone color designation information according to the operation.

[0141] For example, when the operation unit 500 receives selections of multiple operating elements 112, it obtains information specifying the timbre pre-set in the virtual space associated with the operating element 112. Furthermore, when the operation unit 500 receives operations using an operating element for a low-range sound, an operating element for a high-range sound, an operating element for adjusting the volume, an operating element for adjusting the attenuation characteristics, etc., it obtains information specifying the timbre set using these operating elements.

[0142] Although not shown in the figure, GUI 100 can also display timbre-specifying information using, for example, the filter coefficients of FIR filters 511-518 (described later), a schematic waveform, and the like. In this case, upon receiving adjustments to the timbre-specifying information, GUI 100 can also change the display in response to the adjustments. For example, GUI 100 can change the waveform display in response to the adjustments.

[0143] The timbre setting unit 501 sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 based on the timbre-specifying information. For example, if the timbre setting unit 501 receives information specifying an emphasis on the low-frequency range, it sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 to enhance the low-frequency range. Alternatively, if the timbre setting unit 501 receives information specifying an emphasis on the high-frequency range, it sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 to enhance the high-frequency range. The timbre setting unit 501 outputs the set filter coefficients to the FIR filter circuit 51. Furthermore, the timbre setting unit 501 can also set and adjust the sampling frequency and filter length as filter characteristics, in addition to the filter coefficients.

[0144] Furthermore, the timbre setting unit 501 sets the gain value of each tap of the FIR filters 511 to 518 of the FIR filter circuit 51 based on the timbre designation information. The timbre setting unit 501 outputs the set gain value to the FIR filter circuit 51 .

[0145] The plurality of FIR filters 511-518 are filters corresponding to the sound signals SA1-SA8 classified by region. The sound signals SA1-SA8 classified by region are input to the FIR filters 511-518. For example, Figure 10 As shown, the sound signal SA1 for each region is input to FIR filter 511, the sound signal SA2 for each region is input to FIR filter 512, the sound signal SA3 for each region is input to FIR filter 513, and the sound signal SA4 for each region is input to FIR filter 514. The sound signal SA5 for each region is input to FIR filter 515, the sound signal SA6 for each region is input to FIR filter 516, the sound signal SA7 for each region is input to FIR filter 517, and the sound signal SA8 for each region is input to FIR filter 518.

[0146] The multiple FIR filters 511-518 have the same number of taps. For example, the multiple FIR filters 511-518 have 16,000 taps. This number of taps is just an example and can be set based on the resource conditions of the sound signal processing device 10 and the desired timbre accuracy of the initial reflected sound.

[0147] Multiple FIR filters 511-518 perform filtering (convolution operations) on the multiple region-specific sound signals SA1-SA8 using the filter coefficients and gain values ​​set by the timbre setting unit 501. As described above, the multiple FIR filters 511-518 generate filtered region-specific sound signals SA1f-SA8f. For example, FIR filter 511 performs filtering (convolution operations) on the region-specific sound signal SA1 using the filter coefficients and gain values ​​set by the timbre setting unit 501, generating filtered region-specific sound signal SA1f. Similarly, the multiple FIR filters 512-518 generate filtered region-specific sound signals SA2f-SA8f based on the region-specific sound signals SA2-SA8.

[0148] Multiple FIR filters 511-518 output the filtered, region-specific sound signals SA1f-SA8f to multiple LD taps 521-528. For example, FIR filter 511 outputs the filtered, region-specific sound signal SA1f to LD tap 521. Similarly, multiple FIR filters 512-518 output the filtered, region-specific sound signals SA2f-SA8f to multiple LD taps 522-528.

[0149] Furthermore, the timbre-specifying information is not limited to information regarding the importance of the sound range but also includes information for specifying the waveform of the initial reflected sound to have characteristics desired by the user. By using this timbre-specifying information, the sound signal processing device 10 can achieve a wider variety of initial reflected sounds with timbre characteristics that suit the user's preferences.

[0150] [Virtual sound source settings and LDtap settings]

[0151] The virtual sound source setting unit 502 sets a virtual sound source based on the position coordinates of the sound collection point in the reproduction space and the geometric shape of the virtual space.

[0152] Figure 12 This is a flowchart illustrating an example of a virtual sound source setting process. The virtual sound source setting unit 502 obtains the position coordinates of a sound collection point in the broadcast space ( S131 ). For example, the virtual sound source setting unit 502 obtains the position coordinates of the sound collection point in the broadcast space through user input, position detection by a position detection sensor, or the like.

[0153] The virtual sound source setting unit 502 obtains the geometric shape of the virtual space (S132). For example, the virtual sound source setting unit 502 obtains the geometric shape of the virtual space through user input. The geometric shape of the virtual space includes, for example, a coordinate set representing the shape of a wall arranged in the virtual space.

[0154] The virtual sound source setting unit 502 is connected to the GUI 100. When the user selects a desired operating element 112 from the plurality of operating elements 112, the GUI 100 reads and obtains the geometric shape of the virtual space associated with the operating element 112. Furthermore, when the user adjusts the room size (the size of the playback space) using the knob 1131, the GUI 100 obtains the adjusted room size value.

[0155] The virtual sound source setting unit 502 obtains the position coordinates of the geometric shape of the virtual space in which the room size is set based on the various settings obtained by the GUI 100 in the above manner. In addition, the virtual sound source setting unit 502 obtains the position coordinates of the sound source SS and the position coordinates of the sound receiving point RP (the center of the room (the center position of the playback space)). The virtual sound source setting unit 502 uses this obtained information to set the virtual sound source in the following manner. The virtual sound source setting unit 502 makes the coordinate system of the playback space consistent with the coordinate system of the virtual space. The virtual sound source setting unit 502 uses the position coordinates of the sound receiving point in the playback space and the geometric shape of the virtual space, and uses the above Figure 4 (A) Figure 4 (B) The position coordinates of the virtual sound source in the playback space are set (S133).

[0156] Figure 13 (A) Figure 13 (B) is a diagram showing an example of setting each virtual sound source when the geometric shapes are different. Figure 13 (A) is the virtual wall IWL of the quadrilateral, Figure 13 (B) is a hexagonal virtual wall IWLh.

[0157] As described above, if the geometric shape of the virtual space is different, even if the position coordinates of the sound source SSa and the sound receiving point RP do not change, the positional relationship between the sound source SSa and the sound receiving point RP and the virtual wall IWL, and the positional relationship between the sound source SSa and the sound receiving point RP and the virtual wall IWLh will also be different. Figure 13 In case (A), the positions of the virtual sound sources IS1a, IS2a, and IS3a are set and Figure 13 In (B), the positions of the virtual sound sources IS1ah, IS2ah, and IS3ah are different.

[0158] Figure 14 (A) Figure 14 (B) and Figure 14 (C) is a diagram showing an example of setting a virtual sound source. Figure 14 (A) Figure 14 (B) Figure 14 (C) is a diagram showing the plane change of the virtual sound source. Figure 14 (B) shows the Figure 14(A) The case where the position of the sound source SSa relative to the reference point (sound receiving point RP) is the same but the size of the virtual space is different. Figure 14 (C) shows the Figure 14 (A) The size of the virtual space remains the same, but the positional relationship between the reference point of the virtual space and the reference point (receiving point) of the playback space changes (the center of the playback space changes).

[0159] As according to Figure 14 (A) and Figure 14 As can be seen from the comparison results of (B), the size of the virtual space on the playback space ( Figure 14 (A) is recorded using the virtual wall IWL, Figure 14 (B) is recorded with a virtual wall IWLc) is different, so the distance and position relationship between the sound source SSa as the original source of the virtual sound source and the virtual wall are different. Figure 14 In case (A), the positions of the virtual sound sources IS1a, IS2a, and IS3a are set and Figure 14 In case (B), the positions of the virtual sound sources IS1c, IS2c, and IS3c are set differently.

[0160] In addition, according to Figure 14 (A) and Figure 14 As can be seen from the comparison results of (C), the positional relationship between the reference point of the virtual space and the sound receiving point RP changes, and thus the position of the virtual sound source in the playback space (the position of the virtual sound source relative to the sound receiving point RP and the speaker) moves. Figure 14 In case (A), the positions of the virtual sound sources IS1a, IS2a, and IS3a are set and Figure 14 In case (C), the positions of the virtual sound sources IS1as, IS2as, and IS3as are set differently.

[0161] Figure 15 (A) Figure 15 (B) Figure 15 (C) is a diagram showing an example of setting a virtual sound source. Figure 15 (A) Figure 15 (B) Figure 15 (C) is a diagram showing changes in the position of the virtual sound source in the height direction.

[0162] exist Figure 15 (A) and Figure 15 In (B), the height of the ceiling is different. Figure 15 The distance (height) from the virtual wall IWFL of the floor to the virtual wall IWCL of the ceiling shown in (A) and Figure 15The virtual wall WILL shown in (B) has a different distance (height) from the virtual wall IWFL of the floor to the virtual wall IWCLL of the ceiling.

[0163] As according to Figure 15 (A) and Figure 15 As can be seen from the comparison results of (B), the height of the ceiling is different, and thus the distance and positional relationship between the sound source as the original source of the virtual sound source and the virtual walls IWCL and IWCLL of the ceiling are different. Figure 15 In case (A), the position of the virtual sound source IS1Ca is set and Figure 15 In case (B), the position of the virtual sound source IS1CaL is set differently.

[0164] exist Figure 15 (A) and Figure 15 In (C), the shape of the ceiling is different. Figure 15 The shape of the ceiling virtual wall IWCL of the virtual wall IWL shown in (A) Figure 15 The virtual wall IWLx shown in (C) has a different shape from the virtual wall IWCLx of the ceiling.

[0165] As according to Figure 15 (A) and Figure 15 As can be seen from the comparison results of (C), the shape of the ceiling is different, and thus the positional relationship between the sound source as the original source of the virtual sound source and the virtual walls IWCL and IWCLx of the ceiling is different. Figure 15 In case (A), the position of the virtual sound source IS1Ca is set and Figure 15 In case (C), the position of the virtual sound source IS1Cax is set differently.

[0166] As described above, the virtual sound source setting unit 502 can optimally set the position of the virtual sound source in the playback space in accordance with the geometric shape of the virtual space and the positional relationship between the playback space and the virtual space. Consequently, the sound signal processing device 10 can clearly localize the sound image of the initial reflected sound in accordance with the position coordinates of the speakers in the playback space, the geometric shape of the virtual space, and the positional relationship between the playback space and the virtual space.

[0167] The virtual sound source setting unit 502 outputs the position coordinates of the virtual sound sources set for each of the plurality of areas Area 1 to Area 8 to the output speaker setting unit 5201 of the LDtap circuit 52 .

[0168] The output speaker setting unit 5201 sets the virtual sound source IS assigned to each speaker based on the position coordinates of the virtual sound source IS, the position coordinates of the sound collection point RP, and the position coordinates of the plurality of speakers SP1 to SP64. Figure 16 This is a flowchart showing the process of allocating virtual sound sources to speakers.

[0169] The output speaker setting unit 5201 obtains the position coordinates of the virtual sound source from the virtual sound source setting unit 502 (S141). The output speaker setting unit 5201 obtains the position coordinates of the sound collection point in the playback space, for example, through user input (S142). The output speaker setting unit 5201 obtains the position coordinates of the plurality of speakers SP1-SP64, for example, through user input (S143).

[0170] The output speaker setting unit 5201 sets the responsible area of ​​the virtual sound source of each speaker based on the positional relationship between the sound receiving point RP of the reproduction space and the plurality of speakers SP1 to SP64 ( S144 ).

[0171] More specifically, the output speaker setting unit 5201 sets the responsible area of ​​the virtual sound source of each speaker in the following manner. Figure 17 (A) Figure 17 (B) is a diagram showing the concept of allocating virtual sound sources to speakers. Figure 17 (A) shows the concept of allocation using the azimuth angle φ, Figure 17 (B) shows the concept of allocation using the pitch angle θ. In the following description, the speaker SP1 is used as an example, but the output speaker setting unit 5201 also sets the assigned areas for the other speakers SP2 to SP64 using the same method.

[0172] The output speaker setting unit 5201 uses the position coordinates of the sound collection point RP and the position coordinates of the speaker SP1 to calculate the straight line ( Figure 17 (A) is set as the dotted line. Figure 17 As shown in (A), the output speaker setting unit 5201 sets the output speaker relative to the straight line ( Figure 17 The dotted line in (A) sets the azimuth angle φ extending from the sound receiving point RP as the reference point on the plane toward the speaker SP1. The azimuth angle φ is the angle formed in the horizontal direction by the straight line passing through the sound receiving point RP and the speaker SP1. Figure 17 As shown in (B), the output speaker setting unit 5201 sets the output speaker relative to the above-mentioned straight line ( Figure 17 The pitch angle θ is set to extend in the vertical direction perpendicular to the plane (dashed line in (B)). The pitch angle θ is the angle formed in the vertical direction (orthogonal to the horizontal direction) with respect to the straight line passing through the sound pickup point RP and the loudspeaker SP1.

[0173] The output speaker setting unit 5201 sets a space closer to the speaker SP1 than the boundary determined by the azimuth angle φ and the elevation angle θ (the boundary plane determining the horizontal area, the boundary plane determining the vertical area) as the responsible area RGSP1 of the speaker SP1.

[0174] The output speaker setting unit 5201 obtains multiple virtual sound sources IS (in Figure 17 In the case of multiple virtual sound sources ISa-ISg) position coordinates.

[0175] The output speaker setting unit 5201 uses the position coordinates of the multiple virtual sound sources ISa-ISg and the coordinates indicating the responsible region RGSP1 to determine whether the multiple virtual sound sources ISa-ISg are within the responsible region RGSP1. This determination can be achieved using the same method as the grouping of sound sources into regions described above.

[0176] The output speaker setting unit 5201 performs this determination process, thereby Figure 14 A. Figure 14 B. Figure 14 In the case shown in C, it is determined that the plurality of virtual sound sources ISa, ISb, ISc, and ISd are within the responsible region RGSP1, and the plurality of virtual sound sources ISe, ISf, and ISg are outside the responsible region RGSP1.

[0177] The output speaker setting unit 5201 allocates the plurality of virtual sound sources ISa, ISb, ISc, and ISd determined to be within the responsible area RGSP1 to the speaker SP1 (S145).

[0178] The output speaker setting unit 5201 outputs allocation information for the multiple virtual sound sources to the multiple speakers SP1-SP64 to the coefficient setting unit 5202. At this time, the output speaker setting unit 5201 outputs the position coordinates of the sound collection point RP, the position coordinates of the multiple speakers SP1-SP64, and the position coordinates of the multiple virtual sound sources to the coefficient setting unit 5202 along with the allocation information.

[0179] The azimuth angle φ is, for example, 60°, and the elevation angle θ is, for example, 45°. These azimuth angle φ and elevation angle θ are just examples, and may be set or adjusted by, for example, an operation input from a user.

[0180] The coefficient setting unit 5202 sets the tap coefficients assigned to the LDtap 521-528 using the distances between the sound collection point RP and the speakers SP1-SP64 and the distances between the sound collection point RP and the virtual sound source IS. The tap coefficients assigned to the LDtap 521-528 are the gain and delay values ​​of the LDtap 521-528.

[0181] Figure 18 This is a flowchart showing the coefficient setting process of LDtap. Figure 19 (A) Figure 19 (B) is a diagram for explaining the concept of coefficient setting.

[0182] The coefficient setting unit 5202 calculates the distance (speaker distance) between the sound collection point RP and the plurality of speakers SP1 - SP64 using the position coordinates of the sound collection point RP and the position coordinates of the plurality of speakers SP1 - SP64 ( S151 ).

[0183] The coefficient setting unit 5202 calculates the distances (virtual sound source distances) between the sound receiving point RP and the plurality of virtual sound sources IS (S152).

[0184] The coefficient setting unit 5202 compares the speaker distances and the virtual sound source distances for the plurality of speakers SP1-SP64 and the plurality of virtual sound sources IS assigned to the speakers SP1-SP64 (S153). Figure 17 In the example of (A), the speaker distance and the virtual sound source distance are compared for the speaker SP1 and multiple virtual sound sources ISa, ISb, ISc, and ISd.

[0185] If the speaker distance is equal to or smaller than the virtual sound source distance ( S153 : YES), the coefficient setting unit 5202 sets the tap coefficients using the virtual sound source distance as it is ( S154 ).

[0186] For example, in Figure 19 In the case shown in (A), the virtual sound source Isa is farther away from the sound receiving point RP than the speaker SP1, and the virtual sound source distance Lia between the sound receiving point RP and the virtual sound source Isa is greater than the speaker distance Ls1 between the sound receiving point RP and the speaker SP1.

[0187] In this case, the coefficient setting unit 5202 sets the tap coefficients using the distance Da1 between the virtual sound source Isa and the loudspeaker SP1. Specifically, the coefficient setting unit 5202 sets the gain and delay for the virtual sound source Isa based on the distance Da1. The coefficient setting unit 5202 sets the gain smaller as the distance Da1 increases, and sets the delay larger as the distance Da1 increases.

[0188] If the speaker distance is greater than the virtual sound source distance (S153: NO), the coefficient setting unit 5202 determines whether to play the virtual sound source. In other words, the coefficient setting unit 5202 determines whether to play the virtual sound source closer to the sound pickup point than the speaker (S155).

[0189] If a virtual sound source closer to the sound pickup point than the speaker is being played (S155: YES), the coefficient setting unit 5202 moves the position of the virtual sound source (S156). More specifically, the coefficient setting unit 5202 moves the position of the virtual sound source closer to the sound pickup point than the speaker to a position farther away from the sound pickup point than the speaker. In this case, the coefficient setting unit 5202 uses the distance difference between the virtual sound source and the speaker to move the position of the virtual sound source. The coefficient setting unit 5202 uses the position coordinates of the moved virtual sound source to set the tap coefficients (S157).

[0190] For example, in Figure 19 In the case shown in (B), the virtual sound source ISd is closer to the sound receiving point RP than the speaker SP1, and the virtual sound source distance Lid between the sound receiving point RP and the virtual sound source ISd is smaller than the speaker distance Ls1 between the sound receiving point RP and the speaker SP1.

[0191] In this case, the coefficient setting unit 5202 uses the distance difference Dd between the virtual sound source distance Lid and the speaker distance Ls1 to move the virtual sound source ISd. More specifically, the coefficient setting unit 5202 moves the virtual sound source ISd to a position on a straight line passing through the sound collection point RP and the speaker SP1, with the speaker SP1 as a reference and on the side opposite the sound collection point RP. The coefficient setting unit 5202 then uses this distance difference Dd to set the tap coefficients. Specifically, the coefficient setting unit 5202 sets the gain and delay for the virtual sound source ISd based on the distance difference Dd. The greater the distance difference Dd, the smaller the gain, and the greater the delay. While conceptually moving the virtual sound source as described above, the coefficient setting unit 5202 can simply set the tap coefficients based on the distance between the speaker distance and the virtual sound source distance.

[0192] In other words, the coefficient setting unit 5202 only moves the virtual sound sources located between the sound pickup point and the loudspeaker. While it is preferable that virtual sound sources located further outward from the loudspeaker relative to the sound pickup point remain unchanged, this also includes situations where such outward virtual sound sources move within a specified range. For example, even if such outward virtual sound sources move, this is sufficient as long as the distance between the outward virtual sound source and the loudspeaker remains within a specified range. The specified range refers to a range where changes in the initial reflected sound control signal due to movement do not cause discomfort to the audience. If a virtual sound source closer to the sound pickup point than the loudspeaker is not being played (S155: NO), the coefficient setting unit 5202 does not set the tap coefficients for such virtual sound sources.

[0193] The coefficient setting unit 5202 sets the tap coefficients for each speaker SP1-SP64 in a plurality of LDtap nodes. More specifically, the coefficient setting unit 5202 sets the tap coefficients for each speaker SP1-SP64 in LDtap 521 based on the virtual sound source position set in area Area1. Similarly, the coefficient setting unit 5202 sets the tap coefficients for the virtual sound sources assigned to each speaker SP1-SP64 in LDtap 522-528 based on the virtual sound source positions set in areas Area2-Area8.

[0194] The multiple LDtap units 521-528 perform gain processing and delay processing on the filtered, region-specific sound signals SA1f-SA8f according to the set tap coefficients, and output the signals to the summing unit 53. More specifically, the tap coefficients are set according to the virtual sound source positions of the multiple regions and the speaker combinations, as described above. Therefore, the multiple LDtap units 521-528 set tap coefficients for each speaker based on the virtual sound source assigned to that speaker. The multiple LDtap units 521-528 perform gain processing and delay processing on the filtered, region-specific sound signals SA1f-SA8f for each speaker. The multiple LDtap units 521-528 output the gain- and delay-processed signals to each speaker.

[0195] For example, when virtual sound sources ISa, ISb, ISc, and ISd are assigned to speaker SP1, LDtap 521 applies gain processing and delay processing to the filtered, region-specific sound signal SA1f using tap coefficients (gain values ​​and delay amounts) based on the virtual sound sources ISa, ISb, ISc, and ISd. LDtap 521 then outputs this signal to the addition unit 53 as intended for speaker SP1. Multiple LDtap units 522-528 perform the aforementioned processing on the virtual sound sources to which the tap coefficients are assigned.

[0196] The adding unit 53 adds the LDtap-processed signals for the speakers SP1 to SP64 outputted from the LDtap units 521 to 528, respectively, for the speakers SP1 to SP64, and outputs the added signals to the adder 80 as initial reflected sound control signals ER1 to ER64 for the speakers SP1 to SP64.

[0197] By performing the above-described processing, the initial reflected sound control signal generating unit 50 can generate an initial reflected sound control signal having the following characteristics.

[0198] Figure 20 (A) Figure 20(B) is a waveform diagram showing an example of the relationship between the shape of the virtual space and the components of the initial reflected sound control signal achieved by LDtap. Figure 20 (A) shows the case where the virtual space shape is large, Figure 20 (B) shows the case where the virtual space shape is small. Figure 20 (A) Figure 20 (B) shows an example of the components of the initial reflected sound control signal when a plurality of virtual sound sources are set for one speaker.

[0199] If the positional relationship between the playback space and the virtual space does not change, and the positions of the sound pickup points and the speakers do not change, then if the virtual space is large, the distribution of the virtual sound source will spread over a wider range than if the virtual space is small. Figure 20 (A) Figure 20 As shown in (B), the virtual space shape is large, and the components set in LDtap521-528 tend to become smaller, and the distribution range on the time axis also becomes wider.

[0200] As described above, by performing the above-mentioned processing, the initial reflected sound control signal generating unit 50 can set the optimal tap coefficients according to the shape of the virtual space.

[0201] Furthermore, even if the positional relationship between the virtual space and the playback space changes, the speaker position changes, or the sound pickup point changes, the initial reflected sound control signal generating unit 50 can set the optimal tap coefficients in accordance with these changes, just like when the shape of the virtual space changes.

[0202] At this point, the multiple sound sources OBJ1-OBJ96 are optimally distributed to the multiple speakers SP1-SP64 through grouping based on the multiple areas Area1-Area8. Furthermore, the multiple virtual sound sources are optimally set for the multiple speakers SP1-SP64. Therefore, even if the relationship between the virtual space and the reproduction space, the position of the sound pickup point RP, the positions of the multiple speakers SP1-SP64, or the positions of the sound sources OBJ1-OBJ96 change, the sound signal processing device 10 can clearly localize the sound image based on the initial reflected sound in response to these changes.

[0203] Furthermore, in the above-described configuration, even if the virtual sound source IS is closer to the sound pickup point RP than the loudspeaker SP, the initial reflected sound control signal generator 50 can approximately reproduce the components of the initial reflected sound control signal derived from the virtual sound source IS. Therefore, for example, when the number of virtual sound sources set is small relative to the initial reflected sound control signal, the initial reflected sound control signal generator 50 can utilize a virtual sound source closer to the sound pickup point RP than the loudspeaker SP. In this case, the initial reflected sound control signal generator 50, as described above, uses the distance difference between the virtual sound source IS and the loudspeaker SP to reposition the virtual sound source outside the loudspeaker. As described above, the initial reflected sound control signal generator 50 can suppress the unpleasant sensation of initial reflected sound caused by shifting the position of the virtual sound source.

[0204] Furthermore, in the above configuration, the initial reflected sound control signal generator 50 may also set the virtual sound source IS at the position of the loudspeaker SP when the virtual sound source IS is located closer to the sound pickup point RP than the loudspeaker SP. As described above, the initial reflected sound control signal generator 50 can reduce the processing load of moving the virtual sound source IS.

[0205] Furthermore, in the above configuration, the initial reflected sound control signal generator 50 may not use the virtual sound source IS for generating the initial reflected sound control signal if the virtual sound source IS is located closer to the sound receiving point RP than the loudspeaker SP. As described above, the initial reflected sound control signal generator 50 does not need to perform the processing load of moving the virtual sound source IS, thereby reducing the processing load of generating the initial reflected sound control signal.

[0206] Furthermore, in the above-described configuration, the initial reflected sound control signal generator 50 sets the components of the initial reflected sound control signal derived from the virtual sound source and adjusts the timbre using the FIR filters 511-518. The FIR filters 511-518 have the aforementioned number of taps (e.g., 16,000 taps), which is greater than the number of taps in the LDtap 521-528. Furthermore, the time interval between the taps of the FIR filters 511-518 (which depends on the sampling frequency) is shorter than the time interval between the taps of the LDtap 521-528 (which depends on the arrangement of the virtual sound source). Therefore, the components of the initial reflected sound control signal generated by the FIR filters 511-518 are more densely arranged on the time axis than the components of the initial reflected sound control signal generated by the LDtap 521-528. In other words, the resolution (temporal resolution) of the FIR filters 511-518 on the time axis is higher than that of the LDtap 521-528, resulting in a greater number of components per unit time.

[0207] The initial reflected sound control signal generator 50 multiplies the processing of the FIR filters 511 to 518 by the LDtap 521 to 528. Therefore, the initial reflected sound control signal generator 50 can generate initial reflected sound control signals ER1 to ER64 with high resolution on the time axis and more diverse timbre. Figure 21 1 is a diagram schematically showing the waveform of the initial reflected sound control signal generated by the initial reflected sound control signal generating unit 50 .

[0208] like Figure 21 As shown, the initial reflected sound control signal generator 50 can generate an initial reflected sound control signal that allows the initial reflected sound components obtained from the virtual sound source to remain and can handle higher resolution and a wider variety of timbres. In other words, the sound signal processing device 10 can achieve initial reflected sound that ensures clear sound image localization using the initial reflected sound using the virtual sound source and a timbre that matches the user's preference.

[0209] Furthermore, the FIR filter has high resolution. Therefore, for example, in the case of a short pulse sound, such as a sound source, the initial reflection sound control signal obtained solely based on the LDtap component may become coarse, resulting in an unnatural timbre. However, through the above-described structure and processing, the sound signal processing device 10 can suppress such coarseness and unnatural timbre of the initial reflection sound.

[0210] Furthermore, in the above-described configuration, the initial reflected sound control signal generator 50 sets a region assigned to a virtual sound source IS for each loudspeaker SP, and does not assign virtual sound sources IS outside of this region to that loudspeaker SP. As described above, the initial reflected sound control signal generator 50 can suppress excessive generation of initial reflected sound components. Consequently, the sound signal processing device 10 can suppress excessive generation of initial reflected sound, achieving more natural initial reflected sound that matches the virtual space.

[0211] [Generation of reverberation control signal]

[0212] Figure 22 1 is a functional block diagram showing an example of the configuration of the reverberant sound control signal generating unit 70 . Figure 23 This is a flowchart showing an example of a process for generating a reverberant sound control signal.

[0213] like Figure 22 As shown, the reverberant sound control signal generator 70 includes a PEQ 71, an FIR filter circuit 72, a distributor 73, a reverberant sound region setting unit 701, a filter coefficient setting unit 702, a reverberant sound playback speaker setting unit 703, and an operation unit 700. The FIR filter circuit 72 includes a plurality of FIR filters 721-728.

[0214] The reverberant sound area setting unit 701 sets a plurality of reverberant sound areas Arr1-Arr8 for the playback space. More specifically, the reverberant sound area setting unit 701 divides the playback space into a plurality of reverberant sound areas Arr1-Arr8 over the entire circumference of the plane, for example, with the center point Psr of the playback space as a reference (see the following description). Figure 25 ).

[0215] The reverberant sound region setting unit 701 outputs coordinate information indicating the plurality of reverberant sound regions Arr1 to Arr8 to the filter coefficient setting unit 702 and the reverberant sound reproduction speaker setting unit 703 .

[0216] The filter coefficient setting unit 702 sets the filter coefficient for the reverberant sound through user operations, etc. The filter coefficient for the reverberant sound is set, for example, based on the measured results of the impulse responses of different spaces (virtual spaces) reproduced in the playback space. In addition, the filter coefficient for the reverberant sound can also be approximately set using the geometric shape of the virtual space, the material of the wall, etc. In this case, the filter coefficient setting unit 702 uses the coordinate information of each reverberant sound area Arr1-Arr8 to set the filter coefficient for each reverberant sound area Arr1-Arr8.

[0217] The filter coefficient setting unit 702 receives input of the volume and surface area of ​​the virtual space through user operations, etc. The filter coefficient setting unit 702 sets a fading function for the filter coefficient based on parameters such as the volume and surface area of ​​the virtual space.

[0218] More specifically, the filter coefficient setting unit 702 calculates the mean free path (ρ) using the volume V and surface area S of the virtual space. The formula for calculating the mean free path (ρ) is ρ = 4V / S. The mean free path (MPP) refers to the average distance sound travels in a closed space from reflection on a wall to the next reflection. By dividing the MPP by the speed of sound (c0), the average time required for sound to reflect off a wall and then to the next reflection can be calculated.

[0219] The filter coefficient setting unit 702 sets the connection timing tc ( Figure 23 : S231). Specifically, the filter coefficient setting unit 702 sets the connection timing tc using the mean free path ρ, the speed of sound c0, and the number of reflections n. The connection timing tc is calculated using the equation tc = ρ × n / c0.

[0220] As can be seen from this calculation formula, connection timing tc corresponds to the average time required for n reflections in virtual space. When reproducing n initial reflected sounds, this corresponds to the moment when the sound begins to transform into reverberant sound. In other words, connection timing tc corresponds to the time when the components of the initial reflected sound control signal generated by the initial reflected sound control signal generator 50 disappear.

[0221] By performing such processing, the filter coefficient setting unit 702 can optimally set the connection timing tc between the initial reflected sound and the reverberant sound according to the geometric shape of the virtual space.

[0222] The filter coefficient setting unit 702 sets the fade-in function according to the following equation using the connection timing tc: Figure 23 :S232).

[0223] [Formula 1]

[0224]

[0225] In this formula, t is the time elapsed from the occurrence of the direct sound, and K is set according to the following formula.

[0226] [Formula 2]

[0227]

[0228] In addition, in this formula, G REV The gain value of the reverberant sound at time t = 0 can be set by the user. For example, the reverberation time is usually the time required to decay to -60dB, so it can be set to G REV =-60dB, etc.

[0229] The filter coefficient setting unit 702 sets the reverberant sound filter coefficient ( Figure 23 : S233) and output to multiple FIR filters 721-728.

[0230] The reverberant sound generating signal Sr output from the mixer 60 is input to the PEQ 71. The PEQ 71 performs predetermined signal processing on the reverberant sound generating signal Sr and outputs the result to the plurality of FIR filters 721-728.

[0231] The PEQ 71 performs signal processing to adjust the level (signal magnitude), timbre, and other aspects of the reverberant sound generating signal Sr. For example, the PEQ 71 can refer to the volume of the initial reflected sound control signal and adjust the level (signal magnitude) of the reverberant sound generating signal Sr so that the volume of the initial reflected sound and the reverberant sound are approximately the same at the connection timing tc. Furthermore, the PEQ 71 can adjust the timbre and other aspects through settings by the user or the like.

[0232] Multiple FIR filters 721-728 use the reverberant sound filter coefficients to perform filtering processing on the reverberant sound generating signal Sr, thereby generating reverberant sound control signals REVr1-REVr8 differentiated by region. For example, FIR filter 721 uses the reverberant sound filter coefficients for region Arr1 set as reverberant sound to perform convolution operation on reverberant sound generating signal Sr, thereby generating reverberant sound control signal REVr1 differentiated by region for region Arr1. Similarly, FIR filters 722-728 use the reverberant sound filter coefficients for regions Arr2-Arr8 set as reverberant sound to perform convolution operation on reverberant sound generating signal Sr, thereby generating reverberant sound control signals REVr2-REVr8 differentiated by region for regions Arr2-Arr8 ( Figure 23 : S234 ) The plurality of FIR filters 721 - 728 outputs the reverberant sound control signals REVr1 - REVr8 , which are differentiated by region, to the distributor 73 .

[0233] By setting the above-mentioned fade-in function, the reverberation control signal becomes as follows: Figure 24 The waveform shown. Figure 24 Graph showing waveform examples of direct sound, initial reflected sound control signal, and reverberant sound control signal. Figure 24 For convenience, the reverberant sound control signal is shown in the figure by the envelope of each time component. Figure 24 The vertical axis represents dB.

[0234] like Figure 24 As shown in (A), the reverberant sound control signal gradually increases in signal level following a fade-in function from the direct sound output timing to the connection timing tc. More specifically, the reverberant sound control signal's signal level is -60 dBFs at the direct sound output timing, gradually increases until the connection timing tc, and reaches 0 dBFs at the connection timing tc. This level is set based on the signal level at the connection timing tc of the initial reflected sound control signal.

[0235] exist Figure 24In this example, the above-described fade-in function is used to exponentially increase the signal level as the connection timing tc approaches. In other words, the above-described fade-in function has characteristics opposite to the attenuation curve of the reverberant sound control signal without fade-in processing. Furthermore, the characteristics of the level change of the reverberant sound control signal obtained by fade-in processing are not limited to these and can be set to any desired characteristics by appropriately configuring the fade-in function.

[0236] By performing the above-described processing, the reverberant sound control signal generating unit 70 can generate a reverberant sound control signal that accurately reproduces the reverberant sound in the virtual space using the FIR filters 721-728. Furthermore, the reverberant sound control signal gradually increases in level during the period in which the initial reflected sound control signal is present, reaches a peak corresponding to the signal level of the initial reflected sound control signal at connection timing tc, and then decays.

[0237] As described above, the sound signal processing device 10 can smooth the connection between the initial reflected sound control signal generated by multiple LD taps, which reproduces the virtual sound source distribution at multiple sound source positions in the virtual space, and the reverberant sound control signal, using the reverberant sound obtained based on the reverberant sound control signal. Consequently, the sound output from the sound signal processing device 10 and heard by the user is a sound that suppresses the discomfort felt during the connection from the initial reflected sound to the reverberant sound.

[0238] The reverberant sound reproduction speaker setting unit 703 groups the plurality of speakers SP1 to SP64 into reverberant sound areas Arr1 to Arr8 .

[0239] More specifically, the reverberant sound playback speaker setting unit 703 divides the playback space into a plurality of reverberant sound areas Arr1-Arr8, for example, using the center point Psr of the playback space as a reference, along the entire circumference of the plane. The reverberant sound playback speaker setting unit 703 uses the position coordinates of the plurality of speakers SP1-SP64 and the coordinate information representing the plurality of reverberant sound areas Arr1-Arr8 to group the plurality of speakers SP1-SP64 for the plurality of reverberant sound areas Arr1-Arr8. This grouping can be achieved using the same method as the method for grouping the sound source OBJ described above.

[0240] Figure 25 This is a diagram showing an example of area settings for reverberant sound. Figure 25 In order to simplify the description and make it easier to understand, multiple speakers SP1-SP14 are shown. For example, the reverberant sound playback speaker setting unit 703 is as follows Figure 25As shown, the presence of speakers SP6 and SP7 in the reverberant sound area Arr1 is detected, and speakers SP6 and SP7 are grouped into the reverberant sound area Arr1. Similarly, the reverberant sound playback speaker setting unit 703 groups the other speakers SP1-SP5 and SP8-SP14 into multiple reverberant sound areas Arr2-Arr8.

[0241] The reverberant sound reproduction speaker setting unit 703 outputs grouping information of the plurality of speakers SP1 to SP64 for the plurality of reverberant sound areas Arr2 to Arr8 to the distributor 73 .

[0242] Distributor 73 distributes the region-specific reverberant sound control signals REVr1-REVr8 to the plurality of speakers SP1-SP64 using the grouping information from reverberant sound playback speaker setting unit 703. Based on the distribution, distributor 73 outputs the region-specific reverberant sound control signals REVr1-REVr8 as reverberant sound control signals REV1-REV48 for each of the plurality of speakers SP1-SP64.

[0243] For example, based on the grouping information, distributor 73 detects that speaker SP6 and speaker SP7 are grouped in area Arr1. Distributor 73 distributes the region-specific reverberant sound control signal REVr1 for area Arr1 to speaker SP6 and speaker SP7. Distributor 73 outputs the region-specific reverberant sound control signal REVr1 to speaker SP6 as reverberant sound control signal REV6 for speaker SP6. Furthermore, distributor 73 outputs the region-specific reverberant sound control signal REVr1 to speaker SP7 as reverberant sound control signal REV7 for speaker SP7.

[0244] By performing the distribution processing of the reverberant sound control signals REVr1-REVr8 differentiated by region for each region by the distributor 73 as described above, the reverberant sound control signal generating unit 70 can output the optimal reverberant sound control signal to each of the multiple speakers SP1-SP64 according to the configuration of the multiple speakers SP1-SP64.

[0245] [Output Adjustment]

[0246] Figure 26 1 is a functional block diagram showing an example of the configuration of the output adjustment unit 90 . Figure 27 This is a flowchart showing an example of output adjustment processing.

[0247] like Figure 26As shown, the output adjustment unit 90 includes a gain control unit 91, a hysteresis control unit 92, a gain and hysteresis setting unit 901, an operation unit 900, and a display unit 909. The gain control unit 91 includes a plurality of gain control units 9101-9168 corresponding to the plurality of speakers SP1-SP64. The hysteresis control unit 92 includes a plurality of hysteresis control units 9201-9264 corresponding to the plurality of speakers SP1-SP64.

[0248] The operation unit 900 receives the setting of the acoustic parameters of the playback space through the operation input from the user ( Figure 27 : S321). The acoustic parameters of the playback space are parameters used to reproduce a desired sound field in the playback space.

[0249] In this case, the acoustic parameters of the broadcast space are not the gain values ​​or delay amounts of the multiple speakers SP1-SP64, but the weight values ​​indicating the weighting of the sound in the broadcast space in a specific direction and the shape values ​​indicating the expansion of the sound in the broadcast space in a specific direction.

[0250] The weight value is composed of a gain value and a delay amount, including weight values ​​for the front and back of the playback space, the left and right weight values, and the top and bottom weight values. The shape value is composed of a gain value and a delay amount, including a horizontal shape value.

[0251] The display unit 909 includes a GUI. Figure 28 This is a diagram showing an example of a GUI for output adjustment.

[0252] like Figure 28 As shown, the GUI 100A includes a setting display window 111 , an output state display window 115 , and a plurality of operating elements 116 . The plurality of operating elements 116 include a knob 1161 and an adjustment value display window 1162 .

[0253] Multiple operating elements 116 are used to set weighted volume for setting weight values, shape volume for setting shape values, and the like. The weighted volume operating elements 116 include operating elements for setting left and right weights, front and back weights, and top and bottom weights, and also include operating elements for setting gain values ​​and delay amounts. The shape volume operating elements 116 include operating elements for setting expansion, gain values, and delay amounts.

[0254] The output state display window 115 graphically displays the sound expansion and localization achieved by the weight values ​​and shape values ​​set by the plurality of operating elements 116. This allows the user to easily recognize the sound expansion and localization set by the plurality of operating elements 116 as an image.

[0255] The user sets the desired audio parameters (weight values ​​and delay amounts) using GUI 100A on display unit 909. The operation unit 900 receives the settings made using GUI 100A and outputs the settings (weight values ​​and delay amounts for the audio parameters) to gain and lag setting unit 901.

[0256] The gain and lag setting unit 901 sets the gain values ​​and delay amounts for the plurality of speakers SP1 to SP64 based on the weight values ​​and delay amounts of the acoustic parameters. More specifically, the gain and lag setting unit 901 performs the following processing.

[0257] The gain and hysteresis setting unit 901 obtains the position coordinates of the plurality of speakers SP1-SP64 arranged in the playback space (S322). The position coordinates are expressed, for example, in a coordinate system in which the x-axis is set to the left-right direction of the playback space, the y-axis is set to the front-back direction of the playback space, and the z-axis is set to the top-bottom direction.

[0258] The gain and hysteresis setting unit 901 extracts the maximum value and the minimum value of the position coordinates of the plurality of speakers SP1 to SP64 in each axial direction ( S323 ).

[0259] The gain and hysteresis setting unit 901 stores coefficient setting formulas, including, for example, a weight coefficient setting formula for setting weighting in a predetermined direction in the reproduction space and a shape coefficient setting formula for setting weighting in a predetermined direction in the reproduction space.

[0260] The coefficient setting formula for weighting includes a setting formula for the gain value for weighting and a setting formula for the delay amount for weighting. The coefficient setting formula for shape includes a setting formula for the gain value for shape and a setting formula for the delay amount for shape.

[0261] The coefficient setting formula for weights includes a coefficient setting formula for the front-to-back direction for setting the weighting of the front-to-back direction of the playback space, a coefficient setting formula for the left-to-right direction for setting the weighting of the left-to-right direction of the playback space, and a coefficient setting formula for the up-down direction for setting the weighting of the up-down direction of the playback space.

[0262] The coefficient setting formula for the shape includes the coefficient setting formula for the left and right directions of the playback space.

[0263] The coefficient setting formula for the gain value used for the weight is, for example, a linear function obtained by combining the gain value of the set weight value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker for setting the gain value (the speaker of the setting object). It is a formula that determines the gain value in proportion to the difference between the position coordinates of the speaker of the setting object and the minimum value of the position coordinates.

[0264] The coefficient setting formula for the delay amount used for weighting is, for example, a linear function obtained by combining the delay amount of the set weight value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker for setting the delay amount (the speaker of the setting object). It is a formula that determines the delay amount in proportion to the difference between the position coordinates of the speaker of the setting object and the minimum value of the position coordinates.

[0265] The coefficient setting formula for the gain value used for the shape is, for example, a linear function obtained by combining the gain value of the set shape value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker for setting the gain value (the speaker of the setting object). It is a formula that determines the gain value in proportion to the difference between the position coordinates of the speaker of the setting object and the minimum value of the position coordinates.

[0266] The coefficient setting formula for the delay amount used for the shape is, for example, a linear function obtained by combining the delay amount of the set shape value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker for setting the delay amount (the speaker of the setting object). It is a formula that determines the delay amount in proportion to the difference between the position coordinates of the speaker of the setting object and the minimum value of the position coordinates.

[0267] The gain and lag setting unit 901 calculates the gain and delay for each speaker to be set using the set gain and delay (acoustic parameters), the maximum and minimum values ​​of the extracted position coordinates, and the coefficient setting formula ( S324 ).

[0268] By using the above-described processing, the gain and lag setting unit 901 does not need to manually set the gain values ​​and delay amounts of the plurality of speakers SP1 to SP64 arranged in the reproduction space individually, but can automatically calculate and set them using a coefficient setting formula.

[0269] The gain and hysteresis setting unit 901 outputs gain values ​​set for the speakers SP1 to SP64 to the gain control units 9101 to 9164 and delay amounts set for the speakers SP1 to SP64 to the hysteresis control units 9201 to 9264.

[0270] Speaker signals Sat1 to Sat64 corresponding to the plurality of speakers SP1 to SP64 are input from the adder 80 to the plurality of gain control units 9101 to 9164 , respectively.

[0271] Gain control units 9101-9164 control the signal levels of speaker signals Sat1-Sat64 using their respective set gain values, and output these signals to hysteresis control units 9201-9264. For example, gain control unit 9101 controls the signal level of speaker signal Sat1 using the gain value set for gain control unit 9101, and outputs these signals to hysteresis control unit 9201. Similarly, gain control units 9102-9164 control the signal levels of speaker signals Sat2-Sat64 using their respective set gain values, and output these signals to hysteresis control units 9202-9264, respectively.

[0272] The multiple hysteresis control units 9201-9264 control the signal levels of signals input from the multiple gain control units 9101-9164 using the delay amounts set therein, and output the signals to the multiple speakers SP1-SP64. For example, the hysteresis control unit 9201 controls the signal levels of signals input from the gain control unit 9101 using the delay amounts set therein, and outputs the signals to the multiple speakers SP1-SP64. Similarly, the hysteresis control units 9202-9264 control the signal levels of signals input from the gain control units 9102-9164 using the delay amounts set therein, and output the signals to the multiple speakers SP2-SP64, respectively.

[0273] With the above-described configuration, the sound signal processing device 10 can easily achieve a desired sound field corresponding to the set acoustic parameters using the initial reflected sound control signal and the reverberant sound control signal, without requiring the user to master complex settings for each of the multiple speakers. For example, as described above, the sound signal processing device 10 can easily achieve a sound field that produces a Haas effect at a predetermined location within the playback space.

[0274] (Example of realizing the sound field based on output control)

[0275] Figure 29 (A) Figure 29 (B) is a diagram showing a setting example in which positioning and expansion are performed on the rear side of the reproduction space. Figure 29 (A) is a diagram showing an example of setting a gain value and a delay amount. Figure 29 (B) is based on Figure 29 (A) is a diagram showing an overview of the weighting of the sound achieved by the setting. Figure 29 (A) Figure 29 In (B), for the sake of simplicity of explanation and ease of understanding, a case is shown in which 14 speakers SP1 to SP14 are arranged.

[0276] exist Figure 29 (A) Figure 29 In the method shown in (B), the rear-end gain and delay are set as acoustic parameters, for example. Gain and lag setting unit 901 sets the front-end gain and delay to values ​​with opposite signs to those of the rear-end gain and delay. Gain and lag setting unit 901 calculates the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14.

[0277] The gain and lag setting unit 901 calculates the gain values ​​of the 14 speakers SP1-SP14 using the gain values ​​of the rear and front ends, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and the coefficient setting formula for the front and back directions for setting the weighting of the front and back directions of the playback space (for gain value setting).

[0278] In addition, the gain and lag setting unit 901 calculates the delay amount of the 14 speakers SP1-SP14 using the delay amount at the rear end and the front end, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and the coefficient setting formula for the front and rear directions for setting the weighting of the front and rear directions of the playback space (for delay amount setting).

[0279] Through this processing, the sound signal processing device 10 Figure 29 As shown in (A), it is possible to automatically and easily set the sound parameters such that the gain value and delay amount are larger for the speakers closer to the rear of the playback space, and smaller for the speakers closer to the front. As a result, the sound signal processing device 10 can easily realize a sound field that is extensible and localized in the rear of the playback space (see Figure 29 (B)).

[0280] In this description, an example of the front-back direction is shown, but the audio signal processing device 10 can also realize a weighted sound field in the left-right direction and the height direction (up-down direction) in the same manner.

[0281] Figure 30 (A) Figure 30 (B) is a diagram showing a setting example in which the sound spreads in the lateral direction of the reproduction space. Figure 30 (A) is a diagram showing an example of setting a gain value and a delay amount. Figure 30 (B) is based on Figure 30 (A) is a diagram showing the general situation of the expansion of the sound achieved by the setting. Figure 30 (A) Figure 30 In (B), for the sake of simplicity of explanation and ease of understanding, a case where fourteen speakers SP1 to SP14 are arranged is shown.

[0282] exist Figure 30 (A) Figure 30 In the embodiment shown in (B), for example, a value obtained by quantifying the extension of sound (extension setting value) is set as the acoustic parameter. The gain and hysteresis setting unit 901 calculates the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1 to SP14.

[0283] The gain and hysteresis setting unit 901 calculates gain values ​​for the 14 speakers SP1 to SP14 using the digitized sound expansion value, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1 to SP14, and a shape coefficient setting formula (for gain value setting).

[0284] The gain and lag setting unit 901 calculates the delay amounts of the 14 speakers SP1-SP14 using the delay amounts at the rear and front ends, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and a shape coefficient setting formula (for delay amount setting).

[0285] Through this processing, the sound signal processing device 10 Figure 30 As shown in (A), it is easy to set the sound parameters so that the closer the speakers are to the two ends of the playback space, the greater the gain value and delay amount, and the closer the speakers are to the center of the playback space, the smaller the gain value and delay amount. As described above, the sound signal processing device 10 can easily realize a sound field that is extensible in the horizontal direction of the playback space and in which the sound is localized (see Figure 30 (B)).

[0286] Furthermore, by setting the above-mentioned acoustic parameters, the sound signal processing device 10 can not only achieve weighting in the front-to-back direction, weighting in the left-to-right direction, and lateral expansion of the playback space, but can also achieve weighting and expansion in the height direction (up-down direction) of the playback space. For example, Figure 31 This is a diagram schematically showing the spread of sound in a case where the sound spreads in the height direction.

[0287] The sound signal processing device 10 makes the gain value and delay of the speaker SPU on the ceiling side greater than the gain value and delay of the speakers SPL and SPR near the floor surface. As described above, the sound signal processing device 10 can easily realize a sound field that is more expansive and resonant in the direction of the ceiling of the playback space (see Figure 31 ).

[0288] In the above configuration, the output adjustment unit 90 outputs the output signals So1 to So64 to the plurality of speakers SP1 to SP64. However, the audio signal processing device may perform binaural processing on the output signals So1 to So64 before outputting them.

[0289] Figure 32 This is a functional block diagram showing the structure of a sound signal processing device with binaural playback function. Figure 32 As shown, the audio signal processing device 10A with a binaural playback function differs from the above-mentioned audio signal processing device 10 in that it includes an output adjustment unit 90A, a reverberation processing unit 97 , a selection unit 98 , and a binaural processing unit 99 .

[0290] The output adjustment unit 90A generates a plurality of output signals So1 to So64 based on the plurality of speaker signals Sat1 to Sat64 output from the adder 80 , using the same processing as that of the output adjustment unit 90 described above.

[0291] The output adjustment unit 90A can select an output target. This selection is performed, for example, by user input using the aforementioned GUI. More specifically, the GUI displays an operation button that allows selection between speaker output and binaural output, and the output target is selected by operating the operation button.

[0292] When speaker output is selected, output adjustment unit 90A outputs multiple output signals So1-So64 to multiple speakers SP1-SP64 (same processing as output adjustment unit 90 ). When binaural output is selected, output adjustment unit 90A outputs multiple output signals So1-So64 to selection unit 98 .

[0293] The reverberation processing unit 97 receives input of the sound signals S1-S96 from the multiple sound sources OBJ1-OBJ96. The reverberation processing unit 97 adds an initial reflected sound control signal and a reverberation sound control signal to the multiple sound signals S1-S96 and outputs the signals to the selection unit 98. The initial reflected sound control signal for the multiple sound signals S1-S96 is set based on the position coordinates of the multiple sound sources OBJ1-OBJ96. The reverberation processing unit 97 outputs the multiple reverberation-processed sound signals S1'-S96' to the selection unit 98.

[0294] The selection unit 98 receives inputs of a plurality of output signals So1-So64 and a plurality of reverberated sound signals S1'-S96'. The selection unit 98 selects between the plurality of output signals So1-So64 and the reverberated sound signals S1'-S96' through user input, for example, using the aforementioned GUI. More specifically, the GUI displays an operating element for selecting between sounds subjected to acoustic processing by the sound signal processing device 10A and sounds subjected to virtual acoustic processing based on the positional coordinates of the sound sources OBJ1-OBJ96. By operating the operating element, the output target is selected.

[0295] When a sound subjected to acoustic processing by the sound signal processing device 10A is selected, the selection unit 98 selects a plurality of output signals So1-So64 and outputs them to the binaural processing unit 99. When a sound subjected to virtual acoustic processing based on the position coordinates of the sound sources OBJ1-OBJ96 is selected, the selection unit 98 selects a plurality of reverberation-processed sound signals S1'-S96' and outputs them to the binaural processing unit 99.

[0296] The binaural processing unit 99 performs binaural processing on the input sound signals. More specifically, if the plurality of output signals So1-So64 are input, the binaural processing unit 99 performs binaural processing on the plurality of output signals So1-So64. If the plurality of reverberated sound signals S1'-S96' are input, the binaural processing unit 99 performs binaural processing on the plurality of reverberated sound signals S1'-S96'.

[0297] The binaural processing is a process using a head transfer function, and the details are known, so a detailed description of the binaural processing will be omitted.

[0298] The binaural processing unit 99 outputs the two-channel audio signal subjected to the binaural processing.

[0299] As described above, the user can hear the sound generated by the sound signal processing device 10A and the sound to which the virtual reverberation processing based on the position coordinates of the sound sources OBJ1-OBJ96 is applied through binaural playback. Therefore, even without physically constructing the playback space, the user can easily confirm whether the sound processing applied by the sound signal processing device 10A can reproduce the sound of the virtual space by using headphones or the like. The sound processing applied by the sound signal processing device 10A includes, for example, the grouping of sound sources described above, the setting of the initial reflected sound control signal, the setting of the reverberation sound control signal, the setting of the output control, and the like. Moreover, by performing the audio-visual comparison as described above, the user can adjust the settings of the above-mentioned sound processing to reproduce the sound of the virtual space more realistically.

[0300] Furthermore, binaural playback is not limited to headphones and can also be performed through stereo speakers or the like.

[0301] The description of the present embodiment is illustrative in all respects and is not intended to be restrictive. The scope of the present invention is not indicated by the above-described embodiment but by the claims. Furthermore, the scope of the present invention includes all modifications within the meaning and scope equivalent to the claims.

[0302] Description of the label

[0303] 10.10A: Sound signal processing device

[0304] 30: Area Setting Department

[0305] 40: Grouping Department

[0306] 41: Sound source position detection unit

[0307] 42: Area Determination Unit

[0308] 50: Initial reflected sound control signal generation unit

[0309] 51: FIR filter circuit

[0310] 52: LDtap circuit

[0311] 53: Addition processing unit

[0312] 60: Mixer

[0313] 70: Reverberation sound control signal generation unit

[0314] 71: PEQ

[0315] 72: FIR filter circuit

[0316] 73: Allocator

[0317] 80: Adder

[0318] 90, 90A: Output adjustment unit

[0319] 91: Gain control unit

[0320] 92: Hysteresis control unit

[0321] 97: Echo Processing Department

[0322] 98: Selection Department

[0323] 99: Binaural processing unit

[0324] 100, 100A: GUI

[0325] 400: Matrix mixer

[0326] 500: Operation Department

[0327] 501: Tone setting section

[0328] 502: Virtual sound source setting section

[0329] 511-518: FIR filter

[0330] 521-528: LDtap

[0331] 700: Operation Department

[0332] 701: Reverberant sound area setting unit

[0333] 702: Filter coefficient setting unit

[0334] 703: Reverberant sound playback speaker setting unit

[0335] 721-728: FIR filter

[0336] 900: Operation Department

[0337] 901: Gain and hysteresis setting section

[0338] 909: Display unit

[0339] 5201: Output speaker setting unit

[0340] 5202: Coefficient setting unit

[0341] 9101-9164: Gain control unit

[0342] 9201-9264: Hysteresis control unit.

Claims

1. A method for processing a sound signal, wherein: Receive the sound signal from the first sound source, Receive the position of the virtual sound source, the position of the speaker and the position of the receiving point, The virtual sound sources representing the reflected sound of the target acoustic space are classified into a first virtual sound source, a second virtual sound source, and a third virtual sound source. The first virtual sound source is located between the speaker and the sound collecting point. The second virtual sound source represents the first sound source of the reflected sound of the target acoustic space. The second virtual sound source is located outside the speaker. The second sound source represents the second sound source of the reflected sound of the target acoustic space. The third virtual sound source is located at the same position as the speaker. For the first virtual sound source, when the virtual sound source distance between the sound collection point and the first virtual sound source is smaller than the speaker distance between the sound collection point and the speaker, the position of the first virtual sound source is moved to a position on a straight line passing through the sound collection point and the speaker, and away from the distance difference between the first virtual sound source and the speaker on the side opposite to the sound collection point with the speaker as a reference. Based on the distance difference between the first virtual sound source and the speaker, gain processing and delay processing are performed on the sound signal of the first sound source.

2. The sound signal processing method according to claim 1, wherein: A gain value and a delay amount for the first virtual sound source are set based on a difference between a distance between the speaker and the sound collection point and a distance between the moved first virtual sound source and the sound collection point.

3. The sound signal processing method according to claim 1 or 2, wherein: A gain value and a delay amount for the second virtual sound source are set according to the distance between the second virtual sound source and the speaker.

4. A sound signal processing device comprising: a loudspeaker that plays reflected sound from a target acoustic space; and An initial reflected sound control unit receives a sound signal from a first sound source, a position of a virtual sound source, a position of a loudspeaker, and a position of a sound collection point, and classifies virtual sound sources representing reflected sound in a target acoustic space into a first virtual sound source located between the loudspeaker and the sound collection point, a second virtual sound source located outside the loudspeaker, and a third virtual sound source located at the same position as the loudspeaker. If a virtual sound source distance between the sound collection point and the first virtual sound source is less than a loudspeaker distance between the sound collection point and the loudspeaker, the unit moves the position of the first virtual sound source to a position on a straight line passing through the sound collection point and the loudspeaker, and away from the loudspeaker on the opposite side of the sound collection point by a distance difference between the first virtual sound source and the loudspeaker, and performs gain processing and delay processing on the sound signal of the first sound source based on the distance difference between the first virtual sound source and the loudspeaker.

5. The sound signal processing device according to claim 4, wherein: The initial reflected sound control unit sets a gain value and a delay amount for the first virtual sound source based on a difference between a distance between the speaker and the sound collection point and a distance between the moved first virtual sound source and the sound collection point.

6. The sound signal processing device according to claim 4 or 5, wherein: The initial reflected sound control unit sets a gain value and a delay amount for the second virtual sound source according to a distance between the second virtual sound source and the speaker.

7. A recording medium which is a non-volatile computer-readable recording medium and which records a program for causing a computer to execute the following processing: Receive the sound signal from the first sound source, Receive the position of the virtual sound source, the position of the speaker and the position of the receiving point, The virtual sound source representing the reflected sound of the target acoustic space is classified into a first virtual sound source, a second virtual sound source, and a third virtual sound source, wherein the first virtual sound source is located between the speaker and the sound collecting point, the second virtual sound source is located outside the speaker, the second sound source represents the reflected sound of the target acoustic space, and the third virtual sound source is located at the same position as the speaker. For the first virtual sound source, when the virtual sound source distance between the sound collection point and the first virtual sound source is smaller than the speaker distance between the sound collection point and the speaker, the position of the first virtual sound source is moved to a position on a straight line passing through the sound collection point and the speaker, and away from the distance difference between the first virtual sound source and the speaker on the side opposite to the sound collection point with the speaker as a reference. Based on the distance difference between the first virtual sound source and the speaker, gain processing and delay processing are performed on the sound signal of the first sound source.

Citation Information

Patent Citations

  • Acoustic processing device

    CN101010987A

  • Simulation System, Sound Processing Method And Information Storage Medium

    CN107277736A