Sound signal processing method, sound signal processing apparatus, and recording medium

By using virtual sound source filtering and timbre adjustment in virtual space, the problem of poor sound quality of initial reflected sound was solved, achieving more realistic sound quality and clearer sound image positioning.

CN115119102BActive Publication Date: 2026-05-12YAMAHA CORP
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
YAMAHA CORP
Filing Date
2022-03-14
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies struggle to improve the sound quality of the initial reflected sound.

Method used

The virtual sound source in the virtual space is generated and filtered to adjust the timbre of the virtual sound source. The initial reflected sound and the reverberation sound are simulated and controlled respectively, and finally the initial reflected sound control signal is generated.

Benefits of technology

It improves the sound quality of the initial reflected sound, making it more realistic and natural, and achieves clearer sound image localization and richer spatial expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115119102B_ABST
    Figure CN115119102B_ABST
Patent Text Reader

Abstract

The sound signal processing method and the sound signal processing apparatus of the present application improve the sound quality of an initial reflected sound. The sound signal processing method obtains a sound signal of a sound source, applies a first filter process of generating a virtual sound source of a virtual space to the sound signal, applies a second filter process of adjusting the tone color of the virtual sound source to the sound signal, and outputs an initial reflected sound control signal generated by applying the first filter process and the second filter process to the sound signal.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing apparatus for performing prescribed processing on sound input from a sound source. Background Technology

[0002] The technology for controlling reflected sound in sound systems such as those in halls is being put into practical use in various ways.

[0003] For example, the reflected sound generation device described in Patent Document 1 includes a first FIR filter and a second FIR filter. The first FIR filter performs a convolution operation on the speech signal using the first reflected sound parameter to generate first reflected sound data. The second FIR filter performs a convolution operation on the first reflected sound data using the second reflected sound parameter to generate second reflected sound data.

[0004] Thus, the reflected sound generating device described in Patent Document 1 generates a reflected sound consisting of an initial reflected sound and a subsequent echo.

[0005] Patent Document 1: Japanese Patent Application Publication No. 2000-163086

[0006] However, in the existing structure described above, it is difficult to improve the sound quality of the initial reflected sound. Summary of the Invention

[0007] Therefore, one embodiment of the present invention aims to improve the sound quality of the initial reflected sound.

[0008] The sound signal processing method obtains the sound signal from the sound source, performs a first filtering process on the sound signal to generate a virtual sound source in a virtual space, performs a second filtering process on the sound signal to adjust the timbre of the virtual sound source, and outputs an initial reflection control signal generated by the sound signal after the first and second filtering processes.

[0009] The effects of the invention

[0010] Sound signal processing methods can improve the sound quality of the initial reflected sound. Attached Figure Description

[0011] Figure 1 This is a functional block diagram illustrating the structure of an audio system that includes the sound signal processing apparatus according to embodiments of the present invention.

[0012] Figure 2 This is a flowchart of a sound signal processing method according to an embodiment of the present invention.

[0013] Figure 3 It is a diagram representing the discrete waveforms of a sound that includes the usual direct tone, initial reflection tone, and reverberation tone (back reverberation tone).

[0014] Figure 4 (A) Figure 4 (B) is a diagram representing the conceptual designation of a virtual sound source.

[0015] Figure 5 This is a functional block diagram illustrating an example of the structure of the grouping section 40.

[0016] Figure 6 This is a flowchart representing the grouping method of sound sources.

[0017] Figure 7 It is a diagram representing the concept of grouping multiple sound sources into multiple regions.

[0018] Figure 8 (A) is a flowchart illustrating the grouping method using sound sources representing points. Figure 8 (B) is a flowchart representing the grouping method of sound sources at the boundary of the area of ​​use.

[0019] Figure 9 This is a flowchart illustrating an example of a grouping method based on the movement of a sound source.

[0020] Figure 10 This is a functional block diagram illustrating an example of the structure of the initial reflected sound control signal generation unit 50.

[0021] Figure 11 This is a diagram representing an example of a GUI.

[0022] Figure 12 This is a flowchart illustrating an example of the setting and processing of a virtual sound source.

[0023] Figure 13 (A) Figure 13 (B) is a diagram showing the configuration examples of various virtual sound sources with different geometric shapes.

[0024] Figure 14 (A) Figure 14 (B) and Figure 14 (C) is a diagram representing a setting example of a virtual sound source.

[0025] Figure 15 (A) Figure 15 (B) Figure 15 (C) is a diagram representing a setting example of a virtual sound source.

[0026] Figure 16 This is a flowchart illustrating the process of assigning virtual sound sources to loudspeakers.

[0027] Figure 17 (A) Figure 17 (B) is a diagram illustrating the concept of assigning a virtual sound source to a loudspeaker.

[0028] Figure 18 This is a flowchart representing the coefficient setting process of LDtap.

[0029] Figure 19 (A) Figure 19 (B) is a diagram used to illustrate the concept of coefficient setting.

[0030] Figure 20 (A) shows an example of the LDtap coefficient for cases where the virtual space shape is large. Figure 20 (B) shows an example of the LDtap coefficient for the case where the virtual space shape is small.

[0031] Figure 21 It is a diagram showing the waveform of the initial reflected tone control signal generated by the initial reflected tone control signal generation unit 50.

[0032] Figure 22 This is a functional block diagram illustrating an example of the structure of the echo control signal generation unit 70.

[0033] Figure 23 This is a flowchart illustrating an example of the generation and processing of the echo control signal.

[0034] Figure 24 It is a graph showing the waveforms of the direct tone, the initial reflection tone control signal, and the reverberation tone control signal.

[0035] Figure 25 This is a diagram illustrating an example of a zone setting used to represent reverberation.

[0036] Figure 26 This is a functional block diagram illustrating an example of the structure of the output adjustment unit 90.

[0037] Figure 27 This is a flowchart illustrating an example of output adjustment processing.

[0038] Figure 28 This is a diagram illustrating an example of a GUI used for output adjustment.

[0039] Figure 29 (A) Figure 29 (B) is a diagram illustrating a setting example where sound is positioned and extended at the rear of the playback space.

[0040] Figure 30 (A) Figure 30 (B) is a diagram illustrating a setting example where sound is positioned and extended laterally in the playback space.

[0041] Figure 31 It is a diagram that shows the general extent of sound expansion in cases of vertical expansion.

[0042] Figure 32 This is a functional block diagram representing the structure of a sound signal processing device with binaural playback capability. Detailed Implementation

[0043] The sound signal processing method and sound signal processing apparatus according to embodiments of the present invention will be described with reference to the accompanying drawings. Furthermore, in the following embodiments, firstly, an overview of the sound signal processing method and sound signal processing apparatus will be described, and then the specific details of each processing step and each structure will be explained.

[0044] Furthermore, in this embodiment, the playback space is the space in which the user (listener) listens to the sound (direct tone, initial reflection tone, reverberation tone) from the sound source using a speaker or the like. The virtual space is a space with a sound field (audio) different from the playback space, and is a space reproduced (simulated) in the playback space by the initial reflection tone and reverberation tone obtained through that sound field.

[0045] [Brief Structure of a Sound Signal Processing Device]

[0046] Figure 1 This is a functional block diagram illustrating the structure of an audio system that includes the sound signal processing apparatus according to embodiments of the present invention.

[0047] like Figure 1 As shown, the sound signal processing apparatus 10 includes a zone setting unit 30, a grouping unit 40, an initial reflection control signal generation unit 50, a mixer 60, an echo control signal generation unit 70, an adder 80, and an output adjustment unit 90. The sound signal processing apparatus 10 is implemented, for example, by an electronic circuit or a computer or other computational processing device that respectively implements the zone setting unit 30, the grouping unit 40, the initial reflection control signal generation unit 50, the mixer 60, the echo control signal generation unit 70, the adder 80, and the output adjustment unit 90. The portion consisting of the adder 80 and the output adjustment unit 90 corresponds to the "output signal generation unit" of this invention.

[0048] The sound signal processing unit 10 is connected to multiple speakers SP1-SP64. Furthermore, Figure 1 The example shown uses 64 speakers, but the number of speakers is not limited to this.

[0049] The sound signal processing device 10 receives sound signals S1-S96 from multiple sound sources OBJ1-OBJ96. Furthermore, Figure 1 The example shown uses 96 sound sources, but the number of sound sources is not limited to this.

[0050] The region setting unit 30 divides the playback space into multiple regions and sets information (region information) related to the divided regions. The region information includes the position coordinates of the region's boundaries and the position coordinates of the representative point set in the region.

[0051] The region setting unit 30 outputs the region information of the multiple regions Area1 to Area8 to the grouping unit 40. Furthermore, in Figure 1 The diagram shows a configuration with 8 regions, but the number of regions is not limited to this.

[0052] The grouping unit 40 groups sound sources OBJ1-OBJ96 into multiple regions Area1-Area8. Based on the grouping results, the grouping unit 40 generates region-specific sound signals SA1-SA8 for each region Area1-Area8 using the sound signals S1-S96 from sound sources OBJ1-OBJ96. For example, the grouping unit 40 mixes the sound signals from multiple sound sources grouped into region Area1 to generate the region-specific sound signal SA1.

[0053] The grouping unit 40 outputs multiple regionally differentiated sound signals SA1-SA8 to the initial reflected sound control signal generation unit 50. Additionally, the grouping unit 40 outputs sound signals S1-S96 from sound sources OBJ1-OBJ96 to the mixer 60.

[0054] The initial reflection control signal generation unit 50 generates initial reflection control signals ER1-ER64 for each of the multiple speakers SP1-SP64 based on multiple area-divided sound signals SA1-SA8. The initial reflection control signals ER1-ER64 are signals output to each of the speakers SP1-SP64 to simulate the initial reflections of the virtual space in the playback space. The initial reflection control signal generation unit 50 outputs the generated initial reflection control signals ER1-ER64 to the adder 80.

[0055] In general terms (detailed structure and processing will be described later), the initial reflection control signal generation unit 50 uses the positions of the speakers SP1-SP64 arranged in the playback space and the geometry of the virtual space to set the virtual sound source in the playback space. Furthermore, the specific settings for the virtual sound source will be described later. The initial reflection control signal generation unit 50 generates initial reflection control signals ER1-ER64 that simulate the initial reflection in the virtual space using the virtual sound source. At this time, the initial reflection control signal generation unit 50 adjusts the initial reflection control signals ER1-ER64 to the desired timbre.

[0056] Mixer 60 is an additive mixer. Mixer 60 mixes the sound signals S1-S96 from sound sources OBJ1-OBJ96 to generate an echo generation signal Sr. Mixer 60 outputs the echo generation signal Sr to the echo control signal generation unit 70.

[0057] The reverberation control signal generation unit 70 generates reverberation control signals REV1-REV64 for each of the multiple speakers SP1-SP64 based on the reverberation generation signal Sr. The reverberation control signals REV1-REV64 are signals output to each of the speakers SP1-SP64 to simulate the reverberation (rear reverberation) of the virtual space in the playback space. The reverberation control signal generation unit 70 outputs the generated reverberation control signals REV1-REV64 to the adder 80.

[0058] In general terms (detailed structure and processing will be described later), the reverberation control signal generation unit 70 divides the playback space into multiple reverberation setting areas and generates reverberation control signals for each of the multiple reverberation setting areas. The reverberation control signal generation unit 70 assigns multiple speakers SP1-SP64 to the multiple reverberation setting areas. Based on this assignment, the reverberation control signal generation unit 70 sets the reverberation control signal for each of the multiple speakers SP1-SP64 for each reverberation setting area.

[0059] At this time, the echo control signal generation unit 70 sets the connection timing of the initial reflected sound and the echo sound based on the geometric shape of the playback space. During the period before the connection timing, the echo control signal generation unit 70 gradually increases the level (amplitude) of the echo control signal, and during the connection timing and thereafter, it gradually decreases the level (amplitude) of the echo control signal.

[0060] Adder 80 adds the initial reflection control signals and reverberation control signals generated for each of the multiple speakers SP1-SP64 to generate multiple speaker signals Sat1-Sat64. For example, adder 80 adds the initial reflection control signal and the reverberation control signal for speaker SP1 to generate speaker signal Sat1. Adder 80 outputs the multiple speaker signals Sat1-Sat64 to output adjustment unit 90.

[0061] The output adjustment unit 90 performs gain control and hysteresis control on signals Sat1-Sat64 from multiple loudspeakers to generate output signals So1-So64. The output adjustment unit 90 then outputs output signals So1-So64 to multiple loudspeakers SP1-SP64. For example, the output adjustment unit 90 performs gain control and hysteresis control on signal Sat1 from loudspeaker SP1 to generate output signal So1. The output adjustment unit 90 then outputs output signal So1 to loudspeaker SP1.

[0062] (Detailed structure and processing will be described later), in general terms, the output adjustment unit 90 receives input audio parameters of the playback space. Audio parameters include, for example, adjustments to the expansion of the sound space in the width direction, the expansion of the sound space behind the reception point, and the expansion of the sound space in the ceiling direction. Based on the position coordinates of the multiple speakers SP1-SP64 and the audio parameters, the output adjustment unit 90 centrally sets the gain and hysteresis (delay) values ​​of the multiple speaker signals Sat1-Sat64. Centralized setting means that the settings are not performed individually for each speaker; for example, the gain and hysteresis values ​​of each speaker are set by inputting the position coordinates of each speaker into a specific calculation formula shared by all speakers. The output adjustment unit 90 uses the set gain and hysteresis values ​​to perform gain control and hysteresis control on the multiple speaker signals Sat1-Sat64.

[0063] [An overview of sound signal processing methods]

[0064] Figure 2 This is a flowchart of a sound signal processing method according to an embodiment of the present invention. Figure 2 Shown by Figure 1 The sound signal processing method implemented by the sound signal processing device 10. Furthermore, Figure 2 The contents of each process shown are described above. Figure 1 As explained in the description, it will be briefly recorded here.

[0065] (Grouping of sound sources OBJ1-OBJ96)

[0066] The grouping unit 40 groups the multiple sound sources OBJ1-OBJ96 into multiple regions Area1-Area8 respectively (S11).

[0067] (Generation of the initial reflected tone control signal)

[0068] The initial reflection control signal generation unit 50 sets the timbre for the initial reflection for each group (S12). The initial reflection control signal generation unit 50 sets a virtual sound source for each group (S13). Using the timbre and the virtual sound source, the initial reflection control signal generation unit 50 generates initial reflection control signals for each of the multiple speakers SP1-SP64 (S14).

[0069] (Generation of echo control signal)

[0070] Mixer 60 adds the audio signals S1-S96 from multiple sound sources OBJ1-OBJ96 (S21). Echo control signal generation unit 70 sets the connection timing of the initial reflected sound and echo based on the geometry of the playback space (S22). Echo control signal generation unit 70 generates an echo control signal using the set connection timing (S23). Echo control signal generation unit 70 distributes the generated echo control signal to multiple speakers SP1-SP64 based on their position coordinates in the playback space (S24).

[0071] (Output processing to multiple speakers)

[0072] Adder 80 adds the initial reflection control signal and the echo control signal to the multiple loudspeakers SP1-SP64 respectively to generate loudspeaker signals Sat1-Sat64 (S31).

[0073] The output adjustment unit 90 uses audio parameters that enable the positioning of reverberation and the expansion of space in the playback space to generate output signals So1-So64 based on the speaker signals Sat1-Sat64 (S32). The output adjustment unit 90 outputs the output signals So1-So64 to multiple speakers SP1-SP64 (S33).

[0074] By using the above-described structure and processing, the sound signal processing apparatus 10 (sound signal processing method) achieves the following various effects.

[0075] (1) The sound signal processing apparatus 10 (sound signal processing method) groups the sound sources for each region obtained by dividing the playback space, and generates initial reflected sounds, thereby achieving clear sound image localization and rich spatial expansion. At this time, the reverberation sound is constant throughout the playback space, and only the initial reflected sounds vary depending on the position of the sound source. Therefore, for example, when the position of the sound source moves, the movement of the sound from that sound source is smoother.

[0076] (2) The sound signal processing device 10 (sound signal processing method) generates an initial reflected sound control signal using a virtual sound source, thereby enabling more realistic simulation of the initial reflected sound obtained from the geometric shape of the virtual space in the playback space.

[0077] (3) The sound signal processing device 10 (sound signal processing method) adjusts the timbre of the initial reflected tone control signal, thereby eliminating, for example, the unnaturalness of the timbre of the initial reflected tone simulated only by a virtual sound source.

[0078] (4) The sound signal processing device 10 (sound signal processing method) sets the connection timing of the initial reflection control signal and the reverberation control signal according to the geometric shape of the playback space, thereby enabling a smoother and more natural connection from the initial reflection to the reverberation.

[0079] (5) The sound signal processing device 10 (sound signal processing method) adjusts the gain and hysteresis of the loudspeaker signals Sat1-Sat64, which include the initial reflection control signal and the reverberation control signal, thereby enabling the user to achieve the desired sound field in the playback space through easier operation input.

[0080] [Detailed descriptions of each signal processing unit and each process]

[0081] The following describes the specific details of each signal processing unit and each processing step. First, referring to the accompanying drawings, the initial reflected sound, echo sound, and virtual sound source necessary for understanding the invention will be explained.

[0082] [Initial reflected sound and echo]

[0083] Figure 3 It is a diagram representing the discrete waveforms of a sound that typically includes the direct tone, initial reflection, and reverberation (back reverberation). For example, a hall where a performance or content is played has an enclosed space surrounded by walls. If a sound occurs in this enclosed space, the direct tone, initial reflection, and reverberation (back reverberation) reach the reception point.

[0084] A direct tone is a sound that travels directly from the point of origin to the point of reception.

[0085] The initial reflected sound is the sound that, after being reflected by walls, floors, or ceilings at its origin, arrives at the receiver relatively early. Therefore, the initial reflected sound arrives at the receiver after the direct sound. Furthermore, the volume (level) of the initial reflected sound is less than the volume (level) of the direct sound. If the number of reflections is one, it is a first-order reflected sound; if it is n times, it is an nth-order reflected sound. The direction of arrival and the volume of the initial reflected sound at the receiver are largely influenced by the location of the sound's origin.

[0086] The reverberation arrives at the receiver after the initial reflection. An echo is the sound that occurs at the point of origin and reaches the receiver after multiple reflections. That is, an echo is the sound that arrives at the receiver after further reflections and attenuation of the initial reflection. Therefore, the volume (level) of the echo is less than the volume (level) of the initial reflection. Furthermore, the direction of arrival and volume of the echo are less affected by the location of origin compared to the initial reflection.

[0087] [Virtual sound source]

[0088] Figure 4 (A) Figure 4 (B) is a diagram representing the conceptual definition of a virtual sound source. Furthermore, in Figure 4 (A) Figure 4 In (B), for ease of explanation, the concept of setting up a two-dimensional virtual sound source is shown, but the virtual sound source can also be set up in three dimensions using the same concept. That is, in the actual playback space, when the sound sources are not aligned on a plane but are spatially arranged, and the virtual space is set up in a three-dimensional manner, the virtual sound source is set up in three dimensions.

[0089] There are sound sources SS and receivers RP in the playback space. Furthermore, Figure 4 (A) Figure 4 The sound source SS shown in (B) has a different meaning from the sound source OBJ described above; it refers to a sound source that produces ordinary sound. Additionally, a virtual wall IWL is set in the playback space to realize the sound field of the virtual space. The virtual wall IWL is obtained based on the geometry of the virtual space.

[0090] The sound source SS and the pickup point RP exist within a space enclosed by virtual walls IWL. The virtual walls IWL include virtual walls IWL1, IWL2, IWL3, and IWL4. Virtual walls IWL1 and IWL4 are positioned in the first direction of the playback space (…). Figure 4 (A) Figure 4 (B) The sound source SS and the receiver RP are positioned in the middle, with the virtual wall IWL1 positioned closer to the sound source SS than the receiver RP, and the virtual wall IWL4 positioned closer to the receiver RP than the sound source SS. Virtual walls IWL2 and IWL3 are positioned in the second direction of the playback space ( Figure 4 (A) Figure 4 (B) The sound source SS and the receiver RP are sandwiched in the middle. The virtual wall IWL2 is positioned closer to the sound source SS than the receiver RP, and the virtual wall IWL3 is positioned closer to the receiver RP than the sound source SS.

[0091] If virtual walls IWL1, IWL2, IWL3, and IWL4 are real-world walls that reflect sound, then... Figure 4 As shown in (B), the sound emitted from the sound source SS is reflected at virtual walls IWL1, IWL2, and IWL3 before reaching the receiving point RP. Furthermore, in Figure 4 In (B), there is no record of reflection from virtual wall IWL4, but virtual wall IWL4 also produces reflection in the same way as virtual walls IWL1, virtual wall IWL2 and virtual wall IWL3.

[0092] However, virtual walls IWL1, IWL2, IWL3, and IWL4 do not actually exist in the playback space. Therefore, as Figure 4 As shown in (A), the sound signal processing device 10 uses the reflection of sound at the wall as a mirror reflection to set virtual sound source IS1, virtual sound source IS2 and virtual sound source IS3.

[0093] Specifically, the sound signal processing device 10 sets the virtual sound source IS1 at a position symmetrical about the sound source SS line with virtual wall IWL1 as a reference line. The sound signal processing device 10 sets the virtual sound source IS2 at a position symmetrical about the sound source SS line with virtual wall IWL2 as a reference line. The virtual sound source IS3 is set at a position symmetrical about the sound source SS line with virtual wall IWL3 as a reference line. Furthermore, by adjusting the sound power of each virtual sound source IS, the energy loss from reflections at the virtual wall IWL can be simulated.

[0094] By making the settings described above, the sound generated at virtual sound source IS1 is the same as the sound generated at sound source SS and reflected at virtual wall IW1. The sound generated at virtual sound source IS2 is the same as the sound generated at sound source SS and reflected at virtual wall IW2. The sound generated at virtual sound source IS3 is the same as the sound generated at sound source SS and reflected at virtual wall IW3. Furthermore, in Figure 4 (A) Figure 4 (B) does not describe the virtual sound source for virtual wall IWL4, but virtual sound sources can be set for virtual wall IWL4 in the same way as virtual wall IWL1, virtual wall IWL2 and virtual wall IWL3.

[0095] By setting the virtual sound source in the manner described above, the sound signal processing device 10 is able to simulate the initial reflected sound of the virtual space in a playback space without real walls.

[0096] [Structure and processing of group section 40]

[0097] Figure 5This is a functional block diagram illustrating an example of the structure of the grouping section 40. Figure 6 This is a flowchart representing the grouping method of sound sources.

[0098] like Figure 5 As shown, the grouping unit 40 includes a sound source location detection unit 41, a region determination unit 42, and a matrix mixer 400.

[0099] The sound source location detection unit 41 detects the position coordinates of multiple sound sources OBJ1-OBJ96 in the playback space. Figure 6 (S111). For example, the sound source position detection unit 41 detects the position coordinates of sound sources OBJ1-OBJ96 based on user input. Alternatively, the sound source position detection unit 41 may have a position detection sensor for detecting sound sources OBJ1-OBJ96, and detect the position coordinates of sound sources OBJ1-OBJ96 based on the position detected by the position detection sensor.

[0100] The sound source location detection unit 41 outputs the position coordinates of sound sources OBJ1-OBJ96 to the region determination unit 42.

[0101] The region determination unit 42 uses region information from multiple regions Area1-Area8 from the region setting unit 30 and the position coordinates of sound sources OBJ1-OBJ96 from the sound source position detection unit 41 to group sound sources OBJ1-OBJ96 into multiple regions Area1-Area8. Figure 6 (S112). More specifically, the region determination unit 42 groups regions in the following manner.

[0102] Figure 7 This is a diagram representing the concept of grouping multiple sound sources into multiple regions. Furthermore, in Figure 7 In the image, the top side is the front of the hall that serves as the playback space, and the bottom side is the back of the hall.

[0103] The region setting unit 30 sets a reference point Pso for region segmentation based on the playback space. For example, such as Figure 7 As shown, the zone setting unit 30 sets the center position of the hall where the playback space is to be played at the reference point Pso. Furthermore, the zone setting unit 30 can also use a point (position) set by the user as the reference point. For example, the zone setting unit 30 can use a receiver point or similar point set by the user as the reference point.

[0104] The area setting unit 30 sets eight areas, Area1 to Area8, by dividing the entire perimeter of the plane into eight parts, using the reference point Pso for area division as the center. For example, in Figure 7In this case, the area setting unit 30 sets multiple areas Area1, Area2, and Area3 in the lobby (playback space) at positions further forward than the reference point Pso. Additionally, the area setting unit 30 sets Area4 to the left of the lobby facing forward from the reference point Pso, and Area5 to the right of the lobby facing forward from the reference point Pso. Furthermore, the area setting unit 30 sets multiple areas Area6, Area7, and Area8 in the lobby (playback space) at positions further back than the reference point Pso.

[0105] Furthermore, this area setting is just one example; any other setting is possible as long as the entire playback space can be covered by multiple set areas. Additionally, this description illustrates the setting of a planar area, but spatial areas can also be set in the same way. For example, the vertical range of area Area1 is also included within area Area1.

[0106] The area setting unit 30 sets representative points RP1-RP8 for each of the multiple areas Area1-Area8. For example, the area setting unit 30 sets the multiple representative points RP1-RP8 at the center of the multiple areas Area1-Area8. Or, in such cases... Figure 7 In the case of such a radially spreading area, for example, the area setting unit 30 sets the representative point at a predetermined distance from the reference point Pso on a straight line passing through the center of the radially spreading angle. Furthermore, one example of how these representative points are set is that one representative point can be set for each area; any other method is acceptable as long as it reliably performs the grouping of sound sources.

[0107] The area setting unit 30 outputs area information of multiple areas Area1 to Area8 to the area determination unit 42 of the grouping unit 40 and the matrix mixer 400. The area information of the multiple areas Area1 to Area8 includes the position coordinates of representative points RP1 to RP8 of areas Area1 to Area8, the coordinate information of the boundary lines that represent the shapes of areas Area1 to Area8, etc.

[0108] (A method of grouping sound sources into regions using representative points)

[0109] Figure 8 (A) is a flowchart representing a grouping method for sound sources using representative points.

[0110] The region determination unit 42 obtains the position coordinates of representative points RP1-RP8 based on the region information of multiple regions Area1-Area8 (S1121). The region determination unit 42 calculates the distance between the position coordinates of the sound source of the grouped determination object and the position coordinates of the representative points RP1-RP8 (S1122). The region determination unit 42 groups the sound sources into regions containing the representative points that have the shortest distance (S1123).

[0111] For example, in Figure 7 In the example where sound source OBJ1 is used, the region determination unit 42 detects the position coordinates of sound source OBJ1 and obtains the position coordinates of multiple representative points RP1-RP8. Based on the position coordinates of sound source OBJ1 and the position coordinates of the multiple representative points RP1-RP8, the region determination unit 42 calculates the distance between sound source OBJ1 and the multiple representative points RP1-RP8. The region determination unit 42 detects cases where the distance between sound source OBJ1 and representative point RP1 is shorter than the distance between sound source OBJ1 and other representative points RP2-RP8. In other words, the region determination unit 42 detects cases where the distance between sound source OBJ1 and representative point RP1 is the shortest distance. The region determination unit 42 groups sound source OBJ1 into the region Area1 associated with representative point RP1.

[0112] (A method of grouping sound sources into regions using region boundaries)

[0113] Figure 8 (B) is a flowchart representing the grouping method of sound sources at the boundary of the area of ​​use.

[0114] The region determination unit 42 obtains the coordinate information (boundary coordinates) representing the boundary lines of each of the multiple regions Area1 to Area8 based on the region information (S1124). The region determination unit 42 then determines whether the position coordinates of the sound source of the grouped object are inside each of the regions Area1 to Area8 (S1125). For example, the region determination unit 42 uses the Crossing Number Algorithm to determine whether the sound source is inside or outside the region. If the sound source is within the region (S1125: YES), the region determination unit 42 groups the sound source into that region (S1126).

[0115] For example, in Figure 7In the example where sound source OBJ1 is involved, the region determination unit 42 detects the position coordinates of sound source OBJ1 and obtains the coordinate information (boundary coordinates) representing the boundary lines of multiple regions Area1 to Area8. Based on the position coordinates of sound source OBJ1 and the boundary coordinates of the multiple regions Area1 to Area8, the region determination unit 42 determines whether sound source OBJ1 is inside or outside of the multiple regions Area1 to Area8. The region determination unit 42 detects the case where sound source OBJ1 is within region Area1. The region determination unit 42 then groups sound source OBJ1 into region Area1.

[0116] The area determination unit 42 groups the multiple input sound sources OBJ1-OBJ96 into multiple areas Area1-Area8. For example, if it is... Figure 7 For example, the region determination unit 42 groups sound sources OBJ1 and OBJ4 to region Area1, sound source OBJ2 to region Area2, and sound source OBJ3 to region Area5.

[0117] The zone determination unit 42 outputs the grouping information to the matrix mixer 400. The grouping information indicates which sound source is grouped into which zone.

[0118] The matrix mixer 400 generates area-specific sound signals SA1-SA8 for multiple areas Area1-Area8 based on grouping information and using sound signals S1-S96 from multiple sound sources OBJ1-OBJ96. For example, if there are multiple sound sources in an area group, the matrix mixer 400 mixes the sound signals from these multiple sound sources to generate an area-specific sound signal for that area. The matrix mixer 400 outputs the area-specific sound signals for each area to the initial reflection control signal generation unit 50. Furthermore, even if there is only one sound source in an area group, the matrix mixer 400 outputs the sound signal of that sound source as the area-specific sound signal for that area to the initial reflection control signal generation unit 50.

[0119] in the case of Figure 7 For example, Area 1 is grouped with sound sources OBJ1 and OBJ4. The matrix mixer 400 mixes the audio signal S1 from sound source OBJ1 and the audio signal S4 from sound source OBJ4 to generate and output the area-specific audio signal SA1 for Area 1. Additionally, Area 2 is grouped with sound source OBJ2. The matrix mixer 400 outputs the audio signal S2 from sound source OBJ2 as the area-specific audio signal SA2 for Area 2. Furthermore, Area 5 is grouped with sound source OBJ3. The matrix mixer 400 outputs the audio signal S3 from sound source OBJ3 as the area-specific audio signal SA5 for Area 5.

[0120] By implementing the structure and processing described above, the sound signal processing device 10 can group multiple sound sources into multiple regions that divide the sound space, and generate initial reflected sound control signals. As described above, the sound signal processing device 10 can reproduce the initial reflected sound corresponding to the position of the sound source, and can achieve clear sound image localization and rich spatial expansion.

[0121] Furthermore, while the above description does not detail the situation where the sound source moves, the grouping unit 40 performs [operations] when the sound source moves. Figure 9 The processing shown. Figure 9 This is a flowchart illustrating an example of a grouping method based on the movement of a sound source.

[0122] The sound source position detection unit 41 detects the movement of the sound source (S104). The sound source position detection unit 41 detects the movement of the sound source, for example, through user input. Alternatively, the sound source position detection unit 41 continuously detects the position of the sound source using a position detection sensor, thereby detecting the movement of the sound source. Then, the region determination unit 42 regroups the moved sound sources (S105). The sound source position detection unit 41 detects the position coordinates of the moved sound sources and outputs them to the region determination unit 42.

[0123] The area determination unit 42 uses the position coordinates of the moved sound source to group the sound sources into multiple areas Area1-Area8 as described above (S105).

[0124] By performing the processing described above, even if the sound source moves, the sound signal processing device 10 can generate an initial reflected sound control signal corresponding to the position of the moved sound source. As described above, the sound signal processing device 10 can reproduce the changes in the initial reflected sound corresponding to the movement of the sound source, and even if the sound source moves, it can achieve clear sound image localization and rich spatial expansion corresponding to the movement.

[0125] Furthermore, when the sound source moves as described above, the sound signal processing device 10 can perform crossfade processing on the initial reflected sound control signal before the movement and the initial reflected sound control signal after the movement. For example, when the sound source moves, the sound signal processing device 10 gradually decreases the component of the sound signal of that sound source in the area-divided sound signal including the sound source before the movement. On the other hand, the sound signal processing device 10 gradually increases the component of the sound signal of that sound source in the area-divided sound signal including the sound source after the movement.

[0126] By performing the processing described above, the sound signal processing device 10 can suppress the discontinuous changes in the initial reflected sound when the sound source moves. As described above, when the sound source moves, the sound signal processing device 10 can make the initial reflected sound change more smoothly in response to the movement of the sound source.

[0127] Additionally, the matrix mixer 400 outputs the audio signals S1-S96 from multiple sound sources OBJ1-OBJ96 to the mixer 60. As described above, the mixer 60 adds the audio signals S1-S96 to generate an echo generation signal Sr and outputs it to the echo control signal generation unit 70. The echo control signal generation unit 70 uses the echo generation signal Sr to generate echo control signals REV1-REV64.

[0128] Through the processing described above, the reverberation is not affected by the position or movement of the sound source. Therefore, even if the sound source moves, the sound signal processing device 10 can maintain the reverberation in the playback space as constant and reproduce the movement of the sound source more clearly through changes in the initial reflected sound.

[0129] [Generation of the initial reflected tone control signal]

[0130] Figure 10 This is a functional block diagram illustrating an example of the structure of the initial reflected sound control signal generation unit 50. Figure 11 This is a diagram representing an example of a GUI.

[0131] like Figure 10 As shown, the initial reflected sound control signal generation unit 50 includes an FIR filter circuit 51, an LDtap circuit 52, an addition processing unit 53, a timbre setting unit 501, a virtual sound source setting unit 502, and an operation unit 500. The LDtap circuit 52 is a circuit that amplifies and delays the input signal and outputs it. The FIR filter circuit 51 includes multiple FIR filters 511-518. The LDtap circuit 52 includes multiple LDtap filters 521-528, an output speaker setting unit 5201, and a coefficient setting unit 5202. Furthermore, the order of the FIR filter circuit 51 and the LDtap circuit 52 can also be reversed.

[0132] [Timbre adjustment of the initial reflected tone]

[0133] The operation unit 500 receives timbre specification information for the initial reflected tone from the user and outputs it to the timbre setting unit 501. The timbre specification information includes, for example, information specifying the emphasis on the low-frequency range, the emphasis on the high-frequency range, the volume of the initial reflected tone, and the attenuation characteristics of the initial reflected tone (information indicating filtering characteristics).

[0134] As a specific example, the operation unit 500, through, as... Figure 11The GUI 100 (Graphical User Interface) shown is used to receive operations.

[0135] The GUI 100 has a setting display window 111, multiple operating components 112, a knob 1131, and an adjustment value display window 1132.

[0136] The setting display window 111 displays the shape of the virtual wall IWL of the virtual space, which is set by multiple operating elements 112 and knobs 1131. At this time, the setting display window 111 can display the positions of the sound source SS, the speaker SP, the receiver RP, the coordinate axes of the playback space, and the virtual wall IWL together.

[0137] Multiple operation units 112 are associated with pre-defined samples of virtual spaces (various lobbies, rooms, etc.). Furthermore, although the illustration is omitted, multiple operation units 112 are displayed with indexes (e.g., lobbies, etc.) that clearly indicate the samples of virtual spaces associated with each operation unit 112.

[0138] Knob 1131 is used to set the room size of the virtual space. Adjustment value display window 1132 displays the set value of the room size of the virtual space.

[0139] The GUI 100 receives various operations for adjusting tone. For example, the GUI 100 has multiple operation elements 112, including operation elements for the low-frequency range, operation elements for the high-frequency range, operation elements for volume adjustment, operation elements for attenuation characteristic adjustment, etc., and receives operations through these operation elements.

[0140] If a user operates a desired control using the GUI 100, the operation unit 500 detects the operation and sets the specified information for the tone accordingly.

[0141] For example, if the operation unit 500 receives selections from multiple operation elements 112, it obtains the specified timbre information preset in the virtual space associated with that operation element 112. Furthermore, if the operation unit 500 receives operations performed via operation elements for the low-frequency range, high-frequency range, volume adjustment, attenuation characteristic adjustment, etc., it obtains the specified timbre information set via these operation elements.

[0142] Furthermore, although illustrations are omitted, the GUI 100 can also display specified information about the timbre using, for example, the filter coefficients of the FIR filters 511-518 described later, and a rough waveform. In this case, if the GUI 100 receives an adjustment to the specified timbre information, it can also change the display accordingly. For example, the GUI 100 can change the waveform display in response to the adjustment.

[0143] The tone setting unit 501 sets the filtering coefficients of the FIR filters 511-518 of the FIR filter circuit 51 based on the specified tone information. For example, if the tone setting unit 501 receives specified information emphasizing the low-frequency range, it sets the filtering coefficients of the low-frequency range of the FIR filters 511-518 of the FIR filter circuit 51 to enhance the low-frequency range. Similarly, if the tone setting unit 501 receives specified information emphasizing the high-frequency range, it sets the filtering coefficients of the high-frequency range of the FIR filters 511-518 of the FIR filter circuit 51 to enhance the high-frequency range. The tone setting unit 501 outputs the set filtering coefficients to the FIR filter circuit 51. Furthermore, not limited to filtering coefficients, the tone setting unit 501 can also set and adjust the sampling frequency and filter length as filtering characteristics.

[0144] Furthermore, the tone setting unit 501 sets the gain value of each tap of the FIR filters 511-518 in the FIR filter circuit 51 based on the tone specification information. The tone setting unit 501 outputs the set gain value to the FIR filter circuit 51.

[0145] Multiple FIR filters 511-518 are filters corresponding to the region-specific audio signals SA1-SA8, respectively. The region-specific audio signals SA1-SA8 are input to FIR filters 511-518. For example, as Figure 10 As shown, the area-specific audio signal SA1 is input to FIR filter 511, the area-specific audio signal SA2 is input to FIR filter 512, the area-specific audio signal SA3 is input to FIR filter 513, the area-specific audio signal SA4 is input to FIR filter 514, the area-specific audio signal SA5 is input to FIR filter 515, the area-specific audio signal SA6 is input to FIR filter 516, the area-specific audio signal SA7 is input to FIR filter 517, and the area-specific audio signal SA8 is input to FIR filter 518.

[0146] Multiple FIR filters 511-518 have the same number of taps. For example, multiple FIR filters 511-518 have 16,000 taps. Furthermore, this number of taps is just an example and can be set based on the resource conditions of the sound signal processing device 10, the desired accuracy of the timbre of the initial reflected sound to be reproduced, etc.

[0147] Multiple FIR filters 511-518 perform filtering (convolution) on multiple region-specific sound signals SA1-SA8 using filter coefficients and gain values ​​set by the timbre setting unit 501. As described above, the multiple FIR filters 511-518 generate filtered region-specific sound signals SA1f-SA8f. For example, FIR filter 511 performs filtering (convolution) on the region-specific sound signal SA1 using filter coefficients and gain values ​​set by the timbre setting unit 501 to generate a filtered region-specific sound signal SA1f. Similarly, multiple FIR filters 512-518 generate filtered region-specific sound signals SA2f-SA8f separately based on the region-specific sound signals SA2-SA8.

[0148] Multiple FIR filters 511-518 output the filtered, region-specific audio signals SA1f-SA8f to multiple LDtap 521-528. For example, FIR filter 511 outputs the filtered, region-specific audio signal SA1f to LDtap 521. Similarly, multiple FIR filters 512-518 output the filtered, region-specific audio signals SA2f-SA8f to multiple LDtap 522-528.

[0149] Furthermore, the timbre specification information is not limited to the emphasis information of the vocal range, but also includes specification information that sets the waveform of the initial reflected tone to the characteristics desired by the user. By using the timbre specification information as described above, the sound signal processing device 10 can more diversely realize initial reflected tones with timbre corresponding to the user's preferences.

[0150] [Virtual sound source settings and LDtap settings]

[0151] The virtual sound source setting unit 502 sets the virtual sound source based on the position coordinates of the sound receiving point in the playback space and the geometric shape of the virtual space.

[0152] Figure 12 This is a flowchart illustrating an example of the virtual sound source setting process. The virtual sound source setting unit 502 obtains the position coordinates of the sound receiving point in the playback space (S131). For example, the virtual sound source setting unit 502 obtains the position coordinates of the sound receiving point in the playback space through user operation input, position detection by a position detection sensor, etc.

[0153] The virtual sound source setting unit 502 obtains the geometric shape of the virtual space (S132). For example, the virtual sound source setting unit 502 obtains the geometric shape of the virtual space through user operation input, etc. The geometric shape of the virtual space includes a coordinate set, etc., representing the shape of the walls arranged in the virtual space.

[0154] The virtual sound source setting unit 502 is connected to the GUI 100. If the user selects a desired operation 112 from multiple operation 112s, the GUI 100 reads and obtains the geometry of the virtual space associated with that operation 112. In addition, if the user adjusts the room size (size of the playback space) using the knob 1131, the GUI 100 obtains the adjustment value of the room size.

[0155] The virtual sound source setting unit 502 obtains the position coordinates of the geometric shape of the virtual space with the set room size based on the settings obtained by the GUI 100 in the manner described above. Additionally, the virtual sound source setting unit 502 obtains the position coordinates of the sound source SS and the position coordinates of the pickup point RP (the center of the room (the center of the playback space)). Using this information, the virtual sound source setting unit 502 sets the virtual sound source in the following manner: The virtual sound source setting unit 502 aligns the coordinate system of the playback space with the coordinate system of the virtual space. The virtual sound source setting unit 502 uses the position coordinates of the pickup point in the playback space and the geometric shape of the virtual space, and sets the virtual sound source using the aforementioned... Figure 4 (A) Figure 4 (B) The position coordinates of the virtual sound source in the playback space are set according to the concept (S133).

[0156] Figure 13 (A) Figure 13 (B) is a diagram showing the configuration examples of various virtual sound sources with different geometric shapes. Figure 13 (A) is the virtual wall IWL of a quadrilateral. Figure 13 (B) is the virtual wall IWLh of the hexagon.

[0157] As described above, if the geometry of the virtual space is different, then even if the position coordinates of the sound source SSa and the position coordinates of the receiver RP remain unchanged, the positional relationships between the sound source SSa and the receiver RP and the virtual wall IWL, and the positional relationships between the sound source SSa and the receiver RP and the virtual wall IWLh, will also be different. As described above, in Figure 13 The positions of the virtual sound sources IS1a, IS2a, and IS3a are set under condition (A), and in... Figure 13 The virtual sound source positions IS1ah, IS2ah, and IS3ah set in (B) are different.

[0158] Figure 14 (A) Figure 14 (B) and Figure 14 (C) is a diagram representing a setting example of a virtual sound source. Figure 14 (A) Figure 14 (B) Figure 14 (C) is a diagram representing the planar changes of the virtual sound source. Figure 14 (B) shows relative to Figure 14(A) The sound source SSa is in the same position relative to the reference point (receiving point RP), but the size of the virtual space is different. Figure 14 (C) shows relative to Figure 14 (A) The virtual space has the same size, but the positional relationship between the reference point of the virtual space and the reference point (receiver point) of the playback space has changed (the center of the playback space has changed).

[0159] According to Figure 14 (A) and Figure 14 As can be seen from the comparison results in (B), the size of the virtual space in the playback space ( Figure 14 (A) uses virtual wall IWL to record, Figure 14 (B) uses a virtual wall IWLc, which differs from the description therein, thus the distance and positional relationship between the original sound source SSa, which serves as the virtual sound source, and the virtual wall are different. As mentioned above, in Figure 14 The positions of the virtual sound sources IS1a, IS2a, and IS3a are set under condition (A), and in... Figure 14 In case (B), the positions of the virtual sound sources IS1c, IS2c, and IS3c are different.

[0160] In addition, as according to Figure 14 (A) and Figure 14 The comparison results in (C) show that the positional relationship between the reference point and the receiver RP in the virtual space changes, thereby shifting the position of the virtual sound source in the playback space (the position of the virtual sound source relative to the receiver RP and the speaker). As mentioned above, in Figure 14 The positions of the virtual sound sources IS1a, IS2a, and IS3a are set under condition (A), and in... Figure 14 In case (C), the positions of the virtual sound sources IS1as, IS2as, and IS3as are different.

[0161] Figure 15 (A) Figure 15 (B) Figure 15 (C) is a diagram representing a setting example of a virtual sound source. Figure 15 (A) Figure 15 (B) Figure 15 (C) is a diagram showing the change in the position of the virtual sound source in the height direction.

[0162] exist Figure 15 (A) and Figure 15 In (B), the ceiling height is different. That is, Figure 15 (A) shows the distance (height) of the virtual wall IWL from the floor virtual wall IWFL to the ceiling virtual wall IWCL. Figure 15(B) shows a different distance (height) from the virtual wall IWFL on the floor to the virtual wall IWCLL on the ceiling for the virtual wall WILL.

[0163] According to Figure 15 (A) and Figure 15 As the comparison results in (B) show, the different ceiling heights result in different distances and positional relationships between the original sound source (which acts as a virtual sound source) and the virtual walls IWCL and IWCLL of the ceiling. As mentioned above, in Figure 15 (A) The position of the virtual sound source IS1Ca set and in Figure 15 The position of the virtual sound source IS1CaL is different in case (B).

[0164] exist Figure 15 (A) and Figure 15 In (C), the shape of the ceiling is different. That is, Figure 15 (A) shows the shape of the virtual wall IWCL ceiling and Figure 15 The shape of the virtual wall IWCLx on the ceiling of the virtual wall IWLx shown in (C) is different.

[0165] According to Figure 15 (A) and Figure 15 As can be seen from the comparison results in (C), the different shapes of the ceilings result in different positional relationships between the original sound source (which acts as a virtual sound source) and the virtual walls IWCL and IWCLx of the ceiling. As mentioned above, in Figure 15 (A) The position of the virtual sound source IS1Ca set and in Figure 15 The location of the virtual sound source IS1Cax is different in case (C).

[0166] As described above, the virtual sound source setting unit 502 can optimally set the position of the virtual sound source in the playback space in accordance with the geometric shape of the virtual space, the positional relationship between the playback space and the virtual space. As a result, the sound signal processing device 10 can clearly locate the sound image of the initial reflected sound in accordance with the position coordinates of the speaker in the playback space, the geometric shape of the virtual space, and the positional relationship between the playback space and the virtual space.

[0167] The virtual sound source setting unit 502 outputs the position coordinates of the virtual sound sources set for multiple areas Area1 to Area8 to the output speaker setting unit 5201 of the LDtap circuit 52.

[0168] The output speaker setting unit 5201 sets the virtual sound source IS assigned to each speaker based on the position coordinates of the virtual sound source IS, the position coordinates of the receiving point RP, and the position coordinates of the multiple speakers SP1-SP64. Figure 16 This is a flowchart illustrating the process of assigning virtual sound sources to loudspeakers.

[0169] The output speaker setting unit 5201 obtains the position coordinates of the virtual sound source from the virtual sound source setting unit 502 (S141). The output speaker setting unit 5201 obtains the position coordinates of the receiving point in the playback space, for example, through operation input from the user (S142). The output speaker setting unit 5201 obtains the position coordinates of multiple speakers SP1-SP64, for example, through operation input from the user (S143).

[0170] The output speaker setting unit 5201 sets the area of ​​responsibility of the virtual sound source of each speaker (S144) according to the positional relationship between the sound receiving point RP of the playback space and the multiple speakers SP1-SP64.

[0171] More specifically, the output speaker setting unit 5201 sets the area of ​​responsibility for the virtual sound source of each speaker in the following manner. Figure 17 (A) Figure 17 (B) is a diagram illustrating the concept of assigning a virtual sound source to a loudspeaker. Figure 17 (A) illustrates the concept of using the allocation of azimuth angle φ. Figure 17 (B) illustrates the concept of using the pitch angle θ for allocation. Furthermore, while the following explanation uses speaker SP1 as an example, the output speaker setting unit 5201 also sets the responsible area for other speakers SP2-SP64 using the same method.

[0172] The output speaker setting unit 5201 uses the position coordinates of the receiver point RP and the speaker SP1 to set the position coordinates of the straight line passing through the receiver point RP and the speaker SP1. Figure 17 (A) is set using the dashed line. For example... Figure 17 As shown in (A), the output speaker setting unit 5201 sets the output speaker relative to the straight line ( Figure 17 (A) The dashed line indicates that the azimuth angle φ extending from the receiver point RP to the speaker SP1 on the plane is set. The azimuth angle φ is the angle in the horizontal direction relative to the straight line passing through the receiver point RP and the speaker SP1. Additionally, as... Figure 17 As shown in (B), the output speaker setting unit 5201 sets the line relative to the above-mentioned line ( Figure 17 (B) The pitch angle θ is set by extending in a vertical direction orthogonal to the plane (the dashed line in B). The pitch angle θ is the angle in the vertical direction (orthogonal to the horizontal direction) relative to the straight line passing through the receiver point RP and the speaker SP1.

[0173] The output speaker setting unit 5201 sets the space that is closer to the speaker SP1 than the boundary (the boundary surface that determines the horizontal region and the boundary surface that determines the vertical region) determined by the azimuth angle φ and the elevation angle θ as the responsible area RGSP1 of the speaker SP1.

[0174] The output speaker setting unit 5201 acquires multiple virtual sound sources IS (in Figure 17 The position coordinates of multiple virtual sound sources (ISa-ISg) in the case of [the situation].

[0175] The output speaker setting unit 5201 uses the position coordinates of multiple virtual sound sources ISa-ISg and the coordinates representing the responsible region RGSP1 to determine whether the multiple virtual sound sources ISa-ISg are within the responsible region RGSP1. This determination can be achieved using the same method as the grouping of sound sources into regions described above.

[0176] The output speaker setting unit 5201 performs this determination process, thereby, for example, in... Figure 14 A, Figure 14 B. Figure 14 In the case shown in C, it is determined that multiple virtual sound sources ISa, ISb, ISc, and ISd are located within the responsible region RGSP1, while multiple virtual sound sources ISe, ISf, and ISg are located outside the responsible region RGSP1.

[0177] The output speaker setting unit 5201 assigns multiple virtual sound sources ISa, ISb, ISc, ISd that are determined to be within the responsible area RGSP1 to the speaker SP1 (S145).

[0178] The output speaker setting unit 5201 outputs the allocation information of multiple virtual sound sources for multiple speakers SP1-SP64 to the coefficient setting unit 5202. At this time, the output speaker setting unit 5201 outputs the position coordinates of the receiving point RP, the position coordinates of multiple speakers SP1-SP64, the position coordinates of multiple virtual sound sources, and the allocation information to the coefficient setting unit 5202.

[0179] Furthermore, the azimuth angle φ is, for example, 60°, and the elevation angle θ is, for example, 45°. These azimuth angles φ and elevation angles θ are just examples; they can also be set and adjusted through user input.

[0180] The coefficient setting unit 5202 sets the tap coefficients assigned to LDtap521-528 using the distances between the receiver point RP and the multiple loudspeakers SP1-SP64, and the distance between the receiver point RP and the virtual sound source IS. The tap coefficients assigned to LDtap521-528 are the gain value and delay amount of LDtap521-528.

[0181] Figure 18 This is a flowchart representing the coefficient setting process of LDtap. Figure 19 (A) Figure 19 (B) is a diagram used to illustrate the concept of coefficient setting.

[0182] The coefficient setting unit 5202 uses the position coordinates of the receiver point RP and the position coordinates of the multiple speakers SP1-SP64 to calculate the distance (speaker distance) between the receiver point RP and the multiple speakers SP1-SP64 (S151).

[0183] The coefficient setting unit 5202 calculates the distance (virtual sound source distance) between the receiving point RP and the multiple virtual sound sources IS (S152).

[0184] The coefficient setting unit 5202 compares the distance between the speakers SP1-SP64 and the distance between the virtual sound sources IS, which are respectively assigned to these speakers SP1-SP64 (S153). For example, if it is Figure 17 In example (A), the distance between the speaker SP1 and the distance between the virtual sound sources ISa, ISb, ISc, and ISd are compared.

[0185] If the speaker distance is less than or equal to the virtual sound source distance (S153: YES), the coefficient setting unit 5202 directly uses the virtual sound source distance to set the tap coefficient (S154).

[0186] For example, in such Figure 19 In the case shown in (A), the virtual sound source Isa is farther away from the receiver point RP than the speaker SP1, and the virtual sound source distance Lia between the receiver point RP and the virtual sound source Isa is greater than the speaker distance Ls1 between the receiver point RP and the speaker SP1.

[0187] In this case, the coefficient setting unit 5202 sets the tap coefficient using the distance Da1 between the virtual sound source Isa and the speaker SP1. Specifically, the coefficient setting unit 5202 sets the gain value and delay amount for the virtual sound source Isa based on the distance Da1. For the coefficient setting unit 5202, the larger the distance Da1, the smaller the gain value is set, and the larger the distance Da1, the larger the delay amount is set.

[0188] If the distance to the speaker is greater than the distance to the virtual sound source (S153: NO), the coefficient setting unit 5202 determines whether to play the virtual sound source. In other words, the coefficient setting unit 5202 determines whether to play the virtual sound source that is closer to the receiving point than the speaker (S155).

[0189] If a virtual sound source closer to the receiver than the speaker is played (S155: YES), the coefficient setting unit 5202 moves the position of the virtual sound source (S156). More specifically, the coefficient setting unit 5202 moves the position of the virtual sound source from the side closer to the receiver than the speaker to a position farther away from the receiver than the speaker. At this time, the coefficient setting unit 5202 uses the distance difference between the virtual sound source and the speaker to move the position of the virtual sound source. The coefficient setting unit 5202 uses the position coordinates of the moved virtual sound source to set the tap coefficient (S157).

[0190] For example, in such Figure 19 In the case shown in (B), the virtual sound source ISd is closer to the receiver point RP than the speaker SP1, and the virtual sound source distance Lid between the receiver point RP and the virtual sound source ISd is less than the speaker distance Ls1 between the receiver point RP and the speaker SP1.

[0191] In this case, the coefficient setting unit 5202 moves the virtual sound source ISd using the distance difference Dd between the virtual sound source distance Lid and the speaker distance Ls1. More specifically, the coefficient setting unit 5202 moves the virtual sound source ISd to a position on the straight line passing through the receiver point RP and the speaker SP1, with a distance difference of Dd on the opposite side of the receiver point RP, using the speaker SP1 as a reference. Furthermore, the coefficient setting unit 5202 uses this distance difference Dd to set the tap coefficient. Specifically, the coefficient setting unit 5202 sets the gain value and delay amount relative to the virtual sound source ISd based on the distance difference Dd. For the coefficient setting unit 5202, the larger the distance difference Dd, the smaller the gain value is set, and the larger the distance difference Dd, the larger the delay amount is set. Furthermore, conceptually, the virtual sound source is moved as described above, but as a tap coefficient setting process, the coefficient setting unit 5202 only needs to set the tap coefficient based on the distance between the speaker distance and the virtual sound source distance.

[0192] That is, the coefficient setting unit 5202 only moves the virtual sound source located between the receiver point and the speaker. Regarding this, it is preferable not to move virtual sound sources that are further out of the receiver point than the speaker, but this also includes cases where such outer virtual sound sources move within a predetermined range. For example, even if such outer virtual sound sources move, it is acceptable as long as the distance between the outer virtual sound source and the speaker is within a predetermined range. The predetermined range means that the change in the initial reflected sound control signal caused by the movement will not cause discomfort to the viewer. If a virtual sound source closer to the receiver point than the speaker is not played (S155: NO), the coefficient setting unit 5202 does not set a tap coefficient for that virtual sound source.

[0193] The coefficient setting unit 5202 sets the tap coefficients for each speaker SP1-SP64 to multiple LDtaps. More specifically, the coefficient setting unit 5202 sets the tap coefficients for each speaker SP1-SP64 to LDtap 521 based on the virtual sound source positions set in Area 1. Similarly, the coefficient setting unit 5202 sets the tap coefficients of the virtual sound sources assigned to each speaker SP1-SP64 to LDtap 522-528 based on the virtual sound source positions set in multiple areas Area 2-Area 8 respectively.

[0194] Multiple LDtap 521-528s, corresponding to the set tap coefficients, perform gain processing and delay processing on the filtered, region-specific audio signals SA1f-SA8f and output them to the addition processing unit 53. More specifically, the tap coefficients are set as described above, corresponding to the virtual sound source positions of the multiple regions and the combination of each speaker. Therefore, the multiple LDtap 521-528s set tap coefficients for each speaker based on the virtual sound source assigned to that speaker. The multiple LDtap 521-528s perform gain processing and delay processing on the filtered, region-specific audio signals SA1f-SA8f for each speaker. The multiple LDtap 521-528s output the signal with the gain and delay processing applied to each speaker.

[0195] For example, when virtual sound sources ISa, ISb, ISc, and ISd are assigned to the speaker SP1, LDtap 521 performs gain processing and delay processing on the filtered, region-specific sound signal SA1f based on the tap coefficients (gain and delay) of the virtual sound sources ISa, ISb, ISc, and ISd. Furthermore, LDtap 521 outputs this signal to the addition processing unit 53 as the speaker SP1. Multiple LDtaps 522-528 perform the above processing on the virtual sound sources with set tap coefficients.

[0196] The adder processing unit 53 adds the signals processed by the LDtap of each of the multiple LDtaps 521-528 for each of the multiple speakers SP1-SP64 separately. The adder processing unit 53 outputs these added signals as initial reflection control signals ER1-ER64 for each of the multiple speakers SP1-SP64 to the adder 80.

[0197] By performing the processing described above, the initial reflection tone control signal generation unit 50 is able to generate an initial reflection tone control signal having the following characteristics.

[0198] Figure 20 (A) Figure 20(B) is a waveform diagram illustrating an example of the relationship between the shape of the virtual space and the components of the initial reflection tone control signal implemented by LDtap. Figure 20 (A) shows the case where the virtual space has a large shape. Figure 20 (B) illustrates the case where the virtual space has a small shape. Furthermore, Figure 20 (A) Figure 20 (B) shows an example of the components of the initial reflection control signal when multiple virtual sound sources are set for a single loudspeaker.

[0199] Without changing the positional relationship between the playback space and the virtual space, and without changing the positions of the sound pickup points and speakers, if the virtual space is large, the distribution of the virtual sound source will extend over a wider area compared to when the virtual space is small. Therefore, as... Figure 20 (A) Figure 20 As shown in (B), the virtual space has a large shape, and the components set in LDtap521-528 are easy to become smaller, and the distribution range on the time axis also becomes wider.

[0200] As described above, by performing the above processing, the initial reflected sound control signal generation unit 50 can set the optimal tap coefficient according to the shape of the virtual space.

[0201] Furthermore, even if the positional relationship between the virtual space and the playback space changes, the speaker position changes, or the sound pickup point changes, the initial reflection control signal generation unit 50 can set the optimal tap coefficient accordingly, just as the shape of the virtual space changes.

[0202] At this time, the multiple sound sources OBJ1-OBJ96 are optimally allocated to the multiple loudspeakers SP1-SP64 through grouping based on multiple regions Area1-Area8. Furthermore, the multiple virtual sound sources are optimally positioned relative to the aforementioned multiple loudspeakers SP1-SP64. Therefore, for the sound signal processing apparatus 10, even if there are changes in the relationship between the virtual space and the playback space, changes in the position of the receiving point RP, changes in the positions of the multiple loudspeakers SP1-SP64, and changes in the positions of the sound sources OBJ1-OBJ96, the sound image localization based on the initial reflected sound can be clearly achieved in accordance with these changes.

[0203] Furthermore, in the above structure, even if the virtual sound source IS is closer to the receiving point RP than the speaker SP, the initial reflection control signal generation unit 50 can approximately reproduce the components of the initial reflection control signal obtained based on the virtual sound source IS. Therefore, for example, when the number of virtual sound sources is less than the initial reflection control signal, the initial reflection control signal generation unit 50 can utilize a virtual sound source that is closer to the receiving point RP than the speaker SP. In this case, the initial reflection control signal generation unit 50, as described above, uses the distance difference between the virtual sound source IS and the speaker SP to reposition the virtual sound source outside the speaker. As described above, the initial reflection control signal generation unit 50 can suppress the discomfort of the initial reflection sound caused by moving the position of the virtual sound source.

[0204] Furthermore, in the above structure, the initial reflection control signal generation unit 50 can also set the virtual sound source IS to the position of the speaker SP when the virtual sound source IS is located closer to the receiving point RP than the speaker SP. As described above, the initial reflection control signal generation unit 50 can reduce the processing load of moving the virtual sound source IS.

[0205] Furthermore, in the above structure, the initial reflection control signal generation unit 50 can omit the virtual sound source IS from generating the initial reflection control signal when the virtual sound source IS is positioned closer to the receiving point RP than the loudspeaker SP. As described above, the initial reflection control signal generation unit 50 is exempt from the processing load of moving the virtual sound source IS, thus reducing the processing load for generating the initial reflection control signal.

[0206] Furthermore, in the above structure, the initial reflection control signal generation unit 50 sets the components of the initial reflection control signal obtained based on the virtual sound source and adjusts the timbre using FIR filters 511-518. FIR filters 511-518 have the aforementioned number of taps (e.g., 16,000 taps), which is more than the number of taps in LDtap 521-528. Additionally, the time interval between taps in FIR filters 511-518 (depending on the sampling frequency) is shorter than the time interval between taps in LDtap 521-528 (depending on the configuration of the virtual sound source). Therefore, the components of the initial reflection control signal generated by FIR filters 511-518 are more densely arranged on the time axis compared to the components of the initial reflection control signal generated by LDtap 521-528. In other words, the time resolution (temporal resolution) of FIR filters 511-518 is higher than that of LDtap 521-528, and the number of components per unit time is greater.

[0207] Furthermore, the initial reflection control signal generation unit 50 multiplies the processing of FIR filters 511-518 with LDtap 521-528. Therefore, the initial reflection control signal generation unit 50 is able to generate initial reflection control signals ER1-ER64 with high resolution and more diverse timbres on the time axis. Figure 21 This is a diagram showing an overview of the waveform of the initial reflected tone control signal generated by the initial reflected tone control signal generation unit 50.

[0208] like Figure 21 As shown, the initial reflection control signal generation unit 50 can generate an initial reflection control signal that retains the initial reflection components obtained based on the virtual sound source and can handle higher resolution and more diverse timbres. That is, the sound signal processing device 10 can achieve an initial reflection that ensures clear sound image localization using the initial reflection from the virtual sound source and that matches the user's preferences.

[0209] Furthermore, because FIR filters have high resolution, in cases such as short pulses from a sound source, the initial reflected tone control signal obtained solely from the LDtap can sometimes become coarse, resulting in an unnatural timbre. However, through the aforementioned structure and processing, the sound signal processing device 10 can suppress such coarsening and unnatural timbre in the initial reflected tone.

[0210] Furthermore, in the above structure, the initial reflection control signal generation unit 50 sets a responsible area for the virtual sound source IS for each speaker SP, and does not assign virtual sound sources IS outside that area to that speaker SP. As described above, the initial reflection control signal generation unit 50 can suppress the over-generation of initial reflection components. Therefore, the sound signal processing device 10 can suppress the over-generation of initial reflections and achieve a more natural initial reflection that matches the virtual space.

[0211] [Generation of Echo Control Signal]

[0212] Figure 22 This is a functional block diagram illustrating an example of the structure of the echo control signal generation unit 70. Figure 23 This is a flowchart illustrating an example of the generation and processing of the echo control signal.

[0213] like Figure 22 As shown, the reverberation control signal generation unit 70 includes a PEQ 71, an FIR filter circuit 72, a distributor 73, a reverberation zone setting unit 701, a filter coefficient setting unit 702, a reverberation playback speaker setting unit 703, and an operation unit 700. The FIR filter circuit 72 includes multiple FIR filters 721-728.

[0214] The reverberation zone setting unit 701 sets multiple reverberation zones Arr1-Arr8 for the playback space. More specifically, the reverberation zone setting unit 701 sets the reverberation zones by dividing the playback space into multiple reverberation zones Arr1-Arr8 within a plane, for example, using the center point Psr of the playback space as a reference (see below). Figure 25 ).

[0215] The reverberation zone setting unit 701 outputs coordinate information representing multiple reverberation zones Arr1-Arr8 to the filter coefficient setting unit 702 and the reverberation playback speaker setting unit 703.

[0216] The filter coefficient setting unit 702 sets the filter coefficient for reverberation through user operation, etc. The filter coefficient for reverberation is set, for example, by measuring the impulse response of different spaces (virtual spaces) reproduced in the playback space. Alternatively, the filter coefficient for reverberation can be approximately set using the geometry of the virtual space, the material of the walls, etc. In this case, the filter coefficient setting unit 702 sets the filter coefficient for each reverberation region Arr1-Arr8 using the coordinate information of each reverberation region Arr1-Arr8.

[0217] The filter coefficient setting unit 702 receives inputs such as the volume and surface area of ​​the virtual space through user operations. Based on parameters such as the volume and surface area of ​​the virtual space, the filter coefficient setting unit 702 sets a fade-in function for the filter coefficient.

[0218] More specifically, the filter coefficient setting unit 702 calculates the mean free path ρ using the volume V of the virtual space and the surface area S of the virtual space. The formula for calculating the mean free path ρ is ρ = 4V / S. The mean free path refers to the average distance that sound travels in an enclosed space from its reflection at a wall to its next reflection. By dividing the mean free path by the speed of sound c0, the average time required for sound to reflect from a wall to its next reflection can be calculated.

[0219] The filter coefficient setting unit 702 sets the connection timing tc based on the average free path ρ. Figure 23 (S231). Specifically, the filter coefficient setting unit 702 sets the connection timing tc using the mean free path ρ, the speed of sound c0, and the number of reflections n. The formula for calculating the connection timing tc is tc = ρ × n / c0.

[0220] As can be seen from the calculation formula, the connection timing tc corresponds to the average time required for n reflections in the virtual space, and in the case of reproducing the initial reflected sound n times, it corresponds to the moment when it begins to transform into an echo. In other words, the connection timing tc corresponds to the timing at which the components of the initial reflected sound control signal obtained based on the initial reflected sound control signal generation unit 50 described above disappear.

[0221] By performing this processing, the filter coefficient setting unit 702 can optimally set the connection timing tc of the initial reflected sound and the echo sound according to the geometric shape of the virtual space.

[0222] The filter coefficient setting unit 702 uses the connection timing tc to set the fade-in function according to the following formula ( Figure 23 (S232).

[0223] Formula 1

[0224]

[0225] Furthermore, in this formula, t is the elapsed time from the occurrence of the direct sound, and K is set according to the following formula.

[0226] Formula 2

[0227]

[0228] Furthermore, in this formula, G REV This is the gain value of the reverberation at time t=0, which can be set by the user. For example, the reverberation time is usually the time required to decay to -60dB, so it can be set to G. REV = -60dB, etc.

[0229] The filter coefficient setting unit 702 sets the reverb filter coefficient based on the filter coefficient and the fade-in function fin. Figure 23 (S233) and output to multiple FIR filters 721-728.

[0230] The reverb generation signal Sr output from mixer 60 is input to PEQ 71. PEQ 71 performs prescribed signal processing on the reverb generation signal Sr and outputs it to multiple FIR filters 721-728.

[0231] By processing the signal using the PEQ 71, the level (signal magnitude) and timbre of the reverberation generation signal Sr can be adjusted. For example, the PEQ 71 can adjust the level (signal magnitude) of the reverberation generation signal Sr by referring to the volume of the initial reflection control signal, so that the volume of the initial reflection and the reverberation are at the same level during the connection timing tc mentioned above. Furthermore, the PEQ 71 can adjust the timbre and other parameters according to user settings.

[0232] Multiple FIR filters 721-728 filter the reverberation generation signal Sr using reverberation filtering coefficients to generate region-specific reverberation control signals REVr1-REVr8. For example, FIR filter 721 performs a convolution operation on the reverberation generation signal Sr using the reverberation filtering coefficients set for region Arr1, thereby generating region-specific reverberation control signal REVr1 for region Arr1. Similarly, FIR filters 722-728 perform convolution operations on the reverberation generation signal Sr using the reverberation filtering coefficients set for regions Arr2-Arr8, respectively, thereby generating region-specific reverberation control signals REVr2-REVr8 for regions Arr2-Arr8. Figure 23 (S234). Multiple FIR filters 721-728 output reverberation control signals REVr1-REVr8, differentiated by region, to distributor 73.

[0233] By setting the fade-in function described above, the reverb control signal becomes as follows: Figure 24 The waveform shown. Figure 24 This is a graph showing examples of the waveforms of the direct tone, initial reflection tone control signal, and reverberation tone control signal. Furthermore, in Figure 24 In the diagram, for simplicity, the reverberation control signal is illustrated using the envelopes of each time component. Additionally, Figure 24 The vertical axis represents dB.

[0234] like Figure 24 As shown, the reverb control signal level gradually increases with the fade-in function from the output timing of the direct tone to the connection timing tc. More specifically, the reverb control signal level is -60 dBFs at the output timing of the direct tone, gradually increasing until the connection timing tc, where it becomes 0 dBFs. This level is set based on the initial reverb control signal connection timing tc.

[0235] exist Figure 24In the example, the aforementioned fade-in function is used, and the signal level increases exponentially with approach connection timing tc. In other words, the fade-in function described above has the opposite characteristics to the attenuation curve of the echo control signal without fade-in processing. Furthermore, the characteristics of the change in the echo control signal level obtained through fade-in processing are not limited to this, and can be set to the characteristics desired by the user by appropriately setting the fade-in function.

[0236] By performing the processing described above, the reverberation control signal generation unit 70 can generate a reverberation control signal that reproduces the reverberation of the virtual space with high precision using FIR filters 721-728. Furthermore, for the reverberation control signal, in the interval where the initial reflection control signal exists, the signal level gradually increases, reaching a peak corresponding to the signal level of the initial reflection control signal at the connection timing tc, and then attenuating.

[0237] As described above, the sound signal processing device 10 can smoothly connect the initial reflection control signal and the reverberation control signal generated by multiple LDtaps to reproduce the distribution of virtual sound sources at multiple sound source locations in the virtual space through the reverberation obtained based on the reverberation control signal. Therefore, the sound output from the sound signal processing device 10 and heard by the user becomes a sound that suppresses the discomfort when connecting from the initial reflection to the reverberation.

[0238] The reverberation sound playback speaker setting unit 703 groups multiple speakers SP1-SP64 into reverberation sound areas Arr1-Arr8.

[0239] More specifically, the reverberation speaker setting unit 703, for example, divides the playback space into multiple reverberation zones Arr1-Arr8 within a plane, using the center point Psr of the playback space as a reference. The reverberation speaker setting unit 703 uses the position coordinates of multiple speakers SP1-SP64 and coordinate information representing the multiple reverberation zones Arr1-Arr8 to group the multiple speakers SP1-SP64 for each of the multiple reverberation zones Arr1-Arr8. This grouping can be achieved using the same method as the method for grouping the sound source OBJ described above.

[0240] Figure 25 This is a diagram illustrating an example of zone settings used for reverberation. In Figure 25 To simplify the explanation and facilitate understanding, multiple speakers SP1-SP14 are shown. For example, the reverberation playback speaker setting unit 703 is shown... Figure 25As shown, the presence of speakers SP6 and SP7 within the reverberation area Arr1 is detected, and speakers SP6 and SP7 are grouped into the reverberation area Arr1. Similarly, the reverberation playback speaker setting unit 703 also groups the other speakers SP1-SP5 and SP8-SP14 into multiple reverberation areas Arr2-Arr8 respectively.

[0241] The reverberation playback speaker setting unit 703 outputs grouping information for multiple speakers SP1-SP64 in multiple reverberation zones Arr2-Arr8 to the distributor 73.

[0242] The distributor 73 uses the grouping information from the reverberation playback speaker setting unit 703 to distribute the area-divided reverberation control signals REVr1-REVr8 to multiple speakers SP1-SP64. Based on the distribution, the distributor 73 outputs the area-divided reverberation control signals REVr1-REVr8 as reverberation control signals REV1-REV48 for each of the multiple speakers SP1-SP64.

[0243] For example, distributor 73 extracts the case where speaker SP6 and speaker SP7 are grouped in region Arr1 based on the grouping information. Distribute distributor 73 allocates the region-specific reverberation control signal REVr1 for region Arr1 to speaker SP6 and speaker SP7. Distributor 73 outputs the region-specific reverberation control signal REVr1 as the reverberation control signal REV6 for speaker SP6. In addition, distributor 73 outputs the region-specific reverberation control signal REVr1 as the reverberation control signal REV7 for speaker SP7.

[0244] By performing the allocation processing of the reverberation control signals REVr1-REVr8 for each region by the distributor 73 as described above, the reverberation control signal generation unit 70 can output the optimal reverberation control signal to each of the multiple speakers SP1-SP64 according to the configuration of the multiple speakers SP1-SP64.

[0245] [Output Adjustment]

[0246] Figure 26 This is a functional block diagram illustrating an example of the structure of the output adjustment unit 90. Figure 27 This is a flowchart illustrating an example of output adjustment processing.

[0247] like Figure 26As shown, the output adjustment unit 90 includes a gain control unit 91, a hysteresis control unit 92, a gain and hysteresis setting unit 901, an operation unit 900, and a display unit 909. The gain control unit 91 has multiple gain control units 9101-9168 corresponding to the multiple speakers SP1-SP64. The hysteresis control unit 92 has multiple hysteresis control units 9201-9264 corresponding to the multiple speakers SP1-SP64.

[0248] The operation unit 900 receives settings for the audio parameters of the playback space via user input. Figure 27 (S321). The acoustic parameters of the playback space are parameters used to reproduce the desired sound field in the playback space.

[0249] At this time, the audio parameters of the playback space are not the gain values ​​or delay values ​​of the multiple speakers SP1-SP64, but rather the weight values ​​representing the weighted distribution of sound in the playback space in a specified direction, and the shape values ​​representing the expansion of sound in the playback space in a specified direction.

[0250] The weight value consists of gain and delay, including weights for the front and back of the playback space, the left and right sides of the playback space, and the up and down directions of the playback space. The shape value consists of gain and delay, including the horizontal shape value.

[0251] Display unit 909 has a GUI. Figure 28 This is a diagram illustrating an example of a GUI used for output adjustment.

[0252] like Figure 28 As shown, the GUI 100A has a setting display window 111, an output status display window 115, and multiple operating components 116. The multiple operating components 116 include a knob 1161 and an adjustment value display window 1162.

[0253] Multiple operating units 116 are operating units for setting weighted volume, shape volume, etc., for setting weight values. The weighted volume operating units 116 include operating units 116 for setting left / right weight, front / back weight, and up / down weight, and operating units 116 for setting gain values ​​and delay amounts. The shape volume operating units 116 include operating units for setting extension, setting gain values, and setting delay amounts.

[0254] The output status display window 115 graphically illustrates the sound expansion and localization achieved by the weight and shape values ​​set by the multiple operating elements 116. Thus, the user can easily recognize the sound expansion and localization set by the multiple operating elements 116 as an image.

[0255] The user uses the GUI 100A of the display unit 909 to set the audio parameters (weight values ​​and delay amounts) they want to reproduce. The operation unit 900 receives the settings made using the GUI 100A. The operation unit 900 outputs the settings (weight values ​​and delay amounts of the audio parameters) to the gain and hysteresis setting unit 901.

[0256] The gain and hysteresis setting unit 901 sets the gain and delay values ​​for multiple speakers SP1-SP64 based on the weight values ​​and delay values ​​of the audio parameters. More specifically, the gain and hysteresis setting unit 901 performs the following processing.

[0257] The gain and hysteresis setting unit 901 obtains the position coordinates of the plurality of speakers SP1-SP64 arranged in the playback space (S322). The position coordinates are represented, for example, in the following coordinate system, which sets the x-axis in the left-right direction of the playback space, the y-axis in the front-back direction of the playback space, and the z-axis in the up-down direction.

[0258] The gain and hysteresis setting unit 901 extracts the maximum and minimum values ​​of the position coordinates of multiple speakers SP1-SP64 in each axial direction (S323).

[0259] The gain and hysteresis setting unit 901 stores coefficient setting formulas. The coefficient setting formulas include, for example, a weighting coefficient setting formula for setting weights in a specified direction in the playback space, and a shape coefficient setting formula for setting weights in a specified direction in the playback space.

[0260] The coefficient setting formula for weighting includes the formula for setting the gain value and the formula for setting the delay amount for weighting. The coefficient setting formula for shape includes the formula for setting the gain value and the formula for setting the delay amount for shape.

[0261] The weighting coefficient setting formulas include the forward and backward direction coefficient setting formula for weighting the forward and backward directions of the playback space, the left and right direction coefficient setting formula for weighting the left and right directions of the playback space, and the up and down direction coefficient setting formula for weighting the up and down directions of the playback space.

[0262] The coefficient settings for the shape include the coefficient settings for the left and right directions of the playback space.

[0263] The coefficient setting formula for the gain value used for weighting is, for example, a linear function obtained by combining the gain value of the set weight value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker of the setting object) for which the gain value is set. It is a formula that determines the gain value in proportion to the difference between the position coordinates and the minimum value of the position coordinates of the speaker of the setting object.

[0264] The formula for setting the delay amount for weighting is, for example, a linear function obtained by combining the delay amount of the set weight value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker of the set object) for which the delay amount is set. It is a formula that determines the delay amount in proportion to the difference between the position coordinates and the minimum value of the position coordinates of the speaker of the set object.

[0265] The coefficient setting formula for the gain value used for the shape is, for example, a linear function obtained by combining the gain value of the set shape value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker of the set object) for which the gain value is set. It is a formula that determines the gain value in proportion to the difference between the position coordinates and the minimum value of the position coordinates of the speaker of the set object.

[0266] The formula for setting the delay amount for the shape is, for example, a linear function obtained by combining the delay amount of the set shape value, the maximum and minimum values ​​of the extracted position coordinates, and the position coordinates of the speaker (the speaker of the set object) for which the delay amount is set. It is a formula that determines the delay amount in proportion to the difference between the position coordinates and the minimum value of the position coordinates of the speaker of the set object.

[0267] The gain and hysteresis setting unit 901 uses the set gain and delay values ​​(audio parameters), the maximum and minimum values ​​of the extracted position coordinates, and the coefficient setting formula to calculate the gain and delay values ​​for each speaker of the set object (S324).

[0268] By using the processing described above, the gain and hysteresis setting unit 901 can automatically calculate and set the gain and delay values ​​of the multiple speakers SP1-SP64 arranged in the playback space without having to manually set them individually.

[0269] The gain and hysteresis setting unit 901 outputs the gain values ​​set for each of the multiple speakers SP1-SP64 to multiple gain control units 9101-9164. The gain and hysteresis setting unit 901 outputs the delay values ​​set for each of the multiple speakers SP1-SP64 to multiple hysteresis control units 9201-9264.

[0270] Each of the multiple gain control units 9101-9164 receives speaker signals Sat1-Sat64 corresponding to multiple speakers SP1-SP64 from the adder 80.

[0271] Multiple gain control units 9101-9164 use preset gain values ​​to control the signal levels of speaker signals Sat1-Sat64 and output the results to multiple hysteresis control units 9201-9264. For example, gain control unit 9101 uses a gain value set for itself to control the signal level of speaker signal Sat1 and outputs the result to hysteresis control unit 9201. Similarly, gain control units 9102-9164 use their respective preset gain values ​​to control the signal levels of speaker signals Sat2-Sat64 and output the results to hysteresis control units 9202-9264 respectively.

[0272] Multiple hysteresis control units 9201-9264, using preset delay amounts, control the signal levels of signals input from multiple gain control units 9101-9164 and output them to multiple speakers SP1-SP64. For example, hysteresis control unit 9201 uses a preset delay amount to control the signal level of signals input from gain control unit 9101 and outputs it to speaker SP1. Similarly, hysteresis control units 9202-9264, using preset delay amounts respectively, control the signal levels of signals input from gain control units 9102-9164 and output them to speakers SP2-SP64 respectively.

[0273] With the structure described above, the sound signal processing device 10 eliminates the need for the user to be proficient in complex settings for multiple speakers. It can easily achieve the desired sound field corresponding to the set acoustic parameters using the initial reflection control signal and the reverberation control signal. As mentioned above, for example, the sound signal processing device 10 can easily achieve a Haas effect sound field for a specified position within the playback space.

[0274] (An example of implementing a sound field based on output control)

[0275] Figure 29 (A) Figure 29 (B) is a diagram showing a setting example where there is positioning and expansion at the rear of the playback space. Figure 29 (A) is a diagram showing an example of the settings for gain and delay. Figure 29 (B) indicates that it is based on Figure 29 (A) is a diagram outlining the weighted average of the sound settings implemented. Furthermore, in Figure 29 (A) Figure 29 (B) shows a configuration with 14 speakers SP1-SP14, as can be easily understood by the simplified explanation.

[0276] exist Figure 29 (A) Figure 29 In the manner shown in (B), as audio parameters, for example, the gain and delay values ​​of the rear end are set. The gain and hysteresis setting unit 901 sets the gain and delay values ​​of the front end to values ​​with the opposite signs of the gain and delay values ​​of the rear end. The gain and hysteresis setting unit 901 calculates the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14.

[0277] The gain and hysteresis setting unit 901 calculates the gain values ​​of the 14 speakers SP1-SP14 using the gain values ​​of the rear and front ends, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and a coefficient setting formula (for gain value setting) for the front and rear directions of the playback space.

[0278] In addition, the gain and hysteresis setting unit 901 calculates the delay of the 14 speakers SP1-SP14 using the delay amounts of the rear and front ends, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and a coefficient setting formula for the front and rear directions (for delay setting) that sets the weighting of the front and rear directions of the playback space.

[0279] Through this processing, the sound signal processing device 10, as... Figure 29 As shown in (A), the audio parameters can be automatically and easily set such that the gain and delay of speakers further back in the playback space are greater, while the gain and delay of speakers further forward are smaller. Therefore, the audio signal processing device 10 can easily achieve a sound field with expandability and sound localization at the rear of the playback space (see reference). Figure 29 (B)).

[0280] Furthermore, while an example in the front-back direction is shown in this description, the sound signal processing device 10 can also achieve a weighted sound field in the left-right direction and the height direction (up-down direction).

[0281] Figure 30 (A) Figure 30 (B) is a diagram illustrating a setting example where there is a horizontal expansion of sound in the playback space. Figure 30 (A) is a diagram showing an example of the settings for gain and delay. Figure 30 (B) indicates that it is based on Figure 30 (A) is a diagram outlining the expansion of sound achieved through its settings. Furthermore, in Figure 30 (A) Figure 30 (B) shows a configuration with 14 speakers SP1-SP14, as can be easily understood from the simplified explanation.

[0282] exist Figure 30 (A) Figure 30 In the manner shown in (B), as an audio parameter, for example, the value of the expanded sound is set (expansion setting value). The gain and hysteresis setting unit 901 calculates the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14.

[0283] The gain and hysteresis setting unit 901 calculates the gain value of the 14 speakers SP1-SP14 using the value after the sound extension is numerated, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and the coefficient setting formula for shape (for gain value setting).

[0284] In addition, the gain and hysteresis setting unit 901 calculates the delay of the 14 speakers SP1-SP14 using the delay amounts of the rear and front ends, the maximum and minimum values ​​of the position coordinates of the 14 speakers SP1-SP14, and the coefficient setting formula for the shape (for delay amount setting).

[0285] Through this processing, the sound signal processing device 10, as... Figure 30 As shown in (A), audio parameters can be easily set such that the gain and delay are greater for speakers closer to the horizontal ends of the playback space, and smaller for speakers closer to the center of the horizontal space. As described above, the audio signal processing device 10 can easily achieve a sound field with horizontal expansion and sound localization within the playback space (see reference). Figure 30 (B)).

[0286] Furthermore, by setting the aforementioned audio parameters, the audio signal processing device 10 can not only achieve weighted scaling in the front-to-back direction, left-to-right direction, and lateral expansion of the playback space, but also weighted scaling and expansion in the vertical direction (up-down direction) of the playback space. For example, Figure 31 It is a diagram that shows the general extent of sound expansion in cases of vertical expansion.

[0287] The sound signal processing device 10 makes the gain and delay of the ceiling-side speaker SPU greater than the gain and delay of the floor-side speakers SPL and SPR. As described above, the sound signal processing device 10 can easily achieve a sound field with greater expansiveness and reverberation localization in the ceiling direction of the playback space (see reference). Figure 31 ).

[0288] In addition, in the above structure, the output adjustment unit 90 outputs output signals So1-So64 to multiple speakers SP1-SP64. However, the sound signal processing device may also perform binaural processing on the output signals So1-So64 before outputting them.

[0289] Figure 32 This is a functional block diagram representing the structure of a sound signal processing device with binaural playback capability. For example... Figure 32 As shown, the audio signal processing device 10A with binaural playback function differs from the audio signal processing device 10 described above in that it has an output adjustment unit 90A, an echo processing unit 97, a selection unit 98, and a binaural processing unit 99.

[0290] The output adjustment unit 90A generates multiple output signals So1-So64 based on the multiple speaker signals Sat1-Sat64 output from the adder 80 using the same processing as the output adjustment unit 90 described above.

[0291] The output adjustment unit 90A is capable of selecting the output target. The selection of the output target is performed, for example, by user input via the GUI described above. More specifically, the GUI displays an operator that can select speaker output and binaural output, and the output target is selected by operating this operator.

[0292] When speaker output is selected, the output adjustment unit 90A outputs multiple output signals So1-So64 to multiple speakers SP1-SP64 respectively (the same processing as output adjustment unit 90). When binaural output is selected, the output adjustment unit 90A outputs multiple output signals So1-So64 to the selection unit 98.

[0293] The reverberation processing unit 97 receives sound signals S1-S96 from multiple sound sources OBJ1-OBJ96. The reverberation processing unit 97 applies initial reflection control signals and reverberation control signals to the multiple sound signals S1-S96 and outputs them to the selection unit 98. The initial reflection control signals for the multiple sound signals S1-S96 are set based on the position coordinates of the multiple sound sources OBJ1-OBJ96. The reverberation processing unit 97 outputs the reverberated sound signals S1'-S96' to the selection unit 98.

[0294] Multiple output signals So1-So64 and multiple reverberation-processed sound signals S1'-S96' are input to the selection unit 98. The selection unit 98 selects the multiple output signals So1-So64 and the reverberation-processed sound signals S1'-S96', for example, by using the operation input from the user through the GUI described above. More specifically, the GUI displays an operation device that can select the sound that has undergone sound processing by the sound signal processing device 10A and the sound that has undergone virtual sound processing based on the position coordinates of the sound sources OBJ1-OBJ96, and selects the output object by operating the operation device.

[0295] When a sound that has undergone sound processing by the sound signal processing device 10A is selected, the selection unit 98 selects multiple output signals So1-So64 and outputs them to the binaural processing unit 99. When a sound that has undergone virtual sound processing based on the position coordinates of sound sources OBJ1-OBJ96 is selected, the selection unit 98 selects multiple reverberation-processed sound signals S1'-S96' and outputs them to the binaural processing unit 99.

[0296] The binaural processing unit 99 performs binaural processing on the input sound signal. More specifically, if multiple output signals So1-So64 are input, the binaural processing unit 99 performs binaural processing on the multiple output signals So1-So64. If multiple reverberation-processed sound signals S1'-S96' are input, the binaural processing unit 99 performs binaural processing on the multiple reverberation-processed sound signals S1'-S96'.

[0297] Furthermore, binauralization is a process that uses a head transfer function, the details of which are known, so a detailed description of binauralization is omitted.

[0298] The binaural processing unit 99 outputs the two-channel audio signal that has undergone binaural processing.

[0299] As described above, users can hear the sound generated by the sound signal processing device 10A and the sound that has undergone virtual reverberation processing based on the position coordinates of sound sources OBJ1-OBJ96 through binaural playback. Therefore, even without physically constructing a playback space, users can easily verify whether the sound processing performed by the sound signal processing device 10A can reproduce the sound of the virtual space using headphones or the like. The sound processing performed by the sound signal processing device 10A includes, for example, the aforementioned grouping of sound sources, setting of the initial reflection control signal, setting of the reverberation control signal, and setting of output control. Furthermore, by performing audiovisual comparison as described above, users can adjust the settings of the aforementioned sound processing to more realistically reproduce the sound of the virtual space.

[0300] Furthermore, binaural playback is not limited to headphones; it can also be achieved through stereo speakers and other means.

[0301] The description of this embodiment is illustrative in all respects and is not restrictive. The scope of the invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the invention includes all equivalents of the claims and all modifications within that scope.

[0302] Explanation of the label

[0303] 10, 10A: Sound signal processing device

[0304] 30: Regional Setting Department

[0305] 40: Grouping Department

[0306] 41: Sound Source Location Detection Department

[0307] 42: Area Determination Department

[0308] 50: Initial Reflection Sound Control Signal Generation Unit

[0309] 51: FIR filter circuit

[0310] 52: LDtap circuit

[0311] 53: Addition Processing Department

[0312] 60: Mixer

[0313] 70: Echo Control Signal Generation Unit

[0314] 71: PEQ

[0315] 72: FIR filter circuit

[0316] 73: Distributor

[0317] 80: Adder

[0318] 90, 90A: Output Adjustment Section

[0319] 91: Gain Control Section

[0320] 92: Lag Control Department

[0321] 97: Echo Processing Department

[0322] 98: Selection Department

[0323] 99: Diphasic Processing Department

[0324] 100, 100A: GUI

[0325] 400: Matrix Mixer

[0326] 500: Operations Department

[0327] 501: Sound Setting Department

[0328] 502: Virtual Sound Source Setting Department

[0329] 511-518: FIR Filters

[0330] 521-528: LDtap

[0331] 700: Operations Department

[0332] 701: Echo Zone Setting Unit

[0333] 702: Filter Coefficient Setting Section

[0334] 703: Reverberation Sound Playback Speaker Setting Section

[0335] 721-728: FIR Filters

[0336] 900: Operations Department

[0337] 901: Gain and Hysteresis Setting Section

[0338] 909: Display Unit

[0339] 5201: Output speaker setting unit

[0340] 5202: Coefficient Setting Section

[0341] 9101-9164: Gain Control Section

[0342] 9201-9264: Lag Control Department.

Claims

1. A sound signal processing method, wherein, Obtain the sound signal from the sound source. The first filtering process is implemented to generate a virtual sound source in the virtual space. A second filtering process is implemented to adjust the timbre of the initial reflected tone. The system outputs an initial reflection control signal generated from the sound signal after the first and second filtering processes, and outputs the initial reflection control signal to a speaker to simulate the initial reflection of virtual space in the playback space. The second filtering process is applied to the audio signal, and the first filtering process is applied to the audio signal after the second filtering process, or... The first filtering process is applied to the audio signal, and the second filtering process is applied to the audio signal after the first filtering process. The first filtering process uses the geometry of the virtual space and the position of the virtual sound source to set the gain and delay values ​​for the input sound signal. The second filtering process is an FIR filter that performs a convolution operation on the input audio signal. The number of components per unit time generated by the second filtering process is greater than the number of components per unit time generated by the first filtering process.

2. The sound signal processing method according to claim 1, wherein, The time resolution of the second filtering process is higher than that of the first filtering process.

3. The sound signal processing method according to claim 1 or 2, wherein, The components of the initial reflected sound control signal obtained based on the first filtering process are different from the components of the initial reflected sound control signal obtained based on the second filtering process on the time axis.

4. The sound signal processing method according to claim 1 or 2, wherein, The second filtering process can set filtering characteristics including at least one of sampling frequency, filter length, and filter coefficient.

5. The sound signal processing method according to claim 4, wherein, The filtering characteristics of the second filtering process can be set via external operation input.

6. A sound signal processing device, comprising: The sound signal acquisition unit acquires the sound signal from the sound source; The first filtering processing unit performs the first filtering processing to generate the virtual sound source of the virtual space; The second filtering processing unit performs a second filtering process to adjust the timbre of the initial reflected tone; and The initial reflection control signal output unit outputs an initial reflection control signal generated by the sound signal after the first and second filtering processes, and outputs the initial reflection control signal to a speaker to simulate the initial reflection of virtual space in the playback space. The second filtering unit performs the second filtering process on the audio signal, and the first filtering unit performs the first filtering process on the audio signal after the second filtering process has been performed, or... The first filtering unit performs the first filtering process on the audio signal, and the second filtering unit performs the second filtering process on the audio signal after the first filtering process. The first filtering processing unit uses the geometry of the virtual space and the position of the virtual sound source to set the gain and delay values ​​for the input sound signal. The second filtering processing unit is an FIR filter that performs convolution operations on the input audio signal. The number of components per unit time generated by the second filtering process is greater than the number of components per unit time generated by the first filtering process.

7. The sound signal processing apparatus according to claim 6, wherein, The time resolution of the second filtering process is higher than that of the first filtering process.

8. The sound signal processing apparatus according to claim 6 or 7, wherein, The components of the initial reflected sound control signal obtained by the first filtering process based on the first filtering processing unit are different from the components of the initial reflected sound control signal obtained by the second filtering process based on the second filtering processing unit on the time axis.

9. The sound signal processing apparatus according to claim 6 or 7, wherein, The second filtering processing unit can set the filtering characteristics including at least one of the sampling frequency, filtering length, and filtering coefficient.

10. The sound signal processing apparatus according to claim 9, wherein, It has an operation unit that receives operation input of the filtering characteristics of the second filtering process.

11. A recording medium, which is a non-volatile, computer-readable recording medium, having recorded a program that causes a computer to perform the following processes: Obtain the sound signal from the sound source. The first filtering process is implemented to generate a virtual sound source in the virtual space. A second filtering process is implemented to adjust the timbre of the initial reflected tone. The system outputs an initial reflection control signal generated from the sound signal after the first and second filtering processes, and outputs the initial reflection control signal to a speaker to simulate the initial reflection of virtual space in the playback space. The second filtering process is applied to the audio signal, and the first filtering process is applied to the audio signal after the second filtering process, or... The first filtering process is applied to the audio signal, and the second filtering process is applied to the audio signal after the first filtering process. The first filtering process uses the geometry of the virtual space and the position of the virtual sound source to set the gain and delay values ​​for the input sound signal. The second filtering process is an FIR filter that performs a convolution operation on the input audio signal. The number of components per unit time generated by the second filtering process is greater than the number of components per unit time generated by the first filtering process.