Sound signal processing method, sound signal processing device, and recording medium
By classifying and moving the virtual sound source to the speaker position, the problem of unclear sound and image positioning in the existing audio system is solved, and the clear reproduction and spatial expansion of sound and image in the virtual space is achieved, simulating the realistic effect of the initial reflected sound.
Patent Information
- Application Number
- CN202510393554.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2021-03-19
- Filing Date
- 2022-03-14
- Publication Date
- 2025-07-04
AI Technical Summary
Existing audio systems cannot clearly reproduce the sound and image positioning in the virtual space in a space where speakers are provided.
The virtual sound source of reflected sound in the target audio space is classified into a first virtual sound source and a second virtual sound source, the first virtual sound source is located between the speaker and the radio point, and the second virtual sound source is located outside the speaker, and in the case of the first virtual sound source, its position is moved to a speaker position near the virtual sound source for playback, and the processing is performed by the sound signal processing device.
Clear sound image positioning and rich space expansion in the virtual space are realized, which can realistically simulate the initial reflected sound of the virtual space, eliminate the unnaturalness of the tone, and maintain the smoothness of the sound image positioning when the sound source moves.
Smart Images

Figure CN120264190A_ABST
Abstract
Description
[0001] This application is a divisional application of the Chinese national application No. 202210248168.9 (Sound signal processing method, sound signal processing device, and recording medium) filed on March 14, 2022, the content of which is incorporated herein by reference. Technical Field
[0002] One embodiment of the present invention relates to a sound signal processing method and a sound signal processing device for performing a predetermined process on sound input from a sound source. Background Art
[0003] In an audio system in a space such as a hall, sound image localization with respect to a sound source is performed by speakers arranged in the space.
[0004] For example, the speech processing device described in Patent Document 1 outputs the sound of an audio object (sound source) through two or more speakers near the audio object (sound source). At this time, the speech processing device described in Patent Document 1 calculates the gain of the speech signal output to each speaker using the position information and sound image information of the audio object (sound source).
[0005] Patent Document 1: WO 2013 / 208406
[0006] However, in the above-described conventional structure, it is impossible to clearly reproduce the sound image localization in the virtual space in the space where the speakers are arranged. Summary of the Invention
[0007] Therefore, an object of one embodiment of the present invention is to clearly reproduce the sound image localization in the virtual space.
[0008] The sound signal processing method classifies a virtual sound source representing the reflected sound of a target acoustic space into a first virtual sound source and a second virtual sound source. The first virtual sound source is located between the position of the speaker and the position of the sound collection point, and represents the first sound source of the reflected sound of the target acoustic space. The second virtual sound source is located outside the speaker and represents the second sound source of the reflected sound of the target acoustic space. Only in the case of the first virtual sound source, the position of the first virtual sound source is moved to a position where it can be played using the position of the speaker near the first virtual sound source.
[0009] Advantages of the Invention
[0010] The sound signal processing method can clearly reproduce the sound image localization in the virtual space. Brief Description of the Drawings
[0011] Figure 1 It is a functional block diagram showing the structure of an audio system including a sound signal processing device according to an embodiment of the present invention.
[0012] Figure 2 is a flowchart of the sound signal processing method related to the embodiments of the present invention.
[0013] Figure 3 is a diagram showing the discrete waveform of a sound including a normal direct sound, an initial reflection sound, and a reverberation sound (rear reverberation sound).
[0014] Figure 4 (A), Figure 4 (B) is a diagram showing the concept of setting virtual sound sources.
[0015] Figure 5 is a functional block diagram showing an example of the structure of the grouping unit 40.
[0016] Figure 6 is a flowchart showing the method of grouping sound sources.
[0017] Figure 7 is a diagram showing the concept of grouping multiple sound sources into multiple regions.
[0018] Figure 8 (A) is a flowchart showing the method of grouping sound sources using representative points, Figure 8
[0019] (B) is a flowchart showing the method of grouping sound sources using the boundaries of regions.
[0020] Figure 9 is a flowchart showing an example of the method of grouping by the movement of sound sources.
[0021] Figure 10 is a functional block diagram showing an example of the structure of the initial reflection sound control signal generation unit 50.
[0022] Figure 11 is a diagram showing an example of a GUI.
[0023] Figure 12 is a flowchart showing an example of the setting process of virtual sound sources.
[0024] Figure 13 (A), Figure 13 (B) is a diagram showing the setting examples of each virtual sound source when the geometric shapes are different.
[0025] Figure 14 (A), Figure 14 (B) and Figure 14 (C) are diagrams showing the setting examples of virtual sound sources.
[0026] Figure 15 (A), Figure 15 (B), Figure 15(C) is a diagram showing an example of setting a virtual sound source.
[0027] Figure 16 is a flowchart showing the process of assigning virtual sound sources to speakers.
[0028] Figure 17 (A), Figure 17 (B) is a diagram showing the concept of assigning virtual sound sources to speakers.
[0029] Figure 18 is a flowchart showing the coefficient setting process of LDtap.
[0030] Figure 19 (A), Figure 19 (B) is a diagram for explaining the concept of coefficient setting.
[0031] Figure 20 (A) shows an example of the LDtap coefficient in the case of a large virtual space shape, Figure 20 (B) shows an example of the LDtap coefficient in the case of a small virtual space shape.
[0032] Figure 21 is a diagram showing the waveform of the initial reflected sound control signal generated by the initial reflected sound control signal generation unit 50.
[0033] Figure 22 is a functional block diagram showing an example of the structure of the reverberation sound control signal generation unit 70.
[0034] Figure 23 is a flowchart showing an example of the generation process of the reverberation sound control signal.
[0035] Figure 24 is a curve diagram showing waveform examples of the direct sound, the initial reflected sound control signal, and the reverberation sound control signal.
[0036] Figure 25 is a diagram showing an example of the area setting for the reverberation sound.
[0037] Figure 26 is a functional block diagram showing an example of the structure of the output adjustment unit 90.
[0038] Figure 27 is a flowchart showing an example of the output adjustment process.
[0039] Figure 28 is a diagram showing an example of the GUI for output adjustment.
[0040] Figure 29 (A), Figure 29(B) is a diagram showing a setting example of a case where sound is localized and expanded at the rear side of the playback space.
[0041] Figure 30 (A), Figure 30 (B) is a diagram showing a setting example of a case where sound is localized and expanded in the lateral direction of the playback space.
[0042] Figure 31 is a diagram showing an overview of the expansion of sound in a case where expansion in the height direction is provided.
[0043] Figure 32 is a functional block diagram showing the structure of a sound signal processing apparatus having a binaural playback function. DETAILED DESCRIPTION
[0044] A sound signal processing method and a sound signal processing apparatus according to an embodiment of the present invention will be described with reference to the accompanying drawings. In addition, in the following embodiments, first, an overview of the sound signal processing method and the sound signal processing apparatus will be described, and then, the specific details of each process and each structure will be described.
[0045] In addition, in the present embodiment, the playback space is a space in which a user (listener) listens to sound (direct sound, initial reflected sound, reverberant sound) from a sound source using a speaker or the like. The virtual space is a space having a sound field (acoustics) different from the playback space, and is a space in which the initial reflected sound and the reverberant sound obtained through this sound field are reproduced (simulated) in the playback space.
[0046] [Schematic Structure of Sound Signal Processing Apparatus]
[0047] Figure 1 is a functional block diagram showing the structure of an audio system including the sound signal processing apparatus according to an embodiment of the present invention.
[0048] As Figure 1 shown, the sound signal processing apparatus 10 includes a region setting unit 30, a grouping unit 40, an initial reflected sound control signal generation unit 50, a mixer 60, a reverberant sound control signal generation unit 70, an adder 80, and an output adjustment unit 90. The sound signal processing apparatus 10 is implemented, for example, by an arithmetic processing apparatus such as an electronic circuit or a computer that respectively implements the region setting unit 30, the grouping unit 40, the initial reflected sound control signal generation unit 50, the mixer 60, the reverberant sound control signal generation unit 70, the adder 80, and the output adjustment unit 90. The part constituted by the adder 80 and the output adjustment unit 90 corresponds to the "output signal generation unit" of the present invention.
[0049] The sound signal processing apparatus 10 is connected to a plurality of speakers SP1 - SP64. In addition, Figure 1The figure shows a way of using 64 speakers, but the number of speakers is not limited to this.
[0050] The sound signal processing device 10 is input with the sound signals S1 - S96 of a plurality of sound sources OBJ1 - OBJ96. In addition, Figure 1 The figure shows a way of using 96 sound sources, but the number of sound sources is not limited to this.
[0051] The area setting unit 30 divides the playback space into a plurality of areas and sets information related to the divided areas (area information). The area information is the position coordinates determining the boundaries of the areas and the position coordinates of representative points set in the areas.
[0052] The area setting unit 30 outputs the area information of the plurality of set areas Area1 - Area8 to the grouping unit 40. In addition, Figure 1 The figure shows a form in which the areas are set to 8, but the number of areas is not limited to this.
[0053] The grouping unit 40 groups the sound sources OBJ1 - OBJ96 into a plurality of areas Area1 - Area8. Based on the grouping result, the grouping unit 40 generates area - differentiated sound signals SA1 - SA8 for each of the areas Area1 - Area8 using the sound signals S1 - S96 of the sound sources OBJ1 - OBJ96. For example, the grouping unit 40 mixes the sound signals of the plurality of sound sources grouped into area Area1 to generate the area - differentiated sound signal SA1.
[0054] The grouping unit 40 outputs the plurality of area - differentiated sound signals SA1 - SA8 to the initial reflected sound control signal generation unit 50. In addition, the grouping unit 40 outputs the sound signals S1 - S96 of the sound sources OBJ1 - OBJ96 to the mixer 60.
[0055] The initial reflected sound control signal generation unit 50 generates initial reflected sound control signals ER1 - ER64 for each of the plurality of speakers SP1 - SP64 according to the plurality of area - differentiated sound signals SA1 - SA8. The initial reflected sound control signals ER1 - ER64 are signals respectively output to the speakers SP1 - SP64 in order to simulate the initial reflected sound of the virtual space in the playback space. The initial reflected sound control signal generation unit 50 outputs the generated initial reflected sound control signals ER1 - ER64 to the adder 80.
[0056] Roughly (the detailed structure and processing will be described later), the initial reflected sound control signal generation unit 50 sets virtual sound sources (virtual sound sources) on the playback space by using the positions of the speakers SP1 - SP64 arranged in the playback space and the geometric shape of the virtual space. In addition, the specific setting of the virtual sound sources will be described later. The initial reflected sound control signal generation unit 50 generates initial reflected sound control signals ER1 - ER64 that simulate the initial reflected sound in the virtual space by using the virtual sound sources. At this time, the initial reflected sound control signal generation unit 50 performs desired timbre adjustment on the initial reflected sound control signals ER1 - ER64.
[0057] The mixer 60 is an additive mixer. The mixer 60 mixes the sound signals S1 - S96 of the sound sources OBJ1 - OBJ96 to generate a reverberation sound generation signal Sr. The mixer 60 outputs the reverberation sound generation signal Sr to the reverberation sound control signal generation unit 70.
[0058] The reverberation sound control signal generation unit 70 generates reverberation sound control signals REV1 - REV64 for each of the multiple speakers SP1 - SP64 based on the reverberation sound generation signal Sr. The reverberation sound control signals REV1 - REV64 are signals respectively output to the speakers SP1 - SP64 to simulate the reverberation sound (rear reverberation sound) of the virtual space in the playback space. The reverberation sound control signal generation unit 70 outputs the generated reverberation sound control signals REV1 - REV64 to the adder 80.
[0059] Roughly (the detailed structure and processing will be described later), the reverberation sound control signal generation unit 70 divides the playback space into multiple reverberation sound setting regions, and generates reverberation sound control signals for each of the multiple reverberation sound setting regions. The reverberation sound control signal generation unit 70 assigns the multiple speakers SP1 - SP64 to the multiple reverberation sound setting regions. The reverberation sound control signal generation unit 70 sets the reverberation sound control signals for each reverberation sound setting region for the multiple speakers SP1 - SP64 based on this assignment.
[0060] At this time, the reverberation sound control signal generation unit 70 sets the connection timing of the initial reflected sound and the reverberation sound based on the geometric shape of the playback space. The reverberation sound control signal generation unit 70 gradually increases the level (amplitude) of the reverberation sound control signal during the period before the connection timing, and gradually decreases the level (amplitude) of the reverberation sound control signal during the connection timing and subsequent periods.
[0061] The adder 80 adds the initial reverberation control signals and the echo control signals respectively generated for the multiple speakers SP1 - SP64 to generate signals Sat1 - Sat64 for the multiple speakers. For example, the adder 80 adds the initial reverberation control signal for the speaker SP1 and the echo control signal for the speaker SP1 to generate the signal Sat1 for the speaker. The adder 80 outputs the signals Sat1 - Sat64 for the multiple speakers to the output adjustment unit 90.
[0062] The output adjustment unit 90 performs gain control and hysteresis control on the signals Sat1 - Sat64 for the multiple speakers to generate output signals So1 - So64. The output adjustment unit 90 outputs the output signals So1 - So64 to the multiple speakers SP1 - SP64. For example, the output adjustment unit 90 performs gain control and hysteresis control for the speaker SP1 on the signal Sat1 for the speaker to generate the output signal So1. The output adjustment unit 90 outputs the output signal So1 to the speaker SP1.
[0063] Generally (the detailed structure and processing will be described later), the output adjustment unit 90 receives the input of the acoustic parameters of the playback space. The acoustic parameters are, for example, parameters for setting the adjustment of the expansion of the space in the width direction of the sound space, the adjustment of the expansion of the space behind the sound collection point in the sound space, the adjustment of the expansion of the space in the ceiling direction of the sound space, etc. The output adjustment unit 90 centrally sets the gain values and hysteresis amounts (delay amounts) of the signals Sat1 - Sat64 for the multiple speakers based on the position coordinates of the multiple speakers SP1 - SP64 and the acoustic parameters. Centrally setting means that instead of setting separately for each speaker, for example, by inputting the position coordinates of each speaker into a specific calculation formula shared by all speakers, the gain values and hysteresis amounts of each speaker are set. The output adjustment unit 90 performs gain control and hysteresis control on the signals Sat1 - Sat64 for the multiple speakers using the set gain values and hysteresis values.
[0064] [Overview of the sound signal processing method]
[0065] Figure 2 is a flowchart of the sound signal processing method according to an embodiment of the present invention. Figure 2 shows the Figure 1 sound signal processing method implemented by the sound signal processing device 10. In addition, Figure 2 the content of each process shown in Figure 1 has been described in the above
[0066] (Grouping of sound sources OBJ1 - OBJ96)
[0067] The grouping unit 40 groups multiple sound sources OBJ1 - OBJ96 for each of multiple areas Area1 - Area8 (S11).
[0068] (Generation of initial reflected sound control signal)
[0069] The initial reflected sound control signal generation unit 50 sets the timbre for the initial reflected sound for each group (S12). The initial reflected sound control signal generation unit 50 sets virtual sound sources for each group (S13). The initial reflected sound control signal generation unit 50 uses the timbre and the virtual sound sources to generate initial reflected sound control signals for each of the multiple speakers SP1 - SP64 (S14).
[0070] (Generation of reverberation sound control signal)
[0071] The mixer 60 adds the sound signals S1 - S96 of the multiple sound sources OBJ1 - OBJ96 (S21). The reverberation sound control signal generation unit 70 sets the connection timing between the initial reflected sound and the reverberation sound based on the geometric shape of the playback space (S22). The reverberation sound control signal generation unit 70 generates a reverberation sound control signal using the set connection timing (S23). The reverberation sound control signal generation unit 70 distributes the generated reverberation sound control signal to the multiple speakers SP1 - SP64 based on the position coordinates of the multiple speakers SP1 - SP64 in the playback space (S24).
[0072] (Output processing to multiple speakers)
[0073] The adder 80 adds the initial reflected sound control signal and the reverberation sound control signal for each of the multiple speakers SP1 - SP64 to generate speaker signals Sat1 - Sat64 (S31).
[0074] The output adjustment unit 90 generates output signals So1 - So64 based on the speaker signals Sat1 - Sat64 using acoustic parameters that achieve the localization of the reverberation in the playback space and the expansion of the space (S32). The output adjustment unit 90 outputs the output signals So1 - So64 to the multiple speakers SP1 - SP64 (S33).
[0075] By using the above structure and processing, the sound signal processing device 10 (sound signal processing method) obtains the following various effects.
[0076] (1) The sound signal processing device 10 (sound signal processing method) groups sound sources for each region obtained by dividing the playback space, generates initial reflected sounds, thereby enabling clear sound image localization and rich spatial expansion. At this time, the reverberant sound is constant throughout the playback space, and only the initial reflected sound varies depending on the position of the sound source. Therefore, for example, when the position of the sound source moves, the movement of the sound of the sound source is smoother.
[0077] (2) The sound signal processing device 10 (sound signal processing method) generates an initial reflected sound control signal using virtual sound sources, thereby enabling more realistic simulation of the initial reflected sounds obtained based on the geometric shape of the virtual space in the playback space.
[0078] (3) The sound signal processing device 10 (sound signal processing method) performs tone color adjustment of the initial reflected sound control signal, thereby enabling elimination of unnaturalness in the tone color of the initial reflected sound simulated only by virtual sound sources, for example.
[0079] (4) The sound signal processing device 10 (sound signal processing method) sets the connection timing of the initial reflected sound control signal and the reverberant sound control signal according to the geometric shape of the playback space, thereby enabling smoother and more natural connection from the initial reflected sound to the reverberant sound.
[0080] (5) The sound signal processing device 10 (sound signal processing method) collectively adjusts the gain values and hysteresis amounts of the speaker signals Sat1 - Sat64 including the initial reflected sound control signal and the reverberant sound control signal, thereby enabling realization of the sound field desired by the user in the playback space through easier operation input.
[0081] [Specific descriptions of each signal processing unit and each process]
[0082] Hereinafter, specific descriptions of each of the above signal processing units and each process will be described. First, with reference to the drawings, the initial reflected sound, reverberant sound, and virtual sound sources required for understanding the invention will be described.
[0083] [Initial reflected sound and reverberant sound]
[0084] Figure 3 It is a diagram showing the discrete waveform of sound including normal direct sound, initial reflected sound, and reverberant sound (rear reverberant sound). For example, a hall where performance or content is played has an enclosed space surrounded by walls. If a sound occurs in this enclosed space, the direct sound, initial reflected sound, and reverberant sound (rear reverberant sound) reach the sound collection point.
[0085] The direct sound is the sound that directly reaches the sound collection point from the sound generation position.
[0086] The initial reflected sound is the sound that reaches the sound pickup point at an earlier time after the sound generated at the generation position is reflected by the walls, floor, and ceiling. Therefore, the initial reflected sound reaches the sound pickup point after the direct sound. In addition, the volume (level) of the initial reflected sound is less than that of the direct sound. If the number of reflections is 1, it is the first-order reflected sound; if it is n, it is the nth-order reflected sound. The arrival direction and volume of the initial reflected sound at the sound pickup point are greatly affected by the generation position of the sound.
[0087] The reverberant sound reaches the sound pickup point after the initial reflected sound. The reverberant sound is the sound that reaches the sound pickup point after multiple reflections of the sound generated at the generation position. That is, the reverberant sound is the sound that reaches the sound pickup point after the reflected sound is further reflected and attenuated multiple times. Therefore, the volume (level) of the reverberant sound is less than that of the initial reflected sound. Also, the arrival direction and volume of the reverberant sound are less affected by the generation position of the sound compared to the initial reflected sound.
[0088] [Virtual sound source]
[0089] Figure 4 (A), Figure 4 (B) are diagrams showing the concept of setting a virtual sound source. In addition, in Figure 4 (A), Figure 4 (B), for ease of explanation, the concept of setting a two-dimensional virtual sound source is shown, but the virtual sound source can also be set in three dimensions with the same concept. That is, in the actual playback space, when the sound sources are not aligned on a plane but are spatially arranged and the virtual space is set stereoscopically, the virtual sound source is set three-dimensionally.
[0090] In the playback space, there are a sound source SS and a sound pickup point RP. In addition, Figure 4 (A), Figure 4 (B), the sound source SS shown is different in meaning from the sound source OBJ described above and refers to the sound source that generates ordinary sounds. Also, in the playback space, a virtual wall IWL for realizing the sound field of the virtual space is set. The virtual wall IWL is obtained according to the geometric shape of the virtual space.
[0091] The sound source SS and the sound pickup point RP exist in the space surrounded by the virtual wall IWL. The virtual wall IWL has a virtual wall IWL1, a virtual wall IWL2, a virtual wall IWL3, and a virtual wall IWL4. The virtual wall IWL1 and the virtual wall IWL4 are in the first direction in the playback space ( Figure 4 (A), Figure 4(In the longitudinal direction of (B)), it is arranged in such a way that the sound source SS and the sound reception point RP are sandwiched in the middle. The virtual wall IWL1 is arranged on the side closer to the sound source SS than the sound reception point RP, and the virtual wall IWL4 is arranged on the side closer to the sound reception point RP than the sound source SS. The virtual walls IWL2 and IWL3 are arranged in such a way that they sandwich the sound source SS and the sound reception point RP in the second direction of the playback space ( Figure 4 (A), Figure 4 (In the lateral direction of (B)), it is arranged in such a way that the sound source SS and the sound reception point RP are sandwiched in the middle. The virtual wall IWL2 is arranged on the side closer to the sound source SS than the sound reception point RP, and the virtual wall IWL3 is arranged on the side closer to the sound reception point RP than the sound source SS.
[0092] If the virtual walls IWL1, IWL2, IWL3, and IWL4 were walls that reflect sound in reality, then as Figure 4 (shown in (B)), the sound emitted from the sound source SS is reflected at the virtual walls IWL1, IWL2, and IWL3 and reaches the sound reception point RP. In addition, in Figure 4 (B), the reflection from the virtual wall IWL4 is not described, but at the virtual wall IWL4, reflection also occurs in the same way as at the virtual walls IWL1, IWL2, and IWL3.
[0093] However, the virtual walls IWL1, IWL2, IWL3, and IWL4 do not actually exist in the playback space. Therefore, as Figure 4 (shown in (A)), the sound signal processing device 10 sets the virtual sound sources IS1, IS2, and IS3 by regarding the reflection of sound at the wall surface as specular reflection.
[0094] Specifically, the sound signal processing device 10 sets the virtual sound source IS1 at a position symmetric to the sound source SS line with respect to the virtual wall IWL1 as the reference line. The sound signal processing device 10 sets the virtual sound source IS2 at a position symmetric to the sound source SS line with respect to the virtual wall IWL2 as the reference line. The virtual sound source IS3 is set at a position symmetric to the sound source SS line with respect to the virtual wall IWL3 as the reference line. In addition, by adjusting the sound power of each virtual sound source IS, it is possible to simulate the energy loss of the reflection at the virtual wall IWL.
[0095] By making the above settings, the sound generated at the virtual sound source IS1 is the same as the sound generated at the sound source SS and reflected at the virtual wall IW1. The sound generated at the virtual sound source IS2 is the same as the sound generated at the sound source SS and reflected at the virtual wall IW2. The sound generated at the virtual sound source IS3 is the same as the sound generated at the sound source SS and reflected at the virtual wall IW3. In addition, in Figure 4 (A), Figure 4In (B), there is no description of the virtual sound source for the virtual wall IWL4. However, for the virtual wall IWL4, a virtual sound source can also be set in the same way as for the virtual walls IWL1, IWL2, and IWL3.
[0096] The sound signal processing device 10 sets the virtual sound source in the above-described manner, so that the initial reflected sound in the virtual space can be simulated in the playback space in the virtual space where there is no real wall.
[0097] [Structure and Processing of Grouping Unit 40]
[0098] Figure 5 is a functional block diagram showing an example of the structure of the grouping unit 40. Figure 6 is a flowchart showing the method of grouping sound sources.
[0099] As Figure 5 shown, the grouping unit 40 includes a sound source position detection unit 41, a region determination unit 42, and a matrix mixer 400.
[0100] The sound source position detection unit 41 detects the position coordinates of a plurality of sound sources OBJ1 - OBJ96 in the playback space ( Figure 6 : S111). For example, the sound source position detection unit 41 detects the position coordinates of the sound sources OBJ1 - OBJ96 through an operation input from the user. Alternatively, the sound source position detection unit 41 has a position detection sensor for detecting the sound sources OBJ1 - OBJ96, and detects the position coordinates of the sound sources OBJ1 - OBJ96 based on the positions detected by the position detection sensor.
[0101] The sound source position detection unit 41 outputs the position coordinates of the sound sources OBJ1 - OBJ96 to the region determination unit 42.
[0102] The region determination unit 42 groups the sound sources OBJ1 - OBJ96 into a plurality of regions Area1 - Area8 using the region information of the plurality of regions Area1 - Area8 from the region setting unit 30 and the position coordinates of the sound sources OBJ1 - OBJ96 from the sound source position detection unit 41 ( Figure 6 : S112). More specifically, the region determination unit 42 groups in the following manner.
[0103] Figure 7 is a diagram showing the concept of grouping a plurality of sound sources into a plurality of regions. In addition, in Figure 7 , the upper side of the diagram is the front side of the hall as the playback space, and the lower side of the diagram is the rear side of the hall.
[0104] The region setting unit 30 sets a reference point Pso for region division for the playback space. For example, as Figure 7As shown, the area setting unit 30 sets the center position of the hall that realizes the playback space at the reference point Pso. In addition, the area setting unit 30 can also use the point (position) set by the user as the reference point. For example, the area setting unit 30 can use the sound collection point set by the user, etc. as the reference point.
[0105] The area setting unit 30 sets eight areas Area1 - Area8 in such a way that the reference point Pso for area division is the center and the entire circumference on the plane is divided into eight parts. For example, in Figure 7 this case, the area setting unit 30 sets multiple areas Area1, Area2, Area3 at positions in front of the reference point Pso in the hall (playback space). In addition, the area setting unit 30 sets area Area4 at a position to the left when facing the front side of the hall from the reference point Pso, and sets area Area5 at a position to the right when facing the front side of the hall from the reference point Pso. In addition, the area setting unit 30 sets multiple areas Area6, Area7, Area8 at positions behind the reference point Pso in the hall (playback space).
[0106] In addition, the setting of this area is an example, and as long as the entire playback space can be covered by the set multiple areas, other settings are also possible. In addition, this description shows the setting of a planar area, but for a spatial area, the setting can also be carried out in the same way. For example, the vertical range of area Area1 is also included in area Area1.
[0107] The area setting unit 30 sets representative points RP1 - RP8 for the multiple areas Area1 - Area8 respectively. For example, the area setting unit 30 sets the multiple representative points RP1 - RP8 at the center positions of the multiple areas Area1 - Area8. Or, in a case where the areas spread radially as in Figure 7 this, for example, the area setting unit 30 sets the representative point at a position a specified distance from the reference point Pso on the straight line passing through the center of the radially spreading angle. In addition, the setting method of these representative points is an example. For example, one representative point can also be set for one area, and as long as it is a method that can reliably perform the grouping process of sound sources, other methods are also possible.
[0108] The area setting unit 30 outputs the area information of the multiple areas Area1 - Area8 to the area determination unit 42 and the matrix mixer 400 of the grouping unit 40. The area information of the multiple areas Area1 - Area8 is the position coordinates of the representative points RP1 - RP8 of the areas Area1 - Area8, the coordinate information of the boundary line representing the shape of the areas Area1 - Area8, etc.
[0109] (Method for grouping sound sources into regions using representative points)
[0110] Figure 8 (A) is a flowchart showing a method for grouping sound sources using representative points.
[0111] The region determination unit 42 obtains the position coordinates of the representative points RP1 - RP8 based on the region information of the multiple regions Area1 - Area8 (S1121). The region determination unit 42 calculates the distances between the position coordinates of the sound sources to be grouped and the position coordinates of the representative points RP1 - RP8 (S1122). The region determination unit 42 groups the sound sources into the region containing the representative point that becomes the shortest distance (S1123).
[0112] For example, in the case of the sound source OBJ1 in Figure 7 the example, the region determination unit 42 detects the position coordinates of the sound source OBJ1 and obtains the position coordinates of the multiple representative points RP1 - RP8. The region determination unit 42 calculates the distances between the sound source OBJ1 and the multiple representative points RP1 - RP8 respectively based on the position coordinates of the sound source OBJ1 and the position coordinates of the multiple representative points RP1 - RP8. The region determination unit 42 detects the case where the distance between the sound source OBJ1 and the representative point RP1 is shorter than the distances between the sound source OBJ1 and the other representative points RP2 - RP8. In other words, the region determination unit 42 detects the case where the distance between the sound source OBJ1 and the representative point RP1 becomes the shortest distance. The region determination unit 42 groups the sound source OBJ1 into the region Area1 associated with the representative point RP1.
[0113] (Method for grouping sound sources into regions using the boundaries of regions)
[0114] Figure 8 (B) is a flowchart showing a method for grouping sound sources using the boundaries of regions.
[0115] The region determination unit 42 obtains the coordinate information (boundary coordinates) representing the boundary lines of the respective regions Area1 - Area8 based on the region information of the multiple regions Area1 - Area8 (S1124). The region determination unit 42 determines whether the position coordinates of the sound sources to be grouped are inside the respective regions Area1 - Area8 (S1125). For example, the region determination unit 42 uses the Crossing Number Algorithm to determine whether the sound source is inside or outside the region. If the sound source is inside the region (S1125: YES), the region determination unit 42 groups the sound source into that region (S1126).
[0116] For example, in Figure 7In the case of the sound source OBJ1 in the example, the area determination unit 42 detects the position coordinates of the sound source OBJ1 and obtains coordinate information (boundary coordinates) representing the boundary lines of the plurality of areas Area1 - Area8. The area determination unit 42 performs an inside / outside determination of the sound source OBJ1 with respect to the plurality of areas Area1 - Area8 based on the position coordinates of the sound source OBJ1 and the boundary coordinates of the plurality of areas Area1 - Area8. The area determination unit 42 detects the case where the sound source OBJ1 is within the area Area1. The area determination unit 42 groups the sound source OBJ1 into the area Area1.
[0117] The area determination unit 42 groups the input plurality of sound sources OBJ1 - OBJ96 into the plurality of areas Area1 - Area8. For example, if it is Figure 7 the example, the area determination unit 42 groups the sound sources OBJ1 and OBJ4 into the area Area1, groups the sound source OBJ2 into the area Area2, and groups the sound source OBJ3 into the area Area5.
[0118] The area determination unit 42 outputs the grouping information to the matrix mixer 400. The grouping information is information indicating which sound source is grouped into which area as described above.
[0119] Based on the grouping information, the matrix mixer 400 generates area - differentiated sound signals SA1 - SA8 for the plurality of areas Area1 - Area8 using the sound signals S1 - S96 of the plurality of sound sources OBJ1 - OBJ96. For example, if there are multiple sound sources grouped in an area, the matrix mixer 400 mixes the sound signals of these multiple sound sources to generate the area - differentiated sound signal for that area. The matrix mixer 400 outputs the area - differentiated sound signals for each area to the initial reflected sound control signal generation unit 50. In addition, even if there is only 1 sound source grouped in an area, the matrix mixer 400 outputs the sound signal of that sound source as the area - differentiated sound signal for that area to the initial reflected sound control signal generation unit 50.
[0120] If it is Figure 7 the example, the area Area1 is grouped with the sound sources OBJ1 and OBJ4. The matrix mixer 400 mixes the sound signal S1 of the sound source OBJ1 and the sound signal S4 of the sound source OBJ4 to generate and output the area - differentiated sound signal SA1 for the area Area1. In addition, the area Area2 is grouped with the sound source OBJ2. The matrix mixer 400 outputs the sound signal S2 of the sound source OBJ2 as the area - differentiated sound signal SA2 for the area Area2. In addition, the area Area5 is grouped with the sound source OBJ3. The matrix mixer 400 outputs the sound signal S3 of the sound source OBJ3 as the area - differentiated sound signal SA5 for the area Area5.
[0121] By implementing the above-described structure and processing, the sound signal processing apparatus 10 can group a plurality of sound sources for each of a plurality of regions obtained by dividing the sound space, and generate an initial reverberation control signal. As described above, the sound signal processing apparatus 10 can reproduce the initial reverberation corresponding to the position of the sound source, and can achieve clear sound image localization and rich spatial expansion.
[0122] In addition, in the above description, the case where the sound source moves is not shown in detail. However, in the case where the sound source moves, the grouping unit 40 performs Figure 9 the processing shown. Figure 9 is a flowchart showing an example of a method of grouping by the movement of the sound source.
[0123] The sound source position detection unit 41 detects the movement of the sound source (S104). The sound source position detection unit 41 detects the movement of the sound source, for example, by an operation input from the user. Alternatively, the sound source position detection unit 41 continuously detects the sound source position by a position detection sensor, and thereby detects the movement of the sound source. Then, the area determination unit 42 re-groups the sound source after the movement (S105). The sound source position detection unit 41 detects the position coordinates of the sound source after the movement and outputs them to the area determination unit 42.
[0124] The area determination unit 42 uses the position coordinates of the sound source after the movement to perform grouping into a plurality of areas Area1 - Area8 as described above (S105).
[0125] By performing the above-described processing, even if the sound source moves, the sound signal processing apparatus 10 can generate an initial reverberation control signal corresponding to the position of the sound source after the movement. As described above, the sound signal processing apparatus 10 can reproduce the change in the initial reverberation corresponding to the movement of the sound source, and can achieve clear sound image localization and rich spatial expansion corresponding to the movement even when there is movement of the sound source.
[0126] In addition, when the movement of the sound source as described above occurs, the sound signal processing apparatus 10 can perform a crossfade process on the initial reverberation control signal before the movement and the initial reverberation control signal after the movement. For example, when the sound source moves, the sound signal processing apparatus 10 gradually decreases the component of the sound signal of the sound source in the regionally divided sound signal including the sound source before the movement. On the other hand, the sound signal processing apparatus 10 gradually increases the component of the sound signal of the sound source in the regionally divided sound signal including the sound source after the movement.
[0127] By performing the above-described processing, the sound signal processing apparatus 10 can suppress discontinuous changes in the initial reflected sound when the sound source moves. As described above, when the sound source moves, the sound signal processing apparatus 10 can cause the initial reflected sound to change more smoothly in response to the movement of the sound source.
[0128] In addition, the matrix mixer 400 outputs the sound signals S1 - S96 of the plurality of sound sources OBJ1 - OBJ96 to the mixer 60. As described above, the mixer 60 adds the sound signals S1 - S96 to generate a reverberation sound generation signal Sr and outputs it to the reverberation sound control signal generation unit 70. The reverberation sound control signal generation unit 70 uses the reverberation sound generation signal Sr to generate reverberation sound control signals REV1 - REV64.
[0129] By the above-described processing, the reverberation sound is not affected by the position or movement of the sound source. Therefore, even when the sound source moves, the sound signal processing apparatus 10 can keep the reverberation sound in the playback space constant and reproduce the movement of the sound source more clearly through the change in the initial reflected sound.
[0130] [Generation of Initial Reflected Sound Control Signal]
[0131] Figure 10 is a functional block diagram showing an example of the structure of the initial reflected sound control signal generation unit 50. Figure 11 is a diagram showing an example of the GUI.
[0132] As Figure 10 shown, the initial reflected sound control signal generation unit 50 includes an FIR filter circuit 51, an LDtap circuit 52, an addition processing unit 53, a timbre setting unit 501, a virtual sound source setting unit 502, and an operation unit 500. The LDtap circuit 52 is a circuit that amplifies and delays an input signal and outputs it. The FIR filter circuit 51 includes a plurality of FIR filters 511 - 518. The LDtap circuit 52 includes a plurality of LDtaps 521 - 528, an output speaker setting unit 5201, and a coefficient setting unit 5202. In addition, the order of the FIR filter circuit 51 and the LDtap circuit 52 may be reversed.
[0133] [Timbre Adjustment of Initial Reflected Sound]
[0134] The operation unit 500 receives designation information of the timbre added to the initial reflected sound from the user and outputs it to the timbre setting unit 501. The designation information of the timbre is, for example, information for designating emphasis on the bass range, emphasis on the treble range, the volume of the initial reflected sound, the attenuation characteristic of the initial reflected sound, etc. (information indicating filtering characteristics).
[0135] As a specific example, the operation unit 500 by as Figure 11Receives operations for the GUI 100 (Graphical User Interface) shown.
[0136] The GUI 100 has a setting display window 111, a plurality of operation members 112, a knob 1131, and an adjustment value display window 1132.
[0137] The setting display window 111 displays the shape of the virtual wall IWL of the virtual space set through the plurality of operation members 112 and the knob 1131. At this time, the setting display window 111 can display the position of the additionally set sound source SS, the position of the speaker SP, the position of the sound collection point RP, and the coordinate axes of the playback space together with the virtual wall IWL.
[0138] The plurality of operation members 112 are associated with samples (various halls, rooms, etc.) of a preset virtual space. In addition, although not shown in the figure, indexes (such as hall names) clearly indicating the samples of the virtual space associated with each operation member 112 are displayed on the plurality of operation members 112.
[0139] The knob 1131 is used to set the room size of the virtual space. The adjustment value display window 1132 displays the set value of the room size of the virtual space.
[0140] The GUI 100 receives various operations for adjusting the tone color. For example, the GUI 100 has a plurality of operation members 112, an operation member for the bass range, an operation member for the treble range, an operation member for volume adjustment, an operation member for attenuation characteristic adjustment, etc., and receives operations through these operation members.
[0141] If the user operates a desired operation member using the GUI 100, the operation unit 500 detects this operation and sets the specified information of the tone color corresponding to these operations.
[0142] For example, if the operation unit 500 receives the selection of the plurality of operation members 112, it acquires the specified information of the tone color preset in the virtual space associated with the operation member 112. In addition, if the operation unit 500 receives operations performed through the operation member for the bass range, the operation member for the treble range, the operation member for volume adjustment, the operation member for attenuation characteristic adjustment, etc., it acquires the specified information of the tone color set through these operation members.
[0143] In addition, although not shown in the figure, the GUI 100 can also display the specified information of the tone color using, for example, the filter coefficients of the FIR filters 511 - 518 described later, a schematic waveform, etc. In this case, if the GUI 100 receives an adjustment of the specified information of the tone color, it can also change the display accordingly. For example, the GUI 100 can change the display of the waveform corresponding to the adjustment.
[0144] The timbre setting unit 501 sets the filter coefficients of the FIR filters 511-518 of the FIR filter circuit 51 based on the timbre designation information. For example, if the timbre setting unit 501 receives the designation information emphasizing the bass range, it sets the filter coefficients of the low range of the FIR filters 511-518 of the FIR filter circuit 51 to be enhanced. In addition, if the timbre setting unit 501 receives the designation information emphasizing the treble range, it sets the filter coefficients of the high range of the FIR filters 511-518 of the FIR filter circuit 51 to be enhanced. The timbre setting unit 501 outputs the set filter coefficients to the FIR filter circuit 51. In addition, not limited to the filter coefficients, as the filter characteristics, the timbre setting unit 501 can also set and adjust the sampling frequency and the filter length.
[0145] In addition, the timbre setting unit 501 sets the gain values of the respective taps of the FIR filters 511-518 of the FIR filter circuit 51 based on the timbre designation information. The timbre setting unit 501 outputs the set gain values to the FIR filter circuit 51.
[0146] The plurality of FIR filters 511-518 are filters respectively corresponding to the sound signals SA1-SA8 divided by region. The sound signals SA1-SA8 divided by region are input to the FIR filters 511-518. For example, as Figure 10 shown, the sound signal SA1 divided by region is input to the FIR filter 511, the sound signal SA2 divided by region is input to the FIR filter 512, the sound signal SA3 divided by region is input to the FIR filter 513, and the sound signal SA4 divided by region is input to the FIR filter 514. The sound signal SA5 divided by region is input to the FIR filter 515, the sound signal SA6 divided by region is input to the FIR filter 516, the sound signal SA7 divided by region is input to the FIR filter 517, and the sound signal SA8 divided by region is input to the FIR filter 518.
[0147] The plurality of FIR filters 511-518 have the same number of taps. For example, the plurality of FIR filters 511-518 have 16,000 taps. In addition, this number of taps is an example, and it can be set based on the resource conditions of the sound signal processing device 10, the accuracy of the timbre of the initial reflected sound to be reproduced, and the like.
[0148] Multiple FIR filters 511 - 518 perform filter processing (convolution operation) on the multiple regionally divided sound signals SA1 - SA8 respectively according to the filter coefficients and gain values set by the timbre setting unit 501. As described above, the multiple FIR filters 511 - 518 generate the regionally divided sound signals SA1f - SA8f after the filter processing. For example, the FIR filter 511 performs filter processing (convolution operation) on the regionally divided sound signal SA1 according to the filter coefficients and gain values set by the timbre setting unit 501, and generates the regionally divided sound signal SA1f after the filter processing. Similarly, the multiple FIR filters 512 - 518 respectively generate the regionally divided sound signals SA2f - SA8f after the filter processing according to the regionally divided sound signals SA2 - SA8.
[0149] The multiple FIR filters 511 - 518 output the regionally divided sound signals SA1f - SA8f after the filter processing to the multiple LDtaps 521 - 528. For example, the FIR filter 511 outputs the regionally divided sound signal SA1f after the filter processing to the LDtap 521. Similarly, the multiple FIR filters 512 - 518 output the regionally divided sound signals SA2f - SA8f after the filter processing to the multiple LDtaps 522 - 528.
[0150] In addition, the specified information of the timbre is not limited to the emphasized information of the pitch range, but also includes the specified information that sets the waveform of the initial reflected sound to the desired characteristics of the user. By using the specified information of the timbre as described above, the sound signal processing device 10 can more diversely implement the initial reflected sound of the timbre corresponding to the user's preferences.
[0151] [Virtual sound source setting and LDtap setting]
[0152] The virtual sound source setting unit 502 sets the virtual sound source based on the position coordinates of the sound collection point in the playback space and the geometric shape of the virtual space.
[0153] Figure 12 is a flowchart showing an example of the setting process of the virtual sound source. The virtual sound source setting unit 502 acquires the position coordinates of the sound collection point in the playback space (S131). For example, the virtual sound source setting unit 502 acquires the position coordinates of the sound collection point in the playback space through the operation input from the user, the detection of the position by the position detection sensor, etc.
[0154] The virtual sound source setting unit 502 acquires the geometric shape of the virtual space (S132). For example, the virtual sound source setting unit 502 acquires the geometric shape of the virtual space through the operation input from the user, etc. The geometric shape of the virtual space includes the coordinate group representing the shape of the wall arranged in the virtual space, etc.
[0155] The virtual sound source setting unit 502 is connected to the GUI 100. If the user selects a desired operating member 112 from among a plurality of operating members 112, the GUI 100 reads and acquires the geometric shape of the virtual space associated with the operating member 112. Further, if the user adjusts the room size (the size of the playback space) using the knob 1131, the GUI 100 acquires the adjustment value of the room size.
[0156] The virtual sound source setting unit 502 acquires the position coordinates of the geometric shape of the virtual space with the room size set based on each setting acquired by the GUI 100 in the above-described manner. Further, the virtual sound source setting unit 502 acquires the position coordinates of the sound source SS and the position coordinates of the sound collection point RP (the center of the room (the center position of the playback space)). The virtual sound source setting unit 502 sets the virtual sound source in the following manner using the acquired information. The virtual sound source setting unit 502 aligns the coordinate system of the playback space and the coordinate system of the virtual space. The virtual sound source setting unit 502 uses the position coordinates of the sound collection point in the playback space and the geometric shape of the virtual space, and sets the position coordinates (S133) of the virtual sound source in the playback space by using the concepts of Figure 4 (A), Figure 4 (B).
[0157] Figure 13 (A), Figure 13 (B) are diagrams showing setting examples of respective virtual sound sources when the geometric shapes are different. Figure 13 (A) is a quadrilateral virtual wall IWL, Figure 13 (B) is a hexagonal virtual wall IWLh.
[0158] As described above, if the geometric shape of the virtual space is different, even if the position coordinates of the sound source SSa and the sound collection point RP do not change, the positional relationship between the sound source SSa and the sound collection point RP and the virtual wall IWL, and the positional relationship between the sound source SSa and the sound collection point RP and the virtual wall IWLh are different. As described above, the positions of the virtual sound sources IS1a, IS2a, and IS3a set in the case of Figure 13 (A) and the positions of the virtual sound sources IS1ah, IS2ah, and IS3ah set in Figure 13 (B) are different.
[0159] Figure 14 (A), Figure 14 (B), and Figure 14 (C) are diagrams showing setting examples of virtual sound sources. Figure 14 (A), Figure 14 (B), Figure 14 (C) are diagrams showing the planar change of the virtual sound source. Figure 14 (B) shows with respect to Figure 14(A) shows a case where the position of the sound source SSa relative to the reference point (the sound pickup point RP) is the same, but the size of the virtual space is different. Figure 14 (C) shows the case relative to Figure 14 (A) where the size of the virtual space is the same, but the positional relationship between the reference point of the virtual space and the reference point (the sound pickup point) of the playback space has changed (the case where the center of the room in the playback space has changed).
[0160] As can be seen from Figure 14 (A) and Figure 14 (B), the size of the virtual space on the playback space ( Figure 14 (recorded by the virtual wall IWL in (A), Figure 14 (recorded by the virtual wall IWLc in (B)) is different, and thus the distance and positional relationship between the sound source SSa, which is the origin of the virtual sound source, and the virtual wall are different. As described above, the positions of the virtual sound sources IS1a, IS2a, and IS3a set in the case of Figure 14 (A) are different from the positions of the virtual sound sources IS1c, IS2c, and IS3c set in the case of Figure 14 (B).
[0161] In addition, as can be seen from the comparison results of Figure 14 (A) and Figure 14 (C), the positional relationship between the reference point of the virtual space and the sound pickup point RP has changed, and thus the position of the virtual sound source on the playback space (the position of the virtual sound source relative to the sound pickup point RP and the speaker) moves. As described above, the positions of the virtual sound sources IS1a, IS2a, and IS3a set in the case of Figure 14 (A) are different from the positions of the virtual sound sources IS1as, IS2as, and IS3as set in the case of Figure 14 (C).
[0162] Figure 15 (A), Figure 15 (B), Figure 15 (C) are diagrams showing examples of setting virtual sound sources. Figure 15 (A), Figure 15 (B), Figure 15 (C) are diagrams showing the change in the position of the virtual sound source in the height direction.
[0163] In Figure 15 (A) and Figure 15 (B), the height of the ceiling is different. That is, Figure 15 the distance (height) from the virtual wall IWFL on the floor to the virtual wall IWCL on the ceiling shown in Figure 15(B) The distance (height) of the virtual wall WILL from the virtual wall IWFL on the floor to the virtual wall IWCLL on the ceiling is different.
[0164] As according to Figure 15 (A) and Figure 15 (B), as can be seen from the comparison result, the height of the ceiling is different. Thus, the distance and positional relationship between the sound source, which is the origin of the virtual sound source, and the virtual walls IWCL and IWCLL of the ceiling are different. As described above, the position of the virtual sound source IS1Ca set in the case of Figure 15 (A) is different from the position of the virtual sound source IS1CaL set in the case of Figure 15 (B).
[0165] In Figure 15 (A) and Figure 15 (C), the shape of the ceiling is different. That is, Figure 15 the shape of the virtual wall IWCL of the ceiling of the virtual wall IWL shown in Figure 15 (A) is different from the shape of the virtual wall IWCLx of the ceiling of the virtual wall IWLx shown in
[0166] As according to Figure 15 (A) and Figure 15 (C), as can be seen from the comparison result, the shape of the ceiling is different. Thus, the positional relationship between the sound source, which is the origin of the virtual sound source, and the virtual walls IWCL and IWCLx of the ceiling are different. As described above, the position of the virtual sound source IS1Ca set in the case of Figure 15 (A) is different from the position of the virtual sound source IS1Cax set in the case of Figure 15 (C).
[0167] As described above, the virtual sound source setting unit 502 can optimally set the position of the virtual sound source in the playback space corresponding to the geometric shape of the virtual space, the positional relationship between the playback space and the virtual space. Thus, the sound signal processing device 10 can clearly localize the sound image of the initial reflected sound corresponding to the position coordinates of the speakers in the playback space, the geometric shape of the virtual space, and the positional relationship between the playback space and the virtual space.
[0168] The virtual sound source setting unit 502 outputs the position coordinates of the virtual sound sources set for the multiple regions Area1 - Area8 to the output speaker setting unit 5201 of the LDtap circuit 52.
[0169] The output speaker setting unit 5201 sets the virtual sound source IS assigned to each speaker based on the position coordinates of the virtual sound source IS, the position coordinates of the sound collection point RP, and the position coordinates of the multiple speakers SP1 - SP64. Figure 16 It is a flowchart showing the process of assigning the virtual sound source to the speaker.
[0170] The output speaker setting unit 5201 acquires the position coordinates of the virtual sound source from the virtual sound source setting unit 502 (S141). The output speaker setting unit 5201 acquires the position coordinates of the sound collection point in the playback space, for example, through an operation input from the user or the like (S142). The output speaker setting unit 5201 acquires the position coordinates of the plurality of speakers SP1 - SP64, for example, through an operation input from the user or the like (S143).
[0171] The output speaker setting unit 5201 sets the responsible area of the virtual sound source for each speaker according to the positional relationship between the sound collection point RP in the playback space and the plurality of speakers SP1 - SP64 (S144).
[0172] More specifically, the output speaker setting unit 5201 sets the responsible area of the virtual sound source for each speaker in the following manner. Figure 17 (A), Figure 17 (B) is a diagram showing the concept of allocating virtual sound sources to speakers. Figure 17 (A) shows the concept of allocation using the azimuth angle φ, Figure 17 (B) shows the concept of allocation using the elevation angle θ. In addition, the following description is given taking the speaker SP1 as an example, but the output speaker setting unit 5201 sets the responsible area for the other speakers SP2 - SP64 by the same method.
[0173] The output speaker setting unit 5201 uses the position coordinates of the sound collection point RP and the position coordinates of the speaker SP1 to set the straight line passing through the sound collection point RP and the speaker SP1 ( Figure 17 (the dotted line in (A)). As Figure 17 (A) shows, the output speaker setting unit 5201 sets the azimuth angle φ that expands toward the speaker SP1 side with the sound collection point RP as the reference point in the plane with respect to this straight line ( Figure 17 (the dotted line in (A)). The azimuth angle φ is the angle formed in the horizontal direction with respect to the straight line passing through the sound collection point RP and the speaker SP1. In addition, as Figure 17 (B) shows, the output speaker setting unit 5201 sets the elevation angle θ that expands in the up - and - down direction orthogonal to the plane with respect to the above - mentioned straight line ( Figure 17 (the dotted line in (B)). The elevation angle θ is the angle formed in the vertical direction (the direction orthogonal to the horizontal direction) with respect to the straight line passing through the sound collection point RP and the speaker SP1.
[0174] The output speaker setting unit 5201 sets the space closer to the speaker SP1 side than the boundary determined by the azimuth angle φ and the elevation angle θ (the boundary surface determining the horizontal area, the boundary surface determining the vertical area) as the responsible area RGSP1 of the speaker SP1.
[0175] The output speaker setting unit 5201 acquires the position coordinates of a plurality of virtual sound sources IS (in the case of Figure 17 the case of multiple virtual sound sources ISa - ISg).
[0176] The output speaker setting unit 5201 uses the position coordinates of the plurality of virtual sound sources ISa - ISg and the coordinates representing the responsible area RGSP1 to determine whether the plurality of virtual sound sources ISa - ISg are within the responsible area RGSP1. This determination can be achieved by the same method as the grouping of the above sound sources into areas.
[0177] By performing this determination process, the output speaker setting unit 5201, for example, in Figure 14 A, Figure 14 B, Figure 14 the case shown in C, determines that the plurality of virtual sound sources ISa, ISb, ISc, ISd are within the responsible area RGSP1, and the plurality of virtual sound sources ISe, ISf, ISg are outside the responsible area RGSP1.
[0178] The output speaker setting unit 5201 assigns the plurality of virtual sound sources ISa, ISb, ISc, ISd determined to be within the responsible area RGSP1 to the speaker SP1 (S145).
[0179] The output speaker setting unit 5201 outputs the allocation information of the plurality of virtual sound sources for the plurality of speakers SP1 - SP64 to the coefficient setting unit 5202. At this time, the output speaker setting unit 5201 outputs the position coordinates of the sound collection point RP, the position coordinates of the plurality of speakers SP1 - SP64, the position coordinates of the plurality of virtual sound sources, and the allocation information to the coefficient setting unit 5202 together.
[0180] In addition, the azimuth angle φ is, for example, 60°, and the elevation angle θ is, for example, 45°. The angles of these azimuth angle φ and elevation angle θ are an example, and can also be set and adjusted, for example, by an operation input from the user.
[0181] The coefficient setting unit 5202 sets the tap coefficients given to LDtap521 - 528 using the distances between the sound collection point RP and the plurality of speakers SP1 - SP64, and the distances between the sound collection point RP and the virtual sound source IS. The tap coefficients given to LDtap521 - 528 are the gain values and delay amounts of LDtap521 - 528.
[0182] Figure 18 is a flowchart showing the coefficient setting process of LDtap. Figure 19 (A), Figure 19 (B) are diagrams for explaining the concept of coefficient setting.
[0183] The coefficient setting unit 5202 calculates the distances (speaker distances) between the sound collection point RP and the plurality of speakers SP1 - SP64 using the position coordinates of the sound collection point RP and the position coordinates of the plurality of speakers SP1 - SP64 (S151).
[0184] The coefficient setting unit 5202 calculates the distances (virtual sound source distances) between the sound collection point RP and the plurality of virtual sound sources IS (S152).
[0185] The coefficient setting unit 5202 compares the speaker distances and the virtual sound source distances for the plurality of speakers SP1 - SP64 and the plurality of virtual sound sources IS respectively assigned to these speakers SP1 - SP64 (S153). For example, if it is Figure 17 the example of (A), then for the speaker SP1 and the plurality of virtual sound sources ISa, ISb, ISc, ISd, the speaker distances and the virtual sound source distances are compared.
[0186] If the speaker distance is less than or equal to the virtual sound source distance (S153: YES), then the coefficient setting unit 5202 directly uses the virtual sound source distance to set the tap coefficient (S154).
[0187] For example, in the case as Figure 19 shown in (A), the virtual sound source Isa is farther from the sound collection point RP than the speaker SP1, and the virtual sound source distance Lia between the sound collection point RP and the virtual sound source Isa is greater than the speaker distance Ls1 between the sound collection point RP and the speaker SP1.
[0188] In this case, the coefficient setting unit 5202 uses the distance Da1 between the virtual sound source Isa and the speaker SP1 to set the tap coefficient. Specifically, the coefficient setting unit 5202 sets the gain value and the delay amount set for the virtual sound source Isa according to the distance Da1. For the coefficient setting unit 5202, the greater the distance Da1, the smaller the gain value is set, and the greater the distance Da1, the greater the delay amount is set.
[0189] If the speaker distance is greater than the virtual sound source distance (S153: NO), then the coefficient setting unit 5202 determines whether to play the virtual sound source. In other words, the coefficient setting unit 5202 determines whether to play the virtual sound source closer to the sound collection point than the speaker (S155).
[0190] If playback is performed for a virtual sound source closer to the sound collection point than the speaker (S155: YES), the coefficient setting unit 5202 moves the position of the virtual sound source (S156). More specifically, the coefficient setting unit 5202 moves the position of the virtual sound source on the sound collection point side of the speaker to a position farther from the sound collection point than the speaker. At this time, the coefficient setting unit 5202 uses the distance difference between the virtual sound source and the speaker to move the position of the virtual sound source. The coefficient setting unit 5202 sets the tap coefficient using the position coordinates of the moved virtual sound source (S157).
[0191] For example, in the case as Figure 19 (B) shows, the virtual sound source ISd is closer to the sound collection point RP than the speaker SP1, and the virtual sound source distance Lid between the sound collection point RP and the virtual sound source ISd is smaller than the speaker distance Ls1 between the sound collection point RP and the speaker SP1.
[0192] In this case, the coefficient setting unit 5202 moves the virtual sound source ISd using the distance difference Dd between the virtual sound source distance Lid and the speaker distance Ls1. More specifically, the coefficient setting unit 5202 moves the virtual sound source ISd to a position on the straight line passing through the sound collection point RP and the speaker SP1 and having a distance difference of Dd on the opposite side of the sound collection point RP side with the speaker SP1 as a reference. Moreover, the coefficient setting unit 5202 sets the tap coefficient using this distance difference Dd. Specifically, the coefficient setting unit 5202 sets the gain value and the delay amount set with respect to the virtual sound source ISd according to the distance difference Dd. For the coefficient setting unit 5202, the larger the distance difference Dd, the smaller the gain value is set, and the larger the distance difference Dd, the larger the delay amount is set. In addition, conceptually, the virtual sound source is moved as described above, but as a process of setting the tap coefficient, the coefficient setting unit 5202 only needs to set the tap coefficient according to the distance between the speaker distance and the virtual sound source distance.
[0193] That is, the coefficient setting unit 5202 only moves the virtual sound source located between the sound collection point and the speaker. In this regard, it is preferable that the virtual sound source outside the speaker with respect to the sound collection point does not move, but it also includes the case where the virtual sound source outside moves within a specified range. For example, even if the virtual sound source outside moves, as long as the distance between the virtual sound source outside and the speaker is within the specified range, the specified range means a range in which the change in the initial reflected sound control signal due to the movement does not cause discomfort to the audience. If playback is not performed for a virtual sound source closer to the sound collection point than the speaker (S155: NO), the coefficient setting unit 5202 does not set the tap coefficient for the virtual sound source.
[0194] The coefficient setting unit 5202 sets the tap coefficients set for each of the speakers SP1 - SP64 to a plurality of LDtaps. More specifically, the coefficient setting unit 5202 sets the tap coefficients to LDtap 521 for each of the speakers SP1 - SP64 based on the virtual sound source positions set in the area Area1. Similarly, the coefficient setting unit 5202 sets the tap coefficients of the virtual sound sources assigned to each of the speakers SP1 - SP64 to LDtaps 522 - 528 based on the virtual sound source positions set in the plurality of areas Area2 - Area8 respectively.
[0195] Corresponding to the set tap coefficients, the plurality of LDtaps 521 - 528 perform gain processing and delay processing on the area-separated sound signals SA1f - SA8f after the filtering process and output them to the addition processing unit 53. More specifically, as described above, the tap coefficients are set corresponding to the combinations of the virtual sound source positions in the plurality of areas and each speaker. Therefore, the plurality of LDtaps 521 - 528 set the tap coefficients based on the virtual sound sources assigned to each speaker for each speaker. The plurality of LDtaps 521 - 528 perform gain processing and delay processing on the area-separated sound signals SA1f - SA8f after the filtering process for each speaker. The plurality of LDtaps 521 - 528 output the signals on which the gain processing and delay processing have been performed to each speaker.
[0196] For example, in the case where the virtual sound sources ISa, ISb, ISc, and ISd are assigned to the speaker SP1, LDtap 521 performs gain processing and delay processing on the area-separated sound signal SA1f after the filtering process by the tap coefficients (gain values and delay amounts) based on the virtual sound sources ISa, ISb, ISc, and ISd. Moreover, LDtap 521 outputs this signal as the one for the speaker SP1 to the addition processing unit 53. The plurality of LDtaps 522 - 528 perform the above-described processing for the virtual sound sources for which the tap coefficients are set.
[0197] The addition processing unit 53 adds the signals after the LDtap processing for each of the plurality of speakers SP1 - SP64 output from the plurality of LDtaps 521 - 528 for each of the plurality of speakers SP1 - SP64. The addition processing unit 53 outputs these added signals as the initial reflection sound control signals ER1 - ER64 for each of the plurality of speakers SP1 - SP64 to the adder 80.
[0198] By performing the above-described processing, the initial reflection sound control signal generation unit 50 can generate an initial reflection sound control signal having the following characteristics.
[0199] Figure 20 (A), Figure 20(B) is a waveform diagram showing an example of the relationship between the shape of the virtual space and the components of the initial reflected sound control signal implemented by LDtap. Figure 20 (A) shows the case where the virtual space shape is large. Figure 20 (B) shows the case where the virtual space shape is small. In addition, Figure 20 (A), Figure 20 (B) shows an example of the components of the initial reflected sound control signal when multiple virtual sound sources are set for one speaker.
[0200] When the positional relationship between the playback space and the virtual space does not change, and the positions of the sound pickup point and the speakers do not change, if the virtual space shape is large, the distribution of the virtual sound sources extends over a wider range compared to when the virtual space shape is small. Therefore, as Figure 20 (A), Figure 20 (B) shows, when the virtual space shape is large, the components set in LDtap521 - 528 tend to become smaller, and the distribution range on the time axis also becomes wider.
[0201] As described above, by performing the above processing, the initial reflected sound control signal generation unit 50 can set the optimal tap coefficients corresponding to the shape of the virtual space.
[0202] Moreover, even when the positional relationship between the virtual space and the playback space changes, the speaker position changes, or the sound pickup point changes, the initial reflected sound control signal generation unit 50 can set the optimal tap coefficients corresponding to these changes in the same way as in the case of the change in the shape of the virtual space.
[0203] At this time, the multiple sound sources OBJ1 - OBJ96 are optimally allocated to the multiple speakers SP1 - SP64 through grouping based on the multiple regions Area1 - Area8. Moreover, the multiple virtual sound sources are optimally set with respect to the multiple speakers SP1 - SP64. Therefore, for the sound signal processing device 10, even when there are changes in the relationship between the virtual space and the playback space, the position of the sound pickup point RP, the positions of the multiple speakers SP1 - SP64, and the positions of the sound sources OBJ1 - OBJ96, the sound image localization based on the initial reflected sound can be made clear corresponding to these changes.
[0204] In addition, in the above structure, even if the virtual sound source IS is closer to the sound collection point RP than the speaker SP, the initial reflected sound control signal generation unit 50 can approximately reproduce the components of the initial reflected sound control signal obtained based on the virtual sound source IS. Therefore, for example, when the number of set virtual sound sources is small relative to the initial reflected sound control signal, etc., the initial reflected sound control signal generation unit 50 can use a virtual sound source closer to the sound collection point RP than the speaker SP. At this time, as described above, the initial reflected sound control signal generation unit 50 uses the distance difference between the virtual sound source IS and the speaker SP to reposition the virtual sound source outside the speaker. As described above, the initial reflected sound control signal generation unit 50 can suppress the discomfort of the initial reflected sound caused by moving the position of the virtual sound source.
[0205] In addition, in the above structure, when the virtual sound source IS is in a position closer to the sound collection point RP than the speaker SP, the initial reflected sound control signal generation unit 50 may also set the virtual sound source IS to the position of the speaker SP. As described above, the initial reflected sound control signal generation unit 50 can reduce the load of the process of moving the virtual sound source IS.
[0206] Also, in the above structure, when the virtual sound source IS is in a position closer to the sound collection point RP than the speaker SP, the initial reflected sound control signal generation unit 50 may not use the virtual sound source IS for generating the initial reflected sound control signal. As described above, the initial reflected sound control signal generation unit 50 can reduce the load of the generation process of the initial reflected sound control signal without the load of the process of moving the virtual sound source IS.
[0207] In addition, in the above structure, the initial reflected sound control signal generation unit 50 performs setting of the components of the initial reflected sound control signal obtained based on the virtual sound source and tone color adjustment using the FIR filters 511 - 518. The FIR filters 511 - 518 have the above-mentioned number of taps (for example, 16,000 taps), and have a larger number of taps than the LD taps 521 - 528. In addition, the time interval of the taps of the FIR filters 511 - 518 (depending on the sampling frequency) is shorter than the time interval between the taps of the LD taps 521 - 528 (depending on the configuration of the virtual sound source). Therefore, the components of the initial reflected sound control signal generated by the FIR filters 511 - 518 are arranged more densely on the time axis compared to the components of the initial reflected sound control signal generated by the LD taps 521 - 528. In other words, the resolution on the time axis (time resolution) of the FIR filters 511 - 518 is higher than that of the LD taps 521 - 528, and the number of components per unit time increases.
[0208] Moreover, the initial reflected sound control signal generation unit 50 multiplies the processing of the FIR filters 511 - 518 by the LDtaps 521 - 528. Therefore, the initial reflected sound control signal generation unit 50 can generate initial reflected sound control signals ER1 - ER64 with high resolution on the time axis and more diverse timbres. Figure 21 It is a diagram showing an overview of the waveform of the initial reflected sound control signal generated by the initial reflected sound control signal generation unit 50.
[0209] As Figure 21 shown, the initial reflected sound control signal generation unit 50 can generate an initial reflected sound control signal that retains the initial reflected sound component obtained based on the virtual sound source and can handle higher resolution and more diverse timbres. That is, the sound signal processing device 10 can realize an initial reflected sound that ensures clear sound image localization by the initial reflected sound using the virtual sound source and has a timbre that matches the user's preference.
[0210] In addition, since the FIR filter has high resolution, for example, in the case of a short sound source such as a pulse sound, based only on the initial reflected sound component obtained from the LDtap, sometimes the initial reflected sound control signal becomes rough and the timbre becomes unnatural. However, with the above structure and processing, the sound signal processing device 10 can suppress such roughness of the initial reflected sound and unnatural timbre.
[0211] In addition, in the above structure, the initial reflected sound control signal generation unit 50 sets the responsible area of the virtual sound source IS for each speaker SP and does not assign the virtual sound source IS outside this area to the speaker SP. As described above, the initial reflected sound control signal generation unit 50 can suppress the excessive generation of the initial reflected sound component. Therefore, the sound signal processing device 10 can suppress the excessive generation of the initial reflected sound and realize a more natural initial reflected sound that matches the virtual space.
[0212] [Generation of reverberation sound control signal]
[0213] Figure 22 It is a functional block diagram showing an example of the structure of the reverberation sound control signal generation unit 70. Figure 23 It is a flowchart showing an example of the generation process of the reverberation sound control signal.
[0214] As Figure 22 shown, the reverberation sound control signal generation unit 70 includes a PEQ 71, an FIR filter circuit 72, a distributor 73, a reverberation sound area setting unit 701, a filter coefficient setting unit 702, a reverberation sound playback speaker setting unit 703, and an operation unit 700. The FIR filter circuit 72 includes a plurality of FIR filters 721 - 728.
[0215] The reverberation sound use area setting unit 701 sets a plurality of reverberation sound use areas Arr1 - Arr8 for the playback space. More specifically, the reverberation sound use area setting unit 701 sets, for example, with the center point Psr of the playback space as a reference, in the entire circumferential range on the plane, the playback space is divided into a plurality of reverberation sound use areas Arr1 - Arr8 (refer to Figure 25 ).
[0216] The reverberation sound use area setting unit 701 outputs the coordinate information indicating the plurality of reverberation sound use areas Arr1 - Arr8 to the filter coefficient setting unit 702 and the reverberation sound use playback speaker setting unit 703.
[0217] The filter coefficient setting unit 702 sets the filter coefficients for the reverberation sound through user operations or the like. The filter coefficients for the reverberation sound are set, for example, based on the measured results of the impulse responses of different spaces (virtual spaces) reproduced in the playback space. In addition, the filter coefficients for the reverberation sound can also be approximately set using the geometric shape of the virtual space, the material of the wall surface, etc. At this time, the filter coefficient setting unit 702 uses the coordinate information of each reverberation sound use area Arr1 - Arr8 to set the filter coefficients for each reverberation sound use area Arr1 - Arr8.
[0218] The filter coefficient setting unit 702 receives the input of the volume of the virtual space, the surface area of the virtual space, etc. through user operations or the like. The filter coefficient setting unit 702 sets the fade - in function for the filter coefficients based on parameters such as the volume of the virtual space and the surface area of the virtual space.
[0219] More specifically, the filter coefficient setting unit 702 calculates the mean free path ρ using the volume V of the virtual space and the surface area S of the virtual space. The calculation formula for the mean free path ρ is ρ = 4V / S. The mean free path refers to the average transmission distance that sound travels in a closed space from one reflection on the wall surface to the next reflection. By dividing the mean free path by the speed of sound c0, the average time required for sound to be reflected on the wall surface until the next reflection can be calculated.
[0220] The filter coefficient setting unit 702 sets the connection timing tc ( Figure 23 : S231) according to the mean free path ρ. Specifically, the filter coefficient setting unit 702 sets the connection timing tc using the mean free path ρ, the speed of sound c0, and the number of reflections n. The calculation formula for the connection timing tc is tc = ρ×n / c0.
[0221] As can be seen from this calculation formula, the connection timing tc is equivalent to the average time required for n reflections in the virtual space. When reproducing the initial reflected sound n times, it is equivalent to the moment when the conversion to the reverberant sound starts. In other words, the connection timing tc corresponds to the timing when the component of the initial reflected sound control signal obtained by the above-described initial reflected sound control signal generation unit 50 disappears.
[0222] By performing such processing, the filter coefficient setting unit 702 can optimally set the connection timing tc between the initial reflected sound and the reverberant sound according to the geometric shape of the virtual space.
[0223] The filter coefficient setting unit 702 sets the fade-in function according to the following formula using the connection timing tc ( Figure 23 : S232).
[0224]
Equation 1
[0225]
[0226] In addition, in this formula, t is the elapsed time since the direct sound occurred, and K is set according to the following formula.
[0227]
Equation 2
[0228]
[0229] In addition, in this formula, G REV is the gain value of the reverberant sound at time t = 0, which can be set by the user. For example, since the normal reverberation time is the time required to decay to -60 dB, it can be set to G REV = -60 dB, etc.
[0230] The filter coefficient setting unit 702 sets the filter coefficient for the reverberant sound according to the filter coefficient and the fade-in function fin ( Figure 23 : S233) and outputs it to the plurality of FIR filters 721 - 728.
[0231] The reverberant sound generation signal Sr output from the mixer 60 is input to the PEQ 71. The PEQ 71 performs prescribed signal processing on the reverberant sound generation signal Sr and outputs it to the plurality of FIR filters 721 - 728.
[0232] By performing signal processing by the PEQ 71, it is possible to adjust the level (size of the signal), timbre, etc. of the reverberation sound generation signal Sr. For example, the PEQ 71 can refer to the volume of the initial reflected sound control signal, etc., and adjust the level (size of the signal) of the reverberation sound generation signal Sr in such a way that the volume of the initial reflected sound and the volume of the reverberation sound become the same level at the above-mentioned connection timing tc. In addition, the PEQ 71 can adjust the timbre, etc. according to the settings of the user or the like.
[0233] The plurality of FIR filters 721 - 728 perform filtering processing on the reverberation sound generation signal Sr using the reverberation sound filtering coefficients, and generate the regionally divided reverberation sound control signals REVr1 - REVr8. For example, the FIR filter 721 performs a convolution operation on the reverberation sound generation signal Sr using the reverberation sound filtering coefficient set for the region Arr1 of the reverberation sound, thereby generating the regionally divided reverberation sound control signal REVr1 for the region Arr1. Similarly, the FIR filters 722 - 728 respectively perform a convolution operation on the reverberation sound generation signal Sr using the reverberation sound filtering coefficients set for the regions Arr2 - Arr8 of the reverberation sound, thereby generating the regionally divided reverberation sound control signals REVr2 - REVr8 ( Figure 23 : S234). The plurality of FIR filters 721 - 728 output the regionally divided reverberation sound control signals REVr1 - REVr8 to the distributor 73.
[0234] By setting the above-mentioned fade-in function, the reverberation sound control signal becomes as Figure 24 shown in the waveform. Figure 24 is a graph showing waveform examples of the direct sound, the initial reflected sound control signal, and the reverberation sound control signal. In addition, in Figure 24 for convenience, the reverberation sound control signal is illustrated by the envelope of each time component. In addition, Figure 24 the vertical axis represents dB.
[0235] As Figure 24 shown, the reverberation sound control signal gradually increases in signal level along with the fade-in function in the range from the output timing of the direct sound to the connection timing tc. More specifically, the signal level of the reverberation sound control signal is -60 dBFs at the output timing of the direct sound, gradually increases until the connection timing tc, and becomes 0 dBFs at the connection timing tc. This level is set based on the signal level at the connection timing tc of the initial reflected sound control signal.
[0236] In Figure 24In the example, the above-mentioned fade-in function is used to exponentially increase the signal level as it approaches the connection timing tc. In other words, the above-mentioned fade-in function has characteristics opposite to the attenuation curve of the reverberation control signal without fade-in processing. In addition, the characteristics of the change in the level of the reverberation control signal obtained by the fade-in processing are not limited to this, and can also be set to characteristics desired by the user or the like by appropriately setting the fade-in function.
[0237] By performing the above-mentioned processing, the reverberation control signal generation unit 70 can generate a reverberation control signal that accurately reproduces the reverberation in the virtual space using the FIR filters 721 - 728. In addition, for the reverberation control signal, in the section where there is an initial reflection sound control signal, the signal level gradually increases, reaches a peak corresponding to the signal level of the initial reflection sound control signal at the connection timing tc, and then attenuates.
[0238] As described above, the sound signal processing device 10 can make the connection between the initial reflection sound control signal and the reverberation control signal generated by the multiple LDtaps that reproduce the virtual sound source distribution at multiple sound source positions in the virtual space smooth through the reverberation obtained based on the reverberation control signal. Therefore, the sound output from the sound signal processing device 10 and heard by the user becomes a sound that suppresses the discomfort when connecting from the initial reflection sound to the reverberation.
[0239] The reverberation playback speaker setting unit 703 groups the multiple speakers SP1 - SP64 into the reverberation regions Arr1 - Arr8.
[0240] More specifically, the reverberation playback speaker setting unit 703 is set, for example, with the center point Psr of the playback space as a reference, and divides the playback space into multiple reverberation regions Arr1 - Arr8 in the entire circumferential range on the plane. The reverberation playback speaker setting unit 703 groups the multiple speakers SP1 - SP64 for the multiple reverberation regions Arr1 - Arr8 using the position coordinates of the multiple speakers SP1 - SP64 and the coordinate information indicating the multiple reverberation regions Arr1 - Arr8. This grouping can be achieved by the same method as the method for grouping the above-mentioned sound source OBJ.
[0241] Figure 25 It is a diagram showing an example of the area setting for reverberation. In Figure 25 For simplicity of explanation and easy understanding, multiple speakers SP1 - SP14 are shown. For example, the reverberation playback speaker setting unit 703 is as Figure 25As shown, it is detected that the speakers SP6 and SP7 are present in the reverberation sound area Arr1, and the speakers SP6 and SP7 are grouped into the reverberation sound area Arr1. Similarly, the reverberation sound playback speaker setting unit 703 also groups the other speakers SP1 - SP5, SP8 - SP14 into multiple reverberation sound areas Arr2 - Arr8 respectively.
[0242] The reverberation sound playback speaker setting unit 703 outputs the grouping information of the multiple speakers SP1 - SP64 for the multiple reverberation sound areas Arr2 - Arr8 to the distributor 73.
[0243] The distributor 73 uses the grouping information from the reverberation sound playback speaker setting unit 703 to distribute the area - differentiated reverberation sound control signals REVr1 - REVr8 to the multiple speakers SP1 - SP64. Based on the distribution, the distributor 73 outputs the area - differentiated reverberation sound control signals REVr1 - REVr8 as the reverberation sound control signals REV1 - REV48 for the respective multiple speakers SP1 - SP64.
[0244] For example, the distributor 73 extracts the case where the speakers SP6 and SP7 are grouped in the area Arr1 according to the grouping information. The distributor 73 distributes the area - differentiated reverberation sound control signal REVr1 of the area Arr1 to the speakers SP6 and SP7. The distributor 73 outputs the area - differentiated reverberation sound control signal REVr1 as the reverberation sound control signal REV6 for the speaker SP6 to the speaker SP6. In addition, the distributor 73 outputs the area - differentiated reverberation sound control signal REVr1 as the reverberation sound control signal REV7 for the speaker SP7 to the speaker SP7.
[0245] By the distributor 73 performing the distribution process of the area - differentiated reverberation sound control signals REVr1 - REVr8 for each area as described above, the reverberation sound control signal generation unit 70 can output the optimal reverberation sound control signals to the multiple speakers SP1 - SP64 respectively corresponding to the configuration of the multiple speakers SP1 - SP64.
[0246] [Output adjustment]
[0247] Figure 26 It is a functional block diagram showing an example of the structure of the output adjustment unit 90. Figure 27 It is a flowchart showing an example of the output adjustment process.
[0248] As Figure 26As shown, the output adjustment unit 90 includes a gain control unit 91, a hysteresis control unit 92, a gain and hysteresis setting unit 901, an operation unit 900, and a display unit 909. The gain control unit 91 includes a plurality of gain control units 9101 - 9168 corresponding to a plurality of speakers SP1 - SP64. The hysteresis control unit 92 includes a plurality of hysteresis control units 9201 - 9264 corresponding to a plurality of speakers SP1 - SP64.
[0249] The operation unit 900 receives the setting of the acoustic parameters of the playback space through an operation input from the user ( Figure 27 : S321). The acoustic parameters of the playback space are parameters for reproducing a desired sound field in the playback space.
[0250] At this time, the acoustic parameters of the playback space are not the gain values or delay amounts of the plurality of speakers SP1 - SP64 respectively, but the weight (Weight) values representing the weighting of the sound in the playback space in a specified direction and the shape (Shape) values representing the expansion of the specified direction of the sound in the playback space.
[0251] The weight value is composed of a gain value and a delay amount, and includes the weight values in the front and back of the playback space, the weight values on the left and right of the playback space, and the weight values in the up and down directions of the playback space. The shape value is composed of a gain value and a delay amount, and includes the shape value in the horizontal direction.
[0252] The display unit 909 has a GUI. Figure 28 It is a diagram showing an example of the GUI for output adjustment.
[0253] As Figure 28 shown, the GUI 100A has a setting display window 111, an output status display window 115, and a plurality of operation members 116. The plurality of operation members 116 have a knob 1161 and an adjustment value display window 1162.
[0254] The plurality of operation members 116 are operation members for setting the weight volume for setting the weight value, the shape volume for setting the shape value, etc. The operation members 116 for the weight volume each have operation members 116 for setting the left - right weight, the front - back weight, and the up - down weight, and each have an operation member for setting the gain value and an operation member for setting the delay amount. The operation members 116 for the shape volume have an operation member for setting the expansion, an operation member for setting the gain value, and an operation member for setting the delay amount.
[0255] The output status display window 115 graphically and schematically displays the expansion and sense of localization of the sound achieved through the weight value and shape value set by the plurality of operation members 116. Thus, the user can easily recognize the expansion and sense of localization of the sound set by the plurality of operation members 116 as an image.
[0256] The user uses the GUI 100A of the display unit 909 to set the audio parameters (weight values and delay amounts) that the user wants to reproduce. The operation unit 900 receives the settings made using the GUI 100A. The operation unit 900 outputs the setting content (each weight value and each delay amount of the audio parameters) to the gain and lag setting unit 901.
[0257] The gain and lag setting unit 901 sets the gain values and delay amounts for the multiple speakers SP1 - SP64 based on each weight value and each delay amount of the audio parameters. More specifically, the gain and lag setting unit 901 performs the following processing.
[0258] The gain and lag setting unit 901 obtains the position coordinates (S322) of the multiple speakers SP1 - SP64 arranged in the playback space. The position coordinates are represented, for example, in the following coordinate system, in which the x - axis is set in the left - right direction of the playback space, the y - axis is set in the front - back direction of the playback space, and the z - axis is set in the up - down direction.
[0259] The gain and lag setting unit 901 extracts the maximum and minimum values of the position coordinates of the multiple speakers SP1 - SP64 in each axis direction (S323).
[0260] The gain and lag setting unit 901 stores coefficient setting formulas. The coefficient setting formulas include, for example, a coefficient setting formula for weights for setting the weighting in a specified direction in the playback space and a coefficient setting formula for shapes for setting the weighting in a specified direction in the playback space.
[0261] The coefficient setting formula for weights includes a gain - value setting formula for weights and a delay - amount setting formula for weights. The coefficient setting formula for shapes includes a gain - value setting formula for shapes and a delay - amount setting formula for shapes.
[0262] The coefficient setting formula for weights includes a coefficient setting formula for the front - back direction for setting the weighting in the front - back direction of the playback space, a coefficient setting formula for the left - right direction for setting the weighting in the left - right direction of the playback space, and a coefficient setting formula for the up - down direction for setting the weighting in the up - down direction of the playback space.
[0263] The coefficient setting formula for shapes includes a coefficient setting formula for the left - right direction of the playback space.
[0264] The coefficient setting formula for the gain value of weights is, for example, a linear function obtained by combining the gain value of the set weight value, the maximum and minimum values of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for setting the gain value, and is a formula that determines the gain value in proportion to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.
[0265] The coefficient setting formula for the delay amount used for the weight is, for example, a linear function obtained by combining the delay amount of the set weight value, the maximum and minimum values of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the delay amount is set, and is a formula that determines the delay amount in direct proportion to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.
[0266] The coefficient setting formula for the gain value used for the shape is, for example, a linear function obtained by combining the gain value of the set shape value, the maximum and minimum values of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the gain value is set, and is a formula that determines the gain value in direct proportion to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.
[0267] The coefficient setting formula for the delay amount used for the shape is, for example, a linear function obtained by combining the delay amount of the set shape value, the maximum and minimum values of the extracted position coordinates, and the position coordinates of the speaker (the speaker to be set) for which the delay amount is set, and is a formula that determines the delay amount in direct proportion to the difference between the position coordinates of the speaker to be set and the minimum value of the position coordinates.
[0268] The gain and delay setting unit 901 calculates the gain value and the delay amount (acoustic parameters) for each speaker to be set by using the set gain value and delay amount, the maximum and minimum values of the extracted position coordinates, and the coefficient setting formula (S324).
[0269] By using the above-described processing, the gain and delay setting unit 901 can automatically calculate and set the gain value and the delay amount for the multiple speakers SP1 - SP64 arranged in the playback space without manually setting them individually.
[0270] The gain and delay setting unit 901 outputs the gain values set for the multiple speakers SP1 - SP64 to the multiple gain control units 9101 - 9164. The gain and delay setting unit 901 outputs the delay amounts set for the multiple speakers SP1 - SP64 to the multiple delay control units 9201 - 9264.
[0271] Speaker signals Sat1 - Sat64 corresponding to the multiple speakers SP1 - SP64 are input to the multiple gain control units 9101 - 9164 from the adder 80, respectively.
[0272] A plurality of gain control units 9101 - 9164 respectively use the set gain values to control the signal levels of the speaker signals Sat1 - Sat64 and output them to a plurality of hysteresis control units 9201 - 9264. For example, the gain control unit 9101 uses the gain value set for the gain control unit 9101 to control the signal level of the speaker signal Sat1 and outputs it to the hysteresis control unit 9201. Similarly, the gain control units 9102 - 9164 use the gain values set for the gain control units 9102 - 9164 respectively to control the signal levels of the speaker signals Sat2 - Sat64 and output them to the hysteresis control units 9202 - 9264 respectively.
[0273] A plurality of hysteresis control units 9201 - 9264 use the set delay amounts to control the signal levels of the signals input from the plurality of gain control units 9101 - 9164 and output them to a plurality of speakers SP1 - SP64. For example, the hysteresis control unit 9201 uses the delay amount set for the hysteresis control unit 9201 to control the signal level of the signal input from the gain control unit 9101 and outputs it to the speaker SP1. Similarly, the hysteresis control units 9202 - 9264 use the delay amounts set for the hysteresis control units 9202 - 9264 respectively to control the signal levels of the signals input from the gain control units 9102 - 9164 and output them to the speakers SP2 - SP64 respectively.
[0274] With the above structure, the sound signal processing device 10 does not require the user to be good at complex settings for each of the multiple speakers, and can easily achieve a desired sound field corresponding to the set acoustic parameters by using the initial reflected sound control signal and the reverberation sound control signal. As described above, for example, the sound signal processing device 10 can easily achieve a sound field that obtains the Haas effect for a specified position in the playback space.
[0275] (Example of realizing the sound field based on output control)
[0276] Figure 29 (A), Figure 29 (B) are diagrams showing a setting example in which there is localization and expansion at the rear of the playback space. Figure 29 (A) is a diagram showing an example of the setting of the gain value and the delay amount, Figure 29 (B) shows based on Figure 29 (A) is a diagram showing an overview of the weighting of the sound achieved by the setting. In addition, in Figure 29 (A), Figure 29 (B), for easy understanding with a simplified explanation, shows the case where 14 speakers SP1 - SP14 are arranged.
[0277] InFigure 29 (A), Figure 29 In the manner shown in (B), as acoustic parameters, for example, the gain value and the delay amount at the rear end are set. The gain and lag setting unit 901 sets the gain value and the delay amount at the front end to values with opposite signs of the gain value and the delay amount at the rear end. The gain and lag setting unit 901 calculates the maximum and minimum values of the position coordinates of the 14 speakers SP1 - SP14.
[0278] The gain and lag setting unit 901 calculates the gain values of the 14 speakers SP1 - SP14 using the gain values at the rear end and the front end, the maximum and minimum values of the position coordinates of the 14 speakers SP1 - SP14, and the coefficient setting formula (for gain value setting) for the front - rear direction set by weighting the front - rear direction of the playback space.
[0279] In addition, the gain and lag setting unit 901 calculates the delay amounts of the 14 speakers SP1 - SP14 using the delay amounts at the rear end and the front end, the maximum and minimum values of the position coordinates of the 14 speakers SP1 - SP14, and the coefficient setting formula (for delay amount setting) for the front - rear direction set by weighting the front - rear direction of the playback space.
[0280] Through this process, the sound signal processing device 10, as shown in Figure 29 (A), can automatically and easily set the acoustic parameters such that the gain value and the delay amount are larger for the speakers at the rear of the playback space and smaller for the speakers at the front. Thus, the sound signal processing device 10 can easily achieve a sound field with expandability at the rear of the playback space and sound localization (refer to Figure 29 (B)).
[0281] In addition, in this description, an example of the front - rear direction is shown, but the sound signal processing device 10 can also similarly achieve a weighted sound field for the left - right direction and the height direction (up - down direction).
[0282] Figure 30 (A), Figure 30 (B) are diagrams showing a setting example of the case where sound expands horizontally in the playback space. Figure 30 (A) is a diagram showing an example of the setting of the gain value and the delay amount, Figure 30 (B) is a diagram showing an overview of the sound expansion achieved based on the setting in Figure 30 (A). In addition, in Figure 30 (A), Figure 30 (B), as is easily understood from the simplified description, the case where 14 speakers SP1 - SP14 are arranged is shown.
[0283] In Figure 30 (A), Figure 30 In the manner shown in (B), as acoustic parameters, for example, a value obtained by digitizing the sound expansion (expansion setting value) is set. The gain and hysteresis setting unit 901 calculates the maximum and minimum values of the position coordinates of the 14 speakers SP1 - SP14.
[0284] The gain and hysteresis setting unit 901 calculates the gain values of the 14 speakers SP1 - SP14 using the value obtained by digitizing the sound expansion, the maximum and minimum values of the position coordinates of the 14 speakers SP1 - SP14, and the coefficient setting formula for shape (for gain value setting).
[0285] In addition, the gain and hysteresis setting unit 901 calculates the delay amounts of the 14 speakers SP1 - SP14 using the delay amounts of the rear end and the front end, the maximum and minimum values of the position coordinates of the 14 speakers SP1 - SP14, and the coefficient setting formula for shape (for delay amount setting).
[0286] Through this processing, as shown in Figure 30 (A), the sound signal processing device 10 can easily set the acoustic parameters such that the gain value and the delay amount are larger for the speakers closer to the lateral ends of the playback space, and smaller for the speakers closer to the lateral center. As described above, the sound signal processing device 10 can easily achieve a sound field with expandability in the lateral direction of the playback space and sound localization (refer to Figure 30 (B)).
[0287] In addition, by setting the above - mentioned acoustic parameters, the sound signal processing device 10 can not only achieve the weighting in the front - rear direction, the weighting in the left - right direction, and the expansion in the lateral direction of the playback space, but also achieve the weighting and expansion in the height direction (up - down direction) of the playback space. For example, Figure 31 is a diagram showing an overview of the sound expansion in the case of having expansion in the height direction.
[0288] The sound signal processing device 10 makes the gain value and the delay amount of the ceiling - side speaker SPU larger than those of the speakers SPL and SPR closer to the floor surface. As described above, the sound signal processing device 10 can easily achieve a sound field with more expandability in the ceiling direction of the playback space and echo localization (refer to Figure 31 ).
[0289] In addition, in the above structure, the output adjustment unit 90 outputs the output signals So1 - So64 to the plurality of speakers SP1 - SP64. However, the sound signal processing device may also perform binaural processing on the output signals So1 - So64 and then output them.
[0290] Figure 32 It is a functional block diagram showing the structure of a sound signal processing device with a binaural playback function. As Figure 32 shown, the sound signal processing device 10A with a binaural playback function is different from the above-mentioned sound signal processing device 10 in that it has an output adjustment unit 90A, a reverberation processing unit 97, a selection unit 98, and a binaural processing unit 99.
[0291] The output adjustment unit 90A generates a plurality of output signals So1 - So64 using the same processing as the above-mentioned output adjustment unit 90 based on the plurality of speaker signals Sat1 - Sat64 output from the adder 80.
[0292] The output adjustment unit 90A can select the output target. The selection of the output target is performed, for example, by an operation input from the user using the above-mentioned GUI. More specifically, the GUI displays an operation element that can select speaker output and binaural output, and the output target is selected by operating this operation element.
[0293] When speaker output is selected, the output adjustment unit 90A outputs the plurality of output signals So1 - So64 to the plurality of speakers SP1 - SP64 respectively (the same processing as the output adjustment unit 90). When binaural output is selected, the output adjustment unit 90A outputs the plurality of output signals So1 - So64 to the selection unit 98.
[0294] The reverberation processing unit 97 is input with the sound signals S1 - S96 of the plurality of sound sources OBJ1 - OBJ96. The reverberation processing unit 97 adds an initial reflection sound control signal and a reverberation sound control signal to the plurality of sound signals S1 - S96 and outputs them to the selection unit 98. The initial reflection sound control signal for the plurality of sound signals S1 - S96 is set based on the position coordinates of the plurality of sound sources OBJ1 - OBJ96. The reverberation processing unit 97 outputs the plurality of reverberation-processed sound signals S1' - S96' to the selection unit 98.
[0295] Multiple output signals So1 - So64 and multiple reverberation-processed sound signals S1' - S96' are input to the selection unit 98. The selection unit 98 selects the multiple output signals So1 - So64 and the reverberation-processed sound signals S1' - S96' by, for example, operation input from the user using the above-described GUI. More specifically, the GUI displays an operation element that enables selection between the sound subjected to the audio processing of the sound signal processing device 10A and the sound subjected to virtual audio processing based on the position coordinates of the sound sources OBJ1 - OBJ96, and the output target is selected by operating this operation element.
[0296] When the sound subjected to the audio processing of the sound signal processing device 10A is selected, the selection unit 98 selects the multiple output signals So1 - So64 and outputs them to the binaural processing unit 99. When the sound subjected to the virtual audio processing based on the position coordinates of the sound sources OBJ1 - OBJ96 is selected, the selection unit 98 selects the multiple reverberation-processed sound signals S1' - S96' and outputs them to the binaural processing unit 99.
[0297] The binaural processing unit 99 performs binaural processing on the input sound signal. More specifically, if multiple output signals So1 - So64 are input, the binaural processing unit 99 performs binaural processing on the multiple output signals So1 - So64. If multiple reverberation-processed sound signals S1' - S96' are input, the binaural processing unit 99 performs binaural processing on the multiple reverberation-processed sound signals S1' - S96'.
[0298] In addition, binaural processing is processing using the head-related transfer function, and the details are known, so the detailed description of binaural processing is omitted.
[0299] The binaural processing unit 99 outputs the binaural-processed two-channel sound signal.
[0300] As described above, the user can hear the sound generated by the sound signal processing device 10A and the sound subjected to virtual reverberation processing based on the position coordinates of the sound sources OBJ1 - OBJ96 through binaural playback. Therefore, even without physically constructing a playback space, the user can easily confirm whether the audio processing performed by the sound signal processing device 10A can reproduce the audio in the virtual space using headphones or the like. The audio processing performed by the sound signal processing device 10A is, for example, the above-described grouping of sound sources, setting of the initial reflection sound control signal, setting of the reverberation sound control signal, setting of the output control, etc. Moreover, by making the above-described audiovisual comparison, the user can adjust the setting of the above-described audio processing to more realistically reproduce the audio in the virtual space.
[0301] In addition, binaural playback is not limited to headphones and can also be performed through stereo speakers or the like.
[0302] The description of the present embodiment is illustrative in all aspects and is not restrictive. The scope of the present invention is represented not by the above-described embodiments but by the claims. Moreover, the scope of the present invention includes meanings equivalent to the claims and all modifications within the scope.
[0303] Description of Reference Numerals
[0304] 10, 10A: Sound signal processing device
[0305] 30: Region setting unit
[0306] 40: Grouping unit
[0307] 41: Sound source position detection unit
[0308] 42: Region determination unit
[0309] 50: Initial reflection sound control signal generation unit
[0310] 51: FIR filter circuit
[0311] 52: LDtap circuit
[0312] 53: Addition processing unit
[0313] 60: Mixer
[0314] 70: Reverberation sound control signal generation unit
[0315] 71: PEQ
[0316] 72: FIR filter circuit
[0317] 73: Distributor
[0318] 80: Adder
[0319] 90, 90A: Output adjustment unit
[0320] 91: Gain control unit
[0321] 92: Hysteresis control unit
[0322] 97: Reverberation processing unit
[0323] 98: Selection unit
[0324] 99: Binaural processing unit
[0325] 100, 100A: GUI
[0326] 400: Matrix mixer
[0327] 500: Operation Unit
[0328] 501: Tone Color Setting Unit
[0329] 502: Virtual Sound Source Setting Unit
[0330] 511 - 518: FIR Filter
[0331] 521 - 528: LDtap
[0332] 700: Operation Unit
[0333] 701: Reverberation Sound Area Setting Unit
[0334] 702: Filter Coefficient Setting Unit
[0335] 703: Reverberation Sound Playback Speaker Setting Unit
[0336] 721 - 728: FIR Filter
[0337] 900: Operation Unit
[0338] 901: Gain and Hysteresis Setting Unit
[0339] 909: Display Unit
[0340] 5201: Output Speaker Setting Unit
[0341] 5202: Coefficient Setting Unit
[0342] 9101 - 9164: Gain Control Unit
[0343] 9201 - 9264: Hysteresis Control Unit.
Claims
1. A method for processing a sound signal, wherein, it is determined whether a virtual sound source representing a reflected sound of a target acoustic space is within a region having a prescribed azimuth angle and elevation angle with respect to the center position of the acoustic space and the position of a loudspeaker, and only when it is within the region, output control of an initial reflected sound control signal of the virtual sound source is performed by the loudspeaker.
2. The method for processing a sound signal according to claim 1, wherein, the virtual sound source is set based on the shape of the acoustic space, the position of the loudspeaker in the acoustic space, and the position of a sound pickup point.
3. The method for processing a sound signal according to claim 1 or 2, wherein, a gain value and a delay amount of the initial reflected sound control signal are set based on the positional relationship among the virtual sound source, the sound pickup point, and the loudspeaker.
4. The method for processing a sound signal according to any one of claims 1 to 3, wherein, the virtual sound source is classified into a first virtual sound source and a second virtual sound source. The first virtual sound source is located between the loudspeaker and the sound pickup point and represents a first sound source of the reflected sound of the target acoustic space, and the second virtual sound source is located outside the loudspeaker and represents a second sound source of the reflected sound of the target acoustic space, and the method for setting the position of the virtual sound source is made different by the first virtual sound source and the second virtual sound source.
5. The method for processing a sound signal according to claim 4, wherein, only when the virtual sound source representing the reflected sound of the target acoustic space is the first virtual sound source, the position of the first virtual sound source is moved to a position where it can be played using the position of a loudspeaker near the first virtual sound source.
6. A sound signal processing apparatus, comprising: a loudspeaker that plays a reflected sound of a target acoustic space; and an initial reflected sound control unit that determines whether a virtual sound source representing a reflected sound of a target acoustic space is within a region having a prescribed azimuth angle and elevation angle with respect to the center position of the acoustic space and the position of the loudspeaker, and only when it is within the region, performs output control of an initial reflected sound control signal of the virtual sound source by the loudspeaker.
7. The sound signal processing apparatus according to claim 6, wherein, the initial reflected sound control unit sets the virtual sound source based on the shape of the acoustic space, the position of the loudspeaker in the acoustic space, and the position of a sound pickup point.
8. The sound signal processing apparatus according to claim 6 or 7, wherein, the initial reflected sound control unit sets a gain value and a delay amount of the initial reflected sound control signal based on the positional relationship between the virtual sound source and the loudspeaker.
9. The sound signal processing apparatus according to any one of claims 6 to 8, wherein, the initial reflected sound control unit classifies the virtual sound source into a first virtual sound source and a second virtual sound source. The first virtual sound source is located between the loudspeaker and the sound pickup point and represents a first sound source of the reflected sound of the target acoustic space, and the second virtual sound source is located outside the loudspeaker and represents a second sound source of the reflected sound of the target acoustic space, The method of setting the gain value and the delay amount of the initial reflected sound control signal is different through the first virtual sound source and the second virtual sound source.
10. The sound signal processing device according to claim 9, wherein the initial reflected sound control unit moves the position of the first virtual sound source to a position where it can be played using a speaker near the first virtual sound source only when the virtual sound source is the first virtual sound source.
11. A recording medium, which is a non-volatile computer-readable recording medium, recording a program for causing a computer to execute the following processing: Determine whether a virtual sound source representing the reflected sound of a target sound space is within a region having a specified azimuth angle and elevation angle with respect to the center position of the sound space and the position of the speaker, Only when it is within the region, output control of the initial reflected sound control signal of the virtual sound source is performed through the speaker.