Sound signal processing method and sound signal processing device
The sound signal processing method improves realism by using virtual speakers to simulate a larger venue environment with fewer physical speakers, addressing the resource-intensive challenges of traditional speaker installations.
Patent Information
- Application Number
- JP2024191340
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-10-31
- Publication Date
- 2025-11-12
- Estimated Expiration
- 2040-09-09
AI Technical Summary
Installing a large number of speakers and equipment in a venue to improve sound quality requires significant work and resources, such as wiring and manpower.
A sound signal processing method that determines the type of sound signal and generates localized or distributed sound signals using virtual speakers, combining them with real speakers to enhance the sense of realism without the need for extensive equipment.
Enhances the sense of realism by simulating a larger venue environment with fewer physical speakers, allowing for more realistic audio experiences for both live attendees and remote listeners.
Smart Images

Figure 0007768324000001 
Figure 0007768324000002 
Figure 0007768324000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound signal processing method and a sound signal processing device for processing a sound signal. [Background technology]
[0002] Patent Document 1 discloses an acoustic signal compensation device equipped with a compensation speaker that outputs compensation sound to compensate for the sound reproduced from the speaker being masked by background noise and other noise at venues such as public viewings. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2017-200025 Summary of the Invention [Problem to be solved by the invention]
[0004] Installing a large number of speakers and other equipment in a venue improves the sound quality and creates a more realistic feeling. However, increasing the amount of equipment requires more work, such as wiring, securing a power source, and securing manpower.
[0005] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide a sound signal processing method and a sound signal processing device that can improve the sense of realism even with a small amount of equipment. [Means for solving the problem]
[0006] A sound signal processing method acquires a sound signal, determines the type of the sound signal, sets a plurality of virtual speakers, and when the determined type of the sound signal is a first type, generates a first sound signal that has been subjected to localization processing to localize a sound image in any one of the plurality of virtual speakers; when the determined type of the sound signal is a second type, generates a second sound signal that has been subjected to distribution processing to localize a sound image in two or more virtual speakers of the plurality of virtual speakers; adds the first sound signal and the second sound signal to generate a sum signal; and outputs the sum signal to a plurality of real speakers. [Effects of the Invention]
[0007] Users can improve the sense of realism even with less equipment. [Brief explanation of the drawings]
[0008] [Figure 1] 1 is a block diagram showing the configuration of a sound signal processing system 1. FIG. [Figure 2] 2 is a schematic plan view showing an installation state of a plurality of speakers 14A to 14G. FIG. [Figure 3] FIG. 2 is a block diagram showing the configuration of a mixer 11. [Figure 4] FIG. 2 is a block diagram showing the functional configuration of a mixer 11. [Figure 5] 4 is a flowchart showing the operation of the mixer 11. [Figure 6] FIG. 1 is a schematic plan view of a live music venue 70 showing virtual speakers. [Figure 7] 3A and 3B are plan views schematically illustrating output modes of a first sound signal and a second sound signal. [Figure 8] 1 is a plan view schematically showing the audio-visual environment of each listener using an information processing terminal 13. FIG. [Figure 9] 1 is a plan view schematically showing the audio-visual environment of each listener using an information processing terminal 13. FIG. [Figure 10] 1 is a plan view schematically showing the audio-visual environment of each listener using an information processing terminal 13. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0009] 1 is a block diagram showing the configuration of a sound signal processing system 1. The sound signal processing system 1 includes a mixer 11, a plurality of information processing terminals 13, and a plurality of speakers 14A to 14G.
[0010] The mixer 11 and the information processing terminals 13 are installed in different locations, and are connected to each other via the Internet.
[0011] The mixer 11 is connected to a plurality of speakers 14 to 14 G. The mixer 11 and the plurality of speakers 14 A to 14 G are connected via a network cable or an audio cable.
[0012] The mixer 11 is an example of the sound signal processing device of the present invention. The mixer 11 receives sound signals from a plurality of information processing terminals 13 via the Internet, performs panning processing and effect processing, and supplies the sound signals to a plurality of speakers 14A to 14G.
[0013] 2 is a schematic plan view showing the installation of multiple speakers 14A to 14G. The multiple speakers 14A to 14G are installed along the walls of a live music venue 70. In this example, the live music venue 70 is rectangular in plan view. A stage 50 is placed in front of the live music venue 70. Performers perform on stage 50, such as singing or playing music.
[0014] Speaker 14A is installed on the left side of stage 50, and speaker 14B is installed on the right side of stage 50. Speaker 14C is installed on the left side of the center between the front and back of live house 70, and speaker 14D is installed on the right side of the center between the front and back of live house 70. Speaker 14E is installed on the left side at the rear of live house 70, speaker 14F is installed in the center between the left and right at the rear of live house 70, and speaker 14G is installed on the right side at the rear of live house 70.
[0015] A listener L1 is positioned in front of the speaker 14F. The listener L1 listens to the performers' performances and cheers, claps, calls, etc. on the performers. The sound signal processing system 1 outputs the sounds of the other listeners' cheers, claps, calls, etc. to the live music venue 70 via the speakers 14A to 14G. The sounds of the other listeners' cheers, claps, calls, etc. are input to the mixer 11 from the information processing terminal 13. The information processing terminal 13 is a portable information processing device such as a personal computer (PC), a tablet computer, or a smartphone. The user of the information processing terminal 13 is a listener who remotely listens to the performances, such as singing or musical instruments, at the live music venue 70. The information processing terminal 13 acquires the sounds of the listeners' cheers, claps, calls, etc. via a microphone (not shown). Alternatively, the information processing terminal 13 may display icon images such as "cheering," "applause," "calling out," and "bustling" on a display (not shown) and accept a selection operation for these icon images from the listener. When accepting a selection operation for these icon images, the information processing terminal 13 may generate a sound signal corresponding to each icon image and obtain it as the sound of the listener's cheering, applause, or call, etc.
[0016] The information processing terminal 13 transmits sounds such as cheers, applause, or calls from each listener to the mixer 11 via the Internet. The mixer 11 receives the sounds such as cheers, applause, or calls from each listener. The mixer 11 performs panning and effect processing on the received sounds and distributes the sound signals to the multiple speakers 14A to 14G. In this way, the sound signal processing system 1 can deliver sounds such as cheers, applause, or calls from a large number of listeners to the live music venue 70.
[0017] The configuration and operation of the mixer 11 will be described in detail below. Fig. 3 is a block diagram showing the hardware configuration of the mixer 11. Fig. 4 is a block diagram showing the functional configuration of the mixer 11. Fig. 5 is a flowchart showing the operation of the mixer 11.
[0018] The mixer 11 includes a display 101, a user I / F (interface) 102, an audio I / O (Input / Output) 103, a signal processing unit (DSP) 104, a network I / F 105, a CPU 106, a flash memory 107, and a RAM 108. These components are connected via a bus 171.
[0019] The CPU 106 is a control unit that controls the operation of the mixer 11. The CPU 106 performs various operations by reading out a predetermined program stored in a flash memory 107, which is a storage medium, into a RAM 108 and executing the program.
[0020] The program read by CPU 106 does not need to be stored in flash memory 107 within the device itself. For example, the program may be stored in a storage medium of an external device such as a server. In this case, CPU 106 simply reads the program from the server into RAM 108 and executes it each time.
[0021] The signal processing unit 104 is configured with a DSP for performing various signal processing. The signal processing unit 104 receives sound signals related to cheers, applause, calls, etc. from the listeners from the information processing terminal 13 via the network I / F 105.
[0022] The signal processing unit 104 performs panning and effect processing on the received sound signal, and outputs the processed sound signal via the audio I / O 103 to the speakers 14A to 14G.
[0023] As shown in FIG. 4, the CPU 106 and the signal processing unit 104 functionally include an acquisition unit 301 , a determination unit 302 , a setting unit 303 , a localization processing unit 304 , a distribution processing unit 305 , and an addition unit 306 .
[0024] The acquisition unit 301 acquires sound signals related to cheers, applause, calls, etc. from listeners from each of the multiple information processing terminals 13 (S11). Thereafter, the determination unit 302 determines the type of sound signal (S12).
[0025] The types of sound signals include a first type and a second type. The first type includes cheers from individual listeners such as "Go for it!", calling out the performers' names, and exclamations such as "Bravo." In other words, the first type is a sound that can be recognized as the voice of an individual listener without being drowned out by the audience. The second type is a sound that cannot be recognized as the voice of an individual listener and is made by many listeners simultaneously, and includes, for example, applause, singing together, cheers such as "Wow!", and murmuring.
[0026] When the determination unit 302 recognizes a voice such as "Go for it!" or "Bravo" as described above, for example, through a voice recognition process, the determination unit 302 determines that the sound signal is of the first type. When the determination unit 302 does not recognize a voice, the determination unit 302 determines that the sound signal is of the second type.
[0027] The determination unit 302 outputs the sound signals determined to be of the first type to the localization processing unit 304, and outputs the sound signals determined to be of the second type to the distributed processing unit 305. The localization processing unit 304 and the distributed processing unit 305 set a plurality of virtual speakers (S13).
[0028] FIG. 6 is a schematic plan view of the live music venue 70 showing virtual speakers. As shown in FIG. 6, the localization processing unit 304 and the distributed processing unit 305 set up multiple virtual speakers 14N1 to 14N16. The localization processing unit 304 and the distributed processing unit 305 manage the positions of the speakers 14A to 14G and the virtual speakers 14N1 to 14N16 using two-dimensional or three-dimensional Cartesian coordinates with a predetermined position in the live music venue (for example, the center of the stage 50) as the origin. The speakers 14A to 14G are real speakers. Therefore, the coordinates of the speakers 14A to 14G are stored in advance in the flash memory 107 (or a server, not shown). The localization processing unit 304 and the distributed processing unit 305 evenly arrange the virtual speakers 14N1 to 14N16 throughout the entire live music venue 70, as shown in FIG. 6. In the example of FIG. 6, the localization processing unit 304 and the distributed processing unit 305 also set a virtual speaker 14N16 at a position outside the live music venue 70.
[0029] The virtual speaker setting process (S13) does not need to be performed after the sound signal type determination process (S12). The virtual speaker setting process (S13) may be performed before the sound signal acquisition process (S11) or the sound signal type determination process (S12).
[0030] Thereafter, the localization processing unit 304 performs localization processing to generate a first sound signal, and the distributed processing unit 305 performs distributed processing to generate a second sound signal (S14).
[0031] The localization process is a process of localizing a sound image at the position of one of the virtual speakers 14N1 to 14N16. However, the position where the sound image is localized is not limited to the virtual speakers 14N1 to 14N16. When the position where the sound image is localized matches the position of the speakers 14A to 14G, the localization processing unit 304 outputs a sound signal to one of the speakers 14A to 14G.
[0032] The localization position of the first type of sound signal may be set randomly, or the mixer 11 may include a position information receiving unit that receives position information from the listener. The listener operates the information processing terminal 13 to specify the localization position of their own sound. For example, the information processing terminal 13 displays an image that resembles a plan view or a perspective view of the live music venue 70 and receives the localization position from the user. The information processing terminal 13 transmits position information (coordinates) corresponding to the received localization position to the mixer 11. The localization processing unit 304 of the mixer 11 sets a virtual speaker at the coordinates corresponding to the position information received from the information processing terminal 13 and performs processing to localize a sound image at the position of the set virtual speaker.
[0033] The localization processing unit 304 performs panning processing or effect processing in order to localize the sound images at the positions of the virtual speakers 14N1 to 14N16.
[0034] Panning is a process of supplying the same sound signal to multiple speakers among speakers 14A to 14G and controlling the volume of the supplied sound signal to phantom-localize a sound image at the position of a virtual speaker. For example, if the same sound signal with the same volume is supplied to speaker 14A and speaker 14C, the sound image is localized as if the virtual speaker were installed at the center position on the line connecting speaker 14A and speaker 14C. In other words, panning is a process of increasing the volume of the sound signal supplied to speakers closer to the virtual speaker and decreasing the volume of the sound signal supplied to speakers farther from the virtual speaker. Note that in FIG. 6, multiple virtual speakers 14N1 to 14N16 are set on the same plane. However, the localization processing unit 304 can also localize a sound image at a virtual speaker at any position on the three-dimensional coordinate system by supplying the same sound signal to multiple speakers installed at different heights.
[0035] Furthermore, the effect processing includes, for example, processing to apply a delay. If a delay is applied to the sound signals supplied to the real speakers 14A to 14G, the listener will perceive a sound image at a position farther away than the real speakers. Therefore, by applying a delay to the sound signals, the localization processing unit 304 can localize the sound image at a virtual speaker set at a position farther away than the real speakers 14A to 14G.
[0036] The effect processing may also include processing to add reverb. Adding reverb to a sound signal causes a listener to perceive a sound image at a position farther away than the positions of the actual speakers. Therefore, by adding reverb to the sound signal, the localization processing unit 304 can localize the sound image at a virtual speaker that is set at a position farther away than the actual speakers 14A to 14G.
[0037] The effect processing may also include processing to apply frequency characteristics using an equalizer. A listener perceives a sound image not only based on the volume and time differences between the two ears, but also based on the difference in frequency characteristics. Therefore, the localization processing unit 304 can localize a sound image at the position of the set virtual speaker by applying frequency characteristics according to the transfer characteristics from the position of the target virtual speaker to the target listening position (for example, the center of the stage 50).
[0038] On the other hand, the distributed processing is a process of localizing a sound image by distributing it to a plurality of the virtual speakers 14N1 to 14N16. When the position where the sound image is to be localized matches the position of the real speakers 14A to 14G, the distributed processing unit 305 also outputs a sound signal to one of the speakers 14A to 14G.
[0039] The distributed processing unit 305 performs panning processing or effect processing to localize sound images at a plurality of positions of the virtual speakers 14N1 to 14N16. The method of localizing each sound image at any one of the positions of the virtual speakers 14N1 to 14N16 is the same as that of the localization processing unit 304. The distributed processing unit 305 reproduces sounds such as applause, chorus, cheers, or murmurs by localizing sound images distributed among a plurality of virtual speakers.
[0040] In the above example, a sound image is localized in a virtual speaker that is set at a position farther away than the actual speakers 14A to 14G by adding reverb. However, reverb can make a listener perceive a spatial spread of sound. Therefore, the distributed processing unit 305 may perform a process of making a listener perceive a spatial spread such as reverb in addition to the process of localizing sound images in a plurality of virtual speakers.
[0041] Furthermore, it is preferable that the distributed processing unit 305 adjusts the output timing of the sound signals output to the speakers 14A to 14G to shift the timing at which the sounds output from the multiple real speakers reach the listener. This enables the distributed processing unit 305 to further distribute the sound and provide a spatial spread.
[0042] The adder 306 adds the first sound signal that has been localized and the second sound signal that has been distributed (S15). The addition is performed by an adder for each speaker. The adder 306 outputs an added signal obtained by adding the first sound signal and the second sound signal to each of the multiple real speakers (S16).
[0043] In this way, the first sound signal reaches the listener via one of the virtual speakers 14N1 to 14N16 as a sound source. The second sound signal reaches the listener via dispersion from the multiple virtual speakers 14N1 to 14N16. FIG. 7 is a plan view schematically illustrating the output mode of the first sound signal and the second sound signal. As shown in FIG. 7, sounds such as "Bravo" are output from specific virtual speakers. In the example of FIG. 7, sounds such as "Bravo" are output from the virtual speaker 14N3 in the center of the front of the audience seats, the virtual speakers 14N9 on the left and right behind each seat, the virtual speaker 14N12, and the virtual speaker 14N16 in the rear outside the live music venue 70. Applause and cheers such as "Wow!" are output from multiple virtual speakers.
[0044] This allows the performers on stage 50 to hear the voices, applause, cheers, etc. of listeners from locations other than listener L1, enabling the live performance to be performed in a more realistic environment. Also, listener L1 in live house 70 can hear the voices, applause, cheers, etc. of many listeners in the same space, allowing the performers to watch the live performance in a more realistic environment.
[0045] In particular, the sound signal processing method of this embodiment can emit the voices, applause, cheers, etc. of listeners from a larger number of virtual speakers 14N1 to 14N16 than the real speakers 14A to 14G. Therefore, the sound signal processing method of this embodiment can output the voices, applause, cheers, etc. of listeners from various positions even with a small amount of equipment, thereby improving the sense of realism. Furthermore, the sound signal processing method of this embodiment can output the voices, applause, cheers, etc. of listeners by setting the positions of the virtual speakers outside the space of the real venue, simulating the environment of a venue that is even larger than the real space.
[0046] In the above embodiment, an example has been shown in which the sense of realism is improved in the live music venue 70. However, the sound signal processing method of this embodiment can also improve the sense of realism for each listener in a remote location who uses the information processing terminal 13.
[0047] 8, 9, and 10 are plan views schematically showing the audio-visual environments of each listener using information processing terminal 13. In this example, speakers 14FL, 14FR, 14C, 14SL, and 14SR are installed along the walls of living room 75. Living room 75 in this example has a rectangular shape in plan view. Display 55 is placed in front of living room 75. Listener L2 is located in the center of the room. Listener L2 watches the performers' performances displayed on display 55.
[0048] Speaker 14FL is installed to the left of display 55, speaker 14C is installed in front of display 55, and speaker 14FR is installed to the right of display 55. Speaker 14SL is installed at the rear left side of living room 75, and speaker 14SR is installed at the rear right side of living room 75.
[0049] The information processing terminal 13 acquires video and sound related to the performer's performance. For example, in the example of FIG. 2, the mixer 11 acquires sound such as the performer's playing sound or singing sound, and transmits it to the information processing terminal 13.
[0050] Similar to mixer 11, information processing terminal 13 performs signal processing such as panning and effect processing on the acquired sound, and outputs the processed sound signals to speakers 14FL, 14FR, 14C, 14SL, and 14SR. Speakers 14FL, 14FR, 14C, 14SL, and 14SR output sounds related to the performers' performances.
[0051] Furthermore, the information processing terminal 13 acquires sound signals related to cheers, applause, calls, etc. of other listeners from other information processing terminals 13. Like the mixer 11, the information processing terminal 13 determines the type of sound signal and performs localization processing or distribution processing.
[0052] As a result, as shown in FIG. 9, listener L2 can get the same sense of presence as if he or she were in the center of live music venue 70, watching the performers' performance together with a large audience, even in room 75.
[0053] The information processing terminal 13 may include a seat designation receiving unit that receives seat position designation information from a listener. In this case, the information processing terminal 13 changes the contents of the panning process and effect processing based on the seat position designation information. For example, if a listener designates a seat position immediately in front of the stage 50, the information processing terminal 13 sets the listener L2 to a position immediately in front of the stage 50, as shown in FIG. 10, sets multiple virtual speakers, and performs localization processing and distribution processing of sound signals related to cheers, applause, calls, etc. from other listeners. This allows listener L2 to experience the same sense of presence as if he or she were directly in front of the stage 50.
[0054] A provider of a sound signal processing system offers tickets for seats in front of the stage, seats beside the stage, seats in the center of a live music venue, or seats in the back. A user of an information processing terminal 13 purchases a ticket for one of these seat positions. The user can, for example, choose a seat in front of the stage which is more expensive and offers a greater sense of realism, or a seat in the back which is less expensive. The information processing terminal 13 changes the content of the panning process and effect process depending on the seat position selected by the user. This allows the user to experience the same sense of realism as if they were watching a performance in the seat position they purchased. Furthermore, a provider of a sound signal processing method can conduct business in the same way as if they were providing an event in a real space.
[0055] Furthermore, in the sound signal processing method of this embodiment, multiple users may specify the same seat position. For example, multiple users may each specify a seat position immediately in front of the stage 50. In this case, the information processing terminal 13 of each user gives the listener a sense of realism as if they were sitting in a seat immediately in front of the stage 50. This allows multiple listeners to watch a performer's performance with the same sense of realism for one seat. Therefore, providers of sound signal processing methods can provide services that exceed the capacity of an existing space for accommodating spectators.
[0056] The description of the present embodiment is illustrative in all respects and is not restrictive. The scope of the present invention is defined not by the above-described embodiments but by the claims. Furthermore, the scope of the present invention is intended to include all modifications that are equivalent to the claims and fall within the scope thereof.
[0057] For example, in the above embodiment, a sound signal is subjected to voice recognition processing, and if voice is recognized by the voice recognition processing, the type of the sound signal is determined to be the first type. If voice cannot be recognized by the voice recognition processing, the type of the sound signal is determined to be the second type. However, a sound signal may include multiple channels and include additional information (metadata) indicating whether each channel is the first type or the second type. For example, when the information processing terminal 13 receives a selection operation from a listener, such as "cheering," "applause," "calling," or "bustling noise," and generates a corresponding sound signal, the information processing terminal 13 generates a sound signal for a channel corresponding to the selected sound, attaches the additional information, and transmits the sound signal to the mixer 11. In this case, the determination unit 302 of the mixer 11 determines the type of the sound signal for each channel based on the additional information.
[0058] Furthermore, the sound signal may include both a first type and a second type of sound source. In this case, the mixer 11 (or the information processing terminal 13) separates the first type of sound signal and the second type of sound signal into sound sources. The localization processing unit 304 and the distribution processing unit 305 generate a first sound signal and a second sound signal from each of the separated sound signals. Any method for separating the sound sources may be used. For example, as described above, the first type is a speech sound from a specific listener. Therefore, the determination unit 302 separates the first type of sound signal using noise reduction processing that treats the speech sound as a target sound and eliminates other sounds as noise sounds. [Explanation of symbols]
[0059] 1...Sound signal processing system 11...Mixer 13...Information processing terminal 14A~14G...Speakers 14FL, 14FR, 14C, 14SL, 14SR...Speakers 14N1~14N16...Virtual speakers 50...Stage 55...Indicator 70...Live house 75…Room 101...Indicator 102...User I / F 103...Audio I / O 104...signal processing unit 105...Network I / F 106...CPU 107...Flash memory 108...RAM 171...Bus 301…Acquisition Department 302...Judgment section 303...Settings Department 304...Localization processing unit 305...Distributed processing unit 306...Addition section
Claims
1. Accepting a selection of an icon image and generating a sound signal corresponding to the selected icon image; determining the type of the sound signal; Set up multiple virtual speakers generating a first sound signal that has been subjected to localization processing for localizing a sound image to any one of the plurality of virtual speakers when the determined type of the sound signal is a first type; When the determined type of the sound signal is a second type, a second sound signal is generated by performing a distribution process for distributing the sound signal to two or more virtual speakers among the plurality of virtual speakers and localizing a sound image; adding the first sound signal and the second sound signal to generate a sum signal; outputting the summed signal to a plurality of real speakers; Sound signal processing method.
2. the sound signal includes a plurality of channels; determining the type for each channel; The sound signal processing method according to claim 1 .
3. When the sound signal includes both the first type and the second type of sound sources, the sound signals are separated into the first type of sound signal and the second type of sound signal; generating the first sound signal and the second sound signal from the respective separated sound signals; The sound signal processing method according to claim 1 .
4. performing a voice recognition process on the sound signal; determining that the type of the sound signal is the first type when the voice is recognized in the voice recognition processing; determining that the type of the sound signal is the second type when the voice cannot be recognized by the voice recognition processing; The sound signal processing method according to any one of claims 1 to 3.
5. the localization process includes a process of outputting the first sound signal solely to a real speaker when the localization position coincides with the real speaker; The sound signal processing method according to any one of claims 1 to 4.
6. Accept location information from the user, the localization process localizes the first sound signal at a position of the received position information. The sound signal processing method according to any one of claims 1 to 5.
7. The localization processing realizes the virtual speakers by panning processing and effect processing. The sound signal processing method according to any one of claims 1 to 6.
8. Accepts seat location specification information from the user, changing the contents of the panning process and the effect process based on the seat position designation information; The sound signal processing method according to claim 7.
9. The effect processing includes delay, equalizer, or reverb.
9. The sound signal processing method according to claim 7 or 8.
10. the distributed processing includes adjusting an output timing of the second sound signal. The sound signal processing method according to any one of claims 1 to 9.
11. an acquisition unit that receives a selection of an icon image and generates a sound signal corresponding to the selected icon image; a determination unit for determining the type of the sound signal; a setting unit that sets a plurality of virtual speakers; generating a first sound signal that has been subjected to localization processing for localizing a sound image to any one of the plurality of virtual speakers when the determined type of the sound signal is a first type; When the determined type of the sound signal is a second type, a second sound signal is generated by performing a distribution process for distributing the sound signal to two or more virtual speakers among the plurality of virtual speakers and localizing a sound image; adding the first sound signal and the second sound signal to generate a sum signal; a signal processing unit that outputs the summed signal to a plurality of real speakers; A sound signal processing device comprising:
12. the sound signal includes a plurality of channels; The determination unit determines the type for each channel. The sound signal processing device according to claim 11 .
13. a sound source separation unit that separates the sound signals into the first type of sound signal and the second type of sound signal when the sound signals include both the first type and the second type of sound sources; generating the first sound signal and the second sound signal from the respective separated sound signals; The sound signal processing device according to claim 11 .
14. a voice recognition processing unit that performs voice recognition processing on the sound signal; The determination unit determining that the type of the sound signal is the first type when the voice is recognized in the voice recognition processing; determining that the type of the sound signal is the second type when the voice cannot be recognized by the voice recognition processing; The sound signal processing device according to any one of claims 11 to 13.
15. the localization process includes a process of outputting the first sound signal solely to a real speaker when the localization position coincides with the real speaker; The sound signal processing device according to any one of claims 11 to 14.
16. a location information receiving unit that receives location information from a user; the localization process localizes the first sound signal at a position of the received position information. The sound signal processing device according to any one of claims 11 to 15.
17. The localization processing realizes the virtual speakers by panning processing and effect processing. The sound signal processing device according to any one of claims 11 to 16.
18. a seat designation receiving unit that receives seat position designation information from a user; the signal processing unit changes the content of the panning processing and the effect processing based on the seat position designation information. The sound signal processing device according to claim 17.
19. The effect processing includes delay, equalizer, or reverb. The sound signal processing device according to claim 17 or 18.
20. the distributed processing includes adjusting an output timing of the second sound signal. The sound signal processing device according to any one of claims 11 to 19.
Citation Information
Patent Citations
Sound environment generator for onset of sleeping and waking up
JP2011130100A
Acoustic signal compensation device and program thereof
JP2017200025A
Three dimensional audio telephony
US20030044002A1