Sound signal processing method and sound signal processing apparatus

By generating virtual speakers and performing sound image localization and dispersion processing, the problem of inefficient wiring and power supply caused by adding equipment in the venue was solved, achieving a better sense of presence and spatial simulation effect with fewer devices.

CN116034591BActive Publication Date: 2025-10-24YAMAHA CORP
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202180054332.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2020-09-09
Filing Date
2021-08-25
Publication Date
2025-10-24
Estimated Expiration
2041-08-25

AI Technical Summary

Technical Problem

When increasing the number of devices in a venue to improve sound quality and immersiveness, issues such as cabling and power supply lead to inefficiencies.

Method used

By generating virtual speakers and performing sound image localization and dispersion processing, multiple virtual speakers are simulated using a small number of actual speakers, and summed signals are generated and output to enhance the sense of presence.

Benefits of technology

It achieves enhanced immersion with fewer devices, simulates larger spaces, and supports immersive experiences for remote users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116034591B_ABST
    Figure CN116034591B_ABST
Patent Text Reader

Abstract

A sound signal processing method acquires a sound signal, determines a type of the sound signal, sets a plurality of virtual speakers, generates a first sound signal obtained by applying a localization process that localizes a sound image to any one of the plurality of virtual speakers when the determined type of the sound signal is a first type, generates a second sound signal obtained by applying a dispersion process that dispersively localizes a sound image to two or more of the plurality of virtual speakers when the determined type of the sound signal is a second type, generates an added signal by adding the first sound signal and the second sound signal, and outputs the added signal to a plurality of actually existing speakers.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to a sound signal processing method and a sound signal processing apparatus that process a sound signal. BACKGROUND

[0002] In Patent Literature 1, a sound signal compensation apparatus is disclosed that has a compensation speaker that outputs a compensation sound in order to compensate for noise masking of sound played from a speaker by noise such as ambient noise in a venue such as a public viewing.

[0003] PRIOR ART DOCUMENTS

[0004] PATENT LITERATURE

[0005] Patent Literature 1: Japanese Patent Application Publication No. 2017-200025 SUMMARY

[0006] PROBLEMS TO BE SOLVED BY THE INVENTION

[0007] If a large number of speakers and the like are provided in a venue, sound quality is improved and the sense of presence is improved. However, if the number of equipment is increased, it takes time to lay wiring, ensure power supplies, and ensure manpower and the like.

[0008] Therefore, an object of the present application is to provide a sound signal processing method and a sound signal processing apparatus that can improve the sense of presence with fewer pieces of equipment.

[0009] MEANS FOR SOLVING THE PROBLEMS

[0010] The sound signal processing method obtains a sound signal, determines a type of the sound signal, sets a plurality of virtual speakers, generates a first sound signal obtained by implementing localization processing that localizes a sound image to any one of the plurality of virtual speakers when the determined type of the sound signal is a first type, generates a second sound signal obtained by implementing dispersion processing that dispersively localizes a sound image to two or more of the plurality of virtual speakers when the determined type of the sound signal is a second type, adds the first sound signal and the second sound signal to generate an added signal, and outputs the added signal to a plurality of actually existing speakers.

[0011] EFFECTS OF THE INVENTION

[0012] The user can improve the sense of presence with fewer pieces of equipment. BRIEF DESCRIPTION OF DRAWINGS

[0013] Figure 1 is a block diagram showing the configuration of a sound signal processing system 1.

[0014] Figure 2is a top view schematic diagram showing a setting manner of the plurality of speakers 14A to 14G.

[0015] Figure 3 is a block diagram showing a structure of the mixer 11.

[0016] Figure 4 is a block diagram showing a functional structure of the mixer 11.

[0017] Figure 5 is a flowchart showing an action of the mixer 11.

[0018] Figure 6 is a top view schematic diagram showing a small-size performance site 70 of a virtual speaker.

[0019] Figure 7 is a top view schematically showing an output manner of the first sound signal and the second sound signal.

[0020] Figure 8 is a top view schematically showing an audiovisual environment of each listener using the information processing terminal 13.

[0021] Figure 9 is a top view schematically showing an audiovisual environment of each listener using the information processing terminal 13.

[0022] Figure 10 is a top view schematically showing an audiovisual environment of each listener using the information processing terminal 13. DETAILED DESCRIPTION

[0023] Figure 1 is a block diagram showing a structure of a sound signal processing system 1. The sound signal processing system 1 is provided with a mixer 11, a plurality of information processing terminals 13, and a plurality of speakers 14A to 14G.

[0024] The mixer 11 and the plurality of information processing terminals 13 are respectively provided at different sites. The mixer 11 and the plurality of information processing terminals 13 are connected via the Internet.

[0025] The mixer 11 is connected to the plurality of speakers 14A to 14G. The mixer 11 and the plurality of speakers 14A to 14G are connected via a network cable or an audio cable.

[0026] The mixer 11 is an example of a sound signal processing apparatus of the present application. The mixer 11 receives sound signals from the plurality of information processing terminals 13 via the Internet, performs panning processing and effect processing, and supplies the sound signals to the plurality of speakers 14A to 14G.

[0027] Figure 2is a schematic plan view showing the arrangement of the plurality of speakers 14A to 14G. The plurality of speakers 14A to 14G are arranged along the walls of the small performance venue 70. The small performance venue 70 of this example is rectangular in plan view. A stage 50 is arranged in front of the small performance venue 70. On the stage 50, a performer performs a performance such as singing or playing a musical instrument.

[0028] The speaker 14A is arranged on the left side of the stage 50, and the speaker 14B is arranged on the right side of the stage 50. The speaker 14C is arranged on the left side of the front and rear center of the small performance venue 70, and the speaker 14D is arranged on the right side of the front and rear center of the small performance venue 70. The speaker 14E is arranged on the left side of the rear of the small performance venue 70, the speaker 14F is arranged on the left and right center of the rear of the small performance venue 70, and the speaker 14G is arranged on the right side of the rear of the small performance venue 70.

[0029] The listener LI is located in front of the speaker 14F. The listener LI views and listens to the performance of the performer, and cheers, claps, or calls out to the performer, and the like. The sound signal processing system 1 outputs the sound of the cheering, clapping, or calling out to the small performance venue 70 from the other listeners via the speakers 14A to 14G. The sound of the cheering, clapping, or calling out to the small performance venue 70 from the other listeners is input to the mixer 11 from the information processing terminal 13. The information processing terminal 13 is a portable information processing device such as a personal computer (PC), a tablet computer, or a smartphone. The user of the information processing terminal 13 becomes a listener who views and listens to the performance of singing or playing a musical instrument, and the like, at the small performance venue 70 from a remote location. The information processing terminal 13 acquires the sound of the cheering, clapping, or calling out of each listener via a microphone not shown. Alternatively, the information processing terminal 13 can display icon images of "cheering", "clapping", "calling out", and "jeering", and the like, on a display (not shown), and accept a selection operation of these icon images from the listener. The information processing terminal 13 can generate a sound signal corresponding to each icon image if the selection operation of these icon images is accepted, and acquire the sound of the cheering, clapping, or calling out of each listener.

[0030] The information processing terminal 13 transmits the sound of the cheering, clapping, or calling out of each listener to the mixer 11 via the Internet. The mixer 11 receives the sound of the cheering, clapping, or calling out of each listener. The mixer 11 performs sound image processing and effect processing on the received sound, and distributes the sound signal to the plurality of speakers 14A to 14G. Thus, the sound signal processing system 1 can deliver the sound of the cheering, clapping, or calling out of a large number of listeners to the small performance venue 70.

[0031] Hereinafter, the structure and operation of the mixer 11 will be described in detail. Figure 3is a block diagram showing a hardware structure of the mixer 11. Figure 4 is a block diagram showing a functional structure of the mixer 11. Figure 5 is a flowchart showing an action of the mixer 11.

[0032] The mixer 11 is provided with a display 101, a user I / F (interface) 102, an audio I / O (input / output) 103, a signal processing section (DSP) 104, a network I / F 105, a CPU 106, a flash memory 107, and a RAM 108. These structures are connected via a bus 171.

[0033] The CPU 106 is a control section that controls an action of the mixer 11. The CPU 106 reads out a prescribed program stored in the flash memory 107 that is a storage medium to the RAM 108 and executes, thereby performing various actions.

[0034] In addition, the program read out by the CPU 106 does not necessarily have to be stored in the flash memory 107 in the device. For example, the program can be stored in a storage medium of an external device such as a server. In this case, the CPU 106 can read out the program to the RAM 108 and execute each time from the server.

[0035] The signal processing section 104 is constituted by a DSP that performs various signal processing. The signal processing section 104 receives a sound signal related to a cheer, applause, or a call of a listener from the information processing terminal 13 via the network I / F 105.

[0036] The signal processing section 104 performs a sound image processing and an effect processing on the received sound signal. The signal processing section 104 outputs the sound signal after the signal processing to the speaker 14A to the speaker 14G via the audio I / O 103.

[0037] As shown in Figure 4 , the CPU 106 and the signal processing section 104 functionally have a obtaining section 301, a determining section 302, a setting section 303, a positioning processing section 304, a dispersing processing section 305, and an adding section 306.

[0038] The obtaining section 301 obtains a sound signal related to a cheer, applause, or a call of a listener from each of the plurality of information processing terminals 13 (S11). Thereafter, the determining section 302 determines a type of the sound signal (S12).

[0039] The type of the sound signal includes a first type or a second type. The first type includes a cheer of "cheer up" or the like of each of the listeners, a call of a name of a performer, or an exclamation of "well done" or the like. That is, the first type is a sound that can be recognized as a sound of an individual listener without being buried in the audience. The second type is a sound that cannot be recognized as a sound of an individual listener, such as a sound of a large number of listeners simultaneously, including applause, a chorus, or a cheer of "wow" or the like, or a noise.

[0040] The determination section 302 determines the sound signal as the first type in a case where the sound of "cheer up," "well done," or the like as described above is recognized by, for example, a sound recognition process. The determination section 302 determines a sound signal of which the sound is not recognized as the second type.

[0041] The determination section 302 outputs the sound signal determined as the first type to the localization processing section 304 and outputs the sound signal determined as the second type to the dispersion processing section 305. The localization processing section 304 and the dispersion processing section 305 set a plurality of virtual speakers (S13).

[0042] Figure 6 is a top view of a small-size performance venue 70 that represents virtual speakers. As shown in Figure 6 , the localization processing section 304 and the dispersion processing section 305 set a plurality of virtual speakers 14N1 to 14N16. The localization processing section 304 and the dispersion processing section 305 manage the positions of the speakers 14A to 14G and the virtual speakers 14N1 to 14N16 using two-dimensional or three-dimensional orthogonal coordinates with a prescribed position of the small-size performance venue (for example, the center of the stage 50) as the origin. The speakers 14A to 14G are actually existing speakers. Therefore, the coordinates of the speakers 14A to 14G are stored in advance in the flash memory 107 (or a server or the like not shown). As shown in Figure 6 , the localization processing section 304 and the dispersion processing section 305 uniformly arrange the virtual speakers 14N1 to 14N16 in the entire small-size performance venue 70. Further, in the Figure 6 , the localization processing section 304 and the dispersion processing section 305 set the virtual speakers 14N16 also at positions outside the small-size performance venue 70.

[0043] In addition, the setting processing of the virtual speakers (S13) does not necessarily have to be performed after the determination processing of the type of the sound signal (S12). The setting processing of the virtual speakers (S13) can also be performed in advance before the acquisition processing of the sound signal (S11) or the determination processing of the type of the sound signal (S12).

[0044] Thereafter, the positioning processing section 304 performs positioning processing to generate a first sound signal, and the dispersion processing section 305 performs dispersion processing to generate a second sound signal (S14).

[0045] The positioning processing is processing to position a sound image at a position of any one of the virtual speakers 14N1 to 14N16. The position at which the sound image is positioned is not limited to the virtual speakers 14N1 to 14N16. The positioning processing section 304 outputs a sound signal to any one of the speakers 14A to 14G in a case where the position at which the sound image is positioned coincides with the position of the speakers 14A to 14G.

[0046] In addition, the position at which the sound signal of the first type is positioned can also be set randomly, but the mixer 11 can also be provided with a position information reception section that receives position information from a listener. The listener operates the information processing terminal 13 to specify the position at which his or her own sound is positioned. For example, the information processing terminal 13 displays an image that simulates a top view or an oblique view of the small performance venue 70, and receives the position from the user. The information processing terminal 13 transmits position information (coordinates) corresponding to the received position to the mixer 11. The positioning processing section 304 of the mixer 11 sets a virtual speaker at the coordinates corresponding to the position information received from the information processing terminal 13, and performs processing to position a sound image at the position of the set virtual speaker.

[0047] The positioning processing section 304 performs sound image processing or effect processing in order to position a sound image at the positions of the virtual speakers 14N1 to 14N16.

[0048] The sound image processing is processing to supply the same sound signal to a plurality of speakers among the speakers 14A to 14G, and to control the volume of the supplied sound signal, thereby making the sound image illusorily positioned at the position of a virtual speaker. For example, if the same sound signal of the same volume is supplied to the speaker 14A and the speaker 14C, the sound image is positioned as if a virtual speaker is provided at a position in the center of a straight line connecting the speaker 14A and the speaker 14C. That is, the sound image processing is processing to increase the volume of a sound signal supplied to a speaker that is closer to the position of a virtual speaker, and to decrease the volume of a sound signal supplied to a speaker that is farther from the position of the virtual speaker. In addition, in the case where a plurality of virtual speakers 14N1 to 14N16 are set on the same plane, the sound image processing is processing to position a sound image at the position of any one of the virtual speakers 14N1 to 14N16. Figure 6 In the case where a plurality of virtual speakers 14N1 to 14N16 are set on the same plane, the sound image processing is processing to position a sound image at the position of any one of the virtual speakers 14N1 to 14N16.

[0049] Further, the effect processing includes, for example, processing of imparting a delay. If a delay is imparted to a sound signal supplied to the actually existing speakers 14A to 14G, a listener perceives a sound image at a position farther than the actually existing speakers. Thus, the positioning processing section 304 can cause the sound image to be positioned at a virtual speaker set at a position farther than the actually existing speakers 14A to 14G by imparting a delay to the sound signal.

[0050] Further, the effect processing also includes processing of imparting reverberation. If reverberation is imparted to a sound signal, a listener perceives a sound image at a position farther than the position of the actually existing speakers. Thus, the positioning processing section 304 can cause the sound image to be positioned at a virtual speaker set at a position farther than the actually existing speakers 14A to 14G by imparting reverberation to the sound signal.

[0051] Further, the effect processing also includes processing of imparting a frequency characteristic by an equalizer. A listener perceives a sound image not only according to a difference in volume and a difference in time between the two ears, but also according to a difference in frequency characteristic. Thus, the positioning processing section 304 can cause the sound image to be positioned at a virtual speaker set at a position farther than the actually existing speakers 14A to 14G by imparting a frequency characteristic corresponding to a transfer characteristic from the position of the target virtual speaker to the target listening position (e.g., the center of the stage 50).

[0052] On the other hand, the dispersion processing is processing of causing a sound image to be positioned at a plurality of virtual speakers among the virtual speakers 14N1 to 14N16 in a dispersed manner. The dispersion processing section 305 also outputs a sound signal to any one of the speakers 14A to 14G in a case where the position at which the sound image is positioned coincides with the position of the actually existing speakers 14A to 14G.

[0053] The dispersion processing section 305 performs sound image processing or effect processing in order to cause a sound image to be positioned at a plurality of positions of the virtual speakers 14N1 to 14N16. The method of causing each sound image to be positioned at the position of any one of the virtual speakers 14N1 to 14N16 is the same as that of the positioning processing section 304. The dispersion processing section 305 reproduces a sound of applause, chorus, cheers, or a roar, and the like by causing a sound image to be positioned at a plurality of virtual speakers in a dispersed manner.

[0054] Further, the example of causing a sound image to be positioned at a virtual speaker set at a position farther than the actually existing speakers 14A to 14G by imparting reverberation is described in the above. Among them, reverberation enables a listener to perceive a spatial expansion of a sound. Therefore, the dispersion processing section 305 can also perform processing for perceiving a spatial expansion such as reverberation in addition to the processing of causing a sound image to be positioned at a plurality of virtual speakers.

[0055] Further, the dispersion processing section 305 preferably adjusts the output timing of the sound signals output to the speakers 14A to 14G so that the arrival timing of the sound output from the plurality of actually existing speakers to the listener is staggered. Thereby, the dispersion processing section 305 can further disperse the sound and can impart spatiality expansion.

[0056] The addition section 306 adds the first sound signal after the positioning processing and the second sound signal after the dispersion processing (S15). The addition processing is performed by an addition operator of each speaker. The addition section 306 outputs an added signal obtained by adding the first sound signal and the second sound signal to each of the plurality of actually existing speakers (S16).

[0057] As described above, the first sound signal reaches the listener with any one of the virtual speakers 14N1 to 14N16 as a sound source. The second sound signal reaches the listener dispersedly from the plurality of virtual speakers 14N1 to 14N16. Figure 7 is a plan view schematically showing the output manner of the first sound signal and the second sound signal. As shown in Figure 7 , the sound of "Great" and the like is output from a specific virtual speaker. In Figure 7 , the sound of "Great" and the like is output from the virtual speaker 14N3 at the center of the front of the audience seats, the virtual speakers 14N9 at the back of each of the audience seats, the virtual speaker 14N12, and the virtual speaker 14N16 at the back on the outer side of the small performance site 70. The applause and the sound of "Wow" and the like are output from the plurality of virtual speakers.

[0058] Thereby, the performer of the stage 50 can hear the sound of the listener, the applause, the sound of "Wow" and the like from a place other than the listener LI, and can perform live in an environment rich in the sense of presence. Further, the listener LI located at the small performance site 70 can also hear the sound of a large number of listeners, the applause, the sound of "Wow" and the like in the same space, and can view and listen to the live performance in an environment rich in the sense of presence.

[0059] In particular, the sound signal processing method of the present embodiment can emit the sound of the listener, the applause, the sound of "Wow" and the like from the virtual speakers 14N1 to 14N16 more than the actually existing speakers 14A to 14G. Thus, the sound signal processing method of the present embodiment can output the sound of the listener, the applause, the sound of "Wow" and the like from various positions with less equipment, and can improve the sense of presence. Further, the sound signal processing method of the present embodiment can simulate the environment of a larger conference site than the actually existing space by setting the position of the virtual speaker at a position on the outer side of the space of the actually existing conference site, and output the sound of the listener, the applause, the sound of "Wow" and the like.

[0060] In the above-described embodiment, an example of improving the sense of presence in the small performance site 70 is described. However, the sound signal processing method of the present embodiment can also improve the sense of presence for each listener who remotely uses the information processing terminal 13.

[0061] Figure 8 、 Figure 9 and Figure 10 is a plan view schematically showing the audiovisual environment of each listener who uses the information processing terminal 13. In this example, the speaker 14FL, the speaker 14FR, the speaker 14C, the speaker 14SL, and the speaker 14SR are arranged along the wall surface of the living room 75. The living room 75 of this example is rectangular in plan view. The display 55 is arranged in front of the living room 75. The listener L2 is located in the center of the living room. The listener L2 watches the performance of the performer displayed on the display 55.

[0062] The speaker 14FL is arranged on the left side of the display 55, the speaker 14C is arranged in front of the display 55, and the speaker 14FR is arranged on the right side of the display 55. The speaker 14SL is arranged on the left side at the back of the living room 75, and the speaker 14SR is arranged on the right side at the back of the living room 75.

[0063] The information processing terminal 13 acquires the image and the sound involved in the performance of the performer. For example, in the example of Figure 2 , the mixer 11 acquires the sound of the performance of the performer, such as the sound of playing a musical instrument or singing, and transmits it to the information processing terminal 13.

[0064] The information processing terminal 13, like the mixer 11, applies signal processing such as sound image processing and effect processing to the acquired sound, and outputs the sound signal after the signal processing to the speaker 14FL, the speaker 14FR, the speaker 14C, the speaker 14SL, and the speaker 14SR. The speaker 14FL, the speaker 14FR, the speaker 14C, the speaker 14SL, and the speaker 14SR output the sound involved in the performance of the performer.

[0065] Further, the information processing terminal 13 acquires the sound signal involved in the cheering, clapping, or calling of other listeners from other information processing terminals 13. The information processing terminal 13, like the mixer 11, sets the type of the sound signal and performs positioning processing or dispersion processing.

[0066] Thus, as shown in Figure 9 , the listener L2 can obtain the sense of presence as if he or she is located in the center of the small performance site 70 and watches the performance of the performer together with a large number of audience members even in the living room 75.

[0067] The information processing terminal 13 can also have a seat designation reception unit that receives designation information of a seat position from a listener. In this case, the information processing terminal 13 changes the contents of the sound image processing and the effect processing based on the designation information of the seat position. For example, if a listener designates a seat position in front of the stage 50, the information processing terminal 13 sets the listener L2 at a position in front of the stage 50 and sets a plurality of virtual speakers, and performs positioning processing and dispersion processing of sound signals involved in the sound of the other listeners' cheers, applause, or calls, as shown in FIG. 8. Thus, the listener L2 can obtain a sense of presence as if he or she were located in front of the stage 50. Figure 10

[0068] The provider of the sound signal processing system provides tickets for a seat position in front of the stage, a seat position on the side of the stage, a seat position in the center of a small performance venue, or a seat position at the back, or the like. The user of the information processing terminal 13 purchases a ticket for any one of these seat positions. The user can select a seat position in front of the stage for a high price to obtain a high sense of presence, or select a seat position at the back for a low price, for example. The information processing terminal 13 changes the contents of the sound image processing and the effect processing in accordance with the seat position selected by the user. Thus, the user can obtain a sense of presence as if he or she were located in the seat position that he or she purchased to view the performance. Furthermore, the provider of the sound signal processing method can conduct a business equivalent to organizing an event in a space that actually exists.

[0069] Further, in the sound signal processing method of the present embodiment, a plurality of users can designate the same seat position. For example, a plurality of users can each designate a seat position in front of the stage 50. In this case, the information processing terminal 13 of each user imparts a sense of presence as if he or she were located in the seat position in front of the stage 50. Thus, a plurality of listeners can view the performance of the performer with the same sense of presence for one seat. Consequently, the provider of the sound signal processing method can provide a service that exceeds the number of viewers that can be accommodated in a space that actually exists.

[0070] The description of the present embodiment is illustrative in all respects and is not restrictive. The scope of the present application is not shown by the above-described embodiment but by the claims. Further, within the scope of the present application, all modifications within the meaning and range equivalent to the claims are intended to be included.

[0071] ​For example, in the above-described embodiment, the sound recognition processing is performed on the sound signal, and the type of the sound signal is determined as the first type in a case where the sound is recognized by the sound recognition processing, and the type of the sound signal is determined as the second type in a case where the sound cannot be recognized by the sound recognition processing. However, the sound signal can also include a plurality of sound channels, and include additional information (metadata) indicating whether each sound channel is the first type or the second type. For example, in a case where the information processing terminal 13 receives a selection operation of "cheering", "clapping", "calling", "noisemaking", and the like from the listener and generates a sound signal corresponding thereto, the information processing terminal 13 generates a sound signal of a sound channel corresponding to the selected sound, attaches additional information, and transmits the sound signal to the mixer 11. In this case, the determination section 302 of the mixer 11 determines the type of the sound signal on the basis of the additional information for each sound channel.

[0072] In addition, the sound signal can also include sound sources of both the first type and the second type. In this case, the mixer 11 (or the information processing terminal 13) separates the sound sources of the sound signal of the first type and the sound signal of the second type. The positioning processing section 304 and the dispersion processing section 305 generate the first sound signal and the second sound signal from each of the separated sound signals. The method of sound source separation can be any method. For example, as described above, the first type is the utterance sound of a specific listener. Therefore, the determination section 302 separates the sound signal of the first type using noise reduction processing that targets the utterance sound as a target sound and deletes other sounds as noise.

[0073] Label Explanation

[0074] 1 Sound signal processing system

[0075] 11 Mixer

[0076] 13 Information processing terminal

[0077] 14A to 14G Loudspeakers

[0078] 14FL, 14FR, 14C, 14SL, 14SR Loudspeakers

[0079] 14N1 to 14N16 Virtual loudspeakers

[0080] 50 Stage

[0081] 55 Display

[0082] 70 Small concert venue

[0083] 75 Living room

[0084] 101 Display

[0085] 102... user I / F

[0086] 103... audio I / O

[0087] 104... signal processing section

[0088] 105... network I / F

[0089] 106... CPU

[0090] 107... flash memory

[0091] 108... RAM

[0092] 171... bus

[0093] 301... acquisition section

[0094] 302... determination section

[0095] 303... setting section

[0096] 304... positioning processing section

[0097] 305... dispersion processing section

[0098] 306... addition section

Claims

1. A sound signal processing method, acquiring a sound signal containing both a first type and a second type of sound source, determining the type of the sound signal, setting a plurality of virtual speakers, when the determined type of the sound signal is the first type, generating a first sound signal obtained by implementing a positioning process that positions a sound image at any one of the plurality of virtual speakers, when the determined type of the sound signal is the second type, generating a second sound signal obtained by implementing a dispersion process that dispersively positions a sound image at two or more of the plurality of virtual speakers, adding the first sound signal and the second sound signal to generate an added signal, outputting the added signal to a plurality of actually existing speakers.

2. The sound signal processing method according to claim 1, wherein the sound signal contains a plurality of channels, the type is determined for each channel.

3. The sound signal processing method according to claim 1, wherein in a case where the sound signal contains both the first type and the second type of sound source, performing sound source separation on the first type of sound signal and the second type of sound signal, generating the first sound signal and the second sound signal from each of the separated sound signals.

4. The sound signal processing method according to any one of claims 1 to 3, wherein performing a sound recognition process on the sound signal, determining the type of the sound signal to be the first type in a case where a sound is recognized by the sound recognition process, determining the type of the sound signal to be the second type in a case where a sound is not recognized by the sound recognition process.

5. The sound signal processing method according to any one of claims 1 to 3, wherein the positioning process includes a process of individually outputting the first sound signal to an actually existing speaker in a case where the position of the positioning coincides with the actually existing speaker.

6. The sound signal processing method according to any one of claims 1 to 3, wherein receiving position information from a user, the positioning process positions the first sound signal at the position of the received position information.

7. The sound signal processing method according to any one of claims 1 to 3, wherein the positioning process implements the virtual speaker by a sound image process and an effect process.

8. The sound signal processing method according to claim 7, wherein receiving seat position designation information from a user, based on the seat position designation information, changing the contents of the sound image process and the effect process.

9. The sound signal processing method according to claim 7, wherein the effect process includes delay, equalization, or reverb.

10. The sound signal processing method according to any one of claims 1 to 3, wherein the dispersion process includes adjustment of the output timing of the second sound signal.

11. A sound signal processing apparatus, comprising: acquiring a sound signal including both a first type and a second type of sound source; determining the type of the sound signal; setting a plurality of virtual speakers; and generating a first sound signal obtained by implementing a positioning process that positions a sound image at any one of the plurality of virtual speakers, when the determined type of the sound signal is the first type, generating a second sound signal obtained by implementing a dispersion process that dispersively positions a sound image at two or more of the plurality of virtual speakers, when the determined type of the sound signal is the second type, adding the first sound signal and the second sound signal to generate an added signal, and outputting the added signal to a plurality of actually existing speakers.

12. The sound signal processing apparatus according to claim 11, wherein the sound signal includes a plurality of channels, the determining section determines the type for each channel.

13. The sound signal processing apparatus according to claim 11, wherein the sound signal includes both the first type and the second type of sound source, a sound source separating section separates the first type of sound signal and the second type of sound signal, the first sound signal and the second sound signal are generated from each of the separated sound signals.

14. The sound signal processing apparatus according to any one of claims 11 to 13, wherein a sound recognition processing section performs sound recognition processing on the sound signal, the determining section determines the type of the sound signal as the first type when a sound is recognized by the sound recognition processing, and determines the type of the sound signal as the second type when a sound is not recognized by the sound recognition processing.

15. The sound signal processing apparatus according to any one of claims 11 to 13, wherein the positioning process includes a process of individually outputting the first sound signal to an actually existing speaker when the position of the positioning coincides with the actually existing speaker.

16. The sound signal processing apparatus according to any one of claims 11 to 13, wherein a position information receiving section receives position information from a user, the positioning process positions the first sound signal at the position of the received position information.

17. The sound signal processing apparatus according to any one of claims 11 to 13, wherein the positioning process realizes the virtual speaker by a sound image process and an effect process.

18. The sound signal processing apparatus according to claim 17, wherein a seat position specifying receiving section receives specifying information of a seat position from a user, the signal processing section changes the contents of the sound image process and the effect process based on the specifying information of the seat position.

19. The sound signal processing apparatus according to claim 17, wherein The effect processing includes delay, equalization, or reverb.

20. The sound signal processing apparatus according to any one of claims 11 to 13, wherein The dispersion processing includes adjustment of output timing of the second sound signal.

Citation Information

Patent Citations

  • Acoustic signal compensation device and program thereof

    JP2017200025A

  • Stereophonic sound generator, method for generating stereophonic sound, and medium storing stereophonic sound

    JP2000333297A

  • Voice conference system and device

    JP2008177802A