Sound effect mixing device

The sound effect mixing device addresses storage and timing issues by dynamically adding sound effects based on music analysis, enhancing the naturalness and immersion of live venue experiences.

JP2026053568APending Publication Date: 2026-03-25PIONEER IP
View PDF 1 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-12-24
Publication Date
2026-03-25

AI Technical Summary

Technical Problem

Existing sound effect systems for live music venues require large storage capacity for venue-specific sound source data, fail to account for varying timing and intensity of audience reactions, and produce unnatural sound effects due to repetitive ambient noise and synchronized clapping.

Method used

A sound effect mixing device that analyzes music features to determine when and how to add sound effects like applause, ambient noise, and rhythmic clapping, using a combination of sound source data and dynamic output control to mimic live venue atmospheres.

Benefits of technology

Enhances the naturalness of sound effects by adapting to music genre, venue size, and audience reactions, reducing storage needs and providing a more immersive live venue experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026053568000001_ABST
    Figure 2026053568000001_ABST
Patent Text Reader

Abstract

The present invention provides a sound effect mixing device, method, program, and storage medium that add more natural sound effects to music for the listener. [Solution] The process is performed by a sound effect mixing device, and includes the steps of: S301 acquiring feature quantities of a song or selecting a mode; S302 outputting a song and sound effects based on the acquired feature quantities or selected mode; and S303 mixing sound effects with the song.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a sound effect mixing device.

Background Art

[0002] There is a known technique for adding sound effects to music so that the atmosphere of a live venue can be enjoyed. For example, Patent Document 1 discloses a karaoke sound effect system. In this karaoke sound effect setting system, the type of sound effect is set according to the genre of the music, and the output mode of the sound effect (the number of people making handclaps and cheers) is set according to the scale of the selected live venue.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The sounds and volumes of cheers and applause by the audience vary depending on the scale of the venue. Therefore, by preparing sound source data such as cheers for each venue and changing the sound source data used as sound effects according to the music, it becomes possible to make the listener feel more the atmosphere of being at a live venue. However, if sound source data such as cheers for each venue is prepared, a large-capacity storage device is required to store these sound source data. However, Patent Document 1 does not disclose a technique for solving such problems.

[0005] Furthermore, at live music venues, cheers and applause occur from the audience at the beginning and end of songs. The timing of these cheers and applause varies depending on the size of the venue and the genre of music performed. However, Patent Document 1 does not take into account these differences in the timing of cheers and applause.

[0006] Furthermore, live music venues are bustling with activity even when there is no cheering or applause, due to the large number of people gathered there. Therefore, by adding (mixing) these ambient sounds, such as the background noise, to the music, it is possible to recreate the atmosphere of a live music venue. However, simply outputting a single ambient sound source data repeatedly results in the same sound source data being played regularly, which feels more regular and less natural than the ambient sounds that actually occur at a live music venue. Patent Document 1 does not disclose any technology to solve this problem.

[0007] Furthermore, at live concerts, audience members sometimes clap their hands or stomp their feet to the rhythm of the music. However, audience members do not clap or stomp their hands to the rhythm of every song performed. Patent Document 1 includes clapping in every song, and does not disclose any technology to solve this problem.

[0008] Furthermore, at live concerts, audience members attempt to clap along to the beat of the music, but their clapping rarely perfectly aligns with the beat. Therefore, if the sound of clapping is output to precisely coincide with the beat, listeners may perceive the clapping mixed into the music as mechanical and unnatural. However, Patent Document 1 does not disclose any technology to solve this problem.

[0009] One example of a problem that this invention aims to solve is adding more natural-sounding sound effects to music for the listener. [Means for solving the problem]

[0010] To solve the above problems, the invention described in claim 1 includes a music feature acquisition unit that acquires characteristic quantities of a music, and an output determination unit that determines whether or not to mix sound effects into the music based on the characteristic quantities.

[0011] The invention described in claim 6 comprises a music feature acquisition step of acquiring the feature quantities of a musical piece, and an output determination step of determining whether or not to mix sound effects into the musical piece based on the feature quantities.

[0012] The invention described in claim 7 involves having a computer execute the sound effect output method described in claim 6.

[0013] The invention described in claim 8 stores the sound effect output program described in claim 7. [Brief explanation of the drawing]

[0014] [Figure 1] This is a sound effect mixing device 100 according to one embodiment of the present invention. [Figure 2] This figure shows an example of music output by the music output unit 120 and sound effect output by the sound effect output unit 130. [Figure 3] This figure shows an example of processing operation in a sound effect mixing device 100 according to one embodiment of the present invention. [Figure 4] This figure shows an example of music output by the music output unit 120 and sound effect output by the sound effect output unit 130. [Figure 5] This figure shows an example of processing operation in a sound effect mixing device 100 according to one embodiment of the present invention. [Figure 6] This figure shows an example of music output by the music output unit 120 and sound effect output by the sound effect output unit 130. [Figure 7] This figure shows an example of music output by the music output unit 120 and sound effect output by the sound effect output unit 130. [Figure 8]It is a diagram showing an example of a processing operation in the sound effect mixing device 100 according to an embodiment of the present invention. [Figure 9] It is a diagram showing an example of the output of sound effects by the sound effect output unit 130. [Figure 10] It is a diagram showing an example of a processing operation in the sound effect mixing device 100 according to an embodiment of the present invention. [Figure 11] It is a diagram showing an example of a processing operation in the sound effect mixing device 100 according to this embodiment. [Figure 12] It is a diagram showing an example of the output of music by the music output unit 120 and the output of sound effects by the sound effect output unit 130. [Figure 13] It is a diagram showing an example of the output of music by the music output unit 120 and the output of sound effects by the sound effect output unit 130. [Figure 14] It is a diagram showing an example of a processing operation in the sound effect mixing device 100 according to this embodiment.

Mode for Carrying Out the Invention

[0015] The sound effect output device according to an embodiment of the present invention includes a music feature amount acquisition unit that acquires a feature amount of music, and an output determination unit that determines whether to mix sound effects into the music based on the feature amount. By doing so, it becomes possible to add sound effects according to the feature amount of the music such as the genre of the music. As a result, it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener feel more the atmosphere of being at a live venue.

[0016] The feature amount of the music includes a first index that is an index based on the number of beats per unit time, the volume, and a harmony change parameter indicating the change of harmony in the music. The output determination unit may determine whether to mix sound effects into the music based on the first index. By doing so, clapping will be added to the music only when the music is a lively song. As a result, it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener feel more the atmosphere of being at a live venue.

[0017] The feature quantity of the music includes a second index which is an index based on the power for each band and a harmony clarity parameter indicating the clarity of the harmony in the music, and the output determination unit may determine whether to mix sound effects into the music based on the second index. By doing so, the beat will be added to the music only when the music is a lively song. As a result, it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener experience more the atmosphere of being at a live venue.

[0018] The feature quantity of the music includes a first index which is an index based on the number of beats per unit time, the volume, and a harmony change parameter indicating the change of the harmony in the music, and a second index which is an index based on the power for each band and a harmony clarity parameter indicating the clarity of the harmony in the music, and the output determination unit may determine whether to mix sound effects into the music based on the first index and the second index. By doing so, the beat will be added to the music only when the music is a lively song. As a result, it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener experience more the atmosphere of being at a live venue.

[0019] The feature quantity of the music includes the number of beats per unit time, and when the number of beats per unit time is not within a predetermined range, the output determination unit determines not to mix sound effects into the music, and when the number of beats per unit time is within the predetermined range, the output determination unit may determine whether to mix sound effects into the music based on the calculated first index and second index. By doing so, the beat will be added to the music only when the music is a lively song. As a result, it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener experience more the atmosphere of being at a live venue.

[0020] Furthermore, a sound effect output method according to one embodiment of the present invention includes a music feature acquisition step of acquiring the feature quantities of a song, and an output determination step of determining whether or not to mix sound effects into the song based on the feature quantities. In this way, it becomes possible to add sound effects according to the feature quantities of a song, such as the genre of the song, and as a result, it becomes possible to add more natural sound effects to the song for the listener, making it possible to give the listener a better feel of being at a live venue.

[0021] Furthermore, the sound effect output program according to one embodiment of the present invention causes a computer to execute the above-described sound effect output method. In this way, it becomes possible to add more natural sound effects to the music using a computer, and to allow the listener to experience the atmosphere of being at a live venue more vividly.

[0022] Furthermore, a computer-readable storage medium according to one embodiment of the present invention stores the above-mentioned sound effect output program. In this way, the above-mentioned sound effect output program can be distributed independently, in addition to being incorporated into a device, and version upgrades can be easily performed. [Examples]

[0023] <Sound effect mixing device 100> Figure 1 shows a sound effect mixing device 100 according to one embodiment of the present invention. The sound effect mixing device 100 mixes (adds) sound effects to a song and outputs it so that the listener can experience the atmosphere of listening to a song at a live venue. The sound effect mixing device 100 includes a storage unit 110 that stores song data and sound source data for sound effects, a song output unit 120 that outputs the song, a sound effect output unit 130 that outputs the sound effects, and a mixing unit 140 that mixes the song output from the song output unit 120 with the sound effects output from the sound effect output unit 130. The sound of the song mixed with sound effects by the mixing unit 140 is output from a sound output device such as a speaker.

[0024] The memory unit 110 stores music data and sound source data for sound effects. The memory unit 110 is a storage device such as a hard disk or flash memory.

[0025] The music output unit 120 outputs music. The music output unit 120 acquires music data stored in, for example, the storage unit 110, a CD (Compact Disc), or the cloud, generates a music signal from this acquired data, and outputs the generated music signal.

[0026] The sound effect output unit 130 outputs sound effects. For example, the sound effect output unit 130 acquires sound source data for sound effects stored in the memory unit 110, generates a sound effect signal from the acquired sound source data, and outputs the generated sound effect signal.

[0027] Sound effects can be categorized into three types: first, cheers and applause that occur at the beginning and end of a song at a live venue; second, ambient sounds (bustling noises) that are constantly present at a live venue; and third, handclaps that are performed in time with the rhythm and beat of the music at a live venue.

[0028] Figure 2 shows an example of music output by the music output unit 120 and sound effect output by the sound effect output unit 130. In the example shown in Figure 2, the first sound effect (such as cheers or applause) is added to the beginning and end of the music. The second sound effect (ambient sound) is output before the music output begins and continues to be played throughout the music playback. The third sound effect (such as handclaps) is output at the beat positions of the music while the music is being played, as will be described in detail below.

[0029] The mixing unit 140 mixes the music output from the music output unit 120 with the sound effects output from the sound effects output unit 130, and outputs the music with the sound effects mixed in. The mixing unit 140 is, for example, a device that adds up multiple signals and outputs the added signal. It adds the music signal output from the music output unit 120 and the sound effect signal output from the sound effects output unit 130, and outputs the added signal.

[0030] Furthermore, the sound effect mixing device 100 has a control unit 150 that controls the output of music from the music output unit 120 and the output of sound effects from the sound effect output unit 130. The control unit 150 is composed of a computer, for example, a CPU (Central Processing Unit).

[0031] The control unit 150 includes, for example, a music feature acquisition unit 151 that acquires the features of a song, a mode selection unit 152 that selects one mode from among several modes relating to the style and genre of the song, the size of the live venue, etc., and an output control unit 153 that controls the output of the song from the music output unit 120 and the output of sound effects from the sound effect output unit 130 based on the music features acquired by the music feature acquisition unit 151 and the mode selected by the mode selection unit 152. The sound effect output device of the above embodiment includes, for example, a sound effect output unit 130 and a control unit 150.

[0032] The music feature acquisition unit 151 acquires music features. Music features include, for example, the volume of the music, the position of the beats in the music, the number of beats per unit time (e.g., BPM (Beats Per Minute)), the time signature of the music, the number of types of chords used in the music, the number of chords per unit time, the clarity of the chords, the power of each frequency band, and the position of the chorus in the music.

[0033] The music feature acquisition unit 151 may acquire music features by analyzing the music, or the music features obtained from prior analysis may be stored in the memory unit 110 or the cloud, and the music feature acquisition unit 151 may acquire the music features stored in the memory unit 110 or the cloud. Alternatively, the music feature acquisition unit 151 may acquire music features from tag information attached to music data stored in the memory unit 110 or on a CD or the like.

[0034] For example, the output control unit 153 may control the volume of the sound effects output from the sound effect output unit 130 based on the volume of the music acquired by the music feature acquisition unit 151. By doing so, it becomes possible to prevent the volume of the mixed sound effects from becoming too loud or too quiet compared to the volume of the music, making it possible to add more natural sound effects to the music for the listener and allowing the listener to experience the atmosphere of being at a live venue more fully.

[0035] Furthermore, the output control unit 153 may detect the musical style of a song based on the musical feature quantity acquired by the musical feature quantity acquisition unit 151, and control the output of sound effects from the sound effect output unit 130 based on this detected musical style. In this case, for example, the memory unit 110 may store sound effects for large venues such as stadiums, outdoor festivals, and arenas, sound effects for medium-sized venues such as halls and medium-to-large-sized live music venues, and sound effects for small venues such as small live music venues and music bars, and the output control unit 153 may decide whether to output sound effects for large venues, medium-sized venues, or small venues from the sound effect output unit 130 based on the detected musical style. By doing so, sound effects that match the song will be added, making it possible to output more natural sound effects according to the song. As a result, it becomes possible to add more natural sound effects to the song for the listener, making it possible to make the listener feel more like they are at a live venue.

[0036] The mode selection unit 152 selects one mode from among several modes related to the genre of the song, the size of the live venue, etc. At this time, the mode selection unit 152 may select a mode based on user input, or it may select a mode based on the characteristics of the song, the tag information of the song, etc.

[0037] For example, the multiple modes should include modes tailored to the size of the live venue. For instance, modes for large venues, medium venues, and small venues should be provided. The output control unit 153 then determines the sound effect to be output from the sound effect output unit 130 based on the mode selected by the mode selection unit 152 (i.e., if the large-scale mode is selected, the output sound effect for large venues is determined), and controls the sound effect output unit 130 to output this determined sound effect. In this way, sound effects appropriate to the song are added, making it possible to output more natural sound effects according to the song. As a result, it becomes possible to add more natural sound effects to the song for the listener, allowing the listener to experience the atmosphere of being at a live venue more fully.

[0038] For example, multiple modes should include modes prepared for different musical styles and genres. For instance, a mode for upbeat pop music, a mode for calm pop music, a mode for classical music, and a mode for jazz music could be provided. The memory unit 110 would then store sound effects for each mode, and the output control unit 153 would determine the sound effect to be output from the sound effect output unit 130 based on the mode selected by the mode selection unit 152 (i.e., if the upbeat pop music mode is selected, the sound effect to be output would be an upbeat pop music sound effect), and control the sound effect output unit 130 to output this determined sound effect. In this way, sound effects that match the music will be added, making it possible to output more natural sound effects according to the music. As a result, it becomes possible to add more natural sound effects to the music for the listener, allowing the listener to experience the atmosphere of being at a live venue more vividly.

[0039] <Processing operation in sound effect mixing device 100> Figure 3 shows an example of the processing operation in the sound effect mixing device 100 according to this embodiment. The music feature quantity 151 acquires the features of the music, or the mode selection unit 152 selects a mode (step S301). The output control unit 153 outputs the music using the music output unit 120 and outputs sound effects using the sound effect output unit 130 based on the acquired features or the selected mode (step S302). The mixing unit 140 mixes the music output from the music output unit 120 with the sound effects output from the sound effect output unit 130 (step S303).

[0040] <Outputting sound effects> The sounds and volume of audience cheers, applause, ambient noise, and handclaps vary depending on the size of the venue, for example. Therefore, by preparing sound source data for cheers and other sounds specific to each venue size, and changing the sound source data used as sound effects according to the song, it becomes possible to make listeners feel more like they are at a live concert. However, if sound source data for cheers and other sounds specific to each venue size is prepared, a large-capacity storage device will be required to store this sound source data.

[0041] Therefore, in the sound effect mixing device 100 according to one embodiment of the present invention, sound effects are output by mixing sounds generated in multiple regions. Specifically, in this embodiment, the output control unit 153 outputs multiple sound source data as sound effects, each containing a sound generated in one of the multiple regions, via the sound effect output unit 130. At this time, the output control unit 153 determines the output level of the multiple sound source data based on the mode selected by the mode selection unit 152, for example.

[0042] For example, in this embodiment, the output control unit 153 may output, via the sound effect output unit 130, sound source data for nearby sounds including sounds generated in the vicinity of the reference position (first region), and sound source data for distant sounds including sounds generated far from the reference position (second region further from the reference position than the first region), as sound effects. In this case, the output control unit 153 may determine the output level of the output nearby sound source data and the output level of the output distant sound source data based on the mode selected by the mode selection unit 152, for example. Here, the reference position is, for example, the position of the audience at a live venue. Alternatively, the reference position may be the position of the stage at the live venue.

[0043] Therefore, in this embodiment, for example, by setting the output level of the sound source data for distant sounds to be about the same as the output level of the sound source data for nearby sounds, both cheers and applause occurring nearby and cheers and applause occurring far away are added to the song, allowing the listener to experience the atmosphere of the song being performed in a large live venue. Alternatively, by setting the output level of the sound source data for distant sounds to zero, only cheers and applause occurring nearby are added to the song, allowing the listener to experience the atmosphere of the song being performed in a small live venue. In other words, in this embodiment, it is possible to add sound effects suitable for venues of various sizes to a song using a small amount of sound source data. As a result, it is possible to add more natural sound effects to a song using a small amount of sound source data, allowing the listener to experience the atmosphere of being in a live venue more fully.

[0044] The sound effect for which sound source data for nearby sounds and sound source data for distant sounds is prepared can be a first sound effect (such as cheers or applause), a second sound effect (such as ambient sounds), or a third sound effect (such as handclaps). For example, in the example shown in Figure 4, sound source data for nearby sounds and sound source data for distant sounds are output for the first sound effect and the second sound effect.

[0045] For example, by including modes for large venues, medium venues, and small venues among the multiple modes, this embodiment makes it possible to add sound effects to the music according to the size of the venue. For example, when the mode selection unit 152 selects the mode for large venues, the output control unit 153 should control the sound effect output unit 130 so that it outputs the sound source data for distant sounds and the sound source data for nearby sounds at the same output level. Also, for example, when the mode selection unit 152 selects the mode for medium venues, the output control unit 153 should control the sound effect output unit 130 so that the output level of the sound source data for distant sounds is lower than the output level of the sound source data for nearby sounds. When the mode selection unit 152 selects the mode for small venues, the output control unit 153 should control the sound effect output unit 130 so that the output level of the sound source data for distant sounds is set to zero and only the sound source data for nearby sounds is output.

[0046] Alternatively, the memory unit 110 may store multiple types of sound source data as sound source data for nearby sounds, and the output control unit 153 may randomly select the sound source data for nearby sounds output from the sound effect output unit 130 from these multiple types of sound source data. Similarly, the memory unit 110 may store multiple types of sound source data as sound source data for distant sounds, and the output control unit 153 may randomly select the sound source data for distant sounds output from the sound effect output unit 130 from these multiple types of sound source data. By doing so, the sound effects output from the sound effect output unit 130 will not be monotonous, and as a result, it will be possible to add more natural sound effects to the music for the listener, making it possible to give the listener a better sense of being at a live venue.

[0047] Furthermore, the memory unit 110 may be configured to store sound source data for reverberation, and the output control unit 153 may output sound source data for reverberation in addition to sound source data for nearby sounds and sound source data for distant sounds. By doing so, the sound effects added to the music will be closer to the sounds that occur in a live venue, and as a result, it will be possible to add more natural sound effects to the music for the listener, allowing the listener to experience the atmosphere of being in a live venue more fully.

[0048] Figure 5 shows an example of the processing operation in the sound effect mixing device 100 according to this embodiment. The mode selection unit 152 selects a mode (step S501). The output control unit 153 determines the output level of the sound source data for nearby sounds and the output level of the sound source data for distant sounds based on the selected mode (step S502), and the sound effect output unit 120 outputs the sound source data for nearby sounds and the sound source data for distant sounds at these determined output levels (step S503).

[0049] In the above embodiment, the variations of the first sound effect are increased by mixing multiple sound source data. However, the variations of the first sound effect can also be increased by using different sound processing techniques (for example, by changing parameter settings).

[0050] Furthermore, in the above embodiment, the output level of each of the multiple sound source data is changed to output a first sound effect suitable for the size of the venue. However, by changing the output level of each of the multiple sound source data, it is also possible to output a first sound effect suitable for other characteristics of the venue (such as the shape of the venue).

[0051] <Output of the first sound effect (cheers, applause, etc.)> At live concerts, cheers and applause erupt from the audience at the beginning and end of each song. The timing of these cheers and applause varies depending on the size of the venue and the genre of music being performed.

[0052] Therefore, in the sound effect mixing device 100 according to one embodiment of the present invention, the method of adding the first sound effect (cheers, applause) to the music is determined based on the genre of the music. At this time, it is preferable that the genre of the music be acquired by the music feature acquisition unit 151. Also, if modes are provided for each musical style or genre, the genre of the music can be determined by selecting a mode with the mode selection unit 152. Thus, in the sound effect mixing device 100 according to this embodiment, the method of adding the first sound effect (cheers, applause) to the music may be determined based on the set mode (i.e., the mode selected by the mode selection unit).

[0053] For example, in the sound effect mixing device 100 according to this embodiment, the output control unit 153 determines the duration for which the first sound effect overlaps with the beginning of the song, based on the genre of the song or the mode selected by the mode selection unit. Based on this determination, it controls the output of the song by the song output unit 120 and the output of the sound effect by the sound effect output unit 130. In particular, in the sound effect mixing device 100 according to this embodiment, the output control unit 153 determines whether or not to output the first sound effect so as to overlap with the beginning of the song, based on the genre of the song or the mode selected by the mode selection unit. Based on this determination, it controls the output of the song by the song output unit and the output of the sound effect by the sound effect output unit. For example, in this embodiment, the sound effect output unit 130 is controlled to output the first sound effect so as to overlap with the beginning of the song, as shown in Figure 6. Also, in this embodiment, the sound effect output unit 130 is controlled to output the first sound effect before the output of the song starts, so as not to overlap with the song, as shown in Figure 7. In this way, this embodiment makes it possible to output sound effects according to the genre of the song. As a result, it becomes possible to add more natural sound effects to the music, allowing listeners to experience the atmosphere of being at a live concert more vividly.

[0054] Furthermore, the timing of outputting sound effects added to the beginning of a song may be determined based on the song's volume and the speed at which its volume changes.

[0055] Furthermore, in the sound effect mixing device 100 according to this embodiment, the timing for outputting the sound effect to be added to the end of the song is determined based on the time remaining until the end of the song and the volume of the song. For example, as shown in Figures 6 and 7, the first sound effect to be added to the end of the song is output when the level of the song falls below a predetermined volume value after a predetermined time te before the end of the song. In this way, it becomes possible to add the first sound effect, such as cheers or applause, to the song at the same timing as the cheers or applause that audiences make when they feel that the song has ended at a live venue.

[0056] The change in volume at the end of a song varies from song to song. Therefore, the timing of outputting sound effects added to the end of a song may be based on the speed of the volume change in the song. By doing so, it becomes possible to output sound effects that are appropriate for the volume change at the end of the song.

[0057] In this case, the predetermined volume value that serves as the threshold may be determined based on the genre of the music and / or the size of the venue. Doing so makes it possible to add more natural sound effects to the music for the listener, allowing the listener to experience the atmosphere of being at a live venue more fully.

[0058] For example, if music genres are divided into two groups, a first group and a second group, the output control unit 153 may control the music output unit 120 and the sound effect output unit 130 so that when the genre of the music belongs to the first group, the output of the first sound effect begins, as shown in Figure 6. Furthermore, when the genre of the music belongs to the second group, the output control unit 153 may control the music output unit 120 and the music output unit 130 so that the music output begins after the output of the first sound effect by the sound effect output unit 130 has finished, as shown in Figure 7.

[0059] Music in calmer genres is often performed in smaller venues than music in upbeat genres. Therefore, the output control unit 153 may, for example, control the sound effect output unit 140 so that when the genre of the music is included in the first group, it starts outputting the first sound effect after a predetermined time te before the end of the music and when the volume of the music falls below a first volume value, and when the genre of the music is included in the second group, it may control the sound effect output unit 140 so that it starts outputting the first sound effect after a predetermined time te before the end of the music and when the volume of the music falls below a second volume value.

[0060] Alternatively, the memory unit 110 may be configured to store multiple types of sound source data for the first sound effect, and the output control unit 153 may randomly select one sound source data from these multiple types of sound source data each time it outputs the first sound effect, and output this sound source data as the first sound effect. In this way, the cheers and applause will change at the beginning and end of the song, and the cheers and applause will also change from song to song, making the sound effects added to the song more natural. This makes it possible to add more natural sound effects to the song for the listener, and allows the listener to experience the atmosphere of being at a live venue more fully.

[0061] In this case, the memory unit 110 may store multiple types of sound source data for the first sound effect according to the size of the venue, and the output control unit 153 may, each time it outputs the first sound effect, randomly select from the multiple types of sound source data for venues of the size based on the mode selected by the mode selection unit 152, and output this sound source data as the first sound effect. By doing so, cheers and applause appropriate to the size of the venue will be added to the song, making it possible to add more natural sound effects to the song for the listener and allowing the listener to experience the atmosphere of being at a live venue more fully. In this case, the sound source data for the first sound effect may be composed of sound source data for nearby sound sources and sound source data for distant sound sources, as described above.

[0062] In addition to the sound source data for the first sound effect applied to the beginning of the song, a separate sound source data file may be prepared for the first sound effect applied to the end of the song. Furthermore, in this case, multiple types of sound source data may be prepared for both the first sound effect applied to the beginning of the song and the first sound effect applied to the end of the song.

[0063] When outputting songs consecutively, the output control unit 153 should, as shown in Figure 6, configure the song output unit 120 to output the next song after the output of the first sound effect added to the end of the previous song has finished. In this case, the time ti between the previous song and the next song may be constant or may be randomly selected.

[0064] Furthermore, for example, when songs of genres included in the second group are output consecutively, the output control unit 153 may, as shown in Figure 7, output the next song using the song output unit 120 when the output of the first sound effect attached to the end of the previous song has finished. Alternatively, when songs of genres included in the second group are output consecutively, the output control unit 153 may not output the first sound effect using the sound effect output unit 130 while the songs are playing consecutively, and may only output the first sound effect using the sound effect output unit 130 at the end of the last song. In this case, the output control unit 153 may also output songs using the song output unit 120 so that there are no gaps between consecutive songs.

[0065] Furthermore, the output level of the first sound effect may be set to gradually increase from zero for a certain period of time from the start of output of the first sound effect. In other words, the first sound effect may be set to fade in.

[0066] Figure 8 shows an example of the processing operation in the sound effect mixing device 100 according to this embodiment. If the genre of the music is a genre included in the first group (step S801, YES), the output control unit 153 first outputs the music using the music output unit 120 (step S802). When a predetermined time tb has elapsed since the music was output (step S803, YES), the output control unit 153 randomly selects one sound source data from a plurality of types of first sound effect sound source data (step S804), and outputs this sound source data as the first sound effect using the sound effect output unit 130 (step S805). Then, after exceeding a predetermined time te before the end of the song (step S806, YES), when the volume of the song falls below a first volume value (step S807, YES), the output control unit 153 randomly selects one sound source data from multiple types of first sound effect sound source data (step S808), and the sound effect output unit 130 outputs this sound source data as the first sound effect (step S809).

[0067] On the other hand, if the genre of the song is not a genre included in the first group but a genre included in the second group (step S801, NO), the output control unit 153 first randomly selects one sound source data from multiple types of first sound effect sound source data (step S810), and the sound effect output unit 130 outputs this sound source data as the first sound effect (step S811). When the output of the first sound effect is finished (step S812, YES), the output control unit 153 outputs the song using the song output unit 120 (step S813). Then, after exceeding a predetermined time te before the end of the song (step S814, YES), when the volume of the song falls below a second volume value (step S815, YES), the output control unit 153 randomly selects one sound source data from multiple types of first sound effect sound source data (step S816), and the sound effect output unit 130 outputs this sound source data as the first sound effect (step S817).

[0068] Furthermore, the output control unit 153 may determine the volume of the first sound effect output by the sound effect output unit 130 and the transition of the volume of the first sound effect based on the volume of the music and the speed at which the volume of the music changes.

[0069] <Output of the second sound effect (ambient sound)> At a live music venue, many people gather, creating a constant buzz even when there is no cheering or applause. Therefore, by adding (mixing) this ambient noise into the music, it is possible to recreate the atmosphere of a live concert. However, simply outputting a single ambient sound source data repeatedly results in the same sound source data being played regularly, which feels more regular and less natural than the ambient sounds that actually occur at a live concert.

[0070] Therefore, in the sound effect mixing device 100 according to one embodiment of the present invention, when a second sound effect is output continuously, some of the second sound effects that are one before and one after the other are output simultaneously, and the time during which these one before and one after the other second sound effects are output simultaneously (overlap time TO) is set randomly. In other words, in this embodiment, when the sound effect output unit 130 outputs a second sound effect, the output control unit 153 randomly selects an overlap time TO, and controls the sound effect output unit 130 to output the next second sound effect by the selected overlap time TO of the end of the previously output second sound effect, as shown in Figure 9.

[0071] By doing this, even if only one type of sound source data is used for the second sound effect, the timing of the output of the second sound effect will not be monotonous but random, and the regularity will be reduced. As a result, it becomes possible to add more natural ambient sounds to the music for the listener, and it becomes possible to make the listener feel more like they are at a live venue.

[0072] The sound source data for the second sound effect may contain only one type of sound source data, or it may contain multiple types of sound source data. If the sound source data for the second sound effect contains only one type of sound source data, it becomes possible to reduce the capacity of the storage unit 110 stored in the sound source data for the second sound effect.

[0073] On the other hand, if the sound source for the second sound effect includes multiple types of sound source data, it is preferable that each of the second sound effects output by the sound effect output unit be the sound of a sound source randomly selected from the multiple types of sound source data. In other words, the output control unit 153 should randomly select one sound source data from the multiple types of sound source data for the second sound effect and control the sound effect output unit 130 to output this selected sound source data as the second sound effect. By doing so, the regularity is reduced, and as a result, it becomes possible to add more natural ambient sounds to the music for the listener, allowing the listener to experience the atmosphere of being at a live venue more vividly.

[0074] Furthermore, if different modes are available for each venue size, the second sound effect audio source data may be prepared separately for each venue size. The output control unit 153 may then select the audio source data based on the mode selected by the mode selection unit 152, and output this audio source data via the sound effect output unit 130.

[0075] The overlap time TO may be selected from a predetermined range, for example. In this case, the overlap time TO may be selected from different ranges depending on the size of the venue, that is, the mode selected by the mode selection unit 152.

[0076] The output control unit 153 should, during the overlap time TO, gradually decrease the output level of the previous second sound effect to zero, that is, fade it out, and gradually increase the output level of the later second sound effect from zero, that is, fade it in. By doing this, the output level of the sound effects will not change even during the overlap time TO, and as a result, it will be possible to add more natural ambient sounds to the music for the listener, allowing the listener to experience the atmosphere of being at a live venue more fully.

[0077] In the above example, the sound effect output consecutively is the second sound effect (ambient sound), but the sound effect output consecutively may also be the first sound effect (cheers, applause, etc.). In other words, two or more sound source data may be output consecutively as the first sound effect added to the beginning or end of a song.

[0078] Figure 10 shows an example of the processing operation in the sound effect mixing device 100 according to this embodiment. When music playback starts (step S1001, YES), the output control unit 153 outputs a second sound effect using the sound effect output unit 130 (step S1002). If music playback has not finished (step S1003, NO), the output control unit 153 randomly selects an overlap time (step S1004), and outputs a second sound effect using the sound effect output unit 130 at a time before the overlap time of the end of output of the previously outputted second sound effect (step S1005, YES), so that it overlaps with the previously outputted second sound effect by the selected overlap time (step S1002). If music playback has finished (step S1003, YES), the process ends. In the processing operation shown in Figure 10, if multiple songs are played continuously, in step S1003, it is checked whether or not this continuous playback has finished, and when continuous playback has finished, it is determined that music playback has finished.

[0079] <Detection of the third sound effect (such as clapping)> At live concerts, audience members sometimes clap or stomp their feet to the rhythm of the music. However, they don't clap or stomp their feet to every song that's played. For example, they might clap or stomp to upbeat songs.

[0080] Therefore, in the sound effect mixing device 100 according to one embodiment of the present invention, the output control unit 153 determines whether or not to mix (add) a third sound effect (such as clapping) to the song based on the characteristics of the song. In this way, for example, the output control unit 153 can determine whether or not the song is upbeat based on the characteristics of the song, and add a third sound effect such as clapping only to upbeat songs. As a result, it becomes possible to add more natural sound effects to the song for the listener, and to make the listener feel more like they are at a live concert.

[0081] For example, the music feature acquisition unit 151 may acquire the number of beats per unit time, such as BPM, as a feature of the music. The output control unit 153 then determines whether this number of beats per unit time is within a predetermined range. If the number of beats per unit time is not within the predetermined range, it may decide not to mix a third sound effect into the music. If the number of beats per unit time is within the predetermined range, it may decide to mix a third sound effect into the music.

[0082] Alternatively, for example, the music feature acquisition unit 151 may acquire a first index as a feature of the music based on volume (for example, the maximum volume in the music), the number of beats per unit time, and chord change parameters indicating chord changes in the music, and the output control unit 153 may decide whether or not to mix a third sound effect into the music based on the first index. The first index is an index for determining whether or not the music is upbeat, and can be obtained, for example, by weighting and adding the number of beats per unit time, volume, and chord change parameters.

[0083] By looking at the number of beats per unit time, it is possible to determine whether a song has a fast tempo or not. Fast tempo songs are likely to be catchy and have a rhythm that makes you want to tap your feet. Also, songs with a loud volume are likely to be catchy and have the power to overwhelm and engage the audience. Furthermore, songs with large chord changes are likely to be catchy and have the power to excite the audience. Chord change parameters include the number of chords per unit time, the number of different chord types in the song, and parameters obtained by weighting and adding these parameters.

[0084] Therefore, by using the first indicator, it becomes possible to determine whether a song is upbeat, has a fast tempo, loud volume, and many chord changes. For example, if the first indicator is defined such that the faster the tempo, the louder the volume, and the more chord changes a song has, the output control unit 153 will decide to add the third sound effect to the song if the first indicator calculated for the song is below the first threshold, and decide not to add the third sound effect to the song if the first indicator is above the first threshold, and decide not to add the third sound effect to the song if the song is not upbeat. In this way, handclaps will only be added to the song when it is upbeat. As a result, it becomes possible to add more natural sound effects to the song for the listener, and it becomes possible to make the listener feel more like they are at a live concert.

[0085] Alternatively, for example, the music feature acquisition unit 151 may acquire a second index as a feature of the music, based on the power of each frequency band and a chord clarity parameter indicating the clarity of chords in the music, and the output control unit 153 may decide whether or not to mix a third sound effect into the music based on the second index. The second index is an index for determining whether or not the music is upbeat, and can be obtained, for example, by weighting and adding the power ratio between frequency bands (for example, the power ratio between the low frequency range and the high frequency range) or the chord clarity parameter.

[0086] By examining the power levels in each frequency band, it's possible to determine whether a song has a strong rhythm section, such as bass or kick drum. Songs with strong rhythm sections are likely to be upbeat and inviting, making you want to tap your feet. Furthermore, songs with clear chords provide a sense of peace and tranquility, while songs with unclear chords sound noisy and emphasize the rhythm. Therefore, songs with unclear chords are likely to be upbeat and inviting, making you want to tap your feet. As a chord clarity parameter, for example, the harmony clarity of WO2009 / 104269 is a good choice. This harmony clarity value also indicates the clarity of chords in a song.

[0087] Therefore, by using the second indicator, it becomes possible to determine whether a song is a catchy song with a strong rhythm section and unclear chords. For example, if the second indicator is defined such that the stronger the rhythm section and the less clear the chords, the higher the second indicator will be, the output control unit 153 will decide to add the third sound effect to the song if the second indicator calculated for the song is above the second threshold, and decide not to add the third sound effect to the song if the second indicator is below the second threshold, and decide not to add the third sound effect to the song if the song is not catchy. In this way, handclaps will only be added to the song when it is a catchy song. As a result, it becomes possible to add more natural sound effects to the song for the listener, and it becomes possible to make the listener feel more like they are at a live concert.

[0088] In the above example, the sound effect whose inclusion in the song is determined based on the song's features is the third sound effect (handclaps). However, the sound effect whose inclusion is determined based on the above example could also be the first sound effect (cheers, applause, etc.) or the second sound effect (ambient sounds). In other words, the decision of whether or not to add the first sound effect to the beginning or end of the song, or whether or not to add the second sound effect during the song's generation, could also be made based on the song's features.

[0089] Figure 11 shows an example of the processing operation in the sound effect mixing device 100 according to this embodiment. The output control unit 153 checks whether all three conditions are met: the first condition "the BPM of the music is within a predetermined range", the second condition "the first indicator is less than or equal to the first threshold", and the third condition "the second indicator is greater than or equal to the second threshold" (step S1101). If all three conditions are met (step S1101, YES), the output control unit 153 decides to mix the third sound effect into the music and outputs the third sound effect using the sound effect output unit 130 (step S1102). If all three conditions are not met (step S1101, NO), the output control unit 153 decides not to mix the third sound effect into the music (step S1103).

[0090] <Output of the third sound effect (handclaps)> At live concerts, audiences attempt to clap along to the beat of the music, but often their clapping doesn't perfectly align with the beat. Therefore, if the clapping sounds are output to precisely coincide with the beat, listeners may perceive the clapping mixed into the music as mechanical and unnatural.

[0091] Therefore, in the sound effect mixing device 100 according to one embodiment of the present invention, for example, as shown in Figure 12, the output control unit 153 outputs a third sound effect (such as clapping) from the sound effect output unit 120 after a first time t1 has elapsed from the beat position Bp of the song acquired by the song feature acquisition unit 151. For this reason, in this embodiment, the clapping does not perfectly coincide with the beat position. As a result, it becomes possible to add more natural sound effects to the song for the listener, and it becomes possible to make the listener feel more like they are at a live venue.

[0092] The first time point t1 may not be constant but randomly selected. This randomizes the time lag between the beat and the start of the clapping, resulting in a more natural sound effect for the listener and allowing them to experience the atmosphere of a live concert more vividly.

[0093] Furthermore, in the sound effect mixing device 100 according to one embodiment of the present invention, multiple types of sound source data for the third sound effect are prepared, and after a first time has elapsed from the beat position, sound source data randomly selected from the multiple types of sound source data is output. In other words, in this embodiment, the storage unit 110 stores multiple types of sound source data for the third sound effect, and the output control unit 153 outputs sound source data randomly selected from the multiple types of sound source data as the third sound effect, after a first time has elapsed from the beat position of the music acquired by the music feature acquisition unit 151, using the sound effect output unit 120. As a result, in this embodiment, the handclaps that are mixed in are not monotonous, and as a result, it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener feel more like they are at a live venue.

[0094] When recording the clapping of multiple people for each of several types of sound source data, the clapping of the multiple people will never be perfectly synchronized. Therefore, the timing of the clapping among the multiple people will change with each recording. In such cases, the output control unit 153 may output a sound source data randomly selected from the multiple types of sound source data as a third sound effect at the beat position Bp of the song, as shown in Figure 13, via the sound effect output unit 120. Doing so will also prevent all the clapping from perfectly matching the beat position, resulting in a more natural sound effect for the listener and allowing the listener to experience the atmosphere of a live venue more vividly.

[0095] The third sound effect may be applied from the beginning to the end of the song, or, as shown in Figure 13, it may be applied only for a portion of the song. For example, the timing of when to start outputting the third sound effect may be determined based on the time elapsed since the start of the song's output and the song's volume. For example, the output control unit 153 may start outputting the third sound effect at the beat position of the song (or 1 time elapsed since the beat position of the song) after the volume of the song has reached or exceeded the third volume value, which is done after a second time t2 has elapsed since the start of the song's output. In this way, it becomes possible to start adding clapping at the timing when the audience starts to get into the groove of the song, and as a result, it becomes possible to add a more natural sound effect to the song for the listener, allowing the listener to experience the atmosphere of being at a live venue more fully. If the volume of the song reaches or exceeds the third volume value before the second time t2 has elapsed since the start of the song's output, the output control unit 153 may start outputting the third sound effect when the second time t2 has elapsed since the start of the song's output.

[0096] Furthermore, for example, the timing for ending the output of the third sound effect should be determined based on the position where the song changes parts, such as the start of the chorus, and the volume of the song. For example, the song feature acquisition unit 151 can acquire the position where the song changes parts, and the output control unit 153 can stop the output of the third sound effect three times after the acquired position where the song changes parts. In this way, for example, it becomes possible to end the addition of handclaps at a timing when the audience is listening intently to the chorus of the song, and as a result, it becomes possible to add a more natural sound effect to the song for the listener, allowing the listener to experience the atmosphere of being at a live venue more fully. The output of the third sound effect may also be terminated four times after the output of the third sound effect has started.

[0097] Furthermore, when starting the output of the third sound effect, it is best to gradually increase the output level of the third sound effect from zero over time, that is, to fade it in. Similarly, when ending the output of the third sound effect, it is best to gradually decrease the output level of the third sound effect to zero for three hours before it ends, that is, to fade it out. Doing so allows listeners to experience the atmosphere of some audience members' clapping spreading throughout the venue, and the atmosphere of the audience's clapping gradually fading out. As a result, it becomes possible to add more natural sound effects to the song and allow listeners to experience the atmosphere of being at a live venue more vividly.

[0098] The third sound effect may be output at each beat position in the song, as shown in Figure 13, or it may be output at some of the beat positions in the song. For example, at a live concert, the audience will clap along to upbeat songs. That is, if the song is in 4 / 4 time, the audience will clap on the second and fourth beats of each measure. Therefore, if the song is in 4 / 4 time, the third sound effect may be output at the second and fourth beats of each measure. Alternatively, patterned clapping may be incorporated.

[0099] Figure 14 shows an example of the processing operation in the sound effect mixing device 100 according to this embodiment. When a second time has elapsed since the start of the song (step S1401, YES) and the volume of the song is equal to or greater than a third volume value (step S1402, YES), the output of the third sound effect to the beat position of the song is started (step S1403). When a third time has elapsed since the start of the chorus of the song (step S1404, YES), the output of the third sound effect is stopped (step S1405).

[0100] The present invention has been described above with reference to preferred embodiments. Although the present invention has been described with reference to specific examples, various modifications and changes can be made to these examples without departing from the spirit and scope of the invention as described in the claims. [Explanation of Symbols]

[0101] 100 Sound Effects Mixing Device 110 Storage section 120 Music Output Section 130 Sound effect output section 140 Mixing section 150 Control Unit 151 Music Feature Acquisition Unit 152 Mode Selection Section 153 Output Control Unit

Claims

1. A sound effect output device that outputs sound effects to be mixed into the beginning and / or end of a song, A sound effect output device that determines how to add sound effects to a song based on the genre of the song or a set mode.

2. The sound effect output device according to claim 1, which determines the duration for which the sound effect overlaps with the beginning of the song based on the genre of the song or a set mode.

3. The sound effect mixing device according to claim 2, which determines whether or not to output a sound effect that overlaps with the beginning of the song, based on the genre of the song or a set mode.

4. The sound effect output device according to any one of claims 1 to 3, wherein the timing of outputting a sound effect added to the end of the aforementioned song is determined based on the time remaining until the end of the song and the volume of the song.

5. The sound effect output device according to claim 4, wherein the sound effect to be mixed into the end portion of the aforementioned song is output after a predetermined time before the end of the song and when the volume of the song falls below a predetermined volume value.

6. The sound effect output device according to claim 5, wherein the predetermined volume value is determined based on the genre of the music or a set mode.

7. The system includes a music feature acquisition unit that acquires the characteristic quantities of the music, including the genre of the music. The genre of the aforementioned song is determined based on the characteristics of the song acquired by the song feature acquisition unit, as described in any one of claims 1 to 6, for the sound effect output device.

8. The sound effect output device according to any one of claims 1 to 7, wherein the sound effect to be mixed into the beginning and / or end of the aforementioned music is randomly selected from a plurality of types of sound source data for the sound effect.

9. A sound effect output method that outputs sound effects to be mixed into the beginning and / or end of a song, wherein the method determines how to apply the sound effects to the song based on the genre of the song or a set mode.

10. A sound effect output program that causes a computer to execute the sound effect output method described in claim 9.

11. A computer-readable storage medium storing the sound effect output program described in claim 10.

Citation Information

Patent Citations

  • Karaoke effective sound setting system

    JP2016070999A