Sound effect mixing apparatus
The sound effect mixing device analyzes music features to synchronize sound effects with song characteristics, ensuring they are added only to catchy songs with clear beats and appropriate rhythms, thereby enhancing the naturalness of the audio experience.
Patent Information
- Application Number
- JP2025161346
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2026-01-06
AI Technical Summary
Existing sound effect systems fail to add sound effects to music in a manner that sounds natural to listeners, as audience participation is inconsistent and not synchronized with every song.
A sound effect mixing device that includes a music feature acquisition unit to analyze music features such as clarity of beats, number of beats per unit time, and uniformity of volume levels, determining when to mix sound effects like applause or environmental sounds based on these features.
The device adds sound effects only to catchy songs with clear beats and appropriate rhythms, enhancing the naturalness of the audio experience, mimicking a live concert atmosphere.
Smart Images

Figure 2026001116000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a sound effect mixing device. [Background technology]
[0002] There is known a technology that adds sound effects to music to create the atmosphere of a live concert venue. For example, Patent Document 1 discloses a karaoke sound effect system, in which the type of sound effect is set according to the genre of the music, and the output mode of the sound effect (the number of people clapping and cheering) is set according to the size of the selected live concert venue. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Japanese Patent Application Laid-Open No. 2016-70999 Summary of the Invention [Problem to be solved by the invention]
[0004] At live concerts, audience members often clap their hands or stomp their feet to keep rhythm with the music. However, audience members do not clap their hands or stomp their feet to keep rhythm with every song that is played. Patent Document 1 includes clapping for every song, and does not disclose any technology to solve this problem.
[0005] One example of a problem that the present invention aims to solve is how to add sound effects to music that sound more natural to a listener. [Means for solving the problem]
[0006] In order to solve the above problem, the invention described in claim 1 comprises a music feature acquisition unit that acquires features of the music, including the clarity of the beats of the music, and an output determination unit that determines whether or not to mix sound effects into the music, based on the clarity of the beats of the music.
[0007] The invention described in claim 3 has a music feature acquisition unit that acquires features of the music, including the number of beats per unit time of the music and the uniformity of the volume level at the positions of the beats of the music, and an output determination unit that determines whether or not to mix sound effects into the music, based on the number of beats per unit time of the music and the uniformity of the volume level at the positions of the beats of the music.
[0008] The invention described in claim 6 comprises a music feature acquisition unit that acquires features of the music, including the number of beats per unit time of the music, and an output determination unit that determines whether or not to mix sound effects into the music, based on the number of beats per unit time of the music.
[0009] The invention described in claim 6 is a sound effect mixing method executed by a computer, comprising: a music feature acquisition step of acquiring music feature amounts including clarity of the beats of the music; and an output determination step of determining whether or not to mix sound effects into the music based on the clarity of the beats of the music.
[0010] The seventh aspect of the present invention allows a computer to execute the sound effect mixing method of the sixth aspect.
[0011] The invention described in claim 8 stores the sound effect mixing program described in claim 7. [Brief explanation of the drawings]
[0012] [Figure 1] 1 shows a sound effect mixing device 100 according to an embodiment of the present invention. [Figure 2] 10A and 10B are diagrams showing an example of music output by a music output unit 120 and sound effects output by a sound effect output unit 130. FIG. [Figure 3]1 is a diagram showing an example of a processing operation in the sound effect mixing device 100. FIG. [Figure 4] 1 is a diagram showing an example of a processing operation in the sound effect mixing device 100. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0013] A sound effect output device according to one embodiment of the present invention includes a music feature acquisition unit that acquires music feature amounts, including the clarity of the beats of the music, and an output determination unit that determines whether to mix sound effects into the music based on the clarity of the beats of the music. Therefore, in this embodiment, it is possible to add sound effects only to songs that are catchy. As a result, it is possible to add sound effects to music that sound more natural to listeners, allowing listeners to experience the atmosphere of being at a live concert.
[0014] The feature amount of the music may include the number of beats per unit time of the music, and the output determination unit may determine whether to mix sound effects into the music based on the clarity of the beats of the music and the number of beats per unit time of the music. In this way, it is possible to add sound effects only to music with clear beats and a good rhythm.
[0015] A sound effect output device according to one embodiment of the present invention includes a music feature acquisition unit that acquires music feature quantities, including the number of beats per unit time of the music and the uniformity of volume levels at the positions of the beats, and an output determination unit that determines whether to mix sound effects into the music based on the number of beats per unit time of the music and the uniformity of volume levels at the positions of the beats. Therefore, this embodiment can add sound effects only to music with a good rhythm, and can avoid mixing sound effects into music with a "four-on-the-floor" tempo that is too fast. As a result, it is possible to add sound effects to music that sound more natural to listeners, allowing listeners to experience the atmosphere of being at a live concert.
[0016] The feature amount of the music may include clarity of the beats of the music, and the output determination unit may determine whether to mix sound effects into the music based on the clarity of the beats of the music, the number of beats per unit time of the music, and the uniformity of volume levels at the positions of the beats of the music. In this way, it is possible to add sound effects only to music with clear beats and a good rhythm.
[0017] The output determination unit may determine not to mix sound effects into the music if the volume level of the music remains below a predetermined volume level for a predetermined time from the start of the music. In this way, it is possible to prevent sound effects from being mixed into special music, such as music with a long intro.
[0018] A sound effect output device according to one embodiment of the present invention includes a music feature acquisition unit that acquires music feature amounts, including the number of beats per unit time of the music, and an output determination unit that determines whether to mix sound effects into the music based on the number of beats per unit time of the music. Therefore, in this embodiment, it is possible to add sound effects only to songs that are catchy. As a result, it is possible to add sound effects to music that sound more natural to listeners, allowing listeners to experience the atmosphere of being at a live concert.
[0019] The output determination unit may determine not to mix sound effects into the music if the number of beats per unit time of the music is not within a predetermined range. In this way, it becomes possible to avoid mixing sound effects into music that has a tempo that is too fast and makes it difficult to keep the rhythm by clapping or the like.
[0020] Furthermore, a sound effect output method according to one embodiment of the present invention is a computer-executed sound effect mixing method, and includes a music feature acquisition step of acquiring music feature amounts, including the clarity of the beats of the music, and an output determination step of determining whether to mix sound effects into the music based on the clarity of the beats of the music. Therefore, in this embodiment, it is possible to add sound effects only to songs that are catchy. As a result, it is possible to add sound effects to music that sound more natural to listeners, allowing listeners to experience the atmosphere of being at a live concert.
[0021] A sound effect output program according to an embodiment of the present invention causes a computer to execute the above-described sound effect output method, thereby enabling the use of a computer to more effectively give listeners the feeling of being at a live concert venue.
[0022] Furthermore, a computer-readable storage medium according to one embodiment of the present invention stores the above-described sound effect output program, which allows the above-described sound effect output program to be distributed as a standalone program in addition to being incorporated into a device, and makes it easy to perform version upgrades, etc. [Example]
[0023] <Sound effect mixer 100> 1 shows a sound effects mixing device 100 according to an embodiment of the present invention. The sound effects mixing device 100 mixes (adds) sound effects to music and outputs the music so that listeners can experience the atmosphere of listening to music at a live venue. The sound effects mixing device 100 includes a storage unit 110 that stores music data, sound source data for sound effects, etc., a music output unit 120 that outputs music, a sound effects output unit 130 that outputs sound effects, and a mixing unit 140 that mixes the music output from the music output unit 120 with the sound effects output from the sound effects output unit 130. The sound of the music mixed with the sound effects by the mixing unit 140 is output from a sound output device such as a speaker.
[0024] The storage unit 110 stores music data and sound source data for sound effects. The storage unit 110 is a storage device such as a hard disk or flash memory.
[0025] The music output unit 120 outputs music. The music output unit 120 acquires music data stored in the storage unit 110, a CD (Compact Disc), or on the cloud, for example, generates a music signal from the acquired data, and outputs the generated music signal.
[0026] The sound effect output unit 130 outputs sound effects. For example, the sound effect output unit 130 acquires sound source data for sound effects stored in the storage unit 110, generates a sound effect signal from the acquired sound source data, and outputs the generated sound effect signal.
[0027] Sound effects include first sound effects such as cheers and applause that occur at the beginning and end of a song at a live venue, second sound effects such as ambient noise (bustling sounds) that occur constantly at a live venue, and third sound effects such as clapping in time with the rhythm and beat of the song at a live venue.
[0028] 2 is a diagram showing an example of music output by the music output unit 120 and sound effects output by the sound effect output unit 130. In the example shown in FIG. 2, first sound effects (such as cheers and applause) are added during the music or at the beginning and end of the music. Second sound effects (environmental sounds) are output before the music starts to be output and are output throughout the music being played. Third sound effects (such as clapping) are output in synchronization with the beat and tempo of the music while the music is being output, as will be described in detail below.
[0029] The third sound effect may be output in synchronization with all the beats of the music, or may be output in synchronization with only some of the beats of the music. For example, at a live concert, the audience claps their hands to a catchy song. That is, if the song is in 4 / 4 time, the audience claps their hands on the second and fourth beats of each measure. Therefore, if the song is in 4 / 4 time, the third sound effect may be output on the second and fourth beats of each measure. Alternatively, a patterned clapping may be added.
[0030] The mixing unit 140 mixes the music output from the music output unit 120 with the sound effects output from the sound effect output unit 130, and outputs the music mixed with the sound effects. The mixing unit 140 is, for example, a device that adds multiple signals and outputs the added signal, and adds the music signal output from the music output unit 120 and the sound effect signal output from the sound effect output unit 130, and outputs the added signal.
[0031] Furthermore, the sound effect mixing device 100 has a control unit 150 that controls the output of music from the music output unit 120 and the output of sound effects from the sound effect output unit 130. The control unit 150 is configured by, for example, a computer having a CPU (Central Processing Unit) and the like.
[0032] The control unit 150 has, for example, a music feature acquisition unit 151 that acquires the features of the music, a mode selection unit 152 that selects one mode from a plurality of modes related to the melody and genre of the music, the size of the live venue, etc., and an output control unit 153 that controls the output of music from the music output unit 120 and the output of sound effects from the sound effect output unit 130 based on the features of the music acquired by the music feature acquisition unit 151 and the mode selected by the mode selection unit 152.
[0033] The music feature amount acquisition unit 151 acquires music feature amounts. The music feature amounts include, for example, the volume of the music, the position of the beats in the music, the number of beats per unit time (for example, BPM (Beats Per Minute)), the time signature of the music, the clarity of the beats in the music, the uniformity of the volume levels at the positions of the beats in the music, the number of types of chords used in the music, the number of chords per unit time, the clarity of the chords, the power of each band, the position of the chorus in the music, etc.
[0034] The music feature quantity acquisition unit 151 may acquire music feature quantities by analyzing the music, or music feature quantities obtained by a prior analysis may be stored in the storage unit 110 or the cloud, and the music feature quantity acquisition unit 151 may acquire music feature quantities stored in the storage unit 110 or the cloud. Furthermore, the music feature quantity acquisition unit 151 may acquire music feature quantities from tag information attached to music data stored in the storage unit 110, a CD, or the like.
[0035] For example, the output control unit 153 may control the volume of the sound effects output from the sound effect output unit 130 based on the volume of the music acquired by the music feature amount acquisition unit 151. In this way, it is possible to prevent the volume of the mixed sound effects from being too loud or too quiet compared to the volume of the music, making it possible to impart sound effects to the music that are more natural to the listener and allowing the listener to experience the atmosphere of being at a live concert. In addition, the output control unit 153 may detect the melody of the music based on the features of the music acquired by the music feature amount acquisition unit 151, and control the volume of the sound effects output from the sound effect output unit 130 based on the detected melody.
[0036] The output control unit 153 may also detect the level and melody of the music based on the feature quantities of the music acquired by the music feature quantity acquisition unit 151, and control the output of sound effects from the sound effect output unit 130 based on the detected level and melody. In this case, for example, the storage unit 110 may store sound effects for each level and melody of the music, and the output control unit 153 may output the sound effects based on the detected level and melody. For example, the storage unit 110 may store sound effects for large venues such as stadiums, outdoor festivals, and arenas, sound effects for medium-sized venues such as halls and medium- to large-sized live music venues, and sound effects for small venues such as small live music venues and music bars, and the output control unit 153 may determine, based on the detected melody, whether the sound effects output from the sound effect output unit 130 should be large-, medium-, or small-sized venues. This allows sound effects to be added that match the music, making it possible to output more natural sound effects according to the music. As a result, it becomes possible to add sound effects to music that sound more natural to the listener, allowing the listener to experience the atmosphere of being at a live concert.
[0037] The mode selection unit 152 selects one mode from a plurality of modes related to the melody and genre of the music, the size of the live venue, etc. At this time, the mode selection unit 152 may select the mode based on a user input, or may select the mode based on the feature amount of the music, tag information of the music, etc.
[0038] For example, the multiple modes may include modes prepared for different sizes of live music venues. For example, modes for large-scale venues, medium-scale venues, and small-scale venues may be prepared. The output control unit 153 then determines the sound effects to be output from the sound effect output unit 130 based on the mode selected by the mode selection unit 152 (i.e., for example, if a large-scale mode is selected, the output sound effects may be determined to be large-scale sound effects), and may control the sound effect output unit 130 to output the determined sound effects. In this way, sound effects that match the music can be added, making it possible to output more natural sound effects according to the music. As a result, it is possible to add sound effects that sound more natural to the listener, allowing the listener to experience the atmosphere of being at a live music venue.
[0039] For example, the multiple modes may include modes prepared for different musical styles or genres. For example, a mode for upbeat music, a mode for calm music, a mode for classical music, a mode for jazz, and the like may be prepared. The storage unit 110 may then store sound effects for each mode, and the output control unit 153 may determine the sound effects to be output from the sound effect output unit 130 based on the mode selected by the mode selection unit 152 (i.e., when the upbeat music mode is selected, the output sound effects may be determined to be those for upbeat music), and may control the sound effect output unit 130 to output the determined sound effects. In this manner, sound effects suited to the music may be added, making it possible to output more natural sound effects according to the music. As a result, it is possible to add sound effects that are more natural to the listener, allowing the listener to experience the atmosphere of being at a live concert.
[0040] <Processing Operation in Sound Effect Mixing Apparatus 100> 3 is a diagram showing an example of the processing operation of the sound effect mixing device 100 according to this embodiment. The music feature quantity 151 acquires the feature quantity of the music, or the mode selection unit 152 selects a mode (step S301). The output control unit 153 outputs the music from the music output unit 120 and outputs the sound effects from the sound effect output unit 130 based on the acquired feature quantity or the selected mode (step S302). The mixing unit 140 mixes the music output from the music output unit 120 with the sound effects output from the sound effect output unit 130 (step S303).
[0041] <Mixing judgment of third sound effect (e.g., clapping)> At live concerts, audience members often clap their hands or stomp their feet to keep the rhythm of the music. However, audience members do not clap their hands or stomp their feet to keep the rhythm of every song that is played. For example, audience members clap their hands or stomp their feet to keep the rhythm of a catchy song.
[0042] Therefore, in the sound effect mixing apparatus 100 according to an embodiment of the present invention, the output control unit 153 determines whether or not to mix (add) a third sound effect (such as clapping) into a song based on the feature amount of the song. In this way, for example, the output control unit 153 can determine whether or not a song is catchy based on the feature amount of the song, and add a third sound effect such as clapping only to songs that are catchy. As a result, it becomes possible to add sound effects to the song that are more natural to the listener, allowing the listener to experience the atmosphere of being at a live concert.
[0043] Music with a fast tempo is likely to be a catchy song that encourages rhythmic tapping. On the other hand, even if a song has a catchy tempo, if the tempo is too fast, it is difficult to keep the rhythm by clapping, making it difficult to clap. Therefore, it is preferable to mix the third sound effect into the music when the tempo of the music is within a predetermined range. Therefore, for example, the music feature amount acquisition unit 151 may acquire the number of beats per unit time, such as BPM, as a feature amount of the music. The output control unit 153 may then determine whether the number of beats per unit time is within a predetermined range (whether it is greater than a first threshold and less than a second threshold). If the number of beats per unit time is not within the predetermined range (less than the first threshold or greater than the second threshold), the output control unit 153 may decide not to mix the third sound effect into the music. However, if the number of beats per unit time is within the predetermined range (greater than the first threshold and less than the second threshold), the output control unit 153 may decide to mix the third sound effect into the music.
[0044] A song with an unclear beat is likely not a catchy song. Therefore, for example, the song feature amount acquisition unit 151 may acquire the clarity of the beat of the song as a feature amount of the song. Then, the output control unit 153 may decide whether or not to mix the third sound effect into the song based on the clarity of the beat of the song. In this case, for example, the output control unit 153 may decide not to mix the third sound effect into the song if the beat of the song is unclear, and may decide to mix the third sound effect into the song if the beat of the song is clear.
[0045] This makes it possible to add third sound effects, such as clapping, only to songs with a good rhythm. As a result, it becomes possible to add sound effects to songs that sound more natural to the listener, allowing the listener to experience the atmosphere of being at a live concert.
[0046] As an index of the clarity of the beats of a piece of music, for example, the period of the beats of the piece of music or the amplitude of a rhythm that repeats at a period twice the period of the beats of the piece of music can be used. In a catchy piece of music, it is highly likely that the beats are played by a bass drum, snare drum, or the like at the period of the beats of the piece of music or at a period twice the period of the beats of the piece of music. The period of the beats of the piece of music or the amplitude of a rhythm that repeats at a period twice the period of the beats of the piece of music can be detected, for example, by performing frequency resolution processing (e.g., FFT (Fast Fourier Transform) processing) on the envelope of the waveform of the piece of music (see, for example, International Publication No. 2009 / 125489). For example, the output control unit 153 may determine that the beat of the song is clear if the amplitude of the components in the envelope of the waveform of the song near the frequency of the beat of the song and the amplitude of the components near twice the frequency of the beat of the song are greater than a predetermined value, and may determine that the beat of the song is not clear if the amplitude of the components in the envelope of the waveform of the song near the frequency of the beat of the song and the amplitude of the components near twice the frequency of the beat of the song are not greater than a predetermined value.
[0047] Generally, songs with a rhythm such as a bass drum on every beat, known as "four-on-the-floor" songs, are catchy songs. In "four-on-the-floor" songs, handclaps are usually added to every beat. However, in songs with a fast tempo, adding handclaps to every beat can make the song sound hectic. Therefore, it is recommended not to mix a third sound effect into a "four-on-the-floor" song with a fast tempo. For example, "four-on-the-floor" songs generally have a relatively uniform volume level at the beat positions. Therefore, for example, the music feature acquisition unit 151 may acquire, as music feature amounts, the uniformity of the volume level at the beat positions of the music and the number of beats per unit time, such as BPM. Then, the output control unit 153 may determine whether to mix a third sound effect into the music based on the uniformity of the volume level at the beat positions of the music and the number of beats per unit time, such as BPM. In this case, for example, if the volume level at the beat positions of the music is uniform and the number of beats per unit time is greater than the third threshold, the output control unit 153 determines not to mix the third sound effect into the music even if the number of beats per unit time is within the above-mentioned predetermined range.
[0048] By doing this, third sound effects are not added to "four-on-the-floor" songs, which have a tempo that is too fast. As a result, it is possible to add sound effects to songs that sound more natural to the listener, making it possible to give listeners a better sense of the atmosphere of being at a live concert.
[0049] For example, if the music is in 4 / 4 time, the evenness of the volume levels at the beat positions of the music may be calculated using the sum of the volume levels of the 4N-3 (N=1, 2, . . .) beats, the sum of the volume levels of the 4N-2 beats, the sum of the volume levels of the 4N-1 beats, and the sum of the volume levels of the 4N beats. For example, if the value ((B1-B2) / B2) obtained by subtracting the second-largest value B2 from the largest value B1 of the above four volume level sums and dividing the result by the second-largest value B2 is smaller than a predetermined value, the output control unit 153 may determine that the volume levels at the beat positions of the music are even, and if (B1-B2) / B2 is equal to or greater than the predetermined value, it may determine that the volume levels at the beat positions of the music are not even.
[0050] A song whose volume level remains low for a long time at the beginning of the song is likely to be a special song, such as one with a long intro. At live music venues, audience members generally do not clap their hands to keep rhythm while such special songs are being played. Therefore, it is advisable not to mix a third sound effect into such special songs. Therefore, for example, if the volume level of the song remains lower than a predetermined volume level (first volume level) from the start of the song until a predetermined time (first time) has elapsed, the output control unit 153 determines not to mix a third sound effect into the song. In this way, the third sound effect is not added to special songs, such as those with a long intro. As a result, it is possible to add sound effects to the song that sound more natural to the listener, allowing the listener to experience the atmosphere of being at a live music venue.
[0051] In the above, the sound effect that is determined to be mixed into the music based on the feature amount of the music is the third sound effect (claps), but the sound effect that is determined to be mixed may be the first sound effect (cheers, applause, etc.) or the second sound effect (environmental sound). In other words, it may be determined based on the feature amount of the music whether to add the first sound effect during the music or at the beginning or end of the music, or whether to add the second sound effect while the music is being played.
[0052] 4 is a diagram showing an example of the processing operation of the sound effect mixing apparatus 100 according to this embodiment. The output control unit 153 checks whether the number of beats per unit time is greater than a first threshold value and equal to or less than a second threshold value (step S401).
[0053] If the number of beats per unit time is equal to or less than the first threshold or greater than the second threshold (step S401, NO), the output control unit 153 decides not to mix the third sound effect into the music (step S402), and if the number of beats per unit time is greater than the first threshold and equal to or less than the second threshold (step S401, YES), the output control unit 153 checks whether the beats of the music are clear (step S403).
[0054] If the beats of the music are not clear (step S403, NO), the output control unit 153 decides not to mix the third sound effect into the music (step S402), and if the beats of the music are clear (step S403, YES), the output control unit 153 checks whether the volume levels at the beat positions of the music are uniform (step S404).
[0055] If the volume levels at the beat positions of the music are uniform (step S404, YES), output control unit 153 checks whether the number of beats per unit time is greater than a third threshold (e.g., first threshold<third threshold<second threshold) (step S405). If the number of beats per unit time is greater than the third threshold (step S405, YES), it decides not to mix a third sound effect into the music (step S402). If the number of beats per unit time is not greater than the third threshold (step S405, NO), it checks whether the volume level of the music remains lower than the first volume level from the start of the music until a first time has elapsed (step S406).
[0056] If the volume level at the beat positions of the music piece is not uniform (step S404, NO), the output control unit 153 checks whether the volume level of the music piece remains lower than the first volume level from the start of the music piece until the first time period has elapsed (step S406).
[0057] If the volume level of the music remains lower than the first volume level from the start of the music until the first time has elapsed (step S406, YES), output control unit 153 decides not to mix the third sound effect into the music (step S402), and if the volume level of the music becomes equal to or higher than the first volume level from the start of the music until the first time has elapsed (step S406, NO), it decides to mix the third sound effect into the music, and outputs the third sound effect via sound effect output unit 130 (step S407).
[0058] The present invention has been described above in terms of preferred embodiments thereof. While the present invention has been described herein with reference to specific examples, various modifications and variations can be made to these examples without departing from the spirit and scope of the present invention as set forth in the claims. [Explanation of symbols]
[0059] 100 Sound Effects Mixer 110 Storage section 120 Music output section 130 Sound effect output section 140 Mixing section 150 control section 151 Music feature acquisition unit 152 Mode selection section 153 Output control section
Claims
[Claim 1] a music feature amount acquisition unit that acquires a feature amount of the music piece, including clarity of beats of the music piece; an output determination unit that determines whether or not to mix sound effects into the music based on the clarity of the beats of the music.
Citation Information
Patent Citations
Karaoke effective sound setting system
JP2016070999A