Information processing device
The music processing apparatus addresses the challenge of real-time sound effect mixing in streamed music by using a music acquisition processing unit and an effect sound output control unit to output sound effects based on specific conditions, resulting in a more immersive live venue experience for listeners.
Patent Information
- Application Number
- JP2023203555
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-12
AI Technical Summary
Existing techniques for mixing sound effects into streamed music do not consider real-time processing, which is necessary to effectively mix natural sound effects into music being streamed.
A music processing apparatus that includes a music acquisition processing unit to acquire sound information from a music piece and an effect sound output control unit that outputs effect sounds based on predetermined conditions, such as elapsed time since the music volume dropped below a threshold and continued for a specified duration.
This solution allows for real-time mixing of natural sound effects into music, preventing cheers and applause from being mixed in silent parts while allowing them to be mixed in the ending part, thereby enhancing the listener's experience of being at a live venue.
Smart Images

Figure 2025088823000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus.
Background Art
[0002] There is known a technique for mixing sound effects into music so that the atmosphere of a live venue can be enjoyed. For example, Patent Documents 1-4 disclose techniques for mixing natural sound effects for listeners into music.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, streaming playback has been widely used as a method of playing music. In order to mix sound effects into music being streamed, it is necessary to perform the mixing process of sound effects in real time. The techniques disclosed in Patent Documents 1-4 do not consider real-time processing.
[0005] An example of the problem to be solved by the present invention is to mix more natural sound effects into music for listeners.
Means for Solving the Problems
[0006] In order to solve the above problems, the invention according to claim 1 includes a music acquisition processing unit that acquires information related to the sound of a music piece, and an effect sound output control unit that controls the output of an effect sound. The effect sound output control unit outputs the effect sound when a predetermined time has elapsed since the volume value of the music piece became equal to or less than a predetermined threshold value, provided that the state in which the volume value of the music piece is equal to or less than the predetermined threshold value has continued for a predetermined time or more after the volume value of the music piece has become equal to or less than the predetermined threshold value.
[0007] The invention according to claim 7 is an information processing method executed by a computer, including a music acquisition processing step of acquiring information related to the sound of a music piece, and an effect sound output control step of controlling the output of an effect sound. In the effect sound output control step, the effect sound is output when a predetermined time has elapsed since the volume value of the music piece became equal to or less than a predetermined threshold value, provided that the state in which the volume value of the music piece is equal to or less than the predetermined threshold value has continued for a predetermined time or more after the volume value of the music piece has become equal to or less than the predetermined threshold value.
[0008] The invention according to claim 8 is an information processing program for causing a computer to execute the information processing method according to claim 7.
[0009] The invention according to claim 9 is a computer-readable storage medium storing the information processing program according to claim 8.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Embodiments for Carrying Out the Invention
[0011] An information processing apparatus according to an embodiment of the present invention includes a music piece acquisition processing unit that acquires information related to the sound of a music piece, and an effect sound output control unit that controls the output of an effect sound. The effect sound output control unit outputs the effect sound when a predetermined time has elapsed since the volume value of the music piece became equal to or less than a predetermined threshold value, provided that the state where the volume value of the music piece is equal to or less than the predetermined threshold value has continued for a predetermined time or more after the volume value of the music piece has become equal to or less than the predetermined threshold value. The information related to the sound of the music piece may include the audio signal of the music piece and / or streaming data. The information processing apparatus may perform real-time processing. For this reason, in the present embodiment, it is possible to prevent cheers and applause from being mixed in silent parts (breaks) that exist depending on the music piece, while allowing cheers and applause to be mixed in the ending part of the music piece. As a result, in the present embodiment, it is possible to make the listener experience the atmosphere of being at a live venue more.
[0012] The above information processing apparatus further includes a music piece feature quantity analysis unit that analyzes the feature quantity of the music piece, and the predetermined time may be determined based on a value related to the intensity of the beats of the music piece. The predetermined time may be determined to become longer as the intensity of the beats of the music piece becomes stronger. Since music pieces with stronger beats tend to have longer break periods, doing so makes it possible to prevent cheers and applause from being mixed in the breaks.
[0013] An information processing method according to an embodiment of the present invention is an information processing method executed by a computer, and includes a music acquisition processing step of acquiring information related to the sound of a music piece, and an effect sound output control step of controlling the output of an effect sound. In the effect sound output control step, if the volume value of the music piece has been below a predetermined threshold and this state where the volume value of the music piece is below the predetermined threshold has continued for a predetermined time or more, then when the predetermined time has elapsed since the volume value of the music piece became below the predetermined threshold, the effect sound is output. Therefore, in this embodiment, it is possible to prevent cheers and applause from being mixed in silent parts (breaks) that exist depending on the music piece, and it is possible to mix cheers and applause in the ending part of the music piece. As a result, in this embodiment, it is possible to make the listener experience more the atmosphere of being at a live venue.
[0014] An information processing program according to an embodiment of the present invention causes a computer to execute the above information processing method. Therefore, in this embodiment, it is possible to make the listener experience more the atmosphere of being at a live venue.
[0015] A storage medium according to an embodiment of the present invention stores the above information processing program. Therefore, in this embodiment, in addition to being incorporated into a device, the above information processing program can be distributed alone, and it becomes possible to easily perform version updates and the like.
Example
[0016] <Effect sound mixing device 100> FIG. 1 is a diagram showing an effect sound mixing device 100 according to an embodiment of the present invention. The effect sound mixing device 100 includes a music input unit 110, a storage unit 120, an effect sound output unit 130, a mixing unit 140, an external output unit 150, and a control unit 160.
[0017] In the sound effect mixing device 100 according to this embodiment, for example, as shown in FIG. 2, the sound effects output from the sound effect output unit 130 are mixed with the music input from the outside to the music input unit 110 in the mixing unit 140, and the music with the sound effects mixed therein is output to the outside from the external output unit 150. At this time, in the sound effect mixing device 100 according to this embodiment, the storage unit 120 stores sound source data for sound effects, and the control unit 160 controls the output of the sound effects from the sound effect output unit 130 by real-time processing.
[0018] The music input unit 110 receives information regarding the music sound from the outside and outputs the music (for example, the music sound signal) to the mixing unit 140. The information regarding the music sound includes, for example, the music sound signal and streaming data. When the information regarding the music sound is the streaming data of the music, the music input unit 110 converts the streaming data of the music into a music sound signal and outputs it to the mixing unit 140. The music input unit 110 includes, for example, a signal input terminal that receives the input of the music sound signal from a playback device that plays the music, and a communication device that receives the input of the streaming data of the music from the outside. Further, the music input unit 110 may include a voice input device (for example, a microphone) that receives the input of the music sound.
[0019] The storage unit 120 is a storage device that stores information such as a hard disk drive, a solid state drive, and a memory. The storage unit 120 stores sound source data for sound effects.
[0020] The sound effect output unit 130 outputs sound effects (for example, sound effect sound signals) to the mixing unit 140. The sound effect output unit 130, for example, acquires the sound source data for sound effects stored in the storage unit 120, generates a sound effect sound signal from the acquired sound source data, and outputs the generated sound effect sound signal to the mixing unit 140.
[0021] FIG. 3 is a diagram showing an example of a music input to the music input unit 110 and sound effects output from the sound effects output unit 130. In the example shown in FIG. 3, the first sound effect (such as cheers or applause) is output during the music, at the beginning and end of the music. The second sound effect (ambient sound) is continuously output from before the start to after the end of the music. The third sound effect (such as hand clapping) is output during the music in synchronization with the beats and tempo of the music.
[0022] The mixing unit 140 mixes the audio signal of the music output from the music input unit 110 with the audio signal of the sound effects output from the sound effects output unit 130, and outputs an audio signal of the music with the sound effects mixed therein. The mixing unit 140 is, for example, a device that adds a plurality of signals and outputs the added signal, and adds the audio signal of the music output from the music input unit 110 and the audio signal of the sound effects output from the sound effects output unit 130, and outputs the added signal.
[0023] The external output unit 150 outputs the music (that is, the music with the sound effects mixed therein) output by the mixing unit 140 to the outside. The external output unit 150 includes, for example, an audio output device (such as a speaker) that outputs audio, a signal output terminal that outputs an audio signal, and a communication device that transmits information to other devices.
[0024] The control unit 160 controls the output of the sound effects from the sound effects output unit 130. The control unit 160 is an information processing device such as a computer. The control unit 160 of the sound effects output device 100 according to the present embodiment performs real-time processing. That is, the control unit 160 performs real-time processing on information (audio signal or streaming data) regarding the audio of the music input from the outside by the music input unit 110, and controls the output of the sound effects from the sound effects output unit 130 based on the feature amount of the music obtained by the processing.
[0025] FIG. 4 is a diagram showing the control unit 160. The control unit 160 includes a music acquisition processing unit 161, a music feature amount analysis unit 162, and a sound effects output control unit 163.
[0026] The music acquisition processing unit 161 acquires information (audio signal or streaming data) regarding the audio of the music input from the outside by the music input unit 110. At this time, for example, every time information (audio signal or streaming data) regarding the audio of the music is input to the music input unit 110, the music acquisition processing unit 161 sequentially acquires the information (audio signal or streaming data) regarding the audio of the music. That is, the music acquisition processing unit 161 acquires, for example, the audio signal or streaming data of the music input from the outside by the music input unit 110 in real time.
[0027] The music feature quantity analysis unit 162 analyzes the feature quantity of the music based on the information (music audio signal or streaming data) regarding the audio of the music acquired by the music acquisition processing unit 161. The feature quantity of the music includes, for example, the volume value of the music, the position of the beats of the music, the tempo of the music (e.g., BPM (Beats Per Minute)), and the intensity of the beats of the music.
[0028] The music feature quantity analysis unit 162 may analyze the feature quantity of the music by frame processing, for example. At this time, as shown in FIG. 5, it is preferable that a part of the frames overlap. By doing so, for example, it becomes possible to analyze the feature quantity of the music without being affected by minute fluctuations of the audio signal. When the music feature quantity analysis unit 162 performs frame processing, the volume value of the music in each frame is, for example, the RMS (Root Mean Square) value in each frame.
[0029] The sound effect output control unit 163 controls the output of the sound effects from the sound effect output unit 130 based on the feature quantity of the music analyzed by the music feature quantity analysis unit 162.
[0030] The sound effect output control unit 163 controls, for example, the volume value of the sound effects output from the sound effect output unit 130 based on the volume value of the music analyzed by the music feature amount analysis unit 162. By doing so, it becomes possible to prevent the volume of the mixed sound effects from becoming too large or too small compared to the volume of the music, and it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener feel more the atmosphere of being at a live venue.
[0031] Also, as shown in FIG. 4, the control unit 160 may further include a melody analysis unit 164. The melody analysis unit 164 analyzes the melody of the music based on information related to the sound of the music (the sound signal or streaming data of the music) acquired by the music acquisition processing unit 161 and the feature amounts of the music acquired by the music feature amount analysis unit 162. Then, the sound effect output control unit 163 may control the output of the sound effects from the sound effect output unit 130 based on the melody analyzed by this melody analysis unit 164. The sound effect output control unit 163 may, for example, control the volume value of the sound effects output from the sound effect output unit 130. At this time, the storage unit 120 may store sound source data for sound effects prepared for each volume value and melody of the music, and the sound effect output control unit 163 may determine the sound source data for the sound effects output from the sound effect output unit 130 based on the analyzed volume value and melody of the music.
[0032] Also, as shown in FIG. 4, the control unit 160 may further include a mode selection unit 165. The mode selection unit 165 selects one mode from a plurality of modes related to the melody, genre, scale of the live venue, etc. of the music. At this time, the mode selection unit 165 may, for example, select one mode from the plurality of modes based on the analyzed feature amounts and melody of the music. Then, the sound effect output control unit 163 may control the output of the sound effects output from the sound effect output unit 130 based on the mode input by this mode selection unit 165.
[0033] For example, it is preferable that the plurality of modes include modes for each scale of the live venue. The plurality of modes may include, for example, modes for large-scale venues such as stadiums, outdoor festivals, arenas, etc., modes for medium-scale venues such as halls and medium- to large-scale live houses, and modes for small-scale venues such as small live houses and music bars. At this time, for example, the storage unit 120 may store sound effects for large-scale venues, sound effects for medium-scale venues, and sound effects for small-scale venues.
[0034] Also, as shown in FIG. 1, the sound effect mixing device 100 may further include a user input unit 170, and the mode selection unit 165 may select one mode from among the plurality of modes based on the input from this user input unit 170. The user input unit 170 is an input device that receives input of information from a user, such as buttons, keyboards, touch panels, cameras, microphones, etc.
[0035] <Processing Operations in the Control Unit 160> FIG. 6 is a diagram showing an example of the processing operations in the control unit 160. The processing operations shown in FIG. 6 are executed, for example, at predetermined time intervals. When the music feature amount analysis unit 162 performs frame processing, the processing operations shown in FIG. 6 are executed, for example, for each frame. The music acquisition processing unit 161 acquires information regarding the sound of the music input from the outside to the music input unit 110 (for example, the sound signal or streaming data of the music) (step S601). The music feature amount analysis unit 162 and the key analysis unit 164 analyze the feature amount and key of the music based on the information regarding the sound of the music acquired by the music acquisition processing unit 161 (step S602). The sound effect output control unit 163 controls the output of the sound effects from the sound effect output unit 130 based on the feature amount and key of the music analyzed by the music feature amount analysis unit 162 and the key analysis unit 164 (step S603).
[0036] <Output of the First Sound Effects (such as cheers and applause) at the Beginning of the Music> At a live venue, when a piece of music is being played, there may be cheers and applause from the audience at the beginning of the performance. When the sound effect mixing device 100 only receives the input of information regarding the music's audio (audio signal or streaming data) from the playback device and does not receive the input of information regarding the start time of the music, the sound effect mixing device 100 cannot mix cheers and applause into the beginning part of the music based on the start time of the music. Also, the volume and manner of such cheers and applause depend on the melody of the music being played at the live venue.
[0037] Therefore, in this embodiment, the melody analysis unit 164 analyzes the melody of the music based on the information regarding the music's audio from when the volume value of the music exceeds a predetermined threshold (first threshold) until a predetermined time (first time) has elapsed, and the sound effect output control unit 163 outputs a first sound effect based on the melody analyzed by the melody analysis unit 164 when the first time has elapsed since the volume value of the music exceeded the first threshold.
[0038] For this reason, in this embodiment, even when the sound effect mixing device 100 does not receive the input of information regarding the start time of the music, it is possible to mix cheers and applause into the beginning part of the music. Also, in this embodiment, the melody of the music is analyzed using only the information regarding the music's audio until the timing at which the first sound effect (such as cheers and applause) is output (the timing when the first time has elapsed since the volume value of the music exceeded the first threshold). Therefore, in this embodiment, it is possible to mix cheers and applause into the music in real-time.
[0039] Here, the first time is related to, for example, the time it takes for the audience at the live venue to sense the start of the performance. Generally, it takes from 1 second to 1.4 seconds from the start of the performance for the audience at the live venue to sense the start of the performance. Therefore, the first time is preferably, for example, 1 second or more and 1.4 seconds or less.
[0040] Also, the time it takes for the audience at the live venue to sense the start of the performance varies based on the size of the live venue. When the live venue is smaller, the distance between the audience and the performer is closer, and the time it takes for the audience to sense the start of the performance is shorter. Therefore, the plurality of modes may include modes for each scale of the live venue, and the first time may be determined based on the mode selected by the mode selection unit 165. At this time, it is preferable that the first time is determined to increase as the scale of the live venue increases. For example, when the mode for a small-scale venue is selected, the first time is determined to be 1 second; when the mode for a medium-scale venue is selected, the first time is determined to be 1.2 seconds; and when the mode for a large-scale venue is selected, the first time is determined to be 1.4 seconds.
[0041] After the output of the first sound effect (such as cheers or applause) starts, the melody of the music may change. Therefore, the melody analysis unit 164 may analyze the melody of the music based on the information regarding the sound of the music after the first time has elapsed since the volume value of the music exceeded the first threshold. Then, if the melody analyzed by the melody analysis unit 164 changes, the sound effect output control unit 163 may change the sound effect to be output based on the change in the melody.
[0042] For example, when the melody of the music changes from a quiet melody to an intense melody, the sound effect output control unit 163 may output a mixture of the sound effect for the quiet melody and the sound effect for the intense melody.
[0043] Also, when the melody of the music changes from the first melody to the second melody, the sound effect output control unit 163 may change the sound effect to be output from the sound effect for the first melody to the sound effect for the second melody. When changing from the sound effect for the first melody to the sound effect for the second melody, the sound effect output control unit 163 may gradually decrease the output level of the sound effect for the first melody to zero (fade out) and gradually increase the output level of the sound effect for the second melody from zero (fade in).
[0044] <Output of the first sound effect (such as cheers or applause) at the end of the music> At a live venue, when a music piece is being played, cheers or applause from the audience may occur at the end of the performance. When the sound effect mixing device 100 only receives the input of information regarding the music piece's audio (audio signal or streaming data) from the playback device and does not receive the input of information regarding the end time of the music piece, the sound effect mixing device 100 cannot mix cheers or applause into the end part of the music piece based on the end time of the music piece.
[0045] Therefore, in this embodiment, when the volume value of the music piece becomes equal to or less than a predetermined threshold value (second threshold value), the sound effect output control unit 163 outputs the first sound effect (such as cheers or applause) from the sound effect output unit 130.
[0046] At this time, the timing that the audience feels the performance has ended varies depending on the volume value of the music piece. Therefore, in this embodiment, this second threshold value is determined based on a value related to the volume value of the music piece. The value related to the volume value of the music piece is, for example, the most recent average value of the volume value of the music piece. The most recent average value of the volume value of the music piece may be the most recent moving average value of the volume value of the music piece, or may be an average value obtained using a time constant circuit. When the music piece feature amount analysis unit 162 performs frame processing, the most recent moving average value of the volume value of the music piece may be, for example, the average value of the RMS values in the most recent predetermined number (2 or more) of frames.
[0047] When the second threshold is determined based on the most recent average value of the volume value of the music, as shown in FIG. 7, it is preferable that the second threshold be a value obtained by subtracting an offset value from the most recent average value of the volume value of the music. The value obtained by subtracting the offset value from the most recent average value of the volume value of the music may be negative. Therefore, it is preferable to set a minimum value (minimum threshold) for the second threshold. Then, as shown in FIG. 7, if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is greater than or equal to the minimum threshold, the second threshold is set to the value obtained by subtracting the offset value from the most recent average value of the volume value of the music, and if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is less than the minimum threshold, the second threshold is set to the minimum threshold. In FIG. 7, the thin solid line represents the most recent average value of the volume value of the music, the thick solid line represents the second threshold, the one-dot chain line represents the minimum threshold, and the broken line represents the value obtained by subtracting the offset value from the most recent average value of the volume value of the music.
[0048] The determination of the end of the music (whether the volume value of the music has become less than or equal to the second threshold) may be made at the end portion of the music. Therefore, it is preferable that the sound effect output control unit 163 does not determine whether the volume value of the music is less than or equal to the second threshold for a predetermined period (first period) from the start of the music. Here, the start of the music is, for example, the timing when the input of information regarding the voice of the music is started or the timing when the volume value of the music exceeds the first threshold. Also, as will be described in detail below, when the information regarding the voice of the music (voice signal or streaming data) includes two or more pieces of music, the start of the next piece of music among two consecutive pieces of music is the cut-off point between the previous piece of music and the next piece of music.
[0049] <Determination of the end portion of the music> Depending on the music, there may be a silent portion (break). As described above, when a portion where the value has become less than or equal to the second threshold is determined to be the end portion of the music, the cheer or applause at the end portion of the music will also be mixed into this break portion.
[0050] Therefore, in this embodiment, after the volume value of the music becomes equal to or less than the second threshold value and the state where the volume value of the music is equal to or less than the second threshold value continues for a predetermined time (second time), when the second time has elapsed since the volume value of the music became equal to or less than the second threshold value, the sound effect output unit 130 outputs a first sound effect (such as cheers or applause).
[0051] At this time, the second time may be a constant. At this time, the second time may be an integer multiple of the beat of the music, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music.
[0052] Also, the stronger the beat of the music, the longer the break period tends to be. Therefore, the second time may be determined based on a value related to the intensity of the beat of the music. At this time, it is preferable that the second time becomes longer as the intensity of the beat of the music increases.
[0053] At this time, as shown in FIG. 8, the second time may change linearly with respect to the intensity of the beat of the music, or may change stepwise with respect to the intensity of the beat of the music as shown in FIG. 9. When the second time changes stepwise with respect to the intensity of the beat of the music (FIG. 9), the second time at each step may be an integer multiple of the beat of the music, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music.
[0054] The intensity of the beat of the music is determined based on, for example, the volume value of the music. When the music feature amount analysis unit 162 performs frame processing and the volume value of the music is the RMS in each frame, the intensity of the beat of the music is preferably, for example, the maximum value among the rising differences in a predetermined number (two or more) of the most recent frames. Here, the rising difference is the difference between the RMS of the later frame and the RMS of the previous frame when the RMS of the later frame is larger than the RMS of the previous frame in the RMS of two consecutive frames (FIG. 10).
[0055] <Output of the third sound effect (such as a clap)> At a live venue, the audience may keep rhythm by clapping or tapping their feet in time with the music. Therefore, in this embodiment, the music feature quantity analysis unit 162 analyzes the beat positions of the music based on the information regarding the sound of the music acquired by the music acquisition processing unit 161. Then, the sound effect output control unit 163 outputs the third sound effect (such as a clap) at the beat positions analyzed by the music feature quantity analysis unit 162.
[0056] In a playback device that plays back music, operations such as pausing and fast-forwarding may be performed. When the sound effect mixing device 100 only receives the input of information regarding the sound of the music (audio signal or streaming data) from the playback device and does not receive the input of information regarding the operations performed on the playback device, after the timing when operations such as pausing and fast-forwarding are performed, the claps mixed into the music may deviate significantly from the beat positions of the music.
[0057] Therefore, in this embodiment, the music feature quantity analysis unit 162 analyzes not only the beat positions of the music but also the confidence levels of the beat positions based on the information regarding the sound of the music acquired by the music acquisition processing unit 161. The confidence level of a beat position is an index indicating the accuracy of the beat position. Then, the sound effect output control unit 163 gradually reduces the output level of the third sound effect to zero after the beat positions where the confidence level value has fallen below a predetermined confidence level value.
[0058] At this time, it is preferable that the time from when the output level of the sound effect starts to gradually decrease until it becomes zero is within a predetermined time (for example, 10 seconds). By doing so, when the claps mixed into the music deviate from the beat positions of the music, it becomes possible to interrupt the mixing of the claps without the listener feeling unnatural.
[0059] Also, at this time, the sound effect output control unit 163 may output the third sound effect at a constant tempo while gradually decreasing the output level of the third sound effect to zero after the beat position where the confidence value becomes equal to or less than the first confidence value. Here, the constant tempo may be determined based on the interval between the beat position where the confidence value becomes equal to or less than the first confidence value and the immediately preceding beat position. Also, the constant tempo may be determined based on the average value of the intervals between two adjacent beat positions until the confidence value becomes equal to or less than a predetermined confidence value. By doing so, when the beat mixed into the music deviates from the beat position of the music, it becomes possible to interrupt the mixing of the beat without the listener feeling unnatural.
[0060] The analysis of the beat position of the music and the confidence level of the beat position may be performed using, for example, the technique described in Makoto Goto, Yoichi Murakami, "Real-Time Beat Tracking System for Acoustic Signals - Corresponding to Music without Percussion Sounds by Code Change Detection -", Transactions of the Institute of Electronics, Information and Communication Engineers D, 1998 / 2, Vol. J81-D2 pp.227-237.
[0061] <Determination of the break between two pieces of music> The analysis of the feature amount of the music becomes more accurate as the amount of information that can be used is larger. For this reason, when two or more pieces of music are included in the music-related information (audio signal or streaming data) of the music, it is advisable to analyze the feature amount of the next piece of music from the break between the previous piece of music and the next piece of music. At this time, if the timing when the volume value of the next piece of music exceeds the first threshold value is regarded as the break between the previous piece of music and the next piece of music, the information of the next piece of music while the volume value is less than the first threshold value will not be used for the analysis of the feature amount of the next piece of music. Also, if the timing when the volume value of the previous piece of music becomes equal to or less than the second threshold value is regarded as the break between the previous piece of music and the next piece of music, the information of the previous piece of music while the volume value is less than the second threshold value may be used for the analysis of the feature amount of the next piece of music.
[0062] Therefore, when information regarding the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, for example, a third threshold smaller than the second threshold is set. When the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues for a predetermined period (second period), it is advisable to consider that second period as a silent portion between the previous piece of music and the next piece of music. Therefore, when information regarding the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, when the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues for the second period, the music feature quantity analysis unit 162 and the melody analysis unit 164 determine the timing regarding the period during which the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues (for example, the timing when the volume value of the previous piece of music becomes equal to or lower than the third threshold, or the timing after the second period from the timing when the volume value of the previous piece of music becomes equal to or lower than the third threshold) as the cut-off point between the previous piece of music and the next piece of music (that is, the start of the next piece of music), and based on the information regarding the audio of the music after the cut-off point between the previous piece of music and the next piece of music, it is advisable to analyze the feature quantity and melody of the next piece of music.
[0063] In particular, when information regarding the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, when the first effect sound (such as cheers or applause) is output when the first time has elapsed since the volume value of the next piece of music exceeded the first threshold, the melody analysis unit 164 analyzes the melody of the next piece of music based on the information regarding the audio of the music between the cut-off point between the previous piece of music and the next piece of music and the timing when the first time has elapsed since the volume value of the next piece of music exceeded the first threshold, and the effect sound output control unit 163 may output the first effect sound based on the melody.
[0064] Noise may be included in the silent portion between two consecutive pieces of music. Therefore, it is advisable that the effect sound output control unit 163 determines that the volume value of the previous piece of music continues to be equal to or lower than the third threshold even if a momentary increase in volume value (noise) occurs after the volume value of the previous piece of music becomes equal to or lower than the third threshold.
[0065] The present invention has been described above according to its preferred embodiments. Although specific examples have been shown to describe the present invention herein, various modifications and changes can be made to these specific examples without departing from the spirit and scope of the present invention described in the claims.
Explanation of Reference Numerals
[0066] 100 Sound effect mixing device 110 Music input unit 120 Storage unit 130 Sound effect output unit 140 Mixing unit 150 External output unit 160 Control unit 161 Music acquisition processing unit 162 Music feature quantity analysis unit 163 Sound effect output control unit 164 Melody analysis unit 165 Mode selection unit 170 User input unit
Claims
1. A music acquisition processing unit that acquires information related to the audio of a music piece, and a sound effect output control unit that controls the output of sound effects, wherein the sound effect output control unit outputs the sound effect when a state in which the volume value of the music piece has been equal to or less than a predetermined threshold value has continued for a predetermined time or more after the volume value of the music piece has become equal to or less than the predetermined threshold value, that is, when the predetermined time has elapsed since the volume value of the music piece became equal to or less than the predetermined threshold value. An information processing apparatus.
2. further comprising a music feature quantity analysis unit that analyzes a feature quantity of the music piece, wherein the predetermined time is determined based on a value related to the intensity of the beats of the music piece. The information processing apparatus according to Claim 1.
3. wherein the predetermined time is determined to become longer as the intensity of the beats of the music piece becomes stronger. The information processing apparatus according to Claim 2.
4. wherein the sound effect output control unit does not determine whether the volume value of the music piece is equal to or less than the predetermined threshold value for a predetermined period from the start of the music piece. The information processing apparatus according to Claim 1.
5. wherein the information related to the audio of the music piece includes the audio signal of the music piece and / or streaming data. The information processing apparatus according to Claims 1 to 4.
6. wherein the information processing apparatus executes real-time processing. The information processing apparatus according to Claim 5.
7. An information processing method executed by a computer, comprising a music acquisition processing step of acquiring information related to the audio of a music piece, and a sound effect output control step of controlling the output of sound effects, wherein in the sound effect output control step, when a state in which the volume value of the music piece has been equal to or less than a predetermined threshold value has continued for a predetermined time or more after the volume value of the music piece has become equal to or less than the predetermined threshold value, the sound effect is output when the predetermined time has elapsed since the volume value of the music piece became equal to or less than the predetermined threshold value. An information processing method.
8. An information processing program for causing a computer to execute the information processing method according to Claim 7.
9. A computer-readable storage medium storing the information processing program according to Claim 8.
Citation Information
Patent Citations
Sound effect output device
JP2021162707A
Effect sound mixing device
JP2023050570A
Audio output device
WO2023054236A1
Sound effect output device
WO2023054237A1