Information processing device
The information processing apparatus addresses the challenge of real-time sound effect mixing in streaming music by using a music acquisition processing unit and an effect sound output control unit to output sound effects based on music volume, thereby enhancing the listener's experience with a more natural and immersive atmosphere.
Patent Information
- Application Number
- JP2023203554
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-01
- Publication Date
- 2025-06-12
AI Technical Summary
Existing techniques for mixing sound effects into music during streaming playback do not consider real-time processing, resulting in an unnatural atmosphere for listeners.
An information processing apparatus and method that includes a music acquisition processing unit to acquire audio information and an effect sound output control unit to output sound effects when the music volume reaches a predetermined threshold, determined based on the music's volume value, enabling real-time mixing of natural sound effects.
This solution allows for the real-time mixing of natural sound effects into music during streaming playback, enhancing the listener's experience by creating a more immersive atmosphere similar to that of a live venue.
Smart Images

Figure 2025088822000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an information processing apparatus.
Background Art
[0002] Techniques are known for mixing sound effects into music so that the atmosphere of a live venue can be enjoyed. For example, Patent Documents 1-4 disclose techniques for mixing natural sound effects for listeners into music.
Prior Art Documents
Patent Documents
[0003]
Patent Document 1
Patent Document 2
Patent Document 3
Patent Document 4
Summary of the Invention
Problems to be Solved by the Invention
[0004] In recent years, streaming playback has been widely used as a method of playing music. In order to mix sound effects into music being played in a streaming manner, it is necessary to perform the mixing process of sound effects in real time. The techniques disclosed in Patent Documents 1-4 do not consider real-time processing.
[0005] An example of a problem to be solved by the present invention is to mix more natural sound effects into music for listeners.
Means for Solving the Problems
[0006] In order to solve the above problems, the invention according to claim 1 includes a music acquisition processing unit that acquires information related to the audio of a music piece, and an effect sound output control unit that outputs an effect sound when the volume value of the music piece becomes equal to or less than a predetermined threshold value, where the predetermined threshold value is determined based on a value related to the volume value of the music piece.
[0007] The invention according to claim 8 is an information processing method executed by a computer, including a music acquisition processing unit that acquires information related to the audio of a music piece, and an effect sound output control unit that outputs an effect sound when the volume value of the music piece becomes equal to or less than a predetermined threshold value, where the predetermined threshold value is determined based on a value related to the volume value of the music piece.
[0008] The invention according to claim 9 is an information processing program that causes a computer to execute the information processing method according to claim 8.
[0009] The invention according to claim 10 is a computer-readable storage medium that stores the information processing program according to claim 9.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Mode for Carrying Out the Invention
[0011] An information processing apparatus according to an embodiment of the present invention includes a music acquisition processing unit that acquires information related to the sound of a piece of music, and an effect sound output control unit that outputs an effect sound when the volume value of the music becomes equal to or less than a predetermined threshold value. The predetermined threshold value is determined based on a value related to the volume value of the music. The information related to the sound of the music may include the sound signal of the music and / or streaming data. The information processing apparatus may execute real-time processing. The value related to the volume value of the music may be the most recent average value of the volume value of the music. The most recent average value of the volume value of the music may be the most recent moving average value of the volume value of the music. For this reason, in the present embodiment, even when the effect sound mixing device does not receive an input of information related to the end time of the music, it is possible to mix cheers and applause into the end part of the music. Also, in the present embodiment, the timing at which cheers and applause occur is the same as the timing in a live venue. As a result, in the present embodiment, it is possible to make the listener feel more of the atmosphere of being in a live venue.
[0012] If the value obtained by subtracting an offset value from the most recent average value of the volume value of the music is equal to or greater than the minimum threshold value, the predetermined threshold value is the value obtained by subtracting the offset value from the most recent average value of the volume value of the music. If the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is less than the minimum threshold value, it may be the minimum threshold value. By doing so, the timing at which cheers and applause occur is the same as the timing in a live venue, and it becomes possible to make the listener feel more of the atmosphere of being in a live venue.
[0013] The sound effect output control unit may be configured not to determine whether the volume value of the music is equal to or less than a predetermined threshold value for a predetermined period from the start of the music. By doing so, it is possible to reduce the processing amount when outputting sound effects.
[0014] An information processing method according to an embodiment of the present invention is an information processing method executed by a computer, including a music acquisition processing unit that acquires information related to the sound of a music, and a sound effect output control unit that outputs a sound effect when the volume value of the music becomes equal to or less than a predetermined threshold value. The predetermined threshold value is determined based on a value related to the volume value of the music. Therefore, in this embodiment, even when the sound effect mixing device does not receive input of information related to the end time of the music, it is possible to mix cheers and applause into the end part of the music. Further, in this embodiment, the timing at which cheers and applause occur is the same as the timing in a live venue. As a result, in this embodiment, it is possible to make the listener feel more of the atmosphere of being in a live venue.
[0015] An information processing program according to an embodiment of the present invention causes a computer to execute the above information processing method. Therefore, in this embodiment, it is possible to make the listener feel more of the atmosphere of being in a live venue.
[0016] A storage medium according to an embodiment of the present invention stores the above information processing program. Therefore, in this embodiment, the above information processing program can be distributed alone in addition to being incorporated into a device, and it becomes possible to easily perform version updates and the like.
Example
[0017] <Sound effect mixing device 100> FIG. 1 is a diagram showing a sound effect mixing device 100 according to an embodiment of the present invention. The sound effect mixing device 100 includes a music input unit 110, a storage unit 120, a sound effect output unit 130, a mixing unit 140, an external output unit 150, and a control unit 160.
[0018] In the sound effect mixing device 100 according to this embodiment, for example, as shown in FIG. 2, the sound effect output from the sound effect output unit 130 is mixed with the music input from the outside to the music input unit 110 in the mixing unit 140, and the music with the sound effect mixed therein is output to the outside from the external output unit 150. At this time, in the sound effect mixing device 100 according to this embodiment, the storage unit 120 stores the sound source data for the sound effect, and the control unit 160 controls the output of the sound effect from the sound effect output unit 130 by real-time processing.
[0019] The music input unit 110 receives information regarding the music sound from the outside and outputs the music (for example, the music sound signal) to the mixing unit 140. The information regarding the music sound includes, for example, the music sound signal and streaming data of the music. When the information regarding the music sound is the streaming data of the music, the music input unit 110 converts the streaming data of the music into a music sound signal and outputs it to the mixing unit 140. The music input unit 110 includes, for example, a signal input terminal that receives the input of the music sound signal from a playback device that plays the music, and a communication device that receives the input of the streaming data of the music from the outside. Further, the music input unit 110 may include a voice input device (for example, a microphone) that receives the input of the music sound.
[0020] The storage unit 120 is a storage device that stores information such as a hard disk drive, a solid state drive, and a memory. The storage unit 120 stores the sound source data for the sound effect.
[0021] The sound effect output unit 130 outputs a sound effect (for example, a sound effect sound signal) to the mixing unit 140. The sound effect output unit 130, for example, acquires the sound source data for the sound effect stored in the storage unit 120, generates a sound effect sound signal from the acquired sound source data, and outputs the generated sound effect sound signal to the mixing unit 140.
[0022] FIG. 3 is a diagram showing an example of a music input to the music input unit 110 and sound effects output from the sound effects output unit 130. In the example shown in FIG. 3, the first sound effects (such as cheers and applause) are output during the music, at the beginning and end of the music. The second sound effects (environmental sounds) are continuously output from before the start to after the end of the music. The third sound effects (such as clapping) are output during the music in synchronization with the beats and tempo of the music.
[0023] The mixing unit 140 mixes the audio signal of the music output from the music input unit 110 with the audio signal of the sound effects output from the sound effects output unit 130, and outputs an audio signal of the music with the sound effects mixed therein. The mixing unit 140 is, for example, a device that adds a plurality of signals and outputs the added signal, and adds the audio signal of the music output from the music input unit 110 and the audio signal of the sound effects output from the sound effects output unit 130, and outputs the added signal.
[0024] The external output unit 150 outputs the music (that is, the music with the sound effects mixed therein) output by the mixing unit 140 to the outside. The external output unit 150 includes, for example, an audio output device (such as a speaker) that outputs audio, a signal output terminal that outputs an audio signal, and a communication device that transmits information to other devices.
[0025] The control unit 160 controls the output of the sound effects from the sound effects output unit 130. The control unit 160 is an information processing device such as a computer. The control unit 160 of the sound effects output device 100 according to the present embodiment performs real-time processing. That is, the control unit 160 performs real-time processing on information (audio signal or streaming data) regarding the audio of the music input from the outside by the music input unit 110, and controls the output of the sound effects from the sound effects output unit 130 based on the feature amount of the music obtained by the processing.
[0026] FIG. 4 is a diagram showing the control unit 160. The control unit 160 includes a music acquisition processing unit 161, a music feature amount analysis unit 162, and a sound effects output control unit 163.
[0027] The music acquisition processing unit 161 acquires information (audio signal or streaming data) regarding the audio of the music input from the outside by the music input unit 110. At this time, for example, each time information (audio signal or streaming data) regarding the audio of the music is input to the music input unit 110, the music acquisition processing unit 161 sequentially acquires the information (audio signal or streaming data) regarding the audio of the music. That is, the music acquisition processing unit 161 acquires, for example, the audio signal or streaming data of the music input from the outside by the music input unit 110 in real time.
[0028] The music feature amount analysis unit 162 analyzes the feature amount of the music based on the information (music audio signal or streaming data) regarding the audio of the music acquired by the music acquisition processing unit 161. The feature amount of the music includes, for example, the volume value of the music, the position of the beats of the music, the tempo of the music (e.g., BPM (Beats Per Minute)), and the intensity of the beats of the music.
[0029] The music feature amount analysis unit 162 may analyze the feature amount of the music by frame processing, for example. At this time, as shown in FIG. 5, it is preferable that a part of the frames overlap. By doing so, for example, it becomes possible to analyze the feature amount of the music without being affected by minute fluctuations in the audio signal. When the music feature amount analysis unit 162 performs frame processing, the volume value of the music in each frame is, for example, the RMS (Root Mean Square) value in each frame.
[0030] The sound effect output control unit 163 controls the output of the sound effects from the sound effect output unit 130 based on the feature amount of the music analyzed by the music feature amount analysis unit 162.
[0031] The sound effect output control unit 163 controls the volume value of the sound effects output from the sound effect output unit 130 based on, for example, the volume value of the music analyzed by the music feature amount analysis unit 162. By doing so, it becomes possible to prevent the volume of the mixed sound effects from becoming too large or too small compared to the volume of the music, and it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener feel more the atmosphere of being at a live venue.
[0032] Also, as shown in FIG. 4, the control unit 160 may further include a melody analysis unit 164. The melody analysis unit 164 analyzes the melody of the music based on information regarding the sound of the music (the sound signal or streaming data of the music) acquired by the music acquisition processing unit 161 and the feature amounts of the music acquired by the music feature amount analysis unit 162. Then, the sound effect output control unit 163 may control the output of the sound effects from the sound effect output unit 130 based on the melody analyzed by this melody analysis unit 164. The sound effect output control unit 163 may, for example, control the volume value of the sound effects output from the sound effect output unit 130. At this time, the storage unit 120 may store sound source data for sound effects prepared for each volume value and melody of the music, and the sound effect output control unit 163 may determine the sound source data for the sound effects output from the sound effect output unit 130 based on the analyzed volume value and melody of the music.
[0033] Also, as shown in FIG. 4, the control unit 160 may further include a mode selection unit 165. The mode selection unit 165 selects one mode from a plurality of modes related to the melody, genre, scale of the live venue, etc. of the music. At this time, the mode selection unit 165 may, for example, select one mode from the plurality of modes based on the analyzed feature amounts and melody of the music. Then, the sound effect output control unit 163 may control the output of the sound effects output from the sound effect output unit 130 based on the mode input by this mode selection unit 165.
[0034] For example, it is preferable that the plurality of modes include modes for each scale of the live venue. The plurality of modes may include, for example, modes for large-scale venues such as stadiums, outdoor festivals, arenas, etc., modes for medium-scale venues such as halls and medium to large-scale live houses, and modes for small-scale venues such as small live houses and music bars. At this time, for example, the storage unit 120 may store sound effects for large-scale venues, sound effects for medium-scale venues, and sound effects for small-scale venues.
[0035] Also, as shown in FIG. 1, the sound effect mixing device 100 may further include a user input unit 170, and the mode selection unit 165 may select one mode from the plurality of modes based on the input from the user input unit 170. The user input unit 170 is an input device that receives input of information from the user, such as buttons, keyboards, touch panels, cameras, microphones, etc.
[0036] <Processing operation in the control unit 160> FIG. 6 is a diagram showing an example of the processing operation in the control unit 160. The processing operation shown in FIG. 6 is executed, for example, at predetermined time intervals. When the music feature amount analysis unit 162 performs frame processing, the processing operation shown in FIG. 6 is executed, for example, for each frame. The music acquisition processing unit 161 acquires information regarding the sound of the music input from the outside to the music input unit 110 (for example, the sound signal or streaming data of the music) (step S601). The music feature amount analysis unit 162 and the key analysis unit 164 analyze the feature amount and key of the music based on the information regarding the sound of the music acquired by the music acquisition processing unit 161 (step S602). The sound effect output control unit 163 controls the output of the sound effects from the sound effect output unit 130 based on the feature amount and key of the music analyzed by the music feature amount analysis unit 162 and the key analysis unit 164 (step S603).
[0037] <Output of the first sound effect (such as cheers and applause) at the start of the music> At a live venue, when a piece of music is being played, cheers and applause from the audience may occur at the beginning of the performance. When the sound effect mixing device 100 only receives the input of information regarding the music's audio (audio signal or streaming data) from the playback device and does not receive the input of information regarding the start time of the music, the sound effect mixing device 100 cannot mix cheers and applause into the beginning part of the music based on the start time of the music. Also, the volume and manner of such cheers and applause depend on the melody of the music played at the live venue.
[0038] Therefore, in this embodiment, the melody analysis unit 164 analyzes the melody of the music based on the information regarding the music's audio from when the volume value of the music exceeds a predetermined threshold (first threshold) until a predetermined time (first time) has elapsed. The sound effect output control unit 163 outputs a first sound effect based on the melody analyzed by the melody analysis unit 164 when the first time has elapsed since the volume value of the music exceeded the first threshold.
[0039] For this reason, in this embodiment, even when the sound effect mixing device 100 does not receive the input of information regarding the start time of the music, it is possible to mix cheers and applause into the beginning part of the music. Also, in this embodiment, the melody of the music is analyzed using only the information regarding the music's audio until the timing when the first sound effect (such as cheers and applause) is output (the timing when the first time has elapsed since the volume value of the music exceeded the first threshold). Therefore, in this embodiment, it is possible to mix cheers and applause into the music in real-time.
[0040] Here, the first time is related to, for example, the time it takes for the audience at the live venue to sense the start of the performance. Generally, it takes from 1 second to 1.4 seconds from the start of the performance for the audience at the live venue to sense the start of the performance. Therefore, the first time is preferably, for example, 1 second or more and 1.4 seconds or less.
[0041] In addition, the time until the audience at the live venue feels the start of the performance varies based on the size of the live venue. When the live venue is smaller, the distance between the audience and the performer is closer, and the time until the audience feels the start of the performance is shorter. Therefore, the plurality of modes may include modes for each scale of the live venue, and the first time may be determined based on the mode selected by the mode selection unit 165. At this time, it is preferable that the first time is determined to increase as the scale of the live venue increases. For example, when the mode for a small-scale venue is selected, the first time is determined to be 1 second, when the mode for a medium-scale venue is selected, the first time is determined to be 1.2 seconds, and when the mode for a large-scale venue is selected, the first time is determined to be 1.4 seconds.
[0042] After the output of the first sound effect (such as cheers and applause) starts, the melody of the music may change. Therefore, the melody analysis unit 164 may analyze the melody of the music based on information regarding the sound of the music after the first time has elapsed since the volume value of the music exceeded the first threshold. Then, if the melody analyzed by the melody analysis unit 164 changes, the sound effect output control unit 163 may change the sound effect to be output based on the change in the melody.
[0043] For example, when the melody of the music changes from a quiet melody to an intense melody, the sound effect output control unit 163 may output a mixture of the sound effect for the quiet melody and the sound effect for the intense melody.
[0044] Also, when the melody of the music changes from the first melody to the second melody, the sound effect output control unit 163 may change the sound effect to be output from the sound effect for the first melody to the sound effect for the second melody. When changing from the sound effect for the first melody to the sound effect for the second melody, the sound effect output control unit 163 preferably gradually reduces the output level of the sound effect for the first melody to zero (fade out) and gradually increases the output level of the sound effect for the second melody from zero (fade in).
[0045] <Output of the first sound effect (such as cheers and applause) at the end of the music> <In a live venue, when a piece of music is being played, cheers and applause from the audience may occur at the end of the performance. When the sound effect mixing device 100 only receives the input of information (audio signal or streaming data) related to the music sound from the playback device and does not receive the input of information regarding the end time of the music, the sound effect mixing device 100 cannot mix cheers and applause into the end part of the music based on the end time of the music.>
[0046] <Therefore, in this embodiment, when the volume value of the music becomes equal to or less than a predetermined threshold value (the second threshold value), the sound effect output control unit 163 outputs the first sound effect (such as cheers and applause) from the sound effect output unit 130.>
[0047] <At this time, the timing at which the audience feels that the performance has ended varies depending on the volume value of the music. Therefore, in this embodiment, this second threshold value is determined based on a value related to the volume value of the music. The value related to the volume value of the music is, for example, the most recent average value of the volume value of the music. The most recent average value of the volume value of the music may be the most recent moving average value of the volume value of the music, or may be an average value obtained using a time constant circuit. When the music feature amount analysis unit 162 performs frame processing, the most recent moving average value of the volume value of the music is preferably, for example, the average value of the RMS values in the most recent predetermined number (two or more) of frames.>
[0048] When the second threshold is determined based on the most recent average value of the volume value of the music, as shown in FIG. 7, it is preferable that the second threshold be a value obtained by subtracting an offset value from the most recent average value of the volume value of the music. The value obtained by subtracting the offset value from the most recent average value of the volume value of the music may be negative. Therefore, it is preferable to set a minimum value (minimum threshold) for the second threshold. Then, as shown in FIG. 7, if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is equal to or greater than the minimum threshold, the second threshold is set to the value obtained by subtracting the offset value from the most recent average value of the volume value of the music, and if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is less than the minimum threshold, the second threshold is set to the minimum threshold. In FIG. 7, the thin solid line represents the most recent average value of the volume value of the music, the thick solid line represents the second threshold, the one-dot chain line represents the minimum threshold, and the broken line represents the value obtained by subtracting the offset value from the most recent average value of the volume value of the music.
[0049] The determination of the end of the music (whether the volume value of the music has become equal to or less than the second threshold) may be performed at the end portion of the music. Therefore, it is preferable that the sound effect output control unit 163 not determine whether the volume value of the music is equal to or less than the second threshold for a predetermined period (first period) from the start of the music. Here, the start of the music is, for example, the timing at which the input of information regarding the voice of the music is started or the timing at which the volume value of the music exceeds the first threshold. Also, as will be described in detail below, when the information regarding the voice of the music (voice signal or streaming data) includes two or more pieces of music, the start of the next piece of music among two consecutive pieces of music is the cut-off point between the previous piece of music and the next piece of music.
[0050] <Judgment of the end portion of the music> Depending on the music, there may be a silent portion (break). As described above, when a portion that has become equal to or less than the second threshold is determined to be the end portion of the music, the cheer or applause at the end portion of the music will also be mixed into this break portion.
[0051] Therefore, in this embodiment, after the volume value of the music becomes equal to or less than the second threshold value, if the state where the volume value of the music is equal to or less than the second threshold value continues for a predetermined time (second time), when the second time has elapsed since the volume value of the music became equal to or less than the second threshold value, the sound effect output unit 130 outputs a first sound effect (such as cheers or applause).
[0052] At this time, the second time may be a constant. At this time, the second time may be an integer multiple of the beat of the music, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music.
[0053] Also, the stronger the beat of the music, the longer the break period tends to be. Therefore, the second time may be determined based on a value related to the intensity of the beat of the music. At this time, it is preferable that the second time becomes longer as the intensity of the beat of the music increases.
[0054] At this time, as shown in FIG. 8, the second time may change linearly with respect to the intensity of the beat of the music, or may change stepwise with respect to the intensity of the beat of the music as shown in FIG. 9. When the second time changes stepwise with respect to the intensity of the beat of the music (FIG. 9), the second time at each step may be an integer multiple of the beat of the music, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music.
[0055] The intensity of the beat of the music is determined based on, for example, the volume value of the music. When the music feature amount analysis unit 162 performs frame processing and the volume value of the music is the RMS in each frame, the intensity of the beat of the music is preferably, for example, the maximum value among the rising differences in a predetermined number (2 or more) of the most recent frames. Here, the rising difference is the difference between the RMS of the later frame and the RMS of the previous frame when the RMS of the later frame is larger than the RMS of the previous frame in the RMS of two consecutive frames (FIG. 10).
[0056] <Output of the third sound effect (such as clapping)> At a live venue, the audience may keep rhythm by clapping or stomping along with the music. Therefore, in this embodiment, the music feature amount analysis unit 162 analyzes the beat position of the music based on the information regarding the sound of the music acquired by the music acquisition processing unit 161. Then, the sound effect output control unit 163 outputs the third sound effect (such as clapping) at the beat position analyzed by the music feature amount analysis unit 162.
[0057] In a playback device that plays music, operations such as pause and fast forward may be performed. When the sound effect mixing device 100 only receives the input of information regarding the sound of the music (audio signal or streaming data) from the playback device and does not receive the input of information regarding the operations performed on the playback device, after the timing when operations such as pause and fast forward are performed, the clapping mixed in the music may deviate significantly from the beat position of the music.
[0058] Therefore, in this embodiment, the music feature amount analysis unit 162 analyzes not only the beat position of the music but also the confidence level of the beat position based on the information regarding the sound of the music acquired by the music acquisition processing unit 161. The confidence level of the beat position is an index indicating the accuracy of the beat position. Then, the sound effect output control unit 163 gradually reduces the output level of the third sound effect to zero after the beat position where the value of the confidence level becomes less than or equal to a predetermined confidence level value.
[0059] At this time, it is preferable that the time from when the output level of the sound effect starts to gradually decrease until it becomes zero is within a predetermined time (for example, 10 seconds). By doing so, when the clapping mixed in the music deviates from the beat position of the music, it becomes possible to interrupt the mixing of the clapping without the listener feeling unnatural.
[0060] Also, at this time, the sound effect output control unit 163 may output the third sound effect at a constant tempo after the beat position where the confidence value becomes equal to or less than the first confidence value, and gradually reduce the output level of the third sound effect to zero. Here, the constant tempo may be determined based on the interval between the beat position where the confidence value becomes equal to or less than the first confidence value and the beat position immediately before the said beat position. Also, the constant tempo may be determined based on the average value of the intervals between two adjacent beat positions until the confidence value becomes equal to or less than a predetermined confidence value. By doing so, when the tempo mixed into the music deviates from the beat position of the music, it becomes possible to interrupt the mixing of the tempo without the listener feeling unnatural.
[0061] The analysis of the beat position of the music and the confidence level of the said beat position may be performed using, for example, the technology described in Makoto Goto, Yoichi Muraoka, "Real-time Beat Tracking System for Acoustic Signals - Corresponding to Music without Percussion Sounds by Code Change Detection -", Transactions of the Institute of Electronics, Information and Communication Engineers D, 1998 / 2, Vol. J81-D2 pp.227-237.
[0062] <Determination of the break between two pieces of music> The analysis of the feature amount of the music becomes more accurate as the amount of information that can be used is larger. Therefore, when two or more pieces of music are included in the music-related information (voice signal or streaming data) of the music, it is advisable to analyze the feature amount of the next piece of music starting from the break between the previous piece of music and the next piece of music. At this time, if the timing when the volume value of the next piece of music exceeds the first threshold value is regarded as the break between the previous piece of music and the next piece of music, the information of the next piece of music while the volume value is less than the first threshold value will not be used for the analysis of the feature amount of the next piece of music. Also, if the timing when the volume value of the previous piece of music becomes equal to or less than the second threshold value is regarded as the break between the previous piece of music and the next piece of music, the information of the previous piece of music while the volume value is less than the second threshold value may be used for the analysis of the feature amount of the next piece of music.
[0063] Therefore, when information regarding the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, for example, a third threshold smaller than the second threshold is set. When the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues for a predetermined period (second period), it is advisable to consider that second period as a silent section between the previous piece of music and the next piece of music. Therefore, when information regarding the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, when the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues for the second period, the music feature quantity analysis unit 162 and the melody analysis unit 164 determine the timing regarding the period during which the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues (for example, the timing when the volume value of the previous piece of music becomes equal to or lower than the third threshold, or the timing after the second period from the timing when the volume value of the previous piece of music becomes equal to or lower than the third threshold) as the cut-off point between the previous piece of music and the next piece of music (that is, the start of the next piece of music), and based on the information regarding the audio of the music after the cut-off point between the previous piece of music and the next piece of music, it is advisable to analyze the feature quantity and melody of the next piece of music.
[0064] In particular, when information regarding the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, when the first effect sound (such as cheers or applause) is output when the first time has elapsed since the volume value of the next piece of music exceeded the first threshold by the effect sound output control unit 163, the melody analysis unit 164 analyzes the melody of the next piece of music based on the information regarding the audio of the music between the cut-off point between the previous piece of music and the next piece of music and the timing when the first time has elapsed since the volume value of the next piece of music exceeded the first threshold, and the effect sound output control unit 163 may output the first effect sound based on the melody.
[0065] Noise may be included in the silent section between two consecutive pieces of music. Therefore, it is advisable that the effect sound output control unit 163 determines that the volume value of the previous piece of music continues to be equal to or lower than the third threshold even if a momentary increase in volume value (noise) occurs after the volume value of the previous piece of music has become equal to or lower than the third threshold.
[0066] The present invention has been described above according to its preferred embodiments. Although specific examples have been shown here to describe the present invention, various modifications and changes can be made to these specific examples without departing from the spirit and scope of the present invention described in the claims.
Explanation of Reference Numerals
[0067] 100 Sound effect mixing device 110 Music input section 120 Memory section 130 Sound effect output section 140 Mixing section 150 External output section 160 Control section 161 Music acquisition processing section 162 Music feature amount analysis section 163 Sound effect output control section 164 Key analysis section 165 Mode selection section 170 User input section
Claims
1. A music acquisition processing unit that acquires information related to the audio of a music piece, and an effect sound output control unit that outputs an effect sound when the volume value of the music piece becomes equal to or less than a predetermined threshold value, wherein the predetermined threshold value is determined based on a value related to the volume value of the music piece. An information processing apparatus.
2. The information processing apparatus according to claim 1, wherein the value related to the volume value of the music piece is the most recent average value of the volume value of the music piece.
3. The information processing apparatus according to claim 2, wherein the most recent average value of the volume value of the music piece is the most recent moving average value of the volume value of the music piece.
4. The predetermined threshold value is if the value obtained by subtracting an offset value from the most recent average value of the volume value of the music piece is equal to or greater than a minimum threshold value, the value obtained by subtracting the offset value from the most recent average value of the volume value of the music piece, and if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music piece is less than the minimum threshold value, the minimum threshold value. The information processing apparatus according to claim 2.
5. The information processing apparatus according to claim 1, wherein the effect sound output control unit does not determine whether the volume value of the music piece is equal to or less than the predetermined threshold value for a predetermined period from the start of the music piece.
6. The information processing apparatus according to any one of claims 1 to 5, wherein the information related to the audio of the music piece includes the audio signal of the music piece and / or streaming data.
7. The information processing apparatus according to claim 6, wherein the information processing apparatus executes real-time processing.
8. An information processing method executed by a computer, including a music acquisition processing step of acquiring information related to the audio of a music piece, and an effect sound output control step of outputting an effect sound when the volume value of the music piece becomes equal to or less than a predetermined threshold value, wherein the predetermined threshold value is determined based on a value related to the volume value of the music piece. An information processing method.
9. An information processing program for causing a computer to execute the information processing method according to claim 8.
10. A computer-readable storage medium storing the information processing program according to claim 9.
Citation Information
Patent Citations
Sound effect output device
JP2021162707A
Effect sound mixing device
JP2023050570A
Audio output device
WO2023054236A1
Sound effect output device
WO2023054237A1