Information processing device

The information processing apparatus addresses the challenge of real-time sound effect mixing in streamed music by analyzing melodies and outputting effect sounds in sync with the music, thereby enhancing the live venue atmosphere for listeners.

JP2025088821APending Publication Date: 2025-06-12PIONEER IP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023203553
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-12

AI Technical Summary

Technical Problem

Existing techniques for mixing sound effects into streamed music do not consider real-time processing, which is necessary to create a natural atmosphere similar to a live venue.

Method used

An information processing apparatus that includes a music acquisition processing unit, a melody analysis unit, and an effect sound output control unit. This apparatus analyzes the melody of a music piece in real-time based on the sound information when the volume exceeds a threshold, and outputs effect sounds accordingly.

Benefits of technology

Enables real-time mixing of natural sound effects into streamed music, enhancing the listener's experience of being at a live venue by accurately timing cheers and applause with the music.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025088821000001_ABST
    Figure 2025088821000001_ABST
Patent Text Reader

Abstract

To provide an information processing device and information processing method for mixing a sound effect more natural for a listener with a music piece.SOLUTION: An information processing device performs: analyzing a music tone of a music piece based on information related to the sound of the music piece from a time when a volume value of the music piece exceeds a predetermined threshold to a time when a predetermined time elapses; and outputting a sound effect based on the analyzed music tone when the predetermined time elapses after the volume value of the music piece exceeds the predetermined threshold.SELECTED DRAWING: Figure 6
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an information processing apparatus.

Background Art

[0002] Techniques are known for mixing sound effects into music so that the atmosphere of a live venue can be experienced. For example, Patent Documents 1-4 disclose techniques for mixing natural sound effects for listeners into music.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Patent Document 2

Patent Document 3

Patent Document 4

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, streaming playback has been widely used as a method of playing music. In order to mix sound effects into music being streamed, it is necessary to perform the mixing process of the sound effects in real time. The techniques disclosed in Patent Documents 1-4 do not consider real-time processing.

[0005] An example of the problem to be solved by the present invention is to mix more natural sound effects for listeners into music.

Means for Solving the Problems

[0006] In order to solve the above problems, the invention according to claim 1 includes a music acquisition processing unit that acquires information related to the sound of a music piece, a melody analysis unit that analyzes the melody of the music piece, and an effect sound output control unit that controls the output of the effect sound. The melody analysis unit analyzes the melody of the music piece based on the information related to the sound of the music piece from when the volume value of the music piece exceeds a predetermined threshold until a predetermined time has elapsed. The effect sound output control unit outputs an effect sound based on the melody analyzed by the melody analysis unit when the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold.

[0007] The invention according to claim 9 is an information processing method executed by a computer, including a music acquisition processing step of acquiring information related to the sound of a music piece, a melody analysis step of analyzing the melody of the music piece, and an effect sound output control step of controlling the output of the effect sound. In the melody analysis step, the melody of the music piece is analyzed based on the information related to the sound of the music piece from when the volume value of the music piece exceeds a predetermined threshold until a predetermined time has elapsed. In the effect sound output control step, an effect sound based on the melody analyzed in the melody analysis step is output when the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold.

[0008] The invention according to claim 10 is an information processing program for causing a computer to execute the information processing method according to claim 9.

[0009] The invention according to claim 11 is a computer-readable storage medium storing the information processing program according to claim 10.

Brief Description of the Drawings

[0010]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Modes for Carrying Out the Invention

[0011] An information processing apparatus according to an embodiment of the present invention includes a music acquisition processing unit that acquires information related to the sound of a music piece, a melody analysis unit that analyzes the melody of the music piece, and an effect sound output control unit that controls the output of effect sounds. The melody analysis unit analyzes the melody of the music piece based on the information related to the sound of the music piece from when the volume value of the music piece exceeds a predetermined threshold until a predetermined time has elapsed. The effect sound output control unit outputs an effect sound based on the melody analyzed by the melody analysis unit when the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold. The information related to the sound of the music piece may include the sound signal of the music piece and / or streaming data. The information processing apparatus may be configured to perform real-time processing. For this reason, in this embodiment, even when the effect sound mixing device does not receive an input of information related to the start time of the music piece, it is possible to mix cheers and applause into the beginning part of the music piece. Also, in this embodiment, the melody of the music piece is analyzed using only the information related to the sound of the music piece until the timing at which the first effect sound (such as cheers and applause) is output (the timing when the first time has elapsed since the volume value of the music piece exceeded the first threshold). For this reason, in this embodiment, it is possible to mix cheers and applause into the music piece by real-time processing. As a result, in this embodiment, it is possible to make the listener feel more the atmosphere of being at a live venue.

[0012] The above information processing apparatus may further include a mode selection unit that selects one mode from a plurality of modes, and the predetermined time may be determined based on the mode selected by the mode selection unit. The plurality of modes include modes for each scale of the live venue, and the predetermined time may be determined to increase as the scale of the live venue increases. By doing so, the timing at which cheers and applause occur becomes the same as that of the live venue, and it becomes possible to make the listener feel more the atmosphere of being at a live venue.

[0013] The predetermined time may be set to be 1 second or more and 1.4 seconds or less. By doing so, the timing at which cheers and applause occur will be the same as that in a live venue, enabling the listeners to better experience the atmosphere of being in a live venue.

[0014] Based on the information about the sound of the music after the predetermined time has elapsed since the volume value of the music exceeded the predetermined threshold, the melody analysis unit analyzes the melody of the music. If the melody analyzed by the melody analysis unit changes, the sound effect output control unit may change the sound effect to be output based on the change in the melody. By doing so, cheers and applause will change according to the change in the melody of the music, enabling the listeners to better experience the atmosphere of being in a live venue.

[0015] The information about the sound of the music includes two or more pieces of music. When the state where the volume value of the previous piece of music is equal to or less than a third threshold continues for a predetermined period, the melody analysis unit determines the timing of the period during which the volume value of the previous piece of music is equal to or less than the third threshold as the cut-off point between the previous piece of music and the next piece of music. Based on the information about the sound of the music between the cut-off point between the previous piece of music and the next piece of music and the timing when the predetermined time has elapsed since the volume value of the next piece of music exceeded the predetermined threshold, the melody analysis unit analyzes the melody of the next piece of music. By doing so, even when the information about the sound of the music includes two or more pieces of music, it becomes possible to mix cheers and applause into the music in real-time processing, enabling the listeners to better experience the atmosphere of being in a live venue.

[0016] An information processing method according to an embodiment of the present invention is an information processing method executed by a computer, and includes a music acquisition processing step of acquiring information related to the sound of a music piece, a melody analysis step of analyzing the melody of the music piece, and an effect sound output control step of controlling the output of an effect sound. In the melody analysis step, the melody of the music piece is analyzed based on the information related to the sound of the music piece from when the volume value of the music piece exceeds a predetermined threshold until a predetermined time has elapsed. In the effect sound output control step, when the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold, an effect sound based on the melody analyzed in the melody analysis step is output. Therefore, in the present embodiment, even when the effect sound mixing device does not receive input of information regarding the start time of the music piece, it is possible to mix cheers and applause into the beginning part of the music piece. Further, in the present embodiment, the melody of the music piece is analyzed using only the information related to the sound of the music piece until the timing when the first effect sound (such as cheers and applause) is output (the timing when the first time has elapsed since the volume value of the music piece exceeded the first threshold). Therefore, in the present embodiment, it is possible to mix cheers and applause into the music piece by real-time processing. As a result, in the present embodiment, it is possible to make the listener feel more the atmosphere of being in a live venue.

[0017] An information processing program according to an embodiment of the present invention causes a computer to execute the above information processing method. Therefore, in the present embodiment, it is possible to make the listener feel more the atmosphere of being in a live venue.

[0018] A storage medium according to an embodiment of the present invention stores the above information processing program. Therefore, in the present embodiment, the above information processing program can be distributed alone in addition to being incorporated into a device, and it becomes possible to easily perform version updates and the like.

Example

[0019] <Effect sound mixing device 100> FIG. 1 is a diagram showing an effect sound mixing device 100 according to an embodiment of the present invention. The effect sound mixing device 100 includes a music input unit 110, a storage unit 120, an effect sound output unit 130, a mixing unit 140, an external output unit 150, and a control unit 160.

[0020] In the effect sound mixing device 100 according to this embodiment, for example, as shown in FIG. 2, the music input from the outside to the music input unit 110 and the effect sound output from the effect sound output unit 130 are mixed in the mixing unit 140, and the music with the effect sound mixed therein is output to the outside from the external output unit 150. At this time, in the effect sound mixing device 100 according to this embodiment, the storage unit 120 stores sound source data for effect sounds, and the control unit 160 controls the output of the effect sound from the effect sound output unit 130 by real-time processing.

[0021] The music input unit 110 receives information related to the music sound from the outside and outputs the music (for example, the music sound signal) to the mixing unit 140. The information related to the music sound includes, for example, the music sound signal and streaming data. When the information related to the music sound is the streaming data of the music, the music input unit 110 converts the streaming data of the music into a music sound signal and outputs it to the mixing unit 140. The music input unit 110 includes, for example, a signal input terminal that receives the input of the music sound signal from a playback device that plays the music, and a communication device that receives the input of the streaming data of the music from the outside. Further, the music input unit 110 may include a voice input device (for example, a microphone) that receives the input of the music sound.

[0022] The storage unit 120 is a storage device that stores information such as a hard disk drive, a solid state drive, and a memory. The storage unit 120 stores sound source data for effect sounds.

[0023] The sound effect output unit 130 outputs a sound effect (e.g., an audio signal of a sound effect) to the mixing unit 140. The sound effect output unit 130, for example, acquires sound source data for sound effects stored in the storage unit 120, generates an audio signal of a sound effect from the acquired sound source data, and outputs the generated audio signal of the sound effect to the mixing unit 140.

[0024] FIG. 3 is a diagram showing an example of a music input to the music input unit 110 and a sound effect output from the sound effect output unit 130. In the example shown in FIG. 3, the first sound effect (such as cheers and applause) is output during the music, at the beginning and end of the music. The second sound effect (ambient sound) is continuously output from before the start to after the end of the music. The third sound effect (such as beats) is output during the music in synchronization with the beats and tempo of the music.

[0025] The mixing unit 140 mixes the audio signal of the music output from the music input unit 110 with the audio signal of the sound effect output from the sound effect output unit 130, and outputs an audio signal of the music with the sound effect mixed therein. The mixing unit 140 is, for example, a device that adds a plurality of signals and outputs the added signal, adds the audio signal of the music output from the music input unit 110 and the audio signal of the sound effect output from the sound effect output unit 130, and outputs the added signal.

[0026] The external output unit 150 outputs the music output by the mixing unit 140 (i.e., the music with the sound effect mixed therein) to the outside. The external output unit 150 includes, for example, an audio output device (e.g., a speaker) that outputs sound, a signal output terminal that outputs an audio signal, and a communication device that transmits information to other devices.

[0027] The control unit 160 controls the output of the sound effects from the sound effect output unit 130. The control unit 160 is an information processing device such as a computer. The control unit 160 of the sound effect output device 100 according to this embodiment performs real-time processing. That is, the control unit 160 performs real-time processing on information (voice signals and streaming data) related to the sound of the music input from the outside by the music input unit 110, and based on the feature amounts of the music obtained by the processing, controls the output of the sound effects from the sound effect output unit 130.

[0028] Figure 4 is a diagram showing the control unit 160. The control unit 160 includes a music acquisition processing unit 161, a music feature amount analysis unit 162, and a sound effect output control unit 163.

[0029] The music acquisition processing unit 161 acquires information (voice signals and streaming data) related to the sound of the music input from the outside by the music input unit 110. At this time, the music acquisition processing unit 161 sequentially acquires the information (voice signals and streaming data) related to the sound of the music, for example, every time the information (voice signals and streaming data) related to the sound of the music is input to the music input unit 110. That is, the music acquisition processing unit 161 acquires the voice signals and streaming data of the music input from the outside by the music input unit 110 in real time, for example.

[0030] The music feature amount analysis unit 162 analyzes the feature amounts of the music based on the information (the voice signal or streaming data of the music) related to the sound of the music acquired by the music acquisition processing unit 161. The feature amounts of the music include, for example, the volume value of the music, the position of the beats of the music, the tempo of the music (for example, BPM (Beats Per Minute)), and the intensity of the beats of the music.

[0031] The music feature quantity analysis unit 162 may analyze the feature quantity of music by, for example, frame processing. At this time, as shown in FIG. 5, it is preferable that a part of the frames overlap. By doing so, for example, it becomes possible to analyze the feature quantity of music without being affected by minute fluctuations in the audio signal. When the music feature quantity analysis unit 162 performs frame processing, the volume value of the music in each frame is, for example, the RMS (Root Mean Square) value in each frame.

[0032] The sound effect output control unit 163 controls the output of the sound effects from the sound effect output unit 130 based on the feature quantity of the music analyzed by the music feature quantity analysis unit 162.

[0033] The sound effect output control unit 163 controls, for example, the volume value of the sound effects output from the sound effect output unit 130 based on the volume value of the music analyzed by the music feature quantity analysis unit 162. By doing so, it becomes possible to prevent the volume of the mixed sound effects from becoming too large or too small compared to the volume of the music, and it becomes possible to add more natural sound effects to the music for the listener, and it becomes possible to make the listener feel more the atmosphere of being in a live venue.

[0034] Further, as shown in FIG. 4, the control unit 160 may further include a melody analysis unit 164. The melody analysis unit 164 analyzes the melody of a piece of music based on information regarding the sound of the music (audio signal or streaming data of the music) acquired by the music acquisition processing unit 161 and the feature amounts of the music acquired by the music feature amount analysis unit 162. Then, the sound effect output control unit 163 may control the output of the sound effects from the sound effect output unit 130 based on the melody analyzed by the melody analysis unit 164. For example, the sound effect output control unit 163 may control the volume value of the sound effects output from the sound effect output unit 130. At this time, the storage unit 120 may store sound source data for sound effects prepared for each volume value and melody of the music, and the sound effect output control unit 163 may determine the sound source data for the sound effects output from the sound effect output unit 130 based on the volume value and melody of the analyzed music.

[0035] Also, as shown in FIG. 4, the control unit 160 may further include a mode selection unit 165. The mode selection unit 165 selects one mode from a plurality of modes related to the melody, genre, scale of the live venue, etc. of the music. At this time, for example, the mode selection unit 165 may select one mode from the plurality of modes based on the feature amounts and melody of the analyzed music. Then, the sound effect output control unit 163 may control the output of the sound effects output from the sound effect output unit 130 based on the mode input by the mode selection unit 165.

[0036] For example, it is preferable that the plurality of modes include modes for each scale of the live venue. The plurality of modes may include, for example, modes for large-scale venues such as stadiums, outdoor festivals, arenas, etc., modes for medium-scale venues such as halls and medium to large-scale live houses, and modes for small-scale venues such as small-scale live houses and music bars. At this time, for example, the storage unit 120 may store sound effects for large-scale venues, sound effects for medium-scale venues, and sound effects for small-scale venues.

[0037] Also, as shown in FIG. 1, the sound effect mixing device 100 may further include a user input unit 170, and the mode selection unit 165 may select one mode from among a plurality of modes based on an input from this user input unit 170. The user input unit 170 is an input device that receives input of information from a user, such as a button, a keyboard, a touch panel, a camera, a microphone, etc.

[0038] <Processing Operations in the Control Unit 160> FIG. 6 is a diagram showing an example of processing operations in the control unit 160. The processing operations shown in FIG. 6 are executed, for example, at predetermined time intervals. When the music feature amount analysis unit 162 performs frame processing, the processing operations shown in FIG. 6 are executed, for example, for each frame. The music acquisition processing unit 161 acquires information regarding the sound of the music input from the outside to the music input unit 110 (for example, the audio signal or streaming data of the music) (step S601). The music feature amount analysis unit 162 and the key tone analysis unit 164 analyze the feature amount and key tone of the music based on the information regarding the sound of the music acquired by the music acquisition processing unit 161 (step S602). The sound effect output control unit 163 controls the output of the sound effects from the sound effect output unit 130 based on the feature amount and key tone of the music analyzed by the music feature amount analysis unit 162 and the key tone analysis unit 164 (step S603).

[0039] <Output of the First Sound Effect (such as cheers or applause) at the Beginning of the Music> At a live venue, when a music is being played, cheers or applause from the audience may occur at the beginning part of the performance. When the sound effect mixing device 100 only receives input of information regarding the sound of the music (audio signal or streaming data) from the playback device and does not receive input of information regarding the start time of the music, the sound effect mixing device 100 cannot mix cheers or applause into the beginning part of the music based on the start time of the music. Also, the volume and manner of this cheers or applause depend on the key tone of the music played at the live venue.

[0040] Therefore, in this embodiment, the melody analysis unit 164 analyzes the melody of the music based on the information about the sound of the music from when the volume value of the music exceeds a predetermined threshold (the first threshold) until a predetermined time (the first time) elapses, and the sound effect output control unit 163 outputs the first sound effect based on the melody analyzed by the melody analysis unit 164 when the first time elapses after the volume value of the music exceeds the first threshold.

[0041] For this reason, in this embodiment, even when the sound effect mixing device 100 does not receive the input of the information about the start time of the music, it is possible to mix cheers and applause into the beginning part of the music. Also, in this embodiment, the melody of the music is analyzed using only the information about the sound of the music until the timing when the first sound effect (such as cheers and applause) is output (the timing when the first time elapses after the volume value of the music exceeds the first threshold). Therefore, in this embodiment, it is possible to mix cheers and applause into the music by real-time processing.

[0042] Here, the first time is related to, for example, the time until the audience at the live venue feels the start of the performance. Generally, it takes from 1 second to 1.4 seconds from the start of the performance until the audience at the live venue feels the start of the performance. Therefore, the first time is preferably, for example, 1 second or more and 1.4 seconds or less.

[0043] Also, the time it takes for the audience at the live venue to sense the start of the performance varies based on the size of the live venue. When the live venue is smaller, the distance between the audience and the performer is closer, and the time it takes for the audience to sense the start of the performance is shorter. Therefore, the plurality of modes may include modes for each scale of the live venue, and the first time may be determined based on the mode selected by the mode selection unit 165. At this time, it is preferable that the first time is determined to increase as the scale of the live venue increases. For example, when the mode for a small-scale venue is selected, the first time is determined to be 1 second, when the mode for a medium-scale venue is selected, the first time is determined to be 1.2 seconds, and when the mode for a large-scale venue is selected, the first time is determined to be 1.4 seconds.

[0044] After the output of the first sound effects (such as cheers and applause) starts, the melody of the music may change. Therefore, the melody analysis unit 164 may analyze the melody of the music based on the information regarding the sound of the music after the first time has elapsed since the volume value of the music exceeded the first threshold. Then, if the melody analyzed by the melody analysis unit 164 changes, the sound effect output control unit 163 may change the sound effects to be output based on the change in the melody.

[0045] For example, when the melody of the music changes from a quiet melody to an intense melody, the sound effect output control unit 163 may output a mixture of the sound effects for the intense melody and the sound effects for the quiet melody.

[0046] Also, when the melody of the music changes from the first melody to the second melody, the sound effect output control unit 163 may change the sound effects to be output from the sound effects for the first melody to the sound effects for the second melody. When changing from the sound effects for the first melody to the sound effects for the second melody, the sound effect output control unit 163 may gradually decrease the output level of the sound effects for the first melody to zero (fade out) and gradually increase the output level of the sound effects for the second melody from zero (fade in).

[0047] <Output of the first sound effect (such as cheers or applause) at the end of the music> <In a live venue, when a music piece is being played, at the end part of the performance, cheers or applause from the audience may occur. When the sound effect mixing device 100 only receives the input of information (audio signal or streaming data) regarding the music sound from the playback device and does not receive the input of information regarding the end time of the music, the sound effect mixing device 100 cannot mix cheers or applause into the end part of the music based on the end time of the music.>

[0048] <Therefore, in this embodiment, when the volume value of the music becomes equal to or lower than a predetermined threshold value (the second threshold value), the sound effect output control unit 163 outputs the first sound effect (such as cheers or applause) from the sound effect output unit 130.>

[0049] <At this time, the timing when the audience feels that the performance has ended varies depending on the volume value of the music. Therefore, in this embodiment, this second threshold value is determined based on a value related to the volume value of the music. The value related to the volume value of the music is, for example, the most recent average value of the volume value of the music. The most recent average value of the volume value of the music may be the most recent moving average value of the volume value of the music, or may be an average value obtained using a time constant circuit. When the music feature amount analysis unit 162 performs frame processing, the most recent moving average value of the volume value of the music is preferably, for example, the average value of the RMS values in the most recent predetermined number (2 or more) of frames.>

[0050] When the second threshold value is determined based on the most recent average value of the volume value of the music, as shown in FIG. 7, it is preferable that the second threshold value be a value obtained by subtracting an offset value from the most recent average value of the volume value of the music. The value obtained by subtracting the offset value from the most recent average value of the volume value of the music may be negative. Therefore, it is preferable to set a minimum value (minimum threshold value) for the second threshold value. Then, as shown in FIG. 7, if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is equal to or greater than the minimum threshold value, the second threshold value is set to the value obtained by subtracting the offset value from the most recent average value of the volume value of the music, and if the value obtained by subtracting the offset value from the most recent average value of the volume value of the music is less than the minimum threshold value, the second threshold value is set to the minimum threshold value. In FIG. 7, the thin solid line represents the most recent average value of the volume value of the music, the thick solid line represents the second threshold value, the dashed-dotted line represents the minimum threshold value, and the broken line represents the value obtained by subtracting the offset value from the most recent average value of the volume value of the music.

[0051] The determination of the end of the music (the determination of whether the volume value of the music has become equal to or less than the second threshold value) may be made at the end portion of the music. Therefore, it is preferable that the sound effect output control unit 163 does not determine whether the volume value of the music is equal to or less than the second threshold value for a predetermined period (first period) from the start of the music. Here, the start of the music is, for example, the timing at which the input of information regarding the sound of the music is started or the timing at which the volume value of the music exceeds the first threshold value. Also, as will be described in detail below, when the information regarding the sound of the music (sound signal or streaming data) includes two or more pieces of music, the start of the next piece of music among two consecutive pieces of music is the cut-off point between the previous piece of music and the next piece of music.

[0052] <Determination of the end portion of the music> Depending on the music, there may be a silent portion (break). As described above, when a portion that has become equal to or less than the second threshold value is determined to be the end portion of the music, the cheer or applause at the end portion of the music will also be mixed into this break portion.

[0053] Therefore, in this embodiment, after the volume value of the music becomes equal to or less than the second threshold value, if the state where the volume value of the music is equal to or less than the second threshold value continues for a predetermined time (second time), when the second time has elapsed since the volume value of the music became equal to or less than the second threshold value, the sound effect output unit 130 outputs a first sound effect (such as cheers or applause).

[0054] At this time, the second time may be a constant. At this time, the second time may be an integer multiple of the beat of the music, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music.

[0055] Also, the stronger the beat of the music, the longer the break period tends to be. Therefore, the second time may be determined based on a value related to the intensity of the beat of the music. At this time, it is preferable that the second time becomes longer as the intensity of the beat of the music increases.

[0056] At this time, as shown in FIG. 8, the second time may change linearly with respect to the intensity of the beat of the music, or may change stepwise with respect to the intensity of the beat of the music as shown in FIG. 9. When the second time changes stepwise with respect to the intensity of the beat of the music (FIG. 9), the second time at each step may be an integer multiple of the beat of the music, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music.

[0057] The intensity of the beat of the music is determined based on, for example, the volume value of the music. When the music feature amount analysis unit 162 performs frame processing and the volume value of the music is the RMS in each frame, the intensity of the beat of the music is preferably, for example, the maximum value among the rising differences in a predetermined number (2 or more) of the most recent frames. Here, the rising difference is the difference between the RMS of the later frame and the RMS of the previous frame when the RMS of the later frame is larger than the RMS of the previous frame in the RMS of two consecutive frames (FIG. 10).

[0058] <Output of the third sound effect (such as clapping)> At a live venue, the audience may keep rhythm by clapping or tapping their feet in time with the music. Therefore, in this embodiment, the music feature amount analysis unit 162 analyzes the beat position of the music based on the information regarding the sound of the music acquired by the music acquisition processing unit 161. Then, the sound effect output control unit 163 outputs the third sound effect (such as clapping) at the beat position analyzed by the music feature amount analysis unit 162.

[0059] In a playback device that plays back music, operations such as pause and fast forward may be performed. When the sound effect mixing device 100 only receives the input of information regarding the sound of the music (audio signal or streaming data) from the playback device and does not receive the input of information regarding the operations performed on the playback device, after the timing when operations such as pause and fast forward are performed, the clapping mixed in the music may deviate significantly from the beat position of the music.

[0060] Therefore, in this embodiment, the music feature amount analysis unit 162 analyzes, based on the information regarding the sound of the music acquired by the music acquisition processing unit 161, not only the beat position of the music but also the confidence level of the beat position. The confidence level of the beat position is an index indicating the accuracy of the beat position. Then, the sound effect output control unit 163 gradually reduces the output level of the third sound effect to zero after the beat position at which the value of the confidence level becomes equal to or lower than a predetermined confidence level value.

[0061] At this time, it is preferable that the time from when the output level of the sound effect starts to gradually decrease until it becomes zero is equal to or shorter than a predetermined time (for example, 10 seconds). By doing so, when the clapping mixed in the music deviates from the beat position of the music, it becomes possible to interrupt the mixing of the clapping without the listener feeling unnatural.

[0062] Also, at this time, the sound effect output control unit 163 may output the third sound effect at a constant tempo after the beat position where the confidence value becomes equal to or less than the first confidence value, and gradually decrease the output level of the third sound effect to zero. Here, the constant tempo may be determined based on the interval between the beat position where the confidence value becomes equal to or less than the first confidence value and the beat position immediately before the beat position. Also, the constant tempo may be determined based on the average value of the intervals between two adjacent beat positions until the confidence value becomes equal to or less than a predetermined confidence value. By doing so, when the tempo mixed into the music deviates from the beat position of the music, it becomes possible to interrupt the mixing of the tempo without the listener feeling unnatural.

[0063] The analysis of the beat position of the music and the confidence level of the beat position may be performed, for example, using the technique described in Makoto Goto, Yoichi Murakami, "Real-Time Beat Tracking System for Acoustic Signals - Corresponding to Music without Percussion Sounds by Code Change Detection -", Transactions of the Institute of Electronics, Information and Communication Engineers D, 1998 / 2, Vol. J81-D2 pp.227-237.

[0064] <Determination of the break between two pieces of music> The analysis of the feature amount of the music becomes more accurate as the amount of information that can be used is larger. For this reason, when two or more pieces of music are included in the information related to the sound of the music (sound signal or streaming data), it is advisable to analyze the feature amount of the next piece of music from the break between the previous piece of music and the next piece of music. At this time, if the timing when the volume value of the next piece of music exceeds the first threshold value is regarded as the break between the previous piece of music and the next piece of music, the information of the next piece of music while the volume value is less than the first threshold value will not be used for the analysis of the feature amount of the next piece of music. Also, if the timing when the volume value of the previous piece of music becomes equal to or less than the second threshold value is regarded as the break between the previous piece of music and the next piece of music, the information of the previous piece of music while the volume value is less than the second threshold value may be used for the analysis of the feature amount of the next piece of music.

[0065] Therefore, when information related to the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, for example, a third threshold smaller than the second threshold is set. When the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues for a predetermined period (second period), it is advisable to consider that second period as a silent section between the previous piece of music and the next piece of music. Therefore, when information related to the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, when the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues for the second period, the music feature amount analysis unit 162 and the melody analysis unit 164 determine the timing related to the period during which the state where the volume value of the previous piece of music is equal to or lower than the third threshold continues (for example, the timing when the volume value of the previous piece of music becomes equal to or lower than the third threshold, or the timing after the second period from the timing when the volume value of the previous piece of music becomes equal to or lower than the third threshold) as the cut-off point between the previous piece of music and the next piece of music (that is, the start of the next piece of music), and based on the information related to the audio of the music after the cut-off point between the previous piece of music and the next piece of music, it is advisable to analyze the feature amount and melody of the next piece of music.

[0066] In particular, when information related to the audio of a piece of music (audio signal or streaming data) includes two or more pieces of music, when the first effect sound (such as cheers or applause) is output when the first time has elapsed since the volume value of the next piece of music exceeded the first threshold by the effect sound output control unit 163, the melody analysis unit 164 analyzes the melody of the next piece of music based on the information related to the audio of the music between the cut-off point between the previous piece of music and the next piece of music and the timing when the first time has elapsed since the volume value of the next piece of music exceeded the first threshold, and the effect sound output control unit 163 may output the first effect sound based on the melody.

[0067] Noise may be included in the silent section between two consecutive pieces of music. Therefore, it is advisable that the effect sound output control unit 163 determines that the volume value of the previous piece of music continues to be equal to or lower than the third threshold even if a momentary increase in volume value (noise) occurs after the volume value of the previous piece of music has become equal to or lower than the third threshold.

[0068] The present invention has been described above according to its preferred embodiments. Although the present invention has been described with specific examples here, various modifications and changes can be made to these specific examples without departing from the spirit and scope of the present invention described in the claims.

Explanation of Reference Numerals

[0069] 100 Sound effect mixing device 110 Music input section 120 Storage section 130 Sound effect output section 140 Mixing section 150 External output section 160 Control section 161 Music acquisition processing section 162 Music feature quantity analysis section 163 Sound effect output control section 164 Key analysis section 165 Mode selection section 170 User input section

Claims

1. A music acquisition processing unit that acquires information regarding the audio of a music piece, A melody analysis unit that analyzes the melody of the music piece, An effect sound output control unit that controls the output of effect sounds, and having, The melody analysis unit analyzes the melody of the music piece based on the information regarding the audio of the music piece from when the volume value of the music piece exceeds a predetermined threshold until a predetermined time has elapsed, The effect sound output control unit outputs an effect sound based on the melody analyzed by the melody analysis unit when the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold. An information processing apparatus.

2. Further comprising a mode selection unit that selects one mode from among a plurality of modes, The predetermined time is determined based on the mode selected by the mode selection unit. The information processing apparatus according to claim 1.

3. The plurality of modes include modes for each scale of a live venue, The predetermined time is determined to increase as the scale of the live venue increases. The information processing apparatus according to claim 2.

4. The predetermined time is 1 second or more and 1.4 seconds or less. The information processing apparatus according to claim 1.

5. The melody analysis unit analyzes the melody of the music piece based on the information regarding the audio of the music piece after the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold, If the melody analyzed by the melody analysis unit changes, the effect sound output control unit changes the effect sound to be output based on the change in the melody. The information processing apparatus according to claim 1.

6. The information regarding the audio of the music piece includes two or more music pieces, The melody analysis unit, When a state where the volume value of the previous music piece is equal to or less than a third threshold continues for a predetermined period, the timing regarding the period during which the state where the volume value of the previous music piece is equal to or less than the third threshold continues is determined as the cut-off between the previous music piece and the next music piece, Based on the information regarding the audio of the music piece between the cut-off between the previous music piece and the next music piece and the timing when the predetermined time has elapsed since the volume value of the next music piece exceeded the predetermined threshold, the melody of the next music piece is analyzed. The information processing apparatus according to claim 1.

7. The information regarding the audio of the music piece includes the audio signal of the music piece and / or streaming data. The information processing apparatus according to claims 1 to 6.

8. The information processing apparatus according to claim 7, which executes real-time processing.

9. An information processing method executed by a computer, comprising: a music acquisition processing step of acquiring information regarding the sound of a music piece; a melody analysis step of analyzing the melody of the music piece; a sound effect output control step of controlling the output of a sound effect, wherein in the melody analysis step, the melody of the music piece is analyzed based on the information regarding the sound of the music piece from when the volume value of the music piece exceeds a predetermined threshold until a predetermined time has elapsed; and in the sound effect output control step, a sound effect based on the melody analyzed in the melody analysis step is output when the predetermined time has elapsed since the volume value of the music piece exceeded the predetermined threshold. Information processing method.

10. An information processing program for causing a computer to execute the information processing method according to claim 9.

11. A computer-readable storage medium storing the information processing program according to claim 10.

Citation Information

Patent Citations

  • Sound effect output device

    JP2021162707A

  • Effect sound mixing device

    JP2023050570A

  • Audio output device

    WO2023054236A1

  • Sound effect output device

    WO2023054237A1