Information processing device

The information processing apparatus addresses the challenge of real-time sound effect mixing in streaming music by analyzing the beat position and confidence level of music pieces and outputting sound effects accordingly, resulting in a more natural and immersive listening experience.

WO2025115210A1PCT designated stage expired Publication Date: 2025-06-05PIONEER IP
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2023/043088
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-12-01
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

Existing techniques for mixing sound effects into music during streaming playback do not consider real-time processing, resulting in an unnatural experience for listeners.

Method used

An information processing apparatus that includes a music acquisition processing unit, a music feature amount analysis unit, and a sound effect output control unit, which analyzes the beat position and confidence level of a music piece and outputs sound effects in real-time, gradually reducing the output level to zero when the confidence level drops below a predetermined value.

Benefits of technology

This solution enables natural integration of sound effects into music during streaming playback, enhancing the listener's experience by simulating the atmosphere of a live venue, even when the music's tempo deviates from the beat position.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2023043088_05062025_PF_FP_ABST
    Figure JP2023043088_05062025_PF_FP_ABST
Patent Text Reader

Abstract

With the present invention, audio effects that are more natural to a listener are introduced into a piece of music. Beat positions in a piece and a certainty factor of each of the beat positions are analysed, and audio effects are output on the beat positions analysed. The output level of the audio effects is gradually reduced to zero from the beat position where the value of the certainty factor decreased to or below a preset certainty factor value.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device

[0001] The present invention relates to an information processing device.

[0002] There are known techniques for mixing sound effects into music to create the atmosphere of a live concert venue. For example, Patent Documents 1 to 4 disclose techniques for mixing sound effects that sound natural to listeners into music.

[0003] Japanese Patent Application Publication No. 2021-162707 International Publication No. 2023 / 054236 International Publication No. 2023 / 054237 Japanese Patent Application Publication No. 2023-50570

[0004] In recent years, streaming playback has become popular as a method for playing music. In order to mix sound effects with music being streamed, the sound effects must be mixed in real time. The techniques disclosed in Patent Documents 1 to 4 do not take real time processing into consideration.

[0005] One example of a problem that the present invention aims to solve is mixing sound effects that sound more natural to a listener into music.

[0006] In order to solve the above problem, the invention described in claim 1 comprises a music acquisition processing unit that acquires information related to the audio of a music piece; a music feature analysis unit that analyzes the beat positions of the music piece and the certainty of the beat positions based on the information related to the audio of the music piece acquired by the music acquisition processing unit; and a sound effect output control unit that outputs sound effects at the beat positions analyzed by the music feature analysis unit, wherein the sound effect output control unit gradually reduces the output level of the sound effects to zero after the beat position at which the certainty value becomes equal to or less than a predetermined certainty value.

[0007] The invention described in claim 8 is an information processing method executed by a computer, comprising: a music acquisition processing step of acquiring information related to the audio of a music piece; a music feature analysis step of analyzing the beat positions of the music piece and the certainty of the beat positions based on the information related to the audio of the music piece acquired by the music acquisition processing unit; and a sound effect output control step of outputting sound effects at the beat positions analyzed by the music feature analysis unit, wherein in the sound effect output control step, the output level of the sound effects is gradually reduced to zero after the beat position at which the certainty value becomes equal to or less than a predetermined certainty value.

[0008] The invention as set forth in claim 9 is an information processing program that causes a computer to execute the information processing method as set forth in claim 8.

[0009] A tenth aspect of the present invention is a computer-readable storage medium storing the information processing program according to the ninth aspect.

[0010] 1 is a diagram illustrating a sound effect mixing device 100 according to an embodiment of the present invention. FIG. 2 is a diagram illustrating processing in the sound effect mixing device 100. FIG. 3 is a diagram illustrating an example of music input by a music input unit 110 and sound effects output by a sound effect output unit 130. FIG. 4 is a diagram illustrating a control unit 160. FIG. 5 is a diagram illustrating partial overlap of frames. FIG. 6 is a diagram illustrating an example of processing operation in the control unit 160. FIG. 7 is a diagram illustrating an example of the relationship between the most recent average volume value of music and a second threshold. FIG. 8 is a diagram illustrating an example of the relationship between a second time and the strength of the beat of music. FIG. 9 is a diagram illustrating an example of the relationship between a second time and the strength of the beat of music. FIG. 10 is a diagram illustrating a rise difference.

[0011] An information processing device according to one embodiment of the present invention includes a music acquisition processing unit that acquires information about the audio of a piece of music; a music feature analysis unit that analyzes the beat positions of the piece of music and the confidence of the beat positions based on the information about the audio of the piece of music acquired by the music acquisition processing unit; and a sound effect output control unit that outputs sound effects at the beat positions analyzed by the music feature analysis unit. The sound effect output control unit gradually reduces the output level of the sound effects to zero after a beat position at which the confidence value falls below a predetermined confidence value. The information about the audio of the piece of music may include an audio signal and / or streaming data of the piece of music. The information processing device may also perform real-time processing. Therefore, in this embodiment, if handclaps mixed into the music deviate from the beat positions of the music, the mixing of the handclaps can be interrupted without causing the listener to feel unnatural. As a result, this embodiment can more effectively evoke the atmosphere of being at a live concert.

[0012] The sound effect output control unit may output the sound effect at a constant tempo while gradually reducing the output level of the sound effect to zero after the beat position at which the certainty value becomes equal to or less than a predetermined certainty value. The constant tempo may be determined based on the distance between the beat position at which the certainty value becomes equal to or less than the predetermined certainty value and the beat position immediately preceding that beat position. The constant tempo may be determined based on the average value of the distance between two adjacent beat positions until the certainty value becomes equal to or less than the predetermined certainty value. In this way, when handclaps mixed into music deviate from the beat positions of the music, the mixing of handclaps can be interrupted without sounding unnatural to the listener.

[0013] The time it takes for the sound effect output level to decrease gradually to zero may be set to a predetermined time or less. This makes it possible to quickly stop mixing handclaps into music if the handclaps mixed into the music deviate from the beat position of the music.

[0014] An information processing method according to one embodiment of the present invention is an information processing method executed by a computer, and includes: a music acquisition processing step for acquiring information about the audio of a music piece; a music feature analysis step for analyzing the beat positions of the music piece and the confidence levels of the beat positions based on the information about the audio of the music piece acquired by the music acquisition processing unit; and a sound effect output control step for outputting sound effects at the beat positions analyzed by the music feature analysis unit. In the sound effect output control step, the output level of the sound effects is gradually reduced to zero after the beat position at which the confidence level falls below a predetermined confidence level. Therefore, in this embodiment, if handclaps mixed into the music piece deviate from the beat positions of the music piece, the mixing of the handclaps can be stopped without the listener feeling unnatural. As a result, this embodiment can more effectively evoke the atmosphere of being at a live concert.

[0015] An information processing program according to an embodiment of the present invention causes a computer to execute the above-described information processing method, thereby enabling listeners to experience the atmosphere of being at a live concert venue.

[0016] A storage medium according to one embodiment of the present invention stores the information processing program, which allows the information processing program to be distributed as a standalone program rather than being incorporated into a device, making it easy to perform version upgrades, etc.

[0017] 1 is a diagram showing a sound effect mixing device 100 according to an embodiment of the present invention. The sound effect mixing device 100 includes a music input unit 110, a storage unit 120, a sound effect output unit 130, a mixing unit 140, an external output unit 150, and a control unit 160.

[0018] 2, in the sound effect mixing device 100 according to this embodiment, for example, a music piece input from an external device to a music piece input unit 110 is mixed in a mixing unit 140 with sound effects output from a sound effect output unit 130, and the music piece mixed with the sound effects is output to the outside from an external output unit 150. In this case, in the sound effect mixing device 100 according to this embodiment, a storage unit 120 stores sound source data for sound effects, and a control unit 160 controls the output of sound effects from the sound effect output unit 130 by real-time processing.

[0019] The music input unit 110 receives information about the audio of a song from an external device and outputs the song (e.g., an audio signal of the song) to the mixing unit 140. The information about the audio of the song includes, for example, an audio signal of the song and streaming data. If the information about the audio of the song is streaming data of the song, the music input unit 110 converts the streaming data of the song into an audio signal of the song and outputs it to the mixing unit 140. The music input unit 110 includes, for example, a signal input terminal that receives an input of the audio signal of the song from a playback device that plays back the song, and a communication device that receives an input of the streaming data of the song from an external device. The music input unit 110 may also include an audio input device (e.g., a microphone) that receives input of the audio of the song.

[0020] The storage unit 120 is a storage device that stores information such as a hard disk drive, a solid state drive, a memory, etc. The storage unit 120 stores sound source data for sound effects.

[0021] The sound effect output unit 130 outputs sound effects (for example, audio signals of sound effects) to the mixing unit 140. The sound effect output unit 130, for example, acquires sound source data for sound effects stored in the storage unit 120, generates audio signals of sound effects from the acquired sound source data, and outputs the generated audio signals of sound effects to the mixing unit 140.

[0022] 3 is a diagram showing an example of music input to the music input unit 110 and sound effects output from the sound effect output unit 130. In the example shown in FIG. 3, first sound effects (cheers, applause, etc.) are output during the music and at the beginning and end of the music. Second sound effects (environmental sounds) are output continuously from before the music starts until it ends. Third sound effects (claps, etc.) are output in the middle of the music in synchronization with the beat and tempo of the music.

[0023] The mixing unit 140 mixes the audio signal of the music output from the music input unit 110 with the audio signal of the sound effects output from the sound effect output unit 130, and outputs the audio signal of the music mixed with the sound effects. The mixing unit 140 is, for example, a device that adds multiple signals and outputs the added signal, and adds the audio signal of the music output from the music input unit 110 and the audio signal of the sound effects output from the sound effect output unit 130, and outputs the added signal.

[0024] The external output unit 150 outputs the music output by the mixer 140 (i.e., the music mixed with sound effects) to the outside. The external output unit 150 includes, for example, an audio output device (e.g., a speaker) that outputs sound, a signal output terminal that outputs an audio signal, and a communication device that transmits information to another device.

[0025] The control unit 160 controls the output of sound effects from the sound effect output unit 130. The control unit 160 is an information processing device such as a computer. The control unit 160 of the sound effect output device 100 according to this embodiment performs real-time processing. That is, the control unit 160 processes information (audio signals and streaming data) related to the audio of music input from the outside by the music input unit 110 in real time, and controls the output of sound effects from the sound effect output unit 130 based on the feature quantities of the music obtained by this processing.

[0026] 4 is a diagram showing the control unit 160. The control unit 160 has a music acquisition processing unit 161, a music feature amount analysis unit 162, and a sound effect output control unit 163.

[0027] The music acquisition processing unit 161 acquires information (audio signals and streaming data) related to the audio of a song that has been externally input to the music input unit 110. At this time, the music acquisition processing unit 161 sequentially acquires information (audio signals and streaming data) related to the audio of a song, for example, each time the information (audio signals and streaming data) related to the audio of the song is input to the music input unit 110. In other words, the music acquisition processing unit 161 acquires, for example, the audio signals and streaming data of a song that have been externally input to the music input unit 110 in real time.

[0028] The music feature amount analysis unit 162 analyzes the feature amounts of a music piece based on information about the audio of the music piece (music audio signals or streaming data) acquired by the music acquisition processing unit 161. The feature amounts of a music piece include, for example, the volume value of the music piece, the position of the beats of the music piece, the tempo of the music piece (for example, BPM (Beats Per Minute)), and the strength of the beats of the music piece.

[0029] The music feature amount analysis unit 162 may analyze music feature amounts by, for example, frame processing. In this case, it is preferable to have some of the frames overlap, as shown in FIG. 5 . This makes it possible to analyze music feature amounts without being affected by, for example, small fluctuations in the audio signal. When the music feature amount analysis unit 162 performs frame processing, the volume value of the music in each frame is, for example, the RMS (Root Mean Square) value of each frame.

[0030] The sound effect output control unit 163 controls the output of sound effects from the sound effect output unit 130 based on the feature amount of the music analyzed by the music feature amount analysis unit 162 .

[0031] The sound effect output control unit 163 controls the volume value of the sound effects output from the sound effect output unit 130, for example, based on the volume value of the music analyzed by the music feature amount analysis unit 162. In this way, it is possible to prevent the volume of the mixed sound effects from being too loud or too quiet compared to the volume of the music, making it possible to impart sound effects to the music that are more natural to the listener, and enabling the listener to experience the atmosphere of being at a live concert.

[0032] The control unit 160 may further include a music melody analysis unit 164, as shown in FIG. 4 . The music melody analysis unit 164 analyzes the music melody based on information about the audio of the music (music audio signal or streaming data) acquired by the music acquisition processing unit 161 and the music feature quantities acquired by the music feature quantity analysis unit 162. The sound effect output control unit 163 may then control the output of sound effects from the sound effect output unit 130 based on the music melody analyzed by the music melody analysis unit 164. The sound effect output control unit 163 may, for example, control the volume value of the sound effects output from the sound effect output unit 130. In this case, the storage unit 120 may store sound source data for sound effects prepared for each volume value and music melody of the music, and the sound effect output control unit 163 may determine the sound source data for sound effects to be output from the sound effect output unit 130 based on the analyzed volume value and music melody of the music.

[0033] 4, the control unit 160 may further include a mode selection unit 165. The mode selection unit 165 selects one mode from a plurality of modes related to the melody and genre of the music, the size of the live venue, etc. In this case, the mode selection unit 165 may select one mode from the plurality of modes based on, for example, the analyzed features and melody of the music. The sound effect output control unit 163 may then control the output of sound effects from the sound effect output unit 130 based on the mode input by the mode selection unit 165.

[0034] For example, the multiple modes may include modes for different sizes of live venues. The multiple modes may include, for example, a mode for large venues such as stadiums, outdoor festivals, and arenas, a mode for medium-sized venues such as halls and medium- to large-sized live music venues, and a mode for small venues such as small live music venues and music bars. In this case, for example, the storage unit 120 may store sound effects for large venues, sound effects for medium-sized venues, and sound effects for small venues.

[0035] 1, the sound effect mixing apparatus 100 may further include a user input unit 170, and the mode selection unit 165 may select one mode from among a plurality of modes based on input from the user input unit 170. The user input unit 170 is an input device that receives information input from a user, such as a button, keyboard, touch panel, camera, or microphone.

[0036] <Processing Operation in Control Unit 160> FIG. 6 is a diagram showing an example of the processing operation in the control unit 160. The processing operation shown in FIG. 6 is executed, for example, at predetermined time intervals. When the music feature amount analysis unit 162 performs frame processing, the processing operation shown in FIG. 6 is executed, for example, for each frame. The music acquisition processing unit 161 acquires information about the audio of music (e.g., music audio signals or streaming data) input from an external source to the music input unit 110 (step S601). The music feature amount analysis unit 162 and the music melody analysis unit 164 analyze the features and music melody of the music based on the information about the audio of the music acquired by the music acquisition processing unit 161 (step S602). The sound effect output control unit 163 controls the output of sound effects from the sound effect output unit 130 based on the features and music melody analyzed by the music feature amount analysis unit 162 and the music melody analysis unit 164 (step S603).

[0037] <Output of first sound effects (cheers, applause, etc.) at the beginning of a piece of music> At a live music venue, when a piece of music is being played, the audience may cheer or applaud at the beginning of the performance. If the sound effects mixing device 100 receives only information about the audio of the piece of music (audio signals or streaming data) from the playback device and does not receive information about the start time of the piece of music, the sound effects mixing device 100 cannot mix cheers or applause into the beginning of the piece of music based on the start time of the piece of music. Furthermore, the volume and style of the cheers and applause depend on the melody of the piece of music being played at the live music venue.

[0038] Therefore, in this embodiment, the melody analysis unit 164 analyzes the melody of the song based on information about the audio of the song from when the volume value of the song exceeds a predetermined threshold (first threshold) until a predetermined time (first time) has elapsed, and the sound effect output control unit 163 outputs a first sound effect based on the melody analyzed by the melody analysis unit 164 when the first time has elapsed since the volume value of the song exceeded the first threshold.

[0039] Therefore, in this embodiment, even if the sound effects mixing apparatus 100 does not receive input of information regarding the start time of the music, it is possible to mix cheers and applause into the beginning of the music. Also, in this embodiment, the melody of the music is analyzed using only information regarding the audio of the music up to the timing when the first sound effect (cheers, applause, etc.) is output (the timing when the first time period has elapsed since the volume value of the music exceeded the first threshold). Therefore, in this embodiment, it is possible to mix cheers and applause into the music through real-time processing.

[0040] Here, the first time period relates to, for example, the time it takes for an audience at a live venue to realize that a performance has begun. It generally takes 1 to 1.4 seconds from the start of the performance for an audience at a live venue to realize that a performance has begun. Therefore, it is preferable to set the first time period to, for example, 1 second or more and 1.4 seconds or less.

[0041] Furthermore, the time it takes for audience members at a live venue to realize that the performance has begun varies depending on the size of the venue. Smaller venues have a shorter distance between the audience and the performers, and the time it takes for audience members to realize that the performance has begun is shorter. Therefore, the multiple modes may include modes for each size of venue, and the first time may be determined based on the mode selected by the mode selection unit 165. In this case, the first time may be determined to increase as the size of the venue increases. For example, when a small-venue mode is selected, the first time may be set to 1 second; when a medium-sized venue mode is selected, the first time may be set to 1.2 seconds; and when a large-venue mode is selected, the first time may be set to 1.4 seconds.

[0042] The melody of the music may change after the output of the first sound effect (such as cheers or applause) begins. Therefore, the melody analysis unit 164 may analyze the melody of the music based on information about the audio of the music after a first time has elapsed since the volume value of the music exceeded a first threshold. If the melody analyzed by the melody analysis unit 164 has changed, the sound effect output control unit 163 may change the sound effect to be output based on the change in melody.

[0043] For example, when the melody of a piece of music changes from a quiet melody to a more intense melody, the sound effect output control unit 163 may output a mixture of the sound effect for the quiet melody and the sound effect for the more intense melody.

[0044] Furthermore, when the musical melody changes from a first musical melody to a second musical melody, the sound effect output control unit 163 may change the sound effect to be output from the sound effect for the first musical melody to the sound effect for the second musical melody. When changing from the sound effect for the first musical melody to the sound effect for the second musical melody, the sound effect output control unit 163 may gradually decrease the output level of the sound effect for the first musical melody to zero (fade out) and gradually increase the output level of the sound effect for the second musical melody from zero (fade in).

[0045] <Output of first sound effect (cheers, applause, etc.) at the end of a piece of music> At a live music venue, when a piece of music is being played, the audience may cheer or applaud at the end of the performance. If sound effects mixing apparatus 100 receives only input of information related to the audio of the piece of music (audio signals or streaming data) from the playback device and does not receive input of information related to the end time of the piece of music, sound effects mixing apparatus 100 will not be able to mix cheers or applause into the end of the piece of music based on the end time of the piece of music.

[0046] Therefore, in this embodiment, the sound effect output control unit 163 outputs a first sound effect (such as cheers or applause) from the sound effect output unit 130 when the volume value of the music falls below a predetermined threshold value (second threshold value).

[0047] At this time, the timing at which the audience perceives the performance as over varies depending on the volume value of the music. Therefore, in this embodiment, this second threshold is determined based on a value related to the volume value of the music. The value related to the volume value of the music is, for example, the most recent average volume value of the music. The most recent average volume value of the music may be the most recent moving average volume value of the music, or may be an average value obtained using a time constant circuit. When the music feature analysis unit 162 performs frame processing, the most recent moving average volume value of the music may be, for example, the average RMS value of the most recent predetermined number (two or more) of frames.

[0048] When the second threshold is determined based on the most recent average volume value of a song, it is preferable that the second threshold be a value obtained by subtracting an offset value from the most recent average volume value of a song, as shown in FIG. 7 . The value obtained by subtracting the offset value from the most recent average volume value of a song may be negative. Therefore, it is preferable to set a minimum value (minimum threshold) for the second threshold. Then, as shown in FIG. 7 , if the value obtained by subtracting the offset value from the most recent average volume value of a song is equal to or greater than the minimum threshold, the second threshold is set to the value obtained by subtracting the offset value from the most recent average volume value of a song. If the value obtained by subtracting the offset value from the most recent average volume value of a song is less than the minimum threshold, the second threshold is set to the minimum threshold. In FIG. 7 , the thin solid line represents the most recent average volume value of a song, the thick solid line represents the second threshold, the dashed-dotted line represents the minimum threshold, and the dashed line represents the value obtained by subtracting the offset value from the most recent average volume value of a song.

[0049] The determination of the end of a song (determining whether the volume value of the song has become equal to or less than the second threshold) may be performed at the end of the song. Therefore, the sound effect output control unit 163 may not determine whether the volume value of the song is equal to or less than the second threshold for a predetermined period (first period) from the start of the song. Here, the start of a song refers to, for example, the timing when input of information about the audio of the song begins or the timing when the volume value of the song exceeds the first threshold. Furthermore, as described in detail below, if the information about the audio of the song (audio signal or streaming data) includes two or more songs, the start of the next song of two consecutive songs is the boundary between the previous song and the next song.

[0050] <Determining the End of a Song> Some songs may have silent sections (breaks). If a section below the second threshold is determined to be the end of a song as described above, cheers and applause from the end of the song will be mixed into this break.

[0051] Therefore, in this embodiment, if the volume value of the music falls below the second threshold and remains below the second threshold for a predetermined period of time (second time), the sound effect output control unit 163 outputs a first sound effect (such as cheers or applause) from the sound effect output unit 130 when the second time has elapsed since the volume value of the music fell below the second threshold.

[0052] In this case, the second time may be a constant, which may be an integer multiple of the beat of the music piece, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music piece.

[0053] Furthermore, the stronger the beat of a piece of music, the longer the break period tends to be. Therefore, the second time period may be determined based on a value related to the strength of the beat of the piece of music. In this case, it is preferable that the second time period be set to be longer as the strength of the beat of the piece of music becomes stronger.

[0054] In this case, the second time may change linearly with respect to the strength of the beat of the music piece, as shown in Fig. 8, or may change in stages with respect to the strength of the beat of the music piece, as shown in Fig. 9. When the second time changes in stages with respect to the strength of the beat of the music piece (Fig. 9), the second time at each stage may be an integer multiple of the beat of the music piece, or may be a value obtained by adding a predetermined time (for example, 100 msec) to an integer multiple of the beat of the music piece.

[0055] The strength of a beat of a song is determined, for example, based on the volume value of the song. When the song feature amount analysis unit 162 performs frame processing and the volume value of the song is the RMS value for each frame, the strength of the beat of a song may be determined, for example, as the maximum value of the rising differences for a predetermined number (two or more) of frames most recently recorded. Here, the rising difference is the difference between the RMS values ​​of the subsequent frame and the previous frame when the RMS value of the subsequent frame is greater than the RMS value of the previous frame ( FIG. 10 ).

[0056] <Output of third sound effect (handclaps, etc.)> At a live venue, the audience may keep rhythm with the music by clapping their hands or stomping their feet. In this embodiment, the music feature amount analysis unit 162 analyzes the beat positions of the music based on information about the audio of the music acquired by the music acquisition processing unit 161. The sound effect output control unit 163 then outputs a third sound effect (handclaps, etc.) at the beat positions analyzed by the music feature amount analysis unit 162.

[0057] A playback device that plays back music may perform operations such as pausing, fast-forwarding, etc. If the sound effects mixing device 100 receives only input of information about the audio of the music (audio signals and streaming data) from the playback device and does not receive input of information about operations performed on the playback device, the handclaps mixed into the music may become significantly out of sync with the beat positions of the music after the operation such as pausing or fast-forwarding is performed.

[0058] Therefore, in this embodiment, the music feature amount analysis unit 162 analyzes not only the beat positions of the music but also the confidence levels of the beat positions based on information about the audio of the music acquired by the music acquisition processing unit 161. The confidence level of the beat positions is an index that indicates the accuracy of the beat positions. Then, the sound effect output control unit 163 gradually reduces the output level of the third sound effect to zero after the beat position where the confidence level value falls below a predetermined confidence level.

[0059] At this time, it is advisable to set the time from when the output level of the sound effect starts to decrease until it reaches zero to a predetermined time (for example, 10 seconds) or less. By doing so, even if the handclaps mixed into the music deviate from the beat position of the music, it is possible to stop mixing the handclaps without the listener feeling that it is unnatural.

[0060] Furthermore, at this time, the sound effect output control unit 163 may output the third sound effect at a constant tempo while gradually reducing the output level of the third sound effect to zero from the beat position where the certainty value becomes equal to or less than the first certainty value. Here, the constant tempo may be determined based on the distance between the beat position where the certainty value becomes equal to or less than the first certainty value and the beat position immediately preceding that beat position. Alternatively, the constant tempo may be determined based on the average value of the distance between two adjacent beat positions until the certainty value becomes equal to or less than a predetermined certainty value. In this way, when the handclaps mixed into the music deviate from the beat positions of the music, the mixing of the handclaps can be interrupted without the listener feeling that it is unnatural.

[0061] The beat positions of a piece of music and the confidence levels of those beat positions may be analyzed using, for example, the technology described in Goto Masataka and Muraoka Yoichi, "Real-time Beat Tracking System for Acoustic Signals - Adapting to Music Without Percussion Instruments by Detecting Chord Changes," IEICE Transactions on Electronics, Information and Communication Engineers, Vol. J81-D2, pp. 227-237, 1998 / 2.

[0062] <Determining the Break Between Two Songs> The analysis of song features becomes more accurate the more information is available. Therefore, when audio information (audio signals or streaming data) for songs includes two or more songs, it is preferable to analyze the feature of the next song from the break between the previous song and the next song. In this case, if the break between the previous song and the next song is defined as the time when the volume value of the next song exceeds a first threshold, information about the next song while its volume value is below the first threshold will not be used in analyzing the feature of the next song. Furthermore, if the break between the previous song and the next song is defined as the time when the volume value of the previous song falls below a second threshold, information about the previous song while its volume value is below the second threshold may be used in analyzing the feature of the next song.

[0063] Therefore, when the information about the audio of the music (audio signals or streaming data) includes two or more music pieces, for example, a third threshold smaller than the second threshold may be set, and when the volume value of the previous music piece remains equal to or less than the third threshold for a predetermined period (second period), the second period may be determined to be a silent portion between the previous music piece and the next music piece. When the information about the audio of the music (audio signals or streaming data) includes two or more music pieces, the music feature amount analysis unit 162 and the music melody analysis unit 164 may determine, when the volume value of the previous music piece remains equal to or less than the third threshold for the second period, that timing of the period during which the volume value of the previous music piece remains equal to or less than the third threshold (e.g., the timing when the volume value of the previous music piece becomes equal to or less than the third threshold, or the timing the second period after the timing when the volume value of the previous music piece becomes equal to or less than the third threshold) as the boundary between the previous music piece and the next music piece (i.e., the start of the next music piece), and analyze the features and music melody of the next music piece based on the information about the audio of the music piece after the boundary between the previous music piece and the next music piece.

[0064] In particular, when the information relating to the audio of the music (audio signal or streaming data) includes two or more music pieces, when the sound effect output control unit 163 outputs a first sound effect (such as cheers or applause) when a first time has elapsed since the volume value of the next music piece exceeded a first threshold, the melody analysis unit 164 may analyze the melody of the next music piece based on information relating to the audio of the music piece between the boundary between the previous music piece and the next music piece and the timing when the first time has elapsed since the volume value of the next music piece exceeded the first threshold, and the sound effect output control unit 163 may output the first sound effect based on the melody.

[0065] Noise may be included in the silent portion between two consecutive songs. Therefore, the sound effect output control unit 163 may determine that the volume value of the previous song continues to be equal to or less than the third threshold, even if a momentary increase in volume (noise) occurs after the volume value of the previous song falls below the third threshold.

[0066] The present invention has been described above in terms of preferred embodiments thereof. While the present invention has been described herein with reference to specific examples, various modifications and variations can be made to these examples without departing from the spirit and scope of the present invention as set forth in the claims.

[0067] REFERENCE SIGNS LIST 100 Sound effect mixing device 110 Music input unit 120 Storage unit 130 Sound effect output unit 140 Mixing unit 150 External output unit 160 Control unit 161 Music acquisition processing unit 162 Music feature amount analysis unit 163 Sound effect output control unit 164 Music melody analysis unit 165 Mode selection unit 170 User input unit

Claims

1. A music acquisition processing unit that acquires information related to the sound of a music piece, a music feature amount analysis unit that analyzes the beat position of the music piece and the confidence level of the beat position based on the information related to the sound of the music piece acquired by the music acquisition processing unit, and an effect sound output control unit that outputs an effect sound at the beat position analyzed by the music feature amount analysis unit, wherein the effect sound output control unit gradually reduces the output level of the effect sound to zero after the beat position where the value of the confidence level becomes equal to or less than a predetermined confidence level value, an information processing apparatus.

2. The information processing apparatus according to claim 1, wherein the effect sound output control unit gradually reduces the output level of the effect sound to zero while outputting the effect sound at a constant tempo after the beat position where the value of the confidence level becomes equal to or less than a predetermined confidence level value.

3. The information processing apparatus according to claim 2, wherein the constant tempo is determined based on the interval between the beat position where the value of the confidence level becomes equal to or less than a predetermined confidence level value and the beat position immediately before the beat position.

4. The information processing apparatus according to claim 2, wherein the constant tempo is determined based on the average value of the intervals between two adjacent beat positions until the value of the confidence level becomes equal to or less than a predetermined confidence level value.

5. The information processing apparatus according to claim 1, wherein the time from when the output level of the effect sound starts to gradually decrease until it becomes zero is equal to or less than a predetermined time.

6. The information processing apparatus according to claims 1 to 5, wherein the information related to the sound of the music piece includes the sound signal of the music piece and / or streaming data.

7. The information processing apparatus according to claim 6, wherein the information processing apparatus performs real-time processing.

8. An information processing method executed by a computer, comprising: a music acquisition processing step of acquiring information related to the sound of a music piece; a music feature amount analysis step of analyzing the beat position of the music piece and the confidence level of the beat position based on the information related to the sound of the music piece acquired in the music acquisition processing step; and an effect sound output control step of outputting an effect sound at the beat position analyzed in the music feature amount analysis step, wherein in the effect sound output control step, after the beat position where the value of the confidence level becomes equal to or less than a predetermined confidence level value, the output level of the effect sound is gradually reduced to zero, an information processing method.

9. An information processing program for causing a computer to execute the information processing method according to claim 8.

10. A computer-readable storage medium storing the information processing program according to claim 9.

Citation Information

Patent Citations

  • Sound effect output device

    JP2021162707A

  • Effect sound mixing device

    JP2023050570A

  • Audio output device

    WO2023054236A1

  • Sound effect output device

    WO2023054237A1

  • Sound effect outputting device

    JP2021162711A