Audio data processing device, audio data processing method, and program
The audio data processing device and method address the challenge of uniformly matching BPMs in songs with changing BPMs by performing targeted time stretching on beat units, resulting in improved mixing and playback experiences with preserved sound quality.
Patent Information
- Application Number
- PCT/JP2023/039039
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
Existing audio data processing techniques struggle to uniformly match the BPM of songs with changing BPMs, leading to difficulties in mixing songs with a natural listening experience.
An audio data processing device and method that calculates a ratio of the input audio data's beat unit length to a target length, and performs time stretching on each beat or beat unit to align the BPM uniformly, while preserving the natural sound quality of percussion sounds.
The solution effectively uniformizes the BPM of songs with changing BPMs, enhancing the naturalness of mixing and playback, while maintaining the sound quality of percussion elements.
Smart Images

Figure JP2023039039_08052025_PF_FP_ABST
Abstract
Description
Audio data processing device, audio data processing method and program
[0001] The present invention relates to an audio data processing device, an audio data processing method, and a program.
[0002] The FFT (Fast Fourier Transform) method is known as a method for achieving time stretching, which expands or compresses the length of digital audio data on the time axis without changing its pitch. For example, Patent Document 1 discloses an invention related to a pitch shifter that performs an FFT on a frame-by-frame basis and shifts the phase and amplitude of each resulting frequency component to perform pitch shifting. The pitch shifter estimates the actual frequency channel, which is the frequency channel where the frequency component actually exists, by referring to the amplitude, and performs phase correction on the phase after pitch shifting to maintain the phase difference between the actual frequency channel and nearby frequency channels. By performing such phase correction, audio data with a reduced sense of phase shift can be generated.
[0003] Patent Document 2 describes a technology for converting digital audio data into high-quality sound by maintaining the phase relationship between an arbitrary frequency band and its adjacent frequency band under certain conditions when performing phase calculation processing using the FFT method with the technology described above. More specifically, the technology executes a frequency conversion step for converting digital audio data into the amplitude and phase of each frequency component, a spectral peak detection step for detecting a spectral peak at which the amplitude is maximized, a phase difference calculation step for calculating the phase difference between the phase of the spectral peak and the phase of an adjacent frequency band adjacent to the spectral peak, and a phase determination step for determining whether or not to maintain the phase relationship between the spectral peak and the adjacent frequency band based on the phase difference.
[0004] JP 2008-216381 A Japanese Patent No. 6118522 A
[0005] The time stretching technique described above is used, for example, when a DJ mixes two songs together to match the BPMs of the two songs in order to achieve a natural listening experience. However, some songs have BPMs that change over the course of the song, and in such cases, even with time stretching, the BPMs of the two songs only match for a portion of the song, making it difficult to mix them in a way that sounds natural.
[0006] SUMMARY OF THE INVENTION It is therefore an object of the present invention to provide an audio data processing device, an audio data processing method, and a program that can facilitate use in mixing and the like by leveling out the BPM of music that fluctuates during the course of the music.
[0007] [1] An audio data processing device comprising: an analysis processing unit that calculates a ratio of a length in input audio data to a target length for a beat or a beat unit including multiple beats in a piece of music; and a time stretch processing unit that performs time stretching on the input audio data for each beat or beat unit in accordance with the ratio, thereby generating output audio data in which the beats or beat units are played back at the target length. [2] The audio data processing device described in [1], wherein the analysis processing unit sets the target length to an average value of the lengths of the beats or beat units in the input audio data. [3] The audio data processing device described in [1], wherein the analysis processing unit sets the length of the beats or beat units of a first piece of music to the target length for the beats or beat units in a second piece of music different from the first piece of music, and the time stretch processing unit performs time stretching on the input audio data of the second piece of music. [4] The musical piece includes a first part that is phonetically separable, and the analysis processing unit sets the ratio so that the beat or the beat unit reaches the target length by time stretching only the sections of the beat or the beat unit other than the pronunciation section of the first part, and the time stretch processing unit does not perform time stretch processing in the pronunciation section of the first part, but performs time stretch processing in accordance with the ratio in the sections other than the pronunciation section of the first part. An audio data processing device described in any one of [1] to [3].[5] The audio data processing device according to any one of [1] to [3], wherein the music piece includes a first part and a second part that are phonetically separable, and the analysis processing unit, in addition to the ratio, calculates a number of samples such that the beats or the beat units reach the target length by zero-filling or deleting samples in sections of the beats or the beat units other than the utterance sections of the first part, and the time stretch processing unit performs time stretching processing for the audio data of the second part for each beat or beat unit according to the ratio, and for the audio data of the first part, performs zero-filling or sample deletion in sections other than the utterance sections of the first part according to the number of samples without performing time stretching processing, and generates the output audio data by integrating the audio data of the first part and the second part after processing. [6] The audio data processing device according to [4] or [5], wherein the first part is composed of percussion sounds. [7] The audio data processing device according to [6], wherein the percussion sounds include a kick sound. [8] The audio data processing device according to any one of [1] to [7], wherein the analysis processing unit and the time stretch processing unit implement a processing unit that equalizes the BPM in advance without any particular target BPM value, and a processing unit that saves the equalized song. [9] An audio data processing method comprising the steps of: calculating a ratio of a length in input audio data to a target length for a beat in a song or a beat unit including multiple beats; and performing time stretching processing on the input audio data according to the ratio for each beat or beat unit, thereby generating output audio data in which the beat or the beat unit is reproduced at the target length.
[10] A program for causing a computer to implement the functions of calculating a ratio of a length in input audio data to a target length for a beat in a song or a beat unit including multiple beats, and performing time stretching processing on the input audio data according to the ratio for each beat or beat unit, thereby generating output audio data in which the beat or the beat unit is reproduced at the target length.
[0008] FIG. 1 is a diagram showing the overall configuration of a system according to an embodiment of the present invention. FIG. 2 is a block diagram showing a schematic functional configuration of an audio data processing device in the example of FIG. 1. FIG. 3 is a flowchart showing the processing flow of the audio data processing device shown in FIG. 2. FIG. 4 is a diagram conceptually showing an example of time stretch processing in an embodiment of the present invention. FIG. 5 is a diagram for explaining a first example of processing applicable in the time stretch processing according to an embodiment of the present invention. FIG. 6 is a diagram for explaining a first example of processing applicable in the time stretch processing according to an embodiment of the present invention. FIG. 7 is a diagram for explaining a second example of processing applicable in the time stretch processing according to an embodiment of the present invention.
[0009] FIG. 1 is a diagram showing the overall configuration of a system according to an embodiment of the present invention. The system 10 according to this embodiment includes a PC (Personal Computer) 100, a DJ controller 200, and a speaker 300. The PC 100 is a device that stores, processes, and plays audio data. It may be a terminal device such as a tablet or smartphone, not limited to a PC. The PC 100 includes a display 101 that displays information to the user and an input device such as a touch panel or mouse that acquires user input. The DJ controller 200 is connected to the PC 100 via a communication means such as a USB (Universal Serial Bus) and acquires user input related to music playback using a channel fader, crossfader, performance pad, jog dial, and various knobs and buttons. The audio data is played back using, for example, a speaker 300.
[0010] In this embodiment, the PC 100 functions as an audio data processing device in the system 10 described above. For example, the PC 100 performs processing of stored audio data in response to user input during playback of the audio data. Alternatively, the PC 100 may perform processing on the audio data prior to playback and store the processed audio data. In this case, the DJ controller 200 and speakers 300 may not be connected to the PC 100 when the processing is performed. In this embodiment, the PC 100 functions as an audio data processing device. However, in other embodiments, DJ equipment such as a mixer or an all-in-one DJ system (a digital audio player with communication and mixing functions) may function as an audio data processing device. Furthermore, a server connected to the PC or DJ equipment via a network may function as an audio data processing device.
[0011] 2 is a block diagram showing a schematic functional configuration of the audio data processing device in the example of FIG. 1. The PC 100 functioning as the audio data processing device includes an analysis processing unit 120 and a time stretch processing unit 130. These functions are implemented by a processor such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor) operating in accordance with a program. The program is read from the storage or removable recording medium of the PC 100, or downloaded from a server via a network, and loaded into the memory of the PC 100.
[0012] The analysis processing unit 120 calculates the ratio between the length in the input audio data 110 and the target length for a beat in a piece of music or a beat unit containing multiple beats. Here, a beat unit is a unit in a piece of music consisting of multiple beats, such as 2, 3, or 4 beats, such as a measure. Therefore, the analysis processing unit 120 calculates the ratio to the target length in units of, for example, one beat, one bar, or two beats in the case of a 4-beat time signature. The following describes an example in which the ratio to the target length is calculated for each beat, but similar processing can be performed when the unit is one bar or any number of beats.
[0013] The time stretch processing unit 130 performs time stretching on the input audio data 110 for each beat or beat unit in accordance with the ratio calculated by the analysis processing unit 120, thereby generating output audio data 140 in which the beats or beat units are played back at the target length. As described above, the analysis processing unit 120 calculates the ratio of the length of each beat or beat unit in the input audio data 110 to the target length, and therefore, when time stretching is performed in accordance with this ratio, the length of each beat or beat unit in the output audio data 140 is aligned to the target length. In other words, if the BPM of a piece of music fluctuates midway through, the lengths of the beats or beat units will be uneven depending on the position in the music, but by performing the time stretching process as described above, the lengths of the beats or beat units are aligned and the BPM is made uniform.
[0014] Note that detailed explanation of the time stretching process performed by the time stretching processor 130 will be omitted because conventional time stretching techniques can be used for target sections of beats or beat units rather than the entire musical piece. However, as will be described later, in addition to general time stretching, processing may be performed to prevent deterioration in the sound quality of percussion sounds, more specifically kick sounds.
[0015] FIG. 3 is a flowchart showing the processing flow of the audio data processing device shown in FIG. 2 . First, the analysis processing unit 120 detects fluctuations in beat position intervals (step S110). Specifically, for m beat positions pos[i] (i = 0 to m-1) in a song, the interval to the next beat position, T[i] = pos[i+1] - pos[i], is calculated. If T[i] ≠ T[i+1], fluctuations in beat position intervals are detected. Since the beat position interval can be considered the instantaneous BPM for that beat, fluctuations in beat position intervals indicate that the BPM of the song has fluctuated during the song. Note that even in songs with constant BPM, not all beat position intervals T[i] are necessarily of the exact same length. Therefore, fluctuations in beat position intervals may be detected if the variance in the distribution of beat position intervals T[i] exceeds a threshold. If fluctuations in beat position intervals are not detected, conventional time stretching processing may be performed in which the same time stretch rate is set for the entire song, instead of the following steps S120 to S140.
[0016] Next, the analysis processing unit 120 sets a target length of the beat position interval (step S120). The target length may be, for example, the average value T[i] of the beat position intervals T[i] calculated in step S110. ave= (ΣT[i]) / m. Alternatively, when mixing two songs using the DJ controller 200 of the example shown in FIG. 1 , the target length may be set according to the BPM of the song being mixed. For example, when mixing a first song whose BPM does not fluctuate with a second song whose BPM fluctuates, the beat length of the first song can be set to the target beat length of the second song, thereby equalizing the BPM of the second song after time stretching and matching it to the BPM of the first song. In addition to these examples, the target length may be set according to, for example, an arbitrary BPM set by the user. Even if the time stretching process is not performed in real time during mixing, the present invention has the advantage that the BPMs of the two songs can be matched using conventional time stretching techniques by performing the process of the present invention in advance to equalize the BPMs and saving the BPMs in memory, etc. There are two components for realizing the above-mentioned advantages: a "processing unit that equalizes the BPM in advance without any particular target BPM value" and a "processing unit that saves the equalized song," which may be realized by the analysis processing unit 120 and the time stretch processing unit 130. More specifically, in this case, the analysis processing unit 120 sets an arbitrary length (for example, but not limited to, the average length of the beats or beat units in the input audio data) to the target length, and the time stretch processing unit 130 performs time stretching processing so that the length of the beats or beat units becomes the target length, and saves the output audio data generated by the processing in a memory or the like.
[0017] Next, the analysis processing unit 120 calculates the time stretch rate for each beat (step S130). The time stretch rate rate[i] for the i-th inter-beat signal, i.e., the audio signal between beat positions pos[i] and pos[i+1], is the ratio of the length T[i] of the input audio data 110 to the target length set in step S120. For example, when the target length is an average value T ave If the time stretch rate is ave / T[i]. The target length is calculated as the average value T aveIn other cases, the time stretch rate rate[i] is calculated in the same manner.
[0018] Next, the time stretch processing unit 130 performs time stretch processing for each beat on the input audio data 110 in accordance with the control data calculated in steps S110 to S130 (step S140). ave In this case, the time stretch processing unit 130 calculates the length of the ith inter-beat signal, that is, the audio signal between the beat position pos[i] and the beat position pos[i+1], as follows: T[i]×rate[i]=T ave The output audio data 140 is generated by performing time stretching so that the target length is equal to the average value T ave In other cases, the time stretch process is similarly performed using the time stretch rate rate[i].
[0019] For example, when processing is performed to prevent deterioration in the sound quality of percussion sounds, more specifically kick sounds, as described below, a uniform time stretch rate rate[i] is not necessarily applied to the entire inter-beat signal, but the length of the inter-beat signal after performing time stretch processing according to the ratio calculated by the analysis processing unit 120 will still be the target length.
[0020] 4 is a diagram conceptually illustrating an example of time stretch processing according to an embodiment of the present invention. In this embodiment, as described above, different time stretch rates are applied to each beat or beat unit of a song to equalize the BPM of the song. In the illustrated example, a time stretch rate of 76% or 86% is applied to beats whose beat position interval T[i] is longer than the target length (shown as beat position intervals T1 and T2), and a time stretch rate of 100% is applied to beats whose beat position interval T[i] is the same as the target length (shown as beat position interval T3). This results in the length of each beat (beat position interval T3) and the BPM being equalized after time stretch processing.
[0021] The following describes further processing applicable to embodiments of the present invention for preventing deterioration in the sound quality of percussion sounds, more specifically kick sounds. In time stretching, the duration of the drum resonance of percussion sounds, more specifically kick sounds, may change, resulting in perceived deterioration in sound quality. Therefore, the following processing may be performed to prevent such deterioration.
[0022] 5A and 5B are diagrams illustrating a first example of processing applicable to the time stretching processing according to an embodiment of the present invention. When performing time stretching to stretch the interval between beats at a start point at which a kick sound K is included so that the original end point bx becomes the end point bt, if the time Ta of the entire interval between beats including the sounding interval of the kick sound K is stretched to time Tb as shown in FIG. 5A, the sounding interval of the kick sound K is also stretched by Tb / Ta times (kick sound K'). As described above, in this case, sound quality degradation is noticeable.
[0023] Therefore, as shown in Figure 5B, the time stretch rate is set so that the length of the beat becomes the target length by time stretching only the section (length Tc) other than the sound interval of the kick sound K between beats. In this case, where Td is the length of the section other than the sound interval of the kick sound K, the analysis processing unit 120 sets the time stretch rate Td / Tc so that (K + Td) / (K + Tc) = Tb / Ta. The time stretch processing unit 130 does not perform time stretch processing on the sound interval of the kick sound K, but performs time stretch processing to extend the section other than the sound interval of the kick sound K according to the time stretch rate Td / Tc set as described above. This makes it possible to prevent deterioration in the sound quality of the kick sound K while changing the length of the beat in the same way as in Figure 5A.
[0024] Although the above example describes the case where the interval between beats is extended, similarly, when the interval between beats is shortened, deterioration in the sound quality of the kick sound can be prevented by performing time stretching processing that shortens only the time of the interval other than the sound interval of the kick sound. Note that, for example, known techniques for separating parts of a musical piece can be used to detect the kick sound K and identify the sound interval.
[0025] 6 and 7 are diagrams illustrating a second example of processing applicable to the time stretching processing according to an embodiment of the present invention. In the example shown in FIG. 6 , a part separation process 121 is performed on input audio data 110, separating the audio data into a percussion part and a non-percussion part. For the audio data of the non-percussion part, a time stretch rate is calculated for each beat, as described above, and a time stretching process 131 is performed according to the calculated time stretch rate. For the audio data of the percussion part, the analysis process 120 performs a process 122 to calculate the number of samples for the zero interpolation / sample deletion process, and the time stretching process 130 performs a zero interpolation / sample deletion process 132 according to the calculated number of samples. Here, zero interpolation refers to a process of padding (padding) samples indicating no sound (zero samples) between beats. The number of samples for the zero interpolation / sample deletion process is calculated so that the length of the beat will be a separately calculated target length by interpolating or deleting zero samples in sections other than the sound interval of the percussion part within the beat. Therefore, the length of the beat in the audio data of the parts other than the percussion sounds after the time stretch process 131 is the same as the length of the beat in the audio data of the percussion sound part after the zero interpolation / sample deletion process 132. The audio data of each part after the processes is integrated to generate output audio data 140.
[0026] 7 is a diagram showing an example of waveform changes due to the processing shown in FIG. 6. In the example shown, the length of the beat is adjusted to the target length by interpolating zero samples or deleting samples in sections other than the sound intervals (silent sections) of the percussion sound part. For example, beats A and B are longer than the target length, so the beat lengths are extended by executing zero interpolation of 5000 samples and 2000 samples after the silent section, respectively. Beat C is longer than the target length, so the beat length is shortened by deleting 2000 samples from the silent section.
[0027] For example, percussion sounds such as kick sounds generally have a silent section after the sounding interval within a beat, so the length of the beat can be extended or shortened by zero interpolation or sample deletion, as in the above example, which corresponds to time stretching. By performing the above processing, the sounding interval of the percussion sound is not affected by the time stretching process, and deterioration in the sound quality of percussion sounds, including kick sounds, can be prevented.
[0028] In the above example, processing to prevent deterioration in sound quality was performed on a percussion sound or a kick sound part included in a percussion sound, but this is not limited to this example. For a piece of music including first and second parts that are phonetically separable, it is possible to perform time stretching processing outside the sound interval of the first part to set the length of the beat to a target length (example of FIG. 5B ), or to perform time stretching processing on the second part according to a time stretch rate for each beat, while zero-filling or sample deletion is used to set the length of the beat of the first part to a target length (examples of FIGS. 6 and 7 ). In the above example, the first part is a percussion sound or a kick sound.
[0029] 10...system, 100...PC, 101...display, 110...input audio data, 120...analysis processing unit, 130...time stretch processing unit, 140...output audio data, 200...DJ controller, 300...speaker
Claims
1. An audio data processing device comprising: an analysis processing unit that calculates a ratio of the length of input audio data to a target length for a beat or a beat unit including multiple beats in a musical piece; and a time stretch processing unit that performs time stretch processing on the input audio data in accordance with the ratio for each beat or beat unit, thereby generating output audio data in which the beat or beat unit is played back at the target length.
2. The audio data processing device according to claim 1, wherein the analysis processing section sets the target length to an average value of the lengths of the beats or beat units in the input audio data.
3. The audio data processing device of claim 1, wherein the analysis processing unit sets the length of a beat or beat unit of a first piece of music to the target length of the beat or beat unit in a second piece of music different from the first piece of music, and the time stretch processing unit performs time stretch processing on the input audio data of the second piece of music.
4. An audio data processing device as described in any one of claims 1 to 3, wherein the musical piece includes a first part that is phonetically separable, the analysis processing unit sets the ratio so that the beat or the beat unit reaches the target length by time stretching only a section of the beat or the beat unit other than the pronunciation section of the first part, and the time stretch processing unit does not perform time stretch processing in the pronunciation section of the first part, but performs time stretch processing in accordance with the ratio in a section other than the pronunciation section of the first part.
5. An audio data processing device as described in any one of claims 1 to 3, wherein the musical piece includes a first part and a second part which are acoustically separable, the analysis processing unit calculates, in addition to the ratio, a number of samples such that the beat or the beat unit will have the target length by zero-filling or deleting samples of sections of the beat or the beat unit other than the pronunciation section of the first part, and the time stretch processing unit performs time stretch processing on the audio data of the second part for each beat or beat unit in accordance with the ratio, and on the audio data of the first part, does not perform time stretch processing but zero-fills or deletes samples of sections other than the pronunciation section of the first part in accordance with the number of samples, and generates the output audio data by integrating the processed audio data of the first part and the second part.
6. An audio data processing device according to claim 4 or 5, wherein the first part is constituted by percussion sounds.
7. The audio data processing device according to claim 6, wherein the percussion sound includes a kick sound.
8. An audio data processing device as claimed in any one of claims 1 to 7, wherein the analysis processing unit and the time stretch processing unit realize a processing unit that equalizes the BPM in advance without any particular target BPM value, and a processing unit that stores the equalized song.
9. A method for processing audio data comprising the steps of: calculating a ratio of the length of input audio data to a target length for a beat or a beat unit including multiple beats in a musical piece; and generating output audio data in which the beat or the beat unit is played back at the target length by performing a time stretching process on the input audio data for each beat or beat unit in accordance with the ratio.
10. A program for causing a computer to realize the following functions: calculating the ratio of the length of input audio data to a target length for a beat or a beat unit including multiple beats in a piece of music; and generating output audio data in which the beat or beat unit is played back at the target length by executing a time stretch process on the input audio data in accordance with the ratio for each beat or beat unit.
Citation Information
Patent Citations
Recording medium, device and method for data recording, and device and method for data editing
JP2003228963A
Music signal conversion device and music signal conversion program
JP2017032603A
Playback control method, playback control device, and program
WO2019058942A1