Audio signal processing device, audio signal processing method and program
By using cross-correlation functions and weighted synchronous addition on band representative waveforms, the method addresses the limitations of existing audio signal processing technologies, achieving high-quality separation of instrument sounds on the time axis, enhancing sound quality and separation performance.
Patent Information
- Application Number
- JP2024540212
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2022-08-12
- Publication Date
- 2025-08-20
- Estimated Expiration
- 2042-08-12
AI Technical Summary
Existing audio signal processing technologies face limitations in achieving high time and frequency resolution in spectrograms, leading to difficulties in distinguishing simultaneous instrument sounds and resulting in degraded sound quality and S/N ratio, particularly in applications requiring minimal latency like DJ systems.
The method employs cross-correlation functions to detect sound generation positions on the time axis, using band representative waveforms from different frequency bands of kick sounds, and applies weighted synchronous addition to minimize noise from other instruments, thereby enhancing separation performance and sound quality.
This approach allows for high-quality separation of instrument sounds on the time axis, minimizing noise from other instruments and maintaining sound quality, especially in applications requiring precise timing like DJ systems.
Smart Images

Figure 0007727117000001 
Figure 0007727117000002 
Figure 0007727117000003
Abstract
Description
[Technical Field]
[0001] The present invention relates to an audio signal processing device, an audio signal processing method, and a program. [Background technology]
[0002] There is known a technique for extracting any musical sound from a musical piece. For example, Patent Document 1 describes an audio signal processing device including a sound generation position information acquisition unit that acquires sound generation position information indicating the sound generation position of any musical instrument included in the musical piece, a search section identification unit that identifies a search section for searching for a sound generation section of the any musical instrument sound based on the sound generation position information, an extraction unit that extracts an amplitude value at a predetermined position in the search section, and a processing unit that processes audio data included in the search section based on the amplitude value extracted by the extraction unit. [Prior art documents] [Patent documents]
[0003] [Patent Document 1] Patent No. 6263383 Summary of the Invention [Problem to be solved by the invention]
[0004] The technology described in Patent Document 1 obtains a spectrogram of an input piece of music, distinguishes between instrument sounds on the spectrogram data, and separates the sounds on the frequency axis. However, due to the principles of DFT (Digital Fourier Transform), time resolution and frequency resolution are in conflict with each other in a spectrogram, making it impossible to achieve both high time resolution and high frequency resolution. Furthermore, a spectrogram only contains information on power over time and frequency; phase information is unavailable. Due to these fundamental limitations, for example, instrument sounds that are playing simultaneously at the same frequency cannot be distinguished from one another. Therefore, when attempting to remove a specific instrument sound from a piece of music, sound that should not be removed may be removed, resulting in degradation of sound quality.
[0005] Although Patent Document 1 also describes synchronous addition processing, this is to average the results of synchronous subtraction of the spectrogram to reduce errors in the spectrogram shape, rather than synchronously adding the waveforms of the audio data. In the example of Patent Document 1, phase information is not used for the audio data, so even if synchronous addition is performed, the effect of improving the S / N ratio (S: the instrument sound to be extracted, N: other instrument sounds) cannot be obtained.
[0006] For example, in functions and products for DJs, latency is strictly controlled, and time lag must be minimized. Therefore, the time resolution, which affects the final sound position, is set high. However, due to the fundamental limitations of the spectrogram mentioned above, it is not possible to increase the frequency resolution. This technology relies almost entirely on frequency differences to distinguish between instruments, so if the frequency resolution cannot be increased, the result will likely be a deterioration in sound quality and other effects.
[0007] Therefore, an object of the present invention is to provide an audio signal processing device, an audio signal processing method, and a program that can achieve high sound quality and high separation performance by separating instrument sounds on the time axis. [Means for solving the problem]
[0008] [1] A musical piece including a first part and a second part that are phonetically separable, the musical piece including a sound analysis unit that detects the pronunciation position of the second part, The audio analysis unit calculates a cross-correlation function between a representative waveform of the second part extracted from the audio signal of the song and a waveform of the audio signal of the song in a section of length corresponding to the representative waveform as a function of time, and detects the section in which a peak of the cross-correlation function appears as the pronunciation position. [2] The voice analysis unit calculates the cross-correlation function within a predetermined search range, a first sound generation position detection process for detecting a first sound generation position within a first search range using the first representative waveform; a second sound generation position detection process for detecting a second sound generation position in a second search range that is set based on the first sound generation position and is smaller than the first search range, using a second representative waveform; The audio signal processing device according to [1], which executes the above. [3] The above voice analysis unit: In the first sound generation position detection process, a cross-correlation function is calculated between a first band representative waveform extracted from a first band sound signal obtained by extracting a first frequency band from the sound signal of the music piece and a section of the first band sound signal; The audio signal processing device according to [2], wherein the second sound production position detection process calculates a cross-correlation function between a second band representative waveform extracted from a second band audio signal obtained by extracting a second frequency band from the audio signal of the music piece and a section of the second band audio signal. [4] The second part above is composed of a kick sound, the first frequency band is a body resonance band of the kick sound, The audio signal processing device according to [3], wherein the second frequency band is an attack band of the kick sound. [5] An audio signal processing device according to any one of [1] to [4], wherein the representative waveform is generated by synchronously adding waveforms of the pronunciation interval of the second part in the audio signal of the music piece. [6] The audio signal processing device described in [5], wherein the audio analysis unit calculates a provisional cross-correlation function as a function of time between a provisional representative waveform extracted from the audio signal of the song according to a predetermined rule and the waveform of the audio signal of the song in a section of a length corresponding to the provisional representative waveform, and generates the representative waveform by synchronously adding the waveform of the audio signal of the song for the section in which a peak of the provisional cross-correlation function appears. [7] A method for generating a musical piece including a first part and a second part that are phonetically separable, the method including a sound analysis step of detecting a pronunciation position of the second part in the musical piece, The audio signal processing method includes a step of calculating a cross-correlation function of a representative waveform of the second part extracted from the audio signal of the music piece and a waveform of the audio signal of the music piece in a section of length corresponding to the representative waveform, as a function of time, and a step of detecting the section in which a peak of the cross-correlation function appears as the pronunciation position. [8] A sound analysis unit is provided for detecting the pronunciation position of a second part in a musical piece including a first part and a second part that are phonetically separable; The audio analysis unit calculates a cross-correlation function between the representative waveform of the second part extracted from the audio signal of the song and the waveform of the audio signal of the song in a section of length corresponding to the representative waveform, as a function of time, and detects the section in which the peak of the cross-correlation function appears as the pronunciation position. This is a program for causing a computer to function as an audio signal processing device. [Brief explanation of the drawings]
[0009] [Figure 1] 1 is a diagram showing the overall configuration of a system according to an embodiment of the present invention; [Figure 2] 2 is a block diagram showing a schematic functional configuration of the audio signal processing device in the example of FIG. 1. FIG. [Figure 3] 3 is a flowchart showing the overall flow of processing by the voice analysis unit shown in FIG. 2. [Figure 4] FIG. 1 is a diagram schematically illustrating the waveform structure of a kick sound. [Figure 5] 4 is a flowchart showing the process of detecting a sound generation position in milliseconds shown in FIG. 3. [Figure 6] 4 is a flowchart showing the sound generation position detection process in 10 microsecond units shown in FIG. 3. [Figure 7] 7 is a flowchart showing a process for generating a band representative waveform shown in FIGS. 5 and 6. [Figure 8] FIG. 8 is a diagram for conceptually explaining the process of calculating the cross-correlation function shown in FIG. 7. [Figure 9]FIG. 8 is a diagram for conceptually explaining the weighted synchronous addition process shown in FIG. 7. [Figure 10] 7 is a flowchart showing the process of detecting the sound generation position shown in FIGS. 5 and 6. [Figure 11] 11 is a diagram for conceptually explaining the process of detecting the sound generation position shown in FIG. 10. FIG. [Figure 12] 4 is a flowchart showing the Kick representative waveform generation process shown in FIG. 3. [Figure 13] FIG. 10 is a diagram for conceptually explaining an example in which different fade-out curves are applied to different bands. [Figure 14] 4 is a flowchart showing the kick sound removal sound generation process shown in FIG. 3. [Figure 15] 15 is a diagram for conceptually explaining the process of adding reverse-phase signals shown in FIG. 14. FIG. DETAILED DESCRIPTION OF THE INVENTION
[0010] FIG. 1 shows the overall configuration of a system according to one embodiment of the present invention. The system 10 according to this embodiment includes a PC (Personal Computer) 100, a DJ controller 200, and a speaker 300. The PC 100 is a device that stores, processes, and plays audio data. It may be a terminal device such as a tablet or smartphone, not just a PC. The PC 100 includes a display 101 that displays information to the user and an input device such as a touch panel or mouse that receives user input. The DJ controller 200 is connected to the PC 100 via a communication means such as a USB (Universal Serial Bus) and receives user input related to music playback using a channel fader, crossfader, performance pad, jog dial, and various knobs and buttons. The audio data is played back using, for example, a speaker 300.
[0011] In this embodiment, the PC 100 functions as an audio signal processing device in the system 10 described above. For example, the PC 100 performs processing of stored audio data in response to a user's operation input when the audio data is played back. Alternatively, the PC 100 may perform processing on the audio data before playback and store the processed audio data. In this case, the DJ controller 200 and speakers 300 do not need to be connected to the PC 100 when the processing is performed. In this embodiment, the PC 100 functions as an audio signal processing device, but in other embodiments, DJ equipment such as a mixer or an all-in-one DJ system (a digital audio player with communication and mixing functions) may function as an audio signal processing device. Furthermore, a server connected to the PC or DJ equipment via a network may function as an audio signal processing device.
[0012] Fig. 2 is a block diagram showing a schematic functional configuration of the audio signal processing device in the example of Fig. 1. The PC 100 functioning as the audio signal processing device includes an audio analysis unit 120, a display unit 140, a mix processing unit 150, and an operation unit 160. These functions are implemented by a processor such as a CPU (Central Processing Unit) or a DSP (Digital Signal Processor) operating in accordance with a program. The program is read from the storage or removable recording medium of the PC 100, or downloaded from a server via a network, and loaded into the memory of the PC 100.
[0013] The audio analysis unit 120 receives music audio data 110 including a first part and a second part that are acoustically separable. In this embodiment, the first part is a vocal and / or instrumental sound part other than the kick sound, and the second part is a kick sound part. Here, the kick sound is a bass drum sound or a synthesized sound that imitates the sound of a bass drum. The audio analysis unit 120 extracts kick sound-removed audio data 131, kick unit sound data 132, and kick pronunciation data 133 from the music audio data 110 using, for example, a music separation engine. Here, the kick sound-removed audio data 131 is data of audio obtained by removing the kick sound from the music audio data 110, i.e., audio data of the first part. The kick unit sound data 132 is data of the kick sound included in the music audio data 110, i.e., data of the unit sound of the second part (hereinafter also referred to as the kick unit sound). Kick sound generation data 133 is data indicating the sound generation position of a kick sound in music piece audio data 110. The sound generation position is the time position at which a kick sound is generated in music piece audio data 110, and is recorded, for example, as a time code in the music piece or a count in units of bars / beats.
[0014] A unit sound is a sound extracted by taking one pronunciation of the sound of the second part as a unit. In the following description, the waveform of the unit sound is also referred to as a representative waveform of the second part. For example, the audio analysis unit 120 extracts unit sounds by separating the Kick sound part from the music audio data 110, further dividing the Kick sound part into pronunciations, and classifying the pronunciations according to the characteristics of the audio waveform. Multiple unit sounds with different audio waveform characteristics may be extracted. The Kick unit sound data 132 may be, for example, audio data sampled from the Kick sound part, or may be temporal position information at which the unit sound is reproduced in the Kick sound part, or may be audio data of a sample sound similar to the extracted sound, or an identifier of the sample sound.
[0015] The display unit 140 displays information based on the Kick unit sound data 132 or the Kick sound generation data 133 on, for example, the display 101 of the PC 100. Meanwhile, the operation unit 160 acquires user input via an input device, such as a touch panel or mouse, of the PC 100. Specifically, for example, the display unit 140 displays the audio waveform of the music piece (which may be a waveform based on the music piece audio data 110 or a waveform based on the Kick sound-removed audio data 131) and the Kick sound generation position associated with the waveform, and the operation unit 160 acquires a user operation to change the Kick sound generation position to any position within the music piece. Alternatively, the display unit 140 may display the Kick sound placement according to a preset rhythm pattern, and the operation unit 160 may acquire a user operation to select a rhythm pattern. Note that, for example, when changing the Kick sound placement according to a preset rhythm pattern, the Kick sound position may be determined automatically without user operation. In this case, the display unit 140 and operation unit 160 described above may not be included in the functions of the audio signal processing device.
[0016] The mix processing unit 150 generates mixed audio data 170 based on the Kick sound-removed audio data 131 and the Kick unit sound data 132. The mixed audio data 170 is audio data in which the rearranged Kick unit sounds are mixed with the Kick sound-removed audio data 131. The sound generation positions of the Kick unit sounds in the mixed audio data 170 are determined according to a user operation acquired by the operation unit 160 as described above, or an automatically determined rhythm pattern. Here, the sound generation positions of the Kick unit sounds in the mixed audio data 170 may include positions different from the sound generation positions of the Kick sounds in the original music audio data 110.
[0017] FIG. 3 is a flowchart showing the overall flow of processing by the audio analysis unit shown in FIG. 2. As shown in the figure, the audio analysis unit 120 first detects kick sounds in units of sixteenth notes (step S110) and classifies the detected kick sounds (step S120). The detection processing in step S110 is performed using, for example, technology such as that described in International Publication No. 2017 / 168644. The classification processing in step S120 is performed, for example, by calculating a correlation function between the waveforms of the detected kick sounds and clustering them. The processing in steps S110 and S120 is processing for identifying the rough position of a kick sound in units of sixteenth notes (whether or not it exists) and the classification of the kick sound, and is preparation for identifying the onset position and representative waveform of the kick sound with higher accuracy.
[0018] Hereinafter, a loop process is executed for each of the kick sound classifications identified in step S120 (step S130). Specifically, a process for detecting the onset position in millisecond units (step S140), a process for detecting the onset position in 10 microsecond units (step S150), a process for generating a representative kick waveform (step S160), and a process for generating a kick-sound-removed sound (step S170) are executed for each of the identified kick sound classifications.
[0019] Before describing each process, the waveform structure of the kick sound used in this embodiment will be described. FIG. 4 is a diagram schematically illustrating the waveform structure of a kick sound. As shown in the figure, the waveform of a kick sound includes an attack portion (ATTACK) and a body resonance portion (SUSTAIN). Because the frequency bands and durations differ between the attack portion and the body resonance portion, distinguishing between them makes it possible to more accurately detect the kick sound's generation position and generate a representative kick waveform and kick-sound-free audio. Note that this type of waveform structure is not limited to kick sounds, but is also seen in other instrument sounds, such as percussion sounds including drum sounds including hi-hats and snares.
[0020] Fig. 5 is a flowchart showing the process of detecting the onset position in milliseconds shown in Fig. 3. In the process of detecting the onset position in milliseconds, first, the body resonance band (first frequency band) of the kick sound is extracted from the audio signal of the music by processing the audio signal of the music with a 200 Hz low-pass filter (step S141). For the audio signal of the body resonance band (first band audio signal), a process of generating a band representative waveform (step S142) and a process of detecting the onset position (step S143) are executed, thereby making it possible to detect the onset position of the kick sound in milliseconds.
[0021] Fig. 6 is a flowchart showing the process of detecting the sound generation position in 10 microsecond increments shown in Fig. 3. In the process of detecting the sound generation position in 10 microsecond increments, first, the audio signal of the music piece is processed with a 3 kHz high-pass filter to extract the audio signal of the attack band (second frequency band) of the kick sound from the audio signal of the music piece (step S151). For the audio signal of the attack band (second band audio signal), a process of generating a band representative waveform (step S152) and a process of detecting the sound generation position (step S153) are executed, thereby making it possible to detect the sound generation position of the kick sound in 10 microsecond increments.
[0022] FIG. 7 is a flowchart showing the band representative waveform generation process (steps S142 and S152) shown in FIGS. 5 and 6. In the band representative waveform generation process, a tentative band representative waveform is extracted from the band audio signal (the audio signal in the body or attack band of the kick sound) in each process according to a predetermined rule (step S210). For example, if one measure of a musical piece consists of four beats, the second and fourth beats are excluded because the kick sound is likely to occur simultaneously with the snare sound, and the first beat is excluded because the kick sound is likely to occur simultaneously with the cymbal sound. In this case, the waveform of the kick sound on the third beat is extracted as the tentative band representative waveform. If there are multiple kick sounds on the third beat, the kick sound levels for each measure may be histogrammed, and the kick sound on the third beat that is included in the highest frequency class may be selected. If there is no kick sound on the third beat, the tentative band representative waveform may be extracted from the kick sound on the first beat because the cymbal sound has relatively little overlap with the kick sound in the frequency band.
[0023] The process of generating a band representative waveform is performed in a millisecond-based pronunciation position detection process (step S142) and a 10-microsecond-based pronunciation position detection process (step S152). However, since each process is performed on a different band audio signal, even if the pronunciation position of the provisional band representative waveform is determined in step S210 according to a common rule, the provisional band representative waveform in each process is different.
[0024] Next, using the tentative band representative waveform determined in step S210, loop processing is executed for each Kick sound to be processed (step S220). Specifically, for each Kick sound, a cross-correlation function with the tentative band representative waveform is calculated within a predetermined search range (step S230), and a weighted synchronous addition is performed on the waveforms of the band audio signals for the interval where the peak of the cross-correlation function appears (step S240) to generate a band representative waveform of the Kick sound.
[0025] Fig. 8 is a diagram for conceptually explaining the calculation process of the cross-correlation function shown in Fig. 7. In the process of step S230 shown in Fig. 7, a tentative band representative waveform S is calculated from the band audio signal S0 of the music piece. temp Extract the section S1 of length corresponding to the time t, and calculate the waveform of each section as a function f temp (t), f1(t), and in the assumed time relationship, the cross-correlation function φ1(τ)=f temp (t)*f1(t+τ) is calculated, where τ is the amount of delay that was assumed. The amount of delay can be estimated from the position of the peak of the obtained cross-correlation function φ1(τ). In other words, the onset interval of the kick sound in the band audio signal can be identified.
[0026] 9 is a diagram for conceptually explaining the weighted synchronous addition process shown in FIG. 7. As explained above with reference to FIG. 8, the process of identifying the onset interval of the Kick sound in the band audio signal is executed for all of the Kick sounds to be processed that are included in the band audio signal S0. In the process of step S240 shown in FIG. 7, the waveforms of the band audio signal S0 are synchronously added for the onset interval of each Kick sound, thereby obtaining a band representative waveform S p In this embodiment, the waveforms of the sound generation intervals of the respective kick sounds are weighted by weighting coefficients W1, W2, W3, . . . set according to a predetermined rule and then synchronously added.
[0027] Synchronous addition is a technique for obtaining a waveform close to the original signal waveform when waveforms with the same characteristics repeatedly appear in a signal. By aligning the time of each waveform, they are added and averaged, and signals that are not correlated with the signal are reduced by phase cancellation, thereby obtaining a waveform close to the original signal waveform. However, it is difficult to obtain an effect unless the waveform times are aligned accurately. In the synchronous addition process in step S240 shown in Figure 7, the sounding intervals of each kick sound are specified in the previous step S230 as intervals where peaks of the cross-correlation function with the provisional band representative waveform appear, so that the noise-reduced band representative waveform S p can be obtained.
[0028] On the other hand, the weighting coefficients W1, W2, W3, ... used in the synchronous addition are set according to, for example, the position of the sounding interval of the kick sound within the music. For example, as in the generation of the temporary band representative waveform described above, the weighting coefficients may be set so that a low weight is assigned to a kick sound that is likely to occur simultaneously with other drum sounds, and a high weight is assigned to a kick sound that is unlikely to occur simultaneously with other drum sounds. For example, if one measure of a music piece consists of four beats, the weights on the second and fourth beats are the lowest because the kick sound is likely to occur simultaneously with a snare sound, and the weight on the first beat is the next lowest because the kick sound is likely to occur simultaneously with a cymbal sound. In this case, the waveforms of the sounding interval of each kick sound are weighted according to the number of beats and synchronously added. Furthermore, even among the first, second, and fourth beats, the kick sound that is most likely to occur simultaneously with other drum sounds is the common-time beat, but is unlikely to occur simultaneously on the half-time beat. Therefore, a high weight may be assigned to a kick sound located on a backbeat regardless of the number of beats. The weighting ratios of the kick sounds set from this perspective may be, for example, 0.8 / 0.5 / 1.0 / 0.5 for the first / second / third / fourth beats for the downbeats, and 1.0 for the backbeats regardless of the number of beats. In this case, the waveforms of the sounding intervals of each kick sound are weighted according to whether they are downbeats or backbeats, and are added synchronously.
[0029] Fig. 10 is a flowchart showing the sound generation position detection process (steps S143, S153) shown in Fig. 5 and Fig. 6. In the sound generation position detection process, following the band representative waveform generation process described above with reference to Figs. 7 to 9, a loop process is executed for each Kick sound to be processed (step S310). Specifically, a cross-correlation function between the band representative waveform and the waveform of the band audio signal within a predetermined search range is calculated (step S320), and the interval where the peak of the cross-correlation function appears is detected as the sound generation position of each Kick sound (step S330).
[0030] Fig. 11 is a diagram for conceptually explaining the process of detecting the sound generation position shown in Fig. 10. In step S320, a band representative waveform S is generated from the band audio signal S0 of the music piece. pExtract the section S2 of length corresponding to the time t, and calculate the waveform of each section as a function f p (t), f2(t), and in the assumed time relationship, the cross-correlation function φ2(τ)=f p (t)*f2(t+τ) is calculated, where τ is the amount of deviation that was assumed. The amount of deviation can be estimated from the peak position of the obtained cross-correlation function φ2(τ). In other words, the kick sound generation positions P1, P2, P3, ... can be identified.
[0031] In this embodiment, the band representative waveform generation process shown in FIGS. 7 and 8 also uses the tentative band representative waveform S temp The sound interval of the kick sound is identified based on the cross-correlation function φ1(τ) with the frequency band, and the sound position detection process shown in Figs. 10 and 11 also uses the band representative waveform S p The kick sound's position is detected based on the cross-correlation function φ2(τ) with the waveform S. temp is extracted from the band audio signal S0 of the music piece based on a rule, and therefore is a waveform that includes noise such as instrument sounds other than the kick sound. p is a waveform in which noise from instruments other than the Kick sound is reduced by synchronously adding the waveforms of the sounding intervals of the Kick sounds, and more accurately represents the characteristics of the original Kick sound waveform. temp The cross-correlation function φ1(τ) is used to identify the exact kick sound interval, and the waveforms of this interval are synchronously added to obtain the representative waveform S p The cross-correlation function φ2(τ) is used to identify the exact location of the kick sound.
[0032] Here, when the process of generating a band representative waveform and the process of detecting the onset position are performed in a onset position detection process in millisecond units (steps S142 and S143 shown in FIG. 5), because the onset positions of the kick sounds in sixteenth note units have been detected in the previous process (step S110 shown in FIG. 3), the cross-correlation function is calculated within a search range (first search range) of a sixteenth note (approximately 100 milliseconds), for example. As explained with reference to FIG. 4, the waveform of the kick sound's body band does not show a sharp peak, but does have a certain duration, so while the accuracy is lower than in the case of the attack band, which will be explained next, the calculation load can be reduced even when the search range is wide.
[0033] On the other hand, when the generation process of the band representative waveform and the detection process of the onset position are performed in the onset position detection process in 10-microsecond increments (steps S152 and S153 shown in FIG. 6), since the onset position of the kick sound has already been detected in millisecond increments in the previous process (step S140 shown in FIG. 3), the cross-correlation function is calculated, for example, by setting the search range (second search range) to several milliseconds before and after the onset position of the kick sound in millisecond increments. As explained with reference to FIG. 4, the waveform of the attack band of the kick sound exhibits a steep peak but its duration is short, so while accuracy is high, a wide search range increases the calculation load. Therefore, detecting the onset position using a waveform including the attack band from the beginning can increase the calculation load. However, in this embodiment, by first performing the onset position detection process in millisecond increments and narrowing the search range to several milliseconds, it is possible to detect the onset position with high accuracy while reducing the calculation load.
[0034] FIG. 12 is a flowchart showing the Kick representative waveform generation process shown in FIG. 3. In the Kick representative waveform generation process, weighted synchronous addition is performed on waveforms of the Kick sound's sound interval in the audio signal of the music piece, based on the sounding positions detected in 10-microsecond increments in the sounding position detection process (step S150 shown in FIG. 3) in 10-microsecond increments. Similar to the process of step S240 of the band representative waveform generation process shown in FIG. 7, the weighted synchronous addition is a process in which the waveforms of the sounding interval of each Kick sound are weighted by a weighting coefficient set according to a predetermined rule and then synchronously added. Unlike step S240, step S161 of the Kick representative waveform generation process shown in FIG. 12 synchronously adds waveforms of the audio signal of the music piece, including both the body and attack bands (i.e., not a specific frequency band). By performing synchronous addition based on the accurate sounding positions in 10-microsecond increments, a Kick representative waveform can be obtained in which noise from instruments other than the Kick sound is minimized.
[0035] In this embodiment, weighted synchronous addition is performed in the generation of a band representative waveform in the sound generation position detection process in millisecond units (step S142 shown in FIG. 5), the generation of a band representative waveform in the sound generation position detection process in 10 microsecond units (step S152 shown in FIG. 6), and the Kick representative waveform generation process (step S161 shown in FIG. 12). However, since the frequency bands of the audio signals to be synchronously added are different (the body resonance band of the Kick sound in step S142, the attack band of the Kick sound in step S152, and the Kick sound including both the body resonance band and the attack band in step S161), the weighting coefficients used in each weighted synchronous addition process may be the same or different.
[0036] Referring again to FIG. 12, next, a different fade-out curve is applied to each band of the waveform obtained by the weighted synchronous addition in step S161 (step S162). Even after synchronous addition of the waveforms, the resulting waveform does not consist entirely of the kick sound; rather, a small amount of other parts, such as the snare and vocals, remain. Therefore, a filter is applied to remove the remaining snare and vocals and extract the kick sound waveform, such as that shown in FIG. 4. As can be seen from FIG. 4, the kick sound is composed of high frequencies in the attack portion (ATTACK) at the beginning of the sound and low frequencies in the body resonance portion (SUSTAIN) that follows, which is the majority of the period. Because the snare and vocals are mid-range, extracting most of the low frequencies of the kick sound can be achieved by using an LPF with a cutoff frequency at that boundary (around 200 Hz). However, if the entire kick sound waveform were filtered as is, the high frequencies in the attack portion would be lost, as mentioned above. To avoid this, the attack section is filtered through a high-pass filter (for example, a high-pass filter above 640 Hz), while the body section is filtered through a low-pass filter (for example, a low-pass filter below 200 Hz). The section between these is filtered through a band-pass filter between 200 Hz and 640 Hz. Finally, the outputs of these filters are combined to produce the desired kick waveform.
[0037] More specifically, as shown in Fig. 13, the kick waveform obtained by weighted synchronous addition is divided into three band filters, and a different fade-out curve is applied to each band. Specifically, for example, a short (e.g., 20 millisecond) fade-out curve for the attack portion is applied to the signal that passes through a high-pass filter of 640 Hz or higher, a medium-length (e.g., 60 millisecond) fade-out curve for the beginning of the body resonance is applied to the signal that passes through a band-pass filter of 200 Hz to 640 Hz, and a long (e.g., 100 millisecond to 500 millisecond) fade-out curve for the entire body resonance portion is applied to the signal that passes through a low-pass filter of 200 Hz or lower. Note that the length of the fade-out curve for the entire body resonance portion may be adjusted according to the amplitude envelope of the waveform after synchronous addition.
[0038] By performing the above-described Kick representative waveform generation process, it is possible to generate a Kick representative waveform that minimizes noise from other instrument sounds and extracts waveform characteristics common to the Kick sound. In this embodiment, synchronous addition of waveforms is used to separate the Kick sound from other sounds on the time axis, so problems such as deterioration in sound quality due to separation on the frequency axis do not occur, and a high-quality Kick representative waveform can be generated.
[0039] Fig. 14 is a flowchart showing the kick-sound-removed audio generation process shown in Fig. 3. In the kick-sound-removed audio generation process, the kick representative waveform generated by the processes shown in Figs. 12 and 13 is rearranged (step S171) based on the onset positions detected in 10-microsecond increments in the onset position detection process (step S150 shown in Fig. 3) in 10-microsecond increments. This results in an audio signal from which only the kick sound contained in the music piece has been extracted. By adding a reverse-phase signal of this audio signal (kick-sound-rearranged audio signal) to the original music piece audio signal (step S172), it is possible to generate an audio signal of kick-sound-removed audio from which only the kick sound has been removed from the music piece audio signal.
[0040] Fig. 15 is a diagram for conceptually explaining the process of adding the reverse phase signal shown in Fig. 14. The process of step S172 shown in Fig. 14 adds the rearranged kick sound S to the audio signal S of the music. Kick negative phase signal S Kick_Rev Add the rearranged kick sound S Kick corresponds to the kick sound component contained in the audio signal S of the music, then the inverted signal S Kick_Rev By adding the above, the kick sound is cancelled out, and an audio signal containing only sounds other than the kick sound is obtained. At this time, in order to accurately remove only the kick sound component, it is necessary to accurately identify the onset position and waveform of the kick sound. In this embodiment, as described above, the onset position is detected with high accuracy by the onset position detection process in 10-microsecond increments (step S150 shown in FIG. 3), and an accurate kick representative waveform is identified by the synchronous addition process (step S160 shown in FIG. 3) using this accurate onset position. Therefore, it is possible to accurately remove only the kick sound component and generate high-quality kick-free audio in which the sound quality of sounds other than the kick sound is not deteriorated.
[0041] In the embodiment of the present invention described above, the onset position of the kick sound contained in the audio signal of the music piece is identified as the peak position of the cross-correlation function with the band representative waveform. By accurately identifying the onset position of the kick sound in this way, it becomes possible to separate the kick sound from other sounds on the time axis, thereby achieving high sound quality and high separation performance.
[0042] Note that the embodiment of the present invention described above is merely illustrative, and various modifications are possible. For example, in the above embodiment, the audio analysis unit 120 executes a process of extracting a band representative waveform from a band audio signal. However, this process is an example of a process of extracting a representative waveform of a second part from an audio signal of a musical piece. In other embodiments, a representative waveform of a second part may be extracted from an audio signal of a musical piece from which a specific frequency band has not been extracted, and the onset position of the second part may be detected based on a cross-correlation function between the representative waveform and a section of the audio signal of the musical piece. Similarly, the process in which the audio analysis unit 120 generates a band representative waveform by synchronously adding the waveforms of band audio signals is an example of a process of generating a representative waveform of a second part by synchronously adding the audio signal of a musical piece.
[0043] Furthermore, for example, in the above embodiment, the first part of the song is described as a part other than the kick sound, and the second part is described as a kick sound part, but there are no limitations on how the vocal and / or instrument sounds are separated and assigned to the first and second parts. The second part may be any part from which a representative waveform can be extracted, for example, a hi-hat or snare part, or a percussion sound part such as a drum sound that combines a kick sound with a hi-hat or snare. Because it is possible to extract multiple unit sounds with different audio waveform characteristics as described above, the second part may be a drum sound part, and the kick unit sound, as well as the hi-hat and snare unit sounds, may be rearranged.
[0044] Furthermore, for example, in the above embodiment, the detection result of the kick sound generation position by the audio analysis unit 120 and the extraction result of the kick unit sound are used to generate mixed audio data 170 in which the generation position of the kick sound included in the original music piece audio data 110 has been changed. However, in other embodiments, mixed audio data 170 does not necessarily have to be generated. For example, kick unit sound data 132 may be extracted alone and used as a sample sound source for performance. Alternatively, only the kick sound-removed audio data 131 may be output without outputting the kick unit sound data, or the kick sound of a music piece may be replaced with a kick sound of another music piece or a sample sound source based on the kick sound-removed audio data 131 and the kick sound generation data 133. The same applies to the case where the second part described above is a sound other than a kick sound. [Explanation of symbols]
[0045] 10...system, 100...PC, 101...display, 110...music audio data, 120...audio analysis unit, 131...kick sound removed audio data, 132...kick unit sound data, 133...kick pronunciation data, 140...display unit, 150...mix processing unit, 160...operation unit, 170...mixed audio data, 200...DJ controller, 300...speaker
Claims
1. a voice analysis unit that detects a pronunciation position of a second part in a piece of music that includes a first part and a second part that are phonetically separable; The audio analysis unit calculates a cross-correlation function between a representative waveform of the second part extracted from the audio signal of the song and a waveform of the audio signal of the song in a section of length corresponding to the representative waveform as a function of time, and detects the section in which a peak of the cross-correlation function appears as the pronunciation position.
2. the audio analysis unit calculates the cross-correlation function within a predetermined search range; a first sound production position detection process for detecting a first sound production position within a first search range using a first representative waveform; a second sound production position detection process for detecting a second sound production position in a second search range that is set based on the first sound production position and is smaller than the first search range, using a second representative waveform; The audio signal processing apparatus according to claim 1 , wherein the audio signal processing apparatus executes the following:
3. The voice analysis unit In the first sound generation position detection process, a cross-correlation function is calculated between a first band representative waveform extracted from a first band sound signal obtained by extracting a first frequency band from the sound signal of the music piece and a section of the first band sound signal; 3. The audio signal processing device according to claim 2, wherein the second onset position detection process calculates a cross-correlation function between a second band representative waveform extracted from a second band audio signal obtained by extracting a second frequency band from the audio signal of the music piece and a section of the second band audio signal.
4. the second part is composed of a kick sound, the first frequency band is a body resonance band of the kick sound, The audio signal processing device according to claim 3 , wherein the second frequency band is an attack band of the kick sound.
5. The audio signal processing device according to claim 1 , wherein the representative waveform is generated by synchronously adding waveforms of the sound interval of the second part in the audio signal of the music piece.
6. 6. The audio signal processing device according to claim 5, wherein the audio analysis unit calculates a provisional cross-correlation function as a function of time between a provisional representative waveform extracted from the audio signal of the music piece according to a predetermined rule and a waveform of the audio signal of the music piece in a section of a length corresponding to the provisional representative waveform, and generates the representative waveform by synchronously adding the waveform of the audio signal of the music piece for a section in which a peak of the provisional cross-correlation function appears.
7. a voice analysis step of detecting a pronunciation position of a second part in a piece of music including a first part and a second part that are phonetically separable, The audio signal processing method includes a step of calculating a cross-correlation function of a representative waveform of the second part extracted from the audio signal of the music piece and a waveform of the audio signal of the music piece in a section of a length corresponding to the representative waveform, as a function of time, and a step of detecting the section in which a peak of the cross-correlation function appears as the pronunciation position.
8. a voice analysis unit that detects a pronunciation position of a second part in a piece of music that includes a first part and a second part that are phonetically separable; The audio analysis unit calculates a cross-correlation function between a representative waveform of the second part extracted from the audio signal of the song and a waveform of the audio signal of the song in a section of length corresponding to the representative waveform as a function of time, and detects the section in which a peak of the cross-correlation function appears as the pronunciation position. This is a program for causing a computer to function as an audio signal processing device.
Citation Information
Patent Citations
Visual system for distinguishing contact part
JP1987063383A
Method of removing noise, noise removing system and program
JP2003022100A
Information processing system and program
JP2013076887A
Song analysis device and song analysis program
WO2019053766A1