A method for detecting the tempo of music based on a neural network

Through the music beat speed detection method based on neural network, the problem of insufficient accuracy of music beat speed detection in the prior art is solved, and higher accuracy and computing efficiency are achieved.

CN114882905BActive Publication Date: 2025-06-20KUNMING UNIV OF SCI & TECH
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202210374604.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2025-06-20
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

The prior art has insufficient accuracy in music beat speed detection, making it difficult to effectively distinguish different music types and affect the accurate calculation of beat speed.

Method used

The music beat speed detection method based on neural network is used, and the beat speed detection method is detected by detecting music type, signal filtering, frame acquisition envelope, differential processing and moving average processing, and the training data is generated and input into the neural network for training. The final test is obtained.

Benefits of technology

It greatly improves the accuracy of music beat estimation, simplifies the calculation process, and improves the calculation speed and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114882905B_ABST
    Figure CN114882905B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for detecting the tempo of music based on a neural network, belonging to the technical field of audio signal processing. The present invention determines whether it is instrumental music or vocal music according to the spectrogram of the music signal; according to the above judgment result, if it is instrumental music, high-pass filtering is performed, and if it is vocal music, low-pass filtering is performed; after filtering, the signal is framed, and then the maximum value of each frame is taken to synthesize an envelope; first-order difference and second-order difference are performed on the envelope; multiple moving average processes are performed on the difference results; after the moving average process is completed, it is input into the neural network for training, and finally the result of the music beat value is obtained through testing. Most of the algorithms involved in the present invention are performed in the time domain, and a small part involves the frequency domain. Compared with the method of calculating the tempo in the pure frequency domain, this method is simpler and more convenient, and has higher calculation speed and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for detecting the tempo of music based on a neural network, belonging to the technical field of audio signal processing. Background Art

[0002] BPM is the number of beats per minute, and the magnitude of the value represents the speed. It is an important part of a piece of music. The emotional tones of music played or performed with different BPMs are different: slower speeds are mostly lyrical and narrative song types, moderate speeds are mostly cheerful and relaxing song types, and faster speeds are mostly urgent and tense music types. The function of the method for detecting the tempo of music based on a neural network is to accurately calculate the tempo of different music pieces. After obtaining the tempo of the music, further research on music rhythm analysis, music beat tracking, and music genre classification can be carried out.

[0003] The prior art related to the present application is the patent document CN114005464A, which discloses a method, device, computer device, and storage medium for estimating the tempo. The method includes: extracting audio features from the current music; performing autocorrelation processing on the audio features; listing multiple possible options for the number of beats per minute for the current music; generating a characteristic beat array for each possible option of the number of beats per minute; performing cross-correlation processing on the autocorrelation-processed audio features and each characteristic beat array; and based on the cross-correlation processing results, selecting a cross-correlation function with a dynamic range meeting a preset threshold as the estimated tempo result of the current music. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to provide a method for detecting the tempo of music based on a neural network, which can greatly improve the accuracy of music tempo estimation, thereby solving the above problems.

[0005] The technical solution of the present invention is: a method for detecting the tempo of music based on a neural network, and the specific steps are as follows:

[0006] Step1: Detect the music genre, and determine whether it is instrumental music or vocal music according to the spectrogram of the music signal.

[0007] The sequence x[n] represents a one-dimensional music signal. First, perform a Fourier transform on the signal to obtain the amplitude spectrum F(n). Visualize the amplitude spectrum F(n) and detect the group of impulse lines near (0 - 300Hz). If there is an obvious interval between the impulse lines, it is determined that the music genre is instrumental solo. If the interval is not obvious and there are other continuous spectral components attached, it is determined that the music genre is vocal music.

[0008] Step 2: Perform signal filtering. If it is instrumental music, perform high-pass filtering; if it is vocal music, perform low-pass filtering.

[0009] According to the above judgment results, if it is instrumental music, perform high-pass filtering. The cut-off frequency of the high-pass filter is 2400 Hz. If it is vocal music, perform low-pass filtering. The cut-off frequency of the low-pass filter is 1600 Hz.

[0010] Step 3: After filtering, perform signal framing, and then take the maximum value of each frame to synthesize the envelope.

[0011] Frame the time-domain signal. Set the resampling rate of the music signal to 8000 Hz / s. The frame length of framing is frame_length = 2048 points, the frame shift is frame_shift = 512 points, and the expression for calculating the number of frames num is as follows:

[0012]

[0013] Take the maximum value of each frame to form the envelope. The sequence at this time is set as envelope[num], where num is the number of speech frames. However, for the unity of subsequent data, take num = 1000 frames.

[0014] Step 4: Perform first-order and second-order differences on the envelope.

[0015] Perform a first-order difference on the envelope signal envelope[num] to obtain the envelope_1[num] signal. Each impulse line represents a peak. The first-order difference formula is as follows:

[0016] envelope_1[n] = envelope_1[n + 1] - envelope_1[n], (n = 0, 1, 2,... num - 1)

[0017] Perform a second-order difference on the envelope signal envelope[num] to obtain the envelope_2[num] signal. Each impulse line represents a peak. The second-order difference formula is as follows:

[0018] envelope_2[n] = envelope_2[n + 2] - 2 × envelope_2[n + 1] + envelope_2[n], (n = 0, 1, 2,..., num - 2)

[0019] Step 5: Perform multiple moving average processes on the difference results.

[0020] Perform multiple moving average processes on the first-order and second-order difference data. The expression formula is as follows:

[0021]

[0022] mean_1[num] represents the average value of each frame of the first-order difference data envelope_1[num], where num is the number of speech frames. mean_2[num] represents the average value of each frame of the second-order difference data envelope_2[num], and num is the number of speech frames. The same music beat value is set for the data processed by the moving average, which is used as the label for training.

[0023] Step6 After multiple moving average processes are completed, it is input into the neural network for training, and finally the beat speed result is obtained through testing.

[0024] The data of mean_envelope_1[n] and the corresponding label values are input into the neural network model for training to obtain Model 1. The data of mean_envelope_2[n] and the corresponding label values are input into the neural network model for training to obtain Model 2. The data of mean_envelope_1[n], mean_envelope_2[n] and the corresponding label values are mixed and input into the neural network model for training to obtain Model 3. The model with the best parameter effect is selected for testing to obtain the music beat speed value.

[0025] The beneficial effects of the present invention are as follows: Most of the algorithms involved in the present invention are performed in the time domain, and a small part involves the frequency domain. Compared with the method of calculating the beat speed in the pure frequency domain, the present invention is simpler and more convenient, and has higher calculation speed and accuracy. Brief Description of the Drawings

[0026] Figure 1 is the flow schematic diagram of the present invention;

[0027] Figure 2 is the time-domain waveform diagram of the instrument music in the embodiment of the present invention;

[0028] Figure 3 is the time-domain waveform diagram of the vocal music in the embodiment of the present invention;

[0029] Figure 4 is the frequency-domain waveform diagram of the instrument music in the embodiment of the present invention;

[0030] Figure 5 is the frequency-domain waveform diagram of the vocal music in the embodiment of the present invention;

[0031] Figure 6 is the time-domain waveform diagram of the instrument music after high-pass filtering in the embodiment of the present invention;

[0032] Figure 7 is the time-domain waveform diagram of the vocal music after low-pass filtering in the embodiment of the present invention;

[0033] Figure 8 is the time-domain envelope diagram after filtering the instrument music in the embodiment of the present invention;

[0034] Figure 9 is the time-domain envelope diagram after filtering the vocal music in the embodiment of the present invention;

[0035] Figure 10 is the first-order difference diagram of the envelope of the instrument music in the embodiment of the present invention;

[0036] Figure 11 is the second-order difference diagram of the envelope of the instrument music in the embodiment of the present invention;

[0037] Figure 12 is the first-order difference diagram of the envelope of the vocal music in the embodiment of the present invention;

[0038] Figure 13 is the second-order difference diagram of the envelope of the vocal music in the embodiment of the present invention;

[0039] Figure 14 is the moving average diagram after the first-order difference of the instrument music in the embodiment of the present invention;

[0040] Figure 15 is the moving average diagram after the second-order difference of the instrument music in the embodiment of the present invention;

[0041] Figure 16 is the moving average diagram after the first-order difference of the vocal music in the embodiment of the present invention;

[0042] Figure 17 is the moving average diagram after the second-order difference of the vocal music in the embodiment of the present invention;

[0043] Figure 18 is the error diagram of the training process of the first model of the instrument music in the embodiment of the present invention;

[0044] Figure 19 is the error diagram of the training process of the second model of the instrument music in the embodiment of the present invention;

[0045] Figure 20 is the error diagram of the training process of the third model of the instrument music in the embodiment of the present invention;

[0046] Figure 21 is the error diagram of the training process of the first model of the vocal music in the embodiment of the present invention;

[0047] Figure 22 is the error diagram of the training process of the second model of the vocal music in the embodiment of the present invention;

[0048] Figure 23It is the error graph of the training process of the vocal music model three in the embodiments of the present invention. Detailed implementation manners

[0049] The present invention will be further described below in conjunction with the accompanying drawings and detailed implementation manners.

[0050] Embodiment 1: As Figure 1 shown, a method for detecting the tempo of music based on a neural network, the specific steps are as follows:

[0051] Step1: Detect the music type, and judge whether it is instrumental music or vocal music according to the spectrogram of the music signal;

[0052] Step2: Perform signal filtering. If it is instrumental music, perform high-pass filtering. If it is vocal music, perform low-pass filtering;

[0053] Step3: After filtering, perform signal framing, then take the maximum value of each frame and synthesize the envelope;

[0054] Step4: Perform first-order difference and second-order difference on the envelope;

[0055] Step5: Perform multiple moving average processes on the difference results;

[0056] Step6: After multiple moving average processes are completed, input them into the neural network for training, and finally test to obtain the tempo result.

[0057] Each step will be described in detail below.

[0058] First of all, it is necessary to distinguish whether the music signal is of the instrumental music type or the vocal music type, and visualize the graphs of these two major types of music in the time domain. The duration of each piece of music signal is approximately between 15s and 25s. As Figures 2-3 shown, after visualizing the time-domain waveform diagram, the difference is not obvious, resulting in the inability to distinguish the music type. Therefore, it is necessary to perform a fast Fourier transform to the frequency domain to observe them. As Figures 4-5 shown, the difference between them can be seen at this time. Detect the impulse line group near (0 - 300Hz); if there is an obvious interval between the impulse lines, judge that the music type is instrumental solo; if the interval is not obvious and there are other continuous spectral components attached, judge that the music type is vocal music.

[0059] After confirming the music type, signal filtering is required. The sampling frequency of all music signals in this embodiment is 8000Hz. If it is instrumental music, perform high-pass filtering, and the cut-off frequency of the high-pass filter is 2400Hz; if it is vocal music, perform low-pass filtering, and the cut-off frequency of the low-pass filter is 1600Hz, as Figures 6-7 shown.

[0060] After filtering, it is necessary to extract the envelope of the signal. Here, the method of taking the maximum value for each frame is used for envelope extraction. The time-domain signal is framed. The resampling rate of the music signal is set to 8000 Hz / s. The frame length frame_length of the framing is 2048 points, the frame shift frame_shift is 512 points, and the expression for calculating the number of frames num is as follows:

[0061]

[0062] Take the maximum value of each frame to form the envelope. The sequence at this time is set to envelope[num], where num is the number of speech frames. However, for the unity of subsequent data, the value of num is taken as 1000 frames, as Figures 8-9 shown. The envelope diagram at this time is not the envelope diagram of the entire input signal because some tails need to be removed.

[0063] After envelope extraction, the peaks of the signal are very obvious. However, there are many sub-peaks beside the high peaks, which is not conducive to the extraction of the rhythm speed. At this time, performing first- and second-order differences can make the peaks prominent and weaken the sub-peaks. Perform a first-order difference on the envelope signal envelope[num] to obtain the envelope_1[num] signal. Each impulse line represents a peak. The first-order difference formula is as follows:

[0064] envelope_1[n] = envelope_1[n + 1] - envelope_1[n], (n = 0, 1, 2,... num - 1)

[0065] Perform a second-order difference on the envelope signal envelope[num] to obtain the envelope_2[num] signal. Each impulse line represents a peak. The second-order difference formula is as follows:

[0066] envelope_2[n] = envelope_2[n + 2] - 2 × envelope_2[n + 1] + envelope_2[n], (n = 0, 1, 2,..., num - 2) as Figures 10-13 shown.

[0067] The graph after the first- and second-order differences has negative values. Now it is necessary to remove the negative values and only leave the data on the upper half-axis; at the same time, if you want to highlight the highest peak again and weaken the sub-peaks, using multiple moving averages can solve this problem. Perform multiple moving average processing on the first- and second-order difference data. The expression formula is as follows:

[0068]

[0069] mean_1[num] represents the average value of each frame of the first-order difference data envelope_1[num], where num is the number of speech frames, and mean_2[num] represents the average value of each frame of the second-order difference data envelope_2[num], where num is the number of speech frames; as Figures 14-17 shown.

[0070] The same music beat value is set for the data processed by the moving average, which is used as the label for training. There are three categories of training data: the first category is training with the first-order difference data, the second category is training with the second-order difference data, and the third category is training with a mixture of the first- and second-order difference data. As a result, three different training effect diagrams for each music type are obtained. As Figures 18-23 shown, after testing with the model test set data, the model with the best effect is selected for predicting the beat speed value.

[0071] The specific embodiments of the present invention have been described in detail above in conjunction with the accompanying drawings. However, the present invention is not limited to the above embodiments. Within the scope of knowledge possessed by those of ordinary skill in the art, various changes can be made without departing from the gist of the present invention.

Claims

1. A method for detecting the tempo of music based on a neural network, characterized in that: Step1: Detect the music type and determine whether it is instrumental music or vocal music based on the spectrogram of the music signal; Step2: Perform signal filtering. If it is instrumental music, perform high-pass filtering. If it is vocal music, perform low-pass filtering; Step3: After filtering, perform signal framing, then take the maximum value of each frame and synthesize the envelope; Step4: Perform first-order and second-order differences on the envelope; Step5: Perform multiple moving average processes on the difference results; Step6: After multiple moving average processes are completed, input them into the neural network for training, and finally test to obtain the beat speed result; The specific content of Step5 is as follows: Step5.1: Perform multiple moving average processes on the first-order and second-order difference data. The expression formula is as follows: mean_1[num] represents the average value of each frame of the first-order difference data envelope_1[num], where num is the number of speech frames, mean_2[num] represents the average value of each frame of the second-order difference data envelope_2[num], where num is the number of speech frames; The specific content of Step6 is as follows: Step6.1: Input the mean_envelope_1[n] data and the corresponding label values into the neural network model for training to obtain Model 1; Step6.2: Input the mean_envelope_2[n] data and the corresponding label values into the neural network model for training to obtain Model 2; Step6.3: Input the mean_envelope_1[n], mean_envelope_2[n] data and the corresponding label values into the neural network model for training to obtain Model 3; Step6.4: Select the model with the best parameter effect for testing to obtain the music beat speed.

2. The method for detecting the tempo of music based on a neural network according to claim 1, characterized in that The specific content of Step1 is as follows: Step1.1: The sequence x[n] represents a one-dimensional music signal. Perform Fourier transform on the signal to obtain the magnitude spectrum F(n); Step1.2: Visualize the magnitude spectrum F(n) and detect the impulse line group near 0 - 300Hz; Step1.3: If there is an obvious interval between the impulse lines, determine that the music type is instrumental solo. If the interval is not obvious and there are other continuous spectral components, determine that the music type is vocal music.

3. The method for detecting the tempo of music based on a neural network according to claim 1, characterized in that The specific content of Step2 is as follows: Step2.1: Perform classification filtering on the music signal. Let the filtered signal sequence be x_filter[n]; Step2.2: If it is instrumental music, perform high-pass filtering. The cut-off frequency of the high-pass filter is 2400Hz. If it is vocal music, perform low-pass filtering. The cut-off frequency of the low-pass filter is 1600Hz.

4. The method for detecting the tempo of music based on a neural network according to claim 1, characterized in that The specific content of Step3 is as follows: Step3.1: Frame the time-domain signal. The resampling rate of the music signal is set to 8000Hz / s. The frame length of framing is frame_length = 2048 points, the frame shift is frame_shift = 512 points, and the expression for calculating the number of frames num is as follows: Step 3.2: Take the maximum value of each frame to form an envelope. The sequence at this time is set as envelope[num], where num is the number of speech frames.

5. The method for detecting the tempo of music based on a neural network according to claim 1, characterized in that The specific content of Step 4 is as follows: Step 4.1: Perform a first-order difference on the envelope signal envelope[num] to obtain the envelope_1[num] signal. Each impulse line represents a peak. The first-order difference formula is shown as follows: envelope_1[n] = envelope_1[n + 1] - envelope_1[n], n = 0, 1, 2,... num - 1 Step 4.2: Perform a second-order difference on the envelope signal envelope[num] to obtain the envelope_2[num] signal. Each impulse line represents a peak. The second-order difference formula is shown as follows: envelope_2[n] = envelope_2[n + 2] - 2 × envelope_2[n + 1] + envelope_2[n], n = 0, 1, 2,..., num - 2.

Citation Information

Patent Citations

  • Beat speed estimation method and device thereof, computer equipment and storage medium

    CN114005464A

  • Music beat detection method and system

    CN111508457A

  • Audio beat information detection method and device and storage medium

    CN111508526A