Rhythm detection and speed estimation method based on rhythm state space diagram
By using a rhythm state space map-based method in complex music environments, combined with the Hidden Markov model, the difficulties of beat detection and velocity estimation are solved, achieving higher detection accuracy and stability, especially suitable for styles such as classical music and jazz.
Patent Information
- Application Number
- CN202510159726.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2045-02-13
AI Technical Summary
In complex musical environments, especially in classical music and jazz, beat detection and velocity estimation are difficult, and the prior art is difficult to accurately track changes in local performance speeds, and ignores the implicit correlation of velocity changes.
Using a method based on the rhythm state space diagram, the spectral characteristics of the audio data are obtained, the note intensity feature function and velocity spectrum are extracted, and the two-dimensional rhythm state space diagram is established. The state variables are estimated using the Hidden Markov model to obtain the optimal estimation sequence of music beats and velocities.
It improves the beat detection accuracy in complex music environments, can provide more stable detection results in environments with dynamic changes and noise complexity, and is suitable for music styles with complex rhythm changes.
Smart Images

Figure CN119993101A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a beat detection and speed estimation method based on a rhythm state space graph. Background Art
[0002] At present, music beat and speed detection in the field of music are widely used in music recommendation, audio editing, music creation and musicology analysis. Through software, functions such as classification retrieval, style analysis, track alignment and automatic arrangement can be realized, which can improve the music user experience. Beat is the rhythm position that appears periodically in music and can be clearly perceived by the human ear; speed is the frequency of beats, usually marked in beats per minute (BPM), indicating the speed of music performance. Although beat detection and speed estimation are relatively easy to implement in simple music environments, related technologies face greater challenges in complex music environments such as classical music or jazz. These styles of music often lack obvious percussion instruments, and the beat is difficult to identify. In addition, the natural fluctuations of human performances and the dynamic changes of music rhythm (acceleration or deceleration) increase the difficulty of beat detection and speed estimation. In addition, the presence of rests and syncopation may cause the notes to shift or disappear. The mainstream processing methods mainly rely on the note intensity characteristics, and often treat speed estimation as a subsequent task. In the related art, after obtaining the note intensity characteristic function from the signal processing module or the neural network, it is assumed that the speed of the whole song is stable, and the peak picking algorithm is used to select the significant position in the characteristic function as the beat, and the speed is calculated by counting all the beat intervals. The disadvantage of this type of post-processing method is that it is necessary to assume a correct rhythm change interval in advance, and it cannot be accurately tracked when the local performance speed jumps or drifts significantly. There are also some technologies that treat speed estimation as a separate problem, use a neural network to predict the local speed at the current moment, obtain the speed spectrum of the probability distribution of the local speed that changes with time, and extract the local speed through the local maximum value. The disadvantage of this type of method is that it ignores the implicit correlation of speed changes, is prone to unstable rhythm, and cannot obtain the beat position at the same time. Therefore, it is necessary to fuse the note intensity characteristics and the music speed spectrum to provide a joint post-processing method of beat and speed. Summary of the invention
[0003] In order to overcome the above-mentioned shortcomings and deficiencies of the prior art, an object of the present invention is to provide a beat detection and speed estimation method based on a rhythm state space graph.
[0004] The purpose of the present invention is achieved through the following technical solutions:
[0005] A beat detection and speed estimation method based on a rhythm state space graph, comprising:
[0006] Acquire audio data, divide the audio data into frames to obtain multiple data frames and the frame division times corresponding to the data frames, and extract the frequency spectrum features of the audio data;
[0007] Acquire a note intensity characteristic function from the frequency spectrum characteristics;
[0008] Acquire the tempo spectrum of the music from the frequency spectrum features;
[0009] By modeling the joint probability of the note intensity characteristic function and the speed spectrum, a two-dimensional rhythm state space diagram is obtained, wherein the state variables of the two-dimensional rhythm state space diagram are the music beat position and the music speed;
[0010] The hidden Markov model is introduced to estimate the state variables in the two-dimensional rhythm state space, and the optimal estimation sequence of music beat and music speed is obtained to describe the music rhythm.
[0011] Furthermore, the audio data is a monophonic audio sampling point sequence, and a short-time Fourier transform is performed on the data frame to obtain a frequency spectrum feature.
[0012] Furthermore, the note intensity characteristic function is obtained by specifically using a signal processing method or a neural network model to predict the probability of music beats to obtain the note intensity characteristic function, wherein the note intensity characteristic function is specifically a one-dimensional vector O with a length of N, and each element in the one-dimensional vector corresponds to a probability value that a frame of data is a beat.
[0013] Furthermore, the frequency spectrum feature obtains the tempo spectrum of the music, specifically using signal processing or a neural network model to predict the probability distribution of the tempo at all the frame moments.
[0014] Furthermore, the two-dimensional rhythm state space diagram is a fusion of the beat probability and speed probability distribution of the music, specifically, a matrix of size N×M is obtained by multiplying the two probability distributions.
[0015] Furthermore, the hidden Markov model is introduced to estimate the state variables in the two-dimensional rhythm state space diagram, and the optimal estimation sequence of music beat and music speed is obtained to describe the music rhythm, specifically:
[0016] The hidden Markov model is introduced to describe the relationship between state variables and observations, and the problem of maximizing joint probability distribution is transformed into a recursive Bayesian estimation problem.
[0017] Based on the dynamic model of rhythm, assumptions are made on the prior distribution of state variables, and observation information is obtained from the two-dimensional rhythm state space based on the prior. The maximum a posteriori estimation of the state variables is sequentially solved under the Bayesian framework to obtain the optimal state estimation sequence, and further the optimal estimation sequence of music beat and music speed is obtained.
[0018] Further, specifically:
[0019] The state variables include random variables τ and v, where τ represents the beat time and v represents the speed;
[0020] In the existing observation y 1:K Find a state variable sequence x under the condition 1:K , so that the joint distribution p(x 1:K ,y 1:K ) obtains the maximum value, where K is the number of state variables and K is a positive integer;
[0021] According to the first-order Markov property assumption and the observation independence assumption, the state at time k depends only on the state at time k-1, and the observation y k Only depends on the state x at time k k , the joint distribution can be factorized into:
[0022]
[0023] Among them, p(x1) is the prior distribution of the initial state, p(x k |x k-1 ) is the state transition probability, p(y k |x k ) is a two-dimensional rhythm state space diagram;
[0024] The optimal estimation sequence of the beats is solved by maximizing the joint distribution.
[0025] A device based on the beat detection and speed estimation method, comprising:
[0026] Music feature extraction module: used to extract the spectrum features of audio data;
[0027] Beat probability extraction module: used to extract beat detection results and obtain note intensity feature function;
[0028] Speed probability extraction module: used to extract the probability distribution of speed and obtain the speed spectrum;
[0029] Rhythm state space modeling module: used to build a two-dimensional rhythm state space graph;
[0030] State variable estimation module: used to describe the music rhythm by the optimal estimation sequence of music beat and music speed.
[0031] A computer-readable storage medium stores a computer program, wherein the computer program is used to be executed by a processor to implement the method described.
[0032] A computer program product comprises a computer program, wherein the computer program is loaded and executed by a processor to implement the method.
[0033] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0034] (1) By constructing a two-dimensional rhythm state space diagram, the note intensity characteristic function and the velocity spectrum are explicitly combined to achieve unified modeling, realize the synchronous optimization of beat position and velocity, and improve the detection accuracy in complex music environments.
[0035] (2) The introduction of a Bayesian method to jointly track beat and tempo based on maximum a posteriori estimation can provide more stable detection results in dynamically changing and noisy music environments.
[0036] (3) This method is designed for styles with complex rhythmic changes, such as classical music and jazz. It can cope with challenges such as natural fluctuations, tempo drift, and lack of percussion instruments. It is particularly suitable for music styles with irregular rhythms and large tempo changes. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a schematic diagram of the implementation environment of the method of the present invention;
[0038] Figure 2 It is a structural schematic diagram of the method of the present invention. DETAILED DESCRIPTION
[0039] The present invention will be further described in detail below in conjunction with examples, but the embodiments of the present invention are not limited thereto.
[0040] like Figure 1 As shown, an embodiment of the present invention provides a solution implementation environment based on a computer device 10, which includes electronic devices capable of data calculation, processing and storage. The computer device 10 can be a terminal device (such as an edge device, a mobile phone, a computer, an intelligent voice interaction device, a smart home appliance, a vehicle terminal, an aircraft, a game console, a wearable device, a multimedia playback device, an augmented reality device, a virtual reality device, etc.) or a server (such as an independent physical server, a server cluster or a distributed system, and a cloud server that provides basic cloud computing services such as edge computing, cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content distribution networks, big data and artificial intelligence platforms);
[0041] By executing the method provided in this embodiment, the computer device 10 can jointly analyze the beat and speed detection results, and can stably track in a dynamically changing and noisy music environment. It is particularly suitable for styles with complex rhythm changes such as classical music and jazz. The embodiments of the present application can be widely used in scenarios such as music creation, music editing, virtual instrument performance, and online singing platforms.
[0042] like Figure 2 As shown, this embodiment provides a beat detection and speed estimation device based on a rhythm state space diagram, which is suitable for classical music, jazz and other styles with complex rhythm changes. The device has the function of implementing the above method example, and the function can be implemented by hardware or by hardware executing corresponding software. The device can be the computer device 10 introduced above, or it can be set in the computer device 10. Figure 2 As shown, the apparatus 200 may include an audio feature extraction module 210 , a beat probability extraction module 220 , a speed probability extraction module 230 , a rhythm state space modeling module 240 and a state variable estimation module 250 .
[0043] include:
[0044] The audio feature extraction module 210 is used to extract the spectral features of the audio, wherein the audio data is input into the system in mono with a sampling rate of 44.1kHz. If it is multi-channel audio data, all channels are averaged to obtain mono data; if it is non-44.1kHz audio data, 44.1kHz is obtained by up / down sampling. Assuming that the audio data is framed with a window length of 882 and a step length of 441 to obtain a series of data frames, then the interval of each data frame is 10ms, and the corresponding time of each frame is t=10, 20, 30, ... ms. A Hanning window is added to the data frame, and a short-time Fourier transform is performed to obtain a time-frequency spectrum. Optionally, the time-frequency spectrum is mapped to a Mel scale to obtain a Mel inverse spectrum.
[0045] The beat probability extraction module 220 is used to extract the beat detection result and obtain the note intensity characteristic function, which can predict the beat probability of music through a signal processing method or a neural network model. Optionally, the above-mentioned music beat extraction result can be a single vector. For example, if N is the number of frame data, N is a positive integer, the prediction result is a one-dimensional vector O with a length of N, and each element in the one-dimensional vector corresponds to a probability value that a frame data is a beat. In some embodiments, the beat probability detection process can be obtained by performing time domain difference summation, half-wave rectification, and normalization on the spectrum. Beat probability detection can also be implemented through one of the following neural network models: recurrent neural network, convolutional neural network, long short-term memory recurrent neural network, etc., or other models that can predict the probability value of each frame data in music being a beat.
[0046] The speed probability extraction module 230 is used to extract the probability distribution of speed and obtain the speed spectrum. The probability distribution of speed at all the frame moments can be predicted by signal processing or a neural network model. For example, assuming that the number of all possible discrete speed values is M, M is a positive integer, the prediction result is a matrix R (n, m) of size N×M, which represents the probability that the speed value is m at the nth moment in the discrete space. Optionally, the speed probability extraction module can be implemented using an autocorrelation method, wavelet transform, template similarity matching, and convolution, or can be implemented using one of the following neural network models: a recurrent neural network, a convolutional neural network, a long short-term memory recurrent neural network, or other models that can predict the probability distribution of speed for each frame of music.
[0047] The rhythm state space modeling module 240 is used to establish a two-dimensional rhythm state space diagram. Based on the beat probability vector O of the music and the probability distribution matrix R of the speed, the joint probability distribution of the beat and the speed is modeled as a two-dimensional rhythm state space diagram Tg, which is a matrix of size N×M, wherein the state variables include the beat position and the speed. The two-dimensional rhythm state space diagram is a fusion of the probability distribution of the beat probability and the speed of the music, which can be expressed in mathematical form as follows:
[0048] T g (n,m)=O(n)·R(n,m)
[0049] Where n = 1, 2, ..., N, m = 1, 2, ..., M. n and m correspond to the beat position and speed respectively. This method describes the joint characteristics of the rhythm by taking the product of two probability distributions as the value of the state space graph.
[0050] It is further explained that the fusion method of the beat probability vector O and the speed probability distribution matrix R is not limited to multiplication, but can also be fused using other mathematical methods such as linear combination, logical function, power function or polynomial function to more flexibly capture the relationship between rhythm and speed.
[0051] It is further explained that in order to reduce the impact of outliers on the calculation results, smoothing techniques such as Gaussian smoothing, exponential smoothing or Laplace smoothing can be applied to the obtained state space diagram, thereby improving the robustness and accuracy of the model.
[0052] It is further illustrated that the discretization of the rhythm state space graph can be combined with the multi-resolution modeling approach to describe the multi-level characteristics of the rhythm by constructing state graphs at different time and speed scales respectively.
[0053] The state variable estimation module 250 is used to provide an optimal estimate of the state variable and use a Hidden Markov Model (HMM) to describe the music rhythm.
[0054] Specifically, the state variable x contains random variables τ and v, where τ represents the beat time and v represents the speed, which can be represented by the time interval between adjacent beats or by the number of beats per minute (BPM), without any restriction.
[0055] In order to find the best prediction of beat and speed, given the observation y 1:K Find a state sequence x under the condition 1:K , so that the joint distribution p(x 1:K ,y 1:K ) obtains the maximum value, where K is the number of state variables and K is a positive integer. According to the first-order Markov property assumption and the observation independence assumption, the state at time k only depends on the state at time k-1. The observation y k Only depends on the state x at time k k , the joint distribution can be factorized into:
[0056]
[0057] Where p(x1) is the prior distribution of the initial state, p(x k |x k-1 ) is the state transition probability, which is given by the system dynamic model; p(y k |x k ) is the observation model, and the state space graph T g Given.
[0058] By maximizing the joint distribution p(x 1:K ,y 1:K ) to solve the optimal estimate of the beat, which is converted into a recursive solution of the maximum posterior probability p(x k |y k ,x k-1 ), according to the Bayesian formula, it is equivalent to maximizing the product of the transition probability and the observation model in steps p(y k |x k )p(x k |x k-1 ), the derivation process is as follows:
[0059]
[0060] Among them, p(x1) is the prior distribution of the initial state, p(x k |x k-1 ) is the state transition probability, which is given by the system dynamic model; p(y k|x k ) is the observation model, and the state space graph T g Given.
[0061] Due to the nature of the model, the observation y 1:K In fact, it is implicit in the state sequence x 1:K In the above example, we can gradually generate y by considering the observation process as a state-driven generation process. 1:K :First, initialize an initial state x1 and generate the initial observation y1 using the model p(y1|x1). After that, for subsequent time points k=2,3,...,K, according to the state transition model p(x k ∣x k-1 )Predict the next state Then according to the observation model p(y k ∣x k ) and predicted status Generate observation y k .
[0062] Further explanation: the observation model p(y k |x k ) is represented by the two-dimensional rhythm state space graph T g express:
[0063] p(y k |x k )=T g (x k )
[0064] Among them, x k It is represented by the beat time n and speed m in the discretized state space, where n∈[1,N],m∈[1,M]. In order to reduce the amount of calculation and avoid the influence of some abnormal observations, the possible observation range χ is limited. k And renormalize the conditional probability p(y k ∣x k ), realizes the generation mechanism combining geometric constraints and statistical probability, which is very effective in the application scenario of music rhythm analysis with strong spatiotemporal correlation. A possible embodiment is to Inside,
[0065]
[0066] Among them, the normalization factor To T g Normalize to ensure that k The sum of the probabilities generated within is 1, and r is the window radius, which defines the range of allowed observations.
[0067] To further illustrate, assume that the state variable satisfies the Gaussian distribution:
[0068]
[0069] In the framework of the linear Gaussian model, the transition probability of the current state is obtained by the recursive relationship:
[0070]
[0071] μ k =Fμ k-1 ,Σ k =FΣ k-1 F T +Q
[0072] Among them, F is the state transfer matrix, which describes the linear dynamic characteristics of the system, Q is the process noise covariance matrix, which represents the uncertainty introduced during state transfer, and μ k-1 ,Σ k-1 is the mean and covariance of the state at the previous time step. In the relevant embodiment, the state transfer matrix F, the process noise covariance matrix Q, and the mean and covariance μ0,Σ0 of the initial state are not agreed upon. After the above decoding process, the optimal estimated state sequence can be obtained Then we can get the optimal estimation sequence of beat and speed
[0073] Further explanation: the state variables can be extended to include not only the existing beat time τ and velocity v, but also their higher-order derivatives, such as acceleration This enables the model to capture the dynamics of rhythm changes more accurately. At this point, the hidden Markov model can be used to describe the expanded state variables, where the state transition probabilities include τ, v, and The relationship between.
[0074] Further explanation: the model structure is not limited to Markov chains, and other types of probabilistic graphical models such as Bayesian networks can be used. These models represent the causal relationship between state variables through directed acyclic graphs (DAGs) or use dynamic Bayesian networks (DBNs) to process time series data, thereby achieving more flexible state prediction and estimation.
[0075] Further explanation: the dynamic transfer relationship of state variables can be described by linear or nonlinear models. For example, the state transfer equation can adopt a first-order linear dynamic equation, or introduce a nonlinear dynamic equation to model complex rhythm changes, or even a high-order equation to capture more detailed change patterns.
[0076] Further explanation: the optimization objective can be selected as the maximum likelihood estimation (MLE) to maximize the conditional probability of the observed data; or other objective functions, such as minimizing the observation error, can be used to optimize the state estimation process. In addition, the optimization method under the non-Bayesian framework can also be used as an alternative to meet the needs of different application scenarios.
[0077] This embodiment also provides a beat detection and speed estimation method based on a rhythm state space graph, including:
[0078] Acquire audio data, divide the audio data into frames to obtain multiple data frames and the frame division times corresponding to the data frames, and extract the frequency spectrum features of the audio data;
[0079] Acquire a note intensity characteristic function from the frequency spectrum characteristics;
[0080] Acquire the tempo spectrum of the music from the frequency spectrum features;
[0081] By modeling the joint probability of the note intensity characteristic function and the speed spectrum, a two-dimensional rhythm state space diagram is obtained, wherein the state variables of the two-dimensional rhythm state space diagram are the music beat position and the music speed;
[0082] The hidden Markov model is introduced to estimate the state variables in the two-dimensional rhythm state space, and the optimal estimation sequence of music beat and music speed is obtained to describe the music rhythm, specifically:
[0083] In the process of beat detection and speed estimation, a two-dimensional rhythm state space diagram is established and the state variables are recursively estimated in combination with the dynamic transfer equation, wherein the recursive estimation combines the joint probability distribution of the music beat position and the speed;
[0084] Using dynamic modeling technology, the transfer equation can be a linear or nonlinear model, and can be modeled using dynamic relationships of first-order or higher-order derivatives;
[0085] In the state estimation process, the state space diagram is smoothed to eliminate outliers, wherein the smoothing process includes but is not limited to Gaussian smoothing, exponential smoothing or Laplace smoothing;
[0086] By optimizing the objective function, such as maximum a posteriori estimation, maximum likelihood estimation or other error minimization methods, the beat and speed state sequences are decoded and the optimal beat detection results and speed estimation values are output.
[0087] This embodiment further provides a computer-readable storage medium, in which a computer program is stored. The computer program is used to be executed by a processor to implement a beat detection and speed estimation method.
[0088] This embodiment also provides a computer program product, which includes a computer program. The computer program is loaded and executed by a processor to implement the beat detection and speed estimation method.
[0089] The above embodiments are preferred implementation modes of the present invention, but the implementation modes of the present invention are not limited to the embodiments. Any other changes, modifications, substitutions, combinations, and simplifications that do not deviate from the spirit and principles of the present invention should be equivalent replacement methods and are included in the protection scope of the present invention.
Claims
1. A beat detection and speed estimation method based on a rhythm state space graph, characterized in that: include: Acquire audio data, divide the audio data into frames to obtain multiple data frames and the frame division times corresponding to the data frames, and extract the frequency spectrum features of the audio data; Acquire a note intensity characteristic function from the frequency spectrum characteristics; Acquire the tempo spectrum of the music from the frequency spectrum features; By modeling the joint probability of the note intensity characteristic function and the speed spectrum, a two-dimensional rhythm state space diagram is obtained, wherein the state variables of the two-dimensional rhythm state space diagram are the music beat position and the music speed; The hidden Markov model is introduced to estimate the state variables in the two-dimensional rhythm state space, and the optimal estimation sequence of music beat and music speed is obtained to describe the music rhythm.
2. The beat detection and speed estimation method according to claim 1, characterized in that: The audio data is a monophonic audio sampling point sequence, and a short-time Fourier transform is performed on the data frame to obtain a frequency spectrum feature.
3. The beat detection and speed estimation method according to claim 1, characterized in that: The note intensity characteristic function is obtained by using a signal processing method or a neural network model to predict the probability of music beats to obtain the note intensity characteristic function. The note intensity characteristic function is specifically a one-dimensional vector with a length of N, and each element in the one-dimensional vector corresponds to a probability value that a frame of data is a beat.
4. The beat detection and speed estimation method according to claim 1, characterized in that: The frequency spectrum feature obtains the speed spectrum of the music, and specifically uses signal processing or a neural network model to predict the probability distribution of the speed at all the frame moments.
5. The beat detection and speed estimation method according to claim 1, characterized in that: The two-dimensional rhythm state space diagram is a fusion of the beat probability and speed probability distribution of the music. Specifically, the two probability distributions are multiplied to obtain a matrix of size N×M.
6. The beat detection and speed estimation method according to any one of claims 1 to 5, characterized in that: The hidden Markov model is introduced to estimate the state variables in the two-dimensional rhythm state space diagram, and the optimal estimation sequence of music beat and music speed is obtained to describe the music rhythm, specifically: The hidden Markov model is introduced to describe the relationship between state variables and observations, and the problem of maximizing joint probability distribution is transformed into a recursive Bayesian estimation problem. Based on the dynamic model of rhythm, assumptions are made on the prior distribution of state variables, and observation information is obtained from the two-dimensional rhythm state space based on the prior. The maximum a posteriori estimation of the state variables is sequentially solved under the Bayesian framework to obtain the optimal state estimation sequence, and further the optimal estimation sequence of music beat and music speed is obtained.
7. The beat detection and speed estimation method according to claim 6, characterized in that: Specifically: The state variables include random variables τ and v, where τ represents the beat time and v represents the speed; In the existing observation y 1:K Find a state variable sequence x under the condition 1:K , so that the joint distribution p(x 1:K ,y 1:K ) obtains the maximum value, where K is the number of state variables and K is a positive integer; According to the first-order Markov property assumption and the observation independence assumption, the state at time k depends only on the state at time k-1, and the observation y k Only depends on the state x at time k k , the joint distribution can be factorized into: Among them, p(x1) is the prior distribution of the initial state, p(x k |x k-1 ) is the state transition probability, p(y k |x k ) is a two-dimensional rhythm state space diagram; The optimal estimation sequence of the beats is solved by maximizing the joint distribution.
8. A device based on the beat detection and speed estimation method according to any one of claims 1 to 7, characterized in that: include: Music feature extraction module: used to extract the spectrum features of audio data; Beat probability extraction module: used to extract beat detection results and obtain note intensity feature function; Speed probability extraction module: used to extract the probability distribution of speed and obtain the speed spectrum; Rhythm state space modeling module: used to build a two-dimensional rhythm state space graph; State variable estimation module: used to describe the music rhythm by the optimal estimation sequence of music beat and music speed.
9. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, and the computer program is used to be executed by a processor to implement the method according to any one of claims 1 to 7.
10. A computer program product, characterized in that The computer program product comprises a computer program, which is loaded and executed by a processor to implement the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Post-processing method, device and equipment for music beat detection and storage medium
CN117894286A
Detecting Beat Information Using a Diverse Set of Correlations
US20100300271A1