An audio processing method and related device
Through the spectrum graph-based feature extraction and the probability model of the stream Transformer encoder, the problem of accurately capturing beat information in online beat tracking is solved, and efficient beat tracking is achieved in complex backgrounds, improving the accuracy and robustness of the model.
Patent Information
- Application Number
- CN202510421243.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing online beat tracking techniques are difficult to accurately capture key beat information in audio signals, especially in the context of variable music rhythms and complex audio, and traditional loss functions are prone to overfitting the model.
Using a spectrum graph-based feature extraction method, combined with a stream Transformer encoder and probability model, precise tracking of audio beats and strong beats is achieved through feature extraction, beat activation value calculation and probability inference.
Improves the accuracy and robustness of online beat tracking, and can capture key beat information in complex backgrounds, reduces computational complexity, and alleviate the problem of model overfitting.
Smart Images

Figure CN120071875B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of music information retrieval, and particularly relates to an audio processing method and related device. Background Art
[0002] Beat tracking is a key task in the field of Music Information Retrieval (MIR). Its main purpose is to analyze the beat and downbeat sequences in an audio stream to express the rhythm of music. Good beat estimation is beneficial to various downstream tasks of MIR, including music transcription, structure analysis, etc. Moreover, since human perception of music rhythm is related to motor sensitivity, beat tracking can also be applied to more scenarios such as human-computer interaction and music therapy. In addition, tracking and analyzing the beats of music can also help people compose music, automatically generate music, and classify and annotate music.
[0003] Currently, the research directions in the field of music beat tracking can be roughly divided into two categories: offline beat tracking and online beat tracking:
[0004] Offline beat tracking is to detect beats when an audio file is pre-recorded and completely stored; usually, this type of tracking does not require real-time performance, so more complex algorithms can be used for in-depth analysis.
[0005] Online beat tracking refers to detecting and marking the positions of beats in real time during the real-time playback of an audio signal. This method is usually applied to scenarios that require immediate feedback, such as music players, DJ (Disc Jockey) software, real-time synchronization and interactive music systems, etc. Compared with offline beat tracking, online beat tracking has received less attention because online beat tracking faces unique challenges, including only being able to access partial data, the processing speed cannot be too slow, and it is impossible to correct previous detections, and online methods are usually causal, which means they can only use past and present features for speculation.
[0006] There are some problems with existing online beat tracking technologies:
[0007] 1. Existing methods usually extract the basic features of the spectrogram through multiple two-dimensional convolutional layers, but this way is prone to ignoring the potential key beat information in the audio signal, especially in the context of variable music rhythms and complex audio backgrounds.
[0008] 2. In the prior art, the flow Transformer encoder is used to improve the accuracy of online beat tracking. However, the self-attention mechanism in the flow Transformer encoder mainly captures the relationships or dependencies between positions in the input sequence. In the task of beat tracking, it is also necessary to pay attention to the importance of different frames for the entire sequence in order to obtain a more accurate beat sequence.
[0009] 3. In existing beat tracking methods, cross-entropy is usually used as the loss function. Although this method is simple and effective, it is prone to causing the model to be overconfident in predicting certain classes during training, thereby generating the risk of overfitting, especially in the case of unbalanced data or high noise. Summary of the Invention
[0010] The purpose of the present invention is to provide an audio processing method and related device, which solves the problem that the existing method cannot accurately track the beat sequence.
[0011] The present invention is implemented through the following technical solutions:
[0012] The present invention discloses an audio processing method, including the following steps:
[0013] S1. Preprocess the audio file and convert it into a spectrogram;
[0014] S2. Use a pre-constructed feature extraction model to extract features from the spectrogram to obtain feature vectors;
[0015] S3. Use a pre-constructed neural network model based on the flow Transformer encoder to process the feature vectors to obtain beat activation values and strong beat activation values;
[0016] S4. Based on the beat activation values and strong beat activation values, use a probability model to infer the beat sequence and strong beat sequence.
[0017] Further, S1 is specifically:
[0018] Preprocess the input audio file through the short-time Fourier transform method and convert it into a spectrogram.
[0019] Further, in S2, the pre-constructed feature extraction model includes a two-dimensional convolutional layer, a first max-pooling layer, a frequency-time attention module, a second max-pooling layer, and a depthwise separable convolutional layer connected in sequence;
[0020] The process of feature extraction is specifically:
[0021] S2.1. Use the two-dimensional convolutional layer to perform preliminary channel expansion on the spectrogram to obtain a multi-channel time-frequency feature vector;
[0022] S2.2. Reduce the frequency dimension size of the multi-channel time-frequency feature vector using the first max pooling layer to obtain the time-frequency feature vector after the first dimensionality reduction.
[0023] S2.3. Use the frequency-time attention module to weight the time-frequency feature vector after the first dimensionality reduction in the time dimension and frequency dimension, and at the same time expand the number of channels again to obtain the time-frequency feature vector after frequency-time attention.
[0024] S2.4. Reduce the frequency dimension size of the time-frequency feature vector after frequency-time attention using the second max pooling layer to obtain the time-frequency feature vector after the second dimensionality reduction.
[0025] S2.5. Use the depthwise separable convolutional layer to expand the number of channels of the time-frequency feature vector after the second dimensionality reduction to obtain the feature vector.
[0026] Further, in S3, the pre-constructed neural network model based on the flow Transformer encoder includes a flow Transformer encoder, a linear layer, and a non-linear activation function layer connected in sequence.
[0027] The flow Transformer encoder includes an attention layer, a splicing layer, a fully connected layer, and a feed-forward layer connected in sequence. The attention layer is divided into a multi-head self-attention layer and an external attention layer.
[0028] The multi-head self-attention layer and the external attention layer respectively calculate the attention feature vectors, and the two attention feature vectors are spliced through the splicing layer to obtain the spliced feature vector.
[0029] The spliced feature vector is restored through the fully connected layer to obtain the restored feature vector; the restored feature vector has the same dimension as the input feature vector of the flow Transformer encoder.
[0030] Then the deep information of the restored feature vector is extracted by the feed-forward layer, and finally the deep information is fused through the linear layer and the non-linear activation function layer to obtain the beat activation value and the strong beat activation value.
[0031] Further, in S4, the probability model uses an online dynamic Bayesian network.
[0032] Further, the loss functions of the feature extraction model and the neural network model are the label-smoothing cross-entropy loss functions, and the formula of the label-smoothing cross-entropy loss function is as follows:
[0033] ;
[0034] Among them, is the label after label smoothing; is the loss value calculated by the cross-entropy loss function with label smoothing;
[0035] ;
[0036] wherein, represents the target category, including three categories: beat, strong beat, and non-beat; C represents the number of categories; p i represents the prediction probability of the online dynamic Bayesian network for the i-th category, where i is a specific category; represents the label smoothing factor; log() represents the logarithmic function.
[0037] The present invention also discloses an audio processing system, including:
[0038] A preprocessing module for preprocessing an audio file and converting it into a spectrogram;
[0039] A feature extraction module for extracting features from the spectrogram using a pre-constructed feature extraction model to obtain a feature vector;
[0040] An activation value calculation module for processing the feature vector using a pre-constructed neural network model based on a flow Transformer encoder to obtain a beat activation value and a strong beat activation value;
[0041] A prediction module for inferring a beat sequence and a strong beat sequence based on the beat activation value and the strong beat activation value using a probability model.
[0042] Furthermore, the feature extraction module includes a two-dimensional convolutional layer, a first max-pooling layer, a frequency-time attention module, a second max-pooling layer, and a depthwise separable convolutional layer connected in sequence;
[0043] The two-dimensional convolutional layer is used to preliminarily expand the channels of the spectrogram to obtain a multi-channel time-frequency feature vector;
[0044] The first max-pooling layer is used to reduce the frequency dimension size of the multi-channel time-frequency feature vector to obtain a once-dimension-reduced time-frequency feature vector;
[0045] The frequency-time attention module is used to weight the once-dimension-reduced time-frequency feature vector in the time dimension and the frequency dimension, and at the same time expand the number of channels again to obtain a frequency-time attentioned time-frequency feature vector;
[0046] The second max-pooling layer is used to reduce the frequency dimension size of the frequency-time attentioned time-frequency feature vector to obtain a twice-dimension-reduced time-frequency feature vector;
[0047] The depthwise separable convolutional layer is used to further expand the number of channels of the twice-dimension-reduced time-frequency feature vector to obtain a feature vector.
[0048] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the audio processing method are implemented.
[0049] The present invention also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the audio processing method are implemented.
[0050] Compared with the prior art, the present invention has the following beneficial technical effects:
[0051] The present invention discloses an audio processing method. First, the time-domain signal of an audio file is converted into a spectrogram frequency-domain representation form, which can more clearly show the energy distribution of the audio signal at different frequencies; key features that can characterize the audio beat characteristics are extracted from the spectrogram, converting complex spectral information into a low-dimensional feature vector. By extracting the feature vector, the data volume can be reduced, irrelevant information can be removed, and the essential features related to beat tracking can be highlighted, enabling the subsequent neural network model to more attentively learn and analyze the patterns related to beats; a pre-constructed neural network model based on a flow Transformer encoder is used to process the feature vector to obtain a beat activation value and a downbeat activation value. The flow Transformer encoder can capture the long-term and short-term dependencies in the feature vector sequence, while paying attention to the correlations between different moments in the beat signal, learning the complex mapping relationship between the feature vector and the beat activation value and the downbeat activation value. These two values reflect the likelihood that this moment is a beat or a downbeat, thereby realizing the preliminary prediction and positioning of audio beats and downbeats. Finally, based on the beat activation value and the downbeat activation value, the present invention uses a probability model to infer the beat sequence and the downbeat sequence. Since the beat activation value and the downbeat activation value only reflect the probability that each moment is a beat or a downbeat, these values alone may have certain noise and uncertainties; the probability model can comprehensively consider the activation value information of multiple moments and use the method of probability inference to infer the most likely beat sequence and downbeat sequence. The probability model can smooth the activation values, remove some local fluctuations and false predictions, optimize the inference result, and obtain a more accurate and coherent beat sequence and downbeat sequence, ultimately realizing the precise tracking of beats in the audio file.
[0052] Furthermore, when the present invention extracts features from the spectrogram, it first converts the original single-channel or few-channel spectrogram into a multi-channel time-frequency feature vector. The increase in the number of channels means that more types of feature information can be captured, enriching the feature expression ability and providing a more comprehensive feature representation for subsequent processing. Then, it reduces the frequency dimension size of the feature vector to reduce the data volume and computational complexity. The time-frequency feature vector after the first dimensionality reduction is weighted in both the time dimension and the frequency dimension. In the time dimension, it can pay attention to the importance of features at different time points; in the frequency dimension, it can highlight the features within certain specific frequency ranges. Through this weighting operation, the model can pay more attention to important time-frequency features and improve the feature expression ability. The number of channels is expanded again to further increase the diversity and richness of features, enabling better capture of complex information in the spectrogram. The frequency dimension size is reduced again to further reduce the data volume and computational complexity while continuing to highlight important features and improving the compactness and representativeness of features. The number of channels of the time-frequency feature vector after the second dimensionality reduction is further expanded to further enrich the feature expression ability, enabling the neural network model to learn more complex feature representations.
[0053] Furthermore, the neural network model includes a streaming Transformer encoder, a fully connected layer, and a feedforward layer connected in sequence. The streaming Transformer encoder retains the multi-head self-attention layer with relative position encoding and introduces an external attention layer at the same time. The multi-head self-attention layer mainly focuses on the dependencies within the input sequence at the micro level and models local context information. The external attention layer can establish connections between different time frames of the audio signal and macroscopically learn which time steps are most important for the current prediction. The streaming Transformer encoder combines the external attention mechanism, which can not only capture the dependencies between various positions in the input feature vector but also generate more accurate beat activation values by paying attention to the importance of different frames for the entire sequence, thus significantly improving the accuracy of online beat tracking.
[0054] Furthermore, the use of the cross-entropy loss function with label smoothing effectively alleviates the problem that the traditional cross-entropy loss function is prone to causing model overfitting. Especially in the case of data imbalance or the presence of noise, it shows stronger generalization ability, thus improving the robustness and reliability of the model. Description of the Drawings
[0055] Figure 1 is a flowchart of an audio processing method of the present invention;
[0056] Figure 2 is a schematic diagram of an audio processing system of the present invention;
[0057] Figure 3 is an overall network architecture diagram of an audio processing system;
[0058] Figure 4 Schematic diagram of the frequency-time attention module of the present invention;
[0059] Figure 5 Schematic diagram of the depthwise separable convolution of the present invention;
[0060] Figure 6 Schematic diagram of the context chunk processing mechanism in the flow Transformer encoder of the present invention;
[0061] Figure 7 Schematic diagram of the overall structure of the flow Transformer encoder of the present invention;
[0062] Figure 8 Schematic diagram of the self-attention mechanism and the external attention mechanism in the flow Transformer encoder of the present invention; wherein, Figure (a) is the schematic diagram of the self-attention mechanism; Figure (b) is the schematic diagram of the external attention mechanism. Detailed implementation manners
[0063] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further detailed description is given in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all embodiments.
[0064] The components described and illustrated in the accompanying drawings and embodiments of the present invention can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the following drawings is not intended to limit the scope of the claimed invention, but merely represents a selected embodiment of the present invention. Based on the accompanying drawings and embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.
[0065] The features and performance of the present invention are further described in detail below in conjunction with the embodiments.
[0066] Embodiment 1
[0067] As Figure 1 shown, the present invention provides an audio processing method, including the following steps:
[0068] S1. Preprocess the audio file and convert it into a spectrogram;
[0069] S2. Use a pre-constructed feature extraction model to extract features from the spectrogram to obtain feature vectors;
[0070] S3. Process the feature vectors using a pre - constructed neural network model based on a streaming Transformer encoder to obtain beat activation values and strong beat activation values;
[0071] S4. Based on the beat activation values and strong beat activation values, use a probability model to infer and obtain a beat sequence and a strong beat sequence.
[0072] Embodiment 2
[0073] Based on Embodiment 1, S1 is introduced in detail.
[0074] Pre - process the audio file and convert it into a spectrogram, specifically:
[0075] Pre - process the input audio file through the short - time Fourier transform method and convert it into a spectrogram.
[0076] Because the beat information of the audio is closely related to frequency components, energy changes, etc., these beat - related features can be better captured through the spectrogram, preparing for subsequent feature extraction and analysis.
[0077] The principle of the short - time Fourier transform (STFT) is as follows: When processing an audio signal, the audio signal is usually a non - stationary signal that changes over time. The traditional Fourier transform can only give the overall spectral characteristics of the signal and cannot reflect the change of frequency components at different times. The core idea of the short - time Fourier transform is to divide the audio signal into many shorter, approximately stationary time segments, and perform Fourier transform on the signal in each time segment, thereby obtaining the spectral information of the signal at different times.
[0078] Embodiment 3
[0079] Based on Embodiment 1, S2 is introduced in detail.
[0080] The feature vectors include time - domain feature vectors and frequency - domain feature vectors.
[0081] To efficiently extract accurate time - domain features and frequency - domain features from the spectrogram, the present invention adopts a feature extraction method of frequency - temporal attention (FTA) module and depth - separable convolution.
[0082] The core principle of the FTA module is to calculate the attention weights in the frequency domain and time domain to highlight the important feature parts in the data. Specifically, it analyzes the input data in the frequency domain and time domain respectively, calculates the importance scores of each position, and then weights the data according to these scores, thereby enhancing the representation of important features and suppressing unimportant features.
[0083] Specifically, as Figure 3 shown, an audio file is preprocessed and converted into a spectrogram file of size (bs, C, T, F), where bs represents the number of samples in one training, C represents the number of channels, T represents the number of time frames, and F represents the number of frequency bins included in the spectrogram.
[0084] As Figure 3 shown, the pre-constructed feature extraction model includes a two-dimensional convolutional layer, a first max pooling layer, a frequency-time attention module, a second max pooling layer, and a depthwise separable convolutional layer connected in sequence.
[0085] Feature extraction is performed on the spectrogram to obtain a feature vector, specifically:
[0086] First, a two-dimensional convolutional layer is used to expand the input channels of the spectrogram to 32 channels, obtaining a multi-channel time-frequency feature vector;
[0087] Then, a max pooling layer is used to reduce the frequency dimension size of the multi-channel time-frequency feature vector, obtaining a time-frequency feature vector after the first dimensionality reduction;
[0088] Next, the FTA module weights the feature vector in the time dimension and frequency dimension, and at the same time expands the number of channels of the feature map from 32 to 64 to capture richer information in a higher-dimensional feature space, obtaining a time-frequency feature vector after frequency-time attention;
[0089] Continue to use a max pooling layer to reduce the frequency dimension size of the time-frequency feature vector after frequency-time attention, obtaining a time-frequency feature vector after the second dimensionality reduction;
[0090] Finally, through the depthwise separable convolutional layer, the number of channels of the time-frequency feature vector after the second dimensionality reduction is expanded to 256, ensuring that the feature extraction part reaches the optimal in computational efficiency while retaining richer time-frequency features, and the size of the finally output feature vector is (bs, T, 256).
[0091] Among them, the FTA module can dynamically adjust the focus of attention of the feature extraction model in the time and frequency dimensions, enabling the subsequent neural network model to effectively extract and fuse useful spectrogram features from complex inputs and improving the accuracy of beat tracking. The depthwise separable convolution can effectively reduce the parameters and computational amount of the neural network model and improve the efficiency of beat tracking.
[0092] The following are the specific structures of the frequency-time attention module and the depthwise separable convolution.
[0093] The network structure of the frequency-time attention module is as Figure 4 shown, including a frequency domain attention sub-module in the left dotted area part as Figure 4 shown in the middle and a time domain attention sub-module in the right dotted area part asFigure 4 The time-domain attention sub-module in the right dashed-line area in the middle.
[0094] Specifically, the processing process of the frequency-domain attention sub-module is as follows:
[0095] Given the input feature map X ∈ R C×F×T , first, the input feature map X is subjected to average pooling through a max-pooling layer to calculate the distribution of the amplitude along the time axis, obtaining the frequency descriptor f ∈ R C×F :
[0096] ;
[0097] where represents the element in the i-th row and j-th column of the input feature map X, and i and j are positive integers; T represents the number of time frames.
[0098] Then, two one-dimensional convolutional layers are used to learn the correlation in the frequency dimension . For the frequency descriptor f, the process of one-dimensional convolution can be written as:
[0099] ;
[0100] where is the one-dimensional convolution kernel of the l-th layer, is the newly generated feature map, and * is the convolution operator.
[0101] Finally, a Softmax layer is applied to obtain the frequency attention feature map :
[0102] ;
[0103] .
[0104] Similarly, through the same process, the time-domain attention sub-module can also obtain the time attention feature map of the time axis , .
[0105] Meanwhile, in order to learn high-level semantic features, two two-dimensional convolutional layers with convolution kernel sizes of (3 × 3) and (5 × 5) respectively are applied to the input feature map to obtain two new feature maps , and then and are multiplied using matrix multiplication to obtain the final output .
[0106] where ;
[0107] ;
[0108] 。
[0109] Among them, broadcast is a broadcast operation that enables element-wise multiplication of matrices of different shapes to be performed compatibly; respectively represent the frequency feature map and the temporal feature map that have learned high-level semantic features; represents the matrix multiplication operator.
[0110] The structure of the depthwise separable convolutional layer is as Figure 5 shown. In traditional convolutional operations, the convolutional kernel performs convolutional operations in both the spatial dimension and the channel dimension of the input feature map. The convolutional kernel is usually a k×k matrix. Assuming the input feature map has C in channels, the convolutional kernel has C out channels. Thus, the computational complexity of traditional convolutional operations is O(k 2 ×C in ×C out ).
[0111] The depthwise separable convolutional layer divides the convolutional operation into two stages:
[0112] First, perform depthwise convolution. The convolutional kernel of depthwise convolution is in single-channel mode, that is, each channel of the input is processed using an independent convolutional kernel without information interaction between channels. This is a lightweight operation because the convolutional computational complexity of each channel is reduced;
[0113] Then, perform pointwise convolution. Pointwise convolution is to fuse the features obtained from depthwise convolution across channels. Through a 1×1 convolutional kernel, the information between different channels is combined to obtain a new feature map, and the number of output feature maps depends on the number of filters.
[0114] By combining these two stages, both the computational complexity and the number of parameters are greatly reduced. The computational complexity of depthwise convolution is O(k 2 ×C in ), and the computational complexity of pointwise convolution is O(C in ×C out ). Therefore, the overall computational complexity is reduced to O(k 2 ×C in +C in ×C out ), which is much less than traditional convolutional operations.
[0115] In summary, the combination of the FTA module and the depthwise separable convolutional layer can not only retain sufficient feature expression ability but also greatly reduce the computational complexity. In practical applications, especially for complex audio signals such as multi-instrument performances, different music styles, or scenes with strong noise interference, this feature extraction method can achieve a good balance between computational efficiency and performance and meet the requirements of real-time processing.
[0116] By introducing the frequency-time attention module and the depthwise separable convolutional layer, the present invention can more effectively capture key beat information in complex audio backgrounds and changing music rhythms. While reducing the computational cost, this method can also significantly improve the accuracy and robustness of feature extraction.
[0117] Embodiment 4
[0118] Based on Embodiment 1, S3 is introduced in detail: A pre-constructed neural network model based on a flow Transformer encoder is used to process the feature vector to obtain beat activation values and strong beat activation values.
[0119] The neural network model based on the flow Transformer encoder includes a flow Transformer encoder, a linear layer, and a non-linear activation function layer. The non-linear activation function layer uses the Sigmoid activation function. The flow Transformer encoder is responsible for modeling the temporal information of the audio signal. Through the self-attention mechanism and the external attention mechanism, it captures both the microscopic information and the macroscopic information in the sequence. Then, through a linear layer and a non-linear activation function layer, further feature transformation is performed to finally obtain the activation values of beats and strong beats.
[0120] The present invention uses a neural network model based on a flow Transformer encoder to obtain the activation values of beats and strong beats. The Transformer module for offline beat tracking directly processes the entire input sequence, while the online beat tracking of the present invention is based on a context chunk processing mechanism and generates outputs by providing partial audio frames.
[0121] Figure 6 is the basic structure of the context chunk processing mechanism. First, the time-frequency feature vector is used as the input feature, and the sequence is segmented into multiple non-overlapping blocks C = {C1, C2, …, C b , …}, where b is the index of the block, and each block contains N c frames. To alleviate the block boundary effect, that is, the problem of broken context information between blocks, each block C b is extended to a context block: (1) the left sub-block L b : the N b frames in front of C l ; (2) C b; (3) right sub-block R b : C b The subsequent N r frames. Connect them to form a context block [L b , C b , R b , which serves as the input to the stream Transformer encoder. Figure 6 is an example of the case where N c = 2, N l = 1, N r = 1. Among them, [L b-1 , C b-1 , C b-1 , R b-1 represents the input feature vector of the (b - 1)-th block, [L b , C b , C b , R b represents the input feature vector of the b-th block, [L b+1 , C b+1 , C b+1 , R b+1 represents the input feature vector of the (b + 1)-th block, Encoder Layer represents an encoder layer, Z b-1 represents the output of the (b - 1)-th block, Z b represents the output of the b-th block, Z b+1 represents the output of the (b + 1)-th block. To capture the long-term dependencies of the input sequence, each block C b also inherits a context embedding vector c b during its processing, which is obtained from the previous block (i.e., the (b - 1)-th block). This mechanism can transmit information across blocks, thereby alleviating the context break problem between blocks. In the figure, c b-1 represents the context embedding vector of block C b-1 , and c b+1 represents the context embedding vector of block C b+1 . In the network of the present invention, N c , N l , N r take 16, 256, 16 respectively. The encoder layer uses an improved stream Transformer encoder, that is, Figure 7 the network structure shown.
[0122] After the time-frequency feature vector is processed by block division, it enters the calculation of the attention feature vector. As Figure 7As shown, the streaming Transformer encoder includes an attention layer, a concatenation layer, a fully connected layer, and a feed-forward layer connected in sequence. The attention layer is divided into a multi-head self-attention layer and an external attention layer. In each block, the original multi-head self-attention layer with relative position encoding is retained, and at the same time, an external attention layer is introduced to calculate the attention feature vectors respectively. Then, the two attention feature vectors are concatenated through the concatenation layer to obtain the concatenated feature vector. A fully connected layer is used to restore the concatenated feature vector to its original dimension size, obtaining a feature vector with the same dimension as the input feature vector of the streaming Transformer encoder. Then, the feed-forward layer extracts the deep information of the restored feature vector, and finally, the deep information is fused through a linear layer and a non-linear activation function layer to obtain the beat activation value and the strong beat activation value. Among them, the multi-head self-attention layer mainly focuses on the dependence within the input sequence at the micro level and models the local context information. The external attention layer can establish connections between different time frames of the audio signal and macroscopically learn which time steps are most important for the current prediction. The combination of the two can make full use of local information and improve the understanding ability of the neural network model for the beat features in the audio signal.
[0123] The traditional self-attention calculation method relies on generating an attention map by calculating the correlation between the query vector and the key vector. As Figure 8 shown in Figure (a) of [reference], this attention map reflects the degree of attention of each input feature to other features. Subsequently, these attention weights are applied to the value vector to obtain a weighted feature map.
[0124] However, the working principle of the external attention mechanism is different. As Figure 8 shown in Figure (b) of [reference], first, an attention map is generated by calculating the correlation between the query vector and the externally learnable key attention matrix , and then, by multiplying this attention map with the externally learnable value attention matrix , a more refined feature map is finally generated as the output feature. In this mechanism, the externally learnable attention matrix provides an additional information source when calculating the attention, enabling the model to model the dependence between different time steps.
[0125] Among them, these two external attention matrices are implemented through fully connected layers to ensure that they can be optimized in an end-to-end backpropagation manner. It is worth noting that these two attention matrices are independent of individual samples and share parameters across the entire dataset. This sharing mechanism makes the computational complexity linearly related to the number of input features, not only reducing the number of parameters but also enabling the model to capture the most informative parts globally, helping the model to focus more on key time-step features, playing a strong regularization role for the model, and thus enhancing the generalization ability of the attention mechanism.
[0126] Specifically, assume that the size of the input features is [bs, T, dmodel], where dmodel represents the input feature dimension, taking 256. First, it passes through a fully connected layer M k , and the feature dimension changes from dmodel to S (S is the feature dimension after dimensionality reduction, taking 64). Through M k , the 256-dimensional input features can be compressed into a feature space of size 64 to extract information in the low-dimensional space. In this way, the neural network model can focus on more critical beat information and not be disturbed by excessive redundant audio features. Then, a Softmax operation is applied on the second dimension (i.e., the time dimension) of the feature vector to normalize all S-dimensional features at each time step. In this way, each time step will obtain a clear attention allocation weight value, which can distinguish which features contribute more to the rhythm at different time steps. Then, it passes through another fully connected layer M v , which maps the low-dimensional features at each time step back to the original dimension dmodel for subsequent processing by network modules.
[0127] The external attention only calculates the attention through simple linear projection and softmax operations. This lightweight calculation can improve the accuracy without reducing the efficiency. Coupled with the accurate capture of local feature information by self-attention, the model can comprehensively understand the input features.
[0128] Embodiment 5
[0129] Based on Embodiment 1, S4 is introduced in detail: infer the final beat sequence and strong beat sequence using a probability model based on the beat activation value and the strong beat activation value.
[0130] In the post-processing stage, the output of the neural network model based on the flow Transformer encoder will be optimized through an online Dynamic Bayesian Network (DBN).
[0131] Dynamic Bayesian networks are commonly used to process time series data and can smooth and correct the results of beat detection to improve stability and accuracy. An online dynamic Bayesian network is a probabilistic graphical model suitable for dynamic time series data, which can process real-time data streams in an updated manner.
[0132] At each time step, the online DBN only processes the data at the current moment and makes predictions based on the previous historical states. When the feature vector is processed by the neural network model, the probability values of each frame being a beat, a strong beat, or a non-beat can be obtained. Through further processing by the online DBN, the beat (strong beat) tracking results of the entire time series can be obtained.
[0133] Optimize the model by combining the loss function and verify and test the feature extraction model and the neural network model.
[0134] The cross-entropy loss function is a commonly used loss function in classification tasks. Its calculation formula is:
[0135] ;
[0136] where C represents the number of classes, y i is an element of the true label. For the target class being 1 and the other classes being 0, p i is the predicted probability of the model for the i-th class, usually calculated through the softmax function.
[0137] In real music signals, beat tracking usually faces problems such as complex rhythm changes, audio noise, or irregular music rhythms. If hard labels (e.g., directly setting the label to 1 or 0) are used for the classification labels of beats and strong beats, then the model may be too confident in the classification of some time steps, which will cause it to be unable to adapt to the irregularity of the rhythm or noise. Since the online beat tracking task cannot modify the already predicted data, this method of calculating the loss is not suitable for online tasks. The present invention proposes to use a cross-entropy loss function with label smoothing to reduce the overfitting of the model and enhance the robustness of the model.
[0138] The cross-entropy loss with label smoothing is a modification of the traditional cross-entropy loss function, aiming to reduce the overconfidence of the model in a certain class. In label smoothing, the target label is "smoothed" so that the label value of the target class decreases from 1 to , and the label of the non-target class increases from 0 to , where is the label smoothing factor, usually taking a very small value (e.g., 0.1). The formula for the cross-entropy loss function with label smoothing is as follows:
[0139] ;
[0140] Among them, is the label after label smoothing:
[0141] ;
[0142] Among them, represents the target category, including three categories: beat, strong beat, and non-beat; C represents the number of categories, taking C = 3; p i represents the prediction probability of the online dynamic Bayesian network for the i-th category, where i is the specific category; represents the label smoothing factor; log() represents the logarithmic function.
[0143] In this way, label smoothing reduces the absolute weight of the target category label being 1, and at the same time assigns a very small non-zero value to the non-target categories. This change avoids the overconfidence of the online dynamic Bayesian network in a certain category during the training process, making the prediction probability smoother.
[0144] Label smoothing enables the model to maintain a certain degree of flexibility and robustness when facing label noise or mislabeling, and adjust and make predictions more smoothly. If there are some beats mislabeled or the positions of some beats are not clear enough in the training data, the traditional cross-entropy loss will cause the model to overfit these uncertain or inaccurate labels, while the loss function with label smoothing can effectively alleviate the overfitting of the model to the wrong labels by reducing the "hard" constraints of the labels, and improve the generalization ability of the model. And when the dataset is small or the class is imbalanced, label smoothing can also make the prediction distribution of the model more gentle, prompting the model to pay more attention to the features of each category and enhancing the robustness of the model.
[0145] Based on the existing Ballroom dataset, SMC dataset, Hainsworth dataset, and GTZAN dataset, the present invention uses a specially designed loss function to train and iterate the feature extraction model and the neural network model, and finally obtains a robust model.
[0146] Embodiment 6
[0147] As Figure 2 shown, the present invention also discloses an audio processing system, including:
[0148] A preprocessing module for preprocessing the audio file and converting it into a spectrogram;
[0149] A feature extraction module for extracting features from the spectrogram using a pre-constructed feature extraction model to obtain feature vectors;
[0150] An activation value calculation module, configured to process the feature vector by using a pre-constructed neural network model to obtain a beat activation value and a strong beat activation value;
[0151] A prediction module, configured to infer a beat sequence and a strong beat sequence based on the beat activation value and the strong beat activation value by using a probability model.
[0152] Embodiment 7
[0153] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the audio processing method are implemented. Among them, the memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc.; the processor, the network interface, and the memory are interconnected through an internal bus, and the internal bus may be an Industry Standard Architecture bus, a Peripheral Component Interconnect standard bus, an Extended Industry Standard Architecture bus, etc. The bus may be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include a memory and a non-volatile memory, and provide instructions and data to the processor.
[0154] Embodiment 8
[0155] The present invention also discloses a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of the audio processing method are implemented. Specifically, the computer-readable storage medium includes, but is not limited to, for example, a volatile memory and / or a non-volatile memory. The volatile memory may include a random access memory and / or a cache memory, etc. The non-volatile memory may include a read-only memory, a hard disk, a flash memory, an optical disc, a magnetic disk, etc.
[0156] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, optical memories, etc.) containing computer-usable program code.
[0157] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and combinations of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general purpose computers, special purpose computers, embedded processors, or other programmable data processing devices to produce a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0158] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0159] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in the flow Figure 1 one flow or multiple flows and / or blocks Figure 1 means for implementing the functions specified in one block or multiple blocks.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent replacements can still be made to the specific embodiments of the present invention, and any modifications or equivalent replacements that do not depart from the spirit and scope of the present invention should be covered by the protection scope of the present invention.
Claims
1. An audio processing method, characterized in that, It includes the following steps: S1. Preprocess the audio file and convert it into a spectrogram; S2. Use a pre-constructed feature extraction model to extract features from the spectrogram to obtain feature vectors; S3. Use a pre-constructed neural network model based on a flow Transformer encoder to process the feature vectors to obtain beat activation values and strong beat activation values; S4. Based on the beat activation values and strong beat activation values, use a probability model to infer the beat sequence and strong beat sequence; In S2, the pre-constructed feature extraction model includes a two-dimensional convolutional layer, a first max-pooling layer, a frequency-time attention module, a second max-pooling layer, and a depthwise separable convolutional layer connected in sequence; The specific process of feature extraction is as follows: S2.
1. Use a two-dimensional convolutional layer to preliminarily expand the channels of the spectrogram to obtain a multi-channel time-frequency feature vector; S2.
2. Use the first max-pooling layer to reduce the frequency dimension size of the multi-channel time-frequency feature vector to obtain a time-frequency feature vector after the first dimensionality reduction; S2.
3. Use the frequency-time attention module to weight the time-frequency feature vector after the first dimensionality reduction in the time dimension and frequency dimension, and at the same time expand the number of channels again to obtain a time-frequency feature vector after frequency-time attention; S2.
4. Use the second max-pooling layer to reduce the frequency dimension size of the time-frequency feature vector after frequency-time attention to obtain a time-frequency feature vector after the second dimensionality reduction; S2.
5. Use a depthwise separable convolutional layer to expand the number of channels of the time-frequency feature vector after the second dimensionality reduction to obtain a feature vector; In S3, the pre-constructed neural network model based on a flow Transformer encoder includes a flow Transformer encoder, a linear layer, and a non-linear activation function layer connected in sequence; The flow Transformer encoder includes an attention layer, a splicing layer, a fully connected layer, and a feed-forward layer connected in sequence. The attention layer is divided into a multi-head self-attention layer and an external attention layer; The flow Transformer encoder first performs block processing on the feature vectors obtained in S2 based on the context block processing mechanism; The multi-head self-attention layer and the external attention layer respectively calculate the attention feature vectors, and the two attention feature vectors are spliced through the splicing layer to obtain a spliced feature vector; The spliced feature vector is restored through the fully connected layer to obtain a restored feature vector; the dimension of the restored feature vector is the same as that of the input feature vector of the flow Transformer encoder; Then the deep information of the restored feature vector is extracted by the feed-forward layer, and finally the deep information is fused through the linear layer and the non-linear activation function layer to obtain the beat activation value and the strong beat activation value.
2. The audio processing method according to claim 1, characterized in that S1 is specifically: The input audio file is preprocessed by the short-time Fourier transform method and converted into a spectrogram.
3. An audio processing method according to claim 1, characterized in that In S4, the probability model uses an online dynamic Bayesian network.
4. An audio processing method according to claim 1, characterized in that, The loss functions of the feature extraction model and the neural network model are label-smoothing cross-entropy loss functions, and the formula of the label-smoothing cross-entropy loss function is as follows: ; Among them, is the label after label smoothing; is the loss value calculated by the cross-entropy loss function with label smoothing; ; Among them, represents the target category, including three categories: beat, strong beat, and non-beat; C represents the number of categories; p i represents the prediction probability of the online dynamic Bayesian network for the i-th category, where i is a specific category; represents the label smoothing factor; log() represents the logarithmic function.
5. An audio processing system for implementing the audio processing method according to any one of claims 1-4, characterized in that, It includes: A preprocessing module for preprocessing the audio file and converting it into a spectrogram; A feature extraction module, configured to extract features from a spectrogram using a pre-constructed feature extraction model to obtain a feature vector; An activation value calculation module, configured to process the feature vector using a pre-constructed neural network model based on a streaming Transformer encoder to obtain a beat activation value and a downbeat activation value; A prediction module, configured to infer a beat sequence and a downbeat sequence based on the beat activation value and the downbeat activation value using a probability model.
6. An audio processing system according to claim 5, wherein The feature extraction module includes a two-dimensional convolutional layer, a first max pooling layer, a frequency-time attention module, a second max pooling layer, and a depthwise separable convolutional layer connected in sequence; The two-dimensional convolutional layer is configured to perform preliminary channel expansion on the spectrogram to obtain a multi-channel time-frequency feature vector; The first max pooling layer is configured to reduce the frequency dimension size of the multi-channel time-frequency feature vector to obtain a time-frequency feature vector with a first dimensionality reduction; The frequency-time attention module is configured to weight the time-frequency feature vector with a first dimensionality reduction in the time dimension and the frequency dimension, and at the same time expand the number of channels again to obtain a time-frequency feature vector after frequency-time attention; The second max pooling layer is configured to reduce the frequency dimension size of the time-frequency feature vector after frequency-time attention to obtain a time-frequency feature vector with a second dimensionality reduction; The depthwise separable convolutional layer is configured to further expand the number of channels of the time-frequency feature vector with a second dimensionality reduction to obtain a feature vector.
7. A computer device, comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, When the processor executes the computer program, the steps of the audio processing method according to any one of claims 1 to 4 are implemented.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the audio processing method according to any one of claims 1 to 4 are implemented.
Citation Information
Patent Citations
Training method of beat re-shooting joint detection model and beat re-shooting joint detection method
CN114897157A
Deep network line spectrum detection method embedded with attention mechanism
CN118568459A