Audio processing method and related device

By performing spectral transformation and feature extraction on audio files, combined with stream Transformer encoder and probability model, the accuracy problem of online beat tracking technology in complex backgrounds is solved, and accurate tracking of beats in audio files is achieved.

CN120071875AActive Publication Date: 2025-05-30XI AN JIAOTONG UNIV
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510421243.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-05-30
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing online beat tracking techniques are difficult to accurately capture beat information in audio signals, especially in the context of changing musical rhythms and complex audio.

Method used

By converting the audio file into a spectrum graph, a pre-constructed feature extraction model and a neural network model based on the stream Transformer encoder are used to extract and process the spectrum graph to obtain the beat activation value and the strong beat activation value, and the probability model is used to infer the beat sequence and the strong beat sequence.

Benefits of technology

Accurate tracking of beats in audio files is achieved, the accuracy and robustness of online beat tracking is improved, and the ability to effectively capture beat information in complex audio backgrounds.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071875A_ABST
    Figure CN120071875A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of music information retrieval, and particularly relates to an audio processing method and a related device. The audio processing method comprises the following steps: preprocessing an audio file, and converting the audio file into a spectrogram; performing feature extraction on the spectrogram by adopting a pre-constructed feature extraction model to obtain a feature vector; processing the feature vector by adopting a pre-constructed neural network model based on a flow Transform encoder to obtain a beat activation value and a strong beat activation value; and on the basis of the beat activation value and the forced beat activation value, a beat sequence and a forced beat sequence are inferred by adopting a probability model. The problem that a beat sequence cannot be accurately tracked in an existing method is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of music information retrieval, and particularly relates to an audio processing method and related device. Background Art

[0002] Beat tracking is a key task in the field of Music Information Retrieval (MIR). Its main purpose is to analyze the beat and downbeat sequences in an audio stream to express the rhythm of music. Good beat estimation is beneficial to various downstream tasks of MIR, including music transcription, structure analysis, etc. And because human perception of music rhythm is related to motor sensitivity, beat tracking can also be applied to more scenarios such as human-computer interaction and music therapy. In addition, tracking and analyzing the beats of music can also help people compose music, automatically generate music, and classify and annotate music.

[0003] Currently, the research directions in the field of music beat tracking can be roughly divided into two categories: offline beat tracking and online beat tracking: Offline beat tracking detects beats when an audio file is pre-recorded and completely stored; usually, this type of tracking does not require real-time performance, so more complex algorithms can be used for in-depth analysis.

[0004] Online beat tracking refers to detecting and marking the positions of beats in real time during the real-time playback of an audio signal. This method is usually applied to scenarios that require immediate feedback, such as music players, DJ (Disc Jockey) software, real-time synchronization and interactive music systems, etc. Compared with offline beat tracking, online beat tracking has received less attention because online beat tracking faces unique challenges, including only being able to access partial data, the processing speed cannot be too slow, and it is impossible to correct previous detections, and online methods are usually causal, which means they can only use past and present features for speculation.

[0005] There are some problems with existing online beat tracking technologies: 1. Existing methods usually extract the basic features of the spectrogram through multiple two-dimensional convolutional layers, but this way is prone to ignoring the potential key beat information in the audio signal, especially in the context of changing music rhythms and complex audio backgrounds.

[0006] 2. In existing technologies, the streaming Transformer encoder is used to improve the accuracy of online beat tracking. However, the self-attention mechanism in the streaming Transformer encoder mainly captures the relationships or dependencies between various positions in the input sequence. In the task of beat tracking, it is also necessary to pay attention to the importance of different frames for the entire sequence in order to obtain a more accurate beat sequence.

[0007] 3. In existing beat tracking methods, cross-entropy is usually used as the loss function. Although this method is simple and effective, it is prone to causing the model to be overconfident in predicting certain categories during training, thereby generating the risk of overfitting, especially in the case of imbalanced data or high noise. Summary of the Invention

[0008] The purpose of the present invention is to provide an audio processing method and related device, which solves the problem that the existing method cannot accurately track the beat sequence.

[0009] The present invention is realized through the following technical solutions: The present invention discloses an audio processing method, including the following steps: S1. Preprocess the audio file and convert it into a spectrogram; S2. Use a pre-constructed feature extraction model to extract features from the spectrogram to obtain feature vectors; S3. Use a pre-constructed neural network model based on a flow Transformer encoder to process the feature vectors to obtain beat activation values and downbeat activation values; S4. Based on the beat activation values and downbeat activation values, use a probability model to infer the beat sequence and downbeat sequence.

[0010] Further, S1 is specifically: Preprocess the input audio file through the short-time Fourier transform method and convert it into a spectrogram.

[0011] Further, in S2, the pre-constructed feature extraction model includes a two-dimensional convolutional layer, a first max-pooling layer, a frequency-time attention module, a second max-pooling layer, and a depthwise separable convolutional layer connected in sequence; The process of feature extraction is specifically: S2.1. Use a two-dimensional convolutional layer to initially expand the channels of the spectrogram to obtain a multi-channel time-frequency feature vector; S2.2. Use the first max-pooling layer to reduce the frequency dimension size of the multi-channel time-frequency feature vector to obtain a time-frequency feature vector with one-dimensional reduction; S2.3. Use the frequency-time attention module to weight the time-frequency feature vector with one-dimensional reduction in both the time dimension and the frequency dimension, and at the same time expand the number of channels again to obtain a time-frequency feature vector after frequency-time attention; S2.4. Use the second max-pooling layer to reduce the frequency dimension size of the time-frequency feature vector after frequency-time attention to obtain a time-frequency feature vector with two-dimensional reduction; S2.5. Use the depthwise separable convolutional layer to expand the number of channels of the time-frequency feature vector with two-dimensional reduction again to obtain feature vectors.

[0012] Further, in S3, the pre-constructed neural network model based on the flow Transformer encoder includes a flow Transformer encoder, a linear layer, and a non-linear activation function layer connected in sequence; The flow Transformer encoder includes an attention layer, a splicing layer, a fully connected layer, and a feed-forward layer connected in sequence. The attention layer is divided into a multi-head self-attention layer and an external attention layer; The multi-head self-attention layer and the external attention layer respectively calculate attention feature vectors, and the two attention feature vectors are spliced through the splicing layer to obtain a spliced feature vector; The spliced feature vector is restored through the fully connected layer to obtain a restored feature vector; the restored feature vector has the same dimension as the input feature vector of the flow Transformer encoder; Then, the deep information of the restored feature vector is extracted by the feed-forward layer, and finally, the deep information is fused by the linear layer and the non-linear activation function layer to obtain a beat activation value and a strong beat activation value.

[0013] Further, in S4, the probability model uses an online dynamic Bayesian network.

[0014] Further, the loss functions of the feature extraction model and the neural network model are label-smoothing cross-entropy loss functions, and the formula of the label-smoothing cross-entropy loss function is as follows: ; where is the label after label smoothing; is the loss value calculated by the cross-entropy loss function with label smoothing; ; where represents the target category, including three categories: beat, strong beat, and non-beat; C represents the number of categories; p i represents the prediction probability of the online dynamic Bayesian network for the i-th category, and i is a specific category; represents the label smoothing factor; log() represents the logarithmic function.

[0015] The present invention also discloses an audio processing system, including: A preprocessing module for preprocessing an audio file and converting it into a spectrogram; A feature extraction module for extracting features from the spectrogram using a pre-constructed feature extraction model to obtain feature vectors; An activation value calculation module, configured to process the feature vector by using a pre-constructed neural network model based on a streaming Transformer encoder to obtain a beat activation value and a strong beat activation value; A prediction module, configured to infer a beat sequence and a strong beat sequence based on the beat activation value and the strong beat activation value by using a probability model.

[0016] Furthermore, the feature extraction module includes a two-dimensional convolutional layer, a first max pooling layer, a frequency-time attention module, a second max pooling layer, and a depthwise separable convolutional layer connected in sequence; The two-dimensional convolutional layer is configured to perform preliminary channel expansion on the spectrogram to obtain a multi-channel time-frequency feature vector; The first max pooling layer is configured to reduce the frequency dimension size of the multi-channel time-frequency feature vector to obtain a time-frequency feature vector with a first dimensionality reduction; The frequency-time attention module is configured to weight the time-frequency feature vector with a first dimensionality reduction in the time dimension and the frequency dimension, and at the same time expand the number of channels again to obtain a time-frequency feature vector after frequency-time attention; The second max pooling layer is configured to reduce the frequency dimension size of the time-frequency feature vector after frequency-time attention to obtain a time-frequency feature vector with a second dimensionality reduction; The depthwise separable convolutional layer is configured to further expand the number of channels of the time-frequency feature vector with a second dimensionality reduction to obtain a feature vector.

[0017] The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the audio processing method are implemented.

[0018] The present invention also discloses a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the audio processing method are implemented.

[0019] Compared with the prior art, the present invention has the following beneficial technical effects: The present invention discloses an audio processing method. First, the time-domain signal of an audio file is converted into a spectrogram frequency-domain representation form, which can more clearly show the energy distribution of the audio signal at different frequencies; key features that can characterize the audio beat characteristics are extracted from the spectrogram, converting complex spectral information into a low-dimensional feature vector. By extracting the feature vector, the data volume can be reduced, irrelevant information can be removed, and the essential features related to beat tracking can be highlighted, enabling the subsequent neural network model to more focusedly learn and analyze the patterns related to beats; a pre-constructed neural network model based on a flow Transformer encoder is used to process the feature vector to obtain a beat activation value and a downbeat activation value. The flow Transformer encoder can capture the long-term and short-term dependencies in the feature vector sequence, while paying attention to the correlation between different moments in the beat signal, learning the complex mapping relationship between the feature vector and the beat activation value and the downbeat activation value. These two values reflect the likelihood that the current moment is a beat or a downbeat, thereby realizing the preliminary prediction and positioning of audio beats and downbeats. Finally, based on the beat activation value and the downbeat activation value, the present invention uses a probability model to infer the beat sequence and the downbeat sequence. Since the beat activation value and the downbeat activation value only reflect the probability that each moment is a beat or a downbeat, there may be certain noise and uncertainty in these individual values; through the probability model, the activation value information of multiple moments can be comprehensively considered, and the method of probability inference is used to infer the most likely beat sequence and downbeat sequence. The probability model can smooth the activation value, remove some local fluctuations and false predictions, optimize the inference result, and obtain a more accurate and coherent beat sequence and downbeat sequence, ultimately realizing the precise tracking of beats in the audio file.

[0020] Furthermore, when the present invention extracts features from the spectrogram, it first converts the original single-channel or few-channel spectrogram into a multi-channel time-frequency feature vector. The increase in the number of channels means that more types of feature information can be captured, enriching the feature expression ability and providing a more comprehensive feature representation for subsequent processing. Then, it reduces the frequency dimension size of the feature vector to reduce the data volume and computational complexity. The time-frequency feature vector after the first dimensionality reduction is weighted in both the time dimension and the frequency dimension. In the time dimension, it can focus on the feature importance at different time points; in the frequency dimension, it can highlight the features within certain specific frequency ranges. Through this weighting operation, the model can pay more attention to important time-frequency features and improve the feature expression ability. The number of channels is expanded again to further increase the diversity and richness of features, enabling better capture of complex information in the spectrogram; the frequency dimension size is reduced again to further reduce the data volume and computational complexity, while continuing to highlight important features and improving the compactness and representativeness of features; the number of channels of the time-frequency feature vector after the second dimensionality reduction is further expanded to further enrich the feature expression ability, enabling the neural network model to learn more complex feature representations.

[0021] Furthermore, the neural network model includes a streaming Transformer encoder, a fully connected layer, and a feed-forward layer connected in sequence; the streaming Transformer encoder retains the multi-head self-attention layer with relative position encoding and introduces an external attention layer at the same time. The multi-head self-attention layer mainly focuses on the dependencies within the input sequence at the micro level and models local context information; the external attention layer can establish connections between different time frames of the audio signal and macroscopically learn which time steps are most important for the current prediction. The streaming Transformer encoder combines the external attention mechanism, which can not only capture the dependencies between various positions in the input feature vector but also generate more accurate beat activation values by paying attention to the importance of different frames for the entire sequence, thus significantly improving the accuracy of online beat tracking.

[0022] Furthermore, the use of the cross-entropy loss function with label smoothing effectively alleviates the problem that the traditional cross-entropy loss function is prone to causing model overfitting. Especially in the case of data imbalance or the presence of noise, it shows stronger generalization ability, thereby improving the robustness and reliability of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] Figure 1 is a flowchart of an audio processing method of the present invention; Figure 2 is a schematic diagram of an audio processing system of the present invention; Figure 3 is an overall network architecture diagram of an audio processing system; Figure 4Schematic diagram of the frequency-time attention module of the present invention; Figure 5 Schematic diagram of the depthwise separable convolution of the present invention; Figure 6 Schematic diagram of the context chunk processing mechanism in the flow Transformer encoder of the present invention; Figure 7 Schematic diagram of the overall structure of the flow Transformer encoder of the present invention; Figure 8 Schematic diagrams of the self-attention mechanism and the external attention mechanism in the flow Transformer encoder of the present invention; among them, Figure (a) is the schematic diagram of the self-attention mechanism; Figure (b) is the schematic diagram of the external attention mechanism. Detailed implementation manners

[0024] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, rather than to limit the present invention, that is, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0025] The components described and illustrated in the accompanying drawings and embodiments of the present invention can be arranged and designed in various different configurations. Therefore, the detailed description of the embodiments of the present invention provided in the following drawings is not intended to limit the scope of the claimed invention, but only represents a selected embodiment of the present invention. Based on the accompanying drawings and embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present invention.

[0026] The features and performances of the present invention are further described in detail below with reference to the embodiments.

[0027] Embodiment 1 As Figure 1 shown, the present invention provides an audio processing method, including the following steps: S1. Preprocess the audio file and convert it into a spectrogram; S2. Use a pre-constructed feature extraction model to extract features from the spectrogram to obtain feature vectors; S3. Use a pre-constructed neural network model based on a flow Transformer encoder to process the feature vectors to obtain beat activation values and strong beat activation values; S4. Based on the beat activation values and strong beat activation values, use a probability model to infer the beat sequence and strong beat sequence.

[0028] Embodiment 2 Based on Example 1, S1 is introduced in detail.

[0029] Preprocess the audio file and convert it into a spectrogram, specifically: Preprocess the input audio file through the short-time Fourier transform method and convert it into a spectrogram.

[0030] Since the beat information of the audio is closely related to frequency components, energy changes, etc., these beat-related features can be better captured through the spectrogram, preparing for subsequent feature extraction and analysis.

[0031] The principle of the short-time Fourier transform (STFT) is as follows: When processing audio signals, audio signals are usually non-stationary signals that change over time. The traditional Fourier transform can only give the overall spectral characteristics of the signal and cannot reflect the changes in frequency components at different times. The core idea of the short-time Fourier transform is to divide the audio signal into many shorter, approximately stationary time periods, and perform Fourier transform on the signals within each time period, thereby obtaining the spectral information of the signal at different times.

[0032] Example 3 Based on Example 1, S2 is introduced in detail.

[0033] The feature vector includes a time-domain feature vector and a frequency-domain feature vector.

[0034] In order to efficiently extract accurate time-domain features and frequency-domain features from the spectrogram, the present invention adopts a feature extraction method of a frequency-temporal attention (FTA) module and depthwise separable convolution.

[0035] The core principle of the FTA module is to highlight the important feature parts in the data by calculating the attention weights in the frequency domain and time domain. Specifically, it analyzes the input data in the frequency domain and time domain respectively, calculates the importance scores of each position, and then weights the data according to these scores, thereby enhancing the representation of important features and suppressing unimportant features.

[0036] Specifically, as Figure 3 shown, an audio file is preprocessed and converted into a spectrogram file of size (bs, C, T, F), where bs represents the number of samples in one training, C represents the number of channels, T represents the number of time frames, and F represents the number of frequency bins included in the spectrogram.

[0037] As Figure 3As shown, the pre - constructed feature extraction model includes a two - dimensional convolutional layer, a first max - pooling layer, a frequency - time attention module, a second max - pooling layer, and a depth - separable convolutional layer connected in sequence.

[0038] Feature extraction is performed on the spectrogram to obtain a feature vector, specifically: First, a two - dimensional convolutional layer is used to expand the number of input channels of the spectrogram to 32 channels, obtaining a multi - channel time - frequency feature vector; Then, a max - pooling layer is used to reduce the frequency - dimension size of the multi - channel time - frequency feature vector, obtaining a time - frequency feature vector with one - dimensional reduction; Next, through the FTA module, the feature vector is weighted in both the time dimension and the frequency dimension, and at the same time, the number of channels of the feature map is expanded from 32 to 64 to capture richer information in a higher - dimensional feature space, obtaining a time - frequency feature vector after frequency - time attention; Continue to use a max - pooling layer to reduce the frequency - dimension size of the time - frequency feature vector after frequency - time attention, obtaining a time - frequency feature vector with two - dimensional reduction; Finally, through the depth - separable convolutional layer, the number of channels of the time - frequency feature vector with two - dimensional reduction is expanded to 256, ensuring that the feature extraction part reaches the optimal in computational efficiency while retaining richer time - frequency features. The size of the finally output feature vector is (bs, T, 256).

[0039] Among them, the FTA module can dynamically adjust the focus of attention of the feature extraction model in both the time and frequency dimensions, enabling the subsequent neural network model to effectively extract and fuse useful spectrogram features from complex inputs, improving the accuracy of beat tracking. Depth - separable convolution can effectively reduce the parameters and computational amount of the neural network model, improving the efficiency of beat tracking.

[0040] The following are the specific structures of the frequency - time attention module and the depth - separable convolution.

[0041] The network structure of the frequency - time attention module is as Figure 4 shown, including a frequency - domain attention sub - module in the left dotted - line area in Figure 4 and a time - domain attention sub - module in the right dotted - line area in Figure 4 .

[0042] Specifically, the processing process of the frequency - domain attention sub - module is: Given the input feature map X ∈ R C×F×T , first, the input feature map X is subjected to average pooling through a max - pooling layer to calculate the distribution of the amplitude along the time axis, obtaining a frequency descriptor f ∈ R C×F : ; Among them, Denote the element at the \(i\)-th row and \(j\)-th column in the input feature map \(X\), where \(i\) and \(j\) are positive integers; \(T\) represents the number of time frames.

[0043] Then, two 1D convolutional layers are used to learn the correlations in the frequency dimension . For the frequency descriptor \(f\), the process of 1D convolution can be written as: ; where is the 1D convolutional kernel of the \(l\)-th layer, is the newly generated feature map, and \(*\) is the convolution operator.

[0044] Finally, a Softmax layer is applied to obtain the frequency attention feature map : ; .

[0045] Similarly, through the same process, the temporal attention sub-module can also obtain the temporal attention feature map on the time axis , .

[0046] Meanwhile, to learn high-level semantic features, two 2D convolutional layers with kernel sizes of \((3\times3)\) and \((5\times5)\) respectively are applied to the input feature map to obtain two new feature maps , and then and are multiplied using matrix multiplication to obtain the final output .

[0047] where ; ; .

[0048] where, broadcast is a broadcast operation that enables element-wise multiplication of matrices with different shapes to be compatible; represent the frequency feature map and the temporal feature map that have learned high-level semantic features respectively; represents the matrix multiplication operator.

[0049] The structure of the depthwise separable convolutional layer is as shown in Figure 5 . In traditional convolution operations, the convolutional kernel performs convolution operations in both the spatial dimension and the channel dimension of the input feature map. The convolutional kernel is usually a \(k\times k\) matrix. Assume the input feature map has \(C\) in channels, then the convolutional kernel has \(C\) out channels. Thus, the computational complexity of traditional convolution operations is \(O(k\)2 ×C in ×C out )。

[0050] The depthwise separable convolutional layer divides the convolution operation into two stages: First, depthwise convolution is performed. The convolutional kernel of depthwise convolution is in single-channel mode, that is, each channel of the input is processed using an independent convolutional kernel without information interaction between channels. This is a lightweight operation because the convolution computation amount for each channel is reduced; Then, pointwise convolution is performed. Pointwise convolution fuses the features obtained from depthwise convolution across channels. Through a 1×1 convolutional kernel, the information between different channels is combined to obtain a new feature map, and the number of output feature maps depends on the number of filters.

[0051] By combining these two stages, both the computation amount and the number of parameters are greatly reduced. The computation amount of depthwise convolution is O(k 2 ×C in ), and the computation amount of pointwise convolution is O(C in ×C out ). Therefore, the overall computation amount is reduced to O(k 2 ×C in +C in ×C out ), which is much less than that of traditional convolution operations.

[0052] In summary, the combination of the FTA module and the depthwise separable convolutional layer can not only retain sufficient feature expression ability but also greatly reduce the computation amount. In practical applications, especially for complex audio signals, such as multi-instrument performances, different music styles, or scenarios with strong noise interference, this feature extraction method can achieve a good balance between computational efficiency and performance and meet the requirements of real-time processing.

[0053] By introducing the frequency-time attention module and the depthwise separable convolutional layer, the present invention can more effectively capture key beat information in complex audio backgrounds and changing music rhythms. This method can significantly improve the accuracy and robustness of feature extraction while reducing the computational cost.

[0054] Embodiment 4 Based on Embodiment 1, S3 is introduced in detail: The feature vector is processed using a pre-constructed neural network model based on a flow Transformer encoder to obtain a beat activation value and a strong beat activation value.

[0055] The neural network model based on the Flow Transformer encoder includes a Flow Transformer encoder, a linear layer, and a non-linear activation function layer. The non-linear activation function layer uses the Sigmoid activation function. The Flow Transformer encoder is responsible for modeling the temporal information of the audio signal, and captures both the microscopic information and macroscopic information in the sequence through the self-attention mechanism and the external attention mechanism. Then, through a linear layer and a non-linear activation function layer, further feature transformation is performed to finally obtain the activation values of the beats and strong beats.

[0056] The present invention uses a neural network model based on the Flow Transformer encoder to obtain the activation values of the beats and strong beats. The Transformer module for offline beat tracking directly processes the entire input sequence, while the online beat tracking of the present invention is based on a context chunk processing mechanism, and outputs are generated by providing partial audio frames.

[0057] Figure 6 is the basic structure of the context chunk processing mechanism. First, the time-frequency feature vector is used as the input feature, and the sequence is segmented into multiple non-overlapping blocks C = {C 1 , C 2 , …, C b , …}, where b is the index of the block, and each block contains N c frames. To alleviate the block boundary effect, that is, the problem of broken context information between blocks, each block C b is extended to a context block: (1) the left sub-block L b : the N b frames in front of C l ; (2) C b ; (3) the right sub-block R b : the N b frames behind C r . Connect them to form a context block [L b , C b , R b as the input to the Flow Transformer encoder. Figure 6 is an example of the case where N c = 2, N l = 1, N r = 1. Among them, [L b-1 , C b-1 , C b-1 , R b-1 represents the input feature vector of the (b - 1)-th block, and [L b , C b , C b , R b represents the input feature vector of the b-th block, [Lb+1 , C b+1 , C b+1 , R b+1 represents the input feature vector of the (b + 1)-th block, Encoder Layer represents an encoder layer, and Z b-1 represents the output of the (b - 1)-th block, and Z b represents the output of the b-th block, and Z b+1 represents the output of the (b + 1)-th block. To capture the long-term dependencies of the input sequence, each block C b also inherits a context embedding vector c b during its processing, which is obtained from the previous block (i.e., the (b - 1)-th block). This mechanism can transmit information across blocks, thereby alleviating the context break problem between blocks. In the figure, c b-1 represents the context embedding vector of block C b-1 , and c b+1 represents the context embedding vector of block C b+1 . In the network of the present invention, N c , N l , N r are respectively 16, 256, 16. The encoder layer uses an improved flow Transformer encoder, that is, the network structure shown in Figure 7 .

[0058] After the time-frequency feature vector is block-processed, it enters the calculation of the attention feature vector. As shown in Figure 7 , the flow Transformer encoder includes a sequentially connected attention layer, a splicing layer, a fully connected layer, and a feed-forward layer. The attention layer is divided into a multi-head self-attention layer and an external attention layer. In each block, the original multi-head self-attention layer with relative position encoding is retained, and at the same time, an external attention layer is introduced to calculate the attention feature vectors respectively. Then, the two attention feature vectors are spliced through the splicing layer to obtain the spliced feature vector; a fully connected layer is used to restore the spliced feature vector to its original dimension size, obtaining a feature vector with the same dimension as the input feature vector of the flow Transformer encoder; then the deep information of the restored feature vector is extracted by the feed-forward layer, and finally, the deep information is fused through a linear layer and a non-linear activation function layer to obtain the beat activation value and the strong beat activation value. Among them, the multi-head self-attention layer mainly focuses on the dependencies within the input sequence at the micro level and models local context information; the external attention layer can establish connections between different time frames of the audio signal and macroscopically learn which time steps are most important for the current prediction. The combination of the two can make full use of local information and improve the understanding ability of the neural network model for the beat features in the audio signal.

[0059] Traditional self-attention calculation methods rely on generating an attention map by calculating the correlation between query vectors and key vectors. As shown in Figure 8 Figure (a) in

[0060] this attention map reflects the degree of attention of each input feature to other features. Subsequently, these attention weights are applied to the value vectors to obtain a weighted feature map. Figure 8 However, the external attention mechanism works differently. As shown in Figure (b) in it first generates an attention map by calculating the correlation between the query vector and the externally learnable key attention matrix

[0061] and then multiplies this attention map by the externally learnable value attention matrix

[0062] to finally generate a more refined feature map as the output feature. In this mechanism, the externally learnable attention matrices provide an additional source of information during attention calculation, enabling the model to model the dependencies between different time steps. k Specifically, assuming the size of the input features is [bs, T, dmodel], where dmodel represents the input feature dimension and takes 256, it first passes through a fully connected layer M k and the feature dimension changes from dmodel to S (S is the reduced feature dimension and takes 64). By M v the 256-dimensional input features can be compressed into a feature space of size 64 to extract information in the low-dimensional space, so that the neural network model can focus on more critical beat information without being disturbed by excessive redundant audio features. Then, a Softmax operation is applied on the second dimension (i.e., the time dimension) of the feature vector to normalize all S-dimensional features at each time step, so that each time step will obtain a clear attention allocation weight, which can distinguish which features contribute more to the rhythm at different time steps. Then it passes through another fully connected layer M

[0063] External attention only calculates attention through simple linear projection and softmax operations. This lightweight calculation can improve accuracy without sacrificing efficiency. Coupled with the accurate capture of local feature information by self-attention, the model can comprehensively understand the input features.

[0064] Embodiment 5 Based on Embodiment 1, S4 is introduced in detail: According to the beat activation value and the strong beat activation value, a probability model is used to infer the final beat sequence and strong beat sequence.

[0065] In the post-processing stage, the output of the neural network model based on the flow Transformer encoder will be optimized by an online Dynamic Bayesian Network (DBN).

[0066] Dynamic Bayesian Networks are usually used to process time series data and can smooth and correct the results of beat detection to improve stability and accuracy. An online Dynamic Bayesian Network is a probabilistic graphical model suitable for dynamic time series data, which can process real-time data streams in an updated manner.

[0067] At each time step, the online DBN only processes the data at the current moment and makes predictions based on the previous historical states. When the feature vector passes through the processing of the neural network model, the probability values of each frame being a beat, a strong beat, or a non-beat can be obtained. Through further processing by the online DBN, the beat (strong beat) tracking results of the entire time series can be obtained.

[0068] Optimize the model by combining the loss function, and verify and test the feature extraction model and the neural network model.

[0069] The cross-entropy loss function is a commonly used loss function in classification tasks. Its calculation formula is: ; where C represents the number of categories, y i is an element of the true label, which is 1 for the target category and 0 for the remaining categories, and p i is the predicted probability of the model for the i-th category, usually calculated through the softmax function.

[0070] In real music signals, beat tracking usually faces problems such as complex rhythm changes, audio noise, or irregular music rhythms. If hard labels are used for the classification labels of beats and strong beats (for example, directly setting the labels to 1 or 0), then the model may be overly confident in the classification of certain time steps, which can lead to its inability to adapt to rhythm irregularities or noise. Since the online beat tracking task cannot modify the data that has already been predicted, this method of calculating loss is not suitable for online tasks. The present invention proposes to use a cross-entropy loss function with label smoothing to reduce the overfitting of the model and enhance the robustness of the model.

[0071] Label-smoothing cross-entropy loss is a modification of the traditional cross-entropy loss function, aiming to reduce the model's overconfidence in a certain class. In label smoothing, the target label is "smoothed" so that the label value of the target class drops from 1 to and the label of the non-target class rises from 0 to , where is the label-smoothing factor, usually taking a very small value (for example, 0.1). The formula for the label-smoothing cross-entropy loss function is as follows: ; where, is the label after label smoothing: ; where, represents the target class, including three classes: beats, strong beats, and non-beats; C represents the number of classes, taking C = 3; p i represents the prediction probability of the online dynamic Bayesian network for the i-th class, where i is a specific class; represents the label-smoothing factor; log() represents the logarithmic function.

[0072] In this way, label smoothing reduces the absolute weight of the target class label being 1 and gives a very small non-zero value to the non-target class. This change avoids the overconfidence of the online dynamic Bayesian network in a certain class during the training process and makes the prediction probability smoother.

[0073] Label smoothing enables the model to maintain a certain degree of flexibility and robustness when faced with label noise or mislabeling, and adjust and make predictions more smoothly. If there are some beats mislabeled or the positions of some beats are not clear enough in the training data, the traditional cross-entropy loss will cause the model to overfit these uncertain or inaccurate labels. However, the loss function with label smoothing can effectively alleviate the overfitting of the model to incorrect labels by reducing the "hard" constraints of the labels, and improve the generalization ability of the model. Moreover, when the dataset is small or the class is imbalanced, label smoothing can also make the prediction distribution of the model more gentle, prompting the model to pay more attention to the features of each class and enhancing the robustness of the model.

[0074] Based on the existing Ballroom dataset, SMC dataset, Hainsworth dataset, and GTZAN dataset, the present invention uses a specially designed loss function to train and iterate the feature extraction model and the neural network model, and finally obtains a robust model.

[0075] Embodiment 6 As Figure 2 shown, the present invention also discloses an audio processing system, including: A preprocessing module for preprocessing the audio file and converting it into a spectrogram; A feature extraction module for extracting features from the spectrogram using a pre-constructed feature extraction model to obtain feature vectors; An activation value calculation module for processing the feature vectors using a pre-constructed neural network model to obtain beat activation values and downbeat activation values; A prediction module for inferring the beat sequence and the downbeat sequence based on the beat activation values and the downbeat activation values using a probability model.

[0076] Embodiment 7 The present invention also discloses a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the steps of the audio processing method. Among them, the memory may include a memory, such as a high-speed random access memory, and may also include a non-volatile memory, such as at least one disk memory, etc.; the processor, the network interface, and the memory are interconnected through an internal bus, and this internal bus can be an Industry Standard Architecture bus, a Peripheral Component Interconnect standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program can include program code, and the program code includes computer operation instructions. The memory can include a memory and a non-volatile memory, and provides instructions and data to the processor.

[0077] Embodiment 8 The present invention also discloses a computer-readable storage medium storing a computer program, which when executed by a processor implements the steps of the audio processing method. Specifically, the computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include random access memory and / or cache memory, etc. The non-volatile memory may include read-only memory, hard disk, flash memory, optical disc, magnetic disk, etc.

[0078] Those skilled in the art should understand that the embodiments of the present invention may be provided as a method, a system, or a computer program product. Therefore, the present invention may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0079] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks.

[0080] These computer program instructions can also be stored in a computer-readable memory capable of guiding a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks.

[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are performed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks.

[0082] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: modifications or equivalent substitutions can still be made to the specific implementation manners of the present invention, and any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the protection scope of the present invention.

Claims

1. An audio processing method, characterized in that: The following steps are involved: S1, preprocess the audio file and convert it into a spectrogram; S2, using a pre-built feature extraction model to extract features from the spectrum graph to obtain a feature vector; S3, using a pre-built neural network model based on a streaming Transformer encoder to process the feature vector to obtain a beat activation value and a strong beat activation value; S4. Based on the beat activation value and the strong beat activation value, a probability model is used to infer the beat sequence and the strong beat sequence.

2. The audio processing method according to claim 1, characterized in that: S1 is specifically: The input audio file is preprocessed by short-time Fourier transform method and converted into a spectrogram.

3. The audio processing method according to claim 1, characterized in that: In S2, the pre-built feature extraction model includes a two-dimensional convolutional layer, a first maximum pooling layer, a frequency-time attention module, a second maximum pooling layer, and a depth-separable convolutional layer connected in sequence; The feature extraction process is as follows: S2.1, using a two-dimensional convolutional layer to perform preliminary channel expansion on the spectrum map to obtain a multi-channel time-frequency feature vector; S2.2, using the first maximum pooling layer to reduce the frequency dimension of the multi-channel time-frequency feature vector to obtain a time-frequency feature vector with a reduced dimension; S2.3, use the frequency-time attention module to weight the time-frequency feature vector after the first dimension reduction in the time dimension and the frequency dimension, and expand the number of channels again to obtain the time-frequency feature vector after the frequency-time attention; S2.4, using the second maximum pooling layer to reduce the frequency dimension of the time-frequency feature vector after the frequency-reduction attention, and obtaining the second dimension-reduced time-frequency feature vector; S2.

5. Use a deep separable convolutional layer to expand the number of channels of the second dimensionally reduced time-frequency feature vector to obtain a feature vector.

4. The audio processing method according to claim 1, characterized in that: In S3, the pre-built neural network model based on the stream Transformer encoder includes a stream Transformer encoder, a linear layer, and a non-linear activation function layer connected in sequence; The stream Transformer encoder includes an attention layer, a concatenation layer, a fully connected layer, and a feedforward layer connected in sequence, and the attention layer is divided into a multi-head self-attention layer and an external attention layer; The multi-head self-attention layer and the external attention layer calculate the attention feature vectors respectively, and the two attention feature vectors are concatenated through the concatenation layer to obtain the concatenated feature vector; The concatenated feature vector is restored through a fully connected layer to obtain a restored feature vector; the restored feature vector has the same dimension as the input feature vector of the stream Transformer encoder; The deep information of the restored feature vector is then extracted by the feedforward layer, and finally the deep information is fused through the linear layer and the nonlinear activation function layer to obtain the beat activation value and the strong beat activation value.

5. The audio processing method according to claim 1, characterized in that: In S4, the probability model adopts an online dynamic Bayesian network.

6. The audio processing method according to claim 1, characterized in that: The loss function of the feature extraction model and the neural network model is a label-smoothed cross entropy loss function, and the formula of the label-smoothed cross entropy loss function is as follows: ; in, is the label after label smoothing; is the loss value calculated by the cross entropy loss function with label smoothing; ; in, represents the target category, including beat, strong beat and off-beat; C represents the number of categories; p i represents the predicted probability of the online dynamic Bayesian network for the i-th category, where i is the specific category; Represents the label smoothing factor; log() represents the logarithmic function.

7. An audio processing system, characterized in that: include: A preprocessing module is used to preprocess the audio file and convert it into a spectrogram; A feature extraction module is used to extract features from the spectrum graph using a pre-built feature extraction model to obtain a feature vector; An activation value calculation module is used to process the feature vector using a pre-built neural network model based on a streaming Transformer encoder to obtain a beat activation value and a strong beat activation value; The prediction module is used to infer the beat sequence and the strong beat sequence using a probability model based on the beat activation value and the strong beat activation value.

8. An audio processing system according to claim 7, characterized in that: The feature extraction module includes a two-dimensional convolutional layer, a first maximum pooling layer, a frequency-time attention module, a second maximum pooling layer, and a depth-wise separable convolutional layer connected in sequence; The two-dimensional convolution layer is used to perform preliminary channel expansion on the spectrum graph to obtain a multi-channel time-frequency feature vector; The first maximum pooling layer is used to reduce the frequency dimension of the multi-channel time-frequency feature vector to obtain a time-frequency feature vector with a reduced dimension; The frequency-time attention module is used to weight the time-frequency feature vector after the first dimension reduction in the time dimension and the frequency dimension, and expand the number of channels again to obtain the time-frequency feature vector after the frequency-time attention; The second maximum pooling layer is used to reduce the frequency dimension of the time-frequency feature vector after frequency-time attention to obtain a second-dimensional reduced time-frequency feature vector; The depthwise separable convolutional layer is used to further expand the number of channels of the second dimensionally reduced time-frequency feature vector to obtain a feature vector.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the computer program, the steps of the audio processing method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the audio processing method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Single-channel human voice and background voice separation method based on convolutional recurrent neural network

    CN112259120A

  • Training method of beat re-shooting joint detection model and beat re-shooting joint detection method

    CN114154574A

  • Training method of beat re-shooting joint detection model and beat re-shooting joint detection method

    CN114897157A

  • Audio-visual voice noise reduction method based on multi-mode gating lifting model

    CN116013297A

  • Deep network line spectrum detection method embedded with attention mechanism

    CN118568459A