Music Genre Classification Method and Device Based on Self-Supervised and Hybrid Attention Mechanisms

By adopting the self-supervised and mixed attention mechanism methods in music genre classification, a music genre classification model including feature extraction, mixed attention mechanism and deep modeling layer was constructed, and the problems of insufficient accuracy and insufficient robustness in the existing technology were solved, and more efficient music genre classification was achieved.

CN119832883BActive Publication Date: 2025-07-01SHAANXI MINGYUAN CULTURE & ART COMMUNICATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510295073.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-07-01
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

The prior art has problems of insufficient precision, overfitting and insufficient robustness in music genre classification, especially when dealing with music data with varying styles and large data scale.

Method used

A music genre classification method based on self-supervised and mixed attention mechanism is adopted to construct a music genre classification model including feature extraction layer, mixed attention mechanism layer, depth modeling layer and classification layer. The feature extraction layer uses the self-supervised pre-trained wav2vec2.0 model, the hybrid attention mechanism layer combines the multi-head self-attention and channel attention mechanism, and the deep modeling layer uses a stacked Transformer encoding layer and a global average pooling layer.

Benefits of technology

It improves the accuracy and applicability of music genre classification, can more effectively capture the characteristics of complex music segments, reduce overfitting and improve the robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119832883B_ABST
    Figure CN119832883B_ABST
Patent Text Reader

Abstract

The present invention discloses a music genre classification method and device based on self-supervision and hybrid attention mechanism, which relates to the field of music classification, and includes: constructing a music genre classification model and training it to obtain a trained music genre classification model; obtaining audio data to be classified, inputting the audio data to be classified into the trained music genre classification model, the audio data to be classified passes through a feature extraction layer, and a corresponding feature sequence is extracted; inputting the feature sequence into a hybrid attention mechanism layer to obtain hybrid features; inputting the hybrid features into a deep modeling layer, first performing context-dependent modeling through several Transformer encoding layers, and then inputting the output features of the last Transformer encoding layer into a global average pooling layer to obtain a modeling vector; inputting the modeling vector into a classification layer to obtain an audio genre prediction label corresponding to the audio data to be classified. The present invention solves the problems of low accuracy and weak applicability in the existing music genre classification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of music classification, and particularly relates to a music genre classification method and device based on self-supervised and hybrid attention mechanisms. Background Art

[0002] There are diverse music genres, including but not limited to pop music, classical music, jazz, folk music, electronic music, etc. Each genre has its unique instrument configuration, timbre characteristics, rhythm patterns, and performance techniques. The classification of music genres occupies an important position in the field of Music Information Retrieval (MIR), and its applications are extensive, covering multiple aspects such as rhythm detection, music composition, personalized recommendation systems, audio track separation, and instrument recognition.

[0003] With the progress of deep learning theory and computing hardware, researchers have widely attempted to use machine learning algorithms (such as random forests, support vector machines, etc.) and deep learning models (such as convolutional neural network CNN, recurrent neural network RNN, Transformer, etc.) to extract features and model music in the field of music genre classification. Some research works convert audio signals into Mel spectrograms or spectrograms and then apply deep convolutional networks for learning; there are also literatures that use self-attention mechanisms to capture the context information of music sequences. However, for music, which is a highly complex and multi-level signal, traditional or single deep network architectures still face certain limitations in practical applications: on the one hand, music is extremely diverse in dimensions such as melody, harmony, and arrangement, and simple feature extraction is difficult to fully exploit the latent high-order semantic information; on the other hand, if only relying on a single path for modeling, it is often difficult to balance time series information and frequency domain structure, resulting in insufficient accuracy in the recognition of certain styles. This defect is particularly obvious in new types of music with multi-genre blends.

[0004] In order to further improve the accuracy and generalization ability of music genre classification, new ideas such as multi-branch network fusion, attention mechanisms, and data augmentation have emerged in recent years. However, the existing methods still have the following pain points: First, most attention modules are single self-attention or channel attention, making it difficult to flexibly interact between different dimensions and scales; Second, the data augmentation strategies are relatively limited, mostly relying only on random segmentation or simple sampling, and unable to fully utilize time domain, frequency domain, and even multi-modal information; Third, the modeling of audio data still largely depends on supervised data, and when the annotations are insufficient or the distribution deviation is large, the model is prone to overfitting or insufficient robustness. Summary of the Invention

[0005] The purpose of this application is to propose a music genre classification method and device based on self-supervised and hybrid attention mechanisms for the above-mentioned technical problems, aiming to overcome the deficiencies faced by the prior art in dealing with music data with diverse styles, large data scale, and insufficient annotation, and to provide a solution with higher accuracy and stronger applicability for the music genre classification task.

[0006] In the first aspect, the present invention provides a music genre classification method based on self-supervised and hybrid attention mechanisms, including the following steps:

[0007] Construct and train a music genre classification model to obtain a trained music genre classification model. The trained music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer, and a classification layer connected in sequence. The feature extraction layer uses the wav2vec2.0 model pre-trained by self-supervision. The deep modeling layer includes a number of Transformer encoding layers stacked in sequence and a global average pooling layer;

[0008] Obtain the audio data to be classified, and input the audio data to be classified into the trained music genre classification model. The audio data to be classified passes through the feature extraction layer to extract the corresponding feature sequence; the feature sequence is input into the hybrid attention mechanism layer to obtain a hybrid feature; the hybrid feature is input into the deep modeling layer, first undergoes context-dependent modeling through a number of Transformer encoding layers, and then the output feature of the last Transformer encoding layer is input into the global average pooling layer to obtain a modeling vector; the modeling vector is input into the classification layer to obtain the audio genre prediction label corresponding to the audio data to be classified.

[0009] Preferably, the hybrid attention mechanism layer includes a multi-head self-attention layer, a first normalization layer, a channel attention layer, and a second normalization layer. The feature sequence is input into the hybrid attention mechanism layer, first passes through the multi-head self-attention layer to obtain the output feature of the multi-head self-attention layer. The feature sequence is added to the output feature of the multi-head self-attention layer to obtain a first addition result, and then the first addition result is input into the first normalization layer to obtain an intermediate feature. The intermediate feature passes through the channel attention layer to obtain the output feature of the channel attention layer. The intermediate feature and the output feature of the channel attention layer are weighted and fused to obtain a first fusion feature. The feature sequence and the first fusion feature are added to obtain a second addition result, and then the second addition result is input into the second normalization layer to obtain a hybrid feature.

[0010] Preferably, the Transformer encoding layer includes a local-global attention unit, a feed-forward network unit, and a third layer normalization layer. The local-global attention unit includes a local attention module, a global attention module, and a fourth layer normalization layer. The feed-forward network unit includes a first fully-connected layer and a second fully-connected layer. The ReLU activation function is used in the first fully-connected layer. When the first Transformer encoding layer is the current Transformer encoding layer, the mixed features are the output features of the previous Transformer encoding layer. The output features of the previous Transformer encoding layer are input into the current Transformer encoding layer. First, they pass through the local attention module and the global attention module of the local-global attention unit of the current Transformer encoding layer respectively to obtain the output features of the local attention module and the output features of the global attention module. The output features of the local attention module and the output features of the global attention module are weighted and fused to obtain the second fused feature. The output features of the previous Transformer encoding layer are added to the second fused feature to obtain the third addition result. Then, the third addition result is input into the fourth layer normalization layer to obtain the output features of the local-global attention unit of the current Transformer encoding layer, specifically including:

[0011] First, perform the query, key, and value operations in the local attention module, as shown in the following formula:

[0012] , , ;

[0013] Among them, , and are the query weight, key weight, and value weight in the local attention module respectively. , and are the query matrix, key matrix, and value matrix in the local attention module respectively. represents the output features of the (l - 1)-th Transformer encoding layer;

[0014] Calculate the local attention weight using the following formula :

[0015] ;

[0016] Among them, represents the dimension of the key matrix in the local attention module. is the local window centered on the t-th time step, and s is the s-th time step within the local window. denotes the key matrix in the local attention module at the $s$-th time step within the local window, denotes the query matrix in the local attention module at the $t$-th time step;

[0017] The output feature of the local attention module at the $t$-th time step is calculated using the following formula :

[0018] ;

[0019] where, denotes the value matrix in the local attention module at the $s$-th time step within the local window;

[0020] Secondly, perform the query, key, and value operations in the global attention module, as shown in the following formula:

[0021] , , ;

[0022] where, , and are the query weight, key weight, and value weight in the global attention module respectively, , and are the query matrix, key matrix, and value matrix in the global attention module respectively;

[0023] The global attention weight is calculated using the following formula :

[0024] ;

[0025] where, denotes the dimension of the key matrix in the global attention module, $u$ represents the $u$-th time step, denotes the key matrix in the global attention module at the $u$-th time step, denotes the query matrix in the global attention module at the $t$-th time step;

[0026] The output feature of the local attention module at the $t$-th time step is calculated using the following formula :

[0027] ;

[0028] where, denotes the value matrix in the local attention module at the $u$-th time step, $T$ represents the total number of time steps;

[0029] Fuse the output features of the local attention module at the t-th time step and the output features of the global attention module at the t-th time step to obtain the second fused feature at the t-th time step, as shown in the following formula:

[0030] ;

[0031] Among them, represents the second fused feature at the t-th time step, represents the weighted weight;

[0032] Add the second fused feature at the t-th time step and perform residual addition, and pass through the fourth layer normalization layer to obtain the output features of the local-global attention unit of the l-th Transformer encoding layer , as shown in the following formula:

[0033] ;

[0034] Among them, represents the layer normalization operation;

[0035] Input the output features of the local-global attention unit of the current Transformer encoding layer into the feed-forward network unit of the current Transformer encoding layer, and successively pass through the first fully connected layer and the second fully connected layer to obtain the output features of the feed-forward network unit of the current Transformer encoding layer. Input the output features of the feed-forward network unit of the current Transformer encoding layer into the third layer normalization layer of the current Transformer encoding layer to obtain the output features of the current Transformer encoding layer. Take the output features of the current Transformer encoding layer as the output features of the previous Transformer encoding layer and repeat the calculation process of the current Transformer encoding layer above until the output features of the last Transformer encoding layer are obtained.

[0036] Preferably, the classification layer includes a Dropout layer, a third fully connected layer, a fourth fully connected layer, a Softmax function layer, and an argmax function layer connected in sequence. The ReLU activation function is used in the third fully connected layer; the modeling vector passes through the Dropout layer to obtain an adjusted modeling vector. The adjusted modeling vector successively passes through the third fully connected layer, the fourth fully connected layer, and the Softmax function layer to obtain the predicted probability distribution of each music genre. Input the predicted probability distribution of each music genre into the argmax function layer to determine that the music genre corresponding to the maximum predicted probability is the predicted label of the audio genre corresponding to the audio data to be classified.

[0037] Preferably, the wav2vec2.0 model pre-trained by self-supervision includes an encoding network and a context network. During the training process of the music genre classification model, the parameters of the encoding network are fixed, and the parameters of the context network are adjusted. The loss function used during the training process of the music genre classification model is:

[0038] ;

[0039] where represents the loss function, N represents the total number of samples in the training data, q represents the q-th sample, c represents the c-th audio genre, C represents the total number of audio genres, represents the weight of the c-th audio genre, represents the predicted probability that the q-th sample belongs to the c-th audio genre, represents the true label indicating whether the q-th sample belongs to the c-th audio genre.

[0040] Preferably, the process of obtaining the training data used during the training process of the music genre classification model is as follows:

[0041] Collect the original audio data with the true labels of audio genres;

[0042] Randomly crop the original audio data in the time domain to obtain several audio segments;

[0043] Randomly mix and add noise to each audio segment to obtain the noise-added audio segments;

[0044] Perform time-frequency occlusion on the noise-added audio segments to obtain the enhanced audio data;

[0045] Construct the training data from the enhanced audio data and its corresponding true labels of audio genres.

[0046] In a second aspect, the present invention provides a music genre classification device based on self-supervision and hybrid attention mechanism, including:

[0047] A model construction module, configured to construct and train a music genre classification model to obtain a trained music genre classification model. The trained music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer, and a classification layer connected in sequence. The feature extraction layer uses the wav2vec2.0 model pre-trained by self-supervision. The deep modeling layer includes several Transformer encoding layers stacked in sequence and a global average pooling layer;

[0048] A classification module, configured to obtain audio data to be classified, input the audio data to be classified into a trained music genre classification model, the audio data to be classified passes through a feature extraction layer to extract a corresponding feature sequence; input the feature sequence into a hybrid attention mechanism layer to obtain hybrid features; input the hybrid features into a deep modeling layer, first perform context-dependent modeling through several Transformer encoding layers, and then input the output features of the last Transformer encoding layer into a global average pooling layer to obtain a modeling vector; input the modeling vector into a classification layer to obtain an audio genre prediction label corresponding to the audio data to be classified.

[0049] In a third aspect, the present invention provides an electronic device, including one or more processors; a storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, enabling the one or more processors to implement the method described in any implementation manner of the first aspect.

[0050] In a fourth aspect, the present invention provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0051] In a fifth aspect, the present invention provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the method described in any implementation manner of the first aspect.

[0052] Compared with the prior art, the present invention has the following beneficial effects:

[0053] (1) The music genre classification method based on self-supervised and hybrid attention mechanism proposed by the present invention constructs a music genre classification model, and uses the hybrid attention mechanism layer that integrates the multi-head self-attention mechanism and the channel attention mechanism in the music genre classification model to perform feature weighting and fusion in different dimensions, which can capture richer information to improve the discrimination of complex music segments, and further uses the deep modeling layer to perform further modeling and deep semantic refinement on the hybrid features output by the hybrid attention mechanism layer.

[0054] (2) The music genre classification method based on self-supervised and hybrid attention mechanism proposed by the present invention uses the wav2vec2.0 model pre-trained by self-supervision, which can learn general representations on a large amount of unlabeled audio data, thereby providing richer and more robust features for downstream music genre classification and reducing the training difficulty.

[0055] (3) The music genre classification method based on self-supervised and hybrid attention mechanism proposed by the present invention further adopts a diversified data augmentation technique to construct training data, which helps the music genre classification model capture the feature distributions under various time-frequency transformations and noise conditions during the training process, and can effectively improve the applicability and classification accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0057] Figure 1 It is a schematic flowchart of the music genre classification method based on self-supervised and hybrid attention mechanism according to the embodiment of the present application;

[0058] Figure 2 It is a schematic diagram of the music genre classification model of the music genre classification method based on self-supervised and hybrid attention mechanism according to the embodiment of the present application;

[0059] Figure 3 It is a schematic diagram of the hybrid attention mechanism layer of the music genre classification method based on self-supervised and hybrid attention mechanism according to the embodiment of the present application;

[0060] Figure 4 It is a schematic diagram of the Transformer encoding layer of the music genre classification method based on self-supervised and hybrid attention mechanism according to the embodiment of the present application;

[0061] Figure 5 It is a schematic diagram of the music genre classification device based on self-supervised and hybrid attention mechanism according to the embodiment of the present application;

[0062] Figure 6 It is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0063] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the drawings. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0064] Figure 1 It shows a music genre classification method based on self-supervised and hybrid attention mechanism provided by the embodiment of the present application, including the following steps:

[0065] S1. Construct a music genre classification model and train it to obtain a trained music genre classification model. The trained music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer, and a classification layer connected in sequence. The feature extraction layer uses the wav2vec2.0 model pre-trained with self-supervision. The deep modeling layer includes a number of stacked Transformer encoding layers and a global average pooling layer connected in sequence.

[0066] In a specific embodiment, the acquisition process of the training data used in the training process of the music genre classification model is as follows:

[0067] Collect the original audio data with the true labels of audio genres.

[0068] Perform random time-domain cropping on the original audio data to obtain a number of audio segments.

[0069] Perform random mixing and noise addition on each audio segment to obtain the noise-added audio segments.

[0070] Perform time-frequency masking on the noise-added audio segments to obtain the enhanced audio data.

[0071] Construct the training data by combining the enhanced audio data with its corresponding true labels of audio genres.

[0072] Specifically, when constructing the training data used in the training process of the music genre classification model, first perform enhancement processing on the original audio data, specifically including random time-domain cropping, random mixing and noise addition, and time-frequency masking (SpecAugment). By significantly increasing the number of training samples while ensuring the true labels of audio genres remain unchanged, the adaptability of the model to different timbres, noise environments, and music segments is improved. The specific steps are as follows:

[0073] (1) Random time-domain cropping: Sample the original audio data into equal-length segments, and set the sampling rate (sr) to 22050 Hz. The specific method is to randomly select a window of 3 seconds in the original audio data, and select five non-overlapping random positions , and intercept a segment of length at each position to obtain a number of audio segments . Denote the original audio data as , T' represents the maximum time of the original audio data, represents the set of real numbers, then there is:

[0074] ;

[0075] Among them, represents the i-th audio segment, representing the audio data intercepted at a random position in the original audio data as the starting position, with a length of of the audio data, , , . Under the premise of ensuring that the true label of the audio genre remains unchanged, this process generates more training samples.

[0076] (2) Random mixing and noise addition: On the intercepted audio segment , different types or intensities of noise n are superimposed to form a noise-added audio segment, as shown in the following formula:

[0077] ;

[0078] where, represents the i-th noise-added audio segment, is the noise intensity control parameter, which can be adjusted according to the noise type and the target signal-to-noise ratio (SNR). This step can effectively improve the model's adaptability to various environmental noises.

[0079] (3) Time-frequency masking (SpecAugment): To simulate a similar effect of SpecAugment at the time-domain waveform level, a time-domain interval can be randomly selected, and the values in this interval are set to 0. Mathematically, it can be expressed as:

[0080] ;

[0081] where, represents the value of the i-th enhanced audio data at the -th time step, represents the value of the i-th noise-added audio segment at the -th time step. The enhanced audio data contains the effects of multiple enhancements such as random cropping, noise addition, and time-domain masking. Finally, all the enhanced audio data and their true labels y of the audio genre are made into training data, and the music genre classification model is trained using this training data. Since the wav2vec2.0 model is directly used to model the time-domain waveform in the subsequent process, the enhanced audio data is used as the input data of the music genre classification model during the training process.

[0082] Reference Figure 2 , the music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer, and a classification layer. First, the feature extraction layer is constructed, and the feature extraction layer uses the wav2vec2.0 model pre-trained with self-supervision , the model has been pre-trained on a large amount of unlabeled audio data and has the ability to deeply represent the context information of audio sequences.

[0083] Input the enhanced audio data into the wav2vec2.0 model. Using the context representation learned from a large amount of unlabeled audio, perform deep feature extraction on the input audio data to obtain a more robust and generalization-capable feature sequence. The specific steps include: input the i-th enhanced audio data into the encoding network of the wav2vec2.0 model without manually extracting Mel spectrograms or other traditional audio features; use the encoding network (encoder network) of the wav2vec2.0 model composed of multiple convolutional layers to capture short-term features in the time domain and output a latent representation sequence. Since it is still desired to retain the "self-supervised" property during the fine-tuning stage, mask a part of the latent representation sequence to obtain a masked representation sequence, enabling the model to infer the masked frames based on the unmasked information. Input the masked representation sequence into the context network of the wav2vec2.0 model. The context network adopts a transformer structure and outputs a context representation sequence, which is used as the feature sequence output by the feature extraction layer.

[0084] Secondly, construct a hybrid attention mechanism layer that combines multi-head self-attention and channel attention to jointly model the time series dimension and feature channel dimension of the feature sequence output from the feature extraction layer. On the one hand, capture long-range dependencies in the feature sequence; on the other hand, focus on key frequency bands or instrument parts to improve the discrimination of complex music segments.

[0085] Thirdly, construct a deep modeling layer to further perform sequence modeling and deep semantic refinement on the hybrid features output by the hybrid attention mechanism layer, and take into account the simplicity and efficiency of the network design during network design to ensure that the computing power requirements are relatively controllable during deployment.

[0086] Finally, construct a classification layer, input the modeling vector output by the deep modeling layer into the classification layer to obtain the corresponding audio genre prediction label.

[0087] In a specific embodiment, the wav2vec2.0 model pre-trained with self-supervision includes an encoding network and a context network. During the training process of the music genre classification model, the parameters of the encoding network are fixed, and the parameters of the context network are adjusted. The loss function used during the training process of the music genre classification model is:

[0088] ;

[0089] Among them, Let \(L\) denote the loss function, \(N\) denote the total number of samples in the training data, \(q\) denote the \(q\)-th sample, \(c\) denote the \(c\)-th audio genre, and \(C\) denote the total number of audio genres. Let \(w_c\) denote the weight of the \(c\)-th audio genre. Let \(p_{q,c}\) denote the predicted probability that the \(q\)-th sample belongs to the \(c\)-th audio genre. Let \(y_{q,c}\) denote the true label indicating whether the \(q\)-th sample belongs to the \(c\)-th audio genre.

[0090] Specifically, the loss function is constructed based on the true labels and predicted labels of audio genres. The music genre classification model is trained based on the above loss function, and optimizers such as Stochastic Gradient Descent or Adam are used to update the parameters, thereby training the model to obtain a trained music genre classification model. During the training process, since the feature extraction layer uses the self-supervised pre-trained wav2vec2.0 model, the parameters of the encoding network in the wav2vec2.0 model are frozen, and the parameters of the context network in the wav2vec2.0 model are adjusted during the training of the music genre classification model.

[0091] S2: Obtain the audio data to be classified, and input the audio data to be classified into the trained music genre classification model. The audio data to be classified passes through the feature extraction layer to extract the corresponding feature sequence; the feature sequence is input into the hybrid attention mechanism layer to obtain the hybrid features; the hybrid features are input into the deep modeling layer, first passing through several Transformer encoding layers for context-dependent modeling, and then the output features of the last Transformer encoding layer are input into the global average pooling layer to obtain the modeling vector; the modeling vector is input into the classification layer to obtain the predicted label of the audio genre corresponding to the audio data to be classified.

[0092] Specifically, the trained music genre classification model is deployed. In the inference stage, the obtained audio data to be classified is input into the trained music genre classification model. The audio data to be classified is input into the feature extraction layer of the wav2vec2.0 model to extract the corresponding feature sequence. The feature sequence then passes through the hybrid attention mechanism layer and the deep modeling layer to obtain the modeling vector. The modeling vector is input into the classification layer to obtain the corresponding predicted label of the audio genre.

[0093] In a specific embodiment, the hybrid attention mechanism layer includes a multi-head self-attention layer, a first normalization layer, a channel attention layer, and a second normalization layer. The feature sequence is input into the hybrid attention mechanism layer. First, it passes through the multi-head self-attention layer to obtain the output features of the multi-head self-attention layer. The feature sequence is added to the output features of the multi-head self-attention layer to obtain a first addition result. Then, the first addition result is input into the first normalization layer to obtain intermediate features. The intermediate features pass through the channel attention layer to obtain the output features of the channel attention layer. The intermediate features and the output features of the channel attention layer are weighted and fused to obtain a first fusion feature. The feature sequence and the first fusion feature are added to obtain a second addition result. Then, the second addition result is input into the second normalization layer to obtain hybrid features.

[0094] Specifically, referring to Figure 3 , the multi-head self-attention layer adopts a multi-head self-attention mechanism (MHA). The calculation process of the multi-head self-attention mechanism is an existing calculation process, and its details will not be elaborated here. The output features of the multi-head self-attention layer are subjected to residual connection with the feature sequence, and then pass through the first normalization layer to obtain intermediate features. The intermediate features are input into the channel attention layer for the calculation of the channel attention mechanism (SE). The calculation process of the channel attention mechanism is also an existing calculation process, and its details will not be elaborated here. Then, a weighted fusion strategy is adopted to simultaneously introduce the multi-head attention mechanism and the channel attention mechanism to obtain a first fusion feature. Finally, the first fusion feature is subjected to residual connection with the feature sequence and then passes through the second normalization layer to further stabilize the training and accelerate the convergence.

[0095] In a specific embodiment, the Transformer encoding layer includes a local-global attention unit, a feed-forward network unit, and a third layer normalization layer. The local-global attention unit includes a local attention module, a global attention module, and a fourth layer normalization layer. The feed-forward network unit includes a first fully connected layer and a second fully connected layer. The ReLU activation function is used in the first fully connected layer. When the first Transformer encoding layer is the current Transformer encoding layer, the mixed features are the output features of the previous Transformer encoding layer. The output features of the previous Transformer encoding layer are input into the current Transformer encoding layer. First, they pass through the local attention module and the global attention module of the local-global attention unit of the current Transformer encoding layer respectively, to obtain the output features of the local attention module and the output features of the global attention module. The output features of the local attention module and the output features of the global attention module are weighted and fused to obtain the second fused feature. The output features of the previous Transformer encoding layer are added to the second fused feature to obtain the third addition result. Then, the third addition result is input into the fourth layer normalization layer to obtain the output features of the local-global attention unit of the current Transformer encoding layer. The output features of the local-global attention unit of the current Transformer encoding layer are input into the feed-forward network unit of the current Transformer encoding layer, passing through the first fully connected layer and the second fully connected layer in sequence to obtain the output features of the feed-forward network unit of the current Transformer encoding layer. The output features of the feed-forward network unit of the current Transformer encoding layer are input into the third layer normalization layer of the current Transformer encoding layer to obtain the output features of the current Transformer encoding layer. The output features of the current Transformer encoding layer are used as the output features of the previous Transformer encoding layer and the above calculation process of the current Transformer encoding layer is repeated until the output features of the last Transformer encoding layer are obtained.

[0096] Specifically, referring to Figure 4 , the mixed features obtained by the hybrid attention mechanism layer after fusing the multi-head self-attention and the channel / space attention are used as the input data of the deep modeling layer, denoted as , where d represents the number of channels and T represents the total number of time steps. The hybrid features have been enhanced in both the temporal dimension and the channel dimension, and can better represent various elements such as the timbres, rhythms, and melodies of multiple musical instruments in music. In the embodiments of the present application, L stacked Transformer encoding layers are provided. Each Transformer encoding layer includes two unit blocks, namely a local-global attention unit and a feed-forward network unit, and is combined with a third layer normalization layer; the local-global attention unit includes a local attention module, a global attention module, and a fourth layer normalization layer. After the hybrid attention mechanism is completed, the Transformer encoding layer focuses on performing deeper context-dependent modeling on the hybrid features, and fully integrates the instrument sounds and melody structures at different time steps.

[0097] The calculation process of the deep modeling layer is as follows:

[0098] Input the output features of the (l - 1)-th Transformer encoding layer into the l-th Transformer encoding layer, and first pass through the local attention module and the global attention module in the local-global attention unit of the l-th Transformer encoding layer respectively;

[0099] First, perform the query, key, and value operations in the local attention module, as shown in the following formula:

[0100] , , ;

[0101] where, , and are the query weight, key weight, and value weight in the local attention module respectively, , and are the query matrix, key matrix, and value matrix in the local attention module respectively, represents the output features of the (l - 1)-th Transformer encoding layer;

[0102] Calculate the local attention weights using the following formula :

[0103] ;

[0104] where, represents the dimension of the key matrix in the local attention module, is the local window centered on the t-th time step, s is the s-th time step within the local window, represents the key matrix in the local attention module at the s-th time step within the local window, represents the query matrix in the local attention module at the t-th time step;

[0105] The output feature of the local attention module at the t-th time step is calculated using the following formula :

[0106] ;

[0107] where represents the value matrix in the local attention module at the s-th time step within the local window;

[0108] Secondly, perform the query, key, and value operations in the global attention module as shown in the following formula:

[0109] , , ;

[0110] where , and are the query weight, key weight, and value weight in the global attention module respectively, , and are the query matrix, key matrix, and value matrix in the global attention module respectively;

[0111] The global attention weight is calculated using the following formula :

[0112] ;

[0113] where represents the dimension of the key matrix in the global attention module, u represents the u-th time step, represents the key matrix in the global attention module at the u-th time step, represents the query matrix in the global attention module at the t-th time step;

[0114] The output feature of the local attention module at the t-th time step is calculated using the following formula :

[0115] ;

[0116] where represents the value matrix in the local attention module at the u-th time step, T represents the total number of time steps;

[0117] Fuse the output feature of the local attention module at the t-th time step and the output feature of the global attention module at the t-th time step to obtain the second fusion feature at the t-th time step as shown in the following formula:

[0118] ;

[0119] wherein, represents the second fusion feature at the t-th time step, represents the weighted weight;

[0120] Add the second fusion feature at the t-th time step and perform residual addition, and pass through the fourth layer normalization layer to obtain the output feature of the local-global attention unit of the l-th Transformer encoding layer , as shown in the following formula:

[0121] ;

[0122] wherein, represents the layer normalization operation.

[0123] Input the output feature of the local-global attention unit of the l-th Transformer encoding layer into the feed-forward network unit of the l-th Transformer encoding layer, and perform a fully connected mapping on through the first fully connected layer and the second fully connected layer to obtain the output feature of the feed-forward network unit of the l-th Transformer encoding layer, as shown in the following formula:

[0124] ;

[0125] wherein, represents the ReLU activation function, and are the weights of the first fully connected layer and the second fully connected layer respectively, and are the bias terms corresponding to the first fully connected layer and the second fully connected layer respectively.

[0126] Perform residual addition on and to obtain the output feature of the l-th Transformer encoding layer, as shown in the following formula:

[0127] ;

[0128] Pass the output feature of the l-th Transformer encoding layer to the (l + 1)-th Transformer encoding layer until the stacking of the L-th Transformer encoding layer is completed, and finally obtain the output feature of the L-th Transformer encoding layer Input it into the global average pooling layer to map the variable-length sequence into a fixed-length modeling vector as shown in the following formula: ;

[0129] ;

[0130] wherein represents the global average pooling operation

[0131] In a specific embodiment, the classification layer includes a Dropout layer, a third fully-connected layer, a fourth fully-connected layer, a Softmax function layer, and an argmax function layer connected in sequence. The ReLU activation function is adopted in the third fully-connected layer; the modeling vector passes through the Dropout layer to obtain an adjusted modeling vector. The adjusted modeling vector passes through the third fully-connected layer, the fourth fully-connected layer, and the Softmax function layer in sequence to obtain the predicted probability distribution of each music genre. The predicted probability distribution of each music genre is input into the argmax function layer to determine that the music genre corresponding to the maximum predicted probability is the predicted label of the audio genre corresponding to the audio data to be classified

[0132] Specifically, use the modeling vector as the input vector of the classification layer. To improve the training stability or prevent overfitting, first perform Dropout on the input vector of the classification layer through the Dropout layer. Let the probability of Dropout be , then set the j-th element in the input vector of the classification layer to zero with a probability of to obtain the adjusted modeling vector . This operation helps to reduce overfitting caused by insufficient data volume or highly correlated features, and its intensity can be adjusted by the hyperparameter in actual training

[0133] To perform a certain non-linear combination before the network outputs, the third fully-connected layer and the fourth fully-connected layer are set so that the adjusted modeling vector passes through the third fully-connected layer and the fourth fully-connected layer in sequence as shown in the following formula:

[0134] ;

[0135] ;

[0136] wherein and respectively represent the weights and bias terms of the third fully-connected layer represents the output feature of the third connection layer and respectively represent the weights and bias terms of the fourth fully-connected layer, represent the output features of the fourth connection layer.

[0137] Then apply the Softmax function for normalization processing to obtain the predicted probability distribution of each music genre, and finally select the music genre corresponding to the maximum predicted probability through the argmax function as the final music genre prediction label, as shown in the following formula:

[0138] ;

[0139] ;

[0140] Among them, represents the value of the output features of the fourth connection layer corresponding to the audio data f to be classified that belongs to the c-th audio genre, represents the value of the output features of the fourth connection layer corresponding to the audio data f to be classified that belongs to the a-th audio genre, represents the predicted probability that the audio data f to be classified belongs to the c-th audio genre, represents the music genre prediction label corresponding to the audio data f to be classified, represents taking the index of the audio genre with the maximum predicted probability.

[0141] Further referring to Figure 5 , as an implementation of the methods shown in the above figures, this application provides an embodiment of a music genre classification device based on self-supervised and hybrid attention mechanisms. This device embodiment corresponds to Figure 1 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0142] The embodiment of this application provides a music genre classification device based on self-supervised and hybrid attention mechanisms, including:

[0143] A model construction module 1, configured to construct and train a music genre classification model to obtain a trained music genre classification model. The trained music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer, and a classification layer connected in sequence. The feature extraction layer uses a wav2vec2.0 model pre-trained through self-supervision. The deep modeling layer includes a plurality of Transformer encoding layers stacked in sequence and a global average pooling layer;

[0144] The classification module 2 is configured to obtain the audio data to be classified, input the audio data to be classified into the trained music genre classification model. The audio data to be classified passes through the feature extraction layer to extract the corresponding feature sequence; input the feature sequence into the hybrid attention mechanism layer to obtain the hybrid features; input the hybrid features into the deep modeling layer, first perform context-dependent modeling through several Transformer encoding layers, and then input the output features of the last Transformer encoding layer into the global average pooling layer to obtain the modeling vector; input the modeling vector into the classification layer to obtain the predicted audio genre label corresponding to the audio data to be classified.

[0145] Figure 6 This is a schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. As Figure 6 shown, the electronic device of this embodiment includes: a processor 601 and a memory 602; wherein the memory 602 is used to store computer execution instructions; the processor 601 is used to execute the computer execution instructions stored in the memory to implement each step executed by the electronic device in the above embodiment. For details, please refer to the relevant descriptions in the foregoing method embodiments.

[0146] Optionally, the memory 602 can be either independent or integrated with the processor 601.

[0147] When the memory 602 is independently set, the electronic device further includes a bus 603 for connecting the memory 602 and the processor 601.

[0148] The embodiment of the present invention also provides a computer storage medium, in which computer execution instructions are stored. When the processor 601 executes the computer execution instructions, the above method is implemented.

[0149] The embodiment of the present invention also provides a computer program product, including a computer program. When the computer program is executed by the processor 601, the above method is implemented.

[0150] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of modules is only a logical function division. In actual implementation, there may be other division methods. For example, multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of devices or modules can be in electrical, mechanical or other forms.

[0151] The modules described as separate components may or may not be physically separated. The components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected according to actual needs to implement the solution of this embodiment.

[0152] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing unit, or each module may exist physically alone, or two or more modules may be integrated in one unit. The units formed by the above modules can be implemented in the form of hardware or in the form of a combination of hardware and software functional units.

[0153] The integrated modules implemented in the form of software functional modules can be stored in a computer-readable storage medium. The above software functional modules are stored in a storage medium and include several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 601 to execute some steps of the methods in various embodiments of the present application.

[0154] It should be understood that the above processor 601 may be a central processing unit (Central Processing Unit, abbreviated as CPU), or may also be other general-purpose processors, digital signal processors (Digital Signal Processor, abbreviated as DSP), application specific integrated circuits (Application Specific Integrated Circuit, abbreviated as ASIC), etc. The general-purpose processor may be a microprocessor or the processor 601 may also be any conventional processor 601, etc. The steps of the method disclosed in combination with the invention can be directly implemented by the hardware processor 601 or implemented by a combination of hardware and software modules in the processor 601.

[0155] The memory 602 may contain high-speed RAM memory and may also include non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a disk, or an optical disc, etc.

[0156] The bus 603 can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, or the like. The bus 603 can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 603 in the accompanying drawings of this application is not limited to only one bus 603 or one type of bus 603.

[0157] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as a Static Random Access Memory (SRAM), an Electrically Erasable Programmable Read-Only Memory (EEPROM), an Erasable Programmable Read-Only Memory (EPROM), a Programmable Read-Only Memory (PROM), a Read-Only Memory (ROM), a magnetic memory, a flash memory, a magnetic disk, or an optical disk. The storage medium can be any available medium accessible by a general-purpose or special-purpose computer.

[0158] An exemplary storage medium is coupled to the processor 601, enabling the processor 601 to read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 601. The processor 601 and the storage medium can be located in an Application Specific Integrated Circuit (ASIC). Of course, the processor 601 and the storage medium can also exist as discrete components in an electronic device or a master device.

[0159] Those of ordinary skill in the art can understand that all or part of the steps of implementing the above method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including the above method embodiments; and the foregoing storage medium includes various media that can store program codes, such as ROM, RAM, magnetic disks, or optical disks.

[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A music genre classification method based on self-supervision and hybrid attention mechanism, characterized in that: The following steps are involved: Constructing and training a music genre classification model to obtain a trained music genre classification model, wherein the trained music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer and a classification layer connected in sequence, wherein the feature extraction layer adopts a wav2vec2.0 model that has undergone self-supervised pre-training, and the deep modeling layer includes a plurality of Transformer encoding layers stacked in sequence and a global average pooling layer; Acquire audio data to be classified, input the audio data to be classified into the trained music genre classification model, the audio data to be classified passes through the feature extraction layer to extract the corresponding feature sequence; input the feature sequence into the hybrid attention mechanism layer to obtain a hybrid feature, the hybrid attention mechanism layer includes a multi-head self-attention layer, a first normalization layer, a channel attention layer and a second normalization layer, the feature sequence is input into the hybrid attention mechanism layer, first passes through the multi-head self-attention layer to obtain the output feature of the multi-head self-attention layer, the feature sequence is added to the output feature of the multi-head self-attention layer to obtain a first addition result, and then the first addition result is input into the first normalization layer to obtain an intermediate feature, the intermediate feature The feature passes through the channel attention layer to obtain the output feature of the channel attention layer, the intermediate feature and the output feature of the channel attention layer are weightedly fused to obtain a first fused feature, the feature sequence and the first fused feature are added to obtain a second addition result, and then the second addition result is input into the second normalization layer to obtain the mixed feature; the mixed feature is input into the deep modeling layer, first passes through several Transformer encoding layers for context dependency modeling, and then the output feature of the last Transformer encoding layer is input into the global average pooling layer to obtain a modeling vector; the modeling vector is input into the classification layer to obtain the audio genre prediction label corresponding to the audio data to be classified.

2. The music genre classification method based on self-supervision and hybrid attention mechanism according to claim 1, characterized in that: The Transformer coding layer includes a local-global attention unit, a feedforward network unit and a third normalization layer, the local-global attention unit includes a local attention module, a global attention module and a fourth normalization layer, the feedforward network unit includes a first fully connected layer and a second fully connected layer, and the first fully connected layer adopts a ReLU activation function; when the first Transformer coding layer is used as the current Transformer coding layer, the mixed features are used as the output features of the previous Transformer coding layer; the output features of the previous Transformer coding layer are input into the current Transformer coding layer, and are first respectively passed through the current The local attention module and the global attention module of the local-global attention unit of the Transformer coding layer are used to obtain the output features of the local attention module and the output features of the global attention module, and the output features of the local attention module and the output features of the global attention module are weightedly fused to obtain a second fused feature, and the output feature of the previous Transformer coding layer is added to the second fused feature to obtain a third addition result, and the third addition result is input into the fourth normalization layer to obtain the output feature of the local-global attention unit of the current Transformer coding layer, specifically including first performing the query, key, and value operations in the local attention module, as shown in the following formula: in, and are the query weight, key weight, and value weight in the local attention module, respectively. local , K local and V local are the query matrix, key matrix, and value matrix in the local attention module, respectively. (l-1) Represents the output features of the l-1th Transformer encoding layer; The local attention weight is calculated using the following formula Among them, d l represents the dimension of the key matrix in the local attention module, W t is the local window centered at the tth time step, s is the sth time step in the local window, represents the key matrix in the local attention module at the s-th time step within the local window, represents the query matrix in the local attention module at the t-th time step; The output features of the local attention module at the tth time step are calculated using the following formula: Among them, V s local Represents the value matrix in the local attention module at the s-th time step within the local window; Secondly, perform the query, key, and value operations in the global attention module as shown below: in, and are the query weight, key weight, and value weight in the global attention module, respectively. global , K global and V global They are the query matrix, key matrix, and value matrix in the global attention module respectively; The global attention weight is calculated using the following formula Among them, d g represents the dimension of the key matrix in the global attention module, u represents the u-th time step, represents the key matrix in the global attention module at the u-th time step, represents the query matrix in the global attention module at the t-th time step; The output features of the local attention module at the tth time step are calculated using the following formula: in, represents the value matrix in the local attention module at the uth time step, and T represents the total number of time steps; The output features of the local attention module at the t-th time step and the output features of the global attention module at the t-th time step are fused to obtain the second fused features at the t-th time step, as shown in the following formula: Among them, O t represents the second fusion feature of the tth time step, and γ represents the weighted weight; The second fusion feature of the tth time step and H (l-1) Add the residuals and pass through the fourth normalization layer to get the output features of the local-global attention unit of the lth Transformer encoding layer As shown below: Among them, LayerNorm represents the layer normalization operation; The output features of the local-global attention unit of the current Transformer coding layer are input into the feedforward network unit of the current Transformer coding layer, and sequentially pass through the first fully connected layer and the second fully connected layer to obtain the output features of the feedforward network unit of the current Transformer coding layer. The output features of the feedforward network unit of the current Transformer coding layer are input into the third normalization layer of the current Transformer coding layer to obtain the output features of the current Transformer coding layer. The output features of the current Transformer coding layer are used as the output features of the previous Transformer coding layer and the calculation process of the current Transformer coding layer is repeated until the output features of the last Transformer coding layer are obtained.

3. The music genre classification method based on self-supervision and hybrid attention mechanism according to claim 1, characterized in that: The classification layer includes a Dropout layer, a third fully connected layer, a fourth fully connected layer, a Softmax function layer and an argmax function layer connected in sequence, and the ReLU activation function is used in the third fully connected layer; the modeling vector passes through the Dropout layer to obtain an adjusted modeling vector, and the adjusted modeling vector passes through the third fully connected layer, the fourth fully connected layer and the Softmax function layer in sequence to obtain the predicted probability distribution of each music genre, and the predicted probability distribution of each music genre is input into the argmax function layer, and the music genre corresponding to the maximum predicted probability is determined as the audio genre prediction label corresponding to the audio data to be classified.

4. The music genre classification method based on self-supervision and hybrid attention mechanism according to claim 1, characterized in that: The self-supervised pre-trained wav2vec2.0 model includes an encoding network and a context network. During the training of the music genre classification model, the parameters of the encoding network are fixed, and the parameters of the context network are adjusted. The loss function used during the training of the music genre classification model is: Among them, L weighted represents the loss function, N represents the total number of samples in the training data, q represents the qth sample, c represents the cth audio genre, C represents the total number of audio genres, and w c represents the weight of the cth audio genre, p q,c represents the predicted probability that the qth sample belongs to the cth audio genre, y q,c The true label indicating whether the qth sample belongs to the cth audio genre.

5. The music genre classification method based on self-supervision and hybrid attention mechanism according to claim 1, characterized in that: The process of obtaining the training data used in the training process of the music genre classification model is as follows: Collect raw audio data with true audio genre labels; Performing random time-domain cropping on the original audio data to obtain a plurality of audio clips; Randomly mix and add noise to each audio clip to obtain a noisy audio clip; Performing time-frequency masking on the noisy audio clip to obtain enhanced audio data; The enhanced audio data and its corresponding audio genre true labels constitute the training data.

6. A music genre classification device based on self-supervision and hybrid attention mechanism, characterized in that: include: A model building module is configured to build and train a music genre classification model to obtain a trained music genre classification model, wherein the trained music genre classification model includes a feature extraction layer, a hybrid attention mechanism layer, a deep modeling layer and a classification layer connected in sequence, wherein the feature extraction layer adopts a wav2vec2.0 model that has undergone self-supervised pre-training, and the deep modeling layer includes a plurality of Transformer encoding layers stacked in sequence and a global average pooling layer; The classification module is configured to obtain audio data to be classified, input the audio data to be classified into the trained music genre classification model, the audio data to be classified passes through the feature extraction layer to extract the corresponding feature sequence; the feature sequence is input into the hybrid attention mechanism layer to obtain a mixed feature, the hybrid attention mechanism layer includes a multi-head self-attention layer, a first normalization layer, a channel attention layer and a second normalization layer, the feature sequence is input into the hybrid attention mechanism layer, first passes through the multi-head self-attention layer to obtain the output feature of the multi-head self-attention layer, the feature sequence is added to the output feature of the multi-head self-attention layer to obtain a first addition result, and then the first addition result is input into the first normalization layer to obtain an intermediate feature, The intermediate features are passed through the channel attention layer to obtain the output features of the channel attention layer, the intermediate features and the output features of the channel attention layer are weightedly fused to obtain a first fused feature, the feature sequence and the first fused feature are added to obtain a second addition result, and the second addition result is input into the second normalization layer to obtain the mixed feature; the mixed feature is input into the deep modeling layer, first passed through several Transformer encoding layers for context dependency modeling, and then the output features of the last Transformer encoding layer are input into the global average pooling layer to obtain a modeling vector; the modeling vector is input into the classification layer to obtain the audio genre prediction label corresponding to the audio data to be classified.

7. An electronic device comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Music audio classification method based on convolutional recurrent neural network

    CN112199548A

  • Sound source separation method based on shallow feature reactivation and multi-stage mixed attention

    CN114023350A