Music information classification method and system based on data terminal identification
Through the music information classification method based on data head recognition, combined with multi-dimensional feature extraction and multi-modal feature fusion, a hierarchical classification model is constructed, which solves the problems of low music classification accuracy and insufficient adaptability in the existing technology, and achieves more efficient and accurate music classification.
Patent Information
- Application Number
- CN202510639708.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-19
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-19
AI Technical Summary
The prior art has problems in the music classification that artificial labeling is time-consuming, simple audio features cannot fully express music semantics, low classification accuracy, and insufficient adaptability to cross-category music and emerging music types.
The music information classification method based on data head recognition is adopted, and the multi-dimensional feature vector is extracted through time-frequency domain transformation, and the pre-trained deep neural network model is input to extract high-level semantic features, generate data codes, and extract head features through the data head recognition model. Combining high-level semantic features and head features, fusion of multimodal features is carried out, hierarchical classification models are constructed, feature weights are dynamically adjusted, and refined classification is carried out layer by layer.
It significantly improves the accuracy and robustness of music classification, can better adapt to the hierarchical relationships between music types and emerging music types, and provides more refined music management and classification results.
Smart Images

Figure CN120183441A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to information classification technology, and in particular to a music information classification method and system based on data code header recognition. Background Art
[0002] With the rapid development of digital music, the efficient management and accurate classification of a large amount of music data have become an urgent problem to be solved. Traditional music classification methods mainly rely on manual annotation or simple audio feature extraction, and these methods have the following disadvantages: Manual annotation is time-consuming and laborious, and it is difficult to deal with a large amount of data; simple audio features cannot fully express the semantic information of music; the classification accuracy is not high, especially for cross-category music works; there is a lack of adaptability to the evolution of music styles and emerging music types.
[0003] In recent years, with the development of deep learning technology, music feature extraction and classification methods based on neural networks have made certain progress. However, these methods still face the following challenges: The computational complexity of feature extraction is high, and it is difficult to achieve real-time classification; the model generalization ability is limited, and the adaptability to unseen music types is poor; there is a lack of modeling of the hierarchical relationship between music types; the classification results lack interpretability and are difficult to apply to refined music management.
[0004] To solve the above problems, researchers have proposed various improvement methods, such as using transfer learning to improve the model generalization ability, introducing an attention mechanism to capture key features, adopting a hierarchical classification strategy, etc. However, these methods often only solve some problems and lack a unified framework to comprehensively improve the efficiency and accuracy of music classification.
[0005] Therefore, there is an urgent need for a music classification method that can efficiently extract music semantic features, construct a hierarchical classification system, and has good generalization ability and interpretability. Summary of the Invention
[0006] The embodiments of the present invention provide a music information classification method and system based on data code header recognition, which can solve the problems in the prior art.
[0007] In the first aspect of the embodiments of the present invention, A music information classification method based on data code header recognition is provided, including: Performing time-frequency domain conversion on the input audio data to obtain a spectrogram; extracting a multi-dimensional feature vector from the spectrogram, inputting the multi-dimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code of the audio data according to the high-level semantic features; Construct a data code header recognition model, input the generated data code into the data code header recognition model, and extract the header features of the data code; perform a preliminary classification on the audio data according to the header features to obtain a preliminary classification result; establish a music type classification tree, where the music type classification tree includes music type nodes at multiple levels; locate the initial node in the music type classification tree according to the preliminary classification result; Perform multi-modal feature fusion on the high-level semantic features and the header features to obtain a fused feature vector; construct a hierarchical classification model, where the hierarchical classification model includes multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music type classification tree; input the fused feature vector into the hierarchical classification model, and start from the initial node to perform refined classification layer by layer downward; During the classification process at each level, dynamically adjust the feature weights to highlight the key features of the current level; update the music type classification tree according to the classification result, including adding new music type nodes or adjusting the relationships between nodes; output the final music information classification result.
[0008] In an alternative implementation, Input the multi-dimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code for the audio data according to the high-level semantic features includes: Reshape the multi-dimensional feature vector into the Mel spectrogram format; input the Mel spectrogram into a pre-trained VGGish deep neural network model; extract high-level semantic features through the convolutional layer block and the fully connected layer of the VGGish deep neural network model; Quantize each element in the high-level semantic features into an 8-bit unsigned integer; perform bit-plane decomposition on the quantized feature vector to obtain 8 binary bit-planes; calculate the entropy value of each binary bit-plane; sort the binary bit-planes according to the entropy value, and select the top K binary bit-planes with the largest amount of information; Connect the selected K binary bit-planes into a long binary sequence, and compress the long binary sequence using run-length encoding; if the length after compression exceeds the target length, truncate it; if the length after compression is less than the target length, pad it with 0s to the target length; Calculate the checksum of the compressed sequence using the CRC-32 algorithm; add the checksum to the end of the compressed sequence to generate a fixed-length binary data code.
[0009] In an alternative implementation, Construct a data code header recognition model, input the generated data code into the data code header recognition model, and extract the header features of the data code; perform a preliminary classification on the audio data according to the header features to obtain a preliminary classification result, including: The data code header recognition model includes a convolutional neural network and a recurrent neural network; the convolutional neural network includes multiple convolutional kernels of different sizes, and multiple filters are set for each size of convolutional kernel; the recurrent neural network is a bidirectional long short-term memory network; a self-attention mechanism layer is set after the bidirectional long short-term memory network; Convert the data code into a data code sequence represented by a numerical value; input the converted data code sequence into the data code header recognition model; extract multi-scale local features through the convolutional neural network and max pooling operation; input the extracted multi-scale local features into the bidirectional long short-term memory network to obtain a context-related feature representation; use the self-attention mechanism layer to calculate the importance weights of different positions; pass the weighted features through a fully connected layer to obtain the header features of the data code; Perform z-score normalization processing on the header features; use the principal component analysis method to reduce the dimension of the normalized header features, and retain the principal components whose proportion of explained variance reaches a preset threshold; construct a random forest classifier, input the dimension-reduced header features into the random forest classifier to obtain the probability that each sample belongs to each music category; select the category with the highest probability as the preliminary classification result.
[0010] In an alternative embodiment, The method further includes: Use the cross-validation method to evaluate the classification performance, and calculate the accuracy, precision, recall rate, and F1 score; analyze the samples with classification errors to identify the music types that are difficult to classify; according to the analysis results, adjust the parameters of the data code header recognition model and the parameters of the random forest classifier to improve the classification accuracy.
[0011] In an alternative embodiment, Establish a music type classification tree, and the music type classification tree includes music type nodes at multiple levels; locating the initial node in the music type classification tree according to the preliminary classification result includes: Construct a music type classification tree, and the construction of the music type classification tree includes: defining multi-level music categories, and each node contains a unique identifier, a node name, a parent node, a list of child nodes, a feature vector, and the number of samples; collecting labeled music samples, using a clustering algorithm to analyze the samples, and dynamically adjusting the tree structure; calculating the representative feature vector of each node, and the representative feature vector is the mean vector of the high-level semantic features of all music samples of the node; calculating the cosine similarity between all nodes, and constructing a similarity matrix between nodes; Classify the input music sample, and the classification of the input music sample includes: extracting the high-level semantic feature vector of the music sample; using a random forest classifier to perform a preliminary classification on the music sample to obtain the probability distribution of each top-level category; selecting the top N top-level categories with the highest probabilities, and performing a multi-path search for each selected top-level category. The multi-path search includes starting from the top-level node, calculating the similarity between the music sample feature vector and the current node feature vector, selecting the child node with the highest similarity to enter the next layer, and repeating this process until reaching the leaf node or the similarity is lower than the preset threshold; collecting all candidate nodes obtained from the multi-path search, calculating the similarity between the music sample feature vector and the feature vector of each candidate node, and sorting the candidate nodes according to the similarity; selecting the candidate node with the highest similarity as the initial node.
[0012] In an alternative embodiment, Evaluate the classification result and optimize the classification tree. The evaluation of the classification result and the optimization of the classification tree include: Calculating the similarity between the music sample feature vector and the initial node feature vector, comparing the similarity with the preset threshold, and evaluating the classification confidence; if the confidence is lower than the preset threshold, then marking the music sample as "uncertain"; recording the classification result, including the initial node, confidence, and search path; updating the sample quantity and feature vector of the node according to the classification result; regularly recalculating the similarity matrix between nodes to optimize the music type classification tree structure.
[0013] In an alternative embodiment, Construct a hierarchical classification model. The hierarchical classification model includes multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music type classification tree; inputting the fused feature vector into the hierarchical classification model, and starting from the initial node, performing refined classification layer by layer includes: Performing standardization processing on the high-level semantic feature vector and the head feature vector to eliminate the dimensional difference between different features; using the random forest algorithm to evaluate the importance of the standardized features, and calculating the Gini importance index of each feature; selecting the most important feature subset according to the feature importance index; Adopting an attention mechanism for feature fusion, defining an attention weight matrix, using the softmax function to calculate the weight of each feature, and obtaining a fused feature vector; performing principal component analysis dimensionality reduction processing on the fused feature vector to obtain a dimensionality-reduced feature vector; Construct a hierarchical classification model. The hierarchical classification model includes multiple sub-classifiers, each sub-classifier corresponds to a non-leaf node in the music type classification tree, and each sub-classifier consists of a support vector machine, a random forest, and a gradient boosting decision tree; Customize the feature set for each sub-classifier, including the dimensionality-reduced feature vectors and node-specific domain features; train each sub-classifier using the hierarchical cross-validation method and fuse the prediction results of the base models using the Stacking method; design a hierarchical decision-making strategy, define the confidence threshold and node selection function for node selection and termination condition judgment during the classification process; Calculate the feature importance vector of each node and use the SHAP value to quantify the importance of features; design a weight adjustment function to dynamically adjust the feature weights according to the difference in feature importance between the current node and the parent node; When classifying to the next lower layer each time, use the weight adjustment function to update the feature vector; introduce an adaptive mechanism to dynamically adjust the parameters of the weight adjustment function according to the classification performance; input the dimensionality-reduced feature vector into the hierarchical classification model, start from the initial node, and perform refined classification layer by layer downward, dynamically adjusting the feature weights during the classification process at each level; output the final music genre classification result.
[0014] In the second aspect of the embodiments of the present invention, there is provided a music information classification system based on data code header recognition, including: A first unit for performing time-frequency domain conversion on the input audio data to obtain a spectrogram; extracting a multi-dimensional feature vector from the spectrogram, inputting the multi-dimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code of the audio data according to the high-level semantic features; A second unit for constructing a data code header recognition model, inputting the generated data code into the data code header recognition model to extract the header features of the data code; performing preliminary classification on the audio data according to the header features to obtain a preliminary classification result; establishing a music genre classification tree, where the music genre classification tree includes multiple levels of music genre nodes; positioning the initial node in the music genre classification tree according to the preliminary classification result; A third unit for performing multi-modal feature fusion on the high-level semantic features and the header features to obtain a fused feature vector; constructing a hierarchical classification model, where the hierarchical classification model includes multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music genre classification tree; inputting the fused feature vector into the hierarchical classification model, and starting from the initial node, performing refined classification layer by layer downward; A fourth unit for dynamically adjusting the feature weights during the classification process at each level to highlight the key features of the current level; updating the music genre classification tree according to the classification results, including adding new music genre nodes or adjusting the relationships between nodes; outputting the final music information classification result.
[0015] In the third aspect of the embodiments of the present invention, there is provided an electronic device, including: A processor; A memory for storing processor-executable instructions; Wherein, the processor is configured to call the instructions stored in the memory to execute the aforementioned method.
[0016] In a fourth aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the aforementioned method is implemented.
[0017] Through multi-dimensional feature extraction and multi-modal feature fusion, the present invention makes full use of the time-frequency domain information, semantic information, and head features of audio data. Advanced semantic features are extracted by a deep neural network, combined with the head features extracted by a data code head recognition model, to achieve a comprehensive representation of music information. This multi-angle and multi-level feature extraction and fusion method significantly improves the accuracy of music classification. At the same time, by preprocessing the audio data, including noise reduction, framing, and windowing, the adaptability of the system to audio of different qualities is enhanced, and the robustness of classification is improved.
[0018] The present invention constructs a hierarchical music genre classification tree and a corresponding hierarchical classification model. By initially classifying to locate the initial nodes and then performing refined classification layer by layer downward, it simulates the hierarchical cognitive process of human beings for music genres. During the classification process at each level, the feature weights are dynamically adjusted to highlight the key features of the current level, further improving the classification accuracy. This method can not only achieve a fine division of music genres but also handle the fuzzy boundaries and overlapping relationships between music genres, providing more accurate and detailed music classification results for users. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 It is a schematic flowchart of the music information classification method based on data code head recognition according to an embodiment of the present invention; Figure 2 It is a schematic structural diagram of the music information classification system based on data code head recognition according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0021] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.
[0022] Figure 1 The following is a schematic flowchart of a music information classification method based on data code header recognition according to an embodiment of the present invention. As Figure 1 shown, the method includes: Perform time-frequency domain conversion on the input audio data to obtain a spectrogram; extract a multi-dimensional feature vector from the spectrogram, input the multi-dimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generate a data code for the audio data according to the high-level semantic features; Construct a data code header recognition model, input the generated data code into the data code header recognition model to extract the header features of the data code; perform preliminary classification on the audio data according to the header features to obtain a preliminary classification result; establish a music type classification tree, and the music type classification tree includes music type nodes at multiple levels; locate the initial node in the music type classification tree according to the preliminary classification result; Perform multi-modal feature fusion on the high-level semantic features and the header features to obtain a fused feature vector; construct a hierarchical classification model, and the hierarchical classification model includes multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music type classification tree; input the fused feature vector into the hierarchical classification model, and start from the initial node to perform refined classification layer by layer downward; During the classification process of each level, dynamically adjust the feature weights to highlight the key features of the current level; update the music type classification tree according to the classification results, including adding new music type nodes or adjusting the relationships between nodes; output the final music information classification result.
[0023] The specific implementation manner of the music information classification method based on data code header recognition is as follows: First, preprocess the input audio data. The preprocessing includes noise reduction, framing, and windowing. Noise reduction uses spectral subtraction. The audio signal is transformed into the frequency domain, the noise spectrum is estimated, the noise spectrum is subtracted from the original signal spectrum, and then it is transformed back into the time domain to obtain the noise-reduced signal. Framing uses a frame length of 25 ms and a frame shift of 10 ms to cut the audio into a series of short-time frames. Windowing uses the Hamming window function to reduce spectral leakage.
[0024] Perform time-frequency domain conversion on the preprocessed audio data to obtain a spectrogram. Specifically, use the short-time Fourier transform (STFT) to transform each frame of the signal into the frequency domain to obtain a spectrogram. The horizontal axis of the spectrogram represents time, the vertical axis represents frequency, and the color depth represents the energy magnitude.
[0025] Extract multi-dimensional feature vectors from the spectrogram, including Mel Frequency Cepstral Coefficients (MFCC), chroma features, rhythm features, and pitch features. The MFCC extraction uses 13th-order coefficients. The chroma feature extraction is a 12-dimensional chroma vector. The rhythm features include beat intensity and rhythm regularity. The pitch features include fundamental frequency and pitch distribution.
[0026] Input the multi-dimensional feature vectors into a pre-trained deep neural network model to extract high-level semantic features. The deep neural network model adopts a five-layer fully connected network structure, with the number of neurons in the hidden layers being 512, 256, 128, 64 respectively, and the output layer being 32-dimensional high-level semantic features.
[0027] Generate a data code for the audio data based on the high-level semantic features. The data code is a 256-bit binary sequence. The specific method is to quantize the 32-dimensional high-level semantic features into a 256-dimensional binary vector, with each feature represented by 8-bit binary numbers. For example, for the high-level semantic feature vector [0.7, 0.3, 0.5, ..., 0.8], the quantized data code is "1100011101001110 10000000 ... 11001100".
[0028] Construct a data code header recognition model, including a Convolutional Neural Network (CNN) and a Recurrent Neural Network (RNN). The CNN adopts a 3-layer convolutional structure, with the convolutional kernel sizes being 5, 3, 3 respectively, and the number of convolutional kernels being 32, 64, 128 respectively. The RNN adopts a Long Short-Term Memory network (LSTM), with the hidden layer size being 128. The CNN is used to extract local features, and the RNN is used to capture long-range dependencies.
[0029] Input the generated data code into the data code header recognition model to extract the header features of the data code. Specifically, take the first 64 bits of the data code as the input, and after being processed by the CNN and RNN, obtain a 32-dimensional header feature vector.
[0030] Perform a preliminary classification of the audio data based on the header features to obtain a preliminary classification result. Use a softmax classifier to map the 32-dimensional header features to 10 predefined coarse-grained music categories, such as classical, pop, rock, etc. The category with the highest output probability is used as the preliminary classification result.
[0031] Establish a music genre classification tree, including multiple levels of music genre nodes. The classification tree adopts a three-layer structure. The first layer is 10 coarse-grained categories, the second layer is subdivided into 50 subcategories, and the third layer is further subdivided into 200 specific types. For example, under the "classical" category, it can be subdivided into subcategories such as "symphony" and "chamber music", and further into specific types such as "Mozart Symphony" and "Beethoven Symphony".
[0032] Locate the initial node in the music genre classification tree according to the preliminary classification result. For example, if the preliminary classification result is "Classical", locate the "Classical" node in the first layer of the classification tree as the initial node.
[0033] Perform multi-modal feature fusion on the extracted high-level semantic features and the extracted head features to obtain a fused feature vector. The attention mechanism is used for feature fusion to calculate the importance weights of the two features, and the weighted sum is used to obtain a 64-dimensional fused feature vector.
[0034] Construct a hierarchical classification model, including multiple sub-classifiers, each sub-classifier corresponding to a non-leaf node in the music genre classification tree. Each sub-classifier is implemented using a support vector machine (SVM). For example, the sub-classifier corresponding to the "Classical" node is used to further classify "Classical" music into sub-categories such as "Symphony" and "Chamber Music".
[0035] Input the fused feature vector into the hierarchical classification model, and start from the initial node to perform refined classification layer by layer downward. During the classification process at each level, dynamically adjust the feature weights to highlight the key features of the current level. For example, when further classifying "Classical" music, increase the weight of the features related to the instrument composition.
[0036] Update the music genre classification tree according to the classification results, including adding new music genre nodes or adjusting the relationships between nodes. For example, if a large number of samples are classified as "Electronic Classical", a new sub-category node "Electronic Classical" can be added under "Classical".
[0037] Finally, output the music information classification results, including music genre, style, emotion label, and recommendation label. The music genre corresponds to the leaf node in the classification tree, such as "Mozart Symphony". The style label is selected from a predefined style vocabulary, such as "elegant" and "exciting". The emotion label is predicted based on the music features, such as "cheerful" and "melancholy". The recommendation label is generated according to the user's listening history, such as "suitable for work" and "relaxing before bedtime".
[0038] Through the above steps, refined classification of music information based on data code head recognition is achieved, providing strong support for applications such as music retrieval and recommendation. This method combines multiple features and classification techniques, can accurately identify music genres, and provides rich label information.
[0039] In an alternative embodiment, input the multi-dimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generate a data code for the audio data according to the high-level semantic features, and the data code includes a fixed-length binary sequence, including: Reshape the multi-dimensional feature vector into the Mel spectrogram format; input the Mel spectrogram into a pre-trained VGGish deep neural network model; extract high-level semantic feature vectors through the convolutional layer blocks and fully connected layers of the VGGish deep neural network model; Perform L2 normalization on the high-level semantic feature vectors; use principal component analysis to reduce the dimensionality of the normalized feature vectors; quantize the dimension-reduced feature vectors to obtain quantized feature vectors; introduce an attention mechanism to calculate the importance weights of different features in the quantized feature vectors; weight the quantized feature vectors according to the importance weights to obtain enhanced feature vectors; Quantize each element in the enhanced feature vectors into 8-bit unsigned integers; perform bit-plane decomposition on the quantized feature vectors to obtain 8 binary bit-planes; calculate the entropy value of each binary bit-plane; sort the binary bit-planes according to the entropy value and select the top K binary bit-planes with the largest amount of information; Connect the selected K binary bit-planes into a long binary sequence; perform (7,4) Hamming coding on the long binary sequence; use run-length coding to compress the Hamming-coded sequence; if the compressed length exceeds the target length, truncate it; if the compressed length is less than the target length, pad it with 0s to the target length; Calculate the checksum of the compressed sequence using the CRC-32 algorithm; add the checksum to the end of the compressed sequence to generate a fixed-length binary data code; the fixed-length binary data code includes the encoded semantic information and the checksum.
[0040] Reshape the multi-dimensional feature vector into the Mel spectrogram format. The multi-dimensional feature vector usually contains the time and frequency features of the audio signal, and it needs to be converted into a two-dimensional Mel spectrogram for subsequent input into the deep neural network model. For example, reshape a 128-dimensional feature vector into an 8x16 Mel spectrogram.
[0041] Input the Mel spectrogram into a pre-trained VGGish deep neural network model. The VGGish model is an audio feature extractor based on the VGG network structure. It has been pre-trained on a large-scale audio dataset and can effectively extract the high-level semantic features of audio. The model contains multiple convolutional layer blocks and fully connected layers, and more abstract feature representations can be gradually extracted through these layers.
[0042] Extract high-level semantic feature vectors through the convolutional layer blocks and fully connected layers of the VGGish model. Specifically, the Mel spectrogram first passes through 4 convolutional layer blocks, each of which contains 2-3 convolutional layers and 1 max-pooling layer. Then it passes through 2 fully connected layers, and finally a 128-dimensional feature vector is obtained, which is the extracted high-level semantic feature.
[0043] Perform L2 normalization on the high-level semantic feature vectors. L2 normalization can scale the feature vectors to unit length, which helps to eliminate the scale differences between different samples. The specific operation is to divide the feature vector by its L2 norm (Euclidean norm).
[0044] Use principal component analysis (PCA) to reduce the dimensionality of the normalized feature vectors. PCA can find the main directions of data variation, remove redundant information, and at the same time retain the most important features. For example, reduce the 128-dimensional feature vector to 64 dimensions. Quantize the feature vector after dimensionality reduction to obtain a quantized feature vector. Quantization can discretize the continuous feature values, which is beneficial for subsequent coding and compression. Uniform quantization or non-uniform quantization methods can be adopted to map the feature values to a finite integer range, such as 0 - 255.
[0045] Introduce an attention mechanism to calculate the importance weights of different features in the quantized feature vector. The attention mechanism can adaptively assign weights to different features, highlighting the role of important features. Self-attention or other attention calculation methods can be used to obtain a 64-dimensional weight vector. Weight the quantized feature vector according to the importance weights to obtain an enhanced feature vector. Multiply the weight vector and the quantized feature vector element by element to obtain a 64-dimensional enhanced feature vector.
[0046] Quantize each element in the enhanced feature vector into an 8-bit unsigned integer. This step further discretizes the feature values, facilitating subsequent bit-plane decomposition. Specifically, the feature values can be linearly mapped to the integer range of 0 - 255. Perform bit-plane decomposition on the quantized feature vector to obtain 8 binary bit-planes. Split each 8-bit integer into 8 binary bits to obtain 8 64-bit binary sequences, with each sequence corresponding to a bit-plane.
[0047] Calculate the entropy value of each binary bit-plane. The entropy value reflects the richness of information in the bit-plane. The occurrence probabilities of 0 and 1 in each bit-plane can be counted, and then the information entropy can be calculated. Sort the binary bit-planes according to the entropy values, and select the top K binary bit-planes with the largest amount of information. Usually, select the top 4 - 6 bit-planes with the largest entropy values, as these bit-planes contain the most critical information. Connect the selected K binary bit-planes into a long binary sequence. For example, select 4 64-bit bit-planes, and after connection, obtain a 256-bit binary sequence.
[0048] Perform (7,4) Hamming coding on the long binary sequence. Hamming coding can detect and correct single-bit errors, improving the reliability of data. The length of the encoded sequence will increase to 7 / 4 times the original, that is, 448 bits. Use run-length coding to compress the Hamming-coded sequence. Run-length coding can effectively compress consecutive 0s or 1s, reducing the data length. The compressed length is not fixed and depends on the distribution characteristics of the data.
[0049] If the compressed length exceeds the target length, truncation is performed. If it is less than the target length, it is padded with 0s to the target length. Assuming the target length is 512 bits, then corresponding truncation or padding operations need to be performed on the compressed sequence. Use the CRC-32 algorithm to calculate the checksum of the compressed sequence. CRC-32 can detect errors during data transmission, improving data integrity. The checksum is usually 32 bits.
[0050] Add the checksum to the end of the compressed sequence to generate a binary data code of fixed length. The finally obtained data code includes 480 bits of encoded semantic information and 32 bits of CRC checksum, with a total length of 512 bits.
[0051] In this way, the generation process from the multi-dimensional feature vector to the fixed-length binary data code is completed. The generated data code contains both the key semantic information of the audio and has the ability of error detection and correction, and can be used in application scenarios such as audio retrieval and recognition.
[0052] In an alternative embodiment, a data code header recognition model is constructed, and the data code header recognition model includes a convolutional neural network and a recurrent neural network; the generated data code is input into the data code header recognition model to extract the header features of the data code; based on the header features, the audio data is preliminarily classified, and the preliminary classification results include: Construct a data code header recognition model, and the data code header recognition model includes a convolutional neural network and a recurrent neural network; the convolutional neural network includes multiple convolutional kernels of different sizes, and multiple filters are set for each size of convolutional kernel; the recurrent neural network is a bidirectional long short-term memory network; a self-attention mechanism layer is set after the bidirectional long short-term memory network; Receive a fixed-length binary data code sequence, convert the binary data code sequence into a numerical representation; input the converted data code sequence into the data code header recognition model; extract multi-scale local features through the convolutional neural network and max-pooling operations; input the extracted multi-scale local features into the bidirectional long short-term memory network to obtain context-related feature representations; use the self-attention mechanism layer to calculate the importance weights of different positions; pass the weighted features through the fully connected layer to obtain the header features of the data code; Perform z-score normalization on the head features; use the principal component analysis method to reduce the dimension of the normalized head features, and retain the principal components whose proportion of explained variance reaches the preset threshold; construct a random forest classifier, and the parameters of the random forest classifier include the number of decision trees, the maximum tree depth, the minimum number of samples in leaf nodes, and the feature selection criterion; input the dimension-reduced head features into the random forest classifier to obtain the probability that each sample belongs to each music category; select the category with the highest probability as the preliminary classification result.
[0053] According to the method, constructing a data code head recognition model is a key step. This model combines the advantages of convolutional neural network (CNN) and recurrent neural network (RNN), and can effectively extract the local features and sequence features of the data code.
[0054] In specific implementation, first construct the CNN part. The CNN contains convolutional kernels of different sizes, such as 3x3, 5x5, and 7x7, and multiple filters are set for each size of convolutional kernel, such as 64, 128, and 256. In this way, local features of different scales can be extracted. After the convolutional layer, a max pooling layer is connected to reduce the dimension and extract significant features.
[0055] The RNN part uses a bidirectional long short-term memory network (Bi-LSTM). The Bi-LSTM contains two LSTMs, forward and backward, which can consider the context information of the past and the future at the same time. The LSTM unit contains an input gate, a forget gate, and an output gate, which can effectively handle the long-term dependence problem. The number of hidden units of the Bi-LSTM can be set to 128 or 256.
[0056] Add a self-attention mechanism layer after the Bi-LSTM. This layer calculates the correlation between different positions in the sequence to obtain attention weights. This helps the model focus on important feature positions. The dimension of the self-attention layer can be the same as the number of hidden units of the Bi-LSTM.
[0057] When receiving the input binary data code sequence, first convert it into a numerical representation. For example, "0" can be mapped to -1, and "1" can be mapped to 1. Assuming the input sequence length is 1024 bits, a 1024-dimensional vector is obtained after conversion. Input the converted sequence into the CNN for convolution operation. Taking the 3x3 convolutional kernel as an example, a feature map of (1024 - 3 + 1) x 64 = 1022 x 64 is obtained after convolution. After max pooling, the feature dimension is further reduced. Repeat this process for convolutional kernels of different sizes, and finally obtain feature maps of multiple scales.
[0058] The features extracted by the CNN are input into the Bi-LSTM. The Bi-LSTM processes the feature sequence bidirectionally and outputs a new feature sequence containing context information. Assuming the number of hidden units in the Bi-LSTM is 128, the output dimension is 1022x256. The self-attention layer calculates the correlations between positions in the feature sequence to obtain an attention matrix. The attention weights are multiplied by the original features to obtain a weighted feature representation. Finally, through a fully connected layer, the features are mapped to a lower dimension, such as 256 dimensions, to obtain the head feature representation of the data code. This feature is z-score normalized to have a mean of 0 and a variance of 1.
[0059] Principal Component Analysis (PCA) is used to reduce the dimension of the normalized features. Assuming 95% of the variance information is retained, the 256-dimensional features may be reduced to around 64 dimensions.
[0060] A random forest classifier is constructed for preliminary classification. The number of decision trees is set to 100, the maximum tree depth is 10, the minimum number of samples in a leaf node is 5, and the Gini coefficient is used for feature selection. The dimension-reduced features are input into the random forest to obtain the probabilities of the samples belonging to each music category. The category with the highest probability is selected as the preliminary classification result.
[0061] Through the above steps, an end-to-end processing from the original data code to the preliminary classification result is achieved. This method combines deep learning and machine learning techniques, and can effectively extract the key features of the data code, laying a foundation for subsequent fine classification.
[0062] In an optional implementation manner, the method further includes: Using a cross-validation method to evaluate the classification performance, calculating the accuracy, precision, recall, and F1 score; analyzing the samples with classification errors to identify the music types that are difficult to classify; according to the analysis results, adjusting the parameters of the data code head recognition model and the parameters of the random forest classifier to improve the classification accuracy.
[0063] In terms of evaluating and optimizing the classification performance, this implementation manner adopts methods such as cross-validation, multi-index evaluation, and error analysis to comprehensively evaluate the model effect and optimize it targeted. The specific implementation process is as follows: The data set is divided into a training set and a test set with a ratio of 8:2. Then, 5-fold cross-validation is performed on the training set, that is, the training set is evenly divided into 5 parts, and each time 4 of them are taken as training data, and the remaining 1 part is used as validation data, repeating 5 times. This can make full use of limited data and avoid overfitting. In each fold of cross-validation, the data code head recognition model and the random forest classifier are trained using the training data, and then predictions are made on the validation data. Record the true labels and predicted labels of each sample. Take the average of the 5 results and calculate metrics such as accuracy, precision, recall, and F1 score.
[0064] Accuracy reflects the proportion of correct overall classifications. Precision reflects the proportion of positive examples predicted as positive. Recall reflects the proportion of correct predictions of positive examples. F1 score is the harmonic mean of precision and recall. These indicators evaluate classification performance from different perspectives. For example, in the music genre classification task, the following results may be obtained: accuracy 85%, precision 80%, recall 75%, and F1 score 77.5%. This shows that the overall classification effect is good, but there is still room for improvement.
[0065] Next, we analyze the samples that were misclassified. We count all the misclassified samples according to their true labels and predicted labels to generate a confusion matrix. The confusion matrix allows us to intuitively see which categories are easily confused. For example, we found that there are many misclassifications between classical music and jazz. Further analysis of the features of these samples may reveal that they are similar in terms of rhythm, harmony, etc. This shows that the existing features may not be sufficient to distinguish the two genres well.
[0066] After optimization, re-perform cross-validation evaluation. For example, after optimization, the accuracy rate is increased to 88%, the precision rate is 83%, the recall rate is 80%, and the F1 score is 81.5%. This shows that the optimization measures have achieved certain results. Finally, a final evaluation is performed on the test set to verify the generalization ability of the model. If the performance on the test set is similar to the cross-validation result, it means that the model has good generalization ability. Otherwise, there may be an overfitting problem, and the model structure or regularization parameters need to be further adjusted.
[0067] Through this iterative optimization process, the performance of the classification model is continuously improved, enabling it to more accurately identify different types of music. This method is not only applicable to music classification, but can also be extended to other audio classification tasks.
[0068] In an optional implementation, a music type classification tree is established, the music type classification tree comprising multiple levels of music type nodes; and locating an initial node in the music type classification tree according to the preliminary classification result comprises: Constructing a music type classification tree, the construction of the music type classification tree includes: defining a multi-level music category, each node contains a unique identifier, a node name, a parent node, a child node list, a feature vector and a number of samples; collecting labeled music samples, using a clustering algorithm to analyze the samples, and dynamically adjusting the tree structure; calculating a representative feature vector of each node, the representative feature vector is a high-level semantic feature mean vector of all music samples of the node; calculating the cosine similarity between all nodes, and constructing an inter-node similarity matrix; Classify the input music sample. The classification of the input music sample includes: extracting the high-level semantic feature vector of the music sample; using a random forest classifier to perform a preliminary classification on the music sample to obtain the probability distribution of each top-level category; selecting the top N top-level categories with the highest probabilities, and performing a multi-path search on each selected top-level category. The multi-path search includes starting from the top-level node, calculating the similarity between the music sample feature vector and the current node feature vector, selecting the child node with the highest similarity to enter the next layer, and repeating this process until reaching a leaf node or the similarity is lower than a preset threshold; collecting all candidate nodes obtained from the multi-path search, calculating the similarity between the music sample feature vector and the feature vector of each candidate node, and sorting the candidate nodes according to the similarity; selecting the candidate node with the highest similarity as the initial node.
[0069] In a specific implementation, first construct a music genre classification tree. Define multi-level music categories. For example, it can include top-level categories such as pop, rock, classical, jazz, etc. Each top-level category can also contain multiple sub-categories. Define attributes such as a unique identifier, node name, parent node, list of child nodes, feature vector, and number of samples for each node. For example, the pop music node can be defined as: Identifier: POP001; Node name: Pop music; Parent node: ROOT; List of child nodes: [POP002, POP003, POP004]; Feature vector: [0.8, 0.6, 0.3, 0.7]; Number of samples: 10,000.
[0070] Next, collect the labeled music samples. The classified music data can be obtained from major music platforms. Use clustering algorithms such as K-means to analyze the samples, and dynamically adjust the tree structure according to the clustering results. For example, it may be necessary to add new sub-categories or merge categories with high similarity.
[0071] Then calculate the representative feature vector of each node. Extract the high-level semantic features of all music samples of this node, including rhythm, timbre, harmony, etc., and average the feature vectors of all samples to obtain the representative feature vector of this node. For example, the representative feature vector of the pop music node may be [0.75, 0.65, 0.35, 0.68].
[0072] Next, calculate the cosine similarity between all nodes to construct a node similarity matrix. For example, the similarity between the pop music node and other nodes can be: Pop - Rock: 0.82; Pop - Classical: 0.23; Pop - Jazz: 0.56; For the input music sample, first extract its high-level semantic feature vector. Use the trained random forest classifier to perform a preliminary classification on the sample to obtain the probability distribution of each top-level category. For example, the classification result of a certain sample can be: Pop: 0.6; Rock: 0.3; Classical: 0.05; Jazz: 0.05; Select the top N top-level categories with the highest probabilities. For example, select the categories with probabilities greater than 0.1. Here, the two categories of Pop and Rock are selected.
[0073] Perform a multi-path search for each selected top-level category. Taking pop music as an example, start from the pop music node and calculate the cosine similarity between the sample feature vector and the current node feature vector. Suppose the pop music node has 3 child nodes: pop rock, electro-pop, and R&B. Calculate the similarities between the sample and these 3 child nodes as 0.85, 0.72, and 0.63 respectively, and select the pop rock node with the highest similarity to enter the next layer. Repeat this process until reaching a leaf node or the similarity is lower than a preset threshold (such as 0.5).
[0074] Perform the same multi-path search process for the rock category. Finally, collect all candidate nodes obtained from the multi-path search, calculate the similarity between the sample feature vector and each candidate node feature vector, and sort the candidate nodes according to the similarity. For example, the final sorted list of candidate nodes can be: 1. Pop Rock: 0.85; 2. Alternative Rock: 0.78; 3. Electro-Pop: 0.72; 4. Hard Rock: 0.68; Select the pop rock node with the highest similarity as the initial node. In this way, the precise classification of the input music sample is completed, and the most matching node in the music genre classification tree is located.
[0075] This method can accurately classify music samples by constructing a multi-level music genre classification tree, combining the preliminary classification of the random forest and the multi-path search. The dynamic adjustment mechanism of the music genre classification tree ensures that the classification system keeps pace with the times, and the multi-path search strategy improves the accuracy and robustness of the classification. This method can be widely applied in fields such as music recommendation and music retrieval to provide more accurate music services for users.
[0076] In an optional implementation manner, evaluate the classification result and optimize the classification tree. The evaluating the classification result and optimizing the classification tree includes: Calculate the similarity between the music sample feature vector and the initial node feature vector, compare the similarity with a preset threshold to evaluate the classification confidence; if the confidence is lower than the preset threshold, mark the music sample as "uncertain"; record the classification result, including the initial node, confidence, and search path; update the sample quantity and feature vector of the node according to the classification result; regularly recalculate the similarity matrix between nodes to optimize the music type classification tree structure.
[0077] In the specific implementation, the method for evaluating the classification result and optimizing the classification tree may include the following steps: Calculate the similarity between the feature vector of the music sample to be classified and the feature vector of the initial node of the classification tree. The similarity calculation can use the cosine similarity method. After normalizing the two vectors, calculate their dot product. For example, assume the music sample feature vector is [0.3, 0.5, 0.2], and the initial node feature vector is [0.4, 0.4, 0.2], then their similarity is 0.94.
[0078] Compare the calculated similarity with a preset threshold to evaluate the classification confidence. The preset threshold can be set to 0.8. If the similarity is greater than or equal to 0.8, it is considered that the classification confidence is high; if the similarity is less than 0.8, it is considered that the classification confidence is low. For the above example, the similarity 0.94 is greater than the threshold 0.8, so the classification confidence is high.
[0079] Process according to the evaluation result of the confidence. If the confidence is high, continue to classify downward along the classification tree; if the confidence is low, mark the music sample as "uncertain". For the case of high confidence, continue to calculate the similarity between the sample and the child nodes of the current node, and select the child node with the highest similarity as the next classification node, repeating the above process until reaching the leaf node.
[0080] During the classification process, record the classification result, including the initial node, classification nodes at all levels, the final leaf node classified to, the confidence of each node, and the entire search path. For example, a possible record is: initial node - pop music (0.94), secondary node - Chinese pop (0.88), leaf node - Chinese mainland pop (0.92).
[0081] After completing one classification, update the information of the relevant nodes according to the classification result. First, update the sample quantity of each node, adding 1 to the sample numbers of all nodes on the classification path. Then update the feature vector of the node, using the weighted average method, adding the feature vector of the new sample to the original feature vector with a certain weight. For example, the weight can be taken as 0.1, then the updated feature vector is 0.9 times the original vector plus 0.1 times the new sample vector.
[0082] Recalculate the similarity matrix between nodes in the entire classification tree regularly (e.g., after every 1000 songs are classified). Based on the updated similarity matrix, the classification tree structure can be optimized. Specifically, it can include merging nodes with very high similarity, splitting nodes with too many samples, adjusting the hierarchical relationship of nodes, etc. For example, if it is found that the similarity between the "Chinese Pop" and "Hong Kong, Taiwan Pop" nodes reaches 0.95, they can be considered merged into a "Chinese Pop" node.
[0083] Through the above steps, the classification results can be continuously evaluated and the classification tree structure can be optimized, improving the accuracy and efficiency of music genre classification. This method can adapt to the dynamic changes of music genres and continuously improve the classification system as the sample size increases.
[0084] In an optional implementation manner, a hierarchical classification model is constructed. The hierarchical classification model includes multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music genre classification tree; the fused feature vector is input into the hierarchical classification model, and starting from the initial node, refined classification is performed layer by layer downward, including: Normalize the high-level semantic feature vector and the head feature vector to eliminate the dimensional difference between different features; use the random forest algorithm to evaluate the importance of the normalized features and calculate the Gini importance index of each feature; select the most important feature subset according to the feature importance index; Adopt an attention mechanism for feature fusion, define an attention weight matrix, use the softmax function to calculate the weight of each feature, and obtain the fused feature vector; perform principal component analysis dimensionality reduction processing on the fused feature vector to obtain the dimensionality-reduced feature vector; Construct a hierarchical classification model. The hierarchical classification model includes multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music genre classification tree. Each sub-classifier consists of a support vector machine, a random forest, and a gradient boosting decision tree; Customize a feature set for each sub-classifier, including the dimensionality-reduced feature vector and node-specific domain features; use the stratified cross-validation method to train each sub-classifier, and use the Stacking method to fuse the prediction results of the basic models; design a hierarchical decision-making strategy, define a confidence threshold and a node selection function for node selection and termination condition judgment during the classification process; Calculate the feature importance vector of each node, and use the SHAP value to quantify the importance of features; design a weight adjustment function to dynamically adjust the feature weights according to the difference in feature importance between the current node and the parent node; When classifying to the next lower level each time, use a weight adjustment function to update the feature vector; introduce an adaptive mechanism to dynamically adjust the parameters of the weight adjustment function according to the classification performance; input the dimension-reduced feature vector into the hierarchical classification model, and start from the initial node to perform refined classification layer by layer downward, and dynamically adjust the feature weights during the classification process at each level; output the final music genre classification result.
[0085] In the specific implementation manner, first perform standardization processing on the high-level semantic feature vector and the head feature vector to eliminate the dimensional difference between different features. The standardization processing uses the Z-score standardization method, subtracting the mean of each feature and then dividing by the standard deviation.
[0086] Next, use the random forest algorithm to evaluate the importance of the standardized features and calculate the Gini importance index of each feature. The random forest consists of multiple decision trees, and each tree is constructed using a randomly selected subset of features. For each feature, calculate the average decrease in impurity caused when it is used as a splitting node in all decision trees, which is the Gini importance index of the feature. For example, assume there are 10 features, and the Gini importance indices calculated by the random forest algorithm are [0.2, 0.15, 0.1, 0.05, 0.1, 0.08, 0.12, 0.07, 0.06, 0.07].
[0087] According to the feature importance index, select the most important feature subset. A threshold can be set to select features with a Gini importance index greater than the threshold, or select the top k features ranked by the Gini importance index. For example, when setting the threshold to 0.1, the selected feature subset is the 4 features corresponding to [0.2, 0.15, 0.1, 0.12].
[0088] Then, adopt an attention mechanism for feature fusion. Perform principal component analysis dimensionality reduction processing on the fused feature vector to obtain the dimension-reduced feature vector. First, calculate the covariance matrix, then calculate the eigenvalues and eigenvectors of the covariance matrix, select the eigenvectors corresponding to the largest k eigenvalues to form a projection matrix, project the original feature vector into the k-dimensional space, and obtain the dimension-reduced feature vector.
[0089] Construct a hierarchical classification model, including multiple sub-classifiers, and each sub-classifier corresponds to a non-leaf node in the music genre classification tree. Each sub-classifier consists of a support vector machine, a random forest, and a gradient boosting decision tree. The support vector machine uses the RBF kernel function, the random forest contains 100 decision trees, and the gradient boosting decision tree is implemented using the XGBoost algorithm.
[0090] Customize the feature set for each sub-classifier, including the feature vectors after dimensionality reduction and node-specific domain features. For example, for a pop music node, relevant features such as rhythm and harmony can be added; for a classical music node, relevant features such as instruments and timbre can be added.
[0091] Train each sub-classifier using the hierarchical cross-validation method. Divide the dataset into 5 folds and perform 5-fold cross-validation. In each validation, use 4-fold data to train the base model, and the remaining 1-fold data as the validation set. Use the Stacking method to fuse the prediction results of the base models, that is, use the prediction results of the base models as new features to train a meta-classifier (such as logistic regression) to obtain the final prediction result.
[0092] Design a hierarchical decision-making strategy, define a confidence threshold and a node selection function. The confidence threshold is used to determine whether to continue downward classification, and the node selection function is used to select the next sub-node to be classified. For example, the confidence threshold can be set to 0.8, and stop classification when the prediction probability of a certain category is greater than 0.8. The node selection function can select the sub-node with the highest prediction probability to continue classification.
[0093] Calculate the feature importance vector of each node, and use SHAP values to quantify the importance of features. SHAP values reflect the contribution of features to the model prediction, and can be estimated by calculating the difference in the model prediction results before and after the feature is missing. For each feature, calculate the average of the absolute values of its SHAP values to obtain the importance score of the feature.
[0094] Design a weight adjustment function to dynamically adjust the feature weights according to the difference in feature importance between the current node and the parent node. In each classification to the next layer, use the weight adjustment function to update the feature vector. Multiply each feature value in the original feature vector by the corresponding adjusted weight to obtain the updated feature vector.
[0095] Introduce an adaptive mechanism to dynamically adjust the parameters of the weight adjustment function according to the classification performance. For example, a performance evaluation metric (such as F1 score) can be defined, and calculate this metric after each classification. If the metric value increases, increase the adjustment coefficient α; if the metric value decreases, decrease the adjustment coefficient α.
[0096] Input the feature vectors after dimensionality reduction into the hierarchical classification model, start from the initial node, and perform refined classification layer by layer downward. Dynamically adjust the feature weights during the classification process at each level. For example, assume the initial node is "Music", the first-layer classification result is "Pop Music", the second-layer classification result is "Chinese Pop", and the third-layer classification result is "Chinese Male Pop". During the classification process at each layer, the weight adjustment function will be used to update the feature vector to highlight the key features of the current layer.
[0097] Finally, output the final music genre classification result, which is "Mandarin male pop". This result is obtained through multi-level and refined classification, making full use of high-level semantic features and header features, and dynamically adjusting the feature weights during the classification process to improve the accuracy and robustness of the classification.
[0098] Figure 2 FIG. is a schematic structural diagram of a music information classification system based on data code header recognition according to an embodiment of the present invention, as Figure 2 shown, the system includes: A first unit for performing time-frequency domain conversion on the input audio data to obtain a spectrogram; extracting a multi-dimensional feature vector from the spectrogram, inputting the multi-dimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code of the audio data according to the high-level semantic features; A second unit for constructing a data code header recognition model, inputting the generated data code into the data code header recognition model to extract the header features of the data code; performing a preliminary classification on the audio data according to the header features to obtain a preliminary classification result; establishing a music genre classification tree, the music genre classification tree including music genre nodes at multiple levels; positioning an initial node in the music genre classification tree according to the preliminary classification result; A third unit for performing multi-modal feature fusion on the high-level semantic features and the header features to obtain a fused feature vector; constructing a hierarchical classification model, the hierarchical classification model including multiple sub-classifiers, each sub-classifier corresponding to a non-leaf node in the music genre classification tree; inputting the fused feature vector into the hierarchical classification model, and starting from the initial node, performing refined classification layer by layer downward; A fourth unit for dynamically adjusting the feature weights during the classification process at each level to highlight the key features of the current level; updating the music genre classification tree according to the classification result, including adding new music genre nodes or adjusting the relationships between nodes; outputting the final music information classification result.
[0099] In the third aspect of the embodiments of the present invention, There is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; wherein, the processor is configured to call the instructions stored in the memory to execute the method described above.
[0100] In the fourth aspect of the embodiments of the present invention, There is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the method described above is implemented.
[0101] The present invention may be a method, apparatus, system, and / or computer program product. The computer program product may include a computer-readable storage medium having thereon computer-readable program instructions for performing various aspects of the present invention.
[0102] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A music information classification method based on data terminal part recognition, characterized in that: include: Convert the input audio data into time-frequency domain to obtain a spectrum diagram; Extracting a multidimensional feature vector from the spectrogram, inputting the multidimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code of the audio data according to the high-level semantic features; Constructing a data header recognition model, inputting the generated data code into the data header recognition model, and extracting the header features of the data code; Preliminarily classify the audio data according to the head features to obtain a preliminary classification result; establish a music type classification tree, the music type classification tree includes multiple levels of music type nodes; Locating an initial node in the music type classification tree according to the preliminary classification result; The high-level semantic features are multimodally fused with the head features to obtain a fused feature vector; a hierarchical classification model is constructed, wherein the hierarchical classification model includes a plurality of sub-classifiers, each of which corresponds to a non-leaf node in the music type classification tree; the fused feature vector is input into the hierarchical classification model, and refined classification is performed layer by layer starting from the initial node; In the classification process of each level, the feature weights are dynamically adjusted to highlight the key features of the current level; Update the music type classification tree according to the classification results, including adding new music type nodes or adjusting the relationship between nodes; output the final music information classification results.
2. The method according to claim 1, characterized in that Inputting the multidimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code of the audio data according to the high-level semantic features includes: Reshape the multidimensional feature vector into a Mel-spectrogram format; input the Mel-spectrogram into a pre-trained VGGish deep neural network model; extract high-level semantic features through convolutional layer blocks and fully connected layers of the VGGish deep neural network model; Quantize each element in the high-level semantic feature into an 8-bit unsigned integer; perform bit plane decomposition on the quantized feature vector to obtain 8 binary bit planes; calculate the entropy value of each binary bit plane; sort the binary bit planes according to the entropy value, and select the first K binary bit planes with the largest amount of information; Connect the selected K binary bit planes into a long binary sequence, and use run-length coding to compress the long binary sequence; if the compressed length exceeds the target length, truncate it; if the compressed length is less than the target length, fill it with 0 to the target length; The checksum of the compressed sequence is calculated using the CRC-32 algorithm; the checksum is added to the end of the compressed sequence to generate a binary data code of a fixed length.
3. The method according to claim 1, characterized in that Constructing a data header recognition model, inputting the generated data code into the data header recognition model, and extracting the header features of the data code; The audio data is preliminarily classified according to the head features, and the preliminary classification results include: The data terminal part recognition model includes a convolutional neural network and a recurrent neural network; the convolutional neural network includes a plurality of convolution kernels of different sizes, and a plurality of filters are set for each convolution kernel of the size; the recurrent neural network is a bidirectional long short-term memory network; a self-attention mechanism layer is set after the bidirectional long short-term memory network; The data code is converted into a data code sequence represented by numerical values; the converted data code sequence is input into the data header recognition model; multi-scale local features are extracted through the convolutional neural network and maximum pooling operation; the extracted multi-scale local features are input into the bidirectional long short-term memory network to obtain context-related feature representations; the importance weights of different positions are calculated using the self-attention mechanism layer; the weighted features are passed through the fully connected layer to obtain the head features of the data code; The head features are subjected to z-score standardization; the standardized head features are subjected to dimensionality reduction using a principal component analysis method, and the principal components whose explained variance ratio reaches a preset threshold are retained; a random forest classifier is constructed, and the head features after dimensionality reduction are input into the random forest classifier to obtain the probability that each sample belongs to each music category; and the category with the highest probability is selected as the preliminary classification result.
4. The method according to claim 3, characterized in that The method further comprises: The classification performance is evaluated using a cross-validation method to calculate accuracy, precision, recall and F1 score; the misclassified samples are analyzed to identify the types of music that are difficult to classify; based on the analysis results, the parameters of the data center identification model and the parameters of the random forest classifier are adjusted to improve the classification accuracy.
5. The method according to claim 1, characterized in that Establishing a music type classification tree, wherein the music type classification tree includes multiple levels of music type nodes; Locating an initial node in the music type classification tree according to the preliminary classification result comprises: Constructing a music type classification tree, the construction of the music type classification tree includes: defining a multi-level music category, each node contains a unique identifier, a node name, a parent node, a child node list, a feature vector and a number of samples; collecting labeled music samples, using a clustering algorithm to analyze the samples, and dynamically adjusting the tree structure; calculating a representative feature vector of each node, the representative feature vector is a high-level semantic feature mean vector of all music samples of the node; calculating the cosine similarity between all nodes, and constructing an inter-node similarity matrix; Classifying the input music sample, wherein the classifying the input music sample comprises: extracting a high-level semantic feature vector of the music sample; performing preliminary classification on the music sample using a random forest classifier to obtain a probability distribution of each top-level category; selecting the top N top-level categories with the highest probability, and performing a multi-path search on each selected top-level category, wherein the multi-path search comprises starting from a top-level node, calculating the similarity between the feature vector of the music sample and the feature vector of the current node, selecting the child node with the highest similarity to enter the next layer, and repeating this process until a leaf node is reached or the similarity is lower than a preset threshold; Collect all candidate nodes obtained by multi-path search, calculate the similarity between the music sample feature vector and the feature vector of each candidate node, sort the candidate nodes according to the similarity; and select the candidate node with the highest similarity as the initial node.
6. The method according to claim 5, characterized in that Evaluating the classification results and optimizing the classification tree, wherein the evaluating the classification results and optimizing the classification tree comprises: Calculate the similarity between the feature vector of the music sample and the feature vector of the initial node, compare the similarity with a preset threshold, and evaluate the classification confidence; if the confidence is lower than the preset threshold, mark the music sample as "uncertain"; record the classification results, including the initial node, confidence and search path; update the sample quantity and feature vector of the node according to the classification results; regularly recalculate the similarity matrix between nodes to optimize the music type classification tree structure.
7. The method according to claim 1, characterized in that Constructing a hierarchical classification model, the hierarchical classification model includes a plurality of sub-classifiers, each sub-classifier corresponds to a non-leaf node in the music type classification tree; inputting the fused feature vector into the hierarchical classification model, starting from the initial node, and performing refined classification layer by layer, including: Standardize the high-level semantic feature vectors and head feature vectors to eliminate the dimensional differences between different features; use the random forest algorithm to evaluate the importance of the standardized features and calculate the Gini importance index of each feature; select the most important feature subset based on the feature importance index; The attention mechanism is used for feature fusion, the attention weight matrix is defined, and the softmax function is used to calculate the weight of each feature to obtain the fused feature vector; the principal component analysis is performed on the fused feature vector to reduce the dimension and obtain the reduced feature vector; Constructing a hierarchical classification model, wherein the hierarchical classification model includes a plurality of sub-classifiers, each sub-classifier corresponds to a non-leaf node in a music type classification tree, and each sub-classifier is composed of a support vector machine, a random forest, and a gradient boosting decision tree; Customize feature sets for each sub-classifier, including feature vectors after dimension reduction and node-specific domain features; use stratified cross-validation to train each sub-classifier, and use the Stacking method to fuse the prediction results of the basic model; design a hierarchical decision strategy, define confidence thresholds and node selection functions, and use them for node selection and termination condition judgment during the classification process; Calculate the feature importance vector of each node and use SHAP value to quantify the importance of the feature; design a weight adjustment function to dynamically adjust the feature weight according to the difference in feature importance between the current node and the parent node; Each time the classification is performed to the next layer, the feature vector is updated using the weight adjustment function; an adaptive mechanism is introduced to dynamically adjust the parameters of the weight adjustment function according to the classification performance; the feature vector after dimension reduction is input into the hierarchical classification model, and refined classification is performed layer by layer starting from the initial node, and the feature weight is dynamically adjusted during the classification process at each level; and the final music type classification result is output.
8. A music information classification system based on data terminal part recognition, used to implement any of the methods in claims 1-7, characterized in that: include: The first unit is used to convert the input audio data into time-frequency domain to obtain a spectrum diagram; Extracting a multidimensional feature vector from the spectrogram, inputting the multidimensional feature vector into a pre-trained deep neural network model to extract high-level semantic features; generating a data code of the audio data according to the high-level semantic features; The second unit is used to construct a data header recognition model, input the generated data code into the data header recognition model, and extract the header features of the data code; Preliminarily classify the audio data according to the head features to obtain a preliminary classification result; establish a music type classification tree, the music type classification tree includes multiple levels of music type nodes; Locating an initial node in the music type classification tree according to the preliminary classification result; The third unit is used to perform multimodal feature fusion on the high-level semantic features and the head features to obtain a fused feature vector; construct a hierarchical classification model, the hierarchical classification model includes a plurality of sub-classifiers, each sub-classifier corresponds to a non-leaf node in the music type classification tree; input the fused feature vector into the hierarchical classification model, and perform refined classification from the initial node downward layer by layer; The fourth unit is used to dynamically adjust the feature weights during the classification process at each level to highlight the key features of the current level; Update the music type classification tree according to the classification results, including adding new music type nodes or adjusting the relationship between nodes; output the final music information classification results.
9. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to call the instructions stored in the memory to execute the method described in any one of claims 1 to 7.
10. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Music genre identification method and device, equipment and storage medium
CN113450828A
Music genre classification method and system fused with knowledge graph
CN115881160A
Broadcast volume adaptive adjustment method and device, equipment, medium and product
CN119420442A
Book sound classification method using machine learning model of a book handling sounds
KR102259299B1
Inferring emotion from speech in audio data using deep learning
US20240013802A1