Cover Recognition Method, Device and Storage Medium Based on Feature Fusion and Clustering
By adopting feature fusion and clustering methods in cover recognition technology, the problems of insufficient recognition robustness and lack of judgment threshold in the prior art are solved, and more efficient and stable cover recognition performance is achieved.
Patent Information
- Application Number
- CN202211068244.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-31
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-08-31
AI Technical Summary
The existing cover recognition technology has shortcomings in recognition robustness and feature fusion, especially the lack of judgment thresholds in practical application scenarios, resulting in unstable recognition performance.
The cover recognition method based on feature fusion and clustering is adopted. The cover recognition results are determined through the fusion classification feature extraction network, music feature clustering network and binary decision-making network, audio features are extracted and clustered, implicit data feature dimension labels are generated, dimensional information of audio features is enriched, and the cover recognition results are determined through the binary decision-making network.
It improves the robustness of cover recognition, enriches the types of data labels, reduces the difficulty of training of cover recognition models, improves the recognition performance of cover recognition models, and avoids the problem of lack of judgment thresholds.
Smart Images

Figure CN115472181B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of audio processing, and in particular, to a cover song recognition method, device, and storage medium based on feature fusion and clustering. Background Art
[0002] Cover song recognition technology is to retrieve the corresponding cover songs or original factory songs of a given song from a music database, and has always been a research hotspot in the field of music information retrieval. Compared with the original factory songs, the cover songs have changed in rhythm, timbre, pitch, and even structure. The existing public cover song datasets cannot provide more classification label feature information except for the cover version labels. Therefore, the cover song recognition task is difficult and challenging. Traditional cover song recognition technologies include using sequence alignment methods to compare the similarity of two songs to determine whether they are cover versions, and also include using metadata comparison to determine whether they are cover songs; however, the sequence alignment judgment method has poor robustness when encountering non-cover songs with similar music styles, and the metadata comparison judgment method has high requirements for data, and it is difficult to apply to actual scenarios.
[0003] With the development of computer hardware and deep learning, many preprocessing features of audio are two-dimensional features in the time-frequency domain. Therefore, the application of convolutional neural networks to the cover song recognition task has become the mainstream technical solution. There are mainly two technical means in the prior art. One is based on music representation learning technology, which extracts the deep features of the preprocessing features of the input audio. When querying the cover version of a song, the Euclidean distance or cosine similarity is used to calculate the similarity between the query song and the deep features of each song in the music database, and the cover version or original song of the song is queried in the music database by sorting the similarities from large to small. Its training method is to rely on the song types in the public dataset to train the cover song recognition model into a multi-class recognition model while generating the audio deep features. The cover versions in the dataset can be directly determined by the original song version labels; the triplet loss can also be used to train the model to improve the similarity discrimination of the model for dividing cover pairs and non-cover pairs. The music representation learning technology is a single-input model, and a unique structure is added to the model result to improve the invariant features of the model for extracting cover songs for the special music element changes in cover song recognition. This type of method has better results in recognition performance; however, because the query songs in the music database in the actual application scenario may not provide the corresponding cover labels, the song with a large similarity value may not necessarily be its corresponding cover version at this time, which makes it difficult to determine the judgment threshold for cover pairs. Even if an empirical threshold is set, it will lead to inaccurate judgments; on the other hand, there are only cover version labels in the cover song dataset, resulting in single label data and lacking feature space labels that are not restricted by version labels.
[0004] The second is the cover song recognition technology based on the siamese convolutional network. It simultaneously extracts the features of the dual inputs through two branches with shared weights, and performs feature fusion through a fully connected layer or a cross-distance matrix. It takes the siamese network as the main architecture. By simultaneously inputting two songs, it outputs a binary classification result indicating whether the two songs are a cover pair or a non-cover pair. This binary classification model based on the siamese network does not need to rely on special knowledge in the field of cover song recognition and can also avoid the problem of lacking a decision threshold in the above methods in actual application scenarios. However, due to the redundant model structure and excessive model parameters of the siamese network, it is prone to overfitting during training, the actual prediction speed is slow, and this method only extracts spatial domain depth features in the independent branches, lacking the learning of channel dimension differences, and the fusion degree of feature fusion through the fully connected layer is limited. Summary of the Invention
[0005] In view of this, embodiments of the present invention provide a cover song recognition method, device, and storage medium based on feature fusion and clustering to eliminate or improve one or more defects existing in the prior art.
[0006] One aspect of the present invention provides a cover song recognition method based on feature fusion and clustering. The method includes the following steps: Select an original song audio and a to-be-recognized audio as input audios, and respectively extract the audio features of each input audio as input vectors for the cover song recognition model based on feature fusion and clustering; in the cover song recognition model based on feature fusion and clustering, fuse the two extracted audio features along the channel dimension to form a dual-channel fusion feature; use a fusion classification feature extraction network to extract the classification features of the two input audios according to the dual-channel fusion feature; use a music feature clustering network to respectively extract the clustering results of the two audio features; use a binary classification decision network to output the binary classification cover song recognition result of the to-be-recognized audio according to the classification features of the two input audios and the clustering results of each input audio.
[0007] In some embodiments of the present invention, the step of extracting the audio features of each input audio includes: using a pre-trained model to extract the pitch contour features of each input audio as the audio features input to the fusion classification feature extraction network and the music feature clustering network.
[0008] In some embodiments of the present invention, the fusion classification feature extraction network includes a fusion feature extraction structure and a channel separation decision structure; the step of using the fusion classification feature extraction network to extract the classification features of the two input audios includes:
[0009] Fuse the two extracted audio features along the channel dimension to form a dual-channel fusion feature, input the dual-channel fusion feature into the fusion feature extraction structure, and use the fusion feature extraction structure to extract the multi-channel feature map of the dual-channel fusion feature; input the multi-channel feature map into the channel separation decision structure, use the channel separation decision structure to divide the multi-channel feature map into two feature maps of equal size along the channel dimension, obtain the multi-channel cross-feature of the two input audios by calculating the cross-example matrix between the two feature maps in each channel, and extract the classification features of the two input audios according to the multi-channel cross-feature.
[0010] In some embodiments of the present invention, the music feature clustering network includes a convolutional layer and a fully connected layer as the clustering layer. The step of using the music feature clustering network to respectively extract the clustering results of the two audio features includes: using the convolutional layer to respectively extract the depth features of the two audio features, and using the clustering layer to respectively cluster the two extracted depth features to form the clustering results of each input audio.
[0011] In some embodiments of the present invention, the step of using the binary classification decision network to output the binary classification cover song recognition result of the audio to be recognized includes: using the classification features of the two input audios as the input vector of the binary classification decision network, splicing and fusing the clustering results of each input audio with the classification features of the two input audios, and making the clustering results of each input audio participate in the training of the binary classification decision network; outputting the binary classification decision result of whether the two input audios are a cover pair or a non-cover pair through the binary classification decision network, so as to obtain the binary classification cover song recognition result of the audio to be recognized.
[0012] In some embodiments of the present invention, a music feature clustering structure is provided in the music feature clustering network, and the music feature clustering structure completes the training of the music feature clustering network through an autoencoder; the autoencoder includes an encoder and a decoder;
[0013] The training step of the music feature clustering network includes: encoding the audio feature of the input audio through the encoder and clustering to form a clustering result; then reconstructing the clustering result through the decoder to obtain a reconstructed feature corresponding to the audio feature of the input audio; using the error between the audio feature of the input audio and the corresponding reconstructed feature to optimize the clustering loss of the music feature clustering network.
[0014] In some embodiments of the present invention, in the training step of the music feature clustering network, the clustering loss of the music feature clustering network is optimized by using a stochastic gradient descent optimization function.
[0015] In some embodiments of the present invention, the method further includes a training step of the cover song recognition model based on feature fusion and clustering, including: optimizing the classification features of two input audio and the binary classification cover song recognition results through a cross-entropy loss function and an Adam optimization function.
[0016] Another aspect of the present invention provides a cover song recognition device based on feature fusion and clustering, including a processor and a memory. Computer instructions are stored in the memory, and the processor is configured to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the cover song recognition method based on feature fusion and clustering as described above.
[0017] Another aspect of the present invention provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the steps of the cover song recognition method based on feature fusion and clustering as described above.
[0018] The cover song recognition method, device, and storage medium of the present invention. The method extracts the fusion features of the input audio from the spatial dimension through a fusion classification feature extraction network, and analyzes the similarity differences of the fusion features among multiple channels from the channel dimension, enriching the dimensional information of the audio features and improving the robustness of cover song recognition; the music feature clustering network generates invisible data feature dimensional labels for cover song data through feature clustering, enriching the types of data labels, reducing the training difficulty of the cover song recognition model, and improving the recognition performance of the cover song recognition model; the binary classification decision network avoids the problem of lacking a determination threshold in the actual scenario of audio recognition. And in the process of feature fusion, after the audio features in the spatial dimension are extracted, feature fusion is realized through convolutional calculation in the channel dimension, avoiding the limitations of feature fusion and realizing the full-process learnability of feature fusion. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present invention, form a part of this application, and do not limit the present invention.
[0020] Figure 1 It is a flowchart of the full-fusion cover song recognition method based on feature fusion and clustering in the embodiment.
[0021] Figure 2 It is a training structure diagram of the music feature clustering network in the embodiment
[0022] Figure 3 It is a scatter plot of the clustering results of the benchmark test dataset;
[0023] Figure 4It is a schematic diagram showing the influence of the decision threshold on the performance of the precise and scalable version identification model (MOVE) for music motive embedding in the SHS5 dataset;
[0024] Figure 5 It is a schematic diagram showing the influence of the decision threshold on the performance of the precise and scalable version identification model (MOVE) for music motive embedding in the Da-Tacos dataset. Detailed implementation manners
[0025] To make the objectives, technical solutions and advantages of the present invention more clear and understandable, the present invention will be further described in detail below in conjunction with the implementation manners and the drawings. Herein, the illustrative implementation manners of the present invention and their descriptions are used to explain the present invention, but are not intended to limit the present invention.
[0026] Herein, it should also be noted that in order to avoid obscuring the present invention due to unnecessary details, only the structures and / or processing steps closely related to the solution of the present invention are shown in the drawings, while other details less related to the present invention are omitted.
[0027] It should be emphasized that the term "comprising / including" when used herein refers to the presence of features, elements, steps or components, but does not exclude the presence or addition of one or more other features, elements, steps or components.
[0028] Herein, it should also be noted that if not otherwise specified, the term "connection" in this document can refer not only to direct connection, but also to indirect connection with an intermediate.
[0029] In the following, embodiments of the present invention will be described with reference to the drawings. In the drawings, the same reference numerals represent the same or similar components, or the same or similar steps.
[0030] In this embodiment, a cover song recognition method based on feature fusion and clustering is proposed, and its main workflow is as Figure 1 shown, including steps S110 - S140:
[0031] In step S110: Select an original song audio and a song to be recognized as input audios, and extract the audio features of each input audio respectively as the input vectors of the cover song recognition model based on feature fusion and clustering. The cover song recognition model based on feature fusion and clustering includes a fusion classification feature extraction network, a music feature clustering network and a binary classification decision network.
[0032] In the embodiment, it is necessary to first determine the original audio that the audio to be recognized may correspond to, use the audio to be recognized and the corresponding original audio as a pair of input audios, perform feature extraction on the two input audios respectively, and obtain the audio features of each input audio, which are used as the input vectors of the cover song recognition model based on feature fusion and clustering. The audio feature is a two-dimensional vector in the spatial dimension.
[0033] In the embodiment, a pre-trained model is used to extract the pitch contour features of each input audio as the audio features of each input audio, use the audio features as the input vectors of the cover song recognition model based on feature fusion and clustering, and uniformly crop each audio feature to a size of 23*1800.
[0034] In step S120, in the cover song recognition model based on feature fusion and clustering, the two extracted audio features are fused along the channel dimension to form a two-channel fusion feature; a fusion classification feature extraction network is used to extract the classification features of the two input audios according to the two-channel fusion feature;
[0035] In the embodiment, in the cover song recognition model based on feature fusion and clustering, the feature sizes of the input audios are unified, the two audio features are uniformly cropped to a size of 23*1800, and the two audio features with unified feature sizes are weighted combined and mapped between channels through convolution calculation along the channel dimension to obtain the fusion features of the two input audios, that is, the two-channel fusion feature. The two-channel fusion feature is a three-dimensional vector in the spatial dimension and the channel dimension; corresponding to the feature size of the audio feature in the above embodiment, the size of the two-channel fusion feature in this embodiment is 2*23*1800. The convolution calculation formula for the two audio features along the channel dimension is:
[0036]
[0037] Among them, N i represents the number of audio features in the batch processing in the convolution calculation, i represents the serial number of the audio feature in the batch processing of audio features, represents the size of the output channel dimension, j represents the serial number of the output channel dimension, k represents the channel serial number of the input audio feature, C in represents the number of channels of the input audio feature, out(·) represents the output vector of the convolution calculation, bias(·) represents the offset, weight(·) represents the weight convolution calculation of each input audio feature, and input(·) represents the input vector of the convolution calculation.
[0038] In an embodiment, the music fusion classification feature extraction network includes a fusion feature extraction structure and a channel separation decision structure; the fusion feature extraction structure extracts multi-channel fusion features of the dual-channel fusion feature through a convolutional block to obtain a multi-channel feature map of the dual-channel fusion feature; the channel separation decision structure divides the multi-channel feature map into two feature maps of equal size along the channel dimension, and obtains the multi-channel cross features of the two input audio by calculating the cross-example matrix between each channel of the two feature maps, so as to reflect the similarity difference between the two input audio as a cover pair and a non-cover pair, and extracts the classification features of the multi-channel cross features through convolution as the classification features of the two input audio output by the music fusion classification feature extraction network. The classification features are three-dimensional vectors in the spatial dimension and the channel dimension.
[0039] In one embodiment, the feature size of the multi-channel feature map output by the fusion feature extraction structure is 512*H*W, where H represents the height of the multi-channel feature map, W represents the width of the multi-channel feature map, and 512 represents the number of channels. The feature sizes of the two feature maps output by the corresponding channel separation decision structure are both 256*H*W, and the feature size of the obtained multi-channel cross features is 256*H*W.
[0040] Aiming at the problems of structural redundancy and feature fusion limitations in the cover song recognition technology based on the siamese convolutional network in the prior art, the fusion classification feature extraction network in this method realizes the fusion of audio features through convolution in the channel dimension, enriches the dimensional information of audio features, avoids the limitations of feature fusion, improves the robustness of cover song recognition, and realizes the full-process learnability of feature fusion parameters.
[0041] In step S130, the music feature clustering network is used to extract the clustering results of the two audio features respectively.
[0042] In the embodiment, the music feature clustering network includes a convolutional layer and a fully connected layer as the clustering layer. The audio features of each input audio are used as the input of the convolutional layer. The convolutional layer is used to extract the depth features of the two audio features respectively, and the clustering layer is used to cluster the two extracted depth features respectively to form the clustering results of each input audio.
[0043] The music feature clustering network is also provided with a music feature clustering structure for completing the training of the music feature clustering network. The music feature clustering structure completes the training of the music feature clustering network through an autoencoder. The autoencoder includes an encoder and a decoder. The training process of the music feature clustering network is as Figure 2As shown, the audio features of the input audio are used as the input of the encoder. The encoder encodes the audio features through a convolutional layer, and then clusters the encoded audio features through a fully connected layer serving as a clustering layer to form a clustering result of the input audio. Then, the clustering result is used as the input of the decoder, and the decoder reconstructs the clustering result through inverse convolution calculation to obtain a reconstructed feature corresponding to the audio features of the input audio. Using the error between the audio features of the input audio and the corresponding reconstructed features, the stochastic gradient descent optimization function (SGD optimization function) is selected to optimize the clustering loss of the music feature clustering structure, making the audio features input to the encoder as close as possible to the reconstructed features output by the decoder, thereby completing the training of the music feature clustering structure.
[0044] In one embodiment, the trained music feature clustering network is used to extract and cluster the deep features of the audio to be recognized. If the size of the clustering layer is set to 50, it means that in this music feature clustering network, the audio features of two input audios are both clustered into 50 categories in the deep feature space, generating corresponding clustering results.
[0045] In step S140, a binary classification decision network outputs a binary classification cover song recognition result of the audio to be recognized according to the classification features of two input audios and the clustering results of each input audio.
[0046] In the embodiment, the binary classification decision network realizes the binary classification cover song recognition of whether two input audios are a cover song pair or a non-cover song pair through a fully connected layer. The classification features of the two input audios are used as the input vectors of the fully connected layer of the binary classification decision network, and the clustering results of each input audio are concatenated and fused with the classification features of the two input audios, making the clustering results of each input audio participate in the training of the binary classification decision network. In the fully connected layer serving as the binary classification decision network, the product of the three dimensions of the classification features of the two input audios and the sum of the clustering numbers of the clustering results of each input audio are used as the input dimension of the fully connected layer serving as the binary classification decision network. The output dimension of the fully connected layer serving as the binary classification decision network is set to 2, so that the binary classification cover song recognition result of whether two input audios are a cover song pair or a non-cover song pair can be output through the fully connected layer serving as the binary classification decision network, thereby obtaining the binary classification cover song recognition result of the audio to be recognized.
[0047] In the embodiment, the training steps of the above cover song recognition model based on feature fusion and clustering include: optimizing the classification features of two input audios and the binary classification cover song recognition results through a cross-entropy loss function and an Adam optimization function.
[0048] Aiming at the problem that there is a lack of decision thresholds and label data in music representation learning in the existing technology, in this method, by introducing a music feature clustering network, the dimensionality labels of the implicit data features of the cover audio are generated, thus enriching the types of data classification label features; by analyzing the similarity differences between multiple channels from the channel dimension through the music fusion classification feature extraction network, the dimensionality information of the audio features is enriched, and the robustness of the cover recognition model is improved; also, through the binary classification decision network, the certainty of the cover recognition result is determined, thus avoiding the problem that the cover recognition model lacks a decision threshold.
[0049] Regarding the feasibility and effectiveness of the above cover recognition method based on feature fusion and clustering, the following analysis is carried out:
[0050] 1. Based on the above cover recognition method based on feature fusion and clustering, for cover songs with changes in music elements such as timbre, there are no changes in the music style and emotional elements. The cover version and the original version still maintain the principle of similarity in the deep feature space. Taking the clustering of the Da-Tacos dataset as an example, the clustering results of the audio features are analyzed. The clustering results of the Da-Tacos dataset are as Figure 3 shown. The x, y, and z axes in the figure respectively represent the attribute values of the song category, clustering result, and clustering label; it can be seen from the figure that the clustering results of songs with the same cover version number are mostly the same, indicating that the clustering results of the autoencoder conform to the invariant features between the original song and the cover version, and at the same time verifying that it is reasonable for audio feature clustering to provide feature dimension data labels for cover song recognition.
[0051] 2. Aiming at the problem of the lack of feature dimension labels in the existing technology, in order to verify that the method described in the present invention can solve this problem, the influence of three different structures in the present invention on the cover audio recognition performance is verified. In this embodiment, based on the music feature fusion structure in the present invention, the channel separation decision structure and the music feature clustering structure are sequentially added to evaluate the audio cover recognition performance of the three structures. The comparison results are shown in Table 1:
[0052] Table 1 Ablation experiment results table for exploring the channel separation decision structure and the music feature clustering structure
[0053]
[0054] Among them, MAP is a cover song recognition evaluation metric, representing the average of the average classification accuracies; P@10 is a cover song recognition evaluation metric, representing the average number of determined cover song versions among the top 10; MR1 is a cover song recognition evaluation metric, representing the average rank of the first recognized cover song version; baseline represents a music feature fusion structure, CSDS represents a channel separation decision structure, and MFCS represents a music feature clustering structure; SHS-TEST, Covers80, and Da-Tacos all represent a dataset.
[0055] As can be seen from Table 1 above, when only the music feature fusion structure is used for feature extraction, since only the time-frequency dimension features are extracted and the feature analysis between channels is lacking, the model effect is average; when the music feature fusion structure and the channel separation decision structure are used for feature extraction, the model feature extraction and difference analysis are more comprehensive, and significant performance improvements are achieved on the three test sets of SHS-TEST, Covers80, and Da-Tacos; when the music feature fusion structure, the channel separation decision structure, and the music feature clustering structure are used for feature extraction, implicit feature clustering labels can be generated, thus enriching the label types and improving the model recognition performance. It can also be seen from the experimental results that the addition of the music feature clustering structure further improves the cover song recognition performance on the three test datasets.
[0056] 3. Aiming at the problem that the audio representation learning method in the prior art is difficult to be widely promoted in practical applications because of the lack of a determination threshold and only relying on similarity, in order to verify that the method described in the present invention can solve this problem, the influence of the determination threshold on the audio cover song recognition prediction performance of the audio representation learning method and the method described in the present invention is verified: every 0.05 size between 0.1 and 0.9 is used as the determination threshold to explore the influence of the change of the determination threshold on the MOVE model in the music representation learning method and the prediction performance of the binary classification described in the present invention for audio representation features in the actual application scenario: Figure 4 It shows that 100 original songs in the Da-Tacos validation set are used, and the MOVE model using the music representation learning method and the binary classification method described in the present invention are used to predict the audio representation features in the validation set, and the Euclidean distances of 13 cover song versions of the original songs and 19 unrelated songs are calculated; Figure 5 It shows that 77 original songs in the SHS5 validation set are used, and the MOVE model using the music representation learning method and the binary classification method described in the present invention are used to predict the audio representation features in the validation set, and the Euclidean distances of 16 cover song versions of the original songs and 16 unrelated songs are calculated, and the Euclidean distance values are normalized to the range of 0 to 1. The smaller the value, the more similar. Figure 4 and Figure 5Among them, accuracy represents accuracy, precision represents precision, and recall represents regression ability.
[0057] From Figure 4 and Figure 5 It can be seen that the binary classification method in the present invention is not affected by the decision threshold in the application scenario, and the model performance is stable. However, the stability and accuracy of audio representation learning methods such as the MOVE model are directly affected by the decision threshold. Moreover, the decision threshold has different effects on the Da-Tacos and SHS5 data sets. Although the overall performance trend is the same, the ideal decision threshold in SHS5 should be between 0.5 and 0.6, while the ideal decision threshold for the Da-Tacos data set is between 0.4 and 0.5. Therefore, the ideal performance thresholds for different data sources are different, and it is difficult to determine a definite decision threshold using an empirical threshold. Therefore, this method is difficult to apply in actual scenarios.
[0058] 4. To verify the recognition performance of the method described in the present invention, on the cover test data sets of SHS-TEST, Covers80, and Da-Tacos, the test performance of the method described in the present invention is compared with that of other methods, and the comparison results are shown in Table 2:
[0059] Table 2: Performance test results table for exploring various recognition methods
[0060]
[0061] Among them, key-Invariant represents a key-invariant convolutional neural network effectively used for cover song recognition, MulKINet represents a multi-stage key-invariant convolutional neural network for accurate and fast cover song recognition, KDTN represents a neural network based on aggregation learning, CQT-TPPNet represents a convolutional neural network based on a temporal pyramid pooling, SCMM represents an identification model based on a cross-similarity matrix of a multi-level depth sequence, MOVE represents an accurate and scalable version identification model using music motif embedding, and Re-MOVE represents a faster and more accurate music motif embedding cover song recognition model.
[0062] As can be seen from Table 2 above, in the SHS-TEST dataset, the MAP metric of the method described in the present invention successfully breaks through 0.8; in the Covers80 dataset, although the method described in the present invention is not as good as SCMM, compared with other methods, its performance in all aspects has been significantly improved. The method described in the present invention performs best on the Da-Tacos dataset because our method extracts relevant features from the spatial domain and channel dimension by fusing the feature extraction structure and the channel separation decision structure, and the music feature clustering structure can separate the cover version from unrelated works in the high-dimensional feature space. Therefore, as the size of the test dataset increases, its performance is less affected.
[0063] 5. Regarding the cover recognition similarity method based on a multi-level deep sequence cross matrix in the prior art, which also uses a binary classification method for a deep decision network to learn the similarity distribution matrix of a query song and a reference song and determine whether two songs are a cover pair. This solution proposes a multi-level sequence cross matrix based on a siamese structure and feature fusion through a cross matrix. The features after each convolution operation are effectively fused, solving the problem of limited feature fusion degree of the cover recognition method based on a siamese network. However, the redundant weight sharing branches not only fail to solve the problem, but instead make the model structure more complex, with more model parameters and slower prediction speed. When performing comparison and recognition in a large-scale music database, it will consume a large amount of time. Therefore, the method described in the present invention is also advanced and competitive in comparison.
[0064] 6. Regarding the music-driven accurate and scalable cover version recognition solution in the prior art, which also extracts audio depth features in the spatial dimension and channel dimension by combining dilated convolution and improves the summary ability of the model in the time domain information. Although this solution makes an innovation in the channel dimension information in terms of the model structure, it still belongs to a type of music representation learning and still lacks a deterministic decision threshold in the actual application scenario. Therefore, the method described in the present invention is also advanced and competitive in comparison.
[0065] Correspondingly, the present invention also provides a cover recognition device based on feature fusion and clustering. The device includes a computer device, the computer device includes a processor and a memory, the memory stores computer instructions, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method described above.
[0066] An embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing cover song recognition method based on feature fusion and clustering are implemented. The computer-readable storage medium may be a tangible storage medium, such as a random access memory (RAM), memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, register, floppy disk, hard disk, removable storage disk, CD-ROM, or any other form of storage medium well-known in the technical field.
[0067] Those of ordinary skill in the art should understand that the various exemplary components, systems, and methods described in conjunction with the embodiments disclosed herein can be implemented in hardware, software, or a combination of both. Specifically, whether to implement in hardware or software depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention. When implemented in hardware, it can be, for example, an electronic circuit, an application-specific integrated circuit (ASIC), appropriate firmware, a plug-in, a functional card, and so on. When implemented in software, the elements of the present invention are programs or code segments used to execute the required tasks. The program or code segment can be stored in a machine-readable medium or transmitted through a data signal carried in a carrier wave on a transmission medium or a communication link.
[0068] It should be clear that the present invention is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present invention is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order between steps after understanding the spirit of the present invention.
[0069] In the present invention, the features described and / or illustrated for one embodiment can be used in the same or similar manner in one or more other embodiments, and / or combined with the features of other embodiments or replace the features of other embodiments.
[0070] The above are only the preferred embodiments of the present invention and are not used to limit the present invention. For those skilled in the art, various changes and modifications can be made to the embodiments of the present invention. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A cover song recognition method based on feature fusion and clustering, characterized in that, the method comprises the following steps: Select an original song audio and a to-be-recognized audio as input audios, and respectively extract the audio features of each input audio as the input vectors of the cover song recognition model based on feature fusion and clustering; In the cover song recognition model based on feature fusion and clustering, fuse the two extracted audio features along the channel dimension to form a two-channel fusion feature; use the fusion classification feature extraction network to extract the classification features of the two input audios according to the two-channel fusion feature; Use the music feature clustering network to respectively extract the clustering results of the two audio features; Use the binary classification decision network to output the binary classification cover song recognition result of the to-be-recognized audio according to the classification features of the two input audios and the clustering results of each input audio.
2. The method according to claim 1, characterized in that, the step of extracting the audio features of each input audio includes: using a pre-trained model to extract the pitch contour features of each input audio as the audio features input to the fusion classification feature extraction network and the music feature clustering network.
3. The method according to claim 1, characterized in that, the fusion classification feature extraction network includes a fusion feature extraction structure and a channel separation decision structure; the step of using the fusion classification feature extraction network to extract the classification features of the two input audios includes: Fuse the two extracted audio features along the channel dimension to form a two-channel fusion feature, input the two-channel fusion feature into the fusion feature extraction structure, and use the fusion feature extraction structure to extract the multi-channel feature map of the two-channel fusion feature; Input the multi-channel feature map into the channel separation decision structure, use the channel separation decision structure to divide the multi-channel feature map into two feature maps of equal size along the channel dimension, obtain the multi-channel cross features of the two input audios by calculating the cross example matrix between each channel of the two feature maps, and extract the classification features of the two input audios according to the multi-channel cross features.
4. The method according to claim 1, characterized in that, the music feature clustering network includes a convolutional layer and a fully connected layer as the clustering layer, and the step of using the music feature clustering network to respectively extract the clustering results of the two audio features includes: Use the convolutional layer to respectively extract the depth features of the two audio features, and use the clustering layer to respectively cluster the two extracted depth features to form the clustering results of each input audio.
5. The method according to claim 1, characterized in that, the step of using the binary classification decision network to output the binary classification cover song recognition result of the to-be-recognized audio includes: Use the classification features of the two input audios as the input vectors of the binary classification decision network, splice and fuse the clustering results of each input audio with the classification features of the two input audios, and make the clustering results of each input audio participate in the training of the binary classification decision network; Output the binary classification decision result of whether the two input audios are a cover pair or a non-cover pair through the binary classification decision network, so as to obtain the binary classification cover recognition result of the audio to be recognized.
6. The method according to claim 1, wherein, a music feature clustering structure is provided in the music feature clustering network, and the music feature clustering structure completes the training of the music feature clustering network through an autoencoder; the autoencoder includes an encoder and a decoder; The training steps of the music feature clustering network include: encoding the audio features of the input audio through the encoder and clustering to form a clustering result; then reconstructing the clustering result through the decoder to obtain a reconstructed feature corresponding to the audio features of the input audio; using the error between the audio features of the input audio and the corresponding reconstructed features to optimize the clustering loss of the music feature clustering network.
7. The method according to claim 6, wherein, in the training steps of the music feature clustering network, a stochastic gradient descent optimization function is used to optimize the clustering loss of the music feature clustering network.
8. The method according to claim 1, wherein, the method further includes the training steps of the cover recognition model based on feature fusion and clustering, including: optimizing the classification features and binary classification cover recognition results of the two input audios through a cross-entropy loss function and an Adam optimization function.
9. A cover recognition device based on feature fusion and clustering, including a processor and a memory, wherein, computer instructions are stored in the memory, and the processor is used to execute the computer instructions stored in the memory. When the computer instructions are executed by the processor, the device implements the steps of the method according to any one of claims 1 to 8.
10. A computer-readable storage medium, on which a computer program is stored, wherein, when the program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Cover version identification method and device and computer storage medium
CN111445923A
Data fusion methods and systems
WO2008092149A2