Turning and singing recognition model training method, turning and singing recognition method and related device
By training a neural network model of song melody feature vectors, the robustness problem of cover song recognition technology in complex scenarios was solved, and a cover song recognition model that can accurately identify the original song of a cover song was trained with a small amount of data.
Patent Information
- Application Number
- CN202511104564.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-11-04
AI Technical Summary
Existing cover song recognition technologies struggle to ensure robustness when faced with complex cover song adaptations, especially AI-generated cover songs, resulting in poor recognition performance.
By obtaining the melody feature vector of the song and using the initial cover song recognition model for feature processing and optimization, a target cover song recognition model is trained. The neural network is then used to automatically learn melody matching, and the training samples are expanded to cover more complex scenarios.
A more robust cover song recognition model was trained with limited data, which can accurately identify the original song corresponding to the cover song, improving the accuracy of cover song recognition and its ability to adapt to complex scenarios.
Smart Images

Figure CN120895056A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of audio processing, in particular to a cover song identification model training method, a cover song identification method and related devices. BACKGROUND
[0002] Cover song identification technology is an important supplement to song recognition technology. Song recognition technology realizes accurate matching through audio fingerprints. When the recording sample is the original song, it can still be accurately identified even with some background noise. However, when it is necessary to identify some new cover songs or original works, fingerprint matching cannot achieve ideal results, because new works cannot be timely extracted into the database, and because there are differences between cover and original works, making it difficult to match. Therefore, cover song identification technology for fuzzy matching of recording segments has become a key supplement to song recognition.
[0003] Nowadays, cover versions are becoming more and more complex, not only because the threshold for creation is lower, but also because AI-generated cover works are diverse. Actual cover identification scenarios are often more complex. Due to various reasons, the differences between cover and original songs make it difficult to ensure the robustness of cover identification, and the cover identification effect is not good. SUMMARY
[0004] Embodiments of the present application provide a cover song identification model training method, a cover song identification method and related devices. The melody features of the song audio in the training sample are disturbed to obtain more training samples, and a cover song identification model with higher robustness can be trained with less data.
[0005] The first aspect of the embodiments of the present application provides a cover song identification model training method, the method comprising:
[0006] obtaining an initial cover song identification model;
[0007] obtaining a melody feature vector of a target song, a positive example melody feature vector of a song corresponding to the same melody as the target song, and a negative example melody feature vector of a song corresponding to a different melody from the target song;
[0008] performing feature processing on the melody feature vector, the positive example melody feature vector and the negative example melody feature vector through the initial cover song identification model to obtain a melody feature, a positive example melody feature and a negative example melody feature;
[0009] optimizing the initial cover model based on the melody feature vector, the positive example melody feature vector and the negative example melody feature vector until a convergence condition is reached to obtain a target cover song identification model; the target cover song identification model is used to identify the original song corresponding to the cover song audio.
[0010] The second aspect of the embodiment of the present application provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the method of the first aspect when executing the computer program.
[0011] The third aspect of the embodiment of the present application provides a computer storage medium, and the computer storage medium stores instructions, which, when executed on a computer, cause the computer to execute the method of the first aspect.
[0012] The fourth aspect of the embodiment of the present application provides a computer program product, which, when executed on a computer device, causes the computer device to execute the method of the first aspect.
[0013] From the above technical solutions, the embodiment of the present application has the following advantages:
[0014] Rich melody feature data can be obtained through a large number of song library data or real scene singing song data, and the cover identification model can be trained based on the rich melody feature data, so that the model has the ability to distinguish the similarity or difference of melodies of different songs. The model is trained based on song data of various song singing scenes, so as to ensure that the model training can cover and match more complex cover scenes, which helps to improve the robustness of the cover identification of the model. BRIEF DESCRIPTION OF DRAWINGS
[0015] Figure 1 The network framework in the embodiment of the present application is shown in the figure;
[0016] Figure 2 The figure is a flowchart of the cover identification model training method in the embodiment of the present application;
[0017] Figure 3 The figure is a schematic diagram of an exemplary application scenario of the cover identification model training method in the embodiment of the present application;
[0018] Figure 4 The figure is a flowchart of the cover identification method in the embodiment of the present application;
[0019] Figure 5 The figure is a structural diagram of the computer device in the embodiment of the present application. DETAILED DESCRIPTION
[0020] The embodiment of the present application provides a cover identification model training method, a cover identification method and related devices, the melody features of song audio in the training sample are disturbed to obtain more training samples, and a cover identification model with higher robustness can be trained with less data.
[0021] Please refer to Figure 1The network framework in the embodiments of the present application includes:
[0022] The business server 100 and the terminal cluster; the terminal cluster can include: terminal device 200a, terminal device 200b, terminal device 200c,..., terminal device 200n, and the like.
[0023] The business server 100 can be a stand-alone physical server, a server cluster or a distributed system composed of multiple physical servers, a cloud server providing cloud database, cloud service, cloud computing, cloud function, cloud storage, network service, cloud communication, middleware service, domain name service, security service, CDN, and basic cloud computing services such as big data and artificial intelligence platform. The terminal device (including terminal device 200a, terminal device 200b, terminal device 200c,..., terminal device 200n) can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a palm computer, a mobile internet device (MID), a wearable device (such as a smart watch, a smart bracelet, etc.), a smart computer, a smart car, and the like.
[0024] The business server 100 and each terminal device in the terminal cluster can establish a communication connection, and each terminal device in the terminal cluster can also establish a communication connection. In other words, the business server 100 can establish a communication connection with each terminal device in the terminal device 200a, the terminal device 200b, the terminal device 200c,..., and the terminal device 200n. For example, the terminal device 200a and the business server 100 can establish a communication connection. The terminal device 200a and the terminal device 200b can establish a communication connection, and the terminal device 200a and the terminal device 200c can also establish a communication connection. The above communication connection is not limited to the connection mode, which can be directly or indirectly connected through a wired communication mode, or directly or indirectly connected through a wireless communication mode, and the like. The specific application scenario can be determined, and the present application does not limit it here.
[0025] It should be understood that, as Figure 1Each terminal device in the illustrated terminal cluster can be installed with an application client, which, when running in each terminal device, can respectively interact with the service server 100 to enable the service server 100 to receive service data (such as user identity data uploaded by a user through a terminal device) from each terminal device. The application client can be a music playing application, a karaoke software application, a browser application, a social application, an instant messaging application, a live broadcast application, a game application, a short video application, a video application, a shopping application, a novel application, a payment application, or any other application client having the function of displaying data information such as text, images, audio, and video. The specific application client can be determined according to actual application scenarios, and is not limited herein. The application client can be a standalone client or an embedded sub-client integrated in a certain client (e.g., a music application or a karaoke software application). The specific application client can be determined according to actual application scenarios, and is not limited herein.
[0026] Nowadays, cover versions of songs are becoming more and more complex, not only because the threshold for creation is lower, but also because AI-generated cover works are diverse, which poses a severe challenge to the improvement of cover recognition technology. In related solutions,
[0027] The cover recognition technology closest to the present case adopts a pitch matching method, and the process includes: first, cutting the sentences of the to-be-predicted song audio to obtain song segments; extracting the human voice pitch sequence from the song segments, and the pitch sequence extraction algorithm can use pin / pyin / creep technology; after cutting the sentences of the song library songs and extracting the pitch sequence, a search library is obtained; the DTW (dynamic time warping) algorithm is used to calculate the DTW similarity of the pitch sequences of the to-be-predicted song and the song library song, which is used to represent the distance between the two sequences; the matching segment with the highest similarity is recalled as the matching result of the current to-be-predicted segment. If the to-be-recognized song has multiple segments, the matching result of each segment can be obtained, and the song corresponding to the most segments of the to-be-recognized song matched in the search library is taken as the final cover recognition result.
[0028] Using the pitch sequence recognition can directly reflect the singing content. Although different covers have differences in tone, singing environment, and slight modifications, the similarity of the pitch sequence can directly reflect whether two different songs have a cover relationship. However, the actual cover recognition scene is often more complex, and the extracted pitch sequence may not be able to be simply matched to the original song. Due to various reasons, the differences between the cover and the original song need to be learned through a model, and the DTW algorithm recognition method of the prior art cannot learn such differences.
[0029] Based on the above technical problems, the embodiment of the present application proposes a method of encoding the melody feature sequence and training the neural network, so that the melody matching process is automatically learned by the network, and the robustness is stronger than the related scheme.
[0030] In the application scenario of the embodiment of the present application, the cover song recognition technology has a wide range of applications, mainly of which is used in song recognition. For example, in some scenarios, the user can use the song recognition function of the music software to identify a live singing music heard by the user in a bar, a concert, a street performance, etc., or identify a song performed in a live concert. The music software will feed back the identification result to the user based on the running of the cover song recognition function, and the identification result can be the song name, the original singer, the lyricist, etc.
[0031] Since there are a large number of cover version songs in the music platform song library, and a large number of cover songs are added every day, the fingerprint library of song recognition cannot cover all of them, and cover song recognition can make up for this gap and improve the recall rate of song recognition. In addition, cover song recognition is also the key to solving the problem that song recognition cannot identify improvised and rearranged songs. However, cover song recognition is usually very challenging to the robustness of the cover song recognition model due to the complexity of the recording scene, equipment, environment, and rearrangement method. The embodiment of the present application is a method for improving the robustness of the cover song recognition model, which can bring better results to cover song recognition.
[0032] The cover song recognition model training method in the embodiment of the present application will be described below in combination with the network framework of Figure 1 and the application scenario of the embodiment of the present application:
[0033] Please refer to Figure 2 The cover song recognition model training method in the embodiment of the present application includes:
[0034] 201, obtaining an initial cover song recognition model;
[0035] The method of the embodiment can be applied to a computer device, which can be a business server 100 or each terminal device in the network framework shown in Figure 1 The computer device can obtain an initial cover song recognition model to train the model. The initial cover song recognition model can be any deep learning network architecture that can be used to process sequence data, for example, it can be a residual network ResNet (Residual Network) such as ResNet18 network, or a convolutional recurrent neural network (Convolutional Recurrent Neural Network, CRNN).
[0036] 202、obtaining a melody feature vector of a target song, a positive example melody feature vector of a song corresponding to the same melody as the target song, and a negative example melody feature vector of a song corresponding to a different melody than the target song;
[0037] The computer device can obtain a melody feature vector of any target song, a positive example melody feature vector of a song corresponding to the same melody as the target song, and a negative example melody feature vector of a song corresponding to a different melody than the target song. The melody feature vector described above can be obtained by encoding the melody information of a song, for example, the melody information can be the pitch of a note, the duration (time value) of a note, and various feature information about the components of the melody. The melody feature vector can be obtained by encoding the feature information of the melody. The melody feature encoding can be in the form of a vector to represent the feature information of the melody, so that the computer can understand and recognize the melody of the song audio based on the vector form.
[0038] Each target song and its corresponding melody feature vector, its corresponding positive example melody feature vector, and its corresponding negative example melody feature vector can form a set of training data. The training data corresponding to each of the multiple target songs can form multiple sets of training data, which can be used to train the initial cover recognition model.
[0039] 203、performing feature processing on the melody feature vector, the positive example melody feature vector, and the negative example melody feature vector through the initial cover recognition model to obtain melody features, positive example melody features, and negative example melody features;
[0040] After obtaining the melody feature vector, the positive example melody feature vector, and the negative example melody feature vector, the initial cover recognition model will perform feature processing on these feature vectors. The feature processing process can include feature extraction, feature transformation, feature selection, and other operations, aiming to extract more critical information for cover recognition tasks from the original feature vectors, such as more detailed pitch change features, rhythm pattern features, etc. These features can more accurately reflect the subtle differences in song melodies. The melody features, positive example melody features, and negative example melody features after feature processing will be used in the subsequent model training process to learn how to distinguish the melody similarity or difference between different songs.
[0041] 204、optimizing the initial cover model based on the melody features, the positive example melody features, and the negative example melody features until a convergence condition is reached to obtain a target cover recognition model; the target cover recognition model is used to identify the original song corresponding to the cover song audio;
[0042] When the initial cover song identification model obtains the melody feature of the target song, the corresponding positive example melody feature, and the negative example melody feature, and other groups of training data, it can learn the similarities between multiple song audios corresponding to the same melody in the melody feature, such as the similarities between the original song and the corresponding cover song, by using the melody feature of the target song and the corresponding positive example melody feature. In addition, it can learn the differences between multiple song audios corresponding to different melodies in the melody feature by using the melody feature of the target song and the corresponding negative example melody feature.
[0043] Through the learning of the model on the above-mentioned multiple groups of training data, the model is gradually optimized, so that the model has the ability to distinguish the melody similarity or difference between multiple song audios. The optimization process usually involves adjusting the parameters of the model to minimize prediction errors or maximize model performance. This can be achieved through optimization techniques such as backpropagation algorithms. When the performance of the model on the training data is steadily improved and reaches the preset convergence condition, the training process is completed, and the trained target cover song identification model is obtained. The target cover song identification model can be used to identify the original song corresponding to the cover song audio.
[0044] The target cover song identification model has strong generalization ability and can accurately identify the original song corresponding to various cover song audios, even in complex recording scenarios, devices, environments, or adaptation methods. This is due to the model learning rich melody feature data and similarity or difference information between melodies during the training process.
[0045] Therefore, in this embodiment, rich melody feature data can be obtained through a large number of song library data or real scene singing song data. Training the cover song identification model based on rich melody feature data can enable the model to distinguish the melody similarity or difference between different songs. Training the model based on song data from various song singing scenarios ensures that the model training can cover and match more complex cover scenarios, which helps to improve the robustness of the model's cover song identification.
[0046] Based on Figure 2 In an optional implementation of the embodiment shown in the figure, when obtaining the melody feature vectors of different songs, multiple groups of positive example samples and multiple groups of negative example samples can be obtained. Each group of positive example samples includes multiple song audios corresponding to the same song, i.e., the multiple song audios in the positive example sample correspond to the same melody, such as the multiple song audios in the positive example sample, which can be different singing versions of a song (such as the original version and the cover version of a song). Each group of negative example samples includes multiple song audios corresponding to different songs, i.e., the multiple song audios in the negative example sample correspond to different melodies.
[0047] The positive example samples and the negative example samples each include the target song. For example, the positive example samples include the target song and a song with the same melody as the target song. For example, a song has a version sung by a singer A and a version sung by a singer B. The song audio of the two versions can be used as a set of positive example samples.
[0048] The song audio in the positive example samples and the song audio in the negative example samples can be the entire song content of a song or a song segment in a song.
[0049] Further, a melody feature vector corresponding to each song audio in each set of positive example samples and each set of negative example samples can be generated, to obtain a melody feature vector of the target song, the positive example melody feature vector and the negative example melody feature vector. The melody feature vector of each song audio in the positive example samples and the negative example samples is obtained by vectorizing the melody feature of the song audio. The melody feature can represent the melody of the song audio, for example, the time value, pitch and other melody features of a note. Therefore, the melody feature vector obtained by vectorizing the melody feature represents the melody of the song audio in the form of a vector, which is convenient for a computer to understand and recognize the melody of the song audio.
[0050] In addition, the melody feature of the song audio in the positive example samples can be disturbed to generate a disturbance vector corresponding to the song audio in the positive example samples, and the disturbance vector is used as the positive example melody feature vector. Disturbing the melody feature of the song audio in the positive example samples means adjusting the melody feature of the song audio in a fine-tuning manner, but without changing the melody key of the song audio. For example, the disturbance can be a subtle adjustment of the time value, number and pitch of a note in the song audio. By disturbing the melody feature of the song audio in the positive example samples, various complex and diverse song cover scenarios in real scenarios can be simulated, so that the new melody feature obtained by disturbance can match and restore more possible song cover scenarios in real scenarios, thereby expanding the data amount of the song cover samples.
[0051] In the embodiment, the new melody feature obtained by disturbing the melody feature of the song audio can be vectorized to obtain a disturbance vector corresponding to the song audio in the positive example samples, i.e., the disturbance vector represents the new melody feature in the form of a vector, which is convenient for a computer to understand and recognize. Specifically, the disturbance vector can represent the similarity between a cover song and an original song in a melody feature vector in a cover scenario other than the cover scenario corresponding to the positive example samples.
[0052] Therefore, when the initial cover song recognition model receives training data such as multiple sets of positive samples, multiple sets of negative samples, and their corresponding melody feature vectors and perturbation vectors, it can learn the similarity between multiple song audios in the positive samples in terms of melody feature vectors, and the differences between multiple song audios in the negative samples in terms of melody feature vectors, through the melody feature vectors. At the same time, it can also learn more potential similarities between cover songs and original songs in more cover song scenarios through perturbation vectors. That is, compared with learning from positive samples, the model can learn more potential similarities between cover songs and original songs, so that the model can still accurately identify more complex cover song recognition scenarios and ensure the robustness of the model recognition.
[0053] The initial cover song recognition model is trained until it meets the convergence condition, at which point training stops, resulting in the target cover song recognition model. This target model can be used to identify the original song corresponding to a cover song audio.
[0054] In this embodiment, based on the characteristics of cover songs, the melodic features of the song audio in the training samples are perturbed, thereby expanding the training sample size to obtain more training samples. This ensures that the model training can cover and match more complex cover song scenarios, getting rid of the problem of insufficient data in real cover song scenarios. A more robust cover song recognition model can be trained with a small amount of data.
[0055] based on Figure 2 In another optional embodiment of the illustrated embodiment and its above optional implementation, when generating the melody feature vector corresponding to each song audio in each group of positive samples and each song audio in each group of negative samples, the computer device can extract melody features for each song audio in each group of positive samples and each song audio in each group of negative samples, and encode the melody features of each song audio in each group of positive samples and the melody features of each song audio in each group of negative samples to obtain their respective melody feature vectors.
[0056] For example, the corresponding MIDI file can be extracted from each song audio in the positive and negative sample samples. A MIDI file, or Musical Instrument Digital Interface file, contains the duration and pitch of each note in the song. Therefore, by extracting the MIDI file of the song audio, the melodic features of the song audio, such as note duration and pitch, can be obtained.
[0057] The MIDI file of the song audio can be extracted in various ways. For example, a melody track determination method based on a neural network can be used to extract the MIDI file of the song audio by training a neural network model. Alternatively, a Melodia algorithm can be used to separate the fundamental frequency (F0) of the main melody from the song audio based on a functional plug-in corresponding to the algorithm, convert the Hertz to MIDI note numbers, eliminate tremolo noise through median filtering, segment the note sequence according to the pitch change, and output a timestamped discrete MIDI note sequence. The embodiment is not limited to the way of extracting the MIDI file of the song audio.
[0058] After obtaining the MIDI file of the song audio, the sequence of note duration and pitch can be determined from the MIDI file. For example, assume that the sequence of notes represented by the MIDI file of the song audio is [(200, 500, 61), (700, 500, 59), …, (2500, 500, 60)], each element in the sequence represents the content of a note, and the content of the note includes the start time, the note duration, and the pitch value in sequence. Therefore, the two melody features of note duration (time value) and pitch can be extracted from the note sequence, and the two melody features can be encoded into a vector to obtain the melody feature vector corresponding to the song audio.
[0059] Therefore, by extracting the melody features of the song audio, the melody of the song audio can be represented in the form of text, making the melody of the song audio tangible and facilitating the recognition and processing of the melody of the song audio. Encoding the melody features into a vector helps the computer to understand and recognize the melody of the song audio, and facilitates the comparison of melodies between different song audios.
[0060] In the encoding of the melody features of the song audio, for each song audio in each set of positive example samples and each song audio in each set of negative example samples, the various melody features of the song audio are encoded respectively to obtain the melody feature vector corresponding to each melody feature of the song audio. The melody feature vector of each melody feature of the song audio is used as the training data of the initial cover song recognition model. Alternatively, the melody feature vectors of the various melody features of the song audio are spliced to obtain a melody feature vector, and the melody feature vector is used as the training data of the initial cover song recognition model.
[0061] Taking the above example, after extracting the note duration and pitch of the song audio from the MIDI file of the song audio, the note duration and pitch of the song audio can be encoded respectively to obtain a feature vector corresponding to the note duration and a feature vector corresponding to the pitch of the note. The feature vector corresponding to the note duration can be used as training data for training the initial cover song recognition model, and the feature vector corresponding to the pitch of the note can be used as training data for training the initial cover song recognition model, that is, the two can be input into the model for training; alternatively, the feature vector corresponding to the note duration and the feature vector corresponding to the pitch of the note can be spliced first, and the spliced vector can be used as training data for training the initial cover song recognition model.
[0062] Using each melody feature vector corresponding to the melody feature of the song audio as training data for the model can enable the model to learn the similarities or differences of different song audios on each melody feature vector, and the learning degrees of the model for different melody feature vectors will not affect each other. For example, if the vector representation of a certain melody feature is inaccurate, it will affect the learning of the model for the vector of the melody feature, but it will not affect the learning of the model for other melody feature vectors, thereby ensuring the effectiveness and accuracy of the model training.
[0063] Using the vectors of multiple melody features as training data for the model after splicing can enable the model to simultaneously learn the similarities or differences of different song audios on the comprehensive vectors of multiple melody features, and the learned melody feature information is more comprehensive, which is helpful for the model to handle more complex and diversified cover song recognition scenarios and improve the robustness of the model recognition.
[0064] In the present embodiment, there are multiple ways to perturb the melody features of the song audio. In some optional embodiments, a preset time length can be added to or subtracted from the time value of part or all of the notes of the song audio in the positive example sample, and / or a glissando note can be inserted between two adjacent notes of the song audio in the positive example sample according to the time value and pitch of the two adjacent notes. The song audio to be perturbed can be any one or more of the song audios in the positive example sample, which is not limited here.
[0065] The glissando is a smooth sliding effect of the pitch in music, which is the transition and progression between notes, and smoothly slides from one pitch to another, rather than jumps in steps.
[0066] For example, the note duration or the pitch of the notes extracted from the MIDI file can be randomly disturbed. When the note duration is disturbed, for each note in the song audio in the positive example sample, a random value within 20% of the duration of the note can be added or subtracted to increase or decrease the duration of the note by a random value within 20%. Of course, other values than 20% can also be used.
[0067] The random disturbance of the pitch can be to add a new note to achieve a gradual and progressive effect between two notes. For example, between any two adjacent notes in the song audio, a glissando note can be inserted with a certain probability, for example, which can be set to 10%. The pitch value of the glissando note is the average of the pitch values of the two adjacent notes, and the duration value of the glissando note is the average of the duration values of the two adjacent notes. Of course, the aforementioned probability can be other values than 10%, and the pitch value and duration value of the glissando note can also be calculated in other ways based on the pitch values and duration values of the two adjacent notes, which are not limited here.
[0068] Therefore, by repeatedly disturbing the melody features to achieve data augmentation, a large number of training samples can be obtained without relying on a large number of real samples for training, greatly reducing the difficulty of training data collection, and by disturbing the melody features, the training samples of the model can cover and restore more complex cover song recognition scenarios that can exist in real scenarios, which helps to improve the robustness of the model in identifying such scenarios.
[0069] The melody features of the song audio in the positive example sample can be disturbed to obtain the disturbed melody features corresponding to the song audio in the positive example sample, and then the disturbed melody features can be encoded to obtain the disturbed vector corresponding to the song audio in the positive example sample. The disturbed vector can be used as training data for model training.
[0070] Based on Figure 2 In an optional embodiment of the embodiment shown in FIG. 8, the obtained multiple sets of positive example samples and multiple sets of negative example samples can be represented in the form of triplets. Specifically, the computer device can obtain multiple triplet data, where each triplet data includes an anchor song audio (anchor), a positive sample song audio (positive), and a negative sample song audio (negative). The anchor song audio and the positive sample song audio correspond to the same song, and the negative sample song audio and the anchor song audio correspond to different songs.
[0071] The anchor song audio and the positive sample song audio constitute a positive example sample, and the negative sample song audio and the anchor song audio constitute a negative example sample. The anchor song audio can be a song audio of an original song, the positive sample song audio can be a song audio corresponding to a cover version of the original song, and the negative sample song audio can be a song audio of another song different from the original song. Of course, the song audio of the original song can also be used as the anchor song audio, and the song audio of the cover version of the original song can be used as the positive sample song audio.
[0072] For example, the training data used for model training can be composed of a large amount of library data and a small amount of real scene data. For the library data, a large number of original songs and corresponding cover songs can be selected from the library data to construct triplets. The positive examples in the triplets are the original songs and the cover songs, and the negative examples are the original songs and other songs, which can be randomly selected from the library. For real scene data, the positive examples in the triplets need to be manually labeled, and the negative examples are randomly extracted from other real scene data. During the training process of the model, the two batches of data can be mixed to randomly sample the training data of the model.
[0073] During model training, the initial cover identification model can calculate a loss function according to the aforementioned training data, and adjust the parameters of the initial cover identification model according to the calculation result of the loss function, until the convergence condition is met to stop training, and obtain the target cover identification model. The loss function can be a variety of loss functions, for example, when the aforementioned triplet data is used as the model training data, the corresponding loss function can be a triplet loss function. The triplet loss function is used to represent the distance between the anchor song audio and the positive sample song audio in the triplet data, and the distance between the anchor song audio and the negative sample song audio.
[0074] For example, the aforementioned training data can be input into a Resnet18 model, and the model can calculate a triplet loss at the output end of the network according to the training data, which is expressed as follows:
[0075] ;
[0076] Wherein, d(a, p) is the distance (such as cosine distance) between the original song audio anchor and the cover song audio positive, d(a, n) is the distance (such as cosine distance) between the anchor and other song audios negative; margin is a degree coefficient that can be adjusted, used to set the minimum interval between the distance of the anchor song audio and the positive sample song audio and the distance of the anchor song audio and the negative sample song audio. The goal of adjusting the model parameters according to the triplet loss is to make the distance between the original song audio and the cover song audio in the vectorization embedding space as close as possible, and the distance between the original song audio and other song audios in the vectorization embedding space as far as possible.
[0077] Of course, in addition to the above-mentioned triplet loss, the model parameters can also be adjusted based on loss functions such as classification loss or center loss, which are not limited by the present embodiment.
[0078] The following further describes an exemplary application scenario of the embodiment based on the foregoing embodiment, as shown in Figure 3 As shown, multiple song samples such as original songs, cover songs of the original songs, and other songs can be collected, and song segments are extracted from each song sample to obtain anchor song segments anchor and cover song segments positive and other song segments negative. MIDI feature extraction is performed on each segment respectively, and the results of MIDI extraction are shown in the figure, where the numerical values represent pitch values, and the value range is 0-128; the horizontal line represents the time value of the note, including the start time and end time of the note. The pitch and time value sequence is a direct embodiment of the song melody. For a pair of positive example samples, there is melody similarity between the original song and the cover song, so there is a strong correlation between the pitch and time value sequence. Since different people sing the same song, the pronunciation duration may be different, so the time value is long or short, and for the transition between two notes, some singers use the glissando way to transition, as shown in the figure, the last note of the positive segment of the positive example sample relative to the anchor segment, the pitch transitions from 58 to 59, and then to 60.
[0079] After the MIDI melody features are extracted, the melody features of the audio segments can be randomly disturbed, and the melody features of the audio segments and the disturbed melody features are encoded. The vector obtained by encoding is input into the ResNet18 network, and then the model is trained according to the training data and the triplet loss is calculated until the model training is completed when the convergence condition is met.
[0080] The following will be based on the foregoing Figure 1 The network framework shown in the figure and Figure 2 The embodiment of the present application will be described in detail. Please refer to Figure 4One embodiment of the cover song recognition method in this application includes:
[0081] 401. Obtain the target cover song recognition model;
[0082] The method of this embodiment can be applied to a computer device, which may be... Figure 1 The network framework shown includes a service server 100 or various terminal devices. This computer device can be connected to... Figure 2 The computer device described in the illustrated embodiments can be the same device or different devices. In some embodiments, the computer device can implement the cover song recognition method provided in this embodiment by running a computer program. For example, the computer program can be a native program or software module in the operating system; it can also be a native application (APP), that is, a program that needs to be installed in the operating system to run; it can also be a small program, that is, a program that only needs to be downloaded to the browser environment to run; or it can be a small program that can be embedded in any APP. In short, the above-mentioned computer program can be any form of application, module or plugin.
[0083] The target cover song recognition model deployed on computer equipment can be based on Figure 1 The cover song recognition model training method of the embodiment shown is used to train the model.
[0084] 402. Obtain the vectorized representation features corresponding to the melodic features of multiple original song audios; wherein, the multiple original song audios correspond to different songs;
[0085] In this embodiment, the vectorized representation feature (embedding feature) corresponding to the melody feature of each original song audio can be generated by: acquiring the multiple original song audios, extracting melody features from each original song audio, encoding the melody features of the original song audio to obtain the melody feature vector corresponding to each original song audio; and inputting the melody feature vector of each original song audio into the target cover song recognition model to obtain the vectorized representation feature of each original song audio output by the target cover song recognition model.
[0086] The explanation of melodic features and melodic feature vectors can be found in the detailed records already mentioned above, and will not be repeated here.
[0087] The process of generating the vectorized representation features corresponding to the melody features of each original song audio can be performed by the computer device of this embodiment; it can also be performed by other devices and then sent to the computer device of this embodiment, which is not limited here.
[0088] 403. For the target song audio to be identified, extract the melody features of the target song audio and encode the melody features of the target song audio to obtain the melody feature vector corresponding to the target song audio.
[0089] 404. Input the melody feature vector of the target song audio into the target cover song recognition model to obtain the vectorized representation features of the target song audio output by the target cover song recognition model;
[0090] In this embodiment, the embedding features of each original song audio are extracted using a target cover song recognition model, and the embedding features of the target song audio to be recognized are extracted. The melody of the song is represented in a vectorized form, which facilitates the computer to identify and compare the melody differences between the original song and the song to be recognized.
[0091] 405. Calculate the distance between the vectorized representation features of the target song audio and the vectorized representation features of each original song audio;
[0092] 406. The original song to which the original song audio corresponding to the smallest distance belongs is determined as the original song corresponding to the target song audio;
[0093] In this embodiment, the original song audio and the target song audio to be identified can be the entire song content or a segment of the song; there is no limitation here.
[0094] Computer devices can use a target cover song recognition model to build an embedding feature library for multiple original songs. When the embedding features of the song to be recognized are obtained based on the target cover song recognition model, the distance (such as cosine distance) between the embedding features of the song to be recognized and the embedding features of each original song segment in the embedding feature library can be calculated. The song corresponding to the segment with the closest distance is selected as the final cover song recognition result.
[0095] Therefore, in this embodiment, the more robust target cover song recognition model trained by the aforementioned embodiments is used for cover song recognition. It can accurately identify various complex and diverse cover song scenarios in real-world situations. The cover song recognition method of this embodiment can be applied to music platforms to better promote the song recognition function, expand the application scope of song recognition, and improve the user experience.
[0096] The computer device in the embodiments of this application is described below. Please refer to [link / reference]. Figure 5 One embodiment of the computer device in this application includes:
[0097] The computer device 500 can include one or more central processing units (CPU) 501 and a memory 505 in which one or more applications or data are stored.
[0098] The memory 505 can be volatile storage or persistent storage. The programs stored in the memory 505 can include one or more modules, each of which can include a series of instruction operations in the computer device. Further, the central processing unit 501 can be configured to communicate with the memory 505 to execute the series of instruction operations in the memory 505 on the computer device 500.
[0099] The computer device 500 can further include one or more power supplies 502, one or more wired or wireless network interfaces 503, one or more input / output interfaces 504, and / or one or more operating systems, such as Windows Server™, Mac OS X™, Unix™, Linux™, FreeBSD™, etc.
[0100] The central processing unit 501 can execute the operations of the computer device in the embodiments described above, and details are not repeated here. Figure 2 and Figure 4 The computer device in the embodiments described above can execute the operations of the computer device in the embodiments described above, and details are not repeated here.
[0101] The embodiments of the present application also provide a computer storage medium, one of which includes: the computer storage medium stores instructions, and the instructions, when executed on a computer, cause the computer to execute the operations of the computer device in the embodiments described above. Figure 2 and Figure 4 The computer device in the embodiments described above can execute the operations of the computer device in the embodiments described above.
[0102] The embodiments of the present application also provide a computer program product, one of which includes: the computer program product, when running on a computer device, causes the computer device to execute the operations of the computer device in the embodiments described above. Figure 2 and Figure 4 The computer device in the embodiments described above can execute the operations of the computer device in the embodiments described above.
[0103] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, device and unit described above can refer to the corresponding processes in the foregoing method embodiments, which are not repeated here.
[0104] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.
[0105] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.
[0106] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically as a separate unit, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.
[0107] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially, or the part that contributes to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: U disk, mobile hard disk, read-only memory (ROM, read-only memory), random access memory (RAM, random access memory), magnetic disk or optical disk, and various other media that can store program codes.
Claims
1. A method for training a cover song recognition model, characterized in that, The method includes: Obtain the initial cover song recognition model; Obtain the melody feature vector of the target song, the positive example melody feature vector of the song with the same melody as the target song, and the negative example melody feature vector of the song with a different melody than the target song; The initial cover song recognition model is used to perform feature processing on the melody feature vector, the positive example melody feature vector, and the negative example melody feature vector to obtain melody features, positive example melody features, and negative example melody features; The initial cover song model is optimized based on the melody features, the positive example melody features, and the negative example melody features until the convergence condition is met, thereby obtaining the target cover song recognition model; the target cover song recognition model is used to identify the original song corresponding to the cover song audio.
2. The method according to claim 1, characterized in that, The process of obtaining the melody feature vector of the target song, the positive example melody feature vectors of songs with the same melody as the target song, and the negative example melody feature vectors of songs with different melody than the target song includes: Multiple sets of positive sample samples and multiple sets of negative sample samples are obtained. Each set of positive sample samples includes multiple song audios corresponding to the same melody, and each set of negative sample samples includes multiple song audios corresponding to different melody. The multiple song audios of the positive sample samples and the multiple song audios of the negative sample samples all include the target song. Generate melody feature vectors for each song audio in each group of positive samples and for each song audio in each group of negative samples, to obtain the melody feature vector of the target song, the positive melody feature vectors, and the negative melody feature vectors; and / or, The melody features of the song audio in the positive sample are perturbed to generate a perturbation vector corresponding to the song audio in the positive sample, and the perturbation vector is used as the positive melody feature vector.
3. The method according to claim 2, characterized in that, The generation of the melody feature vector corresponding to each song audio in each group of positive samples and each song audio in each group of negative samples includes: Melody features are extracted for each song audio in each group of positive samples and for each song audio in each group of negative samples. The melody features of each song audio in each group of positive samples and the melody features of each song audio in each group of negative samples are encoded to obtain their respective melody feature vectors.
4. The method according to claim 3, characterized in that, The process of encoding the melody features of each song audio in each group of positive samples and the melody features of each song audio in each group of negative samples to obtain their respective corresponding melody feature vectors includes: For each song audio in each group of positive samples and each song audio in each group of negative samples, the various melodic features of the song audio are encoded respectively to obtain a melodic feature vector for each melodic feature of the song audio; and, The melody feature vector of each melody feature of the song audio is used as the training data; or, the melody feature vectors of each of the multiple melody features of the song audio are concatenated, and the concatenated melody feature vector is used as the training data.
5. The method according to any one of claims 2 to 4, characterized in that, The melodic features include the pitch and duration of the notes; The step of perturbing the melodic features of the song audio in the positive example samples to generate the perturbation vector corresponding to the song audio in the positive example samples includes: The melodic features of the song audio in the positive example samples are perturbed to obtain the perturbed melodic features corresponding to the song audio in the positive example samples; wherein, the perturbation includes: Increase or decrease the duration of some or all notes in the song audio in the positive sample by a preset duration, and / or, insert a glissando note between two adjacent notes according to the duration and pitch of two adjacent notes in the song audio in the positive sample; The perturbation melody features are encoded to obtain the perturbation vector corresponding to the song audio in the positive sample.
6. The method according to any one of claims 2 to 4, characterized in that, The acquisition of multiple sets of positive samples and multiple sets of negative samples includes: Multiple triplet data are obtained, wherein each triplet data includes anchor song audio, positive sample song audio, and negative sample song audio, wherein the anchor song audio and the positive sample song audio correspond to the same song, and the negative sample song audio and the anchor song audio correspond to different songs. The anchor song audio and the positive sample song audio constitute the positive sample; the negative sample song audio and the anchor song audio constitute the negative sample.
7. The method according to claim 6, characterized in that, The initial cover song recognition model is trained based on the training data, including: The initial cover song recognition model calculates a loss function based on the training data, and adjusts the parameters of the initial cover song recognition model based on the calculation result of the loss function until the convergence condition is met, at which point training stops, and the target cover song recognition model is obtained. The loss function includes a ternary loss function; the ternary loss function is used to characterize the distance between the anchor song audio and the positive sample song audio in the triplet data, and the distance between the anchor song audio and the negative sample song audio.
8. A method for identifying cover songs, characterized in that, The method includes: A target cover song recognition model is obtained, wherein the target cover song recognition model is trained by the cover song recognition model training method according to any one of claims 1 to 7; Obtain the vectorized representation features corresponding to the melodic features of multiple original song audios; wherein, the multiple original song audios correspond to different songs; For the target song audio to be identified, the melody features of the target song audio are extracted and encoded to obtain the melody feature vector corresponding to the target song audio. The melody feature vector of the target song audio is input into the target cover song recognition model to obtain the vectorized representation features of the target song audio output by the target cover song recognition model; Calculate the distance between the vectorized representation features of the target song audio and the vectorized representation features of each original song audio; The original song whose audio corresponds to the smallest distance is identified as the original song corresponding to the target song audio.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the method as described in any one of claims 1 to 8.
10. A computer storage medium, characterized in that, The computer storage medium stores instructions that, when executed on the computer, cause the computer to perform the method as described in any one of claims 1 to 8.