Chord Recognition Model Training Method, System and Product Based on Semi-Supervised Learning

The semi-supervised chord recognition method improves chord recognition accuracy by preprocessing labeled samples, applying data augmentation, and using adaptive confidence thresholds and contrastive learning to optimize model parameters, effectively addressing the challenge of rare chord identification.

CN120126508BActive Publication Date: 2025-07-15XI AN JIAOTONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510620403.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-07-15
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

In the existing semi-supervised automatic chord recognition algorithm, the recognition accuracy of rare chords is lower than that of common chords, and the overall chord recognition accuracy is reduced, mainly due to the small number of labeled audio samples and the insufficient utilization of labelless audio samples.

Method used

Through a semi-supervised learning method, the supervised model is trained using pre-processed labeled audio samples, the adaptive confidence threshold is calculated, multi-dimensional data augmentation is performed, the pseudo-label sample set and low confidence sample set are divided, and the model parameters are optimized through comparative learning methods, and the model is updated in combination with the backpropagation algorithm.

Benefits of technology

It improves the recognition ability of rare chords, improves the overall chord recognition accuracy, and maintains a high recognition rate of common chords, achieving more efficient label-free audio data utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126508B_ABST
    Figure CN120126508B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, system and product for training a chord recognition model based on semi-supervised learning, belonging to the technical field of chord recognition. The method for training a chord recognition model based on semi-supervised learning provided by the present invention improves the utilization effect of the unlabeled audio sample set through adaptive confidence threshold screening, multi-dimensional data augmentation and hierarchical utilization of pseudo-labels; by introducing a contrastive learning method, deeply mines the low-confidence sample set, captures the subtle differences and internal connections between samples, combines the first loss and the second loss, and optimizes the parameters of the supervised model through the backpropagation algorithm, so that the supervised model can maintain a high recognition rate for common chords while improving the recognition ability for rare chords, thereby achieving the improvement of the overall chord recognition accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of chord recognition, and in particular to a method, system and product for training a chord recognition model based on semi-supervised learning. Background Art

[0002] Currently, according to the different chord label information of the input sample set during the training process, the research directions in the field of automatic chord recognition mainly focus on supervised automatic chord recognition algorithms and semi-supervised automatic chord recognition algorithms. Among them, the supervised automatic chord recognition algorithm focuses on feature optimization and network structure optimization, while the semi-supervised automatic chord recognition algorithm pays more attention to the effective utilization of unlabeled audio data. With the booming development of the Internet music market, unlabeled audio data has become relatively easy to obtain, providing new opportunities for the research of automatic chord recognition. Under this background, semi-supervised learning, as a learning strategy that combines labeled data and unlabeled audio data, has attracted much attention.

[0003] In the existing semi-supervised automatic chord recognition algorithms, the recognition accuracy of rare chords is significantly lower than that of common chords, and there is a serious problem of classification imbalance. This is mainly because the available labeled audio sample sets publicly disclosed in the current chord recognition field are scarce, it is difficult to collect label data and requires extremely strong professional knowledge, resulting in too few labeled audio samples available for training. To address this problem, Marcelo Bortolozzo et al. began to consider expanding the dataset by introducing unlabeled audio sample sets. Using a self-learning method, a teacher model generates pseudo-labels for a large number of unlabeled audio samples, and then a classification-balanced sample set is selected for model training. Although this method does improve the recognition accuracy of rare chords, it inevitably loses the recognition accuracy of common chords, resulting in a decrease in the overall chord recognition accuracy.

[0004] Therefore, how to make full use of unlabeled audio sample data and improve the overall chord recognition accuracy has become a technical problem that needs to be urgently solved by those skilled in the art in the current field. Summary of the Invention

[0005] The purpose of the present invention is to provide a method, system and product for training a chord recognition model based on semi-supervised learning to overcome the problem of the decrease in the overall chord recognition accuracy due to the large use of unlabeled audio samples in the prior art.

[0006] The present invention solves the above technical problems through the following technical solutions:

[0007] A method for training a chord recognition model based on semi-supervised learning includes the following steps:

[0008] Taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A, combining with the true chord category probability distribution of the labeled audio sample set a, screening out the correctly predicted labeled audio sample subset X, and calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X;

[0009] Performing first weak augmentation, second weak augmentation, and strong augmentation on the unlabeled audio sample set w respectively to obtain an unlabeled audio sample set b, an unlabeled audio sample set c, and an unlabeled audio sample set d. Inputting the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain a predicted chord category probability distribution B, a predicted chord category probability distribution C, and a predicted chord category probability distribution D respectively;

[0010] Based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, calculating the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w, and combining with the adaptive confidence threshold for each chord category in the labeled audio sample subset X to divide and obtain a pseudo-labeled sample set I and a low-confidence sample set II;

[0011] Based on the pseudo-labeled sample set I and the predicted chord category probability distribution D, calculating a first loss; based on the contrast learning method, dividing the low-confidence sample set II into positive sample pairs and negative sample pairs, and using a contrast loss function to calculate a second loss; based on the first loss and the second loss, combining with the backpropagation algorithm to update the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.

[0012] A further improvement of the present invention lies in that: taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A specifically includes:

[0013] Performing third weak augmentation on the labeled audio sample set x = {x1, x2, x3,..., x s} to obtain the preprocessed labeled audio sample set a, where the third weak augmentation is to adjust the audio rate, and x1, x2, x3,..., x s are the 1st, 2nd, 3rd,..., s-th samples in the labeled audio sample set x respectively;

[0014] Selecting the cross-entropy loss function Loss x as the loss function of the supervised model, taking the preprocessed labeled audio sample set a as input, training the supervised model to obtain a trained supervised model and a predicted chord category probability distribution A;

[0015] The cross-entropy loss function Lossx Specifically:

[0016]

[0017] Among them, log is the logarithm; s is the total number of samples in the labeled audio sample set a after preprocessing; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, , otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j.

[0018] A further improvement of the present invention lies in: calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X, specifically:

[0019]

[0020] Among them, i is the sample index; j is the chord category index; is the adaptive confidence threshold corresponding to the j-th chord; is the total number of samples of the j-th chord in the labeled audio sample subset X with correct prediction; is the true chord category probability distribution of the i-th sample of the j-th chord in the labeled audio sample subset X with correct prediction; is the predicted chord category probability distribution of the i-th sample of the j-th chord in the labeled audio sample subset X with correct prediction; means element-wise multiplication.

[0021] A further improvement of the present invention lies in: the first weak augmentation is to adjust the audio rate; the second weak augmentation is to add noise; the strong augmentation is time-frequency masking.

[0022] A further improvement of the present invention lies in: assuming the unlabeled audio sample set w = {w1, w2, w3, …, w t}, calculating the mean of the predicted chord category probability distributions of each sample in the unlabeled audio sample set w, specifically:

[0023]

[0024] Among them, w1, w2, w3, …, w t are the 1st, 2nd, 3rd, …, t-th samples in the unlabeled audio sample set w respectively; is the mean of the predicted chord category probability distribution of the t-th sample in the unlabeled audio sample set w; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set b; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set c; are the network parameters of the supervised model.

[0025] A further improvement of the present invention lies in: dividing the pseudo-label sample set I and the low-confidence sample set II by using the adaptive confidence threshold of each chord class in the labeled audio sample subset X, specifically including:

[0026] Successively judge whether the mean of the predicted chord class probability distribution of each sample in the unlabeled audio sample set w is ≥ the adaptive confidence threshold of the corresponding chord class in the labeled audio sample subset X. If the judgment result is yes, divide the corresponding sample into the pseudo-label sample set I; if the judgment result is no, divide the corresponding sample into the low-confidence sample set II.

[0027] A further improvement of the present invention lies in: dividing the low-confidence sample set II into positive sample pairs and negative sample pairs based on the contrast learning method, and calculating the second loss by using the contrast loss function, specifically including:

[0028] Successively judge whether each sample in the low-confidence sample set II is a sample that has undergone the first weak augmentation and the second weak augmentation. If the judgment is yes, divide the corresponding sample into the positive sample pair in the contrast learning method; if the judgment is no, divide the corresponding sample into the negative sample pair in the contrast learning method;

[0029] Based on the similarity difference between the positive sample pair and the negative sample pair, use the contrast loss function to calculate the second loss, specifically:

[0030]

[0031] where, log is the logarithm; is the second loss; , is the similarity between sample g in the unlabeled audio sample set b and sample h in the unlabeled audio sample set c, is the similarity between sample g in the unlabeled audio sample set b and sample k in the unlabeled audio sample set d, T is the matrix transpose, ||B|| is the norm of B, ||C|| is the norm of C, ||D|| is the norm of D, is the adjustable temperature parameter; is the total number of samples in the low-confidence sample set II, , all belong to the low-confidence sample set II. g is used to traverse the samples in the low-confidence sample set II and calculate the similarity-related loss between each sample and other samples. h and k are respectively the sample indices used to represent the samples that form positive sample pairs and negative sample pairs with the g sample.

[0032] The present invention also provides a chord recognition model training system based on semi-supervised learning, including:

[0033] The first module is used to take the preprocessed labeled audio sample set a as input, train a supervised model, obtain the trained supervised model and the predicted chord category probability distribution A, combine with the true chord category probability distribution of the labeled audio sample set a, screen out the subset X of correctly predicted labeled audio samples, and calculate the adaptive confidence threshold for each chord category in the subset X of labeled audio samples;

[0034] The second module is used to perform the first weak augmentation, the second weak augmentation and the strong augmentation on the unlabeled audio sample set w respectively, obtain the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d respectively, and input the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d into the trained supervised model to obtain the predicted chord category probability distribution B, the predicted chord category probability distribution C and the predicted chord category probability distribution D respectively;

[0035] The third module is used to calculate the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, and combine with the adaptive confidence threshold of each chord category in the subset X of labeled audio samples to divide and obtain the pseudo-label sample set I and the low-confidence sample set II;

[0036] The fourth module is used to calculate the first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; divide the low-confidence sample set II into positive sample pairs and negative sample pairs based on the contrast learning method, and calculate the second loss using the contrast loss function; based on the first loss and the second loss, combined with the backpropagation algorithm, update the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.

[0037] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned chord recognition model training method based on semi-supervised learning are implemented.

[0038] The present invention also provides a computer device, including a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned chord recognition model training method based on semi-supervised learning are implemented.

[0039] Compared with the prior art, the positive and progressive effects of the present invention are as follows:

[0040] The chord recognition model training method based on semi-supervised learning provided by the present invention uses a pre-processed labeled audio sample set a to train a supervised model, calculates an adaptive confidence threshold for each chord category based on the correctly predicted subset X of the labeled audio samples; performs three types of enhancement processing on the unlabeled audio sample set w respectively, and inputs the enhanced samples into the trained supervised model to obtain different predicted probability distributions; divides the pseudo-label sample set I and the low-confidence sample set II according to the mean value of the predicted chord category probability distribution of the unlabeled audio sample set w and the adaptive confidence threshold; calculates the first loss based on the pseudo-label sample set I, calculates the second loss for the low-confidence sample set II using the contrastive learning method, combines the two, and updates the model parameters through the backpropagation algorithm to obtain a chord recognition model based on semi-supervised learning. This method improves the utilization effect of the unlabeled audio sample set through adaptive confidence threshold screening, multi-dimensional data enhancement, and hierarchical utilization of pseudo-labels; by introducing the contrastive learning method, deeply mines the low-confidence sample set, captures the subtle differences and internal connections between samples, combines the first loss and the second loss, and optimizes the parameters of the supervised model through the backpropagation algorithm, enabling the supervised model to maintain a high recognition rate for common chords while enhancing the recognition ability for rare chords, thereby achieving an improvement in the overall chord recognition accuracy. Description of the Drawings

[0041] The accompanying drawings in the specification are used to provide a further understanding of the present invention, constitute a part of the present invention, and the schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation to the present invention.

[0042] Figure 1 It is a schematic flow chart of a chord recognition model training method based on semi-supervised learning of the present invention;

[0043] Figure 2 It is a flow block diagram of a chord recognition model training method based on semi-supervised learning of the present invention;

[0044] Figure 3 It is a graph of the comparative experiment results of Embodiment 1 of the present invention. Detailed Embodiments

[0045] Glossary:

[0046] Cross-entropy loss function: A loss function widely used in classification tasks, used to measure the difference between the probability distribution predicted by the model and the true probability distribution, which can intuitively reflect the inconsistency between the model prediction result and the true chord category, thereby effectively guiding the update of the model network parameters and improving the prediction accuracy of the model.

[0047] The following further elaborates on the present invention in conjunction with the accompanying drawings and specific embodiments, which is an explanation rather than a limitation of the present invention.

[0048] Refer to Figure 1 and Figure 2 , a method for training a chord recognition model based on semi - supervised learning, comprising the following steps:

[0049] Taking the pre - processed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and the predicted chord category probability distribution A, combining with the true chord category probability distribution of the labeled audio sample set a, screening out the correctly predicted subset X of labeled audio samples, and calculating the adaptive confidence threshold for each chord category in the subset X of labeled audio samples;

[0050] Performing first weak augmentation, second weak augmentation, and strong augmentation on the unlabeled audio sample set w respectively to obtain the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d, and inputting the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain the predicted chord category probability distribution B, the predicted chord category probability distribution C, and the predicted chord category probability distribution D respectively;

[0051] Based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, calculating the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w, and combining with the adaptive confidence threshold for each chord category in the subset X of labeled audio samples to divide and obtain the pseudo - label sample set I and the low - confidence sample set II;

[0052] Calculating the first loss based on the pseudo - label sample set I and the predicted chord category probability distribution D; based on the contrast learning method, dividing the low - confidence sample set II into positive sample pairs and negative sample pairs, and calculating the second loss using the contrast loss function; based on the first loss and the second loss, and combining with the backpropagation algorithm, updating the network parameters of the trained supervised model to obtain a chord recognition model based on semi - supervised learning.

[0053] Based on the first loss and the second loss, combined with the backpropagation algorithm, the network parameters of the trained supervised model are updated to obtain a chord recognition model based on semi-supervised learning. Essentially, by fusing unlabeled audio data, the initial supervised model is improved to enable it to have the ability of semi-supervised learning. Specifically: The first loss and the second loss are weighted and summed according to a certain ratio to obtain a total loss function; by calculating the gradient information of the total loss with respect to the network parameters of the supervised model, the gradient information is propagated from the output layer to the input layer of the trained supervised model by using the backpropagation algorithm, thereby updating the network parameters of the supervised model; repeating the above process, iteratively optimizing the network parameters of the supervised model, so that the supervised model can continuously improve the accuracy and robustness of chord recognition while using labeled audio data and unlabeled audio data.

[0054] This method improves the utilization effect of the unlabeled audio sample set through adaptive confidence threshold screening, multi-dimensional data augmentation, and hierarchical utilization of pseudo-labels; by introducing a contrastive learning method, deeply mines the low-confidence sample set, captures the subtle differences and internal connections between samples, combines the first loss and the second loss, and optimizes the parameters of the supervised model through the backpropagation algorithm, enabling the supervised model to maintain a high recognition rate for common chords while enhancing the recognition ability for rare chords, thereby achieving an improvement in the overall chord recognition accuracy.

[0055] Specifically, taking the preprocessed labeled audio sample set a as the input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A, specifically including:

[0056] Performing a third weak augmentation on the labeled audio sample set x = {x1, x2, x3, …, x s} to obtain the preprocessed labeled audio sample set a, where the third weak augmentation is to adjust the audio rate, and x1, x2, x3, …, x s are the 1st, 2nd, 3rd, …, s-th samples in the labeled audio sample set x respectively;

[0057] Selecting the cross-entropy loss function Loss x as the loss function of the supervised model, taking the preprocessed labeled audio sample set a as the input, training the supervised model to obtain a trained supervised model and a predicted chord category probability distribution A;

[0058] The cross-entropy loss function Loss x is specifically:

[0059]

[0060] Among them, log is the logarithm; s is the total number of samples in the labeled audio sample set a after preprocessing; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, , otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j.

[0061] Specifically, calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X is specifically as follows:

[0062]

[0063] Among them, i is the sample index; j is the chord category index; is the adaptive confidence threshold corresponding to the j-th chord; is the total number of samples of the j-th chord in the labeled audio sample subset X with correct prediction; is the true chord category probability distribution of the i-th sample of the j-th chord in the labeled audio sample subset X with correct prediction; is the predicted chord category probability distribution of the i-th sample of the j-th chord in the labeled audio sample subset X with correct prediction; is element-wise multiplication.

[0064] Specifically, the first weak augmentation is to adjust the audio rate; the second weak augmentation is to add noise; the strong augmentation is time-frequency masking.

[0065] Specifically, let the unlabeled audio sample set w = {w1, w2, w3, …, w t}, calculating the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w is specifically as follows:

[0066]

[0067] Among them, w1, w2, w3, …, w t are the 1st, 2nd, 3rd, …, t-th samples in the unlabeled audio sample set w respectively; is the mean of the predicted chord category probability distribution of the t-th sample in the unlabeled audio sample set w; is the predicted chord category probability distribution corresponding to the t-th sample in the unlabeled audio sample set b; is the predicted chord category probability distribution corresponding to the t-th sample in the unlabeled audio sample set c; are the network parameters of the supervised model.

[0068] Specifically, the adaptive confidence threshold of each chord category in the combined labeled audio sample subset X is used to divide the pseudo-label sample set I and the low-confidence sample set II, which specifically includes:

[0069] Successively determine whether the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w is ≥ the adaptive confidence threshold of the corresponding chord category in the labeled audio sample subset X. If the judgment result is yes, divide the corresponding sample into the pseudo-label sample set I; if the judgment result is no, divide the corresponding sample into the low-confidence sample set II.

[0070] In a specific embodiment of the present invention, the first loss is calculated using the cross-entropy loss function, and the second loss is calculated using the contrastive loss function.

[0071] Specifically, based on the contrastive learning method, the low-confidence sample set II is divided into positive sample pairs and negative sample pairs, and the contrastive loss function is used to calculate the second loss, which specifically includes:

[0072] Successively determine whether each sample in the low-confidence sample set II is a sample that has undergone the first weak augmentation and the second weak augmentation. If the judgment is yes, divide the corresponding sample into the positive sample pair in the contrastive learning method; if the judgment is no, divide the corresponding sample into the negative sample pair in the contrastive learning method;

[0073] Based on the similarity difference between the positive sample pair and the negative sample pair, the contrastive loss function is used to calculate the second loss, specifically:

[0074]

[0075] where log is the logarithm; is the second loss; , is the similarity between sample g in the unlabeled audio sample set b and sample h in the unlabeled audio sample set c, is the similarity between sample g in the unlabeled audio sample set b and sample k in the unlabeled audio sample set d, T is the matrix transpose, ||B|| is the norm of B, ||C|| is the norm of C, ||D|| is the norm of D, is an adjustable temperature parameter; is the total number of samples in the low-confidence sample set II, , both belong to the low-confidence sample set II, g is used to traverse the samples in the low-confidence sample set II to calculate the similarity-related loss of each sample with other samples; h and k are respectively used to represent the sample indices that form positive and negative sample pairs with the g sample.

[0076] Based on the same inventive concept, the present invention provides a chord recognition model training system based on semi-supervised learning, including:

[0077] The first module is used to take the preprocessed labeled audio sample set a as input, train a supervised model, obtain the trained supervised model and the predicted chord category probability distribution A, combine the true chord category probability distribution of the labeled audio sample set a, screen out the correctly predicted labeled audio sample subset X, and calculate the adaptive confidence threshold for each chord category in the labeled audio sample subset X;

[0078] The second module is used to perform the first weak augmentation, the second weak augmentation, and strong augmentation on the unlabeled audio sample set w respectively, obtain the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d respectively, and input the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain the predicted chord category probability distribution B, the predicted chord category probability distribution C, and the predicted chord category probability distribution D respectively;

[0079] The third module is used to calculate the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, combine the adaptive confidence threshold for each chord category in the labeled audio sample subset X, and divide to obtain the pseudo-label sample set I and the low-confidence sample set II;

[0080] The fourth module is used to calculate the first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; based on the contrast learning method, divide the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculate the second loss using the contrast loss function; based on the first loss and the second loss, combine the backpropagation algorithm to update the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.

[0081] In order to verify the advancedness of the chord recognition model training method based on semi-supervised learning of the present invention, the chord recognition model based on semi-supervised learning obtained by the method of the present invention is tested, and the method of the present invention is compared with the semi-supervised algorithm in the current chord recognition field. The comparative experimental results are shown in Table 1. Among them, method one is to train a supervised model only through a labeled audio sample set, which represents the baseline performance when no semi-supervised strategy is adopted; method two is a semi-supervised chord recognition algorithm based on variational autoencoder proposed in the literature [Semi-supervised neural chord estimation based on a variational autoencoder with latent chord labels and features]; method three is a semi-supervised chord recognition algorithm based on pseudo-label reconstruction of a more balanced data set proposed in the literature [Improving the classification of rare chords with unlabeled data]; method four is a semi-supervised chord recognition algorithm based on pseudo-label reconstruction of a more balanced data set proposed in the literature [Deep Semi-Supervised Learning With Contrastive Learning in Large Vocabulary Automatic Chord The semi-supervised chord recognition algorithm based on contrastive learning and noise student framework proposed in the paper [1] uses contrastive learning to capture the similarity between samples and uses the noise student framework to further improve the stability of the model. In the experiment, 7 key indicators related to chords are used to evaluate the performance of this method and methods 1 to 4. The key indicators are Root (the judgment indicator of the root note), Thirds (the judgment indicator of the root note and the third note), Triads (the judgment indicator of the root note, the third note and the fifth note), MajMin (the judgment indicator of the triad), Sevenths (the judgment indicator of the chord), and Sevenths (the judgment indicator of the chord). The three key indicators are as follows: Root is used to evaluate the accuracy of the root note, Thirds is used to evaluate the accuracy of the root note and the third note, Triads is used to evaluate the accuracy of the root note, the third note and the fifth note, MajMin is used to evaluate the accuracy of the triad, Sevenths is used to evaluate the accuracy of the seventh chord, Tetrads is used to evaluate the accuracy of the ninth chord, and Mirex is used to evaluate the accuracy based on the music information retrieval evaluation rules. The key indicators are all in percentage.

[0082]

[0083] It can be seen that the chord recognition model based on semi-supervised learning provided by the method of the present invention is only slightly inferior to Method 4 in the key metrics of chord recognition. This may be because Method 4 uses a structured algorithm for decomposed chord labels, which helps to more accurately capture and process features related to Thirds, thus achieving better recognition accuracy in the Thirds metric. In other key metrics, including Root, Traids, MajMin, Sevenths, Tetrads, and Mirex, this method has improved compared to other semi-supervised algorithms in the current chord recognition field. This fully demonstrates the superiority of this method in the semi-supervised chord recognition task, which can better utilize unlabeled audio data and improve the overall chord recognition accuracy by learning more potential chord features.

[0084] To better demonstrate the research advantages of semi-supervised algorithms compared to supervised algorithms in the field of chord recognition, the recognition accuracies of the method of the present invention and Method 1 on different chord attributes are visually compared. See Figure 3 , where the meanings of the abscissa are as follows: maj represents major triad; maj7 represents major seventh chord; 7 represents seventh chord; min represents minor triad; min7 represents minor seventh chord; aug represents augmented triad; min6 represents minor sixth chord; hdim7 represents half-diminished seventh chord; dim7 represents diminished seventh chord; dim represents diminished triad; maj6 represents major sixth chord; sus4 represents suspended fourth chord; the meaning of the ordinate is the chord recognition accuracy; Ours represents the chord recognition model based on semi-supervised learning provided by the method of the present invention.

[0085] It can be seen that except for the decrease in recognition accuracy of this method compared to Method 1 in major triads and minor seventh chords, significant improvements have been achieved in the recognition accuracies of the remaining chord attributes. Among them, the major seventh chord has increased by about 14.0%, the seventh chord has increased by 12.0%, and the minor triad has increased by about 10.1%. And this advantage is particularly prominent in rare chord attributes. It can be seen that when using Method 1, the minor sixth chord, half-diminished seventh chord, diminished seventh chord, and suspended fourth chord are not recognized. However, by introducing unlabeled audio data, the method of the present invention successfully overcomes this problem, and both the minor sixth chord and the half-diminished seventh chord have been greatly improved: the minor sixth chord has increased by 29.2%, the half-diminished seventh chord has increased by 10.4%, the diminished seventh chord has increased by 0.6%, and the suspended fourth chord has increased by 0.3%.

[0086] Regarding the situation where the recognition accuracy of the major triad and minor seventh chord decreases when using the method of the present invention, the possible reasons are as follows: Compared with other samples, the major triad and minor seventh chord are more common in this test dataset, and the number of samples is relatively larger. The Baseline algorithm may have achieved a relatively high recognition accuracy after training with a large number of samples, while the introduction of perturbed unlabeled audio samples by this method has affected this accuracy to a certain extent. In summary, the semi-supervised chord recognition algorithm demonstrates strong potential and advantages in the field of chord recognition compared with the supervised algorithm, especially having high practical value and research significance in improving the recognition accuracy of rare chords.

[0087] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the chord recognition model training method based on semi-supervised learning are implemented. Specifically, the computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include RAM (Random Access Memory) and / or cache memory, etc. The non-volatile memory may include ROM (Read Only Memory), hard disk, flash memory, optical disc, magnetic disk, etc.

[0088] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the chord recognition model training method based on semi-supervised learning are implemented. Among them, the memory may contain internal memory, such as high-speed random access memory, and may also include non-volatile memory, such as at least one disk memory, etc.; the processor, network interface, and memory are interconnected through an internal bus, and this internal bus may be an Industry Standard Architecture bus, a Peripheral Component Interconnect standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include internal memory and non-volatile memory and provide instructions and data to the processor.

[0089] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical memory, etc.) that contain computer-usable program code.

[0090] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer device or other programmable data processing devices generate means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0091] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0092] These computer program instructions can also be loaded onto a computer device or other programmable data processing devices, such that a series of operation steps are executed on the computer device or other programmable devices to generate a process implemented by the computer device. Thus, the instructions executed on the computer device or other programmable devices provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0093] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they learn the basic creative concepts. Therefore, the appended claims are intended to be construed to include the preferred embodiments as well as all changes and modifications falling within the scope of the present invention.

[0094] Obviously, those skilled in the art can make various modifications and variations to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.

Claims

1. A method for training a chord recognition model based on semi-supervised learning, characterized in that, Including the following steps: Taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A, combining with the true chord category probability distribution of the labeled audio sample set a, screening out the correctly predicted labeled audio sample subset X, and calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X; Performing first weak augmentation, second weak augmentation, and strong augmentation on the unlabeled audio sample set w respectively to obtain an unlabeled audio sample set b, an unlabeled audio sample set c, and an unlabeled audio sample set d, and inputting the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain a predicted chord category probability distribution B, a predicted chord category probability distribution C, and a predicted chord category probability distribution D respectively; Based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, calculating the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w, and combining with the adaptive confidence threshold for each chord category in the labeled audio sample subset X to divide into a pseudo-label sample set I and a low-confidence sample set II; Let the set of unlabeled audio samples be \(w = \{w_1, w_2, w_3, \ldots, w\) t \}. Calculate the mean of the predicted chord class probability distributions for each sample in the set of unlabeled audio samples \(w\), specifically: Among them, w1, w2, w3, …, w t are the 1st, 2nd, 3rd, …, t-th samples in the unlabeled audio sample set w respectively; is the mean of the predicted chord class probability distribution of the t-th sample in the unlabeled audio sample set w; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set b; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set c; is the network parameter of the supervised model; Calculating a first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; based on the contrast learning method, dividing the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculating a second loss using a contrast loss function; based on the first loss and the second loss, combining with the backpropagation algorithm to update the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning; The taking the preprocessed labeled audio sample set a as input and training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A specifically includes: Perform the third weak augmentation on the labeled audio sample set x = {x1, x2, x3, …, x s}, and obtain the preprocessed labeled audio sample set a. Among them, the third weak augmentation is to adjust the audio rate, and x1, x2, x3, …, x s are the 1st, 2nd, 3rd, …, s-th samples in the labeled audio sample set x respectively; Select the cross-entropy loss function Loss x As the loss function of the supervised model, use the preprocessed labeled audio sample set a as the input to train the supervised model, obtaining a trained supervised model and the predicted chord category probability distribution A; The cross-entropy loss function Loss x Specifically: where log is the logarithm; s is the total number of samples in the preprocessed labeled audio sample set a; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, , otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j.

2. The method for training a chord recognition model based on semi-supervised learning according to claim 1, wherein, The calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X specifically is: where \(i\) is the sample index; \(j\) is the chord category index; is the adaptive confidence threshold corresponding to the \(j\)-th chord category; is the total number of samples of the \(j\)-th chord in the subset \(X\) of labeled audio samples with correct predictions; is the true chord category probability distribution of the \(i\)-th sample of the \(j\)-th chord in the subset \(X\) of labeled audio samples with correct predictions; is the predicted chord category probability distribution of the \(i\)-th sample of the \(j\)-th chord in the subset \(X\) of labeled audio samples with correct predictions; is element-wise multiplication.

3. A method for training a chord recognition model based on semi-supervised learning according to claim 1, characterized in that The first weak augmentation is to adjust the audio rate; the second weak augmentation is to add noise; the strong augmentation is time-frequency masking.

4. A method for training a chord recognition model based on semi-supervised learning according to claim 1, characterized in that The combining with the adaptive confidence threshold for each chord category in the labeled audio sample subset X to divide into a pseudo-label sample set I and a low-confidence sample set II specifically includes: Sequentially determining whether the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w is ≥ the adaptive confidence threshold of the corresponding chord category in the labeled audio sample subset X. If the determination result is yes, dividing the corresponding sample into the pseudo-label sample set I; if the determination result is no, dividing the corresponding sample into the low-confidence sample set II.

5. A method for training a chord recognition model based on semi-supervised learning according to claim 1, characterized in that The based on the contrast learning method, dividing the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculating a second loss using a contrast loss function specifically includes: Sequentially determining whether each sample in the low-confidence sample set II is a sample that has undergone the first weak augmentation and the second weak augmentation. If the determination is yes, dividing the corresponding sample into the positive sample pair in the contrast learning method; if the determination is no, dividing the corresponding sample into the negative sample pair in the contrast learning method; Based on the similarity difference between the positive sample pairs and the negative sample pairs, a contrastive loss function is adopted to calculate the second loss, specifically as follows: where log is the logarithm; is the second loss; , is the similarity between sample g in unlabeled audio sample set b and sample h in unlabeled audio sample set c, is the similarity between sample g in unlabeled audio sample set b and sample k in unlabeled audio sample set d, T is the matrix transpose, ||B|| is the norm of B, ||C|| is the norm of C, ||D|| is the norm of D, is an adjustable temperature parameter; is the total number of samples in low-confidence sample set II, , all belong to low-confidence sample set II, g is used to traverse the samples in low-confidence sample set II and calculate the similarity-related loss of each sample with other samples; h and k are respectively the sample indices used to represent the samples that form positive and negative sample pairs with sample g.

6. A chord recognition model training system based on semi-supervised learning, characterized in that, Including: A first module, which is used to take the preprocessed labeled audio sample set a as input, train a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A, combine the true chord category probability distribution of the labeled audio sample set a, screen out the subset X of the labeled audio samples with correct predictions, and calculate the adaptive confidence threshold for each chord category in the subset X of the labeled audio samples; the specific steps of taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A include: Perform the third weak augmentation on the labeled audio sample set x = {x1, x2, x3, …, x s} to obtain the preprocessed labeled audio sample set a, where the third weak augmentation is to adjust the audio rate, and x1, x2, x3, …, x s are the 1st, 2nd, 3rd, …, s-th samples in the labeled audio sample set x, respectively; Select the cross-entropy loss function Loss x As the loss function of the supervised model, use the preprocessed set a of labeled audio samples as the input to train the supervised model, and obtain the trained supervised model and the predicted chord category probability distribution A; The cross-entropy loss function Loss x Specifically: where log is the logarithm; s is the total number of samples in the preprocessed labeled audio sample set a; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, , otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j; A second module, which is used to perform the first weak augmentation, the second weak augmentation and the strong augmentation on the unlabeled audio sample set w respectively to obtain an unlabeled audio sample set b, an unlabeled audio sample set c and an unlabeled audio sample set d, and input the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d into the trained supervised model to obtain a predicted chord category probability distribution B, a predicted chord category probability distribution C and a predicted chord category probability distribution D respectively; The third module is used to calculate the mean of the predicted chord class probability distribution for each sample in the unlabeled audio sample set w based on the predicted chord class probability distribution B and the predicted chord class probability distribution C, and combine the adaptive confidence threshold for each chord class in the labeled audio sample subset X to divide the pseudo-labeled sample set I and the low-confidence sample set II; let the unlabeled audio sample set w = {w1, w2, w3, …, w t}, and the specific calculation of the mean of the predicted chord class probability distribution for each sample in the unlabeled audio sample set w is as follows: Among them, w1, w2, w3, …, w t are the 1st, 2nd, 3rd, …, t-th samples in the unlabeled audio sample set w respectively; is the mean of the predicted chord class probability distributions of the t-th sample in the unlabeled audio sample set w; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set b; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set c; are the network parameters of the supervised model; A fourth module, which is used to calculate the first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; based on the contrastive learning method, divide the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculate the second loss by using the contrastive loss function; based on the first loss and the second loss, combine the backpropagation algorithm to update the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method for training a chord recognition model based on semi-supervised learning according to any one of claims 1 to 5 are implemented.

8. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method for training a chord recognition model based on semi-supervised learning according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Image classification model training method, image classification method, equipment and medium

    CN115496955A

  • Music gene expression programming-based music evolution method

    CN118571196A