Semi-supervised learning-based chord recognition model training method, system and product
By introducing adaptive confidence threshold screening, multi-dimensional data enhancement and pseudo-label hierarchical utilization in the chord recognition model, combined with the comparison learning method, the problem of low accuracy of rare chord recognition in the existing semi-supervised chord recognition algorithm is solved, and the overall chord recognition accuracy is improved.
Patent Information
- Application Number
- CN202510620403.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-14
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-05-14
AI Technical Summary
The existing semi-supervised automatic chord recognition algorithms have low accuracy in rare chord recognition and reduced overall chord recognition accuracy.
By using the pre-processed labeled audio sample set to train a supervised model, calculate the adaptive confidence threshold of each chord category, and perform multi-dimensional enhancement processing on the labelless audio sample set to generate a pseudo-label sample set and a low confidence sample set. Combined with the backpropagation algorithm to update the model parameters, a chord recognition model based on semi-supervised learning is obtained.
The utilization effect of label-free audio sample sets is improved, the recognition ability of rare chords is improved, the high recognition rate of common chords is maintained, and the overall chord recognition accuracy is improved.
Smart Images

Figure CN120126508A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of chord recognition, and in particular to a method, a system and a product for training a chord recognition model based on semi-supervised learning. Background Art
[0002] Currently, according to the different chord label information of the input sample set during the training process, the research directions in the field of automatic chord recognition mainly focus on supervised automatic chord recognition algorithms and semi-supervised automatic chord recognition algorithms. Among them, the supervised automatic chord recognition algorithm focuses on feature optimization and network structure optimization, while the semi-supervised automatic chord recognition algorithm pays more attention to the effective utilization of unlabeled audio data. With the booming development of the Internet music market, unlabeled audio data has become relatively easy to obtain, providing new opportunities for the research of automatic chord recognition. Under this background, semi-supervised learning, as a learning strategy that combines labeled data and unlabeled audio data, has attracted much attention.
[0003] In the existing semi-supervised automatic chord recognition algorithms, the recognition accuracy of rare chords is significantly lower than that of common chords, and there is a serious problem of classification imbalance. This is mainly because the available labeled audio sample sets publicly disclosed in the current chord recognition field are scarce, it is difficult to collect label data and requires extremely strong professional knowledge, resulting in too few labeled audio samples available for training. To address this problem, Marcelo Bortolozzo et al. began to consider expanding the data set by introducing an unlabeled audio sample set. By using a self-learning method, a teacher model generates pseudo-labels for a large number of unlabeled audio samples, and then a classification-balanced sample set is selected for model training. Although this method does improve the recognition accuracy of rare chords, it inevitably loses the recognition accuracy of common chords, resulting in a decrease in the overall chord recognition accuracy.
[0004] Therefore, how to make full use of unlabeled audio sample data and improve the overall chord recognition accuracy has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0005] The purpose of the present invention is to provide a method, a system and a product for training a chord recognition model based on semi-supervised learning, so as to overcome the problem that the overall chord recognition accuracy is reduced due to the large amount of unlabeled audio samples in the prior art.
[0006] The present invention solves the above technical problems through the following technical solutions: A method for training a chord recognition model based on semi-supervised learning, comprising the following steps: Taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A, combining with the true chord category probability distribution of the labeled audio sample set a, screening out the subset X of labeled audio samples with correct predictions, and calculating the adaptive confidence threshold for each chord category in the subset X of labeled audio samples; Performing first weak augmentation, second weak augmentation, and strong augmentation on the unlabeled audio sample set w respectively to obtain an unlabeled audio sample set b, an unlabeled audio sample set c, and an unlabeled audio sample set d, and inputting the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain a predicted chord category probability distribution B, a predicted chord category probability distribution C, and a predicted chord category probability distribution D respectively; Based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, calculating the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w, and combining with the adaptive confidence threshold for each chord category in the subset X of labeled audio samples to divide and obtain a pseudo-label sample set I and a low-confidence sample set II; Based on the pseudo-label sample set I and the predicted chord category probability distribution D, calculating a first loss; based on the contrast learning method, dividing the low-confidence sample set II into positive sample pairs and negative sample pairs, and using a contrast loss function to calculate a second loss; based on the first loss and the second loss, and combining with the backpropagation algorithm, updating the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.
[0007] A further improvement of the present invention lies in that: taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and a predicted chord category probability distribution A, specifically including: Performing third weak augmentation on the labeled audio sample set x = {x 1 , x 2 , x 3 , …, x s} to obtain the preprocessed labeled audio sample set a, where the third weak augmentation is to adjust the audio rate, and x 1 , x 2 , x 3 , …, x s are the 1st, 2nd, 3rd, …, s-th samples in the labeled audio sample set x respectively; Selecting the cross-entropy loss function Loss x as the loss function of the supervised model, taking the preprocessed labeled audio sample set a as input, training the supervised model to obtain a trained supervised model and a predicted chord category probability distribution A; The cross-entropy loss function Lossx Specifically:
[0008] Among them, log is the logarithm; s is the total number of samples in the labeled audio sample set a after preprocessing; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, , otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j.
[0009] A further improvement of the present invention lies in: calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X, specifically:
[0010] Among them, i is the sample index; j is the chord category index; is the adaptive confidence threshold corresponding to the j-th chord; is the total number of samples of the j-th chord in the labeled audio sample subset X with correct prediction; is the true chord category probability distribution of the i-th sample of the j-th chord in the labeled audio sample subset X with correct prediction; is the predicted chord category probability distribution of the i-th sample of the j-th chord in the labeled audio sample subset X with correct prediction; is element-wise multiplication.
[0011] A further improvement of the present invention lies in: the first weak augmentation is to adjust the audio rate; the second weak augmentation is to add noise; the strong augmentation is time-frequency masking.
[0012] A further improvement of the present invention lies in: assuming the unlabeled audio sample set w = {w 1 , w 2 , w 3 , …, w t}, calculating the mean value of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w, specifically:
[0013] Among them, w 1 , w 2 , w 3 , …, w t are the 1st, 2nd, 3rd, …, t-th samples in the unlabeled audio sample set w respectively; is the mean of the predicted chord class probability distribution for the t-th sample in the unlabeled audio sample set w; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set b; is the predicted chord class probability distribution corresponding to the t-th sample in the unlabeled audio sample set c; are the network parameters of the supervised model.
[0014] A further improvement of the present invention lies in: dividing the pseudo-label sample set I and the low-confidence sample set II by using the adaptive confidence threshold of each chord class in the labeled audio sample subset X, specifically including: Successively determine whether the mean of the predicted chord class probability distribution of each sample in the unlabeled audio sample set w is ≥ the adaptive confidence threshold of the corresponding chord class in the labeled audio sample subset X. If the judgment result is yes, divide the corresponding sample into the pseudo-label sample set I; if the judgment result is no, divide the corresponding sample into the low-confidence sample set II.
[0015] A further improvement of the present invention lies in: dividing the low-confidence sample set II into positive sample pairs and negative sample pairs based on the contrast learning method, and calculating the second loss by using the contrast loss function, specifically including: Successively determine whether each sample in the low-confidence sample set II is a sample that has undergone the first weak augmentation and the second weak augmentation. If the judgment is yes, divide the corresponding sample into the positive sample pairs in the contrast learning method; if the judgment is no, divide the corresponding sample into the negative sample pairs in the contrast learning method; Based on the similarity difference between the positive sample pairs and the negative sample pairs, use the contrast loss function to calculate the second loss, specifically:
[0016] where log is the logarithm; is the second loss; , is the similarity between the sample g in the unlabeled audio sample set b and the sample h in the unlabeled audio sample set c, is the similarity between the sample g in the unlabeled audio sample set b and the sample k in the unlabeled audio sample set d, T is the matrix transpose, ||B|| is the norm of B, ||C|| is the norm of C, ||D|| is the norm of D, is the adjustable temperature parameter; is the total number of samples in the low-confidence sample set II, , both belong to the low-confidence sample set II, g is used to traverse the samples in the low-confidence sample set II to calculate the similarity-related loss of each sample with other samples; h and k are respectively the sample indices used to represent the samples that form positive sample pairs and negative sample pairs with the g sample.
[0017] The present invention also provides a chord recognition model training system based on semi-supervised learning, including: A first module, which is used to take the preprocessed labeled audio sample set a as input, train a supervised model, obtain a trained supervised model and a predicted chord category probability distribution A, combine with the true chord category probability distribution of the labeled audio sample set a, screen out the correctly predicted labeled audio sample subset X, and calculate the adaptive confidence threshold for each chord category in the labeled audio sample subset X; A second module, which is used to perform first weak augmentation, second weak augmentation and strong augmentation on the unlabeled audio sample set w respectively to obtain an unlabeled audio sample set b, an unlabeled audio sample set c and an unlabeled audio sample set d, and input the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d into the trained supervised model to obtain a predicted chord category probability distribution B, a predicted chord category probability distribution C and a predicted chord category probability distribution D respectively; A third module, which is used to calculate the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, and combine with the adaptive confidence threshold of each chord category in the labeled audio sample subset X to divide and obtain a pseudo-label sample set I and a low-confidence sample set II; A fourth module, which is used to calculate a first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; based on the contrast learning method, divide the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculate a second loss using a contrast loss function; based on the first loss and the second loss, combine with the backpropagation algorithm to update the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.
[0018] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the above-mentioned chord recognition model training method based on semi-supervised learning are implemented.
[0019] The present invention also provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the above-mentioned chord recognition model training method based on semi-supervised learning are implemented.
[0020] Compared with the prior art, the positive and progressive effects of the present invention are as follows: The chord recognition model training method based on semi-supervised learning provided by the present invention uses a pre-processed labeled audio sample set a to train a supervised model, calculates the adaptive confidence threshold for each chord category based on the subset X of correctly predicted labeled audio samples; performs three types of enhancement processing on the unlabeled audio sample set w respectively, and inputs the enhanced samples into the trained supervised model to obtain different predicted probability distributions; divides the pseudo-label sample set I and the low-confidence sample set II according to the mean of the predicted chord category probability distribution of the unlabeled audio sample set w and the adaptive confidence threshold; calculates the first loss based on the pseudo-label sample set I, calculates the second loss for the low-confidence sample set II using the contrastive learning method, combines the two, and updates the model parameters through the backpropagation algorithm to obtain a chord recognition model based on semi-supervised learning. This method improves the utilization effect of the unlabeled audio sample set through adaptive confidence threshold screening, multi-dimensional data enhancement, and hierarchical utilization of pseudo-labels; by introducing the contrastive learning method, deeply mines the low-confidence sample set, captures the subtle differences and internal connections between samples, combines the first loss and the second loss, and optimizes the parameters of the supervised model through the backpropagation algorithm, so that the supervised model can improve the recognition ability of rare chords while maintaining a high recognition rate for common chords, thereby achieving the improvement of the overall chord recognition accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] The accompanying drawings in the specification are used to provide a further understanding of the present invention, and constitute a part of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention, and do not constitute an improper limitation of the present invention.
[0022] Figure 1 It is a schematic flow chart of a chord recognition model training method based on semi-supervised learning of the present invention; Figure 2 It is a flow block diagram of a chord recognition model training method based on semi-supervised learning of the present invention; Figure 3 It is a comparative experiment result diagram of Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0023] Glossary: Cross-entropy loss function: A loss function widely used in classification tasks, used to measure the difference between the probability distribution predicted by the model and the true probability distribution, which can intuitively reflect the inconsistency between the model prediction result and the true chord category, thereby effectively guiding the update of the model network parameters and improving the prediction accuracy of the model.
[0024] The following further describes the present invention in detail with reference to the accompanying drawings and specific embodiments, which is an explanation of the present invention rather than a limitation.
[0025] See Figure 1 andFigure 2 , a chord recognition model training method based on semi-supervised learning, comprising the following steps: Taking the preprocessed labeled audio sample set a as input, training a supervised model to obtain a trained supervised model and the predicted chord category probability distribution A, combining with the true chord category probability distribution of the labeled audio sample set a, screening out the correctly predicted labeled audio sample subset X, and calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X; Performing first weak augmentation, second weak augmentation, and strong augmentation on the unlabeled audio sample set w respectively to obtain the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d, and inputting the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain the predicted chord category probability distribution B, the predicted chord category probability distribution C, and the predicted chord category probability distribution D respectively; Based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, calculating the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w, and combining with the adaptive confidence threshold for each chord category in the labeled audio sample subset X to divide and obtain the pseudo-label sample set I and the low-confidence sample set II; Calculating the first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; based on the contrast learning method, dividing the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculating the second loss using the contrast loss function; based on the first loss and the second loss, combining with the backpropagation algorithm, updating the network parameters of the trained supervised model to obtain a chord recognition model based on semi-supervised learning.
[0026] Updating the network parameters of the trained supervised model based on the first loss and the second loss, combining with the backpropagation algorithm, to obtain a chord recognition model based on semi-supervised learning. Essentially, it is to improve the initial supervised model by fusing unlabeled audio data to make it have the ability of semi-supervised learning. Specifically: weighting and summing the first loss and the second loss according to a certain ratio to obtain the total loss function; calculating the gradient information of the total loss with respect to the network parameters of the supervised model, and using the backpropagation algorithm to propagate the gradient information from the output layer of the trained supervised model to the input layer, thereby updating the network parameters of the supervised model; repeating the above process to iteratively optimize the network parameters of the supervised model, so that the supervised model can continuously improve the accuracy and robustness of chord recognition while using labeled audio data and unlabeled audio data.
[0027] This method improves the utilization effect of the unlabeled audio sample set through adaptive confidence threshold screening, multi-dimensional data augmentation, and hierarchical utilization of pseudo-labels. By introducing the contrastive learning method, it deeply mines the low-confidence sample set, captures the subtle differences and internal connections between samples, combines the first loss and the second loss, and optimizes the parameters of the supervised model through the backpropagation algorithm, enabling the supervised model to maintain a high recognition rate for common chords while enhancing the recognition ability for rare chords, thereby achieving an improvement in the overall chord recognition accuracy.
[0028] Specifically, taking the preprocessed labeled audio sample set a as the input, training a supervised model to obtain the trained supervised model and the predicted chord category probability distribution A specifically includes: Performing the third weak augmentation on the labeled audio sample set x = {x 1 , x 2 , x 3 , …, x s} to obtain the preprocessed labeled audio sample set a, where the third weak augmentation is to adjust the audio rate, and x 1 , x 2 , x 3 , …, x s are respectively the 1st, 2nd, 3rd, …, s-th samples in the labeled audio sample set x; Selecting the cross-entropy loss function Loss x as the loss function of the supervised model, taking the preprocessed labeled audio sample set a as the input, training the supervised model to obtain the trained supervised model and the predicted chord category probability distribution A; The cross-entropy loss function Loss x is specifically:
[0029] where log is the logarithm; s is the total number of samples in the preprocessed labeled audio sample set a; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, , otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j.
[0030] Specifically, calculating the adaptive confidence threshold for each chord category in the labeled audio sample subset X specifically includes:
[0031] where \(i\) is the sample index; \(j\) is the chord category index; is the adaptive confidence threshold corresponding to the \(j\)th chord category; is the total number of samples of the \(j\)th chord in the subset \(X\) of labeled audio samples with correct predictions; is the true chord category probability distribution of the \(i\)th sample of the \(j\)th chord in the subset \(X\) of labeled audio samples with correct predictions; is the predicted chord category probability distribution of the \(i\)th sample of the \(j\)th chord in the subset \(X\) of labeled audio samples with correct predictions; is element-wise multiplication.
[0032] Specifically, the first weak augmentation is to adjust the audio rate; the second weak augmentation is to add noise; the strong augmentation is time-frequency masking.
[0033] Specifically, let the unlabeled audio sample set \(w = \{w 1 , w 2 , w 3 , \cdots, w t \}, the calculation of the mean of the predicted chord category probability distributions of each sample in the unlabeled audio sample set \(w\) is specifically as follows:
[0034] where \(w 1 , w 2 , w 3 , \cdots, w t are the 1st, 2nd, 3rd, \cdots, \(t\)th samples in the unlabeled audio sample set \(w\) respectively; is the mean of the predicted chord category probability distribution of the \(t\)th sample in the unlabeled audio sample set \(w\); is the predicted chord category probability distribution corresponding to the \(t\)th sample in the unlabeled audio sample set \(b\); is the predicted chord category probability distribution corresponding to the \(t\)th sample in the unlabeled audio sample set \(c\); are the network parameters of the supervised model.
[0035] Specifically, the combination of the adaptive confidence thresholds of each chord category in the subset \(X\) of labeled audio samples to divide and obtain the pseudo-label sample set \(I\) and the low-confidence sample set \(II\) specifically includes: Successively judge whether the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set \(w\) is \(\geq\) the adaptive confidence threshold of the corresponding chord category in the subset \(X\) of labeled audio samples. If the judgment result is yes, the corresponding sample is divided into the pseudo-label sample set \(I\); if the judgment result is no, the corresponding sample is divided into the low-confidence sample set \(II\).
[0036] In a specific embodiment of the present invention, the first loss is calculated using the cross-entropy loss function, and the second loss is calculated using the contrastive loss function.
[0037] Specifically, for the contrastive learning method, the low-confidence sample set II is divided into positive sample pairs and negative sample pairs, and the contrastive loss function is used to calculate the second loss, which specifically includes: Judging in turn whether each sample in the low-confidence sample set II is a sample that has undergone the first weak augmentation and the second weak augmentation. If the judgment is yes, the corresponding sample is divided into the positive sample pair in the contrastive learning method; if the judgment is no, the corresponding sample is divided into the negative sample pair in the contrastive learning method; Based on the similarity difference between the positive sample pair and the negative sample pair, the contrastive loss function is used to calculate the second loss, specifically:
[0038] where log is the logarithm; is the second loss; , is the similarity between sample g in the unlabeled audio sample set b and sample h in the unlabeled audio sample set c, is the similarity between sample g in the unlabeled audio sample set b and sample k in the unlabeled audio sample set d, T is the matrix transpose, ||B|| is the norm of B, ||C|| is the norm of C, ||D|| is the norm of D, is an adjustable temperature parameter; is the total number of samples in the low-confidence sample set II, , both belong to the low-confidence sample set II, g is used to traverse the samples in the low-confidence sample set II to calculate the similarity-related loss of each sample with other samples; h and k are respectively the sample indices used to represent the samples that form positive and negative sample pairs with the g sample.
[0039] Based on the same inventive concept, the present invention provides a chord recognition model training system based on semi-supervised learning, including: The first module is used to take the preprocessed labeled audio sample set a as input, train a supervised model to obtain a trained supervised model and the predicted chord category probability distribution A, combine the true chord category probability distribution of the labeled audio sample set a, screen out the correctly predicted labeled audio sample subset X, and calculate the adaptive confidence threshold of each chord category in the labeled audio sample subset X; The second module is used to perform the first weak augmentation, the second weak augmentation, and the strong augmentation on the unlabeled audio sample set w respectively, obtaining the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d respectively. Input the unlabeled audio sample set b, the unlabeled audio sample set c, and the unlabeled audio sample set d into the trained supervised model to obtain the predicted chord category probability distributions B, C, and D respectively; The third module is used to calculate the mean of the predicted chord category probability distribution for each sample in the unlabeled audio sample set w based on the predicted chord category probability distributions B and C, and combine the adaptive confidence threshold for each chord category in the labeled audio sample subset X to divide and obtain the pseudo-labeled sample set I and the low-confidence sample set II; The fourth module is used to calculate the first loss based on the pseudo-labeled sample set I and the predicted chord category probability distribution D; based on the contrastive learning method, divide the low-confidence sample set II into positive sample pairs and negative sample pairs, and calculate the second loss using the contrastive loss function; based on the first loss and the second loss, and combined with the backpropagation algorithm, update the network parameters of the trained supervised model to obtain the chord recognition model based on semi-supervised learning.
[0040] To verify the advancement of the chord recognition model training method based on semi-supervised learning of the present invention, the chord recognition model based on semi-supervised learning obtained by the method of the present invention is used for testing, and the method of the present invention is compared with the semi-supervised algorithms in the current chord recognition field. The results of the comparative experiment are shown in Table 1. Among them, Method 1 is to train a supervised model only through a labeled audio sample set, which represents the baseline performance when no semi-supervised strategy is adopted; Method 2 is the semi-supervised chord recognition algorithm based on a variational autoencoder proposed in the literature
Semi-supervised neural chord estimation based on a variational autoencoder with latent chord labels and features
Improving the classification of rare chords with unlabeled data
Deep Semi-Supervised Learning With Contrastive Learning in Large Vocabulary Automatic Chord Recognition
[0041]
[0042] It can be seen that the chord recognition model based on semi-supervised learning provided by the method of the present invention is only slightly inferior to Method 4 in terms of the key metrics of chord recognition. This may be because Method 4 uses a structured algorithm for decomposed chord labels, which helps to more accurately capture and process features related to Thirds, thus achieving better recognition accuracy in the Thirds metric. In other key metrics, including Root, Traids, MajMin, Sevenths, Tetrads, and Mirex, this method has been improved compared to other semi-supervised algorithms in the current chord recognition field. This also fully demonstrates the superiority of this method in the semi-supervised chord recognition task, which can better utilize unlabeled audio data, learn more potential chord features, and improve the overall chord recognition accuracy.
[0043] To better demonstrate the research advantages of semi-supervised algorithms compared to supervised algorithms in the field of chord recognition, the recognition accuracies of the method of the present invention and Method 1 on different chord attributes are visually compared. See Figure 3 , where the meanings of the abscissa are as follows: maj is a major triad; maj7 is a major seventh chord; 7 is a seventh chord; min is a minor triad; min7 is a minor seventh chord; aug is an augmented triad; min6 is a minor sixth chord; hdim7 is a half-diminished seventh chord; dim7 is a diminished seventh chord; dim is a diminished triad; maj6 is a major sixth chord; sus4 is a suspended fourth chord; the meaning of the ordinate is the chord recognition accuracy; Ours represents the chord recognition model based on semi-supervised learning provided by the method of the present invention.
[0044] It can be seen that except for the decrease in the recognition accuracy of the major triad and the minor seventh chord compared to Method 1, the recognition accuracies of the remaining chord attributes have been significantly improved. Among them, the major seventh chord has increased by about 14.0%, the seventh chord has increased by 12.0%, and the minor triad has increased by about 10.1%. And this advantage is particularly prominent in rare chord attributes. It can be seen that when using Method 1, the minor sixth chord, half-diminished seventh chord, diminished seventh chord, and suspended fourth chord are not recognized. However, by introducing unlabeled audio data, the method of the present invention successfully overcomes this problem, and both the minor sixth chord and the half-diminished seventh chord have been greatly improved: the minor sixth chord has increased by 29.2%, the half-diminished seventh chord has increased by 10.4%, the diminished seventh chord has increased by 0.6%, and the suspended fourth chord has increased by 0.3%.
[0045] Regarding the situation where the recognition accuracy of the major triad and minor seventh chord decreases when using the method of the present invention, the possible reasons are as follows: Compared with other samples, the major triad and minor seventh chord are more common in this test dataset, and the number of samples is relatively larger. The Baseline algorithm may have achieved a relatively high recognition accuracy after training with a large number of samples, while the introduction of perturbed unlabeled audio samples by this method has affected this accuracy to a certain extent. To sum up, the semi-supervised chord recognition algorithm shows great potential and advantages in the field of chord recognition compared with the supervised algorithm, especially in improving the recognition accuracy of rare chords, which has high practical value and research significance.
[0046] Based on the same inventive concept, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, the steps of the chord recognition model training method based on semi-supervised learning are implemented. Specifically, the computer-readable storage medium includes, but is not limited to, for example, volatile memory and / or non-volatile memory. The volatile memory may include RAM (Random Access Memory) and / or cache, etc. The non-volatile memory may include ROM (Read Only Memory), hard disk, flash memory, optical disc, magnetic disk, etc.
[0047] Based on the same inventive concept, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the chord recognition model training method based on semi-supervised learning are implemented. Among them, the memory may include internal memory, such as high-speed random access memory, and may also include non-volatile memory, such as at least one disk memory, etc.; the processor, network interface, and memory are interconnected through an internal bus, which may be an Industry Standard Architecture bus, a Peripheral Component Interconnect Standard bus, an Extended Industry Standard Architecture bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. The memory is used to store programs. Specifically, the program may include program code, and the program code includes computer operation instructions. The memory may include internal memory and non-volatile memory and provide instructions and data to the processor.
[0048] Those skilled in the art should understand that the embodiments of the present invention can be provided as a method or a computer program product. Therefore, the present invention can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM (Compact Disc Read-Only Memory), optical memory, etc.) that contain computer-usable program code.
[0049] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer device or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0050] These computer program instructions can also be stored in a computer-readable memory that can direct a computer device or other programmable data processing devices to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0051] These computer program instructions can also be loaded onto a computer device or other programmable data processing devices, such that a series of operation steps are executed on the computer device or other programmable devices to produce a process implemented by the computer device, thereby providing steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.
[0052] Although the preferred embodiments of the present invention have been described, those skilled in the art can make additional changes and modifications once they know the basic creative concepts. Therefore, the appended claims are intended to be construed as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0053] Obviously, those skilled in the art can make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if these modifications and variations of the present invention fall within the scope of the claims of the present invention and their equivalent technologies, the present invention is also intended to include these modifications and variations.
Claims
1. A chord recognition model training method based on semi-supervised learning, characterized in that: The following steps are involved: The preprocessed labeled audio sample set a is used as input to train a supervised model, and the trained supervised model and the predicted chord category probability distribution A are obtained. Combined with the true chord category probability distribution of the labeled audio sample set a, a correctly predicted labeled audio sample subset X is screened out, and an adaptive confidence threshold for each chord category in the labeled audio sample subset X is calculated; Performing the first weak enhancement, the second weak enhancement and the strong enhancement on the unlabeled audio sample set w, respectively, to obtain the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d, respectively, inputting the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d into the trained supervised model, to obtain the predicted chord category probability distribution B, the predicted chord category probability distribution C and the predicted chord category probability distribution D, respectively; Based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, the predicted chord category probability distribution mean of each sample in the unlabeled audio sample set w is calculated, and combined with the adaptive confidence threshold of each chord category in the labeled audio sample subset X, the pseudo-labeled sample set I and the low-confidence sample set II are obtained; Calculate the first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; Based on the contrastive learning method, the low confidence sample set II is divided into positive sample pairs and negative sample pairs, and the contrastive loss function is used to calculate the second loss. Based on the first loss and the second loss, combined with the back propagation algorithm, the network parameters of the trained supervised model are updated to obtain a chord recognition model based on semi-supervised learning.
2. A chord recognition model training method based on semi-supervised learning according to claim 1, characterized in that: The method uses the preprocessed labeled audio sample set a as input, trains the supervised model, and obtains the trained supervised model and the predicted chord category probability distribution A, specifically including: For a labeled audio sample set x={x1,x2,x3,…,x s }Perform a third weak enhancement to obtain a preprocessed labeled audio sample set a, wherein the third weak enhancement is to adjust the audio rate, x1, x2, x3, ..., x s are the 1st, 2nd, 3rd, …, sth samples in the labeled audio sample set x respectively; Select the cross entropy loss function Loss x As the loss function of the supervised model, the preprocessed labeled audio sample set a is taken as input to train the supervised model, and the trained supervised model and the predicted chord category probability distribution A are obtained; The cross entropy loss function Loss x Specifically: Wherein, log is the logarithm; s is the total number of samples in the preprocessed labeled audio sample set a; J is the total number of chord categories; i is the sample index; j is the chord category index; is the true chord category of the i-th sample in the true chord category probability distribution of the labeled audio sample set a. When the true chord category of the i-th sample is j, ,otherwise, ; is the probability that the predicted chord category of the i-th sample in the predicted chord category probability distribution A is j.
3. A chord recognition model training method based on semi-supervised learning according to claim 1, characterized in that: The adaptive confidence threshold of each chord category in the labeled audio sample subset X is calculated as follows: Where i is the sample index; j is the chord category index; is the adaptive confidence threshold corresponding to the j-th type of chord; is the total number of samples of the jth chord in the correctly predicted labeled audio sample subset X; To predict the true chord category probability distribution of the i-th sample of the j-th chord in the correct labeled audio sample subset X; To predict the predicted chord category probability distribution of the i-th sample of the j-th chord in the correct labeled audio sample subset X; Multiply the elements bit by bit.
4. The chord recognition model training method based on semi-supervised learning according to claim 1, characterized in that: The first weak enhancement is adjusting the audio rate; the second weak enhancement is adding noise; and the strong enhancement is time-frequency masking.
5. The chord recognition model training method based on semi-supervised learning according to claim 1 is characterized in that: The unlabeled audio sample set w={w1,w2,w3,…,w t }, the calculation obtains the predicted chord category probability distribution mean of each sample in the unlabeled audio sample set w, specifically: Among them, w1,w2,w3,…,w t are the 1st, 2nd, 3rd, …, tth samples in the unlabeled audio sample set w respectively; is the mean of the predicted chord category probability distribution of the tth sample in the unlabeled audio sample set w; is the predicted chord category probability distribution corresponding to the t-th sample in the unlabeled audio sample set b; is the predicted chord category probability distribution corresponding to the tth sample in the unlabeled audio sample set c; are the network parameters of the supervised model.
6. The chord recognition model training method based on semi-supervised learning according to claim 1, characterized in that: The adaptive confidence threshold of each chord category in the labeled audio sample subset X is combined to obtain a pseudo-label sample set I and a low-confidence sample set II, which specifically includes: It is determined in turn whether the mean of the predicted chord category probability distribution of each sample in the unlabeled audio sample set w is ≥ the adaptive confidence threshold of the corresponding chord category in the labeled audio sample subset X. If the judgment result is yes, the corresponding sample is divided into the pseudo-label sample set I; if the judgment result is no, the corresponding sample is divided into the low-confidence sample set II.
7. The chord recognition model training method based on semi-supervised learning according to claim 1, characterized in that: Based on the contrastive learning method, the low confidence sample set II is divided into positive sample pairs and negative sample pairs, and the contrastive loss function is used to calculate the second loss, which specifically includes: It is judged in turn whether each sample in the low confidence sample set II is a sample that has undergone the first weak enhancement and the second weak enhancement. If it is judged to be yes, the corresponding sample is classified as a positive sample pair in the contrastive learning method; if it is judged to be no, the corresponding sample is classified as a negative sample pair in the contrastive learning method; Based on the similarity difference between the positive sample pair and the negative sample pair, the contrast loss function is used to calculate the second loss, which is: Wherein, log is the logarithm; For the second loss; , is the similarity between sample g in unlabeled audio sample set b and sample h in unlabeled audio sample set c, is the similarity between sample g in unlabeled audio sample set b and sample k in unlabeled audio sample set d, T is matrix transpose, ||B|| is the norm of B, ||C|| is the norm of C, ||D|| is the norm of D, is an adjustable temperature parameter; is the total number of samples in low confidence sample set II, , all belong to the low confidence sample set II, g is used to traverse the samples in the low confidence sample set II and calculate the similarity loss between each sample and other samples; h and k are the sample indices used to represent the positive sample pairs and negative sample pairs with the g sample respectively.
8. A chord recognition model training system based on semi-supervised learning, characterized in that: include: The first module is used to take the preprocessed labeled audio sample set a as input, train the supervised model, obtain the trained supervised model and the predicted chord category probability distribution A, combine the real chord category probability distribution of the labeled audio sample set a, screen out the correctly predicted labeled audio sample subset X, and calculate the adaptive confidence threshold of each chord category in the labeled audio sample subset X; The second module is used to perform the first weak enhancement, the second weak enhancement and the strong enhancement on the unlabeled audio sample set w, respectively, to obtain the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d, and input the unlabeled audio sample set b, the unlabeled audio sample set c and the unlabeled audio sample set d into the trained supervised model to obtain the predicted chord category probability distribution B, the predicted chord category probability distribution C and the predicted chord category probability distribution D, respectively; The third module is used to calculate the predicted chord category probability distribution mean of each sample in the unlabeled audio sample set w based on the predicted chord category probability distribution B and the predicted chord category probability distribution C, and combine the adaptive confidence threshold of each chord category in the labeled audio sample subset X to obtain a pseudo-labeled sample set I and a low-confidence sample set II; The fourth module is used to calculate the first loss based on the pseudo-label sample set I and the predicted chord category probability distribution D; Based on the contrastive learning method, the low confidence sample set II is divided into positive sample pairs and negative sample pairs, and the contrastive loss function is used to calculate the second loss. Based on the first loss and the second loss, combined with the back propagation algorithm, the network parameters of the trained supervised model are updated to obtain a chord recognition model based on semi-supervised learning.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the chord recognition model training method based on semi-supervised learning described in any one of claims 1 to 7 are implemented.
10. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the chord recognition model training method based on semi-supervised learning described in any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Image classification model training method, image classification method, equipment and medium
CN115496955A
Music gene expression programming-based music evolution method
CN118571196A
Training method for semi-supervised learning model, image processing method, and device
US20230196117A1
Identification of electronic books for audiobook publication
US20250086478A1