Method and device for classifying music works, storage medium, and electronic device
By constructing a multi-dimensional feature matrix and a binary classification model, and combining objective and subjective evaluation information, the problem of low accuracy in music work classification in existing technologies has been solved, achieving more efficient and accurate music work classification.
Patent Information
- Application Number
- CN202310412537.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-11
- Publication Date
- 2025-11-04
- Estimated Expiration
- 2043-04-11
AI Technical Summary
Existing methods for classifying musical works only use pitch-related and objective features, failing to effectively utilize subjective features, resulting in low accuracy of classification results.
First and second music work classification models are constructed. Feature matrices are used to filter and classify the music works to be classified. Combined with objective and subjective evaluation information, the classification of three categories of music works is achieved through two binary classification models.
It improves the accuracy and efficiency of music classification, reduces reliance on convolutional neural network models and labor costs, and provides a better user experience.
Smart Images

Figure CN116401595B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present disclosure relate to the field of machine learning, and more particularly, embodiments of the present disclosure relate to a music work classification method, a music work classification device, a computer readable storage medium, and an electronic device. BACKGROUND
[0002] This section is intended to provide background or context to the embodiments of the present disclosure. The description herein is not admitted to be prior art merely by inclusion in this section.
[0003] In the existing music work classification method, the work is directly evaluated objectively from the pitch, rhythm, vibrato, timbre and other aspects of singing dry sound, and the subjective evaluation of the audience obtained on the publishing platform is rarely used, which makes the accuracy of the obtained music work classification result low. SUMMARY
[0004] However, in the prior art, on the one hand, since the music work is classified by using the method of extracting the pitch histogram of the dry sound, only the features related to the pitch are used, and the features reflecting the singing skills such as rhythm are not used, which makes the accuracy of the classification result low; on the other hand, since the music work is classified by using the method of convolutional neural network, only the fast Fourier transform spectrum features and the objective features of the singing work are used, and the subjective features are not used, which makes the accuracy of the classification result low.
[0005] Therefore, there is a great need for an improved music work classification method to classify the first music work to be classified and the second music work to be classified according to the first music work classification model and the second music work classification model, respectively, to obtain music works with the first work category, music works with the second work category, and music works with the third work category, and to improve the accuracy of the obtained music works with the first work category, music works with the second work category, and music works with the third work category on the basis of improving the classification efficiency of the music works.
[0006] In this context, embodiments of the present disclosure aim to provide a music work classification method, a music work classification device, a computer readable storage medium, and an electronic device.
[0007] According to one aspect of the present disclosure, a music work classification method is provided, comprising:
[0008] obtaining a first music work to be classified and first work attribute information of the first music work to be classified, and constructing a first feature matrix according to the first work attribute information;
[0009] perform feature screening on the first feature matrix to obtain a second feature matrix, and input the second feature matrix into a first music work classification model to obtain a first output result;
[0010] extract music works with a first work category from the first to-be-classified music works according to the first output result, and remove the music works with the first work category from the first to-be-classified music works to obtain second to-be-classified music works;
[0011] input a third feature matrix corresponding to the second to-be-classified music works into a second music work classification model to obtain a second output result, and extract music works with a second work category and music works with a third work category from the second to-be-classified music works according to the second output result.
[0012] In an exemplary embodiment of the present disclosure, the first work attribute information includes at least one of first dimension work attribute information, second dimension work attribute information, third dimension work attribute information, and fourth dimension work attribute information;
[0013] The first dimension work attribute information includes objective evaluation information, the second dimension work attribute information includes distribution information of the objective evaluation information, the third dimension work attribute information includes subjective evaluation information, and the third dimension work attribute information includes work quantity information.
[0014] In an exemplary embodiment of the present disclosure, a first feature matrix is constructed according to the first work attribute information, including:
[0015] A first dimension sub-feature is calculated according to the objective evaluation information, and a second dimension sub-feature is calculated according to the distribution information of the objective evaluation information;
[0016] A third dimension sub-feature is calculated according to the subjective evaluation information, and a fourth dimension sub-feature is calculated according to the work quantity information;
[0017] The first feature matrix is constructed according to the first dimension sub-feature, the second dimension sub-feature, the third dimension sub-feature, and the fourth dimension sub-feature.
[0018] In an exemplary embodiment of the present disclosure, the first dimension sub-feature is calculated according to the objective evaluation information, including:
[0019] An overall sentence score of each sentence included in the first to-be-classified music works is calculated according to the objective evaluation information of the each sentence, as well as an average and a median of the objective evaluation information;
[0020] According to the number of glides and the number of trills of each sentence, calculate the average number of glides and the average number of trills of each sentence;
[0021] According to the overall sentence score, the average number and the median number of the objective evaluation information, the average number of glides and the average number of trills of each sentence, generate a first dimension sub-feature.
[0022] In an exemplary embodiment of the present disclosure, the objective evaluation information includes at least one of singing pitch information, singing rhythm information, singing breath information, and singing emotion information;
[0023] According to the objective evaluation information of each sentence included in the first music work to be classified, calculate the overall sentence score of each sentence, and the average number and the median number of the objective evaluation information, including:
[0024] According to the pitch score in the singing pitch dimension, the rhythm score in the singing rhythm dimension, the breath score in the singing breath dimension, and the emotion score in the singing emotion dimension of each sentence included in the first music work to be classified, calculate the overall sentence score of each sentence;
[0025] According to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence, calculate the average number of singing pitch information, the average number of singing rhythm information, the average number of singing breath information, and the average number of singing emotion information;
[0026] According to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence, calculate the median number of singing pitch information, the median number of singing rhythm information, the median number of singing breath information, and the median number of singing emotion information.
[0027] In an exemplary embodiment of the present disclosure, according to the distribution information of the objective evaluation information, calculate the second dimension sub-feature, including:
[0028] According to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence, determine the pitch interval, the rhythm interval, the breath interval, and the emotion interval to which each sentence belongs;
[0029] According to the pitch interval, the rhythm interval, the breath interval, and the emotion interval to which each sentence belongs, calculate the pitch interval probability distribution, the rhythm interval probability distribution, the breath interval probability distribution, and the emotion interval probability distribution of each sentence;
[0030] According to the pitch interval probability distribution, the rhythm interval probability distribution, the breath interval probability distribution, and the emotion interval probability distribution, calculate the second dimension sub-feature.
[0031] In an example embodiment of the present disclosure, the third dimension sub-feature is calculated according to the subjective evaluation information, including:
[0032] Obtaining other music works of the work performer of the first music work to be classified except the first music work to be classified;
[0033] Generating the third dimension sub-feature according to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works.
[0034] In an example embodiment of the present disclosure, the subjective evaluation information includes at least one of work popularity information, work like information, gift information, work sharing information, work playing information, and work comment information.
[0035] The third dimension sub-feature is generated according to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works, including:
[0036] Obtaining work popularity information, work like information, gift information, work sharing information, work playing information, and work comment information of the other music works, and calculating work popularity average, work like quantity average, gift quantity average, work sharing times average, work playing times average, and work comment quantity average of the other music works;
[0037] The third dimension sub-feature is generated according to the work popularity information, work like information, gift information, work sharing information, work playing information, and work comment information of the first music work to be classified, and the work popularity average, work like quantity average, gift quantity average, work sharing times average, work playing times average, and work comment quantity average of the other music works.
[0038] In an example embodiment of the present disclosure, the fourth dimension sub-feature is calculated according to the work quantity information, including:
[0039] Obtaining the number of published works and the number of works deleted after publishing of the work performer of the first music work to be classified;
[0040] Determining the number of published and non-deleted works according to the number of published works and the number of works deleted after publishing, and generating the fourth dimension sub-feature according to the number of published works, the number of works deleted after publishing, and the number of published and non-deleted works.
[0041] In an example embodiment of the present disclosure, the first feature matrix is subjected to feature screening to obtain a second feature matrix, including:
[0042] Correlation coefficients between different dimensions of sub-features included in the first feature matrix are calculated, and feature screening is performed on the different dimensions of sub-features included in the first feature matrix according to the correlation coefficients, to obtain a second feature matrix.
[0043] In an example embodiment of the present disclosure, the method for classifying music works further comprises:
[0044] Original audio data is obtained, and a first category data set, a second category data set and a third category data set are constructed according to original labels of the original audio data;
[0045] The second category data set and the third category data set are merged to obtain a total data set, and a first network model is trained according to the first category data set and the total data set to obtain a first music work classification model;
[0046] A second network model is trained according to the second category data set and the third category data set to obtain a second music work classification model.
[0047] In an example embodiment of the present disclosure, the first music work classification model is used for classifying audio data with a third category, and the second music work classification model is used for classifying audio data with a first category and / or audio data with a second category;
[0048] The first category is a poor category, the second category is a general category, and the third category is an excellent category.
[0049] In an example embodiment of the present disclosure, the first network model is trained according to the first category data set and the total data set to obtain the first music work classification model, comprising:
[0050] A third feature matrix of original audio data included in the first category data set and original audio data included in the total data set is extracted;
[0051] The third feature matrix is subjected to feature screening to obtain a fourth feature matrix, and the fourth feature matrix is input into the first network model to obtain a first prediction result;
[0052] According to the first prediction result and original labels of the original audio data, a first harmonic mean of original audio data included in the first category data set and a second harmonic mean of original audio data included in the total data set are calculated;
[0053] The first network model is adjusted according to a first network model hyperparameter optimization condition that the first harmonic mean is greater than the second harmonic mean, to obtain the first music work classification model.
[0054] In an example embodiment of the present disclosure, the second network model is trained according to the second category data set and the third category data set to obtain a second music work classification model, including:
[0055] A fifth feature matrix of the original audio data included in the second category data and a sixth feature matrix of the original audio data included in the third category data are extracted;
[0056] The fifth feature matrix and the sixth feature matrix are subjected to feature screening to obtain a seventh feature matrix and an eighth feature matrix, and the seventh feature matrix and the eighth feature matrix are input into the second network model to obtain a second prediction result;
[0057] According to the second prediction result and original labels possessed by the original audio data, a third harmonic mean of the original audio data included in the second category data set and a fourth harmonic mean of the original audio data included in the third category data set are calculated;
[0058] The fourth harmonic mean is greater than the third harmonic mean as a hyperparameter tuning condition of the second network model, and the hyperparameters of the second network model are adjusted to obtain a second music work classification model.
[0059] According to an aspect of the present disclosure, a music work classification device is provided, including:
[0060] A first feature matrix construction module is configured to acquire a first to-be-classified music work and first work attribute information of the first to-be-classified music work, and construct a first feature matrix according to the first work attribute information;
[0061] A first music work classification module is configured to perform feature screening on the first feature matrix to obtain a second feature matrix, and input the second feature matrix into a first music work classification model to obtain a first data result;
[0062] A music work elimination module is configured to extract music works with a first work category from the to-be-classified music works according to the first output result, and eliminate music works with the first work category from the first to-be-classified music works to obtain second to-be-classified music works;
[0063] A second music work classification module is configured to input a third feature matrix corresponding to the second to-be-classified music works into a second music work classification model to obtain a second output result, and extract music works with a second work category and music works with a third work category from the second to-be-classified music works according to the second output result.
[0064] In an example embodiment of the present disclosure, according to the first work attribute information, a first feature matrix is constructed, including:
[0065] According to the objective evaluation information, a first dimension sub-feature is calculated, and according to distribution information of the objective evaluation information, a second dimension sub-feature is calculated;
[0066] According to the subjective evaluation information, a third dimension sub-feature is calculated, and according to the work quantity information, a fourth dimension sub-feature is calculated;
[0067] According to the first dimension sub-feature, the second dimension sub-feature, the third dimension sub-feature, and the fourth dimension sub-feature, the first feature matrix is constructed.
[0068] In an example embodiment of the present disclosure, according to the objective evaluation information, a first dimension sub-feature is calculated, including:
[0069] According to the objective evaluation information of each sentence included in the first music work to be classified, an overall sentence score of each sentence, and an average and a median of the objective evaluation information are calculated;
[0070] According to the number of slides and the number of trills of each sentence, an average number of slides and an average number of trills of each sentence are calculated;
[0071] According to the overall sentence score, the average and the median of the objective evaluation information, the average number of slides and the average number of trills of each sentence, a first dimension sub-feature is generated.
[0072] In an example embodiment of the present disclosure, the objective evaluation information includes at least one of singing pitch information, singing rhythm information, singing breath information, and singing emotion information;
[0073] According to the objective evaluation information of each sentence included in the first music work to be classified, an overall sentence score of each sentence, and an average and a median of the objective evaluation information are calculated, including:
[0074] According to the pitch score in the singing pitch dimension, the rhythm score in the singing rhythm dimension, the breath score in the singing breath dimension, and the emotion score in the singing emotion dimension of each sentence included in the first music work to be classified, an overall sentence score of each sentence is calculated;
[0075] According to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence, a pitch average of the singing pitch information, a rhythm average of the singing rhythm information, a breath average of the singing breath information, and an emotion average of the singing emotion information are calculated;
[0076] According to the pitch score, the rhythm score, the breath score and the emotion score of each sentence, a pitch median of the singing pitch information, a rhythm median of the singing rhythm information, a breath median of the singing breath information and an emotion median of the singing emotion information are calculated.
[0077] In an example embodiment of the present disclosure, the second dimension sub-feature is calculated according to the distribution information of the objective evaluation information, including:
[0078] According to the pitch score, the rhythm score, the breath score and the emotion score of each sentence, a pitch interval, a rhythm interval, a breath interval and an emotion interval to which each sentence belongs are determined;
[0079] According to the pitch interval, the rhythm interval, the breath interval and the emotion interval to which each sentence belongs, a pitch interval probability distribution, a rhythm interval probability distribution, a breath interval probability distribution and an emotion interval probability distribution of each sentence are calculated;
[0080] According to the pitch interval probability distribution, the rhythm interval probability distribution, the breath interval probability distribution and the emotion interval probability distribution, the second dimension sub-feature is calculated.
[0081] In an example embodiment of the present disclosure, the third dimension sub-feature is calculated according to the subjective evaluation information, including:
[0082] The singer of the first music work to be classified performs other music works except the first music work to be classified;
[0083] According to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works, the third dimension sub-feature is generated.
[0084] In an example embodiment of the present disclosure, the subjective evaluation information includes at least one of work popularity information, work like information, gift information, work sharing information, work playing information and work comment information.
[0085] According to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works, the third dimension sub-feature is generated, including:
[0086] The work popularity information, the work like information, the gift information, the work sharing information, the work playing information and the work comment information of the other music works are obtained, and the work popularity average, the work like quantity average, the gift quantity average, the work sharing times average, the work playing times average and the work comment quantity average of the other music works are calculated.
[0087] generate the third dimension sub-feature according to the work heat information, the work like information, the gift information, the work sharing information, the work playing information, the work comment information of the first music work to be classified, and the work heat mean value, the work like quantity mean value, the gift quantity mean value, the work sharing times mean value, the work playing times mean value, and the work comment quantity mean value of other music works.
[0088] In an example embodiment of the present disclosure, a fourth dimension sub-feature is calculated according to the work quantity information, including:
[0089] The number of published works and the number of works deleted after publication of the work performer of the first music work to be classified are acquired.
[0090] The number of published and non-deleted works is determined according to the number of published works and the number of works deleted after publication, and the fourth dimension sub-feature is generated according to the number of published works, the number of works deleted after publication, and the number of published and non-deleted works.
[0091] In an example embodiment of the present disclosure, the first feature matrix is subjected to feature screening to obtain a second feature matrix, including:
[0092] The correlation coefficients between the sub-features of different dimensions included in the first feature matrix are calculated, and the sub-features of different dimensions included in the first feature matrix are subjected to feature screening according to the correlation coefficients to obtain a second feature matrix.
[0093] In an example embodiment of the present disclosure, the music work classification device further includes:
[0094] A data set construction module is configured to acquire original audio data, and construct a first category data set, a second category data set, and a third category data set according to the original labels of the original audio data.
[0095] A first model training module is configured to merge the second category data set and the third category data set to obtain a total data set, and train a first network model according to the first category data set and the total data set to obtain a first music work classification model.
[0096] A second model training module is configured to train a second network model according to the second category data set and the third category data set to obtain a second music work classification model.
[0097] In an example embodiment of the present disclosure, the first music work classification model is used for classifying audio data with the third category, and the second music work classification model is used for classifying audio data with the first category and / or audio data with the second category.
[0098] The first category is a poor category, the second category is a general category, and the third category is an excellent category.
[0099] In an example embodiment of the present disclosure, a first music work classification model is obtained by training a first network model according to a first category data set and a total data set, including:
[0100] A third feature matrix of the original audio data included in the first category data set and the original audio data included in the total data set is extracted;
[0101] A fourth feature matrix is obtained by performing feature screening on the third feature matrix, and the fourth feature matrix is input into the first network model to obtain a first prediction result;
[0102] According to the first prediction result and the original label possessed by the original audio data, a first harmonic mean of the original audio data included in the first category data set and a second harmonic mean of the original audio data included in the total data set are calculated;
[0103] The first network model is adjusted according to a hyperparameter tuning condition that the first harmonic mean is greater than the second harmonic mean, to obtain the first music work classification model.
[0104] In an example embodiment of the present disclosure, a second music work classification model is obtained by training a second network model according to a second category data set and a third category data set, including:
[0105] A fifth feature matrix of the original audio data included in the second category data and a sixth feature matrix of the original audio data included in the third category data are extracted;
[0106] A seventh feature matrix and an eighth feature matrix are obtained by performing feature screening on the fifth feature matrix and the sixth feature matrix, and the seventh feature matrix and the eighth feature matrix are input into the second network model to obtain a second prediction result;
[0107] According to the second prediction result and the original label possessed by the original audio data, a third harmonic mean of the original audio data included in the second category data set and a fourth harmonic mean of the original audio data included in the third category data set are calculated;
[0108] The second network model is adjusted according to a hyperparameter tuning condition that the fourth harmonic mean is greater than the third harmonic mean, to obtain the second music work classification model.
[0109] According to an aspect of the example embodiments of the present disclosure, there is provided a computer-readable storage medium having stored thereon a computer program which, when executed by a processor, implements the method of classifying a musical work according to any of the example embodiments described above.
[0110] According to an aspect of the example embodiments of the present disclosure, there is provided an electronic device comprising:
[0111] a processor; and
[0112] a memory for storing executable instructions of the processor;
[0113] wherein the processor is configured to execute the method of classifying a musical work according to any of the example embodiments described above via execution of the executable instructions.
[0114] According to the method of classifying a musical work and the apparatus for classifying a musical work of the embodiments of the present disclosure, the first musical work to be classified and the first work attribute information of the first musical work to be classified can be obtained, and a first feature matrix can be constructed according to the first work attribute information. Then, the first feature matrix can be subjected to feature screening to obtain a second feature matrix, and the second feature matrix can be input into a first musical work classification model to obtain a first output result. Then, the musical work having a first work category can be extracted from the first musical work to be classified according to the first output result, and the musical work having the first work category can be removed from the first musical work to be classified to obtain a second musical work to be classified. Finally, the second musical work to be classified and the second feature matrix can be input into a second musical work classification model to obtain a second output result, and the musical work having a second work category and the musical work having a third work category can be extracted from the second musical work to be classified according to the second output result. Thus, the problem of the accuracy of the classification result caused by using the convolutional neural network model to classify the musical work can be significantly reduced, and the problem of high labor cost caused by the convolutional neural network model needing to rely on expert annotation data can be reduced, thereby providing a better experience for the user. BRIEF DESCRIPTION OF DRAWINGS
[0115] The above and other objects, features and advantages of the example embodiments of the present disclosure will be more apparent from the following detailed description taken in conjunction with the accompanying drawings, in which:
[0116] Figure 1 a flowchart of a method of classifying a musical work according to an example embodiment of the present disclosure is schematically shown;
[0117] Figure 2A model structure example diagram of a random forest is schematically shown according to an example embodiment of the present disclosure;
[0118] Figure 3 An example diagram of a specific training process of a first music piece classification model and a second music piece classification model is schematically shown according to an example embodiment of the present disclosure;
[0119] Figure 4 An example diagram of a frequency distribution of each score group in a song is schematically shown according to an example embodiment of the present disclosure;
[0120] Figure 5 An example diagram of an 11-dimensional feature format of each interval frequency distribution is schematically shown according to an example embodiment of the present disclosure;
[0121] Figure 6 A method flowchart of constructing a first feature matrix according to the first piece attribute information is schematically shown according to an example embodiment of the present disclosure;
[0122] Figure 7 An example diagram of specific feature categories included in the obtained first feature matrix is schematically shown according to an example embodiment of the present disclosure;
[0123] Figure 8 An example diagram of Pearson correlation coefficients between each feature is schematically shown according to an example embodiment of the present disclosure;
[0124] Figure 9 A block diagram of a music piece classification apparatus is schematically shown according to an example embodiment of the present disclosure;
[0125] Figure 10 A computer readable storage medium for storing the above-mentioned music piece classification method is schematically shown according to an example embodiment of the present disclosure;
[0126] Figure 11 An electronic device for implementing the above-mentioned music piece classification method is schematically shown according to an example embodiment of the present disclosure.
[0127] In the drawings, identical or corresponding reference signs indicate identical or corresponding parts. DETAILED DESCRIPTION
[0128] The principles and spirits of the present disclosure will be described below with reference to several example embodiments. It should be understood that these embodiments are given only to enable those skilled in the art to better understand and implement the present disclosure, and do not limit the scope of the present disclosure in any way. On the contrary, these embodiments are provided to make the present disclosure more thorough and complete, and to fully convey the scope of the present disclosure to those skilled in the art.
[0129] Those skilled in the art know that the embodiments of the present disclosure can be implemented as a system, a device, an apparatus, a method or a computer program product. Therefore, the present disclosure can be embodied as a whole hardware, a whole software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.
[0130] According to the embodiments of the present disclosure, a music work classification method, a music work classification device, a computer readable storage medium and an electronic device are provided.
[0131] In this document, it should be understood that the number of any elements in the accompanying drawings is used only for example and not limitation, and any naming is only for distinction and does not have any limiting meaning.
[0132] The principles and spirits of the present disclosure are explained in detail below with reference to several representative embodiments of the present disclosure. SUMMARY
[0134] The applicant finds that the existing singing evaluation or classification technology can directly objectively evaluate the works from the aspects of singing dry tone, rhythm, vibrato, tone color, etc., and less use the subjective evaluation of the audience obtained on the publishing platform. At the same time, in the process of actual application, the classification of music works can be realized based on the following two ways: one implementation way is to use the method of extracting the pitch histogram of dry tone, to capture the shape and density of the histogram through GMM (Gaussian Mixture Models) fitting and K-Means clustering, to extract 7 features related to pitch, and to perform singing quality evaluation without standard pitch line reference; however, this method only uses features related to pitch, without using other objective features reflecting singing skills; and without using features that can measure the subjective evaluation of the audience group. Another implementation way is to build a convolutional neural network (CNN) architecture named Bi-DenseNet, to use the fast Fourier transform (FFT) spectrum to distinguish good and bad singing without standard pitch line reference; however, the model input of this method only uses the FFT spectrum, only containing objective features of the singing works; at the same time, the classification result obtained by this method only contains two categories of good and bad, excluding the singing works between the two, which does not conform to the objective law; further, this method also needs to rely on a large amount of human expert annotated data.
[0135] Based on this, the example embodiments of the present disclosure propose a music work classification method. On the one hand, by obtaining a first to-be-classified music work and first work attribute information of the first to-be-classified music work, a first feature matrix is constructed according to the first work attribute information. Then, feature screening is performed on the first feature matrix to obtain a second feature matrix, and the second feature matrix is input into a first music work classification model to obtain a first output result. Then, music works with a first work category are extracted from the first to-be-classified music work according to the first output result, and the music works with the first work category are removed from the first to-be-classified music work to obtain a second to-be-classified music work. Finally, the second to-be-classified music work and the second feature matrix are input into a second music work classification model to obtain a second output result, and music works with a second work category and music works with a third work category are extracted from the second to-be-classified music work according to the second output result. Since the first to-be-classified music work and the second to-be-classified music work can be classified by the first music work classification model and the second music work classification model to obtain music works of different work categories, the problem of low accuracy of the classification result caused by classifying the to-be-classified music work by one classification model can be avoided. On the other hand, since feature screening is performed before music work classification, the data amount of the second feature matrix is greatly reduced, thereby improving the classification efficiency of music works.
[0136] After introducing the basic principles of the present disclosure, various non-limiting embodiments of the present disclosure will be specifically introduced below.
[0137] Exemplary method
[0138] In the example embodiment, a music work classification method is first provided, which can run on a server, a server cluster, or a cloud server, etc. Of course, those skilled in the art can also run the method of the present disclosure on other platforms according to needs, which is not specially limited in the example embodiment. Specifically, referring to FIG. 1, the music work classification method can include the following steps: Figure 1
[0139] Step S110. Obtain a first to-be-classified music work and first work attribute information of the first to-be-classified music work, and construct a first feature matrix according to the first work attribute information.
[0140] Step S120. Perform feature screening on the first feature matrix to obtain a second feature matrix, and input the second feature matrix into a first music work classification model to obtain a first output result.
[0141] Step S130. Extracting music works with the first work category from the first music works to be classified according to the first output result, and removing the music works with the first work category from the first music works to be classified to obtain second music works to be classified;
[0142] Step S140. Inputting the third feature matrix corresponding to the second music works to be classified into the second music work classification model to obtain a second output result, and extracting music works with the second work category and music works with the third work category from the second music works to be classified according to the second output result.
[0143] In the above music work classification method, the first music works to be classified and the first work attribute information of the first music works to be classified can be obtained, the first feature matrix can be constructed according to the first work attribute information, the second feature matrix can be obtained by performing feature screening on the first feature matrix, and the first output result can be obtained by inputting the second feature matrix into the first music work classification model; then, the music works with the first work category can be extracted from the first music works to be classified according to the first output result, and the music works with the first work category can be removed from the first music works to be classified to obtain the second music works to be classified; finally, the second output result can be obtained by inputting the second music works to be classified and the second feature matrix into the second music work classification model, and the music works with the second work category and the music works with the third work category can be extracted from the second music works to be classified according to the second output result. Without using a single convolutional neural network model to classify all music works, the problem of reducing the accuracy of the classification result caused by using the convolutional neural network model to classify the music works is solved, the problem of high labor cost caused by the convolutional neural network model relying on expert labeled data is reduced, and a better experience is provided for users.
[0144] In the following, the music work classification method of the example embodiments of the present disclosure will be explained and described in combination with the accompanying drawings.
[0145] First, the proper nouns involved in the example embodiments of the present disclosure are explained and described.
[0146] Singing quality: the overall evaluation of the pleasant degree of a piece of music work in terms of pitch, rhythm, singing skill, overall listening feeling, etc.
[0147] Dry sound: pure human voice without accompaniment and post-processing.
[0148] Secondly, the inventive purpose of the exemplary embodiments of this disclosure will be explained and described. Specifically, the music work classification method provided by the exemplary embodiments of this disclosure, when sufficient data is available, can effectively improve the machine's evaluation ability by referring to the subjective evaluations of listeners. This allows for the evaluation of performances based on both objective and subjective evaluations, thereby improving the accuracy of the obtained work classification results. Simultaneously, it can also accurately separate karaoke works of good, average, and poor performance quality based on existing karaoke scoring system data and song popularity data.
[0149] Furthermore, the first and second music work classification models described in the exemplary embodiments of this disclosure will be explained and described. Specifically, the first and second music work classification models described in the exemplary embodiments of this disclosure are obtained by training a first network model and a second network model using original audio data; wherein, the first and second network models described here may include binary classification models; and the binary classification model may include any one of the following: Logistic Regression model, k-Nearest Neighbors model, Decision Trees model, Support Vector Machine (SVM) model, Naive Bayes model, and Random Forest model. Further, taking the Random Forest model as an example of a binary classification model, its corresponding structural example diagram can be found in [reference needed]. Figure 2 As shown. Furthermore, the first music work classification model described herein can be used to classify audio data with a third category, and the second music work classification model is used to classify audio data with a first category and / or audio data with a second category; wherein, the first category is the poor category, the second category is the average category, and the third category is the excellent category. That is, in practical applications, the first music work classification model can be used to extract music works of the poor category first, and then the second music work classification model can be used to extract music works of the average and excellent categories.
[0150] In an example embodiment, the music works with the excellent category described above can be referred to as good (A, Awesome) works; wherein the good works can be embodied in the following aspects: on the one hand, the works have excellent pitch level or only a few notes have slight deviation but do not destroy the overall listening; on the other hand, the singing skills of the works have no obvious defects, and the works meeting the two conditions are referred to as good works; further, the music works with the general category described above can be referred to as mediocre (M, Mediocre) works; wherein the mediocre works can be embodied in the following aspects: on the one hand, the works have excellent pitch level or only a few notes have slight deviation but do not destroy the overall listening, on the other hand, the singing skills of the works have some defects affecting the overall musicality, and this kind of works can also be referred to as works with accurate singing but general; further, the music works with the inferior (I, Inferior) category described above can be embodied in the following aspects: on the one hand, the works have some notes out of tune seriously or almost no pitch concept; on the other hand, the singing skills of the works have great defects destroying the overall musicality, and this kind of works can be referred to as bad works.
[0151] In an example embodiment, referring to Figure 3 , the specific training process of the first music work classification model and the second music work classification model can include the following steps:
[0152] Step S310, obtaining original audio data, and constructing a first category data set, a second category data set and a third category data set according to original labels of the original audio data.
[0153] In the example embodiment, first, the original audio data is obtained; wherein the original audio data described herein can be a singing quality evaluation data set with a plurality of dry tone audio samples (for example, 16748); wherein the labels of the data set have three levels, which are determined by two dimensions of pitch and singing skills; the three levels of labels are good, medium and poor; meanwhile, the data set can also include an artificial annotation data set with a size of 892, and the sample number distribution of the three categories can be shown in Table 1 as follows:
[0154] Table 1 Sample data distribution of small artificial annotation data set
[0155] Category A M I Sample size 334 307 251
[0156] Based on this, the annotated data set can be separated from the original audio data, wherein the annotated data set can constitute a 892x61 labeled feature matrix, and the data to be labeled constitutes a 15856x60 feature matrix. Then, according to the original labels of each dry tone in the annotated data set, a first category (poor) data set, a second category (general) data set, and a third category (excellent) data set are constructed, and then model training is performed based on the first category data set, the second category data set, and the third category data set.
[0157] In step S320, the second category data set and the third category data set are merged to obtain a total data set, and the first network model is trained based on the first category data set and the total data set to obtain a first music work classification model.
[0158] In the example embodiment, first, the second category data set and the third category data set are merged to obtain a total data set. The reason for merging the second category data set and the third category data set is that the first network model described in the example embodiment of the present disclosure is a binary classification model and cannot classify excellent, general, and poor music works at one time. Therefore, by merging, poor music works can be extracted first, and then another binary classification model is used to classify excellent and general works, thereby achieving the purpose of the present application. At the same time, the music work classification method described in the example embodiment of the present disclosure achieves the effect of classifying three different categories of music works based on two binary classification models, which can improve the classification efficiency and model efficiency on the basis of improving the accuracy of the obtained classification results, and reduce the system burden. Moreover, since the technical problem solved by the technical solution is a three-classification problem, but the difference between the I-class samples and the A / M-class samples is large, the problem is decomposed into two simpler binary classification problems, and two binary classification models are used to achieve better comprehensive effect.
[0159] Secondly, the first network model is trained according to the first category data set and the total data set, and then the first music work classification model is obtained. Specifically, the following method can be used: first, a third feature matrix of the original audio data included in the first category data set and the original audio data included in the total data set is extracted; secondly, the third feature matrix is subjected to feature screening to obtain a fourth feature matrix, and the fourth feature matrix is input into the first network model to obtain a first prediction result; then, according to the first prediction result and the original label possessed by the original audio data, a first harmonic mean of the original audio data included in the first category data set and a second harmonic mean of the original audio data included in the total data set are calculated; finally, the hyperparameters of the first network model are adjusted under the condition that the first harmonic mean is greater than the second harmonic mean as the hyperparameter tuning condition of the first network model, and the first music work classification model is obtained.
[0160] In the following, the specific training process of the first music work classification model will be further explained and described. Specifically, first, feature extraction is performed; the specific feature extraction process can include: ① the K song scoring system scores each sentence of lyrics in four dimensions of pitch, rhythm, breath, and emotion, and also counts the number of slides and vibratos in the sentence, and gives a comprehensive score for the whole sentence on this basis; ② score the work according to the score of the sentence; wherein the final score of the work, as well as the overall pitch, rhythm, breath, and emotion scores, are all directly available features; however, the number of slides and vibratos has a strong correlation with the number of sentences of the work, but the number of sentences of the work is irrelevant to the singer's singing ability and does not need to be introduced, so the average number of slides per sentence and the average number of vibratos per sentence are counted to utilize the two features of slides and vibratos, and then a comprehensive objective evaluation feature can be composed; further, in order to more comprehensively utilize the single sentence scoring information, the frequency distribution of each score group in the song can be calculated by taking every 10 points as a left-closed right-open interval and 100 points as a separate group, forming an 11-dimensional feature, such as Figure 4 and Figure 5The same operation is performed on the intonation, rhythm and breath of each sentence, and a 44-dimensional distribution feature is formed for each work. In addition to the objective score of the K-song scoring system, the subjective evaluation of the user can also be used. In practical application, the number of likes, gifts, shares, plays and comments of the corresponding work can be pulled, as well as the heat score of the work. Considering that the scores of other works of the same singer can also reflect the singing level of the singer, the average number of likes, gifts, shares, plays and comments, and the average heat score of all works of the same singer can also be used as the feature of the work. After data normalization, a feature vector with a total of 73 dimensions can be obtained for each work, including 14-dimensional comprehensive objective evaluation features, 44-dimensional distribution features of objective evaluation, 12-dimensional subjective evaluation features based on heat, and 3-dimensional other features.
[0161] Secondly, feature screening is performed. The standard for feature screening is that features with a correlation coefficient less than a preset threshold are deleted, and thus a fourth feature matrix after screening can be obtained. Then, the fourth feature matrix is input into a random forest model, and a first prediction result (which can be a 15856x1 prediction label vector) can be obtained. Finally, the random forest model is trained based on the first prediction result and the original label, and a first music work classification model can be obtained. It should be noted that, since the number of samples in the two categories "A / M class samples" and "I class samples" obtained by classifying the small data set is quite different, in order to avoid the classifier tending to classify the samples as "A / M class", the f1-score (first harmonic mean and second harmonic mean) of "I class" can be maximized to perform the hyperparameter tuning process, and thus the accuracy of the obtained first music work classification model can be improved.
[0162] Step S330, the second network model is trained according to the second category data set and the third category data set, and a second music work classification model is obtained.
[0163] Specifically, the second network model is trained according to the second category data set and the third category data set to obtain the second music work classification model, which can be realized in the following manner: first, a fifth feature matrix of the original audio data included in the second category data and a sixth feature matrix of the original audio data included in the third category data are extracted; second, the fifth feature matrix and the sixth feature matrix are subjected to feature screening to obtain a seventh feature matrix and an eighth feature matrix, and the seventh feature matrix and the eighth feature matrix are input into the second network model to obtain a second prediction result; then, according to the second prediction result and original labels possessed by the original audio data, a third harmonic mean of the original audio data included in the second category data set and a fourth harmonic mean of the original audio data included in the third category data set are calculated; finally, the hyperparameters of the second network model are adjusted under the condition that the fourth harmonic mean is greater than the third harmonic mean to obtain the second music work classification model.
[0164] Hereinafter, the specific training process of the second music work classification model will be explained and described. First, feature extraction and feature screening are performed to obtain the seventh feature matrix and the eighth feature matrix; wherein the specific implementation process of feature extraction and feature screening is similar to the extraction and screening process of the fourth feature matrix described above, and will not be further described here; second, the seventh feature matrix and the eighth feature matrix are input into the random forest model to obtain a second prediction result (the second prediction result can be a prediction label vector of 14401x1); and the random forest model is trained based on the second prediction result and the original label to obtain the second music work classification model. It should be noted here that in the process of training the second music work classification model, first, the small data set manually annotated and the to-be-labeled data are separated, and the class I samples in the small data set and the class I samples labeled by the model in the to-be-labeled data are removed. The complete feature matrix has a size of 15042x37, wherein the A class and the M class of the small data set form a 641x38 labeled feature matrix, and the to-be-labeled data form a 14401x37 feature matrix; at the same time, the small data set after removal is used to divide the training set and the test set for hyperparameter optimization of the classifier; further, since the precision (Precision) and the recall (Recall) of the A class are more concerned in actual application, the f1-score (third harmonic mean and fourth harmonic mean) of the A class is maximized to perform the hyperparameter optimization process, thereby improving the accuracy of the second music work classification model.
[0165] Hereinafter, the specific training process of the second music work classification model will be explained and described. First, feature extraction and feature screening are performed to obtain the seventh feature matrix and the eighth feature matrix; wherein the specific implementation process of feature extraction and feature screening is similar to the extraction and screening process of the fourth feature matrix described above, and will not be further described here; second, the seventh feature matrix and the eighth feature matrix are input into the random forest model to obtain a second prediction result (the second prediction result can be a prediction label vector of 14401x1); and the random forest model is trained based on the second prediction result and the original label to obtain the second music work classification model. It should be noted here that in the process of training the second music work classification model, first, the small data set manually annotated and the to-be-labeled data are separated, and the class I samples in the small data set and the class I samples labeled by the model in the to-be-labeled data are removed. The complete feature matrix has a size of 15042x37, wherein the A class and the M class of the small data set form a 641x38 labeled feature matrix, and the to-be-labeled data form a 14401x37 feature matrix; at the same time, the small data set after removal is used to divide the training set and the test set for hyperparameter optimization of the classifier; further, since the precision (Precision) and the recall (Recall) of the A class are more concerned in actual application, the f1-score (third harmonic mean and fourth harmonic mean) of the A class is maximized to perform the hyperparameter optimization process, thereby improving the accuracy of the second music work classification model. Figures 2-5 to Figure 1The classification method of the music works shown in the middle is further explained and described. Specifically:
[0166] In step S110, a first music work to be classified and first work attribute information of the first music work to be classified are acquired, and a first feature matrix is constructed according to the first work attribute information.
[0167] In the example embodiment, first, a first music work to be classified and first work attribute information of the first music work to be classified are acquired. The first music work to be classified can be from a corresponding music application platform or other video application platform, and the like, which is not specially limited in the example. Meanwhile, the first work attribute information can include first-dimension work attribute information, second-dimension work attribute information, third-dimension work attribute information, and fourth-dimension work attribute information, and the like. The first-dimension work attribute information includes objective evaluation information, the second-dimension work attribute information includes distribution information of the objective evaluation information, the third-dimension work attribute information includes subjective evaluation information, and the fourth-dimension work attribute information includes work quantity information.
[0168] In an example embodiment, the objective evaluation information can include singing pitch information, singing rhythm information, singing breath information, and singing emotion information. That is, the singer of the first music work to be classified has the pitch, rhythm, breath, and emotion when singing the first music work to be classified. Meanwhile, the distribution information of the objective information can include overall scores of each sentence, pitch scores, rhythm scores, emotion scores, and breath scores, and average and median of the pitch scores, rhythm scores, emotion scores, and breath scores. Meanwhile, the distribution information of the objective information also includes the number of glissando and vibrato. Further, the subjective evaluation information can include work popularity information, work like information, gift information, work sharing information, work playing information, and work comment information of the music work to be classified. That is, the subjective evaluation information is the popularity value, like value, gift quantity, sharing times, playing times, and comment quantity corresponding to the first music work to be classified, and the like. Further, the work quantity information can include all work quantity published by the singer of the first music work to be classified on all music application platforms and / or video application platforms, deleted work quantity after publishing, and remaining work quantity after deleting, and the like.
[0169] In an example embodiment, referring to Figure 6 As shown, constructing a first feature matrix according to the first work attribute information can include the following steps:
[0170] Step S610, calculating a first dimension sub-feature according to the objective evaluation information, and calculating a second dimension sub-feature according to distribution information of the objective evaluation information.
[0171] In the example embodiment, first, a first dimension sub-feature is calculated according to the objective evaluation information; specifically, the calculation can be implemented in the following manner: first, the overall sentence score of each sentence included in the first music work to be classified, and the average and median of the objective evaluation information are calculated according to the objective evaluation information of each sentence; second, the average number of slides and the average number of trills of each sentence are calculated according to the number of slides and the number of trills of each sentence; then, the first dimension sub-feature is generated according to the overall sentence score, the average and median of the objective evaluation information, and the average number of slides and the average number of trills of each sentence.
[0172] In an example embodiment, the overall sentence score of each sentence included in the first music work to be classified, and the average and median of the objective evaluation information can be calculated in the following manner: first, the overall sentence score of each sentence is calculated according to the pitch score in the singing pitch dimension, the rhythm score in the singing rhythm dimension, the breath score in the singing breath dimension, and the emotion score in the singing emotion dimension of each sentence included in the first music work to be classified; second, the average number of slides and the average number of trills of each sentence are calculated according to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence; then, the median of the pitch score, the median of the rhythm score, the median of the breath score, and the median of the emotion score are calculated according to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence.
[0173] In the following, the specific calculation process of the first dimension sub-feature will be further explained. Specifically, first, the scores of each sentence of lyrics in the four dimensions of pitch, rhythm, breath, and emotion given by a preset song scoring system can be obtained; then, the number of slides and the number of trills in the sentence of lyrics can be obtained; at the same time, the overall score of the sentence is obtained based on the pitch, rhythm, breath, emotion, and slides and trills, and the comprehensive score is obtained; further, the median and average of the pitch, rhythm, breath, and emotion are calculated, and then the first dimension sub-feature is generated according to the overall sentence score, the average and median of the pitch, rhythm, breath, and emotion, and the average number of slides and the average number of trills of each sentence.
[0174] Secondly, the second-dimension sub-feature is calculated according to the distribution information of the objective evaluation information. Specifically, the following method can be used: firstly, the intonation interval, rhythm interval, breath interval and emotion interval to which each sentence belongs are determined according to the intonation score, rhythm score, breath score and emotion score of each sentence; secondly, the intonation interval probability distribution, rhythm interval probability distribution, breath interval probability distribution and emotion interval probability distribution of each sentence are calculated according to the intonation interval, rhythm interval, breath interval and emotion interval to which each sentence belongs; and then, the second-dimension sub-feature is calculated according to the intonation interval probability distribution, rhythm interval probability distribution, breath interval probability distribution and emotion interval probability distribution. The obtained second-dimension sub-feature can refer to the following formula: Figure 4
[0175] Step S620, the third-dimension sub-feature is calculated according to the subjective evaluation information, and the fourth-dimension sub-feature is calculated according to the work quantity information.
[0176] In the example embodiment, firstly, the third-dimension sub-feature is calculated according to the subjective evaluation information. Specifically, the following method can be used: firstly, the singer of the first music work to be classified is taken to obtain other music works except the first music work to be classified; and secondly, the third-dimension sub-feature is generated according to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works. The third-dimension sub-feature can be generated according to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works by the following method: firstly, the work popularity information, work like information, gift information, work sharing information, work playing information and work comment information of the other music works are obtained, and the work popularity average value, work like quantity average value, gift quantity average value, work sharing times average value, work playing times average value and work comment quantity average value of the other music works are calculated; and secondly, the third-dimension sub-feature is generated according to the work popularity information, work like information, gift information, work sharing information, work playing information and work comment information of the first music work to be classified and the work popularity average value, work like quantity average value, gift quantity average value, work sharing times average value, work playing times average value and work comment quantity average value of the other music works.
[0177] Secondly, the fourth dimension sub-feature is calculated based on the number of works. Specifically, this can be achieved as follows: First, obtain the number of published works and the number of published but deleted works of the singer of the first music work to be classified; second, determine the number of published but not deleted works based on the number of published and deleted works, and generate the fourth dimension sub-feature based on the number of published, deleted, and not deleted works.
[0178] Step S630: Construct the first feature matrix based on the first dimension sub-feature, the second dimension sub-feature, the third dimension sub-feature, and the fourth dimension sub-feature.
[0179] Specifically, after obtaining the first, second, third, and fourth dimension sub-features, normalization can be performed on them. Then, the normalized first, second, third, and fourth dimension sub-features are combined to obtain the first feature matrix. This first feature matrix can be a 73-dimensional feature vector. This 73-dimensional feature vector can include 14 dimensions of comprehensive objective evaluation features (first dimension sub-features), 44 dimensions of objective price distribution features (second dimension sub-features), 12 dimensions of subjective evaluation features based on popularity (third dimension sub-features), and 3 dimensions of other features (fourth dimension sub-features). The size of the resulting first feature matrix can be 16748 * 73. Further details on the specific distribution of the first feature matrix can be found in [reference needed]. Figure 7 As shown.
[0180] Among them, Figure 7 The first feature matrix shown includes the following dimensions: the first dimension of the 14-dimensional sub-features: overall sentence score (1-dimensional) + scores for pitch, rhythm, breath, and emotion (4-dimensional) + mean and median of pitch, rhythm, breath, and emotion (8-dimensional) + number of glissando and vibrato (1-dimensional); the second dimension of the 44-dimensional sub-features includes: distribution of overall sentence score (11-dimensional) + distribution of pitch, rhythm, and breath scores (33-dimensional); the third dimension of the 12-dimensional sub-features includes: number of popularity, likes, gifts, shares, plays, and comments (6-dimensional) + average of popularity, likes, gifts, shares, plays, and comments for all songs (6-dimensional); and the fourth dimension of the 3-dimensional sub-features includes: number of works published by the singer (1-dimensional) + number of works deleted after publication (1-dimensional) + number of works remaining after deletion (1-dimensional).
[0181] In step S120, feature screening is performed on the first feature matrix to obtain a second feature matrix, and the second feature matrix is input into a first music work classification model to obtain a first output result.
[0182] In the example embodiment, first, feature screening is performed on the first feature matrix to obtain a second feature matrix. Specifically, the following manner can be used: the correlation coefficients between the sub-features of different dimensions included in the first feature matrix are calculated, and the sub-features of different dimensions included in the first feature matrix are screened according to the correlation coefficients to obtain the second feature matrix. Specifically, the Pearson correlation coefficients between the sub-features of different dimensions can be calculated; wherein the obtained Pearson correlation coefficients can refer to the Pearson correlation coefficients shown in the following table. Figure 8
[0183] In Figure 8 Among the Pearson correlation coefficients shown, the following sub-features (arranged in order from top to bottom and left to right) can be included: class (category, irrelevant to the present application), value (overall score of each sentence), note (pitch), rhythm (rhythm), breath (breath), enthusiasm (emotion), average_score (average of each sentence), median_score (median of each sentence), median_note (median of pitch), median_rhythm (median of rhythm), median_breath (median of breath), average_portamento (average of portamento), average_vibrato (average of vibrato), zan_user_num (number of user likes), gift_num (number of gifts), share_user_num (number of user shares), effective_play_num (number of work plays), comment_user_num (number of user comments), heat_score (work heat score), num_opus (total number of works), avg_heat (average heat), avg_zan (average like), avg_gift (average gift), avg_share (average share), avg_play (average play), avg_comment (average comment), num_opus_ori (number of remaining works), avg_score (average score of sentences), avg_note (average note), avg_rhythm (average rhythm), avg_breath (average breath), avg_enthusiasm (average enthusiasm), avg_portamento (average portamento), avg_vibrato (average vibrato), and opus_deleted (number of deleted works). It should be noted here that Figure 8 The specific sub-features shown in the above are represented in English, and their specific meanings have been exemplified here. In actual application, other sub-features or other meanings can also be set according to actual needs, and the present example does not specially limit this.
[0184] Meanwhile, in the process of feature screening, the specific screening criteria are:
[0185] |Coef i |≥0.1;
[0186] where Coef i is the correlation coefficient (Pearson correlation coefficient) of the i-th feature and the classification category, and has: -1≤Coef i≤1, i = 1, 2, …, 73; that is, features whose absolute values of Pearson correlation coefficients are between [-0.1, 0.1] need to be deleted; it should be noted here that the first music work classification model combines the A and M classes, and uses the I and A / M classes to calculate the correlation coefficients; the second music work classification model only uses the A and M classes to calculate the correlation coefficients, and deletes all model features that satisfy: -0.1 < Coef i <0.1. In a specific use process, the first music work classification model uses data that deletes 13 features, and the types and quantities of the remaining features are shown in Table 2 below; further, the second music work classification model uses data that deletes 36 features, and the types and quantities of the remaining features can be shown in Table 3 below; wherein, the size of the obtained second feature matrix is 16748x60; the size of the features input into the second music work classification model is 16748x37.
[0187] Table 2
[0188]
[0189] Table 3
[0190]
[0191] It should be noted here that in the example diagram of the Pearson correlation coefficient between the two features shown in FIG. 2, the size of the Pearson correlation coefficient is identified by different colors; the larger the absolute value of the Pearson correlation coefficient, the darker the color value of the Pearson correlation coefficient; the smaller the absolute value of the Pearson correlation coefficient, the lighter the color value of the Pearson correlation coefficient; that is, in the actual application process, different sizes of Pearson correlation coefficients can be identified based on different color identifications, thereby facilitating feature screening, thereby improving feature screening efficiency. Figure 8
[0192] At this point, the specific screening process of the first feature matrix has been completed. Next, the first feature matrix also needs to be input into the first music work classification model to obtain the first output result; wherein, in a specific application process, the specific processing process of the first music work classification model on the first feature matrix can include but is not limited to the following manner: calculating the first linear regression part of the first feature matrix in the first internal node of the first music work classification model using the first music work classification model; then, calculating the output value of the first linear regression part in the leaf node of the first music work classification model to obtain the first output result.
[0193] In step S130, music works with the first work category are extracted from the first music works to be classified according to the first output result, and the music works with the first work category are removed from the first music works to be classified, to obtain second music works to be classified.
[0194] In the example embodiment, first, music works with the first work category (i.e. poor music works) are extracted from the first music works to be classified according to the first output result; then, the second music works to be classified are determined. The difference between the first music works to be classified and the second music works to be classified is that the second music works to be classified only include half of the music works in the first music works to be classified and good music works, and there is no other difference.
[0195] In step S140, a third feature matrix corresponding to the second music works to be classified is input into the second music work classification model to obtain a second output result, and music works with the second work category and music works with the third work category are extracted from the second music works to be classified according to the second output result.
[0196] In the example embodiment, first, the second feature matrix corresponding to the second music works to be classified is determined; then, the second feature matrix corresponding to the second music works to be classified is subjected to feature screening, to obtain the third feature matrix; then, the third feature matrix is input into the second music work classification model, to obtain the second output result. The specific screening process of the third feature matrix is the same as the screening process of the first feature matrix, which will not be further described here.
[0197] Further, after obtaining the second output result, music works with the second work category and music works with the third work category can be determined according to the second output result. It should be noted that, since the second music work classification model is a binary classification model, and the second music work classification model only needs to classify general works and good works, the accuracy of the obtained output result can be improved on the basis of simplifying the classification task of the model.
[0198] It should be noted that the music work classification method described in the example embodiments of the present disclosure can not only be used in the scenario of classifying unclassified music works, but also can be used in labeling data sets; that is, when facing a large amount of data sets to be labeled, a part of them can be manually labeled as samples to train the corresponding model, and then the unmarked data sets are labeled based on the trained model; then the model is debugged again through the labeled data sets, so as to improve the accuracy of the obtained music work classification model.
[0199] So far, the music work classification method disclosed in the example embodiments of the present disclosure has been fully implemented. Based on the foregoing, it can be known that the music work classification method disclosed in the example embodiments of the present disclosure, on the one hand, completes the classification task of karaoke works with good singing quality, general singing quality and poor singing quality with high accuracy through computer algorithms, replacing manual singing quality evaluation of a large number of works; on the other hand, the algorithm uses existing karaoke system score data as objective features and uses the heat data of the karaoke platform as subjective evaluation features, fully utilizing various features that can assist in singing quality evaluation and fully exploiting the value of data; on the other hand, an effective feature screening method is used to reduce the amount of calculation while improving the accuracy of the algorithm.
[0200] Exemplary apparatus
[0201] After introducing the method of the example embodiments of the present disclosure, next, with reference to Figure 9 The music work classification device of the example embodiments of the present disclosure is explained and described. Specifically, referring to FIG. 9, Figure 9 The music work classification device can include a first feature matrix construction module 910, a first music work classification module 920, a music work elimination module 930 and a second music work classification module 940. Wherein:
[0202] The first feature matrix construction module 910 can be used to obtain the first to-be-classified music work and the first work attribute information of the first to-be-classified music work, and construct a first feature matrix according to the first work attribute information;
[0203] The first music work classification module 920 can be used to perform feature screening on the first feature matrix to obtain a second feature matrix, and input the second feature matrix into a first music work classification model to obtain a first data result;
[0204] The music work elimination module 930 can be used to extract music works with a first work category from the to-be-classified music works according to the first output result, and eliminate music works with the first work category from the first to-be-classified music works to obtain second to-be-classified music works;
[0205] The second music work classification module 940 can be used to input the second feature matrix corresponding to the second to-be-classified music work into a second music work classification model to obtain a second output result, and extract music works with a second work category and music works with a third work category from the second to-be-classified music works according to the second output result.
[0206] In an example embodiment of the present disclosure, the first feature matrix is constructed according to the first work attribute information, including: calculating a first dimension sub-feature according to the objective evaluation information, and calculating a second dimension sub-feature according to distribution information of the objective evaluation information; calculating a third dimension sub-feature according to the subjective evaluation information, and calculating a fourth dimension sub-feature according to the work quantity information; and constructing the first feature matrix according to the first dimension sub-feature, the second dimension sub-feature, the third dimension sub-feature, and the fourth dimension sub-feature.
[0207] In an example embodiment of the present disclosure, the first dimension sub-feature is calculated according to the objective evaluation information, including: calculating an overall sentence score of each sentence included in the first to-be-classified music work, and an average and a median of the objective evaluation information according to the objective evaluation information of each sentence; calculating a trill average and a vibrato average of each sentence according to a number of trills and a number of vibratos of each sentence; and generating the first dimension sub-feature according to the overall sentence score, the average and the median of the objective evaluation information, the trill average and the vibrato average of each sentence.
[0208] In an example embodiment of the present disclosure, the objective evaluation information includes at least one of singing pitch information, singing rhythm information, singing breath information, and singing emotion information; and wherein the overall sentence score of each sentence included in the first to-be-classified music work, and the average and the median of the objective evaluation information are calculated according to the objective evaluation information of each sentence, including: calculating an overall sentence score of each sentence according to a pitch score in a singing pitch dimension, a rhythm score in a singing rhythm dimension, a breath score in a singing breath dimension, and an emotion score in a singing emotion dimension of each sentence included in the first to-be-classified music work; calculating a pitch average of the singing pitch information, a rhythm average of the singing rhythm information, a breath average of the singing breath information, and an emotion average of the singing emotion information according to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence; and calculating a pitch median of the singing pitch information, a rhythm median of the singing rhythm information, a breath median of the singing breath information, and an emotion median of the singing emotion information according to the pitch score, the rhythm score, the breath score, and the emotion score of each sentence.
[0209] In an example embodiment of the present disclosure, the second-dimension sub-feature is calculated according to distribution information of the objective evaluation information, including: determining a pitch interval, a rhythm interval, a breath interval and an emotion interval to which each sentence belongs according to the pitch score, the rhythm score, the breath score and the emotion score of each sentence; calculating a pitch interval probability distribution, a rhythm interval probability distribution, a breath interval probability distribution and an emotion interval probability distribution of each sentence according to the pitch interval, the rhythm interval, the breath interval and the emotion interval to which each sentence belongs; and calculating the second-dimension sub-feature according to the pitch interval probability distribution, the rhythm interval probability distribution, the breath interval probability distribution and the emotion interval probability distribution.
[0210] In an example embodiment of the present disclosure, the third-dimension sub-feature is calculated according to the subjective evaluation information, including: obtaining other music works of the singer of the first music work to be classified except the first music work to be classified; and generating the third-dimension sub-feature according to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works.
[0211] In an example embodiment of the present disclosure, the subjective evaluation information includes at least one of work popularity information, work like information, gift information, work sharing information, work playing information and work comment information; and the third-dimension sub-feature is generated according to the subjective evaluation information of the first music work to be classified and the subjective evaluation information of the other music works, including: obtaining work popularity information, work like information, gift information, work sharing information, work playing information and work comment information of the other music works, and calculating work popularity average, work like quantity average, gift quantity average, work sharing times average, work playing times average and work comment quantity average of the other music works; and generating the third-dimension sub-feature according to the work popularity information, the work like information, the gift information, the work sharing information, the work playing information and the work comment information of the first music work to be classified and the work popularity average, the work like quantity average, the gift quantity average, the work sharing times average, the work playing times average and the work comment quantity average of the other music works.
[0212] In an example embodiment of the present disclosure, the fourth-dimension sub-feature is calculated according to the work quantity information, including: obtaining the number of published works and the number of works deleted after being published of the singer of the first music work to be classified; determining the number of published and not deleted works according to the number of published works and the number of works deleted after being published, and generating the fourth-dimension sub-feature according to the number of published works, the number of works deleted after being published and the number of published and not deleted works.
[0213] In an example embodiment of the present disclosure, the feature screening on the first feature matrix to obtain a second feature matrix comprises: calculating correlation coefficients between different dimension sub-features included in the first feature matrix, and performing feature screening on the different dimension sub-features included in the first feature matrix according to the correlation coefficients to obtain the second feature matrix.
[0214] In an example embodiment of the present disclosure, the music work classification device further comprises:
[0215] A data set construction module is configured to acquire original audio data, and construct a first category data set, a second category data set and a third category data set according to original labels of the original audio data.
[0216] A first model training module is configured to merge the second category data set and the third category data set to obtain a total data set, and train a first network model according to the first category data set and the total data set to obtain a first music work classification model.
[0217] A second model training module is configured to train a second network model according to the second category data set and the third category data set to obtain a second music work classification model.
[0218] In an example embodiment of the present disclosure, the first music work classification model is used for classifying audio data with a third category, and the second music work classification model is used for classifying audio data with a first category and / or audio data with a second category; the first category is a poor category, the second category is a general category, and the third category is an excellent category.
[0219] In an example embodiment of the present disclosure, the training of the first network model according to the first category data set and the total data set to obtain the first music work classification model comprises: extracting a third feature matrix of original audio data included in the first category data set and original audio data included in the total data set; performing feature screening on the third feature matrix to obtain a fourth feature matrix, and inputting the fourth feature matrix into the first network model to obtain a first prediction result; calculating a first harmonic mean of the original audio data included in the first category data set and a second harmonic mean of the original audio data included in the total data set according to the first prediction result and original labels of the original audio data; and adjusting hyperparameters of the first network model according to a condition that the first harmonic mean is greater than the second harmonic mean as a hyperparameter tuning condition of the first network model to obtain the first music work classification model.
[0220] In an example embodiment of the present disclosure, the second network model is trained according to the second category data set and the third category data set to obtain the second music work classification model, including: extracting a fifth feature matrix of the original audio data included in the second category data and a sixth feature matrix of the original audio data included in the third category data; performing feature screening on the fifth feature matrix and the sixth feature matrix to obtain a seventh feature matrix and an eighth feature matrix, and inputting the seventh feature matrix and the eighth feature matrix into the second network model to obtain a second prediction result; calculating a third harmonic mean of the original audio data included in the second category data set and a fourth harmonic mean of the original audio data included in the third category data set according to the second prediction result and original labels possessed by the original audio data; and adjusting the hyperparameters of the second network model to obtain the second music work classification model, with the fourth harmonic mean being greater than the third harmonic mean as a hyperparameter tuning condition for the second network model.
[0221] Exemplary storage medium
[0222] After introducing the music work classification method and the music work classification device of the example embodiment of the present disclosure, next, with reference to Figure 10 The storage medium of the example embodiment of the present disclosure is described.
[0223] With reference to Figure 10 The program product 1000 for implementing the above method according to the embodiment of the present disclosure is described, which can adopt a portable compact disc read-only memory (CD-ROM) and include program codes, and can run on a terminal device such as a personal computer. However, the program product of the present disclosure is not limited thereto.
[0224] The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium may, for example, be but is not limited to an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples (non-exhaustive list) of the readable storage medium include an electrical connection having one or more wires, a portable disc, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0225] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A readable signal medium may also be any readable medium other than a readable storage medium.
[0226] Program code for performing the operations of this disclosure can be written in any combination of one or more programming languages, including object-oriented programming languages such as Java and C++, and conventional procedural programming languages such as C or similar languages. The program code can execute entirely on the user's computing device, partially on the user's computing device, partially on a remote computing device, or entirely on a remote computing device or server. In cases involving remote computing devices, the remote computing devices can be connected to the user's computing device via any type of network, including a local area network (LAN) or a wide area network (WAN).
[0227] Exemplary electronic device
[0228] Having described the storage medium of exemplary embodiments of this disclosure, the following references are made. Figure 11 An electronic device according to an exemplary embodiment of the present disclosure will be described.
[0229] Figure 11 The electronic device 1100 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0230] like Figure 11 As shown, the electronic device 1100 is manifested in the form of a general-purpose computing device. The components of the electronic device 1100 may include, but are not limited to: at least one processing unit 1110, at least one storage unit 1120, a bus 1130 connecting different system components (including storage unit 1120 and processing unit 1110), and a display unit 1140.
[0231] The storage unit 1120 stores program code that can be executed by the processing unit 1110, causing the processing unit 1110 to perform the steps described in the "Exemplary Methods" section of this specification according to various exemplary embodiments of this disclosure. For example, the processing unit 1110 can perform actions such as... Figure 1 Steps S110-S140 are shown in the diagram.
[0232] The storage unit 1120 can include a volatile storage unit such as a random access memory (RAM) 11201 and / or a cache memory 11202, and also a non-volatile storage, e.g., a read-only memory (ROM) 11203.
[0233] The storage unit 1120 can also include a program / utility 11204 having a set (at least one) of program modules 11205, including but not limited to an operating system, one or more application programs, other program modules, and program data, each of which can include implementation of a network environment, alone or in combination.
[0234] The bus 1130 can include a data bus, an address bus, and a control bus.
[0235] The electronic device 1100 can also communicate with one or more external devices 800 such as a keyboard or pointing device, through the I / O interface 1150. Additionally, the electronic device 1100 can communicate with one or more networks, such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet, through a network adapter 1160. As illustrated, the network adapter 1160 communicates with the other modules of the electronic device 1100 through the bus 1130. It should be appreciated that although not shown, other hardware and / or software modules could be used in conjunction with the electronic device 1100. For example, microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data archival storage systems, etc., can be used with the electronic device 1100.
[0236] It should be noted that although several modules or sub-modules of the pop-up processing apparatus are mentioned in the foregoing detailed description, such a division is merely exemplary and not mandatory. Indeed, according to embodiments of the present disclosure, features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, features and functions of one unit / module described above can be further divided into a plurality of units / modules.
[0237] It should be noted that although several units / modules or sub-units / modules of the apparatus are mentioned in the foregoing detailed description, such a division is merely exemplary and not mandatory. Indeed, according to embodiments of the present disclosure, features and functions of two or more units / modules described above can be embodied in one unit / module. Conversely, features and functions of one unit / module described above can be further divided into a plurality of units / modules.
[0238] Furthermore, although the operations of the methods of the present disclosure are described in a particular, sequential order, this should not be understood as a requirement or implied that the operations be performed in anything less than the order described, and / or that all be performed, to achieve desirable results. Additionally or alternatively, certain steps can be omitted, combined, performed simultaneously, and / or performed separately from other steps disclosed.
[0239] While the spirit and principles of the present disclosure have been described with reference to several specific implementations, it is to be understood that the present disclosure is not limited to the specific implementations disclosed and that the division into aspects is not meant to imply that features from one aspect cannot be combined with features from another aspect to benefit, but is merely for ease of presentation. The present disclosure is intended to cover various modifications and equivalent arrangements included within the spirit and scope of the claims appended hereto.
Claims
1. A method for classifying musical works, including: Obtain a first musical work to be classified and its first work attribute information. Based on the first work attribute information, construct a first feature matrix, including: calculating a first-dimensional sub-feature based on objective evaluation information, and calculating a second-dimensional sub-feature based on the distribution information of the objective evaluation information; calculating a third-dimensional sub-feature based on subjective evaluation information, and calculating a fourth-dimensional sub-feature based on the number of works; and constructing the first feature matrix based on the first-dimensional sub-feature, the second-dimensional sub-feature, the third-dimensional sub-feature, and the fourth-dimensional sub-feature. The first feature matrix is subjected to feature filtering to obtain a second feature matrix, and the second feature matrix is input into the first music work classification model to obtain the first output result; Based on the first output result, extract musical works with the first work category from the first musical works to be classified, and remove musical works with the first work category from the first musical works to be classified to obtain the second musical works to be classified. The third feature matrix corresponding to the second music work to be classified is input into the second music work classification model to obtain the second output result. Based on the second output result, music works with the second work category and music works with the third work category are extracted from the second music work to be classified.
2. The method for classifying musical works according to claim 1, wherein, The first work attribute information includes at least one of the following: first-dimensional work attribute information, second-dimensional work attribute information, third-dimensional work attribute information, and fourth-dimensional work attribute information; The first dimension of work attribute information includes objective evaluation information; the second dimension includes the distribution information of objective evaluation information; the third dimension includes subjective evaluation information; and the fourth dimension includes the number of works.
3. The method for classifying musical works according to claim 1, wherein, The first dimension sub-features are calculated based on the objective evaluation information, including: Based on the objective evaluation information of each sentence included in the first musical work to be classified, calculate the overall sentence score of each sentence, as well as the mean and median of the objective evaluation information; Calculate the average number of glissando and vibrato for each sentence based on the number of glissando and vibrato in each sentence. The first dimension sub-features are generated based on the overall sentence score, the mean and median of objective evaluation information, the mean number of glissando and the mean number of vibrato in each sentence.
4. The method for classifying musical works according to claim 3, wherein, The objective evaluation information includes at least one of the following: singing pitch information, singing rhythm information, singing breath information, and singing emotion information; Specifically, based on the objective evaluation information of each sentence included in the first musical work to be classified, the overall sentence score of each sentence, as well as the mean and median of the objective evaluation information, are calculated, including: Based on the pitch score, rhythm score, breath score, and emotion score of each sentence in the first musical work to be classified, calculate the overall sentence score for each sentence. The average pitch score, average rhythm score, average breath score, and average emotion score of each sentence are calculated based on the pitch score, rhythm score, breath score, and emotion score of the singing. Based on the pitch score, rhythm score, breath score, and emotion score of each sentence, calculate the median pitch of the singing pitch information, the median rhythm of the singing rhythm information, the median breath of the singing breath information, and the median emotion of the singing emotion information.
5. The method for classifying musical works according to claim 1, wherein, The second dimension sub-feature is calculated based on the distribution information of the objective evaluation information, including: Based on the pitch score, rhythm score, breath score, and emotion score of each sentence, determine the pitch range, rhythm range, breath range, and emotion range to which each sentence belongs. Based on the pitch interval, rhythm interval, breath interval, and emotional interval to which each sentence belongs, calculate the probability distribution of the pitch interval, rhythm interval, breath interval, and emotional interval for each sentence. The second dimension sub-feature is calculated based on the probability distribution of pitch interval, rhythm interval, breath interval, and emotion interval.
6. The method for classifying musical works according to claim 1, wherein, The third-dimensional sub-features are calculated based on the subjective evaluation information, including: Obtain other musical works by the singer of the first musical work to be classified, excluding the first musical work to be classified. The third dimension sub-feature is generated based on the subjective evaluation information of the first musical work to be classified and the subjective evaluation information of the other musical works.
7. The method for classifying musical works according to claim 6, wherein, The subjective evaluation information includes at least one of the following: work popularity information, work like information, gift information, work sharing information, work playback information, and work comment information; The third-dimensional sub-feature is generated based on the subjective evaluation information of the first musical work to be classified and the subjective evaluation information of the other musical works, including: Obtain information on the popularity, likes, gifts, shares, plays, and comments of other music works, and calculate the average popularity, number of likes, number of gifts, number of shares, number of plays, and number of comments for other music works. The third-dimensional sub-feature is generated based on the popularity information, likes information, gift information, sharing information, playback information, and comment information of the first music work to be classified, as well as the average popularity, number of likes, number of gifts, number of shares, number of plays, and number of comments of other music works.
8. The method for classifying musical works according to claim 1, wherein, The fourth dimension sub-feature is calculated based on the quantity information of the works, including: Obtain the number of published works and the number of deleted works by the singer of the first music work to be classified; Based on the number of published works and the number of works deleted after publication, the number of published but not deleted works is determined, and the fourth dimension sub-feature is generated based on the number of published works, the number of works deleted after publication, and the number of published but not deleted works.
9. The method for classifying musical works according to claim 1, wherein, The first feature matrix is subjected to feature filtering to obtain a second feature matrix, including: Calculate the correlation coefficients between the sub-features of different dimensions included in the first feature matrix, and perform feature filtering on the sub-features of different dimensions included in the first feature matrix based on the correlation coefficients to obtain the second feature matrix.
10. The method for classifying musical works according to claim 1, wherein, The classification method for musical works also includes: Obtain the raw audio data, and construct a first category dataset, a second category dataset, and a third category dataset based on the original labels of the raw audio data; The second category dataset and the third category dataset are merged to obtain the total dataset. The first network model is then trained based on the first category dataset and the total dataset to obtain the first music work classification model. The second network model is trained using the second and third category datasets to obtain the second music work classification model.
11. The method for classifying musical works according to claim 10, wherein, The first music work classification model is used to classify audio data with a third category, and the second music work classification model is used to classify audio data with a first category and / or audio data with a second category. The first category is the poor category, the second category is the average category, and the third category is the excellent category.
12. The method for classifying musical works according to claim 10, wherein, The first network model is trained based on the first category dataset and the total dataset to obtain the first music work classification model, including: Extract the third feature matrix of the original audio data included in the first category dataset and the original audio data included in the total dataset; The third feature matrix is subjected to feature filtering to obtain a fourth feature matrix, and the fourth feature matrix is input into the first network model to obtain a first prediction result; Based on the first prediction result and the original labels of the original audio data, calculate the first harmonic mean of the original audio data included in the first category dataset and the second harmonic mean of the original audio data included in the total dataset. The hyperparameters of the first network model are adjusted based on the condition that the first harmonic mean is greater than the second harmonic mean, thus obtaining the first music work classification model.
13. The method for classifying musical works according to claim 10, wherein, The second network model is trained using the second and third category datasets to obtain the second music work classification model, including: Extract the fifth feature matrix of the original audio data included in the second category of data and the sixth feature matrix of the original audio data included in the third category of data; The fifth and sixth feature matrices are subjected to feature filtering to obtain the seventh and eighth feature matrices, and the seventh and eighth feature matrices are input into the second network model to obtain the second prediction result; Based on the second prediction result and the original labels of the original audio data, calculate the third harmonic mean of the original audio data included in the second category dataset and the fourth harmonic mean of the original audio data included in the third category dataset. The hyperparameters of the second network model are adjusted by using the condition that the fourth harmonic mean is greater than the third harmonic mean, thus obtaining the second music work classification model.
14. A musical work classification device, comprising: The first feature matrix construction module is used to obtain the first musical work to be classified and the first work attribute information of the first musical work to be classified, and to construct the first feature matrix based on the first work attribute information, including: calculating the first dimension sub-feature based on objective evaluation information, and calculating the second dimension sub-feature based on the distribution information of objective evaluation information; calculating the third dimension sub-feature based on subjective evaluation information, and calculating the fourth dimension sub-feature based on the number of works information; and constructing the first feature matrix based on the first dimension sub-feature, the second dimension sub-feature, the third dimension sub-feature, and the fourth dimension sub-feature. The first music work classification module is used to perform feature filtering on the first feature matrix to obtain a second feature matrix, and input the second feature matrix into the first music work classification model to obtain the first data result; The music work elimination module is used to extract music works with a first work category from the music works to be classified according to the first output result, and to eliminate music works with the first work category from the first music works to be classified to obtain a second music work to be classified. The second music work classification module is used to input the third feature matrix corresponding to the second music work to be classified into the second music work classification model to obtain the second output result, and extract music works with the second work category and music works with the third work category from the second music work to be classified based on the second output result.
15. The musical work classification device according to claim 14, wherein, The first work attribute information includes at least one of the following: first-dimensional work attribute information, second-dimensional work attribute information, third-dimensional work attribute information, and fourth-dimensional work attribute information; The first dimension of work attribute information includes objective evaluation information; the second dimension includes the distribution information of objective evaluation information; the third dimension includes subjective evaluation information; and the fourth dimension includes the number of works.
16. The musical work classification device according to claim 14, wherein, The first dimension sub-features are calculated based on the objective evaluation information, including: Based on the objective evaluation information of each sentence included in the first musical work to be classified, calculate the overall sentence score of each sentence, as well as the mean and median of the objective evaluation information; Calculate the average number of glissando and vibrato for each sentence based on the number of glissando and vibrato in each sentence. The first dimension sub-features are generated based on the overall sentence score, the mean and median of objective evaluation information, the mean number of glissando and the mean number of vibrato in each sentence.
17. The musical work classification device according to claim 16, wherein, The objective evaluation information includes at least one of the following: singing pitch information, singing rhythm information, singing breath information, and singing emotion information; Specifically, based on the objective evaluation information of each sentence included in the first musical work to be classified, the overall sentence score of each sentence, as well as the mean and median of the objective evaluation information, are calculated, including: Based on the pitch score, rhythm score, breath score, and emotion score of each sentence in the first musical work to be classified, calculate the overall sentence score for each sentence. The average pitch score, average rhythm score, average breath score, and average emotion score of each sentence are calculated based on the pitch score, rhythm score, breath score, and emotion score of the singing. Based on the pitch score, rhythm score, breath score, and emotion score of each sentence, calculate the median pitch of the singing pitch information, the median rhythm of the singing rhythm information, the median breath of the singing breath information, and the median emotion of the singing emotion information.
18. The musical work classification device according to claim 14, wherein, The second dimension sub-feature is calculated based on the distribution information of the objective evaluation information, including: Based on the pitch score, rhythm score, breath score, and emotion score of each sentence, determine the pitch range, rhythm range, breath range, and emotion range to which each sentence belongs. Based on the pitch interval, rhythm interval, breath interval, and emotional interval to which each sentence belongs, calculate the probability distribution of the pitch interval, rhythm interval, breath interval, and emotional interval for each sentence. The second dimension sub-feature is calculated based on the probability distribution of pitch interval, rhythm interval, breath interval, and emotion interval.
19. The musical work classification device according to claim 14, wherein, The third-dimensional sub-features are calculated based on the subjective evaluation information, including: Obtain other musical works by the singer of the first musical work to be classified, excluding the first musical work to be classified. The third dimension sub-feature is generated based on the subjective evaluation information of the first musical work to be classified and the subjective evaluation information of the other musical works.
20. The musical work classification device according to claim 19, wherein, The subjective evaluation information includes at least one of the following: work popularity information, work like information, gift information, work sharing information, work playback information, and work comment information; The third-dimensional sub-feature is generated based on the subjective evaluation information of the first musical work to be classified and the subjective evaluation information of the other musical works, including: Obtain information on the popularity, likes, gifts, shares, plays, and comments of other music works, and calculate the average popularity, number of likes, number of gifts, number of shares, number of plays, and number of comments for other music works. The third-dimensional sub-feature is generated based on the popularity information, likes information, gift information, sharing information, playback information, and comment information of the first music work to be classified, as well as the average popularity, number of likes, number of gifts, number of shares, number of plays, and number of comments of other music works.
21. The musical work classification device according to claim 14, wherein, The fourth dimension sub-feature is calculated based on the quantity information of the works, including: Obtain the number of published works and the number of deleted works by the singer of the first music work to be classified; Based on the number of published works and the number of works deleted after publication, the number of published but not deleted works is determined, and the fourth dimension sub-feature is generated based on the number of published works, the number of works deleted after publication, and the number of published but not deleted works.
22. The musical work classification device according to claim 14, wherein, The first feature matrix is subjected to feature filtering to obtain a second feature matrix, including: Calculate the correlation coefficients between the sub-features of different dimensions included in the first feature matrix, and perform feature filtering on the sub-features of different dimensions included in the first feature matrix based on the correlation coefficients to obtain the second feature matrix.
23. The musical work classification device according to claim 14, wherein, The musical works classification device also includes: The dataset construction module is used to acquire raw audio data and construct a first category dataset, a second category dataset, and a third category dataset based on the original labels of the raw audio data. The first model training module is used to merge the second category dataset and the third category dataset to obtain the total dataset, and to train the first network model based on the first category dataset and the total dataset to obtain the first music work classification model. The second model training module is used to train the second network model based on the second category dataset and the third category dataset to obtain the second music work classification model.
24. The musical work classification device according to claim 23, wherein, The first music work classification model is used to classify audio data with a third category, and the second music work classification model is used to classify audio data with a first category and / or audio data with a second category. The first category is the poor category, the second category is the average category, and the third category is the excellent category.
25. The musical work classification device according to claim 23, wherein, The first network model is trained based on the first category dataset and the total dataset to obtain the first music work classification model, including: Extract the third feature matrix of the original audio data included in the first category dataset and the original audio data included in the total dataset; The third feature matrix is subjected to feature filtering to obtain a fourth feature matrix, and the fourth feature matrix is input into the first network model to obtain a first prediction result; Based on the first prediction result and the original labels of the original audio data, calculate the first harmonic mean of the original audio data included in the first category dataset and the second harmonic mean of the original audio data included in the total dataset. The hyperparameters of the first network model are adjusted based on the condition that the first harmonic mean is greater than the second harmonic mean, thus obtaining the first music work classification model.
26. The musical work classification device according to claim 23, wherein, The second network model is trained using the second and third category datasets to obtain the second music work classification model, including: Extract the fifth feature matrix of the original audio data included in the second category of data and the sixth feature matrix of the original audio data included in the third category of data; The fifth and sixth feature matrices are subjected to feature filtering to obtain the seventh and eighth feature matrices, and the seventh and eighth feature matrices are input into the second network model to obtain the second prediction result; Based on the second prediction result and the original labels of the original audio data, calculate the third harmonic mean of the original audio data included in the second category dataset and the fourth harmonic mean of the original audio data included in the third category dataset. The hyperparameters of the second network model are adjusted by using the condition that the fourth harmonic mean is greater than the third harmonic mean, thus obtaining the second music work classification model.
27. A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for classifying musical works according to any one of claims 1-13.
28. An electronic device comprising: processor; as well as Memory for storing the executable instructions of the processor; The processor is configured to execute the classification method of musical works according to any one of claims 1-13 by executing the executable instructions.
Citation Information
Patent Citations
Song evaluation method, apparatus and device, and storage medium
CN108549641A
Automatic artistic painting classification system and method based on support vector machine
CN112070116A
Model training method, work pushing method and device, electronic equipment and storage medium
CN112464083A
Song quality identification method and device, equipment and storage medium
CN112559794A
Intelligent sound box network flow classification method and system, electronic equipment and storage medium
CN114219008A