A song scoring method based on a time series convolutional neural network

Through the method based on time series convolution neural network, the characteristics of singing audio and accompaniment audio are analyzed, and the problem of the inability to accurately evaluate the singing quality of the cover version in the prior art is solved, achieving a more accurate and intuitive evaluation effect.

CN115713947BActive Publication Date: 2025-05-30HANGZHOU NORMAL UNIVERSITY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211399253.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-09
Publication Date
2025-05-30
Estimated Expiration
2042-11-09

AI Technical Summary

Technical Problem

The existing singing rating system is mainly based on the original singing version and cannot accurately reflect the singing quality of the cover version and the degree of popularity among the audience.

Method used

Using a time series convolutional neural network (TCN)-based approach, the singing quality and audience-like degree of the cover version were predicted by collecting and analyzing the acoustic and physical characteristics of singer singing and accompaniment audio.

Benefits of technology

A more accurate and intuitive assessment of the quality of singing and the degree of audience love is achieved, and the limitations of the original singing version are abandoned, and the objectivity and entertainment of the rating are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115713947B_ABST
    Figure CN115713947B_ABST
Patent Text Reader

Abstract

The present invention relates to a song scoring method based on a time series convolutional neural network. The singing quality scoring model of the present invention takes the time series convolutional neural network TCN as the main body, performs sequence analysis on audio acoustic features and physical features, and explores the potential association between the physical and acoustic feature sequences and the singing quality; multiple TCN residual modules are set up to directly connect the input layer and the output layer to achieve cross-layer transmission of feature information; using acoustic and physical features as inputs and the calculated singing quality score of the corresponding version as the expected output to train the model. The present invention trains the model by collecting public evaluation indicators of different cover versions, abandons the scoring idea based on the original singer, and the output score is closer to the subjective feelings of the public; based on the time series convolutional neural network, the human voice track and the accompaniment track are separated to separately extract acoustic features and physical features, improving the accompaniment independence of the singing score and making the evaluation more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of neural networks and relates to a song scoring method based on a time series convolutional neural network (TCN). Background Art

[0002] With the development of social economy, people's demand for spiritual entertainment is becoming increasingly strong. Singing, as an entertainment method with a low threshold, strong interactivity and high participation rate, is deeply loved by the broad masses of the people. The opening of KTVs everywhere, the various singing TV talent shows and various singing apps can illustrate people's demand for this entertainment activity of singing. In singing TV talent shows, the program team will invite several well-known singers to be on-site judges, give comments on the singing performance of each singer, give scores, so as to determine whether the singer can enter the next link of the program; when singing in a KTV, the KTV's song selection system will give a score after the singing is completed, and will also rank the customer's score in the entire KTV system and compare it with all customers; when singing in a singing app, the app will also give different ratings such as "C", "B", "A", "S", "SS" and "SSS" for each song. It can be seen that in the singing behavior, an important part is to evaluate the singing quality and level of the singer.

[0003] When evaluating the singing quality and level, the evaluation of professionals is obviously the most accurate. However, when singing in KTVs and singing apps, it is impossible to obtain the evaluation of professionals at any time, so various singing scoring systems have come into being. The existing singing scoring systems on the market are mainly based on the original version of the song, measure some characteristics of the singer's voice, such as the difference in frequency, sound intensity, etc. from the original version and factors such as the degree of completion, and obtain an evaluation index. Therefore, the existing singing scoring systems mainly measure the gap between the objective index of the singer's voice and the objective index of the original version to obtain the score. However, the original version of some songs may not be the best sung version, and should not be used as the full score basis when scoring in all cases. Therefore, design a song scoring method based on a time series convolutional neural network (TCN) with the most popular version among all versions of the song including the cover version as the full score basis, predict the degree to which the singer's singing voice frequency and accompaniment audio's acoustic and physical characteristics may be liked by the audience; score the song by extracting the acoustic and physical characteristics of the song audio (output a number between 0 and 10), and predict the degree to which it may be liked by the audience according to the high or low score (the higher the score, the higher the possibility of being liked). So that people can get a more accurate and intuitive understanding of the quality and level of their own singing voices in the singing entertainment activity. Summary of the Invention

[0004] The object of the present invention is to provide a song scoring method based on a time series convolutional neural network.

[0005] The present invention specifically includes the following steps:

[0006] Step 1: Collect and obtain publicly available song cover data from any online music platform and construct a cover song dataset; specifically: collect all evaluation indicators for each cover version of each song. The singing evaluation indicators include the number of comment entries of the song on the online music platform and the number of search result entries on the search engine with the "song name + singer name" as the keyword; the above evaluation indicators constitute the cover song dataset.

[0007] The dataset is collected using python crawler technology. First, analyze the web page of an online music platform to obtain the request interface for the required information; obtain the request result by calling the get() method of the requests library; then analyze the request result, parse the response content through the HTML() method of the lxml.etree library, and finally use the xpath() method to obtain the required information.

[0008] Step 2: Synthesize each evaluation indicator, calculate the evaluation scores of several different cover versions of each song and use it as the prediction target of the singing quality scoring model; construct a training set from each song and its evaluation score, and construct a test set from each song itself. where m x,y represents the y-th cover version of the x-th song; represents the singing evaluation score of the y-th cover version of the x-th song; represents the number of comment entries of the y-th cover version of the x-th song on a certain online music platform; represents the number of search result entries obtained on the j-th search engine with the "song name + singer name" as the keyword for the y-th cover version of the x-th song; b w , b j represents the correction weight at the corresponding position; n x represents the total number of different cover versions of the x-th song.

[0009] Step 3: Separate the vocal track and the accompaniment track of each song, and extract their acoustic and physical characteristics of the audio respectively as the input of the model; the said feature extraction includes extracting the peak frequency, frequency domain expectation, time domain variance, short-time energy feature and Mel frequency cepstral coefficient of the accompaniment track, and also includes extracting the timbre feature, zero-crossing rate, fundamental frequency and sound intensity of the vocal track. Integrate the above acoustic characteristics and physical characteristics into a 128-dimensional input vector as the input of the singing quality scoring model.

[0010] Step 4. Establish a singing quality scoring model based on a time series convolutional neural network:

[0011] The singing quality scoring model takes the time series convolutional neural network TCN as the main body, performs sequence analysis on audio acoustic features and physical features, and explores the potential relationship between the physical and acoustic feature sequences and singing quality; multiple TCN residual modules are set up to directly connect the input layer and the output layer to achieve cross-layer transfer of feature information; the input end and the output end of each TCN residual module are also connected to each other through 1×1 convolution.

[0012] The TCN residual module consists of an input fully connected layer, multiple dilated causal convolutional layers, multiple WeightNorm weight normalization layers, multiple Relu activation layers, multiple Dropout regularization layers, and an output fully connected layer. The WeightNorm weight normalization layer, Relu activation layer, and Dropout regularization layer are respectively arranged after each dilated causal convolutional layer in sequence. The first dilated causal convolutional layer is connected to the input fully connected layer, and the last Dropout regularization layer is connected to the output fully connected layer.

[0013] The input fully connected layer is used to receive the input acoustic features and physical feature sequences and integrate them into a fixed 512-dimensional input feature vector. The dilated causal convolutional layer of the time series convolutional neural network extracts the overall features of the input feature sequence and explores its potential mapping relationship with singing quality. The weight normalization layer normalizes the network weight W. The Relu activation layer is a commonly used non-linear correction unit in the neural network model. The Dropout regularization layer means that each node in the model network has a certain probability of being deleted, which can prevent overfitting of the model network. The fully connected layer is used to integrate various determining factors and give the prediction result.

[0014] The singing quality scoring model takes acoustic and physical features as inputs and the calculated singing quality score of the corresponding version as the expected output. The model is trained using the gradient descent method until convergence, and then the model is verified using the cross-validation method.

[0015] Step 5. Use the acoustic and physical features of any cover as the input of the singing quality scoring model and output the evaluation score of this cover.

[0016] The second object of the present invention is to provide a scoring system for accurately outputting the quality of song covers, which is used to run the trained and completed song cover scoring model based on TCN.

[0017] The third object of the present invention is to provide an electronic device, including a processor and a memory, where the memory stores machine-executable instructions that can be executed by the processor, and the processor executes the machine-executable instructions to implement the scoring system.

[0018] The beneficial effects of the present invention are reflected in:

[0019] By collecting all evaluation indicators of each cover version of each song on public music platforms or search engines, integrating various evaluation indicators, and calculating the evaluation scores of several different cover versions of each song; using the obtained public evaluation indicators to measure the quality of the song, abandoning the previous scoring idea of the singing scoring system based on the original singer, the obtained score is closer to the subjective feelings of the public, which can better improve people's understanding of their own singing level and enhance the entertainment.

[0020] Using a time series convolutional neural network, separating the human voice track and the accompaniment track to separately extract acoustic features and physical features, weakening the influence of the accompaniment, improving the accompaniment independence of the singing score, and making the output evaluation more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a schematic flow chart for the establishment of the cover song dataset of the present invention;

[0022] Figure 2 It is a schematic structural diagram of the singing quality scoring model. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. It should be understood that these embodiments are only used to illustrate the present invention and not to limit the scope of the present invention. The operating methods without specific conditions noted in the following embodiments are usually in accordance with conventional conditions or in accordance with the conditions recommended by the manufacturer.

[0024] A song scoring method based on a time series convolutional neural network specifically includes the following steps:

[0025] Step 1: Collect and obtain public song cover data from any online music platform and construct a cover song dataset; specifically: collect all evaluation indicators of each cover version of each song, and the singing evaluation indicators include the number of comment entries of the song on the online music platform and the number of search result entries on the search engine with the "song name + singer name" as the keyword; the above evaluation indicators constitute the cover song dataset;

[0026] Such as Figure 1As shown in the figure, the dataset is collected using Python web scraping technology. First, by analyzing the web pages of a certain online music platform, the request interface for obtaining the required information is obtained; the request result is obtained by calling the get() method of the requests library; then the request result is analyzed, the response content is parsed by the HTML() method of the lxml.etree library, and finally the required information is obtained using the xpath() method.

[0027] Specifically:

[0028] Collect a large number of playlist IDs from a certain online music platform. According to the playlist IDs, use the playlist information interface to crawl a large number of song names, mainly Chinese songs. According to the crawled song names, call the search interface of a certain online music platform to retrieve a large number of different versions of the same song, and record the IDs of these songs on the online music platform. Manually clean the crawled data, and manually remove inappropriate or incorrect versions. Generally, the following error situations will occur: the search results contain incorrect songs; the search results contain inappropriate versions, such as foreign language versions of songs; the original singer of the song is unknown. For the cleaned songs, crawl the number of comment entries of the songs on a certain online music platform. Crawl the number of search result entries of all songs on the three search engines of Google, Baidu, and Sogou with the keyword "song name singer name". Organize the crawled data in an orderly manner, and obtain the actual singing evaluation index of each song through a calculation formula. Separate the vocal track and the accompaniment track of the songs in the dataset, and use the librosa library and other methods in Python to extract features such as the envelope peak frequency, frequency domain expectation, time domain variance, short-time energy feature, and Mel-frequency cepstral coefficient of the accompaniment audio. At the same time, features such as the timbre feature, zero-crossing rate, fundamental frequency, and sound intensity of the vocal audio also need to be extracted; these features are saved for later use.

[0029] Step 2: Combine various evaluation indicators to calculate the evaluation scores of several different cover versions of each song And use it as the training set label of the prediction model; construct a training set from each song and its evaluation score, and construct a test set from each song itself; Among them, m x,y represents the y-th cover version of the x-th song; represents the singing evaluation score of the y-th cover version of the x-th song; represents the number of comment entries of the y-th cover version of the x-th song on a certain online music platform; represents the number of search result entries obtained by the y-th cover version of the x-th song on the j-th search engine with the keyword "song name singer name"; b w , b j represents the correction weight at the corresponding position; n xRepresents the total number of different cover versions of the x-th song.

[0030] Step 3: Separate the vocal track and the accompaniment track of each song, and extract their acoustic and physical characteristics of the audio respectively as the input of the model;

[0031] Feature extraction includes extracting the peak frequency, frequency domain expectation, time domain variance, short-time energy feature, and Mel-frequency cepstral coefficients of the accompaniment track, etc. At the same time, it also includes extracting features such as timbre feature, zero-crossing rate, fundamental frequency, and sound intensity of the vocal track.

[0032] The extraction step of the peak frequency is to extract the envelope of the music signal, mark the peak points of the envelope contour, and calculate the number of times the envelope peak appears per second of the entire music, which is recorded as the peak frequency and used to measure the strength of the music rhythm.

[0033] The extraction step of the frequency domain expectation is to transform the time domain signal of the audio into a frequency domain signal through Fourier transform and then calculate its expectation value.

[0034] The short-time energy feature is used to characterize the acoustic feature of the sound intensity of the song. By calculating the short-time energy feature in the music information frame to characterize the magnitude of the sound intensity, the larger the short-time energy feature, the more energy is contained within this time interval, and the corresponding sound intensity is greater. Conversely, the smaller the short-time energy feature, the smaller the sound intensity.

[0035] The Mel-frequency cepstral coefficient is an audio feature that takes into account the human auditory characteristics. Its extraction step is to first map the linear spectrum to the Mel non-linear spectrum based on auditory perception, and then transform it to the cepstrum. After the above preprocessing of the music signal, perform a fast Fourier transform on each short-time analysis window to obtain the corresponding spectrum, pass this spectrum through the Mel filter to obtain the Mel spectrum, and perform cepstral analysis on this basis to obtain the feature vector.

[0036] The above-mentioned acoustic and physical characteristics will be integrated into a 128-dimensional input vector as the input of the singing quality scoring model.

[0037] Step 4: Establish a singing quality scoring model based on the time series convolutional neural network as Figure 2 shown; the singing quality scoring model takes the time series convolutional neural network TCN as the main body, performs sequence analysis on the audio acoustic features and physical characteristics, and explores the potential relationship between the physical and acoustic feature sequences and the singing quality; sets multiple TCN residual modules, directly connecting the input layer and the output layer to achieve cross-layer transmission of feature information; there is also a 1×1 convolution connection between the input end and the output end of each TCN residual module; T

[0038] The CN residual module is used to directly connect the input layer and the output layer, connecting the input X sequence with the output F(x) of the convolutional network. Its output is: O = Activation(3 + F(3)); where F(x) is the output of the convolutional layer and Activation(·) is the activation function; to achieve cross-layer transmission of feature information.

[0039] The TCN residual module consists of an input fully connected layer, multiple dilated causal convolutional layers, multiple WeightNorm weight normalization layers, multiple Relu activation layers, multiple Dropout regularization layers, and an output fully connected layer. The WeightNorm weight normalization layer, Relu activation layer, and Dropout regularization layer are respectively arranged after each dilated causal convolutional layer in sequence. The first dilated causal convolutional layer is connected to the input fully connected layer, and the last Dropout regularization layer is connected to the output fully connected layer;

[0040] The input fully connected layer is used to receive the input acoustic features and physical feature sequences and integrate them into a fixed 512-dimensional input feature vector. The dilated causal convolutional layer of the time series convolutional neural network extracts the overall features of the input feature sequence and explores its potential mapping relationship with the singing quality. The weight normalization layer normalizes the network weight W. The Relu activation layer is a commonly used non-linear correction unit in the neural network model. The Dropout regularization layer means that each node in the model network has a certain probability of being deleted, which can prevent overfitting of the model network. The fully connected layer is used to integrate various determining factors and give the prediction result.

[0041] The input fully connected layer is used to receive the input of the accompaniment and human voice audio feature data and map the input audio feature data into an embedding vector with a dimension of 512. This embedding vector is used to represent the comprehensive audio feature information of the input human voice and accompaniment;

[0042] The dilated causal convolutional layer is divided into three parts: convolution, dilation, and causality; the convolution refers to the classic convolution in CNN, which is a sliding operation of the convolution kernel on the data. The convolution kernel performs a sliding operation on the audio feature information mapped by the input fully connected layer to extract the local sequence features of the input audio feature data; the dilation refers to dilated convolution. Dilated convolution allows for interval sampling of the input during convolution, aiming to increase the model receptive field while maintaining the number of model layers. Dilated convolution combines the global high-dimensional features on the basis of classic convolution, making the output feature vector result consider a wider sequence range of input audio features; the causality refers to causal convolution. Causal convolution means that the audio feature data at time t in the i-th layer only depends on the values at time t and before in the (i - 1)-th layer. Causal convolution can discard the reading of future audio feature data during training and is a strict time-constrained model;

[0043] The weight normalization layer normalizes the network weights W. The specific method is to decouple the weight vector w into the parameter vector v and the parameter scalar g in terms of its Euclidean norm and its direction, and then use SGD to optimize these two parameters separately. The specific formula is as follows: where v is the original weight and g is the learnable parameter.

[0044] The Relu activation layer is a commonly used non-linear correction unit in neural network models. Its formula is:

[0045] The Dropout regularization layer means that each node in the model network has a certain probability of being deleted, which can prevent overfitting of the model network.

[0046] The output fully connected layer is used to integrate all determining factors and give the prediction result.

[0047] The singing quality scoring model takes acoustic and physical features as inputs and the calculated singing quality score of the corresponding version as the expected output. The model is trained using the gradient descent method until convergence, and then the cross-validation method is used to verify the model.

[0048] Step 5: Use the acoustic and physical features of any cover as the input of the singing quality scoring model, and output the evaluation score of this cover.

[0049] Specifically: The input acoustic and physical features pass through the input fully connected layer and the TCN convolutional module, and are merged with the TCN residual module to obtain an output vector. After the output fully connected layer with an input dimension of 256 and an output dimension of 1 receives the output feature vectors of the TCN convolutional module and the residual module, it integrates all singing quality determining factors and outputs the prediction result. By passing the output value of the output fully connected layer through the sigmod function and then multiplying by 10, the output song quality score can be obtained. Therefore, the score ranges from 0 to 10, and the higher the score, the higher the song quality.

Claims

1. A song scoring method based on a time series convolutional neural network, characterized in that: Specifically, it includes the following steps: Step 1: Collect and obtain public song cover data from any online music platform and construct a cover song dataset; specifically: collect all evaluation indicators for each cover version of each song. The singing evaluation indicators include the number of comment entries of the song on the online music platform and the number of search result entries in the search engine with "song name singer name" as the keyword; the above evaluation indicators constitute the cover song dataset; Step 2: Calculate the evaluation scores of several different cover versions of each song by integrating various evaluation indicators and use them as the prediction targets of the singing quality scoring model; construct a training set from each song and its evaluation score, and construct a test set from each song itself; where m x,y represents the y-th cover version of the x-th song; represents the singing evaluation score of the y-th cover version of the x-th song; represents the number of comment entries of the y-th cover version of the x-th song on a certain online music platform; represents the number of search result entries obtained by using the keyword "song name singer name" for the y-th cover version of the x-th song on the j-th search engine; b w , b j represents the correction weight at the corresponding position; n x represents the total number of different cover versions of the x-th song; Step 3: Separate the vocal track and accompaniment track of each song, and extract their acoustic and physical characteristics of the audio respectively as the input of the model; the feature extraction includes extracting the peak frequency, frequency domain expectation, time domain variance, short-time energy characteristics and Mel frequency cepstral coefficients of the accompaniment track, and also includes extracting the timbre characteristics, zero-crossing rate, fundamental frequency and sound intensity of the vocal track; integrate the above acoustic and physical characteristics into a 128-dimensional input vector as the input of the singing quality scoring model; Step 4: Establish a singing quality scoring model based on the time series convolutional neural network: The singing quality scoring model takes the time series convolutional neural network TCN as the main body, conducts sequence analysis on the audio acoustic and physical characteristics, and explores the potential relationship between the physical and acoustic feature sequences and the singing quality; sets multiple TCN residual modules, directly connecting the input layer and the output layer to achieve cross-layer transmission of feature information; there is also a 1×1 convolution connection between the input end and the output end of each TCN convolution module; The TCN residual module consists of an input fully connected layer, multiple dilated causal convolution layers, multiple WeightNorm weight normalization layers, multiple Relu activation layers, multiple Dropout regularization layers and an output fully connected layer. The WeightNorm weight normalization layer, Relu activation layer and Dropout regularization layer are respectively set after each dilated causal convolution layer in sequence. The first dilated causal convolution layer is connected to the input fully connected layer, and the last Dropout regularization layer is connected to the output fully connected layer; The input fully connected layer is used to receive the input acoustic and physical feature sequences and integrate them into a fixed 512-dimensional input feature vector; the dilated causal convolution layer of the time series convolutional neural network extracts the overall features of the input feature sequence and explores its potential mapping relationship with the singing quality; the weight normalization layer normalizes the network weight W; the Relu activation layer is a commonly used non-linear correction unit in the neural network model; the Dropout regularization layer means that each node in the model network has a certain probability of being deleted, which can prevent overfitting of the model network; the fully connected layer is used to integrate each determinant and give the prediction result; The singing quality scoring model takes the acoustic and physical characteristics as the input and the calculated singing quality score of the corresponding version as the expected output, trains the network model using the gradient descent method until convergence, and then validates the model using the cross-validation method; Step Five: Use the acoustic and physical characteristics of any cover as the input of the singing quality scoring model, and output the evaluation score of this cover.

2. The song scoring method based on a time series convolutional neural network according to claim 1, characterized in that: The data set described in Step One is collected using python crawler technology. First, analyze the web page of a certain online music platform to obtain the request interface for the required information; obtain the request result by calling the get() method of the requests library; then analyze the request result, parse the response content through the HTML() method of the lxml.etree library, and finally use the xpath() method to obtain the required information.

Citation Information

Patent Citations

  • Evaluating method and device for power system transient stability

    CN108879732A