Training method for chorus song recognition network and evaluation method for chorus songs

By training the chorus song recognition network, using the Mel score and timbre vector, combining the long and short-term memory network and the full connection layer, the problem of simple and single chorus song rating is solved, and the precise evaluation of the chorus user's timbre understanding and singing skills is achieved.

CN116137147BActive Publication Date: 2025-07-11TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310177696.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-28
Publication Date
2025-07-11
Estimated Expiration
2043-02-28

AI Technical Summary

Technical Problem

The existing chorus song rating method is simple and single, and it is impossible to evaluate the chorus tacit understanding and singing skills of chorus users in a targeted manner, resulting in inaccurate evaluation.

Method used

By training the chorus song recognition network, using the Mel score and timbre vector, combining the long and short-term memory network and the full connection layer, the chorus user's timbre characteristics and tacit understanding are extracted, and the loss function and backpropagation algorithm are used for correction, and the chorus' timbre tacit understanding and completeness scores are calculated.

Benefits of technology

It improves the pertinence and accuracy of chorus song evaluation, and can better evaluate the chorus user's timbre understanding and singing skills matching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116137147B_ABST
    Figure CN116137147B_ABST
Patent Text Reader

Abstract

Embodiments of the present invention disclose a training method for a choral song recognition network, an evaluation method for choral songs, and related devices, which are used to improve the accuracy of choral song evaluation. The method of the embodiments of the present invention includes: obtaining a labeled first training sample, where the first training sample includes Mel spectrograms of singing samples of multiple choral users singing multiple choral songs and timbre vectors of multiple choral users; inputting the Mel spectrograms of the singing samples of each choral user singing the same choral song into the choral song recognition network to obtain the output timbre vectors of the corresponding choral users; according to the user labels, obtaining the timbre vectors of each choral user output by the choral song recognition network and the timbre vectors of the same choral user in the first training sample; using a loss function to calculate the difference between the timbre vectors of each choral user output by the choral song recognition network and the timbre vectors of the same choral user in the training sample; and correcting the choral song recognition network according to the difference and the backpropagation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of song data processing, and in particular to a training method for a chorus song recognition network, a chorus song evaluation method and related devices. Background Art

[0002] Whether it is online karaoke software or offline karaoke, chorus is a very important singing scene. Usually the chorus of popular songs is a round-robin superimposed with a small amount of re-singing, that is, one paragraph A and one paragraph B, and some paragraphs are marked as chorus paragraphs, that is, two people sing a sentence at the same time.

[0003] The tracks of two users singing the whole song or the chorus sections, or the duet tracks, are all superimposed according to the built-in section divisions, and the corresponding work scores are also accumulated according to their respective parts, that is, the chorus score is the accumulation of A's singing score and B's singing score. This scoring method is relatively simple and single, and cannot provide targeted evaluation of the chorus users in the chorus scene. Summary of the invention

[0004] The embodiments of the present invention provide a training method for a chorus song recognition network, a chorus song evaluation method and related devices, which are used to first train the recognition network for evaluating chorus songs, and further use the chorus song recognition network to score the tone tacit understanding of chorus users, and use the chorus tacit understanding score to evaluate the chorus songs of the chorus, thereby improving the pertinence and accuracy of the evaluation of the chorus songs of the chorus.

[0005] A first aspect of an embodiment of the present application provides a training method for a chorus song recognition network, comprising:

[0006] Acquire a first labeled training sample, wherein the first training sample includes mel spectra of vocal samples of multiple chorus songs sung by multiple chorus users and timbre vectors of multiple chorus users;

[0007] Inputting the mel spectrum of the voice samples of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network;

[0008] According to the user label, obtaining the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample;

[0009] Calculate the difference between the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample using a loss function;

[0010] Correct the choral song recognition network according to the difference and the backpropagation algorithm.

[0011] Preferably, the choral song recognition network includes a long short-term memory network and a fully connected layer;

[0012] Input the Mel spectrogram of the singing voice samples of each choral user singing the same choral song into the choral song recognition network to obtain the timbre vector of the corresponding choral user output by the choral song recognition network, including:

[0013] Input the Mel spectrogram of the singing voice samples of each choral user singing the same choral song into the long short-term memory network to extract the eigenvalue of the timbre vector of each choral user;

[0014] Input the eigenvalue of the timbre vector of each choral user into the fully connected layer to integrate and output the eigenvalue of the timbre vector of each choral user, and obtain the timbre vector of the corresponding choral user output by the choral song recognition network.

[0015] Preferably, correcting the choral song recognition network according to the difference and the backpropagation algorithm includes:

[0016] Correct the long short-term memory network and the fully connected layer according to the difference and the backpropagation algorithm.

[0017] Preferably, the intra-class distance between the Mel spectrogram vectors of the singing voice samples of the same choral user is less than or equal to a preset reference value, and the inter-class distance between the Mel spectrogram vectors of the singing voice samples of different choral users is greater than the preset reference value.

[0018] Preferably, the method further includes:

[0019] Obtain a labeled second training sample, where the second training sample includes the timbre vectors of multiple choral users singing multiple choral songs, and the timbre tacit understanding scores of multiple choral users in each choral song;

[0020] Concatenate the timbre vectors of the song segments of multiple choral users singing the same choral song to obtain a concatenated timbre vector;

[0021] Input the concatenated timbre vector into a fully connected network to obtain the timbre tacit understanding score of the choristers output by the fully connected network;

[0022] According to the user label, obtain the timbre tacit understanding score of the choristers output by the fully connected network and the timbre tacit understanding score of the same choristers in the second training sample;

[0023] Use a loss function to calculate the difference between the timbre tacit understanding score of the choristers output by the fully connected network and the timbre tacit understanding score of the same choristers in the second training sample;

[0024] Modify the fully connected network according to the difference and the backpropagation algorithm.

[0025] Preferably, the fully connected network includes a feature layer, a pooling layer, and a fully connected layer;

[0026] Input the spliced timbre vectors into the fully connected network to obtain the timbre tacit understanding scores output by the fully connected network, including:

[0027] Input the spliced timbre vectors into the feature layer to extract the timbre feature values of the spliced timbre vectors;

[0028] Input the timbre feature values of the spliced timbre vectors into the pooling layer to perform redundancy removal and dimensionality reduction on the timbre feature values of the spliced timbre vectors;

[0029] Input the timbre feature values of the spliced timbre vectors after redundancy removal and dimensionality reduction into the fully connected layer to integrate and output the feature values of the spliced timbre vectors after redundancy removal and dimensionality reduction, so as to obtain the timbre tacit understanding scores of the chorus singers singing the same chorus song.

[0030] Preferably, modifying the fully connected network according to the difference and the backpropagation algorithm includes:

[0031] Modify the feature layer, the pooling layer, and the fully connected layer according to the difference and the backpropagation algorithm.

[0032] The second aspect of the embodiments of the present application provides an evaluation method for a chorus song, and the method includes:

[0033] Obtain the singing segments corresponding to the chorus singers in the chorus song, where the chorus singers include at least two chorus users;

[0034] Extract the Mel spectrograms of the singing segments corresponding to each chorus user;

[0035] Input the Mel spectrograms of the singing segments corresponding to each chorus user into the chorus song recognition network to obtain the timbre vectors of each chorus user. The chorus song recognition network includes a long short-term memory network and a fully connected layer. The long short-term memory network is used to extract the feature values of the timbre vectors of each chorus user, and the fully connected layer is used to integrate and output the feature values of the timbre vectors of each chorus user;

[0036] Splice the timbre vectors of each chorus user to obtain the spliced timbre vectors;

[0037] Input the spliced timbre vectors into the feature layer of the fully connected network to extract the feature values of the spliced timbre vectors;

[0038] Inputting the eigenvalues ​​of the concatenated timbre vectors into a pooling layer of a fully connected network to remove redundancy and reduce the dimensionality of the eigenvalues ​​of the concatenated timbre vectors;

[0039] The eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction are input into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction, and obtain the timbre tacit understanding scores of the chorus members singing the same chorus song;

[0040] The chorus songs sung by the chorus members were evaluated based on their timbre tacit understanding scores.

[0041] Preferably, the timbre vectors of each chorus user are spliced ​​to obtain a spliced ​​timbre vector, including:

[0042] The timbre vectors of each chorus user are averaged in the time dimension to obtain the timbre vector of each chorus user without the time series feature, and the timbre vectors of each user without the time series feature are concatenated.

[0043] Preferably, the timbre vectors of each chorus user are spliced ​​to obtain a spliced ​​timbre vector, including:

[0044] The timbre vectors of at least two chorus users are superimposed in the time dimension, and then the superimposed timbre vectors are averaged according to the number of chorus users to obtain a concatenated timbre vector containing time series features.

[0045] Preferably, the method further comprises:

[0046] Obtaining the chorus interval of the chorus singers for the same chorus song;

[0047] Calculate the completeness score of each chorus user in the chorus interval;

[0048] The completeness score of the chorus users for the same chorus song is calculated according to the completeness score of each chorus user in the chorus interval.

[0049] Preferably, obtaining the chorus interval of the chorus singers for the same chorus song includes:

[0050] Using a voice activity detection (VAD) method, the chorus start time and the chorus end time of the chorus singers in the same chorus song are obtained;

[0051] The chorus interval of the chorus singers in the same chorus song is determined according to the chorus start time and the chorus end time.

[0052] Preferably, the calculating of the completeness score of each chorus user in the chorus interval includes:

[0053] Use the voice activity detection (VAD) method to obtain the vocal segments of each chorus user within the chorus interval, and the vocal duration of each line of lyrics in each vocal segment;

[0054] Obtain the voice file of the chorus song, and obtain the original singer's vocal duration of each line of lyrics from the voice file;

[0055] Calculate the single-line integrity score of each chorus user within the chorus interval according to the vocal duration of each line of lyrics of each chorus user within the chorus interval and the original singer's vocal duration of the corresponding lyrics;

[0056] Calculate the integrity score of each chorus user within the chorus interval according to the single-line integrity scores of each chorus user for the chorus song.

[0057] Preferably, evaluate the chorus songs of the choristers according to the timbre matching degree scores of the choristers, including:

[0058] Evaluate the chorus songs of the choristers according to the timbre matching degree scores of the choristers for the same chorus song and the integrity scores for the same chorus song.

[0059] Preferably, the evaluation of the chorus songs of the choristers according to the timbre matching degree scores of the choristers further includes:

[0060] Extract the pitch values of each chorus user's singing segment;

[0061] Obtain the voice file of the chorus song, and obtain the original singer's pitch values of each line of lyrics from the voice file;

[0062] Obtain the singing skill matching degree scores of multiple chorus users singing the same chorus song according to the pitch values of each chorus user's singing segment and the original singer's pitch values of the corresponding lyric segments;

[0063] Evaluate the chorus songs of the choristers according to the timbre matching degree scores of the choristers singing the same chorus song and the singing skill matching degree scores of the choristers singing the same chorus song.

[0064] The third aspect of the embodiments of the present application provides a training device for a chorus song recognition network, and the device includes:

[0065] A first acquisition unit, configured to acquire a labeled first training sample, where the first training sample includes Mel spectrograms of singing samples of multiple chorus users singing multiple chorus songs and timbre vectors of multiple chorus users;

[0066] A first input unit for inputting the Mel spectrograms of the singing voice samples of each chorus user singing the same chorus song into a chorus song recognition network to obtain the timbre vectors of the corresponding chorus users output by the chorus song recognition network;

[0067] The first obtaining unit is further configured to obtain, according to the user tags, the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training sample;

[0068] A first calculation unit for calculating the difference between the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training sample by using a loss function;

[0069] A first correction unit for correcting the chorus song recognition network according to the difference and the backpropagation algorithm.

[0070] Preferably, the chorus song recognition network includes a long short-term memory network and a fully connected layer;

[0071] The first input unit is specifically configured to:

[0072] Input the Mel spectrograms of the singing voice samples of each chorus user singing the same chorus song into the long short-term memory network to extract the eigenvalue of the timbre vector of each chorus user;

[0073] Input the eigenvalue of the timbre vector of each chorus user into the fully connected layer to integrate and output the eigenvalue of the timbre vector of each chorus user, so as to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

[0074] The first modification unit is specifically configured to:

[0075] Correct the long short-term memory network and the fully connected layer according to the difference and the backpropagation algorithm.

[0076] Preferably, the intra-class distance between the Mel spectrogram vectors of the singing voice samples of the same chorus user is less than or equal to a preset reference value, and the inter-class distance between the Mel spectrogram vectors of the singing voice samples of different chorus users is greater than the preset reference value.

[0077] Preferably, the device further includes:

[0078] A second obtaining unit for obtaining a labeled second training sample, where the second training sample includes the timbre vectors of multiple chorus users singing multiple chorus songs, and the timbre tacit understanding scores of multiple chorus users in each chorus song;

[0079] A splicing unit, configured to splice the timbre vectors of multiple chorus user song segments that sing the same chorus song to obtain a spliced timbre vector;

[0080] A second input unit, configured to input the spliced timbre vector into a fully connected network to obtain a chorus singer timbre matching score output by the fully connected network;

[0081] The second obtaining unit is further configured to obtain, according to the user label, the chorus singer timbre matching score output by the fully connected network and the timbre matching score of the same chorus singer in the second training sample;

[0082] A second calculation unit, configured to calculate, by using a loss function, a difference between the chorus singer timbre matching score output by the fully connected network and the timbre matching score of the same chorus singer in the second training sample;

[0083] A second correction unit, configured to correct the fully connected network according to the difference and the backpropagation algorithm.

[0084] Preferably, the fully connected network includes a feature layer, a pooling layer, and a fully connected layer;

[0085] The second input unit is specifically configured to:

[0086] Input the spliced timbre vector into the feature layer to extract the timbre feature value of the spliced timbre vector;

[0087] Input the timbre feature value of the spliced timbre vector into the pooling layer to perform redundancy removal and dimensionality reduction on the timbre feature value of the spliced timbre vector;

[0088] Input the timbre feature value of the spliced timbre vector after redundancy removal and dimensionality reduction into the fully connected layer to be used for integrating and outputting the feature values of the spliced timbre vector after redundancy removal and dimensionality reduction, so as to obtain the timbre matching score of the chorus singer singing the same chorus song.

[0089] The second correction unit is specifically configured to:

[0090] Correct the feature layer, the pooling layer, and the fully connected layer according to the difference and the backpropagation algorithm.

[0091] A fourth aspect of the embodiments of the present application provides an evaluation device for a chorus song, where the device includes:

[0092] An obtaining unit, configured to obtain a singing segment corresponding to a chorus singer in a chorus song, where the chorus singer includes at least two chorus users;

[0093] An extraction unit, configured to extract the Mel spectrogram of each chorus user's corresponding singing segment;

[0094] An input unit, used for inputting the mel spectrum of the singing segment corresponding to each chorus user into the chorus song recognition network to obtain the timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used for extracting the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used for integrating and outputting the eigenvalue of the timbre vector of each chorus user;

[0095] A splicing unit, used for splicing the timbre vector of each chorus user to obtain a spliced ​​timbre vector;

[0096] The extraction unit is further used to input the concatenated timbre vector into a feature layer of a fully connected network to extract a feature value of the concatenated timbre vector;

[0097] The input unit is further used to input the eigenvalues ​​of the concatenated timbre vectors into the pooling layer of the fully connected network, so as to perform redundancy removal and dimensionality reduction on the eigenvalues ​​of the concatenated timbre vectors;

[0098] The input unit is further used to input the eigenvalues ​​of the spliced ​​timbre vectors after de-redundancy and dimensionality reduction into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the spliced ​​timbre vectors after de-redundancy and dimensionality reduction, and obtain the timbre tacit understanding scores of the chorus members singing the same chorus song;

[0099] The evaluation unit is used to evaluate the chorus songs sung by the chorus according to the timbre tacit understanding scores of the chorus.

[0100] Preferably, the splicing unit is specifically used for:

[0101] The timbre vectors of each chorus user are averaged in the time dimension to obtain the timbre vector of each chorus user without the time series feature, and the timbre vectors of each user without the time series feature are concatenated.

[0102] Preferably, the splicing unit is specifically used for:

[0103] The timbre vectors of at least two chorus users are superimposed in the time dimension, and then the superimposed timbre vectors are averaged according to the number of chorus users to obtain a concatenated timbre vector containing time series features.

[0104] Preferably, the acquisition unit is further used for:

[0105] Obtaining the chorus interval of the chorus singers for the same chorus song;

[0106] The calculation unit is further configured to calculate the integrity score of each chorus user within the chorus interval; and calculate the integrity score of the choristers for the same chorus song according to the integrity scores of each chorus user within the chorus interval.

[0107] Preferably, the obtaining unit is specifically configured to:

[0108] Obtain the chorus start time and chorus end time of the choristers in the same chorus song by using the voice activity detection (VAD) method;

[0109] Determine the chorus interval of the choristers in the same chorus song according to the chorus start time and the chorus end time.

[0110] Preferably, the calculation unit is specifically configured to:

[0111] Obtain the vocal segments of each chorus user within the chorus interval by using the voice activity detection (VAD) method, and the vocal duration of each line of lyrics in each vocal segment;

[0112] Obtain the voice file of the chorus song, and obtain the original singer's vocal duration of each line of lyrics from the voice file;

[0113] Calculate the single-line integrity score of each chorus user within the chorus interval according to the vocal duration of each line of lyrics of each chorus user within the chorus interval and the original singer's vocal duration of the corresponding line of lyrics;

[0114] Calculate the integrity score of each chorus user within the chorus interval according to the single-line integrity scores of each chorus user for the chorus song.

[0115] The evaluation unit is specifically configured to:

[0116] Evaluate the chorus song of the choristers according to the timbre matching degree score of the choristers for the same chorus song and the integrity score of the same chorus song.

[0117] The evaluation unit is further configured to:

[0118] Extract the pitch values of the singing segments of each chorus user;

[0119] Obtain the voice file of the chorus song, and obtain the original singer's pitch value of each line of lyrics from the voice file;

[0120] Obtain the singing skill matching degree scores of multiple chorus users singing the same chorus song according to the pitch values of the singing segments of each chorus user and the original singer's pitch values of the corresponding lyric segments of the singing segments;

[0121] The chorus songs sung by the chorus are evaluated according to the timbre tacit understanding scores and the singing skills matching scores of the chorus members when singing the same chorus song.

[0122] A fifth aspect of an embodiment of the present application provides a computer device, comprising a processor, which, when executing a computer program stored in a memory, is used to implement the training method of a chorus song recognition network provided in the first aspect of an embodiment of the present application or the evaluation method of a chorus song provided in the second aspect of an embodiment of the present application.

[0123] The sixth aspect of the embodiments of the present application provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, it is used to implement the training method of the chorus song recognition network provided in the first aspect of the embodiments of the present application or the evaluation method of the chorus song provided in the second aspect of the embodiments of the present application.

[0124] It can be seen from the above technical solutions that the embodiments of the present invention have the following advantages:

[0125] An embodiment of the present application obtains a first labeled training sample, wherein the first training sample includes Mel-spectra of singing samples of multiple chorus songs sung by multiple chorus users and timbre vectors of multiple chorus users; the Mel-spectra of the singing sample of each chorus user singing the same chorus song is input into a chorus song recognition network to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network; according to the user label, the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample are obtained; the difference between the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample is calculated using a loss function; and the chorus song recognition network is corrected according to the difference and a back propagation algorithm.

[0126] Because the chorus song recognition network is trained using the first training sample in the embodiment of the present application, and the first training sample in the embodiment of the present application adopts the Mel-spectrogram of the chorus user's singing sample and the timbre vector of the chorus user, the chorus song recognition network can obtain the user's timbre vector after inputting the Mel-spectrogram of the chorus user's singing sample, thereby improving the convenience of the user in obtaining the user's timbre characteristics in the chorus scenario. BRIEF DESCRIPTION OF THE DRAWINGS

[0127] Figure 1 A schematic diagram of an embodiment of a training method for a chorus song recognition network in an embodiment of the present application;

[0128] Figure 2Schematic diagram of the intra-class distance between Mel-spectrum vectors of the same user's singing voice segments and the inter-class distance between Mel-spectrum vectors of different users' singing voice segments in the embodiments of this application;

[0129] Figure 3 Schematic diagram of the process of inputting the singing voice samples of chorus users into the chorus song recognition network in the embodiments of this application;

[0130] Figure 4 Another schematic diagram of the training method of the chorus song recognition network in the embodiments of this application;

[0131] Figure 5 Schematic diagram of the process of inputting the timbre vectors after splicing the choristers into the fully connected network in the embodiments of this application;

[0132] Figure 6 Schematic diagram of an embodiment of the chorus song evaluation method in the embodiments of this application;

[0133] Figure 7 Another schematic diagram of the chorus song evaluation method in the embodiments of this application;

[0134] Figure 8 Schematic diagram of the chorus interval of User A and User B in the same chorus song in the embodiments of this application;

[0135] Figure 9 This application Figure 7 Refined steps of step 709 in the embodiments;

[0136] Figure 10 Another schematic diagram of the chorus song evaluation method in the embodiments of this application;

[0137] Figure 11 Schematic diagram of an embodiment of the training device of the chorus song recognition network in the embodiments of this application;

[0138] Figure 12 Schematic diagram of an embodiment of the chorus song evaluation device in the embodiments of this application. Detailed implementation manners

[0139] The embodiments of the present invention provide a training method, an evaluation method and related devices for a chorus song recognition network, which are used to first train the recognition network for evaluating chorus songs, and further use the chorus song recognition network to score the timbre tacit understanding of chorus users, and use the chorus tacit understanding score to evaluate the chorus songs of choristers, so as to improve the pertinence and accuracy of the evaluation of choristers' chorus songs.

[0140] To enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0141] The terms "first", "second", "third", "fourth", etc. in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described herein can be implemented in an order different from that shown or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units need not be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0142] For ease of understanding, the training method of the chorus song recognition network in the embodiments of the present application will be described below. Please refer to Figure 1 , Figure 1 This is an embodiment of the training method of the chorus song recognition network in the embodiments of the present application:

[0143] 101. Obtain a first training sample with labels, where the first training sample includes Mel spectrograms of singing samples of multiple chorus users singing multiple chorus songs and timbre vectors of multiple chorus users;

[0144] Before training the chorus song recognition network, it is generally necessary to first obtain a training sample. Among them, the training sample can be a training sample with labels or a training sample without labels.

[0145] Specifically, in the embodiments of the present application, a first training sample with labels is used. The first training sample includes Mel spectrograms of singing samples of multiple chorus users singing multiple chorus songs and timbre vectors of multiple chorus users.

[0146] If user A and user B sang "Little Donkey" in chorus, user C and user D sang "Little Swallow" in chorus, user E and user F sang "Catching Loaches" in chorus, etc., then the training sample corresponding to the chorus song "Little Donkey" is the Mel spectrogram of the singing samples of user A and user B singing "Little Donkey", the timbre vector of user A, and the timbre vector of user B. Among them, the singing sample of A comes with the label of user A, the singing sample of B comes with the label of user B, the timbre vector of A comes with the label of user A, and the timbre vector of B comes with the label of user B. The same is true for other chorus songs and will not be elaborated here.

[0147] In addition, the singing sample of user A singing "Little Donkey" can be a singing segment (such as one line of lyrics, two lines of lyrics, or three lines of lyrics) sample, or it can be a singing sample of the whole song. There is no limit on the size form of the singing sample here.

[0148] It should be noted that in order to better train the parameters of the chorus song recognition network and improve the accuracy of the output of the chorus song recognition network, in the embodiments of this application, generally, the intra-class distance between the Mel spectrogram vectors of the singing samples of the same chorus users is less than or equal to a preset reference value, and the inter-class distance between the Mel spectrogram vectors of the singing samples of different chorus users is greater than the preset reference value. For example, the distance between the Mel spectrogram vectors of the singing segments of user A singing "Little Donkey" is less than or equal to the preset reference value, while the distance between the Mel spectrogram vector of the singing segment of user A and the Mel spectrogram vector of the singing segment of user B is greater than the preset reference value, so that the chorus song recognition network can better distinguish the timbre of A and the timbre of B.

[0149] For ease of understanding, Figure 2 a schematic diagram of the intra-class distance between the Mel spectrogram vectors of the singing segments of user A, the intra-class distance between the Mel spectrogram vectors of the singing segments of user B, and the inter-class distance between the Mel spectrogram vector of the singing segment of user A and the Mel spectrogram vector of the singing segment of user B is given.

[0150] 102. Input the Mel spectrogram of the singing sample of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network;

[0151] During the training process, in order to train the parameters of the chorus song recognition network, generally, the Mel spectrogram of the singing sample of each chorus user singing the same chorus song is input into the chorus song recognition network. For example, for the song "Little Donkey" corresponding to step 101, the Mel spectrogram of the singing segment of user A singing "Little Donkey" is input into the chorus song recognition network to obtain the timbre vector of user A output by the chorus song recognition network, and the Mel spectrogram of the singing segment of user B singing "Little Donkey" is input into the chorus song recognition network to obtain the timbre vector of user B output by the chorus song recognition network.

[0152] During the training process, generally, the larger the number of training samples, the more accurate the parameters of the chorus song recognition network are trained. Therefore, the larger the voice sample size of the chorus song samples in this application, the better.

[0153] 103. According to the user tags, obtain the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the training samples;

[0154] Since the training samples in the embodiments of this application carry their own tags, after the chorus recognition network outputs the timbre vectors of each user, the timbre vectors of the same user can be further found from the training samples according to the user tags.

[0155] For example, after the chorus song recognition network outputs the timbre vector of user A, use the tag of user A to find the timbre vector of user A from the training samples to facilitate the execution of step 104.

[0156] 104. Use a loss function to calculate the difference between the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training samples;

[0157] After obtaining the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the training samples, use a loss function to calculate the difference between the timbre vectors of each chorus user and the timbre vectors of the same chorus user in the training samples. For example, use a loss function to calculate the difference between the timbre vector of user A output by the chorus song recognition network and the timbre vector of user A in the training samples.

[0158] Among them, the loss function can be selected as the absolute value loss function, Log logarithmic loss function, square loss function, exponential loss function, etc. according to the specific scenario, and no specific limitation is made here.

[0159] 105. Correct the chorus song recognition network according to the difference and the backpropagation algorithm.

[0160] After calculating the difference according to the loss function, use the difference and the backpropagation algorithm to correct the parameters of the chorus song recognition network.

[0161] Specifically, the process of how to correct the parameters of the chorus song recognition network according to the difference and the backpropagation algorithm is similar to that described in the prior art and will not be elaborated here.

[0162] In the embodiments of the present application, the chorus song recognition network is trained using the first training samples, and the first training samples in the embodiments of the present application adopt the Mel spectrogram of the chorus user's singing voice sample and the timbre vector of the chorus user, so that after the Mel spectrogram of the chorus user's singing voice sample is input into the chorus song recognition network, the timbre vector of the user can be obtained, which improves the convenience of obtaining the user's timbre characteristics in the chorus scenario.

[0163] Based on Figure 1 the above-mentioned embodiments, the structure of the chorus song recognition network will be described in detail below. Specifically, the structure of the chorus song recognition network in the embodiments of the present application includes a long short-term memory network and a fully connected layer. Please refer to Figure 3 , Figure 3 FIG. is a schematic diagram of the process of inputting the chorus user's singing voice sample into the chorus song recognition network:

[0164] Input the Mel spectrogram of each chorus user's singing voice sample singing the same chorus song into the long short-term memory network to extract the eigenvalue of each chorus user's timbre vector, and input the eigenvalue of each chorus user's timbre vector into the fully connected layer to integrate and output the eigenvalue of each chorus user's timbre vector, so as to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

[0165] Specifically, if the Mel spectrogram of user A's singing voice sample singing "Little Donkey" is input into the long short-term memory network, the timbre characteristics of user A in different frequency bands can be extracted, and then the timbre characteristics of user A in different frequency bands are input into the fully connected layer. The fully connected layer will integrate and output the extracted timbre characteristics of user A in different frequency bands, so as to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

[0166] For the specific verification process of the long short-term memory network and the fully connected network, it is similar to the description in the prior art and will not be elaborated here.

[0167] Further, when the structure of the chorus song recognition network includes a long short-term memory network and a fully connected layer, the process of correcting the chorus song recognition network according to the difference and the backpropagation algorithm is also the process of correcting the parameters of the long short-term memory network and the fully connected layer according to the difference and the backpropagation algorithm.

[0168] For the specific verification process of the above correction process, it is similar to the description in the prior art and will not be elaborated here either.

[0169] Based on Figures 1 to 3 the above embodiments, the synthesized song recognition network in the embodiments of the present application, in addition to including a long short-term memory network and a fully connected layer for obtaining the user's timbre vector, also includes a fully connected network for scoring the tacit understanding of the timbre vectors of the chorus users.

[0170] Next, the training process of the fully connected network in the embodiments of the present application will be described. Please refer to Figure 4 Another embodiment of the training method of the chorus song recognition network in the embodiments of the present application includes:

[0171] 401. Obtain labeled second training samples, where the second training samples include timbre vectors of multiple chorus users singing multiple chorus songs, and timbre matching scores of multiple chorus users in each chorus song;

[0172] Furthermore, in addition to obtaining the timbre vectors of each chorus user in the chorus song, the embodiments of the present application can also score the timbre matching degree of multiple chorus users according to the timbre vectors of multiple chorus users.

[0173] Therefore, the embodiments of the present application further train the fully connected network so that the fully connected network can score the timbre matching degree of multiple chorus users.

[0174] The embodiments of the present application first obtain labeled second training samples, where the second training samples include timbre vectors of multiple chorus users singing multiple chorus songs, and timbre matching scores of multiple chorus users in each song.

[0175] For ease of explanation, the foregoing example will be used for explanation below:

[0176] For example, the second training samples can be the timbre vectors of A and B in "Little Donkey" respectively, and the timbre matching scores of A and B in "Little Donkey", the timbre vectors of C and D in "Little Swallow" respectively, and the timbre matching scores of C and D in "Little Swallow", the timbre vectors of E and F in "Catching Loaches" respectively, and the timbre matching scores of C and D in "Catching Loaches".

[0177] Of course, each training sample carries its own user label.

[0178] 402. Concatenate the timbre vectors of the song segments of multiple chorus users singing the same chorus song to obtain the concatenated timbre vector;

[0179] After obtaining the timbre vectors of multiple chorus users in the same chorus song, concatenate the timbre vectors of multiple chorus users to obtain the concatenated timbre vector.

[0180] During splicing, the timbre vector of each chorus user may be averaged in the time dimension to obtain a timbre vector of each chorus user that does not contain timing features, and the timbre vectors of each user that do not contain timing features may be spliced; or the timbre vectors of at least two chorus users may be superimposed in the time dimension, and then the superimposed timbre vectors may be averaged according to the number of chorus users to obtain a spliced ​​timbre vector that contains timing features.

[0181] Here are some examples:

[0182] Assume that the timbre vector of user A in "Little Donkey" is 182*256, where 182 represents the number of segments sung by user A in "Little Donkey", each segment is 1ms, and 256 is the dimension of the vector; the timbre vector of user B in "Little Donkey" is also 182*256.

[0183] Then when splicing, you can first average the timbre vector of A in the time dimension to obtain a 1*256 timbre vector of A, then average the timbre vector of B in the time dimension to obtain a 1*256 timbre vector of B, and then concatenate the timbre vectors of A and B to obtain a 1*512 timbre vector.

[0184] In addition, the following concatenation can also be performed, such as superimposing the timbre vector of user A and the timbre vector of user B to obtain a timbre vector of 2*182*256, and then averaging these two vectors according to the number of users to obtain a timbre vector of 1*182*256.

[0185] Obviously, it can be seen that the timbre vector in the first concatenation method does not contain temporal features, while the timbre vector in the second concatenation method contains temporal features.

[0186] 403. Inputting the concatenated timbre vector into a fully connected network to obtain a chorus tacit understanding score output by the fully connected network;

[0187] By inputting the concatenated timbre vector into the fully connected network, the tacit understanding score of the chorus members output by the fully connected network can be obtained.

[0188] 404. Acquire, according to the user tag, the tacit understanding score of the chorus members output by the fully connected network and the tacit understanding score of the same chorus members in the second training sample;

[0189] Because the second training samples all carry user labels, after obtaining the tacit understanding scores of the chorus members output by the fully connected network, the tacit understanding scores of the chorus members in the second training samples can also be obtained according to the user labels.

[0190] If the timbre vectors of User A and User B after splicing are input into the fully connected network, the compatibility scores of User A and User B output by the fully connected network can be obtained. In addition, according to the user tags, the compatibility scores of User A and User B in the second training sample can also be obtained.

[0191] 405. Calculate the difference between the chorus singer compatibility score output by the fully connected network and the compatibility score of the same chorus singer in the second training sample using a loss function;

[0192] Furthermore, after obtaining the chorus singer compatibility score output by the fully connected network and the compatibility score of the same chorus singer in the second training sample, use a loss function to calculate the difference between the chorus singer compatibility score output by the fully connected network and the compatibility score of the same chorus singer in the second training sample. For example, calculate the difference between the compatibility scores of User A and User B output by the fully connected network and the compatibility scores of User A and User B in the second training sample using a loss function.

[0193] 406. Correct the fully connected network according to the difference and the backpropagation algorithm.

[0194] After obtaining the difference, correct the parameters of the fully connected network according to the difference and the backpropagation algorithm. The specific verification process during the correction process is also similar to the prior art and will not be elaborated here.

[0195] In the embodiment of the present application, the fully connected network is trained using the second training sample, and the second training sample in the embodiment of the present application uses the timbre vectors of chorus users and the timbre compatibility scores of chorus users, so that after the timbre vectors of chorus users are input into the fully connected network, the timbre compatibility scores of chorus users can be obtained, improving the convenience of obtaining timbre compatibility scores for users in the chorus scenario.

[0196] Based on Figure 4 the above-mentioned fully connected network, where the fully connected network includes a feature layer, a pooling layer, and a fully connected layer. The process of inputting the timbre vectors of spliced chorus singers into the fully connected network is described below. Please refer to Figure 5 :

[0197] Input the spliced timbre vector into the feature layer to extract the timbre feature values of the spliced timbre vector; input the timbre feature values of the spliced timbre vector into the pooling layer to reduce redundancy and dimensionality of the timbre feature values of the spliced timbre vector; input the timbre feature values of the spliced timbre vector after redundancy reduction and dimensionality reduction into the fully connected layer to integrate and output the feature values of the spliced timbre vector after redundancy reduction and dimensionality reduction to obtain the timbre compatibility score of the chorus singers singing the same chorus song.

[0198] Specifically, after inputting the spliced timbre vectors into the feature layer (such as splicing the timbre vectors of user A and user B), the feature layer extracts the eigenvalue of the spliced timbre vectors through convolution, that is, the timbre features of the spliced vectors at different frequencies. In order to avoid the problem of large computational complexity caused by redundant extraction of timbre eigenvalues, the embodiment of the present application inputs the timbre eigenvalues of the spliced timbre vectors into the pooling layer to reduce redundancy and dimensionality of the timbre eigenvalues of the spliced timbre vectors. Among them, the operation of the pooling layer can be maximum pooling, average pooling, etc., and the specific operation process of the pooling layer is not limited here.

[0199] After obtaining the timbre eigenvalues of the spliced timbre vectors with redundancy reduced and dimensionality decreased, input them into the fully connected layer to integrate and output the eigenvalues of the spliced timbre vectors with redundancy reduced and dimensionality decreased, so as to obtain the timbre matching degree score of the chorus singers singing the same chorus song (such as the timbre matching degree score of user A and user B).

[0200] Furthermore, when the fully connected network includes a feature layer, a pooling layer and a fully connected layer, when using the difference and backpropagation algorithm to correct the fully connected network, the parameters of the feature layer, the pooling layer and the fully connected layer are corrected by using the difference and backpropagation algorithm.

[0201] The above describes the training of the chorus song recognition network in the embodiment of the present application. Next, the evaluation method of the chorus song in the embodiment of the present application will be described. Please refer to Figure 6 An embodiment of the chorus song evaluation method in the embodiment of the present application includes:

[0202] 601. Obtain the singing segments corresponding to the chorus singers in the chorus song, where the chorus singers include at least two chorus users;

[0203] Specifically, when evaluating the chorus song of the chorus singers, it is necessary to first obtain the singing segments corresponding to the chorus singers in the chorus song, where the chorus singers include at least two chorus users.

[0204] For example, when evaluating the chorus of "Little Donkey" by user A and user B, it is necessary to first obtain the singing segments of user A and user B in the song "Little Donkey". Among them, the singing segment can be a single sentence of lyrics, multiple sentences of lyrics or the whole song, etc., and the size of the singing segment is not specifically limited here.

[0205] 602. Extract the Mel spectrogram of the singing segment corresponding to each chorus user;

[0206] After obtaining the singing segments of each chorus user in the chorus song, further extract the Mel spectrogram of the singing segment corresponding to each chorus user.

[0207] 603. Inputting the mel spectrum of the corresponding singing segment of each chorus user into the chorus song recognition network to obtain the timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user;

[0208] After obtaining the mel-spectrogram of the singing segment of each chorus user, the mel-spectrogram of the corresponding singing segment of each chorus user is input into the chorus song recognition network to obtain the timbre vector of each chorus user.

[0209] Specifically, the process of step 603 is the same as Figure 3 The description of the embodiments is similar and will not be repeated here.

[0210] 604. Splicing the timbre vectors of each chorus user to obtain a spliced ​​timbre vector;

[0211] After the timbre vector of each chorus user is obtained, the timbre vectors of each chorus user are further concatenated to obtain a concatenated timbre vector.

[0212] Specifically, the description of the concatenated timbre vector here is similar to Figure 4 The description of step 402 in the embodiment is similar and will not be repeated here.

[0213] 605. Input the concatenated timbre vector to a feature layer of a fully connected network to extract a feature value of the concatenated timbre vector;

[0214] 606. Inputting the eigenvalues ​​of the concatenated timbre vectors into a pooling layer of a fully connected network to perform redundancy removal and dimensionality reduction on the eigenvalues ​​of the concatenated timbre vectors;

[0215] 607. Inputting the eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction, and obtaining the timbre tacit understanding scores of the chorus members singing the same chorus song;

[0216] It should be noted that the description of step 605 to step 607 is the same as Figure 5 The description of the embodiment is similar and will not be repeated here.

[0217] 608. Evaluate the chorus song based on the chorus members’ timbre tacit understanding scores.

[0218] By obtaining the timbre harmony scores of the chorus members singing the same chorus song, the chorus song of the chorus members can be evaluated according to the timbre harmony scores of the chorus members.

[0219] In the embodiments of the present application, in a chorus scenario, a long short-term memory network and a fully connected layer in a chorus song recognition network are used to obtain the timbre vectors of chorus users. Then, the timbre vectors of the chorus users are concatenated to obtain a concatenated timbre vector. Further, the concatenated timbre vector is input into a fully connected network to obtain the timbre matching score of the choristers, that is, the embodiments of the present application can evaluate the chorus songs of the chorus users through the timbre matching score of the chorus users, thereby improving the pertinence and accuracy of the evaluation of chorus songs.

[0220] Based on Figure 6 In the above-mentioned embodiment, after obtaining the timbre matching score of the choristers for the chorus song, the integrity score of the choristers for the chorus song can be further calculated, and the chorus song can be evaluated using the timbre matching score and the integrity score of the choristers. Please refer to Figure 7 , Figure 7 Another embodiment of the method for evaluating a chorus song:

[0221] 701. Obtain the singing segments corresponding to the choristers in the chorus song, where the choristers include at least two chorus users;

[0222] 702. Extract the Mel spectrograms of the singing segments corresponding to each chorus user;

[0223] 703. Input the Mel spectrograms of the singing segments corresponding to each chorus user into the chorus song recognition network to obtain the timbre vectors of each chorus user. The chorus song recognition network includes a long short-term memory network and a fully connected layer. The long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user;

[0224] 704. Concatenate the timbre vectors of each chorus user to obtain a concatenated timbre vector;

[0225] 705. Input the concatenated timbre vector into the feature layer of the fully connected network to extract the eigenvalue of the concatenated timbre vector;

[0226] 706. Input the eigenvalue of the concatenated timbre vector into the pooling layer of the fully connected network to reduce redundancy and dimension of the eigenvalue of the concatenated timbre vector;

[0227] 707. Input the eigenvalue of the concatenated timbre vector after redundancy reduction and dimension reduction into the fully connected layer of the fully connected network to integrate and output the eigenvalue of the concatenated timbre vector after redundancy reduction and dimension reduction, and obtain the timbre matching score of the choristers singing the same chorus song;

[0228] It should be noted that steps 701 to 707 in the embodiments of the present application are similar to the descriptions of steps 601 to 607 in Figure 6 the embodiments, and will not be elaborated here.

[0229] 708. Obtain the chorus interval of the choristers for the same chorus song;

[0230] In addition to calculating the timbre tacit understanding score of the choristers singing the same chorus song, the embodiments of the present application can also calculate the integrity score of the choristers singing the same chorus song.

[0231] Specifically, when calculating the integrity score of the choristers singing the same chorus song, it is necessary to first obtain the chorus interval of the choristers for the same chorus song.

[0232] Among them, when obtaining the synthesis interval of the choristers for the same chorus song, the voice activity detection (VAD) method can be used to obtain the chorus start time and chorus end time of the choristers in the same chorus song, and then according to the chorus start time and the chorus end time, determine the chorus interval of the choristers within the same chorus song.

[0233] For ease of understanding, Figure 8 a schematic diagram of the chorus interval of user A and user B in the same chorus song is given.

[0234] 709. Calculate the integrity score of each chorus user within the chorus interval;

[0235] After obtaining the chorus interval of the chorus users in the same chorus song, the embodiments of the present application calculate the integrity score of each chorus user within the chorus interval.

[0236] Specifically, the process of calculating the integrity score of each chorus user within the chorus interval will be described in the following embodiments and will not be elaborated here.

[0237] 710. Calculate the integrity score of the choristers for the same chorus song according to the integrity scores of each chorus user within the chorus interval.

[0238] After obtaining the integrity scores of each chorus user within the chorus interval, the integrity score of the choristers for the same chorus song can be obtained.

[0239] The following is an example:

[0240] Suppose there are 6 lines of lyrics in the song "Little Donkey". The integrity score of user A within the chorus interval is 8 points, and the integrity score of user B within the chorus interval is 9 points. Then the integrity score of A and B within the chorus interval is (8 + 9) / 2 = 8.5 points.

[0241] 711. The chorus songs sung by the chorus members are evaluated based on their timbre harmony scores and completeness scores for the same chorus songs.

[0242] After obtaining the timbre harmony scores and the completeness scores of the chorus members for the same chorus song, the chorus song of the chorus members is evaluated according to the timbre harmony scores and the completeness scores of the chorus members for the same chorus song.

[0243] The embodiment of the present application evaluates the chorus songs sung by the chorus members from two dimensions: the tacit understanding score of the chorus members and the completeness score of the chorus members, thereby improving the professionalism and accuracy of the evaluation.

[0244] based on Figure 7 The following describes step 709 in detail. Figure 9 , Figure 9 for Figure 7 The detailed steps of step 709 in the embodiment are as follows:

[0245] 901. Using a voice activity detection (VAD) method, obtain a vocalization segment of each chorus user in the chorus interval, and a vocalization duration of each line of lyrics in each vocalization segment;

[0246] Specifically, after obtaining the chorus interval of the chorus users singing the same chorus song, the voice activity detection (VAD) method is further used to obtain the vocalization segment of each chorus user in the chorus interval, and the vocalization duration of each line of lyrics in each vocalization segment.

[0247] For example, the voice activity detection (VAD) method can be used to obtain the vocal segments of user A in "Little Donkey", and the utterance duration of each line of lyrics of user A in the vocal segments can be detected based on the vocal segments. For example, the vocal segments of user A in "Little Donkey" are the first line of lyrics, the third line of lyrics, and the fifth line of lyrics. The utterance duration of user A in the first line of lyrics is 5ms, the utterance duration in the third line of lyrics is 4ms, and the utterance duration in the fifth line of lyrics is 7ms.

[0248] 902. Obtain a voice file of the chorus song, and obtain the original singer's voice duration of each line of lyrics from the voice file;

[0249] In order to calculate the completeness score of each chorus user within the chorus interval, it is also necessary to obtain the voice file of the chorus song (such as the MIDI file of the chorus song) and obtain the original singing duration of each line of lyrics from the voice file.

[0250] Generally, the MIDI file of a song includes the start time and end time of each line of lyrics. Therefore, the original singing duration of each line of lyrics can be obtained from the MIDI file of the song.

[0251] 903. Calculate the single-line integrity score of each chorus user within the chorus interval based on the singing duration of each line of lyrics of each chorus user within the chorus interval and the original singing duration of the corresponding lyrics.

[0252] After obtaining the singing duration of each line of lyrics of each chorus user within the chorus interval and the original singing duration of the corresponding lyrics, for example, if the singing duration of user A in the first line of lyrics is 5 ms, while the singing duration of the corresponding lyrics in the voice file is 7 ms, it means that user A's singing of the first line of lyrics covers 5 / 7 of the original lyrics' singing duration.

[0253] Furthermore, the single-line integrity score of a user's singing of a song can be calculated according to the following formula. For example, set a coverage coefficient a. If the single-line singing duration coverage rate of the user is greater than or equal to a, then obtain the pre-set single-line integrity score. If the single-line singing duration coverage rate of the user is less than a, then obtain the corresponding single-line integrity score, where both a and β are pre-set custom coefficients:

[0254]

[0255] 904. Calculate the integrity score of each chorus user within the chorus interval based on the single-line integrity score of each chorus user for the chorus song.

[0256] After obtaining the single-line integrity score of each chorus user within the chorus interval, then calculate the integrity score of each chorus user within the chorus interval respectively.

[0257] For example, assume that the singing segments of user A in "Little Donkey" are the first line of lyrics, the third line of lyrics, and the fifth line of lyrics. And the integrity score of user A's first line of lyrics is x1, the integrity score of the third line of lyrics is x2, and the integrity score of the fifth line of lyrics is x3. Then the integrity score of user A within the chorus interval is (x1 + x2 + x3) / 3.

[0258] The above embodiments have described in detail the calculation process of the integrity score of each chorus user within the chorus interval, thereby improving the accuracy of calculating the integrity score of chorus users within the chorus interval.

[0259] Based on Figure 7In the embodiment described above, after obtaining the chorus singers' timbre tacit understanding scores and integrity scores for the chorus song, the chorus singers' singing skills matching scores for the same chorus song can be further calculated, and the chorus song can be evaluated using the chorus singers' timbre tacit understanding scores, integrity scores and singing skills matching scores. Figure 10 , Figure 10 Another embodiment of the chorus song evaluation method:

[0260] 1001. Obtain singing segments corresponding to chorus members in a chorus song, wherein the chorus members include at least two chorus users;

[0261] 1002. Extracting the mel-score of the singing segment corresponding to each chorus user;

[0262] 1003. Inputting the mel spectrum of the singing segment corresponding to each chorus user into a chorus song recognition network to obtain a timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user;

[0263] 1004. Splicing the timbre vectors of each chorus user to obtain a spliced ​​timbre vector;

[0264] 1005. Input the concatenated timbre vector to a feature layer of a fully connected network to extract a feature value of the concatenated timbre vector;

[0265] 1006. Inputting the eigenvalues ​​of the concatenated timbre vectors into a pooling layer of a fully connected network to remove redundancy and reduce the dimensionality of the eigenvalues ​​of the concatenated timbre vectors;

[0266] 1007. Inputting the eigenvalues ​​of the concatenated timbre vectors after redundancy reduction into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the concatenated timbre vectors after redundancy reduction, and obtaining the timbre harmony scores of the chorus members singing the same chorus song;

[0267] It should be noted that steps 1001 to 1007 in the embodiment of the present application are similar to Figure 6 The descriptions of steps 601 to 607 in the embodiment are similar and will not be repeated here.

[0268] 1008. Obtaining the chorus interval of the chorus singers for the same chorus song;

[0269] 1009. Calculate the completeness score of each chorus user in the chorus interval;

[0270] 1010. Calculate the integrity score of the choristers for the same choral song according to the integrity scores of each choral user within the choral interval.

[0271] It should be noted that steps 1008 to 1010 in the embodiments of the present application are similar to those described in Figure 7 the embodiments, and will not be elaborated here.

[0272] 1011. Extract the pitch values of the singing segments of each choral user;

[0273] Furthermore, the embodiments of the present application can also evaluate the choral songs of the choristers from the perspective of singing skills matching. Specifically, in the embodiments of the present application, the choral songs of the choristers are evaluated by matching the pitch melodies of the choral users.

[0274] Specifically, the embodiments of the present application can extract the pitch values of the singing segments of each choral user, where the pitch melody can be calculated by a pitch extraction algorithm (such as PYin).

[0275] Of course, other methods can also be used for singing skills matching, such as note-level matching degree, or reference-free singing skills evaluation based on deep learning, etc., which are not specifically limited here.

[0276] 1012. Obtain the voice file of the choral song, and obtain the original pitch values of each line of lyrics from the voice file;

[0277] In order to obtain the pitch accuracy scores of each choral user, the embodiments of the present application also need to obtain the voice file of the choral song (such as MIDI file) to obtain the original pitch values of each line of lyrics from the voice file.

[0278] 1013. According to the pitch values of the singing segments of each choral user and the original pitch values of the corresponding lyric segments of the singing segments, obtain the singing skills matching degree scores of multiple choral users for the same choral song;

[0279] After obtaining the pitch values of the singing segments of each choral user and the original pitch values corresponding to the singing segments, the singing skills scores of each choral user for the same choral song can be calculated by using a sequence matching method (such as similarity method), such as mapping the difference value between the two sequences to the pitch accuracy score, that is, the singing skills score.

[0280] After obtaining the singing skills scores of each choral user for the same choral song (such as S A and S B ), the singing skills matching degree score S2 between multiple choral users can be calculated by using the following formula:

[0281]

[0282] 1014. Evaluate the chorus song according to the timbre harmony score, the completeness score and the singing skill matching score of the chorus members when singing the same chorus song.

[0283] After obtaining the timbre tacit understanding scores, completeness scores and singing skills matching scores of the chorus members singing the same chorus song, the chorus song of the chorus members can be evaluated based on any two or three of the timbre tacit understanding scores, completeness scores and singing skills matching scores of the chorus members singing the same chorus song.

[0284] In the embodiment of the present application, after calculating the timbre tacit understanding score, completeness score and singing skill matching score of the chorus users for the same chorus song, two or three of the above scores are used to evaluate the chorus song of the chorus users, thereby improving the accuracy of the chorus song evaluation.

[0285] The above describes in detail the training method of the chorus song recognition network and the evaluation method of the chorus song in the embodiment of the present application. Next, the training device of the chorus song recognition network and the evaluation device of the chorus song in the embodiment of the present application are described. Please refer to Figure 11 , an embodiment of a training device for a chorus song recognition network in an embodiment of the present application includes:

[0286] A first acquisition unit 1101 is used to acquire a first labeled training sample, wherein the first training sample includes mel spectra of vocal samples of multiple chorus songs sung by multiple chorus users and timbre vectors of multiple chorus users;

[0287] The first input unit 1102 is used to input the mel spectrum of the singing voice sample of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network;

[0288] The first acquisition unit 1101 is further used to acquire the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample according to the user label;

[0289] A first calculation unit 1103 is used to calculate the difference between the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample by using a loss function;

[0290] The first correction unit 1104 is used to correct the chorus song recognition network according to the difference and back propagation algorithm.

[0291] Preferably, the chorus song recognition network includes a long short-term memory network and a fully connected layer;

[0292] The first input unit 1102 is specifically configured to:

[0293] Input the Mel spectrogram of the singing voice samples of each chorus user singing the same chorus song into the long short-term memory network to extract the eigenvalue of the timbre vector of each chorus user;

[0294] Input the eigenvalue of the timbre vector of each chorus user into the fully connected layer to integrate and output the eigenvalue of the timbre vector of each chorus user, so as to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

[0295] The first correction unit 1104 is specifically configured to:

[0296] Correct the long short-term memory network and the fully connected layer according to the difference and the backpropagation algorithm.

[0297] Preferably, the intra-class distance between the Mel spectrogram vectors of the singing voice samples of the same chorus user is less than or equal to a preset reference value, and the inter-class distance between the Mel spectrogram vectors of the singing voice samples of different chorus users is greater than the preset reference value.

[0298] Preferably, the device further includes:

[0299] A second acquisition unit 1105, configured to acquire a labeled second training sample, where the second training sample includes the timbre vectors of multiple chorus users singing multiple chorus songs, and the timbre tacit understanding scores of multiple chorus users in each chorus song;

[0300] A splicing unit 1106, configured to splice the timbre vectors of the song segments of multiple chorus users singing the same chorus song to obtain a spliced timbre vector;

[0301] A second input unit 1107, configured to input the spliced timbre vector into a fully connected network to obtain the timbre tacit understanding score of the chorus singer output by the fully connected network;

[0302] The second acquisition unit 1105 is further configured to acquire the timbre tacit understanding score of the chorus singer output by the fully connected network and the timbre tacit understanding score of the same chorus singer in the second training sample according to the user label;

[0303] A second calculation unit 1108, configured to calculate the difference between the timbre tacit understanding score of the chorus singer output by the fully connected network and the timbre tacit understanding score of the same chorus singer in the second training sample by using a loss function;

[0304] A second correction unit 1109, configured to correct the fully connected network according to the difference and the backpropagation algorithm.

[0305] Preferably, the fully connected network includes a feature layer, a pooling layer and a fully connected layer;

[0306] The second input unit 1107 is specifically used for:

[0307] Inputting the concatenated timbre vector to the feature layer to extract the timbre feature value of the concatenated timbre vector;

[0308] Inputting the timbre feature value of the concatenated timbre vector into the pooling layer to remove redundancy and reduce the dimension of the timbre feature value of the concatenated timbre vector;

[0309] The timbre eigenvalues ​​of the concatenated timbre vectors after redundancy reduction are input into the fully connected layer for integrating and outputting the eigenvalues ​​of the concatenated timbre vectors after redundancy reduction, so as to obtain the timbre tacit understanding scores of the chorus members singing the same chorus song.

[0310] The second correction unit 1109 is specifically used for:

[0311] The feature layer, the pooling layer and the fully connected layer are modified according to the difference and back propagation algorithm.

[0312] The functions of each unit in the embodiment of the present application are Figures 1 to 5 The description in the embodiment is similar and will not be repeated here.

[0313] Next, the evaluation device for choral songs in the embodiment of the present application is described. Figure 12 In an embodiment of the present application, a chorus song evaluation device includes:

[0314] The acquisition unit 1201 is used to acquire a singing segment corresponding to a chorus member in a chorus song, wherein the chorus member includes at least two chorus users;

[0315] An extraction unit 1202 is used to extract the mel spectrum of the singing segment corresponding to each chorus user;

[0316] An input unit 1203 is used to input the mel spectrum of the singing segment corresponding to each chorus user into a chorus song recognition network to obtain a timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user;

[0317] A splicing unit 1204, configured to splice the timbre vectors of each chorus user to obtain a spliced ​​timbre vector;

[0318] The extraction unit 1202 is further configured to input the spliced timbre vectors into the feature layer of the fully connected network to extract the eigenvalue of the spliced timbre vectors;

[0319] The input unit 1203 is further configured to input the eigenvalue of the spliced timbre vectors into the pooling layer of the fully connected network to perform redundancy reduction and dimensionality reduction on the eigenvalue of the spliced timbre vectors;

[0320] The input unit 1203 is further configured to input the eigenvalue of the spliced timbre vectors after redundancy reduction and dimensionality reduction into the fully connected layer of the fully connected network to integrate and output the eigenvalue of the spliced timbre vectors after redundancy reduction and dimensionality reduction, so as to obtain the timbre tacit understanding score of the choristers singing the same choral song;

[0321] The evaluation unit 1205 is configured to evaluate the choral song of the choristers according to the timbre tacit understanding score of the choristers.

[0322] Preferably, the splicing unit 1204 is specifically configured to:

[0323] Average the timbre vectors of each choral user in the time dimension to obtain the timbre vectors of each choral user without temporal features, and splice the timbre vectors of each user without temporal features.

[0324] Preferably, the splicing unit 1204 is specifically configured to:

[0325] Superimpose the timbre vectors of at least two choral users in the time dimension, and then average the superimposed timbre vectors according to the number of choral users to obtain the spliced timbre vectors including temporal features.

[0326] Preferably, the acquisition unit 1201 is further configured to:

[0327] Obtain the choral interval of the choristers for the same choral song;

[0328] The evaluation device further includes a calculation unit 1206, configured to calculate the integrity score of each choral user within the choral interval; according to the integrity score of each choral user within the choral interval, calculate the integrity score of the choristers for the same choral song.

[0329] Preferably, the acquisition unit 1201 is specifically configured to:

[0330] Use the voice activity detection VAD method to obtain the choral start time and choral end time of the choristers in the same choral song;

[0331] According to the choral start time and the choral end time, determine the choral interval of the choristers within the same choral song.

[0332] Preferably, the computing unit 1206 is specifically configured to:

[0333] Using a voice activity detection (VAD) method, obtaining the vocalization segments of each chorus user in the chorus interval, and the vocalization duration of each line of lyrics in each vocalization segment;

[0334] Obtain a voice file of the chorus song, and obtain the original singer's voice duration of each line of lyrics from the voice file;

[0335] Calculate the single sentence completeness score of each chorus user in the chorus interval according to the occurrence duration of each line of lyrics of each chorus user in the chorus interval and the original singing duration of the corresponding lyrics;

[0336] According to the single sentence completeness score of each chorus user for the chorus song, the completeness score of each chorus user in the chorus interval is calculated.

[0337] The evaluation unit 1205 is specifically used for:

[0338] The chorus songs of the chorus members are evaluated according to the timbre tacit understanding scores and the completeness scores of the chorus members for the same chorus song.

[0339] The evaluation unit 1205 is further used for:

[0340] Extract the pitch value of each chorus user's singing segment;

[0341] Obtain a voice file of a chorus song, and obtain the original singing pitch value of each line of lyrics from the voice file;

[0342] According to the pitch value of each chorus user's singing segment and the original singer's pitch value of the lyrics segment corresponding to the singing segment, the singing skill matching scores of multiple chorus users singing the same chorus song are obtained;

[0343] The chorus songs sung by the chorus are evaluated according to the timbre tacit understanding scores and the singing skills matching scores of the chorus members when singing the same chorus song.

[0344] The functions of each unit in the embodiment of the present application are Figures 6 to 10 The description is similar to that in , and will not be repeated here.

[0345] The above describes the training device of the chorus song recognition network and the evaluation device of the chorus song in the embodiment of the present invention from the perspective of modular functional entities. The following describes the computer device in the embodiment of the present invention from the perspective of hardware processing:

[0346] The computer device is used to implement the functions of the training device for the chorus song recognition network. In an embodiment of the present invention, an embodiment of the computer device includes:

[0347] a processor and a memory;

[0348] The memory is used to store a computer program. When the processor executes the computer program stored in the memory, the following steps can be implemented:

[0349] Obtain a labeled first training sample, where the first training sample includes Mel spectrograms of singing samples of multiple chorus users singing multiple chorus songs and timbre vectors of multiple chorus users;

[0350] Input the Mel spectrograms of the singing samples of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vectors of the corresponding chorus users output by the chorus song recognition network;

[0351] According to the user label, obtain the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training sample;

[0352] Use a loss function to calculate the difference between the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training sample;

[0353] Correct the chorus song recognition network according to the difference and the backpropagation algorithm

[0354] In some embodiments of the present invention, the chorus song recognition network includes a long short-term memory network and a fully connected layer. The processor can also be used to implement the following steps:

[0355] Input the Mel spectrograms of the singing samples of each chorus user singing the same chorus song into the long short-term memory network to extract the eigenvalue of the timbre vector of each chorus user;

[0356] Input the eigenvalue of the timbre vector of each chorus user into the fully connected layer to integrate and output the eigenvalue of the timbre vector of each chorus user, and obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

[0357] In some embodiments of the present invention, the processor can also be used to implement the following steps:

[0358] Correct the long short-term memory network and the fully connected layer according to the difference and the backpropagation algorithm.

[0359] In some embodiments of the present invention, the intra-class distance between the Mel-spectrum vectors of the singing samples of the same chorus user is less than or equal to a preset reference value, and the inter-class distance between the Mel-spectrum vectors of the singing samples of different chorus users is greater than the preset reference value.

[0360] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0361] Obtain labeled second training samples, where the second training samples include timbre vectors of multiple chorus users singing multiple chorus songs, and the timbre tacit understanding scores of multiple chorus users in each chorus song;

[0362] Concatenate the timbre vectors of the song segments of multiple chorus users singing the same chorus song to obtain a concatenated timbre vector;

[0363] Input the concatenated timbre vector into a fully connected network to obtain the timbre tacit understanding score of the chorus users output by the fully connected network;

[0364] According to the user label, obtain the timbre tacit understanding score of the chorus users output by the fully connected network and the timbre tacit understanding score of the same chorus users in the second training sample;

[0365] Use a loss function to calculate the difference between the timbre tacit understanding score of the chorus users output by the fully connected network and the timbre tacit understanding score of the same chorus users in the second training sample;

[0366] Correct the fully connected network according to the difference and the backpropagation algorithm.

[0367] In some embodiments of the present invention, the fully connected network includes a feature layer, a pooling layer, and a fully connected layer. The processor may also be used to implement the following steps:

[0368] Input the concatenated timbre vector into the feature layer to extract the timbre feature values of the concatenated timbre vector;

[0369] Input the timbre feature values of the concatenated timbre vector into the pooling layer to perform redundancy reduction and dimensionality reduction on the timbre feature values of the concatenated timbre vector;

[0370] Input the timbre feature values of the concatenated timbre vector after redundancy reduction and dimensionality reduction into the fully connected layer to integrate and output the feature values of the concatenated timbre vector after redundancy reduction and dimensionality reduction, so as to obtain the timbre tacit understanding score of the chorus users singing the same chorus song.

[0371] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0372] The feature layer, the pooling layer and the fully connected layer are modified according to the difference and back propagation algorithm.

[0373] The computer device is used to implement the functions of the chorus song evaluation device. Another embodiment of the computer device in the embodiment of the present invention includes:

[0374] Acquire singing segments corresponding to chorus members in a chorus song, wherein the chorus members include at least two chorus users;

[0375] Extracting the mel-score of the singing segment corresponding to each chorus user;

[0376] Inputting the mel spectrum of the corresponding singing segment of each chorus user into the chorus song recognition network to obtain the timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user;

[0377] Splicing the timbre vectors of each chorus user to obtain a spliced ​​timbre vector;

[0378] Inputting the concatenated timbre vector into a feature layer of a fully connected network to extract a feature value of the concatenated timbre vector;

[0379] Inputting the eigenvalues ​​of the concatenated timbre vectors into a pooling layer of a fully connected network to remove redundancy and reduce the dimensionality of the eigenvalues ​​of the concatenated timbre vectors;

[0380] The eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction are input into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction, and obtain the timbre tacit understanding scores of the chorus members singing the same chorus song;

[0381] The chorus songs sung by the chorus members were evaluated based on their timbre tacit understanding scores.

[0382] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0383] The timbre vectors of each chorus user are averaged in the time dimension to obtain the timbre vector of each chorus user without the time series feature, and the timbre vectors of each user without the time series feature are concatenated.

[0384] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0385] The timbre vectors of at least two chorus users are superimposed in the time dimension, and then the superimposed timbre vectors are averaged according to the number of chorus users to obtain a concatenated timbre vector containing time series features.

[0386] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0387] Obtaining the chorus interval of the chorus singers for the same chorus song;

[0388] Calculate the completeness score of each chorus user in the chorus interval;

[0389] The completeness score of the chorus users for the same chorus song is calculated according to the completeness score of each chorus user in the chorus interval.

[0390] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0391] Using a voice activity detection (VAD) method, the chorus start time and the chorus end time of the chorus singers in the same chorus song are obtained;

[0392] The chorus interval of the chorus singers in the same chorus song is determined according to the chorus start time and the chorus end time.

[0393] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0394] Using a voice activity detection (VAD) method, obtaining the vocalization segments of each chorus user in the chorus interval, and the vocalization duration of each line of lyrics in each vocalization segment;

[0395] Obtain a voice file of the chorus song, and obtain the original singer's voice duration of each line of lyrics from the voice file;

[0396] Calculate the single sentence completeness score of each chorus user in the chorus interval according to the occurrence duration of each line of lyrics of each chorus user in the chorus interval and the original singing duration of the corresponding lyrics;

[0397] According to the single sentence completeness score of each chorus user for the chorus song, the completeness score of each chorus user in the chorus interval is calculated.

[0398] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0399] The chorus songs of the chorus members are evaluated according to the timbre tacit understanding scores and the completeness scores of the chorus members for the same chorus song.

[0400] In some embodiments of the present invention, the processor may also be used to implement the following steps:

[0401] Extract the pitch values of the singing segments of each chorus user;

[0402] Obtain the voice file of the chorus song, and obtain the original pitch values of each line of lyrics from the voice file;

[0403] According to the pitch values of the singing segments of each chorus user and the original pitch values of the corresponding lyric segments of the singing segments, obtain the singing skill matching degree scores of multiple chorus users singing the same chorus song;

[0404] Evaluate the chorus songs of the chorus singers according to the timbre matching degree scores and the singing skill matching degree scores of the chorus singers singing the same chorus song.

[0405] It can be understood that whether it is on the side of the training device of the chorus song recognition network or the evaluation device of the chorus song, when the processor in the above-described computer device executes the computer program, it can also implement the functions of each unit in the corresponding device embodiments described above, which will not be elaborated here. Exemplarily, the computer program may be divided into one or more modules / units, and the one or more modules / units are stored in the memory and executed by the processor to complete the present invention. The one or more modules / units may be a series of computer program instruction segments capable of performing specific functions, and these instruction segments are used to describe the execution process of the computer program in the training device of the chorus song recognition network / evaluation device of the chorus song. For example, the computer program may be divided into each unit in the above-described evaluation device of the chorus song, and each unit may implement the specific functions described in the corresponding evaluation device of the chorus song as above.

[0406] The computer device may be a computing device such as a desktop computer, a notebook, a palm computer, and a cloud server. The computer device may include but is not limited to a processor and a memory. Those skilled in the art can understand that the processor and the memory are only examples of the computer device, and do not constitute a limitation on the computer device. It may include more or fewer components, or combine certain components, or different components. For example, the computer device may further include input / output devices, network access devices, a bus, etc.

[0407] The processor may be a Central Processing Unit (CPU), or may also be other general-purpose processors, Digital Signal Processors (DSPs), Application Specific Integrated Circuits (ASICs), Field-Programmable Gate Arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc. The processor is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and lines.

[0408] The memory can be used to store the computer program and / or modules. The processor realizes various functions of the computer device by running or executing the computer program and / or modules stored in the memory, and by calling the data stored in the memory. The memory mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc.; the data storage area can store data created according to the use of the terminal, etc. In addition, the memory may include high-speed random access memory, and may also include non-volatile memory, such as hard disks, memory, plug-in hard disks, Smart Media Cards (SMCs), Secure Digital (SD) cards, Flash Cards, at least one magnetic disk storage device, flash memory device, or other volatile solid-state storage devices.

[0409] The present invention also provides a computer-readable storage medium, which is used to implement the functions of the training device of the chorus song recognition network. A computer program is stored thereon. When the computer program is executed by a processor, the processor can be used to perform the following steps:

[0410] Obtain a labeled first training sample, where the first training sample includes Mel spectrograms of singing voices of multiple chorus users singing multiple chorus songs and timbre vectors of multiple chorus users;

[0411] Input the Mel spectrograms of the singing voices of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vectors of the corresponding chorus users output by the chorus song recognition network;

[0412] According to the user labels, obtain the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training sample;

[0413] Calculate the difference between the timbre vectors of each chorus user output by the chorus song recognition network and the timbre vectors of the same chorus user in the first training sample using the loss function;

[0414] Correct the chorus song recognition network according to the difference and the backpropagation algorithm

[0415] In some embodiments of the present invention, the chorus song recognition network includes a long short-term memory network and a fully connected layer. When the computer program is executed by a processor, the processor can also be used to implement the following steps:

[0416] Input the Mel spectrograms of the singing samples of each chorus user singing the same chorus song into the long short-term memory network to extract the eigenvalue of the timbre vector of each chorus user;

[0417] Input the eigenvalues of the timbre vectors of each chorus user into the fully connected layer to integrate and output the eigenvalues of the timbre vectors of each chorus user, and obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

[0418] In some embodiments of the present invention, when the computer program is executed by a processor, the processor can also be used to implement the following steps:

[0419] Correct the long short-term memory network and the fully connected layer according to the difference and the backpropagation algorithm.

[0420] In some embodiments of the present invention, the intra-class distance between the Mel spectrogram vectors of the singing samples of the same chorus user is less than or equal to a preset reference value, and the inter-class distance between the Mel spectrogram vectors of the singing samples of different chorus users is greater than the preset reference value.

[0421] In some embodiments of the present invention, when the computer program is executed by a processor, the processor can also be used to implement the following steps:

[0422] Obtain a labeled second training sample, where the second training sample includes the timbre vectors of multiple chorus users singing multiple chorus songs, and the timbre tacit understanding scores of multiple chorus users in each chorus song;

[0423] Concatenate the timbre vectors of the song segments of multiple chorus users singing the same chorus song to obtain a concatenated timbre vector;

[0424] Input the concatenated timbre vector into a fully connected network to obtain the timbre tacit understanding score of the chorus singer output by the fully connected network;

[0425] Obtain the score of the chorister timbre tacit understanding output by the fully connected network and the score of the timbre tacit understanding of the same chorister in the second training sample according to the user tags;

[0426] Use a loss function to calculate the difference between the score of the chorister timbre tacit understanding output by the fully connected network and the score of the timbre tacit understanding of the same chorister in the second training sample;

[0427] Correct the fully connected network according to the difference and the backpropagation algorithm.

[0428] In some embodiments of the present invention, the fully connected network includes a feature layer, a pooling layer, and a fully connected layer. When the computer program is executed by a processor, the processor can also be used to implement the following steps:

[0429] Input the spliced timbre vectors into the feature layer to extract the timbre feature values of the spliced timbre vectors;

[0430] Input the timbre feature values of the spliced timbre vectors into the pooling layer to reduce the redundancy and dimensionality of the timbre feature values of the spliced timbre vectors;

[0431] Input the timbre feature values of the spliced timbre vectors after redundancy reduction and dimensionality reduction into the fully connected layer to integrate and output the feature values of the spliced timbre vectors after redundancy reduction and dimensionality reduction, so as to obtain the score of the timbre tacit understanding of the choristers singing the same choral song.

[0432] In some embodiments of the present invention, when the computer program is executed by a processor, the processor can also be used to implement the following steps:

[0433] Correct the feature layer, the pooling layer, and the fully connected layer according to the difference and the backpropagation algorithm.

[0434] The present invention also provides another computer-readable storage medium, which is used to implement the functions of the evaluation device for choral songs. A computer program is stored thereon. When the computer program is executed by a processor, the processor can be used to execute the following steps:

[0435] Obtain the singing segments corresponding to the choristers in the choral song, where the choristers include at least two choral users;

[0436] Extract the Mel spectrograms of the singing segments corresponding to each choral user;

[0437] Inputting the mel spectrum of the corresponding singing segment of each chorus user into the chorus song recognition network to obtain the timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user;

[0438] Splicing the timbre vectors of each chorus user to obtain a spliced ​​timbre vector;

[0439] Inputting the concatenated timbre vector into a feature layer of a fully connected network to extract a feature value of the concatenated timbre vector;

[0440] Inputting the eigenvalues ​​of the concatenated timbre vectors into a pooling layer of a fully connected network to remove redundancy and reduce the dimensionality of the eigenvalues ​​of the concatenated timbre vectors;

[0441] The eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction are input into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction, and obtain the timbre tacit understanding scores of the chorus members singing the same chorus song;

[0442] The chorus songs sung by the chorus members were evaluated based on their timbre tacit understanding scores.

[0443] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0444] The timbre vectors of each chorus user are averaged in the time dimension to obtain the timbre vector of each chorus user without the time series feature, and the timbre vectors of each user without the time series feature are concatenated.

[0445] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0446] The timbre vectors of at least two chorus users are superimposed in the time dimension, and then the superimposed timbre vectors are averaged according to the number of chorus users to obtain a concatenated timbre vector containing time series features.

[0447] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0448] Obtaining the chorus interval of the chorus singers for the same chorus song;

[0449] Calculate the completeness score of each chorus user in the chorus interval;

[0450] The completeness score of the chorus users for the same chorus song is calculated according to the completeness score of each chorus user in the chorus interval.

[0451] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0452] Using a voice activity detection (VAD) method, the chorus start time and the chorus end time of the chorus singers in the same chorus song are obtained;

[0453] The chorus interval of the chorus singers in the same chorus song is determined according to the chorus start time and the chorus end time.

[0454] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0455] Using a voice activity detection (VAD) method, obtaining the vocalization segments of each chorus user in the chorus interval, and the vocalization duration of each line of lyrics in each vocalization segment;

[0456] Obtain a voice file of the chorus song, and obtain the original singer's voice duration of each line of lyrics from the voice file;

[0457] Calculate the single sentence completeness score of each chorus user in the chorus interval according to the occurrence duration of each line of lyrics of each chorus user in the chorus interval and the original singing duration of the corresponding lyrics;

[0458] According to the single sentence completeness score of each chorus user for the chorus song, the completeness score of each chorus user in the chorus interval is calculated.

[0459] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0460] The chorus songs of the chorus members are evaluated according to the timbre tacit understanding scores and the completeness scores of the chorus members for the same chorus song.

[0461] In some embodiments of the present invention, when the computer program stored in the computer-readable storage medium is executed by the processor, the processor may also be used to implement the following steps:

[0462] Extract the pitch value of each chorus user's singing segment;

[0463] Obtain the voice file of the choral song, and obtain the original pitch value of each line of lyrics from the voice file;

[0464] According to the pitch value of each choral user's singing segment and the original pitch value of the corresponding lyrics segment of the singing segment, obtain the singing skill matching degree scores of multiple choral users singing the same choral song;

[0465] Evaluate the choral song of the choristers according to the timbre matching degree score and the singing skill matching degree score of the choristers singing the same choral song.

[0466] It can be understood that if the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a corresponding computer-readable storage medium. Based on such an understanding, to implement all or part of the processes in the above corresponding embodiment methods of the present invention, it can also be completed by a computer program instructing relevant hardware. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.

[0467] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed mutual coupling or direct coupling or communication connection can be through some interfaces, and the indirect coupling or communication connection of devices or units can be in electrical, mechanical or other forms.

[0468] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed across multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0469] In addition, in each embodiment of the present invention, each functional unit may be integrated in a processing unit, may exist separately as individual physical units, or two or more units may be integrated in one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0470] As described above, the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A training method for a choral song recognition network, characterized in that, The method comprises: Acquire a first labeled training sample, wherein the first training sample includes mel spectra of vocal samples of multiple chorus songs sung by multiple chorus users and timbre vectors of multiple chorus users; Inputting the mel spectrum of the voice samples of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network; According to the user label, obtaining the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample; Calculate the difference between the timbre vector of each chorus user output by the chorus song recognition network and the timbre vector of the same chorus user in the first training sample using a loss function; The chorus song recognition network is modified according to the difference and back propagation algorithm.

2. The training method according to claim 1, wherein The chorus song recognition network includes a long short-term memory network and a fully connected layer; Inputting the Mel-spectrogram of the voice sample of each chorus user singing the same chorus song into the chorus song recognition network to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network, including: Inputting the Mel spectrum of each chorus user's singing voice sample of the same chorus song into the long short-term memory network to extract the feature value of each chorus user's timbre vector; The characteristic value of the timbre vector of each chorus user is input into the fully connected layer to integrate and output the characteristic value of the timbre vector of each chorus user, so as to obtain the timbre vector of the corresponding chorus user output by the chorus song recognition network.

3. The training method according to claim 2, wherein The chorus song recognition network is modified according to the difference and back propagation algorithm, including: The long short-term memory network and the fully connected layer are modified according to the difference and back propagation algorithm.

4. The training method according to any one of claims 1 to 3, characterized in that The intra-class distance between the mel-spectrogram vectors of the same chorus user singing sample is less than or equal to a preset reference value, and the inter-class distance between the mel-spectrogram vectors of different chorus user singing samples is greater than the preset reference value.

5. The training method according to claim 1, characterized in that The method further comprises: Obtaining a second training sample with a label, wherein the second training sample includes timbre vectors of multiple chorus songs sung by multiple chorus users, and timbre tacit understanding scores of the multiple chorus users in each chorus song; splicing the timbre vectors of multiple chorus user song segments singing the same chorus song to obtain a spliced ​​timbre vector; Inputting the concatenated timbre vector into a fully connected network to obtain a timbre tacit understanding score of the chorus members output by the fully connected network; According to the user tag, obtaining the timbre compatibility score of the chorus singer output by the fully connected network and the timbre compatibility score of the same chorus singer in the second training sample; Calculating the difference between the timbre harmony score of the chorus singer output by the fully connected network and the timbre harmony score of the same chorus singer in the second training sample using a loss function; The fully connected network is modified according to the difference and the back propagation algorithm.

6. The training method according to claim 5, wherein The fully connected network includes a feature layer, a pooling layer and a fully connected layer; The concatenated timbre vector is input into the fully connected network to obtain the timbre harmony score of the chorus output by the fully connected network, including: Inputting the concatenated timbre vector to the feature layer to extract the timbre feature value of the concatenated timbre vector; Inputting the timbre feature value of the concatenated timbre vector into the pooling layer to remove redundancy and reduce the dimension of the timbre feature value of the concatenated timbre vector; The timbre eigenvalues ​​of the concatenated timbre vectors after redundancy reduction are input into the fully connected layer for integrating and outputting the eigenvalues ​​of the concatenated timbre vectors after redundancy reduction, so as to obtain the timbre tacit understanding scores of the chorus members singing the same chorus song.

7. The training method according to claim 6, characterized in that The fully connected network is modified according to the difference and the back propagation algorithm, including: The feature layer, the pooling layer and the fully connected layer are modified according to the difference and back propagation algorithm.

8. A method for evaluating a choral song, characterized in that, The method comprises: Acquire singing segments corresponding to chorus members in a chorus song, wherein the chorus members include at least two chorus users; Extracting the mel-score of the singing segment corresponding to each chorus user; Inputting the mel spectrum of the corresponding singing segment of each chorus user into a chorus song recognition network obtained by the training method of the chorus song recognition network according to any one of claims 1 to 7 to obtain the timbre vector of each chorus user, wherein the chorus song recognition network includes a long short-term memory network and a fully connected layer, the long short-term memory network is used to extract the eigenvalue of the timbre vector of each chorus user, and the fully connected layer is used to integrate and output the eigenvalue of the timbre vector of each chorus user; Splicing the timbre vectors of each chorus user to obtain a spliced ​​timbre vector; Inputting the concatenated timbre vector into a feature layer of a fully connected network to extract a feature value of the concatenated timbre vector; Inputting the eigenvalues ​​of the concatenated timbre vectors into a pooling layer of a fully connected network to remove redundancy and reduce the dimensionality of the eigenvalues ​​of the concatenated timbre vectors; The eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction are input into the fully connected layer of the fully connected network, so as to integrate and output the eigenvalues ​​of the concatenated timbre vectors after de-redundancy and dimensionality reduction, and obtain the timbre tacit understanding scores of the chorus members singing the same chorus song; The chorus songs sung by the chorus members were evaluated based on their timbre tacit understanding scores.

9. The evaluation method according to claim 8, characterized in that The timbre vectors of each chorus user are spliced ​​to obtain a spliced ​​timbre vector, including: The timbre vectors of each chorus user are averaged in the time dimension to obtain the timbre vector of each chorus user without the time series feature, and the timbre vectors of each user without the time series feature are concatenated.

10. The method according to claim 8, wherein The timbre vectors of each chorus user are spliced ​​to obtain a spliced ​​timbre vector, including: The timbre vectors of at least two chorus users are superimposed in the time dimension, and then the superimposed timbre vectors are averaged according to the number of chorus users to obtain a concatenated timbre vector containing time series features.

11. The evaluation method according to claim 8, wherein The method further comprises: Obtaining the chorus interval of the chorus singers for the same chorus song; Calculate the completeness score of each chorus user in the chorus interval; The completeness score of the chorus users for the same chorus song is calculated according to the completeness score of each chorus user in the chorus interval.

12. The evaluation method according to claim 11, wherein Obtaining the chorus intervals of the choristers for the same chorus song, including: Using the Voice Activity Detection (VAD) method to obtain the chorus start time and chorus end time of the choristers in the same chorus song; Determining the chorus intervals of the choristers within the same chorus song according to the chorus start time and the chorus end time.

13. The evaluation method according to claim 11, characterized in that, Calculating the integrity scores of each chorus user within the chorus intervals, including: Using the Voice Activity Detection (VAD) method to obtain the vocal segments of each chorus user within the chorus intervals, and the vocal durations of each line of lyrics in each vocal segment; Obtaining the voice file of the chorus song, and obtaining the original singer's vocal duration of each line of lyrics from the voice file; Calculating the single-line integrity scores of each chorus user within the chorus intervals according to the vocal durations of each line of lyrics of each chorus user within the chorus intervals and the original singer's vocal durations of the corresponding lyrics; Calculating the integrity scores of each chorus user within the chorus intervals according to the single-line integrity scores of each chorus user for the chorus song.

14. The evaluation method according to claim 11, wherein Evaluating the chorus songs of the choristers according to the timbre matching scores of the choristers, including: Evaluating the chorus songs of the choristers according to the timbre matching scores of the choristers for the same chorus song and the integrity scores for the same chorus song.

15. The evaluation method according to any one of claims 8 to 14, characterized in that, The evaluation of the chorus songs of the choristers according to the timbre matching scores of the choristers further includes: Extracting the pitch values of the singing segments of each chorus user; Obtaining the voice file of the chorus song, and obtaining the original singer's pitch values of each line of lyrics from the voice file; Obtaining the singing skill matching scores of multiple chorus users singing the same chorus song according to the pitch values of the singing segments of each chorus user and the original singer's pitch values of the corresponding lyric segments of the singing segments; Evaluating the chorus songs of the choristers according to the timbre matching scores of the choristers singing the same chorus song and the singing skill matching scores of singing the same chorus song.

16. A computer device, comprising a processor, characterized in that, When the processor executes the computer program stored in the memory, it is used to implement the training method of the chorus song recognition network according to any one of claims 1 to 7 or the evaluation method of the chorus song according to any one of claims 8 to 15.

17. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it is used to implement the training method of the chorus song recognition network according to any one of claims 1 to 7 or the evaluation method of the chorus song according to any one of claims 8 to 15.

Citation Information

Patent Citations

  • Voice synthesis model training method and device

    CN111508470A

  • Music structure determination method and device, equipment and medium

    CN112037764A