Sound quality evaluation model determination method, sound quality evaluation method, device and medium

By analyzing the pitch distribution vector and sampling probability of audio samples, filtering the audio sample set and training the sound quality evaluation model, the problem of high cost and low accuracy caused by the reliance on manpower in the prior art is solved, and automated evaluation and higher accuracy are achieved.

CN115631768BActive Publication Date: 2025-05-20TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211340306.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-28
Publication Date
2025-05-20
Estimated Expiration
2042-10-28

AI Technical Summary

Technical Problem

In the prior art, audio quality evaluation relies on manpower, resulting in high costs and low accuracy.

Method used

By obtaining dry sound clips of multiple audio samples, determining the pitch distribution vector and sampling probability, filtering the audio sample set, and using this set to train the sound quality evaluation model to generate a model for automatically evaluating audio quality.

Benefits of technology

It reduces labor costs during the audio quality evaluation process and improves the accuracy of the evaluation, avoiding inaccuracy problems caused by differences in assessor levels.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631768B_ABST
    Figure CN115631768B_ABST
Patent Text Reader

Abstract

The present application provides a method for determining a sound quality assessment model, a sound quality assessment method, a device and a medium, wherein the method for determining a sound quality assessment model comprises: obtaining multiple audio samples, each audio sample comprising a dry sound segment; using the dry sound segment of each audio sample, determining the pitch distribution vector of each audio sample, and determining the sampling probability of each audio sample according to the pitch distribution vector of each audio sample; determining an audio sample set from multiple audio samples according to the sampling probability of each audio sample; using the audio sample set to train an initial sound quality assessment model to obtain a trained sound quality assessment model; the sound quality assessment model is used to determine the audio quality of the input audio. The embodiments of the present application not only do not require human costs in the audio quality assessment process, but also can improve the accuracy of audio quality assessment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular, to a method for determining a sound quality evaluation model, a sound quality evaluation method, a device, and a medium. Background Art

[0002] Currently, the audio quality can be obtained by having evaluators score based on their listening perception. With the continuous increase in audio data, the required human cost is also getting higher and higher. Moreover, the evaluation levels of different evaluators vary, and correspondingly, the accuracy of audio quality evaluation is relatively low.

[0003] Therefore, how to reduce the human cost in the audio quality evaluation process and improve the accuracy of audio quality evaluation is an urgent problem to be solved. Summary of the Invention

[0004] In view of the above technical problems, this application provides a method for determining a sound quality evaluation model and related devices, which can reduce the human cost in the audio quality evaluation process and improve the accuracy of audio quality evaluation.

[0005] On the one hand, an embodiment of this application provides a method for determining a sound quality evaluation model. The method includes: obtaining a plurality of audio samples, each audio sample including a dry voice segment; using the dry voice segment of each audio sample to determine the pitch distribution vector of each audio sample, and determining the sampling probability of each audio sample according to the pitch distribution vector of each audio sample; determining an audio sample set from the plurality of audio samples according to the sampling probability of each audio sample; using the audio sample set to train an initial sound quality evaluation model to obtain a trained sound quality evaluation model; the sound quality evaluation model is used to determine the audio quality of the input audio.

[0006] In an alternative embodiment, determining an audio sample set from the plurality of audio samples according to the pitch distribution vector of each audio sample includes: using the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample; wherein, the degree of concentration and dispersion of each frame of the audio sample in the pitch distribution vector is related to the magnitude of the weight corresponding to the audio sample; using the weight corresponding to each audio sample to determine the sampling probability of each audio sample.

[0007] In an alternative embodiment, using the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample includes: using the pitch distribution vector of each audio sample to determine the pitch distribution matrix corresponding to the plurality of audio samples; determining the weight corresponding to each audio sample according to the pitch distribution matrix and vector B, wherein the dimension of vector B is the same as the dimension of the pitch distribution vector, and all elements in vector B are equal.

[0008] In an alternative embodiment, the audio quality evaluation model includes a feature learning model and an audio quality prediction model. The initial audio quality evaluation model is trained using an audio sample set to obtain a trained audio quality evaluation model, which includes: inputting each audio sample in the audio sample set into the feature learning model to obtain a feature representation of each audio sample, where the feature learning model is trained based on a self-supervised learning method; inputting the feature representation of each audio sample into the audio quality prediction model to obtain the audio quality of each audio sample; and training the audio quality evaluation model using the audio quality of each audio sample to obtain a trained audio quality evaluation model.

[0009] In an alternative embodiment, inputting each audio sample in the audio sample set into the feature learning model to obtain a feature representation of each audio sample includes: for each audio sample in the audio sample set, performing frame addition and windowing processing on the audio sample to extract the Mel spectrum features of the audio sample; and inputting the Mel spectrum features of the audio sample into the feature learning model to obtain a feature representation of the audio sample.

[0010] In an alternative embodiment, the audio quality evaluation model further includes a long short-term memory module layer and a non-linear mapping layer; inputting the feature representation of each audio sample into the audio quality prediction model to obtain the audio quality of each audio sample includes: for each audio sample in the audio sample set, using the long short-term memory module layer to perform modeling processing on the feature representation of the audio sample to obtain a modeled feature representation that does not include features irrelevant to audio quality; using the non-linear mapping layer to perform non-linear mapping processing on the modeled feature representation to obtain a non-linearly mapped feature representation; and inputting the non-linearly mapped feature representation into the audio quality prediction model to obtain the audio quality of each audio sample.

[0011] In an alternative embodiment, the method further includes: inputting each audio sample in the validation set into the trained audio quality evaluation model to output the audio quality of each audio sample in the validation set; determining the audio quality accuracy rate of the validation set based on the audio quality of each audio sample in the validation set; and using the audio quality accuracy rate of the validation set to verify the trained audio quality evaluation model to obtain a verification result.

[0012] On the one hand, an embodiment of the present application provides an audio quality evaluation method, which includes: obtaining an audio to be evaluated; inputting the audio to be evaluated into the trained audio quality evaluation model to obtain the audio quality of the audio to be evaluated output by the trained audio quality evaluation model; where the trained audio quality evaluation model is the trained audio quality evaluation model in the audio quality evaluation model determination method in the previous aspect.

[0013] On the one hand, an embodiment of the present application provides an audio quality evaluation model determination device, which includes:

[0014] An acquisition module, configured to acquire a plurality of audio samples, each audio sample including a dry voice segment; a determination module, configured to determine a pitch distribution vector of each audio sample by using the dry voice segment of each audio sample, and determine a sampling probability of each audio sample according to the pitch distribution vector of each audio sample; the determination module is further configured to determine an audio sample set from the plurality of audio samples according to the sampling probability of each audio sample; a training module, further configured to train a sound quality evaluation model by using the audio sample set to obtain a trained sound quality evaluation model; the sound quality evaluation model is used to determine the audio quality of the input audio.

[0015] On the one hand, an embodiment of the present application provides a sound quality evaluation device, and the device includes:

[0016] An acquisition module, configured to acquire an audio to be evaluated; a processing module, configured to input the audio to be evaluated into the trained sound quality evaluation model to obtain the audio quality of the audio to be evaluated output by the trained sound quality evaluation model; wherein, the trained sound quality evaluation model is the trained sound quality evaluation model in the above-mentioned sound quality evaluation model determination method.

[0017] On the one hand, an embodiment of the present application provides an electronic device, including: a processor, a user interface, a communication interface, and a memory, the processor, the user interface, the communication interface, and the memory are connected to each other, wherein, the memory stores executable program code, and the processor is configured to call the executable program code to execute the method provided by the embodiment of the present application.

[0018] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium, on which a computer program is stored, the computer program includes program instructions, and when the program instructions are executed by a processor, the method provided by the embodiment of the present application is implemented.

[0019] Correspondingly, an embodiment of the present application further provides a computer program product or a computer program, the computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the electronic device executes the method provided by the embodiment of the present application.

[0020] In the embodiments of the present application, a plurality of audio samples are obtained, and each audio sample includes a dry voice segment; using the dry voice segment of each audio sample, the pitch distribution vector of each audio sample is determined, and according to the pitch distribution vector of each audio sample, the sampling probability of the audio sample is determined; an audio sample set is determined from the plurality of audio samples according to the sampling probability of each audio sample; the initial audio quality evaluation model is trained using the audio sample set to obtain a trained audio quality evaluation model; the audio quality evaluation model is used to determine the audio quality of the input audio. On the one hand, the audio quality evaluation model determined by this method can automatically determine the audio quality of the input audio, which can reduce the labor cost in the audio quality evaluation process and improve the accuracy of the audio quality evaluation; on the other hand, the audio sample set used to train this audio quality evaluation model is screened based on the pitch distribution vector of each audio sample, thus avoiding the problem of uneven pitch distribution of the training data caused by the audio samples in the audio sample set being concentrated in a certain range of vocal frequencies, and further improving the accuracy of the audio quality evaluation. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0022] Figure 1 is a flowchart of a method for determining an audio quality evaluation model provided by an embodiment of the present application;

[0023] Figure 2 is a schematic diagram of an audio quality evaluation model provided by an embodiment of the present application;

[0024] Figure 3 is a schematic diagram of an audio quality evaluation method provided by an embodiment of the present application;

[0025] Figure 4 is a schematic diagram of a device for determining an audio quality evaluation model provided by an embodiment of the present application;

[0026] Figure 5 is a schematic diagram of the structure of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0027] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0028] Currently, the audio quality can be scored by evaluators based on their listening perception. With the continuous increase of audio data, the required human cost is also increasing. Moreover, the evaluation levels of different evaluators vary, and correspondingly, the evaluation accuracy of audio quality is relatively low.

[0029] The embodiment of the present application provides a method for determining a sound quality evaluation model. In this method, a plurality of audio samples are obtained, and each audio sample includes a dry sound segment; the pitch distribution vector of each audio sample is determined by using the dry sound segment of each audio sample; the sampling probability of each audio sample is determined according to the pitch distribution vector of each audio sample; an audio sample set is determined from the plurality of audio samples according to the sampling probability of each audio sample; the initial sound quality evaluation model is trained by using the audio sample set to obtain a trained sound quality evaluation model; the sound quality evaluation model is used to determine the audio quality of the input audio. On the one hand, the sound quality evaluation model can automatically determine the audio quality of the input audio, which can reduce the human cost in the audio quality evaluation process and improve the evaluation accuracy of audio quality; on the other hand, the audio sample set used to train the sound quality evaluation model is screened based on the pitch distribution vector of each audio sample, so as to avoid the problem of uneven pitch distribution of training data caused by the audio samples in the audio sample set being concentrated in a certain range of sounding frequencies, thereby improving the evaluation accuracy of sound quality.

[0030] It should be noted that in specific implementation, the sound quality evaluation model in the above method can be an electronic device or a module of an electronic device. The electronic device can be a terminal or a server; among them, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch, a smart vehicle terminal, etc., but is not limited thereto. The server can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, etc., but is not limited thereto. When the electronic device is a server, the above method is processed through the server background, and its processing efficiency is high and the running speed is fast.

[0031] The method for determining the sound quality evaluation model will be elaborated in detail below with reference to the accompanying drawings.

[0032] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for determining a sound quality evaluation model provided by an embodiment of the present application. As shown in Figure 1As shown, the method for determining the sound quality evaluation model may include but is not limited to the following steps:

[0033] S101. Obtain a plurality of audio samples, each audio sample including a dry voice segment;

[0034] Optionally, the plurality of audio samples include a large amount of speech data and singing data, such as data exceeding 120 hours, covering multiple languages and various timbres, so that the sound quality evaluation model trained based on the audio samples can have good generalization ability in speech and singing. In addition, screening out the audio sample set based on the pitch distribution vector of the audio samples is beneficial to making the collected audio samples cover the audio data of 60Hz - 1200Hz, and the human vocalization (i.e., dry voice) frequency can be between 60Hz - 1200Hz. In this relatively wide range, the collected audio sample set can include various acoustic features, thus avoiding the problem of uneven pitch distribution of the training data caused by collecting audio samples (or dry voices) without rules, resulting in the pitch of each audio sample being concentrated between 120Hz - 480Hz. Among them, each audio sample or the audio sample set can also be called training data, and each audio sample can also be called an audio segment.

[0035] S102. Use the dry voice segment of each audio sample to determine the pitch distribution vector of each audio sample;

[0036] S103. Determine the sampling probability of each audio sample according to the pitch distribution vector of each audio sample;

[0037] In an optional implementation manner, determining the sampling probability of each audio sample according to the pitch distribution vector of each audio sample includes: using the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample; wherein, the concentration situation of each frame of the audio sample in the pitch distribution vector is related to the size of the weight corresponding to the audio sample; using the weight corresponding to each audio sample to determine the sampling probability of each audio sample.

[0038] For example, for the audio samples in the audio sample set with a relatively concentrated pitch distribution vector, such as concentrated between the pitches of 120Hz - 480Hz, a lower weight can be set, and for those with a relatively dispersed pitch distribution vector, such as outside the pitches of 120Hz - 480Hz, a higher weight can be set.

[0039] Optionally, using the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample can be: combining the pitch distribution vectors of each of the audio samples into a pitch distribution matrix; determining the weight corresponding to each audio sample according to the pitch distribution matrix and vector B, where vector B can be called an intermediate vector, its dimension is the same as that of the pitch distribution vector, and all elements in vector B are equal.

[0040] Optionally, the elements in vector B can be predefined, or the ratio of the sum of all elements in the pitch distribution matrix to the number of audio samples can be used as the elements in vector B.

[0041] For example, for m audio samples, the pitch distribution vectors of each audio sample are analyzed. Assume that the pitch distribution vectors of each audio sample are vectors with a pitch dimension n of 1200. The element corresponding to each pitch dimension in the pitch distribution vector is the number of frames of that pitch dimension in the corresponding audio sample, and the duration of each frame is 16 milliseconds. Then, these m audio samples can form a pitch distribution matrix A with a dimension of m*n; assume the weight of each audio sample is X i , where i represents the i-th audio sample among the m audio samples, so i can take values from 1 to m. Then, a weight vector X with a dimension of m*1 can be formed; to achieve pitch distribution balance, the weight vector X corresponding to the m audio samples can be obtained through the following formula:

[0042] A T X = B, s.t. X i > 0 (1)

[0043] where "s.t. X i > 0" means that the constraint condition that formula (1) needs to satisfy is that the weights of each audio sample are all greater than 0; among them, vector B is a vector with a dimension of n*1, and each element B in B i is equal. For example, the element B in vector B i can be obtained through the following formula (2):

[0044] B i = sum(A) / m (2)

[0045] where sum(A) represents the sum of all elements in the pitch distribution matrix A, so that the elements in vector B can be obtained. Furthermore, the weights X corresponding to each audio sample in the weight vector X can be solved using formula (1) i .

[0046] S104. Determine an audio sample set from multiple audio samples according to the sampling probability of each audio sample;

[0047] It can be seen that the weight corresponding to each audio sample can be used as the sampling probability of the audio sample. Furthermore, according to the sampling probability of each audio sample, determining an audio sample set from multiple audio samples can make the number of frames of each pitch dimension in the audio sample set similar, thereby achieving the overall pitch distribution balance of the training data.

[0048] S105. Use the audio sample set to train the initial sound quality evaluation model to obtain the trained sound quality evaluation model.

[0049] Optionally, the initial audio quality evaluation model may include a feature learning model and an audio quality prediction model. The initial audio quality evaluation model is trained using an audio sample set to obtain a trained audio quality evaluation model, including: inputting each audio sample in the audio sample set into the feature learning model to obtain a feature representation of each audio sample; inputting the feature representation of each audio sample into the audio quality prediction model to obtain the audio quality of each audio sample; and training the audio quality evaluation model using the audio quality of each audio sample to obtain a trained audio quality evaluation model.

[0050] Among them, the audio quality prediction model processes the feature representation of the audio sample to obtain the audio quality of the audio sample. Optionally, the audio quality prediction model converts the classification result of each frame into a score from 0 to 5, and then calculates the mean value as the final score of the entire audio.

[0051] The feature learning model is trained based on a self-supervised learning method. For example, an unsupervised audio representation learning method (i.e., the feature learning model) is trained on a large amount of unlabeled audio data, so as to abstract the original audio signal into a higher-level feature representation form, and solve the problem of too little data with audio quality annotations. The feature representation is an information representation obtained by abstracting the original audio signal.

[0052] Optionally, inputting each audio sample in the audio sample set into the feature learning model to obtain a feature representation of each audio sample includes: for each audio sample in the audio sample set, performing frame addition and windowing processing on the audio sample to extract the Mel spectrum features of the audio sample; and inputting the Mel spectrum features of the audio sample into the feature learning model to obtain a feature representation of the audio sample.

[0053] For example, for each audio sample in the audio sample set, performing frame addition and windowing processing on the audio sample to extract the Mel spectrum features of the audio sample, and the dimension of the Mel spectrum features is [T, M]; inputting the Mel spectrum features of the audio sample into the feature learning model to obtain a feature representation [T, D] of the audio sample. Among them, T represents the number of frames of the audio sample, M represents the number of filter banks of the Mel spectrum, which is generally set to 80, and D represents the dimension of the feature, which is set to 128 in this embodiment.

[0054] Optionally, the audio quality evaluation model further includes a long short-term memory module layer (Long Short Term Memory, LSTM) and a non-linear mapping layer; the obtaining of the audio quality of each audio sample by inputting the feature representation of each audio sample into the audio quality prediction model includes: for each audio sample in the audio sample set, using the long short-term memory module layer to perform modeling processing on the feature representation of the audio sample to obtain a feature representation after modeling processing, where the feature representation after modeling processing does not contain features irrelevant to audio quality; using the non-linear mapping layer to perform non-linear mapping processing on the feature representation after modeling processing to obtain a feature representation after non-linear mapping processing; using the non-linear mapping layer to perform non-linear mapping processing on the feature representation after modeling processing to obtain a feature representation after non-linear mapping processing; inputting the feature representation after non-linear mapping processing into the audio quality prediction model to obtain the audio quality of each audio sample. It can be seen that in this embodiment, the long short-term memory module layer is used to perform modeling processing on the feature representation of the audio sample, and information irrelevant to audio quality in the feature representation of the audio sample (such as the timbre feature of the audio sample, the pronunciation feature of the audio sample, etc.) is eliminated, so as to avoid the problem of jitter in the audio quality score caused by the timbre feature of the audio sample, the pronunciation feature of the audio sample, etc. In addition, in this embodiment, the non-linear mapping layer is used to perform non-linear mapping processing on the feature representation of the audio sample after modeling processing. Since the processing of the non-linear mapping layer can further eliminate information irrelevant to audio quality, the accuracy of the audio quality evaluation model in obtaining audio quality is improved.

[0055] Optionally, the obtaining of the audio quality of each audio sample by inputting the feature representation after non-linear mapping processing into the audio quality prediction model includes: performing softmax processing on the feature representation after non-linear mapping processing to obtain a classification result of the feature representation; determining an evaluation score for each frame in the audio sample according to the classification result of the feature representation; determining an evaluation score for the audio sample according to the evaluation scores for each frame in the audio sample (for example, obtaining the final score of the entire audio sample by taking the mean). It can be seen that in this embodiment, the classification result (vectorization result) of the feature representation of the audio sample is converted into an evaluation score for the audio sample. In this way, the defect that the evaluation scores obtained by using the trained audio quality evaluation model are concentrated near the mean can be effectively avoided.

[0056] For example, assuming that the classification result of the feature representation of the audio sample is a result of 128 classes, and the evaluation score of the audio sample is from 0 to 5 points, the conversion of the classification result (vectorization result) of the feature representation of the audio sample into the evaluation score of the audio sample can adopt the following formula: floor(x / 5*128), where floor represents the operation of rounding down, and x represents the manual scoring situation of the audio quality of the audio sample, and the value range is 0 to 5.

[0057] In addition, in this application, the sound quality evaluation model obtained in S105 also needs to be verified. That is, the method for determining the sound quality evaluation model further includes: inputting each audio sample in the verification set into the trained sound quality evaluation model to output the audio quality of each audio sample in the verification set; determining the audio quality accuracy rate of the verification set according to the audio quality of each audio sample in the verification set; using the audio quality accuracy rate of the verification set to verify the trained sound quality evaluation model to obtain a verification result. For example, if the verification result is that the verification passes, the audio evaluation model can be used for sound quality evaluation. If the verification result is that the verification fails, the audio evaluation model can continue to be trained using the method described in steps S101 to S105.

[0058] It can be seen that the sound quality evaluation model determined in this application can directly determine the audio quality of the input audio, which is a reference-free sound quality evaluation method. Compared with the reference-based sound quality evaluation method, due to the need for multiple internal iterations of alignment processing and parameter filtering in the algorithm, the problem of long time consumption for audio signal sound quality evaluation, and for the reference-based sound quality evaluation method, such as the sound quality evaluation method of perceptual evaluation of speech quality (PESQ), which gives priority to the sampled frequency band processed and is difficult to meet the evaluation requirements of wider signals, this sound quality evaluation model can directly perform sound quality evaluation on the input audio, with shorter time consumption and also adapting to the evaluation requirements of wider signals. In addition, compared with the sound quality evaluation scheme that compares the spectrogram of the input audio with the defined clean spectrogram, the sound quality evaluation model determined in this method for determining the sound quality evaluation model is trained based on an audio sample set, can learn on audio samples with sound quality labels, and can better fit the sound quality score, making the score more accurate and stable.

[0059] Please refer to Figure 2 , Figure 2 which is a schematic diagram of a sound quality evaluation model provided by an embodiment of this application. Assume that the feature learning model is a wav2vec model, as Figure 2As shown, the input audio samples are screened based on the pitch distribution vector of each audio sample, so as to avoid the problem of uneven pitch distribution of training data caused by the fact that the audio samples in the audio sample set are all concentrated in a certain range of vocal frequencies. For details, please refer to the above text and will not be elaborated here. For each audio sample in the audio sample set, frame addition and windowing processing are performed on the audio sample to extract the Mel-spectrum features of the audio sample; the wav2vec model is used to process the Mel-spectrum features of the audio sample to obtain the feature representation of the audio sample. After multiple rounds of training, the final wav2vec model is obtained; the long short-term memory module (LSTM) layer is used to model and smooth the feature representation of the audio sample obtained by the final wav2vec model, and eliminate the information irrelevant to the sound quality in the feature representation of the audio sample (such as timbre features, pronunciation features, etc.); the dropout layer is used to ignore some of the feature representations in the feature representation of the audio sample after modeling and smoothing with a certain probability (such as 0.7), so as to avoid the model over-relying on certain local features (that is, to avoid the model overfitting to some features of the audio sample); the non-linear mapping layer (Linear&SiLU) composed of a linear layer and a sigmoid linear unit layer (SiLU) is used to process the feature representation of the audio sample after the dropout layer; the dropout layer is used again to ignore some of the feature representations in the feature representation of the audio sample after non-linear mapping processing with a certain probability (such as 0.5); the feature representation of the audio sample after the dropout layer (with a probability of 0.5) is classified by the linear softmax layer (Linear&Softmax) to obtain the initial sound quality evaluation model; the validation set is used to verify the initial sound quality evaluation model multiple times (adjust the parameters of the initial sound quality evaluation model during multiple verification processes, and then determine the parameters of the sound quality evaluation model) to determine the sound quality evaluation model. Among them, the process of verifying the initial sound quality evaluation model is as follows: the multi-class cross-entropy loss function value is obtained according to the audio quality prediction score obtained by the initial sound quality evaluation model and the audio quality score of the corresponding validation set. After multiple rounds of verification, the initial sound quality evaluation model with the smallest multi-class cross-entropy loss function value (that is, the highest accuracy rate of the validation set) is determined as the sound quality evaluation model. The audio quality prediction score is obtained by taking the average of the scores of each frame after the classification processing of the feature representation of the audio sample.

[0060] Please refer to Figure 3 , Figure 3 which is a schematic diagram of a sound quality evaluation method provided by an embodiment of the present application. As Figure 3As shown in the figure, the audio to be evaluated is input into the audio quality evaluation model, and the audio quality of the audio to be evaluated can be output. Among them, the audio quality evaluation model is the audio quality evaluation model determined by the above-mentioned audio quality evaluation model determination method. This audio quality evaluation method can be applied not only to scenarios such as speech synthesis, singing voice synthesis, and audio quality detection, but also to evaluate the quality of synthesis models and monitor the quality on the content production side. It can be seen that this audio quality evaluation method can evaluate the audio quality of the audio to be evaluated through the audio quality evaluation model, thereby reducing the labor cost in the audio quality evaluation process. At the same time, during the audio quality evaluation process, the participation of evaluators is not required, avoiding the problem of inaccurate audio quality evaluation caused by differences in the levels of evaluators, and thus improving the accuracy of audio quality evaluation.

[0061] Please refer to Figure 4 , Figure 4 is a schematic diagram of a device for determining an audio quality evaluation model provided by an embodiment of the present application. As Figure 4 shown, the audio quality evaluation model determination device may include, but is not limited to:

[0062] An obtaining module 401, configured to obtain a plurality of audio samples, each audio sample including a dry voice segment;

[0063] A determining module 402, configured to use the dry voice segment of each audio sample to determine the pitch distribution vector of each audio sample, and determine the sampling probability of each audio sample according to the pitch distribution vector of each audio sample;

[0064] The determining module 402 is further configured to determine an audio sample set from the plurality of audio samples according to the sampling probability of each audio sample;

[0065] A training module 403, configured to train an initial audio quality evaluation model using the audio sample set to obtain a trained audio quality evaluation model; the audio quality evaluation model is used to determine the audio quality of the input audio.

[0066] In an optional implementation manner, the determining module 402 determines the sampling probability of each audio sample according to the pitch distribution vector of each audio sample, specifically: using the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample; among them, the distribution of each frame of the audio sample in each pitch in the pitch distribution vector is related to the size of the weight corresponding to the audio sample.

[0067] In an optional implementation manner, the determining module 402 uses the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample, specifically: combining the pitch distribution vectors of each audio sample into a pitch distribution matrix; determining the weight corresponding to each audio sample according to the pitch distribution matrix and vector B, where the dimension of vector B is the same as the dimension of the pitch distribution vector, and all elements in vector B are equal.

[0068] In an alternative embodiment, the sound quality evaluation model includes a feature learning model and a sound quality prediction model. The training module 403 trains the initial sound quality evaluation model using the audio sample set to obtain the trained sound quality evaluation model, which is specifically used for: inputting each audio sample in the audio sample set into the feature learning model to obtain the feature representation of each audio sample, where the feature learning model is trained based on the self-supervised learning method; inputting the feature representation of each audio sample into the sound quality prediction model to obtain the audio quality of each audio sample; and training the sound quality evaluation model using the audio quality of each audio sample to obtain the trained sound quality evaluation model.

[0069] In an alternative embodiment, the training module 403 inputs each audio sample in the audio sample set into the feature learning model to obtain the feature representation of each audio sample, which is specifically used for: for each audio sample in the audio sample set, performing frame addition and windowing processing on the audio sample to extract the Mel spectrum features of the audio sample; and inputting the Mel spectrum features of the audio sample into the feature learning model to obtain the feature representation of the audio sample.

[0070] In an alternative embodiment, the sound quality evaluation model further includes a long short-term memory module layer and a non-linear mapping layer; the training module 403 inputs the feature representation of each audio sample into the sound quality prediction model to obtain the audio quality of each audio sample, which is specifically used for: for each audio sample in the audio sample set, using the long short-term memory module layer to perform modeling processing on the feature representation of the audio sample to obtain the modeled feature representation, where the modeled feature representation does not include features irrelevant to the sound quality; using the non-linear mapping layer to perform non-linear mapping processing on the modeled feature representation to obtain the non-linearly mapped feature representation; and inputting the non-linearly mapped feature representation into the sound quality prediction model to obtain the audio quality of each audio sample.

[0071] In an alternative embodiment, the sound quality evaluation model determining device further includes a verification module 404, and the verification module 404 is used for: inputting each audio sample in the verification set into the trained sound quality evaluation model and outputting the audio quality of each audio sample in the verification set; determining the audio quality accuracy rate of the verification set according to the audio quality of each audio sample in the verification set; and verifying the trained sound quality evaluation model using the audio quality accuracy rate of the verification set to obtain the verification result.

[0072] Optionally, an embodiment of the present application further provides a sound quality evaluation device, which is used to obtain the audio to be evaluated; input the audio to be evaluated into the trained sound quality evaluation model to obtain the audio quality of the audio to be evaluated output by the trained sound quality evaluation model; where the trained sound quality evaluation model is Figure 1The trained sound quality evaluation model in the described sound quality evaluation model determination method.

[0073] It can be understood that the specific implementation and the beneficial effects that can be achieved of each module in the sound quality evaluation model determination device described in the embodiments of the present application can refer to the description of the foregoing method embodiments, and will not be elaborated herein.

[0074] Please refer to Figure 5 , Figure 5 which is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device described in the embodiments of the present application includes: a processor 501, a user interface 502, a communication interface 503, and a memory 504. Among them, the processor 501, the user interface 502, the communication interface 503, and the memory 504 can be connected through a bus or other means. In the embodiments of the present application, taking the connection through the bus as an example.

[0075] Among them, the processor 501 (or CPU (Central Processing Unit)) is the computing core and control core of the electronic device. It can parse various instructions in the electronic device and process various data of the electronic device. For example: the CPU can be used to parse the power-on and power-off instructions sent by the object to the electronic device and control the electronic device to perform power-on and power-off operations; another example: the CPU can transmit various interaction data between the internal structures of the electronic device, and so on. The user interface 502 is a medium for realizing the interaction and information exchange between the user and the electronic device. Its specific embodiment can include a display screen (Display) for output and a keyboard (Keyboard) for input, etc. It should be noted that the keyboard here can be a physical keyboard, a touch screen virtual keyboard, or a keyboard combining a physical and a touch screen virtual keyboard. The communication interface 503 can optionally include a standard wired interface, a wireless interface (such as Wi-Fi, a mobile communication interface, etc.), and is controlled by the processor 501 to send and receive data. The memory 504 (Memory) is a memory device in the electronic device, used to store programs and data. It can be understood that the memory 504 here can include both the built-in memory of the electronic device and, of course, the extended memory supported by the electronic device. The memory 504 provides a storage space, and this storage space stores the operating system of the electronic device, which can include but is not limited to: Android system, iOS system, Windows Phone system, etc. The present application does not make any limitations in this regard.

[0076] In the embodiments of the present application, the processor 501 executes the following operations by running the executable program code in the memory 504:

[0077] Obtain multiple audio samples, each audio sample including a dry voice segment; use the dry voice segment of each audio sample to determine the pitch distribution vector of each audio sample, and determine the sampling probability of each audio sample according to the pitch distribution vector of each audio sample; determine an audio sample set from the multiple audio samples according to the sampling probability of each audio sample; use the audio sample set to train an initial audio quality evaluation model to obtain a trained audio quality evaluation model; the audio quality evaluation model is used to determine the audio quality of the input audio.

[0078] In an alternative embodiment, the processor 501 determines the sampling probability of each audio sample according to the pitch distribution vector of each audio sample, specifically for: using the pitch distribution vector of each audio sample to determine the weight corresponding to each audio sample; wherein, the distribution of each frame of the audio sample in the pitch distribution vector in each pitch is related to the magnitude of the weight corresponding to the audio sample; using the weight corresponding to each audio sample to determine the sampling probability of each audio sample.

[0079] In an alternative embodiment, the processor 501 uses the pitch distribution vector of each audio sample to determine the sampling probability corresponding to each audio sample, specifically for: combining the pitch distribution vectors of each audio sample into a pitch distribution matrix; determining the weight corresponding to each audio sample according to the pitch distribution matrix and vector B, wherein the dimension of vector B is the same as that of the pitch distribution vector, and all elements in vector B are equal.

[0080] In an alternative embodiment, the audio quality evaluation model includes a feature learning model and an audio quality prediction model. The processor 501 uses the audio sample set to train the initial audio quality evaluation model to obtain a trained audio quality evaluation model, specifically for: inputting each audio sample in the audio sample set into the feature learning model to obtain the feature representation of each audio sample, and the feature learning model is trained based on the self-supervised learning method; inputting the feature representation of each audio sample into the audio quality prediction model to obtain the audio quality of each audio sample; using the audio quality of each audio sample to train the audio quality evaluation model to obtain a trained audio quality evaluation model.

[0081] In an alternative embodiment, the processor 501 inputs each audio sample in the audio sample set into the feature learning model to obtain the feature representation of each audio sample, specifically for: for each audio sample in the audio sample set, perform frame addition and windowing processing on the audio sample to extract the Mel spectrum features of the audio sample; input the Mel spectrum features of the audio sample into the feature learning model to obtain the feature representation of the audio sample.

[0082] In an alternative embodiment, the sound quality evaluation model further includes a long short-term memory module layer and a non-linear mapping layer; the processor 501 inputs the feature representation of each audio sample into the sound quality prediction model to obtain the audio quality of each audio sample, specifically as follows: for each audio sample in the audio sample set, the long short-term memory module layer is used to perform modeling processing on the feature representation of the audio sample to obtain the feature representation after the modeling processing, and the feature representation after the modeling processing does not include features irrelevant to the sound quality; the non-linear mapping layer is used to perform non-linear mapping processing on the feature representation after the modeling processing to obtain the feature representation after the non-linear mapping processing; the feature representation after the non-linear mapping processing is input into the sound quality prediction model to obtain the audio quality of each audio sample.

[0083] In an alternative embodiment, the processor 501 is further configured to: input each audio sample in the validation set into the trained sound quality evaluation model, and output the audio quality of each audio sample in the validation set; determine the audio quality accuracy rate of the validation set according to the audio quality of each audio sample in the validation set; and use the audio quality accuracy rate of the validation set to verify the trained sound quality evaluation model to obtain a verification result.

[0084] In specific implementation, the processor 501, the user interface 502, the communication interface 503, and the memory 504 described in the embodiments of the present application may implement the implementation manners of the electronic device described in the sound quality evaluation model determination method provided in the embodiments of the present application, and may also implement the implementation manners described in the sound quality evaluation model determination device provided in the embodiments of the present application, which will not be elaborated herein.

[0085] In the embodiments of the present application, the processor 501 further performs the following operations by running the executable program code in the memory 504:

[0086] Obtain the audio to be evaluated; input the audio to be evaluated into the trained sound quality evaluation model, and obtain the audio quality of the audio to be evaluated output by the trained sound quality evaluation model; wherein, the trained sound quality evaluation model is Figure 1 the trained sound quality evaluation model in the sound quality evaluation model determination method shown.

[0087] In specific implementation, the processor 501, the user interface 502, the communication interface 503, and the memory 504 described in the embodiments of the present application may implement the implementation manners of the electronic device described in the sound quality evaluation method provided in the embodiments of the present application, and may also implement the implementation manners described in the sound quality evaluation device provided in the embodiments of the present application, which will not be elaborated herein.

[0088] The embodiments of the present application further provide a computer-readable storage medium, which stores a computer program. The computer program includes program instructions. When the program instructions are executed by a processor, the audio quality evaluation method provided by the embodiments of the present application is implemented. For specific implementation manners, reference may be made to the implementation manners provided in the above steps, which will not be elaborated herein.

[0089] The embodiments of the present application further provide a computer program product or a computer program, which includes computer instructions stored in a computer-readable storage medium. The processor of the electronic device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions to enable the electronic device to execute the audio quality evaluation method provided by the embodiments of the present application. For specific implementation manners, reference may be made to the foregoing description, which will not be elaborated herein.

[0090] It should be noted that, for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, some steps may be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0091] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and the storage medium can include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc, etc.

[0092] The foregoing disclosure is only a part of the embodiments of the present application. Of course, the scope of rights of the present application cannot be limited thereby. Therefore, equivalent changes made according to the claims of the present application still fall within the scope covered by the present application.

Claims

1. A method for determining a sound quality assessment model, characterized in that: The method comprises: Obtaining a plurality of audio samples, each of the audio samples comprising a dry sound segment; Determine a pitch distribution vector of each of the audio samples using a dry sound segment of each of the audio samples; Determine the weight corresponding to each audio sample by using the pitch distribution vector of each audio sample; wherein the weight corresponding to the audio sample is related to the distribution of each pitch of each frame of the audio sample in the pitch distribution vector; Determine the sampling probability of each of the audio samples by using the weight corresponding to each of the audio samples; Determine an audio sample set from the plurality of audio samples according to a sampling probability of each of the audio samples; The initial sound quality assessment model is trained using the audio sample set to obtain a trained sound quality assessment model; the sound quality assessment model is used to determine the audio quality of the input audio.

2. The method according to claim 1, characterized in that The step of using the pitch distribution vector of each of the audio samples to determine the weight corresponding to each of the audio samples comprises: Combining the pitch distribution vectors of the respective audio samples into a pitch distribution matrix; The weight corresponding to each of the audio samples is determined according to the pitch distribution matrix and vector B, wherein the dimension of the vector B is the same as the dimension of the pitch distribution vector, and each element in the vector B is equal.

3. The method according to claim 1 or 2, characterized in that: The sound quality assessment model includes a feature learning model and a sound quality prediction model. The initial sound quality assessment model is trained using the audio sample set to obtain a trained sound quality assessment model, including: Inputting each audio sample in the audio sample set into the feature learning model to obtain a feature representation of each audio sample, wherein the feature learning model is trained based on a self-supervised learning method; Inputting the feature representation of each audio sample into the sound quality prediction model to obtain the audio quality of each audio sample; The sound quality assessment model is trained using the audio quality of each audio sample to obtain a trained sound quality assessment model.

4. The method according to claim 3, characterized in that The step of inputting each audio sample in the audio sample set into the feature learning model to obtain a feature representation of each audio sample includes: For each audio sample in the audio sample set, performing frame division and windowing processing on the audio sample to extract Mel spectrum features of the audio sample; The mel spectrum feature of the audio sample is input into the feature learning model to obtain the feature representation of the audio sample.

5. The method according to claim 4, characterized in that The sound quality evaluation model further includes a long short-term memory module layer and a nonlinear mapping layer; the step of inputting the feature representation of each audio sample into the sound quality prediction model to obtain the audio quality of each audio sample includes: For each audio sample in the audio sample set, using the long short-term memory module layer to perform modeling processing on the feature representation of the audio sample to obtain the feature representation after modeling processing, wherein the feature representation after modeling processing does not include features irrelevant to sound quality; Using the nonlinear mapping layer to perform nonlinear mapping on the feature representation after modeling processing to obtain the feature representation after nonlinear mapping processing; The feature representation after nonlinear mapping processing is input into the sound quality prediction model to obtain the audio quality of each audio sample.

6. The method according to claim 1, characterized in that The method further comprises: Input each audio sample in the validation set into the trained sound quality assessment model, and output the audio quality of each audio sample in the validation set; Determining the audio quality accuracy of the verification set according to the audio quality of each audio sample in the verification set; The trained sound quality assessment model is verified using the audio quality accuracy of the verification set to obtain a verification result.

7. A sound quality evaluation method, characterized in that: The method comprises: Obtain the audio to be evaluated; The audio to be evaluated is input into the trained sound quality evaluation model according to any one of claims 1 to 6, and the audio quality of the audio to be evaluated output by the trained sound quality evaluation model is obtained.

8. An electronic device, characterized in that: include: A processor, a user interface, a communication interface and a memory, wherein the processor, the user interface, the communication interface and the memory are interconnected, wherein the memory stores an executable program code, and the processor is used to call the executable program code to execute the method according to any one of claims 1 to 6, or to execute the method according to claim 7.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, wherein the computer program includes program instructions, and when the program instructions are executed by a processor, the processor is caused to execute the method according to any one of claims 1 to 6, or the method according to claim 7.

Citation Information

Patent Citations

  • Speech quality detection model training method and speech quality detection method

    CN112967735A

  • Fundamental frequency prediction method and device

    CN113990346A