A method, apparatus, storage medium, and device for evaluating the quality of synthesized speech.

By combining target audience feature modeling and scoring models, the accuracy problem of synthesized speech quality assessment under different scenarios and user groups is solved, realizing automated and personalized synthesized speech scoring, and reducing cost and time consumption.

CN116206631BActive Publication Date: 2026-04-03IFLYTEK CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing methods for evaluating the quality of synthesized speech are not accurate enough in different usage scenarios and user groups, making it difficult to meet customized needs. Furthermore, subjective evaluation is costly and time-consuming.

Method used

By acquiring the characteristic information of the target audience to build a profile space model, and using a scoring model to predict the target audience's rating of the synthesized speech, a personalized scoring model is constructed, reducing human intervention and achieving automated scoring.

Benefits of technology

It achieves accurate scoring of synthesized speech in different usage scenarios, saving manpower and time costs, improving evaluation efficiency, and adapting to the personalized needs of different user groups.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116206631B_ABST
    Figure CN116206631B_ABST
Patent Text Reader

Abstract

This application discloses a method, apparatus, storage medium, and computing device for quality evaluation of synthesized speech, relating to the field of speech processing technology. The method includes: acquiring first synthesized speech and feature information of a target audience evaluating the first synthesized speech; then, modeling a profile space based on the feature information of the target audience to obtain a profile space for the target audience; and finally, predicting the target audience's rating of the first synthesized speech using a scoring model based on the first synthesized speech and the target audience's profile space. This method models a profile space based on the feature information of the target audience, obtains the evaluator features of a virtual evaluator based on the profile space, and predicts the target audience's rating of the first synthesized speech using a scoring model based on the first synthesized speech and the evaluator features. Even in different usage scenarios, it can automatically construct the corresponding profile space of the target audience to predict the target audience's rating of the synthesized speech, saving time and effort and providing greater personalization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech processing technology, and in particular to a method, apparatus, storage medium, computing device, and computer program product for evaluating the quality of synthesized speech. Background Technology

[0002] With the continuous development of speech technology, especially speech synthesis, speech enhancement, and speech conversion technologies, a large amount of synthesized speech has been generated. The quality of speech synthesis, speech enhancement, or speech conversion systems is usually evaluated by assessing the quality of the synthesized speech output by these systems.

[0003] Quality assessment of synthesized speech can be achieved through subjective scoring by professional testers. However, subjective scoring requires a large number of testers to conduct hearing tests and provide perceptual ratings, making it extremely time-consuming and costly. Therefore, the industry is attempting to use deep learning-based assessment models to predict human subjective ratings of synthesized speech.

[0004] Because individual differences can lead to varying evaluations of the same speech by different testers, current methods typically use the average score of all testers, i.e., the mean opinion score (MOS), as the evaluation result. When training a deep learning-based evaluation model, MOS can be used as the training objective.

[0005] However, different speech synthesis products have different use cases and target different user groups. The individual differences brought about by the user groups make the above methods unsuitable for evaluating the quality of synthesized speech in customized scenarios. In some cases, the accuracy of the above methods will drop significantly and fail to meet business needs. Summary of the Invention

[0006] The main purpose of this application is to provide a method, apparatus, storage medium, and computing device for quality evaluation of synthesized speech, which can achieve quality evaluation of synthesized speech for a specified user group in different usage scenarios, ensure the accuracy of quality evaluation, and meet business needs.

[0007] Firstly, this application provides a method for evaluating the quality of synthesized speech, including:

[0008] Obtain the first synthesized speech and the feature information of the target population that evaluates the first synthesized speech;

[0009] Based on the characteristic information of the target population, a profile space model is performed to obtain the profile space of the target population;

[0010] Based on the first synthesized speech and the target audience's profile space, a rating model is used to predict the target audience's rating of the first synthesized speech.

[0011] In some possible implementations, the method further includes:

[0012] Multiple training samples are obtained, each of which includes a second synthesized speech, feature information of at least one evaluator, and the actual score of the second synthesized speech by the at least one evaluator.

[0013] The scoring model is obtained by training the model using the multiple training samples.

[0014] In some possible implementations, the step of training the model using the plurality of training samples to obtain the scoring model includes:

[0015] The second synthesized speech in the training samples is encoded to obtain speech features, and the feature information of the evaluator in the training samples is encoded to obtain evaluator features;

[0016] The speech features and the evaluator features are concatenated, and the concatenated features are input into a base network. The parameters of the base network are updated based on the predicted score output by the base network and the actual score to obtain the scoring model.

[0017] In some possible implementations, the evaluator features carry the evaluator's identifier.

[0018] In some possible implementations, predicting the target audience's rating of the first synthesized speech using a rating model based on the first synthesized speech and the target audience's profile space includes:

[0019] The virtual evaluator's evaluator characteristics are sampled from the profile space of the target population;

[0020] Based on the speech features of the first synthesized speech and the evaluator features of the virtual evaluator, a rating model is used to predict the target audience's rating of the first synthesized speech.

[0021] In some possible implementations, a profile space model is performed based on the feature information of the target population to obtain the profile space of the target population, including:

[0022] The feature information of the target population is input into the coding network to obtain the evaluator features;

[0023] Based on the characteristics of the evaluators, a probability distribution is fitted to obtain the profile space of the target population.

[0024] In some possible implementations, the characteristic information of the target population includes one or more of the following: age, gender, education level, occupation, preferences, and region.

[0025] Secondly, this application provides a synthesized speech quality evaluation device, comprising:

[0026] The communication module is used to acquire the first synthesized speech and the feature information of the target population that evaluates the first synthesized speech;

[0027] The portrait module is used to model the portrait space based on the feature information of the target population, and obtain the portrait space of the target population.

[0028] The evaluation module is used to predict the target audience's rating of the first synthesized speech based on the first synthesized speech and the target audience's profile space using a rating model.

[0029] In some possible implementations, the device further includes:

[0030] The training module is used to acquire multiple training samples. Each training sample includes a second synthesized speech, feature information of at least one evaluator, and the actual score of the second synthesized speech by the at least one evaluator. The scoring model is obtained by training the model using the multiple training samples.

[0031] In some possible implementations, the training module is specifically used for:

[0032] The second synthesized speech in the training samples is encoded to obtain speech features, and the feature information of the evaluator in the training samples is encoded to obtain evaluator features;

[0033] The speech features and the evaluator features are concatenated, and the concatenated features are input into a base network. The parameters of the base network are updated based on the predicted score output by the base network and the actual score to obtain the scoring model.

[0034] In some possible implementations, the evaluator features carry the evaluator's identifier.

[0035] In some possible implementations, the evaluation module is specifically used for:

[0036] The virtual evaluator's evaluator characteristics are sampled from the profile space of the target population;

[0037] Based on the speech features of the first synthesized speech and the evaluator features of the virtual evaluator, a rating model is used to predict the target audience's rating of the first synthesized speech.

[0038] In some possible implementations, the portrait module is specifically used for:

[0039] The feature information of the target population is input into the coding network to obtain the evaluator features;

[0040] Based on the characteristics of the evaluators, a probability distribution is fitted to obtain the profile space of the target population.

[0041] In some possible implementations, the characteristic information of the target population includes one or more of the following: age, gender, education level, occupation, preferences, and region.

[0042] Thirdly, this application provides a computing device, including: a processor, a memory, and a system bus;

[0043] The processor and the memory are connected via the system bus;

[0044] The memory is used to store one or more programs, the one or more programs including instructions, which, when executed by the processor, cause the processor to perform any of the above-described implementations of the synthetic speech quality evaluation method.

[0045] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a computing device, cause the computing device to perform any of the above-described methods for evaluating the quality of synthesized speech.

[0046] Fifthly, this application also provides a computer program product, which, when run on a computing device, causes the computing device to execute any of the above-described methods for evaluating the quality of synthesized speech.

[0047] As can be seen from the above technical solution, this application has at least the following beneficial effects:

[0048] This application provides a method for evaluating the quality of synthesized speech. Specifically, firstly, a first synthesized speech and feature information of a target group evaluating the first synthesized speech are obtained. Then, a profile space model is performed based on the feature information of the target group to obtain the profile space of the target group. Next, based on the first synthesized speech and the profile space of the target group, a scoring model is used to predict the score of the target group on the first synthesized speech.

[0049] The aforementioned method models a profile space based on the feature information of the target audience evaluating the first synthesized speech, obtaining a profile space for the target audience. Based on this profile space, the evaluator features of the virtual evaluator can be obtained. Using the first synthesized speech and these evaluator features, a scoring model can predict the target audience's rating of the first synthesized speech. This method can automatically construct corresponding profile spaces for different usage scenarios to predict the target audience's rating of the synthesized speech, saving time and effort and providing greater personalization. Furthermore, the scoring model can automatically score the input synthesized speech without human intervention during the evaluation phase, shortening the evaluation cycle and saving labor and time costs. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 A schematic diagram of the framework of a scoring model provided in an embodiment of this application;

[0052] Figure 2 A flowchart illustrating a method for evaluating the quality of synthesized speech provided in this application embodiment;

[0053] Figure 3 A schematic diagram of a voice quality evaluation interface provided in an embodiment of this application;

[0054] Figure 4 A schematic diagram illustrating a spatial modeling process for a target audience profile, provided as an embodiment of this application;

[0055] Figure 5 This is a schematic diagram of a synthesized speech quality evaluation device provided in an embodiment of this application. Detailed Implementation

[0056] The terms "first" and "second" used in the embodiments of this application are for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, a feature defined as "first" or "second" may explicitly or implicitly include one or more of that feature.

[0057] First, some technical terms involved in the embodiments of this application will be introduced.

[0058] Quality assessment refers to measuring the quality of synthesized speech using evaluation metrics. Quality assessment of synthesized speech can be divided into objective assessment and subjective assessment. Objective assessment uses a reference signal and algorithms to evaluate the quality of synthesized speech. In practical applications, the reference signal often contains other interference signals, leading to poor quality assessment results. Subjective assessment involves professional testers (evaluators) scoring the synthesized speech. Compared to objective assessment, subjective assessment results are more reliable.

[0059] Considering the high cost and difficulty in real-time assessment of subjective evaluation, the industry has proposed using objective models to predict and learn subjective human ratings. Specifically, this involves training the model using large-scale crowdsourced listening test sample data. This sample data typically includes speech samples and their corresponding subjective ratings. Early work attempted to use carefully designed handcrafted features to adjust simple statistical models such as linear regression, while recent work uses deep neural networks (DNNs) to extract rich feature representations from raw inputs (such as amplitude spectra, Mel spectra), and then further uses appropriate mapping functions based on these learned feature representations to obtain the corresponding predicted ratings.

[0060] Since the prediction target is a subjective human rating, the rating depends on the evaluator, and different evaluators may have significantly different opinions on the same speech. The industry typically uses the average score of a group of evaluators as the prediction target, ignoring the listener's preferences. However, different speech synthesis products have different use cases and target different user groups. Individual differences within these user groups mean that the above methods are not suitable for evaluating the quality of synthesized speech in customized scenarios. For example, speech synthesis products in the medical field are mainly aimed at doctors, nurses, and teachers. The evaluation of speech synthesized by medical speech synthesis products has limited reference value; in other words, the individual differences among doctors and teachers make the above methods unsuitable for applications in the medical field.

[0061] In view of this, this application provides a method for evaluating the quality of synthesized speech. This method can be executed by a computing device. The computing device can be a terminal, including but not limited to smartphones, smart wearable devices (such as smartwatches), tablets, laptops, or desktop computers. In some examples, the computing device can also be a server. Furthermore, the computing device can be a single computing device or multiple computing devices (e.g., a cluster including multiple computing devices). The computing device is equipped with software, such as a synthesized speech quality evaluation device, and the computing device executes the program code of the software device to perform the synthesized speech quality evaluation method of this application.

[0062] Specifically, the computing device can acquire the first synthesized speech and the feature information of the target group that evaluates the first synthesized speech, then perform profile space modeling based on the feature information of the target group to obtain the profile space of the target group, and then predict the rating of the target group on the first synthesized speech by the target group through a rating model based on the first synthesized speech and the profile space of the target group.

[0063] The above method models a profile space based on the feature information of the target audience evaluating the first synthesized speech, obtaining a profile space for the target audience. Based on this profile space, the evaluator features of the virtual evaluator can be obtained. Based on the first synthesized speech and the aforementioned evaluator features, a scoring model can predict the target audience's rating of the first synthesized speech. This method can automatically construct corresponding profile spaces for different use cases to predict the target audience's rating of the synthesized speech, saving time and effort and providing greater personalization.

[0064] Moreover, the above-mentioned scoring model does not require a new group of evaluators from the target group to score each time during the evaluation phase. In other words, it can automatically score the input synthesized speech without human intervention, which shortens the evaluation cycle and saves manpower and time costs.

[0065] To make the technical solution of this application clearer and easier to understand, the method for evaluating the quality of synthesized speech in the embodiments of this application will be described below with reference to the accompanying drawings.

[0066] The quality assessment method for synthesized speech based on virtual evaluators in a target population proposed in this application is divided into a data collection stage, a model training stage, and a prediction stage. The data collection stage typically requires constructing a sample set including a massive number of evaluators' ratings of the synthesized speech. Additionally, the sample set also collects a large amount of evaluator feature information. The model training stage mainly uses the synthesized speech, evaluator feature information, and evaluators' ratings of the synthesized speech to train a rating model and a profile space modeling model. The testing stage involves applying the models (such as the rating model and the profile space modeling model), inputting the synthesized speech to be evaluated (for ease of description, this application refers to the synthesized speech to be evaluated as the first synthesized speech, and the synthesized speech used for model training in the model training stage as the second synthesized speech), constructing a profile space for a specific target group, and predicting the ratings for the speech.

[0067] The following sections provide a detailed explanation of the data collection phase, model training phase, and prediction phase.

[0068] Phase 1: Data Collection

[0069] In this phase, the computing device acquires multiple training samples through data collection. Each training sample includes a second synthesized speech, feature information from at least one evaluator, and the at least one evaluator's actual score for the second synthesized speech.

[0070] To improve generalization and modeling capabilities, the collected second-generation synthesized speech typically covers a massive amount of spoken text; the speech synthesis systems used include both high-performing and low-performing systems, as well as historical and recent ones. These systems can be commercially available products or open-source systems.

[0071] To model the evaluator profile space, a large number of representative evaluators' ratings of the aforementioned second synthesized speech can be collected. Simultaneously, the corresponding evaluator feature information is collected, which includes, but is not limited to, one or more of the following: age, gender, education level, region, occupation, and preferences.

[0072] These basic features only need to adequately describe the characteristics of the evaluator group; the specific types and number of features are not particularly limited. For features such as age and gender, the number of evaluators should be as balanced as possible. The distribution of evaluators across other feature dimensions should not significantly differ from the distribution in a real population to ensure the representativeness of the training data. Specifically, the dataset should contain at least several thousand second-generation synthesized speech samples. The number of different speech synthesis systems these samples belong to should be balanced, for example, the ratio of samples to different speech synthesis systems should be a predetermined ratio. Each sample can be evaluated by tens of thousands of evaluators to obtain real scores. These scores can be obtained using large-scale crowdsourced listening tests. The final dataset is X = [x1, x2, ..., x...]. N There are N second-synthesized speech words, L = [l1, l2, ..., l K There are K evaluators in total, and each evaluator has L... k =[u1,u2,…,u T There are T basic features in total, and each second synthesized speech has M scores Y from evaluators. n =[y1,y2,…,y M In general, audiometry test designs aim to ensure that K > M to reduce the workload of evaluators and guarantee audiometry quality. Different second-synthesized speech samples may be scored by different evaluators.

[0073] Phase Two: Model Training

[0074] The computing device can use multiple training samples obtained in Phase 1 to train the model and obtain a scoring model. The training samples can be sample pairs formed by the second synthesized speech, the evaluator's feature information, and the evaluator's rating of the second synthesized speech (specifically, represented as a speech-evaluator-rating triple). Traditional methods directly average the ratings from different evaluators to obtain sample pairs (specifically, speech-average rating), ignoring the evaluator's influence, and then inputting them into the scoring model for training. Traditional methods directly reduce the data volume from N×M to 1 / M (i.e., N) of the original. However, this approach cannot fully utilize the learning capabilities of data-driven neural networks; furthermore, the evaluation ratings cannot represent the evaluations of a personalized group. Therefore, this application proposes a speech-evaluator-rating sample pair, using a large amount of evaluator feature information to train the scoring model, and using the scoring model to achieve automatic personalized evaluation of synthesized speech.

[0075] Specifically, the computing device can encode the second synthesized speech in the training samples to obtain speech features, and encode the evaluator's feature information in the training samples to obtain evaluator features. Then, the speech features and the evaluator features are concatenated, and the concatenated features are input into the base network. The parameters of the base network are updated according to the predicted score output by the base network and the actual score to obtain the scoring model.

[0076] The model training process will be explained in detail below with reference to the accompanying diagram.

[0077] See Figure 1 The diagram shows the framework of the scoring model, which includes a speech coding module, an evaluator coding module, and a score prediction module. The model processing flow is divided into three parts: speech feature learning, evaluator coding learning, and score prediction learning. Given an input sample: input synthesized speech x... n The evaluator's characteristic information k Actual evaluation score y m Under these conditions, the following three parts will be explained in detail:

[0078] 1. Speech Feature Learning: The computing device uses a speech coding module to learn the second synthesized speech, such as synthesized speech x. n From raw signal encoding to hidden layer signal representation h x The speech coding module can be implemented using a parametric neural network model or a parameterless feature extraction network.

[0079] 2. Evaluator Encoding Network Learning: The evaluator encoding module is used to learn the evaluator's feature information. k The evaluator's encoded hidden layer feature vector h is obtained by embedding representation. lThis can also be referred to as evaluator features. Specifically, in this embodiment, the evaluator encoding module can take various forms. When the evaluator's feature information is discrete, the computing device can use one-hot encoding to encode the discrete features. When the evaluator's feature information is continuous, such as text-based feature information, the computing device can also directly use word2vec or BERT to embed and encode the evaluator's feature information to obtain the evaluator features. In particular, even different people with completely identical feature information will give different scores for the speech. k An additional reviewer identifier (ID) can be added to identify the reviewer. The reviewer ID can be a unique serial number, a random number, or an encoding, etc.

[0080] 3. Rating Prediction Learning: Based on the speech features h obtained in Part 1 x And the evaluator's hidden features h learned from Part 2 l The features are then concatenated and input into the scoring prediction module to obtain the predicted score. The model training objective can be to minimize the true score y. n and predicted scores The distance. The rating prediction module can include a base network, which can be a deep neural network (DNN). The loss function of this network can be set according to requirements, for example, it can be set to the cross-entropy loss function. The computing device can determine the true rating y using the loss function. n and predicted scores The distance is used to update the parameters of the base network until the distance is minimized. When the distance is minimized, training can stop, and a scoring model can be obtained based on the trained base network.

[0081] Phase 3: Model Application

[0082] The computing device can apply the Phase 2 scoring model to achieve personalized quality assessment of synthesized speech. A detailed explanation follows with reference to the accompanying diagram.

[0083] See Figure 2 The flowchart shown illustrates a method for evaluating the quality of synthesized speech, which includes:

[0084] S202: The computing device acquires the first synthesized speech and the feature information of the target population that evaluates the first synthesized speech.

[0085] The first synthesized speech refers to the synthesized speech to be evaluated. This synthesized speech can be generated in real time by a speech synthesis system. A computing device can call the interface of the speech synthesis system, such as an application programming interface (API), to obtain the first synthesized speech to be evaluated. It should be noted that the first synthesized speech can be one or more speech samples; the embodiments of this application do not limit the number of first synthesized speech samples.

[0086] The target audience's characteristics include one or more of the following: age, gender, education level, occupation, preferences, and region. The computing device can receive the target audience's characteristics configured by the user. For example, the computing device can receive the target audience's characteristics configured by the user through a voice quality assessment interface.

[0087] See Figure 3 The diagram shows a speech quality assessment interface 300. This interface includes a user profile configuration component 302, which comprises age configuration controls 3021, gender configuration controls 3022, education level configuration controls 3023, occupation configuration controls 3024, preference configuration controls 3025, and region configuration controls 3026. Users can configure the characteristic information of a target user group using these controls. For example, a user can configure an age of 20-30 years old, gender as "all" (including all genders), education level as undergraduate, occupation as self-media, and preference as fashion. The speech quality assessment interface 300 may also include a confirm control 304 and a cancel control 306. When the user clicks the confirm control 304, the computing device 200 obtains the characteristic information of the target user group configured by the user. When the user clicks the cancel control 306, the computing device 200 cancels the configuration of the target user group.

[0088] It should be noted that when configuring the feature information of the target audience, some dimensions of feature information can be configured by default. For example, the feature information of gender and preference dimensions can be configured by default. Accordingly, the computing device can sample from the dataset to obtain the default feature information, thereby obtaining complete feature information.

[0089] S204: The computing device performs profile space modeling based on the feature information of the target population to obtain the profile space of the target population.

[0090] Specifically, the computing device can input the feature information of the target population into the encoding network to obtain the evaluator features, and then perform probability distribution fitting based on the evaluator features to obtain the profile space of the target population.

[0091] See Figure 4The flowchart shown illustrates the target audience profile spatial modeling process. In addition to the basic network, the computing device can also input the evaluator's feature information. k The hidden layer feature vector h is obtained by passing it to the encoding network. l Then the computing device can be based on h l We perform profile space modeling to realize the high-dimensional and complex probability distribution P(z|h) of the listener profile space z. l The accurate fitting of the target population profile space can be achieved through probabilistic models. For example, the computing device can use a Gaussian mixture model or a flow model to model the target population profile space. The parameters of the probabilistic model (such as a Gaussian mixture model or a flow model) can be obtained through maximum likelihood estimation.

[0092] S206: The computing device predicts the target audience's rating of the first synthesized speech based on the first synthesized speech and the target audience's profile space using a rating model.

[0093] Specifically, the computing device can sample the evaluator features of the virtual evaluator from the profile space of the target population, and then predict the target population's rating of the first synthesized speech based on the speech features of the first synthesized speech and the evaluator features of the virtual evaluator through a rating model.

[0094] To facilitate understanding, examples will be provided below.

[0095] In this example, the given evaluator's feature information l k The evaluator coding network and the profile space modeling model are input sequentially to obtain the controlled profile space z of the target population. For example... Figure 4 As shown, the computing device can randomly sample from a z-distribution to create the evaluator features of the virtual evaluator, typically the hidden layer feature vectors. For example, if a user wants to obtain ratings from young short video creators on the quality of synthesized speech generated by a speech synthesis system, the target group's features can be configured as follows: occupation "short video creator," age "25," and other features randomly sampled from the basic information of listeners with the occupation of short video creators in the training data. This constructs the feature information for the specified target group. Then, this feature information is input into the evaluator encoding network and the profile space modeling model to obtain the profile space of the 25-year-old short video creator group (target group). Then, the evaluator features of a massive number of virtual evaluators are continuously sampled from this profile space; these evaluator features are hidden layer feature vectors. The computing device will then process the desired evaluation speech x... n Hidden feature vectors of virtual evaluators Input the scoring model to obtain personalized scores for a given speech from a virtual 25-year-old video creator.

[0096] Given the above model, developers of speech synthesis systems, or related technical institutions, can use real-time automated evaluation methods to greatly reduce the time and money costs of manual evaluation during the development of speech synthesis systems. At the same time, they can perform targeted optimization based on the target user group, provide focused high-standard services to specific groups of users, and greatly improve the service experience.

[0097] It should be noted that the above examples are illustrative based on the input being feature information with control conditions. In other possible implementations of this application, the input can also be the evaluator features (portrait feature representation) of a virtual evaluator without control conditions. In other words, the computing device can randomly sample the virtual evaluator profile feature representations of a target number of virtual evaluators based on the existing evaluator profile space. The computing device can also obtain the feature information of a specified target group as control conditions to obtain the profile space distribution of the target cluster, and sample from the profile space distribution of the target cluster to obtain specific virtual evaluator profile feature representations. These high-level feature representations of virtual evaluators with or without control conditions, along with the synthesized speech to be evaluated, are used as input for quality evaluation to obtain simulated subjective scores or personalized scores for the specified target group.

[0098] Based on the methods provided in the embodiments of this application, the embodiments of this application also provide a synthesized speech quality evaluation device corresponding to the above methods. The units / modules described in the embodiments of this application can be implemented in software or hardware. The names of the units / modules do not, in some cases, constitute a limitation on the unit / module itself.

[0099] See Figure 5 This is a schematic diagram of the structure of a synthesized speech quality evaluation device provided in an embodiment of this application. The synthesized speech quality evaluation device 500 includes:

[0100] Communication module 502 is used to acquire the first synthesized speech and the feature information of the target group that evaluates the first synthesized speech;

[0101] The portrait module 504 is used to perform portrait space modeling based on the feature information of the target population to obtain the portrait space of the target population.

[0102] Evaluation module 506 is used to predict the target audience's rating of the first synthesized speech based on the first synthesized speech and the target audience's profile space using a rating model.

[0103] In some possible implementations, the device 500 further includes:

[0104] The training module 508 is used to acquire multiple training samples. Each training sample includes a second synthesized speech, feature information of at least one evaluator, and the actual score of the second synthesized speech by the at least one evaluator. The scoring model is obtained by training the model using the multiple training samples.

[0105] In some possible implementations, the training module 508 is specifically used for:

[0106] The second synthesized speech in the training samples is encoded to obtain speech features, and the feature information of the evaluator in the training samples is encoded to obtain evaluator features;

[0107] The speech features and the evaluator features are concatenated, and the concatenated features are input into a base network. The parameters of the base network are updated based on the predicted score output by the base network and the actual score to obtain the scoring model.

[0108] In some possible implementations, the evaluator features carry the evaluator's identifier.

[0109] In some possible implementations, the evaluation module 506 specifically needs to:

[0110] The virtual evaluator's evaluator characteristics are sampled from the profile space of the target population;

[0111] Based on the speech features of the first synthesized speech and the evaluator features of the virtual evaluator, a rating model is used to predict the target audience's rating of the first synthesized speech.

[0112] In some possible implementations, the portrait module 504 is specifically used for:

[0113] The feature information of the target population is input into the coding network to obtain the evaluator features;

[0114] Based on the characteristics of the evaluators, a probability distribution is fitted to obtain the profile space of the target population.

[0115] In some possible implementations, the characteristic information of the target population includes one or more of the following: age, gender, education level, occupation, preferences, and region.

[0116] This application also provides a computing device, including: a processor, a memory, and a system bus;

[0117] The processor and memory are connected via the system bus;

[0118] The memory is used to store one or more programs, wherein the one or more programs include instructions that, when executed by the processor, cause the processor to perform any of the above-described implementations of the synthesized speech quality assessment method.

[0119] This application also provides a computer-readable storage medium storing instructions that, when executed on a computing device, cause the computing device to perform any of the above-described implementations of the synthesized speech quality evaluation method.

[0120] This application also provides a computer program product that, when run on a computing device, causes the computing device to execute any of the above-described methods for evaluating the quality of synthesized speech.

[0121] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that all or part of the steps in the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, a server, or a network communication device such as a media gateway, etc.) to execute the methods described in various embodiments or some parts of the embodiments of this application.

[0122] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus disclosed in the embodiments, since it corresponds to the method disclosed in the embodiments, the description is relatively simple; relevant parts can be referred to in the method section.

[0123] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0124] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for evaluating the quality of synthesized speech, characterized in that, include: Obtain the first synthesized speech and the feature information of the target population that evaluates the first synthesized speech; Based on the characteristic information of the target population, a profile space model is performed to obtain the profile space of the target population; Based on the first synthesized speech and the target audience's profile space, a rating model is used to predict the target audience's rating of the first synthesized speech.

2. The method according to claim 1, characterized in that, The method further includes: Multiple training samples are obtained, each of which includes a second synthesized speech, feature information of at least one evaluator, and the actual score of the second synthesized speech by the at least one evaluator. The scoring model is obtained by training the model using the multiple training samples.

3. The method according to claim 2, characterized in that, The step of training the model using the multiple training samples to obtain the scoring model includes: The second synthesized speech in the training samples is encoded to obtain speech features, and the feature information of the evaluator in the training samples is encoded to obtain evaluator features; The speech features and the evaluator features are concatenated, and the concatenated features are input into a base network. The parameters of the base network are updated based on the predicted score output by the base network and the actual score to obtain the scoring model.

4. The method according to claim 3, characterized in that, The evaluator's characteristics include the evaluator's identifier.

5. The method according to any one of claims 1 to 4, characterized in that, The step of predicting the target audience's rating of the first synthesized speech based on the first synthesized speech and the target audience's profile space using a rating model includes: The virtual evaluator's evaluator characteristics are sampled from the profile space of the target population; Based on the speech features of the first synthesized speech and the evaluator features of the virtual evaluator, a rating model is used to predict the target audience's rating of the first synthesized speech.

6. The method according to any one of claims 1 to 4, characterized in that, The step of modeling the profile space based on the feature information of the target population to obtain the profile space of the target population includes: The feature information of the target population is input into the coding network to obtain the evaluator features; Based on the characteristics of the evaluators, a probability distribution is fitted to obtain the profile space of the target population.

7. The method according to any one of claims 1 to 4, characterized in that, The target population's characteristics include one or more of the following: age, gender, education level, occupation, preferences, and region.

8. A device for evaluating the quality of synthesized speech, characterized in that, The device includes: The communication module is used to acquire the first synthesized speech and the feature information of the target population that evaluates the first synthesized speech; The portrait module is used to model the portrait space based on the feature information of the target population, and obtain the portrait space of the target population. The evaluation module is used to predict the target audience's rating of the first synthesized speech based on the first synthesized speech and the target audience's profile space using a rating model.

9. A computing device, characterized in that, include: Processor, memory, system bus; The processor and the memory are connected via the system bus; The memory is used to store one or more programs, the one or more programs including instructions that, when executed by the processor, cause the processor to perform the method according to any one of claims 1-7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores instructions that, when executed on a computing device, cause the computing device to perform the method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Recommendation list generation method and apparatus, and electronic device

    CN111159558A

  • Matching support system, matching support method, and matching support program

    JP2020201861A