Speech synthesis model evaluation method and device, and storage medium

By acquiring matching degree evaluations of multiple style instructions and speech data, this method solves the problem that traditional evaluation methods cannot evaluate the style instruction matching degree of novel speech synthesis models, and achieves efficient evaluation of novel speech synthesis models.

CN119889353BActive Publication Date: 2025-12-30CHINA UNITED NETWORK COMM GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510020821.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-12-30
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

Traditional speech synthesis model evaluation methods cannot assess the matching degree between the synthesized speech and style instructions of novel speech synthesis models, resulting in difficult and costly evaluation.

Method used

By acquiring multiple style instructions, the speech data corresponding to each style instruction is determined based on the speech synthesis model under test, and the matching degree between the style instructions and the speech data is evaluated, including the assessment of timbre, emotion, speech rate and pitch features.

Benefits of technology

It enables effective evaluation of novel speech synthesis models, accurately assesses the matching degree between the output speech of the speech synthesis model and style instructions, and reduces evaluation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119889353B_ABST
    Figure CN119889353B_ABST
Patent Text Reader

Abstract

The application provides a speech synthesis model evaluation method and device and a storage medium, relates to the technical field of computers, and can evaluate a new speech synthesis model. The method comprises the following steps: acquiring a plurality of style instructions, wherein the style instructions are used for indicating speech features of speech data to be output; determining speech data corresponding to each style instruction based on a to-be-tested speech synthesis model; and evaluating the to-be-tested speech synthesis model according to matching degrees between the plurality of style instructions and the respective corresponding speech data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and in particular to a method, apparatus and storage medium for evaluating speech synthesis models. Background Technology

[0002] With the development of speech synthesis technology, especially based on large language models, new speech synthesis models with diverse performance dimensions, such as ChatTTS, PromptTTS, and StyleTTS, have emerged, offering multiple speaker roles, emotions, and styles. Traditional speech synthesis models only input content instructions, and the synthesized speech output only needs to conform to those instructions. New speech synthesis models, however, input not only content instructions but also style instructions, requiring the synthesized speech output to meet both content and style requirements.

[0003] Currently, traditional methods for evaluating speech synthesis models use mean opinions score (MOS), comparative mean opinion score (CMOS), or A / B tests. However, these methods can only assess the matching degree between the synthesized speech and the content instructions, but cannot assess the matching degree between the synthesized speech and the style instructions, meaning they cannot evaluate novel speech synthesis models. Summary of the Invention

[0004] This application provides a method, apparatus, and storage medium for evaluating speech synthesis models, which can evaluate novel speech synthesis models.

[0005] To achieve the above objectives, this application adopts the following technical solution:

[0006] In a first aspect, this application provides a method for evaluating a speech synthesis model. The method includes: acquiring multiple style instructions, which are used to indicate the speech features of the speech data to be output; determining the speech data corresponding to each style instruction based on the speech synthesis model under test; and evaluating the speech synthesis model under test according to the matching degree between the multiple style instructions and their respective corresponding speech data.

[0007] In one possible implementation, the speech features include at least one of the following: timbre features; emotion features; speech rate features; and pitch features.

[0008] In one possible implementation, based on the speech synthesis model under test, the speech data corresponding to each style instruction is determined, including: acquiring multiple content instructions, which are used to indicate the text content of the speech data to be output; for each style instruction, the style instruction and any one of the multiple content instructions are input into the speech synthesis model under test to obtain the speech data corresponding to the style instruction, so as to obtain the speech data corresponding to each style instruction.

[0009] In one possible implementation, the method further includes: for each style instruction, obtaining the semantic features of the style instruction and the audio features of the speech data corresponding to the style instruction; and determining the matching degree between the style instruction and the speech data based on the semantic features of the style instruction and the audio features of the speech data.

[0010] In one possible implementation, the matching degree between the style instruction and the speech data is determined based on the semantic features of the style instruction and the audio features of the speech data, including: using the similarity between the semantic features of the style instruction and the audio features of the speech data as the matching degree between the style instruction and the speech data.

[0011] In one possible implementation, the speech synthesis model under test is evaluated based on the matching degree between multiple style instructions and their corresponding speech data. This includes: obtaining evaluation parameters based on the matching degree between multiple style instructions and their corresponding speech data, wherein the evaluation parameters are the average and standard deviation of the matching degree; and evaluating the speech synthesis model under test based on the evaluation parameters and a preset equalization coefficient, wherein the preset equalization coefficient is used to adjust the weights of the average and standard deviation of the matching degree.

[0012] Secondly, this application provides a speech synthesis model evaluation device, which includes: a communication unit and a processing unit; the communication unit is used to acquire multiple style instructions, the style instructions being used to indicate the speech features of the speech data to be output; the processing unit is used to determine the speech data corresponding to each style instruction based on the speech synthesis model under test; the processing unit is also used to evaluate the speech synthesis model under test according to the matching degree between the multiple style instructions and their respective corresponding speech data.

[0013] In one possible implementation, the speech features include at least one of the following: timbre features; emotion features; speech rate features; and pitch features.

[0014] In one possible implementation, the communication unit is further configured to acquire multiple content instructions, which are used to indicate the text content of the speech data to be output; for each style instruction, the processing unit is further configured to input the style instruction and any one of the multiple content instructions into the speech synthesis model under test to obtain the speech data corresponding to the style instruction, thereby acquiring the speech data corresponding to each style instruction.

[0015] In one possible implementation, the method further includes: for each style instruction, the communication unit is further configured to acquire the semantic features of the style instruction and the audio features of the speech data corresponding to the style instruction; the processing unit is further configured to determine the matching degree between the style instruction and the speech data based on the semantic features of the style instruction and the audio features of the speech data.

[0016] In one possible implementation, the processing unit is further configured to use the similarity between the semantic features of the style instruction and the audio features of the speech data as the matching degree between the style instruction and the speech data.

[0017] In one possible implementation, the processing unit is further configured to obtain evaluation parameters based on the matching degree between multiple style instructions and their corresponding speech data, wherein the evaluation parameters are the average value and standard deviation of the matching degree; the processing unit is further configured to evaluate the speech synthesis model under test based on the evaluation parameters and preset equalization coefficients, wherein the preset equalization coefficients are used to adjust the weights of the average value and standard deviation of the matching degree.

[0018] Thirdly, this application provides a speech synthesis model evaluation device, which includes: a processor and a communication interface; the communication interface and the processor are coupled, and the processor is used to run computer programs or instructions to implement the speech synthesis model evaluation method as described in the first aspect and any possible implementation of the first aspect.

[0019] Fourthly, this application provides a computer-readable storage medium storing instructions that, when executed on a terminal, cause the terminal to perform the speech synthesis model evaluation method as described in the first aspect and any possible implementation thereof.

[0020] Fifthly, this application provides a computer program product containing instructions that, when run on a speech synthesis model evaluation device, causes the speech synthesis model evaluation device to perform the speech synthesis model evaluation method as described in the first aspect and any possible implementation thereof.

[0021] In a sixth aspect, this application provides a chip including a processor and a communication interface, the communication interface being coupled to the processor, the processor being used to run computer programs or instructions to implement the speech synthesis model evaluation method as described in the first aspect and any possible implementation thereof.

[0022] Specifically, the chip provided in this application also includes a memory for storing computer programs or instructions.

[0023] The above technical solution brings at least the following beneficial effects: It acquires multiple style instructions, determines the speech data corresponding to each style instruction based on the speech synthesis model under test, and evaluates the speech synthesis model under test based on the matching degree between the multiple style instructions and their corresponding speech data. In other words, the speech synthesis model evaluation method provided in this application evaluates the speech synthesis model based on the matching degree between the style instructions input to the speech synthesis model and the speech data output by the speech synthesis model, thereby enabling the evaluation of novel speech synthesis models. Attached Figure Description

[0024] Figure 1 This is a schematic diagram illustrating the composition of a speech synthesis model evaluation device provided in an embodiment of this application;

[0025] Figure 2 A flowchart illustrating a speech synthesis model evaluation method provided in this application embodiment;

[0026] Figure 3 A flowchart illustrating another speech synthesis model evaluation method provided in this application embodiment;

[0027] Figure 4 A flowchart for determining the matching degree between style instructions and voice data is provided as an embodiment of this application;

[0028] Figure 5 A flowchart for determining an evaluation model is provided as an embodiment of this application;

[0029] Figure 6 A flowchart for determining target style text provided in this application embodiment;

[0030] Figure 7 A schematic diagram illustrating model training as provided in an embodiment of this application;

[0031] Figure 8 A schematic diagram of an audio encoder provided in an embodiment of this application;

[0032] Figure 9 A flowchart illustrating another speech synthesis model evaluation method provided in this application embodiment;

[0033] Figure 10 A flowchart for determining system compliance is provided as an embodiment of this application;

[0034] Figure 11 A flowchart for evaluating a single data point in a PromptTTS model is provided as an embodiment of this application;

[0035] Figure 12 This is a schematic diagram of the structure of a speech synthesis model evaluation device provided in an embodiment of this application. Detailed Implementation

[0036] The speech synthesis model evaluation method, apparatus, and storage medium provided in the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0037] In this article, the term "and / or" is merely a description of the relationship between related objects, indicating that there can be three relationships. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone.

[0038] The terms "first" and "second," etc., used in the specification and drawings of this application are used to distinguish different objects or to distinguish different treatments of the same object, rather than to describe a specific order of objects.

[0039] Furthermore, the terms "comprising" and "having," and any variations thereof, used in the description of this application are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the steps or units listed, but may optionally include other steps or units not listed, or may optionally include other steps or units inherent to such process, method, product, or apparatus.

[0040] It should be noted that in the embodiments of this application, the words "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of the words "exemplary" or "for example" is intended to present the relevant concepts in a specific manner.

[0041] In the description of this application, unless otherwise stated, "a plurality of" means two or more.

[0042] As described in the background section, traditional evaluation methods for speech synthesis models are increasingly ill-suited to the performance characteristics of new speech synthesis models. For example, in traditional manual evaluation methods, subjective MOS can only evaluate the overall score of the synthesized speech, CMOS can only evaluate timbre similarity, and A / B tests can only compare two different speech synthesis models based on preferences. These evaluation methods are all biased towards the overall attributes of the speech synthesis model, making it difficult to accurately evaluate the new characteristics exhibited by new speech synthesis models (i.e., whether the synthesized speech follows the input style instructions), and they also incur high manual and time costs.

[0043] Furthermore, to save on the human resources required for evaluation, various model-assisted evaluation algorithms have been developed. For example, there are multiple model-assisted algorithms for subjective MOS evaluation, represented by UTMOS. These methods collect subjective MOS data from speech and human evaluation, use deep learning models based on pre-trained models such as wav2vec or wavLM for fine-tuning, and then combine the results of multiple models using model ensemble learning methods to ultimately predict subjective MOS.

[0044] However, none of the above methods can solve the evaluation problem of novel speech synthesis models with multiple speaker roles, multiple emotions, and multiple styles. That is, the above methods can only evaluate the matching degree between the speech synthesized by the speech synthesis model and the content instructions, but cannot evaluate the matching degree between the speech synthesized by the speech synthesis model and the style instructions.

[0045] In view of this, embodiments of this application provide a speech synthesis model evaluation method that can adapt to the evaluation needs of the current development of speech synthesis technology. The method includes: acquiring multiple style instructions; determining the speech data corresponding to each style instruction based on the speech synthesis model under test; and evaluating the speech synthesis model under test based on the matching degree between the multiple style instructions and their respective corresponding speech data. In other words, the speech synthesis model evaluation method provided in this application evaluates the speech synthesis model based on the matching degree between the style instructions input to the speech synthesis model and the speech data output by the speech synthesis model, thereby enabling the evaluation of novel speech synthesis models.

[0046] For example, Figure 1 This is a schematic diagram of the composition of a speech synthesis model evaluation device 10 provided in an embodiment of this application. The speech synthesis model evaluation device 10 may include a processor 101 and a bus 102.

[0047] Furthermore, the speech synthesis model evaluation device 10 may also include a communication interface 103 and a memory 104. The processor 101, the memory 104, and the communication interface 103 can be connected via a bus 102.

[0048] The processor 101 can be a central processing unit (CPU), a general-purpose processor, a network processor (NP), a digital signal processor (DSP), a microprocessor, a microcontroller, a programmable logic device (PLD), or any combination thereof. The processor 101 can also be other devices with processing capabilities, such as circuits, devices, or software modules, without limitation.

[0049] Bus 102 is used to transmit information between the components included in the speech synthesis model evaluation device 10.

[0050] Communication interface 103 is used to communicate with other devices or other communication networks. These other communication networks can be Ethernet, radio access network (RAN), wireless local area networks (WLAN), etc. Communication interface 103 can be a module, circuit, communication interface, or any device capable of enabling communication.

[0051] Memory 104 is used to store instructions. These instructions can be computer programs.

[0052] The memory 104 can be a read-only memory (ROM) or other type of static storage device that can store static information and / or instructions; it can also be a random access memory (RAM) or other type of dynamic storage device that can store information and / or instructions; it can also be an electrically erasable programmable read-only memory (EEPROM), a compact disc read-only memory (CD-ROM) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, etc., without limitation.

[0053] It should be noted that the memory 104 can exist independently of the processor 101 or can be integrated with the processor 101. The memory 104 can be used to store instructions, program code, or some data. The memory 104 can be located inside or outside the speech synthesis model evaluation device 10, without limitation. The processor 101 is used to execute the instructions stored in the memory 104 to implement the speech synthesis model evaluation method provided in the following embodiments of this application.

[0054] In one example, processor 101 may include one or more CPUs, such as CPU0 and CPU1 (not shown in the figure).

[0055] As an optional implementation, the speech synthesis model evaluation device 10 includes multiple processors.

[0056] As an optional implementation, the speech synthesis model evaluation device 10 also includes an output device and an input device. For example, the input device is a keyboard, mouse, microphone, or joystick, and the output device is a display screen, speaker, or other similar device.

[0057] It should be noted that the speech synthesis model evaluation device 10 can be a desktop computer, laptop computer, network server, mobile phone, tablet computer, wireless terminal, embedded device, chip system, or other device. Figure 1 Equipment with a similar structure. Furthermore... Figure 1 The composition shown does not constitute a basis for this. Figure 1 The limitations of each device in the process, except Figure 1 In addition to the components shown, Figure 1 The various devices may include more or fewer components than illustrated, or combine certain components, or have different component arrangements.

[0058] In this embodiment of the application, the chip system may be composed of chips or may include chips and other discrete devices.

[0059] Furthermore, the actions, terms, etc., involved in the various embodiments of this application can be referenced interchangeably without limitation. The message names or parameter names in the messages exchanged between the various devices in the embodiments of this application are merely examples, and other names may be used in specific implementations without limitation.

[0060] The speech synthesis model evaluation method provided in the embodiments of this application is described below with reference to the accompanying drawings. The actions, terminology, etc., involved in the various embodiments of this application can be referred to mutually without limitation. The message names or parameter names in the messages between various devices in the embodiments of this application are merely examples; other names may be used in specific implementations without limitation. The actions involved in the various embodiments of this application are merely examples; other names may be used in specific implementations, such as replacing "included in" with "carried on" or "carried in," etc.

[0061] It should be noted that there are no restrictions on the entity that can implement the technical solution in this application.

[0062] like Figure 2 As shown in the embodiments of this application, a method for evaluating speech synthesis models is proposed, the method comprising:

[0063] S201, Obtain multiple style commands.

[0064] The style instruction is used to indicate the speech features of the speech data to be output.

[0065] For example, the aforementioned speech features may include at least one of the following: timbre features, emotion features, speech rate features, or pitch features. Among them, timbre features may include gender features and age features.

[0066] Optionally, the above-mentioned speech features may also include intonation style features or volume features.

[0067] For example, the above-mentioned gender characteristic can be a male voice or a female voice. The above-mentioned age characteristic can be a toddler, child, teenager, young adult, middle-aged, or elderly person. The above-mentioned emotional characteristic can be happiness, sadness, anger, fear, or confusion. The above-mentioned speech rate characteristic can be slow, medium, or fast. The above-mentioned pitch characteristic can be low, medium, or high. The above-mentioned intonation style characteristic can be steady, high-pitched, low, or loud. The above-mentioned volume characteristic can be low, medium, or high.

[0068] For example, the style instruction could be: A young male voice is speaking in a happy and enthusiastic manner, and at a fast pace. The speech characteristics corresponding to this style instruction could include: male voice, young, happy, enthusiastic, and fast pace.

[0069] Another example is that the style instruction described above could be: a middle-aged woman's voice slowly telling a story in a gentle and low tone. The speech features corresponding to this style instruction could include: female voice, middle-aged, gentle, slow speech, and low volume.

[0070] S202. Based on the speech synthesis model under test, determine the speech data corresponding to each style instruction.

[0071] The speech synthesis model under test refers to a novel speech synthesis model that needs to be evaluated. This model can output synthesized speech based on input content instructions and style instructions. Content instructions specify the text content of the speech data to be output. The synthesized speech output by the speech synthesis model under test must not only conform to the text of the content instructions but also meet the requirements of the style instructions, that is, express speech according to the specified speech features.

[0072] In one possible implementation, multiple content instructions are obtained. For each style instruction, the style instruction and any one of the multiple content instructions are input into the speech synthesis model under test to obtain the speech data corresponding to the style instruction and the selected content instruction. In this way, the speech data corresponding to each style instruction can be obtained.

[0073] For example, if the content instruction is "Hello everyone, welcome to the speech synthesis service" and the style instruction is "A young male voice is speaking in a happy and enthusiastic mood, and at a fast pace", then the speech data output by the speech synthesis model under test will be a speech of a young male voice, the content of which is "Hello everyone, welcome to the speech synthesis service", and this speech is presented in a happy and enthusiastic mood, and at a relatively fast pace.

[0074] S203. Evaluate the speech synthesis model under test based on the matching degree between multiple style instructions and their corresponding speech data.

[0075] The matching degree between the style instruction and the corresponding speech data refers to the degree to which the speech features of the speech data match the speech features indicated by the style instruction.

[0076] In one possible implementation, for each style instruction, the semantic features of the style instruction and the audio features of the corresponding speech data are obtained, and the similarity between the semantic features of the style instruction and the audio features of the speech data is used as the matching degree between the style instruction and the speech data.

[0077] Furthermore, the matching degree between each style instruction and its corresponding speech data is used as the unit compliance degree. Unit compliance degree is an indicator used to evaluate the speech synthesis model under test based on a single data point (i.e., a pair of style instructions and speech data). The system compliance degree of the speech synthesis model under test is determined based on the average and standard deviation of all unit compliance degrees. System compliance degree is an indicator used to evaluate the speech synthesis model under test based on multiple data points. A higher system compliance degree indicates better performance of the speech synthesis model under test. For more details, please refer to the following... Figure 9 The embodiments described herein will not be elaborated upon here.

[0078] The speech synthesis model evaluation method provided in this application obtains multiple style instructions, determines the speech data corresponding to each style instruction based on the speech synthesis model under test, and evaluates the speech synthesis model under test based on the matching degree between the multiple style instructions and their respective corresponding speech data. In other words, the embodiments of this application evaluate the speech synthesis model based on the matching degree between the style instructions input to the speech synthesis model and the speech data output by the speech synthesis model, thereby enabling the evaluation of novel speech synthesis models.

[0079] Furthermore, prior to S203, it is necessary to determine the matching degree between each style instruction and its corresponding speech data so that the speech synthesis model under test can be evaluated based on the matching degree between multiple style instructions and their respective speech data. Therefore, as follows... Figure 3 As shown in the embodiments of this application, the speech synthesis model evaluation method may further include the following steps.

[0080] S301. For each style instruction, obtain the semantic features of the style instruction and the audio features of the speech data corresponding to the style instruction.

[0081] In one possible implementation, style instructions and speech data are input into the evaluation model to obtain the style instruction embedding vector E. t (i.e., the semantic features of the style instruction) and the embedding vector E of the speech data a (i.e., the audio characteristics of the speech data).

[0082] It should be noted that the above evaluation model was trained using a contrastive learning approach on the target speech dataset and the target style text dataset. For details, please refer to the following... Figure 5 The embodiments described herein will not be elaborated upon here.

[0083] Another possible implementation involves using a text recognition tool to process the style instructions and obtain their semantic features. Then, an audio recognition tool is used to process the speech data and obtain its audio features.

[0084] S302. Based on the semantic features of style instructions and the audio features of speech data, determine the matching degree between style instructions and speech data.

[0085] In one possible implementation, E in S301 above is combined with... t and E a For example, E a and E t The similarity is used as the degree of matching between style instructions and speech data.

[0086] For example, E a and E t Cosine similarity is used as the matching degree between style instructions and speech data, E a and E t The cosine similarity can be obtained using the following formula 1.

[0087]

[0088] Among them, cos(E) a E t ) represents E a and E t cosine similarity, E a E represents the embedding vector of the speech data. t An embedding vector representing a style directive.

[0089] Of course, E can also be used. a and Et Other similarities, such as the matching degree between style instructions and speech data, are not limited in this application embodiment.

[0090] For example, Figure 4 This is a flowchart illustrating the process of determining the matching degree between style instructions and speech data, as provided in an embodiment of this application. Figure 4 As shown, the speech data is sequentially input into an audio encoder and a multilayer perceptron (MLP) to obtain the speech data embedding vector E. a The style instructions are sequentially input into the text encoder and the multilayer perceptron to obtain the embedding vector E of the style instructions. t Calculate E a and E t The cosine similarity is used to obtain the matching degree between style instructions and speech data.

[0091] In one embodiment, such as Figure 5 As shown, the evaluation model in S301 above can be determined by the following S501 to S503.

[0092] S501. Obtain the target speech dataset.

[0093] One possible implementation involves obtaining the target speech dataset from an open-source speech dataset, such as the multi-modal multi-scene multi-label emotional dialogue database (M3ED).

[0094] Another possible implementation involves using a speech synthesis model to generate a large amount of synthesized speech as the target speech dataset.

[0095] Another possible implementation involves using web crawlers to obtain the target speech dataset from the internet.

[0096] It should be noted that, to improve the accuracy of evaluating speech synthesis models, the distribution of speech data in the target speech dataset should conform to real-world usage scenarios. Specifically, it is necessary to acquire as much diverse speech data as possible. For example, based on gender classification, male and female voices should be acquired. Based on age classification, voices from the elderly, young adults, and children should be acquired. Based on emotion classification, voices expressing various emotions such as happiness, surprise, sadness, disgust / disgust, anger / fear, neutrality, and confusion should be acquired. Based on intonation classification, voices with various intonation styles such as steady, normal, high-pitched, low, soothing, whispering, gentle, playful, serious, loud, cheerful, tense, and confident should be collected. Based on speech rate classification, voices with different speech rates such as fast, medium, and slow should be collected. Based on pitch classification, voices with different pitches such as low, medium, and high should be acquired.

[0097] Understandably, acquiring as many diverse speech data as possible as possible as the target speech dataset makes the speech data in the target speech dataset richer, more comprehensive and representative, thereby improving the accuracy of the evaluation model trained using the target speech dataset.

[0098] S502. Determine the target style text set corresponding to the target speech dataset.

[0099] In one possible implementation, for each target speech data in the target speech dataset, multiple discrete labels are obtained based on multiple speech classification models, and a large language model (LLM) is used to transform the multiple discrete labels into target style text describing natural language, thereby obtaining a target style text set.

[0100] It should be noted that the aforementioned speech classification models include gender classification models, age classification models, emotion classification models, intonation classification models, speech rate classification models, and tone classification models. These different speech classification models collect discrete labels corresponding to the target speech data. Specifically, the collection process for each discrete label is explained in detail below.

[0101] 1-1. Gender tag collection

[0102] Using pre-trained speech models such as wav2vec2-large or wavLM as the base model, the base model is fine-tuned using an open-source speech dataset with gender labels (e.g., the emotional voices database, EmoV_DB) to obtain a gender classification model. The target speech data is then input into this gender classification model to obtain the corresponding gender labels.

[0103] 1-2. Age tag collection

[0104] Age labels were obtained from a portion of the speech data using manual annotation. This age-labeled speech data was then used as training data to fine-tune models such as wav2vec2-large or wavLM, resulting in an age classification model. The target speech data was then input into this age classification model to obtain the corresponding age labels.

[0105] 1-3. Collection of Emotional Tags

[0106] The target speech data is input into an open-source emotion classification model, such as the emotion2vec model, to obtain the corresponding emotion label.

[0107] 1-4. Collection of intonation tags

[0108] Pitch labels were obtained from a portion of the speech data using manual annotation. This pitch-labeled speech data was then used as training data to fine-tune models such as wav2vec2-large or wavLM, resulting in a pitch classification model. The target speech data was then input into this pitch classification model to obtain the corresponding pitch labels.

[0109] 1-5. Collection of speech rate tags

[0110] Using an automatic speech recognition (ASR) model, such as the Whisper model, speech recognition and localization are performed on the target speech data, allowing the calculation of the average number of words spoken per second, thus obtaining the speech rate. Speech rate information for the entire target speech dataset is then statistically analyzed, and based on the distribution of the statistical results, the speech rates are divided into five levels, resulting in corresponding speech rate labels.

[0111] 1-6. Collection of pitch tags

[0112] Use audio feature extraction tools (e.g., WORLD or STRAIGHT) to extract the fundamental frequency (f0) features of the target speech data. Calculate the f0 values ​​for the entire target speech dataset. Based on the distribution of the statistical results, divide the pitch into multiple levels to obtain corresponding pitch labels. For example, based on the statistical results, divide the pitch into three equal levels as "low pitch," "medium pitch," and "high pitch" pitch labels.

[0113] It should be noted that the training data obtained through manual annotation was used in sections 1-2 and 1-4 above. When manually annotating speech data, annotators need to be trained, provided with example speech for each label category and key annotation points, so that they understand the purpose and standards of annotation. Furthermore, the annotated data should be delivered in stages for random checks, allowing for timely retraining of annotators after identifying annotation problems, thereby improving the accuracy of manual annotation.

[0114] Optionally, manually annotated datasets are typically divided into training, testing, and validation sets. The training set can be annotated by one person, while the testing and validation sets can be annotated by three people to further improve the accuracy of the annotation.

[0115] Optionally, following the label collection methods from 1-1 to 1-6, other types of labels can also be collected, such as volume labels. Specifically, if a high-accuracy open-source classification model exists, it can be used directly for label collection. If no high-accuracy open-source classification model exists, relevant open-source speech data or manually labeled training data can be used. Based on the training data, models such as wav2vec2-large or wavLM can be fine-tuned to obtain a speech classification model, which can then be used to collect labels corresponding to the target speech data.

[0116] Furthermore, after obtaining multiple discrete labels, since each label is separate and not a complete natural language description, it is necessary to convert the discrete labels into target-style text that is a natural language description. Specifically, large language models such as Tongyi Qianwen or Llama can be used to convert the discrete labels into target-style text that is a natural language description. The specific process includes steps one and two.

[0117] Step 1: Convert multiple discrete labels into text describing them in natural language.

[0118] In one possible implementation, multiple discrete labels and a first prompt are input into a large language model to obtain a natural language description of the text.

[0119] In one example, the first prompt could be: Please convert the following words into a sentence. The discrete labels could be: male voice, youth, fast speech, happy, and high-spirited. Using this example, the input to the large language model would be: Please convert the following words into a sentence: [male voice, youth, fast speech, happy, high-spirited]. The output of the large language model could be: A young male voice is speaking with a happy and high-spirited mood, and at a very fast pace.

[0120] Furthermore, to obtain diverse target style texts, discrete tags can be adjusted. For example, an elderly male voice can be adjusted to sound like an old man.

[0121] Step 2: Rewrite the text described in natural language to obtain the target style text.

[0122] In one possible implementation, the text described in natural language and the second prompt are input into a large language model to obtain the target style text.

[0123] In one example, the second prompt could be: Please rewrite the following sentence as a synonym. The natural language description could be: A young male voice is speaking with a happy and enthusiastic mood, and at a fast pace. Using this example, the input to the large language model would be: Please rewrite the following sentence as a synonym: [A young male voice is speaking with a happy and enthusiastic mood, and at a fast pace]. The output of the large language model could be: A young man is speaking with a joyful and enthusiastic mood, and at a very fast pace.

[0124] Understandably, using large speech models to rewrite texts described in natural language can yield more diverse target style texts.

[0125] For example, Figure 6 This is a flowchart illustrating how to determine target style text, as provided in an embodiment of this application. Figure 6 As shown, the target speech data is input into gender classification models, age classification models, emotion classification models, intonation classification models, speech rate classification models, and tone classification models to obtain multiple discrete labels. These discrete labels and the first prompt are then input into a large language model to obtain a natural language description text. Finally, this natural language description text and the second prompt are input into the large language model to obtain the target style text.

[0126] S503. Based on the target speech dataset and the target style text set, determine the evaluation model.

[0127] In one possible implementation, the target speech dataset and target style text dataset obtained in S502 above are used to train a model in a contrastive learning paradigm to obtain an evaluation model.

[0128] Optionally, when training the model using the aforementioned target speech dataset and target style text set, the training data is typically divided into a training set, a test set, and a validation set. To improve the accuracy of the constructed evaluation model, the target style text in the test and validation sets can be manually corrected, thereby accurately verifying the performance of the evaluation model.

[0129] For example, Figure 7 This is a schematic diagram illustrating a model training method provided in an embodiment of this application. Figure 7As shown, a dual-tower model is used. The target speech data is sequentially input into an audio encoder and a multilayer perceptron to obtain the embedding vector of the target speech data. The target style text is sequentially input into a text encoder and a multilayer perceptron to obtain the embedding vector of the target style text. The model is optimized using a contrastive learning loss function to obtain the final evaluation model. The audio encoder extracts hidden features from the target speech data. The text encoder extracts hidden features from the target style text. The multilayer perceptron normalizes the extracted hidden features of the target speech data and the target style text to the same feature dimension.

[0130] For example, the contrastive learning loss function mentioned above can refer to the contrastive language image pretraining (CLIP) model, and is expressed as Equation 2 below.

[0131]

[0132] Where L represents the contrastive learning loss function, and N represents the total number of samples. Let represent the embedding vector of the target speech data in the i-th sample pair. Let represent the embedding vector of the target speech data in the j-th sample pair. Let represent the embedding vector of the target style text in the i-th sample pair. Let represent the embedding vector of the target style text in the j-th sample. τ represents the trainable temperature parameter used to scale the loss.

[0133] Understandably, by minimizing the contrastive learning loss function and adjusting the parameters of the audio encoder, text encoder, and multilayer perceptron, the evaluation model can learn how to map speech data and style text into a common embedding space, so that matching speech data and style text have similar vector representations in that space.

[0134] For example, Figure 8 This is a schematic diagram of an audio encoder provided in an embodiment of this application. The audio encoder is a hierarchical token-semantic audio transform (HTS-AT) structure, which is a structure based on a transformer network and combined with a convolutional neural network (CNN). Figure 8As shown, speech data is input into a CNN-based patch embedding layer, then processed through N win transformer blocks, and finally input into a CNN-based token-semantic layer to obtain the hidden features of the speech data. The win transformer block includes a win transformer and a patch merging layer.

[0135] For example, the text encoder in the evaluation model above can be the RoBERTa model. The RoBERTa model is an improved BERT model with a similar network structure to BERT, but with several optimizations in the training method.

[0136] It should be noted that this application does not impose restrictions on the feature dimensions and number of layers, or other hyperparameters, of the transformer architecture used in the aforementioned audio encoder and text encoder.

[0137] In one embodiment, such as Figure 9 As shown, the above S203 can be specifically determined by the following S901 to S902.

[0138] S901. Evaluation parameters are obtained based on the matching degree between multiple style instructions and their corresponding speech data.

[0139] The evaluation parameters are the mean and standard deviation of the matching degree.

[0140] In one possible implementation, the matching degree between N style instructions and their corresponding speech data is used as N single-single compliance degrees, denoted as S_single. i , where i is the number. The average compliance rate of an individual can be represented by the following formula 3, and the standard deviation of the compliance rate of an individual can be represented by the following formula 4.

[0141]

[0142] Where μ represents the average compliance rate of a single monomer. N represents the total number of single monomer compliance rates. S_single i This indicates the compliance of the monomer numbered i.

[0143]

[0144] Where σ represents the standard deviation of single-unit compliance. N represents the total number of single-unit compliances. S_single i This represents the compliance rate of monomer number i. μ represents the average compliance rate of monomers.

[0145] S902. Evaluate the speech synthesis model under test based on the evaluation parameters and preset equalization coefficients.

[0146] The preset balance coefficient is used to adjust the weights of the average and standard deviation of the matching degree.

[0147] In one possible implementation, the system compliance of the speech synthesis model under test is calculated based on the evaluation parameters and preset equalization coefficients, and the system compliance is used as the evaluation result of the speech synthesis model under test.

[0148] Understandably, in order to compare and select different speech synthesis models, or to iterate on speech synthesis models under development, it is necessary not only to evaluate the individual compliance of single speech data and style instructions, but also to evaluate the overall performance of the speech synthesis model. In this way, the system compliance can be used to evaluate the speech synthesis model at the system level.

[0149] For example, the preset balance coefficients mentioned above include a and b. Here, a is used to adjust the weight of the average matching degree, and b is used to adjust the weight of the standard deviation of the matching degree.

[0150] Based on the above examples, the system compliance of the speech synthesis model under test can be represented by the following formula 5.

[0151]

[0152] Where S_system represents the system compliance of the speech synthesis model under test. a represents the weight of the mean. μ represents the mean of individual compliance. b represents the weight of the standard deviation. σ represents the standard deviation of individual compliance.

[0153] It is understandable that since the individual compliance score is calculated based on cosine similarity, and the range of cosine similarity is [-1, 1], the average value μ of the individual compliance score also ranges from [-1, 1]. The constant 1 is added to the denominator to prevent the denominator from being zero in extreme cases.

[0154] It's important to note that changing the values ​​of 'a' and 'b' adjusts the weighting of the mean and standard deviation. This allows the calculated system compliance metric to focus more on the average output performance (mean) or the output stability (standard deviation) of the speech synthesis model under test. For example, when evaluating a speech synthesis model, if the application scenario prioritizes the model's average output performance, increasing the value of 'a' and decreasing the value of 'b' will make the system compliance metric more biased towards models with a larger mean μ. Conversely, if the application scenario prioritizes the model's output stability, increasing the value of 'b' and decreasing the value of 'a' will make the system compliance metric more biased towards models with a smaller standard deviation σ.

[0155] For example, Figure 10This is a flowchart illustrating how to determine system compliance, as provided in an embodiment of this application. Figure 10 As shown, each pair of style instructions and corresponding speech data are input into the evaluation model to obtain multiple individual compliance scores. The mean and standard deviation of the individual compliance scores are calculated, and the values ​​of a and b in the preset equalization coefficients are determined. Then, the system compliance score can be obtained according to Formula 5 above.

[0156] The PromptTTS model will be used as an example below to evaluate the speech synthesis model using the above evaluation method.

[0157] It should be noted that the PromptTTS model is a novel, multi-dimensional, controllable speech synthesis model capable of generating speech with specified gender, speech rate, pitch, volume, and emotion. The PromptTTS model takes two inputs: Style Prompt and Content Prompt, and outputs a segment of human audio. Style Prompt refers to the style instruction mentioned earlier, while Content Prompt refers to the content of the spoken words, i.e., the content instruction mentioned earlier.

[0158] For example, Figure 11 This document provides a flowchart for evaluating a single data point in a PromptTTS model, as an embodiment of this application. Figure 11 As shown, content instructions and style instructions are input into the PromptTTS model to obtain speech data. This speech data and style instructions are then input into the evaluation model to obtain the speech data embedding vector E. a Embedding vector E of style instructions t Calculate E a and E t The cosine similarity is used to obtain the compliance index of speech data and style instructions, namely the individual compliance.

[0159] Optional, Figure 11 The evaluation model can be the previously constructed evaluation model, or a reconstructed evaluation model based on the PromptTTS model. Specifically, a large amount of speech data is synthesized using the PromptTTS model as the target speech dataset. Multiple discrete labels are obtained for each target speech data based on gender classification models, speech rate classification models, pitch classification models, emotion classification models, and volume classification models. The Llama2 model and preset prompts are used to convert these discrete labels into style text, thus obtaining the target style text set. Figure 7 The evaluation model is trained using a comparative learning approach.

[0160] For example, the above preset prompt could be: Please convert the following words into a sentence: [a person, {gender tag}, {speech rate tag}, {tone tag}, {emotion tag}, {volume tag}].

[0161] Furthermore, different content instructions and style instructions are combined and repeated. Figure 11 The process is repeated N times to obtain N individual compliance scores. Based on the relevant descriptions in S901 and S902, the system compliance score of the PromptTTS model can be obtained, and this system compliance score can be used as the evaluation result of the PromptTTS model.

[0162] It is understood that the above-described speech synthesis model evaluation method can be implemented by a speech synthesis model evaluation device. To achieve the above functions, the speech synthesis model evaluation device includes hardware structures and / or software modules corresponding to the execution of each function. Those skilled in the art should readily recognize that, based on the modules and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein, the embodiments disclosed in this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the embodiments disclosed in this application.

[0163] The embodiments disclosed in this application can divide the speech synthesis model evaluation device generated by the above method examples into functional modules. For example, each function can be divided into its own functional module, or two or more functions can be integrated into one processing module. The integrated module can be implemented in hardware or as a software functional module. It should be noted that the module division in the embodiments disclosed in this application is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0164] Figure 12 This is a schematic diagram of a speech synthesis model evaluation device provided in an embodiment of this application. Figure 12 As shown, the speech synthesis model evaluation device 120 can be used to perform... Figure 2 , Figure 3 as well as Figure 9 The speech synthesis model evaluation method shown is illustrated. The speech synthesis model evaluation device 120 includes a communication unit 1201 and a processing unit 1202.

[0165] The communication unit 1201 is used to acquire multiple style instructions, which are used to indicate the speech features of the speech data to be output; the processing unit 1202 is used to determine the speech data corresponding to each style instruction based on the speech synthesis model under test; the processing unit 1202 is also used to evaluate the speech synthesis model under test according to the matching degree between the multiple style instructions and their corresponding speech data.

[0166] In one possible implementation, the speech features include at least one of the following: timbre features; emotion features; speech rate features; and pitch features.

[0167] In one possible implementation, the communication unit 1201 is further configured to acquire multiple content instructions, which are used to indicate the text content of the speech data to be output; for each style instruction, the processing unit 1202 is further configured to input the style instruction and any one of the multiple content instructions into the speech synthesis model to be tested, and obtain the speech data corresponding to the style instruction, so as to obtain the speech data corresponding to each style instruction.

[0168] In one possible implementation, the method further includes: for each style instruction, the communication unit 1201 is further configured to acquire the semantic features of the style instruction and the audio features of the speech data corresponding to the style instruction; the processing unit 1202 is further configured to determine the matching degree between the style instruction and the speech data based on the semantic features of the style instruction and the audio features of the speech data.

[0169] In one possible implementation, the processing unit 1202 is further configured to use the similarity between the semantic features of the style instruction and the audio features of the speech data as the matching degree between the style instruction and the speech data.

[0170] In one possible implementation, the processing unit 1202 is further configured to obtain evaluation parameters based on the matching degree between multiple style instructions and their corresponding speech data, wherein the evaluation parameters are the average value and standard deviation of the matching degree; the processing unit 1202 is further configured to evaluate the speech synthesis model under test based on the evaluation parameters and preset equalization coefficients, wherein the preset equalization coefficients are used to adjust the weights of the average value and standard deviation of the matching degree.

[0171] Through the above description of the embodiments, those skilled in the art will clearly understand that, for the sake of convenience and brevity, only the division of the above functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. The specific working process of the system, device, and unit described above can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0172] This disclosure also provides a computer-readable storage medium storing instructions that, when executed by a processor of an electronic device, enable the electronic device to perform the speech synthesis model evaluation method provided in the embodiments of this disclosure described above.

[0173] This disclosure also provides a computer program product containing instructions that, when run on an electronic device, causes the electronic device to execute the speech synthesis model evaluation method provided in the above-described embodiments of this disclosure.

[0174] The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: electrical connections having one or more wires; portable computer disks; hard disks; random access memory (RAM); read-only memory (ROM); erasable programmable read-only memory (EPROM); registers; hard disks; optical fibers; portable compact disc read-only memory (CD-ROM); optical storage devices; magnetic storage devices; or any suitable combination thereof; or any other form of computer-readable storage medium known in the art. An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Of course, the storage medium may also be a component of the processor. The processor and the storage medium may reside in an application-specific integrated circuit (ASIC). In the embodiments of this application, the computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0175] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions within the technical scope disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of speech synthesis model evaluation, the method comprising: The method comprises: obtaining a plurality of style instructions, the style instructions being used to indicate voice features of voice data required to be output; determining voice data corresponding to each style instruction based on a to-be-tested voice synthesis model; for each style instruction, obtaining semantic features of the style instruction and audio features of voice data corresponding to the style instruction; determining a matching degree between the style instruction and the voice data based on the semantic features of the style instruction and the audio features of the voice data; evaluating the to-be-tested voice synthesis model according to the matching degrees between the plurality of style instructions and the respective corresponding voice data.

2. The method of claim 1, wherein, The voice features comprise at least one of: timbre features; emotion features; speech rate features; pitch features.

3. The method of claim 1, wherein, The determination of the voice data corresponding to each style instruction based on the to-be-tested voice synthesis model comprises: obtaining a plurality of content instructions, the content instructions being used to indicate text content of voice data required to be output; for each style instruction, inputting the style instruction and any content instruction of the plurality of content instructions into the to-be-tested voice synthesis model to obtain voice data corresponding to the style instruction, so as to obtain voice data corresponding to each style instruction.

4. The method of claim 1, wherein, The determination of the matching degree between the style instruction and the voice data based on the semantic features of the style instruction and the audio features of the voice data comprises: taking a similarity between the semantic features of the style instruction and the audio features of the voice data as the matching degree between the style instruction and the voice data.

5. The method of claim 1, wherein, The evaluation of the to-be-tested voice synthesis model according to the matching degrees between the plurality of style instructions and the respective corresponding voice data comprises: obtaining evaluation parameters according to the matching degrees between the plurality of style instructions and the respective corresponding voice data, the evaluation parameters being an average value and a standard deviation of the matching degrees; evaluating the to-be-tested voice synthesis model according to the evaluation parameters and a preset balancing coefficient, the preset balancing coefficient being used to adjust weights of the average value and the standard deviation of the matching degrees.

6. A speech synthesis model evaluation apparatus characterized by comprising: The device comprises a communication unit and a processing unit. The communication unit is configured to obtain a plurality of style instructions, the style instructions being used to indicate voice features of voice data required to be output. For each style instruction, the communication unit is further configured to obtain semantic features of the style instruction and audio features of voice data corresponding to the style instruction. The processing unit is configured to determine a matching degree between the style instruction and the voice data based on the semantic features of the style instruction and the audio features of the voice data. The processing unit is further configured to determine voice data corresponding to each style instruction based on a to-be-tested voice synthesis model. The processing unit is further configured to evaluate the to-be-tested voice synthesis model according to the matching degrees between the plurality of style instructions and the respective corresponding voice data.

7. A speech synthesis model evaluation apparatus characterized by comprising: The device comprises: a memory and a processor; the memory and the processor are coupled; the memory is configured to store instructions executable by the processor; the processor executes the instructions to perform the voice synthesis model evaluation method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer instructions, and when the computer instructions run on a computer, the computer executes the speech synthesis model evaluation method in any one of claims 1-5.

9. A computer program product, characterised in that, The computer program product comprises computer program instructions, and when the computer program instructions are executed by a processor, the speech synthesis model evaluation method in any one of claims 1-5 is implemented.

Citation Information

Patent Citations

  • Synthetic speech evaluation method, device and equipment

    CN113223559A

  • Speech synthesis system evaluation method and device, readable storage medium and terminal equipment

    CN113450768A