Model evaluation method, device and electronic equipment
By clustering the voiceprint features of the speech synthesis model and performing cosine distance statistics, the problem of low evaluation efficiency of personalized speech synthesis models is solved, and a fast and accurate evaluation of audio signal fidelity is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING BAIDU NETCOM SCI & TECH CO LTD
- Filing Date
- 2020-05-21
- Publication Date
- 2026-07-21
AI Technical Summary
The evaluation efficiency of existing personalized speech synthesis models is low, and they cannot efficiently assess the fidelity of audio signals.
By clustering M first voiceprint features to obtain K first central features, and clustering N second voiceprint features to obtain J second central features, the cosine distance between the K first central features and the J second central features is calculated to obtain the first distance, thereby evaluating the fidelity of the speech synthesis model.
It improves the evaluation efficiency of personalized speech synthesis models, enables rapid assessment of the fidelity of large batches of audio signals, reduces model evaluation costs, and improves evaluation accuracy.
Smart Images

Figure CN117476038B_ABST
Abstract
Description
[0001] This invention is a divisional application of the invention application filed on May 21, 2020, with application number 202010437127.5 and title "Model Evaluation Method, Apparatus and Electronic Equipment". Technical Field
[0002] This application relates to data processing technology, particularly to the field of audio data processing technology, specifically to a model evaluation method, apparatus, and electronic device. Background Technology
[0003] Speech synthesis technology is a technique that converts text into audio signals for output. It plays an important role in the field of human-computer interaction and has a wide range of applications. Personalized speech synthesis, on the other hand, uses speech synthesis technology to create audio signals that closely resemble human speech and is currently widely used in areas such as maps and smart speakers.
[0004] Currently, there are many personalized speech synthesis models used to synthesize audio signals. However, the audio reproduction accuracy of these personalized speech synthesis models varies greatly. Therefore, it is crucial to evaluate these personalized speech synthesis models.
[0005] Currently, the evaluation of personalized speech synthesis models typically relies on pre-trained voiceprint verification models to assess the audio fidelity of the synthesized audio, i.e., the similarity between the synthesized audio and the human voice, thereby evaluating the quality of the personalized speech synthesis model. However, since voiceprint verification models usually verify the fidelity of each synthesized audio signal individually, the evaluation efficiency is relatively low. Summary of the Invention
[0006] This application provides a model evaluation method, apparatus, and electronic device.
[0007] According to the first aspect, this application provides a model evaluation method, the method comprising:
[0008] Obtain M first audio signals synthesized using the first speech synthesis model to be evaluated, and obtain N recorded second audio signals;
[0009] Voiceprint extraction is performed on each of the M first audio signals to obtain M first voiceprint features; voiceprint extraction is performed on each of the N second audio signals to obtain N second voiceprint features.
[0010] The M first voiceprint features are clustered to obtain K first central features; the N second voiceprint features are clustered to obtain J second central features.
[0011] The first distance is obtained by calculating the cosine distance between the K first central features and the J second central features;
[0012] Based on the first distance, the first speech synthesis model to be evaluated is evaluated;
[0013] Where M, N, K and J are all positive integers greater than 1, M is greater than K and N is greater than J.
[0014] According to the second aspect, this application provides a model evaluation apparatus, comprising:
[0015] The first acquisition module is used to acquire M first audio signals synthesized using the first speech synthesis model to be evaluated, and to acquire N recorded second audio signals;
[0016] The first voiceprint extraction module is used to extract voiceprints from each of the M first audio signals to obtain M first voiceprint features; and to extract voiceprints from each of the N second audio signals to obtain N second voiceprint features.
[0017] The first clustering module is used to cluster the M first voiceprint features to obtain K first central features; and to cluster the N second voiceprint features to obtain J second central features.
[0018] The first statistical module is used to calculate the cosine distance between the K first central features and the J second central features to obtain the first distance;
[0019] The first evaluation module is used to evaluate the first speech synthesis model to be evaluated based on the first distance;
[0020] Where M, N, K and J are all positive integers greater than 1, M is greater than K and N is greater than J.
[0021] According to a third aspect, this application provides an electronic device, comprising:
[0022] At least one processor; and
[0023] A memory communicatively connected to the at least one processor; wherein,
[0024] The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform any of the methods in the first aspect.
[0025] According to a fourth aspect, this application provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to perform any of the methods in the first aspect.
[0026] According to a fifth aspect, this application provides a computer program product including a computer program that, when executed by a processor, implements any of the methods in the first aspect.
[0027] According to the technology of this application, by clustering M first voiceprint features to obtain K first central features, and clustering N second voiceprint features to obtain J second central features; and by calculating the cosine distance between the K first central features and the J second central features to obtain a first distance, it is possible to evaluate the overall fidelity of M first audio signals synthesized using the first speech synthesis model to be evaluated based on the first distance, thereby improving the evaluation efficiency of the first speech synthesis model to be evaluated. This application solves the problem of low efficiency in evaluating personalized speech synthesis models in the prior art.
[0028] It should be understood that the description in this section is not intended to identify key or essential features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will become readily apparent from the following description. Attached Figure Description
[0029] The accompanying drawings are provided for a better understanding of this solution and do not constitute a limitation of this application. Wherein:
[0030] Figure 1 This is a schematic flowchart of the model evaluation method according to the first embodiment of this application;
[0031] Figure 2 This is a flowchart illustrating the evaluation process for the second speech synthesis model to be evaluated.
[0032] Figure 3 This is one of the structural schematic diagrams of the model evaluation device according to the second embodiment of this application;
[0033] Figure 4 This is a second schematic diagram of the model evaluation device according to the second embodiment of this application;
[0034] Figure 5 This is a block diagram of an electronic device used to implement the model evaluation method of the embodiments of this application. Detailed Implementation
[0035] The following description, in conjunction with the accompanying drawings, illustrates exemplary embodiments of this application, including various details to aid understanding. These should be considered merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this application. Similarly, for clarity and brevity, descriptions of well-known functions and structures are omitted in the following description.
[0036] First Embodiment
[0037] like Figure 1 As shown, this application provides a model evaluation method, including the following steps:
[0038] Step S101: Obtain M first audio signals synthesized using the first speech synthesis model to be evaluated, and obtain N recorded second audio signals.
[0039] In this embodiment, the first speech synthesis model to be evaluated is a personalized speech synthesis model, the purpose of which is to synthesize an audio signal similar to the pronunciation of a real person through the first speech synthesis model to be evaluated, so as to be applied to maps, smart speakers and other fields.
[0040] The first speech synthesis model to be evaluated can be generated by pre-training a first preset model. The first preset model is essentially a model constructed by a first algorithm. The parameter data in the first preset model needs to be trained to obtain the first speech synthesis model to be evaluated.
[0041] Specifically, multiple audio signals recorded by the first user based on a text are used as training samples. For example, 20 or 30 audio signals recorded by the first user based on a text are used as training samples and input into the first preset model to train and obtain parameter data in the first preset model, so as to generate the first speech synthesis model to be evaluated for the first user.
[0042] After generating the first speech synthesis model to be evaluated, a batch of text is used, and the first speech synthesis model to be evaluated by the first user is used to generate a batch of first audio signals. Specifically, each text is input into the first speech synthesis model to be evaluated to output the first audio signal corresponding to that text, ultimately obtaining M first audio signals. At the same time, a batch of second audio signals recorded by the first user is obtained, ultimately obtaining N second audio signals.
[0043] M and N can be the same or different; no specific restrictions are imposed here. To ensure more accurate evaluation results for the first speech synthesis model to be evaluated, M and N are usually relatively large, such as 20 or 30.
[0044] Step S102: Extract voiceprints from each of the M first audio signals to obtain M first voiceprint features; extract voiceprints from each of the N second audio signals to obtain N second voiceprint features.
[0045] There are various ways to extract voiceprints from the first audio signal. For example, traditional statistical methods can be used to extract voiceprints from the first audio signal to obtain statistical features of the first audio signal, which are the first voiceprint features. Alternatively, deep neural networks (DNNs) can be used to extract voiceprints from the first audio signal to obtain DNN voiceprint features of the first audio signal, which are the first voiceprint features.
[0046] Meanwhile, the method for extracting voiceprints from the second audio signal is similar to that for extracting voiceprints from the first audio signal, and will not be elaborated here.
[0047] Step S103: Cluster the M first voiceprint features to obtain K first central features; cluster the N second voiceprint features to obtain J second central features.
[0048] Traditional or novel clustering algorithms can be used to cluster the M first voiceprint features to obtain K first central features. K can be calculated using a clustering algorithm based on the actual cosine distances between any two of the M first voiceprint features.
[0049] For example, a clustering algorithm can group these M first voiceprint features into three, four, five, or even more clusters based on the cosine distance between any two first voiceprint features in each cluster, where K is the number of clusters. Within each cluster, the cosine distance between any two first voiceprint features (the intra-group distance) is less than a preset threshold, while the cosine distance between first voiceprint features in different clusters (the inter-group distance) is greater than another preset threshold.
[0050] After clustering, the first central feature of each class is calculated based on the first voiceprint feature of each class. For example, the first central feature of a class can be the voiceprint feature after averaging multiple first voiceprint features of the class, and finally K first central features are obtained.
[0051] Meanwhile, the method for clustering the N second voiceprint features is similar to the method for clustering the M first voiceprint features, and will not be described in detail here.
[0052] K and J can be the same or different; no specific restrictions are imposed here. Additionally, M, N, K, and J are all positive integers greater than 1, with M greater than K and N greater than J.
[0053] Step S104: Calculate the cosine distance between the K first central features and the J second central features to obtain the first distance.
[0054] For each first central feature, the cosine distance between that first central feature and each of the J second central features can be calculated to obtain the cosine distance corresponding to that first central feature. The cosine distance between two central features can characterize the similarity between the two central features.
[0055] For example, the K first central features are first central feature A1, first central feature A2 and first central feature A3, and the J second central features are second central feature B1, second central feature B2 and second central feature B3. The cosine distances between the first central feature A1 and the second central features B1, B2, and B3 can be calculated to obtain the cosine distances A1B1, A1B2, and A1B3 corresponding to the first central feature A1. Similarly, the cosine distances between the first central feature A2 and the second central features B1, B2, and B3 can be calculated to obtain the cosine distances A2B1, A2B2, and A2B3 corresponding to the first central feature A2. Finally, the cosine distances between the first central feature A3 and the second central features B1, B2, and B3 can be calculated to obtain the cosine distances A3B1, A3B2, and A3B3 corresponding to the first central feature A3. This process ultimately yields multiple cosine distances between the K first central features and the J second central features.
[0056] Then, multiple cosine distances between the K first central features and the J second central features are calculated to obtain a first distance. The calculation of these multiple cosine distances between the K first central features and the J second central features can be done in several ways. For example, the first distance can be obtained by summing these cosine distances, or by averaging these cosine distances.
[0057] Furthermore, since the K first central features are clustered based on the M first voiceprint features, and the J second central features are clustered based on the N second voiceprint features, and since the first distance is statistically obtained based on multiple cosine distances between the K first central features and the J second central features, the first distance can be used to evaluate the similarity between the M first voiceprint features and the N second voiceprint features as a whole.
[0058] In other words, the first distance can be used to evaluate the overall similarity between the M first audio signals and the N second audio signals recorded by a real person, that is, to evaluate the fidelity of the M first audio signals synthesized using the first speech synthesis model to be evaluated. If the first distance is less than a first preset threshold, it indicates that the fidelity of the M first audio signals is good; if the first distance is greater than or equal to the first preset threshold, it indicates that the fidelity of the M first audio signals is poor.
[0059] Step S105: Evaluate the first speech synthesis model to be evaluated based on the first distance.
[0060] Since these M first audio signals are synthesized using the first speech synthesis model to be evaluated, the first distance can be used to evaluate the first speech synthesis model to be evaluated, that is, to evaluate the first speech synthesis model to be evaluated based on the first distance.
[0061] In this embodiment, K first central features are obtained by clustering M first voiceprint features, and J second central features are obtained by clustering N second voiceprint features. The cosine distance between the K first central features and the J second central features is calculated to obtain a first distance. Based on the first distance, the fidelity of M first audio signals synthesized using the first speech synthesis model under test can be evaluated as a whole. This allows for rapid evaluation of the fidelity of a large number of first audio signals, improving the evaluation efficiency of the first speech synthesis model under test.
[0062] Furthermore, compared to existing technologies, this embodiment does not require a voiceprint verification model for model evaluation, thus avoiding the drawback of needing to periodically update the voiceprint verification model and reducing the cost of model evaluation. Simultaneously, during the model evaluation process, multiple first voiceprint features and multiple second voiceprint features are clustered to obtain multiple first central features and multiple second central features, thereby fully considering the personalized characteristics of the audio signal and improving the accuracy of model evaluation.
[0063] Furthermore, since the first speech synthesis model to be evaluated is generated by pre-training the first preset model, and the first preset model is essentially a model constructed by a set of algorithms, this embodiment can also generate first speech synthesis models to be evaluated for multiple users through the first preset model, and evaluate the first preset model by evaluating these users' first speech synthesis models to be evaluated, that is, evaluate the algorithm that constructs the first preset model. Therefore, this embodiment can also improve the evaluation efficiency of personalized speech synthesis algorithms.
[0064] For example, a personalized speech synthesis algorithm is used to construct a first preset model. This first preset model is then used to generate first speech synthesis models for multiple users to be evaluated. These first speech synthesis models for each user are then evaluated. Based on the evaluation results of these first speech synthesis models for multiple users, the first preset model is evaluated. If the first speech synthesis models for most or all users are successfully evaluated, the first preset model is considered successfully evaluated, meaning the personalized speech synthesis algorithm used to construct the first preset model has been successfully evaluated.
[0065] Optionally, the step of calculating the cosine distance between the K first central features and the J second central features to obtain the first distance includes:
[0066] For each first central feature, calculate the cosine distance between the first central feature and each second central feature to obtain J cosine distances corresponding to the first central feature; and sum the J cosine distances corresponding to the first central feature to obtain the sum of cosine distances corresponding to the first central feature.
[0067] The first distance is obtained by summing the cosine distances corresponding to the K first central features.
[0068] In this embodiment, multiple cosine distances between the K first central features and the J second central features are calculated, and these multiple cosine distances are summed to obtain a first distance, which is the total distance between the K first central features and the J second central features. This total distance can characterize the similarity between the M first voiceprint features and the N second voiceprint features as a whole. Therefore, in this embodiment, the similarity between the M first audio signals and the N second audio signals recorded by a real person can be evaluated as a whole based on this total distance, that is, the fidelity of the M first audio signals can be evaluated. This allows for rapid evaluation of the fidelity of a large number of first audio signals, thereby improving the evaluation efficiency of the first speech synthesis model to be evaluated.
[0069] Optionally, evaluating the first speech synthesis model to be evaluated based on the first distance includes:
[0070] If the first distance is less than the first preset threshold, the first speech synthesis model to be evaluated is determined to have been successfully evaluated.
[0071] If the first distance is greater than or equal to the first preset threshold, it is determined that the evaluation of the first speech synthesis model to be evaluated is unsuccessful.
[0072] In this embodiment, when the first distance is less than the first preset threshold, it can be determined that the overall fidelity of the M first audio signals is good, thus confirming that the first speech synthesis model used to synthesize the M first audio signals has been successfully evaluated. When the first distance is greater than or equal to the first preset threshold, it can be determined that the overall fidelity of the M first audio signals is insufficient, thus confirming that the first speech synthesis model used to synthesize the M first audio signals has failed the evaluation and needs improvement.
[0073] The first preset threshold can be set according to the actual situation. In fields where high fidelity of synthesized audio is required, the first preset threshold can be set relatively small.
[0074] Optionally, after acquiring M first audio signals synthesized using the first speech synthesis model to be evaluated, and N recorded second audio signals, the method further includes:
[0075] Obtain T third audio signals synthesized using the second speech synthesis model to be evaluated;
[0076] Voiceprint extraction is performed on each of the T third audio signals to obtain T third voiceprint features;
[0077] Cluster the T third voiceprint features to obtain P third central features;
[0078] The second distance is obtained by calculating the cosine distance between the P third central features and the J second central features;
[0079] Based on the first distance and the second distance, the first speech synthesis model to be evaluated or the second speech synthesis model to be evaluated is evaluated.
[0080] Where T and P are positive integers greater than 1, and T is greater than P.
[0081] In this embodiment, the second speech synthesis model to be evaluated is the speech synthesis model of the first user. The second speech synthesis model to be evaluated is also a personalized speech synthesis model. The purpose is to synthesize an audio signal similar to the pronunciation of a real person through the second speech synthesis model to be evaluated, so as to apply it to the fields of maps, smart speakers and so on.
[0082] The second speech synthesis model to be evaluated can be generated by pre-training a second preset model. This second preset model is essentially a model constructed using a second algorithm. The parameter data in the second preset model needs to be obtained through training to arrive at the second speech synthesis model to be evaluated. The second algorithm can be an upgraded version of the first algorithm, or it can be a competing algorithm of the same type as the first algorithm.
[0083] Specifically, multiple audio signals recorded by the first user based on a text are used as training samples. For example, 20 or 30 audio signals recorded by the first user based on a text are used as training samples and input into the second preset model to train and obtain the parameter data in the second preset model, so as to generate the second speech synthesis model to be evaluated for the first user.
[0084] After generating the second speech synthesis model to be evaluated, a batch of text is used, and the second speech synthesis model to be evaluated by the first user is used to generate a batch of third audio signals. Specifically, each text is input into the second speech synthesis model to be evaluated, so as to output the third audio signal corresponding to the text, and finally T third audio signals are obtained.
[0085] Here, M and T can be the same or different; no specific restrictions are imposed here. To make the evaluation results of the second speech synthesis model to be evaluated more accurate, T is usually relatively large, such as 20 or 30.
[0086] In this embodiment, the method for extracting voiceprints from the third audio signal is similar to the method for extracting voiceprints from the first audio signal. The method for clustering the T third voiceprint features is similar to the method for clustering the M first voiceprint features. The method for calculating the cosine distance between the P third central features and the J second central features is similar to the method for calculating the cosine distance between the K first central features and the J second central features. These methods will not be elaborated further here.
[0087] After calculating the cosine distance between the P third central features and the J second central features to obtain the second distance, the first speech synthesis model to be evaluated or the second speech synthesis model to be evaluated can be evaluated based on the first distance and the second distance.
[0088] Specifically, when the second algorithm is an upgraded version of the first algorithm, it is usually necessary to evaluate the second speech synthesis model to be evaluated. See [link / reference] Figure 2 , Figure 2 This is a flowchart illustrating the evaluation process for the second speech synthesis model to be evaluated, such as... Figure 2 As shown, voiceprint extraction is performed on N second audio signals recorded by the user, M first audio signals synthesized by the first speech synthesis model to be evaluated (i.e., the model currently in use online), and T third audio signals synthesized by the second speech synthesis model to be evaluated (i.e., the model upgraded in this case), to obtain M first voiceprint features, N second voiceprint features, and T third voiceprint features.
[0089] Then, clustering is performed on these three voiceprint features to obtain K first central features, J second central features, and P third central features.
[0090] Next, the cosine distance between the K first central features and the J second central features is calculated to obtain the first distance. At the same time, the cosine distance between the P third central features and the J second central features is calculated to obtain the second distance.
[0091] Finally, the first distance and the second distance are compared. If the second distance is less than the first distance, the fidelity of the T third audio signals synthesized using the second speech synthesis model is determined to be better than the fidelity of the M first audio signals synthesized using the first speech synthesis model. Therefore, the evaluation of the second speech synthesis model is considered successful. Otherwise, the evaluation of the second speech synthesis model is considered unsuccessful, and the second algorithm needs to be upgraded and improved again.
[0092] When the second algorithm is a competing algorithm with the first algorithm, it is usually necessary to evaluate the first speech synthesis model to be evaluated. This is done by comparing the magnitudes of a first distance and a second distance. If the second distance is greater than the first distance, it is determined that the fidelity of the T third audio signals synthesized using the second speech synthesis model is worse than the fidelity of the M first audio signals synthesized using the first speech synthesis model. Therefore, the evaluation of the first speech synthesis model is considered successful. Otherwise, the evaluation of the first speech synthesis model is considered unsuccessful, and the first algorithm needs to be upgraded and improved.
[0093] In this embodiment, by clustering T third voiceprint features, P third central features are obtained; and the cosine distance between the P third central features and J second central features is calculated to obtain a second distance. This allows for a comprehensive evaluation of the fidelity of T third audio signals synthesized using the second speech synthesis model under evaluation, enabling rapid evaluation of the fidelity of a large number of third audio signals and improving the evaluation efficiency of the second speech synthesis model. Simultaneously, by comparing the magnitudes of the first and second distances, the fidelity of the T third audio signals synthesized using the second speech synthesis model under evaluation can be compared with the fidelity of M first audio signals synthesized using the first speech synthesis model under evaluation. This allows for comparison of different personalized speech synthesis algorithms, enabling the evaluation of personalized speech synthesis algorithms and improving algorithm evaluation efficiency.
[0094] Optionally, the cosine distance between any two of the K first central features is greater than a second preset threshold; the cosine distance between any two of the J second central features is greater than a third preset threshold.
[0095] In this embodiment, by setting the cosine distance between any two of the K first central features to be greater than a second preset threshold, and setting the cosine distance between any two of the J second central features to be greater than a third preset threshold, the personalized characteristics of the audio signal are fully considered, thereby improving the accuracy of model evaluation.
[0096] The second and third preset thresholds can be set according to the actual situation. In order to fully consider the personalized characteristics of audio signals and ensure the accuracy of model evaluation, the second and third preset thresholds are usually set as large as possible, that is, the greater the distance between groups, the better.
[0097] It should be noted that the various optional implementation methods in the model evaluation method of this application can be implemented in combination with each other or individually, and this application does not limit this.
[0098] Second Embodiment
[0099] like Figure 3 As shown, this application provides a model evaluation device 300, comprising:
[0100] The first acquisition module 301 is used to acquire M first audio signals synthesized using the first speech synthesis model to be evaluated, and to acquire N recorded second audio signals.
[0101] The first voiceprint extraction module 302 is used to extract voiceprints from each of the M first audio signals to obtain M first voiceprint features; and to extract voiceprints from each of the N second audio signals to obtain N second voiceprint features.
[0102] The first clustering module 303 is used to cluster the M first voiceprint features to obtain K first central features; and to cluster the N second voiceprint features to obtain J second central features.
[0103] The first statistical module 304 is used to calculate the cosine distance between the K first central features and the J second central features to obtain the first distance;
[0104] The first evaluation module 305 is used to evaluate the first speech synthesis model to be evaluated based on the first distance;
[0105] Where M, N, K and J are all positive integers greater than 1, M is greater than K and N is greater than J.
[0106] Optionally, the first statistics module 304 is specifically used to calculate the cosine distance between the first central feature and each second central feature for each first central feature, to obtain J cosine distances corresponding to the first central feature; and to sum the J cosine distances corresponding to the first central feature to obtain the sum of cosine distances corresponding to the first central feature; and to sum the sum of cosine distances corresponding to the K first central features to obtain the first distance.
[0107] Optionally, the first evaluation module 305 is specifically used to determine that the first speech synthesis model to be evaluated has been successfully evaluated when the first distance is less than the first preset threshold; and to determine that the first speech synthesis model to be evaluated has been unsuccessfully evaluated when the first distance is greater than or equal to the first preset threshold.
[0108] Optional, such as Figure 4 As shown, this application also provides a model evaluation device 300, based on Figure 3 The module, model evaluation device 300, further includes:
[0109] The second acquisition module 306 is used to acquire T third audio signals synthesized using the second speech synthesis model to be evaluated;
[0110] The second voiceprint extraction module 307 is used to extract voiceprints from each of the T third audio signals to obtain T third voiceprint features.
[0111] The second clustering module 308 is used to cluster the T third voiceprint features to obtain P third central features;
[0112] The second statistical module 309 is used to calculate the cosine distance between the P third central features and the J second central features to obtain the second distance;
[0113] The second evaluation module 310 is used to evaluate the first speech synthesis model to be evaluated or the second speech synthesis model to be evaluated based on the first distance and the second distance.
[0114] Where T and P are positive integers greater than 1, and T is greater than P.
[0115] Optionally, the cosine distance between any two of the K first central features is greater than a second preset threshold; the cosine distance between any two of the J second central features is greater than a third preset threshold.
[0116] The model evaluation device 300 provided in this application can implement all the processes implemented by the model evaluation device in the above-described model evaluation method embodiments, and can achieve the same beneficial effects. To avoid repetition, it will not be described again here.
[0117] According to embodiments of this application, this application also provides an electronic device and a computer-readable storage medium.
[0118] like Figure 5 The diagram shown is a block diagram of an electronic device according to an embodiment of the model evaluation method of this application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device may also represent various forms of mobile devices, such as personal digital processors, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of the present application described and / or claimed herein.
[0119] like Figure 5 As shown, the electronic device includes one or more processors 501, a memory 502, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components are interconnected via different buses and can be mounted on a common motherboard or otherwise as required. The processors can process instructions executed within the electronic device, including instructions stored in or on memory to display graphical information of a GUI on an external input / output device (such as a display device coupled to the interface). In other embodiments, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple electronic devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 5 Take a processor 501 as an example.
[0120] The memory 502 is the non-transitory computer-readable storage medium provided in this application. The memory stores instructions executable by at least one processor to cause the at least one processor to perform the model evaluation method provided in this application. The non-transitory computer-readable storage medium of this application stores computer instructions for causing a computer to perform the model evaluation method provided in this application.
[0121] Memory 502, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the model evaluation method in the embodiments of this application (e.g., appendix). Figure 3The first acquisition module 301, first voiceprint extraction module 302, first clustering module 303, first statistics module 304, first evaluation module 305, second acquisition module 306, second voiceprint extraction module 307, second clustering module 308, second statistics module 309, and second evaluation module 310 shown in diagram 4. The processor 501 executes various functional applications and data processing of the model evaluation device by running non-transient software programs, instructions, and modules stored in the memory 502, thereby implementing the model evaluation method in the above method embodiments.
[0122] Memory 502 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by the use of the electronic device according to the model evaluation method. Furthermore, memory 502 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 502 may optionally include memory remotely located relative to processor 501, and these remote memories can be connected to the electronic device of the model evaluation method via a network. Examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0123] The electronic device for the model evaluation method may further include an input device 503 and an output device 504. The processor 501, memory 502, input device 503, and output device 504 can be connected via a bus or other means. Figure 5 Taking the example of a connection between China and Israel via a bus.
[0124] Input device 503 can receive input digital or character information, as well as key signal inputs related to user settings and function control of electronic devices for model evaluation methods, such as touch screens, keypads, mice, trackpads, touchpads, joysticks, one or more mouse buttons, trackballs, joysticks, etc. Output device 504 may include display devices, auxiliary lighting devices (e.g., LEDs), and haptic feedback devices (e.g., vibration motors). The display device may include, but is not limited to, liquid crystal displays (LCDs), light-emitting diode (LED) displays, and plasma displays. In some embodiments, the display device may be a touch screen.
[0125] Various implementations of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, application-specific integrated circuits (ASICs), computer hardware, firmware, software, and / or combinations thereof. These various implementations may include: implementations in one or more computer programs that can be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, capable of receiving data and instructions from a storage system, at least one input device, and at least one output device, and transferring data and instructions to the storage system, the at least one input device, and the at least one output device.
[0126] These computational programs (also referred to as programs, software, software applications, or code) include machine instructions for a programmable processor and can be implemented using high-level procedural and / or object-oriented programming languages, and / or assembly / machine languages. As used herein, the terms “machine-readable medium” and “computer-readable medium” refer to any computer program product, device, and / or apparatus (e.g., disk, optical disk, memory, programmable logic device (PLD)) used to provide machine instructions and / or data to a programmable processor, including machine-readable media that receive machine instructions as machine-readable signals. The term “machine-readable signal” refers to any signal used to provide machine instructions and / or data to a programmable processor.
[0127] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and pointing device (e.g., a mouse or trackball) through which the user provides input to the computer. Other types of devices can also be used to provide interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including sound input, voice input, or tactile input).
[0128] The systems and technologies described herein can be implemented in computing systems that include backend components (e.g., as a data server), or computing systems that include middleware components (e.g., an application server), or computing systems that include frontend components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and technologies described herein), or any combination of such backend, middleware, or frontend components. The components of the system can be interconnected via digital data communication of any form or medium (e.g., a communication network). Examples of communication networks include local area networks (LANs), wide area networks (WANs), and the Internet.
[0129] Computer systems can include clients and servers. Clients and servers are generally located far apart and typically interact through communication networks. Client-server relationships are created by computer programs running on the respective computers and having a client-server relationship with each other.
[0130] In this embodiment, K first central features are obtained by clustering M first voiceprint features, and J second central features are obtained by clustering N second voiceprint features. The cosine distance between the K first central features and the J second central features is calculated to obtain a first distance. Based on this first distance, the overall fidelity of M first audio signals synthesized using the first speech synthesis model under evaluation can be assessed. This allows for rapid evaluation of the fidelity of a large batch of first audio signals, improving the evaluation efficiency of the first speech synthesis model under evaluation. Therefore, the above-mentioned technical means effectively solve the problem of low efficiency in evaluating personalized speech synthesis models in the prior art.
[0131] According to embodiments of this application, this application also provides a computer program product, including a computer program, which is stored in a computer-readable storage medium. The computer program product is executed by at least one processor to implement the various processes of the above-described model evaluation method embodiments and can achieve the same technical effects. To avoid repetition, it will not be described again here.
[0132] It should be understood that the various forms of processes shown above can be used to rearrange, add, or delete steps. For example, the steps described in this application can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution disclosed in this application can be achieved, and this is not limited herein.
[0133] The specific embodiments described above do not constitute a limitation on the scope of protection of this application. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the scope of protection of this application.
Claims
1. A model evaluation method, characterized in that, The method includes: Obtain M first audio signals synthesized using the first speech synthesis model to be evaluated, and obtain N recorded second audio signals; Voiceprint extraction is performed on each of the M first audio signals to obtain M first voiceprint features; voiceprint extraction is performed on each of the N second audio signals to obtain N second voiceprint features. The M first voiceprint features are clustered to obtain K first central features; the N second voiceprint features are clustered to obtain J second central features. The first distance is obtained by calculating the cosine distance between the K first central features and the J second central features; Based on the first distance, the first speech synthesis model to be evaluated is evaluated; Where M, N, K and J are all positive integers greater than 1, M is greater than K and N is greater than J; After acquiring M first audio signals synthesized using the first speech synthesis model to be evaluated, and N recorded second audio signals, the method further includes: Obtain T third audio signals synthesized using the second speech synthesis model to be evaluated; Voiceprint extraction is performed on each of the T third audio signals to obtain T third voiceprint features; Cluster the T third voiceprint features to obtain P third central features; The second distance is obtained by calculating the cosine distance between the P third central features and the J second central features; Based on the first distance and the second distance, the first speech synthesis model to be evaluated or the second speech synthesis model to be evaluated is evaluated. Where T and P are positive integers greater than 1, and T is greater than P.
2. The method according to claim 1, characterized in that, The evaluation of the first speech synthesis model to be evaluated based on the first distance includes: If the first distance is less than the first preset threshold, the first speech synthesis model to be evaluated is determined to have been successfully evaluated. If the first distance is greater than or equal to the first preset threshold, it is determined that the evaluation of the first speech synthesis model to be evaluated is unsuccessful.
3. The method according to claim 1, characterized in that, The cosine distance between any two of the K first central features is greater than a second preset threshold; the cosine distance between any two of the J second central features is greater than a third preset threshold.
4. A model evaluation device, characterized in that, The device includes: The first acquisition module is used to acquire M first audio signals synthesized using the first speech synthesis model to be evaluated, and to acquire N recorded second audio signals; The first voiceprint extraction module is used to extract voiceprints from each of the M first audio signals to obtain M first voiceprint features; and to extract voiceprints from each of the N second audio signals to obtain N second voiceprint features. The first clustering module is used to cluster the M first voiceprint features to obtain K first central features; and to cluster the N second voiceprint features to obtain J second central features. The first statistical module is used to calculate the cosine distance between the K first central features and the J second central features to obtain the first distance; The first evaluation module is used to evaluate the first speech synthesis model to be evaluated based on the first distance; Where M, N, K and J are all positive integers greater than 1, M is greater than K and N is greater than J; The device further includes: The second acquisition module is used to acquire T third audio signals synthesized using the second speech synthesis model to be evaluated; The second voiceprint extraction module is used to extract the voiceprint of each of the T third audio signals to obtain T third voiceprint features. The second clustering module is used to cluster the T third voiceprint features to obtain P third central features; The second statistical module is used to calculate the cosine distance between the P third central features and the J second central features to obtain the second distance; The second evaluation module is used to evaluate the first speech synthesis model to be evaluated or the second speech synthesis model to be evaluated based on the first distance and the second distance. Where T and P are positive integers greater than 1, and T is greater than P.
5. The apparatus according to claim 4, characterized in that, The first evaluation module is specifically used to determine that the evaluation of the first speech synthesis model to be evaluated is successful when the first distance is less than the first preset threshold; and to determine that the evaluation of the first speech synthesis model to be evaluated is unsuccessful when the first distance is greater than or equal to the first preset threshold.
6. The apparatus according to claim 4, characterized in that, The cosine distance between any two of the K first central features is greater than a second preset threshold; the cosine distance between any two of the J second central features is greater than a third preset threshold.
7. An electronic device, characterized in that, include: At least one processor; as well as A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor to enable the at least one processor to perform the method of any one of claims 1 to 3.
8. A non-transitory computer-readable storage medium storing computer instructions, characterized in that, The computer instructions are used to cause the computer to perform the method according to any one of claims 1 to 3.
9. A computer program product comprising a computer program that, when executed by a processor, implements the method according to any one of claims 1 to 3.