Methods, devices, equipment, and media for filtering prompt information in large-scale speech synthesis scenarios
By performing static quality checks and semantic similarity filtering on the initial prompt speech database, the problem of poor speech synthesis results caused by randomly selected input prompts was solved, achieving higher quality and semantically matched speech synthesis results.
Patent Information
- Application Number
- CN202510037565.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-09
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-01-09
AI Technical Summary
In existing technologies, randomly selected input prompts result in poor speech synthesis performance, failing to provide rich emotional information and expressive capabilities.
By acquiring an initial prompt speech library, static quality detection is performed, and the initial prompt speech with the highest static quality value is selected as the candidate prompt speech. The target prompt speech that best matches the speech synthesis text is then selected based on semantic similarity.
It improves the accuracy and quality of speech synthesis, ensuring that the output speech synthesis results are more consistent with the semantics and quality of the speech-synthesized text.
Smart Images

Figure CN119889277B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device, and medium for filtering large-scale prompt information in speech synthesis scenarios. Background Technology
[0002] With the development of artificial intelligence technology, large language models have achieved remarkable success in natural language processing, speech synthesis, and human-computer interaction. Currently, most speech synthesis technologies based on large language models employ autoregressive generation, which uses input prompts and text-based tagging to achieve zero-sample speech cloning with just a few seconds of prompting speech. The quality of the input prompts is a key factor in producing high-quality synthesized speech with good tone.
[0003] However, input prompts are usually specially set according to different scenarios. Randomly selected input prompts usually cannot provide rich emotional information and expressive ability, which limits the emotional expression of synthesized speech and thus leads to poor speech synthesis effect. Summary of the Invention
[0004] This invention provides a method, apparatus, device, and medium for filtering large model prompts in speech synthesis scenarios, in order to solve the problem that randomly selected input prompts in the prior art result in poor speech synthesis performance.
[0005] A method for filtering prompt information in a large model for speech synthesis scenarios, including:
[0006] Obtain an initial prompt voice library; the initial prompt voice library includes multiple initial prompt voices;
[0007] Static quality detection is performed on the initial prompt voice to determine the static quality value corresponding to each initial prompt voice;
[0008] Select a first preset number of initial prompt voices with the highest static quality values from all the initial prompt voices as candidate prompt voices;
[0009] Receive a speech synthesis instruction containing speech-synthesized text, and determine the semantic similarity between the speech-synthesized text and each of the candidate prompt speech;
[0010] The candidate prompt voice with the highest semantic similarity is determined as the target prompt voice of the synthesized speech text.
[0011] A large-scale prompt information filtering device for speech synthesis scenarios includes:
[0012] The voice library acquisition module is used to acquire an initial prompt voice library; the initial prompt voice library includes multiple initial prompt voices;
[0013] The static quality detection module is used to perform static quality detection on the initial prompt voice and determine the static quality value corresponding to each initial prompt voice.
[0014] The first prompt voice filtering module is used to select a first preset number of initial prompt voices with the highest static quality values from all the initial prompt voices as candidate prompt voices.
[0015] A similarity detection module is used to receive a speech synthesis instruction containing speech-synthesized text and determine the semantic similarity between the speech-synthesized text and each of the candidate prompt speech.
[0016] The second prompt voice filtering module is used to determine the candidate prompt voice with the highest semantic similarity as the target prompt voice of the speech synthesis text.
[0017] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the large model prompt information filtering method for the above-described speech synthesis scenario.
[0018] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the large model prompt information filtering method for the above-described speech synthesis scenario.
[0019] The aforementioned method, apparatus, device, and medium for selecting prompt information for large-scale speech synthesis scenarios involves static quality detection of initial prompts in a speech database to pre-evaluate their quality and select high-quality initial prompts for use in the large-scale model's speech synthesis task. Simultaneously, it considers the semantic similarity between the initial prompts and the synthesized text, dynamically selecting the target prompt that best matches the semantics of the synthesized text from candidate prompts. Thus, the selected target prompts possess high quality and semantic consistency with the synthesized text, improving the accuracy of speech synthesis and resulting in better performance of the large-scale model's output speech synthesis results. Attached Figure Description
[0020] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0021] Figure 1This is a schematic diagram of an application environment for a large model prompt information filtering method in a speech synthesis scenario according to an embodiment of the present invention;
[0022] Figure 2 This is a flowchart of a method for filtering large model prompt information in a speech synthesis scenario according to an embodiment of the present invention;
[0023] Figure 3 This is a flowchart of step S20 in the large model prompt information filtering method for speech synthesis scenarios according to an embodiment of the present invention;
[0024] Figure 4 This is a principle block diagram of a large model prompt information filtering device for a speech synthesis scenario according to an embodiment of the present invention;
[0025] Figure 5 This is a principle block diagram of the static quality detection module in the large model prompt information filtering device for speech synthesis scenarios according to an embodiment of the present invention;
[0026] Figure 6 This is a schematic diagram of a computer device according to an embodiment of the present invention. Detailed Implementation
[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0028] The method for filtering large model prompt information in speech synthesis scenarios provided in this embodiment of the invention can be applied to, for example... Figure 1 In the application environment shown, specifically, the large model prompt information filtering method for this speech synthesis scenario is applied in a large model prompt information filtering system for the speech synthesis scenario. This large model prompt information filtering system for the speech synthesis scenario includes, for example,... Figure 1The diagram illustrates a client, server, and smart cabinet. The client communicates with the server via a network, and the server and smart cabinet establish a service connection to address problems in existing technologies. The client, also known as the user terminal, refers to the program that provides local services to the client, corresponding to the server. The client can be installed on, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be a standalone server or a server cluster consisting of multiple servers. The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0029] In one embodiment, such as Figure 2 As shown, a method for filtering large model prompt information in speech synthesis scenarios is provided, and this method is applied to... Figure 1 Taking the server in the example, the following steps are included:
[0030] S10: Obtain the initial prompt voice library; the initial prompt voice library includes multiple initial prompt voices.
[0031] Understandably, the initial prompt speech can be speech data from different application scenarios, which can be obtained through web crawling from the internet, speech databases, or intelligent question-and-answer interaction systems. The initial prompt speech can be used to guide large-scale models, enabling them to convert text data into corresponding speech data when performing speech synthesis tasks, achieving voice replication, replication of different speech content from the same speaker, and so on. However, since the initial prompt speech is not subjected to quality detection and analysis during crawling, these initial prompt speeches may have problems such as low volume, unclear audio, excessive noise, missing audio, or inappropriate scene. Therefore, this application aims to first perform quality detection and screening on the initial prompt speech to improve the accuracy of large-scale models performing speech synthesis tasks.
[0032] S20: Perform static quality detection on the initial prompt voice and determine the static quality value corresponding to each initial prompt voice.
[0033] Understandably, the static quality detection in this application evaluates the initial prompt speech from four dimensions: the intensity of emotional expression, speech clarity, naturalness and overall quality, consistency between text and emotion, and the performance of model generation. Specifically, the intensity of emotional expression is evaluated by analyzing the pitch features of the initial prompt speech. Speech clarity, naturalness, and overall quality are evaluated using a pre-defined audio quality evaluation model. The consistency between text and emotion is evaluated using a pre-defined text quality evaluation model. The performance of model generation is evaluated by assessing the synthesized speech generated by the large model based on the initial prompt speech in terms of character error rate, feature similarity, and emotional similarity. The static quality value comprehensively reflects the quality level of the initial prompt speech across multiple dimensions, including the audio itself, emotional features, and consistency between text and speech.
[0034] S30: Select a first preset number of initial prompt voices with the highest static quality value from all the initial prompt voices as candidate prompt voices.
[0035] The first preset number can be set according to needs or the number of initial prompt voices. For example, when there are more than 50 initial prompt voices, the first preset number can be set to 10% of the total number of initial prompt voices (if there are 100 initial prompt voices, the first preset number can be set to 10).
[0036] Specifically, after performing static quality detection on the initial prompt voice and determining the static quality value corresponding to each initial prompt voice, the initial prompt voices are sorted based on the static quality values to form a prompt voice sequence. Then, starting from the first initial prompt voice in the prompt voice sequence, a first preset number of initial prompt voices are selected, and the selected initial prompt voices are used as candidate prompt voices.
[0037] S40: Receive a speech synthesis instruction containing speech-synthesized text, and determine the semantic similarity between the speech-synthesized text and each of the candidate prompt speech.
[0038] Understandably, speech synthesis commands can be sent through the user's client or automatically generated after the user uploads speech-synthesized text. The speech-synthesized text is the text data that the user needs to convert into speech data. For example, the speech-synthesized text could be: "I'm so happy you shared this with me. I'm very excited after listening to your sharing and can't wait to participate in this event!" In this case, the user wants to convert the above speech-synthesized text into speech data.
[0039] Furthermore, since the synthesized text is text-based data and the candidate prompts are speech-based data, to facilitate determining the semantic similarity between the synthesized text and the candidate prompts, the candidate prompts, generated through speech recognition, can first be converted into text representations. Alternatively, when constructing the initial prompt library, each initial prompt can be bound to its corresponding text representation, allowing direct acquisition of that text representation. Thus, after obtaining the candidate speech text corresponding to the candidate prompt, the semantic similarity between the synthesized text and the candidate speech text can be detected. Semantic similarity represents the semantic relationship between the candidate speech text and the synthesized text. Higher semantic similarity indicates a closer semantic relationship between the candidate prompt and the synthesized text; conversely, lower semantic similarity indicates a greater semantic difference between the candidate prompt and the synthesized text.
[0040] S50: The candidate prompt speech with the highest semantic similarity is determined as the target prompt speech of the synthesized speech text.
[0041] Specifically, after determining the semantic similarity between the synthesized speech text and each of the candidate prompt speech, the semantic similarity corresponding to each candidate prompt speech can be compared, thereby determining the candidate prompt speech with the highest semantic similarity as the target prompt speech of the synthesized speech text.
[0042] Furthermore, after determining the target prompt speech, both the speech-synthesized text and the target prompt speech can be simultaneously input into the large model. The target prompt speech serves as a prompt for the large model, allowing it to learn the speech features of the target prompt speech. Based on these features, the large model then performs speech synthesis processing on the speech-synthesized text, outputting the speech synthesis result. In addition, users can submit speech synthesis requests, such as requiring a speech synthesis result with the same speaker features as the target prompt speech, or a speech synthesis result with the same emotional expression as the target prompt speech, etc.
[0043] In this embodiment, static quality detection is performed on the initial prompts in the speech database to pre-evaluate their quality and select high-quality prompts for use in the large-scale model for speech synthesis. Simultaneously, the semantic similarity between the initial prompts and the synthesized text is considered, dynamically selecting the target prompt that best matches the semantics of the synthesized text from the candidate prompts. Thus, the selected target prompts are of high quality and semantically consistent with the synthesized text, improving the accuracy of speech synthesis and resulting in better performance of the large-scale model's output.
[0044] For example, in financial scenarios, it is often necessary to set up AI customer service or develop a digital assistant to answer users' basic questions, and it is desired that the AI customer service or digital assistant can output more human-like voice data. Therefore, large-scale voice data is needed to train the AI customer service or digital assistant, but manually collecting voice data is labor-intensive and time-consuming. In this case, the method of this application can be used to filter suitable large-scale model prompts, thereby enabling the large-scale model to quickly and massively output the required voice data, and using the voice data output by the large-scale model to train the AI customer service or digital assistant.
[0045] Another example is in medical settings, where some patients speak in dialects instead of Mandarin, which can be difficult for some doctors to understand. To accommodate these patients speaking in dialects, a dialect recognition model is often trained to identify specific dialects in the patient's speech data and convert them into Mandarin. Training this model requires collecting a large amount of dialect speech data. The method described in this application filters out suitable large-scale model prompts, enabling the large model to quickly and massively output corresponding dialect speech data. The dialect recognition model is then trained using this output dialect speech data.
[0046] Thus, as can be seen from the two examples above, by using the method of this application to select suitable large model prompts, the large model can be better guided. Whether in the model training task or the speech cloning task, the large model can generate better speech synthesis results by referring to the most matching large model prompts.
[0047] In one embodiment, such as Figure 3 As shown, step S20, namely, performing static quality detection on the initial prompt voice to determine the static quality value corresponding to each initial prompt voice, includes:
[0048] S201. Perform emotion classification processing on the initial prompt voice to determine multiple prompt information emotion groups; each prompt information emotion group contains multiple emotion prompt voices; the emotion prompt voices are initial prompt voices that conform to preset emotion categories.
[0049] Understandably, emotion classification processing involves categorizing the initial prompt voice into different emotional states. This means that after selecting initial prompt voices that match a preset emotion category, each initial prompt voice is further subdivided based on its emotional state, forming multiple emotion groups for the prompt information. The preset emotion categories may include, but are not limited to, sadness, happiness, surprise, panic, etc. Furthermore, initial prompt voices without emotional expression will not be classified into any prompt information emotion group.
[0050] Furthermore, there are many methods for emotion classification of initial prompt speech, such as using features of the pitch shape itself, like slope, curvature, and inflection points. However, these data are relatively simple in describing emotions and cannot capture the subtle differences between different emotional states. As a suprasegmental feature, pitch conveys information over a longer period than segmental features (such as spectral envelope). In this embodiment, classification is performed using the mean and variance of the pitch features of the initial prompt speech, as these two data points are more likely to evoke emotional resonance.
[0051] S202. Perform speech quality detection on the emotional prompt speech in each of the aforementioned prompt information emotion groups, and determine the speech quality score of each of the aforementioned emotional prompt speech.
[0052] Understandably, speech quality detection is used to detect the clarity, naturalness, and overall quality of emotional cues, as well as the consistency between text and emotion. In this embodiment, a pre-defined audio quality assessment model is used to detect the clarity, naturalness, and overall quality of emotional cues; this model can be the DNSMOS tool. A pre-defined text quality assessment model is used to detect the consistency between text and emotion; this model can be the ChatGPT API or a self-trained model.
[0053] Furthermore, after performing emotion classification processing on the initial prompt speech and determining multiple emotion groups of prompt information, the emotional prompt speech is input into a preset audio quality evaluation model. The clarity, naturalness, and overall quality of each emotional prompt speech are determined by the preset audio quality evaluation model, and then a first quality score of the emotional prompt speech is determined based on the clarity, naturalness, and overall quality.
[0054] Further, after performing emotion classification processing on the initial prompt speech and determining multiple emotion groups for the prompt information, the corresponding voice prompt text is obtained, and the first emotion category information corresponding to the emotion prompt speech is determined based on the emotion group of the prompt information corresponding to the emotion prompt speech. The voice prompt text and the first emotion category information are input into a preset text quality assessment model. The preset text quality assessment model determines whether the emotion information reflected by the voice prompt text is the same as the first emotion category information, and outputs a second quality score. Here, the voice prompt text is the text representation of the emotion prompt speech. The first emotion category information expresses the emotional state of the emotion prompt speech. The second quality score measures whether the emotion of the text representation of the emotion prompt speech is coherent and consistent with the pre-classified first emotion category information.
[0055] S203. Select the second preset number of emotional prompts with the highest voice quality scores from all the emotional prompts as the filtering prompts.
[0056] The second preset number can be customized or determined based on the number of emotional prompts obtained through filtering. For example, if the number of emotional prompts is 20, the second preset number can be 30% of the number of emotional prompts, i.e., the second preset number is 6.
[0057] Furthermore, after performing voice quality detection on the emotional prompts in each of the aforementioned emotional groups and determining the voice quality score of each emotional prompt, the emotional prompts are sorted based on the voice quality scores to form an emotional voice sequence. Then, starting from the first emotional prompt in the emotional voice sequence, a second preset number of emotional prompts are selected, and the selected emotional prompts are used as filtering prompts.
[0058] S204. Perform large model prompt detection on the filtering prompt voices to determine the static quality value of each filtering prompt voice.
[0059] Understandably, in speech synthesis systems based on large language models, the focus is not only on evaluating the quality of the prompts but also on recognizing that different models may produce drastically different outputs even when given the same prompt. This difference primarily stems from the prompt tagging strategy. To address this issue, we propose a large model prompt detection method that emphasizes the performance of different models when processing the same prompt, mainly including three components: character error rate, feature similarity, and sentiment similarity.
[0060] In this embodiment, the initial prompt speech is evaluated from four dimensions: the intensity of emotional expression, the clarity of speech, the consistency between text and emotion, and the performance of model generation. This allows the selected candidate prompt speech to improve the accuracy of the large model in performing speech synthesis tasks.
[0061] In one embodiment, step S201, namely, performing emotion classification processing on the initial prompt voice to determine multiple prompt information emotion groups, includes:
[0062] Acquire pitch feature information for each initial prompt speech, and determine the average pitch and pitch variance of each initial prompt speech based on the pitch feature information.
[0063] Understandably, pitch feature information is a feature that predates language production, giving speech its tone and rhythm; that is, pitch feature information is the pitch data of the initial prompt speech at different times. The pitch mean refers to the average pitch of the initial prompt speech at different time points. It can reflect the speaker's general tone or emotion. The pitch variance measures the change in pitch over time. A larger pitch variance may indicate a more vivid or emotional speech, while a smaller pitch variance may indicate a monotonous delivery. Different emotional states are associated with different pitch patterns. For example, the sadness and comfort categories both exhibit relatively low pitch mean and variance, indicating a calmer and lower pitch feature, with sadness being slightly more subdued. On the other hand, emotions like the happiness and surprise categories exhibit higher mean and variance, reflecting a more pronounced emotional intensity.
[0064] Based on the average pitch and pitch variance, all initial prompts are clustered to determine multiple initial emotion groups.
[0065] Specifically, some voice data containing emotional states are pre-classified manually. Then, based on the average pitch and pitch variance of the voice data under each emotional state after classification, the data range of the average pitch and pitch variance under different emotional states is determined. Thus, after determining the average pitch and pitch variance of each initial prompt voice based on the pitch feature information, the data range of the average pitch of the initial prompt voice and the average pitch of each emotional state can be compared, as can the data range of the pitch variance of the initial prompt voice and the pitch variance of each emotional state. This allows initial prompt voices that simultaneously belong to the data range of both the average pitch and the pitch variance to be clustered into that emotional state, classifying them into multiple initial emotional groups.
[0066] The emotional average value of the initial emotional group is determined based on the average pitch value of all initial prompts in the initial emotional group, and the variance average value of the initial emotional group is determined based on the pitch variance value of all initial prompts in the initial emotional group.
[0067] Specifically, after clustering all initial prompt voices based on the average pitch and pitch variance to determine multiple initial emotion groups, the emotional average of each initial emotion group can be determined based on the average of the average pitch of all initial prompt voices in each initial emotion group, and the variance average of each initial emotion group can be determined based on the average pitch variance of all initial prompt voices in the initial emotion group.
[0068] Obtain a strong sentiment threshold group and a weak sentiment threshold group; the strong sentiment threshold group includes a first average threshold and a first variance threshold, and the weak sentiment threshold group includes a second average threshold and a second variance threshold.
[0069] Understandably, when clustering all initial prompts into multiple initial emotion groups, the average and / or variance of some initial emotion groups are relatively smooth, meaning they cannot capture the emotional states at two extremes (such as sadness and excitement, surprise, etc.). In speech synthesis tasks, relatively neutral emotional states do not require special training. Therefore, strong and weak emotion threshold groups are introduced to filter the initial emotion groups, selecting speech data that more directly represents human emotional characteristics. Specifically, the first average threshold is greater than the second average threshold, and the first variance threshold is greater than the second variance threshold.
[0070] The initial emotional groups whose average emotional value is greater than a first average threshold and whose average variance is greater than a first variance threshold are determined as the prompt information emotional groups, and the initial emotional groups whose average emotional value is less than a second average threshold and whose average variance is less than a second variance threshold are determined as the prompt information emotional groups.
[0071] Specifically, after obtaining the strong emotion threshold group and the weak emotion threshold group, the emotional average value is compared with the first average threshold and the second average threshold, and the emotional variance value is compared with the first variance threshold and the second variance threshold. The initial emotion group whose emotional average value is greater than the first average threshold and whose variance value is greater than the first variance threshold is determined as the prompt information emotion group (the emotional prompt voice in this prompt information emotion group is a strong emotion category, such as excitement, happiness, or surprise). The initial emotion group whose emotional average value is less than the second average threshold and whose variance value is less than the second variance threshold is determined as the prompt information emotion group (the emotional prompt voice in this prompt information emotion group is a weaker or calmer emotion category, such as sadness or anxiety). In this embodiment, the initial emotion groups whose emotional average value is less than or equal to the first average threshold, greater than or equal to the second average threshold, and whose emotional variance value is less than or equal to the first variance threshold and greater than or equal to the second variance threshold are eliminated.
[0072] In one embodiment, step S204, namely, performing large-model prompt detection on the filtering prompt voices to determine the static quality value of each filtering prompt voice, includes:
[0073] Obtain a pre-defined large language model and training speech text.
[0074] The preset large language model can be a model with speech synthesis capabilities. The training speech text can be random text or text data from the same scene.
[0075] The training speech text and the filtering prompt speech are input into the preset large language model, so that the preset large language model generates the training synthesized speech corresponding to the training speech text based on the filtering prompt speech.
[0076] Specifically, after obtaining the preset large language model and training speech text, the training speech text and the filtering prompt speech can be input into the preset large language model, so that the large language model generates the training synthesized speech corresponding to the training speech text based on the speech features of the filtering prompt speech.
[0077] Based on the filtering prompt speech and its corresponding trained synthesized speech, a static quality value is determined for each of the filtering prompt speeches.
[0078] Specifically, after generating the training synthesized speech corresponding to the training speech text based on the filtering prompt speech using the preset large language model, the static quality value of each filtering prompt speech is determined based on the filtering prompt speech and its corresponding training synthesized speech.
[0079] In a specific real-time mode, determining the static quality value of each of the filtering prompts based on the filtering prompt speech and its corresponding trained synthesized speech includes:
[0080] Based on the training speech text, the character error rate of the training synthesized speech is determined.
[0081] Understandably, the character error rate reflects the accuracy of the synthesized speech during training relative to the training text. Thus, the text content corresponding to the synthesized speech can be determined, and this text content can be compared with the training text to detect whether there are erroneous characters in the synthesized speech. The character error rate is determined based on the proportion of erroneous characters among all characters in the text content corresponding to the synthesized speech.
[0082] Speaker features are extracted from the filtering prompt speech and its corresponding training synthesized speech, and the feature similarity between the speaker features of the filtering prompt speech and the speaker features of the training synthesized speech is determined.
[0083] Understandably, feature similarity reflects the degree of similarity between the speaker's voice in the selection prompt speech and the speaker's voice in the training synthesized speech. Higher feature similarity indicates that the speaker's voice in the selection prompt speech and the training synthesized speech are more similar, while lower feature similarity indicates that the speaker's voice in the selection prompt speech and the training synthesized speech are more dissimilar. Specifically, the Resemblyzer and WavLM tools can be used to evaluate the feature similarity between the speaker features of the selection prompt speech and the speaker features of the training synthesized speech.
[0084] Based on the emotion group of the prompt information corresponding to the filtering prompt voice, the first emotion category information corresponding to the filtering prompt voice is determined. At the same time, the training synthesized voice is subjected to emotion classification processing to determine the second emotion category information of the training synthesized voice.
[0085] Understandably, the first emotion category information represents the emotion category of the filtering prompt speech, while the second emotion category information represents the emotion category of the synthesized speech during training.
[0086] Based on the first emotion category information and the second emotion category information, the emotional similarity between the filtering prompt speech and the trained synthesized speech is determined.
[0087] Specifically, the emotion2vec model can be used to determine the emotional similarity between the filtering prompt speech and the training synthesized speech based on the first and second emotion category information.
[0088] Based on the character error rate, feature similarity, and sentiment similarity corresponding to the same filtering prompt voice, the static quality value of each filtering prompt voice is determined.
[0089] Specifically, after determining the character error rate, feature similarity, and emotional similarity of the training synthesized speech corresponding to the filtering prompt speech, the static quality value of each filtering prompt speech can be determined.
[0090] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0091] In one embodiment, a large model prompt information filtering device for a speech synthesis scenario is provided, which corresponds one-to-one with the large model prompt information filtering method for a speech synthesis scenario described in the above embodiments. For example... Figure 4 As shown, the large-scale prompt information filtering device for this speech synthesis scenario includes a speech library acquisition module 10, a static quality detection module 20, a first prompt speech filtering module 30, a similarity detection module 40, and a second prompt speech filtering module 50. Detailed descriptions of each functional module are as follows:
[0092] The voice library acquisition module 10 is used to acquire an initial prompt voice library; the initial prompt voice library includes multiple initial prompt voices.
[0093] Static quality detection module 20 is used to perform static quality detection on the initial prompt voice and determine the static quality value corresponding to each initial prompt voice.
[0094] The first prompt voice filtering module 30 is used to select a first preset number of initial prompt voices with the highest static quality values from all the initial prompt voices as candidate prompt voices.
[0095] The similarity detection module 40 is used to receive a speech synthesis instruction containing speech synthesis text and determine the semantic similarity between the speech synthesis text and each of the candidate prompt speech.
[0096] The second prompt voice filtering module 50 is used to determine the candidate prompt voice with the highest semantic similarity as the target prompt voice of the speech synthesis text.
[0097] In one embodiment, such as Figure 5 As shown, the static quality inspection module 20 includes:
[0098] The emotion classification submodule 201 is used to perform emotion classification processing on the initial prompt voice and determine multiple prompt information emotion groups; each prompt information emotion group contains multiple emotion prompt voices; the emotion prompt voices are initial prompt voices that conform to preset emotion categories;
[0099] The quality detection submodule 202 is used to perform voice quality detection on the emotional prompt speech in each of the emotional prompt information groups and determine the voice quality score of each emotional prompt speech.
[0100] The first filtering submodule 203 is used to select the second preset number of emotional prompts with the highest voice quality scores from all the emotional prompts as filtering prompts.
[0101] The large model prompt detection submodule 204 is used to perform large model prompt detection on the filtered prompt speech and determine the static quality value of each of the filtered prompt speech.
[0102] In one embodiment, the emotion classification submodule 201 includes:
[0103] The pitch acquisition unit is used to acquire pitch feature information of each of the initial prompting voices, and determine the average pitch and pitch variance of each of the initial prompting voices based on the pitch feature information.
[0104] A clustering unit is used to cluster all initial prompting voices based on the average pitch and pitch variance values to determine multiple initial emotion groups;
[0105] The average value calculation unit is used to determine the average value of the initial emotion group based on the average pitch value of all initial prompting voices in the initial emotion group, and to determine the average variance value of the initial emotion group based on the average pitch variance value of all initial prompting voices in the initial emotion group.
[0106] A threshold acquisition unit is used to acquire a strong emotion threshold group and a weak emotion threshold group; the strong emotion threshold group includes a first average threshold and a first variance threshold, and the weak emotion threshold group includes a second average threshold and a second variance threshold.
[0107] The threshold comparison unit is used to determine the initial emotion group whose average emotion value is greater than a first average threshold and whose average variance value is greater than a first variance threshold as the prompt information emotion group, and to determine the initial emotion group whose average emotion value is less than a second average threshold and whose average variance value is less than a second variance threshold as the prompt information emotion group.
[0108] In one embodiment, the quality inspection submodule 202 includes:
[0109] An audio quality assessment unit is used to input the emotional prompt speech into a preset audio quality assessment model, so as to determine a first quality score for each of the emotional prompt speech through the preset audio quality assessment model;
[0110] A text quality assessment unit is used to obtain a preset text quality assessment model and determine a second quality score for each of the emotional prompt voices through the preset text quality assessment model.
[0111] A quality detection unit is used to determine the sum of the first quality score and the second quality score as the voice quality score of the emotional prompt voice.
[0112] In one embodiment, the text quality assessment unit includes:
[0113] The emotion category determination subunit is used to obtain the voice prompt text corresponding to the emotion prompt voice, and determine the first emotion category information corresponding to the emotion prompt voice based on the prompt information emotion group corresponding to the emotion prompt voice;
[0114] The text quality assessment subunit is used to input the voice prompt text and the first emotion category information into the preset text quality assessment model, so as to determine the second quality score corresponding to the emotion prompt voice through the preset text quality assessment model.
[0115] In one embodiment, the large model cue detection submodule 204 includes:
[0116] The data acquisition unit is used to acquire a pre-defined large language model and training speech text;
[0117] The speech synthesis unit is used to input the training speech text and the filtering prompt speech into the preset large language model, so as to generate the training synthesized speech corresponding to the training speech text based on the filtering prompt speech through the preset large language model;
[0118] A speech comparison unit is used to determine the static quality value of each of the filtering prompts based on the filtering prompts and the corresponding trained synthesized speech.
[0119] In one embodiment, the voice comparison unit includes:
[0120] The character error rate determination subunit is used to determine the character error rate of the trained synthesized speech based on the training speech text;
[0121] The feature similarity determination subunit is used to extract speaker features from the screening prompt speech and its corresponding training synthesized speech, and to determine the feature similarity between the speaker features of the screening prompt speech and the speaker features of the training synthesized speech.
[0122] The emotion classification subunit is used to determine the first emotion category information corresponding to the filtering prompt voice based on the prompt information emotion group corresponding to the filtering prompt voice, and at the same time to perform emotion classification processing on the training synthesized voice to determine the second emotion category information of the training synthesized voice.
[0123] The emotion similarity determination subunit is used to determine the emotion similarity between the screening prompt speech and the training synthesized speech based on the first emotion category information and the second emotion category information;
[0124] The static quality value determination subunit is used to determine the static quality value of each of the filtering prompt voices based on the character error rate, feature similarity, and emotional similarity corresponding to the same filtering prompt voice.
[0125] Specific limitations regarding the large-model prompt information filtering device for speech synthesis scenarios can be found in the limitations of the large-model prompt information filtering method for speech synthesis scenarios described above, and will not be repeated here. Each module in the aforementioned large-model prompt information filtering device for speech synthesis scenarios can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in a computer device in hardware form, or stored in the memory of a computer device in software form, so that the processor can call and execute the corresponding operations of each module.
[0126] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 6As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computational and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system, computer programs, and the database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The database stores data used in the large-model prompt information filtering method for speech synthesis scenarios described in the above embodiments. The network interface of the computer device communicates with external terminals via a network connection. When the computer program is executed by the processor, it implements a large-model prompt information filtering method for speech synthesis scenarios.
[0127] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the large model prompt information filtering method for speech synthesis scenarios described in the above embodiment.
[0128] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the large model prompt information filtering method for speech synthesis scenarios described in the above embodiments.
[0129] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0130] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0131] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for filtering prompt information in a large model for speech synthesis scenarios, characterized in that, include: Obtain the initial prompt voice library; The initial prompt voice library includes multiple initial prompt voices; Static quality detection is performed on the initial prompt voice to determine the static quality value corresponding to each initial prompt voice; Select a first preset number of initial prompt voices with the highest static quality values from all the initial prompt voices as candidate prompt voices; Receive a speech synthesis instruction containing speech-synthesized text, and determine the semantic similarity between the speech-synthesized text and each of the candidate prompt speech; The candidate prompting voice with the highest semantic similarity is determined as the target prompting voice of the synthesized speech text; The step of performing static quality detection on the initial prompt speech to determine the static quality value corresponding to each initial prompt speech includes: The initial prompt voice is subjected to emotion classification processing to determine multiple prompt information emotion groups; each prompt information emotion group contains multiple emotion prompt voices; the emotion prompt voices are the initial prompt voices that conform to the preset emotion category; Speech quality detection is performed on the emotional prompt speech in each of the aforementioned prompt information emotion groups to determine the speech quality score of each of the aforementioned emotional prompt speech; Select the second preset number of emotional prompts with the highest voice quality scores from all the emotional prompts as the filtering prompts; A large model-based prompt detection is performed on the selected prompt voices to determine the static quality value of each selected prompt voice. The step of performing large-model prompt detection on the selected prompt speech to determine the static quality value of each selected prompt speech includes: Obtain a pre-defined large language model and training speech text; The training speech text and the filtering prompt speech are input into the preset large language model, so that the preset large language model generates the training synthesized speech corresponding to the training speech text based on the filtering prompt speech; Based on the filtering prompt speech and its corresponding trained synthesized speech, a static quality value is determined for each of the filtering prompt speeches.
2. The method for filtering large model prompt information in a speech synthesis scenario as described in claim 1, characterized in that, The initial prompt voice is subjected to emotion classification processing to determine multiple emotion groups for the prompt information, including: Acquire pitch feature information for each initial prompt speech, and determine the average pitch and pitch variance of each initial prompt speech based on the pitch feature information; Based on the average pitch and pitch variance, all initial prompt voices are clustered to determine multiple initial emotion groups; The emotional average value of the initial emotional group is determined based on the average pitch value of all initial prompts in the initial emotional group, and the variance average value of the initial emotional group is determined based on the pitch variance value of all initial prompts in the initial emotional group. Obtain a strong sentiment threshold group and a weak sentiment threshold group; the strong sentiment threshold group includes a first average threshold and a first variance threshold, and the weak sentiment threshold group includes a second average threshold and a second variance threshold; The initial emotional groups whose average emotional value is greater than a first average threshold and whose average variance is greater than a first variance threshold are determined as the prompt information emotional groups, and the initial emotional groups whose average emotional value is less than a second average threshold and whose average variance is less than a second variance threshold are determined as the prompt information emotional groups.
3. The method for filtering large model prompt information in a speech synthesis scenario as described in claim 1, characterized in that, The step of performing voice quality detection on the emotional prompt speech in each of the aforementioned prompt information emotion groups, and determining the voice quality score of each emotional prompt speech, includes: The emotional prompt speech is input into a preset audio quality evaluation model to determine a first quality score for each emotional prompt speech through the preset audio quality evaluation model; Obtain a preset text quality assessment model, and determine a second quality score for each of the emotional prompt voices using the preset text quality assessment model; The sum of the first quality score and the second quality score is determined as the voice quality score of the emotional prompt voice.
4. The method for filtering large model prompt information in a speech synthesis scenario as described in claim 3, characterized in that, The step of obtaining a preset text quality assessment model and determining a second quality score for each emotional prompt speech using the preset text quality assessment model includes: Obtain the voice prompt text corresponding to the emotional prompt voice, and determine the first emotional category information corresponding to the emotional prompt voice based on the emotional group of the prompt information corresponding to the emotional prompt voice; The voice prompt text and the first emotion category information are input into the preset text quality assessment model to determine the second quality score corresponding to the emotion prompt voice through the preset text quality assessment model.
5. The method for filtering large model prompt information in a speech synthesis scenario as described in claim 1, characterized in that, The step of determining the static quality value of each of the filtering prompts and the corresponding trained synthesized speech includes: Based on the training speech text, determine the character error rate of the training synthesized speech; Speaker features are extracted from the filtering prompt speech and its corresponding training synthesized speech, and the feature similarity between the speaker features of the filtering prompt speech and the speaker features of the training synthesized speech is determined. Based on the emotion group of the prompt information corresponding to the filtering prompt voice, the first emotion category information corresponding to the filtering prompt voice is determined, and at the same time, the training synthesized voice is subjected to emotion classification processing to determine the second emotion category information of the training synthesized voice. Based on the first emotion category information and the second emotion category information, the emotion similarity between the filtering prompt speech and the training synthesized speech is determined; Based on the character error rate, feature similarity, and sentiment similarity corresponding to the same filtering prompt voice, the static quality value of each filtering prompt voice is determined.
6. A large-scale prompt information filtering device for speech synthesis scenarios, characterized in that, include: The voice library acquisition module is used to acquire the initial prompt voice library; The initial prompt voice library includes multiple initial prompt voices; The static quality detection module is used to perform static quality detection on the initial prompt voice and determine the static quality value corresponding to each initial prompt voice. The first prompt voice filtering module is used to select a first preset number of initial prompt voices with the highest static quality values from all the initial prompt voices as candidate prompt voices. A similarity detection module is used to receive a speech synthesis instruction containing speech-synthesized text and determine the semantic similarity between the speech-synthesized text and each of the candidate prompt speech. The second prompt voice filtering module is used to determine the candidate prompt voice with the highest semantic similarity as the target prompt voice of the speech synthesis text; The step of performing static quality detection on the initial prompt speech to determine the static quality value corresponding to each initial prompt speech includes: The initial prompt voice is subjected to emotion classification processing to determine multiple prompt information emotion groups; each prompt information emotion group contains multiple emotion prompt voices; the emotion prompt voices are the initial prompt voices that conform to the preset emotion category; Speech quality detection is performed on the emotional prompt speech in each of the aforementioned prompt information emotion groups to determine the speech quality score of each of the aforementioned emotional prompt speech; Select the second preset number of emotional prompts with the highest voice quality scores from all the emotional prompts as the filtering prompts; A large model-based prompt detection is performed on the selected prompt voices to determine the static quality value of each selected prompt voice. The step of performing large-model prompt detection on the selected prompt speech to determine the static quality value of each selected prompt speech includes: Obtain a pre-defined large language model and training speech text; The training speech text and the filtering prompt speech are input into the preset large language model, so that the preset large language model generates the training synthesized speech corresponding to the training speech text based on the filtering prompt speech; Based on the filtering prompt speech and its corresponding trained synthesized speech, a static quality value is determined for each of the filtering prompt speeches.
7. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the large model prompt information filtering method for speech synthesis scenarios as described in any one of claims 1 to 5.
8. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the large model prompt information filtering method for speech synthesis scenarios as described in any one of claims 1 to 5.
Citation Information
Patent Citations
AI voice data analysis processing method and system
CN114898733A
Intelligent guiding system developed based on vehicle-mounted system and intelligent guiding method
CN118200671A