Model generation method, speech synthesis method, related device, medium and product
By obtaining target training samples and using large language models and lightweight models to train and generate speech synthesis models, the problem of user personalized speech synthesis in the prior art is solved, and high-quality and personalized speech generation is achieved, which is suitable for a variety of resource-constrained environments.
Patent Information
- Application Number
- CN202510654130.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-08-15
AI Technical Summary
The existing voice synthesis technology is difficult to meet users' personalized voice needs, including customized requirements in terms of tone, style, emotions, etc.
By obtaining target training samples, including target sample speech data and text that conform to target speech characteristics, using preset models for training, we generate a speech synthesis model that can output a target speech characteristics, and combine training techniques of large language models and lightweight models to achieve personalized speech synthesis.
It realizes personalized voice generation that meets user needs, improves the quality and applicability of voice synthesis, is suitable for resource-constrained environments, and meets the needs of diverse application scenarios.
Smart Images

Figure CN120496494A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of speech synthesis technology, and in particular to a model generation method, a speech synthesis method, related devices, media, and products. Background Art
[0002] With the rapid development of deep learning technology, the field of speech synthesis has made significant progress in recent years. Deep learning models, particularly those based on neural networks, have demonstrated the ability to learn complex speech features and linguistic patterns from large-scale datasets, thereby generating high-quality, natural and fluent speech output. These technological advances have greatly promoted the development of speech synthesis technology in multiple application areas, including but not limited to text-to-speech (TTS), voice assistants, and automatic subtitle generation.
[0003] Users have a growing demand for personalized speech, including customization of voice timbre, style, emotion, and other aspects. However, existing speech synthesis technologies still face challenges in achieving highly personalized speech output, making it difficult to meet users' personalized needs. Summary of the Invention
[0004] To overcome the problems existing in the related art, the present disclosure provides a model generation method, a speech synthesis method, related devices, media and products.
[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a model generation method, comprising: Obtaining target training samples, wherein the target training samples include multiple target data sets, each of the target data sets including target sample speech data that meets target speech characteristics and target sample text corresponding to the target sample speech data; The preset model is trained according to the target training sample to obtain a speech synthesis model, and the speech synthesis model is used to output speech data that meets the target speech characteristics.
[0006] Optionally, obtaining a target training sample includes: Acquire a first training sample, where the first training sample includes multiple sets of first data sets, each of the first data sets includes a first sample text and first sample speech data corresponding to the first sample text; Determining a first data set including first sample speech data meeting the target speech feature as a second data set; A target data set is determined based on the second data set to obtain target training samples.
[0007] Optionally, determining a target data set based on the second data set to obtain a target training sample includes: The second data set is determined as a target data set to obtain target training samples.
[0008] Optionally, determining a target data set based on the second data set to obtain a target training sample includes: Performing text augmentation processing on the first sample text in the second data set to obtain a target sample text; Inputting the target sample text into a pre-trained second model to obtain target sample speech data corresponding to the target sample text, wherein the second model is obtained by training a first model based on the second data set, and the first model is obtained by training a large language model based on the first training sample; A target training sample is obtained according to the target sample text and the target sample speech data.
[0009] Optionally, performing text augmentation processing on the sample text in the second data set to obtain target sample text includes: Performing text augmentation processing on the sample text in the second data set to obtain a second sample text; Invalid text in the second sample text is removed to obtain a target sample text.
[0010] Optionally, training a preset model according to the target training sample to obtain a speech synthesis model includes: The first model is trained according to the target training sample to obtain a speech synthesis model, where the first model is obtained by training a large language model according to the first training sample.
[0011] Optionally, the training of a preset model according to the target training sample to obtain a speech synthesis model includes: The third model is trained according to the target training sample to obtain a speech synthesis model, and the third model is a lightweight model.
[0012] Optionally, the first training sample further includes sample speech features corresponding to each first data set, the sample speech features are determined based on the first sample speech data, and the target speech feature is one or more of the sample speech features; The step of determining the first data set including the first sample speech data meeting the target speech feature as the second data set includes: A first data set including sample speech features that meet the target speech features is determined as a second data set.
[0013] Optionally, the target training samples further include target sample speech features corresponding to each target data set, and a preset model is trained according to the target training samples to obtain a speech synthesis model, including: The target sample text and the target sample speech features are used as model input parameters, and the target sample speech data corresponding to the target sample text is used as a model output parameter. The preset model is trained to obtain a speech synthesis model.
[0014] According to a second aspect of an embodiment of the present disclosure, a speech synthesis method is provided, the speech synthesis method comprising: Get text data; According to the text data and a pre-trained speech synthesis model, speech data corresponding to the text data is output, and the speech data meets the target speech features. The speech synthesis model is generated according to the model generation method described in the first aspect of the embodiment of the present disclosure.
[0015] Optionally, the speech synthesis method is applied to an electronic device, and the speech synthesis model is deployed in the electronic device.
[0016] Optionally, the speech synthesis model is obtained by training the target sample text and the target sample speech features corresponding to the target sample text as model input parameters, and the target sample speech data corresponding to the target sample text as model output parameters; the speech synthesis method further includes: Acquiring speech features of the speech to be synthesized, wherein the speech features include the target speech features; Outputting speech data corresponding to the text data based on the text data and a pre-trained speech synthesis model includes: The text data and the speech features are input into a pre-trained speech synthesis model to obtain speech data corresponding to the text data output by the speech synthesis model.
[0017] According to a third aspect of an embodiment of the present disclosure, a model generation device is provided, wherein the model generation device is used to implement the model generation method described in any one of the first aspects of the embodiments of the present disclosure.
[0018] According to a fourth aspect of an embodiment of the present disclosure, a speech synthesis device is provided, which is used to implement the steps of the speech synthesis method described in any one of the second aspects of the embodiments of the present disclosure.
[0019] According to a fifth aspect of the embodiments of the present disclosure, there is provided an electronic device, including: processor; a memory for storing processor-executable instructions; In which, the processor is configured to execute the instructions so that the electronic device implements the steps of the model generation method as described in the first aspect of the embodiment of this disclosure, and / or the steps of the speech synthesis method as described in the second aspect of the embodiment of this disclosure.
[0020] According to the sixth aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the program instructions are executed by a processor, the steps of the model generation method as described in the first aspect of the embodiments of the present disclosure and / or the steps of the speech synthesis method as described in the second aspect of the embodiments of the present disclosure are implemented.
[0021] According to the seventh aspect of the embodiments of the present disclosure, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the steps of the model generation method as described in the first aspect of the embodiments of the present disclosure, and / or the steps of the speech synthesis method as described in the second aspect of the embodiments of the present disclosure.
[0022] Using this technical solution, a pre-set model is trained using target training samples that match the target speech characteristics, resulting in a speech synthesis model that can output speech data that matches the target speech characteristics. This allows the model to output speech data that meets user needs, satisfying their need for customized speech generation and achieving personalized speech synthesis.
[0023] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0025] Figure 1 The figure is a flowchart of a model generation method according to an exemplary embodiment.
[0026] Figure 2 The figure is a schematic diagram showing a method of generating a target sample text according to an exemplary embodiment.
[0027] Figure 3 It is a schematic diagram showing a method of constructing a first model according to an exemplary embodiment.
[0028] Figure 4 is a schematic diagram showing a method of constructing a second model according to an exemplary embodiment.
[0029] Figure 5 The figure is a flowchart of a speech synthesis method according to an exemplary embodiment.
[0030] Figure 6 The figure is a block diagram of a model generation device according to an exemplary embodiment.
[0031] Figure 7 The figure is a block diagram of a speech synthesis apparatus according to an exemplary embodiment.
[0032] Figure 8 It is a block diagram of an electronic device according to an exemplary embodiment. DETAILED DESCRIPTION
[0033] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0034] It should be noted that all actions of acquiring signals, information or data in the present disclosure are carried out in compliance with the corresponding data protection laws and policies of the country where they are located and with the authorization given by the owner of the corresponding device.
[0035] As mentioned in the background, current speech synthesis technology addresses the growing demand for personalized speech, including customization of voice timbre, style, and emotion. However, existing speech synthesis technology still faces challenges in achieving highly personalized speech output, making it difficult to meet users' personalized needs. In light of this, the present disclosure provides a model generation method, a speech synthesis method, related devices, media, and products.
[0036] Figure 1 FIG. 1 is a flow chart showing a model generation method according to an exemplary embodiment. Figure 1 As shown, the model generation method may include the following steps.
[0037] In step S11, a target training sample is obtained.
[0038] The target training samples may include multiple groups of target data sets, and each target data set may include target sample speech data that meets the target speech characteristics and target sample text corresponding to the target sample speech data.
[0039] In the present disclosure, sample speech data refers to historically synthesized speech data. Speech features refer to features related to speech data, including but not limited to: the speaker's gender, age, timbre, language, emotion, style, and dialect. Accordingly, the target speech feature may be one or more of the speech features related to the speech data. For example, the target speech feature may be the speaker's gender and age, or emotion, or style. Target sample speech data that conforms to the target speech feature means that the speech features in the target sample speech data are consistent with the target speech feature. For example, if the target speech feature is a humorous style, then the target sample speech data with a humorous style is recorded as the target sample speech data that conforms to the target speech feature.
[0040] In one possible approach, an audio screening tool can be used to screen target sample voice data that meets the target voice characteristics from a large amount of voice data. This disclosure does not limit the screening process. After the target sample voice data is screened, the text corresponding to the target sample voice data can be determined as the target sample text. In this way, a target training sample can be obtained, and the target training sample is a sample that meets the target voice characteristics.
[0041] In step S12, the preset model is trained according to the target training sample to obtain a speech synthesis model, which is used to output speech data that meets the target speech characteristics.
[0042] For example, a preset model can be trained using target sample text as a model input parameter and target sample speech data as a model output parameter. A speech synthesis model is obtained when training termination conditions are met. These termination conditions may include: reaching a preset number of training rounds or having a model loss that is less than or equal to a preset value.
[0043] Using this technical solution, a pre-set model is trained using target training samples that match the target speech characteristics, resulting in a speech synthesis model that can output speech data that matches the target speech characteristics. This allows the model to output speech data that meets user needs, satisfying their need for customized speech generation and achieving personalized speech synthesis.
[0044] First, the specific scheme for obtaining target training samples is described.
[0045] In the present disclosure, obtaining a target training sample may include: Obtaining a first training sample, where the first training sample includes multiple sets of first data sets, each of which includes a first sample text and first sample speech data corresponding to the first sample text; Determining a first data set including first sample speech data meeting the target speech feature as a second data set; A target data set is determined based on the second data set to obtain target training samples.
[0046] In one embodiment, a first data set that meets the target voice characteristics can be screened out from the first sample voice data, and the screened out first data set can be determined as the second data set. For example, the first data set that meets the target voice characteristics can be screened out manually or automatically by a device.
[0047] In another embodiment, in addition to including multiple sets of first data sets, the first training sample may also include sample speech features corresponding to each first data set. The sample speech features may include, but are not limited to, gender, age, timbre, language, emotion, style, and dialect. An embodiment of determining the first data set containing the first sample speech data that meets the target speech features as the second data set may include determining the first data set corresponding to the sample speech features that meet the target speech features as the second data set.
[0048] The target speech feature is one or more of the sample speech features corresponding to the first data set. Therefore, the sample speech features include the target speech features, and the first data set corresponding to the sample speech features including the target speech features can be determined as the second data set.
[0049] The sample speech features corresponding to the first data set may be manually annotated control labels, or may be features extracted from the first sample speech data by a speech feature extraction model.
[0050] For example, assuming that the target speech feature is a humorous style, the first data set corresponding to the sample speech feature including the humorous style may be determined as the second data set based on the sample speech feature corresponding to the first data set.
[0051] After obtaining the second data set, a target data set is determined based on the second data set to obtain target training samples.
[0052] In one embodiment, the second data set is determined as the target data set to obtain a target training sample.
[0053] In another embodiment, the training samples may be expanded based on the first sample text in the second data set to enrich the target training samples. In this embodiment, determining the target data set based on the second data set to obtain the target training samples may include: Performing text augmentation processing on the first sample text in the second data set to obtain a target sample text; Inputting the target sample text into a pre-trained second model to obtain target sample speech data corresponding to the target sample text, wherein the second model is obtained by training the first model based on the second data set, and the first model is obtained by training the large language model based on the first training sample; A target training sample is obtained according to the target sample text and target sample speech data.
[0054] After obtaining the second dataset, the first sample text in the second dataset is used as example text. A detailed user prompt is designed and augmented using the GPT (Generative Pre-trained Transformer) model. For example, if the target speech feature is humorous, the prompt can explicitly specify that the generated text be humorous.
[0055] Reference Figure 2 After designing the user prompt, you can input the prompt and sample text into the GPT model to generate augmented text. For example, you can adjust the temperature parameter and top-p sampling strategy in a large language model (e.g., the GPT model) to control the diversity and consistency of the generated augmented text. To improve augmentation efficiency, you can use batch generation, inputting multiple sample texts at once, to quickly accumulate tens of thousands of humorous sentences.
[0056] In accordance with Figure 2 After the augmented text is generated in the manner shown, the augmented text can be determined as the target sample text. However, considering that the generated augmented text may include invalid text, in order to ensure the quality of the target sample text, the target sample text can also be obtained by removing the invalid text in the augmented text.
[0057] For example, performing text augmentation processing on the sample text in the second data set to obtain the target sample text may include: performing text augmentation processing on the sample text in the second data set to obtain the second sample text; and removing invalid text in the second sample text to obtain the target sample text.
[0058] The valid text may refer to a text with incomplete sentences and / or incoherent text.
[0059] In addition, the augmented text can be iteratively optimized to further determine the quality of the target sample text. For example, Figure 2As shown, after obtaining the augmented text, you can also use automated screening tools to filter out high-quality samples from the augmented text, and use these high-quality samples as new sample texts. You can also design new user prompts for these high-quality samples, and then input the new sample texts and new user prompts into the GPT model to further optimize the augmentation effect. At the same time, you can also use multi-threading or distributed computing technology to fully utilize hardware resources to accelerate the augmentation process. Ultimately, through multiple iterations and optimizations, you can quickly obtain tens of thousands of high-quality target sample texts with consistent features to meet the needs of subsequent speech synthesis tasks. The number of iterations can be customized by the user, and this is not limited in this disclosure.
[0060] After the target sample text is obtained through text augmentation processing, since the text obtained through the augmentation processing has no corresponding sample speech data, in this embodiment, target sample speech data corresponding to the target sample text needs to be generated.
[0061] In one embodiment, the target sample text can be input into a pre-trained second model to obtain target sample speech data corresponding to the target sample text. The second model is obtained by training the first model based on the second data set, and the first model is obtained by training the large language model based on the first training sample.
[0062] In one possible approach, a large language model is first trained using a first training sample. For example, the large language model is trained using the first sample text as a model input parameter and the first sample speech data as a model output parameter, thereby obtaining a first model. Next, the first model is trained using a second dataset to obtain a second model. For example, the first sample text in the second dataset is used as a model input parameter and the first sample speech data in the second dataset is used as a model output parameter, thereby obtaining a second model.
[0063] In another possible approach, the first training samples also include sample speech features corresponding to each first data set. When training a large language model using the first training samples, the first sample text and the sample speech features corresponding to the first data set containing the first sample text are used as model input parameters, and the first sample speech data is used as the model output parameter to train the large language model, thereby obtaining a first model. This achieves multimodal input and improves model accuracy.
[0064] Figure 3 FIG. 1 is a schematic diagram showing a method of constructing a first model according to an exemplary embodiment. Figure 3As shown, the first sample text and the sample speech features corresponding to the first data set where the first sample text is located (for example, gender labels, age labels, language labels, emotion labels, style labels, and dialect labels) are converted into text token sequences through word segmentation or subword encoding, the acoustic features of the first sample speech data (such as Mel spectrum) are extracted and converted into acoustic token sequences, the text token sequences are used as model input parameters, and the acoustic token sequences are used as model output parameters to train the large language model. Among them, the large language model usually adopts a Transformer-based architecture, including a text encoder, an acoustic feature decoder, and a vocoder, and is trained by jointly optimizing multiple target loss functions. For example, it can be a loss function for mapping text to acoustic features (such as L1 / L2 loss) and a loss function for controlling label constraints (such as classification loss). When the combined loss error of multiple target losses is less than the preset loss error, the first model is obtained. For example, as Figure 3 As shown, the target loss function can include style loss function, sentiment loss function, and reconstruction loss function.
[0065] After training a large language model to obtain the first model, it can be fine-tuned based on the user's target speech features. This learned knowledge can be transferred to a second model with a smaller dataset through transfer learning. This allows small-dataset training tasks to benefit from the rich linguistic knowledge of the large language model, thereby improving the quality of their speech synthesis. During the optimization process, optimizers such as AdamW and learning rate scheduling strategies are used to ensure stable model convergence. Through this multi-task joint training and refined control, the second model can generate high-quality, diverse speech that meets the specified attributes, meeting the needs of user-customized scenarios.
[0066] For example, Figure 4 FIG. 1 is a schematic diagram showing a method of constructing a second model according to an exemplary embodiment. Figure 4 As shown, first, the target speech features are determined. Then, an audio screening tool is used to screen out a second data set including first sample speech data that meets the target speech features and sample speech features corresponding to the second data set. Next, the first sample text included in the second data set and the sample speech features corresponding to the second data set are converted into a text token sequence, and the first sample speech data included in the second data set is converted into an acoustic token sequence. The text token sequence is used as the input parameter of the first model, and the acoustic token sequence is used as the output parameter of the first model. The first model is trained to obtain the second model.
[0067] Because large language models can capture more speech data and features, generating more natural and fluent speech data, the synthesized speech is closer to natural human language expression, significantly improving the user's listening experience. At the same time, large language models can learn and understand richer speech emotional features, allowing the second model trained on the large language model to better express emotion and style, enhancing emotional communication with the user. Furthermore, using the first training sample to train the large language model to generate the first model gives the first model greater generalization capabilities, allowing it to adapt to different accents and languages, improving the applicability and accuracy of the final speech synthesis model.
[0068] Furthermore, after training a large language model to obtain the first model, the first model can be fine-tuned using a dataset that matches the target speech characteristics, enabling highly customized speech synthesis services to meet diverse user needs. At the same time, the large language model can better understand the context of the speech input and generate more coherent, scenario-appropriate speech output, making the speech synthesis model trained on the large language model more intelligent and user-friendly. This technology not only improves the quality of speech synthesis, but also supports flexible, efficient, and customizable speech generation, providing users with a more natural, smooth, and personalized speech synthesis experience that meets their diverse needs.
[0069] After obtaining the second model in the above manner, the target sample text is input into the pre-trained second model to obtain target sample speech data, and then obtain the target training sample.
[0070] It should be understood that if, during the training process of the second model, the input of the second model includes the sample speech features corresponding to the second data set, then during the process of generating the target sample speech data, it is also necessary to obtain the target sample speech features corresponding to the target sample text. The target sample speech features include the target speech features, and the target speech features are fixed features set by the user, such as humor style. Other speech features in the target sample speech features, except for the target speech features, may vary randomly.
[0071] In one embodiment, the target training samples also include target sample speech features corresponding to each of the target data sets. Training a preset model based on the target training samples to obtain a speech synthesis model may include: using the target sample text and target sample speech features as model input parameters, and using the target sample speech data corresponding to the target sample text as model output parameters, training the preset model to obtain a speech synthesis model.
[0072] In another embodiment, a preset model is trained according to a target training sample to obtain a speech synthesis model, including: training a first model according to the target training sample to obtain a speech synthesis model, wherein the first model is obtained by training a large language model according to the first training sample.
[0073] Since the first model is obtained by training a large language model, when the first model is trained using the target training sample to obtain a speech synthesis model, the speech synthesis model can generate more natural and fluent speech data, making the synthesized speech closer to human natural language expression, significantly improving the user's listening experience.
[0074] However, speech synthesis models require significant memory and computing power to run. Therefore, they are typically deployed in the cloud, which communicates with electronic devices. This means that electronic devices must establish a network connection with the cloud before using these models to synthesize speech. When computing resources and network connectivity are limited, these computationally intensive speech synthesis models cannot be used for speech synthesis.
[0075] Therefore, in another embodiment, a preset model is trained according to a target training sample to obtain a speech synthesis model, including: The third model is trained according to the target training sample to obtain a speech synthesis model. The third model is a lightweight model.
[0076] When the target sample speech data is obtained according to the second model, the target training samples including the target sample speech data can be used to train the lightweight third model. For example, first, the second model can serve as a teacher model to generate high-quality target training samples. The target training samples include the teacher model's knowledge, such as details such as the rhythm, timbre, and style of the speech. Next, a lightweight student model, i.e., the third model, is designed. Using the target sample speech data generated by the teacher model as the training target, the model is trained by minimizing the difference between the student model output and the teacher model output (such as mean squared error or KL divergence). In addition, adversarial training or perceptual loss can be introduced to make the speech generated by the student model sound more natural. Through distillation learning, the student model can approach or even achieve the synthesis effect of the teacher model while significantly reducing parameters and computing resources, thereby achieving efficient speech synthesis in a resource-limited environment.
[0077] By compressing the knowledge of the large-scale second model into a lightweight model through distillation technology, the computational effort is significantly reduced, enabling efficient operation in resource-constrained environments such as embedded systems or in-vehicle devices. Furthermore, lightweight models are more stable in actual use and less susceptible to computational resource constraints, meeting the needs of low-resource scenarios.
[0078] Because the speech synthesis model is trained using a lightweight model, it has a small size and low computational complexity. This enables fast inference and processing on the local device, significantly reducing latency and improving real-time performance, thus meeting the responsiveness requirements of speech synthesis in various scenarios.
[0079] Through knowledge distillation technology, knowledge from a large model is transferred to a smaller model, effectively solving the challenge of achieving high-quality speech synthesis in resource-constrained environments. Using speech data generated by the large model as training samples, the smaller model can learn the speech synthesis techniques and features of the larger model. This significantly reduces the computing resource requirements of the smaller model while maintaining high-quality speech output. This enables the speech synthesis model to operate efficiently in low-power environments, expanding its application scope.
[0080] Furthermore, the integration of model compression and optimization techniques (such as knowledge distillation) further effectively addresses the issues of large model parameters and high computational complexity. These methods reduce model size and improve efficiency without compromising speech synthesis quality, making them more suitable for deployment in resource-constrained scenarios. Through fine-tuning and optimization, the performance and efficiency of small models are further improved under limited computing resources.
[0081] The speech synthesis model generated in the above manner can be applied in different scenarios.
[0082] For example, in an in-vehicle environment, speech synthesis models can provide navigation systems with real-time, natural voice navigation instructions, support offline voice assistant functionality, and generate voice reminders to improve driving safety and convenience. In the smart home and IoT sectors, speech synthesis models can enable local voice control in offline environments and provide customizable voice feedback for smart devices, meeting user demands for privacy and responsiveness.
[0083] For example, in the areas of barrier-free devices and education and training, speech synthesis models can provide offline voice assistance functions such as voice navigation, voice reading, and voice reminders for the visually impaired and those with speech impairments, helping them better access information. At the same time, speech synthesis models can also generate offline voice teaching content for educational devices, such as word reading and text explanations, supporting customizable voice assistant functions to improve learning efficiency. In the entertainment and media field, speech synthesis models can generate offline audiobooks and voice dialogues for game characters, meeting users' needs for listening to books anytime, anywhere and enhancing game immersion.
[0084] In the fields of industrial security and healthcare, speech synthesis models can generate offline voice prompts (such as equipment failure warnings and safe operation instructions) and voice log reports, improving work efficiency and safety. In medical devices, they can also provide offline voice assistant functions, such as voice reminders for medication and voice reports of health data, helping medical staff quickly obtain diagnostic information. The application scenarios of customizable speech synthesis technology in offline environments are constantly expanding, providing users with a more intelligent, convenient, and secure voice interaction experience.
[0085] Figure 5 FIG. 1 is a flow chart of a speech synthesis method according to an exemplary embodiment. Figure 5 As shown, the speech synthesis method may include the following steps.
[0086] In step S51, text data is acquired.
[0087] For example, text data may be obtained from a speech synthesis request input by a user, for example, by performing field parsing on the speech synthesis request to obtain the text data.
[0088] In step S52, based on the text data and the pre-trained speech synthesis model, speech data corresponding to the text data is output, and the speech data conforms to the target speech features.
[0089] Among them, the speech synthesis model is generated according to the model generation method provided by the present disclosure.
[0090] By adopting the above technical solution, the speech synthesis model can effectively generate speech data that meets the target speech characteristics, meet the user's needs for customized speech generation, and achieve personalized speech synthesis effects.
[0091] In the present disclosure, a speech synthesis method is applied to an electronic device, and a speech synthesis model is deployed in the electronic device.
[0092] For example, after training the speech synthesis model according to the above scheme, the speech synthesis model can be saved in the model warehouse, and the configuration file information corresponding to the model can be written, migrated to the device side, and sent to the speech synthesis engine for inference operations to convert the text information into customized voice data.
[0093] This solution allows the speech synthesis model to run locally on the device, reducing reliance on cloud services and overall cloud computing costs. This technical solution not only improves the applicability of speech synthesis in resource-constrained environments but also provides strong support for voice interaction on electronic devices. It can adapt to various business scenarios and promote the implementation and popularization of speech synthesis technology in more practical applications.
[0094] In one embodiment, if the speech synthesis model is trained by using the target sample text and the target sample speech features corresponding to the target sample text as model input parameters, and using the target sample speech data corresponding to the target sample text as model output parameters; the speech synthesis method further includes: Acquiring speech features of the speech to be synthesized, the speech features including target speech features; Generating speech data corresponding to the text data based on the text data and a pre-trained speech synthesis model, including: The text data and speech features are input into a pre-trained speech synthesis model to obtain speech data corresponding to the text data output by the speech synthesis model.
[0095] The dimensions of the speech features may be the same as or different from the dimensions of the target sample speech features. For ease of description, the description is made by taking the example that the dimensions of the speech features may be the same as the dimensions of the target sample speech features. For example, assuming that the target sample speech features include the speaker's age, gender, style, language, emotion, and dialect, the speech features may also include the speaker's age, gender, style, language, emotion, and dialect. Assuming that the target speech feature is a humorous style, the style in the speech features of the speech to be synthesized is a humorous style, and the features of other dimensions may be user-defined features.
[0096] After obtaining the speech features of the speech to be synthesized, the text data and the speech features can be input into a pre-trained speech synthesis model to obtain speech data corresponding to the text data output by the speech synthesis model.
[0097] The following is a complete example of a speech synthesis method. The speech synthesis method includes the following steps: Step 1: The business layer calls the external interface and issues a speech synthesis request; Among them, the business layer may include but is not limited to: car cabin device end, mobile phone device end, audio device end, and TV device end.
[0098] Step 2: The offline speech synthesis engine receives the speech synthesis request and performs field parsing on the information in the speech synthesis request to extract the text data and the speech features corresponding to the text data; Step 3: Create a speaker engine based on the speech features corresponding to the text data; Step 4: The speaker engine calls the speech synthesis framework: Step 4.1: The front-end and back-end inference submodules load the front-end model and speech synthesis model into memory; Step 4.2: The front-end inference submodule calls the front-end model to convert the text data into a phoneme list. Step 4.3: Pass the phoneme list to the backend inference submodule, which loads the speech synthesis model and uses it to perform inference and generate speech data. Step 4.4, performing noise reduction, resampling, encoding and decoding operations on the generated audio data; Step 4.5: Stream the voice data processed in step 4.4 back to the business layer.
[0099] In this way, the speech synthesis model can be run locally on the electronic device to achieve speech synthesis, which minimizes the dependence on cloud services and reduces the overall cost of cloud computing.
[0100] Based on the same inventive concept, the present disclosure also provides a model generating device. Figure 6 FIG. 1 is a block diagram of a model generation device according to an exemplary embodiment. Figure 6 As shown, the model generating device 600 may include: The first acquisition module 601 is configured to acquire target training samples, wherein the target training samples include multiple target data sets, each of which includes target sample speech data that meets the target speech characteristics and target sample text corresponding to the target sample speech data; The training module 602 is configured to train a preset model according to the target training sample to obtain a speech synthesis model, and the speech synthesis model is used to output speech data that meets the target speech characteristics.
[0101] Optionally, the first acquisition module 601 may include: A first acquisition submodule is configured to acquire a first training sample, where the first training sample includes a plurality of first data sets, each of which includes a first sample text and first sample speech data corresponding to the first sample text; A first determining submodule is configured to determine a first data set including first sample speech data meeting the target speech feature as a second data set; The second determination submodule is configured to determine a target data set according to the second data set to obtain a target training sample.
[0102] Optionally, the second determining submodule is configured to: determine the second data set as a target data set to obtain a target training sample.
[0103] Optionally, the second determination submodule is configured to: perform text augmentation processing on the first sample text in the second data set to obtain a target sample text; input the target sample text into a pre-trained second model to obtain target sample speech data corresponding to the target sample text, the second model is obtained by training the first model based on the second data set, and the first model is obtained by training a large language model based on the first training sample; obtain a target training sample based on the target sample text and the target sample speech data.
[0104] Optionally, the second determining submodule is further configured to: perform text augmentation processing on the sample text in the second data set to obtain a second sample text; and remove invalid text in the second sample text to obtain a target sample text.
[0105] Optionally, the training module 602 is configured to: train a first model according to the target training sample to obtain a speech synthesis model, where the first model is obtained by training a large language model according to the first training sample.
[0106] Optionally, the training module 602 is configured to: train a third model according to the target training sample to obtain a speech synthesis model, and the third model is a lightweight model.
[0107] Optionally, the first training sample also includes sample speech features corresponding to each first data set, the sample speech features are determined based on the first sample speech data, and the target speech features are one or more of the sample speech features; the first determination submodule is configured to: determine the first data set corresponding to the sample speech features that meet the target speech features as the second data set.
[0108] Optionally, the target training sample also includes target sample speech features corresponding to each target data set, and the training module 602 is configured to: use the target sample text and the target sample speech features as model input parameters, and use the target sample speech data corresponding to the target sample text as model output parameters, train the preset model, and obtain a speech synthesis model.
[0109] Regarding the model generation device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the model generation method, and will not be elaborated here.
[0110] Based on the same inventive concept, the present disclosure also provides a speech synthesis device. Figure 7 FIG. 1 is a block diagram of a speech synthesis device according to an exemplary embodiment. Figure 7As shown, the speech synthesis device 700 may include: The second acquisition module 701 is configured to acquire text data; The output module 702 is configured to output the speech data corresponding to the text data based on the text data and a pre-trained speech synthesis model, and the speech data meets the target speech features. The speech synthesis model is generated according to the model generation method described in the present disclosure.
[0111] Optionally, the speech synthesis model is deployed in the electronic device.
[0112] Optionally, the speech synthesis model is obtained by training the target sample text and the target sample speech features corresponding to the target sample text as model input parameters, and the target sample speech data corresponding to the target sample text as model output parameters; the speech synthesis device 700 may further include: A third acquisition module is configured to acquire speech features of the speech to be synthesized, wherein the speech features include the target speech features; The output module 702 is configured to: input the text data and the speech features into a pre-trained speech synthesis model to obtain speech data corresponding to the text data output by the speech synthesis model.
[0113] Regarding the speech synthesis device in the above embodiment, the specific manner in which each module performs operations has been described in detail in the embodiment of the speech synthesis method, and will not be elaborated here.
[0114] The present disclosure also provides a computer-readable storage medium having computer program instructions stored thereon, which, when executed by a processor, implement the steps of the model generation method and / or speech synthesis method provided by the present disclosure.
[0115] Figure 8 8 is a block diagram of an electronic device according to an exemplary embodiment. For example, the electronic device 800 may be a vehicle-mounted system, a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0116] Reference Figure 8 , the electronic device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .
[0117] The processing component 802 generally controls the overall operation of the electronic device 800, such as operations associated with display, phone calls, data communications, camera operation, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned model generation method and / or to complete all or part of the steps of the above-mentioned speech synthesis method. In addition, the processing component 802 may include one or more modules to facilitate interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate interaction between the multimedia component 808 and the processing component 802.
[0118] The memory 804 is configured to store various types of data to support operations on the electronic device 800. Examples of such data include instructions for any application or method operating on the electronic device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device, or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.
[0119] The power supply component 806 provides power to the various components of the electronic device 800. The power supply component 806 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the electronic device 800.
[0120] The multimedia component 808 includes a screen that provides an output interface between the electronic device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, it may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, slides, and gestures on the touch panel. The touch sensors can not only sense the boundaries of a touch or slide action, but also detect the duration and pressure associated with the touch or slide action. In some embodiments, the multimedia component 808 includes a front-facing camera and / or a rear-facing camera. When the electronic device 800 is in an operating mode, such as a capture mode or a video mode, the front-facing camera and / or the rear-facing camera can receive external multimedia data. Each front-facing camera and the rear-facing camera can have a fixed optical lens system or have focal length and optical zoom capabilities.
[0121] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 800 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals may be further stored in the memory 804 or transmitted via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.
[0122] The input / output interface 812 provides an interface between the processing component 802 and peripheral interface modules, such as a keyboard, a click wheel, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.
[0123] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the electronic device 800. For example, the sensor assembly 814 can detect the open / closed state of the electronic device 800, the relative positioning of components, such as the display and keypad of the electronic device 800. The sensor assembly 814 can also detect changes in the position of the electronic device 800 or a component of the electronic device 800, the presence or absence of user contact with the electronic device 800, the orientation or acceleration / deceleration of the electronic device 800, and temperature changes of the electronic device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an accelerometer, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0124] The communication component 816 is configured to facilitate wired or wireless communication between the electronic device 800 and other devices. The electronic device 800 can access a wireless network based on a communication standard, such as WiFi, 2G or 3G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0125] In an exemplary embodiment, the electronic device 800 may be implemented by one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components to execute all or part of the steps of the above-mentioned model generation method, and / or all or part of the steps of the speech synthesis method.
[0126] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as memory 804 including instructions. The instructions may be executed by processor 820 of electronic device 800 to perform all or part of the steps of the above-described model generation method and / or all or part of the steps of the speech synthesis method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, optical data storage device, etc.
[0127] In another exemplary embodiment, a computer program product is also provided, which includes a computer program that can be executed by a programmable device, and the computer program has a code portion for executing all or part of the steps of the above-mentioned model generation method and / or all or part of the steps of the speech synthesis method when executed by the programmable device.
[0128] It should be understood that, unless otherwise specifically noted, the features of the various embodiments of the present disclosure described herein may be combined with each other. As used herein, the term "and / or" includes any one of the relevant listed items and any combination of any two or more thereof; similarly, "at least one of" includes any one of the relevant listed items and any combination of any two or more thereof.
[0129] The terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features being referred to. Thus, a feature specified as "first" or "second" may explicitly or implicitly include at least one of such features. In this description, "plurality" means at least two, for example, two, three, etc., unless otherwise specifically defined.
[0130] Furthermore, the word "exemplary" is used herein to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Rather, the use of the word exemplary is intended to present concepts in a concrete manner. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or." That is, unless otherwise specified or clear from the context, "X applies to A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies to A; X applies to B; or X applies to both A and B, then "X applies to A or B" satisfies any of the aforementioned instances. Furthermore, the articles "a" and "an," as used in this application and the appended claims, are generally understood to mean "one or more," unless otherwise specified or clear from the context to refer to the singular form.
[0131] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding this specification and the accompanying drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. With particular regard to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. In addition, although particular features of the present disclosure may have been disclosed with respect to only one of several implementations, such features may be combined with one or more other features of other implementations as may be desired and advantageous for any given or particular application. Furthermore, to the extent that the terms "include," "have," "have," "have," or variations thereof are used in the detailed description or claims, such terms are intended to be inclusive in a manner similar to the term "comprising."
[0132] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the present disclosure being indicated by the appended claims.
[0133] It should be understood that the present disclosure is not limited to the exact structures that have been described above and shown in the drawings, and that various modifications and changes can be made without departing from the scope thereof. The scope of the present disclosure is limited only by the appended claims.
Claims
1. A model generation method, characterized in that: include: Obtaining target training samples, wherein the target training samples include multiple target data sets, each of the target data sets including target sample speech data that meets target speech characteristics and target sample text corresponding to the target sample speech data; The preset model is trained according to the target training sample to obtain a speech synthesis model, and the speech synthesis model is used to output speech data that meets the target speech characteristics.
2. The model generation method according to claim 1, characterized in that The obtaining of target training samples includes: Acquire a first training sample, where the first training sample includes multiple sets of first data sets, each of the first data sets includes a first sample text and first sample speech data corresponding to the first sample text; Determining a first data set including first sample speech data meeting the target speech feature as a second data set; A target data set is determined based on the second data set to obtain target training samples.
3. The model generation method according to claim 2, characterized in that: The determining a target data set according to the second data set to obtain a target training sample includes: The second data set is determined as a target data set to obtain target training samples.
4. The model generation method according to claim 2, characterized in that: The determining a target data set according to the second data set to obtain a target training sample includes: Performing text augmentation processing on the first sample text in the second data set to obtain a target sample text; Inputting the target sample text into a pre-trained second model to obtain target sample speech data corresponding to the target sample text, wherein the second model is obtained by training a first model based on the second data set, and the first model is obtained by training a large language model based on the first training sample; A target training sample is obtained according to the target sample text and the target sample speech data.
5. The model generation method according to claim 4, characterized in that: The performing text augmentation processing on the sample text in the second data set to obtain target sample text includes: Performing text augmentation processing on the sample text in the second data set to obtain a second sample text; Invalid text in the second sample text is removed to obtain a target sample text.
6. The model generation method according to any one of claims 2 to 5, characterized in that: The preset model is trained according to the target training sample to obtain a speech synthesis model, including: The first model is trained according to the target training sample to obtain a speech synthesis model, where the first model is obtained by training a large language model according to the first training sample.
7. The model generation method according to any one of claims 1 to 5, characterized in that: The step of training a preset model according to the target training sample to obtain a speech synthesis model includes: The third model is trained according to the target training sample to obtain a speech synthesis model, and the third model is a lightweight model.
8. The model generation method according to claim 2, characterized in that: The first training sample further includes sample speech features corresponding to each first data set, the sample speech features being determined based on the first sample speech data, and the target speech features being one or more of the sample speech features; The step of determining the first data set including the first sample speech data meeting the target speech feature as the second data set includes: A first data set including sample speech features that meet the target speech features is determined as a second data set.
9. The model generation method according to claim 1, characterized in that: The target training samples also include target sample speech features corresponding to each target data set. The preset model is trained according to the target training samples to obtain a speech synthesis model, including: The target sample text and the target sample speech features are used as model input parameters, and the target sample speech data corresponding to the target sample text is used as a model output parameter. The preset model is trained to obtain a speech synthesis model.
10. A speech synthesis method, characterized in that: The speech synthesis method comprises: Get text data; According to the text data and a pre-trained speech synthesis model, speech data corresponding to the text data is output, and the speech data meets the target speech features. The speech synthesis model is generated according to the model generation method as described in any one of claims 1-9.
11. The speech synthesis method according to claim 10, wherein: The speech synthesis method is applied to an electronic device, and the speech synthesis model is deployed in the electronic device.
12. The speech synthesis method according to claim 10, wherein: The speech synthesis model is obtained by training the target sample text and the target sample speech features corresponding to the target sample text as model input parameters and the target sample speech data corresponding to the target sample text as model output parameters; the speech synthesis method further includes: Acquiring speech features of the speech to be synthesized, wherein the speech features include the target speech features; Outputting speech data corresponding to the text data based on the text data and a pre-trained speech synthesis model includes: The text data and the speech features are input into a pre-trained speech synthesis model to obtain speech data corresponding to the text data output by the speech synthesis model.
13. A model generation device, characterized in that: The model generation device is used to implement the model generation method according to any one of claims 1 to 9.
14. A speech synthesis device, characterized in that: The speech synthesis device is used to implement the speech synthesis method according to any one of claims 10 to 12.
15. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; The processor is configured to execute the instructions to enable the electronic device to implement the steps of the model generation method according to any one of claims 1 to 9, and / or the steps of the speech synthesis method according to any one of claims 10 to 12.
16. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the model generation method according to any one of claims 1 to 9 and / or the steps of the speech synthesis method according to any one of claims 10 to 12 are implemented.
17. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the steps of the model generation method according to any one of claims 1 to 9 and / or the steps of the speech synthesis method according to any one of claims 10 to 12.
Citation Information
Cited By
Personalized voice generation system and method based on multivariable parameters
CN121506091A