Speech synthesis method and device, electronic equipment, storage medium and product

By processing the user's text information and emotional information, and using a pre-trained speech synthesis model to generate speech carrying emotional information, the problem of difficulty in expressing emotions in the prior art speech synthesis is solved, and the naturalness and emotional richness of speech are achieved.

CN120126449APending Publication Date: 2025-06-10XIAOMI EV TECH CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510400463.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

Existing speech synthesis technology is difficult to effectively express user emotions, resulting in the generated speech lacking emotional richness and authenticity.

Method used

By processing the text information and emotional information of the target user, a mark sequence with fused emotional information is obtained, and the mark sequence is processed using the pre-trained target speech synthesis model to generate a target speech spectrum carrying emotional information, and finally a target synthetic speech integrating emotions is generated.

Benefits of technology

The generated speech can not only clearly express the user's semantics, but also truly express the user's emotions, enhancing the naturalness and emotional richness of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120126449A_ABST
    Figure CN120126449A_ABST
Patent Text Reader

Abstract

The invention relates to a speech synthesis method and device, electronic equipment, a storage medium and a product, and relates to the technical field of speech synthesis, and the method comprises the steps: processing text information and emotion information corresponding to a target user, and obtaining a mark sequence; processing the mark sequence through a target speech synthesis model to obtain a target speech spectrum, the target speech synthesis model being obtained by training a basic speech synthesis model based on a plurality of speech training samples, the speech training samples including sample text information, sample emotion information and sample speech; and obtaining a target synthetic speech according to the target speech spectrum. According to the embodiment of the invention, the target synthetic speech fused with emotions can be generated, so that the generated target synthetic speech can express the semantics of the user and can also express the emotions of the user.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of speech synthesis technology, and in particular, to a speech synthesis method, apparatus, electronic device, storage medium, and product. Background Art

[0002] Speech synthesis technology can not only provide more natural and fluent speech output, but also play an important role in multiple fields such as text-to-speech conversion, singing synthesis, music synthesis, lip synchronization synthesis, and virtual human motion synthesis.

[0003] In related technologies, the goal of speech synthesis is to generate clear and natural speech, mainly focusing on the intelligibility and fluency of speech. Summary of the Invention

[0004] To overcome the problems existing in related technologies, the present disclosure provides a speech synthesis method, apparatus, electronic device, storage medium, and product. First, the text information and emotional information corresponding to the target user can be processed to obtain a marked sequence fused with emotional information. Then, the marked sequence can be processed by a pre-trained target speech synthesis model, which is obtained by training a basic speech synthesis model based on multiple speech training samples including sample text information, sample emotional information, and sample speech, so as to obtain a target speech spectrum carrying emotional information. Furthermore, through the target speech spectrum, a target synthetic speech fused with emotion can be generated, so that the generated target synthetic speech can not only express the user's semantics but also express the user's emotion.

[0005] According to a first aspect of an embodiment of the present disclosure, there is provided a speech synthesis method, including: Processing the text information and emotional information corresponding to the target user to obtain a marked sequence; Processing the marked sequence through a target speech synthesis model to obtain a target speech spectrum, where the target speech synthesis model is obtained by training a basic speech synthesis model based on multiple speech training samples, and the speech training samples include sample text information, sample emotional information, and sample speech; Obtaining a target synthetic speech according to the target speech spectrum.

[0006] Optionally, the speech training sample further includes sample additional information, and the processing the text information and emotional information corresponding to the target user to obtain a marked sequence includes: Processing the text information, emotional information, and additional features corresponding to the target user to obtain a marked sequence, where the additional features include at least one of gender, language, age, and target style.

[0007] Optionally, the method further includes: Obtaining a target sign language video including sign language actions; Extract features from the target sign language video to obtain gesture features; Process the gesture features through a target sign language recognition model to obtain the text information. The target sign language recognition model is obtained by training a basic sign language recognition model based on multiple sign language training samples, and the sign language training samples include sample sign language videos and actual sign language texts.

[0008] Optionally, the method further includes: Obtain a target expression image including facial expressions; Extract features from the target expression image to obtain facial features; Process the facial features through a target emotion recognition model to obtain the emotion information. The target emotion recognition model is obtained by training a basic emotion recognition model based on multiple emotion training samples, and the emotion training samples include sample expression images and actual emotion categories.

[0009] Optionally, the method further includes: Extract emotion features from the text information to obtain the emotion information corresponding to the text information.

[0010] Optionally, the method further includes: Based on a preset correspondence between the scene and the emotion, determine the target emotion corresponding to the target scene to obtain the emotion information.

[0011] Optionally, the multiple speech training samples include multiple first training samples corresponding to different timbres, and the target speech synthesis model is obtained through the following steps: Through the multiple first training samples, perform multiple rounds of iterative training on the basic speech synthesis model to obtain the target speech synthesis model.

[0012] Optionally, the multiple speech training samples further include multiple second training samples corresponding to the target timbre; The step of performing multiple rounds of iterative training on the basic speech synthesis model through the multiple first training samples to obtain the target speech synthesis model includes: Through the multiple first training samples, perform multiple rounds of first training on the basic speech synthesis model to obtain an original speech synthesis model; Through the multiple second training samples, perform multiple rounds of second training on the original speech synthesis model to obtain the target speech synthesis model.

[0013] Optionally, the step of performing multiple rounds of first training on the basic speech synthesis model through the multiple first training samples to obtain an original speech synthesis model includes: Process the multiple first training samples to obtain multiple first target training samples. Each first target training sample includes a sample tag sequence and an actual speech spectrum. The sample tag sequence is obtained based on the sample text information and the sample emotion information, and the actual speech spectrum is obtained based on the sample speech. Perform multiple rounds of first training on the basic speech synthesis model using the multiple first target training samples. After each round of first training, obtain the first prediction loss corresponding to the current round of first training based on the predicted speech spectrum obtained in the current round of first training and the actual speech spectrum in the first target training sample corresponding to the current round of first training. Optimize the basic speech synthesis model according to the first prediction loss corresponding to the current round of first training. When the basic speech synthesis model meets the first preset stop condition, stop training to obtain the original speech synthesis model.

[0014] Optionally, the first training sample further includes sample additional information, and the sample additional information includes at least one of gender, language, age, and target style. The sample tag sequence is obtained based on the sample text information, the sample emotion information, and the sample additional information.

[0015] Optionally, the method further includes: Determine the target style corresponding to the target scenario based on the preset correspondence between the scenario and the style.

[0016] Optionally, the step of performing multiple rounds of second training on the original speech synthesis model using the multiple second training samples to obtain the target speech synthesis model includes: Process the multiple second training samples to obtain multiple second target training samples. Each second target training sample includes a sample tag sequence and an actual speech spectrum. The sample tag sequence is obtained based on the sample text information and the sample emotion information, and the actual speech spectrum is obtained based on the sample speech. Perform multiple rounds of second training on the original speech synthesis model using the multiple second target training samples. After each round of second training, obtain the second prediction loss corresponding to the current round of second training based on the predicted speech spectrum obtained in the current round of second training and the actual speech spectrum in the second target training sample corresponding to the current round of second training. Optimize the original speech synthesis model according to the second prediction loss corresponding to the current round of second training. When the original speech synthesis model meets the second preset stop condition, stop training to obtain the target speech synthesis model.

[0017] Optionally, the target sign language recognition model is obtained through the following steps: Process the multiple sign language training samples to obtain multiple target sign language training samples, where the target sign language training samples include sample gesture features and actual sign language texts, and the sample gesture features are obtained by extracting features from the sample sign language videos; Perform multiple rounds of third training on the basic sign language recognition model through the sample gesture features in the multiple target sign language training samples; After each round of third training, obtain the third prediction loss corresponding to this round of third training according to the predicted sign language text obtained in this round of third training and the actual sign language text in the target sign language training sample corresponding to this round of third training; Optimize the basic sign language recognition model according to the third prediction loss corresponding to this round of third training; When the basic sign language recognition model meets the third preset stop condition, stop training to obtain the target sign language recognition model.

[0018] Optionally, the target emotion recognition model is obtained through the following steps: Process the multiple emotion training samples to obtain multiple target emotion training samples, where the emotion training samples include sample facial features and actual emotion categories, and the sample facial features are obtained by extracting features from the sample expression images; Perform multiple rounds of fourth training on the basic emotion recognition model through the sample facial features in the multiple target emotion training samples; After each round of fourth training, obtain the fourth prediction loss corresponding to this round of fourth training according to the predicted emotion category obtained in this round of fourth training and the actual emotion category in the target emotion training sample corresponding to this round of fourth training; Optimize the basic emotion recognition model according to the fourth prediction loss corresponding to this round of fourth training; When the basic emotion recognition model meets the fourth preset stop condition, stop training to obtain the target emotion recognition model.

[0019] According to the second aspect of the embodiments of the present disclosure, a speech synthesis device is provided, including: A processing module, configured to process the text information and emotion information corresponding to the target user to obtain a marked sequence; A synthesis module, configured to process the marked sequence through a target speech synthesis model to obtain a target speech spectrum, where the target speech synthesis model is obtained by training a basic speech synthesis model based on multiple speech training samples, and the speech training samples include sample text information, sample emotion information, and sample speech; The first acquisition module is configured to obtain a target synthesized speech according to the target speech spectrum.

[0020] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: A processor; A memory for storing instructions executable by the processor; Wherein, the processor is configured to implement the steps of the speech synthesis method provided in the first aspect of the present disclosure when executed.

[0021] According to a fourth aspect of the embodiments of the present disclosure, there is provided a computer-readable storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the steps of the speech synthesis method provided in the first aspect of the present disclosure are implemented.

[0022] According to a fifth aspect of the embodiments of the present disclosure, there is provided a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the speech synthesis method provided in the first aspect of the present disclosure are implemented.

[0023] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By processing the text information and emotional information corresponding to the target user, a tag sequence is obtained, and then the target speech spectrum is obtained by processing the tag sequence through a target speech synthesis model. The target speech synthesis model is obtained by training a basic speech synthesis model based on multiple speech training samples, and the speech training samples include sample text information, sample emotional information, and sample speech. Furthermore, according to the target speech spectrum, a target synthesized speech is obtained.

[0024] By first processing the text information and emotional information corresponding to the target user to obtain a tag sequence that integrates emotional information, and then processing the tag sequence through a pre-trained target speech synthesis model. The target speech synthesis model is obtained by training a basic speech synthesis model based on multiple speech training samples including sample text information, sample emotional information, and sample speech, so that a target speech spectrum carrying emotional information can be obtained. Furthermore, through the target speech spectrum, a target synthesized speech integrating emotions can be generated, so that the generated target synthesized speech can not only express the semantics of the user, but also express the emotions of the user. It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.

[0026] Figure 1 It is a training schematic diagram of a target speech synthesis model shown according to an exemplary embodiment.

[0027] Figure 2 It is another training schematic diagram of a target speech synthesis model shown according to an exemplary embodiment.

[0028] Figure 3 It is a training schematic diagram of a target sign language recognition model shown according to an exemplary embodiment.

[0029] Figure 4 It is a training schematic diagram of a target emotion recognition model shown according to an exemplary embodiment.

[0030] Figure 5 It is a flowchart of a speech synthesis method shown according to an exemplary embodiment.

[0031] Figure 6 It is a block diagram of a speech synthesis device shown according to an exemplary embodiment.

[0032] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners

[0033] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements.

[0034] In some embodiments of the present disclosure described below, the described implementation manners do not represent all implementation manners consistent with the present disclosure. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present disclosure as detailed in the appended claims.

[0035] It should be noted that all actions of obtaining signals, information, or data in the present disclosure are carried out on the premise of complying with the corresponding data protection regulations and policies of the country where it is located and obtaining authorization from the owner of the corresponding device.

[0036] Speech synthesis technology can not only provide more natural and fluent speech output, but also play an important role in multiple fields such as text-to-speech conversion, singing synthesis, music synthesis, lip sync synthesis, and virtual human motion synthesis.

[0037] In the related art, the goal of speech synthesis is to generate clear and natural speech, mainly focusing on the intelligibility and fluency of speech.

[0038] In view of the above technical problems, embodiments of the present disclosure provide a speech synthesis method, apparatus, electronic device, storage medium, and product. First, the text information and emotional information corresponding to the target user can be processed to obtain a marked sequence integrating the emotional information. Then, the marked sequence can be processed by a pre-trained target speech synthesis model, which is obtained by training a basic speech synthesis model based on multiple speech training samples including sample text information, sample emotional information, and sample speech, so as to obtain a target speech spectrum carrying emotional information. Furthermore, through the target speech spectrum, a target synthesized speech integrating emotions can be generated, so that the generated target synthesized speech can not only express the user's semantics but also express the user's emotions.

[0039] First, some application scenarios of a speech synthesis method of the present disclosure will be introduced.

[0040] Solving communication barriers: Due to hearing or speech impairments, people with speech disorders often have difficulty communicating smoothly with others. The speech synthesis method of the present disclosure can convert text into speech, helping people with speech disorders express their thoughts through text and allowing the other party to hear clear speech. For example, people with speech disorders can input text on a mobile phone or a dedicated device, and the system will convert it into speech for playback, thus achieving barrier-free communication with the listener. This method greatly shortens the communication distance and improves the communication efficiency.

[0041] Emotional output: Sign language or text usually has difficulty clearly conveying emotions. The speech synthesis method of the present disclosure can analyze emotional information through the user's facial expressions and simulate corresponding emotional speech output. For example, the system can generate speech with emotions such as joy, sadness, or anger according to the user's expression, making the expression of people with speech disorders more vivid and rich. It not only enhances the authenticity of emotional communication but also makes the communication more warm and infectious.

[0042] Promoting social integration: In public places or working environments, the speech synthesis method of the present disclosure can help people with speech disorders better participate in social activities. For example, in online live broadcast platforms or meetings, people with speech disorders can use the device corresponding to the speech synthesis method of the present disclosure to communicate with others in real time, reducing social barriers. In addition, it can also be used in scenarios such as daily shopping and medical treatment to help people with speech disorders integrate into social life more confidently and enhance their interaction with the outside world.

[0043] Educational support: In the field of education, the speech synthesis method of the present disclosure can provide important learning support for students with speech disorders. For example, written teaching materials can be converted into understandable speech to help students better understand the learning content. At the same time, it can also be used for language learning to provide standard pronunciation demonstrations and assist students with speech disorders in mastering language skills. It not only improves the learning efficiency but also enhances the confidence and sense of participation of students with speech disorders.

[0044] The speech synthesis method of the present disclosure provides a brand-new communication method for speech-impaired persons through the function of text-to-speech, and can combine multi-modal information such as expressions and gestures, greatly broadening their communication channels with the outside world. It plays an important role in daily communication, emotional expression, social integration, and educational support, significantly improving the quality of life and social participation of speech-impaired persons.

[0045] Figure 1 It is a training schematic diagram of a target speech synthesis model shown according to an exemplary embodiment. As Figure 1 shown, in a possible implementation manner, multiple speech training samples include multiple first training samples corresponding to different timbres. The target speech synthesis model can be obtained through the following steps: Perform multiple rounds of iterative training on the basic speech synthesis model through multiple first training samples to obtain the target speech synthesis model.

[0046] In this implementation manner, if there is no requirement for timbre, multiple first training samples corresponding to different timbres can be used to perform multiple rounds of iterative training on the basic speech synthesis model, so as to be able to train a target speech synthesis model with rich speech features and emotional expression capabilities.

[0047] In a possible implementation manner, multiple speech training samples further include multiple second training samples corresponding to the target timbre.

[0048] Performing multiple rounds of iterative training on the basic speech synthesis model through multiple first training samples to obtain the target speech synthesis model includes: Performing multiple rounds of first training on the basic speech synthesis model through multiple first training samples to obtain the original speech synthesis model; performing multiple rounds of second training on the original speech synthesis model through multiple second training samples to obtain the target speech synthesis model.

[0049] In this implementation manner, if there is a requirement for timbre, multiple first training samples corresponding to different timbres can be used to perform multiple rounds of iterative training on the basic speech synthesis model first, so as to be able to train an original speech synthesis model with rich speech features and emotional expression capabilities. Then, transfer learning is performed on the basis of the original speech synthesis model, and multiple second training samples corresponding to the target timbre are continued to be used to fine-tune the original speech synthesis model, so as to be able to train a target speech synthesis model that can output the speech spectrum of the target timbre, so as to better meet the user's expectations and improve the user experience.

[0050] In a possible implementation manner, performing multiple rounds of first training on the basic speech synthesis model through multiple first training samples to obtain the original speech synthesis model includes: Process multiple first training samples to obtain multiple first target training samples. Each first target training sample includes a sample tag sequence and an actual speech spectrum. The sample tag sequence is obtained based on the sample text information and sample sentiment information, and the actual speech spectrum is obtained based on the sample speech; use the multiple first target training samples to perform multiple rounds of first training on the basic speech synthesis model; after each round of first training, according to the predicted speech spectrum obtained in this round of first training and the actual speech spectrum in the first target training sample corresponding to this round of first training, obtain the first prediction loss corresponding to this round of first training; optimize the basic speech synthesis model according to the first prediction loss corresponding to this round of first training; when the basic speech synthesis model meets the first preset stop condition, stop training to obtain the original speech synthesis model.

[0051] In this embodiment, multiple first training samples can be obtained. Each first training sample can include sample text information, sample sentiment information, and sample speech. The sample speech in the multiple first training samples can be speeches with different timbres.

[0052] First, the multiple first training samples can be processed, including performing word segmentation on the sample text information to obtain a word segmentation sequence, that is, a Token sequence. Each Token represents a semantic unit, such as a word, sub-word, or character, and special Tokens, such as [CLS], [SEP], are added to mark the sentence boundary or task type. And the speech features of the sample speech can be extracted to obtain the actual speech spectrum. The actual speech spectrum can be a Mel spectrum. The sample sentiment information can be aligned, and the sample speech can also be data-augmented, such as adding noise to the audio and time stretching, to improve the data diversity and the robustness of the trained model. Then, the sample sentiment information can be incorporated into the word segmentation sequence to obtain the sample tag sequence. Then, through the sample tag sequence and the actual speech spectrum corresponding to the same first training sample, the first target training sample corresponding to this first training sample is obtained.

[0053] Furthermore, the multiple first target training samples can be used to perform multiple rounds of first training on the basic speech synthesis model. Among them, the first training can be at least one of deep learning, reinforcement learning, and adversarial learning. In each round of first training, the first target training sample corresponding to each round of training can be obtained, and the basic speech synthesis model can be trained based on the corresponding learning method.

[0054] For any first target training sample, the semantic features and sentiment features of the sample tag sequence can be extracted. The sample tag sequence is mapped to a high-dimensional embedding vector, and positional encoding is added to retain the sequence order information to obtain the sample tag sequence vector. Then, through the sequence-to-sequence architecture or autoregressive generation ability of the basic speech synthesis model, the sample tag sequence vector is transformed into speech features, that is, the predicted speech spectrum is obtained.

[0055] After each round of the first training, based on the predicted speech spectrum obtained from the current round of the first training and the actual speech spectrum in the first target training sample corresponding to the current round of the first training, combined with the loss function corresponding to the learning method, an original loss can be obtained. The original loss can be at least one of mean square error, cross-entropy loss, and adversarial loss. Furthermore, the first prediction loss corresponding to the current round of the first training can be obtained through the original loss.

[0056] After obtaining the first prediction loss, optimization parameters can be obtained through hyperparameter optimization techniques and regularization methods, such as gradient clipping and weight decay. Then, the basic speech synthesis model can be optimized using the optimization parameters to prevent overfitting. When the basic speech synthesis model meets the first preset stop condition, the training is stopped, and the original speech synthesis model can be obtained. The first preset stop condition can be that the number of training rounds reaches the first preset number of rounds or the basic speech synthesis model converges.

[0057] The method of obtaining the target speech synthesis model through multiple first training samples and performing multiple rounds of iterative training on the basic speech synthesis model can refer to the above method of obtaining the original speech synthesis model, which will not be elaborated here.

[0058] In a possible implementation manner, the original speech synthesis model or the target speech synthesis model can also be evaluated using a validation set and a test set.

[0059] In a possible implementation manner, the first training sample further includes sample additional information, and the sample additional information includes at least one of gender, language, age, and target style. The sample tag sequence is obtained based on the sample text information, sample emotion information, and sample additional information.

[0060] In this implementation manner, sample additional information can also be added to obtain the original speech synthesis model or the target speech synthesis model that can output a speech spectrum conforming to the sample additional information. The sample additional information can include at least one of gender, language, age, and target style. When obtaining the sample tag sequence, the sample text information can be segmented to obtain a segmented sequence, and the features corresponding to the sample emotion information and the sample additional information can be added to the segmented sequence to obtain the sample tag sequence.

[0061] In a possible implementation manner, the method further includes: Determine the target style corresponding to the target scene based on the preset correspondence between the scene and the style.

[0062] In this embodiment, different scenarios can correspond to different styles, which can be set in advance by the user to establish a preset correspondence between the scenarios and styles, and then determine the target style corresponding to the target scenario to meet the diverse needs of users.

[0063] For example, in the entertainment live broadcast scenario, speech synthesis usually prefers a lively, positive and optimistic style. This style can significantly enhance the interactive fun with users and make the conversation more vivid and interesting. It can attract users' attention and improve their sense of participation through a natural and smooth intonation, rich rhythm changes, and appropriate humor, teasing or emotional expression. For example, according to the emotional changes of users, exaggerated tones, improvised responses or personalized expressions can be appropriately selected to bring a relaxed and pleasant experience to users. This style can not only meet the entertainment needs of users, but also enhance users' love and stickiness.

[0064] In the education and training scenario, speech synthesis mainly needs to be clear, accurate and professional. It should ensure clear pronunciation, moderate speaking speed, and be able to accurately convey teaching content. At the same time, the tone and intonation of the voice should show professionalism and authority to enhance users' trust in the content. For example, in language learning applications, speech synthesis needs to provide standard pronunciation demonstrations; in online courses, speech synthesis should be able to convey knowledge in a clear and well-organized manner with key points highlighted. This style not only helps users better understand and absorb information, but also improves learning efficiency and experience.

[0065] By customizing styles for different scenarios, it can better meet users' needs and improve user experience.

[0066] Figure 2 It is a training schematic diagram of another target speech synthesis model shown according to an exemplary embodiment. As Figure 2 shown, in a possible implementation, through multiple second training samples, the original speech synthesis model is trained multiple times in the second training to obtain the target speech synthesis model, including: Process multiple second training samples to obtain multiple second target training samples. Each second target training sample includes a sample label sequence and an actual speech spectrum. The sample label sequence is obtained based on the sample text information and sample emotion information, and the actual speech spectrum is obtained based on the sample speech. Use the multiple second target training samples to perform multiple rounds of second training on the original speech synthesis model. After each round of second training, based on the predicted speech spectrum obtained in this round of second training and the actual speech spectrum in the second target training sample corresponding to this round of second training, obtain the second prediction loss corresponding to this round of second training. Optimize the original speech synthesis model according to the second prediction loss corresponding to this round of second training. When the original speech synthesis model meets the second preset stop condition, stop the training to obtain the target speech synthesis model.

[0067] In this embodiment, multiple second training samples can be obtained. Each second training sample can include sample text information, sample emotion information, and sample speech. The sample speech in the multiple second training samples can be the sample speech of the target timbre.

[0068] First, process the multiple second training samples, including performing word segmentation on the sample text information to obtain a word segmentation sequence, that is, a Token sequence. Each Token represents a semantic unit, such as a word, sub-word, or character, and add special Tokens, such as [CLS], [SEP], to mark the sentence boundary or task type. And perform speech feature extraction on the sample speech to obtain the actual speech spectrum of the target timbre. The actual speech spectrum can be a Mel spectrum. Also, perform alignment processing on the sample emotion information, and perform data augmentation on the sample speech, such as adding noise to the audio and time stretching, to improve the data diversity and the robustness of the trained model. Then, integrate the sample emotion information into the word segmentation sequence to obtain the sample label sequence. Further, through the sample label sequence and the actual speech spectrum corresponding to the same second training sample, obtain the second target training sample corresponding to this second training sample.

[0069] Furthermore, perform multiple rounds of second training on the original speech synthesis model through the multiple second target training samples. Among them, the second training can be at least one of transfer learning, contrast learning, and adversarial learning. In each round of second training, obtain the second target training sample corresponding to each round of training, and train the original speech synthesis model based on the corresponding learning method.

[0070] For any second target training sample, extract the semantic features and emotion features of the sample label sequence. The sample label sequence is mapped to a high-dimensional embedding vector, and positional encoding is added to retain the sequence order information to obtain the sample label sequence vector. Then, through the sequence-to-sequence architecture or autoregressive generation ability of the original speech synthesis model, convert the sample label sequence vector into speech features, that is, obtain the predicted speech spectrum.

[0071] After each round of the second training, based on the predicted speech spectrum obtained from the current round of the second training and the actual speech spectrum in the second target training sample corresponding to the current round of the second training, combined with the loss function corresponding to the learning method, the original loss can be obtained. The original loss can be at least one of a transfer loss, a contrast loss, and an adversarial loss. Furthermore, the second prediction loss corresponding to the current round of the second training can be obtained through the original loss.

[0072] After obtaining the second prediction loss, the optimization parameters can be obtained through hyperparameter optimization techniques and regularization methods, such as gradient clipping and weight decay. Then, the original speech synthesis model can be optimized using the optimization parameters to prevent overfitting. When the original speech synthesis model meets the second preset stop condition, the training is stopped, and the target speech synthesis model can be obtained. The second preset stop condition can be that the number of training rounds reaches the second preset number of rounds or the original speech synthesis model converges.

[0073] Through the above method, transfer learning and further fine-tuning can be achieved based on the original speech synthesis model. In this way, the target speech synthesis model can inherit the powerful language ability of the original speech synthesis model, thus having higher quality and naturalness when generating speech with a specific timbre. The core of transfer learning lies in leveraging the general knowledge of the original speech synthesis model to provide better initialization for the target speech synthesis model and reduce the dependence on a large amount of target data.

[0074] In addition, in order to achieve efficient deployment with limited computing resources, model compression and optimization techniques, such as pruning, quantization, and knowledge distillation, can be used to process the target speech synthesis model. These methods can significantly reduce the number of parameters and computational complexity of the target speech synthesis model while maintaining high-quality speech synthesis. Through the comprehensive application of transfer learning, fine-tuning, and model compression, an efficient and personalized speech synthesis solution can be achieved while ensuring the effect.

[0075] Figure 3 It is a training schematic diagram of a target sign language recognition model shown according to an exemplary embodiment. As Figure 3 shown, in a possible implementation manner, the target sign language recognition model can be obtained through the following steps: Process multiple sign language training samples to obtain multiple target sign language training samples. The target sign language training samples include sample gesture features and actual sign language texts. The sample gesture features are obtained by extracting features from the sample sign language videos. Use the sample gesture features in the multiple target sign language training samples to perform multiple rounds of third training on the basic sign language recognition model. After each round of third training, based on the predicted sign language text obtained in this round of third training and the actual sign language text in the target sign language training sample corresponding to this round of third training, obtain the third prediction loss corresponding to this round of third training. Optimize the basic sign language recognition model according to the third prediction loss corresponding to this round of third training. When the basic sign language recognition model meets the third preset stop condition, stop the training to obtain the target sign language recognition model.

[0076] In this embodiment, multiple sign language training samples can be processed to obtain multiple target sign language training samples. Specifically, feature extraction can be performed on the sample sign language videos in each sign language training sample. For example, key frames can be extracted, key points of the hand and body can be detected, data augmentation such as rotation and scaling can be performed, and normalization processing can be carried out to obtain sample gesture features, so as to reduce the data distribution difference and improve the model generalization ability. Thus, the sample gesture features and the actual sign language text corresponding to each sign language training sample are determined as the target sign language training sample corresponding to this sign language training sample.

[0077] Furthermore, use the sample gesture features in the multiple target sign language training samples to perform multiple rounds of third training on the basic sign language recognition model. In each round of third training, the basic sign language recognition model can process the input sample gesture features through a convolutional neural network or a recurrent neural network, extract the spatio-temporal features of the sign language actions, and can combine multi-modal information such as videos, key points, and texts, so as to output the corresponding predicted sign language text.

[0078] After each round of third training, based on the predicted sign language text obtained in this round of third training and the actual sign language text in the target sign language training sample corresponding to this round of third training, combine the cross-entropy loss function or the CTC (Connectionist Temporal Classification) loss to obtain the third prediction loss corresponding to this round of third training.

[0079] Furthermore, optimization parameters can be generated based on the third prediction loss, and the basic sign language recognition model can be optimized through the optimization parameters. When the basic sign language recognition model meets the third preset stop condition, stop the training to obtain the target sign language recognition model. Among them, the third preset stop condition can be that the number of training rounds reaches the third preset number of rounds or the basic sign language recognition model converges.

[0080] Figure 4It is a training schematic diagram of a target emotion recognition model shown according to an exemplary embodiment. As Figure 4 shown, in a possible implementation, the target emotion recognition model can be obtained through the following steps: Process a plurality of emotion training samples to obtain a plurality of target emotion training samples. The emotion training samples include sample facial features and actual emotion categories. The sample facial features are obtained by extracting features from the sample expression images; through the sample facial features in the plurality of target emotion training samples, perform multiple rounds of fourth training on the basic emotion recognition model; after each round of fourth training, according to the predicted emotion category obtained in this round of fourth training and the actual emotion category in the target emotion training sample corresponding to this round of fourth training, obtain the fourth prediction loss corresponding to this round of fourth training; according to the fourth prediction loss corresponding to this round of fourth training, optimize the basic emotion recognition model; when the basic emotion recognition model meets the fourth preset stop condition, stop training to obtain the target emotion recognition model.

[0081] In this implementation, a plurality of emotion training samples can be processed to obtain a plurality of target emotion training samples. Specifically, feature extraction can be performed on the sample expression images in each emotion training sample, such as face detection and alignment, image normalization, data augmentation, such as rotation, cropping, and brightness adjustment, and grayscale or normalization processing, to obtain sample facial features, so as to reduce noise and improve the robustness of the model. Thus, the sample facial features and the actual emotion category corresponding to each emotion training sample are determined as the target emotion training sample corresponding to this emotion training sample.

[0082] Furthermore, multiple rounds of fourth training can be performed on the basic emotion recognition model through the sample facial features in the plurality of target emotion training samples. In each round of fourth training, the basic emotion recognition model can extract the spatial features of facial expressions through a convolutional neural network or a pre-trained model, such as VGG (Visual Geometry Group), ResNet (Residual Network). For the images corresponding to multiple consecutive video frames, 3D CNN (3D Convolutional Neural Network) or a temporal model, such as LSTM (Long Short-Term Memory), Transformer (a neural network architecture), can be combined to capture the dynamic changes of expressions, so as to obtain the predicted emotion category. The predicted emotion category can include one of joy, sadness, and anger.

[0083] After each round of the fourth training, according to the predicted emotion category obtained from the fourth training in this round and the actual emotion category in the target emotion training sample corresponding to the fourth training in this round, combined with the cross-entropy loss function or the mean squared error function, the fourth prediction loss corresponding to the fourth training in this round can be obtained.

[0084] Furthermore, according to the fourth prediction loss corresponding to the fourth training in this round, through regularization techniques such as dropout or weight decay, overfitting can be prevented to obtain the corresponding optimized parameters, and the basic emotion recognition model can be optimized by the optimized parameters. When the basic emotion recognition model meets the fourth preset stop condition, the training is stopped to obtain the target emotion recognition model. Among them, the fourth preset stop condition can be that the number of training rounds reaches the fourth preset number of rounds or the basic emotion recognition model converges.

[0085] Figure 5 is a flowchart of a speech synthesis method shown according to an exemplary embodiment. As Figure 5 shown, the following steps may be included.

[0086] In step S501, the text information and emotion information corresponding to the target user are processed to obtain a labeled sequence.

[0087] In this embodiment, the text information and emotion information that the target user expects to express can be obtained, and then the text information and emotion information corresponding to the target user are processed. For example, the text information is segmented to obtain a segmented sequence, and the emotion information is combined into the segmented sequence to obtain a corresponding labeled sequence carrying emotion information. Among them, the text information can be the text corresponding to the words that the target user wants to express, and the emotion information can be the emotion that the target user wants to express, such as joy, sadness, and anger, etc. The text information and emotion information can be directly input by the target user. The text information and emotion information can also be obtained based on the feature data analysis of the target user. Exemplarily, the feature data of the target user can include the target sign language video of the target user, the target expression image, and the preset correspondence between the preset scenarios and emotions set by the target user in advance.

[0088] In step S502, the labeled sequence is processed by the target speech synthesis model to obtain a target speech spectrum. The target speech synthesis model is obtained by training a basic speech synthesis model based on a plurality of speech training samples. The speech training samples include sample text information, sample emotion information, and sample speech.

[0089] In this embodiment, the labeled sequence can be processed by the pre-trained target speech synthesis model. The target speech synthesis model is obtained by training a basic speech synthesis model based on a plurality of speech training samples including sample text information, sample emotion information, and sample speech, so as to be able to obtain a target speech spectrum carrying emotion information. The target speech spectrum can be a Mel spectrum.

[0090] In step S503, a target synthesized speech is obtained according to the target speech spectrum.

[0091] In this embodiment, after obtaining the target speech spectrum carrying emotional information, a target synthesized speech carrying emotional information can be generated based on the target speech spectrum, so that the generated target synthesized speech can not only express the semantics of the user, but also express the emotion of the user. Among them, a vocoder, such as WaveGlow (Waveform Generation Model) or HiFi-GAN (High-Fidelity Generative Adversarial Network), can be used to generate the target synthesized speech from the target speech spectrum.

[0092] In this embodiment, the text information and emotional information corresponding to the target user are first processed to obtain a labeled sequence integrating emotional information, and then the labeled sequence is processed by a pre-trained target speech synthesis model. The target speech synthesis model is obtained by training a basic speech synthesis model based on a plurality of speech training samples including sample text information, sample emotional information, and sample speech, so as to obtain a target speech spectrum carrying emotional information. Furthermore, based on the target speech spectrum, a target synthesized speech integrating emotion can be generated, so that the generated target synthesized speech can not only express the semantics of the user, but also express the emotion of the user.

[0093] In a possible implementation manner, processing the text information and emotional information corresponding to the target user to obtain a labeled sequence includes: Processing the text information, emotional information, and additional features to obtain a labeled sequence, where the additional features include at least one of gender, language, age, and target style, and the speech training samples include sample text information, sample emotional information, sample additional information, and sample speech.

[0094] In this embodiment, additional features can also be added to be able to generate a target speech spectrum conforming to the additional features, and then obtain a target synthesized speech conforming to the additional features. For example, the additional features include at least one of gender, language, age, and target style, so that the target synthesized speech can carry features such as gender, language, age, and target style, better meeting the needs of users.

[0095] In a possible implementation manner, the method further includes: Obtain a target sign language video including sign language actions; extract features from the target sign language video to obtain gesture features; process the gesture features through a target sign language recognition model to obtain text information, where the target sign language recognition model is obtained by training a basic sign language recognition model based on multiple sign language training samples, and the sign language training samples include sample sign language videos and actual sign language texts.

[0096] In this embodiment, if the user does not input text information, text information can be generated through the target sign language recognition model. Specifically, a basic sign language recognition model can be pre-trained with multiple sign language training samples including sample sign language videos and actual sign language texts to obtain the target sign language recognition model. Furthermore, a target sign language video including the user's sign language actions can be obtained. The target sign language video can be collected by a sensor capturing the user's sign language actions, and the sensor can be a camera. Furthermore, features can be extracted from the target sign language video to obtain gesture features that can express the user's sign language information. Then, the trained target sign language recognition model processes the gesture features to obtain the text information.

[0097] In a possible implementation, the method for obtaining emotional information can be: Obtain a target expression image including facial expressions; extract features from the target expression image to obtain facial features; process the facial features through a target emotion recognition model to obtain emotional information, where the target emotion recognition model is obtained by training a basic emotion recognition model based on multiple emotion training samples, and the emotion training samples include sample expression images and actual emotion categories.

[0098] In this embodiment, if the user does not input emotional information, emotional information can be generated through the target emotion recognition model. Specifically, a basic emotion recognition model can be pre-trained with multiple emotion training samples including sample expression images and actual emotion categories to obtain the target emotion recognition model. Furthermore, a target expression image including the user's facial expressions can be obtained. The target expression image can be a single image or an image corresponding to multiple consecutive video frames. The target expression image can be collected by a sensor capturing the user's facial expressions, and the sensor can be a camera. Furthermore, features can be extracted from the target expression image to obtain facial features that can express the user's emotional information. Then, the trained target emotion recognition model processes the facial features to obtain the emotional information.

[0099] Through the above method, emotional information can be automatically generated based on the user's expression image.

[0100] In a possible implementation, the method further includes: Extract emotional features from the text information to obtain the emotional information corresponding to the text information.

[0101] In this embodiment, text information can be analyzed, and in combination with the context of the text information, the content and target context of the text information can be analyzed. Furthermore, the sentiment tendency of a preset word or phrase in the target context can be identified, thereby obtaining the sentiment information corresponding to the text information.

[0102] In this embodiment, corresponding sentiment information can be automatically generated based on the text information.

[0103] In a possible embodiment, the method further includes: Based on a preset correspondence between a scenario and a sentiment, determining the target sentiment corresponding to the target scenario to obtain sentiment information.

[0104] In this embodiment, the user can preset their own sentiment preferences, that is, set the sentiments corresponding to multiple scenarios, and then a preset correspondence between the scenarios and the sentiments can be established. Furthermore, the current target scenario can be determined, and based on the preset correspondence between the scenario and the sentiment, the target sentiment corresponding to the target scenario can be determined, thereby obtaining the sentiment information.

[0105] Through the above method, diversified acquisition of text information and sentiment information can be achieved. Through external input, multimodal acquisition, text prediction, and personalized settings, the diverse needs of users can be better met.

[0106] Figure 6 It is a block diagram of a speech synthesis device shown according to an exemplary embodiment. Referring to Figure 6 , the speech synthesis device 600 includes a processing module 601, a synthesis module 602, and a first acquisition module 603.

[0107] The processing module 601 is configured to process the text information and sentiment information corresponding to the target user to obtain a marked sequence; The synthesis module 602 is configured to process the marked sequence through a target speech synthesis model to obtain a target speech spectrum, and the target speech synthesis model is obtained by training a basic speech synthesis model based on a plurality of speech training samples, and the speech training samples include sample text information, sample sentiment information, and sample speech; The first acquisition module 603 is configured to obtain a target synthesized speech according to the target speech spectrum.

[0108] Optionally, the speech training sample further includes sample additional information, and the processing module 601 includes: A processing sub-module is configured to process the text information, sentiment information, and additional features corresponding to the target user to obtain a marked sequence, where the additional features include at least one of gender, language, age, and target style, and the speech training samples include sample text information, sample sentiment information, sample additional information, and sample speech.

[0109] Optionally, the speech synthesis device 600 further includes: A first acquisition module, configured to acquire a target sign language video including sign language actions; A first feature extraction module, configured to extract features from the target sign language video to obtain gesture features; A second acquisition module, configured to process the gesture features through a target sign language recognition model to obtain the text information, where the target sign language recognition model is obtained by training a basic sign language recognition model based on a plurality of sign language training samples, and the sign language training samples include sample sign language videos and actual sign language texts.

[0110] Optionally, the speech synthesis device 600 further includes: A second acquisition module, configured to acquire a target expression image including facial expressions; A second feature extraction module, configured to extract features from the target expression image to obtain facial features; A third acquisition module, configured to process the facial features through a target emotion recognition model to obtain the emotion information, where the target emotion recognition model is obtained by training a basic emotion recognition model based on a plurality of emotion training samples, and the emotion training samples include sample expression images and actual emotion categories.

[0111] Optionally, the speech synthesis device 600 further includes: A fourth acquisition module, configured to perform emotion feature extraction on the text information to obtain the emotion information corresponding to the text information.

[0112] Optionally, the speech synthesis device 600 further includes: A first determination module, configured to determine a target emotion corresponding to a target scene based on a preset correspondence between the scene and the emotion to obtain the emotion information.

[0113] Optionally, the plurality of speech training samples include a plurality of first training samples corresponding to different timbres, and the device further includes: A first training module, configured to perform multiple rounds of iterative training on the basic speech synthesis model through the plurality of first training samples to obtain the target speech synthesis model.

[0114] Optionally, the plurality of speech training samples further include a plurality of second training samples corresponding to a target timbre; The first training module includes: A first training sub-module, configured to perform multiple rounds of first training on the basic speech synthesis model through the plurality of first training samples to obtain an original speech synthesis model; A second training sub-module, configured to perform multiple rounds of second training on the original speech synthesis model through the multiple second training samples to obtain the target speech synthesis model.

[0115] Optionally, the first training sub-module includes: A first obtaining unit, configured to process the multiple first training samples to obtain multiple first target training samples, each first target training sample including a sample tag sequence and an actual speech spectrum, the sample tag sequence being obtained based on the sample text information and the sample emotion information, and the actual speech spectrum being obtained based on the sample speech; A first training unit, configured to perform multiple rounds of first training on the basic speech synthesis model through the multiple first target training samples; A first loss unit, configured to, after each round of first training, obtain a first prediction loss corresponding to the current round of first training according to the predicted speech spectrum obtained in the current round of first training and the actual speech spectrum in the first target training sample corresponding to the current round of first training; A first optimization unit, configured to optimize the basic speech synthesis model according to the first prediction loss corresponding to the current round of first training; A second obtaining unit, configured to stop training when the basic speech synthesis model meets a first preset stop condition to obtain the original speech synthesis model.

[0116] Optionally, the first training sample further includes sample additional information, the sample additional information including at least one of gender, language, age, and target style, and the sample tag sequence being obtained based on the sample text information, the sample emotion information, and the sample additional information.

[0117] Optionally, the speech synthesis device 600 further includes: A second determination unit, configured to determine the target style corresponding to the target scenario based on a preset correspondence between the scenario and the style.

[0118] Optionally, the second training sub-module includes: A third obtaining unit, configured to process the multiple second training samples to obtain multiple second target training samples, each second target training sample including a sample tag sequence and an actual speech spectrum, the sample tag sequence being obtained based on the sample text information and the sample emotion information, and the actual speech spectrum being obtained based on the sample speech; A second training unit, configured to perform multiple rounds of second training on the original speech synthesis model through the multiple second target training samples; A second loss unit, configured to obtain a second prediction loss corresponding to the current round of the second training according to the predicted speech spectrum obtained from the current round of the second training and the actual speech spectrum in the second target training sample corresponding to the current round of the second training; A second optimization unit, configured to optimize the original speech synthesis model according to the second prediction loss corresponding to the current round of the second training; A fourth obtaining unit, configured to stop the training when the original speech synthesis model meets a second preset stop condition, and obtain the target speech synthesis model.

[0119] Optionally, the speech synthesis device 600 further includes: A fifth obtaining module, configured to process the multiple sign language training samples to obtain multiple target sign language training samples, where the target sign language training samples include sample gesture features and actual sign language texts, and the sample gesture features are obtained by performing feature extraction on the sample sign language videos; A third training module, configured to perform multiple rounds of third training on the basic sign language recognition model through the sample gesture features in the multiple target sign language training samples; A first loss module, configured to obtain a third prediction loss corresponding to the current round of the third training according to the predicted sign language text obtained from the current round of the third training and the actual sign language text in the target sign language training sample corresponding to the current round of the third training; A first optimization module, configured to optimize the basic sign language recognition model according to the third prediction loss corresponding to the current round of the third training; A sixth obtaining module, configured to stop the training when the basic sign language recognition model meets a third preset stop condition, and obtain the target sign language recognition model.

[0120] Optionally, the speech synthesis device 600 further includes: A sixth obtaining module, configured to process the multiple emotion training samples to obtain multiple target emotion training samples, where the emotion training samples include sample facial features and actual emotion categories, and the sample facial features are obtained by performing feature extraction on the sample expression images; A fourth training module, configured to perform multiple rounds of fourth training on the basic emotion recognition model through the sample facial features in the multiple target emotion training samples; A second loss module, configured to obtain a fourth prediction loss corresponding to the current round of the fourth training according to the predicted emotion category obtained from the current round of the fourth training and the actual emotion category in the target emotion training sample corresponding to the current round of the fourth training; A second optimization module, configured to optimize the basic sentiment recognition model according to the fourth prediction loss corresponding to the fourth round of training; A seventh acquisition module, configured to stop training to obtain the target sentiment recognition model when the basic sentiment recognition model meets a fourth preset stop condition.

[0121] Regarding the voice synthesis device 600 in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.

[0122] The present disclosure also provides a computer-readable storage medium, on which computer program instructions are stored, and when the program instructions are executed by a processor, the steps of the voice synthesis method provided by the present disclosure are implemented.

[0123] Figure 7 It is a block diagram of an electronic device shown according to an exemplary embodiment. For example, the electronic device 700 may be a smart terminal device such as a mobile phone, a computer, a tablet, a vehicle-mounted computer, a robot, or a server.

[0124] Referring to Figure 7 , the electronic device 700 may include one or more of the following components: a first processing component 702, a first memory 704, a first power supply component 706, a multimedia component 708, an audio component 710, a first input / output interface 712, a sensor component 714, and a communication component 716.

[0125] The first processing component 702 generally controls the overall operation of the electronic device 700, such as operations associated with display, telephone calls, data communication, camera operations, and recording operations. The first processing component 702 may include one or more processors 720 to execute instructions to complete all or part of the steps of the above voice synthesis method. In addition, the first processing component 702 may include one or more modules to facilitate the interaction between the first processing component 702 and other components. For example, the first processing component 702 may include a multimedia module to facilitate the interaction between the multimedia component 708 and the first processing component 702.

[0126] The first memory 704 is configured to store various types of data to support the operation of the electronic device 700. Examples of such data include instructions for any application or method operating on the electronic device 700, contact data, phone book data, messages, pictures, videos, and the like. The first memory 704 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.

[0127] The first power supply component 706 provides power to various components of the electronic device 700. The first power supply component 706 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the electronic device 700.

[0128] The multimedia component 708 includes a screen that provides an output interface between the electronic device 700 and the user. In some embodiments, the screen can include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 708 includes a front camera and / or a rear camera. When the electronic device 700 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0129] The audio component 710 is configured to output and / or input audio signals. For example, the audio component 710 includes a microphone (MIC) that is configured to receive external audio signals when the electronic device 700 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the first memory 704 or transmitted via the communication component 716. In some embodiments, the audio component 710 further includes a speaker for outputting audio signals.

[0130] The first input / output interface 712 provides an interface between the first processing component 702 and a peripheral interface module, and the peripheral interface module can be a keyboard, a click wheel, buttons, and the like. These buttons can include, but are not limited to: a home button, a volume button, a start button, and a lock button.

[0131] The sensor assembly 714 includes one or more sensors for providing a status assessment of various aspects for the electronic device 700. For example, the sensor assembly 714 can detect the on / off state of the electronic device 700, the relative positioning of components, such as the display and keypad of the electronic device 700. The sensor assembly 714 can also detect a change in the position of the electronic device 700 or a component of the electronic device 700, the presence or absence of user contact with the electronic device 700, the orientation or acceleration / deceleration of the electronic device 700, and a change in the temperature of the electronic device 700. The sensor assembly 714 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 714 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 714 can also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0132] The communication component 716 is configured to facilitate communication between the electronic device 700 and other devices in a wired or wireless manner. The electronic device 700 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 716 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 716 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0133] In an exemplary embodiment, the electronic device 700 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above-described speech synthesis method.

[0134] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a first memory 704 including instructions, and the above instructions can be executed by a processor 720 of the electronic device 700 to complete the above-described speech synthesis method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0135] In another exemplary embodiment, a computer program product is also provided. The computer program product includes a computer program executable by a programmable device, and the computer program has a code portion for performing the above-described speech synthesis method when executed by the programmable device.

[0136] Those skilled in the art can also understand that the various illustrative logical blocks and steps listed in the embodiments of the present application can be implemented by electronic hardware, computer software, or a combination of both. Whether such a function is implemented by hardware or software depends on the specific application and the design requirements of the entire system. For each specific application, those skilled in the art can use various methods to implement the described function, but such implementation should not be construed as exceeding the scope of protection of the embodiments of the present application.

[0137] In the above detailed description, terms indicating directions or representing positional relationships such as "center", "upper", "lower", "left", "right", etc. are used. Since the components of the described device can be positioned in multiple different orientations, the direction terms can be used for illustrative purposes rather than restrictively. It should be understood that other aspects can be utilized and structural or logical changes can be made without departing from the concept of the present disclosure. Therefore, the following detailed description should not be construed in a limiting sense.

[0138] It should be understood that unless otherwise specifically stated, the features of some embodiments of the various present disclosures described herein can be combined with each other.

[0139] Although terms such as "first", "second", and "third" may be used herein to describe various components, parts, regions, layers, or sections, these components, parts, regions, layers, or sections are not limited to these terms. On the contrary, these terms are only used to distinguish one component, part, region, layer, or section from another. Therefore, without departing from the teachings of the various examples, the first component, part, region, layer, or section mentioned in the examples described herein can also be referred to as the second component, part, region, layer, or section. Additionally, the terms "first" and "second" are only used for descriptive purposes and should not be construed as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of such features. In the description herein, "a plurality of" means at least two, such as two, three, etc., unless otherwise specifically and explicitly defined.

[0140] In addition, as used herein, the term "exemplary" is used to mean serving as an example, instance, or illustration. Any aspect or design described herein as "exemplary" is not necessarily to be construed as advantageous over other aspects or designs. Instead, the use of the term exemplary is intended to present concepts in a concrete fashion. As used herein, the term "or" is intended to mean an inclusive "or" rather than an exclusive "or". That is, unless otherwise specified or clear from the context, "X applies A or B" is intended to mean any of the natural inclusive permutations. That is, if X applies A; X applies B; or X applies both A and B, then "X applies A or B" is satisfied under any of the foregoing instances. Additionally, unless otherwise specified or clear from the context that it is referring to the singular form, the articles "a" and "an" as used in this application and the appended claims are generally understood to mean "one or more".

[0141] Likewise, although the present disclosure has been shown and described with respect to one or more implementations, equivalent variations and modifications will occur to those skilled in the art upon reading and understanding the specification and the drawings. The present disclosure includes all such modifications and variations and is limited only by the scope of the claims. Specifically with respect to the various functions performed by the components described above (e.g., elements, resources, etc.), unless otherwise indicated, the terms used to describe such components are intended to correspond to any component (functionally equivalent) that performs the specific function of the described component, even if not structurally equivalent to the disclosed structure. Additionally, although a particular feature of the present disclosure may have been disclosed with respect to only one of several implementations, such a feature may, as may be desired and advantageous for any given or particular application, be combined with one or more other features of other implementations. Further, with respect to the terms "comprises", "comprising", "has", "having", "includes", or variants thereof as used in the detailed description or the claims, such terms are intended to be inclusive in a manner similar to the term "including".

[0142] Other embodiments of the present disclosure will be readily apparent to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include known common knowledge or conventional technical means in the technical field not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of the present disclosure are pointed out by the appended claims.

[0143] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes may be made without departing from its scope. The scope of the present disclosure is limited only by the appended claims.

Claims

1. A speech synthesis method, characterized in that: include: Process the text information and sentiment information corresponding to the target user to obtain a tag sequence; The target speech synthesis model is used to process the labeled sequence to obtain a target speech spectrum, wherein the target speech synthesis model is obtained by training a basic speech synthesis model based on a plurality of speech training samples, wherein the speech training samples include sample text information, sample emotion information and sample speech; According to the target speech spectrum, a target synthesized speech is obtained.

2. The speech synthesis method according to claim 1, characterized in that: The speech training sample also includes sample additional information; The processing of the text information and sentiment information corresponding to the target user to obtain a tag sequence includes: The text information, emotional information and additional features corresponding to the target user are processed to obtain a tag sequence, wherein the additional features include at least one of gender, language, age and target style.

3. The speech synthesis method according to claim 1, characterized in that: The method further comprises: Acquire a target sign language video including sign language movements; Extracting features from the target sign language video to obtain gesture features; The gesture features are processed by a target sign language recognition model to obtain the text information. The target sign language recognition model is obtained by training a basic sign language recognition model based on multiple sign language training samples. The sign language training samples include sample sign language videos and actual sign language texts.

4. The speech synthesis method according to claim 1, characterized in that: The method further comprises: Acquire a target expression image including a facial expression; Extracting features of the target expression image to obtain facial features; The facial features are processed by a target emotion recognition model to obtain the emotion information. The target emotion recognition model is obtained by training a basic emotion recognition model based on a plurality of emotion training samples. The emotion training samples include sample expression images and actual emotion categories.

5. The speech synthesis method according to claim 1, characterized in that: The method further comprises: Perform sentiment feature extraction on the text information to obtain the sentiment information corresponding to the text information.

6. The speech synthesis method according to claim 1, characterized in that: The method further comprises: Based on the preset corresponding relationship between the scene and the emotion, the target emotion corresponding to the target scene is determined to obtain the emotion information.

7. The speech synthesis method according to claim 1, characterized in that: The plurality of speech training samples include a plurality of first training samples corresponding to different timbres, and the target speech synthesis model is obtained by the following steps: The basic speech synthesis model is trained for multiple rounds of iterations using the multiple first training samples to obtain the target speech synthesis model.

8. The speech synthesis method according to claim 7, characterized in that: The plurality of speech training samples also include a plurality of second training samples corresponding to the target timbre; The step of performing multiple rounds of iterative training on the basic speech synthesis model through the multiple first training samples to obtain the target speech synthesis model includes: Performing multiple rounds of first training on the basic speech synthesis model through the multiple first training samples to obtain an original speech synthesis model; The target speech synthesis model is obtained by performing multiple rounds of second training on the original speech synthesis model through the multiple second training samples.

9. The speech synthesis method according to claim 8, characterized in that: The step of performing multiple rounds of first training on the basic speech synthesis model through the multiple first training samples to obtain an original speech synthesis model includes: Processing the multiple first training samples to obtain multiple first target training samples, each of which includes a sample label sequence and an actual speech spectrum, the sample label sequence is obtained based on the sample text information and the sample emotion information, and the actual speech spectrum is obtained based on the sample speech; Performing multiple rounds of first training on the basic speech synthesis model using the multiple first target training samples; After each round of first training, a first prediction loss corresponding to the first round of training is obtained according to the predicted speech spectrum obtained in the first round of training and the actual speech spectrum in the first target training sample corresponding to the first round of training; Optimizing the basic speech synthesis model according to the first prediction loss corresponding to the first training of this round; When the basic speech synthesis model meets the first preset stop condition, the training is stopped to obtain the original speech synthesis model.

10. The speech synthesis method according to claim 9, characterized in that: The first training sample also includes sample additional information, and the sample additional information includes at least one of gender, language, age, and target style. The sample tag sequence is obtained based on the sample text information, the sample emotion information, and the sample additional information.

11. The speech synthesis method according to claim 10, characterized in that: The method further comprises: Based on a preset correspondence between scenes and styles, the target style corresponding to the target scene is determined.

12. The speech synthesis method according to claim 8, characterized in that: The step of performing multiple rounds of second training on the original speech synthesis model through the multiple second training samples to obtain the target speech synthesis model includes: Processing the plurality of second training samples to obtain a plurality of second target training samples, each second target training sample comprising a sample label sequence and an actual speech spectrum, the sample label sequence being obtained based on the sample text information and the sample emotion information, and the actual speech spectrum being obtained based on the sample speech; Performing multiple rounds of second training on the original speech synthesis model using the multiple second target training samples; After each round of second training, a second prediction loss corresponding to the current round of second training is obtained according to the predicted speech spectrum obtained in the current round of second training and the actual speech spectrum in the second target training sample corresponding to the current round of second training; Optimizing the original speech synthesis model according to the second prediction loss corresponding to the second training in this round; When the original speech synthesis model meets the second preset stop condition, the training is stopped to obtain the target speech synthesis model.

13. The speech synthesis method according to any one of claims 2 to 12, characterized in that: The target sign language recognition model is obtained by the following steps: Processing the multiple sign language training samples to obtain multiple target sign language training samples, wherein the target sign language training samples include sample gesture features and actual sign language texts, and the sample gesture features are obtained by extracting features from the sample sign language video; Performing multiple rounds of third training on the basic sign language recognition model using the sample gesture features in the multiple target sign language training samples; After each round of third training, a third prediction loss corresponding to the third round of training is obtained according to the predicted sign language text obtained in the third round of training and the actual sign language text in the target sign language training sample corresponding to the third round of training; Optimizing the basic sign language recognition model according to the third prediction loss corresponding to the third training of this round; When the basic sign language recognition model satisfies a third preset stop condition, the training is stopped to obtain the target sign language recognition model.

14. The speech synthesis method according to claim 4, characterized in that: The target emotion recognition model is obtained by the following steps: Processing the multiple emotion training samples to obtain multiple target emotion training samples, the emotion training samples comprising sample facial features and actual emotion categories, the sample facial features being obtained by extracting features from the sample expression images; Performing multiple rounds of fourth training on the basic emotion recognition model by using the sample facial features in the multiple target emotion training samples; After each round of fourth training, a fourth prediction loss corresponding to the fourth round of training is obtained according to the predicted emotion category obtained in the fourth round of training and the actual emotion category in the target emotion training sample corresponding to the fourth round of training; Optimizing the basic emotion recognition model according to the fourth prediction loss corresponding to the fourth training of this round; When the basic emotion recognition model satisfies the fourth preset stop condition, the training is stopped to obtain the target emotion recognition model.

15. A speech synthesis device, characterized in that: include: A processing module is configured to process text information and sentiment information corresponding to a target user to obtain a tag sequence; A synthesis module is configured to process the tag sequence through a target speech synthesis model to obtain a target speech spectrum, wherein the target speech synthesis model is obtained by training a basic speech synthesis model based on a plurality of speech training samples, wherein the speech training samples include sample text information, sample emotion information and sample speech; The first acquisition module is configured to obtain the target synthesized speech according to the target speech spectrum.

16. An electronic device, characterized in that: include: processor; a memory for storing processor-executable instructions; Wherein, the processor is configured to implement the steps of the speech synthesis method described in any one of claims 1 to 14.

17. A computer-readable storage medium having computer program instructions stored thereon, characterized in that: When the computer program instructions are executed by a processor, the steps of the speech synthesis method described in any one of claims 1 to 14 are implemented.

18. A computer program product, characterized in that The invention comprises a computer program which, when executed by a processor, implements the steps of the speech synthesis method according to any one of claims 1 to 14.