Speech synthesis method and device, electronic equipment and storage medium

By acquiring a speech synthesis dataset and performing emotion tag token conversion and neural coding model training, the speech synthesis model was optimized, generating high-quality emotionally expressive speech data and solving the problem of inaccurate emotional expression in existing technologies.

CN121838720APending Publication Date: 2026-04-10PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-13
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing technologies, discrete emotion category labels cannot accurately and linearly control the emotion expression of speech synthesis, which affects the accuracy of emotion expression in speech synthesis.

Method used

By acquiring a speech synthesis dataset, token conversion and clustering of emotion indicator tags are performed to generate emotion tag tokens. A neural coding model is used for model training to optimize the target coding model. The speech text and emotion conditions are combined for encoding to generate high-quality emotionally expressive speech data.

Benefits of technology

It improves the accuracy of emotional expression in speech synthesis, generates high-quality, emotionally expressive speech data, and solves the problem that emotional category labels cannot accurately control the emotional expression in speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838720A_ABST
    Figure CN121838720A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a speech synthesis method and device, electronic equipment and a storage medium, belongs to the technical field of artificial intelligence, and is suitable for the fields of financial science and technology and medical science and technology. The method comprises the following steps: acquiring a speech synthesis data set; wherein the speech synthesis data set comprises a speech synthesis sample and an emotion indication label; performing token conversion on the emotion indication tag to obtain an emotion tag token; performing model training on a preset neural coding model based on the speech synthesis sample and the emotion label token to obtain a target coding model; encoding a pre-acquired target synthetic text through the target encoding model to obtain original encoding data; and performing speech synthesis based on the original coded data to obtain target speech data. According to the embodiment of the invention, the emotion expression accuracy of speech synthesis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and is applicable to the fields of financial technology and medical technology, and particularly to a speech synthesis method and apparatus, electronic device and storage medium. Background Technology

[0002] Speech synthesis is a technology that converts text content into natural speech. It can be applied in various scenarios, such as in the financial sector, where it is used to synthesize natural speech in scenarios like intelligent customer service and voice assistants. In the medical technology sector, it is used to synthesize natural speech in scenarios like intelligent customer service, medical guidance, and disease introduction.

[0003] Currently, speech synthesis mainly relies on discrete emotion category labels to control the emotional expression. However, in practical applications, it is impossible to precisely and linearly control the specific expression effects of emotions (such as the intensity of emotions, sense of control, or positivity), which affects the accuracy of emotional expression in speech synthesis.

[0004] Therefore, improving the accuracy of emotional expression in speech synthesis has become an urgent technical problem to be solved. Summary of the Invention

[0005] The main objective of this application is to propose a speech synthesis method, apparatus, electronic device, and storage medium, which aims to solve the technical problem that the emotional expression of speech synthesis cannot be accurately controlled due to the inability to accurately control the emotional expression of speech synthesis by emotional category labels, thereby improving the accuracy of emotional expression in speech synthesis.

[0006] To achieve the above objectives, a first aspect of this application proposes a speech synthesis method, the method comprising: Obtain the speech synthesis dataset; the speech synthesis dataset includes speech synthesis samples and emotion indicator labels; The sentiment indicator tag is tokenized to obtain a sentiment tag token; The preset neural coding model is trained based on the speech synthesis samples and the emotion tag tokens to obtain the target coding model; The target synthetic text is encoded using the target encoding model to obtain the original encoded data; Speech synthesis is performed based on the original encoded data to obtain the target speech data.

[0007] In some embodiments, the step of tokenizing the sentiment indicator tag to obtain a sentiment tag token includes: The sentiment indicator tags are clustered to obtain multiple tag clusters and the three-dimensional center point corresponding to each tag cluster; The target boundary is obtained by performing boundary calculations based on the three-dimensional center points. Based on the target boundary, the sentiment indicator tags are binned to obtain binned tag data; The sentiment indicator tags in the binning tag data are used to perform token calculations to obtain the sentiment tag tokens.

[0008] In some embodiments, the step of calculating the target boundary based on the three-dimensional center point includes: The three-dimensional center points are sorted to obtain a three-dimensional center sequence; A three-dimensional center pair is constructed based on each pair of adjacent three-dimensional center points in the three-dimensional center sequence; The target boundary is obtained by performing boundary calculations based on the three-dimensional center pair.

[0009] In some embodiments, the three-dimensional center pair includes a first three-dimensional center and a second three-dimensional center; the boundary calculation based on the three-dimensional center pair to obtain the target boundary includes: Boundary calculations are performed based on the first three-dimensional center and the second three-dimensional center to obtain the intermediate boundary; Obtain the number of labels in the label cluster corresponding to the first three-dimensional center to get the first label count; Obtain the number of labels in the label cluster corresponding to the second three-dimensional center to get the second label count; Boundary calculations are performed based on the first number of labels, the second number of labels, the first 3D center, and the second 3D center to obtain a weighted boundary. The target boundary is obtained by performing aggregation calculations on the intermediate boundary and the weighted boundary.

[0010] In some embodiments, training a preset neural coding model based on the speech synthesis samples and the emotion tag tokens to obtain a target coding model includes: The speech synthesis sample and the emotion tag token are encoded using the neural coding model to obtain sample encoded data; Speech synthesis is performed based on the sample encoded data to obtain sample speech data; Loss calculation is performed based on the speech synthesis samples and the sample speech data to obtain speech synthesis loss data; The neural coding model is optimized based on the speech synthesis loss data to obtain an initial coding model. The sample speech data is grouped to obtain sample speech pairs; Emotional loss is calculated based on the sample speech pairs to obtain emotional preference loss data; The initial encoding model is optimized based on the sentiment preference loss data to obtain the target encoding model.

[0011] In some embodiments, the target synthesized text includes: speech text, emotional conditions, and speaker information; the process of encoding the pre-acquired target synthesized text using the target coding model to obtain raw encoded data includes: The speech text is text encoded to obtain text encoding features; The emotional conditions are subjected to emotional encoding to obtain emotional encoding features; The original encoded data is obtained by neurally encoding the text encoding features, the emotion encoding features, and the speaker information using the target encoding model.

[0012] In some embodiments, the step of neurally encoding the text encoding features, the emotion encoding features, and the speaker information using the target encoding model to obtain the original encoded data includes: Obtain the label sentiment features corresponding to the sentiment conditions; Attention is calculated based on the sentiment encoding features and the labeled sentiment features to obtain fused sentiment features; The original encoded data is obtained by performing speech encoding on the fused emotional features, the text encoding features, and the speaker information using the target encoding model.

[0013] To achieve the above objectives, a second aspect of this application provides a speech synthesis apparatus, the apparatus comprising: The data acquisition module is used to acquire the speech synthesis dataset; the speech synthesis dataset includes speech synthesis samples and emotion indicator labels; The token conversion module is used to convert the sentiment indicator tag into a sentiment tag token. The model training module is used to train a preset neural coding model based on the speech synthesis samples and the emotion tag tokens to obtain the target coding model. The encoding module is used to encode the pre-acquired target synthetic text using the target encoding model to obtain the original encoded data; The speech data synthesis module is used to perform speech synthesis based on the original encoded data to obtain target speech data.

[0014] To achieve the above objectives, a third aspect of the present application provides an electronic device, the electronic device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in the first aspect.

[0015] To achieve the above objectives, a fourth aspect of the present application provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in the first aspect.

[0016] The speech synthesis method, apparatus, electronic device, and storage medium proposed in this application acquire a dataset containing speech synthesis samples and emotion indicator tags to provide data support for subsequent model training, ensuring data diversity and integrity. The emotion indicator tags are then tokenized to obtain emotion indicator tokens, enabling more accurate extraction of their data features. Next, a pre-defined neural coding model is trained based on the speech synthesis samples and emotion indicator tokens, allowing the model to accurately capture the correlation between speech and emotion, resulting in a target coding model that enhances the model's ability to encode emotional speech. Furthermore, the target coding model encodes the pre-acquired target synthesized text to obtain raw coded data, accurately reflecting the relationship between text emotion and speech features. Finally, speech synthesis is performed based on the raw coded data to obtain high-quality, emotionally expressive target speech data. This solves the technical problem that emotion category tags cannot accurately control the emotional expression of speech synthesis, improving the accuracy of emotional expression in speech synthesis. Attached Figure Description

[0017] Figure 1 This is a flowchart of the speech synthesis method provided in the embodiments of this application; Figure 2 yes Figure 1 The flowchart of step S102 in the document; Figure 3 yes Figure 2 The flowchart of step S202 in the document; Figure 4 yes Figure 3 The flowchart of step S303 in the process; Figure 5 yes Figure 1 The flowchart of step S103 in the process; Figure 6 yes Figure 1 The flowchart of step S104 in the process; Figure 7 yes Figure 6 The flowchart of step S603 in the process; Figure 8 This is a schematic diagram of the speech synthesis device provided in the embodiments of this application; Figure 9 This is a schematic diagram of the hardware structure of the electronic device provided in the embodiments of this application. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.

[0019] It should be noted that although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than the module division in the device or the order in the flowchart. The terms "first," "second," etc., in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence.

[0020] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.

[0021] First, let's analyze some of the terms used in this application: Artificial intelligence (AI) is a new branch of computer science that studies, develops, and applies theories, methods, technologies, and systems to simulate, extend, and expand human intelligence. It aims to understand the essence of intelligence and produce intelligent machines that can react in a way similar to human intelligence. Research in this field includes robotics, speech recognition, image recognition, natural language processing, and expert systems. AI can simulate the information processes of human consciousness and thought. Furthermore, AI utilizes digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceiving the environment, acquiring knowledge, and using that knowledge to achieve optimal results.

[0022] Natural Language Processing (NLP): NLP uses computers to process, understand, and utilize human language (such as Chinese and English). It is a branch of artificial intelligence and an interdisciplinary field of computer science and linguistics, often referred to as computational linguistics. NLP includes syntactic analysis, semantic analysis, and discourse understanding. It is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, intent recognition, information extraction and filtering, text classification and clustering, sentiment analysis, and opinion mining. It involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research, and linguistic research related to language computation.

[0023] Speech synthesis, also known as text-to-speech (TTS), is a technology that converts text content into natural speech. The core process of speech synthesis includes text normalization, syntactic analysis, and prosodic generation. It generates speech files containing numbers and proper nouns through dictionary matching, spell-and-pronunciation rules, or deep learning models (such as WaveNet). It is widely used in various scenarios such as voice interaction in smart devices and accessibility services. For example, in the financial sector, it is used to synthesize natural speech in scenarios such as intelligent customer service and voice assistants. In the medical technology field, it is used to synthesize natural speech in scenarios such as intelligent customer service, medical guidance, and disease introduction.

[0024] In speech synthesis, VAD tags are a three-dimensional continuous coordinate system used in psychology to quantify emotional states, including Valence, Arousal, and Dominance. Valence represents the degree of emotion from negative to positive, such as sadness (-1) to pleasure (+1); Arousal reflects the intensity of emotion from calm to excitement, such as drowsiness (0) to excitement (+1); and Dominance reflects the sense of control from passive to active, such as helplessness (-1) to dominance (+1). VAD tags, through continuous values ​​rather than discrete labels, can accurately describe complex emotions (e.g., "grief and indignation" is characterized by low valence, high arousal, and moderate dominance), providing adjustable acoustic parameters for speech synthesis systems and generating nuanced emotional expressions that match the context.

[0025] The Direct Preference Optimization Loss (DPO) is a loss function that optimizes a policy model directly using human preference data without requiring an explicit reward model.

[0026] The core of the DPO loss is based on the Bradley-Terry model, which is trained using pairwise preference data (preferred and second-best responses) and aims to maximize the probability ratio of the preferred response to the second-best response. This loss updates the model parameters through gradient descent, making the policy more inclined to generate outputs that conform to human preferences.

[0027] Currently, speech synthesis mainly relies on discrete emotion category labels to control the emotional expression. However, in practical applications, it is impossible to precisely and linearly control the specific expression effects of emotions (such as the intensity of emotions, sense of control, or positivity), which affects the accuracy of emotional expression in speech synthesis.

[0028] Based on this, embodiments of this application provide a speech synthesis method and apparatus, electronic device and storage medium, which aim to solve the technical problem that the emotional expression of speech synthesis cannot be accurately controlled due to the inability of emotion category labels, and improve the accuracy of emotional expression in speech synthesis.

[0029] The speech synthesis method, apparatus, electronic device, and storage medium provided in this application are specifically described through the following embodiments. First, the speech synthesis method in this application is described.

[0030] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0031] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0032] The speech synthesis method provided in this application relates to the field of artificial intelligence technology. The speech synthesis method provided in this application can be applied to a terminal, a server, or software running on either a terminal or a server. In some embodiments, the terminal can be a smartphone, tablet, laptop, desktop computer, etc.; the server can be configured as an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application implementing the speech synthesis method, but is not limited to the above forms.

[0033] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. Examples include: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, and distributed computing environments including any of the above systems or devices. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform specific tasks or implement specific abstract data types. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0034] It should be noted that in all specific embodiments of this application, when processing data related to user identity or characteristics, such as user information, user behavior data, user historical data, and user location information, user permission or consent is obtained first. Furthermore, the collection, use, and processing of this data comply with relevant laws, regulations, and standards. In addition, when embodiments of this application require access to sensitive personal information of users, separate permission or consent from the user is obtained through pop-ups or redirection to confirmation pages. Only after obtaining the user's separate permission or consent is the necessary user-related data required for the proper functioning of these embodiments acquired.

[0035] Figure 1 This is an optional flowchart of the speech synthesis method provided in the embodiments of this application. Figure 1 The method may include, but is not limited to, steps S101 to S105.

[0036] Step S101: Obtain the speech synthesis dataset; wherein, the speech synthesis dataset includes speech synthesis samples and emotion indicator labels; Step S102: Perform token conversion on the sentiment indicator tag to obtain a sentiment tag token; Step S103: Train the preset neural coding model based on speech synthesis samples and emotion tag tokens to obtain the target coding model; Step S104: Encode the pre-acquired target synthetic text using the target encoding model to obtain the original encoded data; Step S105: Speech synthesis is performed based on the original encoded data to obtain the target speech data.

[0037] Steps S101 to S105, as illustrated in this embodiment, acquire a dataset containing speech synthesis samples and emotion indicator tags to provide data support for subsequent model training, ensuring data diversity and integrity. The emotion indicator tags are then tokenized to obtain emotion label tokens, enabling more accurate extraction of their data features. Next, a pre-defined neural coding model is trained based on the speech synthesis samples and emotion label tokens, allowing the model to accurately capture the correlation between speech and emotion, resulting in a target coding model that enhances the model's ability to encode emotional speech. Furthermore, the target coding model encodes the pre-acquired target synthesized text to obtain raw coded data, accurately reflecting the relationship between text emotion and speech features. Finally, speech synthesis is performed based on the raw coded data to obtain high-quality, emotionally expressive target speech data. This solves the technical problem of inaccurate control of emotional expression in speech synthesis due to emotion category tags, improving the accuracy of emotional expression in speech synthesis.

[0038] In step S101 of some embodiments, the speech synthesis dataset is a pre-collected large-scale emotional speech dataset used for training and evaluating neural coding models. The speech synthesis dataset includes speech synthesis samples and corresponding emotion indicator labels for the speech synthesis samples. This provides a comprehensive and accurate data foundation for subsequent model training, enabling the model to learn diverse speech features. Detailed emotion labels allow the model to understand the correspondence between different emotions and speech features, thereby more accurately expressing emotions during speech synthesis.

[0039] Specifically, the speech synthesis samples include specific sample speech text and sample speech audio, covering various speech features such as different speakers, different speaking speeds, and different intonations.

[0040] The emotion indicator label is the VAD label, used to describe the emotional state of the speech in the speech synthesis sample. The emotion indicator label includes: Arousal label, Dominance label, and Valence label. Specifically, the Arousal label reflects the degree of emotional activation, such as excitement or calmness; the Dominance label reflects the sense of control over the situation, such as assertiveness or compliance; and the Valence label indicates the positive or negative nature of the emotion, such as positive or negative.

[0041] Please see Figure 2 In some embodiments, step S102 may include, but is not limited to, steps S201 to S204: Step S201: Cluster the sentiment indicator tags to obtain multiple tag clusters and the three-dimensional center point corresponding to each tag cluster. Step S202: Calculate the boundary based on the three-dimensional center point to obtain the target boundary; Step S203: Binning the sentiment indicator labels based on the target boundary to obtain binned label data; Step S204: Token calculation is performed on the sentiment indicator tags in the binning tag data to obtain sentiment tag tokens.

[0042] Steps S201 to S204, as illustrated in this embodiment, involve clustering the sentiment indicator tags to obtain multiple tag clusters and the corresponding three-dimensional centroids for each cluster. This makes the tag distribution more regular and facilitates subsequent processing. Next, boundary calculations are performed based on the three-dimensional centroids to obtain target boundaries, accurately defining the ranges of different tag clusters and providing clear limits for classification. Furthermore, the sentiment indicator tags are binned based on the target boundaries to obtain binned tag data. Token calculations are then performed on the sentiment indicator tags within the binned tag data to obtain sentiment tag tokens. This allows for more efficient extraction and representation of sentiment features, improving the accuracy and efficiency of sentiment analysis and contributing to enhancing the model's ability to learn about sentiment.

[0043] In some embodiments, step S201 includes, but is not limited to, the following steps: The sentiment indicator labels are normalized to obtain normalized labels; The normalized labels are clustered to obtain multiple label clusters and the three-dimensional centroid of each label cluster.

[0044] For each sentiment indicator label (arousal label, dominance label, and valence label), the sentiment indicator label is normalized. This is usually done by mapping the data to intervals such as [0,1] or [-1,1] to make data with different dimensions and value ranges comparable and consistent, resulting in a normalized label. Specifically, methods such as min-max normalization and Z-score normalization can be used for normalization, but are not limited to these.

[0045] Next, clustering algorithms (such as K-Means) are used to cluster the normalized labels, dividing them into different clusters based on the similarity between the labels, and calculating the centroid of each cluster to obtain three-dimensional centroids. This reduces the number of labels and complexity, and the three-dimensional centroids provide a visual understanding of the typical characteristics of each cluster, facilitating subsequent processing and analysis.

[0046] The 3D centroid is a point in 3D space that represents the center of each label cluster. Its coordinates are determined by the average value of the labels in the cluster across the three feature dimensions.

[0047] For example: for arousal labels, clustering can group excitement into one cluster and calmness into another; for dominance labels, dominance can be grouped into one cluster and compliance into another; for valence labels, positivity can be grouped into one cluster and negativity into another.

[0048] Please see Figure 3 In some embodiments, step S202 may include, but is not limited to, steps S301 to S303: Step S301: Sort the three-dimensional center points to obtain a three-dimensional center sequence; Step S302: Construct a three-dimensional center pair based on each pair of adjacent three-dimensional center points in the three-dimensional center sequence; Step S303: Perform boundary calculation based on the three-dimensional center pair to obtain the target boundary.

[0049] Steps S301 to S303, as illustrated in this embodiment, involve sorting the three-dimensional center points according to their numerical values ​​to obtain a three-dimensional center sequence. Next, a three-dimensional center pair is constructed based on each pair of adjacent three-dimensional center points in the three-dimensional center sequence to clarify the relationships. Finally, boundary calculation is performed based on the three-dimensional center pairs, which accurately determines the target boundary of each label cluster, effectively improving the accuracy and efficiency of boundary calculation and enhancing the ability to identify and locate target boundaries.

[0050] In step S301 of some embodiments, the three-dimensional center points are sorted in descending order of their numerical values ​​to obtain a three-dimensional center point sequence.

[0051] In step S301 of some other embodiments, the three-dimensional center points can also be arranged in order of coordinate size.

[0052] It is understandable that the 3D centroids are the center points of each label cluster. After being sorted in descending order, each pair of adjacent 3D centroids represents an adjacent label cluster. Therefore, by constructing 3D centroid pairs for each pair of adjacent 3D centroids in the 3D centroid sequence, the boundaries of adjacent label clusters can be accurately calculated, clarifying the boundaries of adjacent label clusters, thereby more accurately distinguishing various emotions and improving the accuracy of emotion recognition.

[0053] In step S302 of some embodiments, adjacent 3D center points are two 3D center points that are physically next to each other in the 3D center sequence, and their sequence numbers differ by 1. For example, 3D center point number 1 and 3D center point number 2 are considered adjacent 3D center points, and two adjacent 3D center points are combined to form a 3D center pair. If the 3D center sequence has 5 3D center points, there are 4 pairs of 3D center points; if the 3D center sequence has 10 3D center points, there are 9 pairs of 3D center points.

[0054] Specifically, the two three-dimensional center points in the three-dimensional center pair are named the first three-dimensional center and the second three-dimensional center, respectively.

[0055] Please see Figure 4 In some embodiments, step S303 may include, but is not limited to, steps S401 to S405: Step S401: Calculate the boundary based on the first three-dimensional center and the second three-dimensional center to obtain the intermediate boundary; Step S402: Obtain the number of labels in the label cluster corresponding to the first three-dimensional center to obtain the first label count; Step S403: Obtain the number of labels in the label cluster corresponding to the second three-dimensional center to obtain the second label count; Step S404: Based on the number of first labels, the number of second labels, the first 3D center, and the second 3D center, perform boundary calculation to obtain the weighted boundary; Step S405: Perform aggregation calculation on the intermediate boundary and the weighted boundary to obtain the target boundary.

[0056] Steps S401 to S405, as illustrated in this embodiment, involve calculating the intermediate boundary based on the first and second 3D centers to obtain the intermediate boundary. Next, the number of labels in the label clusters corresponding to the first and second 3D centers is obtained, resulting in the first and second label counts. Boundary calculations are then performed based on the first and second label counts, the first and second 3D centers, to obtain the weighted boundary. Finally, the intermediate boundary and the weighted boundary are aggregated to obtain the target boundary. This process more accurately and comprehensively defines the target range, improves the accuracy and rationality of boundary determination, and enhances the learning ability for emotion recognition.

[0057] In step S401 of some embodiments, the values ​​of the first three-dimensional center and the second three-dimensional center are averaged to obtain the intermediate boundary. It can be understood that the intermediate boundary is the midpoint between the first three-dimensional center and the second three-dimensional center. For the specific calculation process, please refer to formula (1): (1); in, The middle boundary, in the subscript Indicates wakefulness tag, Indicates the first For three-dimensional center pairs, As the first three-dimensional center, It serves as the second three-dimensional center.

[0058] Additionally, if it's a dominance label, then the subscript... Replace with If it is a valence label, then the subscript... Replace with .

[0059] In step S404 of some embodiments, boundary calculation is performed based on the number of first labels, the number of second labels, the first three-dimensional center, and the second three-dimensional center to obtain a weighted boundary. The specific calculation process is shown in formula (2): (2); in, For weighted boundaries, This represents the number of labels in the label cluster corresponding to the first 3D center, i.e., the number of the first labels. This refers to the number of labels in the label cluster corresponding to the second three-dimensional center, i.e., the number of second labels.

[0060] In step S405 of some embodiments, the intermediate boundary and weighted boundary are aggregated to obtain the target boundary, and the number of target boundaries is the same as the number of 3D center pairs. If the 3D center sequence has 5 3D center points, there are 4 target boundaries; if the 3D center sequence has 10 3D center points, there are 9 target boundaries. For the specific calculation process, please refer to formula (3): (3); in, For the target boundary, It is an indicator function, in In the case of , take 1, if Take 0.

[0061] Specifically, The variance ratio is used to characterize the difference in distribution shape between the label clusters corresponding to the first 3D center and the label clusters corresponding to the second 3D center. The variance ratio threshold is preset. For the specific calculation process of the variance ratio, please refer to formula (4): (4); in, It is the variance of the label cluster corresponding to the first three-dimensional center; It is the variance of the label clusters corresponding to the second three-dimensional center.

[0062] In one embodiment, a preset variance ratio threshold can be used. Set to 2, when This indicates that one cluster is very concentrated while the other is very dispersed, or that the two clusters have vastly different degrees of dispersion.

[0063] In step S203 of some embodiments, the sentiment indicator tags are binned according to the target boundary to obtain multiple intervals. Each interval corresponds to a binned tag data, and each binned tag data includes at least one sentiment indicator tag. Next, for any sentiment indicator tag in a certain bin label data, a token is calculated based on the target boundary to obtain the sentiment tag token. The specific calculation process is shown in formula (5): (5); in, For emotional tag tokens, This represents the value of a sentiment indicator label within the binning label data.

[0064] In some embodiments, for the arousal level tag, the specific implementations shown in steps S201 to S204, S301 to S303, and S401 to S405 above are performed to obtain the emotion tag token corresponding to the arousal level tag. ; For the dominance tag, perform the specific implementation methods shown in steps S201 to S204, S301 to S303, and S401 to S405 above to obtain the sentiment tag token corresponding to the dominance tag. ; For the valence tag, perform the specific implementation methods shown in steps S201 to S204, S301 to S303, and S401 to S405 above to obtain the sentiment tag token corresponding to the valence tag. ; Finally, the emotional tag tokens corresponding to the arousal level tags will be... Dominance tag corresponding to sentiment tag token Emotional tag tokens corresponding to valence Combine them to obtain the total emotional tag tokens. This serves as the sentiment tag token corresponding to the speech synthesis sample. The sentiment tag token... .

[0065] Please see Figure 5 In some embodiments, step S103 may include, but is not limited to, steps S501 to S507: Step S501: Encode the speech synthesis samples and emotion tag tokens using a neural coding model to obtain sample coding data; Step S502: Speech synthesis is performed based on the sample encoded data to obtain sample speech data; Step S503: Based on the speech synthesis samples and sample speech data, perform loss calculation to obtain speech synthesis loss data; Step S504: Optimize the neural coding model based on the speech synthesis loss data to obtain the initial coding model; Step S505: Group the sample speech data to obtain sample speech pairs; Step S506: Calculate the emotion loss based on the sample speech pairs to obtain the emotion preference loss data; Step S507: Optimize the initial coding model based on the sentiment preference loss data to obtain the target coding model.

[0066] Steps S501 to S507 as illustrated in this embodiment involve: First, encoding the speech synthesis samples and emotion tag tokens using a neural coding model to obtain sample coding data; then, performing speech synthesis based on the sample coding data to obtain sample speech data; calculating the loss based on the speech synthesis samples and sample speech data to obtain speech synthesis loss data, which measures the difference between the synthesized speech and the expected speech; finally, optimizing the neural coding model based on the speech synthesis loss data to obtain an initial coding model, thereby improving the quality of speech synthesis. Next, grouping the sample speech data to obtain sample speech pairs; calculating the emotion loss based on the sample speech pairs to obtain emotion preference loss data, which identifies deficiencies in speech emotion; and finally, optimizing the initial coding model based on the emotion preference loss data to obtain a target coding model, which not only optimizes the speech synthesis effect but also enables the model to better capture and express emotions, comprehensively improving the quality of speech synthesis and the accuracy of emotion expression.

[0067] In step S501 of some embodiments, the speech synthesis sample and the emotion tag token are used as input and fed into the neural coding model. The model extracts and integrates the key features through the calculation and transformation of the multi-layer neural network, and outputs the sample coding data, which facilitates subsequent speech synthesis and emotion analysis processing and can better preserve the original features and emotional information of the speech.

[0068] In step S502 of some embodiments, the obtained sample encoded data is input into a vocoder (such as HiFi-GAN, WaveNet model, etc.), and a speech waveform is generated based on the feature information in the sample encoded data to obtain sample speech data.

[0069] In some embodiments, the number of sample speech data is at least two.

[0070] In step S503 of some embodiments, loss functions such as mean squared error loss function and cross-entropy loss function can be used to calculate the loss of sample speech audio and sample speech data in speech synthesis samples, thereby measuring the difference between sample speech audio and sample speech data and obtaining speech synthesis loss data.

[0071] Next, the gradient of the model parameters of the neural coding model is calculated based on the speech synthesis loss data using the backpropagation algorithm. Then, the parameters of the neural coding model are updated using an optimization algorithm (such as stochastic gradient descent SGD). After multiple iterations, the initial coding model is obtained, which can improve the neural coding model's ability to extract and represent speech features, enabling the generated sample coding data to be better used for speech synthesis and improving the quality of speech synthesis.

[0072] Furthermore, by manually evaluating or using a trained preference predictor, the sample speech data are compared pairwise to form sample speech pairs. The data in each sample speech pair comes from a single speech synthesis sample, which facilitates the comparison of emotional differences between different speech samples, thereby better analyzing the model's performance in emotional expression.

[0073] In step S506 of some embodiments, the DPO loss function is calculated based on two sample speech data from the sample speech pair to obtain emotion preference data, quantifying the model's performance in emotion expression and providing a basis for further model optimization. By introducing DPO preference learning, the model can better capture and express emotional information in speech. The generated speech, while expressing strong emotions, can significantly improve naturalness and auditory comfort, effectively solving the industry challenge of the "intensity-quality" trade-off. The synthesized speech is more in line with human subjective preferences.

[0074] Furthermore, DPO preference learning, by learning human relative judgments of paired samples rather than absolute label values, can reduce the noise impact caused by the subjectivity of the labeler in the original VAD label to a certain extent, enabling the model to learn more robust emotional expression patterns.

[0075] Finally, the initial coding model is optimized based on the emotion preference loss data to obtain the target coding model, which further improves the neural coding model's ability in emotion expression. This enables the speech generated by the model to not only be accurately synthesized but also to better convey the corresponding emotional information, thereby improving the overall quality of speech synthesis.

[0076] In step S104 of some embodiments, the target synthesized text includes: speech text, emotional conditions, and speaker information.

[0077] Speech text is text content presented in text form for speech synthesis. Speech text contains semantic information about the speech to be generated, such as a dialogue or a news article.

[0078] Emotional condition is an identifier or information used to describe the emotional state that speech is intended to express. Emotional condition can be any one of VAD tags, discrete emotion tags, or emotional description text.

[0079] Speaker information is information collected in advance that includes the speaker's identity, gender, age, accent, and other characteristics. This information affects the timbre and intonation of the speech.

[0080] Please see Figure 6 In some embodiments, step S104 may include, but is not limited to, steps S601 to S603: Step S601: Perform text encoding on the speech text to obtain text encoding features; Step S602: Perform emotion encoding on the emotion conditions to obtain emotion encoding features; Step S603: The text encoding features, sentiment encoding features and speaker information are neurally encoded using the target encoding model to obtain the raw encoded data.

[0081] Steps S601 to S603, as shown in the embodiments of this application, involve text encoding of the speech text to convert it into processable text encoding features; performing emotion encoding on the emotion conditions to obtain emotion encoding features and clarifying the emotion information; and then performing neural encoding on the text encoding features, emotion encoding features, and speaker information through a target encoding model to obtain the original encoded data. By integrating multiple factors, the encoded data becomes more comprehensive and richer, which can better guide subsequent speech synthesis and improve the accuracy of synthesized speech in terms of content and emotion expression.

[0082] In step S601 of some embodiments, a pre-trained text encoding model (such as Transformer, BERT, etc.) is used to convert the input speech text into vector representations character by character or word by word. These vectors are combined to form text encoding features, which preserve the key information of the speech text, such as vocabulary, grammatical structure, semantic meaning, etc., and provide an accurate semantic basis for subsequent speech synthesis.

[0083] In step S602 of some embodiments, if the emotional condition is a VAD tag, the arousal tag, dominance tag, and valence tag are calculated using the target boundary and formula (5) obtained in step S202 above, respectively, to obtain the emotional tag token corresponding to the arousal tag. Dominance tag corresponding to sentiment tag token Emotional tag tokens corresponding to valence And combine them to obtain the total emotional tag tokens. Furthermore, the sentiment tag token is used as a sentiment encoding feature.

[0084] If the emotional condition is a discrete emotional label, the discrete emotional label is encoded in a conventional way to obtain the emotional coding feature.

[0085] If the sentiment condition is sentiment description text, a pre-trained RoBERTa model (based on the RoBERTa architecture, supplemented with softmax and normalization layers, and trained using MSE and center distance loss) is used to predict the sentiment description text, generating pseudo-VAD tags, and then generating sentiment tag tokens based on the pseudo-VAD tags. The emotion encoding features were obtained, which will not be elaborated further.

[0086] In step S603 of some embodiments, if the emotion condition is a discrete emotion label, the text encoding features, emotion encoding features and speaker information are directly neurally encoded through the target encoding model. By comprehensively considering the text, emotion and speaker information, more comprehensive and accurate original encoding data is generated, so that the synthesized speech can better meet expectations in terms of semantics, emotion and timbre, thereby improving the quality and emotional accuracy of speech synthesis.

[0087] Please see Figure 7 In some embodiments, if the sentiment condition is a VAD tag or sentiment description text (pseudo-VAD tag), step S603 may include, but is not limited to, steps S701 to S703: Step S701: Obtain the label sentiment features corresponding to the sentiment conditions; Step S702: Attention is calculated based on sentiment encoding features and labeled sentiment features to obtain fused sentiment features; Step S703: The speech is encoded by fusing emotion features, text encoding features and speaker information using a target encoding model to obtain the original encoded data.

[0088] Steps S701 to S703, as illustrated in this embodiment, involve acquiring the labeled emotional features corresponding to the emotional conditions, and performing attention calculations based on the emotional coding features and the labeled emotional features to obtain fused emotional features, thereby enhancing emotional expression. Finally, the fused emotional features, text coding features, and speaker information are encoded using a target coding model. By comprehensively considering text, emotion, and speaker information, more comprehensive and accurate raw coding data is generated, enabling the subsequently synthesized speech to better meet expectations in terms of semantics, emotion, and timbre, thus improving the quality and emotional accuracy of speech synthesis.

[0089] In step S701 of some embodiments, the sentiment label features are obtained as follows: first, it is determined which of the label clusters obtained in step S201 above the VAD label / pseudo-VAD label belongs to, and the target cluster is obtained. Then, the sentiment features corresponding to the target cluster are obtained as the label sentiment features.

[0090] In step S702 of some embodiments, attention is calculated using the label sentiment feature as the query feature and the sentiment encoding feature as the key feature and value feature to obtain the attention sentiment feature. Then, a gating algorithm can be used to calculate the attention sentiment feature to obtain the fused sentiment feature. The fused sentiment feature can enhance the expression of emotion and improve the accuracy and richness of the sentiment feature.

[0091] In step S703 of some embodiments, after obtaining the fused emotional features, the fused emotional features, text coding features and speaker information are encoded into speech using a target coding model. By comprehensively considering the text, emotion and speaker information, more comprehensive and accurate original coding data is generated, so that the subsequently synthesized speech can better meet expectations in terms of semantics, emotion and timbre, thereby improving the quality and emotional accuracy of speech synthesis.

[0092] For example, in a speech synthesis system, the input text "The weather is so nice today" is combined with the emotion feature "pleasant" and the speaker information is the timbre of a young woman. After passing through the target encoding model, the model outputs the raw encoded data, and then the speech synthesizer generates speech based on the raw encoded data, which sounds like a young woman saying "The weather is so nice today" in a pleasant tone.

[0093] In step S105 of some embodiments, the original encoded data is converted into actual speech waveforms using a vocoder (such as HiFi-GAN, WaveNet model, etc.), thereby generating target speech data. The target speech data not only conforms to the text content but also has appropriate emotional expression, improving the accuracy of emotional expression in speech synthesis and making the synthesized speech closer to real human speech.

[0094] The speech synthesis method provided in this application can be applied not only to fintech and medical technology scenarios, but also to smart tourism, online education, e-commerce retail and other scenarios. It uses speech synthesis methods to synthesize speech with accurate emotional expression and improves the naturalness of synthesized speech.

[0095] Please see Figure 8 This application also provides a speech synthesis apparatus that can implement the above-described speech synthesis method. The apparatus includes: The data acquisition module 801 is used to acquire the speech synthesis dataset; wherein, the speech synthesis dataset includes speech synthesis samples and emotion indicator labels; The token conversion module 802 is used to convert the sentiment indicator tag into a sentiment tag token; The model training module 803 is used to train a preset neural coding model based on speech synthesis samples and emotion tag tokens to obtain the target coding model. The encoding module 804 is used to encode the pre-acquired target synthetic text using the target encoding model to obtain the original encoded data; The speech data synthesis module 805 is used to synthesize speech based on the original encoded data to obtain the target speech data.

[0096] The specific implementation of this speech synthesis device is basically the same as the specific implementation of the speech synthesis method described above, and will not be repeated here.

[0097] This application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the above-described speech synthesis method. This electronic device can be any smart terminal, including tablet computers, in-vehicle computers, etc.

[0098] Please see Figure 9 , Figure 9 The hardware structure of an electronic device according to another embodiment is illustrated. The electronic device includes: The processor 901 can be implemented using a general-purpose CPU (Central Processing Unit), microprocessor, application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided in the embodiments of this application. The memory 902 can be implemented as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store the operating system and other application programs. When the technical solutions provided in the embodiments of this specification are implemented through software or firmware, the relevant program code is stored in the memory 902 and is called and executed by the processor 901 using the speech synthesis method of the embodiments of this application. The input / output interface 903 is used to implement information input and output; The communication interface 904 is used to enable communication and interaction between this device and other devices. Communication can be achieved through wired means (such as USB, Ethernet cable, etc.) or wireless means (such as mobile network, WIFI, Bluetooth, etc.). Bus 905 transmits information between various components of the device (e.g., processor 901, memory 902, input / output interface 903, and communication interface 904); The processor 901, memory 902, input / output interface 903, and communication interface 904 are connected to each other within the device via bus 905.

[0099] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described speech synthesis method.

[0100] Memory, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs and non-transitory computer-executable programs. Furthermore, memory may include high-speed random access memory, and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory may optionally include memory remotely located relative to the processor, and these remote memories can be connected to the processor via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0101] The speech synthesis method, apparatus, electronic device, and storage medium provided in this application acquire a dataset containing speech synthesis samples and emotion indicator tags to provide data support for subsequent model training, ensuring data diversity and integrity. The emotion indicator tags are then tokenized to obtain emotion indicator tokens, enabling more accurate extraction of their data features. Next, a pre-set neural coding model is trained based on the speech synthesis samples and emotion indicator tokens, allowing the model to accurately capture the correlation between speech and emotion, resulting in a target coding model that improves the model's ability to encode emotional speech. Furthermore, the target coding model encodes the pre-acquired target synthesized text to obtain raw coded data, accurately reflecting the relationship between text emotion and speech features. Finally, speech synthesis is performed based on the raw coded data to obtain high-quality, emotionally expressive target speech data. This solves the technical problem that emotion category tags cannot accurately control the emotional expression of speech synthesis, improving the accuracy of emotional expression in speech synthesis.

[0102] The embodiments described in this application are for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions provided by the embodiments of this application. As those skilled in the art will know, with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of this application are also applicable to similar technical problems.

[0103] Those skilled in the art will understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of this application, and may include more or fewer steps than shown, or combine certain steps, or different steps.

[0104] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs.

[0105] Those skilled in the art will understand that all or some of the steps in the methods disclosed above, as well as the functional modules / units in the systems and devices, can be implemented as software, firmware, hardware, or suitable combinations thereof.

[0106] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.

[0107] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0108] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of the units described above is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0109] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0110] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0111] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes multiple instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this application. The aforementioned storage medium includes various media capable of storing programs, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0112] The software tools or components not belonging to our company that appear in the embodiments of this application are for illustrative purposes only and do not represent actual use.

[0113] The preferred embodiments of the present application have been described above with reference to the accompanying drawings, but this does not limit the scope of the claims of the present application. Any modifications, equivalent substitutions, and improvements made by those skilled in the art without departing from the scope and substance of the embodiments of the present application shall be within the scope of the claims of the present application.

Claims

1. A speech synthesis method, characterized in that, The method includes: Obtain the speech synthesis dataset; the speech synthesis dataset includes speech synthesis samples and emotion indicator labels; The sentiment indicator tag is tokenized to obtain a sentiment tag token; The preset neural coding model is trained based on the speech synthesis samples and the emotion tag tokens to obtain the target coding model; The target synthetic text is encoded using the target encoding model to obtain the original encoded data; Speech synthesis is performed based on the original encoded data to obtain the target speech data.

2. The method according to claim 1, characterized in that, The step of tokenizing the emotion indicator tag to obtain an emotion tag token includes: The sentiment indicator tags are clustered to obtain multiple tag clusters and the three-dimensional center point corresponding to each tag cluster; The target boundary is obtained by performing boundary calculations based on the three-dimensional center points. Based on the target boundary, the sentiment indicator tags are binned to obtain binned tag data; The sentiment indicator tags in the binning tag data are used to perform token calculations to obtain the sentiment tag tokens.

3. The method according to claim 2, characterized in that, The boundary calculation based on the three-dimensional center point to obtain the target boundary includes: The three-dimensional center points are sorted to obtain a three-dimensional center sequence; A three-dimensional center pair is constructed based on each pair of adjacent three-dimensional center points in the three-dimensional center sequence; The target boundary is obtained by performing boundary calculations based on the three-dimensional center pair.

4. The method according to claim 3, characterized in that, The three-dimensional center pair includes a first three-dimensional center and a second three-dimensional center; the boundary calculation based on the three-dimensional center pair to obtain the target boundary includes: Boundary calculations are performed based on the first three-dimensional center and the second three-dimensional center to obtain the intermediate boundary; Obtain the number of labels in the label cluster corresponding to the first three-dimensional center to get the first label count; Obtain the number of labels in the label cluster corresponding to the second three-dimensional center to get the second label count; Boundary calculations are performed based on the first number of labels, the second number of labels, the first 3D center, and the second 3D center to obtain a weighted boundary. The target boundary is obtained by performing aggregation calculations on the intermediate boundary and the weighted boundary.

5. The method according to any one of claims 1 to 4, characterized in that, The step of training a preset neural coding model based on the speech synthesis samples and the emotion tag tokens to obtain the target coding model includes: The speech synthesis sample and the emotion tag token are encoded using the neural coding model to obtain sample encoded data; Speech synthesis is performed based on the sample encoded data to obtain sample speech data; Loss calculation is performed based on the speech synthesis samples and the sample speech data to obtain speech synthesis loss data; The neural coding model is optimized based on the speech synthesis loss data to obtain an initial coding model. The sample speech data is grouped to obtain sample speech pairs; Emotional loss is calculated based on the sample speech pairs to obtain emotional preference loss data; The initial encoding model is optimized based on the sentiment preference loss data to obtain the target encoding model.

6. The method according to any one of claims 1 to 4, characterized in that, The target synthesized text includes: speech text, emotional conditions, and speaker information; the process of encoding the pre-acquired target synthesized text using the target encoding model to obtain raw encoded data includes: The speech text is text encoded to obtain text encoding features; The emotional conditions are subjected to emotional encoding to obtain emotional encoding features; The original encoded data is obtained by neurally encoding the text encoding features, the emotion encoding features, and the speaker information using the target encoding model.

7. The method according to claim 6, characterized in that, The process of performing neural encoding on the text encoding features, the emotion encoding features, and the speaker information using the target encoding model to obtain the original encoded data includes: Obtain the label sentiment features corresponding to the sentiment conditions; Attention is calculated based on the sentiment encoding features and the labeled sentiment features to obtain fused sentiment features; The original encoded data is obtained by performing speech encoding on the fused emotional features, the text encoding features, and the speaker information using the target encoding model.

8. A speech synthesis device, characterized in that, The device includes: The data acquisition module is used to acquire the speech synthesis dataset; the speech synthesis dataset includes speech synthesis samples and emotion indicator labels; The token conversion module is used to convert the sentiment indicator tag into a sentiment tag token. The model training module is used to train a preset neural coding model based on the speech synthesis samples and the emotion tag tokens to obtain the target coding model. The encoding module is used to encode the pre-acquired target synthetic text using the target encoding model to obtain the original encoded data; The speech data synthesis module is used to perform speech synthesis based on the original encoded data to obtain target speech data.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method of any one of claims 1 to 7.