Spoken speech synthesis method and device, equipment and medium
Through colloquial text rewriting and assisted emotional label processing, the problem of lack of realism and emotional richness of speech synthesis system is solved, and voices that are closer to human natural communication are generated, improving the quality and flexibility of speech synthesis.
Patent Information
- Application Number
- CN202510675452.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-08-12
AI Technical Summary
The existing speech synthesis system cannot flexibly and accurately insert appropriate paralinguistic tags based on text emotions and context, resulting in the lack of realism and emotional richness of synthetic speech and cannot meet the needs of high-quality speech synthesis.
Through colloquial text rewriting, loss calculation, auxiliary emotional labeling and acoustic signal processing, the Meer spectrum is constructed and loss optimization is performed, and the target synthetic speech is finally generated to ensure that the speech is closer to the expression of human natural communication.
It greatly enriches the emotional expression and realism of voice, improves the user satisfaction of voice synthesis products, improves the quality and accuracy of voice synthesis, and enhances the flexibility and practicality of voice synthesis system in different scenarios.
Smart Images

Figure CN120472879A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech synthesis technology, and in particular to a spoken speech synthesis method, device, equipment and medium. Background Art
[0002] Driven by the wave of digitalization, speech synthesis technology has gradually moved from laboratory research to widespread practical application, profoundly changing the interaction mode between people and machines. Especially in the two fields of healthcare and financial technology, which have high requirements for interactive experience, speech synthesis technology is playing an increasingly critical role. For example, in the field of healthcare, combining speech synthesis technology with speech recognition technology can realize an electronic medical record system with voice interaction. Medical staff can dictate medical records and medical orders, and the system will automatically recognize and confirm the information in the form of voice. In the field of financial technology, complex financial knowledge (such as investment strategies and financial market analysis) can be converted into voice explanations through speech synthesis, which helps investors improve their financial literacy.
[0003] However, current traditional speech synthesis systems only output speech corresponding to the input text, and speech synthesis systems generally do not support colloquial rewriting of text. In addition, the speech generated by existing speech synthesis systems is mostly mechanical and cannot flexibly and accurately insert appropriate paralinguistic tags based on the text's emotions and context, making the synthesized speech lack realism and emotional richness.
[0004] Therefore, in the face of the growing demand for high-quality speech synthesis, current spoken speech synthesis methods urgently need to be improved to address the problem that existing methods lack realism and emotional richness. Summary of the Invention
[0005] The present invention provides a method, device, equipment and medium for colloquial speech synthesis. By rewriting colloquial text and introducing auxiliary emotional tags and naturally integrating them with the text, the synthesized speech is made closer to the expression mode in natural human communication, greatly enriching the emotional expressiveness and realism of the speech.
[0006] In a first aspect, a spoken speech synthesis method is provided, comprising:
[0007] Rewrite the original text into a colloquial language according to the preset emotional tags to obtain a colloquial text;
[0008] Calculating the loss between the spoken text and a preset real spoken text to obtain a loss value, and adjusting the spoken text according to the loss value to obtain an adjusted spoken text;
[0009] Constructing an auxiliary emotion tag, performing acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features;
[0010] Determining a Mel spectrum according to the acoustic features of the text, and performing loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum;
[0011] The optimized Mel spectrum is converted into target synthesized speech.
[0012] In a second aspect, a spoken speech synthesis device is provided, comprising:
[0013] A rewriting module is used to rewrite the original text into a colloquial language according to preset emotional tags to obtain a colloquial text;
[0014] a calculation module, configured to perform loss calculation on the colloquial text and a preset real colloquial text to obtain a loss value;
[0015] an adjustment module, configured to adjust the colloquial text according to the loss value to obtain an adjusted colloquial text;
[0016] Building module for constructing auxiliary emotion labels;
[0017] a processing module, configured to perform acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features;
[0018] A determination module, configured to determine a Mel spectrum according to the acoustic features of the text;
[0019] The optimization module is used to perform loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum.
[0020] In a third aspect, a computer device is provided, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the above-mentioned spoken speech synthesis method when executing the computer program.
[0021] In a fourth aspect, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above-mentioned spoken speech synthesis method are implemented.
[0022] In the solution implemented by the above-mentioned method, apparatus, computer device, and storage medium for colloquial speech synthesis, paralinguistic tags are introduced and naturally integrated with the text through emotional and colloquial text rewriting, making the synthesized speech closer to the expression methods used in natural human communication, greatly enriching the emotional expressiveness and realism of the speech, and improving user satisfaction with the speech synthesis product. The model can better capture the overall connection between text, emotion, and speech, avoiding the information loss and inconsistency problems that may be caused by traditional staged processing methods, further improving the quality and accuracy of speech synthesis. At the same time, the flexible generation and integration mechanism of paralinguistic tags enables the synthesized speech to be expressed in a variety of ways according to different contexts and emotional needs. The emotion and tone of the speech are adjusted according to the emotional feedback of different scenarios, providing more considerate services, thereby improving the flexibility and practicality of the speech synthesis system in various application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the description of the embodiments of the present invention. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0024] Figure 1 This is a schematic diagram of an application environment of a spoken speech synthesis method according to an embodiment of the present invention;
[0025] Figure 2 This is a flow chart of a method for synthesizing spoken language speech according to an embodiment of the present invention;
[0026] Figure 3 1 is a structural diagram of a spoken speech synthesis device according to an embodiment of the present invention;
[0027] Figure 4 is a structural diagram of a computer device according to an embodiment of the present invention;
[0028] Figure 5 It is another structural schematic diagram of a computer device in one embodiment of the present invention. DETAILED DESCRIPTION
[0029] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of them. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0030] The embodiment of the present invention provides a method for synthesizing spoken language, which can be applied in Figure 1 In an application environment, the client communicates with the server through a network. The server can rewrite the original text into colloquial language according to a preset emotional tag to obtain a colloquial text; calculate the loss of the colloquial text and the preset real colloquial text to obtain a loss value, adjust the colloquial text according to the loss value to obtain an adjusted colloquial text; construct an auxiliary emotional tag, perform acoustic signal processing on the adjusted colloquial text and the auxiliary emotional tag to obtain text acoustic features; determine the Mel spectrum according to the text acoustic features, and perform loss optimization on the Mel spectrum and the preset real Mel spectrum to obtain an optimized Mel spectrum; convert the optimized Mel spectrum into a target synthesized speech, and feed the target synthesized speech back to the client. The present invention provides a colloquial speech synthesis device. For the target synthesized speech business, by rewriting the colloquial text and introducing auxiliary emotional tags and naturally integrating them with the text, the synthesized speech is closer to the expression mode in natural human communication, which greatly enriches the emotional expressiveness and realism of the speech. The client can be, but is not limited to, various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server side can be implemented by an independent server or a server cluster composed of multiple servers. The present invention will be described in detail below through specific embodiments.
[0031] See also Figure 2 As shown, Figure 2 A flow chart of a method for synthesizing spoken language provided by an embodiment of the present invention includes the following steps:
[0032] S1. Rewrite the original text into a colloquial language according to the preset emotional tags to obtain a colloquial text.
[0033] In the embodiment of the present invention, the colloquial rewriting refers to converting a formal, written original text into a language form that is closer to daily oral expression.
[0034] Specifically, based on pre-set emotional labels (such as positive, negative, neutral, cheerful, sad, etc.), the original more formal, written or complex original text is rewritten into a text that conforms to daily oral expression habits, so that the rewritten colloquial text can reflect the preset emotional color.
[0035] In specific medical and health scenarios, medical information is usually more professional and complex, and patients may find it difficult to understand. Rewriting original medical-related texts, such as condition descriptions and treatment plans, in a colloquial manner according to preset emotional labels can convey information to patients in a gentler and easier-to-understand manner.
[0036] In FinTech scenarios, financial product and service descriptions often involve specialized terminology and complex clauses, which can be difficult for customers to understand. Rewriting these original texts in a more colloquial manner and incorporating elements like friendliness and trust based on emotional tags can help customers more easily understand the characteristics and advantages of financial products.
[0037] In the embodiment of the present invention, rewriting the original text into a colloquial language according to the preset emotional tags to obtain the colloquial text includes:
[0038] Determining a preset emotion tag for describing the emotion characteristics of the original text, and encoding the preset emotion tag to obtain an emotion code;
[0039] Constructing a mapping relationship table between preset emotion tags and emotion codes, identifying emotion features in the original text, and matching the identified emotion features with the preset emotion tags according to the mapping relationship table to determine the emotion tags corresponding to the original text;
[0040] The original text is gated and rewritten according to the matched emotion tag and the emotion code to obtain a colloquial text.
[0041] In an embodiment of the present invention, the construction refers to associating preset emotion tags with corresponding emotion codes through existing methods and rules to form a clear mapping relationship table that can be queried and used. The encoding refers to the process of converting preset emotion tags into digital or vector forms that can be understood and processed by computers. The gated rewriting refers to the process of adjusting text according to emotions by controlling the flow of information.
[0042] Specifically, for original texts, such as movie reviews and product reviews, labels can be determined based on the common emotional tendencies in the field. For example, in movie reviews, there may be emotional labels such as "like", "hate", "touched", and "disappointed" related to the viewing experience; in product reviews, there may be labels such as "satisfied", "dissatisfied", "surprised", and "worried" related to the product experience.
[0043] Specifically, each preset emotion tag is assigned a unique code. The coding method can be numbers, letters, symbols, or a combination thereof. For example, a digital code can be used, with 0 representing negativity, 1 representing neutrality, and 2 representing positivity; or an alphabetic code can be used, such as H for joy and S for sadness. The preset emotion tags are mapped one-to-one to the corresponding emotion codes to form a mapping relationship. This mapping relationship can be presented in the form of a table, with one column for the preset emotion tags and another for the corresponding emotion codes. The constructed mapping relationship table is stored in an appropriate data structure.
[0044] Furthermore, an emotional vocabulary is established, containing various words with emotional tendencies, and each word is labeled with an emotional category. Then, the emotional words appearing in the original text are counted, and the emotional characteristics of the text are inferred based on the emotional categories of these words. The extracted emotional characteristics are compared and matched with the preset emotional labels. Depending on the emotional feature extraction method used, different matching methods may be used. A vocabulary-based method can be used to find the preset emotional label that best matches the emotional characteristics of the original text through a matching process. This label is the emotional label corresponding to the original text.
[0045] Specifically, based on the sentiment label corresponding to the original text previously determined, the corresponding sentiment encoding is searched and obtained from the mapping table. A gating mechanism is then created based on the sentiment encoding. This mechanism can be a gating unit in a neural network model, such as a gating mechanism in a long short-term memory network or a gated recurrent unit. The sentiment encoding is used as input, and the gating unit controls the transmission and adjustment of the original text information during the rewriting process. Specifically, the original text is input into the gated rewriting model. The model performs preprocessing operations such as word segmentation and part-of-speech tagging on the original text, and then maps each word or phrase into a vector space. Under the control of the gating mechanism, the model adjusts the vector representation of the original text based on the sentiment encoding. This may involve operations such as vocabulary selection and replacement, word order adjustment, and sentence structure changes. For example, if the sentiment encoding is positive, the model may replace some neutral words with more positive words, such as changing "okay" to "great"; or add some positive modifiers, such as "very" or "special". At the same time, the word order of the sentence may be adjusted to make the text more positive.
[0046] In an embodiment of the present invention, the encoded emotion tags can be used as input features of the model, which facilitates computer storage, calculation and comparison, thereby realizing various natural language processing tasks based on emotion. Gated rewriting can accurately control the rewriting direction of the original text according to the emotion coding, so that the generated colloquial text is highly consistent with the preset emotion tags.
[0047] In the embodiment of the present invention, the text rewritten in colloquial language is closer to people's daily communication style, which can make information transmission more natural and smooth, making it easier for readers or listeners to understand and accept.
[0048] S2. Calculate the loss between the spoken text and a preset real spoken text to obtain a loss value, and adjust the spoken text according to the loss value to obtain an adjusted spoken text.
[0049] In an embodiment of the present invention, the loss calculation refers to the process of measuring the degree of difference between the colloquial text and the preset real colloquial text, and the adjustment refers to the loss value obtained based on the loss calculation. The process of modifying and optimizing the generated colloquial text is adjustment.
[0050] Specifically, the difference between the generated spoken text and the preset real spoken text is quantified through loss calculation to obtain a loss value; then, based on this loss value, the spoken text is adjusted using a specific optimization method to make it continuously approach the preset real spoken text, and finally an adjusted, higher-quality, and more satisfactory spoken text is obtained.
[0051] Accurate and accessible language is crucial for communication between doctors and patients in healthcare settings. By calculating loss values and adapting colloquial text, doctors or healthcare professionals can better tailor the information they provide to patients' understanding and emotional needs.
[0052] In the specific scenarios of financial technology, when customers inquire about financial products or services, the adjusted colloquial text can make the customer service staff's answers clearer and easier to understand, and at the same time optimize according to the customer's emotional state and the background of the problem, allowing customers to experience more considerate service.
[0053] In the embodiment of the present invention, the step of performing loss calculation on the spoken text and the preset real spoken text to obtain the loss value includes:
[0054] Performing word segmentation processing on the spoken text and the preset real spoken text respectively to obtain spoken word segmentation and real spoken word segmentation;
[0055] Converting the colloquial segmentation and the real colloquial segmentation into a colloquial word vector and a real colloquial word vector respectively;
[0056] The colloquial word vector and the real colloquial word vector are subjected to corresponding word distance calculation to obtain a calculation result for each corresponding word, and the calculation results are weighted averaged to obtain a loss value.
[0057] In an embodiment of the present invention, the word segmentation processing refers to the process of dividing continuous text into individual words or phrases according to certain rules, the conversion refers to the process of converting the segmented words into a vector representation that can be processed by a computer, the corresponding word distance calculation refers to calculating the distance between the colloquial word vector and the corresponding word vector in the real colloquial word vector to measure the degree of difference between them in the vector space, and the weighted average refers to the weighted summation of the distance calculation results of each corresponding word, and then dividing it by the total weight to obtain a comprehensive loss value.
[0058] Specifically, the spoken text is segmented with the preset real spoken text. Common word segmentation methods include dictionary-based word segmentation methods, which construct a dictionary containing a large number of words and match the text with the words in the dictionary to divide the words; at the same time, statistical word segmentation methods can be used, which use a large amount of text data to calculate the probability of word occurrence and the degree of correlation between words to determine the most likely word segmentation method.
[0059] Specifically, the conversion of colloquial and real-world colloquial word segmentations into colloquial word vectors and real-world colloquial word vectors is typically performed using a pre-trained word vector model. For example, the Word2Vec model is a common word vector training model. It uses unsupervised learning on large-scale text data to learn word vector representations based on the context of the words in the text. For a given colloquial and real-world colloquial word segmentation, each word is input into the pre-trained Word2Vec model, which outputs the corresponding word vector.
[0060] Furthermore, the distance calculation methods for corresponding words can be used, such as Euclidean distance and cosine similarity. By calculating the Euclidean distance of each pair of corresponding word vectors, the calculation result of each corresponding word can be obtained. Then, a weight is determined for the distance calculation result of each corresponding word. The weight can be set based on factors such as the importance and frequency of occurrence of the word. For example, some key medical terms or financial professional vocabulary may be given a higher weight. The weight and the corresponding word distance calculation result are weighted averaged to obtain the loss value.
[0061] In the embodiment of the present invention, adjusting the colloquial text according to the loss value to obtain the adjusted colloquial text includes:
[0062] Preprocessing the spoken text and the preset real spoken text;
[0063] Calculating the semantic similarity loss value, sentiment consistency loss value, and language fluency loss value between the preprocessed spoken text and the real spoken text based on the loss value;
[0064] Performing weighted synthesis on the calculated semantic similarity loss value, sentiment consistency loss value, and language fluency loss value to obtain a comprehensive loss value;
[0065] An adjustment strategy is determined according to the comprehensive loss value, and the spoken text is adjusted according to the adjustment strategy to obtain an adjusted spoken text.
[0066] In an embodiment of the present invention, the preprocessing refers to performing operations such as text cleaning, case conversion, word segmentation, and normalization on the spoken text and the preset real spoken text; the calculation refers to measuring the degree of difference between the spoken text and the real spoken text at the semantic level, evaluating the degree of consistency between the spoken text and the real spoken text in emotional expression, and judging the gap between the fluency of the spoken text in language expression and that of the real spoken text.
[0067] Specifically, when calculating semantic similarity, some pre-trained semantic similarity models, such as BERT, can be used. The spoken text and the real spoken text are input into the model respectively. The model will output a value representing the semantic similarity between the two. Usually, this value ranges from 0 to 1. The closer to 1, the more similar the semantics, and the closer to 0, the greater the semantic difference. The semantic similarity loss value can be obtained by subtracting this similarity value from 1. When calculating the sentiment consistency loss value, the sentiment analysis model is used to perform sentiment analysis on the spoken text and the real spoken text to obtain their respective sentiment categories (such as positive, negative, neutral) and sentiment intensity. Then, by comparing the two The loss value is calculated based on the emotional category and intensity of the speaker. If the emotional category is the same and the intensity difference is small, the loss value is low; conversely, the loss value is high. For example, a threshold can be set. When the emotional intensity difference exceeds this threshold, the loss value increases accordingly. The specific calculation method can be designed according to the actual situation. For example, methods such as mean square error can be used to calculate the difference in emotional intensity, or methods such as cross entropy can be used to measure the difference between emotional categories. When calculating the language fluency loss value, the spoken text and the real spoken text are input into the language model respectively to obtain their respective probability values. The difference between the two probability values is then calculated using a certain function as the language fluency loss value. For example, the log-likelihood function can be used to calculate the probability difference, or some methods based on edit distance can be used to measure the differences in grammar and word order of the text.
[0068] Furthermore, a corresponding weight is determined for each loss value. The weight is usually determined based on the task requirements, importance, and the degree of attention paid to different aspects. For example, if semantic similarity is more important, a higher weight, such as 0.5, may be assigned to the semantic similarity loss value; while emotional consistency and language fluency are relatively less important, they may be assigned weights of 0.3 and 0.2 respectively. Each loss value is multiplied by its corresponding weight to obtain a weighted loss value. The weighted loss values are added together to obtain the comprehensive loss value.
[0069] Specifically, if the semantic similarity loss value has a high weight, it indicates that there are significant semantic differences between the spoken text and the real spoken text. Based on the results of the semantic analysis, you can identify the semantically different parts and try to replace some words or adjust the sentence structure to make the semantics of the spoken text closer to the real spoken text. When the emotional consistency weight is high, the emotional expression of the spoken text needs to be adjusted. If the emotion of the spoken text is too positive or negative, while the real spoken text is neutral, the emotional intensity can be adjusted by adding some emotionally alleviating words or changing some words that describe the emotion. For cases where the language fluency loss value is high, adjustments are mainly made in terms of grammar and word order. If there are grammatical errors, such as subject-verb inconsistency and inappropriate word collocation, they should be modified according to grammatical rules. For word order issues, you can refer to the word order pattern of real spoken text or make adjustments based on language habits.
[0070] In an embodiment of the present invention, adjustment can be made to make the spoken text semantically closer to the real spoken text, so that the text can convey information more accurately and avoid misunderstandings caused by semantic ambiguity or deviation. At the same time, it can better conform to the specific context and emotional atmosphere, allowing readers or listeners to better feel the emotions contained in the text. The adjustment of the language fluency loss value can make the language expression of the spoken text more natural and fluent, in line with people's daily language habits, reduce obstacles in the reading or understanding process, and improve the readability and acceptability of the text.
[0071] In the embodiment of the present invention, the calculation and analysis of the loss value during the adjustment process can provide linguists and researchers with detailed information about language use and changes.
[0072] S3. Construct an auxiliary emotion tag, and perform acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features.
[0073] In an embodiment of the present invention, the construction refers to assigning corresponding emotion categories or emotion intensity and other information to the adjusted spoken text according to certain rules or algorithms, and the acoustic signal processing refers to the technology of analyzing, transforming, encoding and other operations on the spoken text and auxiliary emotion tags, aiming to extract useful information from the sound signal or optimize it to meet specific application requirements.
[0074] Specifically, auxiliary emotion tags can be expressed as paralinguistic tags, which usually refer to some characteristic tags that can convey additional information in addition to the language content itself. When adjusting the spoken text and auxiliary emotion tags for acoustic signal processing, the text will first be converted into a speech signal, which can be achieved through speech synthesis technology. Then, a series of processing is performed on the generated speech signal, such as sampling and quantizing the signal, and converting it into a digital signal for easy computer processing.
[0075] In detail, paralinguistic labels include various types, such as voice intonation, speaking speed, volume, pauses and other information. When constructing auxiliary emotional labels, we must first clarify the need to construct relevant paralinguistic labels. Common dimensions include: prosodic features: such as the ups and downs of intonation, the speed of speaking, the position of stress, etc., voice quality: including timbre, sound intensity, sound quality and other aspects of the characteristics, pauses and rhythm: the duration and frequency of pauses, as well as the rhythmic pattern of language, etc., including paralinguistic elements such as (hahaha), breathing sounds, um, uh, pursing lips, swallowing saliva, sighing, etc., which are crucial for emotional transmission in real human communication; first collect a large amount of spoken speech data, which should cover a variety of different scenarios, conversations, etc. The system then uses existing speech analysis technologies and tools, such as speech recognition software and acoustic feature extraction tools, to perform preliminary automatic labeling on the speech data. It then uses digital signal processing and speech analysis technologies to extract acoustic features related to paralinguistic labels from the speech data. It then quantifies the extracted continuous acoustic features and converts them into discrete label values. It then checks the consistency of the annotated paralinguistic labels between different annotators or the same annotator at different times, and verifies whether the correlation between the paralinguistic labels and the linguistic content expressed in the speech is reasonable. Finally, it organizes the verified and optimized paralinguistic labels to form a complete labeling system.
[0076] In specific medical and health scenarios, by analyzing the acoustic characteristics of the patient's voice when communicating with medical staff, it can assist in judging the patient's emotional state, such as anxiety, depression, or pain level.
[0077] In financial scenarios, when customers communicate with customer service staff of financial institutions by phone or use voice-interactive financial products, analyzing the acoustic characteristics of their voice can help us understand their emotional state, such as whether they are satisfied, anxious, or angry. This helps financial institutions promptly identify customer problems and dissatisfaction, take appropriate measures to resolve them, and improve customer service quality and customer loyalty.
[0078] In the embodiment of the present invention, the adjusted colloquial text and the auxiliary emotion tag are subjected to acoustic signal processing to obtain text acoustic features, including:
[0079] Converting the spoken text into speech data, and performing acoustic feature extraction on the speech data to obtain speech acoustic features;
[0080] Performing text feature extraction on the spoken text to obtain spoken text features;
[0081] Mapping the auxiliary emotion label to a preset vector space to obtain an emotion vector corresponding to the auxiliary emotion label;
[0082] The speech acoustic features, spoken text features and sentiment vectors are fused to obtain text acoustic features.
[0083] In an embodiment of the present invention, the text feature extraction refers to extracting information that can represent the characteristics of the spoken text from the spoken text, the mapping refers to establishing a correspondence between the auxiliary emotion tags and the preset vector space, and converting the discrete emotion tags into a continuous vector representation, and the feature fusion refers to merging features from different sources, namely speech acoustic features, spoken text features and emotion vectors, into a unified feature vector.
[0084] Specifically, the text is regarded as a collection of words, the order of the words is ignored, the frequency of each word in the text is counted, and a feature vector is formed. On the basis of the bag-of-words model, the importance of the word in the entire text collection is considered. TF (term frequency) represents the frequency of a word in the current text, and IDF (inverse document frequency) measures the rarity of a word in the entire text collection. Through the TF-IDF algorithm, the TF-IDF value of each word can be calculated as the feature of the text; you can also use pre-trained language models such as BERT, GPT, etc., input spoken text into the model, and the model will automatically learn the feature representation of the text.
[0085] Furthermore, a unique vector is assigned to each auxiliary emotion label in the vector space. A simple method is to use one-hot encoding. Assuming that there are N different emotion labels, each emotion label can be represented by a vector of length N, in which only the position corresponding to the emotion label is 1, and the rest of the positions are 0. For example, for the three emotion labels of "happy", "sad" and "angry", if the current auxiliary emotion label is "happy", then its one-hot encoding vector is [1,0,0].
[0086] Furthermore, the speech acoustic feature vector, spoken text feature vector and emotion vector are directly spliced together in a certain order to form a longer feature vector. According to the importance of different features, corresponding weights are assigned to each feature vector, and then the weighted sum is performed to obtain the fused feature vector. Specifically, a deep learning model such as a multi-layer perceptron (MLP), a convolutional neural network (CNN) or a recurrent neural network (RNN) can be used, and the speech acoustic features, spoken text features and emotion vectors are taken as input to allow the model to automatically learn how to fuse these features. The model will adjust parameters through training to optimize a specific objective function, such as classification accuracy or regression error, so as to obtain the best feature fusion method.
[0087] In an embodiment of the present invention, by extracting speech acoustic features, these physical properties can be converted into digital features that can be processed by computers, thereby providing a basis for subsequent analysis. At the same time, the spoken text features are represented in vector form, which is convenient for fusion and comparison with other text data. Through feature fusion, these multimodal information can be integrated, making full use of the advantages of each modality, and more comprehensively and accurately representing the characteristics of spoken text.
[0088] In the embodiment of the present invention, auxiliary emotion tags can help determine the emotional tone of the text, while acoustic features can reflect information such as the speaker's tone and attitude, all of which help to better understand the meaning and intention of the text in complex language scenarios.
[0089] S4. Determine a Mel spectrum according to the acoustic features of the text, and perform loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum.
[0090] In the embodiment of the present invention, the determining refers to a spectrum representation obtained by analyzing the acoustic features of the text on a Mel-frequency scale, and the loss optimization refers to a process of optimizing the Mel-spectrum according to the difference between the Mel-spectrum and a preset true Mel-spectrum.
[0091] Specifically, based on the given acoustic features of the text, the corresponding Mel spectrum is calculated through a series of signal processing and transformation operations, which usually involve framing, windowing, fast Fourier transform (FFT) and other operations on the speech signal; after calculating the difference between the two, the Mel spectrum is adjusted to be as close as possible to the preset true Mel spectrum to improve the accuracy and reliability of the model or processing results.
[0092] In the embodiment of the present invention, determining the mel spectrum according to the acoustic features of the text includes:
[0093] Performing a nonlinear activation function transformation on the acoustic features of the text to obtain an initial Mel spectrum;
[0094] Mapping the emotion vector into a scaling coefficient, and performing emotion adjustment on the initial Mel spectrum according to the scaling coefficient to obtain an adjusted Mel spectrum;
[0095] Dividing the text acoustic features into multiple groups of frame segments, performing duration analysis on the frame segments to obtain frame segment duration analysis results, and calculating the weights of the frame segments according to the frame segment duration analysis results;
[0096] The adjusted mel spectrum is multiplied by the weight corresponding to the frame segment to obtain a mel spectrum.
[0097] In an embodiment of the present invention, the nonlinear activation function transformation refers to converting the input text acoustic features through a nonlinear function, so that the model can learn more complex feature relationships and increase the expressive ability of the model. The emotional adjustment refers to adjusting the initial Mel spectrum according to the scaling coefficient obtained by the emotional vector mapping to incorporate emotional information so that the Mel spectrum can better reflect the emotional features in the text. The duration analysis refers to analyzing the time length of multiple groups of frame segments divided by the acoustic features of the text to obtain information about each frame segment in the time dimension. The calculation refers to the process of calculating the weight of each frame segment through a certain algorithm based on the frame segment duration analysis results.
[0098] Specifically, the text acoustic feature is a vector x, and a nonlinear activation function is selected, such as the ReLU function, for each element x in the vector x i , apply the ReLU function to transform and get y i =max(0,x i ). After such a transformation, the obtained vector y is the result of the nonlinear activation function transformation, that is, the initial Mel spectrum.
[0099] Furthermore, the emotion vector is mapped into a scaling factor s through a mapping function f. This mapping function can be a simple linear transformation plus an activation function, such as s = g(W·v+b), where v is the emotion vector, W is the weight matrix, b is the bias term, and g is the activation function. Then, each element M in the initial Mel spectrum M is converted to ij Multiply by the scaling factor s to get the adjusted Mel spectrum M′ ij =M ij ×s, thereby achieving emotional adjustment of the initial Mel spectrum.
[0100] Furthermore, the sum of the duration of all frame segments is calculated Then, for each frame segment i, its weight Thus, the weight w i It reflects the proportion of the duration of frame segment i in the total duration. Frame segments with longer durations have higher weights, indicating that they may play a more important role in the entire text acoustic features.
[0101] Furthermore, the adjusted Mel spectrum is multiplied by the weight corresponding to the frame segment, that is, for each element M′ in the adjusted Mel spectrum M′ ij , and compare it with the corresponding frame segment weight w i Multiply to get the final Mel spectrum M″ ij =M′ ij ×w i .
[0102] Specifically, the average value of the square of the difference between the Mel spectrum and the preset true Mel spectrum is calculated. The formula is: where y i is the value of the true Mel spectrum, is the value of the Mel spectrum, N is the number of samples, and the mean absolute error of the two is calculated at the same time to obtain the loss value. In addition, to generate or convert Mel spectrum based on a deep learning model, it is necessary to initialize the parameters of the model first; these parameters are usually the weights and bias items in the model. The initialization method can be random initialization or some predefined methods, such as Xavier initialization, Kaiming initialization, etc. According to the loss value, the back propagation algorithm is used to calculate the gradient of the model parameters. According to the calculated gradient, the optimization algorithm is used to update the parameters of the model. Common optimization algorithms include stochastic gradient descent. Repeat the above steps, continuously perform forward propagation, calculate loss, back propagate and update parameters. The model will continuously adjust the parameters so that the predicted Mel spectrum gradually approaches the preset true Mel spectrum, and finally obtain the optimized Mel spectrum.
[0103] In an embodiment of the present invention, the different frequency components of the initial Mel spectrum are adjusted according to the emotion vector so that the Mel spectrum can better reflect emotion-related features. The weights are calculated based on the results of the frame segment duration analysis, which can enable the model to pay more attention to those frame segments that make important contributions to tasks such as speech recognition or emotion analysis, thereby improving the accuracy and performance of the model.
[0104] In the embodiment of the present invention, the loss optimization process continuously adjusts the model parameters so that the calculated mel spectrum is as close as possible to the preset real mel spectrum, thereby improving the accuracy and reliability of the mel spectrum.
[0105] S5. Convert the optimized Mel spectrum into target synthesized speech.
[0106] In the embodiment of the present invention, the conversion refers to reconstructing speech feature data represented in the frequency domain, such as the Mel spectrum, into a speech signal in the time domain through specific algorithms and techniques, thereby generating audible synthesized speech.
[0107] Specifically, speech synthesis technology is used, such as a vocoder-based method; the vocoder will reconstruct the time domain speech waveform based on characteristic parameters such as the Mel spectrum through a series of signal processing operations such as filtering and waveform generation, and ultimately generate the target synthesized speech with specific timbre, intonation and other characteristics.
[0108] In the embodiment of the present invention, converting the optimized Mel spectrum into a target synthesized speech includes:
[0109] Performing an inverse Mel transform on the optimized Mel spectrum to obtain a linear spectrum;
[0110] Performing spectrum phase processing on the linear spectrum to obtain a phase spectrum;
[0111] Performing an inverse Fourier transform on the phase spectrum to obtain a time domain waveform;
[0112] High-frequency noise is removed from the time domain waveform to obtain target synthesized speech.
[0113] In an embodiment of the present invention, the inverse Mel transform refers to converting the information contained in the Mel spectrum to a linear frequency scale, the spectrum phase processing refers to the process of analyzing and adjusting the phase information of the linear spectrum, the inverse Fourier transform refers to a mathematical transformation that converts a frequency domain signal (such as a spectrum after phase processing) into a time domain signal, and the high-frequency noise removal refers to filtering or attenuating the high-frequency noise signal from the time domain waveform.
[0114] Specifically, a conversion matrix is usually first established based on the conversion relationship between Mel frequencies and linear frequencies. This conversion relationship is nonlinear and is generally calculated using a formula. Then, the Mel spectrum is multiplied by this conversion matrix to obtain the linear spectrum. Phase information is extracted from the linear spectrum, typically by representing the linear spectrum as a complex number and then separating its phase component. Next, various phase processing algorithms, such as minimum phase reconstruction, can be used. These algorithms adjust and optimize the phase based on prior knowledge or statistical properties of the speech signal. For example, the phase may be periodically adjusted based on characteristics such as the pitch period of the speech signal to simulate the phase relationship between harmonics in natural speech.
[0115] Furthermore, for discrete spectral data, the inverse Fourier transform (IDFT) is typically implemented using the inverse discrete Fourier transform (IDFT) algorithm. This algorithm multiplies each frequency component in the frequency domain by a corresponding complex exponential factor and sums all frequency components to obtain each sample point in the time domain. The specific calculation formula involves complex number operations and exponential functions. In actual calculations, to improve computational efficiency, the inverse algorithm of the fast Fourier transform (FFT) is often used to implement the inverse Fourier transform. The inverse Fourier transform converts spectral information into a time-domain waveform.
[0116] Furthermore, appropriate filter parameters, such as the cutoff frequency and filter order, must be selected based on the characteristics of the speech signal and the frequency distribution of the noise. The cutoff frequency is typically chosen outside the effective frequency band of the speech signal to preserve as much useful speech information as possible while filtering out high-frequency noise. The time-domain waveform is then passed through a designed low-pass filter. The filter performs operations such as weighted summation on each sample point, attenuating or retaining signals of different frequencies based on the filter coefficients. This results in the target synthesized speech after removing high-frequency noise.
[0117] In an embodiment of the present invention, the inverse Mel transform converts it back into a linear spectrum, making the spectrum representation more consistent with conventional frequency analysis methods, facilitating subsequent signal processing and analysis within the traditional linear frequency framework; at the same time, the inverse Fourier transform can accurately reconstruct the original time domain speech signal based on all frequency components and their phase information in the frequency domain.
[0118] In the embodiment of the present invention, the optimized Mel spectrum is a representation of speech features. Converting it into target synthesized speech can convert abstract speech features into actual sound signals, meeting user needs in application scenarios such as voice interaction and voice broadcasting.
[0119] It can be seen that in the above scheme, for the target synthetic speech service, the original text is rewritten in colloquial language according to the preset emotional label to obtain a spoken text; the spoken text and the preset real spoken text are subjected to loss calculation to obtain a loss value, and the spoken text is adjusted according to the loss value to obtain an adjusted spoken text; an auxiliary emotional label is constructed, and the adjusted spoken text and the auxiliary emotional label are subjected to acoustic signal processing to obtain text acoustic features; the Mel spectrum is determined according to the text acoustic features, and the Mel spectrum and the preset real Mel spectrum are subjected to loss optimization to obtain an optimized Mel spectrum; the optimized Mel spectrum is converted into the target synthetic speech, and through the colloquial text rewriting, the auxiliary emotional label is introduced and naturally integrated with the text, so that the synthetic speech is closer to the expression mode in natural human communication, which greatly enriches the emotional expressiveness and realism of the speech.
[0120] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0121] In one embodiment, a spoken speech synthesis device is provided, and the spoken speech synthesis device corresponds one-to-one to a spoken speech synthesis method in the above embodiment. Figure 3 As shown, the spoken speech synthesis device includes a rewriting module 101, a calculation module 102, an adjustment module 103, a construction module 104, a processing module 105, a determination module 106, and an optimization module 107. The functional modules are described in detail as follows:
[0122] The rewriting module 101 is used to rewrite the original text into a colloquial language according to the preset emotional tags to obtain a colloquial text;
[0123] A calculation module 102 is configured to perform loss calculation on the colloquial text and a preset real colloquial text to obtain a loss value;
[0124] An adjustment module 103, configured to adjust the spoken text according to the loss value to obtain an adjusted spoken text;
[0125] A construction module 104 is used to construct auxiliary emotion labels;
[0126] A processing module 105 is configured to perform acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features;
[0127] A determination module 106 is configured to determine a Mel spectrum according to the acoustic features of the text;
[0128] The optimization module 107 is configured to perform loss optimization on the mel spectrum and a preset true mel spectrum to obtain an optimized mel spectrum.
[0129] In one embodiment, the rewriting module 101, when rewriting the original text into colloquial language according to the preset emotion tags to obtain the colloquial text, is configured to:
[0130] Determining a preset emotion tag for describing the emotion characteristics of the original text, and encoding the preset emotion tag to obtain an emotion code;
[0131] Constructing a mapping relationship table between preset emotion tags and emotion codes, identifying emotion features in the original text, and matching the identified emotion features with the preset emotion tags according to the mapping relationship table to determine the emotion tags corresponding to the original text;
[0132] The original text is gated and rewritten according to the matched emotion tag and the emotion code to obtain a colloquial text.
[0133] In one embodiment, when calculating the loss between the spoken text and a preset real spoken text to obtain the loss value, the calculation module 102 is configured to:
[0134] Performing word segmentation processing on the spoken text and the preset real spoken text respectively to obtain spoken word segmentation and real spoken word segmentation;
[0135] Converting the colloquial segmentation and the real colloquial segmentation into a colloquial word vector and a real colloquial word vector respectively;
[0136] The colloquial word vector and the real colloquial word vector are subjected to corresponding word distance calculation to obtain a calculation result for each corresponding word, and the calculation results are weighted averaged to obtain a loss value.
[0137] In one embodiment, when adjusting the colloquial text according to the loss value to obtain the adjusted colloquial text, the adjustment module 103 is configured to:
[0138] Preprocessing the spoken text and the preset real spoken text;
[0139] Calculating the semantic similarity loss value, sentiment consistency loss value, and language fluency loss value of the preprocessed spoken text and the real spoken text respectively according to the loss value;
[0140] Performing weighted synthesis on the calculated semantic similarity loss value, sentiment consistency loss value, and language fluency loss value to obtain a comprehensive loss value;
[0141] An adjustment strategy is determined according to the comprehensive loss value, and the spoken text is adjusted according to the adjustment strategy to obtain an adjusted spoken text.
[0142] In one embodiment, when the processing module 105 performs acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features, it is configured to:
[0143] Converting the spoken text into speech data, and performing acoustic feature extraction on the speech data to obtain speech acoustic features;
[0144] Performing text feature extraction on the spoken text to obtain spoken text features;
[0145] Mapping the auxiliary emotion label to a preset vector space to obtain an emotion vector corresponding to the auxiliary emotion label;
[0146] The speech acoustic features, spoken text features and sentiment vectors are fused to obtain text acoustic features.
[0147] In one embodiment, when determining the mel spectrum according to the acoustic features of the text, the determination module 106 is configured to:
[0148] Performing a nonlinear activation function transformation on the acoustic features of the text to obtain an initial Mel spectrum;
[0149] Mapping the emotion vector into a scaling coefficient, and performing emotion adjustment on the initial Mel spectrum according to the scaling coefficient to obtain an adjusted Mel spectrum;
[0150] Dividing the text acoustic features into multiple groups of frame segments, performing duration analysis on the frame segments to obtain frame segment duration analysis results, and calculating the weights of the frame segments according to the frame segment duration analysis results;
[0151] The adjusted mel spectrum is multiplied by the weight corresponding to the frame segment to obtain a mel spectrum.
[0152] In one embodiment, when converting the optimized Mel-spectrogram into target synthesized speech, the optimization module 107 is configured to:
[0153] Performing an inverse Mel transform on the optimized Mel spectrum to obtain a linear spectrum;
[0154] Performing spectrum phase processing on the linear spectrum to obtain a phase spectrum;
[0155] Performing an inverse Fourier transform on the phase spectrum to obtain a time domain waveform;
[0156] High-frequency noise is removed from the time domain waveform to obtain target synthesized speech.
[0157] The present invention provides a spoken speech synthesis device. For a target synthesized speech service, an original text is rewritten in a spoken language according to a preset emotion tag to obtain a spoken text; a loss calculation is performed on the spoken text and a preset real spoken text to obtain a loss value, and the spoken text is adjusted according to the loss value to obtain an adjusted spoken text; an auxiliary emotion tag is constructed, and the adjusted spoken text and the auxiliary emotion tag are subjected to acoustic signal processing to obtain text acoustic features; a mel spectrum is determined according to the text acoustic features, and loss optimization is performed on the mel spectrum and a preset real mel spectrum to obtain an optimized mel spectrum; the optimized mel spectrum is converted into a target synthesized speech, and through the spoken text rewriting, the auxiliary emotion tag is introduced and the speech is naturally integrated with the text, so that the synthesized speech is closer to the expression mode in natural human communication, and the emotional expressiveness and realism of the speech are greatly enriched.
[0158] For the specific definition of a spoken speech synthesis device, please refer to the definition of a spoken speech synthesis method above, which will not be repeated here. The various modules in the above-mentioned spoken speech synthesis device can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of the processor in the computer device in the form of hardware, or can be stored in the memory of the computer device in the form of software, so that the processor can call and execute the operations corresponding to the above modules.
[0159] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4As shown. The computer device includes a processor, a memory, a network interface and a database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile and / or volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external client via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the server side of a spoken speech synthesis method.
[0160] In one embodiment, a computer device is provided. The computer device may be a client, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen, and input device connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a client side of a spoken speech synthesis method.
[0161] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:
[0162] Rewrite the original text into a colloquial language according to the preset emotional tags to obtain a colloquial text;
[0163] Calculating the loss between the spoken text and a preset real spoken text to obtain a loss value, and adjusting the spoken text according to the loss value to obtain an adjusted spoken text;
[0164] Constructing an auxiliary emotion tag, performing acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features;
[0165] Determining a Mel spectrum according to the acoustic features of the text, and performing loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum;
[0166] The optimized Mel spectrum is converted into target synthesized speech.
[0167] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0168] Rewrite the original text into a colloquial language according to the preset emotional tags to obtain a colloquial text;
[0169] Calculating the loss between the spoken text and a preset real spoken text to obtain a loss value, and adjusting the spoken text according to the loss value to obtain an adjusted spoken text;
[0170] Constructing an auxiliary emotion tag, performing acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features;
[0171] Determining a Mel spectrum according to the acoustic features of the text, and performing loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum;
[0172] The optimized Mel spectrum is converted into target synthesized speech.
[0173] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the client side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.
[0174] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).
[0175] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0176] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. If software tools or components other than those of the company appear in the application embodiments, they are merely used for illustration and do not represent actual use. Although the present invention has been described in detail with reference to the above-mentioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above-mentioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.
Claims
1. A spoken speech synthesis method, characterized in that: include: Rewrite the original text into a colloquial language according to the preset emotional tags to obtain a colloquial text; Calculating the loss between the spoken text and a preset real spoken text to obtain a loss value, and adjusting the spoken text according to the loss value to obtain an adjusted spoken text; Constructing an auxiliary emotion tag, performing acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features; Determining a Mel spectrum according to the acoustic features of the text, and performing loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum; The optimized Mel spectrum is converted into target synthesized speech.
2. The method for spoken speech synthesis according to claim 1, wherein: The process of rewriting the original text into a colloquial language according to the preset emotional tags to obtain the colloquial text includes: Determining a preset emotion tag for describing the emotion characteristics of the original text, and encoding the preset emotion tag to obtain an emotion code; Constructing a mapping relationship table between preset emotion tags and emotion codes, identifying emotion features in the original text, and matching the identified emotion features with the preset emotion tags according to the mapping relationship table to determine the emotion tags corresponding to the original text; The original text is gated and rewritten according to the matched emotion tag and the emotion code to obtain a colloquial text.
3. The spoken speech synthesis method according to claim 1, wherein: The step of performing loss calculation on the spoken text and the preset real spoken text to obtain a loss value includes: Performing word segmentation processing on the spoken text and the preset real spoken text respectively to obtain spoken word segmentation and real spoken word segmentation; Converting the colloquial segmentation and the real colloquial segmentation into a colloquial word vector and a real colloquial word vector respectively; The colloquial word vector and the real colloquial word vector are subjected to corresponding word distance calculation to obtain a calculation result for each corresponding word, and the calculation results are weighted averaged to obtain a loss value.
4. The method for synthesizing spoken language according to claim 1, wherein: The step of adjusting the spoken text according to the loss value to obtain the adjusted spoken text includes: Preprocessing the spoken text and the preset real spoken text; Calculating the semantic similarity loss value, sentiment consistency loss value, and language fluency loss value between the preprocessed spoken text and the real spoken text based on the loss value; Performing weighted synthesis on the calculated semantic similarity loss value, sentiment consistency loss value, and language fluency loss value to obtain a comprehensive loss value; An adjustment strategy is determined according to the comprehensive loss value, and the spoken text is adjusted according to the adjustment strategy to obtain an adjusted spoken text.
5. The spoken speech synthesis method according to claim 1, wherein: The step of performing acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features includes: Converting the spoken text into speech data, and performing acoustic feature extraction on the speech data to obtain speech acoustic features; Performing text feature extraction on the spoken text to obtain spoken text features; Mapping the auxiliary emotion label to a preset vector space to obtain an emotion vector corresponding to the auxiliary emotion label; The speech acoustic features, spoken text features and sentiment vectors are fused to obtain text acoustic features.
6. The spoken speech synthesis method according to claim 5, wherein: The determining of the Mel spectrum according to the acoustic features of the text includes: Performing a nonlinear activation function transformation on the acoustic features of the text to obtain an initial Mel spectrum; Mapping the emotion vector into a scaling coefficient, and performing emotion adjustment on the initial Mel spectrum according to the scaling coefficient to obtain an adjusted Mel spectrum; Dividing the text acoustic features into multiple groups of frame segments, performing duration analysis on the frame segments to obtain frame segment duration analysis results, and calculating the weights of the frame segments according to the frame segment duration analysis results; The adjusted mel spectrum is multiplied by the weight corresponding to the frame segment to obtain a mel spectrum.
7. The spoken speech synthesis method according to claim 1, wherein: The step of converting the optimized Mel spectrum into a target synthesized speech comprises: Performing an inverse Mel transform on the optimized Mel spectrum to obtain a linear spectrum; Performing spectrum phase processing on the linear spectrum to obtain a phase spectrum; Performing an inverse Fourier transform on the phase spectrum to obtain a time domain waveform; High-frequency noise is removed from the time domain waveform to obtain target synthesized speech.
8. A spoken speech synthesis device, characterized in that: include: A rewriting module is used to rewrite the original text into a colloquial language according to preset emotional tags to obtain a colloquial text; a calculation module, configured to perform loss calculation on the colloquial text and a preset real colloquial text to obtain a loss value; an adjustment module, configured to adjust the colloquial text according to the loss value to obtain an adjusted colloquial text; Building module for constructing auxiliary emotion labels; a processing module, configured to perform acoustic signal processing on the adjusted colloquial text and the auxiliary emotion tag to obtain text acoustic features; A determination module, configured to determine a Mel spectrum according to the acoustic features of the text; The optimization module is used to perform loss optimization on the Mel spectrum and a preset true Mel spectrum to obtain an optimized Mel spectrum.
9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the steps of the spoken speech synthesis method according to any one of claims 1 to 7 are implemented.
10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the steps of the spoken speech synthesis method according to any one of claims 1 to 7 are implemented.