Speech synthesis method, speech synthesis device, electronic device, and storage medium

By acquiring sample data and using the initial synthesis model for masking and emotion feature detection, a predicted Mel spectrum is generated, which solves the problem of unnatural association between linguistic and non-linguistic information in speech synthesis and achieves more accurate emotion expression.

CN119339705BActive Publication Date: 2025-11-25PING AN TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411456893.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-17
Publication Date
2025-11-25
Estimated Expiration
2044-10-17

AI Technical Summary

Technical Problem

Existing technologies fail to deeply understand the relationship between linguistic and non-linguistic information during speech synthesis, resulting in synthesized speech that is not natural or accurate in expressing emotions.

Method used

By acquiring sample data, including sample text, sample sentiment information, and sample raw Mel spectra, the initial synthesis model is used for masking and sentiment feature detection. Combined with the spectrum generation sub-model, a predicted Mel spectrum is generated. The model parameters are adjusted to generate synthesized speech with target speech expression features and target sentiment features.

Benefits of technology

It achieves more accurate synthesized speech for emotional expression by integrating linguistic and non-linguistic information to generate more natural and vivid speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339705B_ABST
    Figure CN119339705B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a speech synthesis method, a speech synthesis device, an electronic device and a storage medium, and belongs to the technical field of financial science and technology. The method comprises the following steps: performing mask processing on a sample original mel spectrum based on sample emotional information to obtain a mask mel spectrum and a sample mask post-mel spectrum; detecting a sample target emotional feature based on a sample emotional information and a sample original mel spectrum by using an emotional detection sub-model; performing spectrum generation on the sample mask post-mel spectrum, the sample target emotional feature and sample text based on a spectrum generation sub-model to obtain a predicted mel spectrum; and adjusting an initial synthesis model based on the mask mel spectrum, the predicted mel spectrum and the sample target emotional feature to obtain a speech synthesis model, so that the speech synthesis model is used for speech synthesis of target text, target object speech with target speech expression characteristics and target emotional information with target non-verbal emotional characteristics. The embodiment of the application can generate synthesized speech with more accurate emotional expression.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of financial technology, and in particular to a speech synthesis method, a speech synthesis device, an electronic device and a storage medium. BACKGROUND

[0002] Speech synthesis refers to a process of creating a natural language speech meeting requirements according to given input data. At present, intelligent speech technology is usually applied in intelligent telephone customer service, intelligent sales and other task scenarios of financial technology. However, human speech not only contains language information, but also contains some non-language information such as crying, laughing, pause and coughing, which can be used to convey different feelings of the speaker and the intention of communication. Therefore, adding some non-language information in the process of synthesizing speech can make the synthesized speech more natural and lively, and closer to the speech in real life.

[0003] Therefore, the related technology usually splices speech information and non-language information in a certain way to synthesize a piece of speech containing non-language information. However, this method does not deeply understand the association between language information and non-language information, so that the connection between speech information and non-language information in the synthesized speech is not natural enough, thereby reducing the accuracy of emotional expression of the synthesized speech. Therefore, how to generate synthesized speech with more accurate emotional expression has become a technical problem to be solved. SUMMARY

[0004] The main purpose of the embodiments of the present application is to provide a speech synthesis method, a speech synthesis device, an electronic device and a storage medium, which can generate synthesized speech with more accurate emotional expression.

[0005] To achieve the above purpose, a first aspect of the embodiments of the present application provides a speech synthesis method, which comprises:

[0006] obtaining sample data, the sample data comprising sample text, sample emotional information and sample original mel spectrum, the sample emotional information having sample non-language emotional features, and the sample original mel spectrum having sample speech expression features of a sample object;

[0007] inputting the sample text, the sample emotional information and the sample original mel spectrum into an initial synthesis model, the initial synthesis model comprising an emotional detection sub-model and a spectrum generation sub-model;

[0008] masking the sample original mel spectrum based on the sample emotional information to obtain a mask mel spectrum and a sample masked mel spectrum, the mask mel spectrum being used to represent a mel spectrum corresponding to a mask region in the sample original mel spectrum, and the sample masked mel spectrum being used to represent a mel spectrum obtained by masking a mel spectrum value of the mask region in the sample original mel spectrum;

[0009] detecting an emotional feature of the sample masked mel spectrum based on the emotional detection sub-model and the sample emotional information to obtain a sample target emotional feature of the sample masked mel spectrum;

[0010] generating a predicted mel spectrum of the sample object based on the sample masked mel spectrum, the sample target emotional feature, and the sample text by using the spectrum generation sub-model;

[0011] adjusting parameters of the initial synthesis model based on the mask mel spectrum, the predicted mel spectrum, and the sample target emotional feature to obtain a speech synthesis model;

[0012] performing speech synthesis processing on a target text, a target object speech having a target speech expression feature, and target emotional information having a target non-verbal emotional feature based on the speech synthesis model to obtain a target synthesized speech having the target speech expression feature and the target emotional feature.

[0013] In some embodiments, the detecting an emotional feature of the sample masked mel spectrum based on the emotional detection sub-model and the sample emotional information to obtain a sample target emotional feature of the sample masked mel spectrum comprises:

[0014] detecting an emotional feature of the sample emotional information to obtain a sample non-verbal emotional feature, the sample non-verbal emotional feature being used to indicate an emotional feature of the mask mel spectrum;

[0015] detecting an emotional feature of the sample original mel spectrum to obtain a sample verbal emotional feature;

[0016] splicing the sample non-verbal emotional feature and the sample verbal emotional feature to obtain the sample target emotional feature of the sample masked mel spectrum.

[0017] In some embodiments, the initial synthesis model further comprises a phoneme detection sub-model, and the generating a predicted mel spectrum of the sample object based on the sample masked mel spectrum, the sample target emotional feature, and the sample text by using the spectrum generation sub-model comprises:

[0018] performing phoneme detection on the sample original mel spectrum and the sample text based on the phoneme detection sub-model to obtain a sample phoneme sequence;

[0019] performing spectrum generation on the sample masked mel spectrum, the sample target emotional feature and the sample phoneme sequence based on the spectrum generation sub-model to obtain the predicted mel spectrum.

[0020] In some embodiments, the performing spectrum generation on the sample masked mel spectrum, the sample target emotional feature and the sample phoneme sequence based on the spectrum generation sub-model to obtain the predicted mel spectrum comprises:

[0021] aligning the sample masked mel spectrum in a time dimension based on the sample phoneme sequence to obtain a sample aligned mel spectrum;

[0022] performing spectrum generation on the sample aligned mel spectrum, the sample target emotional feature and the sample phoneme sequence based on the spectrum generation sub-model to obtain the predicted mel spectrum.

[0023] In some embodiments, the performing spectrum generation on the sample aligned mel spectrum, the sample target emotional feature and the sample phoneme sequence based on the spectrum generation sub-model to obtain the predicted mel spectrum comprises:

[0024] aligning the sample target emotional feature in a time dimension based on the sample phoneme sequence to obtain a sample aligned emotional feature;

[0025] performing spectrum generation on the sample aligned mel spectrum, the sample aligned emotional feature and the sample phoneme sequence based on the spectrum generation sub-model to obtain the predicted mel spectrum.

[0026] In some embodiments, the initial synthesis model further comprises a vocoder, and the performing parameter adjustment on the initial synthesis model based on the masked mel spectrum, the predicted mel spectrum and the sample target emotional feature to obtain a speech synthesis model comprises:

[0027] performing spectrum loss calculation based on the masked mel spectrum and the predicted mel spectrum to obtain a spectrum loss value, the masked mel spectrum having the sample speech expression feature, the predicted mel spectrum having a predicted speech expression feature, the spectrum loss value being used to represent a difference degree between the sample speech expression feature and the predicted speech expression feature;

[0028] performing spectrum conversion on the predicted mel spectrum based on the vocoder to obtain a predicted speech;

[0029] performing emotional feature extraction on the predicted speech to obtain a predicted emotional feature;

[0030] perform emotion feature loss calculation based on the predicted emotion feature and the sample target emotion feature, to obtain an emotion feature loss value;

[0031] determine a model loss value based on the spectrum loss value and the emotion feature loss value, and perform parameter adjustment on the initial synthesis model based on the model loss value, to obtain the speech synthesis model.

[0032] In some embodiments, the determining the model loss value based on the spectrum loss value and the emotion feature loss value comprises:

[0033] perform spectrum feature extraction on the mask mel spectrum, to obtain a mask spectrum feature;

[0034] perform feature detection on the mask spectrum feature based on the sample original mel spectrum, to obtain a mask feature score, the mask feature score being used to represent an importance degree of the mask spectrum feature in the sample original mel spectrum;

[0035] determine a first loss weight of the spectrum loss value and a second loss weight of the emotion feature loss value based on the mask feature score;

[0036] perform weighted calculation based on the spectrum loss value, the first loss weight, the emotion feature loss value and the second loss weight, to obtain the model loss value.

[0037] To achieve the above object, a second aspect of the embodiment of the present application proposes a speech synthesis device, which comprises:

[0038] an acquisition module, configured to acquire sample data, the sample data comprising sample text, sample emotion information and sample original mel spectrum, the sample emotion information having sample non-language emotion features, and the sample original mel spectrum having sample speech expression features of a sample object;

[0039] an input module, configured to input the sample text, the sample emotion information and the sample original mel spectrum into an initial synthesis model, the initial synthesis model comprising an emotion detection sub-model and a spectrum generation sub-model;

[0040] a mask module, configured to perform mask processing on the sample original mel spectrum based on the sample emotion information, to obtain a mask mel spectrum and a sample mask post-mel spectrum, the mask mel spectrum being used to represent a mel spectrum corresponding to a mask region in the sample original mel spectrum, and the sample mask post-mel spectrum being used to represent a mel spectrum obtained by masking the mel spectrum value of the mask region in the sample original mel spectrum;

[0041] The emotion detection module is configured to perform emotion feature detection on the sample emotion information and the sample original mel spectrum based on the emotion detection sub-model, to obtain sample target emotion features of the sample masked mel spectrum.

[0042] The spectrum generation module is configured to perform spectrum generation on the sample masked mel spectrum, the sample target emotion features and the sample text based on the spectrum generation sub-model, to obtain predicted mel spectrum of the sample object.

[0043] The parameter adjustment module is configured to perform parameter adjustment on the initial synthesis model based on the masked mel spectrum, the predicted mel spectrum and the sample target emotion features, to obtain a speech synthesis model.

[0044] The synthesis module is configured to perform speech synthesis processing on target text, target object speech with target speech expression features and target emotion information with target non-verbal emotion features based on the speech synthesis model, to obtain target synthesized speech with the target speech expression features and the target emotion features.

[0045] To achieve the above object, a third aspect of embodiments of the present application provides an electronic device, the electronic device comprising a memory and a processor, the memory storing a computer program, and the processor implementing the method according to any one of the first aspect of embodiments of the present application when executing the computer program.

[0046] To achieve the above object, a fourth aspect of embodiments of the present application further provides a computer readable storage medium, the computer readable storage medium storing a computer program, and the computer program implementing the method according to any one of the first aspect of embodiments of the present application when executed by a processor.

[0047] The voice synthesis method, the voice synthesis device, the electronic equipment and the storage medium provided in the embodiments of the present application first acquire sample data, the sample data includes sample text, sample emotional information and sample original mel spectrum, the sample emotional information has sample non-verbal emotional characteristics, and the sample original mel spectrum has sample voice expression characteristics of a sample object. Further, the sample text, the sample emotional information and the sample original mel spectrum are input into an initial synthesis model, the initial synthesis model includes an emotional detection sub-model and a spectrum generation sub-model. The voice of the sample object is converted based on a voice conversion sub-model to obtain sample original mel spectrum. The sample original mel spectrum is subjected to mask processing based on the sample emotional information to obtain a mask mel spectrum and a sample masked mel spectrum, the mask mel spectrum is used to represent the mel spectrum corresponding to the mask region in the sample original mel spectrum, and the sample masked mel spectrum is used to represent the mel spectrum obtained by masking the mel spectrum value of the mask region in the sample original mel spectrum. The sample emotional information and the sample original mel spectrum are subjected to emotional feature detection based on the emotional detection sub-model to obtain sample target emotional features of the sample masked mel spectrum. The sample masked mel spectrum, the sample target emotional features and the sample text are subjected to spectrum generation based on the spectrum generation sub-model to obtain predicted mel spectrum of the sample object. Further, the initial synthesis model is subjected to parameter adjustment based on the mask mel spectrum, the predicted mel spectrum and the sample target emotional features to obtain a voice synthesis model. Further, the target text, the voice of the target object having target voice expression characteristics and the target emotional information having target non-verbal emotional characteristics are subjected to voice synthesis processing based on the voice synthesis model to obtain target synthesized voice having target voice expression characteristics and target emotional features. The voice synthesis method provided in the present application can generate natural emotional voice with non-verbal information in combination with the input target emotional information having target non-verbal emotional features. Compared with the related art which only uses a splicing method to generate emotional voice, the voice synthesis model constructed in the present application can fuse language information (i.e., information containing voice content) and non-voice information (i.e., information not containing voice content) in the voice generation process, and can achieve more accurate control of voice emotion according to the provided voice of the target object and the target emotional information, so as to generate synthesized voice with more accurate emotional expression. BRIEF DESCRIPTION OF DRAWINGS

[0048] Figure 1 is a flowchart of the voice synthesis method provided in the embodiments of the present application;

[0049] Figure 2 is Figure 1 is a flowchart of the specific method of step S140 in

[0050] Figure 3 is Figure 1 is a flowchart of the specific method of step S150 in

[0051] Figure 4 is Figure 3 a flow chart of the specific method of step S320 in

[0052] Figure 5 is Figure 4 a flow chart of the specific method of step S420 in

[0053] Figure 6 is Figure 1 a flow chart of the specific method of step S160 in

[0054] Figure 7 is Figure 6 a flow chart of the specific method of step S650 in

[0055] Figure 8 is a module structure block diagram of a speech synthesis device provided by an embodiment of the present application;

[0056] Figure 9 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION

[0057] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application is further described in detail below in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application, and are not used to limit the present application.

[0058] It should be noted that although the functional modules are divided in the device schematic diagram, and the logical order is shown in the flow chart, in some cases, the steps shown or described can be executed in a manner different from the module division in the device or the order in the flow chart. The terms "first", "second", etc. in the specification and claims and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence.

[0059] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which the present application belongs. The terms used herein are only for the purpose of describing the embodiments of the present application, and are not intended to limit the present application.

[0060] First, the terms involved in the present application are analyzed:

[0061] Artificial Intelligence (AI): is a new technical science that studies and develops theories, methods, technologies and application systems for simulating, extending and expanding human intelligence; Artificial intelligence is a branch of computer science, artificial intelligence attempts to understand the essence of intelligence and produce a new intelligent machine that can react in a similar way to human intelligence, including robots, language recognition, image recognition, natural language processing and expert systems. Artificial intelligence can simulate the information process of human consciousness and thinking. Artificial intelligence is also the theory, method, technology and application system of using digital computer or digital computer controlled machine to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain optimal results.

[0062] Natural Language Processing (NLP): NLP uses computers to process, understand and use human language (such as Chinese, English, etc.), NLP is a branch of artificial intelligence and is a cross-discipline of computer science and linguistics, also known as computational linguistics. Natural language processing includes syntax analysis, semantic analysis, discourse understanding, etc. Natural language processing is commonly used in machine translation, handwritten and printed character recognition, speech recognition and text-to-speech conversion, information retrieval, information extraction and filtering, text classification and clustering, public opinion analysis and opinion mining, etc. Natural language processing involves data mining, machine learning, knowledge acquisition, knowledge engineering, artificial intelligence research related to language processing, and language computing related linguistic research.

[0063] Text-To-Speech (TTS): is a technology from text to speech. TTS generally includes two steps: the first step is text processing, mainly converting text into phoneme sequence and marking the start and end time, frequency change and other information of each phoneme; The second step is speech synthesis, which mainly generates speech according to the phoneme sequence (and the marked start and end time, frequency change and other information).

[0064] Speech prosody: also known as prosody, refers to the changes in pitch, duration, intensity, and tone of speech, which is the expression of the rhythm and prosody of speech. Prosody can include intonation, pause and rhythm, and can convey the speaker's emotions, attitudes and information. The role of speech prosody can express emotions (i.e. through changes in pitch, duration, intensity and tone, different emotions can be expressed), distinguish meaning (i.e. in some languages, the same speech unit can express different meanings through prosodic changes), and enhance the beauty of language (i.e. in artistic forms such as poetry and songs, the use of prosody can make language more beautiful and moving).

[0065] Speech synthesis refers to a process of creating a natural language speech meeting requirements according to given input data. At present, intelligent speech technology is usually applied in intelligent telephone customer service, intelligent sales and other task scenarios of financial technology. However, human speech not only contains language information, but also contains some non-language information such as crying, laughing, pause and coughing. The non-language information can be used to convey different feelings of the speaker and the intention of communication. Therefore, adding some non-language information in the process of synthesizing speech can make the synthesized speech more natural and lively, and closer to the speech in real life.

[0066] Therefore, the related technology usually splices speech information and non-language information in a certain way to synthesize a piece of speech containing non-language information. However, this method does not deeply understand the association between language information and non-language information, so that the connection between speech information and non-language information in the synthesized speech is not natural enough, thereby reducing the accuracy of emotional expression of the synthesized speech. Therefore, how to generate synthesized speech with more accurate emotional expression has become a technical problem to be solved.

[0067] Therefore, the speech synthesis method, speech synthesis device, electronic equipment and storage medium provided by the embodiments of the present application can generate synthesized speech with more accurate emotional expression.

[0068] The embodiments of the present application can acquire and process related data based on artificial intelligence technology. Among them, artificial intelligence (AI) is to use digital computers or machine controlled by digital computers to simulate, extend and expand human intelligence, perceive environment, acquire knowledge and use knowledge to obtain the best results.

[0069] The basic technology of artificial intelligence generally includes technologies such as sensors, special artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction system, mechatronics, etc. The software technology of artificial intelligence mainly includes computer vision technology, robot technology, biometric technology, speech processing technology, natural language processing technology and machine learning / deep learning, etc.

[0070] The speech synthesis method provided in the embodiments of the present application relates to the technical field of artificial intelligence. The speech synthesis method provided in the embodiments of the present application can be applied to a terminal, can be applied to a server, and can also be software running in the terminal or the server. In some embodiments, the terminal can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart watch or the like; the server can be a stand-alone server or a cloud server providing cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDN), and basic cloud computing services such as big data and artificial intelligence platforms; and the software can be an application implementing the speech synthesis method, but is not limited to the above forms.

[0071] The present application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld devices or portable devices, tablet devices, multi-processor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, and the like that perform specific tasks or implement specific abstract data types. The present application can also be practiced in a distributed computing environment, in which tasks are performed by remote processing devices connected by a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.

[0072] It should be noted that in each specific embodiment of the present application, when relevant processing needs to be performed on data related to the identity or characteristics of the object, such as object information, object voice data, object historical data and object identity information, the permission or consent of the object will be obtained first, and the collection, use and processing of the data will comply with relevant laws, regulations and standards. In addition, when the embodiments of the present application need to obtain sensitive personal information of the object, the separate permission or separate consent of the object will be obtained through a pop-up window or a jump to a confirmation page, and after obtaining the separate permission or separate consent of the object, the necessary object-related data for enabling the embodiments of the present application to normally run will be obtained.

[0073] Please refer to Figure 1 , Figure 1 is an optional flowchart of the speech synthesis method provided by the embodiments of the present application. In some embodiments of the present application, the speech synthesis method provided by the embodiments of the present application includes but is not limited to steps S110 to S170:

[0074] In step S110, sample data is obtained, and the sample data includes sample text, sample sentiment information, and sample original mel spectrum;

[0075] In step S120, the sample text, the sample sentiment information, and the sample original mel spectrum are input into an initial synthesis model;

[0076] In step S130, the sample original mel spectrum is subjected to mask processing based on the sample sentiment information, to obtain a mask mel spectrum and a sample mask post-mel spectrum;

[0077] In step S140, sentiment feature detection is performed on the sample sentiment information and the sample original mel spectrum based on a sentiment detection sub-model, to obtain sample target sentiment features of the sample mask post-mel spectrum;

[0078] In step S150, spectrum generation is performed on the sample mask post-mel spectrum, the sample target sentiment features, and the sample text based on a spectrum generation sub-model, to obtain predicted mel spectrum of a sample object;

[0079] In step S160, parameter adjustment is performed on the initial synthesis model based on the mask mel spectrum, the predicted mel spectrum, and the sample target sentiment features, to obtain a speech synthesis model;

[0080] In step S170, speech synthesis processing is performed on target text, target object speech having target speech expression features, and target sentiment information having target non-verbal sentiment features based on the speech synthesis model, to obtain target synthesized speech having target speech expression features and target sentiment features.

[0081] In step S110 of some embodiments, the sample data refers to data pre-acquired for training the model. The present application can first acquire a training sample set including a plurality of sample data, and the training sample set can be acquired from an existing synthesized speech data set as needed, or can be collected and labeled in real time according to a specific collection device. For each sample data, the sample data can include sample text, sample sentiment information, and sample object voice. The sample text has text content, so that the generated voice can express the text content. The sample sentiment information has sample non-verbal emotional features, and the sample sentiment information is equivalent to a prompt information that expresses emotional features in a non-verbal information (such as crying, laughing, pausing, coughing, etc., which can be used to convey different feelings of the speaker and the intention of communication). The sample sentiment information can be a voice containing non-verbal emotional features (such as voice with laughter or crying), or a text prompt such as "the generated voice has a crying sound", without specific limitation. The sample object voice is a prompt voice for learning the prosodic features of the speaker (i.e., the sample object), and the sample voice has sample prosodic features of the sample object. The sample prosodic features can represent the changes in pitch, duration, intensity, and tone of the corresponding sample text.

[0082] It should be noted that, for example, in the intelligent overcoming scene of a financial customer, if a customer service self-introduction voice is to be synthesized, the corresponding sample text can be "I am the intelligent customer service BB of AA application", and the sample sentiment information with sample non-verbal emotional features input at this time can be a voice with laughter and no language content. At this time, the sample object voice can be a voice pre-recorded by the sample object. After the voice synthesis model of the present application, a voice with laughter and in the speaking manner of the sample object can be generated, and the content of the voice is the text content of the sample text.

[0083] It should be noted that, since the sample sentiment information is used to learn the emotional features corresponding to the non-verbal information, the voice content contained in the sample sentiment information is not important and will not be described in detail.

[0084] It should be noted that the sample object voice is used to represent a reference voice for voice synthesis of the text content in the sample text, and is used to learn the prosody features of the synthesized speaker. Therefore, the voice content in the sample object voice is not necessarily the same as the text content in the sample text. The sample text can be a text phoneme sequence obtained by a text phoneme conversion model. The text factor conversion model can use a Deep Voice3 model, a Grapheme to Phoneme (G2P) model, etc. In order to improve the efficiency of the model, the voice content of the sample object voice can also be set to be the same as the sample text, so as to guide the training process of the model according to the sample voice as the target output voice of the sample text for voice synthesis.

[0085] It should be noted that different sample object voices expressed by different prosody can correspond to different voice emotions, such as happiness, anger, sadness, excitement, and coquettishness. The voice emotion can be the emotion expressed in a piece of voice, the emotion expressed in a sentence of voice, or the emotion expressed in a word of voice, etc. The embodiments of the present application do not make specific limitations.

[0086] It should be noted that the storage format of the voice in the present application can be MP3 format, CDA format, WAV format, WMA format, RA format, MIDI format, OGG format, APE format or AAC format, etc. The present application does not make limitations.

[0087] It should be noted that the voice synthesis method of the present application can be used to assist in, for example, car radio and announcement, car navigation, electronic dictionary, intelligent phone, voice assistant, electronic book reading, etc. For example, in the electronic book reading of financial technology, the target voice with the target object voice expression feature, the voice content being the text content of the target text, and the target synthesized voice with the target non-verbal emotional feature can be synthesized according to the input target text, the target object voice with the target voice expression feature, and the target emotional information with the target non-verbal emotional feature.

[0088] It should be noted that the present application considers non-verbal cue information, i.e. emotional information with non-verbal emotional features, when training the voice synthesis model and actually synthesizing the voice, not only the emotional features of the text, so that the generated voice is more natural and lively, and closer to the voice in real life.

[0089] It should be noted that before the sample data is input into the initial synthesis model, the application can first convert the sample object voice into a mel spectrum representation based on a voice conversion model, that is, obtain a sample original mel spectrum, which is a spectral feature representing a sound signal. The voice conversion sub-model can be a model based on a mel filter, and the design of the mel filter can be based on the auditory perception of the human ear to different frequencies, convert the frequency linearly into a mel frequency scale, redefine the frequency axis, and make it more consistent with the perception of the human ear.

[0090] It should be noted that the application can divide the frequency range into a plurality of mel filters, and the common number is between 20 to 40, so as to construct a voice conversion sub-model according to a plurality of mel filters (i.e., a mel filter bank). These mel filters can be in a triangular distribution, covering a range from low frequency to high frequency. The specific process of voice conversion is not described in detail.

[0091] It should be noted that the application can set the voice conversion model in the initial synthesis model, so that the input target object voice can be converted into a mel spectrum representation, which is not described in detail.

[0092] In step S120 of some embodiments, the initial synthesis model includes an emotion detection sub-model, a phoneme detection sub-model, and a spectrum generation sub-model. In addition, the initial synthesis model of the application can be constructed based on a flow matching (Flow-matching) model, a large language model (Large Language Model, LLM, which refers to a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text), a machine learning model, etc., which is not limited here.

[0093] The flow matching model is a specific generative model, and its core idea is to learn a reversible transformation to convert the data distribution into a simple distribution, such as a standard normal distribution, so that the given noise can be converted into the desired data using the reversible transformation.

[0094] In step S130 of some embodiments, further, the application can perform mask processing on the sample original mel spectrum. Specifically, a new voice mel spectrum can be generated by masking a part of the original mel spectrum, that is, a masked mel spectrum and a sample masked mel spectrum are obtained. The masked mel spectrum is used to represent the mel spectrum extracted from the masked area of the sample original mel spectrum, and the sample masked mel spectrum is used to represent the mel spectrum obtained by masking the mel spectrum value of the masked area of the sample original mel spectrum.

[0095] It should be noted that the mask processing of the present application can adopt a random masking manner, that is, the mel spectrum values in the selected mask region of the sample original mel spectrum can be replaced by predicted values (such as 0, 1, etc.), but the dimension of the mel spectrum diagram does not change, only part of the values in the mel spectrum diagram changes.

[0096] It should be noted that the initial synthesis model of the present application uses the remaining part of the sample original mel spectrum (i.e., the sample masked mel spectrum) as input to assist the model in generating speech, and finally calculates the error between the generated speech mel spectrum and the masked part of the sample original mel spectrum (i.e., the mask mel spectrum) as a loss for parameter updating.

[0097] In step S140 of some embodiments, further, the sample original mel spectrum and the sample emotional information can be subjected to emotional detection based on an emotional detection sub-model to obtain sample target emotional features of the sample masked mel spectrum. The present application does not extract phoneme sequences and emotional features from the new sample masked mel spectrum, but the phoneme sequences and emotional features extracted from the sample original mel spectrum contain features of the unmasked part, and the phoneme sequences and emotional features extracted from the sample emotional information contain features of the masked part, so that the generated speech can be controlled to be more similar to the original text content, speech prosody features and emotional features.

[0098] It should be noted that the emotional detection sub-model, the phoneme detection sub-model, and the spectrum generation sub-model of the present application can specifically adopt model structures such as Convolutional Neural Network (CNN), Variational Auto Encoder (VAE), Generative Adversarial Networks (GAN), diffusion model, and autoregressive model, and are not limited here.

[0099] Please refer to Figure 2 , Figure 2 is an optional flowchart of step S140 provided by the embodiments of the present application. In some embodiments of the present application, step S140 can specifically include but is not limited to steps S210 to S230:

[0100] Step S210, performing emotional detection on the sample emotional information to obtain sample non-verbal emotional features;

[0101] Step S220, performing emotional detection on the sample original mel spectrum to obtain sample verbal emotional features;

[0102] In step S230, the sample non-language emotion feature and the sample language emotion feature are spliced to obtain a sample target emotion feature.

[0103] In steps S210 to S230 of some embodiments, the sample non-language emotion feature corresponding to the sample emotion information is used to indicate the emotion feature of the masked mel spectrum. The sample non-language emotion feature refers to an emotion feature extracted from the sample emotion information (i.e., non-language cue information). The sample language emotion feature corresponding to the sample original mel spectrum is used to indicate the emotion feature corresponding to the unmasked mel spectrum. Further, the sample non-language emotion feature and the sample language emotion feature can be spliced to obtain a complete sample target emotion feature.

[0104] It should be noted that the sample non-language emotion feature and the sample language emotion feature detected by the present application have different effects. The sample object voice is equivalent to a speaker voice cue, which is used to control what kind of speaker says the finally generated voice, and the sample emotion information is equivalent to a non-language information cue, which is used to control the emotion of speaking. For example, the laughter of different people is different, and there is a clear difference between A object speaking with a smile and B object speaking with a smile, so the object prosody feature of the generated voice can be controlled to be different according to the input object voice. The sample emotion information can control whether the finally generated voice is A speaking with a smile or A speaking with a cry. Therefore, by splicing the sample non-language emotion feature and the sample language emotion feature, a sample target emotion feature can be obtained, which can completely express the contained language emotion feature and non-language emotion feature.

[0105] In step S150 of some embodiments, further, the present application can generate a spectrum based on the sample masked mel spectrum, the sample target emotion feature, and the sample text to obtain a predicted mel spectrum of the sample object.

[0106] Please refer to Figure 3 , Figure 3 is an optional flowchart of step S150 provided by the embodiments of the present application. In some embodiments of the present application, step S150 can specifically include but is not limited to steps S310 to S320:

[0107] In step S310, a phoneme detection sub-model is used to detect the sample original mel spectrum and the sample text to obtain a sample phoneme sequence.

[0108] In step S320, a spectrum generation sub-model is used to generate a spectrum based on the sample masked mel spectrum, the sample target emotion feature, and the sample phoneme sequence to obtain a predicted mel spectrum.

[0109] In step S310 of some embodiments, the present application can input the sample original mel spectrum of the sample object voice and the sample text into the sample text input phoneme detection sub-model to perform automatic speech recognition, so as to detect a complete sample phoneme sequence.

[0110] It should be noted that the present application performs phoneme detection on the sample original mel spectrum and the sample text based on the phoneme detection sub-model to obtain a sample phoneme sequence, which specifically includes: performing phoneme detection on the sample original mel spectrum to obtain a first sample phoneme sub-sequence; performing phoneme detection on the sample text to obtain a second sample phoneme sub-sequence; and performing feature splicing on the first sample phoneme sub-sequence and the second sample phoneme sub-sequence to obtain the sample phoneme sequence. The first sample phoneme sub-sequence corresponds to the phoneme sequence of the mel spectrum in the sample original mel spectrum that is not masked. The second sample phoneme sub-sequence corresponds to the phoneme sequence of the mel spectrum in the sample original mel spectrum that is masked. Further, the first sample phoneme sub-sequence and the second sample phoneme sub-sequence are spliced to obtain a complete sample phoneme sequence.

[0111] It should be noted that the specific process of phoneme detection can be that the sample object voice and the sample text are input into a phoneme detection sub-model (such as a deep learning model) to process the audio signal. The phoneme detection sub-model can divide the sample object voice and the sample text into different phoneme segments according to spectral features and acoustic characteristics. Further, the phoneme detection sub-model can output phoneme labels in each time period, which indicate the phoneme types in the time period. Finally, all the phoneme labels are arranged into a sample phoneme sub-sequence to represent all the phonemes of the sample phoneme sub-sequence corresponding to the sample object voice and the sample text.

[0112] In step S320 of some embodiments, the spectrum generation sub-model for performing spectrum generation can be constructed based on a flow matching model. After obtaining the complete sample phoneme sequence, the second sample phoneme sub-sequence obtained by detection can be used to perform time length prediction to obtain a target voice time length. Further, the sample masked mel spectrum can be aligned in the time dimension based on the target voice time length.

[0113] Please refer to Figure 4 , Figure 4 is an optional flowchart of step S320 provided by the embodiments of the present application. In some embodiments of the present application, step S320 can specifically include but is not limited to steps S410 to S420:

[0114] In step S410, the sample masked mel spectrum is aligned in the time dimension based on the sample phoneme sequence to obtain a sample aligned mel spectrum.

[0115] In step S420, the spectrum generation sub-model is used to generate the predicted mel-spectrogram based on the sample post-masked mel-spectrogram, the sample target emotion feature, and the sample phoneme sequence.

[0116] In some embodiments, in step S410 and step S420, the sample post-masked mel-spectrogram is equivalent to the mel-spectrogram of the sample object, and the complete sample phoneme sequence can be longer than the sample post-masked mel-spectrogram in the time dimension. Therefore, the sample post-masked mel-spectrogram can be padded with a corresponding number of preset elements (such as 0 or 1, etc.) to align with the complete sample phoneme sequence in the time dimension. Then, the spectrum generation sub-model is used to generate the predicted mel-spectrogram based on the sample post-aligned mel-spectrogram, the sample target emotion feature, and the sample phoneme sequence.

[0117] In the above embodiments, the alignment and generation of the mel-spectrogram from the phoneme sequence ensure the accurate association between the phoneme information and the model output. The generated mel-spectrogram not only reflects the speech content, but also tends to the target emotion, so that the finally generated speech is more natural and expressive. Compared with the related art which only uses splicing to generate emotional speech, the speech synthesis model constructed by the present application can perform feature extraction, splicing fusion, alignment, etc. on the language information (i.e., information containing speech content) and non-speech information (i.e., information not containing speech content) during speech generation, so as to realize more precise control of speech emotion, thereby being able to generate synthesized speech with more accurate emotional expression.

[0118] Please refer to Figure 5 , Figure 5 is an optional flowchart of step S420 provided by the embodiments of the present application. In some embodiments of the present application, step S420 can include, but is not limited to, steps S510 to S520.

[0119] In step S510, the sample target emotion feature is aligned in the time dimension based on the sample phoneme sequence to obtain a sample post-aligned emotion feature.

[0120] In step S520, the spectrum generation sub-model is used to generate the predicted mel-spectrogram based on the sample post-aligned mel-spectrogram, the sample post-aligned emotion feature, and the sample phoneme sequence.

[0121] In steps S510 and S520 of some embodiments, in order to ensure that the output predicted mel-spectrogram not only represents correct phoneme information, but also reflects the specified non-language emotional features, the application can also perform time dimension alignment on the sample target emotional features based on the sample phoneme sequence when performing spectrum generation based on the spectrum generation sub-model, to obtain the sample aligned emotional features. Further, the sample aligned emotional features, the sample aligned mel-spectrogram, and the sample phoneme sequence after alignment are fed as input to the model. In this way, the generated predicted mel-spectrogram can take into account both the text content and the emotional features, so that the finally generated speech expression matches the expected target emotion. Wherein, the predicted mel-spectrogram is the mel-spectrogram with the sample non-language emotional features, with the predicted emotional features, and corresponding to the text content of the sample text.

[0122] It should be noted that in order to improve the quality of the generated mel-spectrogram, the application can also perform some post-processing on the output predicted mel-spectrogram, such as noise reduction, smoothing, etc. In addition, the initial synthesis model of the application also includes a vocoder, and after obtaining the predicted mel-spectrogram, the vocoder can be used to generate speech based on the predicted mel-spectrogram to obtain audible predicted speech.

[0123] In step S160 of some embodiments, after obtaining the predicted mel-spectrogram, the application can perform parameter adjustment on the initial synthesis model based on the mask mel-spectrogram, the predicted mel-spectrogram, and the sample target emotional features to train the initial synthesis model until a preset training end condition is reached to obtain a speech synthesis model.

[0124] Please refer to Figure 6 , Figure 6 is an optional flowchart of step S160 provided by the embodiments of the application. In some embodiments of the application, step S160 can specifically include but is not limited to steps S610 to S650:

[0125] Step S610, performing spectrum loss calculation based on the mask mel-spectrogram and the predicted mel-spectrogram to obtain a spectrum loss value;

[0126] Step S620, performing spectrum conversion on the predicted mel-spectrogram based on the vocoder to obtain predicted speech;

[0127] Step S630, performing emotional feature extraction on the predicted speech to obtain predicted emotional features;

[0128] Step S640, performing emotional feature loss calculation based on the predicted emotional features and the sample target emotional features to obtain an emotional feature loss value;

[0129] Step S650, determine the model loss value based on the spectrum loss value and the emotion feature loss value, and adjust the parameters of the initial synthesis model based on the model loss value to obtain the speech synthesis model.

[0130] In step S610 of some embodiments, the mask mel spectrum (i.e., the real mel spectrum) has sample speech expression features of a sample object, the predicted mel spectrum (which is synthesized by the model according to the phoneme sequence and the emotion features) has predicted speech expression features, and the spectrum loss value is used to represent the difference between the sample speech expression features and the predicted speech expression features. The loss function used in the spectrum loss calculation can be mean square error (MSE) or mean absolute error (MAE), etc., to measure the gap between the spectrum generated by the model and the target spectrum.

[0131] In steps S620 and S630 of some embodiments, the vocoder is a model that converts mel spectrum into audio waveform, and common vocoders include WaveGlow, Parallel WaveGAN, or HiFi-GAN, etc. The predicted mel spectrum is input into the vocoder to generate the corresponding audio signal, i.e., the predicted speech, through its internal algorithm. Further, the predicted emotion features are extracted from the predicted speech, so that the emotion consistency between the predicted emotion features extracted from the generated predicted speech and the sample target emotion features can be evaluated.

[0132] It should be noted that the emotion feature extraction of the predicted speech can use a pre-trained emotion recognition model or an audio processing method, such as a feature extractor based on emotion analysis, to extract features such as pitch, tone, intensity, etc., into a feature vector representing the emotion state.

[0133] In step S640 of some embodiments, emotion feature loss calculation is performed based on the predicted emotion features and the sample target emotion features to compare the emotion of the speech generated by the model, and the emotion feature loss value is obtained.

[0134] In step S650 of some embodiments, further, the model loss value is determined based on the spectrum loss value and the emotion feature loss value, and the parameters of the initial synthesis model are adjusted based on the model loss value to obtain the speech synthesis model, i.e., by integrating the spectrum loss and the emotion feature loss, the model parameters are optimized to improve the quality of speech synthesis.

[0135] In the above embodiments, the present application constitutes a complete generation and optimization process from spectrum generation to speech synthesis, and then to emotion feature extraction and loss calculation, to ensure that the generated speech can achieve the expected effect in terms of content and emotional expression. This comprehensive method provides a good framework for emotional speech synthesis, making the generated speech more natural and expressive.

[0136] It should be noted that the preset training end condition can be that the synthesis accuracy of the initial synthesis model is greater than or equal to a preset accuracy threshold, or the current iteration number of the initial synthesis model reaches a preset iteration number, that is, the current adjusted model can generate complete and user feedback compliant speech. Moreover, the synthesis accuracy is calculated according to the emotional similarity of the predicted emotional feature and the sample target emotional feature, and the spectral similarity of the mask mel spectrum and the predicted mel spectrum, and the function for calculating the similarity can be selected according to actual needs, such as cosine similarity calculation, mean square error loss function, cross-entropy loss function, contrast loss function, etc., which is not specifically limited here.

[0137] Please refer to Figure 7 , Figure 7 is an optional flowchart of step S650 provided by the embodiments of the present application. In some embodiments of the present application, step S650 can specifically include but is not limited to steps S710 to S740:

[0138] Step S710, performing spectral feature extraction on the mask mel spectrum to obtain mask spectral features;

[0139] Step S720, performing feature detection on the mask spectral features based on the sample original mel spectrum to obtain mask feature scores;

[0140] Step S730, determining a first loss weight of a spectral loss value and a second loss weight of an emotional feature loss value based on the mask feature scores;

[0141] Step S740, performing weighted calculation based on the spectral loss value, the first loss weight, the emotional feature loss value and the second loss weight to obtain a model loss value.

[0142] In step S710 of some embodiments, in order to obtain more accurate model loss value and improve the speech synthesis capability of the final trained speech synthesis model, the present application can first select a feature extraction algorithm: a spectral feature extraction method (such as Mel Frequency Cepstrum Coefficient (MFCC), a deep learning model (such as CNN or RNN)) can be used for feature extraction to obtain rich spectral feature description. These features can include the energy, frequency distribution and phase information of the spectrum.

[0143] In step S720 of some embodiments, further, a feature matching, similarity calculation (such as cosine similarity) or convolutional neural network method can be used to calculate the similarity between the mask spectral features and the sample original mel spectrum. Through calculation, a score value is obtained, which represents the importance of the mask spectral features in the sample original mel spectrum. A high score indicates that the mask spectral features are more critical for reconstructing or generating the target spectrum.

[0144] In step S730 of some embodiments, further, a range or ratio can be set to normalize the mask feature scores to obtain the first loss weight (for spectral loss) and the second loss weight (for emotion feature loss), and to ensure that the two loss weights can reflect the importance of different features. These weights can be derived through a simple linear relationship, or adjusted through a nonlinear function (such as a sigmoid function).

[0145] In step S740 of some embodiments, further, based on the spectral loss value, the first loss weight, the emotion feature loss value, and the second loss weight, a weighted calculation is performed to obtain the model loss value, that is, the overall loss value of the model is calculated by comprehensively considering the spectral loss and the emotion feature loss, to provide a basis for model optimization. In this way, the specific calculation of the model loss value can be: model loss value = first loss weight x spectral loss value + second loss weight x emotion feature loss value. This weighted calculation ensures that the model can consider both the spectral quality and the emotional performance during the optimization process. The finally calculated model loss value will be used for backward optimization in the training process to guide the model to adjust the parameters and improve the synthesis effect.

[0146] In the above embodiments, when calculating the model loss value, the present application can combine the spectral loss value and the emotion feature loss value, and assign corresponding weights to them to meet different concerns. Further, the backpropagation algorithm and the optimizer (such as Adam or SGD) can be used to update the parameters of the initial synthesis model according to the model loss value, in order to improve the performance of the generated speech, including the speech quality and the emotional consistency. In this way, the adjusted final model is the optimized speech synthesis model, which can more accurately generate speech with the target emotional state.

[0147] In step S170 of some embodiments, in actual applications such as a smart phone, a speech synthesis device for converting text into speech can be installed on the smart phone, and the speech synthesis device is deployed with the speech synthesis model trained in the present application. When detecting an operation of converting target text into synthesized speech with target prosodic features of target speech, the smart phone can generate a speech synthesis service request and send the speech synthesis service request to the speech synthesis device. By responding to the speech synthesis service request, the smart phone extracts the target text, the target object speech with target speech expression features, and the target emotional information with target non-verbal emotional features from the speech synthesis service request using the speech synthesis device. Further, the speech synthesis model can generate a piece of speech content with text information of the target text, a target speech mel spectrum with a target object speech expression feature close to the target object speech and target non-verbal emotional features in the non-verbal information from the features extracted from the three information, and generate the target synthesized speech from the target speech mel spectrum through a vocoder. That is, after inputting the target text, the target object speech, and the target emotional information into the trained speech synthesis model, the speech synthesis model can first perform mel spectrum conversion on the target object speech, and after obtaining the predicted mel spectrum, generate the target synthesized speech through the vocoder.

[0148] It should be noted that the non-company software tools or components appearing in the embodiments of the present application are only examples for introduction and do not represent actual use.

[0149] The speech synthesis method provided in the embodiments of the present application can generate more lively and natural emotional speech with non-verbal information, such as speech segments with laughter or crying. Compared with the related art which only uses splicing to generate emotional speech, the speech synthesis model in the present application can fuse language information and non-verbal information in the speech generation process, provide diversified non-verbal information prompts, achieve more fine control of speech emotion, guide the generation of emotional expression of speech, and have more cordial communication with users in use scenarios, which can not only respond to speech information but also make corresponding feedback to emotional information. In addition, the speech synthesis model in the present application uses a flow matching model for modeling, which can effectively improve the generation speed. Therefore, the speech synthesis model constructed in the present application can combine the mask and speech conversion to fuse language information (i.e., information containing speech content) and non-speech information (i.e., information not containing speech content) in the speech generation process, achieve more fine control of speech emotion according to the provided target object speech and target emotional information, and thus generate synthesized speech with more accurate emotional expression.

[0150] Please refer to Figure 8 , Figure 8is a schematic diagram of a module structure of a speech synthesis device provided by an embodiment of the present application. In some embodiments of the present application, the speech synthesis device can specifically include:

[0151] The acquisition module 810 is configured to acquire sample data, the sample data including sample text, sample emotional information, and sample original mel-frequency spectrum, the sample emotional information having sample non-verbal emotional features, and the sample original mel-frequency spectrum having sample voice expression features of a sample object.

[0152] The input module 820 is configured to input the sample text, the sample emotional information, and the sample original mel-frequency spectrum into an initial synthesis model, the initial synthesis model including an emotional detection sub-model and a spectrum generation sub-model.

[0153] The mask module 830 is configured to perform mask processing on the sample original mel-frequency spectrum based on the sample emotional information to obtain a mask mel-frequency spectrum and a sample post-mask mel-frequency spectrum, the mask mel-frequency spectrum being used to represent mel-frequency spectrum corresponding to a mask region in the sample original mel-frequency spectrum, and the sample post-mask mel-frequency spectrum being used to represent mel-frequency spectrum obtained by masking mel-frequency spectrum values of the mask region in the sample original mel-frequency spectrum.

[0154] The emotional detection module 840 is configured to perform emotional feature detection on the sample emotional information and the sample original mel-frequency spectrum based on the emotional detection sub-model to obtain sample target emotional features of the sample post-mask mel-frequency spectrum.

[0155] The spectrum generation module 850 is configured to perform spectrum generation on the sample post-mask mel-frequency spectrum, the sample target emotional features, and the sample text based on the spectrum generation sub-model to obtain predicted mel-frequency spectrum of the sample object.

[0156] The parameter adjustment module 860 is configured to perform parameter adjustment on the initial synthesis model based on the mask mel-frequency spectrum, the predicted mel-frequency spectrum, and the sample target emotional features to obtain a speech synthesis model.

[0157] The synthesis module 870 is configured to perform speech synthesis processing on target text, target object voice having target voice expression features, and target emotional information having target non-verbal emotional features based on the speech synthesis model to obtain target synthesized voice having target voice expression features and target emotional features.

[0158] It should be noted that the speech synthesis device of the embodiments of the present application is used to execute the speech synthesis method described above, and the speech synthesis device of the embodiments of the present application corresponds to the speech synthesis method described above. The specific training process is described above, and will not be described here.

[0159] The electronic device provided by the embodiment of the present application also includes a memory and a processor. The memory stores a computer program. The processor executes the computer program to implement the voice synthesis method of the embodiment of the present application.

[0160] The electronic device can be any intelligent terminal including a mobile phone, a tablet computer, a personal digital assistant (PDA), a vehicle-mounted computer, and the like.

[0161] The following describes the electronic device of the embodiment of the present application in detail. Figure 9 The electronic device of the embodiment of the present application is described in detail.

[0162] Please refer to Figure 9 , Figure 9 The hardware structure of the electronic device of another embodiment is illustrated, which includes:

[0163] The processor 910 can be implemented in a general central processing unit (CPU), a microprocessor, an application specific integrated circuit (ASIC), or one or more integrated circuits, and the like, and is used to execute related programs to implement the technical solutions provided by the embodiments of the present application.

[0164] The memory 920 can be implemented in a read only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM), and the like. The memory 920 can store an operating system and other application programs. When the technical solutions provided by the embodiments of the present application are implemented by software or firmware, the related program codes are stored in the memory 920 and are called and executed by the processor 910 to implement the voice synthesis method of the embodiment of the present application.

[0165] The input / output interface 930 is used to implement information input and output.

[0166] The communication interface 940 is used to implement the communication interaction between the device and other devices. The communication can be realized by a wired manner (for example, a USB, a network cable, and the like) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, and the like).

[0167] The bus 950 is used to transmit information between various components (for example, the processor 910, the memory 920, the input / output interface 930, and the communication interface 940) of the device.

[0168] The processor 910, the memory 920, the input / output interface 930, and the communication interface 940 are communicatively connected with each other inside the device through a bus 950.

[0169] The computer readable storage medium stores a computer program, and the computer program is executed by the processor to implement the voice synthesis method of the embodiments of the present application.

[0170] The memory, as a non-transitory computer readable storage medium, can be used to store non-transitory software programs and non-transitory computer executable programs. In addition, the memory can include a high-speed random access memory, and can also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state memory device. In some embodiments, the memory can optionally include a memory remotely arranged with respect to the processor, and these remote memories can be connected to the processor through a network. Examples of the above network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.

[0171] The embodiments described in the embodiments of the present application are used to more clearly illustrate the technical solutions of the embodiments of the present application, and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art can know that, with the evolution of technology and the appearance of new application scenarios, the technical solutions provided by the embodiments of the present application are also applicable to similar technical problems.

[0172] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and can include more or fewer steps than the figures shown, or combine certain steps, or different steps.

[0173] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separated, that is, they can be located in one place, or can be distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application.

[0174] Those skilled in the art can understand that all or some of the steps in the above disclosed method, the functional modules / units in the system and the device can be implemented as software, firmware, hardware and their appropriate combinations.

[0175] The terms "first", "second", "third", "fourth", and the like in the description of this application and in the claims hereof, if any, are used for distinguishing between similar elements and not necessarily for describing a particular sequential or chronological order. It is to be understood that the use of the terms so termed herein is solely for the convenience of the reader and does not limit the scope of the application. It is also to be understood that the description and examples in this application are intended to cover all possible combinations where any of the several elements can represent one or more elements.

[0176] It should be understood that, in this application, "at least one" means one or more, "multiple" means two or more. "And / or" is used to describe the relationship between associated objects, which means that there can be three relationships, for example, "A and / or B" can mean: only A, only B, and A and B exist at the same time, where A and B can be singular or plural. The character " / " generally represents an "or" relationship between the associated objects. "At least one of the following" or similar expressions means any combination of these items, including any combination of single or multiple items. For example, at least one of a, b or c, can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.

[0177] In several embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative, for example, the division of units is only a logical function division, and actual implementation can have another division manner, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the coupling or direct coupling or communication connection between the displayed or discussed each other can be indirect coupling or communication connection through some interfaces, devices or units, which can be electrical, mechanical or other forms.

[0178] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or they can be distributed on multiple network units. According to actual needs, part or all of the units can be selected to achieve the purpose of the embodiment of the present application.

[0179] In addition, each function unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.

[0180] When the integrated unit is realized in the form of a software function unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application, essentially or in the form of a contribution to the prior art, or all or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes multiple instructions for causing an electronic device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods of the embodiments of the present application. The aforementioned storage medium includes various media that can store program codes, such as a U disk, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk.

[0181] The preferred embodiments of the embodiments of the present application are described above with reference to the accompanying drawings, and are not intended to limit the scope of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of the embodiments of the present application.

Claims

1. A speech synthesis method, characterized in that, The method includes: Acquire sample data, which includes sample text, sample sentiment information, and sample raw Mel spectrum. The sample sentiment information has sample non-verbal sentiment features, and the sample raw Mel spectrum has sample speech expression features of the sample object. The sample text, the sample sentiment information, and the sample original Mel spectrum are input into the initial synthesis model, which includes a sentiment detection sub-model and a spectrum generation sub-model. Based on the sample sentiment information, the original Mel spectrum of the sample is masked to obtain a masked Mel spectrum and a sample masked Mel spectrum. The masked Mel spectrum is used to characterize the Mel spectrum corresponding to the masked region in the original Mel spectrum of the sample, and the sample masked Mel spectrum is used to characterize the Mel spectrum obtained by masking the Mel spectrum values ​​of the masked region in the original Mel spectrum of the sample. Based on the emotion detection sub-model, emotion features are detected on the sample emotion information and the original sample Mel spectrum to obtain the sample target emotion features of the masked Mel spectrum; the process of detecting emotion features on the sample emotion information and the original sample Mel spectrum to obtain the sample target emotion features of the masked Mel spectrum includes: performing emotion detection on the sample emotion information to obtain the sample non-verbal emotion features, which are used to indicate the emotion features of the masked Mel spectrum; performing emotion detection on the original sample Mel spectrum to obtain the sample verbal emotion features; and concatenating the sample non-verbal emotion features and the sample verbal emotion features to obtain the sample target emotion features of the masked Mel spectrum; Based on the spectrum generation sub-model, the Mel spectrum of the masked sample, the sentiment features of the sample target, and the sample text are used to generate the spectrum of the sample object, thereby obtaining the predicted Mel spectrum of the sample object. Based on the masked Mel spectrum, the predicted Mel spectrum, and the target emotion features of the sample, the parameters of the initial synthesis model are adjusted to obtain a speech synthesis model. Based on the speech synthesis model, speech synthesis processing is performed on the target text, the target object speech with the target speech expression features, and the target emotional information with the target non-linguistic emotional features to obtain target synthesized speech with the target speech expression features and the target emotional features.

2. The method according to claim 1, characterized in that, The initial synthesis model further includes a phoneme detection sub-model. The spectrum generation sub-model generates a spectrum based on the masked Mel spectrum of the sample, the target sentiment features of the sample, and the sample text to obtain the predicted Mel spectrum of the sample object, including: Based on the phoneme detection sub-model, phoneme detection is performed on the original Mel spectrum of the sample and the sample text to obtain the sample phoneme sequence; Based on the spectrum generation sub-model, the Mel spectrum of the masked sample, the target emotion features of the sample, and the phoneme sequence of the sample are used to generate the spectrum, thereby obtaining the predicted Mel spectrum.

3. The method according to claim 2, characterized in that, The process of generating a predicted Mel spectrum based on the spectrum generation sub-model for the masked Mel spectrum of the sample, the target sentiment features of the sample, and the phoneme sequence of the sample includes: Based on the sample phoneme sequence, the sample masked Mel spectrum is aligned in the time dimension to obtain the sample aligned Mel spectrum. Based on the spectrum generation sub-model, the spectral generation is performed on the aligned Mel spectrum of the sample, the target sentiment features of the sample, and the phoneme sequence of the sample to obtain the predicted Mel spectrum.

4. The method according to claim 3, characterized in that, The process of generating the predicted Mel spectrum based on the spectrum generation sub-model for the aligned Mel spectrum of the sample, the target sentiment features of the sample, and the phoneme sequence of the sample includes: Based on the sample phoneme sequence, the target emotional features of the sample are aligned in the time dimension to obtain the aligned emotional features of the sample. Based on the spectrum generation sub-model, the spectral generation is performed on the aligned Mel spectrum of the samples, the aligned sentiment features of the samples, and the sample phoneme sequence to obtain the predicted Mel spectrum.

5. The method according to any one of claims 1 to 4, characterized in that, The initial synthesis model further includes a vocoder. The parameter adjustment of the initial synthesis model based on the masked Mel spectrum, the predicted Mel spectrum, and the sample target emotion features yields a speech synthesis model, including: Spectral loss is calculated based on the masked Mel spectrum and the predicted Mel spectrum to obtain a spectral loss value. The masked Mel spectrum has the sample speech expression features, and the predicted Mel spectrum has the predicted speech expression features. The spectral loss value is used to characterize the degree of difference between the sample speech expression features and the predicted speech expression features. Based on the vocoder, the predicted Mel spectrum is converted to obtain the predicted speech. Emotional features are extracted from the predicted speech to obtain predicted emotional features; Based on the predicted sentiment features and the target sentiment features of the sample, the sentiment feature loss is calculated to obtain the sentiment feature loss value; The model loss value is determined based on the spectral loss value and the emotional feature loss value, and the parameters of the initial synthesis model are adjusted based on the model loss value to obtain the speech synthesis model.

6. The method according to claim 5, characterized in that, The process of determining the model loss value based on the spectral loss value and the sentiment feature loss value includes: The mask Mel spectrum is subjected to spectral feature extraction to obtain the mask spectral features; Based on the original Mel spectrum of the sample, feature detection is performed on the mask spectral features to obtain a mask feature score, which is used to characterize the importance of the mask spectral features in the original Mel spectrum of the sample. Based on the mask feature score, determine the first loss weight of the spectral loss value and the second loss weight of the sentiment feature loss value; The model loss value is obtained by weighting the spectral loss value, the first loss weight, the emotional feature loss value, and the second loss weight.

7. A speech synthesis device, characterized in that, The device includes: The acquisition module is used to acquire sample data, which includes sample text, sample sentiment information, and sample raw Mel spectrum. The sample sentiment information has sample non-verbal sentiment features, and the sample raw Mel spectrum has sample speech expression features of the sample object. The input module is used to input the sample text, the sample sentiment information, and the sample original Mel spectrum into the initial synthesis model, the initial synthesis model including a sentiment detection sub-model and a spectrum generation sub-model; The masking module is used to perform masking processing on the original Mel spectrum of the sample based on the sample sentiment information to obtain the masked Mel spectrum and the sample masked Mel spectrum. The masked Mel spectrum is used to characterize the Mel spectrum corresponding to the masked region in the original Mel spectrum of the sample, and the sample masked Mel spectrum is used to characterize the Mel spectrum obtained by masking the Mel spectrum value of the masked region in the original Mel spectrum of the sample. The sentiment detection module is used to perform sentiment feature detection on the sample sentiment information and the original Mel spectrum of the sample based on the sentiment detection sub-model to obtain the sample target sentiment feature of the masked Mel spectrum. The step of performing sentiment feature detection on the sample sentiment information and the original Mel spectrum of the sample based on the sentiment detection sub-model to obtain the sample target sentiment feature of the masked Mel spectrum includes: performing sentiment detection on the sample sentiment information to obtain the sample non-verbal sentiment feature, which is used to indicate the sentiment feature of the masked Mel spectrum; performing sentiment detection on the original Mel spectrum of the sample to obtain the sample verbal sentiment feature; and concatenating the sample non-verbal sentiment feature and the sample verbal sentiment feature to obtain the sample target sentiment feature of the masked Mel spectrum. The spectrum generation module is used to generate a spectrum based on the masked Mel spectrum of the sample, the target sentiment features of the sample, and the sample text, according to the spectrum generation sub-model, to obtain the predicted Mel spectrum of the sample object. The parameter adjustment module is used to adjust the parameters of the initial synthesis model based on the masked Mel spectrum, the predicted Mel spectrum, and the sample target emotion features to obtain a speech synthesis model. The synthesis module is used to perform speech synthesis processing on the target text, the target object speech with target speech expression features and the target emotional information with target non-linguistic emotional features based on the speech synthesis model, so as to obtain target synthesized speech with the target speech expression features and the target emotional features.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method as described in any one of claims 1 to 6.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 6.

Citation Information

Patent Citations

  • Audio editing method and device, electronic equipment and storage medium

    CN116403564A

  • Digital human expression generation method and device

    CN118657863A