Speech synthesis method and device, electronic equipment and storage medium

Through the combination of feature coding sub-model, prosody control sub-model and prosody generation sub-model, the problem that prosody synthesis methods in the prior art is difficult to accurately control prosody characteristics, and more precise speech generation of specific emotions or styles is achieved, ensuring that the synthetic speech is natural and meets user needs.

CN120299448APending Publication Date: 2025-07-11PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510526884.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-07-11

AI Technical Summary

Technical Problem

Existing pronunciation methods are difficult to achieve precise control of pronunciation characteristics, resulting in the inadequate expression of speech that generates specific emotions or styles.

Method used

Using the combination of feature coding sub-model, pronunciation control sub-model and speech generation sub-model, by obtaining phonological pronunciation sketches and target text, feature extraction, text encoding, vector splicing, pronunciation outline prediction and adjustment are performed, and the target feature vector is finally generated and audio synthesis is performed.

Benefits of technology

It improves the accuracy of voice emotion or style generation, making synthesized voice more in line with user needs and more natural.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120299448A_ABST
    Figure CN120299448A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a speech synthesis method and device, electronic equipment and a storage medium, belongs to the technical field of speech synthesis, and is suitable for the fields of financial science and technology and medical treatment. The method comprises the following steps: performing feature extraction on a voice rhythm sketch based on a feature coding sub-model to obtain a rhythm feature vector; based on the rhythm feature vector, performing text coding on the target text to obtain a text vector; based on the feature coding sub-model, performing vector splicing on the text vector and the rhythm feature vector to obtain a spliced feature vector; based on the rhythm control sub-model and the rhythm feature vector, performing rhythm contour prediction to obtain a rhythm feature optimization vector; based on the rhythm control sub-model and the rhythm feature optimization vector, performing rhythm adjustment on the spliced feature vector to obtain a target feature vector; and based on the voice generation sub-model, performing audio generation on the target feature vector. According to the embodiment of the invention, more accurate specific emotion or style can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of speech synthesis technology, and is applicable to the fields of fintech and healthcare. In particular, it relates to a speech synthesis method, apparatus, electronic device, and storage medium. Background Art

[0002] Speech synthesis refers to the technology of converting text information into voice output. For example, in the field of fintech, a financial intelligent customer service can perform speech synthesis on the reply text to complete the reply to the customer's questions, which can help customers quickly understand financial-related information. For example, in the field of healthcare, a doctor can use a medical voice assistant to perform speech synthesis on a patient's examination report, so that the patient's examination report is output through voice, thereby helping the doctor quickly understand the patient's information.

[0003] Currently, speech synthesis methods usually use a speech synthesis model to analyze a large amount of speech data, learn the prosody patterns and rules in the speech data, and then generate speech based on the learned prosody patterns and rules. However, this method is difficult to achieve precise control of prosody features, resulting in less than ideal performance when generating speech with specific emotions or styles. Therefore, how to generate more precise specific emotions or styles has become a technical problem to be solved urgently. Summary of the Invention

[0004] The main purpose of the embodiments of this application is to propose a speech synthesis method, apparatus, electronic device, and storage medium, aiming to generate more precise specific emotions or styles.

[0005] To achieve the above object, in the first aspect of the embodiments of this application, a speech synthesis method is proposed. The method includes:

[0006] Obtain a speech prosody sketch and obtain a target text;

[0007] Obtain a target speech synthesis model, where the target speech synthesis model includes a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model;

[0008] Based on the feature encoding sub-model, extract features from the speech prosody sketch to obtain a prosody feature vector;

[0009] Based on the feature encoding sub-model and the prosody feature vector, perform text encoding on the target text to obtain a text vector;

[0010] Based on the feature encoding sub-model, splice the text vector and the prosody feature vector to obtain a spliced feature vector;

[0011] Based on the prosody control sub-model and the prosody feature vector, perform prosody contour prediction to obtain an optimized prosody feature vector;

[0012] Based on the prosody control sub-model and the prosody feature optimization vector, perform prosody adjustment on the spliced feature vector to obtain a target feature vector;

[0013] Based on the speech generation sub-model, perform audio generation on the target feature vector.

[0014] In some embodiments, the obtaining of the target speech synthesis model includes:

[0015] Obtain synthetic speech sample data, synthetic text sample data, and prosody sketch sample data;

[0016] Obtain an initial speech synthesis model, where the initial speech synthesis model includes the feature encoding sub-model, an initial prosody control sub-model, and an initial speech generation sub-model;

[0017] Based on the feature encoding sub-model, perform prosody feature extraction on the synthetic speech sample data to obtain synthetic speech prosody features;

[0018] Based on the feature encoding sub-model, perform speech encoding on the synthetic speech sample data to obtain a synthetic speech vector;

[0019] Based on the feature encoding sub-model, perform text encoding extraction on the synthetic text sample data to obtain a synthetic text vector;

[0020] Based on the feature encoding sub-model, perform feature extraction on the prosody sketch sample data to obtain sketch prosody features;

[0021] Based on the initial prosody control sub-model, perform supervised learning on the synthetic speech prosody features and the sketch prosody features to obtain the prosody control sub-model;

[0022] Based on the initial speech generation sub-model, perform supervised learning on the synthetic text vector, the sketch prosody features, and the synthetic speech vector to obtain the speech generation sub-model;

[0023] Merge the feature encoding sub-model, the prosody control sub-model, and the speech generation sub-model to obtain the target speech synthesis model.

[0024] In some embodiments, the based on the initial prosody control sub-model, perform supervised learning on the synthetic speech prosody features and the sketch prosody features to obtain the prosody control sub-model includes:

[0025] Based on the initial prosody control sub-model, perform mapping relationship learning on the synthetic speech prosody features and the sketch prosody features to obtain sketch contour mapping information;

[0026] Based on the sketched contour mapping information, parameter adjustment is performed on the initial prosody control sub-model to obtain the prosody control sub-model.

[0027] In some embodiments, the method of performing supervised learning on the synthesized text vector, the sketched prosody features, and the synthesized speech vector based on the initial speech generation sub-model to obtain the speech generation sub-model includes:

[0028] Based on the feature encoding sub-model, vector merging is performed on the synthesized text vector and the sketched prosody features to obtain a comprehensive feature vector;

[0029] Based on the initial prosody control sub-model and the synthesized speech prosody features, prosody feature adjustment is performed on the comprehensive feature vector to obtain a comprehensively adjusted feature vector;

[0030] Based on the initial speech generation sub-model, feature optimization is performed on the comprehensively adjusted feature vector to obtain an optimized speech feature vector;

[0031] Based on the synthesized speech vector and the optimized speech feature vector, parameter adjustment is performed on the initial speech generation sub-model to obtain the speech generation sub-model.

[0032] In some embodiments, the method of performing text encoding on the target text based on the feature encoding sub-model and the prosody feature vector to obtain a text vector includes:

[0033] Based on the feature encoding sub-model, semantic encoding is performed on the target text to obtain a semantic vector;

[0034] Based on the feature encoding sub-model and the prosody feature vector, phoneme duration prediction is performed on the target text to obtain a phoneme duration feature vector;

[0035] Vector merging is performed on the semantic vector and the phoneme duration feature vector to obtain the text vector.

[0036] In some embodiments, the method of performing phoneme duration prediction on the target text based on the feature encoding sub-model and the prosody feature vector to obtain a phoneme duration feature vector includes:

[0037] Based on the feature encoding sub-model, phoneme duration mapping is performed on the target text to obtain an initial phoneme duration feature vector;

[0038] Based on the feature encoding sub-model, feature parsing is performed on the prosody feature vector to obtain a speech rhythm target;

[0039] Based on the feature encoding sub-model and the speech rhythm target, adjust the duration of the initial phoneme duration feature vector to obtain the phoneme duration feature vector.

[0040] In some embodiments, the obtaining of the speech rhythm target by performing feature parsing on the prosody feature vector based on the speech text and the feature encoding sub-model includes:

[0041] Based on the feature encoding sub-model, perform feature quantization on the prosody feature vector to obtain a prosody quantization feature;

[0042] Based on the feature encoding sub-model, perform normalization processing on the prosody quantization feature to obtain a normalized prosody feature;

[0043] Based on the feature encoding sub-model, perform prosody pattern recognition on the normalized prosody feature to obtain a speech prosody pattern;

[0044] Based on the speech prosody pattern, set the speech rhythm target.

[0045] To achieve the above object, a second aspect of the embodiments of the present application proposes a speech synthesis device, the device includes:

[0046] A data acquisition module, configured to acquire a speech prosody sketch and acquire a target text;

[0047] A model acquisition module, configured to acquire a target speech synthesis model, wherein the target speech synthesis model includes a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model;

[0048] A feature extraction module, configured to perform feature extraction on the speech prosody sketch based on the feature encoding sub-model to obtain a prosody feature vector;

[0049] A text encoding module, configured to perform text encoding on the target text based on the feature encoding sub-model and the prosody feature vector to obtain a text vector;

[0050] A vector splicing module, configured to splice the text vector and the prosody feature vector based on the feature encoding sub-model to obtain a spliced feature vector;

[0051] A feature optimization module, configured to perform prosody contour prediction based on the prosody control sub-model and the prosody feature vector to obtain a prosody feature optimization vector;

[0052] A vector adjustment module, configured to perform prosody adjustment on the spliced feature vector based on the prosody control sub-model and the prosody feature optimization vector to obtain a target feature vector;

[0053] An audio generation module, configured to generate audio based on the target feature vector by using the speech generation sub-model.

[0054] To achieve the above object, a third aspect of the embodiments of the present application provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method described in the first aspect above is implemented.

[0055] To achieve the above object, a fourth aspect of the embodiments of the present application provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the method described in the first aspect above is implemented.

[0056] The speech synthesis method, device, electronic device, and storage medium provided in the present application obtain a speech prosody sketch, obtain a target text, and obtain a target speech synthesis model including a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model, thereby clarifying the input conditions and model conditions for speech synthesis. Secondly, according to the feature encoding sub-model, feature extraction is performed on the speech prosody sketch to obtain a prosody feature vector, and according to the feature encoding sub-model and the prosody feature vector, text encoding is performed on the target text to obtain a text vector. Then, according to the feature encoding sub-model, vector splicing is performed on the text vector and the prosody feature vector to obtain a spliced feature vector, so that the spliced feature vector has both the semantic information of the text vector and the prosody expression ability of the prosody feature vector. Finally, according to the prosody control sub-model and the prosody feature vector, prosody contour prediction is performed to obtain a prosody feature optimization vector, and according to the prosody control sub-model and the prosody feature optimization vector, prosody adjustment is performed on the spliced feature vector to obtain a target feature vector. Based on the speech generation sub-model, audio is generated from the target feature vector to obtain a synthesized audio, making the synthesized speech more in line with user needs, improving the generation accuracy of speech emotion or style. In addition, it can also make the synthesized speech data more natural. Description of the Drawings

[0057] Figure 1 is a flowchart of the speech synthesis method provided by the embodiments of the present application;

[0058] Figure 2 is Figure 1 a flowchart of step S102 in

[0059] Figure 3 is Figure 2 a flowchart of step S207 in

[0060] Figure 4 is Figure 2 a flowchart of step S208 in

[0061] Figure 5 is Figure 1 the flowchart of step S104 in

[0062] Figure 6 is Figure 5 the flowchart of step S502 in

[0063] Figure 7 is Figure 6 the flowchart of step S602 in

[0064] Figure 8 is the structural schematic diagram of the speech synthesis device provided by the embodiments of the present application;

[0065] Figure 9 is the hardware structural schematic diagram of the electronic device provided by the embodiments of the present application. Detailed implementation manners

[0066] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0067] It should be noted that although the functional modules are divided in the device schematic diagram and the logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order from the module division in the device or the order in the flowchart. Terms such as "first", "second", etc. in the description, claims and the above drawings are used to distinguish similar objects and do not necessarily need to describe a specific order or sequence.

[0068] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0069] First, several nouns involved in the present application are analyzed:

[0070] Speech synthesis system: A speech synthesis system is a technical system that converts text information into speech output through computer technology. A speech synthesis system usually includes modules such as a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model. The feature encoding sub-model is responsible for extracting semantic information from the text, the prosody control sub-model is used to adjust the prosody features of the speech (such as pitch, rhythm, intensity, etc.), and the speech generation sub-model converts these features into the final speech waveform. The speech synthesis system can imitate the natural fluency of human speech and is widely used in fields such as fintech and healthcare, such as in scenarios like intelligent customer service and voice assistants.

[0071] Speech synthesis refers to the technology of converting text information into voice output. For example, in the field of fintech, financial intelligent customer service can perform speech synthesis on the reply text to complete the reply to customers' questions, which can help customers quickly understand financial-related information. For example, in the medical field, doctors can use medical voice assistants to perform speech synthesis on patients' examination reports, so that the patients' examination reports are output through voice, thereby helping doctors quickly understand patients' information.

[0072] Currently, speech synthesis methods usually use speech synthesis models to analyze a large amount of speech data, learn the prosody patterns and rules in the speech data, and then generate speech based on the learned prosody patterns and rules. However, this method is difficult to achieve precise control of prosody features, resulting in less-than-ideal performance when generating speech with specific emotions or styles. Therefore, how to generate more precise specific emotions or styles has become a technical problem to be solved urgently.

[0073] Based on this, the embodiments of this application provide a speech synthesis method, device, electronic device, and storage medium, aiming to generate more precise specific emotions or styles.

[0074] The speech synthesis method, device, electronic device, and storage medium provided by the embodiments of this application are specifically described through the following embodiments. First, the speech synthesis method in the embodiments of this application is described.

[0075] The embodiments of this application can acquire and process relevant data based on artificial intelligence technology. Among them, artificial intelligence (AI) is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results.

[0076] Artificial intelligence basic technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, big data processing technology, operation / interaction systems, and mechatronics. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, robotics, biometric technology, speech processing technology, natural language processing technology, and machine learning / deep learning.

[0077] The speech synthesis method provided by the embodiments of this application relates to the field of speech synthesis technology and is applicable to the fields of fintech and healthcare. The speech synthesis method provided by the embodiments of this application can be applied to a terminal, or to a server side, or can be software running on a terminal or a server side. In some embodiments, the terminal can be a smart phone, a tablet computer, a laptop computer, a desktop computer, etc.; the server side can be configured as an independent physical server, or can be configured as a server cluster or a distributed system composed of multiple physical servers, or can also be configured as a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms; the software can be an application that implements the speech synthesis method, etc., but is not limited to the above forms.

[0078] This application can be used in many general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and so on. This application can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. This application can also be practiced in a distributed computing environment where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules can be located in local and remote computer storage media including storage devices.

[0079] It should be noted that in each specific embodiment of this application, when it comes to performing relevant processing based on data related to the user's identity or characteristics, such as user information, user behavior data, user historical data, and user location information, etc., the user's permission or consent will be obtained first. Moreover, the collection, use, and processing of these data will comply with relevant laws, regulations, and standards. In addition, when the embodiments of this application need to obtain the user's sensitive personal information, the user's separate permission or separate consent will be obtained through methods such as pop-up windows or redirecting to a confirmation page. After clearly obtaining the user's separate permission or separate consent, the necessary user-related data for the normal operation of the embodiments of this application will be obtained.

[0080] Figure 1 is an optional flowchart of the speech synthesis method provided by the embodiments of this application and is applicable to a speech synthesis system. Figure 1 The method in may include but is not limited to steps S101 to S108.

[0081] Step S101, obtain a speech prosody sketch and obtain the target text;

[0082] Step S102, obtain the target speech synthesis model, where the target speech synthesis model includes a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model;

[0083] Step S103, based on the feature encoding sub-model, extract features from the speech prosody sketch to obtain a prosody feature vector;

[0084] Step S104, based on the feature encoding sub-model and the prosody feature vector, perform text encoding on the target text to obtain a text vector;

[0085] Step S105, based on the feature encoding sub-model, splice the text vector and the prosody feature vector to obtain a spliced feature vector;

[0086] Step S106, based on the prosody control sub-model and the prosody feature vector, perform prosody contour prediction to obtain an optimized prosody feature vector;

[0087] Step S107, based on the prosody control sub-model and the optimized prosody feature vector, perform prosody adjustment on the spliced feature vector to obtain a target feature vector;

[0088] Step S108, based on the speech generation sub-model, generate audio from the target feature vector.

[0089] Steps S101 to S108 shown in the embodiments of the present application, by obtaining a speech prosody sketch, obtaining the target text, and obtaining the target speech synthesis model including a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model, clarify the input conditions and model conditions for speech synthesis. Secondly, according to the feature encoding sub-model, extract features from the speech prosody sketch to obtain a prosody feature vector, and according to the feature encoding sub-model and the prosody feature vector, perform text encoding on the target text to obtain a text vector. Then, according to the feature encoding sub-model, splice the text vector and the prosody feature vector to obtain a spliced feature vector, so that the spliced feature vector has both the semantic information of the text vector and the prosody expression ability of the prosody feature vector. Finally, according to the prosody control sub-model and the prosody feature vector, perform prosody contour prediction to obtain an optimized prosody feature vector, and according to the prosody control sub-model and the optimized prosody feature vector, perform prosody adjustment on the spliced feature vector to obtain a target feature vector, and based on the speech generation sub-model, generate audio from the target feature vector to obtain the synthesized audio, making the synthesized speech more in line with user needs, improving the generation accuracy of speech emotion or style. In addition, it can also make the synthesized speech data more natural.

[0090] In step S101 of some embodiments, a speech prosody sketch refers to a graphical representation containing prosody information such as pitch, rhythm, and intensity. For example, in the intelligent customer service scenario in the fintech field, when a customer consults a financial product, the speech prosody sketch of the customer service voice prompt may include information such as a moderate speech rate and a stable intonation. In the speech rehabilitation training scenario in the medical field, when a patient conducts speech practice, the speech prosody sketch of the practice speech may include information such as a gradually rising intonation. The target text refers to the content that needs to be synthesized into speech. In the financial product introduction scenario in the fintech field, the target text may be the instruction manual of a financial product. In the doctor's order issuing scenario in the medical field, the target text may be the doctor's order information written by a doctor.

[0091] In the embodiments of the present application, the speech synthesis system can obtain the target text by receiving the text input by the user. Similarly, the user can obtain the text prosody feature by setting prosody information such as pitch, rhythm, and intensity, and then input the text prosody feature into the speech synthesis system to generate a graphical representation, thereby obtaining the speech prosody sketch.

[0092] In step S102 of some embodiments, the target speech synthesis model refers to a model that converts the target text and the speech prosody sketch into speech data. The feature encoding sub-model refers to the model part that extracts features from the speech prosody sketch and the target text and encodes them. The prosody control sub-model refers to the model part that adjusts the prosody features of the synthesized speech. The speech generation sub-model refers to the model part that converts speech features into speech data.

[0093] In the embodiments of the present application, an initial speech synthesis model can be constructed, and the initial speech synthesis model can be trained, and the model parameters can be adjusted so that the initial speech synthesis model is converted into a target speech synthesis model suitable for converting text and prosody sketches into speech data.

[0094] Specifically, please refer to Figure 2 , in some embodiments, step S102 may include but is not limited to steps S201 to S209:

[0095] Step S201, obtaining synthetic speech sample data, synthetic text sample data, and prosody sketch sample data;

[0096] Step S202, obtaining an initial speech synthesis model, where the initial speech synthesis model includes a feature encoding sub-model, an initial prosody control sub-model, and an initial speech generation sub-model;

[0097] Step S203, based on the feature encoding sub-model, extracting prosody features from the synthetic speech sample data to obtain synthetic speech prosody features;

[0098] Step S204: Based on the feature encoding sub-model, perform speech encoding on the synthesized speech sample data to obtain a synthesized speech vector;

[0099] Step S205: Based on the feature encoding sub-model, perform text encoding extraction on the synthesized text sample data to obtain a synthesized text vector;

[0100] Step S206: Based on the feature encoding sub-model, perform feature extraction on the prosody sketch sample data to obtain sketch prosody features;

[0101] Step S207: Based on the initial prosody control sub-model, perform supervised learning on the synthesized speech prosody features and sketch prosody features to obtain a prosody control sub-model;

[0102] Step S208: Based on the initial speech generation sub-model, perform supervised learning on the synthesized text vector, sketch prosody features, and synthesized speech vector to obtain a speech generation sub-model;

[0103] Step S209: Merge the feature encoding sub-model, prosody control sub-model, and speech generation sub-model to obtain a target speech synthesis model.

[0104] In steps S201 and S202 of some embodiments, the synthesized speech sample data refers to the speech data used to train the speech synthesis model, and the synthesized speech sample data includes speech content and the prosody features corresponding to the speech content. The synthesized text sample data refers to the text data used to train the speech synthesis model. The prosody sketch sample data refers to the prosody sketch data used to train the speech synthesis model, and the prosody sketch sample data includes graphical representations of prosody features such as pitch, rhythm, and intensity. The initial speech synthesis model refers to an untrained speech synthesis model. The initial prosody control sub-model refers to an untrained model part for prosody adjustment and optimization of speech data. The initial speech generation sub-model refers to an untrained model part for converting speech features into audio data.

[0105] In the embodiments of the present application, a text editor can be used to organize the text data to obtain synthesized text sample data, and a data collection tool can be used to annotate the speech data to obtain synthesized speech sample data. Further, according to the mapping relationship between the synthesized speech sample data and the synthesized text sample data, prosody sketch sample data can be manually drawn.

[0106] In another embodiment of the present application, an initial speech synthesis model can be obtained by selecting a model framework suitable for speech synthesis and then splicing the model framework.

[0107] In steps S203 to S206 of some embodiments, the synthetic speech prosody features refer to the prosody features extracted from the synthetic speech sample data, such as pitch, rhythm, intensity, etc. The synthetic speech vector refers to the vector form of the synthetic speech sample data. The synthetic text vector refers to the vector form of the synthetic text sample data. The sketch prosody features refer to the prosody feature vectors in the prosody sketch sample data.

[0108] In the embodiments of the present application, the prosody features of the synthetic speech sample data and the prosody sketch sample data can be extracted according to the feature encoding sub-model in the initial speech synthesis model, and the synthetic speech prosody features and the sketch prosody features are obtained respectively. Then, according to the encoding function of the feature encoding sub-model, the synthetic speech sample data is converted into a synthetic speech vector, and the synthetic text sample data is converted into a synthetic text vector.

[0109] In step S207 of some embodiments, the initial prosody control sub-model adjusts its own parameters by learning the feature mapping relationship from the sketch prosody features to the synthetic speech prosody features, so as to obtain the prosody control sub-model.

[0110] Specifically, please refer to Figure 3 In some embodiments, step S207 may include but is not limited to steps S301 to S302:

[0111] Step S301: Based on the initial prosody control sub-model, learn the mapping relationship between the synthetic speech prosody features and the sketch prosody features to obtain the sketch contour mapping information;

[0112] Step S302: Based on the sketch contour mapping information, adjust the parameters of the initial prosody control sub-model to obtain the prosody control sub-model.

[0113] In step S301 of some embodiments, the sketch contour mapping information refers to the feature mapping relationship between the sketch prosody features and the synthetic speech prosody features.

[0114] In the embodiments of the present application, the initial prosody control sub-model determines the mapping relationship between the synthetic speech prosody features and the sketch prosody features through the data mapping relationship between the synthetic speech sample data and the prosody sketch sample data, and then learns the process of converting the sketch prosody features into the synthetic speech prosody features, and can learn the sketch contour mapping information.

[0115] In step S302 of some embodiments, during the process of learning the sketch contour mapping information, the initial prosody control sub-model can gradually adjust its own parameters according to the learned knowledge, so that the initial prosody control sub-model is converted into a prosody control sub-model capable of converting the sketch prosody features into the synthetic speech prosody features.

[0116] In steps S301 to S302 illustrated in the embodiments of the present application, the initial prosody control sub-model learns the mapping relationship between the prosody features of the synthesized speech and the sketch prosody features to obtain the sketch contour mapping information. Then, based on the sketch contour mapping information, the parameters of the initial prosody control sub-model are adjusted to obtain the prosody control sub-model, improving the accuracy of the prosody control sub-model and enabling the speech synthesis model to generate more natural and user-requested audio data during the synthesis of speech.

[0117] In step S208 of some embodiments, the synthesized text vector and the sketch prosody features are vector-merged to obtain a comprehensive feature vector. The initial prosody control sub-model can then adjust its own parameters by learning the process of adjusting the comprehensive feature vector according to the prosody features of the synthesized speech, thereby obtaining the speech generation sub-model.

[0118] Specifically, please refer to Figure 4 , in some embodiments, step S208 may include but is not limited to steps S401 to S404:

[0119] Step S401: Based on the feature encoding sub-model, the synthesized text vector and the sketch prosody features are vector-merged to obtain a comprehensive feature vector;

[0120] Step S402: Based on the initial prosody control sub-model and the prosody features of the synthesized speech, the prosody features of the comprehensive feature vector are adjusted to obtain a comprehensively adjusted feature vector;

[0121] Step S403: Based on the initial speech generation sub-model, the comprehensively adjusted feature vector is feature-optimized to obtain an optimized speech feature vector;

[0122] Step S404: Based on the synthesized speech vector and the optimized speech feature vector, the parameters of the initial speech generation sub-model are adjusted to obtain the speech generation sub-model.

[0123] In step S401 of some embodiments, the comprehensive feature vector refers to the vector representation after the merger of the synthesized text vector and the sketch prosody features. It should be noted that the comprehensive feature vector may include the semantic information of the synthesized text vector and the prosody information of the sketch prosody features.

[0124] In the embodiments of the present application, the synthesized text vector and the sketch prosody features can be vector-concatenated to obtain the comprehensive feature vector, thereby improving the semantic expression ability and prosody expression ability of the comprehensive feature vector.

[0125] In step S402 of some embodiments, the initial prosody control sub-model modifies the sketch prosody feature part in the above comprehensive feature vector according to the input synthetic speech prosody features, so that the sketch prosody feature part in the comprehensive feature vector is closer to the synthetic speech prosody features, thereby improving the prosody expression ability of the comprehensive adjustment feature vector.

[0126] In step S403 of some embodiments, after obtaining the comprehensive adjustment feature vector, the initial speech generation sub-model can be used to denoise the comprehensive adjustment feature vector, thereby reducing the influence of noise on the synthetic speech, realizing the feature optimization of the comprehensive adjustment feature vector, and obtaining an optimized speech feature vector.

[0127] In step S404 of some embodiments, by comparing the differences between the optimized speech feature vector and the synthetic speech vector, speech feature loss data is obtained, and then according to the speech feature loss data, the parameters of the initial speech generation sub-model are adjusted to obtain a speech generation sub-model.

[0128] In steps S401 to S404 illustrated in the embodiments of the present application, the feature encoding sub-model is used to merge the synthetic text vector and the sketch prosody features to obtain a comprehensive feature vector, and then the initial prosody control sub-model and the synthetic speech prosody features are used to adjust the prosody features of the comprehensive feature vector to obtain a comprehensive adjustment feature vector. Further, the initial speech generation sub-model is used to optimize the features of the comprehensive adjustment feature vector to obtain an optimized speech feature vector. Finally, according to the synthetic speech vector and the optimized speech feature vector, the parameters of the initial speech generation sub-model are adjusted to obtain a speech generation sub-model, which can improve the generalization ability of the speech generation sub-model and the accuracy of converting speech features into audio data, thereby improving the accuracy of speech synthesis.

[0129] In step S209 of some embodiments, after obtaining the trained feature encoding sub-model, prosody control sub-model and speech generation sub-model, the feature encoding sub-model, prosody control sub-model and speech generation sub-model can be spliced according to the pre-set model splicing order and model splicing position to obtain a target speech synthesis model.

[0130] In steps S201 to S209 illustrated in the embodiments of the present application, by using the obtained synthetic speech sample data, synthetic text sample data and prosody sketch sample data to perform modular training on the initial speech synthesis model, the feature encoding sub-model, prosody control sub-model and speech generation sub-model are obtained, and then the feature encoding sub-model, prosody control sub-model and speech generation sub-model are merged to obtain a target speech synthesis model, which can improve the accuracy of the target speech synthesis and make the speech synthesis effect better.

[0131] In step S103 of some embodiments, the prosody feature vector refers to a vector extracted from the speech prosody sketch, representing prosody features such as pitch, rhythm, and intensity. For example, features such as a slow speaking speed and a steady intonation.

[0132] In the embodiments of the present application, the feature encoding sub-model can obtain prosody-related features by extracting information such as the phase and amplitude of the speech prosody sketch, and then perform processing such as Fourier transform on the information such as the phase and amplitude to obtain prosody feature vectors such as pitch, rhythm, and intensity.

[0133] In step S104 of some embodiments, the text vector refers to the vector form of the target text, and the text vector can represent the semantic information of the target text.

[0134] In the embodiments of the present application, when the feature encoding sub-model performs text encoding on the target text, it can make the vector representation of the target text more accurate by identifying the duration of each phoneme in the target text.

[0135] Specifically, please refer to Figure 5 , in some embodiments, step S104 may include but is not limited to steps S501 to S503:

[0136] Step S501: Based on the feature encoding sub-model, perform semantic encoding on the target text to obtain a semantic vector;

[0137] Step S502: Based on the feature encoding sub-model and the prosody feature vector, perform phoneme duration prediction on the target text to obtain a phoneme duration feature vector;

[0138] Step S503: Merge the semantic vector and the phoneme duration feature vector to obtain a text vector.

[0139] In step S501 of some embodiments, the semantic vector refers to the vector representation of the semantic information contained in the target text.

[0140] In the embodiments of the present application, the feature encoding sub-model can be used to perform word embedding on the target text, so as to map each phrase in the target text to a high-dimensional vector space to obtain a phrase embedding vector. Further, position information is added to each phrase so that the feature encoding sub-model can identify the order of the phrases in the sentence of the target text. Secondly, the position information of each phrase is encoded to obtain a phrase position encoding vector. Finally, the embedding vector of the phrase is merged with the corresponding phrase position encoding vector to obtain a semantic vector representing the semantic information of the target text.

[0141] In step S502 of some embodiments, the phoneme duration feature vector refers to the duration feature vector of each phoneme in the target text.

[0142] In the embodiments of the present application, according to the mapping relationship between the phoneme categories and duration features pre-learned by the feature encoding sub-model, each phoneme in the target text can be mapped to an initial phoneme duration feature vector, and then the prosody feature vector is parsed to obtain the speech rhythm target. Finally, using the feature encoding sub-model, according to the speech rhythm target, the initial phoneme duration feature vector of each phoneme is adjusted to obtain the phoneme duration feature vector.

[0143] Specifically, please refer to Figure 6 , in some embodiments, step S502 may include but is not limited to steps S601 to S603:

[0144] Step S601, based on the feature encoding sub-model, perform phoneme duration mapping on the target text to obtain an initial phoneme duration feature vector;

[0145] Step S602, based on the feature encoding sub-model, perform feature parsing on the prosody feature vector to obtain the speech rhythm target;

[0146] Step S603, based on the feature encoding sub-model and the speech rhythm target, perform duration adjustment on the initial phoneme duration feature vector to obtain the phoneme duration feature vector.

[0147] In step S601 of some embodiments, the initial phoneme duration feature vector refers to a mapping vector obtained by performing duration feature mapping on each phoneme in the target text according to the mapping relationship between the phoneme types and duration features learned by the feature encoding sub-model.

[0148] In the embodiments of the present application, the feature encoding sub-model can learn the phoneme duration mapping relationship between each phoneme type and duration feature during model training, and according to the learned phoneme duration mapping relationship, map each phoneme in the target text to the duration feature vector space, thereby obtaining the initial phoneme duration feature vector.

[0149] In step S602 of some embodiments, the speech rhythm target refers to the rhythm feature that the speech data synthesized from the target text is expected to achieve. For example, when the target text is "I am very happy today", the speech rhythm target may be that the intonation of the two words "happy" is lighter and the speech speed is slower.

[0150] In the embodiments of the present application, by quantifying the prosody feature vector, the prosody quantization feature is obtained, and then the prosody quantization feature is normalized to obtain the normalized prosody feature. Finally, using the feature encoding sub-model, the normalized prosody feature is used for prosody pattern recognition to obtain the speech rhythm target.

[0151] Specifically, please refer to Figure 7, in some embodiments, step S602 may include but is not limited to steps S701 to S704:

[0152] Step S701, based on the feature encoding sub-model, perform feature quantization on the prosody feature vector to obtain a prosody quantization feature;

[0153] Step S702, based on the feature encoding sub-model, perform normalization processing on the prosody quantization feature to obtain a normalized prosody feature;

[0154] Step S703, based on the feature encoding sub-model, perform prosody pattern recognition on the normalized prosody feature to obtain a speech prosody pattern;

[0155] Step S704, based on the speech prosody pattern, set a speech rhythm target.

[0156] In steps S701 and S702 of some embodiments, the prosody quantization feature refers to the numerical representation of the prosody feature vector. For example, a fast speaking speed can be the value 3, a medium speaking speed can be the value 2, and a slow speaking speed can be the value 1. The normalized prosody feature is a form of the prosody feature vector obtained by shrinking the prosody quantization feature for easy understanding by the model.

[0157] In the embodiments of the present application, according to the prosody information in the prosody feature vector, the prosody feature vector can be converted into a corresponding numerical representation, that is, the prosody quantization feature can be obtained. Further, to facilitate the calculation of the prosody quantization feature by the feature encoding sub-model, by performing normalization processing on the prosody quantization feature, a normalized prosody feature can be obtained.

[0158] In steps S703 and S704 of some embodiments, the speech prosody pattern refers to the speech prosody operation model represented by the normalized prosody feature.

[0159] In the embodiments of the present application, by using the feature encoding sub-model to identify the normalized prosody feature, the speech prosody model included in the prosody feature vector can be easily obtained. Further, according to the rhythm features mentioned in the speech prosody pattern, the speech rhythm target of the target text is set.

[0160] In steps S701 to S704 shown in the embodiments of the present application, according to the feature encoding sub-model, the prosody feature vector is feature-quantized to obtain the prosody quantization feature, and then according to the feature encoding sub-model, the prosody quantization feature is normalized to obtain the normalized prosody feature, which is convenient for the feature encoding sub-model to read the prosody feature vector, thereby improving the feature parsing speed of the feature encoding sub-model for the prosody feature vector. Further, according to the feature encoding sub-model, the normalized prosody feature is subjected to prosody pattern recognition to obtain the speech prosody pattern, and then according to the speech prosody pattern, the speech rhythm target is set, so that the pronunciation of the speech data synthesized from the target text satisfies the prosody features of the speech prosody sketch, thereby enabling the target speech synthesis model to generate more accurate speech data with a specific emotion or style.

[0161] In step S603 of some embodiments, after obtaining the speech rhythm target, according to features such as the speech rate and intonation mentioned in the speech rhythm target, the initial phoneme duration feature vector of each phoneme of the target text can be adjusted to obtain the phoneme duration feature vector. For example, when the speech rhythm target is a slow speech rate, the pronunciation duration of each phoneme of the target text can be increased.

[0162] In steps S601 to S603 shown in the embodiments of the present application, according to the feature encoding sub-model, the target text is subjected to phoneme duration mapping to obtain the initial phoneme duration feature vector, and then according to the feature encoding sub-model, the prosody feature vector is subjected to feature parsing to obtain the speech rhythm target. Finally, according to the feature encoding sub-model and the speech rhythm target, the initial phoneme duration feature vector is adjusted in duration to obtain the phoneme duration feature vector, so that the pronunciation of the speech data synthesized from the target text satisfies the prosody features of the speech prosody sketch.

[0163] In step S503 of some embodiments, by concatenating the semantic vector and the phoneme duration feature vector, the vector combination of the semantic vector and the phoneme duration feature vector can be performed in time, so that a text vector containing the phoneme duration feature can be obtained.

[0164] In steps S501 to S503 shown in the embodiments of the present application, according to the feature encoding sub-model, the target text is semantically encoded to obtain the semantic vector, and then according to the feature encoding sub-model and the prosody feature vector, the target text is subjected to phoneme duration prediction to obtain the phoneme duration feature vector. Finally, the semantic vector and the phoneme duration feature vector are vector-combined to obtain the text vector, so that when the text vector is converted into speech data, the phrase pronunciation can be performed according to the phoneme duration feature in the text vector, thereby improving the rhythm of the synthesized speech and ensuring that the phrase pronunciation time of the synthesized speech can satisfy the prosody features of the speech prosody sketch.

[0165] In step S105 of some embodiments, the spliced feature vector refers to the vector representation formed after splicing the text vector and the prosody feature vector. For example, in the field of fintech, the spliced feature vector can be the vector obtained by splicing the text vector of the financial product introduction text and the prosody feature vector. In the medical field, the spliced feature vector can be the vector obtained by splicing the medical advice text vector and the prosody feature vector.

[0166] In the embodiments of the present application, after the feature encoding sub-model generates the text vector and the prosody feature vector, it can splice the text vector and the prosody feature vector to generate a spliced feature vector that contains the semantic information of the target text and the prosody information of the speech prosody sketch.

[0167] In step S106 of some embodiments, the prosody feature optimization vector refers to a prosody feature vector that is more natural and more in line with user requirements.

[0168] In the embodiments of the present application, the prosody control sub-model can map the prosody feature vector to the optimization vector space according to the sketch contour mapping information learned during model training, so as to obtain the prosody feature constant optimization vector.

[0169] In step S107 of some embodiments, the target feature vector refers to the spliced feature vector that has adjusted the prosody feature vector. It should be noted that the prosody of the speech generated by the target feature vector is more natural and more in line with user requirements.

[0170] In the embodiments of the present application, after the prosody control sub-model generates the prosody feature optimization vector, the prosody control sub-model adjusts the prosody feature in the spliced feature vector according to the prosody feature optimization vector, so that the adjusted target feature vector is more natural and more in line with user requirements, thereby ensuring that the speech synthesis model can generate a more accurate specific emotion or style.

[0171] In step S108 of some embodiments, the synthesized audio refers to the audio data output by the target text according to the prosody features in the speech prosody sketch. For example, when the target text is "I am very happy today". The synthesized audio is the voice data of "I am very happy today" with a brisk intonation and a lively speaking speed.

[0172] In the embodiments of the present application, after obtaining the target feature vector, the speech generation sub-model can denoise the target feature vector through a diffusion model to optimize the features of the target feature vector, so as to obtain an optimized feature vector. Then, the decoder is used to decode the optimized feature vector to obtain speech features. Finally, the vocoder is used to convert the speech features into a speech waveform, and the synthesized audio is generated according to the speech waveform.

[0173] This application obtains a speech prosody sketch, obtains a target text, and obtains a target speech synthesis model including a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model, thereby clarifying the input conditions and model conditions for speech synthesis. Secondly, according to the feature encoding sub-model, feature extraction is performed on the speech prosody sketch to obtain a prosody feature vector, and according to the feature encoding sub-model and the prosody feature vector, text encoding is performed on the target text to obtain a text vector. Then, according to the feature encoding sub-model, vector concatenation is performed on the text vector and the prosody feature vector to obtain a concatenated feature vector, such that the concatenated feature vector has both the semantic information of the text vector and the prosody expression ability of the prosody feature vector. Finally, according to the prosody control sub-model and the prosody feature vector, prosody contour prediction is performed to obtain a prosody feature optimization vector, and according to the prosody control sub-model and the prosody feature optimization vector, prosody adjustment is performed on the concatenated feature vector to obtain a target feature vector. Based on the speech generation sub-model, audio generation is performed on the target feature vector to obtain a synthesized audio, making the synthesized speech more in line with user needs, improving the generation accuracy of speech emotion or style. In addition, it can also make the synthesized speech data more natural.

[0174] Please refer to Figure 8 , this embodiment of the application also provides a speech synthesis device that can implement the above speech synthesis method. The device includes:

[0175] A data acquisition module 801, configured to obtain a speech prosody sketch and obtain a target text;

[0176] A model acquisition module 802, configured to obtain a target speech synthesis model, where the target speech synthesis model includes a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model;

[0177] A feature extraction module 803, configured to perform feature extraction on the speech prosody sketch based on the feature encoding sub-model to obtain a prosody feature vector;

[0178] A text encoding module 804, configured to perform text encoding on the target text based on the feature encoding sub-model and the prosody feature vector to obtain a text vector;

[0179] A vector concatenation module 805, configured to perform vector concatenation on the text vector and the prosody feature vector based on the feature encoding sub-model to obtain a concatenated feature vector;

[0180] A feature optimization module 806, configured to perform prosody contour prediction based on the prosody control sub-model and the prosody feature vector to obtain a prosody feature optimization vector;

[0181] A vector adjustment module 807, configured to perform prosody adjustment on the concatenated feature vector based on the prosody control sub-model and the prosody feature optimization vector to obtain a target feature vector;

[0182] The audio generation module 808 is used to generate audio for the target feature vector based on the speech generation sub-model.

[0183] The specific implementation manner of this speech synthesis device is basically the same as the specific embodiments of the above speech synthesis method, and will not be elaborated here.

[0184] An embodiment of this application also provides an electronic device. The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above speech synthesis method is implemented. This electronic device can be any intelligent terminal including a tablet computer, an in-vehicle computer, etc.

[0185] Please refer to Figure 9 , Figure 9 which illustrates the hardware structure of an electronic device in another embodiment. The electronic device includes:

[0186] A processor 901, which can be implemented in ways such as a general-purpose CPU (Central Processing Unit), a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is used to execute relevant programs to implement the technical solutions provided by the embodiments of this application;

[0187] A memory 902, which can be implemented in forms such as a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 902 can store an operating system and other application programs. When implementing the technical solutions provided by the embodiments of this specification through software or firmware, the relevant program codes are stored in the memory 902, and the processor 901 is used to call and execute the speech synthesis method of the embodiments of this application;

[0188] An input / output interface 903, which is used to implement information input and output;

[0189] A communication interface 904, which is used to implement communication interaction between this device and other devices, and can implement communication through a wired manner (such as USB, network cable, etc.) or through a wireless manner (such as a mobile network, WIFI, Bluetooth, etc.);

[0190] A bus 905, which transmits information between various components of the device (such as the processor 901, the memory 902, the input / output interface 903, and the communication interface 904);

[0191] Among them, the processor 901, the memory 902, the input / output interface 903, and the communication interface 904 are communicatively connected to each other inside the device through the bus 905.

[0192] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the above voice synthesis method.

[0193] As a non-transitory computer-readable storage medium, the memory can be used to store non-transitory software programs and non-transitory computer-executable programs. In addition, the memory may include high-speed random access memory and may also include non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory may optionally include memories remotely disposed relative to the processor, and these remote memories may be connected to the processor through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0194] The voice synthesis method, voice synthesis device, electronic device, and storage medium provided by the embodiments of the present application obtain a voice prosody sketch, obtain a target text, obtain a target voice synthesis model including a feature encoding sub-model, a prosody control sub-model, and a voice generation sub-model, then extract features from the voice prosody sketch according to the feature encoding sub-model to obtain a prosody feature vector, encode the target text based on the feature encoding sub-model and the prosody feature vector to obtain a text vector, splice the text vector and the prosody feature vector based on the feature encoding sub-model to obtain a spliced feature vector. Secondly, perform prosody contour prediction based on the prosody control sub-model and the prosody feature vector to obtain an optimized prosody feature vector, adjust the prosody of the spliced feature vector based on the prosody control sub-model and the optimized prosody feature vector to obtain a target feature vector, and finally, generate an audio based on the voice generation sub-model for the target feature vector.

[0195] The embodiments described in the embodiments of the present application are for more clearly illustrating the technical solutions of the embodiments of the present application and do not constitute a limitation on the technical solutions provided by the embodiments of the present application. Those skilled in the art will know that with the evolution of technology and the emergence of new application scenarios, the technical solutions provided by the embodiments of the present application are equally applicable to similar technical problems.

[0196] Those skilled in the art can understand that the technical solutions shown in the figures do not constitute a limitation on the embodiments of the present application, and may include more or fewer steps than shown in the figures, or combine some steps, or different steps.

[0197] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0198] Those of ordinary skill in the art can understand that all or some of the steps in the methods disclosed above, and the functional modules / units in systems and devices, can be implemented as software, firmware, hardware, and their appropriate combinations.

[0199] The terms "first", "second", "third", "fourth", etc. (if any) in the specification of this application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0200] It should be understood that in this application, "at least one (item)" means one or more, and "multiple" means two or more. "And / or" is used to describe the association relationship of associated objects and indicates that three relationships can exist. For example, "A and / or B" can mean: only A exists, only B exists, and both A and B exist at the same time. Among them, A and B can be singular or plural. The character " / " generally means that the associated objects before and after are in an "or" relationship. "At least one (one)" or similar expressions below refer to any combination of these items, including any combination of single item (one) or plural items (ones). For example, at least one (one) of a, b, or c can mean: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, c can be single or multiple.

[0201] In several embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. The displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of devices or units can be in electrical, mechanical or other forms.

[0202] The units described above as separate components may or may not be physically separated. The components shown as units may or may not be physical units, that is, they can be located in one place, or can be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0203] In addition, in each embodiment of the present application, each functional unit can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0204] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes multiple instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in each embodiment of the present application. And the aforementioned storage medium includes: USB flash drive, mobile hard disk, read-only memory (ROM for short), random access memory (RAM for short), magnetic disk or optical disc and other various media that can store programs.

[0205] The preferred embodiments of the embodiments of the present application have been described above with reference to the drawings, but this does not limit the scope of rights of the embodiments of the present application. Any modifications, equivalent replacements, and improvements made by those skilled in the art without departing from the scope and essence of the embodiments of the present application shall be within the scope of rights of the embodiments of the present application.

Claims

1. A voice synthesis method, characterized in that, The method includes: Obtaining a speech prosody sketch and obtaining a target text; Obtaining a target speech synthesis model, where the target speech synthesis model includes a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model; Based on the feature encoding sub-model, performing feature extraction on the speech prosody sketch to obtain a prosody feature vector; Based on the feature encoding sub-model and the prosody feature vector, performing text encoding on the target text to obtain a text vector; Based on the feature encoding sub-model, concatenating the text vector and the prosody feature vector to obtain a concatenated feature vector; Based on the prosody control sub-model and the prosody feature vector, performing prosody contour prediction to obtain an optimized prosody feature vector; Based on the prosody control sub-model and the optimized prosody feature vector, performing prosody adjustment on the concatenated feature vector to obtain a target feature vector; Based on the speech generation sub-model, performing audio generation on the target feature vector.

2. The method according to claim 1, characterized in that, The obtaining of the target speech synthesis model includes: Obtaining synthetic speech sample data, synthetic text sample data, and prosody sketch sample data; Obtaining an initial speech synthesis model, where the initial speech synthesis model includes the feature encoding sub-model, an initial prosody control sub-model, and an initial speech generation sub-model; Based on the feature encoding sub-model, performing prosody feature extraction on the synthetic speech sample data to obtain a synthetic speech prosody feature; Based on the feature encoding sub-model, performing speech encoding on the synthetic speech sample data to obtain a synthetic speech vector; Based on the feature encoding sub-model, performing text encoding extraction on the synthetic text sample data to obtain a synthetic text vector; Based on the feature encoding sub-model, performing feature extraction on the prosody sketch sample data to obtain a sketch prosody feature; Based on the initial prosody control sub-model, performing supervised learning on the synthetic speech prosody feature and the sketch prosody feature to obtain the prosody control sub-model; Based on the initial speech generation sub-model, performing supervised learning on the synthetic text vector, the sketch prosody feature, and the synthetic speech vector to obtain the speech generation sub-model; Merging the feature encoding sub-model, the prosody control sub-model, and the speech generation sub-model to obtain the target speech synthesis model.

3. The method according to claim 2, wherein The performing of supervised learning on the synthetic speech prosody feature and the sketch prosody feature based on the initial prosody control sub-model to obtain the prosody control sub-model includes: Based on the initial prosody control sub-model, performing mapping relationship learning on the synthetic speech prosody feature and the sketch prosody feature to obtain sketch contour mapping information; Based on the sketch contour mapping information, adjusting the parameters of the initial prosody control sub-model to obtain the prosody control sub-model.

4. The method according to claim 2, wherein The performing of supervised learning on the synthetic text vector, the sketch prosody feature, and the synthetic speech vector based on the initial speech generation sub-model to obtain the speech generation sub-model includes: Based on the feature encoding sub-model, merge the synthetic text vector and the sketch prosody features to obtain a comprehensive feature vector; Based on the initial prosody control sub-model and the synthetic speech prosody features, adjust the prosody features of the comprehensive feature vector to obtain a comprehensively adjusted feature vector; Based on the initial speech generation sub-model, optimize the features of the comprehensively adjusted feature vector to obtain an optimized speech feature vector; Based on the synthetic speech vector and the optimized speech feature vector, adjust the parameters of the initial speech generation sub-model to obtain the speech generation sub-model.

5. The method according to claim 1, wherein The text encoding of the target text based on the feature encoding sub-model and the prosody feature vector to obtain a text vector includes: Based on the feature encoding sub-model, perform semantic encoding on the target text to obtain a semantic vector; Based on the feature encoding sub-model and the prosody feature vector, predict the phoneme duration of the target text to obtain a phoneme duration feature vector; Merge the semantic vector and the phoneme duration feature vector to obtain the text vector.

6. The method according to claim 5, wherein The predicting the phoneme duration of the target text based on the feature encoding sub-model and the prosody feature vector to obtain a phoneme duration feature vector includes: Based on the feature encoding sub-model, perform phoneme duration mapping on the target text to obtain an initial phoneme duration feature vector; Based on the feature encoding sub-model, perform feature parsing on the prosody feature vector to obtain a speech rhythm target; Based on the feature encoding sub-model and the speech rhythm target, adjust the duration of the initial phoneme duration feature vector to obtain the phoneme duration feature vector.

7. The method according to claim 6, characterized in that The performing feature parsing on the prosody feature vector based on the feature encoding sub-model to obtain a speech rhythm target includes: Based on the feature encoding sub-model, perform feature quantization on the prosody feature vector to obtain a prosody quantization feature; Based on the feature encoding sub-model, perform normalization processing on the prosody quantization feature to obtain a normalized prosody feature; Based on the feature encoding sub-model, perform prosody pattern recognition on the normalized prosody feature to obtain a speech prosody pattern; Based on the speech prosody pattern, set the speech rhythm target.

8. A voice synthesis device, characterized in that, The device includes: A data acquisition module for acquiring a speech prosody sketch and acquiring a target text; A model acquisition module for acquiring a target speech synthesis model, where the target speech synthesis model includes a feature encoding sub-model, a prosody control sub-model, and a speech generation sub-model; A feature extraction module for extracting features from the speech prosody sketch based on the feature encoding sub-model to obtain a prosody feature vector; A text encoding module for encoding the target text based on the feature encoding sub-model and the prosody feature vector to obtain a text vector; A vector splicing module for splicing the text vector and the prosody feature vector based on the feature encoding sub-model to obtain a spliced feature vector; A feature optimization module, configured to perform prosody contour prediction based on the prosody control sub-model and the prosody feature vector to obtain a prosody feature optimization vector; A vector adjustment module, configured to perform prosody adjustment on the spliced feature vector based on the prosody control sub-model and the prosody feature optimization vector to obtain a target feature vector; An audio generation module, configured to generate audio based on the target feature vector by using the speech generation sub-model.

9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the speech synthesis method according to any one of claims 1 to 7 is implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the speech synthesis method according to any one of claims 1 to 7 is implemented.