Speech synthesis method, device, computer device, storage medium and product

By integrating reference audio feature information and text units, and using a multi-channel duration prediction network for time matching processing, the problem of unnatural and mechanical speech transition in traditional speech synthesis technology is solved, and a more natural and fluent speech synthesis effect is achieved.

CN114333758BActive Publication Date: 2025-07-11TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111302064.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-11-04
Publication Date
2025-07-11
Estimated Expiration
2041-11-04

AI Technical Summary

Technical Problem

In traditional speech synthesis technology, waveform splicing method leads to unnatural speech transitions and poor speech synthesis effect, while statistical parameter methods have a mechanical sense and are not good in synthesis.

Method used

By obtaining the text of the speech to be synthesized and determining the speech type, the reference audio feature information is fused with the text unit, the audio duration information is predicted using the multiple-channel duration prediction network, and the duration matching process is performed, and the speech synthesis is finally performed, preserving the tone and rhythm of different speech types.

Benefits of technology

The speech synthesis effect of different pronunciation types is improved, making the synthesized pronunciation more natural and fluent, and retaining the tone and rhythm characteristics of their respective pronunciation types.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114333758B_ABST
    Figure CN114333758B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a speech synthesis method, apparatus, computer device, storage medium, and product. By obtaining the text of the speech to be synthesized and determining the type of speech to be synthesized; fusing the reference audio feature information corresponding to the speech type with the text units in the text to obtain text speech feature information; determining a target duration prediction network according to the speech type; predicting the audio duration information corresponding to the text units according to the target duration prediction network and the text speech feature information; performing duration matching processing on the text speech feature information according to the audio duration information to obtain the duration-matched text speech feature information; and performing speech synthesis processing according to the duration-matched text speech feature information to obtain the target speech. This solution can extract accurate text speech feature information, and use the corresponding duration prediction network according to the speech type, so that the synthesized target speech retains information such as the timbre and prosody of the speech type, improving the speech synthesis effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a voice synthesis method, apparatus, computer device, storage medium, and product. Background Art

[0002] Voice synthesis technology converts text into corresponding audio content through certain rules or model algorithms, also known as Text to Speech (TTS). Its function is to convert the text information generated by the computer itself or input externally into understandable and fluent speech and read it out. Traditional voice synthesis technologies mainly rely on waveform splicing methods or statistical parameter methods. The splicing method requires collecting all the waveforms corresponding to the pronunciation units in advance and obtaining the corresponding speech through waveform splicing. The statistical parameter method requires first modeling the spectral characteristic parameters of the existing audio, constructing the mapping relationship between the text sequence and the speech features, and generating a parameter synthesizer. When an input text is provided, the text sequence is first mapped to the corresponding audio features, and the corresponding speech is input according to the audio features.

[0003] Among them, the waveform splicing method needs to collect a large amount of audio for each voice type to cover all pronunciation units, and due to splicing, the synthesized speech has an unnatural transition and poor voice synthesis effect. The statistical parameter method can avoid collecting a large amount of audio, but due to the mapping method, the synthesized speech has a strong mechanical feeling and poor synthesis effect. Summary of the Invention

[0004] Embodiments of this application provide a voice synthesis method, apparatus, computer device, storage medium, and product, which can improve the synthesis effect of voice synthesis.

[0005] A voice synthesis method provided by an embodiment of this application includes:

[0006] Obtain the text of the voice to be synthesized and determine the voice type to be synthesized;

[0007] Fuse the reference audio feature information corresponding to the voice type with the text units in the text to obtain the text voice feature information corresponding to each text unit;

[0008] Determine the corresponding target duration prediction network from multiple duration prediction networks according to the voice type;

[0009] Predict the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text voice feature information;

[0010] For each text unit, perform duration matching processing on the text speech feature information corresponding to the text unit according to the audio duration information to obtain the matched text speech feature information corresponding to each text unit;

[0011] Perform speech synthesis processing according to the matched text speech feature information corresponding to each text to obtain the target speech of the speech type.

[0012] Correspondingly, an embodiment of the present application further provides a speech synthesis device, including:

[0013] An acquisition unit, configured to acquire the text of the speech to be synthesized and determine the speech type to be synthesized;

[0014] A feature fusion unit, configured to fuse the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text speech feature information corresponding to each text unit;

[0015] A network determination unit, configured to determine the corresponding target duration prediction network from multiple duration prediction networks according to the speech type;

[0016] A duration prediction unit, configured to predict the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text speech feature information;

[0017] A matching processing unit, configured to perform duration matching processing on the text speech feature information corresponding to each text unit according to the audio duration information to obtain the matched text speech feature information corresponding to each text unit;

[0018] A speech synthesis unit, configured to perform speech synthesis processing according to the matched text speech feature information corresponding to each text to obtain the target speech of the speech type.

[0019] In one embodiment, the feature fusion unit includes:

[0020] A first feature extraction subunit, configured to perform text feature extraction on the text units included in the text to obtain text feature information;

[0021] A feature fusion subunit, configured to fuse the reference audio feature information corresponding to the speech type with the text feature information of the text to obtain the text speech feature information corresponding to each text unit.

[0022] In one embodiment, the feature fusion unit includes:

[0023] An audio acquisition subunit, configured to acquire the reference audio corresponding to the speech type;

[0024] A second feature extraction subunit, configured to perform audio feature extraction on the reference audio according to text units in the text, so as to obtain reference audio feature information.

[0025] In one embodiment, the speech synthesis unit includes:

[0026] A feature processing network determination subunit, configured to determine a target feature processing network from multiple feature processing networks according to the speech type;

[0027] A feature decoding subunit, configured to perform feature decoding processing on the matched text speech feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text;

[0028] A speech synthesis subunit, configured to perform speech synthesis processing according to the acoustic feature information to obtain the target speech of the text with respect to the speech type.

[0029] In one embodiment, the feature decoding subunit includes:

[0030] A decoding module, configured to perform preliminary decoding processing on the matched text speech feature information to obtain decoded feature information;

[0031] A feature conversion module, configured to perform feature conversion processing on the decoded feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text.

[0032] In one embodiment, the matching processing unit includes:

[0033] An upsampling subunit, configured to perform upsampling processing on the text speech feature information corresponding to each text unit according to the audio duration information for each text unit, so as to obtain at least one text speech feature information corresponding to the text unit;

[0034] A feature information determination subunit, configured to obtain the matched text speech feature information according to the at least one text speech feature information corresponding to the text unit.

[0035] In one embodiment, the feature fusion unit is further configured to:

[0036] Perform fusion processing on the reference audio feature information corresponding to the speech type and text units in the text based on a fusion feature extraction network to obtain text speech feature information corresponding to each text unit.

[0037] In one embodiment, the speech synthesis device includes:

[0038] A sample acquisition unit for acquiring text samples of the speech to be synthesized in at least one training sample set, where the text samples correspond to audio samples;

[0039] A feature fusion training unit for fusing the text samples with the audio samples based on an initial fusion feature extraction network to obtain text-speech feature information of the text samples with respect to text units;

[0040] A network determination training unit for determining a target initial duration prediction network that matches the training sample set from the multiple initial duration prediction networks;

[0041] A duration prediction training unit for predicting audio duration information corresponding to each text unit in the text samples according to the target initial duration prediction network and the text-speech feature information;

[0042] A duration matching unit for performing duration matching processing on the text-speech feature information according to the audio duration information for each text unit to obtain the matched text-speech feature information;

[0043] A decoding unit for performing feature decoding processing on the matched text-speech feature information to obtain acoustic feature information corresponding to the text samples;

[0044] A speech synthesis training unit for training the initial fusion feature extraction network and the initial target duration prediction network respectively based on the acoustic feature information corresponding to the text samples and the acoustic feature information corresponding to the audio samples to obtain a fusion feature extraction network and a target duration prediction network.

[0045] In one embodiment, the acquisition unit includes:

[0046] An acquisition subunit for acquiring the initial text of the speech to be synthesized;

[0047] A regularization subunit for performing text regularization processing on the initial text according to the text units to obtain the text.

[0048] Correspondingly, an embodiment of the present application further provides a computer device, including a memory and a processor; the memory stores a computer program, and the processor is used to run the computer program in the memory to execute any speech synthesis method provided by the embodiment of the present application.

[0049] Correspondingly, an embodiment of the present application further provides a computer-readable storage medium, which is used to store a computer program, and the computer program is loaded by a processor to execute any speech synthesis method provided by the embodiment of the present application.

[0050] Correspondingly, an embodiment of the present application further provides a computer program product, including a computer program / instructions, wherein when the computer program / instructions are executed by a processor, any voice synthesis method provided by the embodiment of the present application is implemented.

[0051] In an embodiment of the present application, by obtaining the text of the voice to be synthesized and determining the type of voice to be synthesized; fusing the reference audio feature information corresponding to the voice type with the text units in the text to obtain the text voice feature information corresponding to each text unit; determining the corresponding target duration prediction network from multiple duration prediction networks according to the voice type; predicting the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text voice feature information; for each text unit, performing duration matching processing on the text voice feature information corresponding to the text unit according to the audio duration information to obtain the matched text voice feature information corresponding to each text unit; performing voice synthesis processing according to the matched text voice feature information corresponding to each text, to obtain the target voice of the voice type.

[0052] This solution obtains the reference audio feature information corresponding to different voice types and the text of the voice to be synthesized for fusion processing, can accurately extract the text voice feature information corresponding to the text, determines the audio duration information according to the duration prediction network corresponding to different voice types, so that the synthesized target voice retains audio information such as timbre and rhythm corresponding to different voice types, and improves the voice synthesis effect of different voice types. Description of the Drawings

[0053] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other drawings can be obtained according to these drawings.

[0054] Figure 1 is a scenario diagram of the voice synthesis method provided by the embodiment of the present application;

[0055] Figure 2 is a flowchart of the voice synthesis method provided by the embodiment of the present application;

[0056] Figure 3 is a schematic diagram of a multi-channel fusion voice synthesis system provided by the embodiment of the present application;

[0057] Figure 4 is another flowchart of the voice synthesis method provided by the embodiment of the present application;

[0058] Figure 5 is a schematic diagram of the voice synthesis interface provided by the embodiment of the present application;

[0059] Figure 6 It is a schematic diagram of a voice synthesis device provided by an embodiment of the present application;

[0060] Figure 7 It is a schematic diagram of the structure of a computer device provided by an embodiment of the present application. Specific embodiments

[0061] Next, the technical solutions in the embodiments of the present application will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts shall fall within the protection scope of the present application.

[0062] An embodiment of the present application provides a voice synthesis method, device, computer device, and computer-readable storage medium. The voice synthesis device can be integrated in a computer device, and the computer device can be a server or a terminal device, etc.

[0063] Among them, the terminal may include a mobile phone, a wearable intelligent device, a tablet computer, a notebook computer, a personal computer (PC), and an in-vehicle computer, etc.

[0064] Among them, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0065] For example, as Figure 1 shown, the terminal sends the text of the voice to be synthesized and the specified voice type to be synthesized to the computer device. The computer device fuses the reference audio feature information corresponding to the voice type with the text units in the text to obtain the text voice feature information corresponding to each text unit; determines the corresponding target duration prediction network from multiple duration prediction networks according to the voice type; predicts the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text voice feature information; for each text unit, performs duration matching processing on the text voice feature information corresponding to the text unit according to the audio duration information to obtain the matched text voice feature information corresponding to each text unit; performs voice synthesis processing according to the matched text voice feature information corresponding to each text to obtain the target voice of the voice type.

[0066] This solution obtains the reference audio feature information corresponding to different voice types and the text of the speech to be synthesized for fusion processing, can accurately extract the text voice feature information corresponding to the text, and determines the audio duration information according to the duration prediction network corresponding to different voice types, so that the synthesized target voice retains audio information such as timbre and rhythm corresponding to different voice types, and improves the speech synthesis effect of different voice types.

[0067] The following will be described in detail respectively. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments.

[0068] This embodiment will be described from the perspective of a speech synthesis device. The speech synthesis device can be specifically integrated in a computer device, which can be a server or a terminal device, etc.

[0069] An embodiment of this application provides a speech synthesis method, as Figure 2 shown, the specific process of this speech synthesis method is as follows:

[0070] 101. Obtain the text of the speech to be synthesized and determine the voice type to be synthesized.

[0071] Among them, the text of the speech to be synthesized can include the text that needs to be synthesized into speech. For example, it can include pinyin (t, o, n, g, 1), phonemes (t, o, ng), or characters (tong), etc. Speech synthesis is also called Text to Speech (TTS), and its function is to convert the text of the speech to be synthesized generated by the computer device or input externally into an understandable and fluent speech and read it out.

[0072] Among them, the voice type can include multiple timbre types. For example, it can include voice types such as news broadcasting, artificial customer service, and cute kids.

[0073] For example, specifically, the computer device can obtain the text of the speech to be synthesized from a database or a blockchain, etc., and can also receive the text of the speech to be synthesized sent by the client in response to the user's input operation; determine the voice type to which the target voice to be synthesized belongs according to the storage location of the speech to be synthesized or the corresponding label, and can also be that the client responds to the user's selection operation of the voice type and sends information corresponding to the voice type to the computer device to instruct the computer device to determine the voice type to be synthesized.

[0074] Under normal circumstances, records are usually made in text form. For example, "I am a cat" is usually input instead of "wo3shi4yi4zhi1mao1". Therefore, the obtained text in text form can be converted into text in units of specified text units. For example, "I am a cat" is converted into "wo3shi4yi4zhi1mao1". That is, in one embodiment, the step of "obtaining the text of the speech to be synthesized" can specifically include:

[0075] Obtain the initial text of the speech to be synthesized;

[0076] Perform text regularization processing on the initial text according to the text unit to obtain the text of the speech to be synthesized.

[0077] Among them, the initial text can be the text obtained by the computer that has not been processed, and can include pinyin, phonemes, or text in text form.

[0078] Among them, the text unit can be the pronunciation unit of the initial text of the speech to be synthesized. For example, it can be pinyin, phoneme, or character.

[0079] Among them, the text regularization processing can include converting the initial text into text in a specified text format.

[0080] For example, specifically, it can be to obtain the initial text of the speech to be synthesized and convert the initial text according to the specified text unit. For example, if the specified text unit is pinyin, the initial text is converted into text represented in pinyin form; if the specified text unit is phoneme, the initial text is converted into text represented in phoneme form; if the specified text unit is text, the initial text is converted into text represented in text form.

[0081] Optionally, the initial text can also be subjected to text cleaning, redundant information in the initial text can be removed, the content that needs to be synthesized into speech can be screened out, and processing such as adding pauses can be performed according to the semantic information of the initial text.

[0082] Performing text regularization processing on the initial text can improve the unity of the obtained text in form and format, and improve the convenience and efficiency of fusing the text unit of the text and the reference audio feature information.

[0083] 102. Perform fusion processing on the reference audio feature information corresponding to the speech type and the text unit in the text to obtain the text speech feature information corresponding to each text unit.

[0084] Among them, the reference audio feature information can include feature information representing audio features such as timbre and prosody corresponding to the speech type.

[0085] Among them, the text voice feature information may be feature information representing the acoustic features corresponding to the text units.

[0086] For example, specifically, it may be to obtain the reference audio feature information corresponding to the voice type, and fuse the reference audio feature information and the feature information corresponding to the text units in the text by means such as feature addition or feature multiplication to obtain the text voice feature information of each text unit in the text.

[0087] The feature information corresponding to the text units can be obtained by performing text feature extraction on the text units. That is, in one embodiment, the step of "fusing the reference audio feature information corresponding to the voice type with the text units in the text to obtain the text voice feature information corresponding to each text unit" may specifically include:

[0088] Performing text feature extraction on the text units included in the text to obtain text feature information;

[0089] Fusing the reference audio feature information corresponding to the voice type with the text feature information of the text to obtain the text voice feature information corresponding to each text unit.

[0090] For example, specifically, it may be to perform text feature extraction on the text units included in the text by means such as one-hot encoding or feature embedding to obtain the text feature information of the text corresponding to the text units.

[0091] Adding the text feature information and the reference audio feature information to fuse the text feature information and the audio feature information to obtain the text voice feature information corresponding to each text unit.

[0092] For each voice type, the reference audio feature information may be preset, or the reference audio corresponding to the voice type may be obtained, and the reference audio feature information may be obtained according to the reference audio. That is, in one embodiment, before the step of "fusing the reference audio feature information corresponding to the voice type with the text feature information of the text to obtain the text voice feature information corresponding to each text unit", it may specifically further include:

[0093] Obtaining the reference audio corresponding to the voice type;

[0094] Performing audio feature extraction on the reference audio according to the text units in the text to obtain the reference audio feature information.

[0095] Among them, the reference audio may include the audio corresponding to at least one voice type.

[0096] For example, specifically, it can be to obtain the audio corresponding to at least one voice type as the reference audio, extract the audio features of the reference audio, obtain the audio features such as the timbre and prosody corresponding to each text unit in the text, and obtain the reference audio features corresponding to the text units.

[0097] Optionally, the above process of fusing the reference audio feature information and the text units to obtain the text voice feature information can be processed through a fusion feature network. That is, in one embodiment, the step of "fusing the reference audio feature information corresponding to the voice type with the text units in the text to obtain the text voice feature information corresponding to each text unit" can specifically include:

[0098] Based on the fusion feature extraction network, the reference audio feature information corresponding to the voice type is fused with the text units in the text to obtain the text voice feature information corresponding to each text unit.

[0099] Among them, the fusion feature extraction network can include a text feature extraction model for extracting text features from the text and an audio feature extraction model for extracting audio features from the reference audio.

[0100] For example, specifically, it can be to extract text features from the text through the fusion feature extraction network to obtain the text feature information corresponding to the text, and extract audio features from the reference audio to obtain the reference audio feature information, and perform feature addition processing on the text feature information and the audio feature information to obtain the text voice feature information.

[0101] The fusion feature extraction network is obtained through pre-training. That is, in one embodiment, the voice synthesis method provided by the embodiments of the present application may specifically further include:

[0102] Obtain the text samples of the speech to be synthesized in at least one training sample set, and the text samples correspond to audio samples;

[0103] Based on the initial fusion feature extraction network, the text samples and the audio samples are fused to obtain the text voice feature information of the text samples with respect to the text units;

[0104] Determine the target initial duration prediction network that matches the training sample set from multiple initial duration prediction networks;

[0105] According to the target initial duration prediction network and the text voice feature information, predict the audio duration information corresponding to each text unit in the text sample;

[0106] For each text unit, perform duration matching processing on the text voice feature information according to the audio duration information to obtain the matched text voice feature information;

[0107] Perform feature decoding processing on the voice feature information of the matched text to obtain the acoustic feature information corresponding to the text sample;

[0108] Based on the acoustic feature information corresponding to the text sample and the acoustic feature information corresponding to the audio sample, train the initial fusion feature extraction network and the initial target duration prediction network respectively to obtain the fusion feature extraction network and the target duration prediction network.

[0109] Among them, the training sample set may include at least one text sample, each text sample corresponds to a segment of audio sample, and the voice types of the audio included in different training sample sets are different. For example, the training sample set A may include news-related audio samples and corresponding text samples, and the training sample set B may include customer service-related audio samples and corresponding text samples, etc.

[0110] Among them, the initial fusion feature extraction network may be a feature extraction network that has not been trained well.

[0111] For example, specifically, it may be to obtain the text sample of the speech to be synthesized in at least one training sample set, and the audio sample corresponding to the text sample, extract the text feature of the text sample through the initial fusion feature extraction network to obtain the text feature information corresponding to the text sample, and extract the audio feature of the audio sample through the initial fusion feature extraction network to obtain the audio feature information corresponding to the audio sample, and perform fusion processing on the text feature information and the audio feature information to obtain the text voice feature information of the text sample regarding the text unit.

[0112] The fusion feature extraction network can extract features from the text samples and audio samples in different training sample sets, that is, different training sample sets jointly train the initial fusion feature extraction network. Jointly training different training sample sets on the initial fusion feature extraction network can improve the ability of the fusion feature extraction network to extract text features and audio features.

[0113] If different training sample sets use different feature extraction networks, the feature extraction ability of the feature extraction network trained based on low-quality training sample sets (for example, training sample sets with a small number of samples, training sample sets with poor recording quality resulting in low-quality audio samples, and training sample sets with low-quality audio samples due to the speaker's non-standard pronunciation) is poor. By using the fusion feature extraction network, it can make full use of the high-quality samples in the high-quality training sample sets to assist the fusion feature extraction network in accurately extracting the text voice feature information in the low-quality training sample sets.

[0114] Training the initial fusion feature extraction network with a training sample set containing audio samples of different speech types can improve the feature extraction ability of the fusion feature extraction network for different training sample sets. To preserve the speech styles of different training sample sets, each training sample set can correspond to a duration prediction network, which is used to predict the audio duration information matching the speech style of the training sample set according to the text speech feature information of the training sample set.

[0115] Determine the corresponding target duration prediction network from multiple target duration prediction networks according to the training sample set corresponding to the text sample. Input the text audio feature information obtained by fusing the text sample and the audio sample into the target duration prediction network, and predict the audio duration information of each text unit through the target duration prediction network based on the text speech feature information.

[0116] Perform duration matching processing on the text speech feature information according to the duration information corresponding to each text unit to obtain the duration-matched text speech feature information.

[0117] Perform feature decoding processing on the duration-matched text speech feature information to obtain the predicted acoustic feature information. Use the acoustic feature information corresponding to the audio sample as the sample label, and train the initial fusion feature network based on the error between the predicted acoustic feature information and the acoustic feature information corresponding to the audio sample to obtain the fusion feature extraction network, and train the target initial duration prediction network corresponding to the training sample set to obtain the target duration prediction network.

[0118] Optionally, performing feature decoding processing on the duration-matched text speech feature information can be achieved through a fusion preliminary decoding network and a multi-path feature processing network. For example, perform preliminary decoding on the duration-matched text speech feature information through the fusion preliminary decoding network to obtain the decoded feature information, and then perform secondary decoding and post-processing on the decoded feature information through the target feature processing network corresponding to the training sample set to obtain the acoustic features. At the same time, train the fusion preliminary decoding network and the target feature processing network based on the error between the predicted acoustic feature information and the acoustic feature information corresponding to the audio sample to obtain the trained fusion preliminary decoding network and the target feature processing network.

[0119] Among them, the specific implementation processes of the duration matching processing and the feature decoding processing refer to the relevant descriptions in the subsequent embodiments and will not be elaborated here.

[0120] 103. Determine the corresponding target duration prediction network from the multi-path duration prediction networks according to the speech type.

[0121] Among them, the multi-channel duration prediction network may include multiple duration prediction networks. The duration prediction network can predict the audio duration information of each text unit according to the text audio feature information, so that the synthesized target speech retains the speech style of the speech type, for example, timbre and rhythm, etc.

[0122] For example, specifically, different speech types may correspond to different duration prediction networks, and the duration prediction network corresponding to the speech type is determined from the multi-channel duration prediction network.

[0123] 104. Predict the audio duration information corresponding to the text unit in the text according to the target duration prediction network and the text speech feature information.

[0124] Among them, the audio duration information may include the pronunciation length of the text unit, for example, 1s, 500ms, or 5 audio frames, etc.

[0125] For example, specifically, the text speech feature information can be used as the input of the target duration prediction network, and the target duration prediction network predicts the audio duration information corresponding to each text unit based on the fused feature information. For the same text unit, the required audio duration information is different for different speech types. The target duration prediction network corresponding to the speech type can make the synthesized target speech conform to the timbre and rhythm, etc. corresponding to the speech type.

[0126] 105. For each text unit, perform duration matching processing on the text speech feature information corresponding to the text unit according to the audio duration information to obtain the matched text speech feature information corresponding to each text unit.

[0127] For example, specifically, each text feature information can correspond to one audio frame. Determine the pronunciation duration required for each text unit according to the audio duration information, copy the text speech feature information corresponding to the text unit, or adjust the text speech feature information to obtain similar text speech feature information, so that the audio frame matching the audio duration information can be obtained according to the text feature information corresponding to the text unit.

[0128] The duration matching processing of the text speech feature information can be implemented by upsampling. That is, in one embodiment, the step "for each text unit, perform duration matching processing on the text speech feature information corresponding to the text unit according to the audio duration information to obtain the matched text speech feature information corresponding to each text unit" may specifically include:

[0129] For each text unit, perform upsampling processing on the text speech feature information corresponding to the text unit according to the audio duration information to obtain at least one text speech feature information corresponding to the text unit;

[0130] Obtain the matched text - speech feature information based on at least one text - speech feature information corresponding to the text unit.

[0131] For example, specifically, it can be determined according to the audio duration information how many audio frames each text unit corresponds to, and the text - speech feature information corresponding to the text unit is copied to obtain multiple text - speech feature information, and the number of text - speech feature information corresponding to the text unit is equal to the number of audio frames indicated by the audio duration information.

[0132] Upsampling processing is performed on the text - speech feature information corresponding to each text unit to obtain the matched text - speech feature information of each text unit in the text.

[0133] 106. Perform speech synthesis processing according to the matched text - speech feature information corresponding to each text to obtain the target speech of the speech type.

[0134] For example, specifically, it can be to perform feature decoding processing according to the matched text - speech feature information to obtain the acoustic feature information corresponding to each text unit, and the vocoder converts the acoustic feature information into the target speech. The vocoder (Vocoder) is derived from the abbreviation of Voice Encoder, also known as a speech signal analysis and synthesis system, and its function is to convert acoustic features into sounds.

[0135] Different speech types correspond to different duration prediction networks. Correspondingly, corresponding feature processing networks can be set for different speech types to perform feature decoding processing on the matched text - speech feature information, which can better convert the matched text - speech feature information into acoustic feature information. That is, in one embodiment, the step of "performing speech synthesis processing according to the matched text - speech feature information corresponding to each text to obtain the target speech of the speech type" can specifically include:

[0136] Determine the target feature processing network from the multiple - path feature processing network according to the speech type;

[0137] Perform feature decoding processing on the matched text - speech feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text;

[0138] Perform speech synthesis processing according to the acoustic feature information to obtain the target speech of the text regarding the speech type.

[0139] Among them, the multiple - path feature processing network can include the feature processing network corresponding to each speech type.

[0140] For example, specifically, it can be to determine the target feature processing network corresponding to the speech type from the multiple - path feature processing network, and perform feature decoding processing on the matched text - speech feature information through the target processing network to obtain the acoustic feature information corresponding to the text.

[0141] The vocoder performs speech synthesis processing based on the acoustic feature information to obtain a target speech of the speech type.

[0142] Optionally, the feature processing network may include a secondary decoding network and a post-processing network. The secondary decoding network is used to perform feature decoding processing on the text speech feature information after matching to obtain initial acoustic feature information, and the post-processing network is used to perform processing such as smoothing adjustment on the initial acoustic feature information to obtain acoustic feature information.

[0143] If there is a situation in the multi-branch feature processing network where the feature decoding of the feature processing network of this branch is unstable due to insufficient training samples or low sample quality, therefore, before performing feature decoding through the feature processing network of the branch, preliminary decoding processing may be performed first to obtain a stable decoding result for the feature processing network of the branch, which can reduce the poor decoding effect caused by insufficient training samples and low sample quality. That is, in one embodiment, the step of "performing feature decoding processing on the text speech feature information after matching through the target feature processing network to obtain the acoustic feature information corresponding to the text" may specifically include:

[0144] Performing preliminary decoding processing on the text speech feature information after matching to obtain decoded feature information;

[0145] Performing feature conversion processing on the decoded feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text.

[0146] For example, specifically, it may be to perform processing and decoding on the text speech feature information after matching to obtain decoded feature information, and then perform secondary decoding on the decoded feature information through the target feature processing network, and convert the decoded feature information obtained by the secondary decoding into the corresponding acoustic feature information.

[0147] As can be seen from the above, the computer device in the embodiment of the present application obtains the text of the speech to be synthesized and determines the type of speech to be synthesized; fuses the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text-speech feature information corresponding to each text unit; determines the corresponding target duration prediction network from multiple duration prediction networks according to the speech type; predicts the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text-speech feature information; for each text unit, performs duration matching processing on the text-speech feature information corresponding to the text unit according to the audio duration information to obtain the text-speech feature information after matching corresponding to each text unit; performs speech synthesis processing according to the text-speech feature information after matching corresponding to each text to obtain the target speech of the speech type. In this solution, the reference audio feature information corresponding to different speech types and the text of the speech to be synthesized are fused and processed, so that the text-speech feature information corresponding to the text can be accurately extracted, and the audio duration information is determined according to the duration prediction network corresponding to different speech types, so that the synthesized target speech retains audio information such as timbre and rhythm corresponding to different speech types, and improves the speech synthesis effect of different speech types.

[0148] Based on the above embodiments, further detailed descriptions will be given by way of examples below.

[0149] This embodiment will be described from the perspective of a speech synthesis device, which can be specifically integrated in a computer device, and the computer device can be a server.

[0150] In the embodiment of the present application, the fusion feature extraction network, multiple duration prediction networks, and feature processing network are integrated into a multi-channel fusion speech synthesis system. The structure of the multi-channel fusion speech synthesis system is as Figure 3 shown. The multi-channel fusion speech synthesis system can include a fusion part and a multi-channel part. Fusion part: The texts and training sample sets of different speech types share a fusion feature synthesis network and a fusion preliminary decoding network. Multi-channel part: The texts and training sample sets of different speech types correspond to independent duration prediction networks and feature processing networks.

[0151] A speech synthesis method provided by an embodiment of the present application, as Figure 4 shown, the specific process of the speech synthesis method can be as follows:

[0152] 201. The server obtains the text sample and the corresponding audio sample of the speech to be synthesized in at least one training sample set.

[0153] Among them, the training sample set can be a voice library, which can include audio samples of specific voice types and corresponding text samples. For example, voice library A can include audio samples of news and corresponding text samples, and voice library B can include audio samples of customer service and corresponding text samples, etc. There are varying degrees of differences among multiple voice libraries in terms of text richness and recording quality. The text samples of voice library A usually cover rich content, and the audio samples are generally high-quality professional studio recordings. Such voice libraries have good stability for the trained neural network. For voice library B, the professionalism of the speakers is relatively low, and the text content is mostly for customer service scenarios, so the sample quality of voice library B is low.

[0154] For example, specifically, it can be to obtain the text samples of the speech to be synthesized in at least one training sample set, and the audio samples corresponding to the text samples.

[0155] 202. The server performs feature extraction and fusion processing on the text samples and audio samples through the initial fusion feature extraction network of the multi-channel fusion speech synthesis system to obtain text-speech feature information.

[0156] For example, specifically, it can be to perform text feature extraction on the text samples through the initial fusion feature extraction network of the multi-channel fusion speech synthesis system to obtain the text feature information corresponding to the text samples, and to perform audio feature extraction on the audio samples through the initial fusion feature extraction network to obtain the audio feature information corresponding to the audio samples, and then fuse the text feature information and the audio feature information to obtain the text-speech feature information of the text samples regarding text units.

[0157] Training the initial fusion feature extraction network by fusing multiple training sample sets, and using both high-quality and low-quality voice libraries for training can solve problems such as incomplete text coverage and low recording quality in the low-quality voice library, which may cause network instability, and improve the ability of the fusion feature extraction network to extract text features and audio features.

[0158] 203. The server predicts the audio duration information corresponding to each text unit in the text sample based on the text-speech feature information through the target initial duration prediction network corresponding to the training sample set.

[0159] For example, in order to retain audio information such as the timbre and rhythm of the speakers in different training sample sets, each training sample set can correspond to a duration prediction network, which is used to predict the audio duration information matching the speech style of the training sample set according to the text-speech feature information of the training sample set.

[0160] Determine the corresponding target duration prediction network from multiple target duration prediction networks according to the training sample set corresponding to the text sample. Input the text-audio feature information obtained by fusing the text sample and the audio sample into the target duration prediction network, and predict the audio duration information of each text unit based on the text-audio feature information through the target duration prediction network.

[0161] 204. The server performs upsampling processing on the text-speech feature information according to the audio duration information to obtain the matched text-speech feature information corresponding to the text sample.

[0162] For example, specifically, it can be to determine the number of audio frames corresponding to each text unit in the text sample according to the audio duration information, copy the text-speech feature information corresponding to the text unit to obtain multiple text-speech feature information, and the number of text-speech feature information corresponding to the text unit is equal to the number of audio frames indicated by the audio duration information.

[0163] Perform upsampling processing on the text-speech feature information corresponding to each text unit to obtain the matched text-speech feature information of each text unit in the text sample.

[0164] 205. The server performs preliminary decoding processing on the matched text-speech feature information through the fusion preliminary decoding network of the multi-channel fusion speech synthesis system to obtain the decoded feature information.

[0165] Among them, the fusion preliminary decoding network may include a preliminary decoder.

[0166] For example, specifically, it can be to perform decoding processing on the matched text-speech feature information through the fusion preliminary decoding network of the multi-channel fusion speech synthesis system to obtain the decoded feature information. Sharing the preliminary decoding network by different training sample sets helps to obtain a stable preliminary decoding result for the feature processing network to use, and avoid problems such as unstable decoding network and incorrect decoding result caused by low-quality samples in the training sample set. By training the preliminary decoding network by combining high-quality training samples and low-quality training samples, the initial decoding network can input a stable decoding result, and the feature processing network can obtain a stable output result without training with a large number of samples.

[0167] 206. The server performs feature decoding processing on the decoded feature information through the target feature processing network corresponding to the training sample set to obtain the acoustic feature information corresponding to the text sample.

[0168] For example, specifically, it can be to determine a target feature processing network corresponding to a training sample set from a multi-channel feature processing network, perform secondary decoding processing on the matched text-speech feature information through the secondary decoding network of the target processing network to obtain initial acoustic feature information, and then use a post-processing network to perform smoothing adjustment and other processing on the initial acoustic feature information to obtain acoustic feature information.

[0169] 207. The server trains each network in the multi-channel fusion speech synthesis system based on the acoustic feature information corresponding to the text sample and the acoustic feature information corresponding to the audio sample, and obtains a trained multi-channel fusion speech synthesis system.

[0170] For example, perform feature decoding processing on the matched text-speech feature information to obtain predicted acoustic feature information, use the acoustic feature information corresponding to the audio sample as a sample label, and based on the error between the predicted acoustic feature information and the acoustic feature information corresponding to the audio sample, train the fusion feature extraction network, target duration prediction network, preliminary decoding network, and target feature processing network included in the branch corresponding to the text sample in the multi-channel fusion speech synthesis system to obtain a trained multi-channel fusion speech synthesis system.

[0171] For example, the text sample a comes from the training sample set A, and the training sample set corresponds to the duration prediction network I and the feature processing network I. Then, based on the error between the predicted acoustic feature information and the acoustic feature information corresponding to the audio sample, train the fusion feature extraction network, preliminary decoding network, duration prediction network I, and feature processing network I.

[0172] 208. The server receives a speech synthesis request sent by the client, obtains the text of the speech to be synthesized according to the speech synthesis request, and determines the type of speech to be synthesized.

[0173] For example, as Figure 5 shown, the client provides a speech synthesis interface. The speech synthesis interface can include a text input area and a speech type selection control. The speech synthesis interface also includes a speech synthesis control. The client responds to the user's confirmation operation on the speech synthesis control, obtains the text in the text input area as the text of the speech to be synthesized, and determines the type of speech to be synthesized according to the selection operation on the speech type selection control, and sends a speech synthesis request to the server.

[0174] The server determines the text of the speech to be synthesized according to the speech synthesis request, and determines the type of speech to be synthesized.

[0175] 209. The server performs feature extraction and fusion processing on the text and audio through the fusion feature extraction network to obtain text-speech feature information.

[0176] For example, specifically, text feature extraction can be performed on text units included in the text by means of one-hot encoding or feature embedding, etc., to obtain text feature information corresponding to the text about the text units. At least one audio corresponding to the voice type is obtained as a reference audio, audio feature extraction is performed on the reference audio, and audio features such as timbre and prosody corresponding to each text unit in the text are obtained to obtain reference audio features corresponding to the text units. Then, the text feature information and the audio feature information are subjected to feature addition processing to obtain text voice feature information.

[0177] 210. The server performs upsampling processing on the text voice feature information through a target duration prediction network corresponding to the voice type to obtain the matched text voice feature information.

[0178] For example, specifically, the number of audio frames corresponding to each text unit can be determined according to the audio duration information, and the text voice feature information corresponding to the text unit is copied to obtain multiple pieces of text voice feature information. The number of pieces of text voice feature information corresponding to the text unit is equal to the number of audio frames indicated by the audio duration information.

[0179] Upsampling processing is performed on the text voice feature information corresponding to each text unit to obtain the matched text voice feature information of each text unit in the text.

[0180] 211. The server performs preliminary decoding processing on the matched text voice feature information through a preliminary decoding network to obtain decoded feature information.

[0181] For example, specifically, processing and decoding processing can be performed on the matched text voice feature information to obtain decoded feature information.

[0182] 212. The server performs feature decoding processing on the decoded feature information through a target feature processing network corresponding to the voice type to obtain acoustic feature information corresponding to the text.

[0183] For example, specifically, a target feature processing network corresponding to the voice type can be determined from a multi-channel feature processing network, and secondary decoding processing is performed on the matched text voice feature information through a secondary decoding network of the target processing network to obtain initial acoustic feature information, and then the initial acoustic feature information is subjected to smoothing adjustment and other processing through a post-processing network to obtain acoustic feature information corresponding to the text.

[0184] 213. The server converts the acoustic feature information into a target voice through a vocoder and returns it to the client.

[0185] Voice synthesis processing is performed based on the acoustic feature information through a vocoder to obtain a target voice of the voice type, and the target voice is sent to the client.

[0186] As can be seen from the above, in the server of the embodiment of the present application, the text sample and the corresponding audio sample of the speech to be synthesized are obtained from at least one training sample set; the initial fusion feature extraction network of the multi-channel fusion speech synthesis system is used to extract and fuse the features of the text sample and the audio sample to obtain text-speech feature information; through the target initial duration prediction network corresponding to the training sample set, the audio duration information corresponding to each text unit in the text sample is predicted based on the text-speech feature information; the text-speech feature information is upsampled according to the audio duration information to obtain the matched text-speech feature information corresponding to the text sample; through the fusion preliminary decoding network of the multi-channel fusion speech synthesis system, the matched text-speech feature information is preliminarily decoded to obtain the decoded feature information; through the target feature processing network corresponding to the training sample set, the decoded feature information is feature decoded to obtain the acoustic feature information corresponding to the text sample; based on the acoustic feature information corresponding to the text sample and the acoustic feature information corresponding to the audio sample, each network in the multi-channel fusion speech synthesis system is trained to obtain the trained multi-channel fusion speech synthesis system. The server receives a speech synthesis request sent by the client, obtains the text of the speech to be synthesized according to the speech synthesis request, and determines the type of speech to be synthesized; the fusion feature extraction network is used to extract and fuse the features of the text and the audio to obtain text-speech feature information; the target duration prediction network corresponding to the speech type is used to upsample the text-speech feature information to obtain the matched text-speech feature information; the preliminary decoding network is used to preliminarily decode the matched text-speech feature information to obtain the decoded feature information; the target feature processing network corresponding to the speech type is used to feature decode the decoded feature information to obtain the acoustic feature information corresponding to the text; the acoustic feature information is converted into the target speech through the vocoder and returned to the client.

[0187] This solution adopts a multi-channel fusion speech synthesis system. By using the fusion feature extraction network and the fusion preliminary decoding network, the characteristics of wide text coverage and high recording quality of the high-quality training sample set can be fully utilized to improve the synthesis accuracy and stability of the low-quality training sample set. The corresponding duration prediction network is adopted according to different speech types, so that the synthesized target speech retains the audio information such as timbre and rhythm corresponding to its respective training sample set. In addition, in actual use, no redundant calculation needs to be introduced, and the calculation amount is the same as that of the model trained with a single voice library.

[0188] To facilitate better implementation of the speech synthesis method provided by the embodiment of the present application, a speech synthesis device is also provided in an embodiment. The meanings of the nouns are the same as those in the above speech synthesis method, and the specific implementation details can be referred to the description in the method embodiment.

[0189] The speech synthesis device can be specifically integrated into a computer device, such as Figure 6 As shown, the speech synthesis device may include: an acquisition unit 301, a feature fusion unit 302, a network determination unit 303, a duration prediction unit 304, a matching processing unit 305, and a speech synthesis unit 306, specifically as follows:

[0190] (1) Acquisition unit 301: It is used to acquire the text of the speech to be synthesized and determine the type of speech to be synthesized.

[0191] In one embodiment, the acquisition unit 301 may include an acquisition subunit and a regularization subunit. Specifically:

[0192] Acquisition subunit: It is used to acquire the initial text of the speech to be synthesized;

[0193] Regularization subunit: It is used to perform text regularization processing on the initial text according to text units to obtain the text.

[0194] (2) Feature fusion unit 302: It is used to fuse the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text speech feature information corresponding to each text unit.

[0195] In one embodiment, the feature fusion unit 302 may include a first feature extraction subunit and a feature fusion subunit. Specifically:

[0196] First feature extraction subunit: It is used to perform text feature extraction on the text units included in the text to obtain text feature information;

[0197] Feature fusion subunit: It is used to fuse the reference audio feature information corresponding to the speech type with the text feature information of the text to obtain the text speech feature information corresponding to each text unit.

[0198] In one embodiment, the feature fusion unit 302 may include an audio acquisition subunit and a second feature extraction subunit. Specifically:

[0199] Audio acquisition subunit: It is used to acquire the reference audio corresponding to the speech type;

[0200] Second feature extraction subunit: It is used to perform audio feature extraction on the reference audio according to the text units in the text to obtain the reference audio feature information.

[0201] In one embodiment, the feature fusion unit 302 may also be used to:

[0202] Based on the fusion feature extraction network, fuse the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text speech feature information corresponding to each text unit.

[0203] In one embodiment, the speech synthesis device may further include a sample acquisition unit, a feature fusion training unit, a network determination training unit, a duration prediction training unit, a duration matching unit, a decoding unit, and a speech synthesis training unit. Specifically:

[0204] Sample acquisition unit: configured to acquire text samples of the speech to be synthesized in at least one training sample set, and the text samples correspond to audio samples;

[0205] Feature fusion training unit: configured to perform fusion processing on the text samples and the audio samples based on an initial fusion feature extraction network to obtain text-speech feature information of the text samples with respect to text units;

[0206] Network determination training unit: configured to determine a target initial duration prediction network that matches the training sample set from multiple initial duration prediction networks;

[0207] Duration prediction training unit: configured to predict audio duration information corresponding to each text unit in the text samples according to the target initial duration prediction network and the text-speech feature information;

[0208] Duration matching unit: configured to perform duration matching processing on the text-speech feature information according to the audio duration information for each text unit to obtain the matched text-speech feature information;

[0209] Decoding unit: configured to perform feature decoding processing on the matched text-speech feature information to obtain acoustic feature information corresponding to the text samples;

[0210] Speech synthesis training unit: configured to train the initial fusion feature extraction network and the initial target duration prediction network respectively based on the acoustic feature information corresponding to the text samples and the acoustic feature information corresponding to the audio samples to obtain a fusion feature extraction network and a target duration prediction network.

[0211] (3) Network determination unit 303: configured to determine a corresponding target duration prediction network from multiple duration prediction networks according to the speech type.

[0212] (4) Duration prediction unit 304: configured to predict audio duration information corresponding to text units in the text according to the target duration prediction network and the text-speech feature information.

[0213] (5) Matching processing unit 305: configured to perform duration matching processing on the text-speech feature information corresponding to each text unit according to the audio duration information to obtain the matched text-speech feature information corresponding to each text unit.

[0214] In one embodiment, the matching processing unit 305 may include an upsampling subunit and a feature information determination subunit. Specifically:

[0215] Upsampling subunit: For each text unit, perform upsampling processing on the text speech feature information corresponding to the text unit according to the audio duration information to obtain at least one piece of text speech feature information corresponding to the text unit;

[0216] Feature information determination subunit: Obtain the text speech feature information after matching according to at least one piece of text speech feature information corresponding to the text unit.

[0217] (6) Speech synthesis unit 306: Perform speech synthesis processing according to the text speech feature information after matching corresponding to each text to obtain the target speech of the speech type.

[0218] In one embodiment, the speech synthesis unit 306 may include a feature processing network determination subunit, a feature decoding subunit, and a speech synthesis subunit. Specifically:

[0219] Feature processing network determination subunit: Determine the target feature processing network from multiple feature processing networks according to the speech type;

[0220] Feature decoding subunit: Perform feature decoding processing on the text speech feature information after matching through the target feature processing network to obtain the acoustic feature information corresponding to the text;

[0221] Speech synthesis subunit: Perform speech synthesis processing according to the acoustic feature information to obtain the target speech of the text regarding the speech type.

[0222] In one embodiment, the feature decoding subunit may include a decoding module and a feature conversion module. Specifically:

[0223] Decoding module: Perform preliminary decoding processing on the text speech feature information after matching to obtain the decoded feature information;

[0224] Feature conversion module: Perform feature conversion processing on the decoded feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text.

[0225] As can be seen from the above, the voice synthesis device according to the embodiment of the present application obtains the text of the voice to be synthesized and determines the voice type to be synthesized through the obtaining unit 301; the feature fusion unit 302 fuses the reference audio feature information corresponding to the voice type with the text units in the text to obtain the text voice feature information corresponding to each text unit; the network determination unit 303 determines the corresponding target duration prediction network from multiple duration prediction networks according to the voice type; the duration prediction unit 304 predicts the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text voice feature information; the matching processing unit 305 performs duration matching processing on the text voice feature information corresponding to each text unit according to the audio duration information to obtain the text voice feature information after matching corresponding to each text unit; finally, the voice synthesis unit 306 performs voice synthesis processing according to the text voice feature information after matching corresponding to each text to obtain the target voice of the voice type. This solution obtains the reference audio feature information corresponding to different voice types and the text of the voice to be synthesized for fusion processing, can accurately extract the text voice feature information corresponding to the text, determines the audio duration information according to the duration prediction network corresponding to different voice types, so that the synthesized target voice retains audio information such as timbre and rhythm corresponding to different voice types, and improves the voice synthesis effect of different voice types.

[0226] The embodiment of the present application further provides a computer device, which may be a terminal or a server, such as Figure 7 shown, which shows a schematic structural diagram of the computer device involved in the embodiment of the present application. Specifically:

[0227] The computer device may include components such as a processor 1001 with one or more processing cores, a memory 1002 with one or more computer-readable storage media, a power supply 1003, and an input unit 1004. Those skilled in the art can understand that Figure 7 the computer device structure shown in

[0228] The processor 1001 is the control center of the computer device, connecting various parts of the entire computer device through various interfaces and circuits. By running or executing software programs and / or modules stored in the memory 1002, and by invoking the data stored in the memory 1002, it executes various functions of the computer device and processes data, thereby exercising overall control over the computer device. Optionally, the processor 1001 may include one or more processing cores; preferably, the processor 1001 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interface, computer programs, etc., and the modem processor mainly processes wireless communications. It can be understood that the above-mentioned modem processor may not be integrated into the processor 1001 either.

[0229] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002. The memory 1002 mainly includes a program storage area and a data storage area. Among them, the program storage area can store the operating system, computer programs required for at least one function (such as the sound playback function, image playback function, etc.); the data storage area can store data created according to the use of the computer device. In addition, the memory 1002 can include high-speed random access memory, and can also include non-volatile memory, such as at least one magnetic disk storage device, flash memory device, or other non-volatile solid-state storage devices. Correspondingly, the memory 1002 can also include a memory controller to provide the processor 1001 with access to the memory 1002.

[0230] The computer device also includes a power supply 1003 that powers each component. Preferably, the power supply 1003 can be logically connected to the processor 1001 through a power management system, thereby implementing functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 1003 can also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0231] The computer device may also include an input unit 1004, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.

[0232] Although not shown, the computer device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 1001 in the computer device will, according to the following instructions, load the executable files corresponding to the processes of one or more computer programs into the memory 1002, and the processor 1001 will run the computer programs stored in the memory 1002 to implement various functions as follows:

[0233] Obtain the text of the speech to be synthesized and determine the speech type to be synthesized;

[0234] Fuse the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text-speech feature information corresponding to each text unit;

[0235] Determine the corresponding target duration prediction network from multiple duration prediction networks according to the speech type;

[0236] Predict the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text-speech feature information;

[0237] For each text unit, perform duration matching processing on the text-speech feature information corresponding to the text unit according to the audio duration information to obtain the post-matching text-speech feature information corresponding to each text unit;

[0238] Perform speech synthesis processing according to the post-matching text-speech feature information corresponding to each text to obtain the target speech of the speech type.

[0239] For the specific implementation of each of the above operations, reference may be made to the previous embodiments, which will not be elaborated here.

[0240] As can be seen from the above, the computer device according to the embodiments of the present application can obtain the reference audio feature information corresponding to different speech types and the text of the speech to be synthesized for fusion processing, can accurately extract the text-speech feature information corresponding to the text, and determine the audio duration information according to the duration prediction network corresponding to different speech types, so that the synthesized target speech retains audio information such as timbre and rhythm corresponding to different speech types, improving the speech synthesis effect of different speech types.

[0241] According to one aspect of the present application, there is provided a computer program product or a computer program, the computer program product or the computer program including computer instructions, the computer instructions being stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementation manners in the above embodiments.

[0242] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0243] Therefore, an embodiment of the present application provides a computer-readable storage medium storing a computer program that can be loaded by a processor to execute any one of the voice synthesis methods provided by the embodiments of the present application.

[0244] For the specific implementation of each of the above operations, reference may be made to the previous embodiments and will not be elaborated herein.

[0245] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0246] Since the computer program stored in the computer-readable storage medium can execute any one of the voice synthesis methods provided by the embodiments of the present application, the beneficial effects achievable by any one of the voice synthesis methods provided by the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated herein.

[0247] The above has introduced in detail a voice synthesis method, apparatus, computer device, and computer-readable storage medium provided by the embodiments of the present application. Specific examples are used herein to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A voice synthesis method, characterized in that, Including: Obtaining the text of the speech to be synthesized and determining the type of speech to be synthesized; Fusing the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text-speech feature information corresponding to each text unit; Determining the corresponding target duration prediction network from multiple duration prediction networks according to the speech type; Predicting the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text-speech feature information; For each text unit, performing duration matching processing on the text-speech feature information corresponding to the text unit according to the audio duration information to obtain the text-speech feature information after matching corresponding to each text unit; Performing speech synthesis processing according to the text-speech feature information after matching corresponding to each text to obtain the target speech of the speech type; The fusing the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text-speech feature information corresponding to each text unit includes: Based on a fusion feature extraction network, fusing the reference audio feature information corresponding to the speech type with the text units in the text to obtain the text-speech feature information corresponding to each text unit; The method further includes: obtaining the text samples of the speech to be synthesized in at least one training sample set, and the text samples corresponding to audio samples; Based on an initial fusion feature extraction network, fusing the text samples with the audio samples to obtain the text-speech feature information of the text samples with respect to text units; Determining a target initial duration prediction network matching the training sample set from multiple initial duration prediction networks; Predicting the audio duration information corresponding to each text unit in the text samples according to the target initial duration prediction network and the text-speech feature information; For each text unit, performing duration matching processing on the text-speech feature information according to the audio duration information to obtain the text-speech feature information after matching; Performing feature decoding processing on the text-speech feature information after matching to obtain the acoustic feature information corresponding to the text samples; Based on the acoustic feature information corresponding to the text samples and the acoustic feature information corresponding to the audio samples, training the initial fusion feature extraction network and the target initial duration prediction network respectively to obtain a fusion feature extraction network and a duration prediction network corresponding to the training sample set.

2. The method according to claim 1, characterized in that, The fusing the reference audio feature information corresponding to the speech type with the text feature information of the text to obtain the text-speech feature information corresponding to each text unit includes: Performing text feature extraction on the text units included in the text to obtain text feature information; Fusing the reference audio feature information corresponding to the speech type with the text feature information of the text to obtain the text-speech feature information corresponding to each text unit.

3. The method according to claim 2, characterized in that, Before the fusing the reference audio feature information corresponding to the speech type with the text feature information of the text to obtain the text-speech feature information corresponding to each text unit, the method further includes: Obtain the reference audio corresponding to the voice type; Extract audio feature information from the reference audio according to the text units in the text to obtain reference audio feature information.

4. The method according to claim 1, wherein The voice synthesis processing according to the matched text voice feature information corresponding to each text to obtain the target voice of the voice type includes: Determine the target feature processing network from multiple feature processing networks according to the voice type; Perform feature decoding processing on the matched text voice feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text; Perform voice synthesis processing according to the acoustic feature information to obtain the target voice of the text with respect to the voice type.

5. The method according to claim 4, wherein The performing feature decoding processing on the matched text voice feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text includes: Perform preliminary decoding processing on the matched text voice feature information to obtain decoded feature information; Perform feature conversion processing on the decoded feature information through the target feature processing network to obtain the acoustic feature information corresponding to the text.

6. The method according to claim 1, wherein For each text unit, the performing duration matching processing on the text voice feature information corresponding to the text unit according to the audio duration information to obtain the matched text voice feature information corresponding to each text unit includes: For each text unit, perform upsampling processing on the text voice feature information corresponding to the text unit according to the audio duration information to obtain at least one text voice feature information corresponding to the text unit; Obtain the matched text voice feature information according to at least one text voice feature information corresponding to the text unit.

7. The method according to any one of claims 1-6, characterized in that, The obtaining the text of the voice to be synthesized includes: Obtain the initial text of the voice to be synthesized; Perform text regularization processing on the initial text according to the text unit to obtain the text of the voice to be synthesized.

8. A voice synthesis device, characterized in that, Includes: An obtaining unit, configured to obtain the text of the voice to be synthesized and determine the voice type to be synthesized; A feature fusion unit, configured to perform fusion processing on the reference audio feature information corresponding to the voice type and the text units in the text to obtain the text voice feature information corresponding to each text unit; A network determination unit, configured to determine the corresponding target duration prediction network from multiple duration prediction networks according to the voice type; A duration prediction unit, configured to predict the audio duration information corresponding to the text units in the text according to the target duration prediction network and the text voice feature information; A matching processing unit, configured to perform duration matching processing on the text voice feature information corresponding to each text unit according to the audio duration information to obtain the matched text voice feature information corresponding to each text unit; A voice synthesis unit, configured to perform voice synthesis processing according to the matched text voice feature information corresponding to each text to obtain the target voice of the voice type; The feature fusion unit is configured to: based on a fusion feature extraction network, fuse the reference audio feature information corresponding to the speech type with text units in the text to obtain text-speech feature information corresponding to each text unit; The feature fusion unit is further configured to: obtain text samples of speech to be synthesized in at least one training sample set, where the text samples correspond to audio samples; Based on an initial fusion feature extraction network, fuse the text samples with the audio samples to obtain text-speech feature information of the text samples with respect to text units; Determine a target initial duration prediction network that matches the training sample set from multiple initial duration prediction networks; According to the target initial duration prediction network and the text-speech feature information, predict audio duration information corresponding to each text unit in the text samples; For each text unit, perform duration matching processing on the text-speech feature information according to the audio duration information to obtain duration-matched text-speech feature information; Perform feature decoding processing on the duration-matched text-speech feature information to obtain acoustic feature information corresponding to the text samples; Based on the acoustic feature information corresponding to the text samples and the acoustic feature information corresponding to the audio samples, train the initial fusion feature extraction network and the target initial duration prediction network respectively to obtain a fusion feature extraction network and a duration prediction network corresponding to the training sample set.

9. A computer device, characterized in that, It includes a memory and a processor; the memory stores a computer program, and the processor is configured to run the computer program in the memory to execute the speech synthesis method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium is used to store a computer program, and the computer program is loaded by a processor to execute the speech synthesis method according to any one of claims 1 to 7.

11. A computer program product, comprising a computer program / instructions, characterized in that, When the computer program / instructions are executed by the processor, the steps in the speech synthesis method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Voice synthesis method and device

    CN107705783A

  • Speech recognition method and device, server and computer readable storage medium

    CN112802461A