Speech synthesis method, device, electronic device and storage medium

By encoding the phoneme sequence of the chapter text and combining the language sense characteristics of each sentence for pronunciation synthesis, the problem of rhythm and emotional incoherence in pronunciation synthesis of chapter texts in the prior art is solved, and the naturalness and user experience of pronunciation synthesis are improved.

CN114267330BActive Publication Date: 2025-05-13UNIV OF SCI & TECH OF CHINA +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202111659164.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-30
Publication Date
2025-05-13
Estimated Expiration
2041-12-30

AI Technical Summary

Technical Problem

When the existing pronunciation synthesis method synthesizes chapter texts, it is easy to have problems such as rhythm and emotional inconsistency between the previous and next sentences, which affects the user experience.

Method used

By encoding the phoneme sequence of the chapter text, the phonetic characteristics of the overall modeling of the chapter text are obtained, and the speech synthesis is combined with the language sense characteristics of each sentence to ensure the coherence of the synthesized speech at the speech levels such as rhythm and emotion.

Benefits of technology

It realizes the coherence of synthetic pronunciation in terms of rhythm, emotion, etc., and improves the naturalness and user experience of synthetic pronunciation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114267330B_ABST
    Figure CN114267330B_ABST
Patent Text Reader

Abstract

The present invention provides a speech synthesis method, device, electronic device and storage medium, wherein the method comprises: determining a text phoneme sequence of a text to be synthesized; encoding the text phoneme sequence to obtain the phonetic features of the text; performing speech synthesis based on the phonetic features to obtain the synthesized speech of the text. The method, device, electronic device and storage medium provided by the present invention encode the text phoneme sequence of the text to obtain the phonetic features for modeling the text as a whole, and perform speech synthesis based on the phonetic features, thereby ensuring the coherence of the synthesized speech at the level of rhythm, emotion and other sense of language, and improving the naturalness of the synthesized speech.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer technology, and in particular to a speech synthesis method, device, electronic equipment and storage medium. Background Art

[0002] Text to Speech (TTS) is a technology that converts text into speech. Existing speech synthesis methods based on deep learning are mainly divided into two categories: autoregressive speech synthesis methods and non-autoregressive speech synthesis methods.

[0003] The above two types of speech synthesis methods both perform well when synthesizing a single sentence. However, for a paragraph text containing multiple sentences, the above two types of speech synthesis methods need to splice the speech synthesized independently for each sentence into one segment of speech, which may easily lead to incoherence in the rhythm and emotion of the previous and next sentences, affecting the user experience. Summary of the invention

[0004] The present invention provides a speech synthesis method, device, electronic device and storage medium, which are used to solve the problem of incoherent speech synthesis in the prior art.

[0005] The present invention provides a speech synthesis method, comprising:

[0006] Determining a text phoneme sequence of a text to be synthesized;

[0007] Encoding the text phoneme sequence to obtain the phonetic features of the text;

[0008] Speech synthesis is performed based on the phonetic features to obtain synthesized speech of the passage text.

[0009] According to a speech synthesis method provided by the present invention, the speech synthesis is performed based on the phonetic features to obtain the synthesized speech of the passage text, including:

[0010] Based on the phonetic features and the language sense features of each sentence in the text, speech synthesis is performed to obtain the synthesized speech of the text.

[0011] According to a speech synthesis method provided by the present invention, the sense of language features of each sentence in the passage text are determined based on the following steps:

[0012] Based on the sample language sense features of each sentence in the sample text, extract the language sense of each sentence in the text to obtain the language sense features of each sentence in the text;

[0013] The sample language sense feature is obtained by extracting the language sense feature of the real speech corresponding to the sample chapter text.

[0014] According to a speech synthesis method provided by the present invention, based on the sample language sense features of each sentence in the sample text, language sense extraction is performed on each sentence in the text to obtain the language sense features of each sentence in the text, including:

[0015] Performing semantic extraction on each sentence in the text of the article to obtain semantic features of each sentence in the text of the article;

[0016] Based on the semantic sense conversion relationship, the semantic features of each sentence in the text are converted into sense of language to obtain the sense of language features of each sentence in the text;

[0017] The semantic sense-of-language conversion relationship is determined based on the sample semantic features and sample sense-of-language features of each sentence in the sample chapter text.

[0018] According to a speech synthesis method provided by the present invention, the sample sense of language feature is determined based on the following steps:

[0019] Encoding the acoustic features of the real speech corresponding to the sample passage text to obtain the speech features of the real speech;

[0020] Based on the speech-sense conversion relationship, the speech features are converted into a sense of language to obtain sample sense of language features of each sentence in the sample chapter text;

[0021] The speech-sense conversion relationship is obtained by comparative learning based on the sentence-level features of each sentence in the speech feature, taking the local features of each sentence in the speech feature as positive example points and taking the local features of other sentences in the speech feature as negative example points.

[0022] According to a speech synthesis method provided by the present invention, the speech synthesis is performed based on the phonetic features and the sense of language features of each sentence in the text to obtain the synthesized speech of the text, including:

[0023] The phonetic features and the sense of language features of each sentence in the text are integrated in units of sentences to obtain integrated features of each sentence in the text;

[0024] Based on the fusion features of each sentence in the passage text, speech synthesis is performed to obtain the synthesized speech of the passage text.

[0025] According to a speech synthesis method provided by the present invention, encoding the text phoneme sequence to obtain the phonetic features of the text includes:

[0026] Encoding the text phoneme sequence to obtain a phoneme-level vector of the text;

[0027] Based on the phoneme-level vector, predict the duration of each phoneme in the text phoneme sequence;

[0028] Based on the duration of each phoneme in the text phoneme sequence, the phoneme-level vector is upsampled to obtain the phonetic feature.

[0029] The present invention also provides a speech synthesis device, comprising:

[0030] A phoneme determination unit, used for determining a text phoneme sequence of a text to be synthesized;

[0031] A chapter encoding unit, used for encoding the chapter phoneme sequence to obtain the phonetic features of the chapter text;

[0032] The speech synthesis unit is used to perform speech synthesis based on the phonetic features to obtain the synthesized speech of the passage text.

[0033] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of any of the above-mentioned speech synthesis methods when executing the computer program.

[0034] The present invention also provides a non-transitory computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above-mentioned speech synthesis methods are implemented.

[0035] The speech synthesis method, device, electronic device and storage medium provided by the present invention encode the text phoneme sequence of the text to obtain the phonetic features for the overall modeling of the text, and perform speech synthesis based on the obtained features, thereby ensuring the coherence of the synthesized speech at the level of rhythm, emotion and other sense of language, and improving the naturalness of the synthesized speech. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly described below. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0037] Figure 1 It is a flow chart of the speech synthesis method provided by the present invention;

[0038] Figure 2 It is a flow chart of the method for extracting language sense features provided by the present invention;

[0039] Figure 3It is a flow chart of the sample language sense feature extraction method provided by the present invention;

[0040] Figure 4 It is a structural schematic diagram of the language sense feature extraction model provided by the present invention;

[0041] Figure 5 is a flow chart of step 120 in the method for extracting language sense features provided by the present invention;

[0042] Figure 6 It is a structural schematic diagram of the speech synthesis system provided by the present invention;

[0043] Figure 7 It is a structural schematic diagram of the speech synthesis device provided by the present invention;

[0044] Figure 8 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0045] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0046] At present, speech synthesis methods based on deep learning can be divided into two categories: autoregressive speech synthesis methods and non-autoregressive speech synthesis methods. The autoregressive speech synthesis method adopts the classic encoder-decoder (ED) framework, in which the encoder encodes the input language features, the decoder predicts the acoustic features frame by frame in an autoregressive manner, and the encoder and decoder use the attention mechanism to align the sequence. The Tacotron model is the main representative of this type of method. Compared with the autoregressive speech synthesis method, the non-autoregressive speech synthesis method also uses the ED framework and the encoder has the same function. The difference is that its decoder uses a non-autoregressive method to generate the entire acoustic feature sequence at the same time, and no longer uses an unstable attention mechanism. Instead, an additional duration model is added to use the duration predicted by the duration model to upsample the encoder output sequence to the same length as the acoustic feature sequence. FastSpeech is the main representative of this type of method.

[0047] However, both the autoregressive synthesis method and the non-autoregressive synthesis method model a single sentence, that is, perform speech synthesis on a sentence basis. Since each sentence cannot see the information of the adjacent sentences when modeling a single sentence, its generated state is relatively random, and the spliced ​​speech is prone to have inconsistent rhythm and emotion between the upper and lower sentences, which affects the user experience. In particular, when the above-mentioned speech synthesis method is applied to the recording of audio novels, the problems of inconsistent rhythm and emotion between the upper and lower sentences will cause the recorded audio novel to have a situation similar to the situation where the first sentence is high-spirited and the second sentence is low-spirited, resulting in a very poor user experience.

[0048] In view of the above problems, an embodiment of the present invention provides a speech synthesis method. Figure 1 It is a flow chart of the speech synthesis method provided by the present invention, such as Figure 1 As shown, the method includes:

[0049] Step 110: determine the text phoneme sequence of the text to be synthesized.

[0050] Specifically, the chapter text to be synthesized is a text containing multiple sentences. The chapter text can be the text of a paragraph or the text of the entire chapter containing multiple paragraphs. The chapter text can be directly input by the user, or can be obtained by collecting images through image collection devices such as scanners, mobile phones, cameras, and performing OCR (Optical Character Recognition) on the images, or can be obtained by crawling the Internet, which is not specifically limited in the embodiment of the present invention.

[0051] The chapter phoneme sequence is a phoneme-level text sequence for the entire chapter text. The chapter phoneme sequence includes the phoneme-level text of each sentence in the chapter text. Specifically, the phoneme-level text of each sentence in the chapter text can be obtained by concatenating the phoneme-level text of each sentence in the chapter text according to the arrangement order of each sentence in the chapter text. Here, the phoneme-level text can be obtained by converting the corresponding text into phonemes. For example, the entire chapter text can be converted into phonemes in units of characters to obtain the chapter phoneme sequence.

[0052] Step 120: Encode the text phoneme sequence to obtain the phonetic features of the text.

[0053] Specifically, conventional speech synthesis can be divided into two parts: encoding and decoding, where encoding is the extraction of phonetic features from text, and decoding is the decoding of speech based on phonetic features. Considering that the current speech synthesis method based on sentences usually encodes the phoneme-level text of a single sentence when encoding, the contextual relationship between other sentences in the text and the sentence is ignored.

[0054] To address this problem, in the embodiment of the present invention, the text phoneme sequence is encoded to achieve the extraction of phonetic features of the text. Since the text phoneme sequence covers the phoneme-level text of all sentences in the text, the encoding process based on the text phoneme sequence can refer to the global information in the text, thereby obtaining the phonetic features obtained by modeling the text as a whole. Compared with the phonetic features obtained by modeling a single sentence in the related art, it can better reflect the global information of the text, and the coherence between the phonetic features involved in each sentence is stronger.

[0055] Step 130 , performing speech synthesis based on the phonetic features to obtain synthesized speech of the passage text.

[0056] Specifically, based on step 120, the overall modeling of the text is used to obtain the phonetic features for speech synthesis, that is, the phonetic features are decoded to obtain the synthesized speech of the text. Since the phonetic features in the embodiment of the present invention refer to the global information of the text, the synthesized speech is more continuous and natural in terms of rhythm, emotion and other language sense, which can effectively overcome the problem of rhythmic and emotional incoherence in the speech synthesized based on the text.

[0057] The speech synthesis method provided in the embodiment of the present invention encodes the chapter phoneme sequence of the chapter text to obtain the phonetic features for modeling the chapter text as a whole, and performs speech synthesis based on the obtained phonetic features, which can ensure the coherence of the synthesized speech at the level of rhythm, emotion and other sense of language, and improve the naturalness of the synthesized speech.

[0058] Considering that the modeling capability is relatively limited, modeling the entire text of the passage can certainly enhance the coherence between the previous and next sentences and avoid the listener's feeling of sudden changes between sentences, but the quality of the synthesized speech still needs to be further optimized. Based on the above embodiment, step 130 includes:

[0059] Based on the phonetic features and the language sense features of each sentence in the text, speech synthesis is performed to obtain the synthesized speech of the text.

[0060] Specifically, when performing speech synthesis on a passage text, not only the phonetic features of the passage text can be referred to, but also the sense of language features of each sentence in the passage text can be combined. The sense of language features of each sentence here are used to reflect the rhythm, emotional trend and other sense of language features of the corresponding sentence in the passage text. The sense of language features can be obtained by performing sentiment analysis on the text of each sentence in the passage text, or by performing sentiment analysis on the text of each sentence in the passage text as a whole. The embodiment of the present invention does not make specific limitations on this.

[0061] Before performing speech synthesis, the phonetic features and language sense features of each sentence in the phonetic features of the passage text can be fused, and the fused features of each sentence can be applied to speech synthesis. Alternatively, in the process of speech decoding based on the phonetic features of the passage text, the parameters applied when performing speech decoding on the corresponding sentence can be adjusted based on the language sense features of each sentence in the passage text, that is, the speech decoding of the phonetic features is guided by the language sense features, thereby obtaining the synthesized speech of the passage text.

[0062] The method provided by the embodiment of the present invention combines the language sense features of each sentence in the text of a passage during the speech synthesis process, and guides the rhythm and emotional trend of the synthesized speech through the sentence-level language sense features, so that the speech synthesis can further enhance the long-term coherence of the passage speech on the basis of ensuring local coherence based on passage modeling, that is, the synthesized passage speech can have a rhythm and emotional fluctuation similar to that of real human speech.

[0063] Based on any of the above embodiments, in step 130, the sense of language features of each sentence in the chapter text are determined based on the following steps:

[0064] Based on the sample language sense features of each sentence in the sample text, extract the language sense of each sentence in the text to obtain the language sense features of each sentence in the text;

[0065] The sample language sense feature is obtained by extracting the language sense feature of the real speech corresponding to the sample chapter text.

[0066] Specifically, for the extraction of the sense of language of each sentence in the text, the sense of language of each sentence in the text can be obtained by mapping the mapping relationship between each sentence in the sample text and its sample sense of language features. The mapping relationship here can be specifically reflected in the sense of language extraction model obtained through model training, or in the sense of language extraction rules obtained through association mining, which is not specifically limited in the embodiment of the present invention.

[0067] Furthermore, considering that the language sense features reflect the characteristics of rhythm, emotion, etc., compared with extracting language sense features from text, the language sense features extracted from speech can more realistically and appropriately express the characteristics of real people reading. Before obtaining the sample language sense features of each sentence in the sample chapter text, the embodiment of the present invention first collects the real speech of the sample chapter text. Here, the real speech of the sample chapter text is the speech recorded when a real person reads the sample chapter text. The sample language sense features obtained by extracting language sense features from the real speech learn the characteristics of real people reading, and the characteristics of rhythm, emotion, etc. represented are more realistic, vivid and natural. On this basis, by applying each sentence in the sample chapter text and its sample language sense features, the language sense of each sentence in the chapter text is extracted, which can further improve the authenticity and reliability of the language sense features of each sentence in the chapter text in expressing information such as rhythm and emotion.

[0068] Based on any of the above embodiments, Figure 2 It is a flow chart of the method for extracting language sense features provided by the present invention. Figure 2 As shown, in step 130, the sense of language features of each sentence in the text of the chapter are determined based on the following steps:

[0069] Step 210: perform semantic extraction on each sentence in the text of the chapter to obtain the semantic features of each sentence in the text of the chapter.

[0070] Here, semantic extraction is performed on each sentence in the text of the article. Specifically, semantic extraction can be performed on each sentence in the text of the article independently, or semantic extraction can be performed on the entire text of the article based on the context of each sentence in the text of the article, so as to obtain the semantic features of each sentence in the text of the article.

[0071] Furthermore, semantic extraction can be achieved through the BERT (Bidirectional Encoder Representation from Transformers) model in the field of natural language processing, or by applying other language models with encoding capabilities, such as the Encoder in the Transformer model. Taking the BERT model as an example, the BERT model itself has the ability to understand text, and has strong modeling capabilities, and can output high-dimensional vectors containing semantic information. The input of the BERT model is a word-level text sequence of the text of the passage, and the output is a high-dimensional encoding vector of the same scale containing semantic information. The word-level text sequence referred to here is composed of the text of each sentence in the passage text, and its form can be " <cls>Sentence 1 <sep> <cls>Sentence 2 <sep> … <cls>Sentence <sep>", the output encoding vector is of the same scale as the input. In the embodiment of the present invention, the output of each sentence of the model can be directly taken <cls>The vector corresponding to the label position is used as the semantic feature of each sentence.

[0072] Step 220, based on the semantic sense conversion relationship, the semantic features of each sentence in the text are converted into a sense of language to obtain the sense of language features of each sentence in the text;

[0073] The semantic sense-of-language conversion relationship is determined based on the sample semantic features and sample sense-of-language features of each sentence in the sample chapter text.

[0074] Specifically, after obtaining the semantic features of each sentence in the chapter text, the semantic features of each sentence can be converted based on the semantic sense-to-speech conversion relationship, thereby obtaining the sense-of-speech features of each sentence. The semantic sense-to-speech conversion relationship here can be understood as a part of the mapping relationship referred to in the above embodiment, that is, the above mapping relationship can be divided into two parts: the text semantic conversion relationship and the semantic sense-to-speech conversion relationship. The semantic sense-to-speech conversion relationship here can be reflected as a sense-of-speech extraction model obtained through supervised training of sample semantic features and sample sense-of-speech features of each sentence in the sample chapter text, or it can be reflected as a sense-of-speech extraction rule obtained by association mining of sample semantic features and sample sense-of-speech features of each sentence in the sample chapter text. The embodiment of the present invention does not specifically limit this.

[0075] Here, the semantic sense-of-language conversion relationship can constitute a sense-of-language extraction model together with the module used for semantic extraction in step 210, or it can be used as a sense-of-language extraction model independently of the module used for semantic extraction in step 210. When a neural network is used to represent the semantic sense-of-language conversion relationship, its network structure can be an LSTM (Long Short-Term Memory) plus a linear projection layer, or it can be other structures that can achieve mapping relationship representation. The embodiment of the present invention does not specifically limit this. Accordingly, in the case where the semantic sense-of-language conversion relationship and the module used for semantic extraction in step 210 constitute a sense-of-language extraction model together, training can be performed based on each sentence in the sample chapter text and its sample sense-of-language features as samples. In the case where the semantic sense-of-language conversion relationship is used as a sense-of-language extraction model independently of the module used for semantic extraction in step 210, training can be performed based on the semantic features of each sentence in the sample chapter text and its sample sense-of-language features as samples.

[0076] Based on any of the above embodiments, Figure 3 is a flow chart of the sample language sense feature extraction method provided by the present invention, such as Figure 3 As shown, the sample language sense feature is determined based on the following steps:

[0077] Step 310, encoding the acoustic features of the real speech corresponding to the sample passage text to obtain the speech features of the real speech;

[0078] Step 320, based on the speech-sense conversion relationship, convert the speech feature into a sense of language to obtain sample sense of language features of each sentence in the sample chapter text;

[0079] The speech-sense conversion relationship is obtained by comparative learning based on the sentence-level features of each sentence in the speech feature, taking the local features of each sentence in the speech feature as positive example points and taking the local features of other sentences in the speech feature as negative example points.

[0080] Specifically, for the sample passage text, the real speech corresponding to the sample passage text can be first obtained, and then the sample language sense features can be extracted based on the real speech. Furthermore, the acoustic features of the real speech can be first obtained. The acoustic features here can be obtained by extracting the real speech through fast Fourier transform FFT after framing and windowing, such as Mel Frequency Cepstrum Coefficient (MFCC) features or Perceptual Linear Predictive (PLP) features. Subsequently, the acoustic features of the real speech can be further extracted to obtain frame-level speech features.

[0081] On this basis, the extracted speech features of the real speech can be converted based on the speech emotion conversion relationship to obtain the language sense features of the real speech, that is, the sample language sense features of each sentence in the sample passage text corresponding to the real speech.

[0082] Here, the speech emotion conversion relationship can be reflected as a language sense extraction model obtained through training, or as a language sense extraction rule obtained through association mining. Considering that in the stage of acquiring the speech emotion conversion relationship, the language sense characteristics of the real speech are unknown, the determination of the speech emotion conversion relationship is achieved through comparative learning in the embodiment of the present invention.

[0083] The key to contrastive learning is how to select appropriate sets of positive and negative examples to assist the anchor point in learning useful information. Specifically, in the embodiment of the present invention, the anchor point is the sample language feature c of each sentence in the sample chapter text. sent , the sample sense features of each sentence in the sample text c sent It is determined based on the sentence-level features of each sentence in the speech features of the real speech corresponding to the sample passage text. The sentence-level features referred to here are the frame-level speech features of the sentence obtained by dividing the speech features of the real speech into sentences. Considering the sentence-level information and the local information in the sentence-level information, that is, the sentence-level features of each sentence and the local features of each sentence should be highly correlated, the sample sense of language features obtained based on the sentence-level features of each sentence should also be highly correlated with the local features of the corresponding sentence.

[0084] Based on this, the sample sense feature c of any sentence in the sample text is sent In this case, we can randomly select local features from the sentence-level features of the sentence as positive example points, and select local features from the sentence-level features of other sentences as negative example points for comparative learning, so as to obtain a speech-sense conversion relationship that can realize language sense conversion.

[0085] The method provided by the embodiment of the present invention extracts the sense of speech by using the speech-sense conversion relationship obtained through comparative learning, which helps to improve the representation ability of the sample sense of speech features, thereby improving the fidelity of speech synthesis.

[0086] Based on any of the above embodiments, Figure 4 is a structural diagram of the language sense feature extraction model provided by the present invention, such as Figure 4 As shown in the figure, the real speech is expressed in the form of speech waveform. By extracting the acoustic features of the real speech, the acoustic features of the real speech can be obtained, that is, the acoustic features of the real speech shown in the figure, t-2 、x t-1 、x t , …, where x t is the acoustic feature of the tth frame. The encoder in the figure can perform the acoustic feature encoding of step 310 to obtain the speech features of the real speech. The speech features here are also at the frame level, that is, as shown in the figure..., z t-2 、z t-1 、z t ,…. Assuming that the t-2th frame to the t+3th frame corresponds to a sentence, the feature extractor in the figure can be used to extract z t-2 To z t+3 The sentence-level features of the sentence are transformed into language sense, so as to obtain the sample language sense feature c of the sentence sent .

[0087] Here, the sense of language conversion is performed through the feature extractor, which can be specifically to further extract the sentence-level features, and then average the vectors in a sentence obtained by feature extraction as the sample sense of language feature c of the abstracted sentence sent .

[0088] Accordingly, for the language sense feature extraction model training, the anchor sample language sense feature c sent Positive example points It can be derived from the sentence-level feature z of the clause t-2 To z t+3 The local features randomly selected from t The counterexample point can be a local feature in the sentence-level features of other clauses. For example, multiple local features in other clauses can be randomly selected to construct a set of counterexample points. For example, we can randomly select 300 z from other clauses. t As a set of counterexample points.

[0089] Specifically, in the process of contrastive learning, InfoNCE loss can be used as a loss function to drive the update of the language sense feature extraction model. InfoNCE loss is shown in the following formula:

[0090]

[0091] Where, L N That is, InfoNCE loss, f(c sent ,z t )=exp(c sent ·z t ). After the model training converges, the acoustic features of the real speech are input into the language sense feature extraction model to obtain the sample language sense features of each sentence in the sample passage text corresponding to the real speech output by the model.

[0092] Based on any of the above embodiments, step 130 includes:

[0093] The phonetic features and the sense of language features of each sentence in the text are integrated in units of sentences to obtain integrated features of each sentence in the text;

[0094] Based on the fusion features of each sentence in the passage text, speech synthesis is performed to obtain the synthesized speech of the passage text.

[0095] Specifically, the phonetic features obtained by modeling the entire text of the passage, as well as the language features that reflect the rhythm, emotional trend, etc. of the sentences in the passage text, can be fused on a sentence basis. For example, the phonetic features of each sentence can be located from the phonetic features of the passage text on a sentence basis, and the phonetic features of each sentence can be concatenated with the language features of the corresponding sentence, and the concatenated features can be used as the fused features of each sentence. Alternatively, the concatenated features of each sentence can be re-encoded through a bidirectional LSTM or RNN (Recurrent Neural Network), and the re-encoded features can be used as the fused features of each sentence.

[0096] After obtaining the fusion features of each sentence in the text of the passage, speech synthesis can be achieved by performing speech decoding on the fusion features of each sentence, thereby obtaining the synthesized speech of the text of the passage.

[0097] Based on any of the above embodiments, Figure 5 is a flow chart of step 120 in the method for extracting language sense features provided by the present invention, such as Figure 5 As shown, step 120 includes:

[0098] Step 121, encoding the text phoneme sequence to obtain the phoneme-level vector of the text;

[0099] Step 122, predicting the duration of each phoneme in the text phoneme sequence based on the phoneme-level vector;

[0100] Step 123: up-sample the phoneme-level vector based on the duration of each phoneme in the text phoneme sequence to obtain the phonetic feature.

[0101] Specifically, the process of encoding the text phoneme sequence can be achieved in a non-autoregressive manner. First, a network such as self-attention, multi-layer self-attention or multi-head self-attention can be used to perform nonlinear encoding on the text phoneme sequence to extract the phonetic features of the text phoneme sequence, thereby obtaining the phoneme-level vector of the text, which can be recorded as memory.

[0102] Subsequently, the duration of each phoneme in the text phoneme sequence can be predicted for the encoded phoneme-level vector. Specifically, the phoneme-level vector can be further feature encoded through a network in the form of LSTM, bidirectional LSTM or RNN, and the features obtained by the further feature encoding are then used to predict the duration, thereby obtaining the duration of each phoneme in the text phoneme sequence in the synthesized speech, that is, the duration of each phoneme in the text phoneme sequence.

[0103] On this basis, the vector of each phoneme in the phoneme-level vector can be upsampled based on the duration of each phoneme in the text phoneme sequence. After upsampling, the frame length reflected by the vector of each phoneme corresponds to the duration of each phoneme, thereby obtaining the phonetic features at the frame level. For example, there are 3 phonemes in the text phoneme sequence, and the phoneme-level vector memory is represented as [h1, h2, h3], where the duration corresponding to each phoneme is [2, 3, 2]. Then the output after upsampling is copied, that is, the phonetic features can be [h1, h1, h2, h2, h2, h3, h3]. It should be noted that if the predicted duration is likely not an integer, the duration needs to be rounded off. For example, if the predicted duration of each phoneme is [2.5, 4.3, 2.7], it needs to be rounded off to [3, 4, 3] before upsampling.

[0104] The method provided by the embodiment of the present invention realizes speech synthesis with higher generation efficiency and stability through a non-autoregressive approach.

[0105] Based on any of the above embodiments, Figure 6 Schematic diagram of the structure of the speech synthesis system provided by the present invention. Figure 6 As shown in the figure, speech synthesis needs to rely on three parts: non-autoregressive acoustic module, abstract encoding module and BERT prediction module. Figure 6 In the figure, the dotted arrows are only effective during training, the dashed arrows are only effective during application, the solid arrows are effective both during training and application, and the solid line of the double arrows represents the error loss during training.

[0106] Among them, the non-autoregressive acoustic module, whose main function is to construct the mapping relationship between the input chapter phoneme sequence and the output acoustic features, contains four modules: encoder, duration prediction module, upsampling module and decoder. The input of the non-autoregressive acoustic module is a sequence of the phoneme-level text features of the entire chapter, that is, the chapter phoneme sequence. The encoder performs nonlinear encoding to obtain the phoneme-level vector of the chapter text, which is recorded as memory. The phoneme-level vector memory predicts the duration of each phoneme through the duration prediction module, and calculates the error with the actual duration of each phoneme in the real speech to drive the learning of the duration prediction module. At the same time, the phoneme-level vector memory and the phoneme duration (the actual duration during training and the predicted duration output by the duration prediction module during application) are input into the upsampling module for expansion to obtain the frame-level phonetic features of the same scale as the acoustic features. Finally, after concatenating the frame-level phonetic features and language sense features (sample language sense features during training and language sense features output by the BERT prediction module during application), the acoustic features are predicted by the decoder, and the error between the predicted acoustic features and the true acoustic features is used to drive the learning of the encoder and decoder in the non-autoregressive acoustic module.

[0107] The main function of the abstract coding module is to extract the rhythm, emotion and other language features of each sentence from the acoustic features, so that the generated speech is closer to real people. The main structure includes a feature extraction layer and a pooling layer. The input of the abstract coding module is the acoustic features of the real speech. After being encoded by the feature extraction layer, it is downsampled to the vector of each sentence in the real speech through the pooling layer, which is called the sample language feature of each sentence. The sample language features generated by this module will be spliced ​​with the frame-level phonetic features in the non-autoregressive acoustic module to jointly guide the generation of acoustic features. It should be noted that the abstract coding module is learned in advance by contrastive learning, and the abstract coding module is fixed when training the non-autoregressive acoustic module.

[0108] BERT prediction module. Since the real sample language features are extracted from the acoustic features, a tool is needed to predict the language features of each sentence when synthesizing speech. The BERT prediction module is used to perform this function. The BERT prediction module includes a BERT module and an autoregressive prediction module. The BERT module is pre-trained with a large number of corpus data and then fixed. The input of the BERT prediction module is the chapter text, that is, the word-level text of the entire chapter. After BERT, the word-level encoding vector can be obtained, and the word-level encoding vector of each sentence can be obtained. <cls>The vector corresponding to the label is used as the semantic feature of each sentence. The language sense features of each sentence are modeled through the autoregressive prediction module. The error between the predicted language sense features and the actual sample language sense features is used to drive the learning of the autoregressive prediction module.

[0109] The method provided by the embodiment of the present invention performs speech synthesis through a non-autoregressive acoustic module, thereby ensuring the stability and high efficiency of speech synthesis; and in the process of speech synthesis, the entire text is jointly modeled, so that each sentence in the text can see the information of multiple sentences before and after during modeling, which helps to improve the coherence between the upper and lower sentences during the synthesis of the entire paragraph. In addition, the rhythm and emotional trend of the text speech are controlled through large-scale sentence-level language sense features, ensuring the rationality of the rhythm and emotion of each sentence in the synthesized speech.

[0110] Based on any of the above embodiments, Figure 7 is a structural schematic diagram of the speech synthesis device provided by the present invention, such as Figure 7 As shown, the device comprises:

[0111] The phoneme determination unit 710 is used to determine the text phoneme sequence of the text to be synthesized;

[0112] A chapter encoding unit 720, used for encoding the chapter phoneme sequence to obtain the phonetic features of the chapter text;

[0113] The speech synthesis unit 730 is used to perform speech synthesis based on the phonetic features to obtain the synthesized speech of the passage text.

[0114] The speech synthesis device provided in the embodiment of the present invention encodes the chapter phoneme sequence of the chapter text to obtain the phonetic features for the overall modeling of the chapter text, and performs speech synthesis based on the obtained phonetic features, which can ensure the coherence of the synthesized speech at the level of rhythm, emotion and other sense of language, and improve the naturalness of the synthesized speech.

[0115] Based on any of the above embodiments, the speech synthesis unit 730 is used for:

[0116] Based on the phonetic features and the language sense features of each sentence in the text, speech synthesis is performed to obtain the synthesized speech of the text.

[0117] Based on any of the above embodiments, the device further includes:

[0118] A language sense extraction unit, used to extract the language sense of each sentence in the sample text based on the sample language sense features of each sentence in the sample text, so as to obtain the language sense features of each sentence in the sample text;

[0119] The sample language sense feature is obtained by extracting the language sense feature of the real speech corresponding to the sample chapter text.

[0120] Based on any of the above embodiments, the language sense extraction unit is used to:

[0121] Performing semantic extraction on each sentence in the text of the article to obtain semantic features of each sentence in the text of the article;

[0122] Based on the semantic sense conversion relationship, the semantic features of each sentence in the text are converted into sense of language to obtain the sense of language features of each sentence in the text;

[0123] The semantic sense-of-language conversion relationship is determined based on the sample semantic features and sample sense-of-language features of each sentence in the sample chapter text.

[0124] Based on any of the above embodiments, the device further includes:

[0125] A sample speech sense acquisition unit, used for encoding the acoustic features of the real speech corresponding to the sample passage text to obtain the speech features of the real speech;

[0126] Based on the speech-sense conversion relationship, the speech features are converted into a sense of language to obtain sample sense of language features of each sentence in the sample chapter text;

[0127] The speech-sense conversion relationship is obtained by comparative learning based on the sentence-level features of each sentence in the speech feature, taking the local features of each sentence in the speech feature as positive example points and taking the local features of other sentences in the speech feature as negative example points.

[0128] Based on any of the above embodiments, the speech synthesis unit 730 is used for:

[0129] The phonetic features and the sense of language features of each sentence in the text are integrated in units of sentences to obtain integrated features of each sentence in the text;

[0130] Based on the fusion features of each sentence in the passage text, speech synthesis is performed to obtain the synthesized speech of the passage text.

[0131] Based on any of the above embodiments, the chapter encoding unit 720 includes:

[0132] Encoding the text phoneme sequence to obtain a phoneme-level vector of the text;

[0133] Based on the phoneme-level vector, predict the duration of each phoneme in the text phoneme sequence;

[0134] Based on the duration of each phoneme in the text phoneme sequence, the phoneme-level vector is upsampled to obtain the phonetic feature.

[0135] Figure 8 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 8 As shown, the electronic device may include: a processor 810, a communication interface 820, a memory 830 and a communication bus 840, wherein the processor 810, the communication interface 820 and the memory 830 communicate with each other through the communication bus 840. The processor 810 may call the logic instructions in the memory 830 to execute the speech synthesis method, which includes:

[0136] Determining a text phoneme sequence of a text to be synthesized;

[0137] Encoding the text phoneme sequence to obtain the phonetic features of the text;

[0138] Speech synthesis is performed based on the phonetic features to obtain synthesized speech of the passage text.

[0139] In addition, the logic instructions in the above-mentioned memory 830 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk and other media that can store program codes.

[0140] On the other hand, the present invention further provides a computer program product, the computer program product comprising a computer program stored on a non-transitory computer-readable storage medium, the computer program comprising program instructions, and when the program instructions are executed by a computer, the computer can execute the speech synthesis method provided by the above methods, the method comprising:

[0141] Determining a text phoneme sequence of a text to be synthesized;

[0142] Encoding the text phoneme sequence to obtain the phonetic features of the text;

[0143] Speech synthesis is performed based on the phonetic features to obtain synthesized speech of the passage text.

[0144] In another aspect, the present invention further provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the computer program is executed by a processor to perform the above-mentioned speech synthesis method, the method comprising:

[0145] Determining a text phoneme sequence of a text to be synthesized;

[0146] Encoding the text phoneme sequence to obtain the phonetic features of the text;

[0147] Speech synthesis is performed based on the phonetic features to obtain synthesized speech of the passage text.

[0148] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0149] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0150] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.< / cls> < / cls> < / sep> < / cls> < / sep> < / cls> < / sep> < / cls>

Claims

1. A speech synthesis method, characterized in that: include: Determining a text phoneme sequence of a text to be synthesized; Encoding the text phoneme sequence to obtain phonetic features for overall modeling of the text; Perform speech synthesis based on the phonetic features to obtain synthesized speech of the passage text; The encoding of the text phoneme sequence to obtain the phonetic features of the text includes: Encoding the text phoneme sequence to obtain a phoneme-level vector of the text; Based on the phoneme-level vector, predict the duration of each phoneme in the text phoneme sequence; Based on the duration of each phoneme in the text phoneme sequence, upsampling the phoneme-level vector to obtain the phonetic feature; The performing speech synthesis based on the phonetic features to obtain the synthesized speech of the passage text includes: Based on the phonetic features and the language sense features of each sentence in the text, speech synthesis is performed to obtain the synthesized speech of the text, and the language sense features reflect the characteristics of the rhythm and emotional trend of the corresponding sentence in the text.

2. The speech synthesis method according to claim 1, characterized in that: The sense of language features of each sentence in the text of the passage are determined based on the following steps: Based on the sample language sense features of each sentence in the sample text, extract the language sense of each sentence in the text to obtain the language sense features of each sentence in the text; The sample language sense feature is obtained by extracting the language sense feature of the real speech corresponding to the sample chapter text.

3. The speech synthesis method according to claim 2, characterized in that: The method of extracting the sense of language of each sentence in the sample text based on the sample sense of language of each sentence in the sample text to obtain the sense of language of each sentence in the sample text includes: Performing semantic extraction on each sentence in the text of the article to obtain semantic features of each sentence in the text of the article; Based on the semantic sense conversion relationship, the semantic features of each sentence in the text are converted into sense of language to obtain the sense of language features of each sentence in the text; The semantic sense-of-language conversion relationship is determined based on the sample semantic features and sample sense-of-language features of each sentence in the sample chapter text.

4. The speech synthesis method according to claim 2, characterized in that: The sample language sense features are determined based on the following steps: Encoding the acoustic features of the real speech corresponding to the sample passage text to obtain the speech features of the real speech; Based on the speech-sense conversion relationship, the speech features are converted into a sense of language to obtain sample sense of language features of each sentence in the sample chapter text; The speech-sense conversion relationship is obtained by comparative learning based on the sentence-level features of each sentence in the speech feature, taking the local features of each sentence in the speech feature as positive example points and taking the local features of other sentences in the speech feature as negative example points.

5. The speech synthesis method according to claim 1, characterized in that: The method of performing speech synthesis based on the phonetic features and the sense of language features of each sentence in the text to obtain the synthesized speech of the text includes: The phonetic features and the sense of language features of each sentence in the text are integrated in units of sentences to obtain integrated features of each sentence in the text; Based on the fusion features of each sentence in the passage text, speech synthesis is performed to obtain the synthesized speech of the passage text.

6. A speech synthesis device, characterized in that: include: A phoneme determination unit, used for determining a text phoneme sequence of a text to be synthesized; A chapter encoding unit, used for encoding the chapter phoneme sequence to obtain phonetic features for modeling the chapter text as a whole; A speech synthesis unit, used for performing speech synthesis based on the phonetic features to obtain synthesized speech of the passage text; The chapter coding unit is specifically used for: Encoding the text phoneme sequence to obtain a phoneme-level vector of the text; Based on the phoneme-level vector, predict the duration of each phoneme in the text phoneme sequence; Based on the duration of each phoneme in the text phoneme sequence, upsampling the phoneme-level vector to obtain the phonetic feature; The speech synthesis unit is specifically used for: Based on the phonetic features and the language sense features of each sentence in the text, speech synthesis is performed to obtain the synthesized speech of the text, and the language sense features reflect the characteristics of the rhythm and emotional trend of the corresponding sentence in the text.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the steps of the speech synthesis method according to any one of claims 1 to 5 are implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the speech synthesis method according to any one of claims 1 to 5 are implemented.

Citation Information

Patent Citations

  • Speech synthesis method and device and electronic equipment

    CN113628610A