Acoustic feature determination method and apparatus, electronic device, and speech synthesis system
Patent Information
- Application Number
- CN202311868745.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-28
- Publication Date
- 2026-09-25
- Estimated Expiration
- 2043-12-28
AI Technical Summary
但是,将整个语句的声学特征均预测得到后,再将声学特征输出并转换成语音,会使得确定声学特征的响应时间较长,进而影响语音合成的响应速度
[0045]本申请提出的声学特征确定方法,获取待转换语句对应的短句集合中的各个短句对应的音素韵律特征;其中,短句集合中包括对待转换语句进行语句切分后得到的至少两个短句;对短句集合中的各个短句对应的音素韵律特征进行上下文特征提取处理,并从提取的特征中确定出第一短句对应的历史音素韵律特征和未来音素韵律特征;其中,第一短句为短句集合中的任意一个短句;基于第一短句对应的音素韵律特征、历史音素韵律特征和未来音素韵律特征,预测第一短句对应的声学特征,声学特征用于合成与第一短句对应的语音。采用本申请的技术方案,对待转换语句切分成的每个短句的声学特征进行依次预测,相比于对待转换语句的声学特征进行整句预测,能够实现声学特征的流式预测,减少确定声学特征的响应时长,进而提高语音合成的响应速度。
Smart Images

Figure CN117877463B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of speech synthesis technology, and in particular to a method, apparatus, electronic device and speech synthesis system for determining acoustic features. Background Technology
[0002] A speech synthesis system typically consists of three parts: a front-end, an acoustic model, and a vocoder. The front-end mainly performs functions such as text processing, text-to-phoneme conversion, and prosodic pause prediction; the acoustic model mainly converts phonemes into acoustic features directly related to speech; and the vocoder converts acoustic features into speech.
[0003] Existing acoustic models typically employ a non-streaming structure, using the entire sentence as the basic processing unit. Based on the phonemic prosodic features of the entire sentence, they predict all acoustic features of the sentence before outputting and converting them into speech. However, predicting all acoustic features of the entire sentence before outputting and converting them into speech results in a longer response time for determining the acoustic features, thus affecting the response speed of speech synthesis. Summary of the Invention
[0004] Based on the above requirements, this application proposes an acoustic feature determination method, apparatus, electronic device, and speech synthesis system, which can reduce the response time for determining acoustic features, thereby improving the response speed of speech synthesis.
[0005] To achieve the above objectives, this application proposes the following technical solution:
[0006] According to a first aspect of the embodiments of this application, an acoustic feature determination method is provided, comprising:
[0007] Obtain the phonemic prosodic features of each phrase in the set of phrases corresponding to the statement to be converted; wherein, the set of phrases includes at least two phrases obtained after segmenting the statement to be converted;
[0008] The phoneme prosodic features corresponding to each short sentence in the short sentence set are subjected to context feature extraction processing, and the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence are determined from the extracted features; wherein, the first short sentence is any short sentence in the short sentence set.
[0009] Based on the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features corresponding to the first short sentence, the acoustic features corresponding to the first short sentence are predicted; the acoustic features are used to synthesize the speech corresponding to the first short sentence.
[0010] Optionally, contextual feature extraction is performed on the phoneme prosodic features corresponding to each short phrase in the phrase set, and the historical and future phoneme prosodic features corresponding to the first short phrase are determined from the extracted features, including:
[0011] Contextual feature extraction is performed on the phoneme prosodic features corresponding to each short sentence in the short sentence set to obtain the contextual phoneme prosodic features corresponding to the short sentence set.
[0012] Based on the features corresponding to historical short sentences in the contextual phoneme prosodic features, the historical phoneme prosodic features corresponding to the first short sentence are determined; the historical short sentence is the short sentence preceding the first short sentence in the short sentence set.
[0013] Based on the features corresponding to future sentences in the context phoneme prosodic features, the future phoneme prosodic features corresponding to the first short sentence are determined; the future short sentence is the short sentence following the first short sentence in the short sentence set.
[0014] Optionally, contextual feature extraction is performed on the phoneme prosodic features corresponding to each short phrase in the short phrase set, and the historical and future phoneme prosodic features corresponding to the first short phrase are determined from the extracted features. Based on the phoneme prosodic features, historical and future phoneme prosodic features corresponding to the first short phrase, the acoustic features corresponding to the first short phrase are predicted, including:
[0015] The phonemic prosodic features corresponding to each short sentence in the short sentence set and the phonemic prosodic features corresponding to the first short sentence are both input into the pre-trained first acoustic model to obtain the acoustic features corresponding to the first short sentence.
[0016] The first acoustic model performs contextual feature extraction on the phoneme prosodic features corresponding to each short sentence in the short sentence set, and determines the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence from the extracted features. Based on the phoneme prosodic features, historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence, the acoustic features corresponding to the first short sentence are predicted.
[0017] Optionally, the training process of the first acoustic model includes:
[0018] A set of sample phrases corresponding to the sample conversion statement is pre-collected, and the phoneme prosodic features corresponding to each sample phrase in the sample phrase set are determined; the sample phrase set includes at least two sample phrases obtained after sentence segmentation of the sample conversion statement;
[0019] The phonemic prosodic features corresponding to each sample phrase in the sample phrase set and the phonemic prosodic features corresponding to the first sample phrase are input into the first acoustic model to obtain the first sample prediction information corresponding to the first sample phrase. Additionally, the phonemic prosodic features corresponding to each sample phrase in the sample phrase set are input into a pre-trained second acoustic model to obtain the second sample prediction information corresponding to the sample phrase set. The first sample phrase is any one of the sample phrases in the sample phrase set.
[0020] Based on the first sample prediction information and the second sample prediction information, the model parameters of the first acoustic model are adjusted.
[0021] Optionally, the first sample prediction information includes: the first phoneme prosodic coding feature and the first sample acoustic feature corresponding to the first sample phrase; the second sample prediction information includes: the second phoneme prosodic coding feature and the second sample acoustic feature corresponding to the sample phrase set.
[0022] Wherein, the first phoneme prosodic coding feature is obtained by the encoder in the first acoustic model based on the phoneme prosodic features corresponding to the first sample phrase, the sample historical phoneme prosodic features, and the sample future phoneme prosodic features; the first sample acoustic feature is obtained by the decoder in the first acoustic model by decoding the first phoneme prosodic coding feature; the second phoneme prosodic coding feature is obtained by the encoder in the second acoustic model based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set; the second sample acoustic feature is obtained by the decoder in the second acoustic model by decoding the second phoneme prosodic coding feature.
[0023] Optionally, based on the first sample prediction information and the second sample prediction information, the model parameters of the first acoustic model are adjusted, including:
[0024] Based on the loss function between the first phoneme prosodic coding feature and the second phoneme prosodic coding feature, and the loss function between the first sample acoustic feature and the second sample acoustic feature, the model parameters of the first acoustic model are adjusted.
[0025] Optionally, the training process of the second acoustic model includes:
[0026] Acquire first sample recording data from at least two speakers, the first sample recording data including: first sample recording text and first sample audio;
[0027] The phoneme prosodic features corresponding to the first sample recording text are input into the second acoustic model to obtain the first predicted acoustic features corresponding to the first sample recording text.
[0028] The model parameters of the second acoustic model are adjusted based on the loss function between the first predicted acoustic features and the real acoustic features corresponding to the first sample audio.
[0029] Optionally, the training process of the second acoustic model further includes:
[0030] Acquire second sample recording data of the target speaker, the second sample recording data including: second sample recording text and second sample audio;
[0031] The phoneme prosodic features corresponding to the second sample recording text are input into the second acoustic model to obtain the second predicted acoustic features corresponding to the second sample recording text.
[0032] Based on the loss function between the second predicted acoustic features and the real acoustic features corresponding to the second sample audio, the model parameters of the second acoustic model are fine-tuned.
[0033] According to a second aspect of the embodiments of this application, an acoustic feature determination apparatus is provided, comprising:
[0034] The feature acquisition module is used to acquire the phoneme prosodic features of each short sentence in the short sentence set corresponding to the sentence to be converted; wherein, the short sentence set includes at least two short sentences obtained after sentence segmentation of the sentence to be converted;
[0035] The context feature extraction module is used to perform context feature extraction processing on the phoneme prosodic features corresponding to each short sentence in the short sentence set, and to determine the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence from the extracted features; wherein, the first short sentence is any short sentence in the short sentence set.
[0036] An acoustic feature prediction module is used to predict the acoustic features corresponding to the first short sentence based on the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features corresponding to the first short sentence; the acoustic features are used to synthesize speech corresponding to the first short sentence.
[0037] According to a third aspect of the embodiments of this application, an electronic device is provided, including: a memory and a processor;
[0038] The memory is connected to the processor and is used to store programs;
[0039] The processor is used to implement the above-described acoustic feature determination method by running a program in the memory.
[0040] According to a fourth aspect of the embodiments of this application, a storage medium is provided, on which a computer program is stored, and when the computer program is executed by a processor, it implements the above-described acoustic feature determination method.
[0041] According to a fifth aspect of the embodiments of this application, a speech synthesis system is provided, including: a front-end device, an acoustic feature determination device, and a vocoder;
[0042] The front-end device is used to determine the phonemic prosodic features of each short sentence in the short sentence set corresponding to the sentence to be converted; the short sentence set includes at least two short sentences obtained by segmenting the sentence to be converted.
[0043] The acoustic feature determination device is used to determine the acoustic features corresponding to the short sentences in the short sentence set by performing the above-described acoustic feature determination method;
[0044] The vocoder is used to determine the speech of the short sentences corresponding to the short sentences in the short sentence set based on the acoustic features corresponding to the short sentences in the short sentence set.
[0045] The acoustic feature determination method proposed in this application obtains the phoneme prosodic features of each sentence in a set of short sentences corresponding to the sentence to be converted. The set of short sentences includes at least two sentences obtained after segmenting the sentence to be converted. Contextual feature extraction is performed on the phoneme prosodic features of each sentence in the set of short sentences, and the historical and future phoneme prosodic features of the first sentence are determined from the extracted features. The first sentence can be any sentence in the set of short sentences. Based on the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features of the first sentence, the acoustic features corresponding to the first sentence are predicted. These acoustic features are used to synthesize the speech corresponding to the first sentence. By adopting the technical solution of this application, the acoustic features of each sentence segmented into the sentence to be converted are predicted sequentially. Compared with whole-sentence prediction of the acoustic features of the sentence to be converted, this method enables streaming prediction of acoustic features, reduces the response time for determining acoustic features, and thus improves the response speed of speech synthesis. Attached Figure Description
[0046] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of this application. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0047] Figure 1A flowchart illustrating an acoustic feature determination method provided in an embodiment of this application;
[0048] Figure 2 A schematic diagram of the processing flow for training the first acoustic model provided in an embodiment of this application;
[0049] Figure 3 A schematic diagram of the structure of the first acoustic model and the second acoustic model provided in the embodiments of this application;
[0050] Figure 4 A schematic diagram of a process for training a second acoustic model provided in an embodiment of this application;
[0051] Figure 5 A schematic diagram of another process for training a second acoustic model provided in an embodiment of this application;
[0052] Figure 6 This is a schematic diagram of the structure of an acoustic feature determination device provided in an embodiment of this application;
[0053] Figure 7 This is a schematic diagram of the structure of a speech synthesis system provided in an embodiment of this application;
[0054] Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0055] The technical solution of this application is applicable to speech synthesis application scenarios, specifically in the application scenario of acoustic feature prediction in speech synthesis. By adopting the technical solution of this application, the acoustic features of each short sentence segmented into the sentence to be converted are predicted sequentially. Compared with whole-sentence prediction of the acoustic features of the sentence to be converted, this enables streaming prediction of acoustic features, reduces the response time for determining acoustic features, and thus improves the response speed of speech synthesis.
[0056] Speech synthesis is a method of converting text into speech (Text-to-Speech, TTS), typically using a speech synthesis system to synthesize the speech corresponding to the text. A speech synthesis system usually consists of three parts: a front-end, an acoustic model, and a vocoder. The front-end mainly performs functions such as text processing, text-to-phoneme conversion, and prosodic pause prediction; the acoustic model mainly converts phonemes into acoustic features directly related to speech; and the vocoder converts the acoustic features into speech.
[0057] Existing acoustic models typically employ a non-streaming architecture, using the entire sentence as the basic processing unit. Based on the phonemic prosodic features of the entire sentence, they predict all acoustic features before outputting and converting them into speech. However, predicting all acoustic features of the entire sentence before outputting and converting them into speech results in a long response time for determining acoustic features, thus affecting the response speed of speech synthesis. This is especially true in device-oriented speech synthesis systems (such as mobile phones, dictionary pens, in-vehicle systems, or other embedded systems), where limited computing power and memory resources further reduce the speed of acoustic feature prediction for the entire sentence's phonemic prosodic features, increasing the response time for determining acoustic features and ultimately reducing the overall response speed of speech synthesis in device-oriented systems.
[0058] Based on this, this application proposes an acoustic feature determination method. This technical solution can sequentially predict the acoustic features of each short sentence segmented into which the sentence to be converted is to be converted, thereby realizing the streaming prediction of acoustic features and solving the problems of long response time and low response speed of speech synthesis in the prior art.
[0059] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0060] Exemplary methods
[0061] See Figure 1 As shown in the figure, this application proposes a method for determining acoustic features. The method includes:
[0062] S101. Obtain the phoneme prosodic features of each short sentence in the set of short sentences corresponding to the sentence to be converted.
[0063] In this embodiment, the text statement to be synthesized is taken as the statement to be converted. The statement to be converted is segmented into at least two short sentences, and all the short sentences obtained after segmentation are combined into a short sentence set corresponding to the statement to be converted. That is, the short sentence set includes all the short sentences after the statement to be converted is segmented. For example, the statement to be converted can be segmented according to the punctuation marks in the statement to be converted. If the statement to be converted is "Navigation begins, starting point: Outer Ring West Road, destination: Wangjing North Road, total length: approximately 36 kilometers.", the statement to be converted can be segmented into four short sentences: "Navigation begins," "starting point: Outer Ring West Road," "destination: Wangjing North Road," and "total length: approximately 36 kilometers." Then, the short sentence set corresponding to the statement to be converted will contain the above four short sentences. If the statement between two punctuation marks is too long, the pause position in the statement can also be predicted for sentence segmentation.
[0064] This embodiment segments the sentence to be converted into a set of short sentences using the method described above. Then, it processes the text of each short sentence to determine its corresponding phonemic prosodic features. Specifically, it performs text processing operations such as word segmentation, annotation, and grammatical analysis to convert the text of the short sentences into phonemes and predicts the prosodic pauses of the short sentences, thereby determining their phonemic prosodic features. This embodiment can directly obtain the phonemic prosodic features of each short sentence in the set of short sentences corresponding to the sentence to be converted, determined by the front-end device in the speech synthesis system.
[0065] S102. Perform context feature extraction on the phoneme prosodic features corresponding to each short sentence in the short sentence set, and determine the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence from the extracted features.
[0066] In this embodiment, after obtaining the phonemic prosodic features corresponding to each short phrase in the short phrase combination, in order to improve the prediction accuracy of the acoustic features of each short phrase, it is necessary to combine the context features of each short phrase into which the sentence to be converted is segmented, thereby expanding the receptive field when predicting the acoustic features of the short phrases. That is, the phonemic prosodic features corresponding to each short phrase in the short phrase set are processed by context feature extraction, so that features that combine the context information of the phonemic prosodic features can be extracted. That is, the features extracted after context feature extraction all contain the phonemic prosodic features of the context of the short phrase. The short phrase in the short phrase set whose acoustic features need to be predicted is taken as the first short phrase. The historical phonemic prosodic features corresponding to the first short phrase are determined from the features extracted by context feature extraction as the context features of the phonemic prosodic features corresponding to the first short phrase. The future phonemic prosodic features corresponding to the first short phrase are determined from the features extracted by context feature extraction as the context features of the phonemic prosodic features corresponding to the first short phrase. In this embodiment, the first phrase is any phrase in the phrase set. If the acoustic features of the phrases are predicted in the order they appear in the phrase set, then all phrases before the first phrase have had their acoustic features predicted, while the first phrase and all phrases after it have not. In this embodiment, a neural network can be used to construct a context feature extraction module to extract context features of the phoneme prosodic features corresponding to each phrase in the phrase set. To reduce computational load and improve the efficiency of context feature extraction, a lightweight neural network can be used, preferably a bidirectional LSTM network or an Attention network.
[0067] Furthermore, this step specifically includes:
[0068] First, the contextual features of the phonemic prosody features corresponding to each short sentence in the short sentence set are extracted to obtain the contextual phonemic prosody features corresponding to the short sentence set.
[0069] This embodiment performs contextual feature extraction processing on the phoneme prosodic features corresponding to each short sentence in the short sentence set to obtain the contextual phoneme prosodic features corresponding to the short sentence set. The contextual phoneme prosodic features include the features corresponding to each short sentence, and the features corresponding to each short sentence combine the phoneme prosodic features of the context of that short sentence.
[0070] Second, based on the features corresponding to historical short sentences in the context phoneme prosodic features, the historical phoneme prosodic features corresponding to the first short sentence are determined.
[0071] In this embodiment, the phrase in the phrase set for which acoustic features need to be predicted is designated as the first phrase. Phrases preceding the first phrase in the phrase set are designated as historical phrases. Features corresponding to the historical phrases are extracted from the contextual phoneme prosodic features corresponding to the phrase set. Then, based on the features corresponding to the historical phrases in the contextual phoneme prosodic features, the historical phoneme prosodic features corresponding to the first phrase are determined. Specifically, the historical phoneme prosodic features corresponding to the first phrase can be the set of features corresponding to the historical phrases in the contextual phoneme prosodic features, or they can be a feature fusion of the features corresponding to the historical phrases in the contextual phoneme prosodic features, with the fused features serving as the historical phoneme prosodic features. The feature fusion of the features corresponding to the historical phrases in the contextual phoneme prosodic features can be achieved by concatenating the features together, or by performing dot product calculations on the features corresponding to the historical phrases in the contextual phoneme prosodic features.
[0072] Third, based on the features corresponding to future sentences in the context phoneme prosodic features, determine the future phoneme prosodic features corresponding to the first short sentence.
[0073] In this embodiment, the phrases following the first phrase in the phrase set are considered as future phrases. Features corresponding to the future phrases are extracted from the contextual phoneme prosodic features corresponding to the phrase set. Then, based on the features corresponding to the future phrases in the contextual phoneme prosodic features, the future phoneme prosodic features corresponding to the first phrase are determined. Specifically, the future phoneme prosodic features corresponding to the first phrase can be the set of features corresponding to the future phrases in the contextual phoneme prosodic features, or they can be a feature fusion of the features corresponding to the future phrases in the contextual phoneme prosodic features, with the fused features serving as the future phoneme prosodic features. The feature fusion of the features corresponding to the future phrases in the contextual phoneme prosodic features can be achieved by concatenating the features together, or by performing dot product calculations on the features corresponding to the future phrases in the contextual phoneme prosodic features.
[0074] S103. Based on the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features corresponding to the first short sentence, predict the acoustic features corresponding to the first short sentence.
[0075] The phonemic prosodic features, historical phonemic prosodic features, and future phonemic prosodic features corresponding to the first short phrase are determined through the above steps. A pre-learned acoustic feature prediction method is then used to predict the acoustic features corresponding to the first short phrase. Once the acoustic features corresponding to the first short phrase are predicted, they are transmitted to the vocoder of the speech synthesis system. The vocoder converts the acoustic features into speech, achieving speech synthesis of the first short phrase. The acoustic feature prediction method can employ deep learning, using pre-collected labeled training samples for iterative training to learn a relatively accurate acoustic feature prediction capability. Furthermore, when predicting the acoustic features corresponding to the first short phrase, not only the phonemic prosodic features of the first short phrase are utilized, but also the phonemic prosodic features of the context of the first short phrase are combined, which improves the understanding of the short phrase and thus enhances the accuracy of the acoustic feature prediction for the first short phrase. Furthermore, based on the order of the short sentences in the short sentence set, the acoustic features of each short sentence are predicted sequentially. Since the content of a short sentence is less than the content of the entire sentence to be converted, the prediction time for the acoustic features of a short sentence is necessarily less than the prediction time for the acoustic features of the entire sentence to be converted. Therefore, compared with the non-streaming prediction of acoustic features in the prior art where the entire sentence is input and output, the streaming prediction method of predicting and outputting the acoustic features of each short sentence sequentially in this embodiment has a shorter first response time, reducing the response time for determining acoustic features. Since the acoustic features are directly converted into speech using a vocoder after they are determined, speech synthesis is achieved. Therefore, the response time for determining acoustic features is reduced, thereby improving the response speed of speech synthesis.
[0076] As described above, the acoustic feature determination method proposed in this application obtains the phoneme prosodic features of each short sentence in a set of short sentences corresponding to the sentence to be converted. The set of short sentences includes at least two short sentences obtained after segmenting the sentence to be converted. Contextual feature extraction is performed on the phoneme prosodic features of each short sentence in the set of short sentences, and the historical and future phoneme prosodic features of the first short sentence are determined from the extracted features. The first short sentence is any short sentence in the set of short sentences. Based on the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features of the first short sentence, the acoustic features corresponding to the first short sentence are predicted. These acoustic features are used to synthesize speech corresponding to the first short sentence. By adopting the technical solution of this embodiment, the acoustic features of each short sentence segmented from the sentence to be converted are predicted sequentially. Compared to whole-sentence prediction of the acoustic features of the sentence to be converted, this method enables streaming prediction of acoustic features, reduces the response time for determining acoustic features, and thus improves the response speed of speech synthesis.
[0077] As an optional implementation, another embodiment of this application discloses steps S102-S103, which involve extracting contextual features from the phoneme prosodic features corresponding to each short phrase in the short phrase set, determining the historical and future phoneme prosodic features corresponding to the first short phrase from the extracted features, and predicting the acoustic features corresponding to the first short phrase based on the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features corresponding to the first short phrase. Specifically, this includes:
[0078] The phonemic prosodic features corresponding to each short phrase in the short phrase set and the phonemic prosodic features corresponding to the first short phrase are all input into the pre-trained first acoustic model to obtain the acoustic features corresponding to the first short phrase.
[0079] Specifically, in this embodiment, a first acoustic model can be pre-built and trained. The phonemic prosodic features of each short sentence in the set of short sentences corresponding to the sentence to be converted, and the phonemic prosodic features of the first short sentence in the set of short sentences, are all input into the first acoustic model. The first acoustic model can perform context feature extraction processing on the phonemic prosodic features of each short sentence in the set of short sentences, and determine the historical phonemic prosodic features and future phonemic prosodic features of the first short sentence from the extracted features. Then, based on the phonemic prosodic features, historical phonemic prosodic features, and future phonemic prosodic features of the first short sentence, the acoustic features corresponding to the first short sentence are predicted.
[0080] The first acoustic model includes a context feature extraction module, an encoder, and a decoder. The context feature extraction module extracts context features from the phonemic prosodic features of each phrase in the phrase set, obtaining the context phonemic prosodic features corresponding to the phrase set. Based on the features corresponding to historical phrases within these context phonemic prosodic features, it determines the historical phonemic prosodic features corresponding to the first phrase. Based on the features corresponding to future phrases within these context phonemic prosodic features, it determines the future phonemic prosodic features corresponding to the first phrase. The encoder encodes the phonemic prosodic features, historical phonemic prosodic features, and future phonemic prosodic features corresponding to the first phrase, obtaining the phonemic prosodic encoded features corresponding to the first phrase. The decoder decodes the phonemic prosodic encoded features corresponding to the first phrase, predicting the acoustic features corresponding to the first phrase.
[0081] For training the first acoustic model, sample sentences and their corresponding sample speech can be pre-collected. The sample sentences are then segmented to obtain at least two sample phrases, forming a set of sample phrases corresponding to the sample sentences. This embodiment requires acoustic feature extraction from the sample speech corresponding to the sample sentences, and determining the acoustic features of each sample phrase from the extracted acoustic features. It also requires determining the phoneme prosodic features of each sample phrase in the sample phrase set. The determination of the phoneme prosodic features of the sample phrases is the same as the determination of the phoneme prosodic features of the phrases in the phrase set in the above embodiment, and will not be repeated here.
[0082] The phoneme prosodic features corresponding to each sample phrase in the sample phrase set and the phoneme prosodic features corresponding to the first sample phrase are all input into the first acoustic model to obtain the predicted acoustic features corresponding to the first sample phrase. Then, the loss function between the predicted acoustic features corresponding to the first sample phrase and the acoustic features of the sample speech corresponding to the sample phrase is used to adjust the model parameters of the first acoustic model. The first acoustic model is iteratively trained in the above manner to improve the acoustic feature prediction ability of the first acoustic model.
[0083] While the training method described above can improve the acoustic feature prediction capability of the first acoustic model, it requires a high-performing and relatively complex deep network model. This means the first acoustic model requires significant computational and memory resources for acoustic feature prediction. For device-side speech synthesis systems, which typically have limited computational and memory resources, this may be insufficient to support the first acoustic model, thus impacting its efficiency. Therefore, the training method provided in the next embodiment can also be used to train the first acoustic model.
[0084] As an optional implementation, see [link to implementation details]. Figure 2 and Figure 3 As shown, in another embodiment of this application, the training process of the first acoustic model is disclosed, including:
[0085] S201. Pre-collect a set of sample short sentences corresponding to the sample conversion statement, and determine the phoneme prosodic features corresponding to each sample short sentence in the sample short sentence set.
[0086] This embodiment pre-collects sample conversion statements, then segments the sample conversion statements to obtain at least two sample short sentences. All the sample short sentences obtained from the segmentation of the sample conversion statements are combined into a sample short sentence set. Furthermore, it is necessary to determine the phonemic prosodic features corresponding to each sample short sentence in the sample short sentence set. The method for determining the phonemic prosodic features corresponding to the sample short sentences is the same as the method for determining the phonemic prosodic features of each short sentence in the short sentence set corresponding to the statement to be converted in the previous embodiment, and will not be repeated in this embodiment. Figure 3 As shown, S represents the set of sample short sentences corresponding to the sample transformation statement, and S1, S2, S3, and S4 are sample short sentences in the set of sample short sentences S.
[0087] S202. Input the phoneme prosodic features corresponding to each sample short sentence in the sample short sentence set and the phoneme prosodic features corresponding to the first sample short sentence into the first acoustic model to obtain the first sample prediction information corresponding to the first sample short sentence. Also, input the phoneme prosodic features corresponding to each sample short sentence in the sample short sentence set into the pre-trained second acoustic model to obtain the second sample prediction information corresponding to the sample short sentence set.
[0088] This embodiment first needs to determine the first sample sentence in the sample sentence set. The first sample sentence is any one of the sample sentences in the sample sentence set. For example... Figure 3 As shown, the third sample phrase S3 in the sample phrase set is the first sample phrase. In this embodiment, the phoneme prosodic features corresponding to each sample phrase in the sample phrase set and the phoneme prosodic features corresponding to the first sample phrase are all input into the first acoustic model. The first acoustic model can output the first sample prediction information corresponding to the first sample phrase. The first sample prediction information includes the first phoneme prosodic coding features and the first sample acoustic features corresponding to the first sample phrase.
[0089] Specifically, the context feature extraction module in the first acoustic model performs context feature extraction on the phoneme prosodic features corresponding to each sample sentence in the sample sentence set, obtaining the context phoneme prosodic features corresponding to the sample sentence set. Then, based on the features corresponding to the sample historical sentences in the context phoneme prosodic features corresponding to the sample sentence set, the sample historical phoneme prosodic features corresponding to the first sample sentence are determined. Based on the features corresponding to the sample future sentences in the context phoneme prosodic features corresponding to the sample sentence set, the sample future phoneme prosodic features corresponding to the first sample sentence are determined. Here, the sample historical sentences are the sample sentences preceding the first sample sentence in the sample sentence set, and the sample future sentences are the sample sentences following the first sample sentence in the sample sentence set. The context feature extraction module transmits the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first sample phrase to the encoder in the first acoustic model. The encoder in the first acoustic model encodes the phoneme prosodic features, historical phoneme prosodic features, and future phoneme prosodic features corresponding to the first sample phrase to obtain the first phoneme prosodic encoding feature corresponding to the first sample phrase. This first phoneme prosodic encoding feature is then transmitted to the decoder in the first acoustic model. The decoder decodes the first phoneme prosodic encoding feature to obtain the first sample acoustic feature (e.g., ...) corresponding to the first sample phrase. Figure 3 Y1 in the first acoustic model). The first phoneme prosodic coding features corresponding to the first sample short phrase output by the encoder in the first acoustic model and the first sample acoustic features corresponding to the first sample short phrase output by the decoder in the first acoustic model are used as the first sample prediction information corresponding to the first sample short phrase.
[0090] This embodiment also pre-trains a second acoustic model, which is a pre-trained model capable of predicting the acoustic features of sentences. This second acoustic model is a relatively complex and deep model with good performance, such as a deep neural network. The trained second acoustic model thus possesses good acoustic feature prediction capabilities. This embodiment can utilize the second acoustic model to perform model distillation on the first acoustic model, thereby improving the performance of the first acoustic model while simplifying it and reducing its dependence on computational and memory resources. This allows the first acoustic model to have a faster response speed even when set up in a device-based speech synthesis system, improving the response speed of speech synthesis. Therefore, the first acoustic model distilled using the second acoustic model has similar acoustic model prediction performance to the second acoustic model, and also has a simplified model structure.
[0091] Therefore, in this embodiment, the phoneme prosodic features corresponding to each sample phrase in the sample phrase set are input into a pre-trained second acoustic model. This second acoustic model can output second sample prediction information corresponding to the first sample phrase. The second sample prediction information includes the second phoneme prosodic encoding features and the second sample acoustic features (such as...) corresponding to the first sample phrase. Figure 3 (Y2 in the second acoustic model). Specifically, the encoder in the second acoustic model encodes the phoneme prosodic features corresponding to each sample phrase in the sample phrase set, obtaining the second phoneme prosodic encoded features corresponding to the sample phrase set. These second phoneme prosodic encoded features are then input into the decoder in the second acoustic model. The decoder in the second acoustic model decodes the second phoneme prosodic encoded features to obtain the second acoustic features corresponding to the sample phrase set. The second phoneme prosodic encoded features output by the encoder in the second acoustic model and the second sample acoustic features output by the decoder in the second acoustic model are used as the second sample prediction information.
[0092] This embodiment trains the model parameters of the first acoustic model based on the prediction information of the first sample and the prediction information of the second sample, thereby enabling model distillation of the second acoustic model using the second acoustic model. Specifically, based on the loss function between the first phoneme prosodic coding features and the second phoneme prosodic coding features (such as... Figure 3 L1 in the equation, and the loss function between the acoustic features of the first sample and the acoustic features of the second sample (e.g., L1 in the equation). Figure 3 In L2), the model parameters of the first acoustic model are adjusted. Since both the second phoneme prosodic coding feature and the second sample acoustic feature are features corresponding to the sample phrase set, both the second phoneme prosodic coding feature and the second sample acoustic feature include the features corresponding to each sample phrase in the sample phrase set. The loss function between the first phoneme prosodic coding feature and the second phoneme prosodic coding feature is the loss function calculated between the first phoneme prosodic coding feature corresponding to the first sample phrase and the feature corresponding to the first sample phrase in the second phoneme prosodic coding feature. The loss function between the first sample acoustic feature and the second sample acoustic feature is the loss function calculated between the first sample acoustic feature corresponding to the first sample phrase and the feature corresponding to the first sample phrase in the second sample acoustic feature.
[0093] Furthermore, after the encoder in the first acoustic model outputs the first phoneme prosodic coding feature, it can also predict the duration of each phoneme based on the first phoneme prosodic coding feature as the first duration information. Similarly, after the encoder in the second acoustic model outputs the second phoneme prosodic coding feature, it can also predict the duration of each phoneme based on the second phoneme prosodic coding feature as the second duration information. In this embodiment, the first sample prediction information may include the first duration information, and the second sample prediction information may include the second duration information. This embodiment can also adjust the model parameters of the first acoustic model based on the loss function between the first duration information and the second duration information.
[0094] As an optional implementation, see [link to implementation details]. Figure 4 As shown, in another embodiment of this application, the training process of the second acoustic model is disclosed, including:
[0095] S401. Obtain first sample recording data from at least two speakers.
[0096] This embodiment pre-collects first sample recording data from at least two speakers. The first sample recording data includes a first sample recording text and a corresponding first sample audio. Since the first sample recording data from a single speaker may be limited, the second acoustic model trained using only the first sample recording data from one speaker may lack sufficient acoustic feature prediction capability. Therefore, this embodiment can collect first sample recording data from multiple speakers to ensure a sufficient number of samples for training the second acoustic model, thereby improving the acoustic feature prediction capability of the second acoustic model.
[0097] S402. Input the phoneme prosodic features corresponding to the first sample recording text into the second acoustic model to obtain the first predicted acoustic features corresponding to the first sample recording text.
[0098] This embodiment needs to determine the phonemic prosodic features corresponding to the first sample audio text. The method used is the same as that used in the above embodiment to determine the phonemic prosodic features corresponding to each short sentence in the set of short sentences to be converted. This embodiment will not repeat the details.
[0099] The phoneme prosodic features corresponding to the first sample recorded text are input into the second acoustic model, and the second acoustic model predicts the first predicted acoustic features corresponding to the first sample recorded text.
[0100] S403. Based on the loss function between the first predicted acoustic features and the real acoustic features corresponding to the first sample audio, adjust the model parameters of the second acoustic model.
[0101] In this embodiment, acoustic features are extracted from the first sample audio corresponding to the first sample recorded text to determine the true acoustic features corresponding to the first sample audio. Then, a loss function is calculated between the first predicted acoustic features corresponding to the first sample recorded text and the true acoustic features corresponding to the first sample audio. The model parameters of the second acoustic model are adjusted with the objective of maximizing the similarity between the first predicted acoustic features corresponding to the first sample recorded text and the true acoustic features corresponding to the first sample audio. The second acoustic model is iteratively trained using all pre-acquired first sample recorded data to obtain the trained second acoustic model.
[0102] As an optional implementation, see [link to implementation details]. Figure 5 As shown in another embodiment of this application, the training process of the second acoustic model is disclosed, which further includes:
[0103] S501. Obtain the second sample recording data of the target speaker.
[0104] This embodiment pre-collects second sample recording data of the target speaker, wherein the second sample recording data includes a second sample recording text and a second sample audio corresponding to the second sample recording text. The first sample recording data pre-collected in the above embodiment may also include the second sample recording data in this embodiment.
[0105] S502. Input the phoneme prosodic features corresponding to the second sample recording text into the second acoustic model to obtain the second predicted acoustic features corresponding to the second sample recording text.
[0106] This embodiment needs to determine the phonemic prosodic features corresponding to the second sample audio text. The method used is the same as that used in the above embodiment to determine the phonemic prosodic features corresponding to each short sentence in the set of short sentences to be converted. This embodiment will not repeat the details.
[0107] The phoneme prosodic features corresponding to the second sample recording text are input into the second acoustic model, and the second acoustic model predicts the second predicted acoustic features corresponding to the second sample recording text.
[0108] S503. Based on the loss function between the second predicted acoustic features and the real acoustic features corresponding to the second sample audio, fine-tune the model parameters of the second acoustic model.
[0109] In this embodiment, acoustic features are extracted from the second sample audio corresponding to the second sample recorded text to determine the true acoustic features corresponding to the second sample audio. Then, the loss function between the second predicted acoustic features corresponding to the second sample recorded text and the true acoustic features corresponding to the second sample audio is calculated. The model parameters of the second acoustic model are fine-tuned with the goal of maximizing the similarity between the second predicted acoustic features corresponding to the second sample recorded text and the true acoustic features corresponding to the second sample audio.
[0110] To ensure that the speech synthesized using the acoustic features predicted by the second acoustic model matches the speech habits of the target speaker, this embodiment requires fine-tuning the already trained second acoustic model using the second sample recording data corresponding to the target speaker. Since the sample recording data for a single speaker is limited, directly using the second sample recording data of the target speaker to train the untrained second acoustic model may result in insufficient acoustic feature prediction capability due to insufficient training samples. Therefore, this solution, following the previous embodiment, first initializes the second acoustic model using the first sample recording data of multiple speakers to improve its acoustic feature prediction capability. Then, in this embodiment, it fine-tunes the already initialized second acoustic model using the second sample recording data of the target speaker. This ensures that the second acoustic model possesses both high acoustic feature prediction capability and that the speech synthesized using the acoustic features predicted by the second acoustic model matches the speech habits of the target speaker.
[0111] Exemplary device
[0112] Accordingly, embodiments of this application also provide an acoustic feature determination device, see [link to relevant documentation]. Figure 6 As shown, the device includes:
[0113] The feature acquisition module 100 is used to acquire the phoneme prosodic features of each short sentence in the short sentence set corresponding to the sentence to be converted; wherein, the short sentence set includes at least two short sentences obtained after the sentence to be converted is segmented.
[0114] The context feature extraction module 110 is used to perform context feature extraction processing on the phoneme prosodic features corresponding to each short sentence in the short sentence set, and to determine the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence from the extracted features; wherein, the first short sentence is any short sentence in the short sentence set;
[0115] The acoustic feature prediction module 120 is used to predict the acoustic features corresponding to the first short sentence based on the phoneme prosodic features, historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence; the acoustic features are used to synthesize the speech corresponding to the first short sentence.
[0116] As can be seen from the above description, the acoustic feature determination device proposed in this application predicts the acoustic features of each short sentence segmented into the sentence to be converted in sequence. Compared with the whole sentence prediction of the acoustic features of the sentence to be converted, it can realize the streaming prediction of acoustic features, reduce the response time of determining acoustic features, and thus improve the response speed of speech synthesis.
[0117] As an optional implementation, another embodiment of this application discloses a context feature extraction module 110, specifically used for:
[0118] Contextual feature extraction is performed on the phoneme prosodic features corresponding to each short sentence in the short sentence set to obtain the contextual phoneme prosodic features corresponding to the short sentence set.
[0119] Based on the features corresponding to historical short sentences in the contextual phoneme prosodic features, the historical phoneme prosodic features corresponding to the first short sentence are determined; the historical short sentence is the short sentence preceding the first short sentence in the short sentence set.
[0120] Based on the features corresponding to future sentences in the context phoneme prosodic features, the future phoneme prosodic features corresponding to the first short sentence are determined; the future short sentence is the short sentence following the first short sentence in the short sentence set.
[0121] As an optional implementation, another embodiment of this application discloses a context feature extraction module 110 and an acoustic feature prediction module 120, specifically used for:
[0122] The phonemic prosodic features corresponding to each short phrase in the short phrase set and the phonemic prosodic features corresponding to the first short phrase are all input into the pre-trained first acoustic model to obtain the acoustic features corresponding to the first short phrase.
[0123] The first acoustic model extracts contextual features from the phonemic prosodic features of each short sentence in the short sentence set, and determines the historical and future phonemic prosodic features of the first short sentence from the extracted features. Based on the phonemic prosodic features, historical and future phonemic prosodic features of the first short sentence, the acoustic features of the first short sentence are predicted.
[0124] As an optional implementation, another embodiment of this application discloses an acoustic feature determination device, which further includes: a sample acquisition module, a model processing module, and a first model adjustment module.
[0125] The sample acquisition module is used to pre-acquire a set of sample short sentences corresponding to the sample conversion statement, and to determine the phoneme prosodic features of each sample short sentence in the sample short sentence set; the sample short sentence set includes at least two sample short sentences obtained after the sample conversion statement is segmented.
[0126] The model processing module is used to input the phoneme prosodic features corresponding to each sample phrase in the sample phrase set and the phoneme prosodic features corresponding to the first sample phrase into the first acoustic model to obtain the first sample prediction information corresponding to the first sample phrase; and to input the phoneme prosodic features corresponding to each sample phrase in the sample phrase set into the pre-trained second acoustic model to obtain the second sample prediction information corresponding to the sample phrase set; the first sample phrase is any sample phrase in the sample phrase set;
[0127] The first model adjustment module is used to adjust the model parameters of the first acoustic model based on the prediction information of the first sample and the prediction information of the second sample.
[0128] As an optional implementation, another embodiment of this application discloses that the first sample prediction information includes: the first phoneme prosodic coding feature and the first sample acoustic feature corresponding to the first sample phrase; the second sample prediction information includes: the second phoneme prosodic coding feature and the second sample acoustic feature corresponding to the sample phrase set.
[0129] Specifically, the first phoneme prosodic coding feature is obtained by the encoder in the first acoustic model based on the phoneme prosodic features corresponding to the first sample phrase, the sample historical phoneme prosodic features, and the sample future phoneme prosodic features; the first sample acoustic feature is obtained by the decoder in the first acoustic model by decoding the first phoneme prosodic coding feature; the second phoneme prosodic coding feature is obtained by the encoder in the second acoustic model based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set; and the second sample acoustic feature is obtained by the decoder in the second acoustic model by decoding the second phoneme prosodic coding feature.
[0130] As an optional implementation, another embodiment of this application discloses a first model adjustment module, specifically used for:
[0131] The model parameters of the first acoustic model are adjusted based on the loss function between the first phoneme prosodic coding features and the second phoneme prosodic coding features, as well as the loss function between the first sample acoustic features and the second sample acoustic features.
[0132] As an optional implementation, another embodiment of this application discloses an acoustic feature determination device, which further includes a second model adjustment module.
[0133] The sample acquisition module is also used to acquire first sample recording data from at least two speakers. The first sample recording data includes: first sample recording text and first sample audio.
[0134] The model processing module is also used to input the phoneme prosodic features corresponding to the first sample recorded text into the second acoustic model to obtain the first predicted acoustic features corresponding to the first sample recorded text.
[0135] The second model adjustment module is used to adjust the model parameters of the second acoustic model based on the loss function between the first predicted acoustic features and the real acoustic features corresponding to the first sample audio.
[0136] As an optional implementation, another embodiment of this application discloses that the sample acquisition module is further used to acquire second sample recording data of the target speaker, the second sample recording data including: second sample recording text and second sample audio;
[0137] The model processing module is also used to input the phoneme prosodic features corresponding to the second sample recording text into the second acoustic model to obtain the second predicted acoustic features corresponding to the second sample recording text.
[0138] The second model adjustment module is also used to fine-tune the model parameters of the second acoustic model based on the loss function between the second predicted acoustic features and the real acoustic features corresponding to the second sample audio.
[0139] The acoustic feature determination device provided in this embodiment belongs to the same concept as the acoustic feature determination method provided in the above embodiments of this application. It can execute the acoustic feature determination method provided in any of the above embodiments of this application and has the corresponding functional modules and beneficial effects for executing the acoustic feature determination method. Technical details not described in detail in this embodiment can be found in the specific processing content of the acoustic feature determination method provided in the above embodiments of this application, and will not be repeated here.
[0140] Exemplary System
[0141] Optionally, embodiments of this application also provide a speech synthesis system, see [link to relevant documentation]. Figure 7 As shown, the speech synthesis system includes a front-end device 200, an acoustic feature determination device 210, and a vocoder 220. Data transmission can be achieved between the front-end device 200 and the acoustic feature determination device 210, and data transmission can be achieved between the acoustic feature determination device 210 and the vocoder 220.
[0142] The front-end device 210 is used to determine the phonemic prosodic features of each short sentence in the short sentence set corresponding to the sentence to be converted; the short sentence set includes at least two short sentences obtained by segmenting the sentence to be converted.
[0143] The acoustic feature determination device 220 is used to determine the acoustic features corresponding to the short sentences in the short sentence set by executing the acoustic feature determination method of the above embodiment;
[0144] The vocoder 230 is used to determine the speech of the short sentences corresponding to the short sentences in the short sentence set based on the acoustic features corresponding to the short sentences in the short sentence set. Specifically, after the acoustic feature determining device 220 determines the acoustic features corresponding to each short sentence, it directly transmits the acoustic features of that short sentence to the vocoder 230. The vocoder 230 directly synthesizes the speech of the short sentence corresponding to that short sentence, thereby enabling the sequential output of the speech of each short sentence. This eliminates the need to synthesize the speech of the entire sentence to be converted before outputting the whole sentence, achieving streaming prediction of acoustic features and streaming synthesis of speech, reducing the response time for determining acoustic features, and improving the response speed of speech synthesis.
[0145] Exemplary electronic devices
[0146] Another embodiment of this application also provides an electronic device, see [link to relevant documentation] Figure 8 As shown, the device includes:
[0147] Memory 300 and processor 310;
[0148] The memory 300 is connected to the processor 310 and is used to store programs;
[0149] The processor 310 is configured to implement the acoustic feature determination method disclosed in any of the above embodiments by running the program stored in the memory 300.
[0150] Specifically, the aforementioned electronic device may also include: a bus, a communication interface 320, an input device 330, and an output device 340.
[0151] The processor 310, memory 300, communication interface 320, input device 330, and output device 340 are interconnected via a bus. Among them:
[0152] A bus can include a pathway for transmitting information between various components of a computer system.
[0153] The processor 310 can be a general-purpose processor, such as a general-purpose central processing unit (CPU), a microprocessor, etc., or an application-specific integrated circuit (ASIC), or one or more integrated circuits used to control the execution of the program of the present invention. It can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), an off-the-shelf programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.
[0154] Processor 310 may include a main processor, as well as a baseband chip, modem, etc.
[0155] The memory 300 stores a program that executes the technical solution of this invention, and may also store an operating system and other key business functions. Specifically, the program may include program code, which includes computer operation instructions. More specifically, the memory 300 may include read-only memory (ROM), other types of static storage devices capable of storing static information and instructions, random access memory (RAM), other types of dynamic storage devices capable of storing information and instructions, disk storage, flash memory, etc.
[0156] Input device 330 may include a device for receiving user input data and information, such as a keyboard, mouse, camera, scanner, light pen, voice input device, touch screen, pedometer, or gravity sensor.
[0157] Output device 340 may include devices that allow information to be output to a user, such as a display screen, printer, speaker, etc.
[0158] The communication interface 320 may include a device that uses any transceiver to communicate with other devices or communication networks, such as Ethernet, Radio Access Network (RAN), Wireless Local Area Network (WLAN), etc.
[0159] The processor 310 executes the program stored in the memory 300 and calls other devices, which can be used to implement the various steps of any of the acoustic feature determination methods provided in the above embodiments of this application.
[0160] Exemplary computer program products and storage media
[0161] In addition to the methods and apparatus described above, embodiments of this application may also be computer program products, which include computer program instructions that, when executed by a processor, cause the processor to perform the steps in the acoustic feature determination methods according to various embodiments of this application described in the "Exemplary Methods" section of this specification.
[0162] The computer program product can be written in any combination of one or more programming languages to perform the operations of the embodiments of this application. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on the user's computing device, partially on the user's computing device, as a standalone software package, partially on the user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0163] Furthermore, embodiments of this application may also be storage media storing a computer program, which is executed by a processor in the steps of the acoustic feature determination method according to various embodiments of this application described in the "Exemplary Methods" section above.
[0164] For the foregoing method embodiments, in order to simplify the description, they are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, because according to this application, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to this application.
[0165] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For apparatus embodiments, since they are basically similar to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0166] The steps in the methods of the various embodiments of this application can be adjusted, merged, or deleted in order according to actual needs, and the technical features described in each embodiment can be replaced or combined.
[0167] The modules and sub-modules in the various embodiments of the present application's devices and terminals can be merged, divided, and deleted according to actual needs.
[0168] It should be understood that the disclosed terminals, devices, and methods can be implemented in other ways, given the several embodiments provided in this application. For example, the terminal embodiments described above are merely illustrative. For instance, the division of modules or sub-modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple sub-modules or modules may be combined or integrated into another module, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or modules, and may be electrical, mechanical, or other forms.
[0169] The modules or submodules described as separate components may or may not be physically separate. The components that constitute a module or submodule may or may not be physical modules or submodules; that is, they may be located in one place or distributed across multiple network modules or submodules. Some or all of the modules or submodules can be selected to achieve the purpose of this embodiment's solution, depending on actual needs.
[0170] Furthermore, the functional modules or sub-modules in the various embodiments of this application can be integrated into one processing module, or each module or sub-module can exist physically separately, or two or more modules or sub-modules can be integrated into one module. The integrated modules or sub-modules described above can be implemented in hardware or in the form of software functional modules or sub-modules.
[0171] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0172] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software unit executed by a processor, or a combination of both. The software unit can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0173] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0174] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for determining acoustic features, characterized in that, include: Obtain the phonemic prosodic features of each phrase in the set of phrases corresponding to the statement to be converted; wherein, the set of phrases includes at least two phrases obtained after segmenting the statement to be converted; The phonemic prosodic features corresponding to each short sentence in the short sentence set and the phonemic prosodic features corresponding to the first short sentence are all input into the pre-trained first acoustic model to obtain the acoustic features corresponding to the first short sentence; the first short sentence is any short sentence in the short sentence set. The first acoustic model performs context feature extraction processing on the phoneme prosodic features corresponding to each short sentence in the short sentence set, and determines the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence from the extracted features. Based on the phoneme prosodic features, historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence, the acoustic features corresponding to the first short sentence are predicted. The training of the first acoustic model is based on adjusting the model parameters of the first acoustic model by using the first sample prediction information corresponding to the first sample phrase output by the first acoustic model and the second sample prediction information corresponding to the sample phrase set output by the pre-trained second acoustic model. The first sample prediction information is obtained by the first acoustic model through acoustic prediction based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set and the phoneme prosodic features corresponding to the first sample phrase. The second sample prediction information is obtained by the second acoustic model through acoustic prediction based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set. The sample phrase set includes at least two sample phrases obtained by segmenting the pre-collected sample converted sentences. The first sample phrase is any one of the sample phrases in the sample phrase set.
2. The method according to claim 1, characterized in that, The phoneme prosodic features corresponding to each short phrase in the phrase set are subjected to context feature extraction processing, and the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short phrase are determined from the extracted features, including: Contextual feature extraction is performed on the phoneme prosodic features corresponding to each short sentence in the short sentence set to obtain the contextual phoneme prosodic features corresponding to the short sentence set. Based on the features corresponding to historical short sentences in the contextual phoneme prosodic features, the historical phoneme prosodic features corresponding to the first short sentence are determined; the historical short sentence is the short sentence preceding the first short sentence in the short sentence set. Based on the features corresponding to future short sentences in the context phoneme prosodic features, the future phoneme prosodic features corresponding to the first short sentence are determined; the future short sentence is the short sentence following the first short sentence in the short sentence set.
3. The method according to claim 1, characterized in that, The training process of the first acoustic model includes: A set of sample phrases corresponding to the sample conversion statement is pre-collected, and the phoneme prosodic features corresponding to each sample phrase in the sample phrase set are determined; the sample phrase set includes at least two sample phrases obtained after sentence segmentation of the sample conversion statement; The phonemic prosodic features corresponding to each sample phrase in the sample phrase set and the phonemic prosodic features corresponding to the first sample phrase are input into the first acoustic model to obtain the first sample prediction information corresponding to the first sample phrase. Additionally, the phonemic prosodic features corresponding to each sample phrase in the sample phrase set are input into a pre-trained second acoustic model to obtain the second sample prediction information corresponding to the sample phrase set. The first sample phrase is any one of the sample phrases in the sample phrase set. Based on the first sample prediction information and the second sample prediction information, the model parameters of the first acoustic model are adjusted.
4. The method according to claim 3, characterized in that, The first sample prediction information includes: the first phoneme prosodic coding feature and the first sample acoustic feature corresponding to the first sample phrase; the second sample prediction information includes: the second phoneme prosodic coding feature and the second sample acoustic feature corresponding to the sample phrase set; Wherein, the first phoneme prosodic coding feature is obtained by the encoder in the first acoustic model based on the phoneme prosodic features corresponding to the first sample phrase, the sample historical phoneme prosodic features, and the sample future phoneme prosodic features; the first sample acoustic feature is obtained by the decoder in the first acoustic model by decoding the first phoneme prosodic coding feature; the second phoneme prosodic coding feature is obtained by the encoder in the second acoustic model based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set; the second sample acoustic feature is obtained by the decoder in the second acoustic model by decoding the second phoneme prosodic coding feature.
5. The method according to claim 4, characterized in that, Based on the first sample prediction information and the second sample prediction information, the model parameters of the first acoustic model are adjusted, including: Based on the loss function between the first phoneme prosodic coding feature and the second phoneme prosodic coding feature, and the loss function between the first sample acoustic feature and the second sample acoustic feature, the model parameters of the first acoustic model are adjusted.
6. The method according to claim 3, characterized in that, The training process of the second acoustic model includes: Acquire first sample recording data from at least two speakers, the first sample recording data including: first sample recording text and first sample audio; The phoneme prosodic features corresponding to the first sample recording text are input into the second acoustic model to obtain the first predicted acoustic features corresponding to the first sample recording text. The model parameters of the second acoustic model are adjusted based on the loss function between the first predicted acoustic features and the real acoustic features corresponding to the first sample audio.
7. The method according to claim 6, characterized in that, The training process of the second acoustic model also includes: Acquire second sample recording data of the target speaker, the second sample recording data including: second sample recording text and second sample audio; The phoneme prosodic features corresponding to the second sample recording text are input into the second acoustic model to obtain the second predicted acoustic features corresponding to the second sample recording text. Based on the loss function between the second predicted acoustic features and the real acoustic features corresponding to the second sample audio, the model parameters of the second acoustic model are fine-tuned.
8. An acoustic feature determination device, characterized in that, include: The feature acquisition module is used to acquire the phoneme prosodic features of each short sentence in the short sentence set corresponding to the sentence to be converted; wherein, the short sentence set includes at least two short sentences obtained after sentence segmentation of the sentence to be converted; The acoustic feature prediction module is used to input the phonemic prosodic features corresponding to each short sentence in the short sentence set and the phonemic prosodic features corresponding to the first short sentence into a pre-trained first acoustic model to obtain the acoustic features corresponding to the first short sentence; the first short sentence is any short sentence in the short sentence set. The first acoustic model performs context feature extraction processing on the phoneme prosodic features corresponding to each short sentence in the short sentence set, and determines the historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence from the extracted features. Based on the phoneme prosodic features, historical phoneme prosodic features and future phoneme prosodic features corresponding to the first short sentence, the acoustic features corresponding to the first short sentence are predicted. The training of the first acoustic model is based on adjusting the model parameters of the first acoustic model by using the first sample prediction information corresponding to the first sample phrase output by the first acoustic model and the second sample prediction information corresponding to the sample phrase set output by the pre-trained second acoustic model. The first sample prediction information is obtained by the first acoustic model through acoustic prediction based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set and the phoneme prosodic features corresponding to the first sample phrase. The second sample prediction information is obtained by the second acoustic model through acoustic prediction based on the phoneme prosodic features corresponding to each sample phrase in the sample phrase set. The sample phrase set includes at least two sample phrases obtained by segmenting the pre-collected sample converted sentences. The first sample phrase is any one of the sample phrases in the sample phrase set.
9. An electronic device, characterized in that, include: Memory and processor; The memory is connected to the processor and is used to store programs; The processor is configured to implement the acoustic feature determination method as described in any one of claims 1 to 7 by running a program in the memory.
10. A storage medium, characterized in that, The storage medium stores a computer program, which, when executed by a processor, implements the acoustic feature determination method as described in any one of claims 1 to 7.
11. A speech synthesis system, characterized in that, include: Front-end equipment, acoustic feature determination equipment, and vocoder; The front-end device is used to determine the phonemic prosodic features of each short sentence in the set of short sentences corresponding to the sentence to be converted. The set of short sentences includes at least two short sentences obtained by segmenting the sentence to be converted. The acoustic feature determination device is used to determine the acoustic features corresponding to the short sentences in the short sentence set by executing the acoustic feature determination method according to any one of claims 1 to 7; The vocoder is used to determine the speech of the short sentences corresponding to the short sentences in the short sentence set based on the acoustic features corresponding to the short sentences in the short sentence set.
Citation Information
Patent Citations
Speech synthesis method and device and electronic equipment
CN114267321A