Rhythmical pause information determination method and device

A determination method and rhythm technology, applied in the determination of prosody pause information, can solve problems such as inconsistency in prosody training data, and achieve the effects of improving synthesis fluency, improving prosody rhythm, and natural synthesis effect.

CN105225658AActive Publication Date: 2016-01-06BAIDU ONLINE NETWORK TECH (BEIJIBG) CO LTD
5 Cites 18 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Publication Date
2016-01-06

Smart Images

  • Figure 1
    Figure 1
  • Figure 2
    Figure 2
  • Figure 3
    Figure 3
Patent Text Reader

Abstract

The invention provides a rhythmical pause information determination method and a device. The rhythmical pause information determination method comprises the steps of extracting rhythmical prediction features of a to-be-synthesized text; selecting a self-adaptive rhythmical prediction model corresponding to a selected speaker; and inputting the rhythmical prediction features of the to-be-synthesized text into the self-adaptive rhythmical prediction model corresponding to the selected speaker so as to determine the rhythmical pause information of the to-be-synthesized text. According to the technical scheme of the invention, the problem that the rhythmical training data for an acoustic model and the rhythmical training data for a rhythmical model are inconsistent can be solved. Meanwhile, the rhythmical rhyme and the synthetic fluency are improved. Moreover, the self-adaptive rhythmical prediction model corresponding to the selected speaker is adopted, so that the synthetic effect in a multi-speaker handover situation is more natural.
Need to check novelty before this filing date? Find Prior Art

Description

technical field

[0001] The invention relates to the technical field of speech synthesis, in particular to a method and device for determining prosodic pause information. Background technique

[0002] The purpose of speech synthesis is to convert text into speech and play it to users, and the goal is to achieve the effect of live text broadcasting. An important module in the speech synthesis link is to predict the prosodic pauses of the text to be synthesized, and then generate synthetic speech according to the predicted prosodic pauses.

[0003] Currently, prosody prediction in speech synthesis is implemented based on statistical machine learning, and the process includes preparing training data, training prosody prediction models, and performing prosody predictions based on the trained models.

[0004] However, in the prior art, the prosodic pause pattern trained in the prosody prediction model does not match the prosodic pause pattern trained in the acoustic model. The r...

Examples

Embodiment Construction

[0022] Embodiments of the present invention are described in detail below, examples of which are shown in the drawings, wherein the same or similar reference numerals designate the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the figures are exemplary only for explaining the present invention and should not be construed as limiting the present invention. On the contrary, the embodiments of the present invention include all changes, modifications and equivalents coming within the spirit and scope of the appended claims.

[0023] figure 1 It is a flowchart of an embodiment of the method for determining rhythmic pause information in the present invention, such as figure 1 As shown, the method for determining the prosodic pause information may include:

[0024] Step 101, extracting the prosody prediction features of the text to be synthesized.

[0025] Specifically, extracting the prosodic ...