A method, apparatus and application for prosodic annotation

CN115831092BActive Publication Date: 2026-08-14DINGFU NEW POWER (BEIJING) INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-03
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

”但是,在说话人根据标注的文本进行录制的时候,往往存在两个问题:1、很难完全按照根据语义划分好的韵律标注进行朗读;2、在进行韵律朗读的时候,把握不好各个韵律停顿的时长,例如#1和#2的短停顿

Benefits of technology

[0056]In the above embodiments of this application, since the first text is prosodic-annotated text and includes multiple prosodic tags, the pause duration of each prosodic tag in the first speech data recorded by a specific speaker based on the prosodic tags of the first text depends on the specific speaker's speaking habits. Next, the duration of each prosodic tag is statistically analyzed to obtain statistical data on the duration of each prosodic tag. Since the statistical data on the duration of each prosodic tag is related to the specific speaker's speaking habits, the range of values ​​for the duration of each prosodic tag determined based on the statistical data also matches the specific speaker's speaking habits. Then, the specific speaker records second speech data based on a second text that has not been prosodic-annotated. The duration of each pause in the second speech data matches the specific speaker's speaking habits. Furthermore, since the range of values ​​for the duration of each prosodic tag matches the specific speaker's speaking habits, the prosody in the prosodic-annotated second text can accurately match the prosody of the second speech data based on the range of values ​​for the duration of each prosodic tag and the duration of each pause. The speech synthesis model is trained using prosodic-matched speech data and labeled second text, enabling the model to accurately learn prosodic features. This allows the synthesized speech to clearly reflect the pauses represented by each prosodic label.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115831092B_ABST
    Figure CN115831092B_ABST
Patent Text Reader

Abstract

This application provides a prosodic annotation method, apparatus, and application that enables precise matching of the prosodicity of recorded speech audio with the prosodicity of annotated text. The method includes: acquiring first speech data recorded by a specific speaker based on a first text annotated with prosodic markers, the first text including multiple prosodic tags, each representing a different duration of pause; calculating the duration of each prosodic tag among the multiple prosodic tags based on the first speech data and the first text to obtain statistical data on the duration of each prosodic tag; determining the value range of the duration of each prosodic tag based on the statistical data corresponding to each prosodic tag; acquiring second speech data recorded by a specific speaker based on a second text that has not been annotated with prosodic markers; obtaining the duration of each pause in the second speech data based on the second speech data; and annotating the second text with prosodic markers based on the value range of the duration of each prosodic tag and the duration of each pause.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of natural language processing technology, and in particular to a prosodic annotation method, apparatus and application. Background Technology

[0002] In speech synthesis technology, prosody is represented by pauses in synthesized speech. To make intelligent voice interaction more human-like, current text-to-speech (TTS) neural network models typically need to learn the prosodic features of audio to make the synthesized speech more natural and fluent. Based on the length and position of pauses, prosodic labels #1, #2, #3, and #4 are used to represent different pauses. #1 represents the boundary of a prosodic word, indicating a short pause; #2 represents the boundary of a prosodic phrase, indicating a prolonged sound or a short pause; #3 represents a semantically complete, more obvious pause and a drop in intonation; and #4 represents the end of a sentence, a label indicating the end of the sentence corresponding to each number.

[0003] Currently, the mainstream prosodic annotation scheme is based on semantics. Taking sentence 71 of the BiaoBei dataset (a public dataset) as an example, "Erjing lives outside the North Fifth Ring Road and goes to the Huatang Shopping Mall in Asian Games Village for work." According to semantics, the annotated text is "Erjing lives #2 outside the North Fifth Ring Road #3 and goes to #1 Asian Games Village #2 Huatang Shopping Mall #4 for work." However, when speakers record based on the annotated text, two problems often arise: 1. It is difficult to read aloud exactly according to the semantically defined prosodic annotations; 2. When reading aloud prosody, it is difficult to grasp the duration of each prosodic pause, such as the short pauses between #1 and #2. Therefore, a mismatch occurs between the prosodic of the recorded audio and the prosodic of the annotated text. Using this prosodic mismatched audio and annotated text to train a speech synthesis model prevents the model from learning prosodic features, resulting in synthesized speech that does not clearly reflect the pauses represented by the prosodic labels.

[0004] Therefore, how to obtain and record prosodic text that matches the prosodic rhythm of speech audio has become an urgent problem to be solved. Summary of the Invention

[0005] This application provides a prosody annotation method, apparatus, and application that enables precise matching of the prosody of recorded audio and the prosody of annotated text.

[0006] Firstly, this application provides a prosodic annotation method, including:

[0007] Acquire first speech data recorded by a specific speaker based on a first text. The first text is a text with prosodic annotation. The first text includes multiple prosodic tags, and different prosodic tags represent different durations of pauses.

[0008] Based on the first speech data and the first text, the duration of each prosodic tag in multiple prosodic tags is calculated to obtain statistical data on the duration of each prosodic tag.

[0009] Based on the statistical data corresponding to each prosodic tag, determine the range of values ​​for the duration of each prosodic tag;

[0010] Acquire second speech data recorded by a specific speaker based on a second text, wherein the second text is unprosodic text;

[0011] The duration of each pause in the second speech data is obtained based on the second speech data, and the second text is prosodicly annotated according to the range of the duration of each prosodic tag and the duration of each pause.

[0012] In one example, the method further includes: if the prosodic tags in the first text do not match the pauses in the first speech data, redetermining the prosodic tags in the first text based on the pauses in the first speech data.

[0013] In one example, based on the statistical data corresponding to each prosodic tag, the range of values ​​for the duration of each prosodic tag is determined, including:

[0014] Based on the statistical data corresponding to each prosodic label, determine the mean and standard deviation of the duration of each prosodic label;

[0015] Based on the normal distribution analysis method, the range of values ​​for the duration of each prosodic tag is determined according to the mean and standard deviation of the duration of each prosodic tag.

[0016] In one example, the range of duration values ​​for each prosodic tag is determined based on the mean and standard deviation of the duration for each prosodic tag, including:

[0017] Based on the mean and standard deviation of the duration of each prosodic tag, determine the difference between the mean and standard deviation of the duration of each prosodic tag, and the sum of the mean and standard deviation of the duration of each prosodic tag.

[0018] Based on the difference and sum values ​​corresponding to each prosodic tag, determine the range of values ​​for the duration of each prosodic tag.

[0019] In one example, the duration of each pause in the second speech data is obtained based on the second speech data, and prosodic annotation is performed on the second text according to the value range of the duration of each prosodic tag and the duration of each pause, including:

[0020] The second text is processed into pinyin to obtain pinyin text;

[0021] Based on the Montreal Forced Alignment (MFA) algorithm, the start and end times of each pause are obtained according to the second speech data and the pinyin text. Each pause indicates that there is an interval between the end time of the preceding phoneme and the start time of the following phoneme.

[0022] The duration of each pause is determined based on the start and end times of each pause;

[0023] Based on the duration of each pause and the range of values ​​for the duration of each prosodic tag, determine the prosodic tag corresponding to each pause.

[0024] The second text is prosodicly annotated based on the prosodic tag corresponding to each pause.

[0025] In one example, based on the statistical data corresponding to each prosodic tag, the range of values ​​for the duration of each prosodic tag is determined, including:

[0026] Based on the statistical data corresponding to each prosodic label, determine the maximum and minimum values ​​of the statistical data corresponding to each prosodic label;

[0027] The range of duration values ​​for each prosodic tag is determined based on the maximum and minimum values ​​corresponding to each prosodic tag.

[0028] Secondly, this application provides a prosody annotation device, comprising:

[0029] The speech data acquisition module is used to acquire the first speech data recorded by a specific speaker based on the first text. The first text is a text that has been punctuated with prosody. The first text includes multiple prosody tags, and different prosody tags represent different durations of pauses.

[0030] The prosodic tag duration determination module is used to calculate the duration of each prosodic tag among multiple prosodic tags based on the first speech data and the first text, so as to obtain statistical data on the duration of each prosodic tag.

[0031] The prosodic tag duration determination module is also used to determine the range of values ​​for the duration of each prosodic tag based on the statistical data corresponding to each prosodic tag.

[0032] The voice data acquisition module is also used to acquire second voice data recorded by a specific speaker based on a second text, where the second text is text that has not been prosodic.

[0033] The prosodic annotation module is used to obtain the duration of each pause in the second speech data based on the second speech data, and to perform prosodic annotation on the second text based on the value range of the duration of each prosodic tag and the duration of each pause.

[0034] In one example, the device also includes:

[0035] The prosody tag fine-tuning module is used to redetermine the prosody tags in the first text based on the pauses in the first speech data if the prosody tags in the first text do not match the pauses in the first speech data.

[0036] In one example, based on the statistical data corresponding to each prosodic tag, the range of values ​​for the duration of each prosodic tag is determined, including:

[0037] Based on the statistical data corresponding to each prosodic label, determine the mean and standard deviation of the duration of each prosodic label;

[0038] Based on the normal distribution analysis method, the range of values ​​for the duration of each prosodic tag is determined according to the mean and standard deviation of the duration of each prosodic tag.

[0039] In one example, the range of duration values ​​for each prosodic tag is determined based on the mean and standard deviation of the duration for each prosodic tag, including:

[0040] Based on the mean and standard deviation of the duration of each prosodic tag, determine the difference between the mean and standard deviation of the duration of each prosodic tag, and the sum of the mean and standard deviation of the duration of each prosodic tag.

[0041] Based on the difference and sum values ​​corresponding to each prosodic tag, determine the range of values ​​for the duration of each prosodic tag.

[0042] In one example, the duration of each pause in the second speech data is obtained based on the second speech data, and prosodic annotation is performed on the second text according to the value range of the duration of each prosodic tag and the duration of each pause, including:

[0043] The second text is processed into pinyin to obtain pinyin text;

[0044] Based on the Montreal Forced Alignment (MFA) algorithm, the start and end times of each pause are obtained according to the second speech data and the pinyin text. Each pause indicates that there is an interval between the end time of the preceding phoneme and the start time of the following phoneme.

[0045] The duration of each pause is determined based on the start and end times of each pause;

[0046] Based on the duration of each pause and the range of values ​​for the duration of each prosodic tag, determine the prosodic tag corresponding to each pause.

[0047] The second text is prosodicly annotated based on the prosodic tag corresponding to each pause.

[0048] In one example, based on the statistical data corresponding to each prosodic tag, the range of values ​​for the duration of each prosodic tag is determined, including:

[0049] Based on the statistical data corresponding to each prosodic label, determine the maximum and minimum values ​​of the statistical data corresponding to each prosodic label;

[0050] The range of duration values ​​for each prosodic tag is determined based on the maximum and minimum values ​​corresponding to each prosodic tag.

[0051] Thirdly, this application provides an application for obtaining text through prosodic annotation, based on the method described in the above embodiments to obtain second text with prosodic annotation, the method comprising:

[0052] Based on each text in the second document and the first association relationship, determine the corresponding sequence number for each text. The first association relationship is used to associate texts and sequence numbers.

[0053] Obtain the Mel-frequency cepstral coefficient features of the second speech data;

[0054] The training dataset is determined based on the duration of each prosodic tag in the prosodic-annotated second text, the sequence number of each text, and the Mel frequency cepstral coefficient features of the second speech data.

[0055] Train a speech synthesis model based on the training dataset to obtain the trained speech synthesis model.

[0056] In the above embodiments of this application, since the first text is prosodic-annotated text and includes multiple prosodic tags, the pause duration of each prosodic tag in the first speech data recorded by a specific speaker based on the prosodic tags of the first text depends on the specific speaker's speaking habits. Next, the duration of each prosodic tag is statistically analyzed to obtain statistical data on the duration of each prosodic tag. Since the statistical data on the duration of each prosodic tag is related to the specific speaker's speaking habits, the range of values ​​for the duration of each prosodic tag determined based on the statistical data also matches the specific speaker's speaking habits. Then, the specific speaker records second speech data based on a second text that has not been prosodic-annotated. The duration of each pause in the second speech data matches the specific speaker's speaking habits. Furthermore, since the range of values ​​for the duration of each prosodic tag matches the specific speaker's speaking habits, the prosody in the prosodic-annotated second text can accurately match the prosody of the second speech data based on the range of values ​​for the duration of each prosodic tag and the duration of each pause. The speech synthesis model is trained using prosodic-matched speech data and labeled second text, enabling the model to accurately learn prosodic features. This allows the synthesized speech to clearly reflect the pauses represented by each prosodic label. Attached Figure Description

[0057] To more clearly illustrate the technical solution of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0058] Figure 1 This is a schematic flowchart illustrating a prosody annotation method provided in an embodiment of this application;

[0059] Figure 2 This is a schematic diagram of a syllable sequence text provided in an embodiment of this application;

[0060] Figure 3 This is a schematic diagram of a phoneme sequence text provided in an embodiment of this application;

[0061] Figure 4 This is a schematic diagram of an example of a voice audio file and annotated text structure provided in an embodiment of this application;

[0062] Figure 5 This is a schematic diagram of a prosody annotation device provided in an embodiment of this application. Detailed Implementation

[0063] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having like or similar functions throughout. The embodiments described below by referring to the drawings are exemplary and are only used to explain the present application and should not be construed as a limitation to the present application. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.

[0064] Those skilled in the art of the present technology can understand that, unless specifically stated otherwise, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the term "comprising" used in the specification of the present application means that there are the described features, integers, steps, operations, elements and / or components, but does not exclude the existence or addition of one or more other features, integers, steps, operations, elements, components and / or their groups. It should be understood that when we say that an element is "connected" or "coupled" to another element, it can be directly connected or coupled to other elements, or there may also be intermediate elements. In addition, the "connection" or "coupling" used herein may include wireless connection or wireless coupling. The phrase "and / or" used herein includes all or any unit and all combinations of one or more associated listed items.

[0065] Currently, the prosody annotation scheme mainly annotates prosody through semantics. However, when a speaker records a speech audio according to the text with prosody annotation, they usually cannot read it completely according to the prosody annotation divided by semantics, but will unconsciously read it according to their own reading habits, and it is also impossible for the speaker to grasp the pause duration of each prosody label. Therefore, there is a problem that the prosody of the recorded speech audio does not match the prosody of the annotated text. Using such speech audio with mismatched prosody and the annotated text to train a speech synthesis model will cause the model to fail to learn the prosody features, resulting in the synthesized speech not being able to clearly show the pauses represented by each prosody label.

[0066] To facilitate the understanding of the solutions in the present application, some technical concepts will be briefly introduced below:

[0067] Phoneme: A phoneme is the smallest speech unit divided according to the natural attributes of speech. In Chinese, usually the pronunciation of one Chinese character is one syllable, and the basic syllables in Mandarin are composed of one to four phonemes according to certain combination rules. For example, for "你" (nǐ), two different phonemes "n" and "i" can be divided; for "好" (hǎo), three phonemes "h, α, o" can be divided.

[0068] Normal distribution: Also known as the "normal distribution", and also known as the Gaussian distribution, which was first obtained by A. De Moivre in the asymptotic formula for the binomial distribution. C.F. Gauss derived it from another perspective when studying measurement errors. P.S. Laplace and Gauss studied its properties. It is a very important probability distribution in the fields of mathematics, physics, and engineering, and has a significant influence in many aspects of statistics. The normal distribution curve is bell-shaped, low at both ends, high in the middle, and symmetric about the y-axis. Because its curve is bell-shaped, it is often called the bell curve. If the random variable X follows a normal distribution with a mathematical expectation of μ and a variance of σ^2, it is denoted as N(μ, σ^2). Its probability density function is determined by the expectation μ of the normal distribution, and its standard deviation σ determines the amplitude of the distribution. When μ = 0 and σ = 1, the normal distribution is the standard normal distribution.

[0069] Figure 1 This is a schematic flowchart of a prosody annotation method provided by an embodiment of the present application. To solve the above problems, an embodiment of the present application provides a prosody annotation method, which will be described below in conjunction with Figure 1 to illustrate this method.

[0070] S110, obtain the first speech data recorded by a specific speaker according to the first text.

[0071] Among them, the first text is a text that has been prosody-annotated according to semantics. The first text includes multiple prosody tags, and each prosody tag represents a pause between two word segments. Different prosody tags in the multiple prosody tags represent pauses of different durations.

[0072] For example, use prosody tags #1, #2, #3, and #4 to represent different pauses. Among them, #1 is the boundary of a prosodic word, representing a short pause; #2 is the boundary of a prosodic phrase, representing a lengthened sound or a short pause; #3 represents a more obvious pause and a falling intonation when the semantics is complete; #4 represents the end of a sentence, which is the annotation at the end of each numbered sentence. The prosody-annotated first text is "Erjing's family #2 lives #1 outside the North Fifth #1 Ring Road #3, and has to go #2 to #1 the Asian Games Village #2 Huatang #1 Mall #4 for work." It can be seen that there is a prosody tag between the word segments "Erjing's family" and "lives", indicating that after reading the first word segment, a pause of duration #2 should be made before reading the second word segment.

[0073] In one example, the method further includes: if the prosody tags in the first text do not match the pauses in the first speech data, re-determine the prosody tags in the first text according to the pauses in the first speech data.

[0074] For example, the mismatch between prosodic tags in the first text and pauses in the first speech data includes the following situations:

[0075] (1) A prosodic label was marked at a certain position in the first text, but the first speech data did not pause at that position;

[0076] (2) There is no prosodic label at a certain position in the first text, while there is a pause in the first speech data at that position;

[0077] (3) The absolute difference between the specified duration of the prosodic label marked at a certain position of the first text and the duration of the pause of the first speech data at that position exceeds the first threshold. The first threshold can be set according to actual needs.

[0078] The methods for re-determining prosodic labels based on the above situations include:

[0079] If a prosodic label is marked at a certain position in the first text, and the first speech data does not pause at that position, the prosodic label in the first text is re-determined, including: deleting the prosodic label at that position in the first text.

[0080] If a certain position in the first text is not marked with a prosodic label, but the first speech data has a pause at that position, the prosodic label in the first text is re-determined, including: adding a prosodic label corresponding to the pause at that position in the first text.

[0081] If the absolute difference between the specified duration of the prosodic label at a certain position in the first text and the duration of the pause in the first speech data at that position exceeds a first threshold, the prosodic label in the first text is re-determined, including: modifying the prosodic label at that position in the first text to a correct prosodic label, wherein the absolute difference between the correct prosodic label and the duration of the pause in the first speech data at that position does not exceed the first threshold.

[0082] In the above method, this application modifies the prosodic tags in the first text that do not match the pauses in the first speech data, so as to obtain a first text that is more in line with the speaking habits of a specific speaker, so as to obtain accurate statistical data on the duration of each prosodic tag in the future.

[0083] S120: Based on the first speech data and the first text, calculate the duration of each prosodic tag to obtain statistical data on the duration of each prosodic tag.

[0084] Taking the first text above, “Erjingjia #2 lives #1 outside the North Fifth Ring Road #3, goes to #1 Asian Games Village #2 Huatang #1 shopping mall #4 for work #2” as an example, the statistical data of the duration of rhythm tag #1 includes: 0.1s, 0.09s, 0.12s; the statistical data of the duration of rhythm tag #2 includes: 0.5s, 0.48s, 0.57s; the statistical data of the duration of rhythm tag #3 includes: 0.8s; and the statistical data of the duration of rhythm tag #4 includes: 1s.

[0085] In one example, the first text is processed into pinyin, resulting in pinyin text with prosodic tags. Analyzing the pinyin text and the first speech data using the Montreal Forced Aligner (MFA) algorithm yields syllable sequence text and phoneme sequence text. For example, Figure 2 As shown, the syllable sequence text records the start time (Xmin), end time (Xmax), and content (word or text) of each syllable, as well as the start time (Xmin), end time (Xmax), and content (word or text) of each prosodic tag. For example... Figure 3 As shown, the phoneme sequence text records the start time (Xmin), end time (Xmax), and content (word or text) of each phoneme, as well as the start time (Xmin), end time (Xmax), and content (word or text) of each prosodic tag. The duration of each prosodic tag (the difference between the end time and the start time) is obtained from the phoneme sequence text or syllable sequence text. Then, the duration of each prosodic tag is statistically analyzed to obtain statistical data on the duration of each prosodic tag.

[0086] S130, based on the statistical data corresponding to each prosodic tag, determine the range of values ​​for the duration of each prosodic tag.

[0087] In one example, the mean and standard deviation of the duration of each prosodic label are first determined based on the statistical data corresponding to each prosodic label. Then, based on the normal distribution analysis method, the range of values ​​for the duration of each prosodic label is determined based on the mean and standard deviation of the duration of each prosodic label.

[0088] Furthermore, based on normal distribution analysis, the method for determining the range of duration values ​​for each prosodic tag includes: first, determining the difference between the mean and standard deviation of the duration of each prosodic tag, and the sum of the mean and standard deviation of the duration of each prosodic tag, based on the mean and standard deviation of the duration of each prosodic tag; then, determining the range of duration values ​​for each prosodic tag based on the difference and sum corresponding to each prosodic tag. The minimum value of this range is the difference corresponding to each prosodic tag, and the maximum value of this range is the sum corresponding to each prosodic tag.

[0089] For example, the duration of rhythm tag #1 has a mean of 70ms, a standard deviation of 40, a difference of 30ms between the mean and the standard deviation, and a sum of 110ms between the mean and the standard deviation. Therefore, the duration of this rhythm tag can range from 30ms to 110ms.

[0090] In the above method, since the statistical data corresponding to each prosodic label conforms to a normal distribution, the range of values ​​for the duration of each prosodic label can be quickly and accurately estimated by using the mean and standard deviation of the duration of each prosodic label based on the characteristics of the normal distribution curve.

[0091] In one example, the method for determining the range of duration values ​​for each prosodic tag further includes: determining the maximum and minimum values ​​of the statistical data corresponding to each prosodic tag based on the statistical data corresponding to each prosodic tag; and determining the range of duration values ​​for each prosodic tag based on the maximum and minimum values ​​corresponding to each prosodic tag. The minimum value of this range is the minimum value of the statistical data corresponding to each prosodic tag, and the maximum value of this range is the maximum value of the statistical data corresponding to each prosodic tag.

[0092] In the above method, the range of duration of each prosodic label is directly determined based on the maximum and minimum values ​​of the statistical data corresponding to each prosodic label. The calculation method is simple and fast, but its accuracy is not as good as the range of duration determined based on the normal distribution.

[0093] S140, acquire second speech data recorded by a specific speaker based on a second text, wherein the second text is a text that has not been prosodic.

[0094] The second text is significantly larger than the first text. The speaker who recorded the second audio data is the same person who recorded the first audio data.

[0095] S150, Prosodic annotation is performed on the second text based on the range of duration values ​​for each prosodic tag and the second speech data.

[0096] Specifically, the duration of each pause in the second speech data is obtained based on the second speech data, and the second text is prosodicly annotated according to the range of the duration of each prosodic tag and the duration of each pause.

[0097] In one example, the second text is first processed into pinyin to obtain pinyin text. Then, based on the MFA algorithm, the start and end times of each pause are obtained from the second speech data and the pinyin text. Each pause represents the interval between the end time of the preceding phoneme and the start time of the following phoneme. Next, the duration of each pause is determined based on its start and end times, and the prosodic label corresponding to each pause is determined based on the range of values ​​for the duration of each pause and the duration of each prosodic label. Finally, the second text is prosodically annotated based on the prosodic labels corresponding to each pause.

[0098] In the above method, MFA can analyze and segment speech audio and pinyin text, thereby aligning the speech audio and pinyin text. During the generation of phoneme sequence text or syllable sequence text by MFA, since the pinyin text lacks prosodic tags to indicate pauses, if a phoneme has a gap (i.e., a pause between two phonemes), it will be represented as "Xmin=…,Xmax=…,word=''". The start time Xmin and end time Xmax of the pause can be determined by recognizing the keyword "word=''". The duration of each pause is Xmax-Xmin. Then, based on the duration of each pause and the range of the duration of each prosodic tag, the prosodic tag corresponding to each pause is determined. Finally, based on the prosodic tag corresponding to each pause, the keyword "word='prosodic tag'" is re-determined, and the new keyword replaces the original keyword at the corresponding position, generating a new MFA phoneme sequence text "Xmin=…,Xmax=…,word='prosodic tag'". Finally, the second text is prosodicly annotated based on the position and content of the prosodic tags in the new MFA phoneme sequence text. It should be understood that the above scheme is not limited to the duration of pauses obtained based on the MFA algorithm; other algorithms that can achieve the same effect can also be used.

[0099] This scheme identifies the start and end times of pauses in the generated text of MFA to obtain the duration of each pause. Then, by combining the duration of each pause with the range of duration of each prosodic tag, it accurately and quickly determines the prosodic tag corresponding to each pause. Finally, it can perform prosodic annotation on the second text based on the accurate prosodic tags.

[0100] As can be seen from the above embodiments, since the first text is prosodic-annotated text and includes multiple prosodic tags, the pause duration of each prosodic tag in the first speech data recorded by a specific speaker based on the prosodic tags of the first text depends on the specific speaker's speaking habits. Next, the duration of each prosodic tag is statistically analyzed to obtain statistical data on the duration of each prosodic tag. Since the statistical data on the duration of each prosodic tag is related to the specific speaker's speaking habits, the range of values ​​for the duration of each prosodic tag determined based on this statistical data also matches the specific speaker's speaking habits. Then, the specific speaker records second speech data based on a second text that has not been prosodic-annotated. The duration of each pause in the second speech data matches the specific speaker's speaking habits. Furthermore, since the range of values ​​for the duration of each prosodic tag matches the specific speaker's speaking habits, the prosody in the prosodic-annotated second text can accurately match the prosody of the second speech data based on the range of values ​​for the duration of each prosodic tag and the duration of each pause. The speech synthesis model is trained using prosodic-matched speech data and labeled second text, enabling the model to accurately learn prosodic features. This allows the synthesized speech to clearly reflect the pauses represented by each prosodic label.

[0101] Based on the prosodic annotation method described above, this application also provides an application of prosodic annotation to obtain text, the method of which includes:

[0102] First, based on each text in the second text and the first association relationship, the sequence number corresponding to each text is determined. The first association relationship is used to associate texts and sequence numbers. Next, the Mel-frequency cepstral coefficients (MFCC) features of the second speech data are obtained. Then, based on the duration of each prosodic tag in the prosodic-annotated second text, the sequence number of each text, and the MFCC features of the second speech data, the training dataset is determined. Finally, a speech synthesis model is trained based on the training dataset to obtain the trained speech synthesis model. The trained speech synthesis model is then used for speech synthesis.

[0103] Because the duration of each prosodic tag in the prosodic-annotated second text can accurately match the pause duration of a specific speaker, a training dataset is determined based on the duration of each prosodic tag in the prosodic-annotated second text, the corresponding text number, and the Mel-frequency cepstral coefficient features of the second speech data. The speech synthesis model trained on this dataset can accurately learn the prosodic features of a specific speaker, thus enabling the synthesized speech to clearly reflect the pauses represented by each prosodic tag, and these pauses match the speaking habits of the specific speaker.

[0104] The above method will be described below with reference to specific embodiments. For example, method 200 includes:

[0105] S210, acquire first speech data recorded by a specific speaker based on the first text.

[0106] Specifically, a specific speaker A is designated and asked to record 500 to 1000 speech entries (first speech data) based on a publicly available dataset (first text) that has already been prosodic-annotated.

[0107] S220, fine-tune the prosodic tags of the first text to obtain prosodic-annotated text that matches the first speech data.

[0108] Specifically, it is determined whether the prosodic tags in the first text match the pauses in the first speech data. If they do not match, the prosodic tags in the first text are re-determined based on the pauses in the first speech data.

[0109] For example, if the first text includes "Zhang San #2 Hello #4, Are you there #4", but the corresponding audio does not pause between "Zhang San" and "Hello", then the text should be changed to "Zhang San Hello #4, Are you there #4".

[0110] The final generated file includes an audio file in ".wav" format to store the initial speech data, and a first text file in ".txt" format after fine-tuning the prosodic tags. The file structure is as follows: Figure 4 As shown.

[0111] S230, based on the MFA algorithm, calculates the duration of each prosodic tag according to the first speech data and the first text to obtain statistical data on the duration of each prosodic tag.

[0112] Specifically, the first text is first processed into pinyin to obtain pinyin text. MFA analyzes and segments the first speech data and the corresponding pinyin text to align the first speech data and the pinyin text, resulting in a phoneme sequence text. The duration of each prosodic tag is then obtained from the phoneme sequence text.

[0113] For example, according to the phoneme sequence text, 'Xmin=2.08, Xmax=2.18, word=#1' means that the pause of #1 lasted from 2.08 seconds to 2.18 seconds, for a total of about 0.1 seconds. Through the analysis results of MFA, we can obtain the statistical data on the duration of #1, #2, #3 and #4.

[0114] S240, determine the range of values ​​for the duration of each prosodic tag.

[0115] Specifically, the mean and standard deviation are calculated based on the duration statistics of #1, #2, #3, and #4. Then, based on the fact that the duration statistics of each prosodic label conform to a normal distribution, the range of duration values ​​for #1, #2, #3, and #4 is (mean - standard deviation) to (mean + standard deviation). This range of values ​​for each prosodic label corresponds to the speaking habits of a specific speaker.

[0116] S250, instructs a specific speaker A to record second speech data based on a second text that has not been prosodic.

[0117] Specifically, the second text is large-scale data.

[0118] S260, based on the MFA algorithm, obtains phoneme sequence text from large-scale second text and second speech data, and determines the duration of each pause in the phoneme sequence text.

[0119] Specifically, in the MFA alignment process, for example Figure 3 If there is a gap between phonemes, it will be represented as "Xmin=…,Xmax=…,word=''". By recognizing the keyword "word=''", the start time Xmin and end time Xmax of the pause are determined, and the duration of the pause is further determined as (Xmax-Xmin). In this way, the duration of all pauses in the phoneme sequence text is determined.

[0120] S270, perform prosodic annotation on the second text.

[0121] Specifically, the range of values ​​for the prosodic tag corresponding to each pause is determined based on its duration, thus identifying the prosodic tag for each pause. Then, based on the prosodic tag for each pause, the keyword "word = '#1 (or #2 or #3 or #4)'" is redefined, and this new keyword replaces the original keyword at the corresponding position, generating a new phoneme sequence text "Xmin = ..., Xmax = ..., word = '#1 (or #2 or #3 or #4)'". Finally, based on the position and content of the prosodic tags #1, #2, #3, or #4 in the new phoneme sequence text, the corresponding prosodic tag content is annotated at the corresponding position in the second text, thereby completing the prosodic annotation of the second text.

[0122] In one example, method 200 also includes:

[0123] S280, Based on the text in the second document and the first association relationship, determine the corresponding sequence number in the text;

[0124] S290, Obtain the MFCC features of the second speech data;

[0125] S2100, the training dataset is determined based on the duration of each prosodic tag of the second text after prosodic annotation, the sequence number of each text, and the Mel frequency cepstral coefficient features of the second speech data.

[0126] S2110, train the speech synthesis model based on the training dataset, and obtain the trained speech synthesis model.

[0127] In the above embodiments, firstly, a specific speaker records first speech data based on a first text that has already been prosodicly annotated. Since the specific speaker cannot pause exactly according to the annotated prosodic tags, but rather records according to their own speaking habits, the pause duration in the first speech data depends on the specific speaker's speaking habits (i.e., the specific speaker's prosodic characteristics). Therefore, the value range determined based on the duration of each prosodic tag statistically analyzed from the first speech data can accurately match the pause duration of the specific speaker. Specifically, since the statistical data of the duration of each prosodic tag conforms to a normal distribution, the value range can be quickly and accurately determined based on the mean and standard deviation of the duration of each prosodic tag. Next, the same specific speaker records second speech data based on a second text that has not been prosodicly annotated. The MFA algorithm is used to quickly analyze the pause duration in the second speech data, and the prosodic tags corresponding to each pause duration are determined based on the value range of the duration of each prosodic tag. The prosodicity of the second text after prosodic annotation using these prosodic tags can accurately match the prosodicity of the second speech data. Because the duration of each prosodic tag in the prosodic-annotated second text can accurately match the pause duration of a specific speaker, a training dataset is determined based on the duration of each prosodic tag in the prosodic-annotated second text, the corresponding text sequence number, and the Mel-frequency cepstral coefficient features of the second speech data. The speech synthesis model trained on this dataset can accurately learn the prosodic features of a specific speaker, thus enabling the synthesized speech to clearly reflect the pauses represented by each prosodic tag, and these pauses match the speaker's speaking habits.

[0128] Based on the above prosody annotation method, this application also provides a prosody annotation device, such as... Figure 5 As shown, the device includes:

[0129] The voice data acquisition module 310 is used to acquire first voice data recorded by a specific speaker based on a first text. The first text is a text that has been marked with prosody. The first text includes multiple prosody tags, and different prosody tags represent different durations of pauses.

[0130] The prosodic tag duration determination module 320 is used to calculate the duration of each prosodic tag among multiple prosodic tags based on the first speech data and the first text, so as to obtain statistical data on the duration of each prosodic tag.

[0131] The prosody tag duration determination module 320 is also used to determine the range of values ​​for the duration of each prosody tag based on the statistical data corresponding to each prosody tag.

[0132] The voice data acquisition module 320 is also used to acquire second voice data recorded by a specific speaker based on a second text, wherein the second text is text that has not been prosodic.

[0133] The prosody annotation module 330 is used to obtain the duration of each pause in the second speech data based on the second speech data, and to perform prosody annotation on the second text based on the value range of the duration of each prosody tag and the duration of each pause.

[0134] In one example, the device also includes:

[0135] The prosody tag fine-tuning module 340 is used to redetermine the prosody tags in the first text based on the pauses in the first speech data if the prosody tags in the first text do not match the pauses in the first speech data.

[0136] Other implementations and effects of this device are described in Method 100 and Method 200, and will not be repeated here.

[0137] The basic principles of this application have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this application are merely examples and not limitations, and should not be considered as essential features of each embodiment of this application. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the application to the necessity of employing the aforementioned specific details for implementation.

[0138] It should be understood that although the steps in the flowcharts of the accompanying figures are shown sequentially as indicated by the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowcharts of the accompanying figures may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times, and their execution order is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.

[0139] The block diagrams of devices, apparatuses, devices, and systems involved in this application are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.

[0140] It should also be noted that in the apparatus, equipment, and methods of this application, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions of this application.

[0141] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this application. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein can be applied to other aspects without departing from the scope of this application. Therefore, this application is not intended to be limited to the aspects shown herein, but rather to be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0142] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this application to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.

Claims

1. A method for prosodic annotation, characterized in that, include: Acquire first speech data recorded by a specific speaker based on a first text, wherein the first text is a text with prosodic annotation, and the first text includes multiple prosodic tags, wherein different prosodic tags represent different durations of pauses. Based on the first speech data and the first text, the duration of each prosodic tag in the plurality of prosodic tags is calculated to obtain statistical data on the duration of each prosodic tag. Based on the statistical data corresponding to each prosodic tag, determine the range of values ​​for the duration of each prosodic tag; Acquire second speech data recorded by the specific speaker based on a second text, wherein the second text is text that has not been prosodic; The duration of each pause in the second speech data is obtained based on the second speech data, and prosodic annotation is performed on the second text based on the range of values ​​for the duration of each prosodic tag and the duration of each pause. The step of obtaining the duration of each pause in the second speech data based on the second speech data and performing prosodic annotation on the second text based on the range of values ​​for the duration of each prosodic tag and the duration of each pause includes: The second text is processed into pinyin to obtain pinyin text; Based on the Montreal Forced Alignment (MFA) algorithm, the start and end times of each pause are obtained according to the second speech data and the pinyin text. Each pause indicates that there is an interval between the end time of the preceding phoneme and the start time of the following phoneme. The duration of each pause is determined based on the start and end times of each pause; Based on the duration of each pause and the range of values ​​for the duration of each prosodic tag, determine the prosodic tag corresponding to each pause; The second text is prosodicly annotated based on the prosodic tag corresponding to each pause.

2. The method according to claim 1, characterized in that, The method further includes: If the prosodic tags in the first text do not match the pauses in the first speech data, the prosodic tags in the first text are re-determined based on the pauses in the first speech data.

3. The method according to claim 1 or 2, characterized in that, The step of determining the range of duration values ​​for each prosodic tag based on the statistical data corresponding to each prosodic tag includes: Based on the statistical data corresponding to each prosodic label, determine the mean and standard deviation of the duration of each prosodic label; Based on the normal distribution analysis method, the range of values ​​for the duration of each prosodic tag is determined according to the mean and standard deviation of the duration of each prosodic tag.

4. The method according to claim 3, characterized in that, The step of determining the range of duration values ​​for each prosodic tag based on the mean and standard deviation of the duration of each prosodic tag includes: Based on the mean and standard deviation of the duration of each prosodic tag, determine the difference between the mean and standard deviation of the duration of each prosodic tag, and the sum of the mean and standard deviation of the duration of each prosodic tag. The range of duration values ​​for each prosodic tag is determined based on the difference and sum values ​​corresponding to each prosodic tag.

5. The method according to claim 1 or 2, characterized in that, The step of determining the range of duration values ​​for each prosodic tag based on the statistical data corresponding to each prosodic tag includes: Based on the statistical data corresponding to each prosodic label, determine the maximum and minimum values ​​of the statistical data corresponding to each prosodic label; The duration range of each prosodic tag is determined based on the maximum and minimum values ​​corresponding to each prosodic tag.

6. A prosody marking device, characterized in that, include: The speech data acquisition module is used to acquire first speech data recorded by a specific speaker based on a first text. The first text is a text with prosodic annotation. The first text includes multiple prosodic tags, and different prosodic tags represent different durations of pauses. The prosodic tag duration determination module is used to calculate the duration of each prosodic tag among the plurality of prosodic tags based on the first speech data and the first text, so as to obtain statistical data on the duration of each prosodic tag. The prosody tag duration determination module is also used to determine the range of values ​​for the duration of each prosody tag based on the statistical data corresponding to each prosody tag. The voice data acquisition module is also used to acquire second voice data recorded by the specific speaker based on the second text, wherein the second text is text that has not been prosodic. The prosodic annotation module is used to obtain the duration of each pause in the second speech data based on the second speech data, and to perform prosodic annotation on the second text based on the value range of the duration of each prosodic tag and the duration of each pause. The step of obtaining the duration of each pause in the second speech data based on the second speech data, and performing prosodic annotation on the second text based on the value range of the duration of each prosodic tag and the duration of each pause, includes: The second text is processed into pinyin to obtain pinyin text; Based on the Montreal Forced Alignment (MFA) algorithm, the start and end times of each pause are obtained according to the second speech data and the pinyin text. Each pause indicates that there is an interval between the end time of the preceding phoneme and the start time of the following phoneme. The duration of each pause is determined based on the start and end times of each pause; Based on the duration of each pause and the range of values ​​for the duration of each prosodic tag, determine the prosodic tag corresponding to each pause; The second text is prosodicly annotated based on the prosodic tag corresponding to each pause.

7. The apparatus according to claim 6, characterized in that, The device further includes: The prosody tag fine-tuning module is used to redetermine the prosody tags in the first text based on the pauses in the first speech data if the prosody tags in the first text do not match the pauses in the first speech data.

8. The apparatus according to claim 6 or 7, characterized in that, The step of determining the range of duration values ​​for each prosodic tag based on the statistical data corresponding to each prosodic tag includes: Based on the statistical data corresponding to each prosodic label, determine the mean and standard deviation of the duration of each prosodic label; Based on the normal distribution analysis method, the range of values ​​for the duration of each prosodic tag is determined according to the mean and standard deviation of the duration of each prosodic tag.

9. A method for obtaining text through prosodic annotation, characterized in that the method... include: Acquire first speech data recorded by a specific speaker based on a first text, wherein the first text is a text with prosodic annotation, and the first text includes multiple prosodic tags, wherein different prosodic tags represent different durations of pauses. Based on the first speech data and the first text, the duration of each prosodic tag in the plurality of prosodic tags is calculated to obtain statistical data on the duration of each prosodic tag. Based on the statistical data corresponding to each prosodic tag, determine the range of values ​​for the duration of each prosodic tag; Acquire second speech data recorded by the specific speaker based on a second text, wherein the second text is text that has not been prosodic; The duration of each pause in the second speech data is obtained based on the second speech data, and prosodic annotation is performed on the second text based on the range of values ​​for the duration of each prosodic tag and the duration of each pause. The step of obtaining the duration of each pause in the second speech data based on the second speech data and performing prosodic annotation on the second text based on the range of values ​​for the duration of each prosodic tag and the duration of each pause includes: The second text is processed into pinyin to obtain pinyin text; Based on the Montreal Forced Alignment (MFA) algorithm, the start and end times of each pause are obtained according to the second speech data and the pinyin text. Each pause indicates that there is an interval between the end time of the preceding phoneme and the start time of the following phoneme. The duration of each pause is determined based on the start and end times of each pause; Based on the duration of each pause and the range of values ​​for the duration of each prosodic tag, determine the prosodic tag corresponding to each pause; The second text is punctuated with prosodic tags corresponding to each pause. Based on each text in the second text and the first association relationship, the sequence number corresponding to each text is determined, and the first association relationship is used to associate text and sequence number; Obtain the Mel frequency cepstral coefficient features of the second speech data; The training dataset is determined based on the duration of each prosodic tag in the prosodic-annotated second text, the sequence number corresponding to each text, and the Mel frequency cepstral coefficient features of the second speech data. A speech synthesis model is trained based on the training dataset to obtain the trained speech synthesis model.

Citation Information

Patent Citations

  • Speech annotation method and device

    CN109300468A

  • Rhythm marking method, device and equipment

    CN109326281A