Speech synthesis method, apparatus, medium, and electronic device
By identifying the type of laughter and generating multiple audio information sequences in speech synthesis technology, the problem of the difference between synthesized laughter and real human speech is solved, achieving a more realistic laughter synthesis effect.
Patent Information
- Application Number
- CN202111653397.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-30
- Publication Date
- 2025-11-11
- Estimated Expiration
- 2041-12-30
AI Technical Summary
Existing speech synthesis technologies produce laughter that differs significantly from human speech, making it difficult to achieve personalized and diverse laughter synthesis, thus affecting realism and accuracy.
By determining the type of laughter in the target text, and based on the pre-annotated audio information sequences of various laughter types, the laughter audio information of the target text is generated, including the target phoneme sequence, the target tone sequence, and the prosody sequence, and then synthesized using a speech synthesis model.
It improves the realism and diversity of laughter synthesis, reduces the difference between laughter and real human speech in speech synthesis, and enhances the accuracy of speech synthesis and its ability to fit real-world application scenarios.
Smart Images

Figure CN114255738B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of speech processing, and more specifically, to a speech synthesis method, apparatus, medium, and electronic device. Background Technology
[0002] Currently, speech synthesis technology has brought great convenience to people's lives. For example, TTS (Text To Speech) can intelligently convert text into natural speech streams through neural networks.
[0003] In the aforementioned text-to-speech conversion process, the synthesis of laughter is typically based on the phonetic annotation of the text's pinyin, thereby achieving speech synthesis of laughter. However, in practical applications, the vocalization methods and structural composition of different types of text may differ significantly from spoken language, and the speech synthesis in related technologies still differs considerably from human speech. Summary of the Invention
[0004] This section is provided to briefly introduce the concepts, which will be described in detail in the Detailed Description section later. This section is not intended to identify key or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.
[0005] In a first aspect, this disclosure provides a speech synthesis method, the method comprising:
[0006] Based on the received text to be processed, the target text in the text to be processed is determined, wherein the target text is the text used to represent the synthesis of laughter;
[0007] Based on the target text and the text to be processed, a target audio information sequence corresponding to the target text is determined, wherein the target audio information sequence is determined based on pre-annotated audio information sequences under various laughter types, and the audio information sequence under each laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence;
[0008] Based on the target audio information sequence and the speech synthesis model, laughter audio information of the target text is generated to obtain the audio information of the text to be processed.
[0009] Secondly, this disclosure provides a speech synthesis apparatus, the apparatus comprising:
[0010] The first determining module is used to determine the target text in the received text to be processed, wherein the target text is text used to represent laughter synthesis.
[0011] The second determining module is used to determine the target audio information sequence corresponding to the target text based on the target text and the text to be processed. The target audio information sequence is determined based on audio information sequences under pre-annotated multiple laughter types. Each audio information sequence under the laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence.
[0012] The generation module is used to generate laughter audio information of the target text based on the target audio information sequence and the speech synthesis model, so as to obtain the audio information of the text to be processed.
[0013] Thirdly, this disclosure provides a computer-readable medium having a computer program stored thereon, which, when executed by a processing device, implements the steps of the method described in the first aspect.
[0014] Fourthly, this disclosure provides an electronic device, comprising:
[0015] A storage device on which one or more computer programs are stored;
[0016] One or more processing means are configured to execute the one or more computer programs in the storage device to implement the steps of the method described in the first aspect.
[0017] In the above technical solution, based on the received text to be processed, the target text representing laughter synthesis is determined. Then, based on the target text and the text to be processed, the target audio information sequence corresponding to the target text is determined. Subsequently, based on the target audio information sequence and the speech synthesis model, the laughter audio information of the target text is generated to obtain the audio information of the text to be processed. Therefore, through this technical solution, personalized speech synthesis for paralinguistic sounds such as laughter can be performed during the text-to-speech synthesis process. When synthesizing laughter, for the target text corresponding to the laughter synthesis, the corresponding audio information sequence can be selected from a variety of pre-annotated laughter types. This allows for the use of multiple vocalization methods and compositional structures in the laughter synthesis, reducing the difference between the laughter obtained in the speech synthesis and the speech formed by a real person, effectively improving the realism of the speech synthesis. Simultaneously, by pre-annotating multiple laughter types, the diversity of laughter annotation can be increased, as can the diversity of laughter synthesis, improving the accuracy of the speech synthesis method and better aligning with real-world application scenarios.
[0018] Other features and advantages of this disclosure will be described in detail in the following detailed description section. Attached Figure Description
[0019] The above and other features, advantages, and aspects of the embodiments of this disclosure will become more apparent from the accompanying drawings and the following detailed description. Throughout the drawings, the same or similar reference numerals denote the same or similar elements. It should be understood that the drawings are schematic, and the originals and elements are not necessarily drawn to scale. In the drawings:
[0020] Figure 1 Here is a flowchart of a speech synthesis method provided according to one embodiment of the present disclosure;
[0021] Figure 2 This is a flowchart illustrating an exemplary implementation of determining the target audio information sequence corresponding to the target text based on the target text and the text to be processed.
[0022] Figure 3 This is a schematic diagram of a training audio sequence;
[0023] Figure 4 This is a block diagram of a speech synthesis apparatus provided according to one embodiment of the present disclosure;
[0024] Figure 5 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation
[0025] Embodiments of this disclosure will now be described in more detail with reference to the accompanying drawings. While some embodiments of this disclosure are shown in the drawings, it should be understood that this disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Rather, these embodiments are provided to provide a more thorough and complete understanding of this disclosure. It should be understood that the accompanying drawings and embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of protection of this disclosure.
[0026] It should be understood that the steps described in the method embodiments of this disclosure may be performed in different orders and / or in parallel. Furthermore, the method embodiments may include additional steps and / or omit the steps shown. The scope of this disclosure is not limited in this respect.
[0027] The term "comprising" and its variations as used herein are open-ended inclusions, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments". Definitions of other terms will be given in the description below.
[0028] It should be noted that the concepts of "first" and "second" mentioned in this disclosure are used only to distinguish different devices, modules or units, and are not used to limit the order of functions performed by these devices, modules or units or their interdependencies.
[0029] It should be noted that the terms "a" and "a plurality of" used in this disclosure are illustrative rather than restrictive, and those skilled in the art should understand that, unless otherwise expressly indicated in the context, they should be understood as "one or more".
[0030] The names of messages or information exchanged between multiple devices in the embodiments of this disclosure are for illustrative purposes only and are not intended to limit the scope of such messages or information.
[0031] Figure 1 The diagram shown is a flowchart of a speech synthesis method according to an embodiment of this disclosure. Figure 1 As shown, the method includes:
[0032] In step 11, the target text in the received text to be processed is determined, wherein the target text is the text used to represent the synthesis of laughter.
[0033] The text to be processed can be text used for speech synthesis, such as an article or a dialogue.
[0034] As an example, the target text in the text to be processed can be determined by keyword matching. For instance, in a real-world application scenario, laughter is usually represented by a series of words such as "ha", "he", "xi", and "hey" in a text. This can be detected by keywords. Once the consecutive keywords are identified, the identified consecutive keywords are determined to be the target text.
[0035] As another example, a text prediction model can be used to determine the target text within the text to be processed. For instance, the text can be pre-annotated, specifically the parts of the text corresponding to laughter. A neural network can then be trained based on these annotated samples to obtain a trained predictive text model. This neural network can employ commonly used network models in the field, and the training method will not be elaborated upon here. This text prediction model can then predict the target text corresponding to the positions in the input text where laughter synthesis is needed.
[0036] In step 12, the target audio information sequence corresponding to the target text is determined based on the target text and the text to be processed. The target audio information sequence is determined based on the audio information sequences of various pre-annotated laughter types. Each laughter type audio information sequence includes a target phoneme sequence, a target tone sequence, and a prosody sequence.
[0037] In this embodiment, various laughter type annotations can be pre-made according to the vocalization modes of laughter that may occur in the actual scenario, that is, the phoneme sequences, tone sequences, and prosody sequences under various laughter types are pre-annotated. Thus, when performing laughter synthesis, the corresponding audio information sequences can be selected from the pre-annotated various laughter types, which can improve the diversity of laughter annotation, and at the same time, can also improve the diversity of laughter synthesis and the matching degree with the actual application scenario.
[0038] Among them, a phoneme is the smallest speech unit divided according to the natural attributes of speech. Analyzed according to the pronunciation actions in a syllable, one action constitutes one phoneme. Phonemes are divided into two major categories: vowels and consonants. For Chinese Mandarin, phonemes include initials (initials are consonants used before finals, which together with finals form a complete syllable) and finals (i.e., vowels). For example, the phonemes corresponding to "nihao" are "nihao". Tones are used to represent the changes in the pitch and intonation of sounds. Prosody is used to indicate where to pause when reading a text.
[0039] In step 13, according to the target audio information sequence and the speech synthesis model, the laughter audio information of the target text is generated to obtain the audio information of the text to be processed.
[0040] For example, for the target text in the text to be processed, the target audio information sequence can be obtained through the above method, so as to obtain the laughter audio information of the target text; for the other texts in the text to be processed except the target text, speech synthesis can be performed according to the TTS technology in the related art to obtain the audio information corresponding to the other texts, and then the audio information corresponding to the whole text to be processed can be obtained by combining the laughter audio information.
[0041] In the above technical solution, based on the received text to be processed, the target text representing laughter synthesis is determined. Then, based on the target text and the text to be processed, the target audio information sequence corresponding to the target text is determined. Subsequently, based on the target audio information sequence and the speech synthesis model, the laughter audio information of the target text is generated to obtain the audio information of the text to be processed. Therefore, through this technical solution, personalized speech synthesis for paralinguistic sounds such as laughter can be performed during the text-to-speech synthesis process. When synthesizing laughter, for the target text corresponding to the laughter synthesis, the corresponding audio information sequence can be selected from a variety of pre-annotated laughter types. This allows for the use of multiple vocalization methods and compositional structures in the laughter synthesis, reducing the difference between the laughter obtained in the speech synthesis and the speech formed by a real person, effectively improving the realism of the speech synthesis. Simultaneously, by pre-annotating multiple laughter types, the diversity of laughter annotation can be increased, as can the diversity of laughter synthesis, improving the accuracy of the speech synthesis method and better aligning with real-world application scenarios.
[0042] In one possible embodiment, an exemplary implementation of determining the target audio information sequence corresponding to the target text based on the target text and the text to be processed is as follows: Figure 2 As shown, this step may include:
[0043] In step 21, the context text corresponding to the target text in the text to be processed is determined.
[0044] The context text corresponding to the target text can be the complete sentence to which the target text belongs. After obtaining the text to be processed, it is usually necessary to process it into sentences. After determining the corresponding target text, the complete sentence to which the target text belongs can be obtained as the context text, so as to analyze the scene type corresponding to the target text.
[0045] In step 22, the scene type corresponding to the target text is determined based on the context text, wherein each scene type corresponds to at least one candidate audio information sequence, and the candidate audio information sequence includes at least one audio information sequence under the laughter type.
[0046] The scene type can be used to represent the language scenario for language synthesis. For example, the scene type can include, but is not limited to, various types such as awkward, reluctant, helpless, emotionless, generally happy, very happy, and proud. It can be set according to the actual application scenario, and this disclosure does not limit it. It should be noted that a corresponding candidate audio information sequence is set under each scene type. The candidate audio information sequence is used to represent the audio features of laughter synthesis under that scene type. For example, one or more of the pre-annotated audio information sequences under various laughter types can be combined to obtain the candidate audio information sequence under a scene type.
[0047] As an example, various statements can be pre-labeled with scene types. The labeled statements can then be used as training samples to train a classification model, resulting in a trained scene classification model. This classification model can be a commonly used multi-class classification model in the field. The statements in the training samples can be used as input, and the corresponding scene type labels can be used as the target output to train the multi-class classification model. The specific training process will not be elaborated here. Therefore, the context text can be input into the scene classification model, and the classification output by the scene classification model can be taken as the scene type.
[0048] As another example, the implementation of determining the scene type corresponding to the target text based on the context text is as follows, and this step may include:
[0049] The narrator's role information corresponding to the context text is determined. In practical applications, the context containing laughter synthesis is typically dialogue between characters in a text, such as a novel or script. Therefore, in this embodiment, the narrator's role information corresponding to the context text can be determined, i.e., the characteristic information of the character speaking the context text. The characteristic information of each character in the text can be obtained through character descriptions, such as character descriptions in a novel or introductions in a script. Thus, after determining the narrator corresponding to the context text, the narrator's role information can be directly obtained based on the narrator's characteristic information. For example, the narrator's role information can be used to characterize whether the narrator is a protagonist or antagonist.
[0050] Then, the scene is classified according to the context text and the narrator's role information, and the determined classification is identified as the scene type.
[0051] Similarly, in this step, various statements and their corresponding character features can be pre-labeled with scene types. The labeled statements are then used as training samples to train the classification model, resulting in a trained scene classification model. This classification model can be a commonly used multi-classification model in this field. The input can be the statements and their corresponding character features from the training samples, with the scene type label corresponding to the statement as the target output. The specific training process will not be elaborated here. Thus, the context text and narrator character information can be input into the scene classification model, and the classification output by the scene classification model can be taken as the scene type.
[0052] Therefore, in this embodiment, the scene type corresponding to the target text can be determined by simultaneously combining the content of the context text and the narrator's character information (i.e., character features) corresponding to the context text. This improves the accuracy of the determined scene type and provides accurate data support for subsequent speech synthesis of the laughter corresponding to the target text. It also improves the matching between the synthesized laughter and the scene of the text content to a certain extent, further enhancing the realism of the speech synthesis.
[0053] In step 23, the target audio information sequence corresponding to the target text is generated based on the candidate audio information sequence under the scene type.
[0054] Therefore, through the above technical solution, the scene to which the target text belongs can be predicted based on the context information corresponding to the target text. Thus, the target audio information sequence corresponding to the target text can be generated based on the candidate audio information sequence under the determined scene type, so that the target audio information sequence corresponding to the target text matches the scene to which the target text belongs. While performing personalized speech synthesis on the target text, the diversity of laughter synthesis is further improved, which fits the paralinguistic pronunciation of spoken dialogue in actual application scenarios and improves the degree of speech synthesis's representation of text content.
[0055] As an example, the applicant made the following findings by analyzing laughter audio from real-world application scenarios:
[0056] 1. Laughter at the beginning or middle of a sentence usually includes a restorative inhalation (the vocal cords may or may not vibrate during the phonation), i.e., offset in the table below. This inhalation ensures that the speaker can continue speaking after taking a breath, while laughter at the end of a sentence generally does not have an inhalation.
[0057] 2. Laughter may begin with a vowel initiation segment, i.e., onset in the table below.
[0058] 3. The vowel segment in the previous laughter unit is often influenced by the consonant segment in the next laughter unit, with the second half being aspirated and the formant changing; the beginning of laughter may have a release phase like a plosive, producing a straight line.
[0059] Therefore, in this embodiment of the disclosure, the candidate audio information sequences for various scene types pre-set based on the above-mentioned laughter features are partially configured as follows:
[0060] Scene type Candidate audio information sequence Number of laugh syllables Awkward hn>、he>、ha> Odd or even reluctantly he-、ha> Odd or even have no choice hx one No emotional tendency hn>、he>、ha>、ha-、h Odd or even Generally happy hn>, hn-, he>, ha>, hx, hy> Single, double, multiple Very happy hn>、hn<、he>、he<、ha>、hx、hy> many proud hy-、hy< Single, double, multiple
[0061] The phoneme sequences of the pre-annotated audio information sequences for various laughter types are represented as follows:
[0062]
[0063]
[0064] As shown in the table above, the various laughter types include laughter types corresponding to the inhalation phase and laughter types corresponding to the exhalation phase.
[0065] Among them, the laughter types in the inspiratory segment include the first laughter type corresponding to vocal cord vibration (as marked vd in the table above) and the second laughter type corresponding to vocal cord non-vibration (as marked uvd in the table above);
[0066] The laughter types in the exhalation phase include vocal laughter types corresponding to the vocalization phase and initiation laughter types corresponding to the initiation phase. Each syllable in the phoneme sequence of the vocal laughter type is based on consonants and vowels, while each syllable in the phoneme sequence of the initiation laughter type is based on vowels. The vocal laughter types can be represented as types 1-8 in the table above, and the initiation laughter types can be represented as variants 1 and 2 in the table above. That is, the laughter types corresponding to variants 1 and 2 are used to represent the vowel initiation phase present at the beginning of laughter.
[0067] It should be noted that the classification of laughter types shown in the table above is an illustrative example and does not limit this disclosure. Other laughter types and labels can be determined based on the above classification and actual application scenarios to improve the precision and diversity of laughter types, thereby improving the diversity of laughter synthesis and enhancing the anthropomorphism and realism of the synthesized laughter.
[0068] Among the determined target audio information sequences, different phonemes corresponding to the same syllable have the same tone and prosody. The tones include rising tone, flat tone, and falling tone. In the audio information sequence, the symbol after the phoneme sequence is used to represent tone information. Among them, ">" is used to represent the falling tone, "-" is used to represent the flat tone, and "<" is used to represent the rising tone. Thus, after determining the scene type, the target audio information sequence corresponding to the target text can be determined according to the candidate audio information sequence under this scene type. For example, if the determined scene type is "reluctant", the target audio information sequence corresponding to the target text can be determined according to the candidate audio information sequence "he-, ha>" corresponding to "reluctant". If the target text is "haha", "ha>" can be selected from the candidate audio information sequence "he-, ha>" to further determine the phoneme sequence and tone sequence, so as to obtain the target phoneme sequence and target tone sequence "ha>ha>".
[0069] At the same time, according to the corresponding number of laughter syllables including "single, double", the prosody features are further determined, which can be generated according to the default mode of this scene type or according to the actual syllable number. For example, the target text can be regarded as a prosody phrase, and a pause at the prosody phrase level needs to be made after the target text. Thus, the laughter can be characterized by various phoneme sequences and tone sequences, and the laughter features corresponding to various scenes can be obtained, making the synthesis of laughter speech more in line with the pronunciation mode of real people and improving the accuracy of speech synthesis.
[0070] In a possible embodiment, the exemplary implementation manner of generating the target audio information sequence corresponding to the target text according to the candidate audio information sequence under the scene type is as follows. The steps may include:
[0071] Determine the phoneme sequence and tone sequence corresponding to each syllable included in the target text according to the candidate audio information sequence corresponding to the scene type. The method of determining the phoneme sequence and tone sequence in this step is similar to that above and will not be elaborated here.
[0072] Determine the target audio information sequence corresponding to the target text according to the phoneme sequence and tone sequence corresponding to each syllable included in the target text and the number of syllables in the target text.
[0073] As an example, when the number of syllables in the target text is less than a preset threshold, the target audio information sequence corresponding to the target text can be determined based on the phoneme sequence and tone sequence corresponding to each syllable in the target text, as determined from the candidate audio information sequence. For example, in this step, if the phoneme sequence and tone sequence corresponding to each syllable in the target text determined from the candidate audio information sequence are "ha>", "ha>", and "uvd", and the number of syllables is less than the threshold, then the sequences corresponding to each of the above syllables can be concatenated to obtain the target phoneme sequence and target tone sequence corresponding to the target text as follows: "ha>ha>uvd". The method for determining the prosodic sequence is similar to that described above and will not be repeated here.
[0074] As another example, when the number of syllables in the target text is greater than or equal to a preset threshold, the phoneme sequence and tone sequence corresponding to each syllable in the target text can be selected based on the number of syllables to determine the target audio information sequence corresponding to the target text. For example, in a novel text, the text "hahahahahahahahahahaha" can represent loud laughter. In actual speech synthesis, the speech will not be produced based on the number of syllables in the text. For such long texts, the selection can be made based on the phoneme sequence and tone sequence corresponding to each syllable, combined with the number of syllables in the target text. For example, in this scenario, a preset threshold of 6 can be set. For such long texts, it is possible to select 3, 4, or 5 syllables as the number of syllables in the target audio information sequence. Specifically, the target phoneme sequence and target tone can be determined based on the phoneme sequence and tone corresponding to each syllable.
[0075] Therefore, by using the above technical solution, while generating diverse audio information sequences for the target text, the number of syllables contained in the target text can be further combined to further improve the accuracy of the target audio information sequence corresponding to the determined target text and its adaptability to actual application scenarios, thereby further improving the realism of speech synthesis and enhancing the user experience.
[0076] In one possible embodiment, the speech synthesis model is obtained in the following way:
[0077] Obtain training samples, wherein each training sample includes a training audio sequence and a labeled sequence of the training audio sequence, the labeled sequence including a training phoneme sequence, a training tone sequence and a training prosody sequence corresponding to the training audio sequence.
[0078] The training phoneme sequence, training tone sequence, and training prosody sequence can be obtained by manual annotation based on the laughter features and phoneme sequences described above.
[0079] As an example, the annotation sequence of the training audio sequence in each of the training samples is determined based on the spectrogram of the training audio and the audio information sequences under the pre-annotated multiple laughter types, where the laughter types include the laughter type corresponding to the inhalation segment and the laughter type corresponding to the exhalation segment. Among them, the audio information sequence under the laughter type corresponding to the inhalation segment is annotated in the inhalation segment determined based on the spectrogram, and the audio information sequence under the laughter type corresponding to the exhalation segment is annotated in the exhalation segment determined based on the spectrogram. Among them, the pre-annotated multiple laughter types have been described in detail above. It should be noted that the laughter types corresponding to the table above are for illustrative purposes and do not limit the number of laughter types in the present disclosure, etc.
[0080] Exemplarily, as Figure 3 shown is a schematic diagram of the spectrogram of a training audio sequence. In Figure 3 the example, it includes 3 audio waveforms, where each audio waveform corresponds to a syllable. Through the above syllable sequence, the training text can be annotated to obtain its corresponding training phoneme sequence, training tone sequence, and training prosody sequence "ha>ha>uvd, prosodic phrase pause". Among them, the pre-annotated multiple laughter types in the above table information can be used for manual annotation to obtain annotation information. In the first-layer exhalation segment, the phonetically similar Chinese characters "haha" are annotated, and in the inhalation segment, sp is marked to indicate a short pause in the sentence. In the second-layer exhalation segment, the training phoneme sequence and training tone sequence are annotated, and in the inhalation segment, uvd is marked to indicate that the vocal cords do not vibrate.
[0081] For each training audio sequence, the vectors corresponding to the annotation sequences of the training audio sequence are concatenated to obtain a concatenated vector, and the concatenated vector is input into the encoder of the preset model to obtain the feature vector corresponding to the training audio sequence. Among them, the annotation sequence can be converted into a vector by performing vectorization (embedding) processing on the training phoneme sequence, training tone sequence, and training prosody sequence corresponding to the training audio sequence. Exemplarily, the sequential concatenation method can be adopted, such as concatenation based on the concat() function, to obtain the concatenated vector. Among them, the preset model can be the Tacotron model, which is an end-to-end speech synthesis framework, and the preset model includes an encoder, an attention module, and a decoder.
[0082] The feature vector is input into the attention module of the preset model to obtain the context vector corresponding to the feature vector; the context vector is input into the decoder of the preset model to obtain the synthesized audio information corresponding to the training audio sequence. Exemplarily, the synthesized audio information can be a frame sequence of the Mel spectrogram (Mel).
[0083] Specifically, inputting the feature vector into the attention module for computation allows for greater focus on laughter-related features during speech synthesis, thereby improving the accuracy of the synthesized audio information to some extent when decoding based on context vectors. The specific processing methods of the attention module and encoder can employ computational techniques commonly used in Tacotron models in this field, and will not be elaborated upon here.
[0084] Based on the synthesized audio information and the target audio information extracted from the training audio sequence, the target loss of the preset model is determined, and the preset model is trained based on the target loss to obtain the speech synthesis model.
[0085] For example, features can be extracted from the training audio sequence, such as the Mel spectrum extracted from the training audio sequence, as the target audio information. Then, the Mean Square Error (MSE) can be calculated based on this synthesized audio information and the target audio information to obtain the target loss of the preset model. As an example, if the target loss is greater than a loss threshold, the parameters of the preset model can be optimized and updated using the Adam optimizer based on the target loss to achieve training of the preset model.
[0086] Therefore, by combining the phoneme sequence, tone sequence, and prosody sequence corresponding to the training audio sequence for speech synthesis, the accuracy of speech synthesis can be improved to a certain extent, the scope of application of the speech synthesis model can be broadened, and the adaptation and support for multiple pronunciation modes corresponding to laughter can be enhanced.
[0087] Accordingly, generating laughter audio information of the target text based on the target audio information sequence and the speech synthesis model can be achieved by inputting the target audio information sequence into the speech synthesis model, and then processing the speech synthesis model to output audio information.
[0088] In one possible embodiment, the method may further include:
[0089] The audio information of the text to be processed is input into a vocoder to obtain the corresponding speech information, and then the speech information is output. For example, the vocoder can be used to generate a time-domain waveform, i.e., speech information, based on a predicted Mel-spectrum frame sequence. This allows for the further generation of speech information to be output to the user, providing reading convenience and ensuring that the output speech information closely matches the user's actual reading style, thus enhancing the user experience.
[0090] Based on the same inventive concept, this disclosure also provides a speech synthesis device, such as... Figure 4 As shown, the device 10 includes:
[0091] The first determining module 100 is used to determine the target text in the received text to be processed, wherein the target text is text used to represent laughter synthesis.
[0092] The second determining module 200 is used to determine the target audio information sequence corresponding to the target text based on the target text and the text to be processed, wherein the target audio information sequence is determined based on pre-annotated audio information sequences under multiple laughter types, and the audio information sequence under each laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence.
[0093] The generation module 300 is used to generate laughter audio information of the target text based on the target audio information sequence and the speech synthesis model, so as to obtain the audio information of the text to be processed.
[0094] Optionally, the speech synthesis model is obtained in the following way:
[0095] Obtain training samples, wherein each training sample includes a training audio sequence and a labeled sequence of the training audio sequence, the labeled sequence including a training phoneme sequence, a training tone sequence and a training prosody sequence corresponding to the training audio sequence;
[0096] For each training audio sequence, the vectors corresponding to each labeled sequence of the training audio sequence are concatenated to obtain a concatenated vector, and the concatenated vector is input into the encoder of the preset model to obtain the feature vector corresponding to the training audio sequence.
[0097] The feature vector is input into the attention module of the preset model to obtain the context vector corresponding to the feature vector;
[0098] The context vector is input into the decoder of the preset model to obtain the synthesized audio information corresponding to the training audio sequence;
[0099] Based on the synthesized audio information and the target audio information extracted from the training audio sequence, the target loss of the preset model is determined, and the preset model is trained based on the target loss to obtain the speech synthesis model.
[0100] Optionally, the labeled sequence of the training audio sequence in each training sample is determined based on the spectrogram of the training audio and the pre-labeled audio information sequences under various laughter types. The laughter types include laughter types corresponding to the inhalation segment and laughter types corresponding to the exhalation segment. The inhalation segment determined based on the spectrogram is labeled with audio information sequences under the laughter types corresponding to the inhalation segment, and the exhalation segment determined based on the spectrogram is labeled with audio information sequences under the laughter types corresponding to the exhalation segment.
[0101] Optionally, in the target audio information sequence, different phonemes corresponding to the same syllable have the same tone and rhythm, and the tone includes rising tone, level tone and falling tone.
[0102] Optionally, the second determining module includes:
[0103] The first determining submodule is used to determine the context text in the text to be processed that corresponds to the target text;
[0104] The second determining submodule is used to determine the scene type corresponding to the target text based on the context text, wherein each scene type corresponds to at least one candidate audio information sequence, and the candidate audio information sequence includes at least one audio information sequence under the laughter type;
[0105] The generation submodule is used to generate the target audio information sequence corresponding to the target text based on the candidate audio information sequence under the scene type.
[0106] Optionally, the second determining submodule includes:
[0107] The third determining submodule is used to determine the narrator role information corresponding to the context text;
[0108] The fourth determination submodule is used to classify the scene based on the context text and the narrator's role information, and to determine the determined classification as the scene type.
[0109] Optionally, the generation submodule includes:
[0110] The fifth determining submodule is used to determine the phoneme sequence and tone sequence corresponding to each syllable contained in the target text based on the candidate audio information sequence corresponding to the scene type;
[0111] The sixth determining submodule is used to determine the target audio information sequence corresponding to the target text based on the phoneme sequence and tone sequence corresponding to each syllable contained in the target text, as well as the number of syllables in the target text.
[0112] Optionally, the multiple laughter types include laughter types corresponding to the inhalation phase and laughter types corresponding to the exhalation phase.
[0113] Among them, the laughter types in the inspiratory phase include the first laughter type corresponding to vocal cord vibration and the second laughter type corresponding to vocal cord non-vibration;
[0114] The types of laughter during the exhalation phase include vocal laughter corresponding to the vocalization phase and initiation laughter corresponding to the initiation phase. Each syllable in the phoneme sequence of the vocal laughter type is based on consonants and vowels, while each syllable in the phoneme sequence of the initiation laughter type is based on vowels.
[0115] The following is for reference. Figure 5 This diagram illustrates a structural schematic of an electronic device 600 suitable for implementing embodiments of the present disclosure. The terminal devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0116] like Figure 5 As shown, electronic device 600 may include a processing device (e.g., a central processing unit, a graphics processor, etc.) 601, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 602 or a program loaded from storage device 608 into random access memory (RAM) 603. RAM 603 also stores various programs and data required for the operation of electronic device 600. Processing device 601, ROM 602, and RAM 603 are interconnected via bus 604. Input / output (I / O) interface 605 is also connected to bus 604.
[0117] Typically, the following devices can be connected to I / O interface 605: input devices 606 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 607 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 608 including, for example, magnetic tapes, hard disks, etc.; and communication devices 609. Communication device 609 allows electronic device 600 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 5 An electronic device 600 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0118] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 609, or installed from a storage device 608, or installed from a ROM 602. When the computer program is executed by the processing device 601, it performs the functions defined in the methods of embodiments of this disclosure.
[0119] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0120] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any currently known or future-developed networks.
[0121] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0122] The aforementioned computer-readable medium carries one or more programs that, when executed by the electronic device, cause the electronic device to: determine target text in the received text to be processed, wherein the target text is text used to represent laughter synthesis; determine a target audio information sequence corresponding to the target text based on the target text and the text to be processed, wherein the target audio information sequence is determined based on pre-annotated audio information sequences under multiple laughter types, each of the laughter types including a target phoneme sequence, a target tone sequence, and a prosody sequence; and generate laughter audio information of the target text based on the target audio information sequence and a speech synthesis model to obtain the audio information of the text to be processed.
[0123] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0124] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0125] The modules described in the embodiments of this disclosure can be implemented in software or in hardware. The names of the modules do not necessarily limit the module itself; for example, the first determining module can also be described as "a module that determines the target text in the received text to be processed."
[0126] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0127] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0128] According to one or more embodiments of this disclosure, Example 1 provides a speech synthesis method, wherein the method includes:
[0129] Based on the received text to be processed, the target text in the text to be processed is determined, wherein the target text is the text used to represent the synthesis of laughter;
[0130] Based on the target text and the text to be processed, a target audio information sequence corresponding to the target text is determined, wherein the target audio information sequence is determined based on pre-annotated audio information sequences under various laughter types, and the audio information sequence under each laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence;
[0131] Based on the target audio information sequence and the speech synthesis model, laughter audio information of the target text is generated to obtain the audio information of the text to be processed.
[0132] According to one or more embodiments of this disclosure, Example 2 provides the method of Example 1, wherein the speech synthesis model is obtained in the following manner:
[0133] Obtain training samples, wherein each training sample includes a training audio sequence and a labeled sequence of the training audio sequence, the labeled sequence including a training phoneme sequence, a training tone sequence and a training prosody sequence corresponding to the training audio sequence;
[0134] For each training audio sequence, the vectors corresponding to each labeled sequence of the training audio sequence are concatenated to obtain a concatenated vector, and the concatenated vector is input into the encoder of the preset model to obtain the feature vector corresponding to the training audio sequence.
[0135] The feature vector is input into the attention module of the preset model to obtain the context vector corresponding to the feature vector;
[0136] The context vector is input into the decoder of the preset model to obtain the synthesized audio information corresponding to the training audio sequence;
[0137] Based on the synthesized audio information and the target audio information extracted from the training audio sequence, the target loss of the preset model is determined, and the preset model is trained based on the target loss to obtain the speech synthesis model.
[0138] According to one or more embodiments of this disclosure, Example 3 provides the method of Example 2, wherein the labeled sequence of the training audio sequence in each training sample is determined based on the spectrogram of the training audio and the pre-labeled audio information sequences under multiple laughter types, the laughter types including laughter types corresponding to the inhalation segment and laughter types corresponding to the exhalation segment, wherein the inhalation segment determined based on the spectrogram is labeled with the audio information sequence corresponding to the laughter type of the inhalation segment, and the exhalation segment determined based on the spectrogram is labeled with the audio information sequence corresponding to the laughter type of the exhalation segment.
[0139] According to one or more embodiments of this disclosure, Example 4 provides the method of Example 1, wherein in the target audio information sequence, different phonemes corresponding to the same syllable have the same tone and rhythm, the tone including rising tone, level tone and falling tone.
[0140] According to one or more embodiments of this disclosure, Example 5 provides the method of Example 1, wherein,
[0141] The step of determining the target audio information sequence corresponding to the target text based on the target text and the text to be processed includes:
[0142] Determine the context text in the text to be processed that corresponds to the target text;
[0143] The scene type corresponding to the target text is determined based on the context text, wherein each scene type corresponds to at least one candidate audio information sequence, and the candidate audio information sequence includes at least one audio information sequence under the laughter type;
[0144] Generate the target audio information sequence corresponding to the target text based on the candidate audio information sequence under the scene type.
[0145] According to one or more embodiments of this disclosure, Example 6 provides the method of Example 5, wherein determining the scene type corresponding to the target text based on the context text includes:
[0146] Determine the narrator role information corresponding to the context text;
[0147] The scene is classified based on the context text and the narrator's role information, and the determined classification is identified as the scene type.
[0148] According to one or more embodiments of this disclosure, Example 7 provides the method of Example 5, wherein generating the target audio information sequence corresponding to the target text based on the candidate audio information sequence under the scene type includes:
[0149] Based on the candidate audio information sequence corresponding to the scene type, determine the phoneme sequence and tone sequence corresponding to each syllable contained in the target text;
[0150] The target audio information sequence corresponding to the target text is determined based on the phoneme sequence and tone sequence corresponding to each syllable contained in the target text, as well as the number of syllables in the target text.
[0151] According to one or more embodiments of this disclosure, Example 8 provides the method of Example 1, wherein the multiple laughter types include laughter types corresponding to the inhalation phase and laughter types corresponding to the exhalation phase.
[0152] Among them, the laughter types in the inspiratory phase include the first laughter type corresponding to vocal cord vibration and the second laughter type corresponding to vocal cord non-vibration;
[0153] The types of laughter during the exhalation phase include vocal laughter corresponding to the vocalization phase and initiation laughter corresponding to the initiation phase. Each syllable in the phoneme sequence of the vocal laughter type is based on consonants and vowels, while each syllable in the phoneme sequence of the initiation laughter type is based on vowels.
[0154] According to one or more embodiments of this disclosure, Example 9 provides a speech synthesis apparatus, wherein the apparatus includes:
[0155] The first determining module is used to determine the target text in the received text to be processed, wherein the target text is text used to represent laughter synthesis.
[0156] The second determining module is used to determine the target audio information sequence corresponding to the target text based on the target text and the text to be processed. The target audio information sequence is determined based on audio information sequences under pre-annotated multiple laughter types. Each audio information sequence under the laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence.
[0157] The generation module is used to generate laughter audio information of the target text based on the target audio information sequence and the speech synthesis model, so as to obtain the audio information of the text to be processed.
[0158] According to one or more embodiments of the present disclosure, Example 10 provides a computer-readable medium having a computer program stored thereon, wherein the program, when executed by a processing device, implements the steps of the method described in any of Examples 1-8.
[0159] According to one or more embodiments of this disclosure, Example 11 provides an electronic device, which includes:
[0160] A storage device on which one or more computer programs are stored;
[0161] One or more processing means are configured to execute the one or more computer programs in the storage device to implement the steps of the method described in any of the examples 1-8.
[0162] The above description is merely a preferred embodiment of this disclosure and an explanation of the technical principles employed. Those skilled in the art should understand that the scope of this disclosure is not limited to technical solutions formed by specific combinations of the above-described technical features, but should also cover other technical solutions formed by arbitrary combinations of the above-described technical features or their equivalents without departing from the above-described concept. For example, technical solutions formed by substituting the above features with (but not limited to) technical features disclosed in this disclosure that have similar functions.
[0163] Furthermore, while the operations are described in a specific order, this should not be construed as requiring these operations to be performed in the specific order shown or in a sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, while several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of this disclosure. Certain features described in the context of individual embodiments may also be implemented in combination in a single embodiment. Conversely, various features described in the context of a single embodiment may also be implemented individually or in any suitable sub-combination in multiple embodiments.
[0164] Although the subject matter has been described using language specific to structural features and / or methodological logic, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or actions described above. Rather, the specific features and actions described above are merely illustrative forms of implementing the claims. Regarding the apparatus in the above embodiments, the specific manner in which the various modules perform their operations has been described in detail in the embodiments relating to the method, and will not be elaborated upon here.
Claims
1. A speech synthesis method, characterized in that, The method includes: Based on the received text to be processed, the target text in the text to be processed is determined, wherein the target text is the text used to represent the synthesis of laughter; Based on the target text and the text to be processed, a target audio information sequence corresponding to the target text is determined, wherein the target audio information sequence is determined based on pre-annotated audio information sequences under various laughter types, and the audio information sequence under each laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence; Based on the target audio information sequence and the speech synthesis model, laughter audio information of the target text is generated to obtain the audio information of the text to be processed; The step of determining the target audio information sequence corresponding to the target text based on the target text and the text to be processed includes: Determine the context text in the text to be processed that corresponds to the target text; The scene type corresponding to the target text is determined based on the context text, wherein each scene type corresponds to at least one candidate audio information sequence, and the candidate audio information sequence includes at least one audio information sequence under the laughter type; Generate the target audio information sequence corresponding to the target text based on the candidate audio information sequence under the scene type.
2. The method according to claim 1, characterized in that, The speech synthesis model is obtained through the following methods: Obtain training samples, wherein each training sample includes a training audio sequence and a labeled sequence of the training audio sequence, the labeled sequence including a training phoneme sequence, a training tone sequence and a training prosody sequence corresponding to the training audio sequence; For each training audio sequence, the vectors corresponding to each labeled sequence of the training audio sequence are concatenated to obtain a concatenated vector, and the concatenated vector is input into the encoder of the preset model to obtain the feature vector corresponding to the training audio sequence. The feature vector is input into the attention module of the preset model to obtain the context vector corresponding to the feature vector; The context vector is input into the decoder of the preset model to obtain the synthesized audio information corresponding to the training audio sequence; Based on the synthesized audio information and the target audio information extracted from the training audio sequence, the target loss of the preset model is determined, and the preset model is trained based on the target loss to obtain the speech synthesis model.
3. The method according to claim 2, characterized in that, The labeled sequence of the training audio sequence in each training sample is determined based on the spectrogram of the training audio and the pre-labeled audio information sequences under various laughter types. The laughter types include laughter types corresponding to the inhalation segment and laughter types corresponding to the exhalation segment. The inhalation segment determined based on the spectrogram is labeled with audio information sequences corresponding to the laughter types of the inhalation segment, and the exhalation segment determined based on the spectrogram is labeled with audio information sequences corresponding to the laughter types of the exhalation segment.
4. The method according to claim 1, characterized in that, In the target audio information sequence, different phonemes corresponding to the same syllable have the same tone and rhythm, and the tone includes rising tone, level tone and falling tone.
5. The method according to claim 1, characterized in that, Determining the scene type corresponding to the target text based on the context text includes: Determine the narrator role information corresponding to the context text; The scene is classified based on the context text and the narrator's role information, and the determined classification is identified as the scene type.
6. The method according to claim 1, characterized in that, The step of generating the target audio information sequence corresponding to the target text based on the candidate audio information sequence under the scene type includes: Based on the candidate audio information sequence corresponding to the scene type, determine the phoneme sequence and tone sequence corresponding to each syllable contained in the target text; The target audio information sequence corresponding to the target text is determined based on the phoneme sequence and tone sequence corresponding to each syllable contained in the target text, as well as the number of syllables in the target text.
7. The method according to claim 1, characterized in that, The various laughter types include laughter types corresponding to the inhalation phase and laughter types corresponding to the exhalation phase. Among them, the laughter types in the inspiratory phase include the first laughter type corresponding to vocal cord vibration and the second laughter type corresponding to vocal cord non-vibration; The types of laughter during the exhalation phase include vocal laughter corresponding to the vocalization phase and initiation laughter corresponding to the initiation phase. Each syllable in the phoneme sequence of the vocal laughter type is based on consonants and vowels, while each syllable in the phoneme sequence of the initiation laughter type is based on vowels.
8. A speech synthesis device, characterized in that, The device includes: The first determining module is used to determine the target text in the received text to be processed, wherein the target text is the text used to represent the synthesis of laughter. The second determining module is used to determine the target audio information sequence corresponding to the target text based on the target text and the text to be processed. The target audio information sequence is determined based on audio information sequences under pre-annotated multiple laughter types. Each audio information sequence under the laughter type includes a target phoneme sequence, a target tone sequence, and a prosody sequence. The generation module is used to generate laughter audio information of the target text based on the target audio information sequence and the speech synthesis model, so as to obtain the audio information of the text to be processed; The second determining module includes: The first determining submodule is used to determine the context text in the text to be processed that corresponds to the target text; The second determining submodule is used to determine the scene type corresponding to the target text based on the context text, wherein each scene type corresponds to at least one candidate audio information sequence, and the candidate audio information sequence includes at least one audio information sequence under the laughter type; The generation submodule is used to generate the target audio information sequence corresponding to the target text based on the candidate audio information sequence under the scene type.
9. A computer-readable medium having a computer program stored thereon, characterized in that, When executed by the processing device, the program implements the steps of the method described in any one of claims 1-7.
10. An electronic device, characterized in that, include: A storage device on which one or more computer programs are stored; One or more processing means are configured to execute the one or more computer programs in the storage device to implement the steps of the method according to any one of claims 1-7.
Citation Information
Patent Citations
Voice translation method and device
CN101727904A
Speech synthesis method and related equipment
CN108962217A
Speech synthesis method and device, electronic equipment and readable storage medium
CN112270920A
Pitch model production device, method and pitch model production program
CN1664922A