Audio synthesis method, system, electronic device, and computer-readable storage medium

By inserting colloquial content and pauses into the text and combining them with prosodic features to convert it into audio, the problem of insufficient anthropomorphism in audio synthesis is solved, achieving a more natural and intimate audio synthesis effect.

CN118135988BActive Publication Date: 2026-04-10IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-12-14
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing human-computer interaction processes, the degree of anthropomorphism in audio synthesis is insufficient, resulting in audio synthesis expressions that are not natural or vivid enough.

Method used

By inserting colloquial expressions and pauses into the text to be processed, the prosodic features of the colloquial text are obtained. Based on the prosodic features and pauses, the text is converted into target audio, thereby improving the anthropomorphism of the audio synthesis.

Benefits of technology

By adding conversational content and more natural pauses, the anthropomorphism of the audio synthesis is enhanced, making the audio more human and approachable during playback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118135988B_ABST
    Figure CN118135988B_ABST
Patent Text Reader

Abstract

The application discloses an audio synthesis method, system, electronic equipment and computer readable storage medium. The method comprises the following steps: in response to obtaining a text to be processed, inserting a spoken expression content into the text to be processed to obtain a spoken text; wherein the spoken expression content at least comprises a spoken new content and a spoken pause interval; obtaining a prosody feature of the spoken text, and obtaining a prosody pause interval of the spoken text based on the prosody feature; and converting the spoken text into a target audio based on the spoken pause interval and the prosody pause interval. Through the above manner, the humanization degree of audio synthesis can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of data processing, in particular to an audio synthesis method and system, an electronic device and a computer readable storage medium. BACKGROUND

[0002] With the development of intelligent devices, more and more intelligent devices can support human-computer interaction. During human-computer interaction, text data is usually synthesized into audio for feedback. However, in the existing human-computer interaction process, the text data is usually only converted into stiff audio, resulting in insufficient humanization of audio synthesis. Therefore, how to improve the humanization of audio synthesis has become a problem to be solved. SUMMARY

[0003] The technical problem solved by the present application is to provide an audio synthesis method, system, electronic device and computer readable storage medium, which can improve the humanization of audio synthesis.

[0004] To solve the above technical problem, the first aspect of the present application provides an audio synthesis method, comprising: in response to obtaining a to-be-processed text, inserting oral expression content into the to-be-processed text to obtain an oral text; wherein the oral expression content at least includes oral new content and oral pause interval; obtaining prosodic features of the oral text, and obtaining prosodic pause intervals of the oral text based on the prosodic features; and converting the oral text into target audio based on the oral pause interval and the prosodic pause interval.

[0005] To solve the above technical problem, the second aspect of the present application provides an audio synthesis system, comprising: a conversion module, an analysis module and a generation module. In response to obtaining a to-be-processed text, the conversion module is used to insert oral expression content into the to-be-processed text to obtain an oral text; wherein the oral expression content at least includes oral new content and oral pause interval; the analysis module is used to obtain prosodic features of the oral text, and obtain prosodic pause intervals of the oral text based on the prosodic features; and the generation module is used to convert the oral text into target audio based on the oral pause interval and the prosodic pause interval.

[0006] To solve the above technical problem, the third aspect of the present application provides an electronic device, which comprises a memory and a processor coupled to each other. The memory stores program instructions, and the processor is configured to execute the program instructions to implement the audio synthesis method of the first aspect.

[0007] To solve the above technical problems, the fourth aspect of the present application provides a computer readable storage medium, which stores program instructions, and the program instructions are executed by a processor to implement the audio synthesis method of the first aspect.

[0008] The above scheme, when obtaining the to-be-processed text, inserts the oral expression content into the to-be-processed text, converts the to-be-processed text into the oral text, wherein the oral expression content at least includes the oral new content and the oral pause interval, therefore, the oral text can include more contents that increase when oral expression and the pause that is adopted when oral expression, obtains the prosody feature of the converted oral text, makes the oral new content able to affect the prosody feature, thereby obtaining the prosody pause interval corresponding to the oral text based on the prosody feature, makes the prosody pause interval more oral, based on the oral pause interval and the prosody pause interval, performs audio synthesis on the oral text including the oral new content, converts the oral text into the target audio, makes the target audio include the oral new content and the pause interval more oral when playing, improves the humanization degree of the audio synthesis. BRIEF DESCRIPTION OF DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor. Among them:

[0010] Figure 1 is a flowchart of an embodiment of the audio synthesis method of the present application;

[0011] Figure 2 is a flowchart of another embodiment of the audio synthesis method of the present application;

[0012] Figure 3 is a structural schematic diagram of an embodiment of the audio synthesis system of the present application;

[0013] Figure 4 is a structural schematic diagram of an embodiment of the electronic device of the present application;

[0014] Figure 5 is a structural schematic diagram of an embodiment of the computer readable storage medium of the present application. DETAILED DESCRIPTION

[0015] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0016] In this paper, the terms "system" and "network" are often used interchangeably. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. Additionally, the character " / " generally indicates that the preceding and following related objects have an "or" relationship. Furthermore, "many" in this paper means two or more.

[0017] The audio synthesis method provided in this application is used to synthesize audio with a high degree of anthropomorphism. Its execution subject is a processor capable of processing text and audio. The processor can be a smart terminal, such as a smart terminal capable of human-computer interaction with users. This application does not impose specific restrictions on this.

[0018] Please see Figure 1 , Figure 1 This is a flowchart illustrating one embodiment of the audio synthesis method of this application, which includes:

[0019] S101: In response to obtaining the text to be processed, insert colloquial expressions into the text to be processed to obtain colloquial text, wherein the colloquial expressions include at least the newly added colloquial content and the colloquial pause interval.

[0020] Specifically, when the text to be processed is obtained, colloquial expressions are inserted into the text to convert it into colloquial text. The colloquial expressions include at least the newly added colloquial content and colloquial pause intervals.

[0021] It is understandable that conversational texts can include more conversational expressions, as well as the pauses that are used in conversational expressions.

[0022] In one embodiment, when the text to be processed is obtained, it is input into a pre-trained conversion model, which then inserts colloquial expressions into the text, resulting in colloquial text output by the conversion model. The conversion model is trained using written text and its corresponding colloquial expressions.

[0023] In one embodiment, when the text to be processed is obtained, a first prompt text is constructed based on the text to be processed, and the first prompt text is input into the intelligent analysis model to obtain the conversational text fed back by the intelligent analysis model. The first prompt text is used to prompt the intelligent analysis model to perform conversational conversion.

[0024] Optionally, the intelligent analysis model is a large language model, which may include, but is not limited to, deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTM), and generative pre-trained Transformer models. No specific restrictions are placed on the specific construction and deployment of the large language model.

[0025] Furthermore, the colloquial expression content includes newly added colloquial content and colloquial pause intervals. The newly added colloquial content includes at least various colloquial expressions, and the colloquial pause intervals include pause intervals of various lengths. In addition, the colloquial expression content may also include colloquial adjustment content, which is used to replace parts of the text to be processed, replacing more formal expressions with colloquial adjustment content.

[0026] It should be noted that the colloquial expressions include at least modal particles, filler words, and repetitive words, and the onomatopoeic audio includes at least laughter, inhalation sounds, and smacking sounds.

[0027] S102: Obtain the prosodic features of the colloquial text and obtain the prosodic pause intervals of the colloquial text based on the prosodic features.

[0028] Specifically, the prosodic features of the converted colloquial text are obtained, so that the newly added colloquial content can affect the prosodic features. Based on the prosodic features, the corresponding prosodic pause intervals of the colloquial text are obtained, making the prosodic pause intervals more colloquial.

[0029] In one embodiment, preset characters in the colloquial text are converted into matching text forms, and the adjusted text is divided into basic units, wherein the preset characters include at least numbers, abbreviations and currency symbols, and the basic units include at least words and punctuation marks. The prosodic features of the colloquial text are determined based on the divided text, and the colloquial text is segmented based on the prosodic features to obtain the prosodic pause intervals of the colloquial text.

[0030] In an embodiment, syntax, discourse structure and information structure of the oralized text are obtained, prosodic features of the oralized text are determined based on the syntax, discourse structure and information structure, the oralized text is segmented based on the prosodic features, and prosodic pause intervals of the oralized text are obtained.

[0031] It can be understood that the oralized pause intervals inserted in the oralized text are irrelevant to the text content, and the oralized pause intervals can be shielded when the prosodic features of the oralized text are obtained, but the oralized new content has been added in the oralized text, which can affect the prosodic features in the process of extracting the prosodic features, so that the prosodic features are more inclined to the oralized expression mode.

[0032] In an implementation scenario, the oralized text is input into a text analysis model to obtain prosodic pause intervals fed back by the text analysis model. The text analysis model is trained by using a plurality of written texts and oralized expression texts corresponding to at least part of the written texts.

[0033] S103: converting the oralized text into target audio based on the oralized pause intervals and the prosodic pause intervals.

[0034] Specifically, the oralized text including the oralized new content is synthesized into audio based on the oralized pause intervals and the prosodic pause intervals, and the oralized text is converted into target audio, so that the target audio can include the oralized new content and the pause intervals are more oralized when playing.

[0035] In an embodiment, the oralized pause intervals and the prosodic pause intervals are fused to obtain text pause intervals of the oralized text, phoneme sequences corresponding to the oralized text are obtained based on the text pause intervals and words in the oralized text, and the phoneme sequences are converted into target audio corresponding to the oralized text.

[0036] In an embodiment, the oralized text is divided into a plurality of fields based on the oralized pause intervals and the prosodic pause intervals, phoneme sequences corresponding to the plurality of fields are determined, and the phoneme sequences are converted into target audio corresponding to the oralized text.

[0037] In an implementation scenario, the phoneme sequences are input into an end-to-end audio conversion model to obtain target audio output by the audio conversion model. The audio conversion model is trained by using phoneme sequence samples and conversion audios corresponding to the phoneme sequence samples.

[0038] In an application scenario, the text to be processed is a reply text generated by a user when the user performs human-computer interaction with an intelligent terminal to reply to the user, and the text to be processed is converted into target audio with a higher degree of personification after oralization and output, so that the user feels more real and kind when performing human-computer interaction.

[0039] In another application scenario, the to-be-processed text is a text obtained from user input information, the user input information includes at least one of text, voice and image, the user input information is finally converted into the to-be-processed text, and the to-be-processed text is converted into the target audio with a higher degree of personification after being converted into oral language and output, so that the original expression in the to-be-processed text is adjusted into an oral language expression and converted into playable audio, for example, the originally stiff written expression of the to-be-processed text is converted into the target audio and played, thereby improving the acceptability when a user listens.

[0040] Optionally, the to-be-processed text can be a written knowledge text of a student, and after the to-be-processed text is converted into the target audio, the student can listen to the converted target audio to better learn the knowledge points. The to-be-processed text can also be a story text of a child, and after the to-be-processed text is converted into the target audio, the child can listen to the converted target audio to improve the sense of closeness of story listening. In addition, the to-be-processed text in other specific application scenarios is not enumerated here.

[0041] The above scheme, when the to-be-processed text is obtained, inserts oral language expression content in the to-be-processed text, converts the to-be-processed text into oral language text, wherein the oral language expression content at least includes oral language newly added content and oral language pause interval, therefore, the oral language text can contain more content that will be increased when oral language expression is performed, and the pause that will be adopted when oral language expression is performed, the prosodic features of the converted oral language text are obtained, so that the oral language newly added content can affect the prosodic features, so that the prosodic pause interval corresponding to the oral language text is obtained based on the prosodic features, so that the prosodic pause interval is also more oral language, and based on the oral language pause interval and the prosodic pause interval, the oral language text including the oral language newly added content is synthesized into audio, so that the oral language text is converted into the target audio, so that the target audio can include the oral language newly added content and the pause interval is also more oral language when playing, and the degree of personification of audio synthesis is improved.

[0042] Please refer to Figure 2 , Figure 2 is a flowchart of another embodiment of the audio synthesis method of the present application, which comprises:

[0043] S201: in response to obtaining the to-be-processed text, inputting the to-be-processed text into a text conversion model to obtain oral language text fed back by the text conversion model, wherein the text conversion model is trained by using parallel text pairs, the parallel text pairs include first training text and second training text, and the second training text inserts oral language expression content compared with the first training text.

[0044] Specifically, when the to-be-processed text is obtained, the to-be-processed text is input into the text conversion model to obtain a colloquial text fed back by the text conversion model, so as to convert the to-be-processed text into the colloquial text, wherein the text conversion model is obtained by pre-training using a parallel text pair, and the parallel text pair includes a first training text and a corresponding second training text.

[0045] It should be noted that the second training text is inserted with colloquial expression content compared with the first training text, and the colloquial expression content at least includes colloquial new content and colloquial pause interval, so that the expression of the second training text is more colloquial than that of the first training text, and the inserted colloquial expression content in the second training text can be used as a training label to supervise the training of the text conversion model.

[0046] Further, the text conversion model can be used to efficiently insert colloquial expression content in the to-be-processed text, thereby effectively converting the to-be-processed text into a colloquial text.

[0047] It can be understood that the first training text included in the parallel text pair is usually written expression, and part of the first training text has partial colloquial characteristics, and the second training text is colloquial expression generated by inserting colloquial expression content in the first training text. The text conversion model trained using the parallel text pair can convert the written expression text into the colloquial expression text, and convert the text with partial colloquial characteristics into the completely colloquial expression text.

[0048] It should be noted that the colloquial new content at least includes at least one of an interjection, a filler, a repetition and a pseudo-sound, and the colloquial pause interval includes pause intervals of multiple lengths; wherein each type of colloquial new content is matched with a new start tag before it, and each type of colloquial new content is matched with a new end tag after it, and each type of colloquial pause interval is matched with a pause tag.

[0049] Specifically, in addition to the interjection, the filler, the repetition and the pseudo-sound, the colloquial new content can also include other new content such as a baby sound, and the present application does not make specific limitations thereon. The pseudo-sound can include laughter, inhalation, coughing and smacking, and the present application does not make specific limitations thereon.

[0050] Further, when the colloquial new content is inserted in the text, each type of colloquial new content is inserted with a matching new start tag before it, and each type of colloquial new content is inserted with a new end tag after it, for example, the new start tag of the filler is fib, the new start tag of the filler is fie, the new start tag of the interjection is yqb, the new start tag of the filler is yqe, and other specific tags will not be listed one by one.

[0051] It can be understood that the oral pause interval at least includes the stuttering, short pause, long pause and super long pause, a plurality of pause intervals of time length, and each oral pause interval corresponds to a matching pause label, for example, the pause label of the short pause is spa, the pause label of the long pause is lpa, and other specific labels will not be listed one by one in this application.

[0052] For convenience of description, taking a parallel text pair as an example, the first training sample is in written style and specifically includes: It is basically bean milk and fried dough sticks for breakfast. The second training sample is in oral style and includes: <fib> um <fie> It is basically <spa> bean milk and fried dough sticks for breakfast. Therefore, the newly added starting label and the newly added ending label can distinguish and locate the newly added oral content, the pause label can distinguish and locate the oral pause content, and the text conversion model can complete the training through rule definition or adaptive learning, so that the trained text conversion model can insert oral expression content in the to-be-processed text and set a matching label at the corresponding position.

[0053] In an embodiment, the text conversion model includes an encoder and a decoder, and the first training text passes through the encoder and the decoder to obtain the first predicted text. The oral expression content inserted in the second training text is used as the training label of the text conversion model.

[0054] Specifically, the training process of the text conversion model includes: inputting the first training text into the encoder to obtain the first encoded text output by the encoder, inputting the first encoded text into the decoder to obtain the first predicted text output by the decoder, comparing the first predicted text with the second training text, determining the first training loss based on the oral expression content included in the second training text and the predicted expression content inserted in the first predicted text, adjusting the parameters of the encoder and the decoder based on the first training loss, until the first convergence condition is met, and obtaining the trained text conversion model. Therefore, the text conversion model can be a model constructed based on the encoder and the decoder, and obtained through supervised training, so that the trained text conversion model can complete the oral conversion of the to-be-processed text and improve the controllability of the oral conversion process.

[0055] In another embodiment, the text conversion model includes a prompt construction unit and an intelligent analysis model, the prompt construction unit is used to construct a prompt text based on at least the first training text and input the intelligent analysis model, the prompt text is used to prompt the intelligent analysis model to output the second predicted text, and when the prompt text still includes part of the parallel text pair, the parallel text pair is used as an example sample. The oral expression content inserted in the second training text is used as the training label of the text conversion model.

[0056] Specifically, the prompt construction unit in the text conversion model is configured to construct a prompt text, the prompt text including at least the first training text and a requirement for oralization conversion, the prompt text is fed to the intelligent analysis model, the second predicted text is fed back by the intelligent analysis model, the second predicted text is compared with the second training text, the second training loss is determined based on the oralization expression content included in the second training text and the predicted expression content inserted in the second predicted text, the intelligent analysis model is fine-tuned based on the second training loss, and the fine-tuned intelligent analysis model is obtained until the second convergence condition is met. The trained text conversion model. Therefore, the text conversion model can be obtained by fine-tuning the intelligent analysis model through supervised training, so that the trained text conversion model can complete oralization conversion of the to-be-processed text and reduce the consumption of the terminal for oralization conversion.

[0057] In an implementation scenario, the fine-tuning of the intelligent analysis model can adopt a zero-sample way to construct a prompt text. For example, the first training sample is a written style, which specifically includes: like we had soy milk and fried dough sticks for breakfast. The prompt text constructed by the zero-sample way includes: please convert the text {like we had soy milk and fried dough sticks for breakfast} to oralization and set the label at the corresponding position.

[0058] In another implementation scenario, the fine-tuning of the intelligent analysis model can adopt a few-sample way to construct a prompt text. The prompt text constructed by the few-sample way includes: the written expression text is super partial to dips for southern cuisine. The output transcribed oral expression text is <fib> um <fie> southern cuisine, <fib> er <fie> <std> super partial to dips <yqb> ah <yqe> what. Please transcribe the following written text according to the above instructions: {like we had soy milk and fried dough sticks for breakfast}. Therefore, when the prompt text further includes a part of the parallel text pair, the parallel text pair serves as an example sample, which can improve the learning accuracy and convergence efficiency of the intelligent analysis model.

[0059] S202: Obtain the prosodic feature of the oralization text, and obtain the prosodic pause interval of the oralization text based on the prosodic feature.

[0060] Specifically, the prosodic feature of the oralization text is extracted, and the prosodic pause interval of the oralization text is determined by using the prosodic feature.

[0061] In an implementation, the oralization text is input into a text analysis model to obtain the prosodic pause interval fed back by the text analysis model; wherein the text analysis model is configured to perform text regularization on the oralization text and perform word segmentation on the regularized oralization text, determine the prosodic feature based on the segmented oralization text, and obtain the prosodic pause interval of the oralization text based on the prosodic feature. The text analysis model is trained by using the first training text.

[0062] Specifically, the conversion from the text to be processed to the target audio is divided into two stages, the first stage is the oralization conversion of the text to be processed by the text conversion model to obtain the oralized text, and the second stage is the analysis of the oralized text by the text analysis model to obtain the prosodic pause interval and complete the audio conversion subsequently.

[0063] Further, the text analysis model is trained by the first training text, and the text analysis model focuses on the mining of prosodic features and the ability to determine the prosodic pause interval based on the prosodic features. The trained text analysis model can regularize the oralized text and segment the regularized oralized text. Based on the segmented oralized text, the prosodic features are determined, and the prosodic pause interval of the oralized text is obtained based on the prosodic features.

[0064] It should be noted that the preset characters in the oralized text are converted to matching text forms during text regularization, and the adjusted text is divided into basic units during segmentation. The preset characters include at least numbers, abbreviations, and currency symbols, and the basic units include at least words and punctuation marks. The prosodic features of the oralized text are determined based on the divided text, and the oralized text is separated based on the prosodic features to obtain the prosodic pause interval of the oralized text, thereby improving the accuracy of the prosodic pause interval.

[0065] S203: Convert the oralized text to the target audio based on the oralized pause interval and the prosodic pause interval.

[0066] Specifically, based on the oralized pause interval and the prosodic pause interval, the oralized text including the oralized new content is synthesized into audio to convert the oralized text into the target audio.

[0067] In an embodiment, based on the oralized pause interval and the prosodic pause interval, the text pause interval of the oralized text is determined; based on the oralized text and its corresponding text pause interval, the phoneme sequence corresponding to the oralized text is obtained; and based on the phoneme sequence, the target audio corresponding to the oralized text is obtained.

[0068] Specifically, the oralized pause interval and the prosodic pause interval are fused to obtain the text pause interval corresponding to the oralized text, so that the text pause interval combines the oralization features and the prosodic features, improves the accuracy of the pause in the oralized text, divides the oralized text using the text pause interval to obtain the segmented words, converts the words into phoneme sequences, i.e., represents pronunciation with phonemes, and thus performs audio conversion based on the phoneme sequence to obtain the target audio corresponding to the oralized text, so that the target audio has a high degree of personification.

[0069] In an implementation scenario, the acoustic features are obtained from the phoneme sequence, where the acoustic features at least include a fundamental frequency and spectral features, the acoustic features are converted into continuous speech sample audio signals to obtain the target audio corresponding to the spoken text.

[0070] In a specific implementation scenario, the mapping relationship between the phoneme sequence and the acoustic features is obtained by a feature extraction model, and the feature extraction model can include at least one of a recurrent neural network (RNN), a long short-term memory (LSTM), and a generative pre-training transformer. The present application does not make specific limitations on the algorithm used by the vocoder to convert the acoustic features to the target audio.

[0071] Alternatively, the conversion of the phoneme sequence to the target audio can also be obtained by using an end-to-end audio conversion model, where the audio conversion model is trained using phoneme sequence samples and their corresponding converted audio.

[0072] In an implementation scenario, a plurality of prosodic pause intervals correspond to a plurality of interval duration levels, and a plurality of spoken pause intervals match at least part of the interval duration levels; based on the spoken pause intervals and the prosodic pause intervals, the text pause intervals of the spoken text are determined, including: based on the interval duration level corresponding to the spoken pause interval and the position in the spoken text, and the interval duration level corresponding to the prosodic pause interval and the position in the spoken text, the text pause intervals of the spoken text are obtained.

[0073] Specifically, a plurality of prosodic pause intervals match respective interval duration levels, and a plurality of spoken pause intervals can match at least part of the interval duration levels, the position of the spoken pause interval in the spoken text and the interval duration level matched by the spoken pause interval are labeled, the position of the prosodic pause interval in the spoken text and the interval duration level matched by the prosodic pause interval are labeled, the text pause intervals of the spoken text are obtained, the spoken pause interval and the prosodic pause interval are fused with each other, the text pause intervals with uniform levels are obtained, and the target audio conversion has higher accuracy.

[0074] In a specific application scenario, the plurality of prosodic pause intervals at least include five levels of interval duration levels, wherein the interval duration levels at least include word level L1, phrase level L2, prosodic pause L3, punctuation L4 and paragraph pause L5 from short to long, and the spoken pause intervals include short pause, long pause and super-long pause, wherein the short pause corresponds to L2, the long pause corresponds to L3, and the super-long pause corresponds to L4. In other specific application scenarios, the number of interval duration levels and the specific interval duration levels matched by the spoken pause intervals can be customized, and the present application does not make specific limitations.

[0075] It should be noted that the implementation or implementation scenario described in any of the above embodiments is not limited to a single embodiment, and different implementation manners can be combined with each other, and the present application does not make specific limitations.

[0076] Different from the above embodiments, the conversion from the to-be-processed text to the target audio is divided into two stages, the first stage is the spoken conversion of the to-be-processed text by the text conversion model to obtain the spoken text, and the second stage is the analysis of the spoken text by the text analysis model to obtain the prosodic pause interval and complete the audio conversion subsequently, wherein the construction and training of the text conversion model includes a plurality of implementation manners, and the trained text conversion model can insert spoken expression content in the to-be-processed text and set a matching label at the corresponding position, and then fuse the spoken pause interval and the prosodic pause interval to obtain the phoneme sequence, i.e., to express pronunciation with phonemes, so as to perform audio conversion based on the phoneme sequence to obtain the target audio corresponding to the spoken text, so that the target audio has a higher degree of personification.

[0077] Please refer to Figure 3 , Figure 3 is a structural schematic diagram of an embodiment of the audio synthesis system of the present application. The audio synthesis system 30 includes a conversion module 301, an analysis module 302 and a generation module 303. In response to obtaining the to-be-processed text, the conversion module 301 is used to insert spoken expression content in the to-be-processed text to obtain the spoken text; wherein the spoken expression content at least includes spoken new content and spoken pause interval; the analysis module 302 is used to obtain the prosodic features of the spoken text, and obtain the prosodic pause interval of the spoken text based on the prosodic features; the generation module 303 is used to convert the spoken text into the target audio based on the spoken pause interval and the prosodic pause interval.

[0078] The above scheme, when the to-be-processed text is obtained, the conversion module 301 inserts the oral expression content into the to-be-processed text, converts the to-be-processed text into the oral text, wherein the oral expression content at least includes oral new content and oral pause interval, therefore, the oral text can contain more content that will be increased when oral expression occurs, and the pause that will be adopted when oral expression occurs, the analysis module 302 obtains the prosody feature of the converted oral text, so that the oral new content can affect the prosody feature, thereby obtaining the prosody pause interval corresponding to the oral text based on the prosody feature, so that the prosody pause interval is also more oral, and the generation module 303 performs audio synthesis on the oral text including the oral new content based on the oral pause interval and the prosody pause interval, converts the oral text into the target audio, so that the target audio can include the oral new content and the pause interval is also more oral when playing, and the degree of personification of audio synthesis is improved.

[0079] Optionally, the conversion module 301 is further configured to input the to-be-processed text into a text conversion model to obtain oral text fed back by the text conversion model; wherein the text conversion model is trained by using parallel text pairs, and each parallel text pair includes a first training text and a second training text, and the second training text is inserted with oral expression content compared with the first training text.

[0080] Optionally, the text conversion model includes an encoder and a decoder, and the first training text is input into the encoder and the decoder to obtain a first predicted text; or the text conversion model includes a prompt construction unit and an intelligent analysis model, the prompt construction unit is configured to construct a prompt text based on at least the first training text and input the prompt text into the intelligent analysis model, the prompt text is used to prompt the intelligent analysis model to perform oral conversion and output a second predicted text, and when the prompt text further includes part of the parallel text pair, the parallel text pair is used as an example sample; wherein the oral expression content inserted in the second training text is used as a training label of the text conversion model.

[0081] Optionally, the oral new content at least includes at least one of an emotional word, a filler word, a repeated word and a simulated sound audio, and the oral pause interval includes pause intervals of multiple time lengths; wherein each oral new content is matched with a new start label, each oral new content is matched with a new end label, and each oral pause interval is matched with a pause label.

[0082] Optionally, the analysis module 302 is further configured to input the spoken text into a text analysis model to obtain prosodic pause intervals fed back by the text analysis model; the text analysis model is configured to perform text normalization on the spoken text, perform word segmentation on the normalized spoken text, determine prosodic features based on the segmented spoken text, and thus obtain prosodic pause intervals of the spoken text based on the prosodic features; and the text analysis model is trained based on the first training text.

[0083] Optionally, the generation module 303 is further configured to determine text pause intervals of the spoken text based on the spoken pause intervals and the prosodic pause intervals; obtain a phoneme sequence corresponding to the spoken text based on the spoken text and the text pause intervals corresponding to the spoken text; and obtain target audio corresponding to the spoken text based on the phoneme sequence.

[0084] Optionally, the prosodic pause intervals correspond to multiple interval duration levels, and the multiple spoken pause intervals match at least part of the interval duration levels; and the generation module 303 is further configured to obtain the text pause intervals of the spoken text based on interval duration levels corresponding to the spoken pause intervals and positions of the spoken pause intervals in the spoken text, and interval duration levels corresponding to the prosodic pause intervals and positions of the prosodic pause intervals in the spoken text.

[0085] Please refer to Figure 4 , Figure 4 is a structural schematic diagram of an embodiment of an electronic device of the present application. The electronic device 40 comprises a memory 401 and a processor 402 coupled with each other. The memory 401 stores program instructions (not labeled). The processor 402 is configured to execute the program instructions to implement the audio synthesis method in any of the above embodiments. For related content, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0086] The above scheme can improve the humanization degree of audio synthesis.

[0087] Please refer to Figure 5 , Figure 5 is a structural schematic diagram of an embodiment of a computer readable storage medium of the present application. The computer readable storage medium 50 stores program instructions 500. The program instructions 500 are executed by a processor to implement the audio synthesis method in any of the above embodiments. For related content, please refer to the detailed description of the above method embodiments, which will not be repeated here.

[0088] The above scheme can improve the humanization degree of audio synthesis.

[0089] It should be noted that the units described as separate components may or may not be physically separate, and the components displayed as units may or may not be physical units, i.e. may be located in one place, or may be distributed to multiple network units. Part or all of the units can be selected according to actual needs to achieve the purpose of the present embodiment scheme.

[0090] In addition, each functional unit in each embodiment of the present application can be integrated in one processing unit, or each unit can be physically present alone, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software functional unit.

[0091] The integrated unit, if realized in the form of a software functional unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical scheme of the present application essentially or the part of the prior art that contributes to the technical scheme or the whole or part of the technical scheme can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the method of each embodiment of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various program code storage media.

[0092] The above is only the embodiment of the present application, and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation, or direct or indirect application in other related technical fields, is also included in the patent protection scope of the present application.

Claims

1. An audio synthesis method, characterized by, The method comprises the following steps: In response to obtaining the text to be processed, inserting oral expression content into the text to be processed to obtain oral text; wherein the oral expression content at least includes oral new content and oral pause interval, and the oral expression content further includes oral adjustment content, which is used to replace part of the content in the text to be processed; Obtain the prosodic features of the oral text, and obtain the prosodic pause interval of the oral text based on the prosodic features; Convert the oral text into target audio based on the oral pause interval and the prosodic pause interval; specifically, determine the text pause interval of the oral text based on the oral pause interval and the prosodic pause interval; obtain the phoneme sequence corresponding to the oral text based on the oral text and its corresponding text pause interval; obtain the target audio corresponding to the oral text based on the phoneme sequence; a plurality of prosodic pause intervals correspond to a plurality of interval duration levels, and at least part of the interval duration levels match a plurality of oral pause intervals; The method further comprises the following steps: Determine the text pause interval of the oral text based on the interval duration level corresponding to the oral pause interval and the position in the oral text, and the interval duration level corresponding to the prosodic pause interval and the position in the oral text; the interval duration level includes at least word level L1, phrase level L2, prosodic pause L3, punctuation L4 and paragraph pause L5 from short to long; the oral pause interval includes short pause, long pause and super-long pause, short pause corresponds to L2, long pause corresponds to L3, and super-long pause corresponds to L4.

2. The method of claim 1, wherein, The method further comprises the following steps: Input the text to be processed into a text conversion model to obtain oral text fed back by the text conversion model; wherein the text conversion model is trained using parallel text pairs, and the parallel text pairs include a first training text and a second training text, and the second training text has inserted oral expression content compared with the first training text.

3. The method of claim 2, wherein, The text conversion model includes an encoder and a decoder, and the first training text is input into the encoder and the decoder to obtain a first predicted text; or, The text conversion model includes a prompt construction unit and an intelligent analysis model, the prompt construction unit is used to construct a prompt text based on at least the first training text and input the intelligent analysis model, the prompt text is used to prompt the intelligent analysis model to perform oral conversion and output a second predicted text, and when the prompt text further includes part of the parallel text pairs, the parallel text pairs are used as example samples; The oral expression content inserted in the second training text is used as a training label of the text conversion model.

4. The method according to any one of claims 1 to 3, characterized in that, The oralized new content at least includes various oralized expression words, and the various oralized expression words at least include mood words, filler words and repetition words. Each of the oralized new content corresponds to a matching new start tag, each of the oralized new content corresponds to a new end tag, and each of the oralized pause interval corresponds to a matching pause tag.

5. The method of claim 2, wherein, The method further includes: obtaining a prosody feature of the oralized text; and obtaining a prosody pause interval of the oralized text based on the prosody feature. The method further includes: inputting the oralized text into a text analysis model to obtain the prosody pause interval fed back by the text analysis model; wherein the text analysis model is used to perform text regularization on the oralized text and perform word segmentation on the regularized oralized text, determine the prosody feature based on the segmented oralized text, and obtain the prosody pause interval of the oralized text based on the prosody feature; and the text analysis model is trained by using the first training text.

6. An audio synthesis system characterized by, The method further includes: The conversion module is configured to insert oralized expression content into the to-be-processed text to obtain an oralized text in response to obtaining the to-be-processed text; wherein the oralized expression content at least includes oralized new content and oralized pause intervals, and the oralized expression content further includes oralized adjustment content used to replace part of the content in the to-be-processed text. The analysis module is configured to obtain a prosody feature of the oralized text and obtain a prosody pause interval of the oralized text based on the prosody feature. The generation module is configured to convert the oralized text into a target audio based on the oralized pause interval and the prosody pause interval; specifically, the generation module is configured to determine a text pause interval of the oralized text based on the oralized pause interval and the prosody pause interval, obtain a phoneme sequence corresponding to the oralized text based on the oralized text and the text pause interval corresponding to the oralized text, and obtain the target audio corresponding to the oralized text based on the phoneme sequence; and a plurality of prosody pause intervals correspond to a plurality of interval length levels, and at least part of the interval length levels match the oralized pause intervals. The method further includes: The method further includes: determining the text pause interval of the oralized text based on the interval length level corresponding to the oralized pause interval and the position of the oralized pause interval in the oralized text, and the interval length level corresponding to the prosody pause interval and the position of the prosody pause interval in the oralized text; and the interval length levels at least include a word level L1, a phrase level L2, a prosody pause L3, a punctuation symbol L4 and a paragraph pause L5 from short to long, the oralized pause intervals include a short pause, a long pause and an ultra-long pause, the short pause corresponds to L2, the long pause corresponds to L3, and the ultra-long pause corresponds to L4.

7. An electronic device, comprising: The method further includes: a memory and a processor coupled to each other, the memory storing program instructions, and the processor configured to execute the program instructions to implement the audio synthesis method according to any one of claims 1-5.

8. A computer-readable storage medium having stored thereon program instructions, wherein, The program instructions, when executed by a processor, implement the audio synthesis method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Voice synthesis method and device, electronic equipment and readable storage medium

    CN112735379A

  • Speech synthesis method and related device

    CN116778903A

  • Speech synthesis method and device, equipment and storage medium

    CN117174074A