Audio synthesis method and system, electronic device, and computer readable storage medium

By inserting oral expression content into the text and combining pronunciation characteristics, the problem of insufficient anthropomorphism of audio synthesis in the prior art is solved, and a more natural and realistic audio synthesis effect is achieved.

WO2025123652A1PCT designated stage expired Publication Date: 2025-06-19IFLYTEK CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/103424
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-14
Filing Date
2024-07-03
Publication Date
2025-06-19

AI Technical Summary

Technical Problem

During the existing human-computer interaction process, the degree of anthropomorphism of audio synthesis is insufficient and cannot be effectively improved.

Method used

By inserting colloquial expression content into the pending text, the pronunciation characteristics of the colloquial text are obtained, and the pronunciation pause interval is determined based on these characteristics, and the text is converted into the target audio in combination with the colloquial pause interval.

Benefits of technology

It improves the degree of anthropomorphism of audio synthesis, so that the generated audio contains more colloquial new content and natural pause intervals during playback, enhancing the realism of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024103424_19062025_PF_FP_ABST
    Figure CN2024103424_19062025_PF_FP_ABST
Patent Text Reader

Abstract

An audio synthesis method and system (30), an electronic device (40), and a computer readable storage medium (50). The method comprises: in response to obtaining a text to be processed, inserting spoken expression content into the text to be processed to obtain a spoken text, wherein the spoken expression content at least comprises spoken newly added content and a spoken pause interval (S101); acquiring a prosodic feature of the spoken text, and obtaining a prosodic pause interval of the spoken text on the basis of the prosodic feature (S102); and converting the spoken text into a target audio on the basis of the spoken pause interval and the prosodic pause interval (S103). The method can improve the personification degree of audio synthesis.
Need to check novelty before this filing date? Find Prior Art

Description

Audio synthesis method, system, electronic device and computer-readable storage medium

Technical field

[0001] The present application relates to the field of data processing technology, and in particular to an audio synthesis method, system, electronic device, and computer-readable storage medium. [Background Technology]

[0002] With the development of smart devices, more and more devices support human-computer interaction. This interaction often requires synthesizing text data into audio for feedback. However, existing human-computer interaction processes often simply convert text data into rigid audio, resulting in insufficiently humanized audio synthesis. Therefore, improving the humanization of audio synthesis has become a pressing issue.

[0003] [Summary of the invention]

[0004] The main technical problem solved by this application is to provide an audio synthesis method, system, electronic device and computer-readable storage medium, which can improve the degree of anthropomorphism of audio synthesis.

[0005] In order to solve the above technical problems, the first aspect of the present application provides an audio synthesis method, comprising: in response to obtaining a text to be processed, inserting colloquial expression content into the text to be processed to obtain a colloquial text; wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals; obtaining the rhythmic features of the colloquial text, and obtaining the rhythmic pause intervals of the colloquial text based on the rhythmic features; and converting the colloquial text into target audio based on the colloquial pause intervals and the rhythmic pause intervals.

[0006] In order to solve the above technical problems, the second aspect of the present application provides an audio synthesis system, including: a conversion module, an analysis module and a generation module. In response to obtaining a text to be processed, the conversion module is used to insert colloquial expression content into the text to be processed to obtain a colloquial text; wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals; the analysis module is used to obtain the rhythmic features of the colloquial text and obtain the rhythmic pause intervals of the colloquial text based on the rhythmic features; the generation module is used to convert the colloquial text into target audio based on the colloquial pause intervals and the rhythmic pause intervals.

[0007] In order to solve the above technical problems, the third aspect of the present application provides an electronic device, which includes a memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the audio synthesis method described in the first aspect above.

[0008] In order to solve the above technical problems, the fourth aspect of the present application provides a computer-readable storage medium on which program instructions are stored. When the program instructions are executed by a processor, the audio synthesis method described in the first aspect is implemented.

[0009] The above scheme, when obtaining the text to be processed, inserts colloquial expression content into the text to be processed, and converts the text to be processed into colloquial text, wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals. Therefore, the colloquial text can contain more content that will be added during colloquial expression, as well as pauses that will be used during colloquial expression. The rhythmic features of the converted colloquial text are obtained, so that the colloquial new content can affect the rhythmic features, and thus the rhythmic pause intervals corresponding to the colloquial text are obtained based on the rhythmic features, so that the rhythmic pause intervals are also more colloquial. Based on the colloquial pause intervals and the rhythmic pause intervals, the colloquial text including the colloquial new content is audio synthesized, and the colloquial text is converted into target audio, so that the target audio can include the colloquial new content and the pause intervals are also more colloquial when played, thereby improving the degree of anthropomorphism of the audio synthesis.

Brief Description of the Drawings

[0010] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. Among them:

[0011] FIG1 is a flow chart of an embodiment of an audio synthesis method of the present application;

[0012] FIG2 is a flow chart of another embodiment of the audio synthesis method of the present application;

[0013] FIG3 is a schematic structural diagram of an audio synthesis system according to an embodiment of the present invention;

[0014] FIG4 is a schematic structural diagram of an embodiment of an electronic device of the present application;

[0015] FIG5 is a schematic structural diagram of an embodiment of a computer-readable storage medium of the present application. [Specific implementation method]

[0016] The following will be combined with the accompanying drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0017] The terms "system" and "network" are often used interchangeably in this document. The term "and / or" is simply a description of an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " generally indicates that the related objects are in an "or" relationship. Furthermore, "multiple" in this document refers to two or more than two.

[0018] The audio synthesis method provided in this application is used to synthesize audio with a high degree of anthropomorphism. Its execution subject is a processor capable of processing text and audio. The processor may belong to an intelligent terminal, such as an intelligent terminal capable of human-computer interaction with a user. This application does not impose any specific restrictions on this.

[0019] Please refer to FIG1 , which is a flow chart of an embodiment of an audio synthesis method of the present application. The method includes:

[0020] S101: In response to obtaining a text to be processed, inserting colloquial expression content into the text to be processed to obtain a colloquial text, wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals.

[0021] Specifically, when a text to be processed is obtained, colloquial expression content is inserted into the text to be processed to convert the text to be processed into a colloquial text, wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals.

[0022] It is understandable that colloquial texts can contain more content when they are able to include more colloquial expressions, as well as the pauses that are used in colloquial expressions.

[0023] In one embodiment, when a text to be processed is obtained, the text to be processed is input into a pre-trained conversion model, so that the conversion model inserts colloquial expressions into the text to be processed, thereby obtaining a colloquial text output by the conversion model. The conversion model is trained using written text and its corresponding colloquial expression text.

[0024] In one embodiment, when a text to be processed is obtained, a first prompt text is constructed based on the text to be processed, and the first prompt text is input into the intelligent analysis model to obtain a colloquial text fed back by the intelligent analysis model, wherein the first prompt text is used to prompt the intelligent analysis model to perform colloquial conversion.

[0025] Optionally, the intelligent analysis model is a large language model, which may include but is not limited to deep neural networks (DNNs), convolutional neural networks (CNNs), recurrent neural networks (RNNs), long short-term memory networks (LSTM), and generative pre-trained Transformer models, etc. No specific restrictions are imposed on the specific construction and deployment of the large language model.

[0026] Furthermore, the colloquial expression content includes colloquial additions and colloquial pauses. The colloquial additions include at least various types of colloquial expression words, and the colloquial pauses include pauses of various lengths. Furthermore, the colloquial expression content may also include colloquial adjustments. The colloquial adjustments are used to replace portions of the text to be processed, replacing more formal expressions with the colloquial adjustments.

[0027] It should be noted that various types of colloquial expressions include at least modal particles, filler words and repeated words, and onomatopoeia audio includes at least laughter, inhalation and lip smacking sounds.

[0028] S102: Acquire rhythmic features of the spoken text, and obtain rhythmic pause intervals of the spoken text based on the rhythmic features.

[0029] Specifically, the rhythmic features of the converted colloquial text are obtained so that the newly added colloquial content can affect the rhythmic features, thereby obtaining the rhythmic pause intervals corresponding to the colloquial text based on the rhythmic features, making the rhythmic pause intervals more colloquial.

[0030] In one embodiment, preset characters in a spoken text are converted into a matching text form, and the adjusted text is divided into basic units, wherein the preset characters include at least numbers, abbreviations, and currency symbols, and the basic units include at least words and punctuation marks. The rhythmic features of the spoken text are determined based on the divided text, and the spoken text is separated based on the rhythmic features to obtain the rhythmic pause intervals of the spoken text.

[0031] In one embodiment, the syntax, discourse structure and information structure of the spoken text are obtained, the prosodic features of the spoken text are determined based on the syntax, discourse structure and information structure, the spoken text is segmented based on the prosodic features, and the prosodic pause intervals of the spoken text are obtained.

[0032] It is understandable that the colloquial pause intervals inserted in the colloquial text have nothing to do with the text content, and the colloquial pause intervals can be blocked when obtaining the rhythmic features of the colloquial text. However, the newly added colloquial content has been added to the colloquial text, which can affect the rhythmic features in the process of extracting the rhythmic features, making the rhythmic features more inclined to the colloquial expression.

[0033] In one implementation scenario, a spoken text is input into a text analysis model to obtain prosodic pause intervals as feedback from the text analysis model, wherein the text analysis model is trained using multiple written texts and spoken texts corresponding to at least some of the written texts.

[0034] S103: Converting the spoken text into target audio based on the spoken pause intervals and the prosodic pause intervals.

[0035] Specifically, based on the spoken pause intervals and rhythmic pause intervals, audio synthesis is performed on the spoken text including the newly added spoken content, and the spoken text is converted into target audio so that the target audio can include the newly added spoken content and the pause intervals are more spoken when played.

[0036] In one embodiment, the spoken pause intervals and the prosodic pause intervals are fused to obtain the text pause intervals of the spoken text, and based on the text pause intervals and the words in the spoken text, a phoneme sequence corresponding to the spoken text is obtained, and the phoneme sequence is converted into a target audio corresponding to the spoken text.

[0037] In one embodiment, the spoken text is divided into multiple fields based on spoken pause intervals and prosodic pause intervals, phoneme sequences corresponding to the multiple fields are determined, and the phoneme sequences are converted into target audio corresponding to the spoken text.

[0038] In one implementation scenario, a phoneme sequence is input into an end-to-end audio conversion model to obtain a target audio output by the audio conversion model, wherein the audio conversion model is trained using phoneme sequence samples and their corresponding converted audio.

[0039] In one application scenario, the text to be processed is the reply text generated when the user interacts with the smart terminal to reply to the user. The text to be processed is converted into colloquial language and output as a target audio with a high degree of anthropomorphism, making the user feel more real and intimate during the human-computer interaction.

[0040] In another application scenario, the text to be processed is a text obtained from information input by the user, and the information input by the user includes at least one of text, voice and image. The information input by the user is finally converted into the text to be processed, and the text to be processed is converted into colloquial language and output as target audio with a high degree of anthropomorphism, so that the original expression in the text to be processed is adjusted to a colloquial expression and converted into a playable audio. For example, the text to be processed, which was originally expressed in awkward written form, is converted into target audio and played, thereby improving the user's acceptability when listening.

[0041] Optionally, the text to be processed can be a student's written knowledge text. When the text to be processed is converted into a target audio, it can be convenient for students to listen to the converted target audio and better learn the knowledge points. The text to be processed can also be a children's story text. When the text to be processed is converted into a target audio, it can be convenient for children to listen to the converted target audio and improve the intimacy of listening to the story. In addition, the application of the text to be processed in other specific application scenarios will not be listed one by one.

[0042] The above scheme, when obtaining the text to be processed, inserts colloquial expression content into the text to be processed, and converts the text to be processed into colloquial text, wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals. Therefore, the colloquial text can contain more content that will be added during colloquial expression, as well as pauses that will be used during colloquial expression. The rhythmic features of the converted colloquial text are obtained, so that the colloquial new content can affect the rhythmic features, and thus the rhythmic pause intervals corresponding to the colloquial text are obtained based on the rhythmic features, so that the rhythmic pause intervals are also more colloquial. Based on the colloquial pause intervals and the rhythmic pause intervals, the colloquial text including the colloquial new content is audio synthesized, and the colloquial text is converted into target audio, so that the target audio can include the colloquial new content and the pause intervals are also more colloquial when played, thereby improving the degree of anthropomorphism of the audio synthesis.

[0043] Please refer to FIG2 , which is a flow chart of another embodiment of the audio synthesis method of the present application, the method comprising:

[0044] S201: In response to obtaining a text to be processed, inputting the text to be processed into a text conversion model to obtain a colloquial text fed back by the text conversion model, wherein the text conversion model is obtained by training using a parallel text pair, the parallel text pair including a first training text and a second training text, and the second training text has colloquial expression content inserted compared to the first training text.

[0045] Specifically, when the text to be processed is obtained, the text to be processed is input into the text conversion model to obtain the spoken text fed back by the text conversion model, thereby converting the text to be processed into the spoken text, wherein the text conversion model is obtained by pre-training using parallel text pairs, and the parallel text pairs include a first training text and its corresponding second training text.

[0046] It should be noted that compared with the first training text, the second training text is inserted with colloquial expressions, and the colloquial expressions include at least colloquial new content and colloquial pause intervals. Therefore, the second training text is more colloquial than the first training text. The colloquial expressions inserted in the second training text can be used as training labels for supervised training of the text conversion model.

[0047] Furthermore, the text conversion model can be used to efficiently insert colloquial expressions into the text to be processed, thereby effectively completing the conversion from the text to be processed to the colloquial text.

[0048] It can be understood that the first training text included in the parallel text pair is usually a written expression, wherein part of the first training text has some spoken features, and the second training text is the spoken expression generated after inserting the spoken expression content into the first training text. The text conversion model obtained by training the parallel text pair can convert the text of written expression into the text of spoken expression, and convert the text with some spoken features into the text of complete spoken expression.

[0049] It should be noted that the newly added colloquial content includes at least one of modal particles, filler words, repeated words and onomatopoeic audio, and the colloquial pause intervals include pause intervals of various time lengths; wherein, each newly added colloquial content is preceded by a matching newly added start tag, each newly added colloquial content is followed by a matching newly added end tag, and each colloquial pause interval is accompanied by a matching pause tag.

[0050] Specifically, in addition to modal particles, filler words, repeated words and onomatopeia audio, the new spoken content may also include other new content such as erhua sounds, and this application does not impose specific restrictions on this. Onomatopeia audio may include laughter, inhalation, coughing and lip smacking sounds, and this application does not impose specific restrictions on this.

[0051] Furthermore, when new colloquial content is inserted into the text, a matching new start tag is inserted before each new colloquial content, and a new end tag is inserted after each new colloquial content. For example, the new start tag of the filler word is fib, the new start tag of the filler word is fie, the new start tag of the modal particle is yqb, and the new start tag of the filler word is yqe. Other specific tags are not listed one by one in this application.

[0052] It can be understood that the spoken pause intervals include at least pauses, short pauses, long pauses and extra-long pauses, with a total of pause intervals of various lengths. Each spoken pause interval corresponds to a matching pause label, for example, the pause label for a short pause is spa, and the pause label for a long pause is lpa. Other specific labels are not listed one by one in this application.

[0053] For illustration, using a parallel text pair as an example, the first training sample is written in a written style, specifically including: "For example, in the morning, we have soy milk, fried dough sticks, and steamed buns." The second training sample is spoken in a colloquial style, and after inserting colloquial expressions, it includes: "<fib> um <fie>, like in the morning, <fib>, basically <fie> is <spa>, soy milk, fried dough sticks, and steamed buns." Therefore, adding start and end tags can distinguish and locate the newly added colloquial content, while pause tags can distinguish and locate the colloquial pauses. The text conversion model can be trained through rule definition or adaptive learning, so that the trained text conversion model can insert colloquial expressions into the processed text and set matching labels at the corresponding locations.

[0054] In one embodiment, the text conversion model includes an encoder and a decoder, and the first training text is passed through the encoder and the decoder to obtain the first predicted text. The colloquial expressions inserted into the second training text serve as training labels for the text conversion model.

[0055] Specifically, the training process of the text conversion model includes: inputting a first training text into an encoder to obtain a first encoded text output by the encoder, inputting the first encoded text into a decoder to obtain a first predicted text output by the decoder, comparing the first predicted text with the second training text, determining a first training loss based on the colloquial expression content included in the second training text and the predicted expression content inserted in the first predicted text, adjusting the parameters of the encoder and decoder based on the first training loss until a first convergence condition is met, and obtaining a trained text conversion model. Therefore, the text conversion model can be a model constructed based on the encoder and decoder and obtained through supervised training, so that the trained text conversion model can complete the colloquial conversion of the text to be processed and improve the controllability of the colloquial conversion process.

[0056] In another embodiment, the text conversion model includes a prompt construction unit and an intelligent analysis model. The prompt construction unit is configured to construct a prompt text based on at least a first training text and input the prompt text into the intelligent analysis model. The prompt text is configured to prompt the intelligent analysis model to perform a colloquial conversion and output a second predicted text. When the prompt text also includes some parallel text pairs, the parallel text pairs serve as example samples. The colloquial expressions inserted into the second training text serve as training labels for the text conversion model.

[0057] Specifically, the prompt construction unit in the text conversion model is used to construct a prompt text, which includes at least a first training text and a colloquial conversion requirement. When the prompt text also includes some parallel text pairs, the parallel text pairs serve as example samples that match the colloquial conversion requirement. After the prompt text is given to the intelligent analysis model, the intelligent analysis model feeds back a second predicted text, compares the second predicted text with the second training text, and determines a second training loss based on the colloquial expression content included in the second training text and the predicted expression content inserted in the second predicted text. Based on the second training loss, the intelligent analysis model is fine-tuned until the second convergence condition is met to obtain a trained text conversion model. Therefore, the text conversion model can be obtained by fine-tuning the intelligent analysis model through supervised training, so that the trained text conversion model can complete the colloquial conversion of the text to be processed, reducing the terminal's consumption for colloquial conversion.

[0058] In one implementation scenario, fine-tuning the intelligent analysis model can use a zero-shot approach to construct prompt text. For example, the first training sample is written in a specific style, including: "For example, we have soy milk, fried dough sticks, and steamed buns in the morning." The prompt text constructed using the zero-shot approach may include: "Please convert the text {For example, we have soy milk, fried dough sticks, and steamed buns in the morning} into colloquial language and set labels at the corresponding positions."

[0059] In another implementation scenario, fine-tuning the intelligent analysis model can use a few-shot approach to construct prompt text. The few-shot approach can include the following: For a written expression of Southern cuisine, "I really prefer dipping sauces." The transcribed spoken text is <fib> um <fie>. For Southern cuisine, <fib> er <fie> <std> "I really prefer dipping sauces <yqb> ah <yqe>." Based on the above instructions, transcribe the following written text: {For example, we have soy milk, fried dough sticks, and steamed buns in the morning}. Therefore, when the prompt text also includes some parallel text pairs, using the parallel text pairs as example samples can improve the learning accuracy and convergence efficiency of the intelligent analysis model.

[0060] S202: Acquire rhythmic features of the spoken text, and obtain rhythmic pause intervals of the spoken text based on the rhythmic features.

[0061] Specifically, the prosodic features of the spoken text are extracted, and the prosodic pause intervals of the spoken text are determined using the prosodic features.

[0062] In one embodiment, a spoken text is input into a text analysis model to obtain a rhythmic pause interval as feedback from the text analysis model; wherein the text analysis model is used to regularize the spoken text and segment the regularized spoken text, determine the rhythmic features based on the segmented spoken text, and thereby obtain the rhythmic pause interval of the spoken text based on the rhythmic features, and the text analysis model is trained using a first training text.

[0063] Specifically, the conversion from the text to be processed to the target audio is divided into two stages. The first stage is that the text conversion model converts the text to be processed into colloquial language to obtain colloquial text. The second stage is that the text analysis model parses the colloquial text to obtain the rhythmic pause intervals and completes the audio conversion subsequently.

[0064] Furthermore, the text analysis model is trained using the first training text. The text analysis model focuses on the mining of rhythmic features and the ability to determine rhythmic pause intervals based on the rhythmic features. The trained text analysis model can regularize the spoken text and segment the regularized spoken text, determine the rhythmic features based on the segmented spoken text, and thus obtain the rhythmic pause intervals of the spoken text based on the rhythmic features.

[0065] It should be noted that when regularizing the text, the preset characters in the spoken text are converted into a matching text form, and when segmenting the words, the adjusted text is divided into basic units, wherein the preset characters include at least numbers, abbreviations and currency symbols, and the basic units include at least words and punctuation marks. The rhythmic features of the spoken text are determined based on the divided text, and the spoken text is separated based on the rhythmic features to obtain the rhythmic pause intervals of the spoken text, thereby improving the accuracy of the rhythmic pause intervals.

[0066] S203: Converting the spoken text into target audio based on the spoken pause intervals and the prosodic pause intervals.

[0067] Specifically, based on the spoken pause intervals and the prosodic pause intervals, audio synthesis is performed on the spoken text including the spoken new content, and the spoken text is converted into the target audio.

[0068] In one embodiment, the text pause interval of the spoken text is determined based on the spoken pause interval and the prosodic pause interval; the phoneme sequence corresponding to the spoken text is obtained based on the spoken text and its corresponding text pause interval; and the target audio corresponding to the spoken text is obtained based on the phoneme sequence.

[0069] Specifically, the spoken pause intervals and the rhythmic pause intervals are integrated to obtain the text pause intervals corresponding to the spoken text, so that the text pause intervals are obtained by combining the spoken features and the rhythmic features, thereby improving the accuracy of pauses in the spoken text, and the spoken text is divided using the text pause intervals to obtain the divided words, which are converted into phoneme sequences, that is, pronunciation is represented by phonemes, and audio conversion is performed based on the phoneme sequence to obtain the target audio corresponding to the spoken text, so that the target audio has a higher degree of anthropomorphism.

[0070] In one implementation scenario, acoustic features are obtained from a phoneme sequence, where the acoustic features include at least fundamental frequency and spectral features. The acoustic features are converted into continuous speech sampling point audio signals to obtain target audio corresponding to the spoken text.

[0071] In one specific implementation scenario, the mapping relationship between the phoneme sequence and the acoustic features is obtained by a feature extraction model, which may include at least one of a recurrent neural network (RNN), a long short-term memory (LSTM), and a generative pre-trained transformer. This application does not impose specific restrictions on this. The conversion of acoustic features to target audio is completed by a vocoder. This application does not impose specific restrictions on the algorithm used by the vocoder.

[0072] Optionally, the conversion between the phoneme sequence and the target audio can also be obtained using an end-to-end audio conversion model, wherein the audio conversion model is trained using phoneme sequence samples and their corresponding converted audio.

[0073] In one implementation scenario, multiple prosodic pause intervals correspond to multiple interval duration levels, and multiple spoken pause intervals match at least some of the interval duration levels; based on the spoken pause intervals and the prosodic pause intervals, the text pause intervals of the spoken text are determined, including: based on the interval duration levels corresponding to the spoken pause intervals and their positions in the spoken text, and the interval duration levels corresponding to the prosodic pause intervals and their positions in the spoken text, the text pause intervals of the spoken text are obtained.

[0074] Specifically, multiple prosodic pause intervals are matched with their own interval duration levels, and multiple spoken pause intervals can be matched with at least some of the interval duration levels. The positions of the spoken pause intervals in the spoken text and the interval duration levels of the spoken pause intervals are marked, and the positions of the prosodic pause intervals in the spoken text and the interval duration levels of the prosodic pause intervals are marked to obtain the text pause intervals of the spoken text, so that the spoken pause intervals and the prosodic pause intervals are integrated with each other to obtain text pause intervals with unified grade standards, so that the target audio conversion has higher accuracy.

[0075] In a specific application scenario, the multiple prosodic pause intervals include at least five levels of interval duration, where the interval duration levels, from short to long, include at least word level L1, phrase level L2, prosodic pause L3, punctuation mark L4, and paragraph pause L5. The spoken pause intervals include short pauses, long pauses, and extra-long pauses, where short pauses correspond to L2, long pauses correspond to L3, and extra-long pauses correspond to L4. In other specific application scenarios, the number of interval duration levels and the specific interval duration levels that the spoken pause intervals match can be customized, and this application does not impose specific restrictions on this.

[0076] It should be noted that the implementation methods or implementation scenarios described in any of the above embodiments are not limited to a single embodiment, and different implementation methods can be combined with each other, and this application does not impose specific restrictions on this.

[0077] Different from the above embodiments, the conversion from the text to be processed to the target audio is divided into two stages. The first stage is that the text conversion model converts the text to be processed into colloquial language to obtain colloquial text. The second stage is that the text analysis model parses the colloquial text to obtain rhythmic pause intervals and subsequently completes the audio conversion. Among them, the construction and training of the text conversion model include multiple implementation methods, and the trained text conversion model can insert colloquial expressions into the text to be processed and set matching labels at corresponding positions, and then fuse the colloquial pause intervals with the rhythmic pause intervals to obtain a phoneme sequence, that is, using phonemes to represent pronunciation, thereby performing audio conversion based on the phoneme sequence to obtain the target audio corresponding to the colloquial text, so that the target audio has a high degree of anthropomorphism.

[0078] Please refer to Figure 3, which is a structural diagram of an embodiment of the audio synthesis system of the present application. The audio synthesis system 30 includes: a conversion module 301, an analysis module 302 and a generation module 303. In response to obtaining the text to be processed, the conversion module 301 is used to insert colloquial expression content into the text to be processed to obtain a colloquial text; wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals; the analysis module 302 is used to obtain the rhythmic features of the colloquial text and obtain the rhythmic pause intervals of the colloquial text based on the rhythmic features; the generation module 303 is used to convert the colloquial text into the target audio based on the colloquial pause intervals and the rhythmic pause intervals.

[0079] In the above scheme, when the text to be processed is obtained, the conversion module 301 inserts colloquial expression content into the text to be processed and converts the text to be processed into colloquial text, wherein the colloquial expression content at least includes colloquial new content and colloquial pause intervals. Therefore, the colloquial text can contain more content that will be added during colloquial expression, as well as pauses that will be used during colloquial expression. The analysis module 302 obtains the rhythmic features of the converted colloquial text, so that the colloquial new content can affect the rhythmic features, thereby obtaining the rhythmic pause intervals corresponding to the colloquial text based on the rhythmic features, making the rhythmic pause intervals more colloquial. The generation module 303 performs audio synthesis on the colloquial text including the colloquial new content based on the colloquial pause intervals and the rhythmic pause intervals, and converts the colloquial text into target audio, so that the target audio can include the colloquial new content and the pause intervals are more colloquial when played, thereby improving the degree of anthropomorphism of the audio synthesis.

[0080] Optionally, the conversion module 301 is also used to input the text to be processed into a text conversion model to obtain a colloquial text fed back by the text conversion model; wherein the text conversion model is obtained by training using parallel text pairs, the parallel text pairs include a first training text and a second training text, and the second training text has colloquial expression content inserted compared to the first training text.

[0081] Optionally, the text conversion model includes an encoder and a decoder, and the first training text is passed through the encoder and the decoder respectively to obtain a first predicted text; or, the text conversion model includes a prompt construction unit and an intelligent analysis model, the prompt construction unit is used to construct a prompt text based on at least the first training text and input it into the intelligent analysis model, and the prompt text is used to prompt the intelligent analysis model to perform colloquial conversion and output a second predicted text. When the prompt text also includes some parallel text pairs, the parallel text pairs are used as example samples; wherein, the colloquial expression content inserted in the second training text is used as a training label for the text conversion model.

[0082] Optionally, the colloquial new content includes at least one of modal particles, filler words, repeated words and onomatopoeic audio, and the colloquial pause intervals include pause intervals of various time lengths; wherein, each colloquial new content is preceded by a matching new start tag, each colloquial new content is followed by a matching new end tag, and each colloquial pause interval is accompanied by a matching pause tag.

[0083] Optionally, the analysis module 302 is also used to input the spoken text into the text analysis model to obtain the rhythmic pause intervals fed back by the text analysis model; wherein the text analysis model is used to perform text regularization on the spoken text and to perform word segmentation on the regularized spoken text, and to determine the rhythmic features based on the spoken text after word segmentation, thereby obtaining the rhythmic pause intervals of the spoken text based on the rhythmic features, and the text analysis model is trained using the first training text.

[0084] Optionally, the generation module 303 is also used to determine the text pause interval of the spoken text based on the spoken pause interval and the prosodic pause interval; obtain the phoneme sequence corresponding to the spoken text based on the spoken text and its corresponding text pause interval; and obtain the target audio corresponding to the spoken text based on the phoneme sequence.

[0085] Optionally, a plurality of interval duration levels correspond to each prosodic pause interval, and a plurality of spoken pause intervals match at least some of the interval duration levels; the generation module 303 is further configured to obtain the text pause interval of the spoken text based on the interval duration levels corresponding to the spoken pause intervals and their positions in the spoken text, and the interval duration levels corresponding to the prosodic pause intervals and their positions in the spoken text.

[0086] Please refer to Figure 4, which is a schematic diagram of the structure of an electronic device according to one embodiment of the present application. The electronic device 40 includes a memory 401 and a processor 402, which are coupled to each other. The memory 401 stores program instructions (not labeled), and the processor 402 is configured to execute the program instructions to implement the audio synthesis method of any of the above-mentioned embodiments. For a detailed description of the relevant content, please refer to the detailed description of the above-mentioned method embodiments, and will not be repeated here.

[0087] The above solution can improve the degree of anthropomorphism of audio synthesis.

[0088] Please refer to Figure 5, which is a schematic diagram of the structure of one embodiment of a computer-readable storage medium of the present application. This computer-readable storage medium 50 stores program instructions 500. When executed by a processor, program instructions 500 implement the audio synthesis method described in any of the aforementioned embodiments. For a detailed description of the relevant content, please refer to the detailed description of the aforementioned method embodiments and will not be repeated here.

[0089] The above solution can improve the degree of anthropomorphism of audio synthesis.

[0090] It should be noted that the units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of this embodiment.

[0091] In addition, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0092] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of each embodiment method of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0093] The above description is only an implementation method of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the description and drawings of this application, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An audio synthesis method, wherein: The audio synthesis method comprises: In response to obtaining the text to be processed, inserting colloquial expression content into the text to be processed to obtain a colloquial text; wherein the colloquial expression content at least includes colloquial additional content and colloquial pause intervals; Acquiring a rhythmic feature of the spoken text, and obtaining a rhythmic pause interval of the spoken text based on the rhythmic feature; The spoken text is converted into a target audio based on the spoken pause interval and the prosodic pause interval.

2. The method according to claim 1, wherein: The step of inserting colloquial expressions into the text to be processed to obtain a colloquial text includes: The text to be processed is input into a text conversion model to obtain a spoken text fed back by the text conversion model; wherein the text conversion model is obtained by training with a parallel text pair, the parallel text pair includes a first training text and a second training text, and the second training text has the spoken expression content inserted therein compared to the first training text.

3. The method according to claim 2, wherein: The text conversion model includes an encoder and a decoder, and the first training text is passed through the encoder and the decoder to obtain a first predicted text; The colloquial expression content inserted into the second training text serves as a training label for the text conversion model.

4. The method according to claim 2, wherein: The text conversion model includes a prompt construction unit and an intelligent analysis model, wherein the prompt construction unit is used to construct a prompt text based on at least the first training text and input the prompt text into the intelligent analysis model, and the prompt text is used to prompt the intelligent analysis model to perform colloquial conversion and output a second predicted text; The colloquial expression content inserted into the second training text serves as a training label for the text conversion model.

5. The method according to claim 4, wherein: The prompt text at least includes the first training text and the colloquial conversion requirement. When the prompt text also includes part of the parallel text pairs, the parallel text pairs serve as example samples that match the colloquial conversion requirement.

6. The method according to claim 2, wherein: The step of acquiring the prosodic features of the spoken text and obtaining the prosodic pause interval of the spoken text based on the prosodic features includes: The spoken text is input into a text analysis model to obtain the prosodic pause interval fed back by the text analysis model; wherein the text analysis model is used to perform text regularization on the spoken text and to perform word segmentation on the regularized spoken text, and the prosodic features are determined based on the spoken text after word segmentation, thereby obtaining the prosodic pause intervals of the spoken text based on the prosodic features, and the text analysis model is trained using the first training text.

7. The method according to claim 1, wherein: The colloquial additional content includes at least one of modal particles, filler words, repeated words and onomatopoeic audio, and the colloquial pause intervals include pause intervals of various time lengths.

8. The method according to claim 7, wherein: Each of the colloquial new contents is preceded by a matching new start tag, each of the colloquial new contents is followed by a matching new end tag, and each of the colloquial pause intervals is accompanied by a matching pause tag.

9. The method according to claim 1, wherein: The converting the spoken text into a target audio based on the spoken pause interval and the prosodic pause interval comprises: Determining a text pause interval of the spoken text based on the spoken pause interval and the prosodic pause interval; Based on the spoken text and the corresponding text pause intervals, obtaining a phoneme sequence corresponding to the spoken text; Based on the phoneme sequence, the target audio corresponding to the spoken text is obtained.

10. The method according to claim 9, wherein: A plurality of the prosodic pause intervals correspond to a plurality of interval duration levels, and a plurality of the spoken pause intervals match at least some of the interval duration levels.

11. The method according to claim 10, wherein: The step of determining the text pause interval of the spoken text based on the spoken pause interval and the prosodic pause interval includes: Based on the interval duration level corresponding to the spoken pause interval and the position in the spoken text, and the interval duration level corresponding to the prosodic pause interval and the position in the spoken text, the text pause interval of the spoken text is obtained.

12. The method according to claim 10, wherein: The multiple rhythmic pause intervals include at least five levels of interval duration levels, and the interval duration levels include at least word level, phrase level, rhythmic pause, punctuation mark and paragraph pause from short to long, and the spoken pause intervals include short pauses, long pauses and extra-long pauses; wherein the short pause corresponds to the phrase level, the long pause corresponds to the rhythmic pause, and the extra-long pause corresponds to the punctuation mark.

13. An audio synthesis system, characterized in that: include: A conversion module, in response to obtaining the text to be processed, is used to insert colloquial expression content into the text to be processed to obtain a colloquial text; wherein the colloquial expression content at least includes colloquial additional content and colloquial pause intervals; An analysis module, used for acquiring the prosodic features of the spoken text, and obtaining the prosodic pause intervals of the spoken text based on the prosodic features; A generating module is used to convert the spoken text into a target audio based on the spoken pause interval and the prosodic pause interval.

14. An electronic device, characterized in that: include: A memory and a processor coupled to each other, wherein the memory stores program instructions, and the processor is used to execute the program instructions to implement the audio synthesis method as described in any one of claims 1-12.

15. A computer-readable storage medium having program instructions stored thereon, characterized in that: When the program instructions are executed by a processor, the audio synthesis method according to any one of claims 1 to 12 is implemented.

Citation Information

Patent Citations

  • Voice output method and device, electronic equipment and computer readable storage medium

    CN111489752A

  • Speech synthesis method and related device

    CN116778903A

  • Text processing method and system for speech synthesis

    CN117037769A

  • Speech synthesis method and device, equipment and storage medium

    CN117174074A

  • Audio synthesis method and system, electronic equipment and computer readable storage medium

    CN118135988A