A speech synthesis method and synthesis system

By applying text algorithms and TTS technology in the field of online literature, we can automatically generate audio works with multiple characters, emotions, and scenarios, solving the problems of long production cycles and high costs in existing technologies. This achieves efficient and low-cost speech synthesis and improves the user experience.

CN116092472BActive Publication Date: 2026-02-27SHANGHAI YUEWEN INFORMATION TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202211710199.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-29
Publication Date
2026-02-27
Estimated Expiration
2042-12-29

AI Technical Summary

Technical Problem

Existing speech synthesis technology suffers from problems such as long voice selection cycle, insufficient text preprocessing, lack of emotion and special effects in audio synthesis, and long production cycle and high cost when generating audio works with multiple characters, emotions, and scenes, resulting in a complex process and low efficiency.

Method used

Automatic text preprocessing is performed using text algorithms based on the online literature domain. Combined with TTS speech technology, structured picture books are generated through character mining, text structuring, emotion and special effects sound scene recognition. Audio material tags are trained using BERT deep model to achieve automatic speech synthesis for multiple characters, emotions and scenes.

Benefits of technology

It improves the efficiency and quality of speech synthesis, reduces production cycle and costs, and can automatically generate high-quality commercial audio works, covering more categories of timbres, and enhancing the user's listening experience and immersion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116092472B_ABST
    Figure CN116092472B_ABST
Patent Text Reader

Abstract

The application provides a speech synthesis method, comprising the following steps: step one, preparing an audio material library, dividing the audio in the audio material library into multiple categories according to types, and labeling the audio; step two, training the label of the audio material obtained in step one based on a Bert deep model to obtain the word vector corresponding to each audio material; step three, inputting a text segment of a novel, performing paragraph structural analysis on the text segment to generate a structured picture book; step four, obtaining the audio candidate with the highest similarity in the structured picture book according to semantic approximation; step five, calling an applicable TTS engine, fusing the candidate audio in the audio material library matched, and performing speech output according to the required output structure information. As a new type of technology, the application can fuse the TTS engine with the specific scene or the gender and emotion of the character in the network text, convey more infectious information to the user, and improve the commercial value of the related product.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of speech synthesis, and relates to a speech synthesis method and a synthesis system. BACKGROUND

[0002] The market of audio books has increased by 35%+ year by year in the past 3 years and is still in the growth period; various online reading platforms focus on the audio book track and attract and cultivate user habits through higher quality and more types of audio content to expand the market. At present, the audiobook market is highly competitive, and various audiobook software is flourishing. First, the audiobook system shares knowledge by means of the penetrating power of sound, so that users can make full use of fragmented time. Second, the audiobook system converts dry text into lively and vivid voice reading by means of the rhythm of the intonation, so that users can listen to books like music. Although the current audiobook system technology is already very mature, how to immerse users in the plot of the novel and make it more vivid and immersive has become a difficult problem to improve the user experience.

[0003] There is no similar technology or implementation scheme in the industry at present. The conventional method is to rely on manual speculation of the emotions, actions, etc. of the characters in the novel, and to combine different scenes to perform real person dubbing. For example: on a cold and windy night, she is crying alone in the room. The dubbing here is that the female protagonist speaks with crying sound under the background sound of cold wind howling.

[0004] For example, in the existing field of speech synthesis, such as the speech synthesis method in Chinese invention patent CN109523986A, the steps are as follows:

[0005] 1. Acquiring text information and determining the characters in the text information and the text content of each character;

[0006] 2. Performing character recognition on the text content of each character to determine the character attribute information of each character;

[0007] 3. According to the character attribute information of each character, acquiring the voice actors corresponding to each character, and then manually confirming;

[0008] 4. Generating multi-character synthesized speech according to the text information and the voice actors corresponding to the characters of the text information.

[0009] The main problems of the existing method are:

[0010] (1) Voice selection: after character mining, the appropriate voice needs to be selected from the candidate voice, which is time-consuming;

[0011] (2) Text preprocessing: no pre-processing is performed on the text, and corrections are made for some wrong characters, etc.;

[0012] (3) Audio synthesis single: text synthesis lacks emotion, special effects, scene and other characteristics;

[0013] (4) Longer production cycle, high cost: due to the need for human dubbing, the cycle is longer, the output of the dubbing volume is small and the cost is high;

[0014] (5) Complex and inefficient: because it needs to be targeted to the current scene and the action and emotion of the character, first, the description of the current scene in the novel and the action or emotion of the character are manually identified, and then the background sound that meets the scene description is added. When the character produces a certain action or emotional fluctuation, the real person reading also changes accordingly. In this way, the whole process becomes extremely complex. SUMMARY

[0015] In order to solve the problems existing in the prior art, the purpose of the present application is to provide a voice synthesis method and synthesis system.

[0016] The present application can automatically generate the dubbing corresponding to the plot according to the plot of the novel in the network novel, call the required TTS engine type, and fuse the current scene and the emotion of the character. That is, the present application uses the text algorithm deposited in the network novel field to automatically preprocess the novel text, and automatically generates a structured picture book with multiple roles, multiple emotions and multiple scenes. At the same time, the structured picture book (script picture book) is combined with TTS voice technology to produce high-quality commercial audio works.

[0017] In the present application, the text algorithm deposited in the network novel field includes a network novel character name mining model, a text structured picture book generation model, a dialogue text emotion recognition model, and a special effect sound scene recognition model.

[0018] The network novel character name mining model is used to extract the character names, gender, age and other characteristics of the characters appearing in the novel text, and provides role information support for the text structuring process.

[0019] The text structured picture book generation model is used to disassemble the text into sentences and extract the dependency relationship between the characters and the sentences to provide text structure for audio synthesis.

[0020] The dialogue text emotion recognition model and the special effect sound scene recognition model are used to classify the text emotion and extract the position of the special effect sound scene through the language model, and finally improve the audio listening experience.

[0021] The TTS voice technology in the present application refers to the voice synthesis (Text To Speech) technology, that is, converting the text appearing in the computer into natural and fluent voice output.

[0022] In order to intelligently output high-quality commercial audio works through natural language processing capabilities and audio synthesis capabilities, the present application provides a speech synthesis method, comprising the following steps:

[0023] Step one, prepare an audio material library, divide the audio in the audio material library into multiple categories according to the type, and label the audio;

[0024] Step two, based on the label of the audio material obtained in step one, obtain the word vector corresponding to each audio material; specifically, the label of the audio material is obtained through the Bert deep model to obtain its text semantic vector (i.e. word vector), and in subsequent use, the matching is performed through vector similarity calculation.

[0025] Step three, input the text segment of the novel, perform paragraph structural analysis on the text segment, and generate a structured sketchbook;

[0026] Step four, according to semantic approximation, obtain the audio candidate with the highest similarity in the structured sketchbook for different types of words and the audio material library;

[0027] Step five, call the applicable TTS engine, fuse the candidate audio in the audio material library matched, and perform voice output according to the required output structure information.

[0028] In step one, the type of the audio includes: scene class audio and emotion class audio corresponding to different characters and genders;

[0029] And / or,

[0030] The audio material library is derived from historical accumulated audio, and audio denoising operation is uniformly performed;

[0031] And / or,

[0032] The labeling of the audio refers to labeling each type of scene represented in the audio with a corresponding scene label, and labeling the emotion corresponding to different characters and genders represented in the audio with a gender and emotion label.

[0033] In step two, the generation of the word vector corresponding to the audio material specifically includes the following steps:

[0034] Step 1, train a Bert deep semantic model containing corresponding semantic information in each layer; this model contains several network layers;

[0035] Step 2, input the audio material text label, and extract the Bert output layer word vector.

[0036] The length of the output word vector is 784.

[0037] The step three further includes text content preprocessing, including the following steps:

[0038] Step 31, input original text content;

[0039] Step 32, proofread the original text content using a correction model;

[0040] Step 33, analyze the characters in the proofread text content;

[0041] Step 34, generate formatted text containing characters and plots for structured sketchbook.

[0042] In step 31, the original text content includes novels, literary works, historical biographies.

[0043] In step 32, the proofreading includes text correction, semantic replacement, content dialogue symbol regularization, and non-text filtering; the text correction is to preliminarily recall text errors through an open source correction pycorrect scheme, and then filter through a high-frequency error word set based on tens of millions of historical web text statistics to ensure that the recalled error words are common errors in web text.

[0044] The correction model combines open source correction schemes and web text field correction sets.

[0045] In the present application, text correction greatly reduces easy-to-mistake words and improves speech synthesis accuracy. The accuracy of easy-to-mistake words is improved to 90%. Symbol regularization and non-text filtering mainly filter non-main content in novels, and rules are deposited through massive non-main text in the text.

[0046] Specifically,

[0047] The text correction refers to correcting the wrong words in the text; the semantic replacement refers to, for example: "a small dot is moving towards this side" where "between" will be corrected to "only see"; the content dialogue symbol regularization or non-text filtering is to delete the text symbols that are not related to the text, for example: "……ps: tomorrow add more", which will filter out this part of the text, improving the listener's experience after speech synthesis.

[0048] In step 33, when analyzing the text role, role mining, role alias alignment, role relationship attribute establishment, role attribute identification and role tone matching need to be performed; the role mining refers to end-to-end model learning by constructing a named entity recognition task and fusing a data sample enhancement strategy to achieve the purpose of extracting a novel character name; the role alias alignment refers to associating other appellations of a role to the main appellation of the role, for example, a role name is Gu Nu'er, and a small name of Nu Bao appears in the text, and the role name alignment will associate the two role names; the role relationship attribute establishment refers to constructing a deep semantic role relationship model by constructing a role relationship and a text combination sample, dividing a paragraph of a novel, and then using a role relationship model (a language model based on bert) to determine the relationship between roles; the role attribute identification refers to identifying the age, gender and other attributes of a role appearing in an article by constructing a language model; and the role tone matching refers to matching the extracted role attributes with a tone library to select suitable tone information of the role. The tone library is artificially collected and contains role tone, speech speed and tone of a role.

[0049] In the role mining part, the end-to-end model mainly refers to a NER task realized by a sequence labeling model based on bem, the model mainly includes a multi-layer transform+, the model inputs a novel chapter text, and extracts the role name and gender information of a text segment.

[0050] In the role alias alignment part, the role alias alignment mainly refers to a classification model based on bert, and two role information is input to determine whether they belong to the same alias.

[0051] In the role relationship attribute establishment part, the role relationship attribute establishment mainly refers to a classification model based on bert, and a novel text containing role information is input to output the personality classification of the role.

[0052] The role attribute identification refers to a prompt model based on Bert, role dialogue text and role name information are input, and role attribute information is extracted, for example, the age, gender and other attributes of a role appearing in a text are identified.

[0053] The model constructed in the application mainly adopts a sequence labeling based NER (named entity recognition) task, a novel chapter text is input, and the role name and gender information of a text segment are extracted, and then a role screening strategy is used to extract the frequency of a novel role and the chapter segment where the role appears, so as to provide sufficient screening conditions for role tone identification.

[0054] The role screening strategy refers to, for the role information mined by the model, adopting a regular matching mode to match the role name in the novel text, and counting the frequency of the role and the chapters where the role appears, and filtering the role names with low frequency of appearance or irregular naming.

[0055] The data sample enhancement refers to role name interchanging, and refers to, in the training sample in the named entity recognition task, if the input text paragraph contains multiple role names (which can be used as an identification label), the labels of the training data can be interchanged, so that a single training sample can be expanded into multiple samples, greatly improving the data amount of the training sample.

[0056] The role mining method used in the application can achieve a role extraction accuracy of 97% in the web field.

[0057] For example, inputting the novel 'Da Feng Dacengren', the corresponding role names such as Xu Qian, Wei Yuan and Li Suan are extracted.

[0058] In step three of the application, the text segment is converted into a text dialogue script with multiple roles, multiple emotions and multiple scenes by performing role dialogue recognition, text emotion recognition and special effect scene mining on the input text segment, and a structured script is generated.

[0059] The application converts chapters into dialogues through role dialogue dependency relationship and anaphora resolution text processing, uses Bert to build a text emotion classification model to realize fine-grained emotion classification of the text, and gives the text emotion attributes, extracts the positions of special effects in the novel segment through keywords and scene semantics, and finally creates a structured script, including the following methods:

[0060] Step 35, the chapters of the formatted text containing roles and plot text obtained in advance are disassembled and converted into dialogues;

[0061] Step 36, the dialogue role recognition, text emotion recognition and special effect scene recognition processing are performed on the text content converted into dialogues obtained in step 35;

[0062] Step 37, the final structured script is generated.

[0063] The anaphora resolution text processing in the application refers to, based on the chapter dimension of the novel, the chapter text is split into a plurality of sentences through semantic rules, the dialogue sentence and its corresponding context information are input into a semantic-based role dialogue relationship matching model, the correct dialogue subject is matched for each dialogue sentence, the model can also identify multiple names of the character role (for example, personal pronouns and nicknames), and finally the multiple names can be resolved into a unique role name, simplifying the complexity of role voice matching.

[0064] The semantic-based role dialogue relationship matching model in the application refers to a multiple-choice model based on BERT, which inputs current sentence text, a context segment where the sentence is located and a role list that may appear in the context. The roles in the list are spliced with the context segment respectively, and multiple spliced data are input into the BERT model to finally perform a classification task on the multiple spliced segments to achieve the purpose of matching role dialogue dependency relationships.

[0065] The text sentiment classification model refers to a text sentiment classification model based on BERT, which inputs dialogue center sentences and context information on the basis of novel chapter text splitting, performs a classification task of multiple sentiment categories, and the current text sentiment recognition model mostly adopts a model for judging the high and low of positive and negative emotions. In order to achieve the purpose of reflecting diversified emotions, the application adopts seven kinds of emotion categories, i.e. calm, angry, happy, sad, surprised, frightened and narrative, to identify the emotion types of dialogue segments of characters in novels, so that diversified emotional colors can be added to the audio.

[0066] In step 35 of the application, the chapter disassembly refers to disassembling a formatted text into role dialogue sentences and narrative sentences according to a dialogue disassembly strategy.

[0067] The role dialogue sentence refers to a sentence in which a character asks a dialogue in the text, and the rest of the text is expressed as a narrative sentence.

[0068] The dialogue disassembly strategy refers to a strategy for disassembling a text according to the writing habits of an author and a regular matching mode of the text.

[0069] In step 36, the dialogue role recognition refers to mapping dialogue text in a chapter text to a corresponding role through a dialogue role recognition module. The model structure in the dialogue role recognition module adopts a language model structure, and a role decision structure based on context is constructed.

[0070] The language model structure in the application refers to a deep learning language model based on BERT, which inputs dialogue text and identifies dialogue characters. The role decision structure refers to structuring disassembly of a text, dialogue role relationship recognition and finally converting into a structured role decision structure of a script by combining a disassembly strategy and a language model.

[0071] After the role dialogue recognition module model architecture is constructed, input sample X={X1, X2,..., Xn}, Xn represents the nth data, Xn=(Cn, Qn, [Choice1, Choice2,..., Choicem], label), m represents a total of m candidate role names, Qn represents a dialogue text sentence, referred to as a center sentence, Cn is a context fragment before and after the center sentence, the dynamic range of which changes based on the model recognizable character length L, a given dynamic window range is used to make the character length of the Cn fragment len(Cn)<L, and the window is reduced when the character is too long; [Choice1, Choice2,..., Choicem] are the candidate roles appearing in the fragment, and label is the real role serial number corresponding to the actual sample;

[0072] A training module in the model is constructed, the training module is based on a language model, the language model is an encoding part of the whole model, denoted as LM, and the model structure of the training module is as follows:

[0073] M_role=softmax(concat(Class(LM([Cn, Qn, Choice1]),..., Class(LM([Cn, Qn, Choicem])),

[0074] Wherein, m represents a total of m options, wherein Class is a score corresponding to the current combined text output by the LM, concat is used to splice the scores corresponding to multiple candidate answers, s o ftmax obtains m class outputs; the output result is fitted with label;

[0075] The content in the training module is repeated, and when the accuracy of the model after training reaches a 90% measurement index, the optimal model is saved, and the chapter text is predicted;

[0076] Input original text is denoted as H={H1, H2,..., Hn},

[0077] Wherein, Hn=(Cn, Qn, [Choice1, Choice2,..., Choicem]), the output result R=Choice[Max_index(M_role(Batch(H)))] is obtained, wherein Max_index is the serial number of the role with the maximum probability in the candidate roles corresponding to the center sentence; and finally, the item with the highest model score in the candidate list [Choice1, Choice2,..., Choicem] is output as the output result.

[0078] In step 36 of the present application, the text sentiment recognition refers to recognizing the text sentiment of the dialogue through a sentiment algorithm model, and the model adopts a text classification structure based on a language model.

[0079] The text classification structure of the present application is based on a deep semantic text classification model realized by BERT, and the dialogue is divided into multiple sentiment types.

[0080] After constructing the text sentiment recognition module, input the sample D={D1, D2,..., Dn}, Dn represents the nth data, Dn=(Cn, label), Cn is a text with sentiment, including a main sentence, hereinafter referred to as a center sentence, and context fragments before and after the center sentence, and label is a sentiment label; the serial number in the sentiment classification set is used to represent the sentiment classification set E=[E1, E2,..., Em];

[0081] A training module for text sentiment recognition is constructed; the language model is used as the encoder part of the overall model, denoted as LM; the overall model is as follows:

[0082] M_emotion=softmax(LM(C1, C2,....Cn))

[0083] s o After the output of the ftmax function, the category score output is obtained, which is fitted with the label;

[0084] The above model training steps are repeated, when the model training accuracy reaches 95%, the best model is saved, and the sentiment of the dialogue sentence in the main text is predicted; the input original text C={c1, c2,..., cn} and the output result are:

[0085] R=Choice[Max_index(M-emotion(Batch(C)))]

[0086] Wherein, Max_index is the serial number of the probability maximum role in the center sentence corresponding to the sentiment classification set, and the item with the highest model score in the candidate list [E1, E2,..Em] of the final output sentiment classification set is taken as the output result.

[0087] In step 36 of the present application, the special effect scene recognition refers to understanding the text semantics, extracting the position of the special effect sound in the sentence, and adding the corresponding special effect audio (such as wind sound, footstep sound, etc.) in the sound effect synthesis step. This step is mainly realized based on a sequence labeling model of BERT, and the training sample uses a text sentence plus a special effect sound and its position as an input value, and the output form of the trained model is the starting position of the special effect sound in the text and the corresponding special effect sound category, realizing the position marking of the special effect sound at the text level, and then the special effect sound audio can be added in the audio fusion process.

[0088] The special effect scene recognition of the application is realized by using a text word granularity classification model to locate scene words in the text by a special effect scene mining module, and the scene in which a special effect needs to be added is obtained; the model adopts a text classification structure based on a language model.

[0089] The training data is input into the special effect sound model, and model training is performed, and when a 90% measurement index is reached, the optimal model is saved:

[0090] M_special=NER(LM(Cn))

[0091] The original text T={T1, T2,..., Tn} is input into the model, and the output result (start, end, special_index) = Choice [(M_special(Batch(T))] is extracted; wherein Choice represents threshold screening of the score output by the model, and the special-index in the output is the special effect sound sequence number.

[0092] In step four of the application, the semantic approximate matching is performed between the role, emotion and scene in the structured picture book generated in step three and the word vector of the audio material label mapping in the audio material library, so as to obtain the candidate audio material corresponding to each part of the structured picture book, and the most suitable audio material is selected by manual optimization.

[0093] In step four of the application, the structured picture book tone matching includes the following steps:

[0094] Step 41, the special effect scene text, role characteristics and novel label characteristics in the structured picture book are matched; wherein the novel label characteristics come from the classification attributes of novels in the self-built novel label library; the special effect scene text, role characteristics and novel label characteristics are respectively matched with special effect audio, role tone library and background audio to obtain special effect audio, role audio and background audio.

[0095] Step 42, after the matching is completed, the tone of the role is synthesized.

[0096] Step 43, audio format of special effects, synthesized character sound, background sound, and output the formatted audio.

[0097] The special effect audio refers to using a special effect audio library with more than 30 categories and more than 500 types, and the special effect audio library is a self-built library, wherein the materials come from the Internet, the positions of the special effect texts in the text segment are extracted through keyword and scene semantic extraction, and accurate track is performed;

[0098] The character tone library matching refers to extracting character personality characteristics through syntax dependency analysis and user interactive comments, and assigning the most suitable tone in the tone library to the character through deep semantic similarity matching; wherein the information of the character personality characteristics comes from the character mining part in the text content preprocessing process, and the gender, age and other characteristics of the characters in the article are extracted through the NER task. After completing the character relationship construction and character attribute extraction, the tone assigned to each character is as follows: for example, XQ (young and smart) and Chenfeng (young and energetic) are automatically matched in semantics.

[0099] The background sound matching refers to matching the background music library through the book classification information in the book library, improving the audio background atmosphere, and strengthening the sense of immersion.

[0100] In step 42 of the present application, the character tone synthesis refers to inputting the character associated text and the character tone information through the text-to-speech interface, and outputting the audio of the text.

[0101] In step 43 of the present application, the audio format unification refers to using uniform frequency processing to unify the character tone, special effect sound and background sound in the sense of hearing.

[0102] The syntax dependency analysis refers to disassembling a sentence into a dependency syntax tree, describing the dependency relationship between words in the sentence, or indicating the syntactic collocation relationship between words in the sentence, and the collocation relationship is associated with semantics.

[0103] The tone library content matched by the present application is rich, which can adapt to characters of different categories and different personalities, and can enrich and expand more tones.

[0104] In the tone matching process of the present application, the text structured content is combined into a speech synthesis markup language (SSML) tag, a TTS engine is used for speech synthesis, and high-quality audio with multiple tones and multiple emotions is realized.

[0105] In step five of the present application, a Chinese TTS engine and / or an English TTS engine matched with the content of the text segment are called, and corresponding audio is accompanied when the character emotion changes and / or a certain action occurs in the text scene, and speech output is performed according to the required output speed.

[0106] The audio fusion method comprises the following steps:

[0107] Step 51, audio multi-track fusion is performed on the obtained formatted audio comprising special effect sound, character sound and background sound;

[0108] Step 52, audio system is unified after multi-track audio fusion;

[0109] Step 53, the audio of unified system is output to obtain complete audio of the audio book.

[0110] After the audio fusion, the file is stored in the cloud for downstream distribution. The audio multi-track fusion refers to synthesizing multiple audio tracks into a new audio containing all the audio tracks, so that the special effect sound and the background sound are integrated into the character sound to form a complete chapter audio output. The audio system unification refers to setting the synthesized audio to a unified audio format, maintaining the uniformity of audio frequency and sound quality.

[0111] The above method mainly includes automatic picture book generation and AI voice synthesis, and the two parts can be replaced by the following two ways:

[0112] 1) Artificial picture book import + AI voice synthesis: import artificial picture book, automatically align artificial picture book and automatic picture book format, access AI voice synthesis, and realize automatic synthesis of artificial script;

[0113] 2) Automatic picture book generation + real person and AI mixed recording: automatically create a novel picture book, the system synthesizes multiple voice colors and multiple emotions for each sentence, and replaces part of the AI recorded voice with real person recorded voice to realize mixed recording of real person and AI.

[0114] The key technical points of the present application mainly include two points:

[0115] 1. Web NLP processing technology. Based on the exploration of web text technology precipitation and frontier new technology, character atlas mining, chapter structuring, text emotion classification, scene recognition and semantic matching.

[0116] 2. AI voice library. High-quality voice adaptation for various categories of novel characters in the web field is realized to achieve high-quality multi-voice AI quality broadcasting.

[0117] The present application also proposes an application of a text content preprocessing method in natural language text processing.

[0118] The application further provides a text content preprocessing system for realizing the text content preprocessing method, which comprises a role information mining module, a role relationship construction module, a role voice tone matching module and a role information storage module; the role information mining module is used for mining role attributes to output attribute information including a role name, a role gender and an age; the role relationship construction module is used for aligning role aliases and constructing role relationship attributes; the role voice tone matching module is used for matching role corresponding voice tones according to role attributes, and the information includes a speech speed and a tone of the role voice tone; and the role information storage module is used for storing the output role information in a corresponding database table to construct a role information query system for subsequent structured picture books.

[0119] The application further provides an application of the structured picture book generation method in text structured processing.

[0120] The application further provides a structured picture book generation system for realizing the structured picture book generation method, which comprises a text chapter disintegration module, a dialogue role identification module, a text emotion identification module, a special effect scene identification module and a structured picture book output module.

[0121] The application further provides an application of the voice tone matching method in text voice tone matching.

[0122] The application further provides a voice tone matching system for realizing the voice tone matching method, which comprises a role voice tone matching module, a special effect sound matching module, a scene sound matching module and an audio format unified processing module.

[0123] The application further provides an application of the audio fusion method in audio fusion of an audio book.

[0124] The application further provides a voice tone matching system for realizing the voice tone matching method, which comprises an audio volume normalization module and an audio fusion module.

[0125] The application has the following beneficial effects:

[0126] (1) High output efficiency

[0127] The application realizes full-automatic processing of novel text picture book creation through natural language processing technology from role mining, voice tone matching, chapter structuring to special effect scene identification. Compared with the current manual picture book creation process, the application is more efficient, and the time is saved by 30 times.

[0128] (2) High voice actor richness

[0129] The application maintains a timbre library applicable to multiple categories, which can cover most of the timbre characteristics of novel character roles. And with the popularization of timbre migration technology, the platform timbre library is continuously expanding new timbers. Compared with artificial dubbing, first, the application can cover more categories of timbers, second, AI timbers can concurrently synthesize multiple books, and the time consumption of single book synthesis is reduced by 10 times (300 days->30 days) compared with artificial. If the same timbers are used in multiple books, the time consumption of artificial is limited to the time consumption of the sound engineer, which is longer.

[0130] In the era of fierce competition in the audiobook market and various audiobook software, how to immerse users in the plot of the novel and make them more vivid and empathetic has great commercial value. As a new technology, the application can integrate the TTS engine with the specific scene or gender and emotion of the character in the web novel to convey more infectious information to the user and improve the commercial value of related products. BRIEF DESCRIPTION OF DRAWINGS

[0131] Figure 1 A flowchart of a voice synthesis method in the prior art.

[0132] Figure 2 A flowchart of a voice synthesis method of the application.

[0133] Figure 3 A character map schematic diagram of the application.

[0134] Figure 4 A scene timbre schematic diagram of the application.

[0135] Figure 5 A module diagram involved in the method running process of the application.

[0136] Figure 6 A step diagram executed in the content preprocessing module of the application.

[0137] Figure 7 A step diagram executed in the structured sketchbook module of the application.

[0138] Figure 8 A step diagram executed in the timbre matching module of the application.

[0139] Figure 9 A step diagram executed in the audio fusion module of the application.

[0140] Figure 10 An example diagram in an embodiment of the application. DETAILED DESCRIPTION

[0141] The application will be further described in detail in combination with the following specific embodiments and drawings. The process, conditions, experimental methods and the like for implementing the application are the general knowledge and common sense in the art, and the application does not have special limitations.

[0142] The application realizes a voice synthesis method. The method can call a required TTS engine type according to a novel plot in a network novel, and automatically generate a corresponding dubbing for the plot by fusing a current scene and a character emotion.

[0143] 1. Prepare a material library. The audio types in the material library are divided into two types:

[0144] ① Scene type audio, such as audio of howling wind and sound of water flow. The audio of this type is used as background sound.

[0145] ② Emotion type audio corresponding to different character genders, such as audio of female crying and male laughing.

[0146] Label the above audio types.

[0147] 2. Train the material label. Use a Bert deep model to train and obtain a word vector corresponding to a label of a scene type and an emotion type material.

[0148] 3. Input a novel text segment, use parsing technology to structure it, and specially, separately process the following several types of structured content, as follows:

[0149] (1) Scene words, such as howling wind, dark and windy night, and flowing water.

[0150] (2) Gender words, such as she, he, sister, mother, and Mr.

[0151] (3) Emotion description words of a character, such as crying and laughing.

[0152] 4. Use semantic similarity to obtain the audio with the highest similarity in the material library for the structured picture book obtained by deconstruction. The scene type words and the audio in the material library are given candidates based on semantic similarity, and the most suitable audio for this part is given based on artificial optimization.

[0153] 5. Call a required TTS engine, and when a character emotion changes or a certain action is generated, the corresponding audio is accompanied by sound, and then voice output is performed according to a required output speed. The TTS engine is mainly a Chinese TTS engine and an English TTS engine.

[0154] More specifically, structured storyboards refer to the system's conversion of novel chapters into text dialogue scripts with multiple characters, emotions, and scenes. This function mainly consists of a character dialogue recognition module, a text emotion recognition module, and a special effects scene mining module. An example of a structured storyboard is shown below. Figure 10 As shown.

[0155] The character dialogue recognition module refers to mapping the dialogue text in the chapter's main text to the corresponding characters in the novel, for example... Figure 10 The sentences with IDs 2, 4, and 10 shown are identified as Xu Xinnian, Xu Qi'an, and Xu Xinnian, respectively, as the initiators of the dialogue. The model structure can adopt a language model structure to construct a context-based role decision structure.

[0156] The text sentiment recognition module refers to the identification of emotions in dialogue text using a self-developed sentiment algorithm model, currently supporting emotions such as happiness, sadness, calmness, anger, surprise, fear, and narration. The model structure can also adopt a text classification structure based on a language model.

[0157] The special effects scene mining module refers to the localization of scene words in text through a text word granularity classification model. For example, "walking away quickly." outputs [2, walking sound]. Currently, it supports nearly a thousand scene sounds. The model structure can also adopt a text classification structure based on a language model.

[0158] The steps of destructuring are as follows:

[0159] Step 1: Construct the model architecture for the role dialogue recognition module. Input samples, X = {X1, X2, ..., Xn}, where Xn represents the nth data item, Xn = (cn, Qn, [Choice1, Choice2, ..., Choicem], label), where m represents the number of candidates, Qn represents the main sentence of the dialogue (hereinafter referred to as the center sentence), Cn is the dynamic range context segment before and after the center sentence, [Choice1, Choice2, ..., Choicem] are the role candidates appearing within this segment, e.g., [Xu Qi'an, Xu Xinnian], and label is the actual role number corresponding to the actual sample. Figure 10 For example, in China:

[0160] Qn means "Please wait a moment."

[0161] Cn states, "Xu Qi'an's logical reasoning ability was unparalleled in his previous life, making him a top student in his grade. In the past, Xu Xinnian wouldn't have paid him any attention, but considering this parting between the brothers might be their last, he agreed to his brother's final request, whispering, 'Wait a moment.' He quickly left. His footsteps disappeared into the corridor, and Xu Qi'an sat down against the railing, his heart filled with complex emotions."

[0162] [Choice1, Choice2,..., Choicem] is [Xu Qian, Xu Xinian]

[0163] label is 1.

[0164] Step 2: Construct a language model-based training module. The language model is the encoder part of the overall model, denoted as LM. The overall model is as follows

[0165] M-role = softmax(concat(Class(LM([Cn, Qn, Choice1]),..., Class(LM([Cn, Qn, Choicem])))

[0166] m represents a total of m options, where Class is the score corresponding to the current combination of text output by LM, concat is to concatenate the scores corresponding to multiple candidate answers, s o ftmax gets m class outputs. Fit it with label.

[0167] Step 3: Repeat step 2, save the optimal model when the measurement index is reached (accuracy 90%), and predict the chapter text. Input the original text H = {H1, H2,..…, Hn}

[0168] Where Hn = (Cn, Qn, [Choice1, Choice2,..., Choicem]), the output result R = Choice[Max_index(M_role(Batch(H)))], where Max_index is the probability maximum role sequence number in the candidate role corresponding to the center sentence.

[0169] By using the language model as the encoder, learning based on web text corpus samples greatly reduces the difficulty of the model in understanding the text semantics. For example:

[0170] After the injured Xu Pingzhi went to the hospital, Xu Qian ran over to Li Ru and said, "xx", the model result: Xu Qian

[0171] After the injured Xu Pingzhi ran over to Li Ru and said, "xx". The model result: Xu Pingzhi

[0172] In complex scenes where multiple characters interact, the model can accurately identify the dialogue sender.

[0173] Step 4: Build a text sentiment recognition module, input sample D = {D1, D2,..., Dn}, Dn represents the nth data, Dn = (Cn, label), Cn is a text with emotion, including the main sentence, hereinafter referred to as the center sentence, and the context fragments before and after the center sentence, label is the emotion label represented by the serial number in the emotion classification set, emotion classification set E = [E1, E2,....., Em], for example: [narration, angry, happy...].

[0174] Take Figure 10 for example,

[0175] Cn is: I want to crack the case... Xu Qian said in a low voice: \n"I want to know what happened, even if I die, I won't give up" Directly say crack the case, Xu Xinian will probably think he has a brain, so Xu Qian changed his words

[0176] Label is 2

[0177] Step 5: Build a training module for text sentiment recognition. Language model as the encoder part of the overall model, denoted as LM. The overall model is as follows

[0178] M_emotion = softmax (LM (C1, C2,... Cn))

[0179] s o ftmax after getting the category score output, make it fit with label

[0180] Step 6: Repeat step 5, when the model training accuracy reaches 95%, save the best model and predict the sentiment of the dialogue sentence in the text. Input the original text C = {c1, c2,..., cn};

[0181] Output result:

[0182] R = Choice [Max_index (M_emotion (Batch (C)))]

[0183] Where Max_index is the serial number of the probability maximum role in the center sentence corresponding to the emotion classification set.

[0184] Step 7: Build special effect sound recognition module model architecture: input sample D = {D1, D2,..., Dn}, Dn represents the nth data, Dn = (Cn, (start, end, label)), Cn represents the text containing special effect sound, start and end represent the initial position and end position of the special effect sound respectively; label is the special effect sound category label. Use the serial number in the special effect sound category library to represent, and the special effect sound category library is represented as E = [E1, E2,..…., Em], for example: [footsteps, door opening sound……], the data is represented as follows:

[0185] Cn = 'woman's hoarse cry echoed in the ear, Ye Chen opened his eyes suddenly, and his back was covered with cold sweat.'

[0186] (start, end, label) = (0, 6, '94'), where label = 94 represents the serial number of the cry in the special effect sound library

[0187] Input the training data into the special effect sound model, perform model training, and save the optimal model when the measurement index is reached (accuracy 90%).

[0188] M—special = NER(LM(Cn))

[0189] Step 8: Special effect sound recognition: input the original text T = {T1, T2,..., Tn} into the model, and extract the output result (start, end, special—index) = Choice[(M—special(Batch(T)))] where Choice represents threshold screening of the model output scoTe, and the special—index in the output is the special effect sound serial number.

[0190] Step 9: Voice matching: after the structured sketchbook is generated, the article is disassembled into sentence granularity text structure, and then the structured sketchbook will be matched with voice, including role voice matching, special effect sound voice matching, and audio background sound matching. After voice matching is completed, the sentence granularity text is audio synthesized to generate multiple audio segments.

[0191] In role voice matching, the role gender, age, personality and other characteristics are extracted based on role mining. The role information is input into the role voice library for matching, where the role voice library is a database of different role characteristics corresponding to voice constructed by artificial annotation and collected data.

[0192] Special effect sound matching is to extract the starting position of special effect sound in the text after the text is recognized by the special effect sound recognition model, and to calculate the special effect sound fusion time position according to the following formula after the text audio is generated and merged, and the same way can be used for background music fusion.

[0193] special—local = (seconds—symbol_num * 0.7) * (start—symbol_length) / (words—length—symbol_num) + symbol_length_before_special * 0.7

[0194] + duration_all

[0195] + duration_all

[0196] Wherein, special—local represents the special effect sound insertion audio time position, seconds represents the total duration of the current sentence audio, symbol_num represents the number of symbols in the current sentence, start represents the starting position of the special effect sound in the text, words—length represents the total number of characters in the sentence, symbol_length_before_special represents the number of symbols before the special effect position, and duration_all represents the number of paragraphs in the current paragraph.

[0197] Step 10: Audio fusion: The multi-paragraph sentence granularity audio is fused into chapter granularity multi-track audio, and the audio format is unified to form complete chapter audio. After the synthesis is completed, the audio can be uploaded to the storage cloud and downstream distribution is performed.

[0198] The protection scope of the present application is not limited to the above embodiments. Changes and advantages that can be thought of by those skilled in the art without departing from the spirit and scope of the present application are included in the present application, and are protected by the appended claims.

Claims

1. A speech synthesis method characterized by, Comprise the following steps: Step one, prepare the audio material library, divide the audio in the audio material library into multiple categories according to the type, and label the audio; Step two, based on the label of the audio material obtained in step one, obtain the word vector corresponding to each audio material based on the Bert deep model training; Step three, input the text segment of the novel, and perform paragraph structural analysis on the text segment to generate a structured picture book; In step three, the text segment is converted into a text dialogue script with multiple roles, multiple emotions and multiple scenes through role dialogue recognition, text emotion recognition and special effect scene mining, and a structured picture book is generated; In step three, the text segment is converted into a dialogue through role dialogue dependency relationship and anaphora resolution text processing; a text fine-grained emotion classification model is constructed using Bert to realize text fine-grained emotion classification and assign text emotion attributes; the positions of the special effects in the text segment are extracted through keyword and scene semantic extraction, and finally a structured picture book is created; comprising the following steps: Step 35, the formatted text containing the role and plot text obtained in advance is chaptered and adapted into a dialogue; Step 36, the text content adapted into a dialogue obtained in step 35 is processed by dialogue role recognition, text emotion recognition and special effect scene recognition; In step 36, the special effect scene recognition refers to the realization of scene word positioning in the text through a special effect scene mining module using a text word granularity classification model, obtaining the scenes in the text that need to add scene special effects later; the model uses a text classification structure based on a language model; In the special effect scene recognition process, a special effect sound recognition module model architecture is constructed: input sample D={D1, D2, …, Dn}, Dn represents the nth data, Dn=(Cn, (start, end, label)), Cn represents the text containing special effect sound, start and end represent the initial position and end position of the special effect sound respectively; label is the special effect sound category label; represented by the serial number in the special effect sound category library, the special effect sound category library is represented as Te=[Te1, Te2, …, Tem]; The training data is input into the special effect sound model, and the model training is performed, and when the 90% measurement index is reached, the optimal model is saved: M_special=NER(LM(Cn)), The original text T={T1, T2, …, Tn} is input into the model, and the output result (start, end, special_index)=Choice[(M_special(Batch(T)))] is extracted; wherein, Choice represents threshold screening of the score output by the model, and special_index is the special effect sound serial number; Step 37, generate the final structured picture book; Step four, according to semantic approximation, obtain the highest similarity audio candidate of different types of words in the structured picture book and the audio material library; Step five, call the applicable TTS engine, fuse the matched candidate audio in the audio material library, and output the voice according to the required output structure information.

2. The speech synthesis method of claim 1, wherein, In step one, the type of audio includes: scene type audio and different characters, gender corresponding emotional type audio; And / or, The audio material library is derived from historical accumulated audio, and audio denoising operation is uniformly performed; And / or, The audio is labeled, which means that each type of scene in the audio is labeled with a corresponding scene label, and different characters, gender corresponding emotions in the audio are labeled with gender, emotion labels.

3. The speech synthesis method of claim 1, wherein, In step two, the generation of the word vector corresponding to the audio material specifically includes the following steps: Step 1, training each layer of Bert deep semantic model containing corresponding semantic information; Step 2, input audio material text label, extract Bert output layer word vector.

4. The speech synthesis method of claim 3, wherein, The length of the word vector corresponding to different audio materials in step two is 784.

5. The speech synthesis method of claim 1, wherein, The step three also includes text content preprocessing, including the following steps: Step 31, input original text content; Step 32, correct the original text content using the error correction model; Step 33, analyze the role in the corrected text content; Step 34, generate formatted text containing roles and plots for structured sketchbook.

6. The speech synthesis method of claim 5, wherein, In step 32, the correction includes text error correction, semantic replacement, content dialogue symbol regularization and non-text filtering; the error correction model means that the initial recall of text errors is carried out through the open source error correction pycorrect scheme, and the high frequency error word set is filtered through the historical web text statistics, so as to ensure that the recalled error words are common errors in web text.

7. The speech synthesis method of claim 6, wherein, The text error correction means correcting the wrong words in the text; The semantic replacement means replacing the words with semantic errors with words with correct semantics; The content dialogue symbol regularization and non-text filtering means deleting the characters unrelated to the text in the original text content.

8. The speech synthesis method of claim 5, wherein, In step 33, when analyzing the text role, role mining, role alias alignment, role relationship attribute establishment, role attribute recognition and role tone matching are carried out; The role mining means that the end-to-end model learning is carried out through the construction of named entity recognition task and the fusion of data sample enhancement strategy to achieve the purpose of extracting novel role name; The role alias alignment means associating other appellations of the role to the main appellation of the role; The role relationship attribute establishment means constructing a deep semantic role relationship model by constructing role relationship and text combination samples, dividing the paragraphs of the novel, and then using the role relationship model to judge the relationship between the roles; The role attribute recognition means recognizing the age and gender attributes of the role in the text by constructing a language model; The role tone matching means matching the extracted role attributes with the tone library to select the suitable tone information of the role; The tone library is collected by artificial collection, including role tone and role speed and tone.

9. The speech synthesis method of claim 8, wherein, In the role mining, the end-to-end model refers to the NER task realized by the sequence labeling model based on Bert, the model composition is multi-layer transform+, the novel chapter text is input into the model, and the role name and gender information of the text segment are extracted; And / or, In the role alias alignment, a BERT-based classification model is used to input two role information to determine whether they belong to the same alias. And / or, In the role relationship attribute establishment, a BERT-based classification model is used to input the novel text containing role information, and the output is the personality classification of the role. And / or, In the role attribute identification, a BERT-based prompt model is used to input the role dialogue text and the role name information, and extract the role attribute information.

10. The speech synthesis method of claim 1, wherein, In step 35, the chapter disassembly refers to disassembling the formatted text into role dialogue sentences and narrative sentences according to the dialogue disassembly strategy. The role dialogue sentence refers to the sentence of the dialogue between the characters in the text, and the rest of the text is expressed as a narrative sentence. The dialogue disassembly strategy refers to a strategy for splitting the text according to the author's writing habits and text regular matching. And / or, In step 36, the dialogue role identification refers to mapping the dialogue text in the chapter text to the corresponding role through the dialogue role identification module. The model structure in the dialogue role identification module adopts a language model structure, which constructs a context-based role decision structure.

11. The speech synthesis method of claim 10, wherein, The language model structure refers to a deep learning language model based on BERT, which inputs dialogue text to identify dialogue characters. The role decision structure refers to a structured disassembly of the text based on the disassembly strategy and the language model, dialogue role relationship identification, and finally converted into a structured storyboard role decision structure. And / or, After building the role dialogue identification module model architecture, input the sample X = {X1, X2, …, Xn}, Xn represents the nth data, Xn = (Cn, Qn, [Choice1, Choice2, …, Choicem], label), m represents a total of m candidate role names, Qn represents the dialogue text sentence, called the center sentence, Cn is the context fragment before and after the center sentence, the dynamic range refers to the length of the context fragment, which changes based on the model recognizable character length L, given the dynamic window range, so that the length of the Cn fragment len(Cn) < L, if the character is too long, the window is reduced; [Choice1, Choice2, …, Choicem] is the role candidate appearing in the segment, and label is the real role serial number corresponding to the actual sample; The training module in the model is based on the language model, which is the encoding part of the overall model, denoted as LM. The model structure of the training module is as follows: M_role = softmax(concat(Class(LM([Cn, Qn, Choice1]), …, Class(LM([cn, Qn, Choice])), Where m represents a total of m options, Class is the score corresponding to the current combination of text output by LM, concat is to concatenate the scores of multiple candidate answers, and softmax gets m class outputs; so that the output result is fitted with label. Repeat the content in the training module, save the optimal model when the accuracy of the model after training reaches 90% of the measurement index, and predict the chapter text; Input the original text as H={H1, H2, …, Hn}, Where Hn=(Cn, Qn, [Choice1, Choice2, …, Choicem]), the output result R=Choice[Max_index(M_role(Batch(H)))], and Max_index is the serial number of the role with the maximum probability in the candidate role corresponding to the center sentence; the item with the highest model score in the final output candidate list [Choice1, Choice2, …, Choicem] is taken as the output result.

12. The speech synthesis method of claim 1, wherein, In step 36, the text sentiment recognition refers to identifying the text sentiment of the dialogue through a sentiment algorithm model, and the model adopts a text classification structure based on a language model.

13. The speech synthesis method of claim 12, wherein, The text classification structure is a deep semantic text classification model realized based on bert, which divides the dialogue into multiple emotional types. And / or, After constructing the text sentiment recognition module, input the sample D={D1, D2, …, Dn}, Dn represents the nth data, Dn=(Cn, label), Cn is a text with emotion, including a text sentence, hereinafter referred to as a center sentence, and a dynamic range context segment before and after the center sentence, and label is an emotional label; represented by the serial number in the emotion classification set, emotion classification set E=[E1, E2, …, Em]; Construct a training module for text sentiment recognition; the language model is the encoder part of the overall model, denoted as LM; the overall model is as follows: M_emotion=softmax(LM(C1, C2, …, Cn)), The softmax function outputs the category score output, which is fitted with the label; Repeat the above model training steps, and when the model training accuracy reaches 95%, save the best model to predict the emotion of the dialogue sentence in the text; input the original text C={c1, c2, …, cn}, and the output result is: R=Choice[Max_index(M_emotion(Batch(C)))]; Where Max_index is the serial number of the role with the maximum probability in the center sentence corresponding to the emotion classification set, and the item with the highest model score in the final output candidate list [E1, E2, …, Em] of the emotion classification set is taken as the output result.

14. The speech synthesis method of claim 1, wherein, In step four, according to the role, emotion, and scene in the structured picture book generated in step three, the semantic approximate matching is performed on the word vectors of the audio material label mapping in the audio material library to obtain the candidate audio materials corresponding to each part of the structured picture book, and the most suitable audio material is selected manually.

15. The speech synthesis method of claim 1, wherein, In step four, the structured picture book tone matching includes the following steps: Step 41, special effect scene text, character features, novel label features in the structured picture book; wherein the novel label features come from the classification attributes of the novel in the self-built novel label library; the special effect scene text, character features, novel label features are respectively matched with special effect audio, character voice library, and background audio to obtain special effect audio, character voice, and background audio; Step 42, after matching, the voice of the character is synthesized; Step 43, the special effect audio, synthesized character voice, and background audio are formatted, and the formatted audio is output.

16. The speech synthesis method of claim 15, wherein, the special effect audio refers to the use of a special effect audio library with more than 30 categories and more than 500 types, the special effect audio library is a self-built library, wherein the materials come from the Internet, the positions of the special effect texts in the text segment are extracted through keywords and scene semantic extraction, and accurate track matching is performed; the character voice library matching refers to assigning the most suitable voice in the voice library to the character by analyzing the syntax dependency and mining the character personality characteristics from the user interactive comments; wherein the character personality characteristics information comes from the character mining part in the text content preprocessing process, and the gender, age, and other features of the characters in the article are extracted through the NER task; the background audio matching refers to matching the background music library through the book classification information in the book library to improve the audio background atmosphere and enhance the sense of immersion; and / or, in step 42, the character voice synthesis refers to inputting the character associated text and the character voice information through the text-to-speech interface to output the audio of the text segment; and / or, in step 43, the audio formatting unification refers to using uniform frequency processing to unify the character voice, special effect audio, and background audio in terms of listening experience.

17. The speech synthesis method of claim 1, wherein, In step five, the Chinese TTS engine and / or English TTS engine matched with the content of the text segment are called, and corresponding audio occurs when the character's emotion changes and / or a certain action occurs in the text scene, and the speech is output at the required output speed; and / or, In step five, the fusion of audio includes the following steps: Step 51, the formatted audio obtained including special effect audio, character voice, and background audio is fused into multiple audio tracks; Step 52, the fused multi-track audio is unified in audio format; Step 53, the unified audio is output to obtain a complete audio book audio.

18. The speech synthesis method of claim 17, wherein, The audio multi-track fusion refers to combining multiple audio tracks into a new audio containing all audio tracks, so that the special effect audio and background audio are integrated into the character voice to form a complete chapter audio output; The audio format unification refers to setting the synthesized audio to a unified audio format to maintain the uniformity of audio frequency and sound quality.

19. A system for implementing the speech synthesis method according to any one of claims 1 to 18, characterized in that, The system includes a text input module, a text content preprocessing module, a structured picture book module, a voice matching module, an audio fusion module, and a speech output module; wherein, the text input module is used to input the text segment of the novel; the text content preprocessing module is used to generate formatted text; the structured picture book module is used to perform paragraph structured analysis on the text segment to generate a structured picture book; The timbre matching module is used for matching the mined character attribute with a timbre library to select the timbre information suitable for the character; The audio fusion module is used for audio multi-track fusion of the formatted audios of the special effect sound, the character sound and the background sound to obtain complete audio of the audio book; The voice output module is used for outputting the final audio.

Citation Information

Patent Citations

  • Voice synthesis method, apparatus and device and storage medium

    CN109523986A

  • Video special effect adding method and device

    CN104703043A

  • Text-to-speech conversion method, device, electronic equipment and storage medium

    CN112765971A