A speech synthesis method and system fusing semantic information

By constructing a speech synthesis model that integrates semantic information, the problem of explicit prosodic representation extraction error is solved, and the prosodic naturalness and speaker similarity of speech synthesis are improved.

CN116469368BActive Publication Date: 2025-12-12GUANGZHOU JIUSI INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310386199.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-11
Publication Date
2025-12-12
Estimated Expiration
2043-04-11

AI Technical Summary

Technical Problem

Existing techniques for extracting explicit prosodic representations are prone to errors and cannot connect different prosodic representations, resulting in low prosodic naturalness.

Method used

A speech synthesis model integrating semantic information is constructed, including a phoneme encoder, a word encoder, a word-phoneme attention, an encoder, a variable adapter, and a Mel spectrum decoder. By preprocessing and training the speech data, the error in prosodic representation extraction is reduced and the naturalness of the prosodicity is improved.

Benefits of technology

By incorporating semantic information into the speech synthesis model, errors in explicit prosodic modeling are reduced, and the prosodic naturalness and speaker similarity of the synthesized speech are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116469368B_ABST
    Figure CN116469368B_ABST
Patent Text Reader

Abstract

The application provides a speech synthesis method based on fused semantic information, comprising the following steps: data preparation, collecting speech data and corresponding text, pre-processing the data to extract phoneme sequence, duration, energy and fundamental frequency data; model construction, constructing a speech synthesis model fusing semantic information; model training, inputting phoneme sequence, duration, energy and fundamental frequency data corresponding to real audio to train the model to obtain a trained model; model testing, evaluating and scoring the trained model from two aspects of synthesized speech quality and similarity with a speech speaker, and if the scoring result meets the condition, the model is reserved. The technical scheme of the application can reduce the error in the extraction of rhythm representation in explicit rhythm modeling, and improve the rhythm naturalness.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of speech synthesis, in particular to a speech synthesis method and system fusing semantic information. BACKGROUND

[0002] Speech synthesis technology is also known as text-to-speech (TTS), which aims to convert human understandable fluent speech by inputting text into a device. This technology is an important part of human-computer interaction and has been widely used in voice assistants, voice navigation, intelligent robots, etc. In recent years, the development of deep neural networks has made end-to-end speech synthesis methods the mainstream of research, and the synthesized speech of many excellent models has been very close to human voice. However, most of the synthesized speech still has problems such as mechanical sound, unclear pause, and lack of expressiveness.

[0003] In order to synthesize speech closer to life, in the current mainstream method, prosody modeling is a relatively effective method. The purpose of prosody modeling is to enable the model to learn the ability to predict prosodic representations. There are two different methods of prosody modeling. For the predicted or extracted prosodic representations, if there is an actual physical meaning and can be obtained directly by processing the data, for example, duration, energy, and fundamental frequency, it is called explicit prosody modeling. If there is no obvious physical meaning and needs to be obtained through an additional feature extractor, it is called implicit prosody modeling. The other is explicit prosody modeling, which introduces duration, energy, and fundamental frequency to improve the expressiveness of synthesized speech, but still has the problem that explicit prosodic representation extraction is prone to errors and cannot link different prosodic representations. The prosodic naturalness still has room for improvement. SUMMARY

[0004] The embodiments of the present application provide a related technical solution of a method and device for speech synthesis based on fusing semantic information, to solve the technical problem of explicit prosodic representation extraction being prone to errors and being unable to link different prosodic representations, and low prosodic naturalness in the prior art.

[0005] The embodiments of the present application provide a method for speech synthesis based on fusing semantic information, which includes the following steps:

[0006] Data preparation, collecting speech data and corresponding text, pre-processing the data to extract phoneme sequence, duration, energy, and fundamental frequency data;

[0007] Model construction, constructing a speech synthesis model fusing semantic information;

[0008] Model training, inputting the phoneme sequence, duration, energy, and fundamental frequency data corresponding to the real audio to train the model to obtain a trained model;

[0009] The model test evaluates and scores the trained model from two aspects of synthesized speech quality and similarity to a speech speaker, and if the scoring result meets a condition, the model is used for speech synthesis.

[0010] Further, the synthesis model comprises a phoneme encoder, a word encoder, a word-phoneme attention, an encoder, a variable adapter, and a mel-spectrogram decoder; the phoneme encoder is connected to the word encoder, the phoneme encoder and the word encoder are respectively connected to the word-phoneme attention, the word-phoneme attention is connected to the encoder, the encoder is connected to the variable adapter, and the variable adapter is connected to the mel-spectrogram decoder.

[0011] Further, the Chinese dataset AISHELL-3 and the English dataset LibriTTS are collected and arranged, and the Chinese dataset AISHELL-3 and the English dataset LibriTTS are preprocessed respectively, the Chinese text corresponding to the AISHELL-3 audio is processed to obtain a text sequence in pinyin.

[0012] Further, the Adam optimizer is used in the model training.

[0013] Further, the scoring comprises obtaining scores of synthesized audios of 10 different speakers, and the scores of the 10 persons are averaged again to obtain a final model score.

[0014] The embodiment of the application further provides a device for speech synthesis based on fused semantic information, comprising:

[0015] A data collection module collects speech data and corresponding text, and pre-processes the data to extract phoneme sequences, durations, energies, and fundamental frequency data;

[0016] A model construction module constructs a speech synthesis model fused with semantic information;

[0017] A model training module inputs phoneme sequences, durations, energies, and fundamental frequency data corresponding to real audios to train the model to obtain a trained model;

[0018] A model test module evaluates and scores the trained model from two aspects of synthesized speech quality and similarity to a speech speaker, and if the scoring result meets a condition, the model is used for speech synthesis.

[0019] Further, the synthesis model comprises a phoneme encoder, a word encoder, a word-phoneme attention, an encoder, a variable adapter, and a mel-spectrogram decoder; the phoneme encoder is connected to the word encoder, the phoneme encoder and the word encoder are respectively connected to the word-phoneme attention, the word-phoneme attention is connected to the encoder, the encoder is connected to the variable adapter, and the variable adapter is connected to the mel-spectrogram decoder.

[0020] Further, the scoring includes obtaining a synthesized audio score of 10 different speakers, and averaging the scores of the 10 persons again to obtain a final model score.

[0021] The embodiment of the present application also provides a computer device, comprising a memory and a processor, the memory stores a computer program, and the processor implements the steps of the foregoing method when executing the computer program.

[0022] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the foregoing method.

[0023] The embodiment provided by the present application has at least the following beneficial effects:

[0024] The present application provides a speech synthesis method based on fused semantic information, comprising the following steps: data preparation, collecting speech data and corresponding text, pre-processing the data to extract phoneme sequence, duration, energy and fundamental frequency data; model construction, constructing a speech synthesis model fused with semantic information; model training, inputting the phoneme sequence, duration, energy and fundamental frequency data corresponding to the real audio to train the model to obtain a trained model; model testing, evaluating and scoring the trained model from two aspects of synthesized speech quality and speech speaker similarity, and if the scoring result meets the condition, using the model for speech synthesis. The technical solution of the present application can reduce the error in the extraction of prosody representation in explicit prosody modeling, and improve the prosody naturalness. BRIEF DESCRIPTION OF DRAWINGS

[0025] The accompanying drawings, which are included to provide a further understanding of the present application, form a part of the present application and illustrate the illustrative embodiments of the present application and together with the description serve to explain the present application. In the drawings:

[0026] Figure 1 A schematic diagram of a speech synthesis method based on fused semantic information provided by the embodiment of the present application;

[0027] Figure 2 A module schematic diagram of a speech synthesis model fused with semantic information provided by the embodiment of the present application;

[0028] Figure 3 A module schematic diagram of word-phoneme attention fused with semantic information provided by the embodiment of the present application;

[0029] Figure 4 A schematic diagram of an improved FFT structure in an encoder provided by the embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in combination with specific embodiments of the present application and corresponding drawings. Obviously, the described embodiments are only some of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.

[0031] As shown in the Figures 1-4 The embodiment of the present application provides a speech synthesis method based on fused semantic information, which comprises the following steps:

[0032] Data preparation, collecting speech data and corresponding text, pre-processing the data to extract phoneme sequence, duration, energy, and fundamental frequency data;

[0033] Collect and arrange the Chinese data set AISHELL-3 and the English data set LibriTTS, pre-process the two data sets, perform pinyinization processing on the Chinese text corresponding to the AISHELL-3 audio to obtain a text sequence in pinyin, and mark the tones as 1, 2, 3, 4, and light tone 5. Collect and select suitable phoneme dictionaries for two kinds of speech, and convert the pinyin text sequence and the English text sequence into phoneme sequences according to the phoneme dictionary.

[0034] Process the obtained audio, first, according to the parameters in Table 1, resample the speech data to convert its sampling rate to 22050HZ, pre-emphasize the resampled speech data by 0.97, then perform a short-time Fourier transform with a frame shift of 256, a window length and a frame length of 1024, use a mel filter bank with a number of 80 mel filters to raise from a minimum frequency of 0 to a maximum frequency of 8000, which is consistent with the used HiFi-GAN vocoder.

[0035] Table 1 Audio parameters

[0036]

[0037]

[0038] Model construction, construct a speech synthesis model fused with semantic information;

[0039] As shown in the Figure 2As shown, the speech synthesis model fused with semantic information includes a phoneme encoder, a word encoder, a word-phoneme attention, an encoder, a variance adaptor, and a mel-spectrogram decoder

[0040] The phoneme encoder is composed of an Embedding part and a standard FFT block, which takes phonemes as input, encodes them into phoneme sequences using one-hot encoding, and converts the input into 256-dimensional hidden features through the Embedding operation. The hidden features are processed by the FFT block to obtain the latent variable X p .

[0041] The structure of the word encoder is similar to that of the phoneme encoder and is composed of a standard FFT block. The purpose of the word encoder is to process the phoneme sequence through additional word boundary information to obtain the semantic information in the sentence. Most models only take explicit features such as phonemes as input, and such models cannot extract the syntactic structure information in the text. Therefore, extracting the syntactic information of the sentence is of great significance for synthesizing audio with high expressiveness. The word encoder takes the output X p of the phoneme encoder and the word boundary information as input, first concatenates the two inputs, performs word-level pooling, and then processes the pooled features through a standard FFT to obtain the latent variable X w .

[0042] The word-phoneme attention has a structure as shown in Figure 3 It is composed of a duration predictor, a length regulator, and a multi-head attention. The input of the word-phoneme attention includes the word boundary information (word Boundary), the output X p of the phoneme encoder, and the output X w of the word encoder. The duration predictor takes the word Boundary and X p as input and is composed of two one-dimensional convolutions, a ReLU activation function, a layer normalization, and a linear layer. The duration predictor projects the hidden state at each time step to a scalar to obtain the predicted phoneme time. The length regulator takes X w as input and is set up in the same structure as FastSpeech, which is used to control the speech rate and partial rhythm. The multi-head attention has two heads and takes the output Q of the length regulator and X pThe copied K and V are used as input to obtain the latent variable X. The formula for the word-phoneme attention part is shown below:

[0043] X D =DP(X) P +word)

[0044] X L =LR(X w +X D )

[0045] X = attention(Q, K, V)

[0046] In the formula X p For the output of the phoneme encoder; X w This is the output of the word encoder, where `word` represents word boundary information; `DP` represents the correlation calculation of the duration predictor, and `X` represents the output of the word encoder. D The output of the duration predictor; LR is the correlation operation of the length adjuster; X L Q is the output of the length adjuster; Q is the output of X. L Replication; K, V are X P Copying.

[0047] The encoder consists of four improved FFT blocks, and the improved FFT structure is as follows: Figure 4 As shown, the output X of the phoneme encoder p With the output X of the word encoder w After fusion via WP Attention, a latent variable X containing textual features is obtained. The encoder takes the latent variable X as input, makes three copies named Q, K, and V, and feeds them into a Multi-HeadAttention with a hidden dimension of 256 to obtain the latent variable X. a X a After linear transformation, the result is added to X, and after layer normalization, the first hidden feature X is obtained. 第一隐藏特征 X 第一隐藏特征 The hidden feature X is obtained by feeding it into the newly added bidirectional GRU (Bi-GRU). GRU In this way, the shortcomings of self-attention in focusing on local features can be compensated for, thereby improving the ability to model local dependencies. Then X... GRU After processing by a feedforward neural network (FFN), and X GRU The sums are then normalized to obtain the final output variable X. E The FFN consists of a 1D convolution with a kernel size of 9 and a latent variable dimension of 1024; a ReLU activation function; and a 1D convolution with a kernel size of 1 and a latent variable dimension of 256. The encoder formula is shown below:

[0048] X = WP(X p + X W + word)

[0049] X a = attention(Q, K, V)

[0050] X 第一隐藏特征 = LN(X + Linear(X a ))

[0051] X GRU = BiGRU(X 第一隐藏特征 )

[0052] X E = LN(X GRU + FFN(X GRU ))

[0053] where word represents word boundary information, X p represents phoneme features; X w represents text features; WP represents the relevant operation of word-phoneme attention; X represents the fused features; attention represents the relevant operation of Multi-HeadAttention, X a represents the hidden features after attention operation; Linear represents linear processing; LN represents layer normalization; X 第一隐藏特征 represents the first hidden features; BiGRU represents the relevant operation of bidirectional GRU; X GRU represents the hidden features after bidirectional GRU processing; FFN represents the relevant operation of feedforward neural network; X E represents the output of the final FFT block, that is, the output of the Encoder.

[0054] Variable adapter and mel-spectrogram decoder. Following the settings of the baseline model, the variable adapter is composed of a duration predictor, a pitch predictor, and an energy predictor. The structures of the three predictors are the same, and each is composed of one-dimensional convolution, ReLU activation function, LN, Dropout, one-dimensional convolution, ReLU activation function, LN, Dropout, linear layer 9 parts of linear. The mel-spectrogram decoder is composed of 6 standard FFT blocks.

[0055] Model training, input the phoneme sequence, duration, energy, and fundamental frequency data corresponding to the real audio to train the model to obtain a trained model; the input of the model is the real audio, the corresponding phoneme sequence, duration, energy, and fundamental frequency data, the Adam optimizer is used in the training, and the beta values are 0.9 and 0.98. The mean square error (MSE) is used to calculate the loss of the predicted mel spectrum and the real mel spectrum, and the MAE loss is used for the duration, energy, and fundamental frequency information, and the weight coefficient is 1. The training data is 218 speakers, 85 hours of Chinese data set AISHELL-3, and 2456 speakers, 585 hours of English data set LibriTTS (randomly selecting 400 people, about 100 hours), the data set is divided into a training set, a validation set, and a test set, the model is trained by using the training sets of the two data sets, the batch size during the training is 64, the audio data in each batch is randomly selected, the vocoder uses the commonly used trained HiFi-GAN model, and the models on the two data sets are converged after 100,000 steps of iterative training to obtain the trained model. When the model predicts and generates audio, the duration, energy, and fundamental frequency are the predicted data obtained according to the text and the phoneme sequence.

[0056] Model testing, the trained model is evaluated and scored from two aspects of synthetic speech quality and similarity to the speech speaker, if the scoring result meets the condition, the model is used for speech synthesis. First, 5-10 audios of a person in the model test set are randomly selected, 1 audio is marked as a standard audio, 15-20 audios of the same speaker are synthesized by using random sentences, and all the audios except the standard audio are mixed. Ten native speakers are invited, and they are told to use the international standard 5-point scoring system, from 0 to 5, in order: very poor, the emotion is completely inappropriate to the target emotion, and the emotional expressiveness is extremely poor; poor, the emotion is basically appropriate to the target emotion, and the emotional expressiveness is extremely poor; medium, the emotion is relatively appropriate to the target emotion, and the emotional expressiveness is relatively good; good, the emotion is relatively appropriate to the target emotion, and the emotional expressiveness is sufficient; excellent, the emotion is appropriate to the target emotion, and the emotional expressiveness is outstanding; and every 0.5 points is an interval. First, the standard audio is listened to, and then the mixed audio is scored according to the standard audio, and the scores of the generated audios are averaged. The process is repeated to obtain the scores of the synthetic audios of 10 different speakers. The scores of the 10 people are averaged again to obtain the final model score.

[0057] The embodiment of the application also provides a device for voice synthesis based on fused semantic information, comprising:

[0058] A data collection module collects voice data and corresponding text, and pre-processes the data to extract phoneme sequences, durations, energies, and fundamental frequency data;

[0059] A model construction module constructs a voice synthesis model fused with semantic information.

[0060] a model training module, which trains a model by inputting phoneme sequences, time lengths, energy, and fundamental frequency data corresponding to real audio to obtain a trained model;

[0061] a model testing module, which evaluates and scores the trained model from two aspects of synthesized speech quality and similarity to a speech speaker, and retains the model if the scoring result meets a condition.

[0062] Further, the synthesized model comprises a phoneme encoder, a word encoder, a word-phoneme attention, an encoder, a variable adapter, and a mel-spectrum decoder; the phoneme encoder is connected to the word encoder, the phoneme encoder and the word encoder are respectively connected to the word-phoneme attention, the word-phoneme attention is connected to the encoder, the encoder is connected to the variable adapter, and the variable adapter is connected to the mel-spectrum decoder.

[0063] Further, a Chinese dataset AISHELL-3 and an English dataset LibriTTS are collected and processed, the Chinese text corresponding to the AISHELL-3 audio is processed by pinyinization to obtain a text sequence in pinyin.

[0064] Further, the model training uses an Adam optimizer.

[0065] Further, the scoring comprises obtaining scores of synthesized audio of 10 different speakers, and averaging the scores of the 10 persons to obtain a final model score.

[0066] The embodiment of the present application also provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps of the foregoing method when executing the computer program.

[0067] The embodiment of the present application also provides a computer readable storage medium, which stores a computer program, and the computer program is characterized in that the computer program is executed by a processor to implement the steps of the foregoing method.

[0068] It should be further noted that the terms “comprising”, “including”, or any other variant thereof are intended to cover non-exclusive inclusion, so that processes, methods, articles, or devices including a series of elements not only include those elements, but also include other elements not explicitly listed, or further include elements inherent in such processes, methods, articles, or devices. Without more limitations, the element defined by the statement “comprising a” does not exclude the presence of additional identical elements in the process, method, article, or device including the element.

[0069] Those skilled in the art will appreciate that embodiments of the present application can be devised for a method, a system, or a computer program product. Accordingly, the present application can take the form of an entirely hardware embodiment, an entirely software embodiment or an embodiment combining software and hardware aspects. Furthermore, the present application can take the form of a computer program product on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROMs, optical storage devices, etc.) embodying computer-readable program code thereon for use by or in connection with an instruction execution system. For the purposes of this description, a computer-usable or computer readable storage medium can be any apparatus that can contain, store, communicate, propagate, or transport the program for use by or in connection with the instruction execution system, apparatus, or device. The medium can be electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system (or apparatus or device) or a propagation medium. The computer-usable or computer readable program code can be downloaded from an Internet website, server, or other remote source via a network or a data stream communication path. From the Internet website, server, or other remote source, the code can be downloaded into the instruction execution system, apparatus, or device where execution of the same can take place. The present application is directed to any number and type of computer-usable storage media, apparatuses, and devices self-evidently known to one of ordinary skill in the art.

[0070] The foregoing is merely illustrative of the principles of the application and various modifications can be made by persons skilled in the art. The present application is not limited to the embodiments described above, but encompasses any and all modifications within the scope and equivalents of the following claims.

Claims

1. A method of speech synthesis based on fusion of semantic information, characterized by, The method comprises the following steps: data preparation, collecting voice data and corresponding text, pre-processing the data to extract phoneme sequence, duration, energy and fundamental frequency data; Model construction involves building a speech synthesis model that integrates semantic information. This model includes a phoneme encoder, a word encoder, a word-phoneme attention mechanism, an encoder, a variable adapter, and a Mel-spectrum decoder. The phoneme encoder is connected to the word encoder, and both are connected to the word-phoneme attention mechanism. The word-phoneme attention mechanism is connected to the encoder, which in turn is connected to the variable adapter, which is then connected to the Mel-spectrum decoder. The word-phoneme attention mechanism consists of a duration predictor, a length adjuster, and a multi-head attention mechanism. The input to the word-phoneme attention mechanism includes word boundary information and the output X of the phoneme encoder. P With the output X of the word encoder W As input; the duration predictor uses word boundary information, X p As input, it consists of two one-dimensional convolutions, a ReLU activation function, layer normalization, and a linear layer. The duration predictor projects the hidden state at each time step onto a scalar to obtain the predicted phoneme time; the length adjuster uses X... W As input, the architecture is consistent with FastSpeech, used to control speech rate and partial prosody; multi-head attention has two heads, one for the output Q of the length modulator, and the other for the output X. P The copied K and V are used as input to obtain the latent variable X; Model training: inputting the phoneme sequence, duration, energy and fundamental frequency data corresponding to the real audio into the model to train the model and obtain a trained model; Model testing: evaluating and scoring the trained model from two aspects of synthesized voice quality and voice speaker similarity, and if the scoring result meets the condition, using the model for voice synthesis.

2. The method of speech synthesis based on fusion semantic information according to claim 1, characterized in that, The Chinese data set AISHELL-3 and the English data set LibriTTS are collected and arranged, and the Chinese data set AISHELL-3 and the English data set LibriTTS are pre-processed respectively, the Chinese text corresponding to the AISHELL-3 audio is processed by pinyin, and a text sequence with pinyin as the text is obtained.

3. The method of speech synthesis based on fusion semantic information according to claim 1, characterized in that, The Adam optimizer is used in the model training.

4. The method of speech synthesis based on fusion semantic information according to claim 1, characterized in that, The scoring includes obtaining the synthesized audio scores of 10 different speakers, and averaging the scores of the 10 persons again to obtain the final model score.

5. An apparatus for speech synthesis based on fusion of semantic information, characterized by, The method comprises the following steps: A data collection module is configured to collect voice data and corresponding text, pre-process the data to extract phoneme sequence, duration, energy and fundamental frequency data; The model building module constructs a speech synthesis model that integrates semantic information. This model includes a phoneme encoder, a word encoder, a word-phoneme attention module, an encoder, a variable adapter, and a Mel-spectrum decoder. The phoneme encoder is connected to the word encoder, and both are connected to the word-phoneme attention module. The word-phoneme attention module is connected to the encoder, which in turn is connected to the variable adapter, which is connected to the Mel-spectrum decoder. The word-phoneme attention module consists of a duration predictor, a length adjuster, and a multi-head attention mechanism. The input to the word-phoneme attention module includes word boundary information and the output X of the phoneme encoder. P With the output X of the word encoder W As input; the duration predictor uses word boundary information, X p As input, it consists of two one-dimensional convolutions, a ReLU activation function, layer normalization, and a linear layer. The duration predictor projects the hidden state at each time step onto a scalar to obtain the predicted phoneme time; the length adjuster uses X... W As input, the architecture is consistent with FastSpeech, used to control speech rate and partial prosody; multi-head attention has two heads, one for the output Q of the length modulator, and the other for the output X. P The copied K and V are used as input to obtain the latent variable X; A model training module is configured to input the phoneme sequence, duration, energy and fundamental frequency data corresponding to the real audio into the model to train the model and obtain a trained model; A model testing module is configured to evaluate and score the trained model from two aspects of synthesized voice quality and voice speaker similarity, and if the scoring result meets the condition, use the model for voice synthesis.

6. The apparatus for speech synthesis based on fusion semantic information according to claim 5, wherein, The scoring includes obtaining the synthesized audio scores of 10 different speakers, and averaging the scores of the 10 persons again to obtain the final model score. 7.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-6 when the computer program is executed by the processor. The processor executes the computer program to realize the steps of the method in any one of claims 1 to 4.

8. A computer-readable storage medium, characterized in that, A computer program product is provided, and a computer program is stored in the computer program product, wherein the computer program is executed by a processor to realize the steps of the method in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Voice recognition method, device, system and language switching system

    CN104637482A

  • Rhythm-controllable Chinese and English mixed speech synthesis method and system

    CN112802450A