Artificial intelligence-based speech synthesis method and device, computer equipment and medium

By employing an AI-based speech synthesis method, utilizing feature extraction, stress and pause predictors, and combining them with a prosody predictor, the problem of low accuracy of prosodic attribute features in existing technologies is solved, thereby improving the accuracy and naturalness of speech synthesis.

CN116580698BActive Publication Date: 2026-05-01PING AN TECH (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
PING AN TECH (SHENZHEN) CO LTD
Filing Date
2023-06-16
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing speech synthesis methods struggle to accurately reflect prosodic features such as pauses and stresses in text, resulting in low accuracy of synthesized speech.

Method used

An AI-based speech synthesis method is adopted. By using a trained feature extraction model, stress predictor, pause predictor and prosody predictor, the stress, pause and prosody tags of the text feature vector are predicted and combined respectively, and the text stress vector, pause vector and prosody vector are output for phoneme conversion and speech synthesis.

Benefits of technology

It improves the accuracy of prosodic and emotional representation in synthesized speech, enhances the expressiveness and naturalness of speech, and improves the accuracy of speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116580698B_ABST
    Figure CN116580698B_ABST
Patent Text Reader

Abstract

The application is suitable for the technical field of speech synthesis, and particularly relates to a speech synthesis method and device based on artificial intelligence, computer equipment and a medium. The application extracts a text feature vector of a target text through a feature extraction model, predicts the text feature vector through a stress predictor, outputs a stress prediction vector, adds the stress prediction vector to the text feature vector to obtain a text stress vector, predicts the text feature vector through a pause predictor, outputs a pause prediction vector, adds the pause prediction vector to the text feature vector to obtain a text pause vector, predicts the text stress vector and the text pause vector through a prosody predictor, outputs a text prosody vector, matches the text prosody vector with a phoneme sequence of the target text, obtains a phoneme sequence with prosody labels, performs speech conversion on the phoneme sequence with prosody labels, obtains synthesized speech, and through the prediction of stress, pause and prosody, the expressiveness, naturalness and accuracy of the synthesized speech are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of speech synthesis technology, and in particular to a speech synthesis method, apparatus, computer equipment and medium based on artificial intelligence. Background Technology

[0002] The goal of speech synthesis technology is to convert text information into speech signals. There is a one-to-many mapping between text and speech data. Speech data not only contains the corresponding text information, but also the corresponding prosodic information. Modeling the prosodic information is crucial for synthesizing natural and expressive speech.

[0003] Existing speech synthesis methods generally extract text features from the target text and predict prosodic attributes such as pitch and duration of the synthesized speech based on these features. The text features and prosodic attributes are then concatenated and decoded to synthesize the target speech. However, speech contains information such as pauses and stresses that affect prosodic attributes, while text features focus on representing textual content and cannot accurately reflect pauses and stresses in speech from the target text. Therefore, when the aforementioned methods directly extract prosodic attributes from text features, the accuracy of these features is low, resulting in low accuracy of the synthesized speech.

[0004] Therefore, in the field of speech synthesis technology, how to improve the accuracy of synthesized speech has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a speech synthesis method, apparatus, computer device and medium based on artificial intelligence to solve the problem of low accuracy of speech synthesized by existing speech synthesis methods.

[0006] In a first aspect, embodiments of the present invention provide an artificial intelligence-based speech synthesis method, the speech synthesis method comprising:

[0007] The target text is acquired, and the trained feature extraction model extracts features from the target text, outputting a text feature vector.

[0008] The text feature vector is used by a trained stress predictor to predict stress labels and output a stress prediction vector. The stress prediction vector is then added to the text feature vector to obtain the text stress vector.

[0009] The text feature vector is used by a trained pause predictor to predict pause labels and output a pause prediction vector. The pause prediction vector is then added to the text feature vector to obtain the text pause vector.

[0010] The text stress vector and the text pause vector are used by a trained prosodic predictor to predict prosodic labels and output text prosodic vectors.

[0011] The target text is converted into phonemes to obtain a phoneme sequence corresponding to the target text. The text prosody vector is matched with the phoneme sequence to obtain a phoneme sequence with prosody tags. The phoneme sequence with prosody tags is converted into speech to obtain synthesized speech.

[0012] Secondly, embodiments of the present invention provide an artificial intelligence-based speech synthesis device, the speech synthesis device comprising:

[0013] The feature extraction module is used to acquire target text. The trained feature extraction model extracts features from the target text and outputs a text feature vector.

[0014] The stress prediction module is used to predict stress labels by the trained stress predictor on the text feature vector, output stress prediction vector, and add stress prediction vector and text feature vector to obtain text stress vector.

[0015] The pause prediction module is used to predict pause labels by the trained pause predictor on the text feature vector, output a pause prediction vector, and add the pause prediction vector to the text feature vector to obtain the text pause vector.

[0016] The prosody prediction module is used to predict prosody tags by using the text stress vector and the text pause vector through a trained prosody predictor, and outputs the text prosody vector.

[0017] The speech synthesis module is used to perform phoneme conversion on the target text to obtain a phoneme sequence corresponding to the target text, match the text prosody vector with the phoneme sequence to obtain a phoneme sequence with prosody tags, and perform speech conversion on the phoneme sequence with prosody tags to obtain synthesized speech.

[0018] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech synthesis method as described in the first aspect.

[0019] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the speech synthesis method as described in the first aspect.

[0020] The beneficial effects of this invention compared to existing technologies are as follows: By acquiring target text, a trained feature extraction model extracts features from the target text, outputting a text feature vector. This text feature vector is then used by a trained stress predictor to predict stress labels, outputting a stress prediction vector. The stress prediction vector is added to the text feature vector to obtain a text stress vector. The text feature vector is then used by a trained pause predictor to predict pause labels, outputting a pause prediction vector. This pause prediction vector is added to the text feature vector to obtain a text pause vector. Finally, the text stress vector and text pause vector are used by a trained prosody predictor to predict prosody labels, outputting a text prosody vector. This process is repeated for each text segment. The target text is phoneme-converted to obtain the corresponding phoneme sequence. The text prosody vector is matched with the phoneme sequence to obtain a phoneme sequence with prosody tags. The phoneme sequence with prosody tags is then converted into speech to obtain synthesized speech. The text feature vector is predicted with stress and pause tags by a trained stress predictor and pause predictor, and stress prediction vector and pause prediction vector are output to predict stress and pause features in the synthesized speech. These are used as the basis for prosody prediction to obtain the text prosody vector. By improving the accuracy of prosodic emotion representation, the expressiveness and naturalness of the synthesized speech are improved, thereby improving the accuracy of the synthesized speech. Attached Figure Description

[0021] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of an application environment for an artificial intelligence-based speech synthesis method provided in Embodiment 1 of the present invention;

[0023] Figure 2 This is a flowchart illustrating an artificial intelligence-based speech synthesis method provided in Embodiment 1 of the present invention.

[0024] Figure 3 This is a schematic diagram of the model structure of an artificial intelligence-based speech synthesis method provided in Embodiment 1 of the present invention;

[0025] Figure 4 This is a schematic diagram of the structure of an artificial intelligence-based speech synthesis device provided in Embodiment 2 of the present invention;

[0026] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Detailed Implementation

[0027] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.

[0028] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0029] It should also be understood that the term "and / or" as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0030] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."

[0031] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0032] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0033] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.

[0034] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.

[0035] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0036] To illustrate the technical solution of the present invention, specific embodiments are described below.

[0037] The first embodiment of this invention provides an artificial intelligence-based speech synthesis method, which can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud computing devices, and personal digital assistants (PDAs). The server can be implemented using a standalone server or a server cluster consisting of multiple servers.

[0038] See Figure 2 This is a flowchart illustrating an artificial intelligence-based speech synthesis method provided in Embodiment 1 of the present invention. The speech synthesis method described above can be applied to... Figure 1 In a client application, the speech synthesis method may include the following steps:

[0039] Step S201: Obtain the target text. The trained feature extraction model extracts features from the target text and outputs the text feature vector.

[0040] The target text can be text information represented in text form or text information represented in phoneme form. The goal of speech synthesis technology is to convert text information into speech signals. The synthesized speech signals are based on text information and based on extracted and predicted prosodic information, providing highly human-like, fluent and natural speech synthesis services. It is widely used in many fields such as the Internet, finance, medical care, and education.

[0041] See Figure 3 For the target text to be synthesized, the trained feature extraction model extracts features from the target text and outputs a text feature vector, which serves as the basis for speech synthesis.

[0042] For example, due to the complexity and diversity of financial business, simple tasks such as a large number of consultation and after-sales services can seriously occupy the energy and time of business personnel, reducing their work efficiency and quality. However, adopting intelligent conversation methods based on automatic speech synthesis can save a lot of labor costs. At the same time, the quality of customer service can be improved by effectively controlling the content of voice conversations. Therefore, speech synthesis technology can play an important auxiliary role in the financial field.

[0043] This embodiment takes the application of speech synthesis technology in the financial field as an example. Correspondingly, the target text can be the target speech related to the financial field. The target speech to be synthesized is obtained, and the trained feature extraction model first extracts features from the target speech and outputs the corresponding speech feature vector as the basis for speech synthesis.

[0044] Optionally, the trained feature extraction model includes an embedding layer and a spectral analysis layer;

[0045] The trained feature extraction model extracts features from the target text, and the output text feature vector includes:

[0046] The embedding layer performs text embedding on the target text and outputs text embedding features.

[0047] The spectral analysis layer performs spectral analysis on the text embedding features and outputs text feature vectors.

[0048] The trained feature extraction model includes an embedding layer and a spectrum analysis layer. The embedding layer can embed the input target text and output text embedding features to capture the relationships between words in the text in a high-dimensional space. The spectrum analysis layer can perform spectrum analysis on the input text embedding features and output text feature vectors to represent the text features of the target text in the frequency domain.

[0049] This embodiment performs text embedding and spectrum analysis on the target text based on the embedding layer and the spectrum analysis layer, and extracts text feature vectors to represent the text features of the target text in high-dimensional space and frequency domain, thereby improving the accuracy of the text feature vectors extracted from the target text.

[0050] The steps described above—obtaining the target text, training the feature extraction model to extract features from the target text, and outputting the text feature vector—extract the text feature vector of the target text as the basis for speech synthesis, thereby improving the rationality of speech synthesis.

[0051] In step S202, the text feature vector is used by the trained stress predictor to predict stress labels and output stress prediction vector. The stress prediction vector is added to the text feature vector to obtain the text stress vector.

[0052] Among them, stress includes grammatical stress and logical stress. Grammatical stress is the emphasis placed on a component according to the grammatical structure of a sentence; logical stress is used to express the speaker's thoughts and the emphasis of the meaning and emotion to be expressed. It plays an emphasizing role based on the speaker's intention and is not bound by the rules of grammatical stress. Therefore, ensuring the correctness of stress in synthesized speech is of great significance for the accurate representation of the phonological grammatical structure and prosodic emotion of the target speech.

[0053] Since grammatical stress is based on the grammatical structure of a sentence, it can be determined by analyzing the grammatical structure of the target text. However, logical stress is not bound by grammatical rules and cannot be determined by analyzing the grammatical structure of the target text.

[0054] Therefore, see Figure 3 In this embodiment, based on the obtained text feature vector of the target text, the trained stress predictor predicts the stress label of the text feature vector and outputs the stress prediction vector, which is used to predict the stress features in the speech to be synthesized, so as to improve the accuracy of the representation of prosodic emotion in the speech to be synthesized.

[0055] Then, the accent prediction vector and the text feature vector representing the text content information are added together to obtain the text accent vector, which simultaneously represents the text content information and accent feature information, serving as the basis for speech synthesis to improve the accuracy of prosodic emotion representation in the synthesized speech.

[0056] This embodiment takes the application of speech synthesis technology in the financial field as an example. Correspondingly, the trained stress predictor predicts stress labels from the speech feature vector and outputs a speech stress prediction vector, which predicts the stress features in the speech to be synthesized, thereby improving the accuracy of the representation of prosodic emotion in the synthesized speech. Then, the speech stress prediction vector and the speech feature vector representing the speech content information are added together to obtain the speech stress vector, which simultaneously represents the speech content information and the stress feature information.

[0057] Optionally, the trained accent predictor includes a first convolutional layer, a first fully connected layer, a first normalization layer, and a first feature transformation layer;

[0058] The text feature vector is used by a trained stress predictor to predict stress labels, and the output stress prediction vector includes:

[0059] The first convolutional layer extracts features from the text feature vector and outputs an accent feature vector;

[0060] The first fully connected layer classifies the accent feature vectors and outputs the first accent probability vector;

[0061] The first normalization layer normalizes the first accent probability vector and outputs the second accent probability vector.

[0062] The first feature transformation layer performs feature transformation on the second accent probability vector and outputs an accent prediction vector.

[0063] In order to predict accent labels from text feature vectors, this embodiment obtains a trained accent predictor, which includes a first convolutional layer, a first fully connected layer, a first normalization layer, and a first feature transformation layer.

[0064] Specifically, the first convolutional layer extracts stress features from the text feature vector and outputs a stress feature vector. The first fully connected layer classifies the stress feature vector and outputs a first stress probability vector, which represents the first probability of belonging to each stress label. The first normalization layer normalizes the first stress probability vector, normalizing the first probability to a preset probability range, and outputs a second stress probability vector, which represents the second probability of belonging to each stress label. The first feature transformation layer transforms the input second stress probability vector and outputs a lower-dimensional stress prediction vector to reduce the data dimensionality in the model and improve speech synthesis efficiency.

[0065] This embodiment predicts the stress labels of the text feature vector based on the first convolutional layer, the first fully connected layer, the first normalization layer, and the first feature transformation layer, and outputs the stress prediction vector to predict the stress features in the speech to be synthesized, thereby improving the accuracy of the representation of prosodic emotion in the speech to be synthesized.

[0066] The above text feature vectors are used to predict stress labels using a trained stress predictor, outputting a stress prediction vector. The stress prediction vector is then added to the text feature vector to obtain the text stress vector. Based on the trained stress predictor, stress labels are predicted on the text feature vectors, and a stress prediction vector is output. This is used to predict stress features in the synthesized speech, thereby improving the accuracy of prosodic emotion representation in the synthesized speech.

[0067] Step S203: The text feature vector is used by the trained pause predictor to predict pause labels and output a pause prediction vector. The pause prediction vector is added to the text feature vector to obtain the text pause vector.

[0068] Among these, pauses are grammatical pauses and logical pauses. Grammatical pauses can be pauses with punctuation or pauses within a sentence that follow its grammatical structure. Logical pauses, on the other hand, are pauses made by the speaker due to physiological or linguistic needs to emphasize the meaning of a word, occurring within a sentence, at the end of a sentence, within a sentence group, or between paragraphs, and are not bound by grammatical pause rules. Therefore, ensuring the correctness of pauses in synthesized speech is of great significance for the accurate representation of the phonological and grammatical structure and prosodic emotion of the target speech.

[0069] Since grammatical pauses are based on the grammatical structure of sentences, they can be determined by analyzing the grammatical structure of the target text. Logical pauses, on the other hand, are not bound by grammatical rules and cannot be determined by analyzing the grammatical structure of the target text.

[0070] Therefore, see Figure 3 In this embodiment, based on the text feature vector of the target text, the trained pause predictor predicts the pause labels of the text feature vector and outputs the pause prediction vector, which is used to predict the pause features in the speech to be synthesized, so as to improve the accuracy of the representation of prosodic emotion in the speech to be synthesized.

[0071] Then, the pause prediction vector and the text feature vector representing the text content information are added together to obtain the text pause vector, which simultaneously represents the text content information and the pause feature information, serving as the basis for speech synthesis to improve the accuracy of prosodic emotion representation in the synthesized speech.

[0072] This embodiment takes the application of speech synthesis technology in the financial field as an example. Correspondingly, the trained pause predictor predicts pause labels from the speech feature vector and outputs a speech pause prediction vector, which predicts the pause features in the speech to be synthesized, thereby improving the accuracy of the representation of prosodic emotion in the speech to be synthesized. Then, the speech pause prediction vector and the speech feature vector representing the speech content information are added together to obtain the speech pause vector, which simultaneously represents the speech content information and the pause feature information.

[0073] Optionally, the trained pause predictor includes a second convolutional layer, a second fully connected layer, a second normalization layer, and a second feature transformation layer.

[0074] The text feature vectors are used by a trained pause predictor to predict pause labels, and the output pause prediction vector includes:

[0075] The second convolutional layer extracts features from the text feature vector and outputs a pause feature vector;

[0076] The second fully connected layer classifies the pause feature vector and outputs the first pause probability vector;

[0077] The second normalization layer normalizes the first pause probability vector and outputs the second pause probability vector.

[0078] The second feature transformation layer performs feature transformation on the second pause probability vector and outputs a pause prediction vector.

[0079] In order to predict pause labels from text feature vectors, this embodiment obtains a trained pause predictor, including a second convolutional layer, a second fully connected layer, a second normalization layer, and a second feature transformation layer.

[0080] Specifically, the second convolutional layer extracts pause features from the text feature vector and outputs a pause feature vector. The second fully connected layer classifies the pause feature vector and outputs a first pause probability vector, which represents the first probability of belonging to each pause label. The second normalization layer normalizes the first pause probability vector, normalizing the first probability to a preset probability range, and outputs a second pause probability vector, which represents the second probability of belonging to each pause label. The second feature transformation layer transforms the input second pause probability vector and outputs a pause prediction vector with lower dimensionality, thereby reducing the data dimensionality in the model and improving speech synthesis efficiency.

[0081] This embodiment predicts pause labels on text feature vectors based on the second convolutional layer, the second fully connected layer, the second normalization layer, and the second feature transformation layer, and outputs pause prediction vectors to predict pause features in the speech to be synthesized, thereby improving the accuracy of prosodic emotion representation in the speech to be synthesized.

[0082] The above text feature vectors are used to predict pause labels using a trained pause predictor, which outputs a pause prediction vector. The pause prediction vector is then added to the text feature vector to obtain the text pause vector. Based on the trained pause predictor, pause labels are predicted on the text feature vectors, and a pause prediction vector is output. This is used to predict pause features in the synthesized speech, thereby improving the accuracy of prosodic emotion representation in the synthesized speech.

[0083] In step S204, the text stress vector and text pause vector are used by the trained prosody predictor to predict the prosody label and output the text prosody vector.

[0084] In speech synthesis tasks, in addition to extracting text information as the content basis, it is also necessary to extract prosodic information as the prosodic basis in order to improve the expressiveness and naturalness of synthesized speech.

[0085] Text stress vectors and text pause vectors can predict stress and pause information in synthesized speech, respectively, and predict the speaker's intent and emotional emphasis. Therefore, see [link to relevant documentation]. Figure 3 In this embodiment, the trained prosodic predictor predicts the prosodic labels of the text stress vector and the text pause vector, and outputs the text prosodic vector to improve the expressiveness and naturalness of the synthesized speech.

[0086] Among these, prosodic attributes such as pitch, energy, and duration can influence the tone, thoughts, and emotions expressed by the speaker, and are of great significance in demonstrating the expressiveness and naturalness of synthesized speech. Therefore, the prosodic predictor in this embodiment can be set according to actual conditions to predict one or more prosodic attributes such as pitch, energy, or duration, so as to accurately represent the prosodic attribute information of synthesized speech and improve the accuracy of synthesized speech.

[0087] This embodiment takes the application of speech synthesis technology in the financial field as an example. Correspondingly, the trained prosody predictor predicts the prosody tags of the speech stress vector and speech pause vector, and outputs the speech prosody vector to improve the expressiveness and naturalness of the synthesized speech.

[0088] Optionally, the prosody predictor includes a pitch predictor;

[0089] The text stress vector and text pause vector are used by a trained prosodic predictor to predict prosodic labels, and the output text prosodic vector includes:

[0090] The text stress vector and text pause vector are used by a pitch predictor to predict the pitch, and the output text pitch vector is used as the text prosody vector.

[0091] In this embodiment, the tone, thoughts and emotions that the speaker wants to express can be represented by pitch prediction. Therefore, the prosody predictor in this embodiment includes a pitch predictor, which predicts the pitch of the input text stress vector and text pause vector, and outputs the text pitch vector as the text prosody vector.

[0092] This embodiment uses a pitch predictor to predict the pitch of the input text stress vector and text pause vector, and outputs the text pitch vector as the text prosody vector, which effectively represents the tone, thoughts and emotional information that the speaker wants to express, and improves the naturalness and accuracy of the synthesized speech.

[0093] Optionally, the prosody predictor also includes a duration predictor. The text stress vector and text pause vector are used by the trained prosody predictor to predict prosody tags, and the output text prosody vector includes:

[0094] The text stress vector and text pause vector are used by a pitch predictor to predict the pitch, and the output text pitch vector is then generated.

[0095] The text stress vector and text pause vector are used by the duration predictor to predict the duration, and the output text duration vector is then generated.

[0096] By fusing the text pitch vector and the text duration vector, a text prosody vector is obtained.

[0097] In order to improve the ability to represent the tone, thoughts and emotions that the speaker wants to express, the prosody predictor in this embodiment also includes a duration predictor.

[0098] Correspondingly, the pitch predictor predicts the pitch of the text stress vector and the text pause vector, and outputs the text pitch vector. The duration predictor predicts the duration of the text stress vector and the text pause vector, and outputs the text duration vector. Then, the text pitch vector and the text duration vector are fused to obtain the text prosodic vector, so as to improve the naturalness and accuracy of the synthesized speech by combining pitch and duration.

[0099] This embodiment combines text pitch vectors and text duration vectors to obtain text prosodic vectors, which improves the naturalness and accuracy of synthesized speech.

[0100] Optionally, the prosody predictor includes an energy predictor. Correspondingly, the text stress vector and text pause vector are used by the trained prosody predictor to predict prosody tags, and the output text prosody vector includes:

[0101] The text stress vector and text pause vector are used by the energy predictor to predict pitch, and the output text energy vector is used as the text prosody vector.

[0102] Optionally, the prosody predictor includes a duration predictor. Correspondingly, the text stress vector and text pause vector are used by the trained prosody predictor to predict prosody tags, and the output text prosody vector includes:

[0103] The text stress vector and text pause vector are used by the energy predictor to predict the duration, and the output text duration vector is used as the text prosody vector.

[0104] Optionally, the prosody predictor includes a pitch predictor and an energy predictor. Correspondingly, the text stress vector and text pause vector are used by the trained prosody predictor to predict prosody tags, and the output text prosody vector includes:

[0105] The text stress vector and text pause vector are used by a pitch predictor to predict the pitch, and the output text pitch vector is then generated.

[0106] The text stress vector and text pause vector are used by the energy predictor to predict the duration, and the text energy vector is output.

[0107] By fusing the text pitch vector and the text energy vector, a text prosody vector is obtained.

[0108] Optionally, the prosody predictor includes a duration predictor and an energy predictor. Correspondingly, the text stress vector and text pause vector are used by the trained prosody predictor to predict prosody tags, and the output text prosody vector includes:

[0109] The text stress vector and text pause vector are used by the duration predictor to predict the pitch, and the output text pitch vector is then used.

[0110] The text stress vector and text pause vector are used by the energy predictor to predict the duration, and the text energy vector is output.

[0111] By fusing the text duration vector and the text energy vector, a text prosody vector is obtained.

[0112] Optionally, the prosody predictor includes a pitch predictor, a duration predictor, and an energy predictor. Correspondingly, the text stress vector and text pause vector are used by the trained prosody predictor to predict prosody tags, and the output text prosody vector includes:

[0113] The text stress vector and text pause vector are used by a pitch predictor to predict the pitch, and the output text pitch vector is then generated.

[0114] The text stress vector and text pause vector are used by the duration predictor to predict the pitch, and the output text pitch vector is then used.

[0115] The text stress vector and text pause vector are used by the energy predictor to predict the duration, and the text energy vector is output.

[0116] The text pitch vector, text duration vector, and text energy vector are fused to obtain the text prosody vector.

[0117] The above steps, which predict prosodic labels for text stress vectors and text pause vectors using a trained prosodic predictor and output text prosodic vectors, accurately represent the prosodic attribute information of synthesized speech, thereby improving the expressiveness and naturalness of synthesized speech.

[0118] Step S205: Phoneme conversion is performed on the target text to obtain the phoneme sequence of the corresponding target text. The text prosody vector is matched with the phoneme sequence to obtain the phoneme sequence with prosody tags. The phoneme sequence with prosody tags is converted into speech to obtain synthesized speech.

[0119] Among them, a phoneme is the smallest unit of speech that is divided according to the natural attributes of speech. It is analyzed based on the articulation action in a syllable, and one action can constitute a phoneme.

[0120] Because different languages ​​can have the same word with different pronunciations, it is necessary to convert the target text from text form to phoneme form. See [link to relevant documentation]. Figure 3 The target text is converted into phonemes to obtain the corresponding phoneme sequence, thereby improving the consistency between the target text and the synthesized speech and thus improving the accuracy of the synthesized speech.

[0121] The phoneme sequence contains text content information. Therefore, the text prosody vector is matched with the phoneme sequence to obtain a phoneme sequence with prosody labels, which simultaneously represents the content information of the target text and the prosody information of the predicted synthesized speech. Then, by performing speech conversion on the phoneme sequence with prosody labels, the synthesized speech can be obtained, thus completing the speech synthesis task.

[0122] This embodiment takes the application of speech synthesis technology in the financial field as an example. Correspondingly, the target speech is converted into phonemes to obtain the phoneme sequence corresponding to the target speech, so as to improve the consistency between the target speech and the synthesized speech. The speech prosody vector is matched with the phoneme sequence to obtain a phoneme sequence with prosody tags, so as to simultaneously represent the content information of the target speech and the prosody information of the predicted synthesized speech. Then, by converting the phoneme sequence with prosody tags into speech, the synthesized speech can be obtained, thus completing the speech synthesis task in the financial field and assisting business personnel to improve work efficiency and quality.

[0123] The above steps involve phoneme conversion of the target text to obtain a corresponding phoneme sequence, matching the text prosody vector with the phoneme sequence to obtain a phoneme sequence with prosody tags, and performing speech conversion on the phoneme sequence with prosody tags to obtain synthesized speech. Matching the text prosody vector with the phoneme sequence of the target text to obtain a phoneme sequence with prosody tags simultaneously represents the content information of the target text and the prosodic information of the predicted synthesized speech, thereby improving the naturalness and expressiveness of the synthesized speech and thus improving the accuracy of the synthesized speech.

[0124] In this embodiment of the invention, target text is acquired, and a trained feature extraction model extracts features from the target text, outputting a text feature vector. The text feature vector is then used by a trained stress predictor to predict stress labels, outputting a stress prediction vector. This stress prediction vector is added to the text feature vector to obtain a text stress vector. The text feature vector is then used by a trained pause predictor to predict pause labels, outputting a pause prediction vector. This pause prediction vector is added to the text feature vector to obtain a text pause vector. Finally, the text stress vector and text pause vector are used by a trained prosodic predictor to predict prosodic labels, outputting a text prosodic vector. Phoneme conversion is then performed on the target text. The process involves obtaining the phoneme sequence of the target text, matching the text prosody vector with the phoneme sequence to obtain a phoneme sequence with prosody labels, performing speech conversion on the phoneme sequence with prosody labels to obtain synthesized speech, and using trained stress predictors and pause predictors to predict stress and pause labels on the text feature vector, outputting stress prediction vectors and pause prediction vectors to predict stress and pause features in the synthesized speech, and using these as the basis for prosody prediction to obtain the text prosody vector. By improving the accuracy of prosodic emotion representation, the expressiveness and naturalness of the synthesized speech are improved, thereby improving the accuracy of the synthesized speech.

[0125] Corresponding to the speech synthesis method in the above embodiments, Figure 4 A structural block diagram of an artificial intelligence-based speech synthesis device provided in Embodiment 2 of the present invention is given. For ease of explanation, only the parts related to the embodiments of the present invention are shown.

[0126] See Figure 4 The speech synthesis device includes:

[0127] Feature extraction module 41 is used to acquire target text. The trained feature extraction model extracts features from the target text and outputs text feature vectors.

[0128] The stress prediction module 42 is used to predict stress labels by the trained stress predictor on the text feature vector, output stress prediction vector, and add stress prediction vector and text feature vector to obtain text stress vector.

[0129] The pause prediction module 43 is used to predict pause labels by the trained pause predictor on the text feature vector, output the pause prediction vector, and add the pause prediction vector and the text feature vector to obtain the text pause vector.

[0130] The prosody prediction module 44 is used to predict the prosody tags of the text stress vector and text pause vector using the trained prosody predictor, and output the text prosody vector.

[0131] The speech synthesis module 45 is used to convert the target text into phonemes to obtain the corresponding phoneme sequence of the target text, match the text prosody vector with the phoneme sequence to obtain the phoneme sequence with prosody tags, and convert the phoneme sequence with prosody tags into speech to obtain synthesized speech.

[0132] Optionally, the trained feature extraction model includes an embedding layer and a spectral analysis layer, and the aforementioned feature extraction module 41 includes:

[0133] The text embedding submodule is used by the embedding layer to perform text embedding on the target text and output text embedding features.

[0134] The spectrum analysis submodule is used by the spectrum analysis layer to perform spectrum analysis on the text embedding features and output the text feature vector.

[0135] Optionally, the trained accent predictor includes a first convolutional layer, a first fully connected layer, a first normalization layer, and a first feature transformation layer. The accent prediction module 42 includes:

[0136] The accent feature extraction submodule is used to extract features from the text feature vector in the first convolutional layer and output the accent feature vector.

[0137] The first classification submodule is used by the first fully connected layer to classify the accent feature vector and output the first accent probability vector;

[0138] The first normalization submodule is used to normalize the first accent probability vector in the first normalization layer and output the second accent probability vector.

[0139] The first feature transformation submodule is used to perform feature transformation on the second accent probability vector by the first feature transformation layer and output the accent prediction vector.

[0140] Optionally, the trained pause predictor includes a second convolutional layer, a second fully connected layer, a second normalization layer, and a second feature transformation layer. The pause prediction module 43 includes:

[0141] The pause feature extraction submodule is used by the second convolutional layer to extract features from the text feature vector and output the pause feature vector.

[0142] The second classification submodule is used by the second fully connected layer to classify the pause feature vector and output the first pause probability vector.

[0143] The second normalization submodule is used by the second normalization layer to normalize the first pause probability vector and output the second pause probability vector.

[0144] The second feature transformation submodule is used by the second feature transformation layer to perform feature transformation on the second pause probability vector and output the pause prediction vector.

[0145] Optionally, the prosody predictor includes a pitch predictor, and the prosody prediction module 44 mentioned above includes:

[0146] The first pitch prediction submodule is used to predict the pitch of the text stress vector and the text pause vector through the pitch predictor, and outputs the text pitch vector as the text prosody vector.

[0147] Optionally, the prosody predictor also includes a duration predictor, and the prosody prediction module 44 mentioned above includes:

[0148] The second pitch prediction submodule is used to predict the pitch of the text stress vector and the text pause vector through the pitch predictor, and outputs the text pitch vector.

[0149] The duration prediction submodule is used to predict the duration of text stress vectors and text pause vectors using a duration predictor, and outputs a text duration vector.

[0150] The vector fusion submodule is used to fuse the text pitch vector and the text duration vector to obtain the text prosody vector.

[0151] It should be noted that the information interaction and execution process between the above modules are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.

[0152] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 3 of the present invention. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown in the diagram), a memory, and a computer program stored in the memory and executable on at least one processor, which, when executed by the processor, implements the steps in any of the above-described speech synthesis method embodiments.

[0153] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.

[0154] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0155] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of the computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.

[0156] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.

[0157] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.

[0158] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.

[0159] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0160] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0161] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0162] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A speech synthesis method based on artificial intelligence, characterized in that, The speech synthesis method includes: The target text is acquired, and the trained feature extraction model extracts features from the target text, outputting a text feature vector. The text feature vector is used by a trained stress predictor to predict stress labels and output a stress prediction vector. The stress prediction vector is then added to the text feature vector to obtain the text stress vector. The text feature vector is used by a trained pause predictor to predict pause labels and output a pause prediction vector. The pause prediction vector is then added to the text feature vector to obtain the text pause vector. The text stress vector and the text pause vector are used by a trained prosodic predictor to predict prosodic labels and output text prosodic vectors. The target text is converted into phonemes to obtain a phoneme sequence corresponding to the target text. The text prosody vector is matched with the phoneme sequence to obtain a phoneme sequence with prosody tags. The phoneme sequence with prosody tags is converted into speech to obtain synthesized speech. The trained accent predictor includes a first convolutional layer, a first fully connected layer, a first normalization layer, and a first feature transformation layer; The text feature vector is used by a trained stress predictor to predict stress labels, and the output stress prediction vector includes: The first convolutional layer extracts features from the text feature vector and outputs an accent feature vector; The first fully connected layer classifies the accent feature vector and outputs a first accent probability vector, wherein the first accent probability vector is used to represent the first probability that the corresponding accent feature vector belongs to each accent label; The first normalization layer normalizes the first accent probability vector and outputs the second accent probability vector; The first feature transformation layer performs feature transformation on the second accent probability vector and outputs an accent prediction vector.

2. The speech synthesis method according to claim 1, characterized in that, The trained feature extraction model includes an embedding layer and a spectral analysis layer; The trained feature extraction model extracts features from the target text, and the output text feature vector includes: The embedding layer performs text embedding on the target text and outputs text embedding features; The spectral analysis layer performs spectral analysis on the text embedding features and outputs a text feature vector.

3. The speech synthesis method according to claim 1, characterized in that, The trained pause predictor includes a second convolutional layer, a second fully connected layer, a second normalization layer, and a second feature transformation layer. The text feature vector is used by a trained pause predictor to predict pause labels, and the output pause prediction vector includes: The second convolutional layer extracts features from the text feature vector and outputs a pause feature vector; The second fully connected layer classifies the pause feature vector and outputs a first pause probability vector; The second normalization layer normalizes the first pause probability vector and outputs the second pause probability vector; The second feature transformation layer performs feature transformation on the second pause probability vector and outputs a pause prediction vector.

4. The speech synthesis method according to claim 1, characterized in that, The prosody predictor includes a pitch predictor; The text stress vector and the text pause vector are used by a trained prosodic predictor to predict prosodic labels, and the output text prosodic vector includes: The text accent vector and the text pause vector are used by the pitch predictor to predict the pitch, and the output text pitch vector is used as the text prosody vector.

5. The speech synthesis method according to claim 4, characterized in that, The prosody predictor also includes a duration predictor; The text stress vector and the text pause vector are used by a trained prosodic predictor to predict prosodic labels, and the output text prosodic vector includes: The text accent vector and the text pause vector are used by the pitch predictor to predict the pitch and output the text pitch vector. The text stress vector and the text pause vector are used by the duration predictor to predict the duration and output the text duration vector. The text pitch vector and the text duration vector are fused to obtain the text prosody vector.

6. A speech synthesis device based on artificial intelligence, characterized in that, The speech synthesis device includes: The feature extraction module is used to acquire target text. The trained feature extraction model extracts features from the target text and outputs a text feature vector. The stress prediction module is used to predict stress labels by the trained stress predictor on the text feature vector, output stress prediction vector, and add stress prediction vector and text feature vector to obtain text stress vector. The pause prediction module is used to predict pause labels by the trained pause predictor on the text feature vector, output a pause prediction vector, and add the pause prediction vector to the text feature vector to obtain the text pause vector. The prosody prediction module is used to predict prosody tags by using the text stress vector and the text pause vector through a trained prosody predictor, and outputs the text prosody vector. The speech synthesis module is used to perform phoneme conversion on the target text to obtain a phoneme sequence corresponding to the target text, match the text prosody vector with the phoneme sequence to obtain a phoneme sequence with prosody tags, and perform speech conversion on the phoneme sequence with prosody tags to obtain synthesized speech. The trained accent predictor includes a first convolutional layer, a first fully connected layer, a first normalization layer, and a first feature transformation layer. The accent prediction module includes: The accent feature extraction submodule is used to extract features from the text feature vector by the first convolutional layer and output the accent feature vector; The first classification submodule is used to classify the accent feature vector by the first fully connected layer and output a first accent probability vector, wherein the first accent probability vector is used to represent the first probability that the corresponding accent feature vector belongs to each accent label; The first normalization submodule is used to normalize the first accent probability vector by the first normalization layer and output the second accent probability vector. The first feature transformation submodule is used to perform feature transformation on the second accent probability vector by the first feature transformation layer and output the accent prediction vector.

7. The speech synthesis device according to claim 6, characterized in that, The trained feature extraction model includes an embedding layer and a spectral analysis layer, and the feature extraction module includes: The text embedding submodule is used by the embedding layer to perform text embedding on the target text and output text embedding features; The spectrum analysis submodule is used by the spectrum analysis layer to perform spectrum analysis on the text embedding features and output the text feature vector.

8. A computer device, characterized in that, The computer device includes a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the speech synthesis method as described in any one of claims 1 to 5.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the speech synthesis method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Speech synthesis method and device, electronic equipment and storage medium

    CN110782870A

  • Training method and device for rhythm generation model

    CN110782880A

  • Voice style migration synthesis method and device, electronic equipment and storage medium

    CN116129869A