Speech synthesis method, apparatus, device, storage medium, and program product
By using two-layer sub-models to process semantic features and acoustic features in the speech synthesis technology, the problem of low authenticity of speech synthesis in the prior art is solved, and a speech synthesis effect that is more suitable for human voices is achieved.
Patent Information
- Application Number
- PCT/CN2024/115324
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-27
- Filing Date
- 2024-08-29
- Publication Date
- 2025-06-05
AI Technical Summary
Existing speech synthesis technologies are difficult to generate speech that is comparable to real-person pronunciations, and the synthesized audio has poor authenticity.
By obtaining semantic features and acoustic features, inputting them into the two-layer sub-model in the speech synthesis model for processing, generating synthetic audio that is the same as the reference speech timbre.
The authenticity of speech synthesis is improved, so that the generated audio is more in line with the actual voice characteristics of the object corresponding to the selected tone.
Smart Images

Figure CN2024115324_05062025_PF_FP_ABST
Abstract
Description
Speech synthesis method, device, equipment, storage medium and program product
[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on November 27, 2023, with application number 2023116038296 and application name “Speech synthesis method, device, equipment, storage medium and program product”, the entire contents of which are incorporated by reference into this application. Technical Field
[0002] The embodiments of the present application relate to the field of speech synthesis technology, and in particular to speech synthesis. Background Art
[0003] With the advancement of deep learning, speech synthesis technology has made tremendous progress. Realistic, natural speech synthesis technology has been applied to voice interaction systems such as mobile phone voice assistants, smart speakers, and in-car computers. At the same time, user demand for speech synthesis technology is increasing, and the technical requirements are also increasing. Users not only expect synthesized speech to be comparable to real people, but also to have a variety of voices, even the voices of family and friends.
[0004] In related technologies, during speech synthesis, a text sequence is input into a trained speech synthesis model, and the speech synthesis model automatically generates synthesized audio based on the text; specifically, after inputting a text sequence, the speech synthesis model first maps the text sequence to the corresponding audio features, and then converts the audio features into sounds that we can understand, that is, into synthesized audio.
[0005] However, the above method can only simulate real speech in a rigid way. The generated synthetic audio does not fit the characteristics of the actual voice of the person, and the authenticity of the synthetic audio is poor.
[0006] Summary of the Invention
[0007] This application provides a speech synthesis method, apparatus, device, storage medium, and program product. The technical solutions are as follows:
[0008] According to one aspect of the present application, a speech synthesis method is provided, the method comprising:
[0009] Acquiring semantic features and acoustic features, wherein the semantic features are used to represent features of text information corresponding to the target audio to be synthesized, and the acoustic features are features of acoustic information corresponding to a reference voice, wherein the reference voice refers to the voice of the object corresponding to the selected timbre;
[0010] Embedding the semantic features into the acoustic features through a first-layer sub-model in the speech synthesis model to obtain intermediate acoustic features, wherein the first-layer sub-model is used to embed the semantic features into the acoustic features, and the intermediate acoustic features are used to represent features obtained after embedding the semantic features into the acoustic features;
[0011] Inputting the intermediate acoustic features into a second-layer sub-model in the speech synthesis model to obtain audio synthesis features, wherein the second-layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features, and the audio synthesis features are used to represent features corresponding to the target audio to be synthesized;
[0012] A synthesized audio having the same timbre as the reference speech is generated based on the audio synthesis feature.
[0013] According to one aspect of the present application, a method for training a speech synthesis model is provided, the method comprising:
[0014] Obtaining sample semantic features, sample acoustic features, and sample audio, wherein the sample semantic features are used to represent features of text information corresponding to the target audio to be synthesized, and the sample acoustic features are features of acoustic information corresponding to the sample reference speech, where the sample reference speech refers to the speech of the object corresponding to the selected timbre;
[0015] Embedding the sample semantic features into the sample acoustic features through the first layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features, wherein the first layer sub-model is used to embed the sample semantic features into the sample acoustic features, and the sample intermediate acoustic features are used to represent features obtained after embedding the sample semantic features into the sample acoustic features;
[0016] Inputting the intermediate acoustic features of the sample into a second-layer sub-model in the speech synthesis model to obtain audio synthesis features, wherein the second-layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features of the sample, and the audio synthesis features are used to represent features corresponding to the target audio to be synthesized;
[0017] Generate a synthesized audio with the same timbre as the sample reference speech based on the audio synthesis feature;
[0018] Calculating a training loss of the speech synthesis model based on the sample audio and the synthesized audio;
[0019] Model parameters of the speech synthesis model are updated according to the training loss.
[0020] According to one aspect of the present application, a speech synthesis device is provided, the device comprising:
[0021] an acquisition module, configured to acquire semantic features and acoustic features, wherein the semantic features are used to represent features of text information corresponding to the target audio to be synthesized, and the acoustic features are features of acoustic information corresponding to a reference voice, wherein the reference voice refers to the voice of the object corresponding to the selected timbre;
[0022] a feature processing module, configured to embed the semantic features into the acoustic features through a first-layer sub-model in the speech synthesis model to obtain intermediate acoustic features, wherein the first-layer sub-model is configured to embed the semantic features into the acoustic features, and the intermediate acoustic features are configured to represent features obtained after embedding the semantic features into the acoustic features;
[0023] The feature processing module is used to input the intermediate acoustic features into the second-layer sub-model in the speech synthesis model to obtain audio synthesis features, the second-layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features, and the audio synthesis features are used to represent features corresponding to the target audio to be synthesized;
[0024] A generation module is used to generate a synthesized audio with the same timbre as the reference speech based on the audio synthesis feature.
[0025] According to one aspect of the present application, a device for training a speech synthesis model is provided, the device comprising:
[0026] an acquisition module, configured to acquire sample semantic features, sample acoustic features, and sample audio, wherein the sample semantic features represent features of semantic text information corresponding to the target audio to be synthesized, and the sample acoustic features are features of acoustic information corresponding to the sample reference speech, wherein the sample reference speech refers to the speech of the object corresponding to the selected timbre;
[0027] a feature processing module, configured to embed the sample semantic features into the sample acoustic features through a first-layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features, wherein the first-layer sub-model is configured to embed the sample semantic features into the sample acoustic features, and the sample intermediate acoustic features are configured to represent features obtained after embedding the sample semantic features into the sample acoustic features;
[0028] The feature processing module is used to input the intermediate acoustic features of the sample into the second-layer sub-model in the speech synthesis model to obtain audio synthesis features, the second-layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features of the sample, and the audio synthesis features are used to represent features corresponding to the target audio to be synthesized;
[0029] A generating module, configured to generate a synthesized audio having the same timbre as the sample reference speech based on the audio synthesis feature;
[0030] A calculation module, configured to calculate a training loss of the speech synthesis model based on the sample audio and the synthesized audio;
[0031] An updating module is used to update the model parameters of the speech synthesis model according to the training loss.
[0032] According to another aspect of the present application, a computer device is provided, which includes: a processor and a memory, wherein at least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the speech synthesis method described above, or the training method of the speech synthesis model described above.
[0033] According to another aspect of the present application, a computer storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the speech synthesis method described above, or the training method of the speech synthesis model described above.
[0034] According to another aspect of the present application, a computer program product is provided, which includes a computer program stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes the speech synthesis method described above, or the training method of the speech synthesis model described above.
[0035] The beneficial effects of the technical solution provided by this application include at least:
[0036] By obtaining semantic features and acoustic features corresponding to the reference speech; inputting the semantic features and acoustic features into the first layer sub-model of the speech synthesis model to perform feature embedding to obtain intermediate acoustic features; inputting the intermediate acoustic features into the second layer sub-model of the speech synthesis model to perform speech synthesis to obtain audio synthesis features; decoding the audio synthesis features to obtain synthesized audio with the same timbre as the reference speech. This application processes semantic features and acoustic features through two layers of sub-models in the speech synthesis model, so that the finally generated audio synthesis features can learn both semantic features and acoustic features. Compared with the method of rigidly imitating real speech, the sound output by the speech synthesis model in this method is more consistent with the actual vocal characteristics of the object corresponding to the selected timbre, thereby improving the authenticity of speech synthesis. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] FIG1 is a schematic diagram of a speech synthesis method provided by an exemplary embodiment of the present application;
[0038] FIG2 is a schematic diagram of the architecture of a computer system provided by an exemplary embodiment of the present application;
[0039] FIG3 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0040] FIG4 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application;
[0041] FIG5 is a schematic diagram of obtaining intermediate acoustic features provided by an exemplary embodiment of the present application;
[0042] FIG6 is a schematic diagram of obtaining audio synthesis features provided by an exemplary embodiment of the present application;
[0043] FIG7 is a framework diagram of a training system for generating a speech synthesis model and training a speech synthesis model provided by an exemplary embodiment of the present application;
[0044] FIG8 is a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application;
[0045] FIG9 is a flow chart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application;
[0046] FIG10 is a block diagram of a speech synthesis apparatus provided by an exemplary embodiment of the present application;
[0047] FIG11 is a block diagram of a training apparatus for a speech synthesis model provided by an exemplary embodiment of the present application;
[0048] FIG12 is a schematic structural diagram of a computer device provided by an exemplary embodiment of the present application. DETAILED DESCRIPTION
[0049] To make the objectives, technical solutions, and advantages of the present application more clear, the embodiments of the present application will be further described in detail below with reference to the accompanying drawings. Exemplary embodiments will be described in detail herein, with examples shown in the accompanying drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. Instead, they are merely examples of devices and methods consistent with certain aspects of the present application, as detailed in the appended claims.
[0050] The terms used in this disclosure are for the purpose of describing specific embodiments only and are not intended to limit the disclosure. As used in this disclosure and the appended claims, the singular forms "a," "an," "the," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise. It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0051] It should be understood that although the terms first, second, third, etc. may be used in the present disclosure to describe various information, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other.
[0052] An embodiment of the present application provides a schematic diagram of a speech synthesis method, as shown in FIG1 . The method can be executed by a computer device, which can be a terminal or a server, and a speech synthesis model is provided in the computer device.
[0053] Exemplarily, the computer device acquires semantic features and acoustic features; the computer device inputs the semantic features and acoustic features into a first-layer sub-model in a speech synthesis model to obtain intermediate acoustic features, and the first-layer sub-model is used to embed the semantic features into the acoustic features; the computer device inputs the intermediate acoustic features into a second-layer sub-model in the speech synthesis model to obtain audio synthesis features, and the second-layer sub-model is used to synthesize audio synthesis features based on the intermediate acoustic features; the computer device obtains target audio having a timbre that matches the reference speech based on the audio synthesis features.
[0054] The semantic feature is used to represent the feature of the semantic information corresponding to the target audio to be synthesized; or, the semantic feature is used to represent the feature of the text information corresponding to the target audio to be synthesized.
[0055] Optionally, the semantic feature is obtained in at least one of the following ways, but is not limited thereto:
[0056] By performing speech recognition on the speech signal corresponding to the target audio, the text content is obtained; by performing feature extraction on the text content, the semantic features are obtained;
[0057] Obtaining semantic features by performing text recognition and feature extraction on the text content; optionally, the text content includes at least one of a phrase, sentence, paragraph, or chapter composed of text and / or symbols. Optionally, the language of the text content is not limited to Chinese and English, and can be any one or more languages;
[0058] By directly extracting features from speech signals, semantic features representing the semantics of the speech signals are obtained.
[0059] Acoustic features are features of acoustic information corresponding to the reference speech.
[0060] Optionally, the acoustic features include at least one of recording environment features, timbre features, and rhythm duration features, but are not limited thereto and are not specifically limited in this embodiment of the present application.
[0061] The recording environment features are used to characterize the characteristics of the reference speech recording environment; the timbre features are used to characterize the timbre characteristics of the subject corresponding to the selected timbre; and the prosodic duration features are used to characterize the prosodic characteristics of the reference speech. For example, the prosodic duration features are used to characterize the tone, duration, pitch, and other characteristics of the subject corresponding to the selected timbre when speaking, or the prosodic duration features are used to characterize the intonation characteristics of the subject corresponding to the selected timbre when speaking.
[0062] The reference voice refers to the voice of the object corresponding to the selected timbre.
[0063] Alternatively, the reference voice may be the voice of the subject corresponding to the selected timbre; alternatively, the reference voice may be the voice of the subject corresponding to the selected timbre while reading a text; alternatively, the reference voice may be an audio segment containing the voice of the subject corresponding to the selected timbre, but is not limited thereto. For example, the reference voice may be a recording of person A reading a text.
[0064] The intermediate acoustic features are used to represent the features obtained by embedding the semantic features into the acoustic features.
[0065] Audio synthesis features are used to represent features corresponding to the target audio.
[0066] As shown in Figure 1 , a computer device obtains reference speech 10 and inputs it into an acoustic feature extraction network to extract features, thereby obtaining initial acoustic features 20. The computer device then quantizes initial acoustic features 20 to obtain acoustic features 30. Initial acoustic features 20 before quantization are a one-dimensional matrix, while acoustic features 30 after quantization are a two-dimensional matrix.
[0067] The initial acoustic features 20 refer to continuous features extracted from the reference speech 10 .
[0068] For example, the initial acoustic feature 20 before quantization can be expressed as [a1, a2, a3], and the acoustic feature 30 after quantization is a three-dimensional matrix, which can be expressed as:
[0069] It should be noted that the dimension of a matrix refers to the number of rows in the matrix.
[0070] After obtaining the quantized acoustic feature 30, the computer device adds the feature values of the same column dimension in the acoustic feature 30 to obtain a one-dimensional acoustic feature 50; the computer device splices the semantic feature 40 and the one-dimensional acoustic feature 50 to obtain a spliced feature, where the spliced feature refers to a feature obtained by splicing the semantic feature 40 and the one-dimensional acoustic feature 50; the computer device inputs the spliced feature into the first-layer sub-model 60 to obtain an intermediate acoustic feature 70.
[0071] For example, the computer device adds the eigenvalues of the first column dimension in the acoustic feature 30, for example, numerically adds a11, a12, and a13 to obtain the first acoustic eigenvalue A1 in the one-dimensional acoustic feature 50; similarly, numerically adds a21, a22, and a23 to obtain the second acoustic eigenvalue A2 in the one-dimensional acoustic feature 50; and numerically adds a31, a32, and a33 to obtain the third acoustic eigenvalue A3 in the one-dimensional acoustic feature 50. Therefore, the obtained one-dimensional acoustic feature 50 can be expressed as [A1, A2, A3].
[0072] In some embodiments, the semantic feature 40 can be expressed as [s1, s2, s3]. The computer device concatenates the semantic feature 40 and the one-dimensional acoustic feature 50 to obtain a concatenated feature. The obtained concatenated feature can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the splicing feature into the first layer sub-model 60 to obtain the first intermediate acoustic feature value h1 in the intermediate acoustic feature 70; the computer device inputs the first audio synthesis feature value and the splicing feature corresponding to the generated first intermediate acoustic feature value h1 into the first layer sub-model 60 to predict the intermediate acoustic feature value to obtain the second intermediate acoustic feature value h2, wherein the first audio synthesis feature value is predicted by inputting the first intermediate acoustic feature value into the second layer sub-model 80; and so on, until the number of output intermediate acoustic feature values is equal to the number of feature values in the one-dimensional acoustic feature; the computer device merges the generated intermediate acoustic feature values to obtain the intermediate acoustic feature 70, which can be expressed as [h1, h2, h3, Eos], wherein Eos is used to represent the terminator.
[0073] After obtaining the intermediate acoustic feature 70, the computer device sequentially inputs the intermediate acoustic feature values in the intermediate acoustic feature 70 into the second-layer sub-model 80 to predict the audio synthesis feature values in the audio synthesis feature 90 to obtain the audio synthesis feature 90; the computer device decodes the audio synthesis feature 90 to obtain the synthesized audio 100 with the same timbre as the reference speech 10.
[0074] In some embodiments, the computer device inputs the first intermediate acoustic feature value h1 in the intermediate acoustic feature 70 into the second layer sub-model 80 to predict the audio synthesis feature value in the audio synthesis feature 90, and obtains the first audio synthesis feature value in the audio synthesis feature 90; the computer device inputs the generated first audio synthesis feature value and the splicing feature into the first layer sub-model 60, and obtains the second intermediate acoustic feature value h2; the computer device inputs the generated second intermediate acoustic feature value h2 into the second layer sub-model 80 to perform speech synthesis, and obtains the second audio synthesis feature value in the audio feature 90; and so on, until no intermediate acoustic feature value is input into the second layer sub-model 80, the computer device merges the generated audio synthesis feature values to obtain the audio synthesis feature 90.
[0075] In summary, the method provided in this embodiment obtains semantic features and acoustic features corresponding to the reference speech; inputs the semantic features and acoustic features into the first layer sub-model of the speech synthesis model to perform feature embedding to obtain intermediate acoustic features; inputs the intermediate acoustic features into the second layer sub-model of the speech synthesis model to perform speech synthesis to obtain audio synthesis features; and decodes the audio synthesis features to obtain target audio that matches the timbre of the reference speech. This application processes semantic features and acoustic features through two layers of sub-models in the speech synthesis model, so that the finally generated audio synthesis features can learn both semantic features and acoustic features. Compared with the method of rigidly imitating real speech, the sound output by the speech synthesis model in this method is more consistent with the actual vocal characteristics of the object corresponding to the selected timbre, thereby improving the authenticity of speech synthesis.
[0076] 2 shows a schematic diagram of the architecture of a computer system provided by an embodiment of the present application. The computer system may include: a terminal 100 and a server 200.
[0077] The terminal 100 may be an electronic device such as a mobile phone, a tablet computer, an in-vehicle terminal (car computer), a wearable device, a personal computer (PC), an in-vehicle terminal, an aircraft, an unmanned vending terminal, or the like. The terminal 100 may be installed with a client that runs a target application. The target application may be an application that references speech synthesis or another application that provides speech synthesis functionality, and this application does not limit this. In addition, this application does not limit the form of the target application, including but not limited to an application (Application, App) installed in the terminal 100, a mini-program, etc., and may also be in the form of a web page.
[0078] Server 200 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud computing services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data. Server 200 can be the backend server of the target application described above, used to provide backend services to the client of the target application.
[0079] The terminal 100 and the server 200 may communicate with each other via a network, such as a wired or wireless network.
[0080] In the speech synthesis method and speech synthesis model training method provided in the embodiments of the present application, the execution subject of each step can be a computer device, and the computer device refers to an electronic device with data calculation, processing and storage capabilities. Taking the implementation of the scheme shown in Figure 2 as an example, the speech synthesis method and the speech synthesis model training method can be executed by the terminal 100 (such as the client of the target application installed and running in the terminal 100 executes the speech synthesis method and the speech synthesis model training method), or the speech synthesis method and the speech synthesis model training method can be executed by the server 200, or the terminal 100 and the server 200 interact and cooperate to execute, and this application does not limit this.
[0081] FIG3 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be a terminal or a server, and a speech synthesis model is provided in the computer device. The method includes:
[0082] Step 302: Acquire semantic features and acoustic features.
[0083] The semantic feature is used to represent the semantic information corresponding to the target audio to be synthesized; or, the semantic feature is used to represent the text information corresponding to the target audio to be synthesized.
[0084] Optionally, the method of obtaining semantic features includes: obtaining text content by performing speech recognition on the speech signal corresponding to the target audio; obtaining semantic features by performing feature extraction on the text content; or obtaining semantic features by performing text recognition and feature extraction on the text in the text content; or obtaining semantic features used to represent the semantics of the speech signal by directly performing feature extraction on the speech signal.
[0085] Optionally, the text content includes at least one of a phrase, a sentence, a paragraph, and a chapter composed of characters and / or symbols, etc. Optionally, the language of the text content is not limited to Chinese and English, and can be any one or more languages.
[0086] Acoustic features are used to represent the acoustic information corresponding to the reference speech.
[0087] Optionally, the acoustic features include at least one of recording environment features, timbre features, and rhythm duration features, but are not limited thereto and are not specifically limited in this embodiment of the present application.
[0088] The recording environment features are used to characterize the characteristics of the reference speech recording environment; the timbre features are used to characterize the timbre characteristics of the subject corresponding to the selected timbre; and the prosodic duration features are used to characterize the prosodic characteristics of the reference speech. For example, the prosodic duration features are used to characterize the tone, duration, pitch, and other characteristics of the subject corresponding to the selected timbre when speaking, or the prosodic duration features are used to characterize the intonation characteristics of the subject corresponding to the selected timbre when speaking.
[0089] The reference voice refers to the voice of the object corresponding to the selected timbre. It should be noted that the object can be a physical object, such as a designated user, or a virtual object, such as an intelligent voice assistant, a virtual character in a game, etc., and this application does not limit this.
[0090] Alternatively, the reference voice may be the voice of the subject corresponding to the selected timbre; alternatively, the reference voice may be the voice of the subject corresponding to the selected timbre while reading a text; alternatively, the reference voice may be an audio segment containing the voice of the subject corresponding to the selected timbre, but is not limited thereto. For example, the reference voice may be a recording of person A reading a text.
[0091] Optionally, the reference speech may be a speech of 3 seconds in length.
[0092] In some embodiments, the acoustic features may be obtained by extracting features through an acoustic encoder (also referred to as an acoustic feature extraction network) in a speech synthesis model.
[0093] Alternatively, the acoustic encoder can be implemented using a convolutional neural network (CNN), a transformer neural network (Transformer), or a convolution-enhanced transformer neural network (Conformer). For example, the acoustic encoder can be implemented using a four-layer transformer neural network (Transformer). Of course, the acquisition of acoustic features is not limited to this implementation.
[0094] In some embodiments, semantic features can be obtained by extracting features from a semantic encoder in a speech synthesis model.
[0095] Optionally, the semantic encoder can be implemented by a self-supervised learning (SSL) encoder and a k-means clustering algorithm. Of course, the acquisition of semantic features is not limited to this implementation.
[0096] In some embodiments, the computer device may splice the reference speech and the speech signals corresponding to the target audio to obtain a spliced speech signal, and the computer device may extract features from the spliced speech signal to obtain semantic features and acoustic features.
[0097] Optionally, speech signals corresponding to a reference speech of a first time length and a target audio of a second time length are obtained; the computer device splices the speech signals corresponding to the reference speech and the target audio to obtain a spliced speech signal; the computer device extracts features of the reference speech of the first time length through an acoustic feature extraction network to obtain acoustic features; the computer device extracts features of the speech signal corresponding to the target audio of the second time length through a semantic feature extraction network to obtain semantic features.
[0098] For example, a speech signal corresponding to a 3-second reference speech and a 7-second target audio is obtained; the computer device splices the 3-second reference speech and the 7-second speech signal to obtain a 10-second spliced speech signal; the computer device extracts features from the 3-second reference speech through an acoustic feature extraction network to obtain acoustic features; the computer device extracts features from the 7-second speech signal through a semantic feature extraction network to obtain semantic features. The computer device finally obtains synthesized audio based on the semantic features and acoustic features, that is, a synthesized audio that can express the voice of the object corresponding to the selected timbre through a short 3-second reference speech. For example, if character A wants to imitate character B reading an article, he only needs to obtain 3 seconds of character B's speech as the reference speech and the speech signal of character A reading an article. The computer device splices the 3-second character B's speech and the speech signal of character A reading an article and inputs them into the speech synthesis model provided in the embodiment of the present application, and then character A can be obtained reading an article in the voice of character B.
[0099] Step 304: embed the semantic features into the acoustic features through the first layer sub-model in the speech synthesis model to obtain intermediate acoustic features.
[0100] The intermediate acoustic features are the features obtained by embedding the semantic features into the acoustic features.
[0101] Embedding semantic features into acoustic features means that the acoustic features can learn the semantic information or text information in the semantic features through the first-layer sub-model, that is, the text information that the target audio wants to express can be obtained.
[0102] The speech synthesis model includes a first layer sub-model and a second layer sub-model. The first layer sub-model and the second layer sub-model may have the same model structure, but differ in at least one of their network parameters, execution tasks, and number of network layers.
[0103] The first layer sub-model is used to embed semantic features into acoustic features.
[0104] Optionally, the first-layer sub-model may adopt at least one of the attention network Transformer, the pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and the recurrent neural network (RNN), but is not limited to this. The embodiments of the present application do not make specific limitations on this.
[0105] Step 306: Input the intermediate acoustic features into the second-layer sub-model in the speech synthesis model to obtain audio synthesis features.
[0106] The audio synthesis feature is the feature corresponding to the target audio.
[0107] The second layer sub-model is used to synthesize audio synthesis features based on the intermediate acoustic features.
[0108] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited to this, and the embodiments of the present application do not make specific limitations on this.
[0109] Step 308: Generate target audio that matches the timbre of the reference speech based on the audio synthesis features.
[0110] The target audio is an audio recording that mimics the voice of the subject (the human voice in the reference audio) with the selected timbre, expressing the text message intended by the target audio. The timbre of the voice expressed in the generated target audio matches the timbre of the subject in the reference audio. This match means that the timbre of the text message expressed in the target audio is identical or similar to the timbre of the subject in the reference audio.
[0111] Exemplarily, the computer device decodes the audio synthesis features to obtain target audio that matches the timbre of the reference speech, that is, to obtain synthesized audio that expresses text information using the human voice of the reference speech.
[0112] In summary, the method provided in this embodiment obtains semantic features and acoustic features corresponding to the reference speech; inputs the semantic features and acoustic features into the first layer sub-model of the speech synthesis model to perform feature embedding to obtain intermediate acoustic features; inputs the intermediate acoustic features into the second layer sub-model of the speech synthesis model to perform speech synthesis to obtain audio synthesis features; and decodes the audio synthesis features to obtain target audio that matches the timbre of the reference speech. This application processes semantic features and acoustic features through two layers of sub-models in the speech synthesis model, so that the finally generated audio synthesis features can learn both semantic features and acoustic features. Compared with the method of rigidly imitating real speech, the sound output by the speech synthesis model in this method is more consistent with the actual vocal characteristics of the object corresponding to the selected timbre, thereby improving the authenticity of speech synthesis.
[0113] FIG4 is a flow chart of a speech synthesis method provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be a terminal or a server, and a speech synthesis model is provided in the computer device. The method includes:
[0114] Step 402: Acquire semantic features and acoustic features.
[0115] The semantic feature is used to represent the semantic information corresponding to the target audio; or, the semantic feature is used to represent the text information corresponding to the target audio.
[0116] Acoustic features are features of the acoustic information corresponding to the reference speech. These acoustic features can represent the vocal characteristics, reading characteristics, or characteristics of the recording environment of the reference speech corresponding to the selected timbre.
[0117] In some embodiments, the computer device obtains a reference speech; the computer device inputs the reference speech into an acoustic feature extraction network in a speech synthesis model to extract features and obtain acoustic features.
[0118] Optionally, the computer device inputs the reference speech into an acoustic feature extraction network to extract features to obtain initial acoustic features; the computer device quantifies the initial acoustic features to obtain acoustic features.
[0119] Optionally, the computer device performs Fourier transform on the speech signal corresponding to each time granularity in the reference speech to obtain a mel spectrum corresponding to the reference speech; the computer device inputs the mel spectrum into an acoustic feature extraction network to extract features to obtain initial acoustic features; the computer device quantizes the initial acoustic features to obtain acoustic features.
[0120] The initial acoustic features refer to the continuous features extracted from the reference speech.
[0121] Because the initial acoustic features are continuous, they need to be quantized to facilitate subsequent calculations. The initial acoustic features before quantization are a one-dimensional matrix, while the quantized acoustic features are an n-dimensional matrix, where n is a positive integer greater than 1. The quantized acoustic features are more expressive than the initial acoustic features before quantization, meaning that a multidimensional matrix can more fully represent the acoustic features.
[0122] For example, the initial acoustic feature before quantization can be expressed as [a1, a2, a3], and the acoustic feature after quantization is a three-dimensional matrix, which can be expressed as:
[0123] The steps of quantizing the initial acoustic feature into the acoustic feature include: the computer device quantizes the mth acoustic eigenvalue in the initial acoustic feature to obtain the first quantized acoustic eigenvalue of the mth column dimension in the acoustic feature, where m is a positive integer; the computer device quantizes the residual between the first quantized acoustic eigenvalue and the mth acoustic eigenvalue to obtain the second quantized acoustic eigenvalue in the mth column dimension; the computer device quantizes the residual between the kth quantized acoustic eigenvalue and the k-1th quantized acoustic eigenvalue to obtain the k+1th quantized acoustic eigenvalue in the mth column dimension, where k is a positive integer greater than 1; the computer device merges all the quantized acoustic eigenvalues quantized in the mth column dimension to obtain the quantized acoustic eigenvalue of the mth column dimension in the acoustic feature; and repeats the above steps to merge the quantized acoustic eigenvalues of each column dimension to obtain the acoustic feature.
[0124] For example, the computer device quantizes the first acoustic eigenvalue a1 in the initial acoustic feature to obtain the first quantized acoustic eigenvalue a11 of the first column dimension in the acoustic feature; the computer device quantizes the residual between the first acoustic eigenvalue a1 and the first quantized acoustic eigenvalue a11 of the first column dimension to obtain the second quantized acoustic eigenvalue a12 in the first column dimension; the computer device quantizes the residual between the second acoustic eigenvalue a12 and the first quantized acoustic eigenvalue a11 of the first column dimension to obtain the third quantized acoustic eigenvalue a13 in the first column dimension; all the quantized acoustic eigenvalues in the first column dimension are merged to obtain the acoustic feature of the first column dimension in the acoustic feature; the above steps are performed on other column dimensions to obtain the acoustic features of other column dimensions in the acoustic feature; and the quantized acoustic eigenvalues of each column dimension are merged to obtain the acoustic feature.
[0125] By determining the next quantized acoustic eigenvalue in the column based on the difference between adjacent quantized acoustic eigenvalues in the same column, not only can the low-dimensional initial acoustic features be quickly expanded to high-dimensional acoustic features, but the contextual relationship in the reference speech can also be effectively expressed in the acoustic features, thereby improving the expressive richness and quantization quality of the acoustic features.
[0126] Optionally, the number of column dimensions of the acoustic features can be manually set as needed.
[0127] In some embodiments, the method of obtaining the reference speech includes at least one of the following situations:
[0128] 1. The computer device receives the reference voice. For example, the terminal is the terminal that initiates the audio recording. The terminal records the audio and uses the audio as the reference voice after the recording is completed.
[0129] 2. The computer device obtains the reference speech from the stored database.
[0130] It is worth noting that the above method of obtaining reference speech is only an illustrative example and is not limited to this embodiment of the present application.
[0131] Step 404: Concatenate the semantic features and the one-dimensional acoustic features to obtain concatenated features; input the concatenated features into the first-layer sub-model to obtain intermediate acoustic features.
[0132] Concatenated features refer to features obtained by concatenating semantic features and one-dimensional acoustic features.
[0133] The acoustic feature is a feature of the acoustic information corresponding to the reference speech. The acoustic feature is an n-dimensional matrix, where n is a positive integer greater than 1.
[0134] Exemplarily, the computer device adds the feature values of the same column dimension in the acoustic feature to obtain a one-dimensional acoustic feature.
[0135] For example, the acoustic feature is a three-dimensional matrix and can be expressed as: The computer device adds the eigenvalues of the first column dimension in the acoustic feature. For example, a11, a12, and a13 are numerically added to obtain the first acoustic eigenvalue A1 in the one-dimensional acoustic feature. Similarly, a21, a22, and a23 are numerically added to obtain the second acoustic eigenvalue A2 in the one-dimensional acoustic feature. A31, a32, and a33 are numerically added to obtain the third acoustic eigenvalue A3 in the one-dimensional acoustic feature. Thus, the obtained one-dimensional acoustic feature can be expressed as [A1, A2, A3].
[0136] In some embodiments, the semantic feature can be expressed as [s1, s2, s3], and the computer device splices the semantic feature and the one-dimensional acoustic feature to obtain a spliced feature. The obtained spliced feature can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the concatenated feature into the first layer sub-model to predict the intermediate acoustic feature value in the intermediate acoustic feature, and obtains the first intermediate acoustic feature value in the intermediate acoustic feature; the computer device inputs the i-1th audio synthesis feature value and the concatenated feature corresponding to the i-1th intermediate acoustic feature value into the first layer sub-model to predict the intermediate acoustic feature value, and obtains the i-th intermediate acoustic feature value, and the i-1th audio synthesis feature value is obtained by inputting the i-1th intermediate acoustic feature value into the second layer sub-model for prediction; the computer device loops the previous step until the number of intermediate acoustic feature values outputted is equal to the number of feature values in the one-dimensional acoustic feature, for example, N; the computer device merges the N intermediate acoustic feature values to obtain the intermediate acoustic feature, where i is a positive integer greater than 1.
[0137] For example, as shown in the schematic diagram of obtaining intermediate acoustic features in FIG5 , the computer device adds the feature values of the same column dimension in acoustic feature 501 to obtain a one-dimensional acoustic feature 503. Semantic feature 502 can be expressed as [s1, s2, s3]. The computer device concatenates semantic feature 502 and one-dimensional acoustic feature 503 to obtain a concatenated feature 507. The concatenated feature 507 can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the splicing feature 507 into the first layer sub-model 504 to obtain the first intermediate acoustic feature value h1 in the intermediate acoustic feature 505; the computer device inputs the first audio synthesis feature value corresponding to the generated first intermediate acoustic feature value h1 and the splicing feature 507 into the first layer sub-model 504 to predict the intermediate acoustic feature value, and obtains the second intermediate acoustic feature value h2, wherein the first audio synthesis feature value is predicted by inputting the first intermediate acoustic feature value h1 into the second layer sub-model 506; the computer device inputs the second audio synthesis feature value corresponding to the generated second intermediate acoustic feature value h2 and the splicing feature 507 into the first layer sub-model 504 to predict the intermediate acoustic feature value, and obtains To the third intermediate acoustic eigenvalue h3, wherein the second audio synthesis eigenvalue is predicted by inputting the second intermediate acoustic eigenvalue h2 into the second layer sub-model 506; and so on, until the number of intermediate acoustic eigenvalues output by the first layer sub-model 504 is equal to the number of eigenvalues in the one-dimensional acoustic feature 503. For example, the one-dimensional acoustic feature 503 can be expressed as [Bos, A1, A2, A3], which includes 3 eigenvalues. When the first layer sub-model 504 outputs 3 intermediate acoustic eigenvalues, the first layer sub-model 504 ends the prediction; the computer device merges the generated intermediate acoustic eigenvalues to obtain the intermediate acoustic feature 505, which can be expressed as [h1, h2, h3, Eos].
[0138] By feeding back the generated audio synthesis feature value to the intermediate acoustic feature value determination stage, the determination accuracy of the intermediate acoustic feature can be improved and the expression ability of the intermediate acoustic feature value can be enhanced.
[0139] The speech synthesis model includes a first layer sub-model and a second layer sub-model. The first layer sub-model and the second layer sub-model may have the same model structure, but differ in at least one of their network parameters, execution tasks, and number of network layers.
[0140] The first layer sub-model is used to embed semantic features into acoustic features.
[0141] Optionally, the first-layer sub-model may adopt at least one of the attention network Transformer, the pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and the recurrent neural network (RNN), but is not limited to this. The embodiments of the present application do not make specific limitations on this.
[0142] The semantic feature values in the semantic feature are discrete numbers. In order to facilitate the calculation of the first-layer sub-model, the semantic feature values need to be mathematically processed to convert them into vectors. The vectorization processing formula of the semantic feature can be expressed as: E(s t1 )=E s (s t1 )+PE g (t)
[0143] Among them, E(s t1 ) refers to the quantized semantic features, s t1 Refers to the semantic feature value, E s is the embedding function for semantic feature values, PE g Used to represent the position embedding function of the first layer sub-model, t1 is used to represent the time point or time position, 1≤t1≤T1, and T1 is used to represent the time length corresponding to the semantic feature.
[0144] Similarly, the acoustic eigenvalues in the one-dimensional acoustic features are discrete numbers. To facilitate the calculation of the first-layer sub-model, the one-dimensional acoustic features need to be mathematically processed to convert them into vectors. The vectorization processing formula of the one-dimensional acoustic features can be expressed as:
[0145] Among them, E(a t2 ) refers to the quantized one-dimensional acoustic feature, refers to the acoustic eigenvalue, refers to the one-dimensional acoustic eigenvalue, E s is the embedding function for the acoustic eigenvalues, PE g It is used to represent the position embedding function of the first layer sub-model, t2 is used to represent the time point or time position, 1≤t2≤T2, and T2 is used to represent the time length corresponding to the one-dimensional acoustic feature.
[0146] The calculation formula of the first layer sub-model can be expressed as: t =GlobalTransformer(E(s t1 ),E(a t2 ))=GlobalTransformer(s1,…,s T1,a1,…,a T2 )
[0147] Among them, E(s t1 ) refers to the quantized semantic features, E(a t2 ) refers to the quantized one-dimensional acoustic feature, h t Refers to the intermediate acoustic features, GlobalTransformer is used to represent the first layer sub-model, 1≤t≤T1+T2.
[0148] Step 406: Input the intermediate acoustic feature values in the intermediate acoustic feature into the second layer sub-model in sequence to obtain audio synthesis features.
[0149] Audio synthesis features are used to represent features corresponding to the target audio.
[0150] The second layer sub-model is used to synthesize audio synthesis features based on acoustic features.
[0151] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited to this, and the embodiments of the present application do not make specific limitations on this.
[0152] Exemplarily, the computer device sequentially inputs the intermediate acoustic feature values in the intermediate acoustic features into the second-layer sub-model to obtain audio synthesis features.
[0153] Exemplarily, the computer device inputs the first intermediate acoustic feature value in the intermediate acoustic feature into the second-layer sub-model, and predicts the first audio synthesis feature value in the audio synthesis feature, where the audio synthesis feature value refers to the feature value in the audio synthesis feature; the computer device inputs the jth intermediate acoustic feature value generated by the j-1th audio synthesis feature value into the second-layer sub-model, and predicts the jth audio synthesis feature value, where the jth intermediate acoustic feature value is predicted by inputting the j-1th audio synthesis feature value into the first-layer sub-model; the previous step is repeated until the number of output audio synthesis feature values is equal to the number N of intermediate acoustic feature values in the intermediate acoustic feature; the computer device merges N audio synthesis feature values to obtain the audio synthesis feature.
[0154] For example, as shown in the schematic diagram of obtaining audio synthesis features in FIG6 , the computer device inputs the first intermediate acoustic feature value h1 in the intermediate acoustic feature 602 into the second-layer sub-model 603 to predict the audio synthesis feature value in the audio synthesis feature 604, and obtains the first audio synthesis feature value in the audio synthesis feature 604. The first audio synthesis feature value can be expressed as: [m11, m12, m13]; the computer device inputs the generated first audio synthesis feature value and the concatenation feature into the first-layer sub-model 601 to obtain the second intermediate acoustic feature value h2; the computer device inputs the generated second intermediate acoustic feature value h2 into the second-layer sub-model 603 to predict the audio synthesis feature value, and obtains the second audio synthesis feature value. The second audio synthesis feature value can be expressed as: [m21, m22, m23]; the computer device inputs the generated first audio synthesis feature value, the second audio synthesis feature value and the splicing feature into the first layer sub-model 601 to obtain the third intermediate acoustic feature value h3; the computer device inputs the generated third intermediate acoustic feature value h3 into the second layer sub-model 603 to predict the audio synthesis feature value to obtain the third audio synthesis feature value, which can be expressed as: [m31, m32, m33]; and so on, until the number of output audio synthesis feature values is equal to the number of intermediate acoustic feature values in the intermediate acoustic feature 602 or until no intermediate acoustic feature value is input into the second layer sub-model 603, the computer device merges the generated audio synthesis feature values to obtain the audio synthesis feature 604.
[0155] The acoustic feature values to be input into the second-layer sub-model are discrete numbers. To facilitate the calculation of the second-layer sub-model, the acoustic features need to be mathematically processed to convert them into vectors. The vectorization formula of the acoustic features can be expressed as:
[0156] in, It refers to the quantized acoustic features. Refers to the acoustic eigenvalue, 1≤q≤D, q is the dimension of the acoustic eigenvalue, D is the total dimension of the acoustic feature, E a is the embedding function for the acoustic eigenvalues, PE l It is used to represent the position embedding function of the second-layer sub-model, and t is used to represent the time point or time position.
[0157] The calculation formula of the second-layer sub-model can be expressed as:
[0158] in, It refers to the quantized acoustic features, h t refers to the intermediate acoustic features, Refers to the acoustic eigenvalue, LocalTransformer is used to represent the second-layer sub-model, 1≤t≤T2.
[0159] Step 408: Generate target audio that matches the timbre of the reference speech based on the audio synthesis features.
[0160] The synthesized audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the reference speech) to express the desired text information.
[0161] Exemplarily, the computer device decodes the audio synthesis features to obtain target audio having the same or similar timbre as the reference speech, that is, to obtain synthesized audio expressed using the human voice of the reference speech.
[0162] In some embodiments, the audio synthesis features can be obtained by feature decoding by a decoder in a speech synthesis model.
[0163] Alternatively, the decoder can be implemented using a convolutional neural network (CNN), a transformer neural network (Transformer), or a convolution-enhanced transformer neural network (Conformer). For example, the decoder can be implemented using a six-layer transformer neural network (Transformer), although the decoder structure is not limited to this implementation.
[0164] In order to verify the synthesis effect of the speech synthesis method provided in the embodiment of the present application, the present application compares the synthesis effect of the speech synthesis model provided in the embodiment of the present application with the synthesis effect of the model in the related art. The evaluation indicators selected in the embodiment of the present application are word error rate (WER), speech similarity (SPK) and speech quality (DNSMOS). These three indicators reflect the synthesis advantages of the speech synthesis model provided in the embodiment of the present application, as shown in Table 1.
[0165] Table 1 Comparison of synthesis effects of speech synthesis models
[0166] Among them, there are three groups of experiments. Related model one and related model two refer to the models in the related technology, and the speech synthesis model refers to the model provided in the embodiment of the present application. It can be seen from the table that: the speech or audio synthesized by the speech synthesis model has the lowest word error rate, the speech or audio synthesized by the speech synthesis model has the highest speech similarity, and the speech or audio synthesized by the speech synthesis model has the best speech quality.
[0167] In summary, the method provided in this embodiment obtains semantic features and acoustic features corresponding to the reference speech; inputs the semantic features and acoustic features into the first layer sub-model in the speech synthesis model to perform feature embedding to obtain intermediate acoustic features; inputs the intermediate acoustic features into the second layer sub-model in the speech synthesis model to perform speech synthesis to obtain audio synthesis features; and decodes the audio synthesis features to obtain synthesized audio with the same timbre as the reference speech. This application processes semantic features and acoustic features through two layers of sub-models in the speech synthesis model, so that the finally generated audio synthesis features can learn both semantic features and acoustic features. Compared with the method of rigidly imitating real speech, the sound output by the speech synthesis model in this method is more consistent with the actual vocal characteristics of the object corresponding to the selected timbre, thereby improving the authenticity of speech synthesis.
[0168] The method provided in this embodiment embeds semantic features and acoustic features into features through the first-layer sub-model in the speech synthesis model, so that the semantic features to be imitated are fused with the acoustic features, and based on the semantic features to be expressed and the acoustic features to be imitated, the actual vocal characteristics of the object corresponding to the selected timbre are generated, which improves the realism of the speech synthesis.
[0169] The method provided in this embodiment uses the second-layer sub-model in the speech synthesis model to restore the acoustic information of the intermediate acoustic features output by the first-layer sub-model, and sends the restored results to the first-layer sub-model again for prediction. After cyclic prediction, the audio synthesis features are finally obtained. The semantic features and acoustic features are processed in a cyclic processing manner, so that the finally generated audio synthesis feature skills can learn the semantic features and fully learn the acoustic features. Compared with the method of rigidly imitating real speech, the sound output by the speech synthesis model in this method is more in line with the actual vocal characteristics of the object corresponding to the selected timbre, thereby improving the authenticity of speech synthesis.
[0170] The method provided in this embodiment extracts and quantifies features of the acquired reference speech so that the acoustic features can more fully represent the acoustic characteristics, making the speaker's voice more consistent with the actual vocal characteristics of the object corresponding to the selected timbre, thereby improving the realism of speech synthesis.
[0171] The method provided in this embodiment integrates semantic features and acoustic features through a two-layer sub-model, which not only reduces the computational cost but also effectively learns the interactive relationship between semantic features and acoustic features.
[0172] The training method of the speech synthesis model involved in the present application can be implemented based on the training system of the speech synthesis model. The scheme includes a speech synthesis model training system generation stage and a speech synthesis model training stage. Figure 7 is a framework diagram of a speech synthesis model training system generation and speech synthesis model training shown in an exemplary embodiment of the present application. As shown in Figure 7, in the speech synthesis model training system generation stage, the speech synthesis model training system generation device 710 obtains the speech synthesis model training system through a pre-set training sample data set, and then generates the speech synthesis model training result based on the speech synthesis model training system. In the speech synthesis model training stage, the speech synthesis model training device 720 processes the input audio signal based on the speech synthesis model training system to obtain the training result of the speech synthesis model.
[0173] Among them, the above-mentioned speech synthesis model training system generation device 710 and speech synthesis model training device 720 can be computer devices. For example, the computer device can be a fixed computer device such as a personal computer or a server, or the computer device can also be a mobile computer device such as a tablet computer or an e-book reader.
[0174] Optionally, the speech synthesis model training system generation device 710 and the speech synthesis model training device 720 can be the same device, or the speech synthesis model training system generation device 710 and the speech synthesis model training device 720 can also be different devices. Furthermore, when the speech synthesis model training system generation device 710 and the speech synthesis model training device 720 are different devices, the speech synthesis model training system generation device 710 and the speech synthesis model training device 720 can be the same type of device, such as the speech synthesis model training system generation device 710 and the speech synthesis model training device 720 can both be servers; or the speech synthesis model training system generation device 710 and the speech synthesis model training device 720 can also be different types of devices, such as the speech synthesis model training device 720 can be a personal computer or terminal, while the speech synthesis model training system generation device 710 can be a server, etc. The embodiments of the present application do not limit the specific types of the speech synthesis model training system generation device 710 and the speech synthesis model training device 720.
[0175] The above embodiment describes the speech synthesis method. Next, the training method of the speech synthesis model will be described.
[0176] FIG8 is a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be a terminal or a server, and the computer device is provided with a speech synthesis model. The method includes:
[0177] Step 802: Obtain sample semantic features, sample acoustic features, and sample audio corresponding to the sample semantic features.
[0178] The sample semantic feature is used to represent semantic information corresponding to the sample audio; or, the sample semantic feature is used to represent text information corresponding to the sample audio.
[0179] Optionally, the method of obtaining sample semantic features includes: obtaining text content by performing speech recognition on the speech signal corresponding to the sample audio; obtaining sample semantic features by performing feature extraction on the text content; or obtaining sample semantic features by performing text recognition and feature extraction on the text in the text content; or obtaining sample semantic features used to represent the semantics of the speech signal by directly performing feature extraction on the speech signal.
[0180] Optionally, the text content includes at least one of a phrase, a sentence, a paragraph, and a chapter composed of characters and / or symbols, etc. Optionally, the language of the text content is not limited to Chinese and English, and can be any one or more languages.
[0181] The sample acoustic features are features of the acoustic information corresponding to the sample reference speech.
[0182] Optionally, the sample acoustic features include at least one of recording environment features, timbre features, and rhythm duration features, but are not limited thereto and are not specifically limited in this embodiment of the present application.
[0183] The recording environment feature is used to characterize the characteristics of the recording environment of the sample reference speech; the timbre feature is used to characterize the timbre characteristics of the subject corresponding to the selected timbre; and the prosodic duration feature is used to characterize the prosodic characteristics of the sample reference speech. For example, the prosodic duration feature is used to characterize the tone, duration, pitch, and other characteristics of the subject corresponding to the selected timbre when speaking, or the prosodic duration feature is used to characterize the intonation characteristics of the subject corresponding to the selected timbre when speaking.
[0184] The sample reference voice refers to the voice of the object corresponding to the selected timbre.
[0185] Optionally, the sample reference voice is the voice of the subject corresponding to the selected timbre; or, the sample reference voice is the voice of the subject corresponding to the selected timbre while reading a text; or, the sample reference voice is an audio segment containing the voice of the subject corresponding to the selected timbre, but is not limited thereto. For example, the sample reference voice is a recording of person A reading a text.
[0186] Optionally, the sample reference speech may be a speech of 3 seconds in length.
[0187] The sample audio refers to the audio obtained by using the object corresponding to the selected timbre for text information or semantic information.
[0188] In some embodiments, the acoustic features of the samples may be obtained by extracting features through an acoustic encoder (also called an acoustic feature extraction network) in a speech synthesis model.
[0189] Alternatively, the acoustic encoder can be implemented using a convolutional neural network (CNN), a transformer neural network (Transformer), or a convolution-enhanced transformer neural network (Conformer). For example, the acoustic encoder can be implemented using a four-layer transformer neural network (Transformer). Of course, the acquisition of acoustic features is not limited to this implementation.
[0190] In some embodiments, the semantic features of the samples can be obtained by extracting features through a semantic encoder in a speech synthesis model.
[0191] Optionally, the semantic encoder can be implemented by a self-supervised learning (SSL) encoder and a k-means clustering algorithm. Of course, the acquisition of semantic features is not limited to this implementation.
[0192] In some embodiments, the computer device may splice the sample reference speech and the speech signals corresponding to the sample audio to obtain a spliced speech signal, and the computer device may extract features from the spliced speech signal to obtain sample semantic features and sample acoustic features.
[0193] Optionally, a speech signal corresponding to a sample reference speech of a first time length and a sample audio of a second time length is obtained; the computer device splices the speech signals corresponding to the sample reference speech and the sample audio to obtain a spliced speech signal; the computer device extracts features of the reference speech of the first time length through an acoustic feature extraction network to obtain sample acoustic features; the computer device extracts features of the speech signal corresponding to the sample audio of the second time length through a semantic feature extraction network to obtain sample semantic features.
[0194] For example, a speech signal corresponding to a 3-second sample reference speech and a 7-second sample audio is obtained; the computer device splices the 3-second sample reference speech and the 7-second speech signal to obtain a 10-second spliced speech signal; the computer device extracts features from the 3-second sample reference speech through an acoustic feature extraction network to obtain sample acoustic features; the computer device extracts features from the 7-second speech signal through a semantic feature extraction network to obtain sample semantic features. The computer device ultimately obtains synthesized audio based on the sample semantic features and the sample acoustic features, that is, a synthesized audio that can express the voice of the object corresponding to the selected timbre through a short 3-second sample reference speech. For example, if character A wants to imitate character B reading an article, he only needs to obtain 3 seconds of character B's speech as a reference speech and the speech signal of character A reading an article. The computer device splices the 3-second character B's speech and the speech signal of character A reading an article and inputs them into the speech synthesis model provided in the embodiment of the present application, and then character A can be obtained reading an article in the voice of character B.
[0195] Step 804: embed the sample semantic features into the sample acoustic features through the first layer sub-model in the speech synthesis model to obtain the sample intermediate acoustic features.
[0196] The sample intermediate acoustic features are used to represent the features obtained by embedding the sample semantic features into the sample acoustic features.
[0197] Embedding the sample semantic features into the sample acoustic features means that the sample acoustic features can learn the semantic information or text information in the sample semantic features through the first layer sub-model, that is, the text information that the sample audio wants to express can be obtained.
[0198] The speech synthesis model includes a first layer sub-model and a second layer sub-model. The first layer sub-model and the second layer sub-model have the same model structure, but differ in at least one of their network parameters, execution tasks, and number of network layers.
[0199] The first layer sub-model is used to embed semantic features into acoustic features.
[0200] Optionally, the first-layer sub-model may adopt at least one of the attention network Transformer, the pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and the recurrent neural network (RNN), but is not limited to this. The embodiments of the present application do not make specific limitations on this.
[0201] Step 806: Input the intermediate acoustic features of the sample into the second-layer sub-model in the speech synthesis model to obtain audio synthesis features.
[0202] The audio synthesis feature is used to represent the features corresponding to the synthesized audio.
[0203] The second layer sub-model is used to synthesize audio synthesis features based on the intermediate acoustic features of the samples.
[0204] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited to this, and the embodiments of the present application do not make specific limitations on this.
[0205] Step 808: Generate synthesized audio with the same timbre as the sample reference speech based on the audio synthesis features.
[0206] The synthesized audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the sample reference voice) to express the desired text information.
[0207] Exemplarily, the computer device decodes the audio synthesis feature to obtain synthesized audio having the same timbre as the sample reference speech, that is, obtains synthesized audio expressed using the human voice of the sample reference speech.
[0208] Step 810: Calculate the training loss of the speech synthesis model based on the sample audio and the synthesized audio.
[0209] Exemplarily, the computer device calculates the training loss of the speech synthesis model based on the sample audio and the synthesized audio.
[0210] Training loss refers to the difference between the input and output of the speech synthesis model. The performance of the speech synthesis model is measured by training loss.
[0211] Step 812: Update the model parameters of the speech synthesis model according to the training loss.
[0212] Exemplarily, the computer device updates the model parameters of the speech synthesis model according to the training loss.
[0213] Model parameter updating refers to updating the network parameters in the speech synthesis model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but is not limited to this, and the embodiments of the present application do not limit this.
[0214] In summary, the method provided in this embodiment obtains sample semantic features, sample acoustic features and sample audio; embeds sample semantic features into sample acoustic features through the first layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features; inputs sample intermediate acoustic features and sample acoustic features into the second layer sub-model in the speech synthesis model to obtain audio synthesis features; generates synthesized audio with the same timbre as the sample reference speech based on the audio synthesis features; calculates the training loss of the speech synthesis model based on the sample audio and the synthesized audio; and updates the model parameters of the speech synthesis model according to the training loss. This application processes semantic features and acoustic features so that the audio synthesis feature skills finally generated can not only learn semantic features, but also fully learn acoustic features, and trains the speech synthesis model through the difference between the sample audio and the synthesized audio. Based on this, the synthesis effect of the speech synthesis model can be improved.
[0215] FIG9 is a flowchart of a method for training a speech synthesis model provided by an exemplary embodiment of the present application. The method can be executed by a computer device, which can be a terminal or a server, and the computer device is provided with a speech synthesis model. The method includes:
[0216] Step 902: Obtain sample semantic features, sample acoustic features, and sample audio corresponding to the sample semantic features.
[0217] The sample semantic feature is used to represent the feature of the semantic information corresponding to the sample audio; or, the sample semantic feature is used to represent the feature of the text information corresponding to the sample audio.
[0218] Optionally, the method of obtaining sample semantic features includes: obtaining text content by performing speech recognition on speech signals; obtaining sample semantic features by performing feature extraction on text content; or obtaining sample semantic features by performing text recognition and feature extraction on text in text content; or obtaining sample semantic features used to represent the semantics of the speech signal by directly performing feature extraction on the speech signal.
[0219] The sample acoustic features are features of the acoustic information corresponding to the sample reference speech. The sample acoustic features can represent the vocal characteristics, reading characteristics, or characteristics of the recording environment of the reference speech corresponding to the selected timbre.
[0220] Optionally, the sample acoustic features include at least one of recording environment features, timbre features, and rhythm duration features, but are not limited thereto and are not specifically limited in this embodiment of the present application.
[0221] The recording environment feature is used to characterize the characteristics of the recording environment of the sample reference speech; the timbre feature is used to characterize the timbre characteristics of the subject corresponding to the selected timbre; and the prosodic duration feature is used to characterize the prosodic characteristics of the reference speech. For example, the prosodic duration feature is used to characterize the tone, duration, pitch, and other characteristics of the subject corresponding to the selected timbre when speaking, or the prosodic duration feature is used to characterize the intonation characteristics of the subject corresponding to the selected timbre when speaking.
[0222] The sample reference voice refers to the voice of the object corresponding to the selected timbre.
[0223] Optionally, the sample reference voice is the voice of the subject corresponding to the selected timbre; or, the sample reference voice is the voice of the subject corresponding to the selected timbre while reading a text; or, the sample reference voice is an audio segment containing the voice of the subject corresponding to the selected timbre, but is not limited thereto. For example, the sample reference voice is a recording of person A reading a text.
[0224] In some embodiments, the computer device obtains a sample reference speech; the computer device inputs the sample reference speech into an acoustic feature extraction network in a speech synthesis model to extract features and obtain acoustic features.
[0225] Optionally, the computer device inputs the sample reference speech into an acoustic feature extraction network to extract features to obtain initial sample acoustic features; the computer device quantifies the initial sample acoustic features to obtain sample acoustic features.
[0226] The initial sample acoustic features refer to the continuous features extracted from the sample reference speech.
[0227] Because the initial sample acoustic features are continuous, they need to be quantized to facilitate subsequent calculations, thereby obtaining the sample acoustic features. The initial sample acoustic features before quantization are a one-dimensional matrix, while the quantized sample acoustic features are an n-dimensional matrix, where n is a positive integer greater than 1. The quantized sample acoustic features are more expressive than the initial sample acoustic features before quantization, meaning that a multidimensional matrix can more fully represent the sample acoustic features.
[0228] For example, the initial sample acoustic feature before quantization can be expressed as [a1, a2, a3], and the sample acoustic feature after quantization is a three-dimensional matrix, which can be expressed as:
[0229] The steps of quantizing the initial sample acoustic feature into the sample acoustic feature include: the computer device quantizes the mth acoustic eigenvalue in the initial sample acoustic feature to obtain the first quantized acoustic eigenvalue of the mth column dimension in the sample acoustic feature, where m is a positive integer; the computer device quantizes the residual between the first quantized acoustic eigenvalue and the mth acoustic eigenvalue to obtain the second quantized acoustic eigenvalue in the mth column dimension; the computer device quantizes the residual between the kth quantized acoustic eigenvalue and the k-1th quantized acoustic eigenvalue to obtain the k+1th quantized acoustic eigenvalue in the mth column dimension, where k is a positive integer greater than 1; the computer device merges the first k+1 quantized acoustic eigenvalues in the mth column dimension to obtain the quantized acoustic eigenvalue of the mth column dimension in the sample acoustic feature; and repeats the above steps to merge the quantized acoustic eigenvalues of each column dimension to obtain the sample acoustic feature.
[0230] For example, the computer device quantizes the first acoustic eigenvalue a1 in the initial sample acoustic feature to obtain the first quantized acoustic eigenvalue a11 of the first column dimension in the sample acoustic feature; the computer device quantizes the residual between the first acoustic eigenvalue a1 and the first quantized acoustic eigenvalue a11 of the first column dimension to obtain the second quantized acoustic eigenvalue a12 in the first column dimension; the computer device quantizes the residual between the second acoustic eigenvalue a12 and the first quantized acoustic eigenvalue a11 of the first column dimension to obtain the third quantized acoustic eigenvalue a13 in the first column dimension; the quantized acoustic eigenvalues in the first column dimension are merged to obtain the acoustic feature of the first column dimension in the sample acoustic feature; the above steps are performed on other column dimensions to obtain the acoustic features of other column dimensions in the sample acoustic feature; the quantized acoustic eigenvalues of each column dimension are merged to obtain the sample acoustic feature.
[0231] Optionally, the dimensions of the sample acoustic features can be manually set as needed.
[0232] In some embodiments, the method of obtaining the sample reference speech includes at least one of the following situations:
[0233] 1. The computer device receives the sample reference voice. For example, the terminal is the terminal that initiates the audio recording. The terminal records the audio and uses the audio as the sample reference voice after the recording is completed.
[0234] 2. The computer device obtains sample reference speech from the stored database.
[0235] It is worth noting that the above method of obtaining sample reference speech is only an illustrative example and is not limited to this embodiment of the present application.
[0236] Step 904: Concatenate the sample semantic features and the one-dimensional sample acoustic features to obtain sample concatenated features; input the sample concatenated features into the first-layer sub-model to obtain sample intermediate acoustic features.
[0237] Sample splicing features refer to the features obtained by splicing sample semantic features and one-dimensional sample acoustic features.
[0238] The sample acoustic feature is a feature of the acoustic information corresponding to the sample reference speech. The sample acoustic feature is an n-dimensional matrix, where n is a positive integer greater than 1.
[0239] Exemplarily, the computer device adds the feature values of the same column dimension in the sample acoustic feature to obtain a one-dimensional sample acoustic feature.
[0240] For example, the acoustic feature of a sample is a three-dimensional matrix, which can be expressed as: The computer device adds the eigenvalues of the first column dimension in the sample acoustic feature. For example, a11, a12, and a13 are numerically added to obtain the first acoustic eigenvalue A1 in the one-dimensional sample acoustic feature. Similarly, a21, a22, and a23 are numerically added to obtain the second acoustic eigenvalue A2 in the one-dimensional sample acoustic feature. A31, a32, and a33 are numerically added to obtain the third acoustic eigenvalue A3 in the one-dimensional sample acoustic feature. Thus, the obtained one-dimensional sample acoustic feature can be expressed as [A1, A2, A3].
[0241] In some embodiments, the sample semantic feature can be expressed as [s1, s2, s3], and the computer device splices the sample semantic feature and the one-dimensional sample acoustic feature to obtain the sample splicing feature. The obtained sample splicing feature can be expressed as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the sample splicing feature into the first layer sub-model to predict the first sample intermediate acoustic feature value in the sample intermediate acoustic feature; the computer device inputs the i-1th audio synthesis feature value and the splicing feature corresponding to the i-1th sample intermediate acoustic feature value into the first layer sub-model to predict the intermediate acoustic feature value to obtain the i-th sample intermediate acoustic feature value, and the i-1th audio synthesis feature value is obtained by inputting the i-1th sample intermediate acoustic feature value into the second layer sub-model for prediction; the computer device loops the previous step until the number of sample intermediate acoustic feature values outputted is equal to the number N of feature values in the one-dimensional sample acoustic feature, where i is a positive integer greater than 1 and less than N; the computer device merges the N sample intermediate acoustic feature values to obtain the sample intermediate acoustic feature, where i is a positive integer greater than 1.
[0242] For example, the computer device adds the feature values of the same column dimension in the sample acoustic feature to obtain a one-dimensional sample acoustic feature. The sample semantic feature can be represented as [s1, s2, s3]. The computer device concatenates the sample semantic feature and the one-dimensional sample acoustic feature to obtain a sample concatenation feature. The obtained sample concatenation feature can be represented as [Bos, s1, s2, s3, Bos, A1, A2, A3], where Bos is used to represent the start symbol. The computer device inputs the sample splicing feature into the first layer sub-model to obtain the first sample intermediate acoustic feature value in the sample intermediate acoustic feature; the computer device inputs the first audio synthesis feature value and the sample splicing feature corresponding to the generated first sample intermediate acoustic feature value into the first layer sub-model to predict the sample intermediate acoustic feature value, and obtains the second sample intermediate acoustic feature value, wherein the first audio synthesis feature value is obtained by inputting the first sample intermediate acoustic feature value into the second layer sub-model for prediction; the computer device inputs the second audio synthesis feature value and the sample splicing feature corresponding to the generated second sample intermediate acoustic feature value into the first layer sub-model to predict the sample intermediate acoustic feature value, and obtains To the third intermediate acoustic eigenvalue, wherein the second audio synthesis eigenvalue is obtained by inputting the second intermediate acoustic eigenvalue into the second layer sub-model for prediction; and so on, until the number of sample intermediate acoustic eigenvalues output by the first layer sub-model is equal to the number of eigenvalues in the one-dimensional sample acoustic feature, for example, the one-dimensional sample acoustic feature can be expressed as [Bos, A1, A2, A3], which includes 3 eigenvalues. When the first layer sub-model 504 outputs 3 sample intermediate acoustic eigenvalues, the first layer sub-model ends the prediction; the computer device merges the generated sample intermediate acoustic eigenvalues to obtain the sample intermediate acoustic feature, which can be expressed as [h1, h2, h3, Eos].
[0243] The speech synthesis model includes a first layer sub-model and a second layer sub-model. The first layer sub-model and the second layer sub-model have the same model structure, but differ in at least one of their network parameters, execution tasks, and number of network layers.
[0244] The first layer sub-model is used to embed the sample semantic features into the sample acoustic features.
[0245] Optionally, the first-layer sub-model may adopt at least one of the attention network Transformer, the pre-trained language representation model (Bidirectional Encoder Representation from Transformers, BERT), and the recurrent neural network (RNN), but is not limited to this. The embodiments of the present application do not make specific limitations on this.
[0246] Step 906: Input the sample intermediate acoustic feature values in the sample intermediate acoustic feature into the second layer sub-model in sequence to obtain the audio synthesis feature.
[0247] The audio synthesis feature is used to represent the features corresponding to the synthesized audio.
[0248] The second layer sub-model is used to synthesize audio synthesis features based on the intermediate acoustic features of the samples.
[0249] Optionally, the second-layer sub-model may adopt at least one of the attention network Transformer, BERT, and RNN, but is not limited to this, and the embodiments of the present application do not make specific limitations on this.
[0250] Exemplarily, the computer device sequentially inputs the sample intermediate acoustic feature values in the sample intermediate acoustic feature into the second layer sub-model to obtain the audio synthesis feature.
[0251] Exemplarily, the computer device inputs the first sample intermediate acoustic feature value in the sample intermediate acoustic feature into the second layer sub-model, and predicts the first audio synthesis feature value in the audio synthesis feature, where the audio synthesis feature value refers to the feature value in the audio synthesis feature; the computer device inputs the jth sample intermediate acoustic feature value generated by the j-1th audio synthesis feature value into the second layer sub-model to predict the audio synthesis feature value, and obtains the jth audio synthesis feature value, where the jth sample intermediate acoustic feature value is predicted by inputting the j-1th sample audio synthesis feature value into the first layer sub-model; the previous step is repeated until the number of output audio synthesis feature values is equal to the number N of sample intermediate acoustic feature values in the sample intermediate acoustic feature, where j is a positive integer greater than 1 and less than N; the computer device merges N audio synthesis feature values to obtain the audio synthesis feature, where j is a positive integer greater than 1.
[0252] Exemplarily, the computer device inputs the first sample intermediate acoustic feature value in the sample intermediate acoustic feature into the second layer sub-model to predict the audio synthesis feature value in the audio synthesis feature, and obtains the first audio synthesis feature value in the audio synthesis feature; the computer device inputs the generated first audio synthesis feature value and the sample splicing feature into the first layer sub-model to obtain the second sample intermediate acoustic feature value; the computer device inputs the generated second sample intermediate acoustic feature value into the second layer sub-model to predict the audio synthesis feature value, and obtains the second audio synthesis feature value; the computer device inputs the generated first audio synthesis feature value, the second audio synthesis feature value and the sample splicing feature into the first layer sub-model to obtain the third sample intermediate acoustic feature value; the computer device inputs the generated third sample intermediate acoustic feature value into the second layer sub-model to predict the audio synthesis feature value, and obtains the third audio synthesis feature value; and so on, until the number of output audio synthesis feature values is equal to the number of intermediate acoustic feature values in the intermediate acoustic feature or until no sample intermediate acoustic feature value is input into the second layer sub-model, the computer device merges the generated audio synthesis feature values to obtain the audio synthesis feature.
[0253] Step 908: Generate synthesized audio that matches the sample reference speech timbre based on the audio synthesis features.
[0254] The synthesized audio is an audio that imitates the voice of the object corresponding to the selected timbre (the human voice in the reference speech) to express the desired text information.
[0255] Exemplarily, the computer device decodes the audio synthesis feature to obtain synthesized audio having the same timbre as the sample reference speech, that is, obtains synthesized audio expressed using the human voice of the sample reference speech.
[0256] In some embodiments, the audio synthesis features can be obtained by feature decoding by a decoder in a speech synthesis model.
[0257] Alternatively, the decoder can be implemented using a convolutional neural network (CNN), a transformer neural network (Transformer), or a convolution-enhanced transformer neural network (Conformer). For example, the decoder can be implemented using a six-layer transformer neural network (Transformer), although the decoder structure is not limited to this implementation.
[0258] Step 910: Calculate the training loss of the speech synthesis model based on the sample audio and the synthesized audio.
[0259] Exemplarily, the computer device calculates the training loss of the speech synthesis model based on the sample audio and the synthesized audio.
[0260] Training loss refers to the difference between the input and output of the speech synthesis model. The performance of the speech synthesis model is measured by training loss.
[0261] Step 912: Update the model parameters of the speech synthesis model according to the training loss.
[0262] Exemplarily, the computer device updates the model parameters of the speech synthesis model according to the training loss.
[0263] Model parameter updating refers to updating the network parameters in the speech synthesis model, or updating the network parameters of each network module in the model, or updating the network parameters of each network layer in the model, but is not limited to this, and the embodiments of the present application do not limit this.
[0264] Based on the loss function value, the model parameters of the first-layer sub-model and the second-layer sub-model in the speech synthesis model are updated using the loss function value as a training indicator until the loss function value converges, thereby obtaining a trained speech synthesis model.
[0265] The convergence of the loss function value means that the loss function value no longer changes, or the error difference between two adjacent iterations during speech synthesis model training is less than a preset value, or the number of training times of the speech synthesis model reaches at least one of the preset times, but is not limited to this, and the embodiments of the present application are not limited to this.
[0266] Optionally, the target condition satisfied by training may be that the number of training iterations of the initial model reaches a target number, and the technician can pre-set the number of training iterations. Alternatively, the target condition satisfied by training may be that the loss value meets a target threshold condition, such as a loss value less than 0.00001, but is not limited to this and is not limited to this embodiment of the present application.
[0267] In summary, the method provided in this embodiment obtains sample semantic features, sample acoustic features and sample audio; embeds sample semantic features into sample acoustic features through the first layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features; inputs sample intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain audio synthesis features; generates synthesized audio with the same timbre as the sample reference speech based on the audio synthesis features; calculates the training loss of the speech synthesis model based on the sample audio and the synthesized audio; and updates the model parameters of the speech synthesis model according to the training loss. This application processes semantic features and acoustic features so that the audio synthesis feature skills finally generated can not only learn semantic features, but also fully learn acoustic features, and trains the speech synthesis model through the difference between the sample audio and the synthesized audio. Based on this, the synthesis effect of the speech synthesis model can be improved.
[0268] FIG10 shows a schematic diagram of the structure of a speech synthesis device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:
[0269] Acquisition module 1001 is used to acquire semantic features and acoustic features, wherein the semantic features are used to represent features of text information corresponding to the target audio, and the acoustic features are features of acoustic information corresponding to the reference speech, wherein the reference speech is the speech of the object corresponding to the selected timbre;
[0270] A feature processing module 1002 is configured to embed the semantic features into the acoustic features through a first-layer sub-model in a speech synthesis model to obtain intermediate acoustic features, wherein the first-layer sub-model is configured to embed the semantic features into the acoustic features, and the intermediate acoustic features are configured to represent features obtained after embedding the semantic features into the acoustic features;
[0271] The feature processing module 1002 is further configured to input the intermediate acoustic features into a second-layer sub-model in the speech synthesis model to obtain audio synthesis features, wherein the second-layer sub-model is configured to synthesize the audio synthesis features based on the intermediate acoustic features, and the audio synthesis features are configured to represent features corresponding to the target audio;
[0272] The generating module 1003 is configured to generate a synthesized audio having the same timbre as the reference speech based on the audio synthesis feature.
[0273] In some embodiments, the feature processing module 1002 is further used to add the feature values of the same column dimension in the acoustic feature to obtain a one-dimensional acoustic feature; splice the semantic feature and the one-dimensional acoustic feature to obtain a spliced feature, where the spliced feature refers to a feature obtained by splicing the semantic feature and the one-dimensional acoustic feature; and input the spliced feature into the first layer sub-model to obtain the intermediate acoustic feature.
[0274] In some embodiments, the feature processing module 1002 is further used to input the splicing feature into the first layer sub-model to predict the intermediate acoustic feature value in the intermediate acoustic feature, and obtain the first intermediate acoustic feature value in the intermediate acoustic feature, where the intermediate acoustic feature refers to the feature predicted by the first layer sub-model based on the splicing feature, and the intermediate acoustic feature value refers to the feature value in the intermediate acoustic feature; input the i-1th audio synthesis feature value corresponding to the i-1th intermediate acoustic feature value and the splicing feature into the first layer sub-model to predict the intermediate acoustic feature value, and obtain the i-th intermediate acoustic feature value, where the i-1th audio synthesis feature value is input into the second layer sub-model for prediction; repeat the previous step until the number of the output intermediate acoustic feature values is equal to the number of feature values in the one-dimensional acoustic feature; merge the i-th intermediate acoustic feature value with the previous i-1 intermediate acoustic feature values to obtain the intermediate acoustic feature value, where i is a positive integer greater than 1.
[0275] In some embodiments, the feature processing module 1002 is further configured to sequentially input the intermediate acoustic feature values in the intermediate acoustic feature into the second layer sub-model to obtain the audio synthesis feature.
[0276] In some embodiments, the feature processing module 1002 is also used to input the first intermediate acoustic feature value in the intermediate acoustic feature into the second layer sub-model to predict the first audio synthesis feature value in the audio synthesis feature, where the audio synthesis feature value refers to the feature value in the audio synthesis feature; input the jth intermediate acoustic feature value generated by the j-1th audio synthesis feature value into the second layer sub-model to predict the audio synthesis feature value to obtain the jth audio synthesis feature value, where the jth intermediate acoustic feature value is predicted by inputting the j-1th audio synthesis feature value into the first layer sub-model; loop the previous step until the number of the output audio synthesis feature values is equal to the number N of the intermediate acoustic feature values in the intermediate acoustic feature, where j is a positive integer greater than 1 and less than N; merge the N audio synthesis feature values to obtain the audio synthesis feature.
[0277] In some embodiments, the acquisition module 1001 is further configured to acquire the reference speech; extract features from the reference speech to obtain the acoustic features.
[0278] In some embodiments, the device also includes a computing module 1004, which is further used to input the reference speech into the acoustic feature extraction network to extract features and obtain initial acoustic features, where the initial acoustic features refer to continuous features extracted from the reference speech; and quantify the initial acoustic features to obtain the acoustic features.
[0279] In some embodiments, the calculation module 1004 is also used to quantize the mth acoustic eigenvalue in the initial acoustic feature to obtain the first quantized acoustic eigenvalue of the mth column dimension in the acoustic feature, where m is a positive integer; quantize the residual between the first quantized acoustic eigenvalue and the mth acoustic eigenvalue to obtain the second quantized acoustic eigenvalue in the mth column dimension; quantize the residual between the kth quantized acoustic eigenvalue and the k-1th quantized acoustic eigenvalue to obtain the k+1th quantized acoustic eigenvalue in the mth column dimension, where k is a positive integer greater than 1; merge all quantized acoustic eigenvalues in the mth column dimension to obtain the quantized acoustic eigenvalue of the mth column dimension in the acoustic feature; repeat the above steps to merge the quantized acoustic eigenvalues of each column dimension to obtain the acoustic feature.
[0280] FIG11 shows a schematic diagram of the structure of a speech synthesis device provided by an exemplary embodiment of the present application. The device can be implemented as all or part of a computer device through software, hardware, or a combination of both. The device includes:
[0281] Acquisition module 1101 is configured to acquire sample semantic features, sample acoustic features, and sample audio, wherein the sample semantic features represent features of text information corresponding to the target audio, and the sample acoustic features are features of acoustic information corresponding to a sample reference speech, wherein the sample reference speech refers to the speech of an object corresponding to the selected timbre;
[0282] A feature processing module 1102 is configured to embed the sample semantic features into the sample acoustic features through a first-layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features, wherein the first-layer sub-model is configured to embed the sample semantic features into the sample acoustic features, and the sample intermediate acoustic features are configured to represent features obtained after embedding the sample semantic features into the sample acoustic features;
[0283] The feature processing module 1102 is further configured to input the intermediate acoustic features of the sample into a second-layer sub-model in the speech synthesis model to obtain audio synthesis features, wherein the second-layer sub-model is configured to synthesize the audio synthesis features based on the intermediate acoustic features of the sample, and the audio synthesis features are configured to represent features corresponding to the target audio;
[0284] A generating module 1103 is configured to generate a synthesized audio having the same timbre as the sample reference speech based on the audio synthesis feature;
[0285] A calculation module 1104 is configured to calculate a training loss of the speech synthesis model based on the sample audio and the synthesized audio;
[0286] The updating module 1105 is used to update the model parameters of the speech synthesis model according to the training loss.
[0287] In some embodiments, the feature processing module 1102 is further used to add the feature values of the same column dimension in the sample acoustic feature to obtain a one-dimensional sample acoustic feature; splice the sample semantic feature and the one-dimensional sample acoustic feature to obtain a sample splicing feature, where the sample splicing feature refers to a feature obtained by splicing the sample semantic feature and the one-dimensional sample acoustic feature; and input the sample splicing feature into the first layer sub-model to obtain the sample intermediate acoustic feature.
[0288] In some embodiments, the feature processing module 1102 is also used to input the sample splicing feature into the first layer sub-model to predict the first sample intermediate acoustic feature value in the sample intermediate acoustic feature, where the intermediate acoustic feature refers to the feature predicted by the first layer sub-model based on the splicing feature, and the sample intermediate acoustic feature value refers to the feature value in the sample intermediate acoustic feature; input the i-1th audio synthesis feature value corresponding to the i-1th sample intermediate acoustic feature value and the sample splicing feature into the first layer sub-model to predict the sample intermediate acoustic feature value to obtain the i-th sample intermediate acoustic feature value, where the i-1th audio synthesis feature value is predicted by inputting the i-1th sample intermediate acoustic feature value into the second layer sub-model; repeat the previous step until the number of the sample intermediate acoustic feature values outputted is equal to the number N of feature values in the one-dimensional sample acoustic feature, where i is a positive integer greater than 1 and less than N; merge N sample intermediate acoustic feature values to obtain the sample intermediate acoustic feature, where i is a positive integer greater than 1.
[0289] In some embodiments, the feature processing module 1102 is further configured to sequentially input the sample intermediate acoustic feature values in the sample intermediate acoustic feature into the second layer sub-model to obtain the audio synthesis feature.
[0290] In some embodiments, the feature processing module 1102 is also used to input the first sample intermediate acoustic feature value in the sample intermediate acoustic feature into the second layer sub-model to predict the first audio synthesis feature value in the audio synthesis feature; input the jth sample intermediate acoustic feature value generated by the j-1th audio synthesis feature value into the second layer sub-model to predict the audio synthesis feature value, and obtain the jth audio synthesis feature value in the audio synthesis feature, and the jth sample intermediate acoustic feature value is predicted by inputting the j-1th sample audio synthesis feature value into the first layer sub-model; loop the previous step until the number of the output audio synthesis feature values is equal to the number N of the sample intermediate acoustic feature values in the sample intermediate acoustic feature, j is a positive integer greater than 1 and less than N; merge N audio synthesis feature values to obtain the audio synthesis feature, and j is a positive integer greater than 1.
[0291] In some embodiments, the acquisition module 1101 is further used to acquire the sample reference speech; input the sample reference speech into the acoustic feature extraction network in the speech synthesis model to extract features, thereby obtaining the sample acoustic features.
[0292] In some embodiments, the computing module 1104 is also used to input the sample reference speech into the acoustic feature extraction network to extract features and obtain initial sample acoustic features, where the initial sample acoustic features refer to continuous features extracted from the sample reference speech; and quantify the initial sample acoustic features to obtain the sample acoustic features.
[0293] In some embodiments, the calculation module 1104 is further used to quantize the mth acoustic eigenvalue in the initial sample acoustic feature to obtain the first quantized acoustic eigenvalue of the mth column dimension in the sample acoustic feature, where m is a positive integer; quantize the residual between the first sample quantized acoustic eigenvalue and the mth sample acoustic eigenvalue to obtain the second sample quantized acoustic eigenvalue in the mth column dimension; quantize the residual between the kth quantized acoustic eigenvalue and the k-1th quantized acoustic eigenvalue to obtain the k+1th quantized acoustic eigenvalue in the mth column dimension, where k is a positive integer greater than 1; merge all the quantized acoustic eigenvalues in the mth column dimension to obtain the quantized acoustic eigenvalue of the mth column dimension in the sample acoustic feature; repeat the above steps to merge the quantized acoustic eigenvalues of each column dimension to obtain the sample acoustic feature.
[0294] FIG12 shows a block diagram of a computer device 1200 according to an exemplary embodiment of the present application. The computer device can be implemented as the server in the above-mentioned solution of the present application. The computer device 1200 includes a central processing unit (CPU) 1201, a system memory 1204 including a random access memory (RAM) 1202 and a read-only memory (ROM) 1203, and a system bus 1205 connecting the system memory 1204 and the central processing unit 1201. The computer device 1200 also includes a mass storage device 1206 for storing an operating system 1209, application programs 1210, and other program modules 1211.
[0295] The mass storage device 1206 is connected to the central processing unit 1201 via a mass storage controller (not shown) connected to the system bus 1205. The mass storage device 1206 and its associated computer-readable media provide non-volatile storage for the computer device 1200. In other words, the mass storage device 1206 may include a computer-readable medium (not shown) such as a hard disk or a compact disc read-only memory (CD-ROM) drive.
[0296] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer-readable instructions, data structures, program modules or other data. Computer storage media include RAM, Erasable Programmable Read Only Memory (EPROM), Electronically Erasable Programmable Read-Only Memory (EEPROM), flash memory or other solid-state storage technology, CD-ROM, Digital Versatile Disc (DVD) or other optical storage, tape cassettes, magnetic tape, disk storage or other magnetic storage devices. Of course, those skilled in the art will appreciate that the computer storage media are not limited to the above-mentioned ones. The above-mentioned system memory 1204 and mass storage device 1206 can be collectively referred to as memory.
[0297] According to various embodiments of the present disclosure, the computer device 1200 may also be connected to a remote computer on a network such as the Internet for operation. That is, the computer device 1200 may be connected to a network 1208 via a network interface unit 1207 connected to the system bus 1205, or the network interface unit 1207 may be used to connect to other types of networks or remote computer systems (not shown).
[0298] The memory also includes at least one computer program, which is stored in the memory. The central processing unit 1201 implements all or part of the steps in the speech synthesis method or speech synthesis model training method shown in the above-mentioned embodiments by executing the at least one program.
[0299] An embodiment of the present application also provides a computer device, which includes a processor and a memory, wherein the memory stores at least one program, and the at least one program is loaded and executed by the processor to implement the speech synthesis method or speech synthesis model training method provided by the above-mentioned method embodiments.
[0300] An embodiment of the present application also provides a computer-readable storage medium, which stores at least one computer program, and the at least one computer program is loaded and executed by a processor to implement the speech synthesis method or speech synthesis model training method provided by the above-mentioned method embodiments.
[0301] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes to implement the speech synthesis method or speech synthesis model training method provided in the above-mentioned method embodiments.
[0302] It is understandable that in the specific implementation methods of this application, the data involved, historical data, and portraits and other data related to user data processing related to user identity or characteristics, when the above embodiments of this application are applied to specific products or technologies, need to obtain user permission or consent, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.
[0303] It should be noted that, unless otherwise expressly defined herein, all terms used in the claims are to be interpreted according to their ordinary meaning in the technical field. Unless otherwise expressly stated, all references to "an element, device, component, device, step, etc." are to be interpreted openly as referring to at least one instance of an element, device, component, device, step, etc. Unless expressly stated otherwise, the steps of any method disclosed herein do not have to be performed in the exact order disclosed.
[0304] It should be understood that the term "plurality" used herein refers to two or more. "And / or" describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can mean: A exists alone, A and B exist simultaneously, or B exists alone. The character " / " generally indicates an "or" relationship between the associated objects.
[0305] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, which may be a read-only memory, a disk, or an optical disk, etc.
[0306] The above description is merely an optional embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent switches, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
Claims
1. A speech synthesis method, the method being executed by a computer device, the method comprising: Acquire semantic features and acoustic features, wherein the semantic features are used to represent text information corresponding to the target audio to be synthesized, and the acoustic features are used to represent acoustic information corresponding to a reference voice, wherein the reference voice refers to the voice of the object corresponding to the selected timbre; Embedding the semantic feature into the acoustic feature through a first layer sub-model in a speech synthesis model to obtain an intermediate acoustic feature, wherein the first layer sub-model is used to embed the semantic feature into the acoustic feature; Inputting the intermediate acoustic features into a second layer sub-model in the speech synthesis model to obtain audio synthesis features, wherein the second layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features; The target audio that matches the timbre of the reference voice is generated based on the audio synthesis feature.
2. According to the method of claim 1, embedding the semantic features into the acoustic features through the first layer sub-model in the speech synthesis model to obtain the intermediate acoustic features comprises: Adding the feature values of the same column dimension in the acoustic feature to obtain a one-dimensional acoustic feature; splicing the semantic feature and the one-dimensional acoustic feature to obtain a spliced feature, where the spliced feature refers to a feature obtained by splicing the semantic feature and the one-dimensional acoustic feature; The concatenated features are input into the first layer sub-model to obtain the intermediate acoustic features.
3. The method according to claim 2, wherein inputting the concatenated features into the first layer sub-model to obtain the intermediate acoustic features comprises: Inputting the concatenated features into the first layer sub-model to predict the intermediate acoustic feature values in the intermediate acoustic features, and obtaining the first intermediate acoustic feature value in the intermediate acoustic features; Inputting the i-1th audio synthesis feature value corresponding to the i-1th intermediate acoustic feature value and the concatenated feature into the first layer sub-model to predict the intermediate acoustic feature value, thereby obtaining the i-th intermediate acoustic feature value, wherein the i-1th audio synthesis feature value is obtained by inputting the i-1th intermediate acoustic feature value into the second layer sub-model for prediction; The previous step is repeated until the number of the intermediate acoustic eigenvalues outputted is equal to the number N of eigenvalues in the one-dimensional acoustic feature, where i is a positive integer greater than 1 and less than N; The N intermediate acoustic feature values are combined to obtain the intermediate acoustic feature.
4. According to the method of claim 1, the step of inputting the intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain the audio synthesis features comprises: The intermediate acoustic feature values in the intermediate acoustic feature are sequentially input into the second layer sub-model to obtain the audio synthesis feature.
5. According to the method of claim 4, the step of sequentially inputting the intermediate acoustic feature values in the intermediate acoustic feature into the second layer sub-model to obtain the audio synthesis feature comprises: Inputting a first intermediate acoustic feature value in the intermediate acoustic features into the second layer sub-model to predict and obtain a first audio synthesis feature value in the audio synthesis features; Inputting the j-th intermediate acoustic eigenvalue generated by the j-1-th audio synthesis eigenvalue into the second layer sub-model to predict the j-th audio synthesis eigenvalue, wherein the j-th intermediate acoustic eigenvalue is predicted by inputting the j-1-th audio synthesis eigenvalue into the first layer sub-model; Repeat the previous step until the number of the output audio synthesis feature values is equal to the number N of the intermediate acoustic feature values in the intermediate acoustic feature, where j is a positive integer greater than 1 and less than N; The N audio synthesis feature values are combined to obtain the audio synthesis feature.
6. The method according to any one of claims 1 to 5, wherein the acoustic feature is determined by: Acquiring the reference speech; Extract features from the reference speech to obtain the acoustic features.
7. The method according to claim 6, wherein extracting features from the reference speech to obtain the acoustic features comprises: Inputting the reference speech into the acoustic feature extraction network to extract features and obtain initial acoustic features, wherein the initial acoustic features refer to continuous features extracted from the reference speech; The initial acoustic feature is quantified to obtain the acoustic feature.
8. The method according to claim 7, wherein quantifying the initial acoustic features to obtain the acoustic features comprises: quantizing the mth acoustic feature value in the initial acoustic feature to obtain the first quantized acoustic feature value of the mth column dimension in the acoustic feature, where m is a positive integer; quantizing a residual between the first quantized acoustic eigenvalue and the m-th acoustic eigenvalue to obtain a second quantized acoustic eigenvalue in the m-th column dimension; quantizing a residual between the kth quantized acoustic eigenvalue and the k-1th quantized acoustic eigenvalue to obtain the k+1th quantized acoustic eigenvalue in the mth column dimension, where k is a positive integer greater than 1; Merging all quantized acoustic feature values obtained by quantization in the m-th column dimension to obtain the quantized acoustic feature value of the m-th column dimension in the acoustic feature; Repeat the above steps to combine the quantized acoustic feature values of each column dimension to obtain the acoustic feature.
9. A method for training a speech synthesis model, the method being executed by a computer device, the method comprising: Acquire sample semantic features, sample acoustic features and sample audio, wherein the sample semantic features are used to represent text information corresponding to the sample audio, and the sample acoustic features are used to represent acoustic information corresponding to the sample reference voice, wherein the sample reference voice refers to the voice of the object corresponding to the selected timbre; Embedding the sample semantic features into the sample acoustic features through a first layer sub-model in the speech synthesis model to obtain sample intermediate acoustic features, wherein the first layer sub-model is used to embed the sample semantic features into the sample acoustic features; Inputting the intermediate acoustic features of the sample into a second-layer sub-model in the speech synthesis model to obtain audio synthesis features, wherein the second-layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features of the sample; Generate a synthesized audio that matches the timbre of the sample reference speech based on the audio synthesis feature; Calculating the training loss of the speech synthesis model based on the sample audio and the synthesized audio; The model parameters of the speech synthesis model are updated according to the training loss.
10. The method according to claim 9, wherein embedding the sample semantic features into the sample acoustic features through the first layer sub-model in the speech synthesis model to obtain the sample intermediate acoustic features comprises: Adding the feature values of the same column dimension in the sample acoustic feature to obtain a one-dimensional sample acoustic feature; Splicing the sample semantic feature and the one-dimensional sample acoustic feature to obtain a sample splicing feature, where the sample splicing feature refers to a feature obtained by splicing the sample semantic feature and the one-dimensional sample acoustic feature; The sample concatenation features are input into the first layer sub-model to obtain the sample intermediate acoustic features.
11. The method according to claim 10, wherein inputting the sample concatenation feature into the first layer sub-model to obtain the sample intermediate acoustic feature comprises: Inputting the sample concatenation feature into the first layer sub-model to predict the sample intermediate acoustic feature value in the sample intermediate acoustic feature, to obtain the first sample intermediate acoustic feature value in the sample intermediate acoustic feature; Inputting the i-1th audio synthesis feature value corresponding to the i-1th sample intermediate acoustic feature value and the sample concatenation feature into the first layer sub-model to predict the sample intermediate acoustic feature value, thereby obtaining the i-th sample intermediate acoustic feature value, wherein the i-1th audio synthesis feature value is obtained by inputting the i-1th sample intermediate acoustic feature value into the second layer sub-model for prediction; Repeat the previous step until the number of the outputted intermediate acoustic feature values of the sample is equal to the number N of the feature values in the one-dimensional sample acoustic feature, where i is a positive integer greater than 1 and less than N; The N sample intermediate acoustic feature values are combined to obtain the sample intermediate acoustic feature.
12. The method according to claim 9, wherein the step of inputting the sample intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain the audio synthesis features comprises: The sample intermediate acoustic feature values in the sample intermediate acoustic feature are sequentially input into the second layer sub-model to obtain the audio synthesis feature.
13. The method according to claim 12, wherein the step of sequentially inputting the sample intermediate acoustic feature values in the sample intermediate acoustic feature into the second layer sub-model to obtain the audio synthesis feature comprises: Inputting a first sample intermediate acoustic feature value in the sample intermediate acoustic features into the second layer sub-model to predict a first audio synthesis feature value in the audio synthesis features; Inputting the j-th sample intermediate acoustic feature value generated by the j-1-th audio synthesis feature value into the second layer sub-model to predict the audio synthesis feature value, thereby obtaining the j-th audio synthesis feature value in the audio synthesis feature, wherein the j-th sample intermediate acoustic feature value is obtained by inputting the j-1-th sample audio synthesis feature value into the first layer sub-model for prediction; The previous step is repeated until the number of the output audio synthesis feature values is equal to the number N of the sample intermediate acoustic feature values in the sample intermediate acoustic feature, where j is a positive integer greater than 1 and less than N; The N audio synthesis feature values are combined to obtain the audio synthesis feature.
14. The method according to any one of claims 9 to 13, wherein the acoustic characteristics of the sample are determined by: Acquiring the sample reference speech; Extract features from the sample reference speech to obtain the sample acoustic features.
15. The method according to claim 14, wherein extracting features from the sample reference speech to obtain the sample acoustic features comprises: Inputting the sample reference speech into the acoustic feature extraction network to extract features, thereby obtaining initial sample acoustic features, wherein the initial sample acoustic features refer to continuous features extracted from the sample reference speech; The initial sample acoustic feature is quantified to obtain the sample acoustic feature.
16. A speech synthesis device, comprising: An acquisition module, used to acquire semantic features and acoustic features, wherein the semantic features are used to represent text information corresponding to the target audio to be synthesized, and the acoustic features are used to represent acoustic information corresponding to the reference voice, wherein the reference voice refers to the voice of the object corresponding to the selected timbre; A feature processing module, used for embedding the semantic feature into the acoustic feature through a first layer sub-model in a speech synthesis model to obtain an intermediate acoustic feature, wherein the first layer sub-model is used for embedding the semantic feature into the acoustic feature; The feature processing module is used to input the intermediate acoustic features into the second layer sub-model in the speech synthesis model to obtain audio synthesis features, and the second layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features; A generation module is used to generate a synthesized audio with a timbre that matches the reference speech based on the audio synthesis feature.
17. A training device for a speech synthesis model, the device comprising: An acquisition module, used to acquire sample semantic features, sample acoustic features and sample audio, wherein the sample semantic features are used to represent semantic text information corresponding to the target audio to be synthesized, and the sample acoustic features are used to represent acoustic information corresponding to the sample reference speech, wherein the sample reference speech refers to the speech of the object corresponding to the selected timbre; A feature processing module, used for embedding the sample semantic feature into the sample acoustic feature through a first layer sub-model in the speech synthesis model to obtain a sample intermediate acoustic feature, wherein the first layer sub-model is used for embedding the sample semantic feature into the sample acoustic feature; The feature processing module is used to input the intermediate acoustic features of the sample into the second layer sub-model in the speech synthesis model to obtain audio synthesis features, and the second layer sub-model is used to synthesize the audio synthesis features based on the intermediate acoustic features of the sample; A generating module, configured to generate a synthesized audio conforming to the timbre of the sample reference speech based on the audio synthesis feature; A calculation module, used for calculating the training loss of the speech synthesis model based on the sample audio and the synthesized audio; An updating module is used to update the model parameters of the speech synthesis model according to the training loss.
18. A computer device, comprising: A processor and a memory, wherein at least one computer program is stored in the memory, and at least one of the computer programs is loaded and executed by the processor to implement the speech synthesis method as described in any one of claims 1 to 8, or the training method of the speech synthesis model as described in any one of claims 9 to 15.
19. A computer storage medium, wherein at least one computer program is stored in the computer-readable storage medium, and the at least one computer program is loaded and executed by a processor to implement the speech synthesis method as described in any one of claims 1 to 8, or the training method of the speech synthesis model as described in any one of claims 9 to 15.
20. A computer program product, comprising a computer program, wherein the computer program is stored in a computer-readable storage medium; the computer program is read and executed from the computer-readable storage medium by a processor of a computer device, so that the computer device executes the speech synthesis method as described in any one of claims 1 to 8, or the training method of a speech synthesis model as described in any one of claims 9 to 15.
Citation Information
Patent Citations
Voice synthesis method, model training method and device, and computer equipment
CN109036375A
Speech synthesis method and device
CN116052640A
Cross-statement conditional coherence voice editing method, system and terminal
CN116189653A
Speech synthesis method and device, storage medium and electronic equipment
CN116312476A
Speech generation method and device based on pre-training language model, equipment and medium
CN116364055A