Audio synthesis method and device, storage medium and electronic device

By using a weight distribution network based on discretization-based hybrid logic distribution structure, the problem of instability of the audio synthesis model is solved, the accuracy and stability of the audio synthesis are achieved, and the problems of missed reading and repeated reading are avoided.

CN113763922BActive Publication Date: 2025-08-12TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202110517152.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-05-12
Publication Date
2025-08-12
Estimated Expiration
2041-05-12

AI Technical Summary

Technical Problem

The existing audio synthesis model is unstable, resulting in low accuracy of synthetic audio and problems of missed reading and repeated reading.

Method used

A weight distribution network based on discretization-based hybrid logical distribution structure is adopted to convert the text sequence into an abstract feature sequence, and a context vector is generated through monotonic constraints, and the audio spectrum information matching the context vector is obtained, and the target audio is finally synthesized.

Benefits of technology

Ensure the monotonic stability of the audio synthesis process, avoid omissions and repetition, and improve the accuracy and stability of audio synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113763922B_ABST
    Figure CN113763922B_ABST
Patent Text Reader

Abstract

The present invention discloses an audio synthesis method and apparatus, a storage medium, and an electronic device. The method comprises: obtaining a text sequence to be processed; converting the text sequence into an abstract feature sequence; inputting the abstract feature sequence into a weight distribution network with a monotonicity constraint to obtain a context vector corresponding to the abstract feature sequence, wherein the weight distribution network is constructed based on a discretized hybrid logic distribution structure; obtaining audio spectrum information that matches the context vector; and synthesizing target audio that matches the text sequence using the audio spectrum information. The present invention solves the technical problem of low synthesized audio accuracy caused by unstable audio synthesis models.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computers, and in particular to an audio synthesis method and device, a storage medium, and an electronic device. Background Art

[0002] Nowadays, in order to improve the efficiency of human-computer interaction, more and more applications or businesses are beginning to use synthesized audio to provide users with customized auxiliary services, such as using synthesized audio to broadcast news to users or provide users with map navigation services, thereby freeing users' hands and eliminating the need for users to input interactive control commands on touch screen devices.

[0003] However, the audio synthesis technology currently provided by related technologies not only requires the use of a large amount of corpus to train the audio synthesis model, but the trained audio synthesis model often has problems such as missed reading and repeated reading, making it difficult to ensure the stability of the synthesized audio, and it is also difficult to ensure the accuracy of the synthesized audio.

[0004] To address the above-mentioned problems, no effective solutions have been proposed so far. Summary of the Invention

[0005] Embodiments of the present invention provide an audio synthesis method and apparatus, a storage medium, and an electronic device to at least solve the technical problem of low synthesized audio accuracy caused by an unstable audio synthesis model.

[0006] According to one aspect of an embodiment of the present invention, an audio synthesis method is provided, comprising: obtaining a text sequence to be processed; converting the text sequence into an abstract feature sequence; inputting the abstract feature sequence into a weight distribution network with a monotonicity constraint to obtain a context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a decentralized hybrid logical distribution structure; obtaining audio spectrum information matching the context vector; and synthesizing target audio matching the text sequence using the audio spectrum information.

[0007] According to another aspect of an embodiment of the present invention, an audio synthesis device is also provided, including: a first acquisition module for acquiring a text sequence to be processed; a conversion module for converting the above text sequence into an abstract feature sequence; an input module for inputting the above abstract feature sequence into a weight distribution network with a monotonicity constraint to obtain a context vector corresponding to the above abstract feature sequence, wherein the above weight distribution network is a network constructed based on a discretized hybrid logical distribution structure; a second acquisition module for acquiring audio spectrum information matching the above context vector; and a synthesis module for using the above audio spectrum information to synthesize target audio matching the above text sequence.

[0008] According to another aspect of the embodiments of the present invention, a computer-readable storage medium is provided, in which a computer program is stored. The computer program is configured to execute the above-mentioned audio synthesis method when running.

[0009] According to another aspect of an embodiment of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the audio synthesis method through the computer program.

[0010] In an embodiment of the present invention, a text sequence is converted into an abstract feature sequence, and the abstract feature sequence is converted into a context vector using a weight distribution network with a monotonicity constraint condition. Then, audio spectrum information matching the context vector is obtained, and the target audio is generated using the audio spectrum information. A context vector with a monotonicity constraint condition is generated through a weight distribution network, so that the context vector with a monotonicity constraint condition is applied in the process of spectrum acquisition and audio generation, thereby avoiding omissions, repetitions, direction errors and other problems in the audio synthesis process, and achieving the purpose of ensuring the monotonic stability of the audio synthesis in the direction, thereby achieving the technical effect of ensuring the accuracy and stability of the audio synthesis model in the audio synthesis process, and thus solving the technical problem of low accuracy of synthesized audio caused by the instability of the audio synthesis model. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] The drawings described herein are used to provide a further understanding of the present invention and constitute a part of this application. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation of the present invention. In the drawings:

[0012] Figure 1 is a schematic diagram of an application environment of an optional audio synthesis method according to an embodiment of the present invention;

[0013] Figure 2 is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0014] Figure 3 is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0015] Figure 4 is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0016] Figure 5 is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0017] Figure 6is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0018] Figure 7 is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0019] Figure 8 is a schematic structural diagram of an optional acoustic model according to an embodiment of the present invention;

[0020] Figure 9 is a flowchart of an optional audio synthesis method according to an embodiment of the present invention;

[0021] Figure 10 is a schematic structural diagram of an optional audio adversarial generative network according to an embodiment of the present invention;

[0022] Figure 11 is a schematic structural diagram of an optional audio synthesis device according to an embodiment of the present invention;

[0023] Figure 12 FIG. 4 is a schematic structural diagram of an optional electronic device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0024] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0025] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0026] According to one aspect of an embodiment of the present invention, an audio synthesis method is provided. Optionally, the audio synthesis method can be applied to, but is not limited to, Figure 1In the environment shown, the terminal device 102 exchanges data with the server 122 via the network 110. The server 122 runs a database 124 and a processing engine 126. The processing engine 126 obtains the data stored in the database 124 and processes the data.

[0027] The terminal device 102 collects the text sequence and sends the text sequence to the server 122 through the network 110. The processing engine 126 in the server 122 executes S102 to S110 in sequence. The text sequence to be processed is obtained from the database 124. When the text sequence is obtained, the text sequence is converted into an abstract feature sequence. When the abstract feature sequence is obtained, the abstract feature sequence is input into a weight distribution network with a monotonicity constraint to obtain a context vector corresponding to the abstract feature sequence. The weight distribution network is a network constructed based on a discretized hybrid logical distribution structure. When the context vector is obtained, the audio spectrum information matching the context vector is obtained. The target audio matching the text sequence is synthesized using the obtained audio spectrum information.

[0028] The server 122 sends the synthesized target audio to the terminal device 102 through the network 110, thereby obtaining the target audio that matches the text sequence.

[0029] Optionally, in this embodiment, the above-mentioned terminal device can be a terminal device configured with a target client, which can include but is not limited to at least one of the following: a mobile phone (such as an Android phone, an IOS phone, etc.), a laptop computer, a tablet computer, a PDA, an MID (Mobile Internet Devices), a PAD, a desktop computer, a smart TV, etc. The target client can be a client that can collect text sequences and play target audio, and is not limited to a video client, an instant messaging client, a browser client, an educational client, etc. The above-mentioned network can include but is not limited to: a wired network, a wireless network, wherein the wired network includes: a local area network, a metropolitan area network and a wide area network, and the wireless network includes: Bluetooth, WIFI and other networks that realize wireless communication. The above-mentioned server can be a single server, or it can be a server cluster composed of multiple servers, or a cloud server. The above is only an example, and this embodiment does not impose any limitation on this.

[0030] As an optional implementation, Figure 2 As shown, the above audio synthesis method includes:

[0031] S202, obtaining a text sequence to be processed;

[0032] S204, converting the text sequence into an abstract feature sequence;

[0033] S206, inputting the abstract feature sequence into a weight distribution network with a monotonicity constraint to obtain a context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a discretized hybrid logic distribution structure;

[0034] S208, obtaining audio spectrum information matching the context vector;

[0035] S210: synthesize target audio that matches the text sequence using the audio spectrum information.

[0036] Optionally, the text sequence can be a text sequence obtained by converting the acquired audio corpus into text. Generating target audio that matches the text sequence can be achieved by using a target acoustic model and an audio adversarial network to generate audio that matches the audio corpus corresponding to the text sequence in terms of sound quality. Sound quality matching is not limited to similarity in timbre.

[0037] Optionally, the target acoustic model is not limited to obtaining matching audio spectrum information based on a text sequence and synthesizing the target audio based on the audio spectrum information using an audio adversarial generative network. The target acoustic model is not limited to a model based on a sequence-to-sequence (Seq2seq) structure, including a feature extraction network, a weight allocation network, and a target spectrum generation network connected in sequence. The audio adversarial generative network is not limited to a synthesizer using a generative adversarial network (GAN) framework.

[0038] Optionally, the text sequence is not limited to being converted into an abstract feature sequence by the feature extraction network in the target acoustic model. The feature extraction network is not limited to including a phoneme feature converter and a content encoder.

[0039] Optionally, phoneme feature conversion is not limited to converting a text sequence into a phoneme sequence containing linguistic features, and the feature information in the linguistic features is not limited to including: Chinese phonemes, English phonemes, Chinese vowel tones, word boundaries, phrase boundaries, and sentence boundaries.

[0040] Optionally, the content encoder is not limited to converting a phoneme sequence corresponding to a text sequence into an abstract feature sequence. The content encoder is not limited to being a second preprocessing network and a residual connection network, and the phoneme sequence processed by the second preprocessing network is input into the residual connection network. The residual connection network can be a network model composed of a set of one-dimensional convolutional layers, a high-speed network, and a bidirectional GRU network, and is used to improve the accuracy of converting the phoneme sequence into the abstract feature sequence.

[0041] Optionally, the weight distribution network is not limited to mapping an abstract feature sequence into a context vector containing context information. The weight distribution network can be, but is not limited to, a discretized mixture of logistic (MOL) attention model network. The discretized MOL distribution is used to ensure that the context output by the weight distribution network has a monotonicity constraint.

[0042] Optionally, a target spectrum generation network in the target acoustic model is used to obtain audio spectrum information that matches the context vector. The audio spectrum information is not limited to mel-spectrogram information. A mel-spectrogram (Mel) is a spectrum obtained by Fourier transforming an acoustic signal and then applying a mel-scale transformation.

[0043] In an embodiment of the present application, a text sequence is converted into an abstract feature sequence, and the abstract feature sequence is converted into a context vector using a weight distribution network with a monotonicity constraint. Then, audio spectrum information matching the context vector is obtained, and the target audio is generated using the audio spectrum information. A context vector with a monotonicity constraint is generated through a weight distribution network, and the context vector with a monotonicity constraint is applied to the process of spectrum acquisition and audio generation. This avoids problems such as omissions, repetitions, and directional errors during the audio synthesis process, and achieves the purpose of ensuring the monotonic stability of the audio synthesis in direction, thereby achieving the technical effect of ensuring the accuracy and stability of the audio synthesis model during the audio synthesis process, and thus solving the technical problem of low accuracy of synthesized audio caused by the instability of the audio synthesis model.

[0044] As an optional implementation manner, the above-mentioned obtaining audio spectrum information matching the context vector includes:

[0045] In a target spectrum generation network connected to the weight allocation network, frame spectrum information of one or at least two audio frames matching the context vector is obtained, wherein a timer is configured in the target spectrum generation network, and the timer is used to segment the audio spectrum information generated in the target spectrum generation network to generate frame spectrum information corresponding to each audio frame.

[0046] Optionally, the target spectrum generation network can be a Mel-spectrum Residual Network (SNR) consisting of a first preprocessing subnetwork and a multi-layer long short-term memory network. The first preprocessing subnetwork, as a preprocessing network in the target spectrum generation network, processes the context vector in the first preprocessing subnetwork and then inputs it into the multi-layer long short-term memory network to obtain frame spectrum information.

[0047] Alternatively, taking the example of a target spectrum generation network comprising two layers of long short-term memory networks, the target spectrum generation network comprises a first preprocessing subnetwork, a first long short-term memory network, and a second long short-term memory network. The mel-spectrogram information output by the first preprocessing subnetwork is used as input to the second long short-term memory network, and the mel-spectrogram information output by the long short-term memory network is used as input for the frame spectrum information of the current frame to construct a mel-spectrogram residual network. By constructing a multi-layer mel-spectrogram residual network using multiple layers of long short-term memory networks, the accuracy of frame spectrum information generation is improved through residual connections.

[0048] Optionally, the timer may be a fully connected network for predicting a stop token. When a stop token appears, the generated audio spectrum information is segmented to generate frame spectrum information corresponding to each audio frame.

[0049] Optionally, when generating the frame spectrum information of the current audio frame, the target spectrum generation network uses the frame spectrum information of the previous frame as input to generate the frame spectrum information of the current frame. The frame spectrum information of the previous frame is not limited to being input into the first preprocessing subnetwork to be input into the target spectrum generation network.

[0050] In an embodiment of the present application, a multi-layer long short-term memory network is used to generate a Mel spectrum residual network structure in a target spectrum generation network, thereby improving the accuracy of frame spectrum information generation during the frame spectrum information generation process, thereby improving the accuracy of target audio synthesis.

[0051] As an optional implementation, Figure 3 As shown, in the target spectrum generation network connected to the weight allocation network, obtaining frame spectrum information of one or at least two audio frames matching the context vector includes:

[0052] The following operations are performed in sequence in the target spectrum generation network to generate frame spectrum information:

[0053] S302, obtaining the context vector currently received from the weight distribution network and reference frame spectrum information of the previous audio frame before the current audio frame to be generated;

[0054] S304: Input the context vector and the reference frame spectrum information into the first preprocessing sub-network and the multi-layer long short-term memory network to generate current frame spectrum information of the current audio frame.

[0055] Optionally, the reference frame spectrum information is the frame spectrum information of the previous audio frame before the current audio frame. The frame spectrum information of the previous audio frame is used as the reference frame spectrum information and the context vector is used as the input of the first preprocessing sub-network, so that the target spectrum generation network generates the frame spectrum information of the current frame through the first preprocessing sub-network and the multi-layer long short-term memory network.

[0056] As an optional implementation, Figure 4 As shown, before obtaining the text sequence to be processed, it also includes:

[0057] S402, constructing an initial acoustic model, wherein the initial acoustic model includes: a feature extraction network for extracting features, an initial weight allocation network, and an initial spectrum generation network;

[0058] S404: Pre-train the initial acoustic model using the first sample corpus until a first generation convergence condition is met to obtain a reference acoustic model, wherein the first generation convergence condition indicates that a difference between the generated audio spectrum information and the corresponding label spectrum information is less than a first threshold:

[0059] S406, using the second sample corpus to train the reference acoustic model until a second generation convergence condition is obtained to obtain a target acoustic model, wherein the second generation convergence condition indicates that the difference between the generated audio spectrum information and the corresponding label spectrum information is less than a second threshold, the number of the second sample corpus is less than the number of the first sample corpus, and the target acoustic model includes a feature extraction network, a weight allocation network and a target spectrum generation network that have completed training.

[0060] Optionally, the first sample corpus may be a multilingual corpus containing multiple objects. The fact that the number of the second sample corpus is smaller than the number of the first sample corpus may be that the number of objects contained in the second sample corpus is smaller than the number of objects contained in the first sample corpus. For example, the first sample corpus may contain a large number of objects, such as a corpus of fifty people, and the second sample corpus may be a corpus of the target object. By using the first sample corpus to train the initial acoustic model into a universal reference acoustic model, and then using the second sample corpus of the target object to perform targeted training on the reference acoustic model, a target acoustic model is obtained.

[0061] Optionally, pre-training the initial acoustic model is not limited to optimizing parameters in the feature extraction network, the initial weight assignment network, and the initial spectrum generation network included in the initial acoustic model. Optimizing the parameters included in the initial acoustic model is not limited to optimizing the initial parameters using a stochastic gradient descent (SGD) algorithm.

[0062] Optionally, the second threshold is smaller than the first threshold. Parameters of the reference acoustic model are optimized using the second sample corpus, such that the difference between the generated audio spectrum information and the label spectrum information changes, thereby making the accuracy of the frame spectrum information generated by the target acoustic model higher than the frame spectrum information generated by the reference acoustic model.

[0063] In an embodiment of the present application, the initial acoustic model is trained using a first sample corpus with a larger amount of corpus to obtain a reference acoustic model, and the reference acoustic model is trained using a second sample corpus with a smaller amount of corpus to obtain a targeted target acoustic model. This allows the reference acoustic model to have universal applicability while enabling fine-tuning of the corpus on the reference acoustic model. Thus, by training based on a corpus with fewer target objects, the target acoustic model can be obtained based on the reference acoustic model, reducing the corpus required to obtain the target acoustic model.

[0064] As an optional implementation, Figure 5 As shown, in the process of training the reference acoustic model using the second sample corpus, the following steps are also included:

[0065] S502, obtaining a spectrum training result obtained each time during the training process of the reference acoustic model;

[0066] S504: When the spectrum training result indicates that the model parameters in the reference acoustic model should be adjusted, the network parameters in the weight allocation network and the spectrum generation network that are in a non-frozen state are updated, and the network parameters in the feature extraction network that is in a frozen state are maintained.

[0067] Optionally, when a reference acoustic model is obtained, network parameters included in the feature extraction network are adjusted to a frozen state. In the frozen state, the network parameters remain unchanged and are not updated and optimized following the training of the reference acoustic model.

[0068] Optionally, the spectrum training result indicates that the model parameters need to be adjusted to achieve a second generation convergence condition, where the difference between the audio spectrum information generated by the current spectrum generation network and the corresponding label spectrum information is greater than or equal to a second threshold.

[0069] In an embodiment of the present application, by freezing the network parameters in the feature extraction network of the reference acoustic model, when training the reference acoustic model, when adjusting the model parameters, only the network parameters in the weight allocation network and the spectrum generation network are optimized and updated. When the reference acoustic model is obtained, the parameters in the feature extraction network can be fixed to improve the stability of the transfer training. At the same time, the training efficiency of the target acoustic model can be improved by reducing the number of parameters required for training adjustment, thereby improving the synthesis efficiency of the target audio.

[0070] As an optional implementation, Figure 6 As shown, after building the initial acoustic model, it also includes:

[0071] S602, obtaining a pronunciation representation vector of a sound source object;

[0072] S604, during the training of the reference acoustic model, the pronunciation representation vector of the sound source object is added to the second preprocessing network in the feature extraction network being trained, the residual connection network in the feature extraction network being trained, the gated loop structure in the weight allocation network being trained, and the multi-layer long short-term memory network of the spectrum generation network being trained.

[0073] Optionally, the pronunciation representation vector of the sound source object can be the pronunciation representation vector (speaker embedding) of the target object corresponding to the second sample corpus used to train the reference acoustic model. The pronunciation representation vector is added in each parameter optimization process of the reference acoustic model.

[0074] In an embodiment of the present application, in the process of optimizing the parameters of the reference acoustic model, the pronunciation characteristics of the sound source object are integrated into the parameter optimization process in combination with the pronunciation representation vector of the sound source object. The pronunciation representation vector is combined in multiple processes to avoid inputting the pronunciation representation vector into only one process, which causes the pronunciation representation vector to be diluted as the model is calculated in depth. In the training process of the reference acoustic model, the feature intervention of the pronunciation representation vector is performed multiple times to fully reflect the pronunciation characteristics related to the sound source object, so that the target acoustic model obtained by training is more matched with the sound source object in terms of timbre similarity. While improving the timbre similarity matching, the training speed of the reference acoustic model can be accelerated based on the sound source feature vector, thereby improving the training efficiency of the target acoustic model and improving the generation efficiency of the target audio.

[0075] As an optional implementation, Figure 7 As shown, when the text sequence is a Chinese text sequence, after building the initial acoustic model, it also includes:

[0076] S702, obtaining the tone features of the Chinese text sequence;

[0077] S704 , in the process of training the reference acoustic model, adding the pitch feature to the network structure after the second preprocessing network in the feature extraction network being trained.

[0078] Optionally, when the text sequence includes a Chinese text sequence, a tone feature of the Chinese text sequence is obtained, where the tone feature is used to indicate the pronunciation tone of the Chinese text.

[0079] Optionally, the pitch features are used as input features of the residual connection network in the feature extraction network, and are input into the residual connection network in the feature extraction network, and the residual connection network is used to input the pitch features into each level of the reference acoustic model.

[0080] In an embodiment of the present application, the tonal features of Chinese are input into the residual connection network following the second preprocessing network, thereby avoiding the influence of the second preprocessing network on the tonal features when processing noise and improving the noise resistance of the tonal features. At the same time, the tonal features can also be input into different network layers of the reference acoustic model through the residual connection network. During the training of different layers, the processing accuracy of Chinese tones is improved, thereby obtaining a target acoustic model with higher tonal processing accuracy. By improving the tonal accuracy of the target acoustic model, the pronunciation accuracy of the target audio generation is improved.

[0081] The target acoustic model obtained by training is not limited to Figure 8 As shown in the figure, the feature extraction network includes a text-to-linguistic convertor for converting text sequences into phoneme sequences and a content encoder. The content encoder includes a second preprocessing network (Pre-net) and a residual connection network (res-CBHG). The weight allocation network includes an attention weight model (MOL-attention) and a gated recurrent control unit (GRU). The target spectrum generation network includes a first preprocessing network (Pre-net) and a two-layer long short-term memory network (res-LSTM).

[0082] The text sequence "setence" is input into the text-to-linguistic convertor, which converts the text sequence into a phoneme sequence composed of phoneme features. The converted phoneme sequence is then input into the content encoder. The second preprocessing network in the content encoder, Pre-net, combines the input pronunciation representation vector "speaker embedding" to preprocess the phoneme sequence. The preprocessed phoneme sequence and pitch features are then input into a residual connection network. The pronunciation representation vector "speaker embedding" is then input into the residual connection network again, and the residual connection network is processed to obtain an abstract feature sequence. The resulting abstract feature sequence is then input into the weight allocation network. The pronunciation representation vector "speaker embedding", the predicted mel spectrum information of the previous frame, and the previously obtained context vector "Context" are simultaneously input into the gated recurrent control unit (GRU). This generates the current context vector "Context" which is monotonic and contains contextual features, as output by the discretized hybrid logistic distributed attention model (MOL-attention).

[0083] The current context vector Context and the predicted mel spectrum information of the previous frame are input into the first pre-processing network (pre-net) of the target spectrum generation network for preprocessing. The output of the first pre-processing network (pre-net) is used as the input of the second-layer long short-term memory network, and the output of the first-layer long short-term memory network is used as the input for generating the current mel spectrum information (previous mel). This constructs a target spectrum generation network with a residual connection structure. Furthermore, the target spectrum generation network includes a timer consisting of a stop prediction unit and a time-delayed short-term memory network (TD-LSTM) to indicate the completion of generating the current mel spectrum information (previous mel).

[0084] The target acoustic model utilizes a monotonicity-constrained mixed discrete distribution attention weight model (MOL-attention) to output context vectors for text sequences, enhancing the monotonicity and stability of the context vector output. Residual connection structures are added to both the content encoder and the spectrum generation model. This ensures the accuracy and stability of the mel-spectrogram output during gradient descent parameter optimization of the initial and reference acoustic models. This also accelerates the convergence of the acoustic model and improves its training efficiency. By fully integrating pronunciation feature vectors into multiple network structures of the acoustic model, the training speed of the acoustic model is accelerated while preventing dilution of the pronunciation feature vectors during training. This improves the timbre similarity of the audio and thus the accuracy of the target audio. Furthermore, pitch features are incorporated into the residual connection network of the content encoder to improve the pitch accuracy of the target audio. The residual connection network effectively ensures pitch accuracy during transfer learning.

[0085] Furthermore, during sound model training, the initial acoustic model is first trained based on a large number of predictions to obtain a reference acoustic model with transferability. The transferable parameters in the reference acoustic model are then frozen and fine-tuned based on a small number of predictions of the source object to obtain the target acoustic model. This enables the prediction and generation of audio information based on a small amount of source object corpus. This improves the stability and accuracy of the acoustic model, thereby increasing its training efficiency and, consequently, the accuracy and efficiency of target audio synthesis.

[0086] As an optional implementation, synthesizing target audio that matches the text sequence using audio spectrum information includes:

[0087] The audio spectrum information is input into an audio adversarial generation network to obtain the target audio, wherein the audio adversarial generation network includes a generation subnetwork for generating audio and a discrimination subnetwork for discrimination. The discrimination subnetwork includes: a phase discrimination subnetwork for discriminating phase information in the audio spectrum information, and a period discrimination subnetwork for discriminating period information in the audio spectrum information.

[0088] Optionally, the audio adversarial generative network includes a generator subnetwork and a discriminator subnetwork. The discriminator subnetwork is used to discriminate the audio synthesized by the generator subnetwork, and output the target audio if the synthesized audio passes the discrimination of the discriminator subnetwork.

[0089] Optionally, the generation subnetwork may be, but is not limited to, an encoding network-decoding network structure. The encoding network is composed of a convolutional neural network and is used to extract spectral features contained in the input audio spectrum information. The decoding network is composed of a deconvolutional neural network and is used to extract spectral features obtained by the encoding network according to the constraints corresponding to the loss function.

[0090] Optionally, the loss function may include, but is not limited to, a multi-resolution Fourier transform loss (STFT loss) and a multi-resolution mel-spectrogram loss. The loss function of the generative subnetwork is obtained by weighted summing the two loss functions.

[0091] Optionally, the discriminant subnetwork may be, but is not limited to, a fully convolutional neural network, configured to determine the probability of similarity between the input spectrum information and the label spectrum information. The discriminant subnetwork includes a phase discriminant subnetwork for determining the probability of similarity between phase information contained in the audio spectrum information and a period discriminant subnetwork for determining the probability of similarity between period information contained in the audio spectrum information.

[0092] In the embodiment of the present application, by setting up a generative subnetwork and a discriminative subnetwork, the adversarial nature of the generative subnetwork and the discriminative subnetwork is further enhanced. During the process of synthesizing the target audio from the spectrum information, the spectrum is discriminatively synthesized, integrating frequency discrimination and period discrimination. By discriminating the synthesized audio in multiple dimensions, the accuracy of the synthesized target audio is improved. At the same time, the two loss functions are used to accelerate the convergence speed of the audio adversarial generative network and improve the synthesis efficiency of the target audio.

[0093] As an optional implementation, Figure 9 As shown, before obtaining the text sequence to be processed, it also includes:

[0094] S902, using positive sample audio pairs and negative sample audio pairs to perform cross-adversarial training on the initial audio adversarial generation network until convergence conditions are reached, wherein the positive sample audio pairs include the audio to be distinguished and the labeled audio, and the negative sample audio pairs include the reference audio and the labeled audio generated in the generation subnetwork based on the audio spectrum information of the audio to be distinguished;

[0095] S904, during the training process, using the positive sample audio pairs to train the discriminant subnetwork until the trained discriminant subnetwork meets a first discrimination condition, wherein the first discrimination condition indicates that the discriminant subnetwork recognizes that the audio to be discriminated is the labeled audio with a first confidence level greater than a third threshold;

[0096] S906, saving the network parameters of the discriminant sub-network;

[0097] S908, using negative sample audio pairs to train the initial generative sub-network until the discriminative sub-network reaches a second discrimination condition, wherein the second discrimination condition indicates that the second confidence level of the discriminative sub-network in identifying the reference audio as the labeled audio is greater than a fourth threshold, and adjusting the network parameters of the generative sub-network in training based on the Fourier transform loss between the reference audio and the labeled audio, and the Mel spectrum residual loss between the reference audio and the labeled audio.

[0098] Optionally, the training of the generative subnetwork and the discriminative subnetwork is repeated multiple times. Each training in the training process is not limited to: first, the discriminative subnetwork is trained, a positive sample audio pair is input into the discriminative subnetwork, and the ability to distinguish between the audio to be distinguished and the label audio is enhanced through comprehensive discrimination of periodic information and phase information. The parameters of the discriminative subnetwork are adjusted based on periodic information discrimination and phase information discrimination. When the output result of the discriminative subnetwork indicates that the probability of the audio to be distinguished is greater than the third threshold, the discrimination parameters of the current discriminative subnetwork are fixed, and the generative subnetwork is trained. The reference audio generated by the generative subnetwork is used as the audio to be distinguished, and the network parameters of the generative subnetwork are updated and optimized using the Fourier transform loss and the Mel spectrum residual loss, so that the generated reference audio is input into the current discriminative subnetwork as the audio to be distinguished, and the output result obtained indicates that the probability of the reference audio being the label audio is greater than the fourth threshold. When the conditions are met, the network parameters of the current generative subnetwork are fixed, and the discriminative subnetwork is trained again, thus looping.

[0099] The network model of the audio adversarial network is not limited to Figure 10As shown in the figure, the audio adversarial network includes a generator subnetwork and a predicted audio discriminator subnetwork. The loss function in the generator subnetwork includes a multi-resolution Fourier transform loss (multi-resolution STFT loss) and a multi-resolution mel-spectrogram loss. The predicted audio discriminator subnetwork includes a phase-aware frequency discriminator subnetwork and a multi-period discriminator subnetwork.

[0100] The Mel spectrum information is input into the generator sub-network, and the generated audio is input into the predicted audio discriminant sub-network for discrimination. If the discriminant sub-network discriminates, the generated audio is used as the target audio.

[0101] In the embodiment of the present application, by introducing multi-resolution STFT loss and multi-resolution mel-spectrogram loss into the generation subnetwork, the convergence speed of the generation subnetwork is accelerated, and the efficiency of target audio generation is improved. At the same time, a multi-period discriminator is combined with a phase-aware frequency discriminator, and phase information is used together with amplitude information as phase discrimination. Combined with period information discrimination, a discriminant subnetwork is constructed to perform adversarial discrimination on the generated audio, thereby improving the accuracy of target audio generation.

[0102] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should be aware that the present invention is not limited by the order of the actions described, because according to the present invention, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0103] According to another aspect of the embodiments of the present invention, there is also provided an audio synthesis device for implementing the above audio synthesis method. Figure 11 As shown, the device includes:

[0104] A first acquisition module 1102 is used to acquire a text sequence to be processed;

[0105] A conversion module 1104, configured to convert a text sequence into an abstract feature sequence;

[0106] An input module 1106 is configured to input the abstract feature sequence into a weight distribution network with a monotonicity constraint to obtain a context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a discretized hybrid logic distribution structure;

[0107] A second acquisition module 1108 is configured to acquire audio spectrum information matching the context vector;

[0108] The synthesis module 1110 is configured to synthesize target audio that matches the text sequence using audio spectrum information.

[0109] Optionally, the second acquisition module 1108 is also used to: obtain frame spectrum information of one or at least two audio frames matching the context vector in a target spectrum generation network connected to the weight distribution network, wherein a timer is configured in the target spectrum generation network, and the timer is used to segment the audio spectrum information generated in the target spectrum generation network to generate frame spectrum information corresponding to each audio frame.

[0110] Optionally, the above-mentioned second acquisition module 1108 is also used to obtain frame spectrum information of one or at least two audio frames matching the context vector in a target spectrum generation network connected to the weight distribution network, including: performing the following operations in sequence in the target spectrum generation network to generate frame spectrum information: obtaining the context vector currently received from the weight distribution network, and the reference frame spectrum information of the previous audio frame before the current audio frame to be generated; inputting the context vector and the reference frame spectrum information into the first preprocessing sub-network and the multi-layer long short-term memory network to generate the current frame spectrum information of the current audio frame.

[0111] Optionally, the audio generation device further includes a first training module, configured to:

[0112] Constructing an initial acoustic model, wherein the initial acoustic model includes: a feature extraction network for extracting features, an initial weight allocation network, and an initial spectrum generation network;

[0113] Pre-training the initial acoustic model using the first sample corpus until a first generation convergence condition is met to obtain a reference acoustic model, wherein the first generation convergence condition indicates that a difference between the generated audio spectrum information and the corresponding label spectrum information is less than a first threshold:

[0114] The reference acoustic model is trained using the second sample corpus until a second generation convergence condition is obtained to obtain a target acoustic model, wherein the second generation convergence condition indicates that the difference between the generated audio spectrum information and the corresponding label spectrum information is less than a second threshold, the number of the second sample corpus is less than the number of the first sample corpus, and the target acoustic model includes a feature extraction network, a weight allocation network, and a target spectrum generation network that have completed training.

[0115] Optionally, the above-mentioned first training module is also used to: obtain the spectrum training results obtained each time during the training process of the reference acoustic model using the second sample corpus for training the reference acoustic model; when the spectrum training results indicate that the various model parameters in the reference acoustic model should be adjusted, update the network parameters in the weight allocation network and the spectrum generation network in a non-frozen state, and maintain the network parameters in the feature extraction network in a frozen state.

[0116] Optionally, the above-mentioned first training module is also used to: obtain the pronunciation representation vector of the sound source object after constructing the initial acoustic model; in the process of training the reference acoustic model, add the pronunciation representation vector of the sound source object to the second preprocessing network in the feature extraction network being trained, the residual connection network in the feature extraction network being trained, the gated loop structure in the weight allocation network being trained, and the multi-layer long short-term memory network of the spectrum generation network being trained.

[0117] Optionally, the above-mentioned first training module is also used to: when the text sequence is a Chinese text sequence, after constructing the initial acoustic model, obtain the tone features of the Chinese text sequence; in the process of training the reference acoustic model, add the tone features to the network structure after the second preprocessing network in the feature extraction network under training.

[0118] Optionally, the above-mentioned synthesis module 1110 is also used to input the audio spectrum information into the audio adversarial generation network to obtain the target audio, wherein the audio adversarial generation network includes a generation subnetwork for generating audio and a discrimination subnetwork for discrimination, and the discrimination subnetwork includes: a phase discrimination subnetwork for discriminating the phase information in the audio spectrum information, and a period discrimination subnetwork for discriminating the periodic information in the audio spectrum information.

[0119] Optionally, the audio generation device further includes a second training module, configured to:

[0120] Using positive sample audio pairs and negative sample audio pairs to perform cross-adversarial training on the initial audio adversarial generation network until convergence conditions are reached, wherein the positive sample audio pairs include the audio to be discriminated and the labeled audio, and the negative sample audio pairs include the reference audio and the labeled audio generated in the generation subnetwork based on the audio spectrum information of the audio to be discriminated;

[0121] During the training process, the discriminant subnetwork is trained using the positive sample audio pairs until the trained discriminant subnetwork meets a first discrimination condition, wherein the first discrimination condition indicates that the discriminant subnetwork recognizes that the audio to be discriminated is the labeled audio with a first confidence greater than a third threshold;

[0122] Save the network parameters of the discriminant sub-network;

[0123] An initial generative subnetwork is trained using negative sample audio pairs until the discriminative subnetwork reaches a second discrimination condition, wherein the second discrimination condition indicates that the discriminative subnetwork identifies the reference audio as the labeled audio with a second confidence greater than a fourth threshold, wherein network parameters of the generative subnetwork in training are adjusted based on the Fourier transform loss between the reference audio and the labeled audio, and the Mel spectrum residual loss between the reference audio and the labeled audio.

[0124] In an embodiment of the present application, a text sequence is converted into an abstract feature sequence, and the abstract feature sequence is converted into a context vector using a weight distribution network with a monotonicity constraint. Then, audio spectrum information matching the context vector is obtained, and the target audio is generated using the audio spectrum information. A context vector with a monotonicity constraint is generated through a weight distribution network, and the context vector with a monotonicity constraint is applied to the process of spectrum acquisition and audio generation. This avoids problems such as omissions, repetitions, and directional errors during the audio synthesis process, and achieves the purpose of ensuring the monotonic stability of the audio synthesis in direction, thereby achieving the technical effect of ensuring the accuracy and stability of the audio synthesis model during the audio synthesis process, and thus solving the technical problem of low accuracy of synthesized audio caused by the instability of the audio synthesis model.

[0125] According to another aspect of the embodiment of the present invention, an electronic device for implementing the above audio synthesis method is also provided. The electronic device may be Figure 1 The terminal device or server shown in FIG. This embodiment is described by taking the electronic device as a server as an example. Figure 12 As shown, the electronic device includes a memory 1202 and a processor 1204. The memory 1202 stores a computer program, and the processor 1204 is configured to execute the steps in any of the above method embodiments through the computer program.

[0126] Optionally, in this embodiment, the electronic device may be located in at least one network device among a plurality of network devices of a computer network.

[0127] Optionally, in this embodiment, the processor may be configured to execute the following steps through a computer program:

[0128] S1, obtain the text sequence to be processed;

[0129] S2, converts the text sequence into an abstract feature sequence;

[0130] S3, inputting the abstract feature sequence into a weight distribution network with monotonicity constraints to obtain the context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a discretized hybrid logic distribution structure;

[0131] S4, obtaining audio spectrum information matching the context vector;

[0132] S5, uses audio spectrum information to synthesize target audio that matches the text sequence.

[0133] Alternatively, those skilled in the art will appreciate that Figure 12 The structure shown is for illustration only, and the electronic device may also be a smart phone (such as an Android phone, an iOS phone, etc.), a tablet computer, a PDA, a mobile Internet device (MID), a PAD, or other terminal devices. Figure 12 It does not limit the structure of the electronic device. For example, the electronic device may also include Figure 12 More or fewer components (such as network interfaces, etc.) as shown in, or with Figure 12 Different configurations shown.

[0134] Among them, the memory 1202 can be used to store software programs and modules, such as the program instructions / modules corresponding to the audio synthesis method and device in the embodiment of the present invention. The processor 1204 executes various functional applications and data processing by running the software programs and modules stored in the memory 1202, that is, realizing the above-mentioned audio synthesis method. The memory 1202 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory. In some instances, the memory 1202 may further include a memory remotely located relative to the processor 1204, and these remote memories may be connected to the terminal via a network. Examples of the above-mentioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof. Among them, the memory 1202 can be used specifically, but not limited to, to store information such as text sequences and target audio. As an example, if Figure 12 As shown, the memory 1202 may include, but is not limited to, the first acquisition module 1102, conversion module 1104, input module 1106, second or dark area module 1108, and synthesis module 1110 of the audio synthesis device. Furthermore, the memory 1202 may include, but is not limited to, other modules and units of the audio synthesis device, which will not be described in detail in this example.

[0135] Optionally, the transmission device 1206 is configured to receive or send data via a network. Specific examples of the network may include a wired network and a wireless network. In one embodiment, the transmission device 1206 includes a network interface controller (NIC), which can be connected to other network devices and a router via a network cable to communicate with the Internet or a local area network. In one embodiment, the transmission device 1206 is a radio frequency (RF) module, which is configured to communicate with the Internet wirelessly.

[0136] In addition, the electronic device further includes: a display 1208 for displaying the text sequence; and a connection bus 1210 for connecting various module components in the electronic device.

[0137] In other embodiments, the terminal device or server may be a node in a distributed system, wherein the distributed system may be a blockchain system, and the blockchain system may be a distributed system formed by connecting multiple nodes through network communication. The nodes may form a peer-to-peer (P2P) network, and any computing device, such as a server, terminal, or other electronic device, may become a node in the blockchain system by joining the peer-to-peer network.

[0138] According to one aspect of the present application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations of the aforementioned audio synthesis aspects. The computer program is configured to, when executed, perform the steps of any of the aforementioned method embodiments.

[0139] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:

[0140] S1, obtain the text sequence to be processed;

[0141] S2, converts the text sequence into an abstract feature sequence;

[0142] S3, inputting the abstract feature sequence into a weight distribution network with monotonicity constraints to obtain the context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a discretized hybrid logic distribution structure;

[0143] S4, obtaining audio spectrum information matching the context vector;

[0144] S5, uses audio spectrum information to synthesize target audio that matches the text sequence.

[0145] Optionally, in this embodiment, a person of ordinary skill in the art may understand that all or part of the steps in the various methods of the above embodiments may be completed by instructing the hardware related to the terminal device through a program, and the program may be stored in a computer-readable storage medium, which may include: a flash drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.

[0146] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0147] If the integrated units in the above embodiments are implemented in the form of software functional units and sold or used as independent products, they can be stored in the above-mentioned computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes a number of instructions for causing one or more computer devices (such as personal computers, servers, or network devices) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0148] In the above embodiments of the present invention, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0149] In the several embodiments provided in this application, it should be understood that the disclosed client can be implemented in other ways. Among them, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of units or modules, and can be electrical or other forms.

[0150] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0151] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0152] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principles of the present invention. These improvements and modifications should also be regarded as within the scope of protection of the present invention.

Claims

1. An audio synthesis method, characterized in that: include: Obtain the text sequence to be processed and the pronunciation representation vector of the target sound source object; Converting the text sequence into a phoneme sequence; Inputting the phoneme sequence into a content encoder to convert the phoneme sequence into an abstract feature sequence, wherein the content encoder includes a first preprocessing network and a residual connection network, the first preprocessing network is used to preprocess the phoneme sequence, and the residual connection network is used to generate the abstract feature sequence based on the preprocessed phoneme sequence and the pronunciation representation vector; Inputting the abstract feature sequence into a weight distribution network connected to the content encoder and having a monotonicity constraint to obtain a context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a discretized hybrid logic distribution structure; Inputting the context vector into a target spectrum generation network connected to the weight allocation network to generate audio spectrum information matching the context vector, wherein the target spectrum generation network is a Mel spectrum residual network including a second preprocessing network and a multi-layer long short-term memory network, wherein the second preprocessing network is used to preprocess the context vector, and the multi-layer long short-term memory network is used to generate the audio spectrum information based on the preprocessed context vector and the pronunciation representation vector; The audio spectrum information is used to synthesize target audio that matches the text sequence.

2. The method according to claim 1, characterized in that Inputting the context vector into a target spectrum generation network connected to the weight allocation network to generate audio spectrum information matching the context vector includes: In a target spectrum generation network connected to the weight allocation network, frame spectrum information of one or at least two audio frames matching the context vector is obtained, wherein the target spectrum generation network is configured with a timer, and the timer is used to segment the audio spectrum information generated in the target spectrum generation network to generate the frame spectrum information corresponding to each audio frame.

3. The method according to claim 2, characterized in that In a target spectrum generation network connected to the weight allocation network, obtaining frame spectrum information of one or at least two audio frames matching the context vector includes: The following operations are performed in sequence in the target spectrum generation network to generate the frame spectrum information: Obtaining the context vector currently received from the weight distribution network and reference frame spectrum information of a previous audio frame before the current audio frame to be generated; The context vector and the reference frame spectrum information are input into the second preprocessing network and the multi-layer long short-term memory network to generate current frame spectrum information of the current audio frame.

4. The method according to claim 1, wherein Before obtaining the text sequence to be processed, the method further includes: Constructing an initial acoustic model, wherein the initial acoustic model includes: a feature extraction network for extracting features, an initial weight allocation network, and an initial spectrum generation network; Pre-training the initial acoustic model using the first sample corpus until a first generation convergence condition is met to obtain a reference acoustic model, wherein the first generation convergence condition indicates that a difference between the generated audio spectrum information and the corresponding label spectrum information is less than a first threshold: The reference acoustic model is trained using a second sample corpus until a second generation convergence condition is obtained to obtain a target acoustic model, wherein the second generation convergence condition indicates that the difference between the generated audio spectrum information and the corresponding label spectrum information is less than a second threshold, the number of the second sample corpus is less than the number of the first sample corpus, and the target acoustic model includes a trained feature extraction network, the weight allocation network, and the target spectrum generation network.

5. The method according to claim 4, characterized in that The process of training the reference acoustic model using the second sample corpus further includes: Obtaining a spectrum training result obtained each time during the training process of the reference acoustic model; When the spectrum training result indicates that the various model parameters in the reference acoustic model should be adjusted, the network parameters in the weight allocation network and the spectrum generation network in a non-frozen state are updated, and the network parameters in the feature extraction network in a frozen state are maintained.

6. The method according to claim 4, characterized in that After constructing the initial acoustic model, the method further includes: Get the pronunciation representation vector of the sound source object; During the training of the reference acoustic model, the pronunciation representation vector of the sound source object is added to the first preprocessing network in the feature extraction network being trained, the residual connection network in the feature extraction network being trained, the gated loop structure in the weight allocation network being trained, and the multi-layer long short-term memory network of the spectrum generation network being trained.

7. The method according to claim 4, characterized in that In the case where the text sequence is a Chinese text sequence, after constructing the initial acoustic model, the method further includes: Obtaining the tone features of the Chinese text sequence; During the training of the reference acoustic model, the pitch feature is added to the network structure after the first preprocessing network in the feature extraction network under training.

8. The method according to claim 1, characterized in that The synthesizing target audio matching the text sequence by using the audio spectrum information includes: The audio spectrum information is input into an audio adversarial generation network to obtain the target audio, wherein the audio adversarial generation network includes a generation subnetwork for generating audio and a discrimination subnetwork for discrimination, and the discrimination subnetwork includes: a phase discrimination subnetwork for discriminating phase information in the audio spectrum information, and a period discrimination subnetwork for discriminating periodic information in the audio spectrum information.

9. The method according to claim 8, characterized in that Before obtaining the text sequence to be processed, the method further includes: Performing cross-adversarial training on the initial audio adversarial generation network using positive sample audio pairs and negative sample audio pairs until convergence conditions are reached, wherein the positive sample audio pairs include the audio to be discriminated and the labeled audio, and the negative sample audio pairs include the reference audio generated in the generation subnetwork based on the audio spectrum information of the audio to be discriminated and the labeled audio; During the training process, the discriminant subnetwork is trained using the positive sample audio pair until the trained discriminant subnetwork meets a first discrimination condition, wherein the first discrimination condition indicates that the discriminant subnetwork recognizes that the audio to be discriminated is the labeled audio with a first confidence level greater than a third threshold; Saving the network parameters of the discriminant subnetwork; The generative subnetwork is trained using the negative sample audio pair until the discriminative subnetwork reaches a second discrimination condition, wherein the second discrimination condition indicates that the second confidence level of the discriminative subnetwork in identifying the reference audio as the label audio is greater than a fourth threshold, and the network parameters of the generative subnetwork in training are adjusted according to the Fourier transform loss between the reference audio and the label audio, and the Mel spectrum residual loss between the reference audio and the label audio.

10. An audio synthesis device, characterized in that: include: A first acquisition module is used to acquire a text sequence to be processed; The device is further configured to obtain a pronunciation representation vector of a target sound source object and convert the text sequence into a phoneme sequence; a conversion module, configured to input the phoneme sequence into a content encoder to convert the phoneme sequence into an abstract feature sequence, wherein the content encoder includes a first preprocessing network and a residual connection network, the first preprocessing network being configured to preprocess the phoneme sequence, and the residual connection network being configured to generate the abstract feature sequence based on the preprocessed phoneme sequence and the pronunciation representation vector; An input module, configured to input the abstract feature sequence into a weight distribution network connected to the content encoder and having a monotonicity constraint, to obtain a context vector corresponding to the abstract feature sequence, wherein the weight distribution network is a network constructed based on a discretized hybrid logic distribution structure; a second acquisition module, configured to input the context vector into a target spectrum generation network connected to the weight allocation network to generate audio spectrum information matching the context vector, wherein the target spectrum generation network is a Mel spectrum residual network including a second preprocessing network and a multi-layer long short-term memory network, wherein the second preprocessing network is configured to preprocess the context vector, and the multi-layer long short-term memory network is configured to generate the audio spectrum information based on the preprocessed context vector and the pronunciation representation vector; A synthesis module is used to synthesize target audio that matches the text sequence using the audio spectrum information.

11. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein the method according to any one of claims 1 to 9 is executed when the program is executed.

12. An electronic device comprising a memory and a processor, characterized in that: A computer program is stored in the memory, and the processor is configured to execute the method according to any one of claims 1 to 9 through the computer program.

Citation Information

Patent Citations

  • Multi-speaker voice separation method based on voiceprint features and generative adversarial learning

    CN111128197A

  • Speech synthesis model training method and device, equipment and storage medium

    CN112365876A

  • Audio signal generation method, device and equipment and storage medium

    CN112712812A

  • Synthesizing speech from text using neural networks

    US20200051583A1