A Speech Synthesis Method of Vocoder Based on Semi-Flow Model

By adopting a vocoder based on a semi-stream model in the speech synthesis technology, combining autoregressive streaming and standardized streaming algorithms, the problem of high quality and real-time in the existing technology is solved, efficient and fast speech synthesis is achieved, and computing resource requirements are reduced.

CN114464159BActive Publication Date: 2025-05-30TONGJI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210054963.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-01-18
Publication Date
2025-05-30
Estimated Expiration
2042-01-18

AI Technical Summary

Technical Problem

The existing speech synthesis technology is difficult to balance between high quality and real-time, especially under high sampling rate conditions, the computing resources and synthesis speed are limited, and the training of vocoders based on neural networks is difficult and the convergence speed is slow.

Method used

A vocoder based on a semi-stream model is adopted, combining autoregressive streams and normalized stream algorithms to build a deep neural network model of multi-layer Flow layer and Scale layer. The audio data is converted into Mel spectrum through a preprocessing module, and the likelihood loss function is used for training.

Benefits of technology

It improves the quality and speed of speech synthesis, shortens the convergence time of training, and reduces the demand for computing resources, meeting the requirements of practical applications for efficient speech synthesis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114464159B_ABST
    Figure CN114464159B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for synthesizing speech of a vocoder based on a semi-flow model, including: obtaining original audio data to be synthesized and loading it into a pre-constructed and trained vocoder based on a semi-flow model to obtain a synthesized speech waveform; the vocoder based on the semi-flow model includes a basic model based on a semi-flow, and the basic model based on the semi-flow includes a plurality of sequentially spliced Flow layers, each Flow layer includes a semi-flow model layer and a convolutional network layer connected in sequence, and the semi-flow model layer is composed of a combination of an autoregressive flow algorithm and a normalizing flow algorithm. Compared with the prior art, the present invention can improve the quality of the synthesized speech to a certain extent, while accelerating the speed of the synthesized speech and the convergence speed during training, and reducing certain computing resources.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of speech synthesis, and in particular to a vocoder speech synthesis method based on a semi-flow model. Background Art

[0002] With the increasing frequency of human-machine voice interaction, how to efficiently synthesize high-quality speech has attracted more and more attention. Minor changes in speech quality or latency have a great impact on the user experience. However, high-quality real-time speech synthesis remains a challenging task. Speech synthesis requires generating high-dimensional audio samples with high long-term dependencies. Humans are very sensitive to such dependencies in audio samples. In addition to the quality challenge, real-time speech synthesis also faces many problems such as limited generation speed and computational resources. When the audio sampling rate is less than 16 kHz, the perceived speech quality will decrease significantly, and higher sampling rates will produce higher-quality speech. However, in most cases, users require audio synthesized at a much faster rate than 16 kHz. For example, when synthesizing speech on a remote server, strict interactivity requirements mean that speech must be synthesized quickly at a sampling rate far exceeding real-time requirements.

[0003] Currently, the most advanced speech synthesis models are based on neural networks. Text-to-speech synthesis is usually divided into two steps: the first step is to convert text into time-aligned features, such as mel spectrograms, F0 features, or other linguistic features. The second step is to convert these time-aligned features into audio samples. The neural network model used in the second step is usually called a vocoder, which is computationally challenging and also has a great impact on the quality of the synthesized speech. Most current neural network-based vocoders are autoregressive, which means they place future audio samples on top of previous samples to build a long-term correlation model. The implementation and training of these methods are relatively simple. However, they are inherently serial and thus cannot fully utilize parallel processors such as GPUs or TPUs. Such autoregressive models usually have difficulty synthesizing speech at a speed exceeding 16 kHz without sacrificing the quality of the synthesized audio.

[0004] Therefore, relevant alternative technologies have been developed. Currently, there are three neural network-based models that can synthesize speech in a non-autoregressive manner: Parallel WaveNet, Clarinet, and MCNN for spectrogram inversion. These technologies can synthesize audio at speeds exceeding 500 kHz on GPUs. However, these models are more difficult to train and implement compared to autoregressive models. At the same time, all three methods require composite loss functions to improve audio quality or solve the mode collapse problem. In addition, Parallel WaveNet and Clarinet require two networks: a student network and a teacher network. Their student networks use inverse autoregressive flows. Although inverse autoregressive flow networks can run in parallel during inference, their autoregressive nature makes the model computationally inefficient. To overcome this shortcoming, these networks use the teacher network to train the student network to make the synthesized speech highly realistic. However, these methods are difficult to replicate and deploy because they are difficult to converge during training.

[0005] In subsequent research, people gradually adopted flow-based models to build vocoders. Flow-based models were proposed in RealNVP and Glow and can be used in generative tasks such as image generation and speech synthesis. WaveGlow was the first to apply flow-based models to the speech synthesis task. It is easy to implement and train, and is trained using only a single network and a likelihood loss function. In addition, this model can synthesize speech at a frequency exceeding 500 kHz on an NVIDIA V100 GPU without loss of audio quality. However, this model has a large number of parameters, so it requires a large amount of computing resources, and at the same time, it converges slowly during training and requires a lot of time to converge. Summary of the Invention

[0006] The purpose of the present invention is to overcome the defects of the above-mentioned existing technologies and provide a speech synthesis method for a vocoder based on a semi-flow model, which solves the shortcomings of insufficient computing power of traditional flow models, slow convergence speed, large number of model parameters, slow synthesis speed, and poor generation quality of traditional flow-based vocoders, and meets the requirements of actual speech synthesis applications for neural vocoders.

[0007] The purpose of the present invention can be achieved by the following technical solutions:

[0008] A speech synthesis method for a vocoder based on a semi-flow model, comprising: obtaining the original audio data to be synthesized, and loading it into a pre-constructed and trained vocoder based on a semi-flow model to obtain the synthesized speech waveform;

[0009] The vocoder based on the semi-flow model includes a basic semi-flow model, and the basic semi-flow model includes a plurality of sequentially connected Flow layers. Each Flow layer includes a semi-flow model layer and a convolutional network layer connected in sequence. The semi-flow model layer is composed of a combination of an autoregressive flow algorithm and a normalizing flow algorithm.

[0010] Further, the mapping relationship between the high-dimensional input vector x and the high-dimensional output vector y in the semi-flow model layer is as follows:

[0011] x = (x 1 , x 2 ), y 0 = 0

[0012] (s 1 , t 1 ) = g(m(x 1 , y 0 ))

[0013] y 1 = s 1 ⊙x 1 + t 1

[0014] (s 2 , t 2 ) = g(m(x 2 , y 1 ))

[0015] y 2 = s 2 ⊙x 2 + t 2

[0016] y = (y 1 , y 2 )

[0017] In the formula, x 1 and x 2 represent the first and second halves of x, y 0 is the constant vector 0, g and m are functions or neural networks, m and g can be any transformation, s 1 , s 2 , u 1 , u 2 are affine factors, ⊙ represents the Hadamard product, y 1 and y 2 represent the first and second halves of y.

[0018] Further, four of the Flow layers form a Scale layer. The basic semi-flow model includes a plurality of Scale layers. The Scale layer directly selects a vector of half the dimension as the output and inputs the other half into the next Scale layer.

[0019] Further, the number of the Flow layers is 12, and the convolutional network layer is a 1×1 convolutional network.

[0020] Further, the training process of the vocoder based on the semi-flow model includes:

[0021] A preprocessing module is set before the basic model based on the semi-flow, and this preprocessing module is used to convert the input audio data into Mel spectrograms;

[0022] Obtain a training set and a test set, load the training set into the basic model based on the semi-flow, convert it into Mel spectrograms through the preprocessing module, and then synthesize a speech waveform through the basic model based on the semi-flow, so as to perform model training;

[0023] Invert the trained basic model based on the semi-flow, convert the data in the test set into Mel spectrograms, and then load them into the inverted basic model based on the semi-flow to restore them into speech waveforms, so as to evaluate the quality of the synthesized speech and determine whether the basic model based on the semi-flow is trained.

[0024] Further, the preprocessing module includes a Fourier transform sub-module, and this Fourier transform sub-module uses the short-time Fourier transform to convert the audio data into Mel spectrograms.

[0025] Further, the preprocessing module further includes a pre-emphasis sub-module, and this pre-emphasis sub-module is used to enhance the energy of the high-frequency part of the audio, and the output end of the pre-emphasis sub-module is connected to the Fourier transform sub-module;

[0026] The processing expression of the pre-emphasis sub-module is:

[0027] y(n) = x(n) - αy(n - 1)

[0028] In the formula, x(n) is the nth sampling point of the original audio, y(n) is the nth sampling point of the pre-emphasized audio, α is the pre-emphasis coefficient, and the value of α is between 0.9 and 1.0.

[0029] Further, the loss function in the model training process is:

[0030]

[0031] In the formula, x is the input data during model training, z(x) is the function from x to z during model training, σ 2 is the assumed variance of the Gaussian distribution, #coupling is the number of semi-flow layers included in the model, s j1 is the first affine factor in the jth semi-flow layer, s j2is the second affine factor in the semi-flow of the j-th layer, #conv is the number of 1×1 convolutional networks included in the model, and W k is the weight matrix of the 1×1 convolutional network of the k-th layer.

[0032] Furthermore, the evaluation metrics for the synthetic speech quality include one or more of PESQ, MOS, STOI, and MCD.

[0033] Furthermore, the data in the training set and the test set are both obtained from a speech synthesis dataset, which includes one or more of LibriSpeech, AiShell-3, CSMSC, and LJSpeech.

[0034] Compared with the prior art, the present invention has the following advantages:

[0035] (1) The present invention proposes a semi-flow model that combines the advantages of normalizing flow and autoregressive flow. In the semi-flow, the output of the second half is associated with both the output and the input of the first half, and at the same time, the output of the first half is obtained through an affine transformation of the input of the first half, thereby improving the computational performance of the model; a deep neural network model based on the semi-flow is used to model the voice characteristics of the speaker, so as to restore the corresponding Mel spectrogram to a speech waveform approximating the real human voice. This method can improve the quality of the synthetic speech to a certain extent, while accelerating the speed of the synthetic speech and the convergence speed during training, and reducing a certain amount of computing resources.

[0036] (2) In the basic model of the present invention based on the semi-flow, four Flow layers form a Scale layer, including multiple Scale layers. The Scale layer directly selects a vector of half the dimension as the output, and the other half is input to the next Scale layer; the multi-scale architecture can extract relevant features early and improve the computational efficiency of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 is the algorithm flowchart of a vocoder speech synthesis method based on a semi-flow model provided in an embodiment of the present invention;

[0038] Figure 2 is the model architecture diagram of a vocoder speech synthesis method based on a semi-flow model provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0039] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are some, but not all, of the embodiments of the present invention. Components of the embodiments of the present invention described and illustrated herein generally may be arranged and designed in a variety of different configurations.

[0040] Therefore, the detailed description of the embodiments of the present invention provided in the accompanying drawings is not intended to limit the scope of the claimed invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0041] It should be noted that like reference numerals and letters denote like items in the following drawings. Therefore, once an item is defined in one drawing, it does not need to be further defined and explained in subsequent drawings.

[0042] Embodiment 1

[0043] This embodiment provides a vocoder speech synthesis method based on a semi-flow model, including: obtaining original audio data to be synthesized, and loading it into a pre-constructed and trained vocoder based on a semi-flow model to obtain a synthesized speech waveform;

[0044] The vocoder based on the semi-flow model includes a basic model based on a semi-flow. The basic model based on a semi-flow includes a plurality of Flow layers connected in series. Each Flow layer includes a semi-flow model layer and a convolutional network layer connected in sequence. The semi-flow model layer is composed of a combination of an autoregressive flow algorithm and a normalizing flow algorithm.

[0045] The training process of the vocoder based on the semi-flow model includes:

[0046] A preprocessing module is set before the basic model based on a semi-flow. The preprocessing module is used to convert the input audio data into a Mel spectrogram;

[0047] Obtain a training set and a test set, load the training set into the basic model based on a semi-flow, convert it into a Mel spectrogram through the preprocessing module, and then synthesize a speech waveform through the basic model based on a semi-flow, thereby performing model training;

[0048] Invert the trained basic model based on a semi-flow, convert the data in the test set into a Mel spectrogram, and then load it into the inverted basic model based on a semi-flow to restore it to a speech waveform, thereby evaluating the quality of the synthesized speech and determining whether the basic model based on a semi-flow is trained.

[0049] The model construction, training, and testing processes in this embodiment will be specifically described below.

[0050] (1) Combine the autoregressive flow algorithm F AR and the normalizing flow algorithm F Norm to obtain the semi-flow model F Semi , making it have both the high computational performance of the autoregressive flow and the simplicity of the normalizing flow:

[0051] (1-1) In the autoregressive flow algorithm F AR , the high-dimensional input vector x is transformed through autoregressive transformation to obtain the high-dimensional output vector y, and the mapping relationship between the two is as follows:

[0052] x = (x 1 , x 2 , x 3 ... x i ...)

[0053] (s i , u i ) = g(x 1:i-1 )

[0054] y i = s i x i + u i

[0055] y = (y 1 , y 2 , y 3 ... y i ...)

[0056] where x i and y i respectively represent the i-th elements of x and y, and g can be any function or neural network for calculating the two affine factors s i and u i . It is not difficult to see that in the autoregressive flow, the i-th output element is related to the previous i - 1 input elements. Similarly, if the i-th output element is related to the previous i - 1 output elements, the inverse autoregressive flow algorithm F IAR can be obtained. At this time, the calculation method of the affine factors changes to (s i , u i ) = g(y 1:i-1 );

[0057] (1-2) In the normalizing flow algorithm F Norm , the mapping relationship between the high-dimensional input vector x and the high-dimensional output vector y is as follows:

[0058] x = (x 1 , x 2)

[0059] y 1 = x 1

[0060] (s, u) = g(x 1 )

[0061] y 2 = s ⊙ x 2 + u

[0062] y = (y 1 , y 2 )

[0063] where x 1 and x 2 represent the first and second halves of x, and y 1 and y 2 represent the first and second halves of y. g can be any function or neural network for calculating the two affine factors s and u, and ⊙ represents the Hadamard product. The first half of the input is directly used as the output, and the second half of the input is transformed through an affine transformation to obtain the other part of the output. This structure is also called an affine coupling layer;

[0064] (1 - 3) obtains the semi - flow algorithm F by combining (inverse) autoregressive flow and normalizing flow algorithms Semi , and in F Semi the mapping relationship between the high - dimensional input vector x and the high - dimensional output vector y is as follows:

[0065] x = (x 1 , x 2 ), y 0 = 0

[0066] (s 1 , t 1 ) = g(m(x 1 , y 0 ))

[0067] y 1 = s 1 ⊙ x 1 + t 1

[0068] (s 2 , t 2 ) = g(m(x 2 , y 1 ))

[0069] y 2 = s 2 ⊙ x 2 + t 2

[0070] y = (y1 , y 2 )

[0071] where x 1 and x 2 represent the first and second halves of x, and y 0 is the constant vector 0, m and g can be any transformation, and s 1 , s 2 , u 1 , u 2 is the affine factor, ⊙ represents the Hadamard product, and y 1 and y 2 represent the first and second halves of y. In the semi-flow, the output of the second half is related to both the output and input of the first half, and at the same time, the output of the first half is obtained through an affine transformation of the input of the first half, thus improving the computational performance of the model.

[0072] (2) The semi-flow algorithm can be used as a separate network layer in a neural network. By combining it with a 1×1 convolutional network layer, the basic model of the semi-flow-based vocoder can be obtained:

[0073] (2-1) In the semi-flow-based vocoder, to improve computational efficiency, m is defined as a simple addition transformation, and g is defined as a neural network similar to WaveNet, with 8 hidden layers, a channel size of 128, and a convolutional kernel size of 3. Its calculation formula is as follows:

[0074] z = tanh(W f,k * x) ⊙ σ(W g,k * x)

[0075] where x and z represent the input and output of this network layer respectively, * represents the convolutional operation, ⊙ represents the Hadamard product, σ represents the sigmoid function, k is the layer index, f and g represent filters and gates, and W is a learnable convolutional filter. The affine factor in the semi-flow is obtained through this formula;

[0076] (2-2) The basic model of the semi-flow-based vocoder consists of 12 Flow layers. Each Flow layer contains a semi-flow algorithm layer and a 1×1 convolutional network layer, and the convolutional network layer comes after the semi-flow algorithm layer. The convolutional network layer is used to shuffle the channel order of the intermediate process vector;

[0077] (2-3) Four Flow layers form a Scale layer as a group. The Flow layers in the same Scale layer have the same structure, and different Scale layers are combined in a multi-scale architecture. The multi-scale architecture can extract relevant features early and improve the computational efficiency of the model;

[0078] Table 1

[0079]

[0080] (2 - 4) Totally includes three Scale layers. In the first Scale layer, the dimensions of the input and the vectors in the intermediate process are 12. After passing through each Scale layer, half of the dimensional vectors are directly selected as the output, and the other half are input into the next Scale layer. That is, the vector dimension in the first Scale layer is 12, the vector dimension in the second Scale layer is 6, and the vector dimension in the third Scale layer is 4, as shown in Table 1.

[0081] (3) By adding a pre - processing module in front of the basic model of the semi - flow - based vocoder, a semi - flow - based vocoder can be obtained. The pre - processing module consists of two parts: pre - emphasis and Fourier transform:

[0082] (3 - 1) After the training audio is input into the semi - flow - based vocoder, it will first pass through the pre - emphasis module. In this module, the energy of the high - frequency part of the audio will be enhanced, and a difference equation is used for processing:

[0083] y(n) = x(n) - αy(n - 1)

[0084] Where x(n) represents the n - th sampling point of the original audio, y(n) represents the n - th sampling point of the pre - emphasized audio, α is the pre - emphasis coefficient, which can take values between 0.9 and 1.0, and the preferred value is 0.95. The pre - emphasis module can improve the quality of the audio synthesized by the model;

[0085] (3 - 2) After pre - emphasis, the audio will first be converted into a spectrogram through Fourier transform with a window size of 1024, a frame shift of 256, and 1024 filter banks. Then, these spectrograms are multiplied by 80 Mel filters to obtain the Mel spectrogram. The Mel spectrogram is a spectrogram under the Mel scale, and the conversion formula between the Mel scale and Hertz is:

[0086]

[0087] (4) The pre - processing module and the basic model of the semi - flow - based vocoder together constitute the semi - flow - based vocoder. The pre - processing module is used during the training of the model, and when generating audio, the pre - processing module is not used, and the trained basic model of the semi - flow - based vocoder is directly used to generate audio:

[0088] (4-1) When training a semi-flow based vocoder, it is first necessary to process the existing dataset. Select the CSMSC Chinese standard female voice speech database as the basic database for training, and form 45 groups of small sample datasets from it. Each group of datasets contains a training set and a test set. Each training set contains 50 audio data randomly selected from CSMSC, with a total duration of about 5 minutes. Each test set contains 5000 audio data randomly selected from CSMSC. The audio contained in the training sets of different small samples does not repeat, and the audio within each training set only appears in that training set;

[0089] (4-2) When training a semi-flow based vocoder, 45 groups of small sample datasets are needed to train 45 groups of models. For each training, the batch size is set to 6 during training, and the number of iterations is 3000. The initial learning rate is set to 4e -4 , and then an adaptive learning rate adjustment strategy is adopted. After every 1000 iterations, the learning rate is reduced to one-fourth of the original;

[0090] (4-3) During training, the semi-flow based vocoder will convert the audio in the training dataset into Mel spectrograms. The initial sampling rate of the input audio is 22050Hz. After inputting into the model, each audio will be truncated into a vector of a fixed length. The segment length can take any value not exceeding the audio length, and the preferred value is 16384. Next, the audio will be input into the preprocessing module to obtain the preprocessed input vector, and then this vector will be input into the neural network model;

[0091] (4-4) During training, the relationship between the preprocessed input vector x ′ and the output vector y in terms of the likelihood function is:

[0092]

[0093] where p θ represents the probability density, J represents the Jacobian determinant, and f i represents the i-th layer network in the model. The neural network is trained by finding the maximum likelihood or minimizing the negative log-likelihood.

[0094] During training, it is assumed that y follows a zero-mean spherical Gaussian distribution, that is

[0095]

[0096] Therefore, the probability density of y is

[0097] For the semi-flow layer, its Jacobian determinant s 1 is related to the absolute value of s 2 , and it is as follows:

[0098]

[0099] For the 1×1 convolutional network layer, its calculation formula is:

[0100]

[0101] Where W represents the weight matrix. Therefore, the Jacobian determinant of the 1×1 convolutional network layer is only related to W, as shown below:

[0102]

[0103] In summary, the likelihood function of the semi-flow based vocoder is:

[0104]

[0105] In the formula, x is the input data during model training, z(x) is the function from x to z during model training, σ 2 is the assumed variance of the Gaussian distribution, #coupling is the number of semi-flow layers included in the model, s j1 is the first affine factor in the j-th semi-flow layer, s j2 is the second affine factor in the j-th semi-flow layer, # conv is the number of 1×1 convolutional networks included in the model, W k is the weight matrix of the k-th 1×1 convolutional network.

[0106] This function can be used as the loss function during training;

[0107] (4-5) When testing the semi-flow based vocoder, the mel spectrogram of the test audio is restored to the speech waveform, and the MOS value of the synthesized audio is measured to evaluate the quality.

[0108] (4-5-1) Convert the audio in the test set to the mel spectrogram using the method described in (3-2);

[0109] (4-5-2) Since each layer of the network in the semi-flow based vocoder is reversible, for each dataset in the 45 datasets, each trained model will be inverted during testing, and the mel spectrogram converted from the test set will be input, so as to restore it to the speech waveform;

[0110] The test set of 45 groups of small sample data sets is converted into Mel spectrograms with 80 dimensions using the short-time Fourier transform, with a sampling rate of 22050 Hz, a filter length of 1024, and a window size of 1024. Then, the trained model is used to restore the generated Mel spectrograms to waveforms for testing, and evaluation metrics are used to score the results. Optional evaluation metrics include PESQ, MOS, STOI, and MCD. MOS, i.e., the Mean Opinion Score, is preferred. MOS can be obtained by manual or neural network evaluation. Neural networks include MOSNet, MTL-MOSNet, etc. (4-5-3) The MOS value is the mean opinion score and is usually used to evaluate speech quality and is scored manually. MOSNet is a deep neural network that can automatically measure the MOS value and can solve the problem of consuming human and time resources in traditional MOS evaluation methods. The pre-trained MOSNet is used to evaluate the MOS value of the synthesized speech.

[0111] The present invention is further illustrated by specific experiments as follows:

[0112] Experimental conditions and scoring criteria: In this experiment, the Chinese Standard Mandarin Speech Corpus, a Chinese standard female voice library, is used. Table 2 lists the main information of this database. The measurement indicators mainly include audio quality, synthesis speed, and convergence speed. The audio quality is measured using the Mean Opinion Score (MOS), i.e., the mean opinion score, with a value range of 0-5, and the higher the score, the better the quality. The synthesis speed is measured in samples / s, i.e., the number of samples that can be synthesized per second. The convergence speed is measured by the number of iterations required to reach convergence. When the change rate between adjacent samples is less than a certain threshold, it means the model has converged.

[0113] Table 2 Main information of the database

[0114]

[0115] Experiment 1: Evaluate the quality of the synthesized audio. In this experiment, 45 semi-flow-based models are first trained using 45 groups of data sets, and then the corresponding audio is synthesized for the test set of each group of data sets. Next, MOSNet is used to evaluate the MOS value of each audio, and then a 95% confidence interval is used to show the audio quality. As a comparative experiment, autoregressive flow models and normalizing flow models are used as comparative models. In addition, the reference audio is also shown in the results. The experimental results are shown in Table 3. It can be seen that the semi-flow-based model has the highest MOS value.

[0116] Table 3 Audio quality evaluation

[0117] Model MOS Ground Truth 3.754±0.007 <![CDATA[F Norm > 3.324±0.001 <![CDATA[F AR > 2.785±0.001 <![CDATA[F Semi > 3.416±0.001

[0118] Experiment 2: Evaluation of audio synthesis speed. The first model among the 45 models trained in Experiment 1 was used for testing. 5000 test samples in the first dataset were synthesized, and the total required time was recorded. Then, the number of samples synthesized per second was calculated using "total audio duration × sampling rate / total synthesis duration". This experiment was conducted on a workstation with a single 2080Ti and a Raspberry Pi 4b respectively. As a comparative experiment, the autoregressive flow model and the normalizing flow model were used as comparative models. The experimental results are shown in Table 4. It can be seen that the models based on semi-flows have the fastest synthesis speed values on both devices.

[0119] Table 4 Audio Synthesis Speed Evaluation

[0120] Model Workstation Raspberry pi 4B <![CDATA[F Norm > 405k 4.4k <![CDATA[F AR > 139k failed <![CDATA[F Semi > 522k 5.1k

[0121] Experiment 3: Evaluation of the convergence speed of the model. The first model among the 45 models trained in Experiment 1 was used for testing. The change curve of the loss during training was recorded, and the change rate of the loss between adjacent points was calculated. When the change rate is less than the threshold, the model is considered to have converged. As a comparative experiment, the autoregressive flow model and the normalizing flow model were used as comparative models. The experimental results are shown in Table 5. It can be seen that the models based on semi-flows have the fastest convergence speed.

[0122] Table 5 Audio Synthesis Speed Evaluation

[0123] Model Step <![CDATA[F Norm > 7778 <![CDATA[F AR > 3826 <![CDATA[F Semi > 3700

[0124] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative work. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field of the present invention through logical analysis, reasoning, or limited experiments based on the concept of the present invention on the basis of the prior art shall fall within the protection scope determined by the claims.

Claims

1. A method for synthesizing speech of a vocoder based on a semi - flow model, characterized in that, it includes: Obtain the original audio data to be synthesized, and load it into a pre - constructed and trained vocoder based on a semi - flow model to obtain the synthesized speech waveform; The vocoder based on the semi - flow model includes a basic model based on a semi - flow. The basic model based on a semi - flow includes a plurality of Flow layers connected in series. Each Flow layer includes a semi - flow model layer and a convolutional network layer connected in sequence. The semi - flow model layer is composed of the combination of an autoregressive flow algorithm and a normalizing flow algorithm; Four of the said Flow layers constitute a Scale layer. The basic model based on a semi - flow includes a plurality of Scale layers. The Scale layer selects vectors of half of the dimensions directly as the output and inputs the other half into the next Scale layer; The mapping relationship between the high - dimensional input vector x and the high - dimensional output vector y in the semi - flow model layer is: x = (x 1 , x 2 ), y 0 = 0 (s 1 ,t 1 ) = g(m(x 1 ,y 0 )) y 1 = s 1 ⊙x 1 + t 1 (s 2 ,t 2 ) = g(m(x 2 ,y 1 )) y 2 = s 2 ⊙ x 2 + t 2 y = (y 1 , y 2 ) where x 1 and x 2 represent the first and second halves of x, y 0 is the constant vector 0, g and m are functions or neural networks, s 1 , s 2 , u 1 , u 2 are affine factors, ⊙ represents the Hadamard product, y 1 and y 2 represent the first and second halves of y.

2. The method for synthesizing speech of a vocoder based on a semi - flow model according to claim 1, characterized in that, The number of the Flow layers is 12, and the convolutional network layer is a 1×1 convolutional network.

3. The method for synthesizing speech of a vocoder based on a semi - flow model according to claim 1, characterized in that, The training process of the vocoder based on the semi - flow model includes: Set a pre - processing module in front of the basic model based on a semi - flow. This pre - processing module is used to convert the input audio data into Mel - spectrogram; Obtain a training set and a test set. Load the training set into the basic model based on a semi - flow, convert it into Mel - spectrogram through the pre - processing module, and then synthesize a speech waveform through the basic model based on a semi - flow, so as to carry out model training; Invert the trained basic model based on a semi - flow, convert the data in the test set into Mel - spectrogram, and then load it into the inverted basic model based on a semi - flow to restore it to a speech waveform, so as to evaluate the quality of the synthesized speech and determine whether the basic model based on a semi - flow is trained.

4. The method for synthesizing speech of a vocoder based on a semi - flow model according to claim 3, characterized in that, The pre - processing module includes a Fourier transform sub - module. This Fourier transform sub - module uses the short - time Fourier transform to convert the audio data into Mel - spectrogram.

5. The method for synthesizing speech of a vocoder based on a semi - flow model according to claim 4, characterized in that, The pre - processing module also includes a pre - emphasis sub - module. This pre - emphasis sub - module is used to emphasize the energy of the high - frequency part of the audio. The output end of the pre - emphasis sub - module is connected to the Fourier transform sub - module; The processing expression of the pre - emphasis sub - module is: y(n) = x(n)-αy(n - 1) In the formula, x(n) is the nth sampling point of the original audio, y(n) is the nth sampling point of the pre - emphasized audio, α is the pre - emphasis coefficient, and the value of α is between 0.9 and 1.

0.

6. The method for synthesizing speech of a vocoder based on a semi - flow model according to claim 5, characterized in that, The loss function during the model training process is: where \(x\) is the input data during model training, \(z(x)\) is the function from \(x\) to \(z\) during model training, \(\sigma\) 2 is the assumed variance of the Gaussian distribution, #coupling is the number of half - flow layers included in the model, \(s\) j1 is the first affine factor in the \(j\) - th half - flow, \(s\) j2 is the second affine factor in the \(j\) - th half - flow, #conv is the number of \(1\times1\) convolutional networks included in the model, \(W\) k is the weight matrix of the \(k\) - th \(1\times1\) convolutional network.

7. The method for synthesizing speech of a vocoder based on a semi - flow model according to claim 3, Characterized in that, The evaluation metrics for the quality of the synthesized speech include one or more of PESQ, MOS, STOI, and MCD.

8. A vocoder speech synthesis method based on a semi-stream model according to claim 3, Characterized in that, The data in the training set and the test set are both obtained from a speech synthesis dataset, and the speech synthesis dataset includes one or more of LibriSpeech, AiShell-3, CSMSC, and LJSpeech.