Phoneme Alignment Model Training Method, Computer Device, and Computer Storage Medium

Through the phoneme alignment model training method, the convolutional structure and phoneme sequence processing are used to generate accurate phoneme vectors, which solves the problem of poor training effect of singing vocal synthesis model caused by manual labeling errors, and achieves more efficient singing vocal synthesis model training.

CN115910032BActive Publication Date: 2025-06-10TENCENT MUSIC ENTERTAINMENT TECH (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211557817.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-06
Publication Date
2025-06-10
Estimated Expiration
2042-12-06

AI Technical Summary

Technical Problem

The existing singing synthesis model relies on the duration and location of phonemes manually marked during training, resulting in labeling errors and poor training results.

Method used

Through the phoneme alignment model training method, acoustic features are extracted using convolutional structures, combined with phoneme sequences and position sequences, internal product calculations and SoftMax processing are performed, accurate phoneme vectors are generated, and the initial acoustic model is input to train target acoustic feature parameters.

Benefits of technology

It has got rid of the limitations of manual annotation, improves the accuracy of phoneme alignment, and improves the training effect of the singing synthesis model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115910032B_ABST
    Figure CN115910032B_ABST
Patent Text Reader

Abstract

An embodiment of the present application discloses a method for training a phoneme alignment model, a computer device, and a computer storage medium. Acoustic feature parameters are input into a first convolutional structure to obtain first convolutional features. A phoneme sequence vector is generated according to the phoneme sequence of each phoneme. The phoneme sequence vectors of every three adjacent phonemes of the original audio are input into a second convolutional structure to obtain second convolutional features. The calculation result of the inner product of the first convolutional features and the second convolutional features is subjected to SoftMax calculation to obtain a weight vector. The phoneme sequence vectors of every three adjacent phonemes of the original audio are weighted according to the weight vector to obtain phoneme vectors. The conditional vector obtained by adding the phoneme vectors and the position sequence is input into an initial acoustic model, so that the initial acoustic model is trained according to the conditional vector to obtain a target acoustic model. The accuracy requirements for manually annotating the phoneme positions and durations are reduced, enabling the phonemes to more accurately correspond to the duration of the audio, thereby improving the training effect of the singing synthesis model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Embodiments of the present application relate to the field of speech synthesis, and more particularly to a method for training a phoneme alignment model, a computer device, and a computer storage medium. Background Art

[0002] In recent years, speech synthesis technology has made great progress, and the synthesized speech has approached the level of real human pronunciation in terms of sound quality and naturalness. Compared with speech synthesis technology, the progress of singing synthesis technology is relatively slow. Singing synthesis technology has many application scenarios, such as song adaptation, harmony generation, and virtual singers. Existing solutions mainly train a singing synthesis model and use this singing synthesis model to output synthesized singing. During the training process of the singing synthesis model, it is necessary to train according to the phonemes of the audio training samples, and for a singing synthesis model, the duration and position of the phonemes are crucial.

[0003] Existing solutions only manually mark the duration and position corresponding to the phonemes of the audio training samples, but manual marking is based on human subjective consciousness, and there may be cases of marking errors, resulting in inaccurate manual marking results, which in turn affects the training effect of the singing synthesis model. Summary of the Invention

[0004] Embodiments of the present application provide a method for training a phoneme alignment model, a computer device, and a computer storage medium for accurately aligning the duration and position of each phoneme of an audio.

[0005] In a first aspect of the embodiments of the present application, a method for training a phoneme alignment model is provided, and the method includes:

[0006] Obtain the acoustic feature parameters of the original audio, and obtain the phoneme sequence and position sequence of each phoneme of the original audio;

[0007] Input the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional feature output by the first convolutional structure;

[0008] Generate a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio, and input the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional feature output by the second convolutional structure;

[0009] Perform an inner product calculation on the first convolutional feature and the second convolutional feature to obtain an inner product calculation result;

[0010] Perform SoftMax calculation on the inner product calculation result to obtain a weight vector, and weight the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors;

[0011] Add the phoneme vector to the position sequence to obtain a conditional vector, input the conditional vector into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model, and stop training when the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio satisfies the convergence condition to obtain the target acoustic model.

[0012] The second aspect of the embodiments of the present application provides a computer device, and the method includes:

[0013] An acquisition unit, configured to acquire the acoustic feature parameters of the original audio, and acquire the phoneme sequence and the position sequence of each phoneme of the original audio;

[0014] A feature extraction unit, configured to input the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional features output by the first convolutional structure;

[0015] A generation unit, configured to generate a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio;

[0016] The feature extraction unit is further configured to input the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional features output by the second convolutional structure;

[0017] A calculation unit, configured to perform an inner product calculation on the first convolutional features and the second convolutional features to obtain an inner product calculation result;

[0018] The calculation unit is further configured to perform SoftMax calculation on the inner product calculation result to obtain a weight vector, and weight the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors;

[0019] A training unit, configured to add the phoneme vector to the position sequence to obtain a conditional vector, input the conditional vector into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model, and stop training when the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio satisfies the convergence condition to obtain the target acoustic model.

[0020] In the third aspect of the embodiments of the present application, a computer device is provided, including a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method in the foregoing first aspect is implemented.

[0021] In the fourth aspect of the embodiments of the present application, a computer storage medium is provided. Instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer executes the method in the foregoing first aspect.

[0022] As can be seen from the above technical solutions, the embodiments of the present application have the following advantages:

[0023] In this embodiment, the computer device inputs the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional features output by the first convolutional structure, generates a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio, inputs the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional features output by the second convolutional structure, performs an inner product calculation on the first convolutional features and the second convolutional features to obtain an inner product calculation result, performs a SoftMax calculation on the inner product calculation result to obtain a weight vector, weights the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors, adds the phoneme vectors to the position sequence to obtain a conditional vector, inputs the conditional vector into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model, and stops training when the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio satisfies the convergence condition to obtain the target acoustic model. Therefore, it gets rid of the previous limitation of manually annotating the phoneme positions and durations, enables the phonemes to more accurately correspond to the duration of the audio, and thus improves the training effect of the singing synthesis model. Description of the Drawings

[0024] Figure 1 It is a schematic flowchart of a method for training a phoneme alignment model in an embodiment of the present application;

[0025] Figure 2 It is another schematic flowchart of a method for training a phoneme alignment model in an embodiment of the present application;

[0026] Figure 3 It is a schematic structural diagram of a phoneme alignment model in an embodiment of the present application;

[0027] Figure 4 It is a schematic structural diagram of an initial acoustic model in an embodiment of the present application;

[0028] Figure 5 It is a schematic structural diagram of a computer device in an embodiment of the present application;

[0029] Figure 6 Another structural schematic diagram of the computer device in the embodiment of the present application. Detailed implementation manners

[0030] The embodiment of the present application provides a method for training a phoneme alignment model, a computer device, and a computer storage medium, which are used to accurately align the duration and position of each phoneme of an audio.

[0031] Please refer to Figure 1 , an embodiment of the method for training a phoneme alignment model in the embodiment of the present application includes:

[0032] 101. Obtain the acoustic feature parameters of the original audio, and obtain the phoneme sequence and position sequence of each phoneme of the original audio;

[0033] The method in this embodiment can be applied to a computer device, which can be a server, a terminal, or other computer devices capable of performing data processing. When the computer device is a terminal, it can be a personal computer (PC), a desktop computer, or other terminal devices; when the computer device is a server, it can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud databases, cloud computing, and big data and artificial intelligence platforms.

[0034] The computer device can obtain the original audio for training the singing synthesis model, obtain the acoustic feature parameters of the original audio, and obtain the phoneme sequence and position sequence of each phoneme of the original audio. The acoustic feature parameters of the original audio refer to the specific parameters of the acoustic features of the audio, where the acoustic features refer to the physical quantities representing the acoustic characteristics of speech, and are also the general term for the acoustic manifestations of various elements of sound, such as the energy concentration area, formant frequency, formant intensity, and bandwidth representing timbre, as well as the duration, fundamental frequency, and average speech power representing the prosodic characteristics of speech.

[0035] The phoneme sequence of each phoneme of the original audio refers to the sequence formed by multiple identical phonemes after each phoneme is expanded. The purpose of phoneme expansion is to make the length of the phoneme sequence the same as the length of the sequence of acoustic feature parameters. Each element in the position sequence of the phonemes can reflect the position of each phoneme in the phoneme sequence.

[0036] 102. Input the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional feature output by the first convolutional structure;

[0037] After obtaining the acoustic feature parameters of the original audio, input the acoustic feature parameters into the first convolutional structure of the phoneme alignment model. The first convolutional structure extracts features from the acoustic feature parameters and outputs the feature extraction result, that is, the first convolutional feature.

[0038] 103. Generate a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio, and input the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional feature output by the second convolutional structure;

[0039] The computer device generates a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio, and inputs the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model. The second convolutional structure extracts features from the phoneme sequence vectors of three adjacent phonemes and outputs the feature extraction result, that is, the second convolutional feature.

[0040] 104. Perform an inner product calculation on the first convolutional feature and the second convolutional feature to obtain an inner product calculation result;

[0041] 105. Perform a SoftMax calculation on the inner product calculation result to obtain a weight vector, and weight the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors;

[0042] After obtaining the first convolutional feature and the second convolutional feature, perform an inner product calculation on the first convolutional feature and the second convolutional feature to obtain an inner product calculation result, and perform a SoftMax calculation on this inner product calculation result to obtain a weight vector. Weight the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors.

[0043] 106. Add the phoneme vector and the position sequence to obtain a conditional vector, and input the conditional vector into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model. When the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio satisfies the convergence condition, stop training to obtain the target acoustic model;

[0044] The phoneme vectors of each phoneme of the original audio are added to the position sequence of the phonemes to obtain a conditional vector. The conditional vector is input into the initial acoustic model. The initial acoustic model generates and outputs target acoustic feature parameters according to the conditional vector, and adjusts the model parameters according to the relationship between the output target acoustic feature parameters and the acoustic feature parameters of the original audio. When the relationship between the output target acoustic feature parameters and the acoustic feature parameters of the original audio meets the convergence condition, the model training is stopped, and the target acoustic model is obtained. The target acoustic model can be used to synthesize audio according to the acoustic feature parameters of the audio, such as synthesizing singing according to the acoustic feature parameters of singing.

[0045] In this embodiment, the computer device inputs the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional feature output by the first convolutional structure, generates the phoneme sequence vector of each phoneme according to the phoneme sequence of each phoneme of the original audio, and inputs the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional feature output by the second convolutional structure. The first convolutional feature and the second convolutional feature are subjected to inner product calculation to obtain the inner product calculation result. The SoftMax calculation is performed on the inner product calculation result to obtain the weight vector. The phoneme sequence vectors of every three adjacent phonemes of the original audio are weighted according to the weight vector to obtain the phoneme vector. The phoneme vector is added to the position sequence to obtain the conditional vector. The conditional vector is input into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model. When the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio meets the convergence condition, the training is stopped, and the target acoustic model is obtained. Therefore, it gets rid of the previous limitation of manually annotating the phoneme position and duration, making the phonemes more accurately correspond to the duration of the audio, thereby improving the training effect of the singing synthesis model.

[0046] The following will further describe the embodiments of the present application in detail on the basis of the foregoing Figure 1 illustrated embodiments. Please refer to Figure 2 , another embodiment of the phoneme alignment model training method in the embodiments of the present application includes:

[0047] 201. Obtain the acoustic feature parameters of the original audio, and obtain the phoneme sequence and position sequence of each phoneme of the original audio;

[0048] In this embodiment, the acoustic feature parameters of the original audio may specifically include spectral envelope parameters (SP) and aperiodic signals (AP). A vocoder can be used to extract the fundamental frequency, spectral envelope, and aperiodic signals from the original audio. Among them, the vocoder can be configured with the DIO algorithm and use this algorithm to extract the fundamental frequency feature parameters of the original audio; and configure the CheapTrick algorithm, input the extracted fundamental frequency and the waveform of the original audio into the CheapTrick algorithm to obtain the spectral envelope SP feature parameters output by the CheapTrick algorithm; and configure the D4C algorithm, input the fundamental frequency, spectral envelope SP, and the waveform of the original audio into the D4C algorithm to obtain the aperiodic signals output by the D4C algorithm. The fundamental frequency, spectral envelope, and aperiodic signals can be used to restore the original audio through a speech synthesis algorithm.

[0049] Among them, the vocoder can be a WORLD vocoder, or a STRAIGHT vocoder, a GriffimLim vocoder, etc. The specific type of the vocoder is not limited.

[0050] In this embodiment, to obtain the phoneme sequence of each phoneme in the original audio, one implementation method can be to determine the number of audio frames of the original audio corresponding to each phoneme according to the pre-annotation information, generate a copy of each phoneme in the original audio, and the number of copies is the number of audio frames of the original audio corresponding to the phoneme. The copies of the phoneme constitute the phoneme sequence of the phoneme.

[0051] Among them, the pre-annotation information represents the number of audio frames corresponding to each phoneme in the original audio, that is, the duration corresponding to each phoneme, and it can be given manually, that is, manually pre-annotated.

[0052] For example, for Chinese singing synthesis, the lyric text is generally in Chinese characters. Chinese characters cannot directly represent the pronunciation situation, so a text front-end tool is needed to convert Chinese characters into pinyin form. However, pinyin also cannot directly correspond to the pronunciation situation. For example, in pinyin, the "y" and "w" in "yu" and "wu" are not pronounced, so it is necessary to further parse pinyin into phoneme form. Each phoneme corresponds to a pronunciation situation, and each phoneme corresponds to several frames of acoustic feature parameters. To form a one-to-one mapping relationship between phonemes and acoustic feature parameters, it is necessary to expand the phonemes. For example, if the pre-annotation information indicates that the number of audio frames of a certain phoneme corresponding to the original audio is 3, then this phoneme needs to be repeated 3 times to generate 3 copies of this phoneme, and the 3 copies of this phoneme constitute the phoneme sequence of this phoneme.

[0053] In a preferred implementation of this embodiment, obtaining the position sequence of each phoneme of the original audio may be that the computer device determines the number of audio frames of the original audio corresponding to each phoneme according to the pre-annotation information, and generates a position serial number identifier for each phoneme of the original audio. The number of position serial number identifiers is the number of audio frames of the original audio corresponding to the phoneme, and the position serial number identifiers of the phoneme constitute the position sequence.

[0054] For example, the position serial number identifier can be represented in the form of a fraction, where the numerator represents the position of each element in the phoneme sequence in the phoneme sequence, and the denominator can represent the total number of elements in the phoneme sequence. For example, if the pre-annotation information indicates that the number of audio frames of the original audio corresponding to the phoneme "a" in Chinese pinyin is N, then its phoneme sequence can be expressed as "a 1 , a 2 , …, a N ", and the corresponding position sequence can be expressed as "1 / N, 2 / N, … N / N", that is, each element in the position sequence can represent the position serial number of each element in the phoneme sequence in the phoneme sequence, so as to strengthen the position information of each phoneme in the phoneme sequence and improve the pronunciation quality.

[0055] 202. Input the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional feature output by the first convolutional structure;

[0056] In this embodiment, the acoustic feature parameters of the original audio include the spectral envelope and the aperiodic signal. When using the first convolutional structure to obtain the first convolutional feature, the vector dimension of the acoustic feature parameters of each frame of the original audio can be determined according to the dimension of the spectral envelope of the original audio and the dimension of the aperiodic signal, and the acoustic feature parameters of T frames of the original audio are input into the first convolutional structure. Thus, the first convolutional structure outputs the first convolutional feature according to the vector dimension of the acoustic feature parameters of each frame of the original audio and the number of channels of the first convolutional structure, where T is a positive integer greater than 2.

[0057] This embodiment provides a phoneme alignment model, which is mainly used to process the inevitable errors in data annotation. For example, the structure of the phoneme alignment model is as Figure 3 shown, which includes a first convolutional structure Conv1 and a second convolutional structure Conv2, as well as a MatMul structure for inner product calculation, a Scale structure for performing a scaling operation, and a SoftMax structure for performing a SoftMax calculation.

[0058] In one embodiment, the input of the first convolutional structure Conv1 is the phonetic feature, which is the acoustic feature parameter obtained in the foregoing steps, such as the 60-dimensional spectral envelope of the original audio and the 4-dimensional aperiodic signal. Therefore, the vector dimension of the acoustic feature parameter of each frame of the original audio is [1, 64]. Assuming that the number of frames of the original video is T, the dimension of the acoustic feature parameter of the original video is [T, 64]. If the convolutional kernel size of the first convolutional structure is 5, the stride is 1, and the number of output channels is 128, after the acoustic feature parameter of the original video is input to the first convolutional structure, the first convolutional feature with the output dimension of [T, 128] can be obtained. For the convenience of subsequent matrix operations, the dimension of the first convolutional feature can be expanded to [T, 1, 128].

[0059] 203. Generate a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio, and input the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional feature output by the second convolutional structure;

[0060] When using the second convolutional structure to obtain the second convolutional feature, the vector dimension of the phoneme sequence vector of each phoneme of the original audio can be determined according to the dimension of the spectral envelope and the dimension of the aperiodic signal of the original audio, and the phoneme sequence vectors of every three adjacent phonemes of the original audio are input into the second convolutional structure, so that the second convolutional structure outputs the second convolutional feature according to the vector dimension of the phoneme sequence vector of each phoneme of the original audio and the number of channels of the second convolutional structure.

[0061] The input of the second convolutional structure Conv2 is the acoustic feature, which is the phoneme sequence vector of every three adjacent phonemes in the original audio. For example, word embedding operations can be performed on each phoneme of the original audio to obtain the phoneme sequence vector of each phoneme. Therefore, the dimension of the phoneme sequence vector corresponding to each phoneme of the original audio is [1, 3, 64]. After the feature extraction of the second convolutional structure Conv2, the second convolutional feature with the dimension of [T, 3, 128] can be obtained.

[0062] For example, Prev sequence and Post sequence can be constructed, corresponding to the previous phoneme and the next phoneme of each phoneme respectively. For example, in the lyrics "that's me", the pronunciation phonemes are [n, a, sh, i, uo], and the pre-annotation information indicates that the number of frames corresponding to each phoneme is 2, 4, 3, 4, 3 respectively. Then, the phoneme sequence of each phoneme and the Prev sequence and Post sequence can be obtained, as shown in Table 1 specifically.

[0063] Table 1

[0064] sp sp n n n n a a a sh sh sh sh i i i n n a a a a sh sh sh i i i i uo uo uo a a sh sh sh sh i i i uo uo uo uo sp sp sp

[0065] In Table 1, "sp" represents silence. The middle phoneme in each column is the current phoneme, the top phoneme is the previous phoneme of the current phoneme, and the bottom phoneme is the next phoneme of the current phoneme. It can be seen from the middle row of Table 1 that the phoneme "n" corresponds to 2 audio frames, so it has 2 copies; the phoneme "a" corresponds to 4 audio frames, so it has 4 copies... And the copies of each phoneme correspond to the copies of the previous phoneme and the next phoneme. Then, the first row of Table 1 constitutes the Prev sequence, and the third row constitutes the Post sequence. Therefore, as shown in Table 1, every three adjacent phonemes can form a phoneme sequence. For example, "spna" in the second column of Table 1 forms 1 phoneme sequence, "spna" in the first column forms 1 phoneme sequence, "nash" in the third column forms 1 phoneme sequence...

[0066] 204. Calculate the inner product of the first convolutional feature and the second convolutional feature to obtain the inner product calculation result;

[0067] 205. Perform SoftMax calculation on the inner product calculation result to obtain a weight vector, and weight the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors;

[0068] After obtaining the first convolutional feature and the second convolutional feature, the inner product of the two can be calculated to obtain the inner product calculation result.

[0069] Continuing with the above example, after obtaining the first convolutional feature [T, 1, 128] and the second convolutional feature [T, 3, 128], perform the inner product calculation MatMul on the two to obtain the inner product calculation result, whose dimension is [T, 1, 3].

[0070] In a preferred embodiment, in order to prevent the inner product calculation result from being too large, a scaling operation scale can be performed on the inner product calculation result. For example, the inner product calculation result can be divided by the square root of the feature dimension, where the feature dimension is 128 here.

[0071] After that, the scaled result of the inner product calculation result can be subjected to SoftMax calculation to obtain a weight vector, and the phoneme vectors of each phoneme sequence of the original audio are weighted according to this weight vector to obtain phoneme vectors.

[0072] Continuing with the above example, after performing SoftMax calculation on the inner product calculation result of [T, 1, 3], the weight vector weight can be obtained, which represents the similarity between the current acoustic feature parameters and the adjacent three phonemes. Then, the phoneme sequence vectors of the adjacent three phonemes are weighted according to the weight vector weight to obtain a phoneme vector with a dimension of [T, 128].

[0073] 206. Add the phoneme vector to the position sequence to obtain a conditional vector, and input the conditional vector into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model. When the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio satisfies the convergence condition, stop training to obtain the target acoustic model;

[0074] In this embodiment, after obtaining the phoneme vector and the position sequence of each phoneme of the original audio, the phoneme vector and the position sequence can be added to obtain a conditional vector. Continuing with the above example, add the phoneme vector [T, 128] and the position sequence [T, 128] to obtain a conditional vector Conditional Input with a dimension of [T, 128], and use this conditional vector as the input of the initial acoustic model. The initial acoustic model performs feature extraction based on the input conditional vector and uses 64-dimensional acoustic features (60-dimensional spectral envelope parameters and 4-dimensional aperiodic signals) as output information for supervised learning. When the relationship between the target acoustic feature parameters output by the initial acoustic model and the acoustic feature parameters of the original audio satisfies the convergence condition, stop training. The relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio can be represented by a loss function. This loss function can be the mean squared error, and the optimizer can be Adam, and the learning rate can be set to 1e-5.

[0075] In a preferred implementation, the structure of the initial acoustic model can be as Figure 4 shown, which consists of a series of convolutional layers (Conv) and normalization layers (LayerNorm), where Add represents adding features, Split represents splitting the features into two equal parts, Mul represents multiplying features, and the number of stacked layers is M. Generally, M is taken as 8. 1x1 represents a convolution with a convolution kernel size of 1.

[0076] In another preferred embodiment of this embodiment, after obtaining the target acoustic model, the target acoustic model can be used to synthesize singing voices. For example, the target audio to be processed can be input into the target acoustic model. Then, based on the model structure and the model parameters of each model structure obtained through pre-training, the target acoustic model extracts the acoustic feature parameters of several audio frames corresponding to each phoneme in the target audio, and synthesizes the singing voice data corresponding to the target audio according to the acoustic feature parameters of several audio frames corresponding to each phoneme in the target audio. Among them, the acoustic feature parameters of several audio frames corresponding to each phoneme in the target audio can include the fundamental frequency, spectral envelope, and aperiodic signal. Based on these three acoustic feature parameters, the singing voice data corresponding to the target audio can be synthesized. Since the foregoing training process of the target acoustic model enables the target acoustic model to accurately mark the position and pronunciation duration of each phoneme in the target audio, the synthesis effect of the finally synthesized singing voice data is better, and the quality of singing voice synthesis is improved.

[0077] The method for training the phoneme alignment model in the embodiment of the present application is described above. Next, the computer device in the embodiment of the present application will be described. Please refer to Figure 5 , an embodiment of the computer device in the embodiment of the present application includes:

[0078] An acquisition unit 501, configured to acquire the acoustic feature parameters of the original audio, and acquire the phoneme sequence and position sequence of each phoneme of the original audio;

[0079] A feature extraction unit 502, configured to input the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain a first convolutional feature output by the first convolutional structure;

[0080] A generation unit 503, configured to generate a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio;

[0081] The feature extraction unit 502 is further configured to input the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain a second convolutional feature output by the second convolutional structure;

[0082] A calculation unit 504, configured to perform an inner product calculation on the first convolutional feature and the second convolutional feature to obtain an inner product calculation result;

[0083] The calculation unit 504 is further configured to perform a SoftMax calculation on the inner product calculation result to obtain a weight vector, and weight the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain a phoneme vector;

[0084] A training unit 505, configured to add the phoneme vector and the position sequence to obtain a conditional vector, input the conditional vector into an initial acoustic model to obtain target acoustic feature parameters output by the initial acoustic model, and stop training when the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio satisfies a convergence condition, so as to obtain a target acoustic model.

[0085] In a preferred implementation manner of this embodiment, the generating unit 503 is specifically configured to perform word embedding operations on each phoneme of the original audio to obtain a phoneme sequence vector for each phoneme.

[0086] In a preferred implementation manner of this embodiment, the obtaining unit 501 is specifically configured to determine the number of audio frames of the original audio corresponding to each phoneme in the original audio according to pre-annotation information; generate a copy of each phoneme of the original audio, where the number of copies is the number of audio frames of the original audio corresponding to the phoneme, and the copies of the phoneme constitute a phoneme sequence of the phoneme.

[0087] In a preferred implementation manner of this embodiment, the obtaining unit 501 is specifically configured to determine the number of audio frames of the original audio corresponding to each phoneme in the original audio according to pre-annotation information; generate a position serial number identifier for each phoneme of the original audio, where the number of position serial number identifiers is the number of audio frames of the original audio corresponding to the phoneme, and the position serial number identifiers of the phoneme constitute the position sequence.

[0088] In a preferred implementation manner of this embodiment, the computer device further includes:

[0089] A scaling unit 506, configured to perform a scaling operation on the inner product calculation result to obtain a scaled result of the inner product calculation result;

[0090] The calculating unit 504 is specifically configured to perform a SoftMax calculation on the scaled result of the inner product calculation result to obtain a weight vector.

[0091] In this embodiment, the operations performed by each unit in the computer device are similar to those described in the foregoing Figures 1 to 2 illustrated embodiment, and will not be described herein again.

[0092] In this embodiment, the computer device inputs the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain the first convolutional features output by the first convolutional structure, generates the phoneme sequence vectors of each phoneme according to the phoneme sequence of each phoneme of the original audio, inputs the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain the second convolutional features output by the second convolutional structure, performs an inner product calculation on the first convolutional features and the second convolutional features to obtain the result of the inner product calculation, performs a SoftMax calculation on the result of the inner product calculation to obtain a weight vector, weights the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors, adds the phoneme vectors to the position sequence to obtain a conditional vector, inputs the conditional vector into the initial acoustic model to obtain the target acoustic feature parameters output by the initial acoustic model, and stops training when the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio meets the convergence condition, thereby obtaining the target acoustic model. Therefore, it gets rid of the limitations of manually annotating the phoneme positions and durations in the past, enabling the phonemes to more accurately correspond to the duration of the audio, thus improving the training effect of the singing synthesis model.

[0093] The computer device in the embodiments of the present application will be described below. Please refer to Figure 6 , an embodiment of the computer device in the embodiments of the present application includes:

[0094] The computer device 600 may include one or more central processing units (CPUs) 601 and a memory 605, and one or more applications or data are stored in the memory 605.

[0095] Among them, the memory 605 may be volatile storage or persistent storage. The program stored in the memory 605 may include one or more modules, and each module may include a series of instruction operations on the computer device. Further, the central processing unit 601 may be configured to communicate with the memory 605 and execute a series of instruction operations in the memory 605 on the computer device 600.

[0096] The computer device 600 may further include one or more power supplies 602, one or more wired or wireless network interfaces 603, one or more input / output interfaces 604, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSD TM, etc.

[0097] The central processing unit 601 may execute the operations performed by the computer device in the foregoing Figures 1 to 2 illustrated embodiment, which will not be elaborated herein specifically.

[0098] An embodiment of the present application also provides a computer storage medium. One embodiment includes: instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer is caused to perform the operations performed by the computer device in the foregoing Figures 1 to 2 illustrated embodiment.

[0099] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0100] In several embodiments provided by the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other can be indirect couplings or communication connections through some interfaces, devices, or units, and can be in electrical, mechanical, or other forms.

[0101] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0102] In addition, the functional units in each embodiment of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0103] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs.

Claims

1. A method for training a phoneme alignment model, characterized in that, the method comprises: obtaining acoustic feature parameters of the original audio, and obtaining a phoneme sequence and a position sequence of each phoneme of the original audio; inputting the acoustic feature parameters of the original audio into the first convolutional structure of the phoneme alignment model to obtain a first convolutional feature output by the first convolutional structure; generating a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio, and inputting the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure of the phoneme alignment model to obtain a second convolutional feature output by the second convolutional structure; performing an inner product calculation on the first convolutional feature and the second convolutional feature to obtain an inner product calculation result; performing a SoftMax calculation on the inner product calculation result to obtain a weight vector, and weighting the phoneme sequence vectors of every three adjacent phonemes of the original audio according to the weight vector to obtain phoneme vectors; adding the phoneme vectors and the position sequence to obtain a conditional vector, and inputting the conditional vector into an initial acoustic model to obtain target acoustic feature parameters output by the initial acoustic model, and stopping training when the relationship between the target acoustic feature parameters and the acoustic feature parameters of the original audio meets a convergence condition to obtain a target acoustic model.

2. The method according to claim 1, characterized in that, the generating a phoneme sequence vector for each phoneme according to the phoneme sequence of each phoneme of the original audio comprises: performing a word embedding operation on each phoneme of the original audio to obtain a phoneme sequence vector for each phoneme.

3. The method according to claim 1, characterized in that, the obtaining the phoneme sequence of each phoneme of the original audio comprises: determining the number of audio frames of the original audio corresponding to each phoneme according to pre-annotation information; generating copies of each phoneme of the original audio, the number of copies being the number of audio frames of the original audio corresponding to the phoneme, and the copies of the phoneme constituting the phoneme sequence of the phoneme.

4. The method according to claim 1, characterized in that, the obtaining the position sequence of each phoneme of the original audio comprises: determining the number of audio frames of the original audio corresponding to each phoneme according to pre-annotation information; generating a position serial number identifier for each phoneme of the original audio, the number of position serial number identifiers being the number of audio frames of the original audio corresponding to the phoneme, and the position serial number identifiers of the phoneme constituting the position sequence.

5. The method according to claim 1, characterized in that, the method further comprises: performing a scaling operation on the inner product calculation result to obtain a scaled result of the inner product calculation result; the performing a SoftMax calculation on the inner product calculation result to obtain a weight vector comprises: performing a SoftMax calculation on the scaled result of the inner product calculation result to obtain a weight vector.

6. The method according to claim 1, characterized in that, the acoustic feature parameters include a spectral envelope and an aperiodic signal; Inputting the acoustic feature parameters of the original audio into a first convolutional structure of a phoneme alignment model to obtain first convolutional features output by the first convolutional structure includes: Determining the vector dimension of the acoustic feature parameters of each frame of the original audio according to the dimension of the spectral envelope and the dimension of the aperiodic signal; Inputting the acoustic feature parameters of T frames of the original audio into the first convolutional structure, so that the first convolutional structure outputs the first convolutional features according to the vector dimension of the acoustic feature parameters of each frame of the original audio and the number of channels of the first convolutional structure, where T is a positive integer greater than 2.

7. The method according to claim 1, wherein, the acoustic feature parameters include a spectral envelope and an aperiodic signal; Inputting the phoneme sequence vectors of every three adjacent phonemes of the original audio into a second convolutional structure of the phoneme alignment model to obtain second convolutional features output by the second convolutional structure includes: Determining the vector dimension of the phoneme sequence vector of each phoneme of the original audio according to the dimension of the spectral envelope and the dimension of the aperiodic signal; Inputting the phoneme sequence vectors of every three adjacent phonemes of the original audio into the second convolutional structure, so that the second convolutional structure outputs the second convolutional features according to the vector dimension of the phoneme sequence vector of each phoneme of the original audio and the number of channels of the second convolutional structure, where T is a positive integer greater than 2.

8. The method according to any one of claims 1 to 7, wherein, after obtaining the target acoustic model, the method further includes: Inputting the target audio to be processed into the target acoustic model, so that the target acoustic model extracts the acoustic feature parameters of several audio frames corresponding to each phoneme in the target audio, and synthesizes the singing data corresponding to the target audio according to the acoustic feature parameters of several audio frames corresponding to each phoneme in the target audio.

9. A computer device, including a memory and a processor, the memory stores a computer program, wherein, when the processor executes the computer program, the method according to any one of claims 1 to 8 is implemented.

10. A computer storage medium, wherein, instructions are stored in the computer storage medium, and when the instructions are executed on a computer, the computer executes the method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Voice synthesis method and device based on attention mechanism

    CN109767752A

  • System and method for automatically generating musical output

    CN111213200A