A singing synthesis method, device and storage medium for generating personalized timbre
By establishing an acoustic feature training model and fine-tuning it using the Transformer structure and speaker embedding module, the problems of high tone customization cost and weak generalization ability in the existing technology are solved, and the effect of generating personalized singing tones under small batch data is achieved.
Patent Information
- Application Number
- CN202210434225.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-24
- Publication Date
- 2025-05-02
- Estimated Expiration
- 2042-04-24
AI Technical Summary
The existing singing synthesis technology requires large batches of data to retrain the model when customizing tone, resulting in weak generalization ability, high cost, and long training time, making it difficult to find a balance between parameters and sound quality.
By obtaining historical acoustic feature data, establish an acoustic feature training model, input set acoustic feature data for preprocessing, form a phoneme data sequence, and fine-tune it using the Transformer structural model and speaker embedding module, and introduce a conditional normalization unit to generate a personalized tone.
It realizes customizing personalized singing tones under small batch data, improves the generalization ability of the model, reduces costs and training time, and finds a balance between parameters and sound quality.
Smart Images

Figure CN114724539B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of speech signal processing and artificial intelligence technology, and in particular to a singing synthesis method, device and storage medium for generating personalized timbre. Background Art
[0002] In recent years, with the continuous development of artificial intelligence, its technology has been applied in various fields. Artificial intelligence has more and more application scenarios in entertainment and education. Singing synthesis is the application of artificial intelligence in the field of singing, which can not only reduce the cost of music creation and music education, but also improve efficiency, thereby promoting the development of the singing industry. In the prior art, singing synthesis technology is to synthesize one or more clearer singing timbres through large amounts of data. However, this will cause a series of problems. On the one hand, if you want to complete the timbre customization, you need to retrain a new model with large amounts of data, but the new model cannot obtain the fine-grained information of acoustic features, resulting in weak model generalization ability and increased costs for the customization party; on the other hand, the new model training time is long, and there is no good way to find a balance between parameters and sound quality, which results in an increase in the memory storage and service costs of the service provider. In response to the above problems, we designed a singing synthesis method, device and storage medium for generating personalized timbre. Summary of the invention
[0003] The purpose of the present invention is to provide a singing synthesis method, device and storage medium for generating personalized timbre, which are used to solve the above-mentioned technical problems.
[0004] The embodiments of the present invention are implemented by the following technical solutions:
[0005] A singing synthesis method for generating personalized timbre, comprising the following contents:
[0006] Acquire historical acoustic feature data, establish an acoustic feature training model, train the acoustic feature training model using the historical acoustic feature data, and obtain a trained acoustic feature training model;
[0007] The set acoustic feature data is input, and after preprocessing, a phoneme data sequence is obtained. The phonemes are expanded according to the duration of the phonemes to form an expanded phoneme sequence. The expanded phoneme sequence is processed to make it consistent with the length of the set acoustic feature data, and then integrated and input into the trained acoustic feature training model for calculation to obtain a spectrogram. The spectrogram is synthesized through a vocoder to complete the generation of personalized timbre, wherein the phoneme data sequence includes the duration and pitch of each phoneme.
[0008] Optionally, the historical acoustic feature data includes singing audio, stress, rhythm, duration and environmental noise.
[0009] Optionally, the acoustic feature training model is specifically a Transformer structure model, and each Transformer block in the Transformer structure model includes a self-attention network and a feedforward network.
[0010] Optionally, the acoustic feature training model is preset with a speaker embedding module, and the speaker embedding module is used to obtain fine-grained data in the acoustic features.
[0011] Optionally, the acoustic feature training model further introduces a conditional normalization unit, and its calculation formula is as follows:
[0012]
[0013]
[0014] in, and are matrices, E s Speaker embedding module;
[0015] The self-attention network and the feedforward network are normalized through the conditional normalization unit to complete the fine-tuning of the Transformer structure model.
[0016] Optionally, the preprocessing process of setting the acoustic feature data is specifically: embedding the set acoustic feature data into a dense vector of the same dimension to obtain a vector sequence, and then performing computational superposition on the vector sequence and the position code, and after passing through multiple one-dimensional convolutional networks, obtaining a phoneme data sequence.
[0017] A singing synthesis device for generating personalized timbre, comprising:
[0018] Memory for storing computer programs;
[0019] A processor is used to implement the steps of a singing synthesis method for generating personalized timbre as described in any one of the above when executing the computer program.
[0020] A readable storage medium having a computer program stored thereon, wherein when the computer program is executed by a processor, the steps of a singing synthesis method for generating a personalized timbre as described in any one of the above are implemented.
[0021] The technical solution of the embodiment of the present invention has at least the following advantages and beneficial effects:
[0022] The present invention has reasonable design and simple structure, and achieves the purpose of generating personalized timbre by adding a speaker embedding module and introducing a conditional normalization layer in the decoder to fine-tune some parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] Figure 1 A schematic flow chart of a singing synthesis method for generating personalized timbre provided by the present invention. DETAILED DESCRIPTION
[0024] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0025] In the present invention, acoustic features are divided into two dimensions: one is singing audio; the other is acoustic conditions at the phoneme level, including stress, rhythm, and time environment noise, etc. Because using small batches of data to customize personalized timbre singing will cause overfitting and insufficient model generalization for some features. Therefore, it is necessary to first train an acoustic model with a large batch of singing data so that the decoder can predict the timbre of singing under different acoustic conditions based on these acoustic information.
[0026] In addition, the present invention also includes a music score encoder: the phonemes, duration and pitch of the music score are taken as input, the position code is embedded together with the music data and passed through multiple Transformer layers to obtain the output result of the encoder.
[0027] Variance adapter: The result is input into a duration processor composed of multiple layers of CNN, Linear, etc., and the hidden sequence of each phoneme is obtained to provide variance information including duration, pitch and energy, and the encoder vector sequence is expanded according to its information.
[0028] Mel-spectrogram decoder: The expanded vector sequence is input into the decoder, and the positional encoding and the input vector sequence are passed through multiple Transformer layers and Linear layers to obtain the Mel-spectrogram input to the vocoder, and finally the vocoder is used to synthesize the singing.
[0029] like Figure 1 As shown, the present invention provides one embodiment, and the specific contents are as follows:
[0030] Music scores usually include elements such as phonemes, duration, and pitch, which are necessary input elements for singing. The song is converted into a phoneme sequence. Each word in the song is decomposed into multiple phonemes, and the pitch is converted into pitch values according to the standards of music theory knowledge. Duration is the number of frames of each phoneme.
[0031] The three input factors are embedded into dense vectors of the same dimension, superimposed with the position encoding operation, and encoded through multiple one-dimensional convolutional networks.
[0032] Since there is not enough data for inputting customized singer’s timbre, timbre, rhythm and recording environment to predict the target timbre, the generalization ability is poor during the model adaptation process. Therefore, the speaker embedding is used to capture the rich acoustic features in the adapted sound, and some parameters in the model are trained from acoustic features of different granularities to improve the generalization ability of the model during the training stage. A singer’s acoustic model is first trained using a large amount of data, and the singer dimension trains models of conditions such as stress, rhythm and time environment noise on phonemes to ensure that the singing timbre of a small batch of data can be inferred. The acoustic model models the acoustic conditions at the singing audio and phoneme levels respectively. As the input of the Mel spectrogram decoder, the decoder can predict the singing timbre under different acoustic conditions based on these acoustic information.
[0033] The different granularities mentioned above are represented as follows: singing level, which is the fine-grained acoustic conditions presented by the speaker in each sentence; phoneme level, which is the finer-grained acoustic conditions of each phoneme in a sentence, which needs to be established through the speakembedding module.
[0034] Obtain the hidden sequence of phonemes, which hides the duration and pitch of each phoneme. Expand the phoneme sequence according to the duration of each phoneme (for example, if a phoneme lasts for three seconds, we will copy the phoneme three times to expand the phoneme sequence), and the pitch element also forms a sequence corresponding to the expanded phoneme sequence. After that, the integrated output is the sequence feature aligned with the phoneme feature sequence (with the same length), so it is necessary to align the acoustic features with the phoneme sequence in advance, and then take the average of the acoustic features corresponding to the phonemes to facilitate conversion into the corresponding spectrogram.
[0035] The model is basically built on the Transformer structure, with a self-attention network and a feedforward network in each Transformer block. After applying normalization to the self-attention network and the feedforward network in the encoder, the learnable scaleγ and biasβ can effectively affect the hidden activations and the final prediction results. A small conditional network determines the scale and bias vector in the layer normalization according to the corresponding speaker characteristics, and fine-tunes this conditional network. The conditional network consists of two simple linear layers and Composition, E s For the speaker embedding module, we only need to fine-tune two matrices and Normalization is done at each conditional layer of the decoder and singer embeddings according to the following formula:
[0036]
[0037]
[0038] Calculate each scale to get scaleγ and biasβ, and use a small conditional network to determine the normalized scale and bias vectors, and input the acoustic features of the corresponding speaker. Only two simple linear layers are used, the input is the speaker embedding, and the output is the predicted γ and β. By changing the parameters of the normalization operation in the decoder, the model can be indirectly adjusted to achieve the purpose of customizing personalized singing with small batches of data. The learnable scaleγ and biasβ can effectively affect the hidden activation and the final prediction results.
[0039] The above are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A singing synthesis method for generating personalized timbre, characterized in that: It includes the following: Acquire historical acoustic feature data, establish an acoustic feature training model, train the acoustic feature training model using the historical acoustic feature data, and obtain a trained acoustic feature training model; Input the set acoustic feature data, obtain a phoneme data sequence after preprocessing, expand the phoneme according to the duration of the phoneme to form a phoneme expansion sequence, process the phoneme expansion sequence to make it consistent with the length of the set acoustic feature data, and then integrate and input it into the trained acoustic feature training model for calculation to obtain a spectrogram, synthesize the spectrogram through a vocoder to complete the generation of personalized timbre, wherein the phoneme data sequence includes the duration and pitch of each phoneme; It also includes a music score encoder: it takes the phonemes, duration and pitch of the music score as input, and passes the position encoding together with the music data embedding through multiple Transformer layers to obtain the encoder output; Variance adapter: The result is input into a duration processor composed of multiple layers of CNN and Linear to obtain the variance information including duration, pitch and energy of each phoneme hidden sequence, and the encoder vector sequence is expanded according to its information; Mel-spectrogram decoder: The expanded vector sequence is input into the decoder, and the positional code and the input vector sequence are passed through multiple Transformer layers and Linear layers to obtain the Mel-spectrogram input to the vocoder, and finally the vocoder is used to synthesize the singing; Music scores usually include phonemes, duration, and pitch elements, which are necessary input elements for singing. The song is converted into a phoneme sequence. Each word in the song is decomposed into multiple phonemes. The pitch is converted into pitch values according to the standards of music theory knowledge. The duration is the number of frames of each phoneme. The three input factors are embedded into dense vectors of the same dimension, superimposed with position encoding operations, and encoded through multiple one-dimensional convolutional networks; By capturing the rich acoustic features in the adapted sound through speaker embedding, some parameters in the model are trained from acoustic features of different granularities to improve the generalization ability of the model during the training phase. A singer's acoustic model is first trained using a large batch of data. The singer's dimension trains the model of stress, rhythm, and time environment noise conditions on the phonemes to ensure that the singing timbre of a small batch of data can be inferred. The acoustic model models the acoustic conditions at the singing audio and phoneme levels respectively, which are used as the input of the Mel-spectrogram decoder so that the decoder can predict the singing timbre under different acoustic conditions based on these acoustic information. The different granularities mentioned above are represented as follows: singing level, which is the fine-grained acoustic conditions presented by the speaker in each sentence; phoneme level, which is the fine-grained acoustic conditions of each phoneme in a sentence, which needs to be established through the speak embedding module; Obtain the hidden sequence of phonemes. The duration and pitch of each phoneme are hidden in the hidden sequence of phonemes. The phoneme sequence is expanded according to the duration of each phoneme. The pitch element also forms a sequence corresponding to the expanded phoneme sequence. After that, the integrated output is the sequence feature aligned with the phoneme feature sequence. The acoustic feature is pre-aligned with the phoneme sequence, and then the acoustic features corresponding to the phoneme are averaged to facilitate conversion into the corresponding spectrogram. The model is basically built on the Transformer structure. There is a self-attention network and a feedforward network in each Transformer block. After applying normalization to the self-attention network and the feedforward network in the encoder, the learnable scaleγ and biasβ can effectively affect the hidden activation and the final prediction result. The small conditional network determines the scale and bias vector in the layer normalization according to the corresponding speaker characteristics, and fine-tunes this conditional network. The conditional network consists of two simple linear layers and Composition, E s For the speaker embedding module, fine-tune two matrices and Normalization is done at each conditional layer of the decoder and singer embeddings according to the following formula: Calculate each scale to get scaleγ and biasβ, and use a small conditional network to determine the normalized scale and bias vectors, and input the acoustic features of the corresponding speaker. Only two simple linear layers are used, the input is speakerembedding, and the output is predicted γ and β. By changing the parameters of the normalization operation in the decoder, the model is indirectly adjusted to achieve the purpose of customizing personalized singing with small batches of data. The learnable scaleγ and biasβ can effectively affect the hidden activation and the final prediction results.
2. A singing synthesis device for generating personalized timbre, characterized in that: include: Memory for storing computer programs; A processor is used to implement the steps of a singing synthesis method for generating personalized timbre as described in claim 1 when executing the computer program.
3. A readable storage medium, characterized in that: The readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the singing synthesis method for generating personalized timbre as claimed in claim 1 are implemented.
Citation Information
Patent Citations
Audio synthesis method and device, computer equipment and storage medium
CN114360492A