Audio synthesis method, device, computer-readable storage medium and electronic device

By obtaining text and music score features in audio synthesis, performing time prediction and acoustic encoding processing, and using a decoding network with hierarchical progressive training, the problem of low stability and naturalness extraction of singing parameters in the prior art is solved, and a synthetic song sound audio with high naturalness and stability is achieved.

CN113838443BActive Publication Date: 2025-06-24TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202110815643.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-07-19
Publication Date
2025-06-24
Estimated Expiration
2041-07-19

AI Technical Summary

Technical Problem

In the audio synthesis, the stability and naturalness of extracting singing parameters are low, resulting in poor naturalness of the synthesized singing, making it difficult to take into account both pronunciation stability and expressiveness.

Method used

By obtaining the text features and music score features of the target lyrics, the duration prediction process is performed, the acoustic encoding is generated, and the acoustic encoding is gradually decoded using a hierarchical progressive training decoding network to obtain a high natural target Mel spectrum, and finally generate a synthetic song sound audio.

Benefits of technology

It effectively improves the naturalness of the sound audio of the synthetic song, while taking into account pronunciation stability and expressiveness, reducing singing noise, ensuring stable pronunciation and excellent expressiveness.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN113838443B_ABST
    Figure CN113838443B_ABST
Patent Text Reader

Abstract

The present application discloses an audio synthesis method, apparatus, computer-readable storage medium, and electronic device, which relate to the field of artificial intelligence. The method includes: obtaining text features of target lyrics and score features of a target score; performing duration prediction processing based on the text features and score features to obtain predicted phoneme durations corresponding to each phoneme in the target lyrics; performing acoustic encoding processing on the text features and the score features according to the predicted phoneme durations to generate an acoustic encoding; using at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain a target Mel spectrogram; and generating a synthesized singing audio corresponding to the target lyrics and the target score based on the target Mel spectrogram. The present application effectively improves the naturalness of the synthesized singing audio, while taking into account pronunciation stability and expressiveness.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and particularly relates to an audio synthesis method, apparatus, computer-readable storage medium, and electronic device. Background Art

[0002] Audio synthesis is a technology that converts lyrics and sheet music into singing audio. With the development of the demand for smart life, audio synthesis plays an extremely important role in promoting the implementation of many smart life scenarios.

[0003] Currently, in the related art, there is a solution to extract singing parameters from lyrics and sheet music and synthesize singing audio based on the singing parameters. In the related art, the stability of the process of extracting singing parameters is relatively low, the naturalness of the extracted singing parameters is relatively low, and more singing parameters that cause noise are extracted, resulting in poor naturalness of the synthesized singing and it is difficult to balance pronunciation stability and expressiveness. Summary of the Invention

[0004] Embodiments of this application provide an audio synthesis method and related apparatus, which can effectively improve the naturalness of the synthesized singing audio and at the same time balance pronunciation stability and expressiveness.

[0005] To solve the above technical problems, the embodiments of this application provide the following technical solutions:

[0006] According to an embodiment of this application, an audio synthesis method includes: obtaining the text features of the target lyrics and the score features of the target score; performing duration prediction processing based on the text features and score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics; performing acoustic encoding processing on the text features and score features according to the predicted phoneme duration to generate an acoustic encoding; using at least two cascaded decoding networks trained in a hierarchical progressive manner to perform progressive decoding processing on the acoustic encoding to obtain a target Mel spectrogram; generating the synthesized singing audio corresponding to the target lyrics and the target score based on the target Mel spectrogram.

[0007] According to an embodiment of this application, an audio synthesis apparatus includes: an obtaining module, configured to obtain the text features of the target lyrics and the score features of the target score; a duration prediction module, configured to perform duration prediction processing based on the text features and score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics; an acoustic encoding module, configured to perform acoustic encoding processing on the text features and score features according to the predicted phoneme duration to generate an acoustic encoding; a cascaded decoding module, configured to use at least two cascaded decoding networks trained in a hierarchical progressive manner to perform progressive decoding processing on the acoustic encoding to obtain a target Mel spectrogram; a synthesis module, configured to generate the synthesized singing audio corresponding to the target lyrics and the target score based on the target Mel spectrogram.

[0008] In some embodiments of the present application, the obtaining module includes: a first conversion unit configured to perform feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features; a second conversion unit configured to perform feature conversion processing on the notes, note durations, and slurs in the target musical score to generate the musical score features.

[0009] In some embodiments of the present application, the first conversion unit is configured to: perform feature conversion processing on each phoneme and the phoneme type information corresponding to each phoneme in the target lyrics to generate the phoneme features of each phoneme; determine, from all the phonemes in the target lyrics, the phonemes to be prolonged with the phoneme types of finals and single finals; perform prolongation processing on the phoneme features of the phonemes to be prolonged to obtain the prolonged phoneme features; and generate the text features based on the prolonged phoneme features and the phoneme features of each phoneme.

[0010] In some embodiments of the present application, the second conversion unit is configured to: determine the notes and the number of phonemes corresponding to each syllable in the target lyrics; determine the syllable duration of each syllable according to the note durations of the notes corresponding to each syllable; evenly distribute the syllable duration of each syllable according to the number of phonemes corresponding to each syllable to obtain the phoneme duration of each phoneme in the target lyrics; and perform feature conversion processing on the notes, the phoneme duration of each phoneme, and the slurs in the target musical score to generate the musical score features.

[0011] In some embodiments of the present application, the duration prediction module includes: a first alignment unit configured to align the features in the text features and the musical score features in the order of the phonemes in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; a bidirectional encoding unit configured to perform bidirectional long short-term memory encoding processing on the song-lyric features corresponding to each phoneme in the target lyrics to obtain the long short-term memory features; and a duration prediction unit configured to perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme duration of each phoneme in the target lyrics.

[0012] In some embodiments of the present application, the duration prediction unit is configured to: use the trained duration model to perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme duration of each phoneme in the target text; the apparatus further includes a first training unit configured to: jointly train a preset duration model using a first objective function for predicting phoneme duration and a second objective function for predicting syllable duration to obtain the trained duration model.

[0013] In some embodiments of the present application, the acoustic encoding module includes: a second corresponding unit configured to align the features in the text feature and the musical score feature according to the phoneme order in the target lyrics to obtain the song and lyrics features corresponding to each phoneme in the target lyrics; a self-attention encoding unit configured to perform self-attention encoding processing based on the song and lyrics features corresponding to each phoneme in the target lyrics to obtain a self-attention encoding; and an expansion processing unit configured to perform expansion processing on the self-attention encoding according to the predicted phoneme duration to obtain the acoustic encoding.

[0014] In some embodiments of the present application, the self-attention encoding includes a sub-attention encoding corresponding to each phoneme in the target lyrics; the expansion processing unit includes: a replication subunit configured to perform feature replication processing on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme to obtain a replication encoding corresponding to each phoneme; and a combination subunit configured to generate the acoustic encoding based on the sub-attention encoding and the replication encoding corresponding to each phoneme.

[0015] In some embodiments of the present application, the replication subunit is configured to: determine the syllable duration of each syllable in the target lyrics and the phonemes corresponding to each syllable; perform scaling processing on the predicted phoneme duration of the phonemes corresponding to each syllable according to the syllable duration of each syllable to obtain a scaled phoneme duration corresponding to each phoneme, wherein the scaled phoneme duration corresponding to a phoneme of the initial consonant phoneme type is less than a predetermined duration; and perform feature replication processing on the sub-attention encoding corresponding to each phoneme based on the scaled phoneme duration corresponding to each phoneme.

[0016] In some embodiments of the present application, the cascaded decoding module includes: a self-attention cascaded decoding unit configured to perform self-attention decoding processing on the acoustic encoding in sequence by using at least two cascaded self-attention decoding networks with hierarchical progressive training to obtain a decoded Mel spectrogram; and a target Mel spectrogram generation unit configured to generate the target Mel spectrogram based on the decoded Mel spectrogram.

[0017] In some embodiments of the present application, the target Mel spectrogram generation unit is configured to: perform convolution processing on the input feature sequence corresponding to the decoded Mel spectrogram to obtain a convolution feature sequence; perform a full connection operation processing after splicing the input feature sequence and the convolution feature sequence to obtain a full connection feature sequence; and perform bidirectional recursive feature extraction processing on the full connection feature sequence to obtain a smoothed spectrogram feature sequence, so as to generate the target Mel spectrogram.

[0018] In some embodiments of the present application, the device further includes a second training unit, configured to: add a loss function for predicting the target Mel spectrogram to each self-attention decoding network in at least two cascaded self-attention decoding networks; and perform hierarchical progressive training on the at least two cascaded self-attention decoding networks based on the added loss function.

[0019] According to another embodiment of the present application, a computer-readable storage medium stores a computer program, which, when executed by a processor of a computer, causes the computer to execute the method described in the embodiments of the present application.

[0020] According to another embodiment of the present application, an electronic device includes: a memory storing a computer program; and a processor reading the computer program stored in the memory to execute the method described in the embodiments of the present application.

[0021] According to another embodiment of the present application, a computer program product or a computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, causing the computer device to execute the methods provided in various alternative implementations described in the embodiments of the present application.

[0022] In the embodiments of the present application, text features of the target lyrics and score features of the target score are obtained; duration prediction processing is performed based on the text features and the score features to obtain predicted phoneme durations corresponding to each phoneme in the target lyrics; acoustic encoding processing is performed on the text features and the score features according to the predicted phoneme durations to generate an acoustic encoding; an at least two-cascaded decoding network with hierarchical progressive training is used to perform progressive decoding processing on the acoustic encoding to obtain a target Mel spectrogram; and a synthetic singing audio corresponding to the target lyrics and the target score is generated based on the target Mel spectrogram.

[0023] In this way, during audio synthesis, independent duration prediction with respect to the acoustic encoding is performed based on the text features and the score features, and high-accuracy predicted phoneme durations can be stably obtained. Furthermore, acoustic encoding processing performed on the text features and the score features according to the predicted phoneme durations can obtain an acoustic encoding that can highly accurately represent the rhythm of the singing voice. Further, by using an at least two-cascaded decoding network with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding, a target Mel spectrogram with high naturalness can be obtained. Furthermore, a synthetic singing audio with high naturalness can be generated only through the target Mel spectrogram, which is a singing voice parameter. The singing voice noise is effectively reduced, and at the same time, the singing voice pronunciation is stable and has excellent expressiveness. Thus, the naturalness of the synthetic singing audio is effectively improved, while taking into account pronunciation stability and expressiveness. Description of the Drawings

[0024] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those skilled in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0025] Figure 1 The schematic diagram of the system 100 to which the embodiments of the present application can be applied is shown.

[0026] Figure 2 The schematic diagram of another system to which the embodiments of the present application can be applied is shown.

[0027] Figure 3 The flowchart of the audio synthesis method according to an embodiment of the present application is shown.

[0028] Figure 4 The flowchart of the feature acquisition method according to an embodiment of the present application is shown.

[0029] Figure 5 The flowchart of predicting the phoneme duration according to an embodiment of the present application is shown.

[0030] Figure 6 The flowchart of the acoustic coding process according to an embodiment of the present application is shown.

[0031] Figure 7 The architecture diagram of an audio synthesis system applying the embodiments of the present application is shown.

[0032] Figure 8 The block diagram of the audio synthesis device according to an embodiment of the present application is shown.

[0033] Figure 9 The block diagram of the electronic device according to an embodiment of the present application is shown. Detailed implementation manners

[0034] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative efforts belong to the scope of protection of the present application.

[0035] Figure 1 The schematic diagram of the system 100 to which the embodiments of the present application can be applied is shown. As Figure 1As shown, the system 100 may include a server 101 and a terminal 102. The server 101 and the terminal 102 may be directly or indirectly connected by wireless communication, and this application does not make special restrictions here.

[0036] The server 101 may be an independent physical server, or a server cluster or distributed system composed of multiple physical servers. It may also be a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms.

[0037] In one implementation of this example, the server 101 is a cloud server, and the server 101 may provide artificial intelligence cloud services, such as providing artificial intelligence cloud services for audio synthesis. The so-called artificial intelligence cloud service is generally also referred to as AIaaS (AI as a Service, Chinese for "AI as a Service"). This is a current mainstream service method for artificial intelligence platforms. Specifically, the AIaaS platform will split several common AI services and provide independent or packaged services in the cloud. This service model is similar to opening an AI-themed mall: all developers can access and use one or more artificial intelligence services provided by the platform through the API interface, and some senior developers can also use the AI framework and AI infrastructure provided by the platform to deploy and operate their own exclusive cloud artificial intelligence services.

[0038] The terminal 102 may be any device, including but not limited to mobile phones, computers, intelligent voice interaction devices, smart home appliances, vehicle terminals, VR / AR devices, smart watches, and computers, etc.

[0039] In one implementation of this example, the server 101 may obtain the text features of the target lyrics and the music score features of the target music score; perform duration prediction processing based on the text features and the music score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics; perform acoustic encoding processing on the text features and the music score features according to the predicted phoneme duration to generate an acoustic encoding; use at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain the target Mel spectrogram; generate the synthesized song audio corresponding to the target lyrics and the target music score based on the target Mel spectrogram.

[0040] Among them, the server 101 may obtain the music score information including the target lyrics and the target music score from the terminal 102.

[0041] Figure 2 Shows a schematic diagram of another system 200 to which the embodiments of the present application can be applied. As Figure 2As shown, the system 200 can be a distributed system formed by connecting a client 201 and multiple nodes 202 in the form of network communication.

[0042] Taking the distributed system as a blockchain system as an example, refer to Figure 2 , Figure 2 which is an optional schematic structural diagram of the distributed system 200 provided by the embodiments of the present application applied to a blockchain system. It is formed by multiple nodes 202 and a client 201, and a peer-to-peer (P2P) network is formed among the nodes. The P2P protocol is an application layer protocol running on top of the Transmission Control Protocol (TCP). In a distributed system, any machine such as a server can join and become a node 202 (each node 202 can be, for example, Figure 1 the server 101 in

[0043] Refer to Figure 2 the functions of each node in the blockchain system shown, and the functions involved include:

[0044] 1) Routing, which is a basic function of a node and is used to support communication between nodes.

[0045] In addition to the routing function, a node can also have the following functions:

[0046] 2) Application, which is used to be deployed in the blockchain, implement specific services according to actual business requirements, record the data related to the implemented functions to form record data, carry a digital signature in the record data to indicate the source of the task data, and send the record data to other nodes in the blockchain system. When other nodes verify the source and integrity of the record data successfully, the record data is added to a temporary block.

[0047] For example, the services implemented by the application include:

[0048] 2.1) Wallet, which is used to provide the function of conducting electronic currency transactions, including initiating a transaction (that is, sending the transaction record of the current transaction to other nodes in the blockchain system. After other nodes verify successfully, as a response to acknowledging the validity of the transaction, the record data of the transaction is deposited into the temporary block of the blockchain; of course, the wallet also supports querying the remaining electronic currency in the electronic currency address;

[0049] 2.2) Shared ledger, which is used to provide functions such as storage, query, and modification of account data, sends the recorded data of operations on the account data to other nodes in the blockchain system. After other nodes verify its validity, as a response to acknowledging the validity of the account data, the recorded data is stored in a temporary block, and can also send a confirmation to the node that initiated the operation.

[0050] 2.3) Smart contract, a computerized protocol that can execute the terms of a certain contract, implemented by code deployed on the shared ledger for execution under certain conditions. According to actual business requirements, the code is used to complete automated transactions, such as querying the logistics status of the goods purchased by the buyer and transferring the buyer's electronic currency to the merchant's address after the buyer signs for the goods; of course, smart contracts are not limited to executing contracts for transactions, but can also execute contracts for processing received information.

[0051] 3) Blockchain, including a series of blocks (Block) that are sequentially connected in the order of generation. Once a new block is added to the blockchain, it will not be removed again. The block records the recorded data submitted by nodes in the blockchain system.

[0052] In an implementation of this example, node 202 can obtain the text features of the target lyrics and the score features of the target score; perform duration prediction processing based on the text features and score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics; perform acoustic encoding processing on the text features and score features according to the predicted phoneme duration to generate acoustic encoding; use at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain the target Mel spectrogram; generate the synthesized audio corresponding to the target lyrics and target score based on the target Mel spectrogram.

[0053] Among them, the client 201 can send the score information including the target lyrics and target score to the node 202.

[0054] Figure 3 Schematically shows a flowchart of an audio synthesis method according to an embodiment of the present application. The execution subject of this audio synthesis method can be any terminal, such as Figure 1 the shown server 101, terminal 102 or the terminal corresponding to the node 202 or client 201 as shown in Figure 2 the figure.

[0055] As shown in Figure 3 the figure, this audio synthesis method may include steps S310 to S350.

[0056] Step S310, obtaining the text features of the target lyrics and the score features of the target score;

[0057] Step S320: Perform duration prediction processing based on text features and musical score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics.

[0058] Step S330: Perform acoustic encoding processing on the text features and musical score features according to the predicted phoneme duration to generate an acoustic encoding.

[0059] Step S340: Use at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain the target Mel spectrogram.

[0060] Step S350: Generate the synthesized song audio corresponding to the target lyrics and the target musical score based on the target Mel spectrogram.

[0061] The following describes the specific processes of the steps performed during audio synthesis.

[0062] In step S310, obtain the text features of the target lyrics and the musical score features of the target musical score.

[0063] In the implementation manner of this example, the text features are the features of the lyric information in the song, such as the phoneme sequence after conversion of the target lyrics. The musical score features are the features of the melody information in the song, such as note, note duration, beat, legato, sustain, etc. The target lyrics can be the lyrics corresponding to the synthesized song audio, the target musical score can be the musical score corresponding to the synthesized song audio, and the target lyrics and the target text can be included in the musical score corresponding to the synthesized song audio.

[0064] For the target lyrics, phoneme information (such as phoneme type, etc.) can be extracted according to the pronunciation of each character in the target lyrics, and then the phoneme information is subjected to feature conversion to obtain text features; for the target musical score, musical score information such as the notes, note durations, and legato lines marked for each character in the target lyrics can be extracted from the target musical score, and then the musical score information is subjected to feature conversion to obtain musical score features.

[0065] In one embodiment, refer to Figure 4 , step S310, obtaining the text features of the target lyrics and the musical score features of the target musical score, includes:

[0066] Step S410: Perform feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate text features; step S420: Perform feature conversion processing on the notes, note durations, and legato lines in the target musical score to generate musical score features.

[0067] A phoneme is the smallest unit of speech, such as h, etc. Phoneme type information may include initial consonants, final consonants, etc. In one example, phoneme type information includes initial consonants, final consonants, and simple finals. A musical note is a musical symbol, which may include symbols commonly used in musical scores. Musical notes can express different characteristics of sounds, such as pitch, volume, etc. A slur (i.e., a legato line) represents performance information. A slur can connect several musical notes with different pitches together, indicating that these musical notes should be played smoothly and continuously. The duration of a musical note is the playing duration of the musical note.

[0068] Feature conversion processing can convert each phoneme, musical note, etc. into a corresponding unique feature vector by looking up a vector dictionary, and then generate text features and musical score features.

[0069] By performing feature conversion on the phonemes and phoneme type information in the target lyrics to construct text features, and performing feature conversion on the musical notes, note durations, and slurs in the target musical score to construct musical score features, the musical score features (i.e., musical features) and text features (i.e., Chinese pinyin features) can be finely constructed, further improving the pronunciation stability and singing tone stability of the synthesized singing voice audio. Among them, using slur information and phoneme type information can further improve the pronunciation stability and singing tone stability of the synthesized singing voice audio compared with the case of not using them.

[0070] In one embodiment, step S410 of performing feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features includes:

[0071] Performing feature conversion processing on each phoneme in the target lyrics and the phoneme type information corresponding to each phoneme to generate the phoneme features of each phoneme; determining the phonemes to be sustained with the phoneme types of finals and simple finals from all the phonemes in the target lyrics; performing a sustaining process on the phoneme features of the phonemes to be sustained to obtain sustained phoneme features; and generating text features based on the sustained phoneme features and the phoneme features of each phoneme.

[0072] Performing feature conversion processing on each phoneme in the target lyrics and the phoneme type information corresponding to each phoneme. For example, performing feature conversion on the phoneme A and the phoneme type information of phoneme A to obtain the feature vector A1 corresponding to phoneme A and the feature vector A2 corresponding to the phoneme type information of phoneme A. A1 and A2 are the phoneme features of phoneme A.

[0073] The phonemes to be sustained can be determined according to the sustaining marks marked for the musical notes corresponding to the characters in the target musical score. The sustaining mark, such as a slur, also called an isophone tie (often simply called a tie), is a connecting arc added between adjacent musical notes with the same pitch. If a sustaining mark is marked for the musical note of a certain character, the phoneme corresponding to this character is the phoneme to be sustained, and it needs to be sustained according to the sustaining duration corresponding to the sustaining mark.

[0074] Perform a sustain process on the phoneme features of the sustain phoneme, that is, copy the phoneme features of the sustain phoneme according to the sustain duration corresponding to the sustain mark to obtain the copied sustain phoneme features. Finally, arrange the sustain phoneme features and the phoneme features of each phoneme in the order of phonemes in the target lyrics (that is, the phoneme features of the sustain phoneme can include the initial phoneme features and the sustain phoneme features) to obtain text features.

[0075] In this way, by only performing sustain on the finals and single finals and not on the initials, the applicant finds that the obtained text features can further improve the pronunciation stability and singing tone stability of the synthesized singing audio.

[0076] In one embodiment, in step S420, perform feature conversion processing on the notes, note durations, and slurs in the target musical score to generate musical score features, including:

[0077] Determine the notes and the number of phonemes corresponding to each syllable in the target lyrics; determine the syllable duration of each syllable according to the note durations of the notes corresponding to each syllable; evenly distribute the syllable duration of each syllable according to the number of phonemes corresponding to each syllable to obtain the phoneme duration of each phoneme in the target lyrics; perform feature conversion processing on the notes, the phoneme duration of each phoneme, and the slurs in the target musical score to generate musical score features.

[0078] Each syllable can correspond to one character. The syllable duration of each syllable is the sum of the note durations of the notes corresponding to each syllable. Each syllable can correspond to multiple phonemes. For each syllable, divide the syllable duration of the syllable by the number of phonemes corresponding to the syllable for even distribution to obtain the phoneme duration of each phoneme corresponding to the syllable, and further obtain the phoneme duration of each phoneme in the target lyrics.

[0079] Finally, perform feature conversion processing on the notes, the phoneme duration of each phoneme, and the slurs in the target musical score. For example, perform feature conversion on the notes, phoneme duration, and slur corresponding to phoneme A to obtain the feature vector A3 corresponding to the note, the feature vector A4 corresponding to the phoneme duration, and the feature vector A5 corresponding to the slur. A3, A4, and A5 are the melody features of phoneme A, and the set of all melody features is the musical score features.

[0080] In this way, when constructing the musical score features, through the method of even distribution, the musical score features can further improve the pronunciation stability and singing tone stability of the synthesized singing audio.

[0081] In step S320, perform duration prediction processing based on the text features and the musical score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics.

[0082] In the implementation of this example, the predicted phoneme duration is the reference singing duration during singing predicted for each phoneme. Feature encoding is performed on the text features and musical score features to obtain the encoded features. Duration prediction can be performed based on the encoded features to obtain the predicted phoneme duration corresponding to each phoneme. Independent duration prediction processing is performed based on the text features and musical score features, and the method of sharing features and sharing encoders with subsequent acoustic encoding is not adopted, so that the predicted phoneme duration with high accuracy can be stably obtained. In one example, an independent pre-trained duration model can be used to perform duration prediction processing based on the text features and musical score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics. The duration model is, for example, Figure 7 the bidirectional long short-term memory (BLSTM) model shown.

[0083] In one embodiment, refer to Figure 5 , step S320, perform duration prediction processing based on the text features and musical score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics, including:

[0084] Step S510, align the features in the text features and musical score features in the order of phonemes in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; step S520, perform bidirectional long short-term memory encoding processing on the song-lyric features corresponding to each phoneme in the target lyrics to obtain long short-term memory features; step S530, perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics.

[0085] The text features may include phoneme features corresponding to each phoneme, and the musical score features may include melody features corresponding to each phoneme. Aligning the phoneme features and melody features corresponding to the phoneme gives the song-lyric features corresponding to the phoneme.

[0086] Performing bidirectional long short-term memory encoding processing on the sequence of song-lyric features corresponding to all phonemes to obtain long short-term memory features, and then, performing classification prediction based on the long short-term memory features, for example, performing a fully connected operation based on the long short-term memory features and then performing classification, to obtain the predicted phoneme duration and confidence corresponding to each phoneme, and the predicted phoneme duration with high accuracy can be stably obtained.

[0087] In one example, the song-lyric features corresponding to each phoneme in the target lyrics can be input into a pre-trained independent bidirectional long short-term memory (BLSTM) model, and the bidirectional long short-term memory network in the bidirectional long short-term memory model is used to perform bidirectional long short-term memory encoding processing on the song-lyric features corresponding to each phoneme in the target lyrics to obtain long short-term memory features; a fully connected network and a classifier are used to perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics.

[0088] In one embodiment, in step S530, duration prediction processing is performed based on long short-term memory features to obtain the predicted phoneme duration of each phoneme in the target text, including:

[0089] Using the trained duration model, duration prediction processing is performed based on long short-term memory features to obtain the predicted phoneme duration of each phoneme in the target text; and the training of the duration model includes: using the first objective function for predicting phoneme duration and the second objective function for predicting syllable duration to jointly train the preset duration model to obtain the trained duration model.

[0090] The trained duration model is, for example Figure 7 the pre-trained independent bidirectional long short-term memory (BLSTM) model 701 shown in the figure. Using the trained duration model, the song lyrics features corresponding to each phoneme in the target lyrics can be encoded by bidirectional long short-term memory to obtain long short-term memory features; duration prediction processing is performed based on the long short-term memory features to obtain the predicted phoneme duration of each phoneme in the target lyrics.

[0091] Among them, referring to Figure 7 , when training the duration model: introducing the first objective function 703 for predicting phoneme duration and the second objective function 702 for predicting syllable duration to jointly train the preset duration model. The first objective function and the second objective function can be the mean squared error function (MSE: Mean Squared Error).

[0092] Specifically, the target phoneme duration of each phoneme in the training sample can be predicted using the duration model. Then, based on concatenating the target phoneme durations of the phonemes corresponding to each syllable in the training sample, the target syllable duration of the syllable is obtained. Then, based on the first objective function, the first prediction error (such as the squared difference) between the predicted target phoneme duration of the duration model and the true phoneme duration calibrated by the phoneme in the training sample is determined. Based on the second objective function, the second prediction error (such as the squared difference) between the predicted target syllable duration of the duration model and the true syllable duration calibrated by the syllable in the training sample is determined. The first prediction error and the second prediction error are simultaneously used as the optimization objectives of the duration model to adjust the parameters in the duration model until the first prediction error and the second prediction error both meet the requirements (such as being less than a predetermined threshold) to obtain the trained duration model.

[0093] In this way, the first objective function for predicting phoneme duration and the second objective function for predicting syllable duration are introduced simultaneously to jointly train the preset duration model. The predicted phoneme duration is obtained by using the trained duration model. In subsequent steps, the acoustic encoding generated by acoustically encoding the text features and musical score features based on the predicted phoneme duration will have rhythm transformations at different scales based on the phoneme level and syllable level, which can effectively enhance the coherence of syllables or entire sentences in the singing voice and further improve the naturalness of the singing rhythm.

[0094] In step S330, the text features and musical score features are acoustically encoded according to the predicted phoneme duration to generate acoustic encoding.

[0095] In the implementation manner of this example, feature acoustic encoding is performed based on the text features and musical score features, and rhythm expansion processing is performed on the feature encoding (such as self-attention encoding) output after feature acoustic encoding based on the predicted phoneme duration to obtain acoustic encoding with a complete rhythm. Furthermore, the acoustic encoding obtained by acoustically encoding the text features and musical score features according to the predicted phoneme duration can accurately represent the singing rhythm.

[0096] Among them, a pre-trained acoustic model can be used to acoustically encode the text features and musical score features according to the predicted phoneme duration to generate acoustic encoding. Specifically, it can be based on the self-attention mechanism encoder 704 (SA Encoder) in the acoustic model as shown in Figure 7 to perform feature acoustic encoding processing, and perform expansion processing based on the upsampling network 705 (upsampling) in the acoustic model.

[0097] In one embodiment, referring to Figure 6 , step S330, acoustically encoding the text features and musical score features according to the predicted phoneme duration to generate acoustic encoding includes:

[0098] Step S610, aligning the features in the text features and musical score features in the order of phonemes in the target lyrics to obtain the music and lyrics features corresponding to each phoneme in the target lyrics; step S620, performing self-attention encoding processing based on the music and lyrics features corresponding to each phoneme in the target lyrics to obtain self-attention encoding; step S630, performing expansion processing on the self-attention encoding according to the predicted phoneme duration to obtain acoustic encoding.

[0099] The text features may include phoneme features corresponding to each phoneme, and the musical score features may include melody features corresponding to each phoneme. Aligning the phoneme features and melody features corresponding to the phoneme gives the music and lyrics features corresponding to the phoneme.

[0100] Self-attention encoding processing is performed on the song lyrics feature sequences corresponding to all phonemes to obtain self-attention encoding. The self-attention encoding may include sub-attention encoding corresponding to each phoneme. Then, the sub-attention encoding corresponding to each phoneme in the self-attention encoding can be extended according to the predicted phoneme duration to obtain the acoustic encoding of the complete rhythm.

[0101] Among them, the sub-attention encoding corresponding to each phoneme in the self-attention encoding is extended according to the predicted phoneme duration to obtain the sub-acoustic encoding corresponding to each phoneme. For example, in the extension process, the phoneme A corresponds to the sub-attention encoding a, and the predicted phoneme duration corresponding to the phoneme A is 3. At this time, the sub-attention encoding a can be extended into 3 copies to obtain aaa, and aaa is the sub-acoustic encoding corresponding to the phoneme A. Further, the number of extended copies can be adjusted according to a predetermined coefficient. For example, if the predetermined coefficient is 2, then the duration is 2 * 3 = 6, and the phoneme A is extended into 6 copies to obtain aaaaaa, and aaaaaa is the sub-acoustic encoding corresponding to the phoneme A. Finally, the set of sub-acoustic encodings of all phonemes is the obtained acoustic encoding.

[0102] Among them, referring to Figure 7 , the self-attention mechanism encoder 704 (SAEncoder) in the pre-trained acoustic model can be used to perform self-attention encoding processing based on the song lyrics features corresponding to each phoneme in the target lyrics to obtain self-attention encoding; then, the upsampling network 705 (upsampling) in the acoustic model is used to extend the self-attention encoding according to the predicted phoneme duration to obtain the acoustic encoding.

[0103] In one embodiment, the self-attention encoding includes sub-attention encoding corresponding to each phoneme in the target lyrics; step S630, extending the self-attention encoding according to the predicted phoneme duration to obtain the acoustic encoding, including:

[0104] Performing feature replication processing on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme to obtain the replication encoding corresponding to each phoneme; generating the acoustic encoding based on the sub-attention encoding and the replication encoding corresponding to each phoneme.

[0105] The extended processing is implemented through feature replication processing. For example, phoneme A corresponds to sub-attention encoding a, and the predicted phoneme duration corresponding to phoneme A is 3. At this time, the sub-attention encoding a can be replicated 2 times to obtain replicated encoding aa, and the set aaa of the sub-attention encoding and the replicated encoding is the sub-acoustic encoding corresponding to phoneme A. Further, the number of replications can be adjusted according to a predetermined coefficient. For example, if the predetermined coefficient is 2, then the duration is 2 * 3 = 6, and phoneme A is replicated 5 times to obtain aaaaa, and the set aaaaaa of the sub-attention encoding and the replicated encoding is the sub-acoustic encoding corresponding to phoneme A. Finally, the set of sub-acoustic encodings of all phonemes is the obtained acoustic encoding.

[0106] In one embodiment, feature replication processing is performed on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme, including:

[0107] Determine the syllable duration of each syllable in the target lyrics and the phoneme corresponding to each syllable; perform scaling processing on the predicted phoneme duration of the phoneme corresponding to each syllable according to the syllable duration of each syllable to obtain the scaled phoneme duration corresponding to each phoneme, where the scaled phoneme duration corresponding to a phoneme with the phoneme type of initial consonant is less than the predetermined duration; perform feature replication processing on the sub-attention encoding corresponding to each phoneme based on the scaled phoneme duration corresponding to each phoneme.

[0108] The syllable duration of each syllable is the syllable duration marked for each syllable in the target musical score. Perform scaling processing on the predicted phoneme duration of the phoneme corresponding to each syllable according to the syllable duration of each syllable. For example, the syllable duration corresponding to syllable M is 200, the phonemes corresponding to syllable M are, for example, A and B, the predicted phoneme duration corresponding to A is 70, and the predicted phoneme duration corresponding to B is 120. At this time, 200 / (70 + 120) = 1.05. Then, 1.05 can be used as the scaling ratio (where the scaling ratio can be adjusted according to the actual situation), and the predicted phoneme duration 70 corresponding to A can be scaled to 70 * 1.05 = 73.5 (i.e., the scaled phoneme duration), and the predicted phoneme duration 120 corresponding to B can be scaled to 120 * 1.05 = 126 (i.e., the scaled phoneme duration). Among them, if the phoneme type of B is an initial consonant and the predetermined duration is 125, then the scaled phoneme duration corresponding to B is restricted to the predetermined duration 125.

[0109] Furthermore, performing feature replication processing on the sub-attention encoding corresponding to each phoneme based on the scaled phoneme duration corresponding to each phoneme can make the synthesized singing voice be precisely aligned with the corresponding accompaniment. At the same time, in order to prevent the duration of the initial consonant from being too long after scaling, the maximum scaled phoneme duration of the initial consonant is restricted, for example, restricted to 125 milliseconds, so that the initial consonant can sound more natural on the notes with a longer duration.

[0110] In step S340, at least two cascaded decoding networks with hierarchical progressive training are used to perform progressive decoding processing on the acoustic encoding to obtain the target Mel spectrogram.

[0111] In the implementation of this example, hierarchical progressive training means adding a loss function for predicting the target Mel spectrogram to each layer of the at least two cascaded decoding networks (that is, adding a loss function between each layer of the decoding network and the true Mel spectrogram of the training samples), forming an iterative loss function for the decoding network, and enabling hierarchical progressive training of the at least two cascaded decoding networks to obtain at least two cascaded decoding networks with hierarchical progressive training. The at least two cascaded decoding networks have a fast convergence speed and can decode Mel spectrograms with effectively improved naturalness.

[0112] The loss function added to each layer of the decoding network can be the mean absolute error function (MAE: Mean Absolute Error), making the absolute value of the error between the Mel spectrogram output by each layer of the decoding network and the true Mel spectrogram the optimization target. The smaller the absolute value of the error, the more accurate the decoded Mel spectrogram.

[0113] Among them, at least two cascaded decoding networks with hierarchical progressive training can be located in the acoustic model as the decoder in the acoustic model.

[0114] In one embodiment, step S340, using at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain the target Mel spectrogram, includes:

[0115] Using at least two cascaded self-attention decoding networks with hierarchical progressive training to perform self-attention decoding processing on the acoustic encoding in sequence to obtain the decoded Mel spectrogram; generating the target Mel spectrogram based on the decoded Mel spectrogram.

[0116] The decoding network uses a self-attention decoding network, which can perform self-attention decoding processing on the acoustic encoding based on the self-attention mechanism in sequence. The Mel spectrogram output by the last layer of the self-attention decoding network is the decoded Mel spectrogram, and the quality of the decoded Mel spectrogram is high. Among them, refer to Figure 7 , at least two cascaded self-attention decoding networks with hierarchical progressive training can be located in the acoustic model (Acoustic model), and the at least two cascaded self-attention decoding networks form the self-attention mechanism decoder 706 (SA Decoder) in the acoustic model. In one example, the self-attention decoding network includes 3 layers.

[0117] In one embodiment, generating the target Mel spectrogram based on the decoded Mel spectrogram includes:

[0118] Perform convolution processing on the input feature sequence corresponding to the decoded Mel spectrogram to obtain a convolutional feature sequence; perform a fully connected operation on the concatenation of the input feature sequence and the convolutional feature sequence to obtain a fully connected feature sequence; perform bidirectional recursive feature extraction on the fully connected feature sequence to obtain a smoothed spectral feature sequence, so as to generate the target Mel spectrogram.

[0119] After decoding the Mel spectrogram, perform smoothing processing on the decoded Mel spectrogram to generate the target Mel spectrogram, making the target Mel spectrogram smoother and of better quality. Among them, the smoothing process can be performed by a Mel spectrogram post-processing network. The post-processing network can be composed of a CBHG (Convolution Bank + Highway network + bidirectional Gated Recurrent Unit) module. This module can include a convolutional layer (Convolution Bank), a highway network, and a bidirectional recurrent neural network (bidirectional Gated Recurrent Unit). It is possible to perform convolution processing on the input feature sequence corresponding to the decoded Mel spectrogram based on the convolutional layer to obtain a convolutional feature sequence; after concatenating the input feature sequence and the convolutional feature sequence, perform a fully connected operation based on the highway network to obtain a fully connected feature sequence; finally, perform bidirectional recursive feature extraction on the fully connected feature sequence based on the bidirectional recurrent neural network to obtain a smoothed spectral feature sequence. The smoothed spectral feature sequence is the spectral feature sequence corresponding to the target Mel spectrogram. Refer to Figure 7 , the post-processing network 707 can be located in the acoustic model. The target Mel spectrogram output by the post-processing network is the final Mel spectrogram output by the acoustic model. When training the acoustic model, an output loss function 708 can be added between the target Mel spectrogram output for the training sample and the true Mel spectrogram (GT mel spectrogram) corresponding to the training sample for training.

[0120] In this way, a smoother and better-quality target Mel spectrogram can be obtained, with an effective improvement compared to the simple smoothing effect using a deep convolutional network.

[0121] In one embodiment, it further includes hierarchical progressive training of at least two cascaded self-attention decoding networks, including: adding a loss function for predicting the target Mel spectrogram to each self-attention decoding network in at least two cascaded self-attention decoding networks; based on the added loss function, performing hierarchical progressive training on at least two cascaded self-attention decoding networks.

[0122] The hierarchical progressive training is to add a loss function for predicting the target Mel spectrogram to each self-attention decoding network in at least two cascaded self-attention decoding networks (that is, to add a loss function between each self-attention decoding network and the true Mel spectrogram of the training samples), so as to form an iterative loss function 709 as shown in Figure 7 Figure 2, and hierarchical progressive training can be performed on at least two cascaded self-attention decoding networks to obtain at least two cascaded self-attention decoding networks with hierarchical progressive training. The at least two cascaded self-attention decoding networks have a fast convergence speed and can decode Mel spectrograms with effectively improved naturalness.

[0123] In step S350, a synthesized song audio corresponding to the target lyrics and the target music score is generated based on the target Mel spectrogram.

[0124] In the implementation manner of this example, based on the target Mel spectrogram with high naturalness obtained in the foregoing steps, the target Mel spectrogram can be directly converted into the singing waveform of the synthesized song audio (the target Mel spectrogram can be converted into the corresponding singing waveform based on a vocoder such as the MelGAN model). Furthermore, a synthesized song audio with high naturalness can be generated only through the singing parameter of the target Mel spectrogram, the singing noise is effectively reduced, and at the same time, the singing pronunciation is stable and has excellent expressiveness. Compared with the related art that requires obtaining more acoustic parameters such as the Mel spectrogram and the fundamental frequency (F0), the synthesized song audio either has no obvious improvement in expressiveness, or there are a small number of misjudgments of voiceless and voiced sounds or unclear pronunciation problems.

[0125] In this way, based on steps S310 to S350, during audio synthesis, independent duration prediction relative to acoustic coding can be performed based on text features and music score features, and high-accuracy predicted phoneme durations can be stably obtained. Furthermore, acoustic coding processing of the text features and music score features is performed according to the predicted phoneme durations to obtain acoustic coding that can highly accurately represent the singing rhythm. Further, progressive decoding processing is performed on the acoustic coding by using at least two cascaded decoding networks with hierarchical progressive training to obtain a target Mel spectrogram with high naturalness. Furthermore, a synthesized song audio with high naturalness can be generated only through the singing parameter of the target Mel spectrogram, the singing noise is effectively reduced, and at the same time, the singing pronunciation is stable and has excellent expressiveness. Thereby, the naturalness of the synthesized song audio is effectively improved, and at the same time, the pronunciation stability and expressiveness are taken into account.

[0126] To facilitate better implementation of the audio synthesis method provided in the embodiments of the present application, the embodiments of the present application further provide an audio synthesis device based on the above audio synthesis method. The meanings of the nouns are the same as those in the above audio synthesis method, and the specific implementation details can refer to the description in the method embodiments. Figure 8The block diagram of an audio synthesis device according to an embodiment of the present application is shown. Figure 8 The block diagram of an audio synthesis device according to another embodiment of the present application is shown.

[0127] As Figure 8 shown, the audio synthesis device 800 may include an acquisition module 810, a duration prediction module 820, an acoustic encoding module 830, a cascaded decoding module 840, and a synthesis module 850.

[0128] The acquisition module 810 may be used to acquire the text features of the target lyrics and the musical score features of the target musical score; the duration prediction module 820 may be used to perform duration prediction processing based on the text features and musical score features to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics; the acoustic encoding module 830 may be used to perform acoustic encoding processing on the text features and the musical score features according to the predicted phoneme duration to generate an acoustic encoding; the cascaded decoding module 840 may be used to perform progressive decoding processing on the acoustic encoding by using at least two cascaded decoding networks with hierarchical progressive training to obtain a target Mel spectrogram; the synthesis module 850 may be used to generate a synthesized song audio corresponding to the target lyrics and the target musical score based on the target Mel spectrogram.

[0129] In some embodiments of the present application, the acquisition module 810 includes: a first conversion unit, configured to perform feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features; a second conversion unit, configured to perform feature conversion processing on the notes, note durations, and ties in the target musical score to generate the musical score features.

[0130] In some embodiments of the present application, the first conversion unit is configured to: perform feature conversion processing on each phoneme and the phoneme type information corresponding to each phoneme in the target lyrics to generate the phoneme features of each phoneme; determine, from all the phonemes in the target lyrics, the phonemes to be prolonged with the phoneme types of finals and simple finals; perform prolongation processing on the phoneme features of the phonemes to be prolonged to obtain prolonged phoneme features; and generate the text features based on the prolonged phoneme features and the phoneme features of each phoneme.

[0131] In some embodiments of the present application, the second conversion unit is configured to: determine the notes and the number of phonemes corresponding to each syllable in the target lyrics; determine the syllable duration of each syllable according to the note durations of the notes corresponding to each syllable; evenly distribute the syllable duration of each syllable according to the number of phonemes corresponding to each syllable to obtain the phoneme duration of each phoneme in the target lyrics; and perform feature conversion processing on the notes in the target musical score, the phoneme duration of each phoneme, and the ties to generate the musical score features.

[0132] In some embodiments of the present application, the duration prediction module 820 includes: a first alignment unit configured to align the features in the text feature and the music score feature in the order of phonemes in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; a bidirectional encoding unit configured to perform bidirectional long short-term memory encoding on the song-lyric features corresponding to each phoneme in the target lyrics to obtain long short-term memory features; and a duration prediction unit configured to perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations of each phoneme in the target lyrics.

[0133] In some embodiments of the present application, the duration prediction unit is configured to: use a trained duration model to perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations of each phoneme in the target text; the apparatus further includes a first training unit configured to: jointly train a preset duration model using a first objective function for predicting phoneme durations and a second objective function for predicting syllable durations to obtain the trained duration model.

[0134] In some embodiments of the present application, the acoustic encoding module 830 includes: a second correspondence unit configured to align the features in the text feature and the music score feature in the order of phonemes in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; a self-attention encoding unit configured to perform self-attention encoding on the song-lyric features corresponding to each phoneme in the target lyrics to obtain self-attention encoding; and an expansion processing unit configured to perform expansion processing on the self-attention encoding according to the predicted phoneme durations to obtain the acoustic encoding.

[0135] In some embodiments of the present application, the self-attention encoding includes sub-attention encoding corresponding to each phoneme in the target lyrics; the expansion processing unit includes: a replication subunit configured to perform feature replication on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme to obtain replicated encoding corresponding to each phoneme; and a combination subunit configured to generate the acoustic encoding based on the sub-attention encoding and the replicated encoding corresponding to each phoneme.

[0136] In some embodiments of the present application, the replication subunit is configured to: determine the syllable duration of each syllable and the phonemes corresponding to each syllable in the target lyrics; scale the predicted phoneme duration of the phonemes corresponding to each syllable according to the syllable duration of each syllable to obtain the scaled phoneme duration corresponding to each phoneme, wherein the scaled phoneme duration corresponding to a phoneme of an initial consonant type is less than a preset duration; and perform feature replication on the sub-attention encoding corresponding to each phoneme based on the scaled phoneme duration corresponding to each phoneme.

[0137] In some embodiments of the present application, the cascaded decoding module 840 includes: a self-attention cascaded decoding unit, configured to perform self-attention decoding processing on the acoustic encoding in sequence by using at least two cascaded self-attention decoding networks with hierarchical progressive training to obtain the decoded Mel spectrogram; and a target Mel spectrogram generation unit, configured to generate the target Mel spectrogram based on the decoded Mel spectrogram.

[0138] In some embodiments of the present application, the target Mel spectrogram generation unit is configured to: perform convolution processing on the input feature sequence corresponding to the decoded Mel spectrogram to obtain a convolution feature sequence; perform a fully connected operation processing after concatenating the input feature sequence and the convolution feature sequence to obtain a fully connected feature sequence; and perform bidirectional recursive feature extraction processing on the fully connected feature sequence to obtain a smoothed spectral feature sequence, so as to generate the target Mel spectrogram.

[0139] In some embodiments of the present application, the device further includes a second training unit, configured to: add a loss function for predicting the target Mel spectrogram to each self-attention decoding network in at least two cascaded self-attention decoding networks; and perform hierarchical progressive training on the at least two cascaded self-attention decoding networks based on the added loss function.

[0140] In this way, based on the audio synthesis device 800, during audio synthesis, independent duration prediction can be performed on the acoustic encoding based on text features and musical score features, and high-accuracy predicted phoneme durations can be stably obtained. Furthermore, through the acoustic encoding processing of the text features and musical score features according to the predicted phoneme durations, an acoustic encoding that can highly accurately represent the singing rhythm can be obtained. Further, by performing progressive decoding processing on the acoustic encoding by using at least two cascaded decoding networks with hierarchical progressive training, a target Mel spectrogram with high naturalness can be obtained. Furthermore, a synthetic singing audio with high naturalness can be generated only through the target Mel spectrogram, the singing noise is effectively reduced, and at the same time, the singing pronunciation is stable and has excellent expressiveness. Thus, the naturalness of the synthetic singing audio is effectively improved, while taking into account pronunciation stability and expressiveness.

[0141] It should be noted that although several modules or units of a device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present application, the features and functions of the two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0142] In addition, an embodiment of the present application further provides an electronic device, which can be a terminal or a server, such as Figure 9As shown, it shows a schematic structural diagram of an electronic device involved in an embodiment of the present application. Specifically:

[0143] The electronic device may include a processor 901 with one or more processing cores, a memory 902 with one or more computer-readable storage media, a power supply 903, an input unit 904, and other components. Those skilled in the art can understand that Figure 9 the structural diagram of the electronic device shown in does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements. Among them:

[0144] The processor 901 is the control center of the electronic device. It uses various interfaces and circuits to connect all parts of the entire computer device. By running or executing software programs and / or modules stored in the memory 902, and by calling the data stored in the memory 902, it executes various functions of the computer device and processes data, thereby performing an overall detection of the electronic device. Optionally, the processor 901 may include one or more processing cores; preferably, the processor 901 may integrate an application processor and a modem processor. Among them, the application processor mainly processes the operating system, user interfaces, and application programs, etc., and the modem processor mainly processes wireless communication. It can be understood that the above-mentioned modem processor may not be integrated into the processor 901 either.

[0145] The memory 902 can be used to store software programs and modules. The processor 901 executes various functional applications and data processing by running the software programs and modules stored in the memory 902. The memory 902 may mainly include a program storage area and a data storage area. Among them, the program storage area can store the operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the computer device. In addition, the memory 902 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. Correspondingly, the memory 902 may also include a memory controller to provide the processor 901 with access to the memory 902.

[0146] The electronic device further includes a power supply 903 that powers each component. Preferably, the power supply 903 can be logically connected to the processor 901 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 903 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.

[0147] The electronic device may further include an input unit 904, which may be configured to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.

[0148] Although not shown, the electronic device may further include a display unit and the like, which will not be elaborated here. Specifically, in this embodiment, the processor 901 in the electronic device will load the executable files corresponding to the processes of one or more computer programs into the memory 902 according to the following instructions, and the processor 901 will run the computer programs stored in the memory 902 to implement various functions. For example, the processor 901 may execute:

[0149] Obtain the text features of the target lyrics and the music score features of the target music score; perform duration prediction processing based on the text features and music score features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics; perform acoustic encoding processing on the text features and the music score features according to the predicted phoneme durations to generate an acoustic encoding; use at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain a target Mel spectrogram; generate a synthesized song audio corresponding to the target lyrics and the target music score based on the target Mel spectrogram.

[0150] In one embodiment, when obtaining the text features of the target lyrics and the music score features of the target music score, the processor 901 may execute: perform feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features; perform feature conversion processing on the notes, note durations and ties in the target music score to generate the music score features.

[0151] In one embodiment, when performing feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features, the processor 901 may execute: perform feature conversion processing on each phoneme and the phoneme type information corresponding to each phoneme in the target lyrics to generate the phoneme features of each phoneme; determine the phonemes to be prolonged with the phoneme types of finals and simple finals from all the phonemes in the target lyrics; perform prolongation processing on the phoneme features of the phonemes to be prolonged to obtain prolonged phoneme features; generate the text features based on the prolonged phoneme features and the phoneme features of each phoneme.

[0152] In one embodiment, when performing feature transformation processing on the notes, note durations, and slurs in the target musical score to generate the musical score features, the processor 901 may execute: determining the notes and the number of phonemes corresponding to each syllable in the target lyrics; determining the syllable duration of each syllable according to the note durations of the notes corresponding to each syllable; evenly distributing the syllable duration of each syllable according to the number of phonemes corresponding to each syllable to obtain the phoneme durations of each phoneme in the target lyrics; performing feature transformation processing on the notes, the phoneme durations of each phoneme, and the slurs in the target lyrics to generate the musical score features.

[0153] In one embodiment, when performing duration prediction processing based on the text features and the musical score features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics, the processor 901 may execute: aligning the features in the text features and the musical score features according to the phoneme order in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; performing bidirectional long short-term memory encoding processing on the song-lyric features corresponding to each phoneme in the target lyrics to obtain long short-term memory features; performing duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics.

[0154] In one embodiment, when performing duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations corresponding to each phoneme in the target text, the processor 901 may execute: using the trained duration model to perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations corresponding to each phoneme in the target text; the processor 901 may also execute: using the first objective function for predicting phoneme durations and the second objective function for predicting syllable durations to jointly train the preset duration model to obtain the trained duration model.

[0155] In one embodiment, when performing acoustic encoding processing on the text features and the musical score features according to the predicted phoneme durations to generate the acoustic encoding, the processor 901 may execute: aligning the features in the text features and the musical score features according to the phoneme order in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; performing self-attention encoding processing based on the song-lyric features corresponding to each phoneme in the target lyrics to obtain the self-attention encoding; expanding the self-attention encoding according to the predicted phoneme durations to obtain the acoustic encoding.

[0156] In one embodiment, the self-attention encoding includes sub-attention encodings corresponding to each phoneme in the target lyrics; when extending the self-attention encoding according to the predicted phoneme duration to obtain the acoustic encoding, the processor 901 may execute: performing feature replication processing on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme to obtain a replicated encoding corresponding to each phoneme; generating the acoustic encoding based on the sub-attention encoding and the replicated encoding corresponding to each phoneme.

[0157] In one embodiment, when performing feature replication processing on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme, the processor 901 may execute: determining the syllable duration of each syllable and the phonemes corresponding to each syllable in the target lyrics; performing scaling processing on the predicted phoneme duration of the phonemes corresponding to each syllable according to the syllable duration of each syllable to obtain a scaled phoneme duration corresponding to each phoneme, wherein the scaled phoneme duration corresponding to a phoneme of the initial consonant phoneme type is less than a predetermined duration; performing feature replication processing on the sub-attention encoding corresponding to each phoneme based on the scaled phoneme duration corresponding to each phoneme.

[0158] In one embodiment, when performing progressive decoding processing on the acoustic encoding by using at least two hierarchically cascaded decoding networks with hierarchical progressive training to obtain the target Mel spectrogram, the processor 901 may execute: performing self-attention decoding processing on the acoustic encoding in sequence by using at least two hierarchically cascaded self-attention decoding networks with hierarchical progressive training to obtain the decoded Mel spectrogram; generating the target Mel spectrogram based on the decoded Mel spectrogram.

[0159] In one embodiment, when generating the target Mel spectrogram based on the decoded Mel spectrogram, the processor 901 may execute: performing convolution processing on the input feature sequence corresponding to the decoded Mel spectrogram to obtain a convolutional feature sequence; performing a fully connected operation processing after concatenating the input feature sequence and the convolutional feature sequence to obtain a fully connected feature sequence; performing bidirectional recursive feature extraction processing on the fully connected feature sequence to obtain a smoothed spectrogram feature sequence, so as to generate the target Mel spectrogram.

[0160] In one embodiment, the processor 901 may further execute: adding a loss function for predicting the target Mel spectrogram to each layer of the at least two hierarchically cascaded self-attention decoding networks; performing hierarchical progressive training on the at least two hierarchically cascaded self-attention decoding networks based on the added loss function.

[0161] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by a computer program, or by controlling related hardware through a computer program. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0162] For this reason, the embodiments of the present application further provide a computer-readable storage medium, in which a computer program is stored, and the computer program can be loaded by a processor to execute the steps in any one of the methods provided by the embodiments of the present application.

[0163] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disk, etc.

[0164] Since the computer program stored in the computer-readable storage medium can execute the steps in any one of the methods provided by the embodiments of the present application, the beneficial effects achievable by the methods provided by the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be elaborated here.

[0165] According to one aspect of the present application, there is provided a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the methods provided in the various alternative implementation manners in the above embodiments of the present application.

[0166] After considering the specification and practicing the disclosed embodiments herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.

[0167] It should be understood that the present application is not limited to the embodiments described above and shown in the drawings, and various modifications and changes can be made without departing from its scope.

Claims

1. An audio synthesis method, characterized in that, Including: Obtaining the text features of the target lyrics and the music score features of the target music score; Using the trained duration model, performing duration prediction processing based on the text features and music score features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics; wherein, the trained duration model is obtained by jointly training a preset duration model with a first objective function for predicting phoneme durations and a second objective function for predicting syllable durations; Performing acoustic encoding processing on the text features and the music score features according to the predicted phoneme durations to generate acoustic encodings; Using at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encodings to obtain the target mel spectrogram; Generating a synthesized song audio corresponding to the target lyrics and the target music score based on the target mel spectrogram.

2. The method according to claim 1, wherein The obtaining the text features of the target lyrics and the music score features of the target music score includes: Performing feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features; Performing feature conversion processing on the notes, note durations and ties in the target music score to generate the music score features.

3. The method according to claim 2, wherein The performing feature conversion processing on the phonemes and phoneme type information in the target lyrics to generate the text features includes: Performing feature conversion processing on each phoneme and the phoneme type information corresponding to each phoneme in the target lyrics to generate the phoneme features of each phoneme; Determining the phonemes to be prolonged with the phoneme type of finals from all the phonemes in the target lyrics; wherein, the finals include single finals; Performing prolonging processing on the phoneme features of the phonemes to be prolonged to obtain prolonged phoneme features; Generating the text features based on the prolonged phoneme features and the phoneme features of each phoneme.

4. The method according to claim 2, wherein The performing feature conversion processing on the notes, note durations and ties in the target music score to generate the music score features includes: Determining the notes and the number of phonemes corresponding to each syllable in the target lyrics; Determining the syllable duration of each syllable according to the note duration of the note corresponding to each syllable; Evenly distributing the syllable duration of each syllable according to the number of phonemes corresponding to each syllable to obtain the phoneme duration of each phoneme in the target lyrics; Performing feature conversion processing on the notes, the phoneme duration of each phoneme and the ties in the target lyrics to generate the music score features.

5. The method according to claim 1, wherein The performing duration prediction processing based on the text features and music score features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics includes: Aligning the features in the text features and the music score features according to the phoneme order in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; Performing bidirectional long short-term memory encoding processing on the song-lyric features corresponding to each phoneme in the target lyrics to obtain long short-term memory features; Performing duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics.

6. The method according to claim 5, characterized in that, The performing duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations corresponding to each phoneme in the target lyrics includes: Using the trained duration model, perform duration prediction processing based on the long short-term memory features to obtain the predicted phoneme durations of each phoneme in the target lyrics.

7. The method according to claim 1, wherein The acoustic encoding process for the text features and the musical score features according to the predicted phoneme durations to generate an acoustic encoding includes: Align the features in the text features and the musical score features in the order of phonemes in the target lyrics to obtain the song-lyric features corresponding to each phoneme in the target lyrics; Perform self-attention encoding processing based on the song-lyric features corresponding to each phoneme in the target lyrics to obtain a self-attention encoding; Perform an expansion process on the self-attention encoding according to the predicted phoneme durations to obtain the acoustic encoding.

8. The method according to claim 7, characterized in that, The self-attention encoding includes the sub-attention encodings corresponding to each phoneme in the target lyrics; The expansion process on the self-attention encoding according to the predicted phoneme durations to obtain the acoustic encoding includes: Perform feature replication processing on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme to obtain the replicated encoding corresponding to each phoneme; Generate the acoustic encoding based on the sub-attention encoding and the replicated encoding corresponding to each phoneme.

9. The method according to claim 8, characterized in that, The feature replication processing on the sub-attention encoding corresponding to each phoneme according to the predicted phoneme duration corresponding to each phoneme includes: Determine the syllable duration of each syllable in the target lyrics and the phonemes corresponding to each syllable; Perform scaling processing on the predicted phoneme durations of the phonemes corresponding to each syllable according to the syllable duration of each syllable to obtain the scaled phoneme durations corresponding to each phoneme, where the scaled phoneme durations corresponding to phonemes of the initial consonant phoneme type are less than a predetermined duration; Perform feature replication processing on the sub-attention encoding corresponding to each phoneme based on the scaled phoneme duration corresponding to each phoneme.

10. The method according to claim 1, characterized in that The method of using at least two cascaded decoding networks with hierarchical progressive training to perform progressive decoding processing on the acoustic encoding to obtain the target Mel spectrogram includes: Using at least two cascaded self-attention decoding networks with hierarchical progressive training to perform self-attention decoding processing on the acoustic encoding in sequence to obtain the decoded Mel spectrogram; Generate the target Mel spectrogram based on the decoded Mel spectrogram.

11. The method according to claim 10, characterized in that, The generating the target Mel spectrogram based on the decoded Mel spectrogram includes: Perform convolution processing on the input feature sequence corresponding to the decoded Mel spectrogram to obtain a convolution feature sequence; Perform a fully connected operation processing after splicing the input feature sequence and the convolution feature sequence to obtain a fully connected feature sequence; Perform bidirectional recursive feature extraction processing on the fully connected feature sequence to obtain a smoothed spectrogram feature sequence to generate the target Mel spectrogram.

12. The method according to claim 10, wherein The method further includes: Adding a loss function for predicting the target Mel spectrogram to each self-attention decoding network in at least two cascaded self-attention decoding networks; Based on the added loss function, perform hierarchical progressive training on the at least two cascaded self-attention decoding networks.

13. An audio synthesis device, characterized in that, Including: An acquisition module for acquiring the text features of the target lyrics and the musical score features of the target musical score; A duration prediction module, configured to use a trained duration model to perform duration prediction processing based on the text features and musical score features, so as to obtain the predicted phoneme duration corresponding to each phoneme in the target lyrics; wherein, the trained duration model is obtained by jointly training a preset duration model using a first objective function for predicting phoneme duration and a second objective function for predicting syllable duration; An acoustic encoding module, configured to perform acoustic encoding processing on the text features and the musical score features according to the predicted phoneme duration to generate an acoustic encoding; A cascaded decoding module, configured to perform progressive decoding processing on the acoustic encoding by using at least two cascaded decoding networks trained in a hierarchical progressive manner to obtain a target Mel spectrogram; A synthesis module, configured to generate a synthesized song audio corresponding to the target lyrics and the target musical score based on the target Mel spectrogram.

14. A computer-readable storage medium, characterized in that, A computer program is stored thereon, and when the computer program is executed by a processor of a computer, the computer is caused to execute the method according to any one of claims 1 to 12.

15. An electronic device, characterized in that, Comprising: A memory storing a computer program; A processor, reading the computer program stored in the memory to execute the method according to any one of claims 1 to 12.

16. A computer program product, characterized in that, A computer program product includes computer instructions, and the computer instructions are stored in a computer-readable storage medium; a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Song synthesis method and device, readable medium and electronic equipment

    CN111583900A