Method and device for tone conversion, electronic equipment and product
By extracting semantic features and multimodal information through a self-attention-based diffusion model for timbre conversion, the problem of poor timbre conversion results in existing technologies is solved, achieving timbre conversion with high similarity and high accuracy, and simplifying the model training process.
Patent Information
- Application Number
- CN202410606075.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-11-18
AI Technical Summary
In existing timbre conversion technologies, the timbre of the converted audio has a low similarity to the prompt audio, and the sound quality is unsatisfactory. This is because the model's expressive ability is insufficient and the semantic features of the audio to be converted contain timbre information, which makes it impossible to effectively ignore this information in the converted audio.
A self-attention-based diffusion model is adopted. By extracting the semantic features of the audio to be converted and the prompt audio, the self-attention mechanism is used to generate the converted acoustic features. The timbre is converted by combining multimodal and multiscale information to generate high-quality converted audio.
It improves the timbre similarity and pronunciation accuracy between the converted audio and the prompt audio, reduces model training time and cost, and improves user experience.
Smart Images

Figure CN120977323A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present disclosure relates generally to the field of artificial intelligence, and more particularly, to a method, apparatus, electronic device and product for timbre conversion. BACKGROUND
[0002] Timbre conversion is a technique that changes the timbre characteristics of a sound to make it sound like another sound. Timbre conversion can be used in video production, audiobook production, movie dubbing, and other audio-related fields. In some scenarios, timbre conversion can simply adjust the pitch and sound quality of audio.
[0003] For example, a user may want to attract the audience and create interesting content by changing the sound of his own speech when using a video clip application to create a short video. In this scenario, the creator of the video wants to convert his own voice using the clip application. SUMMARY
[0004] Embodiments of the present disclosure provide a method, apparatus, electronic device and product for timbre conversion.
[0005] In a first aspect of embodiments of the present disclosure, a method for timbre conversion is provided. The method includes determining semantic features of to-be-converted audio, the to-be-converted audio having an original timbre. The method further includes obtaining prompt audio, the prompt audio having a target timbre different from the original timbre. The method further includes generating converted acoustic features based on the semantic features of the to-be-converted audio and the prompt audio, using a self-attention-based diffusion model. In addition, the method further includes generating converted audio based on the converted acoustic features, the converted audio being audio in which the timbre of the to-be-converted audio is converted to the target timbre.
[0006] In a second aspect of embodiments of the present disclosure, an apparatus for timbre conversion is provided. The apparatus includes a semantic feature determination module configured to determine semantic features of to-be-converted audio, the to-be-converted audio having an original timbre. The apparatus further includes a prompt audio obtaining module configured to obtain prompt audio, the prompt audio having a target timbre different from the original timbre. The apparatus further includes an acoustic feature generation module configured to generate converted acoustic features based on the semantic features of the to-be-converted audio and the prompt audio, using a self-attention-based diffusion model. In addition, the apparatus further includes a converted audio generation module configured to generate converted audio based on the converted acoustic features, the converted audio being audio in which the timbre of the to-be-converted audio is converted to the target timbre.
[0007] In a third aspect of embodiments of the present disclosure, an electronic device is provided. The electronic device includes one or more processors; and a storage storing one or more programs, when executed by the one or more processors, cause the one or more processors to implement a method for timbre conversion. The method includes determining semantic features of audio to be converted, the audio to be converted having an original timbre. The method further includes obtaining prompt audio, the prompt audio having a target timbre different from the original timbre. The method further includes generating converted acoustic features using a self-attention based diffusion model based on the semantic features of the audio to be converted and the prompt audio. In addition, the method further includes generating converted audio based on the converted acoustic features, the converted audio being audio in which the timbre of the audio to be converted is converted into the target timbre.
[0008] In a fourth aspect of embodiments of the present disclosure, a computer program product is provided. The computer program product is tangibly stored on a non-transitory computer readable medium and includes machine executable instructions that, when executed, cause a machine to implement a method for timbre conversion. The method includes determining semantic features of audio to be converted, the audio to be converted having an original timbre. The method further includes obtaining prompt audio, the prompt audio having a target timbre different from the original timbre. The method further includes generating converted acoustic features using a self-attention based diffusion model based on the semantic features of the audio to be converted and the prompt audio. In addition, the method further includes generating converted audio based on the converted acoustic features, the converted audio being audio in which the timbre of the audio to be converted is converted into the target timbre.
[0009] The Summary is provided to introduce a selection of concepts in a simplified form that are further described below in the of the Invention. This Summary is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. BRIEF DESCRIPTION OF DRAWINGS
[0010] The above and other features, advantages and aspects of embodiments of the present disclosure will become more apparent upon reading of the following detailed description, taken in conjunction with the accompanying drawings, wherein:
[0011] Figure 1 A schematic diagram illustrating an example environment in which some embodiments of the present disclosure can be implemented is shown;
[0012] Figure 2 A flowchart illustrating a method for timbre conversion according to some embodiments of the present disclosure is shown;
[0013] Figure 3A schematic diagram showing an example process of implementing timbre conversion with a self-attention based diffusion model according to some embodiments of the present disclosure is shown;
[0014] Figure 4 A schematic diagram showing an example architecture of a self-attention based diffusion model according to some embodiments of the present disclosure is shown;
[0015] Figures 5A to 5C A schematic diagram showing an example process of training a self-attention based diffusion model in two stages according to some embodiments of the present disclosure is shown;
[0016] Figure 6 A block diagram of an apparatus for timbre conversion according to some embodiments of the present disclosure is shown; and
[0017] Figure 7 A block diagram of an electronic device according to some embodiments of the present disclosure is shown.
[0018] In all the drawings, the same or similar reference numerals denote the same or similar elements. DETAILED DESCRIPTION
[0019] It can be understood that the data involved in the technical solutions of the present disclosure (including but not limited to the data itself, the acquisition or use of the data) should comply with the requirements of the relevant laws and regulations and relevant provisions.
[0020] It can be understood that before using the technical solutions disclosed in the embodiments of the present disclosure, the type of personal information involved in the present disclosure, the scope of use, the scenario of use, etc. should be informed to the user and the authorization of the user should be obtained through appropriate means according to relevant laws and regulations.
[0021] For example, when receiving the active request of the user, the prompt information is sent to the user to explicitly prompt the user that the operation requested to be executed will need to acquire and use the personal information of the user. Thus, the user can voluntarily choose whether to provide the personal information to the electronic device, application program, server or storage medium, etc. software or hardware that executes the operation of the technical solutions of the present disclosure according to the prompt information.
[0022] As an optional but non-limiting implementation manner, in response to receiving the active request of the user, the manner of sending the prompt information to the user may, for example, be the manner of a pop-up window, and the prompt information may, for example, be presented in the form of text in the pop-up window. In addition, the pop-up window may, for example, also carry a selection control for the user to select “agree” or “disagree” to provide the personal information to the electronic device.
[0023] It can be understood that the above notification and user authorization process is only illustrative, and does not limit the implementation manner of the present disclosure, and other manners that meet the relevant laws and regulations can also be applied to the implementation manner of the present disclosure.
[0024] It can be understood that the timbre involved in the embodiments of the present disclosure is an existing or authorized timbre in the timbre library.
[0025] Embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms, and should not be interpreted as being limited to the embodiments set forth herein, but rather these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are for exemplary purposes only and are not intended to limit the scope of protection of the present disclosure.
[0026] In the description of the embodiments of the present disclosure, the term "comprising" and its similar terms are understood to be open-ended, i.e., "including but not limited to". The term "based on" is understood to be "at least partially based on". The term "one embodiment" or "the embodiment" is understood to be "at least one embodiment". The terms "first", "second", and the like can refer to different or the same objects unless explicitly stated otherwise. Other explicit and implicit definitions can also be included below.
[0027] As described above, in some scenarios, a user can provide a specific voice as a prompt audio, and hope that the model can remember the timbre of the prompt audio. Then, the user can provide another audio (e.g., his own voice) as the audio to be converted, and hope that the model can convert the timbre of the audio to be converted to the timbre of the prompt audio. Compared with some techniques for generating a specific timbre based on text, in this scenario, the content of the converted audio (e.g., including text content, speech rate, intonation, duration, etc.) can be the same as that of the audio to be converted, and only the timbre is converted to the timbre in the prompt audio. In addition, compared with some scenarios that provide several specific timbres for the user to choose from, in this scenario, the user is allowed to provide any prompt audio and perform timbre conversion without the need for model training for the audio. It should be understood that the prompt audio and its timbre are authorized to be used.
[0028] In some related technologies related to timbre conversion, a deep neural network or a generative adversarial network can be used to implement timbre conversion. However, in these related technologies, the similarity of the timbre of the converted audio to the timbre of the prompt audio is low, and the sound quality of the converted audio is also not satisfactory. The reasons for these problems include insufficient expression ability of the model, and the semantic features of the audio to be converted contain some timbre information of the audio to be converted, which cannot be ignored when training the model, thereby resulting in insufficient similarity of the converted audio to the timbre of the prompt audio.
[0029] To this end, embodiments of this disclosure provide a scheme for timbre conversion using a self-attention-based diffusion model. In this scheme, an audio segment to be converted and a prompt audio segment are acquired. The goal is to convert the timbre of the audio segment to be converted (e.g., the user's own timbre) into the timbre of the prompt audio segment without altering the content of the audio segment to be converted. The scheme then determines the semantic features of the audio segment to be converted, and based on these semantic features and the prompt audio segment, uses a self-attention-based diffusion model to generate converted acoustic features. Finally, the scheme generates converted audio based on the converted acoustic features, and the timbre of the converted audio segment is converted into the timbre of the prompt audio segment.
[0030] This approach enhances the model's expressive power, thereby increasing the timbre similarity between the converted audio and the prompt audio, and also improving the pronunciation accuracy of the converted audio. Furthermore, this method allows for timbre conversion based solely on a prompt audio clip without pre-training the model for the timbre of the prompt audio. This reduces the time and cost of training the model and makes timbre conversion convenient for users, thus improving the user experience.
[0031] Figure 1 A schematic diagram of an example environment 100 that may be implemented according to some embodiments of the present disclosure is shown. For example... Figure 1 As shown, environment 100 includes audio to be converted 102 and prompt audio 108. Audio to be converted 102 is a speech by speaker 104 (e.g., the user themselves), with the timbre being the timbre of speaker 104's voice. Audio to be converted 102 also includes content 106, which may include information such as the text, tone, and speech rate of audio to be converted 102. Prompt audio 108 is a speech by speaker 110 (e.g., a movie character), with the timbre being the timbre of speaker 110's voice, and as described above, this timbre is an authorized timbre. Prompt audio 108 also includes content 112, which may include information such as the text, tone, and speech rate of prompt audio 108.
[0032] In environment 100, semantic features 114 can be extracted from the audio to be converted 102. Semantic features refer to features extracted from the audio signal that can express the meaning of the audio content. For example, semantic features can indicate the text, tone, speech rate, etc., of the audio signal. In environment 100, semantic features can be features extracted in various ways, such as the HuBERT model, the BEST-RQ model, models based on ASR bottleneck features, and other convolutional neural networks or recurrent neural networks.
[0033] After extracting the semantic features 114 from the audio to be converted 102, the timbre conversion model 116 can generate converted acoustic features 118 based on the semantic features 114 and the prompt audio 108. Acoustic features can indicate various physical properties of a sound, for example, the acoustic features can indicate the timbre, frequency, articulation, loudness, etc. of an audio signal.
[0034] In embodiments of the present disclosure, the timbre conversion model 116 can be a diffusion model based on self-attention. A diffusion model is a kind of generative model, which is commonly used in image generation tasks. The workflow of this model includes two processes: a forward process and a reverse process. In the forward process, the model adds noise to the data, making the data more random; the reverse process uses the trained model to denoise the noisy data multiple times to recover the clean data. The diffusion model can therefore generate high-quality and detailed data.
[0035] The representative of self-attention mechanism is the Transformer model, so the diffusion model based on self-attention can be the Transformer Diffusion model. The self-attention mechanism can calculate the attention score of each element in the sequence to other elements, and decide which part of the input sequence should pay more attention to when generating each output element according to these scores. The self-attention mechanism allows the model to consider all elements in the sequence when processing data, effectively capturing long-range dependencies in the data.
[0036] Combining the generation ability of the diffusion model with the self-attention mechanism of the Transformer architecture, the timbre conversion model 116 can utilize the context information of the entire original acoustic features in the generation process to generate the target acoustic features. This method can improve the accuracy, authenticity and timbre similarity of the generated converted acoustic features to the prompt audio 108.
[0037] In the environment 100, after generating the converted acoustic features 118, the vocoder 120 can generate converted audio 122 based on the converted acoustic features 118. The vocoder 120 can be any technology that can synthesize audio based on acoustic features, for example, a linear predictive coder, a phase vocoder, a channel vocoder, etc. The content 124 of the generated converted audio 122 is the same as the content 106 of the audio to be converted 102, and its timbre (i.e. the timbre of the speaker 126) is the same as the timbre of the prompt audio 108 (i.e. the timbre of the speaker 110), thereby achieving the conversion of the timbre of the audio to be converted 102 to the timbre of the prompt audio 108 while keeping the audio content unchanged.
[0038] In this way, the timbre similarity between the converted audio 122 and the prompt audio 108 can be improved, and the pronunciation accuracy of the converted audio 122 can also be improved. In addition, in this way, the timbre conversion can be implemented only by providing a piece of prompt audio 108 without pre-training a timbre conversion model for the speaker 110, thereby reducing the time and cost for training the model, and also enabling the user to conveniently perform the timbre conversion, thereby improving the user experience.
[0039] Figure 2 A flowchart of a method 200 for timbre conversion is shown according to some embodiments of the present disclosure. At block 202, the method 200 can determine semantic features of a to-be-converted audio, the to-be-converted audio having an original timbre. For example, in the environment 100 as shown, the to-be-converted audio 102 can be obtained, the to-be-converted audio 102 having a timbre of the speaker 104 (also referred to as an original timbre) and including content 106, which may, for example, include text, intonation, speech rate, etc. information of the to-be-converted audio 102. In the environment 100, semantic features 114 can be extracted from the to-be-converted audio 102 using any technology, which may, for example, indicate the text, intonation, speech rate, etc. information of the to-be-converted audio 102. Figure 1
[0040] At block 204, the method 200 can obtain a prompt audio, the prompt audio having a target timbre different from the original timbre. For example, in the environment 100 as shown, the prompt audio 108 can be obtained, the prompt audio 108 having a timbre of the speaker 110 (also referred to as a target timbre) and including content 112, which may, for example, include text, intonation, speech rate, etc. information of the prompt audio 108. As described above, the target timbre is an authorized usable timbre by the speaker 110 or a related subject. Figure 1
[0041] At block 206, the method 200 can generate converted acoustic features based on the semantic features of the to-be-converted audio and the prompt audio using a self-attention-based diffusion model. For example, in the environment 100 as shown, the converted acoustic features 118 can be generated based on the semantic features 114 of the to-be-converted audio 102 and the prompt audio 108 using the timbre conversion model 116, wherein the timbre conversion model 116 is a self-attention-based diffusion model. Figure 1
[0042] At block 208, the method 200 can generate a converted audio based on the converted acoustic features, the converted audio being an audio whose timbre of the to-be-converted audio is converted to the target timbre. For example, in the environment 100 as shown, the converted audio 122 can be generated based on the converted acoustic features 118 using the timbre conversion model 116, wherein the converted audio 122 is an audio whose timbre of the to-be-converted audio 102 is converted to the target timbre of the speaker 110. Figure 1 In the illustrated environment 100, the converted acoustic features 118 can be used to generate converted audio 122 using a vocoder 120. The vocoder 120 can be any technology that generates audio based on acoustic features. The converted audio 122 has a timbre of the speaker 126 that is the same as the timbre of the speaker 110 of the prompt audio 108. In addition, the converted audio 122 includes content 124 that is the same as the content 106 of the audio to be converted 102.
[0043] In this way, the timbre similarity between the converted audio and the prompt audio can be improved, and the pronunciation accuracy of the converted audio can also be improved. In addition, in this way, the timbre conversion can be implemented only by providing a piece of prompt audio without pre-training a timbre conversion model for the speaker, thereby reducing the time and cost for training the model, and also enabling the user to conveniently perform the timbre conversion, thereby improving the user experience.
[0044] In some embodiments, to further improve the timbre similarity between the converted audio and the prompt audio and the pronunciation accuracy of the converted audio, the converted acoustic features can be generated using information of multiple modalities and information across different scales in a single modality of the audio to be converted and the prompt audio in the process of timbre conversion. In some embodiments, when generating the converted acoustic features, text embeddings associated with the prompt text of the prompt audio and the original text of the audio to be converted, semantic embeddings associated with the prompt semantic features of the prompt audio and the original semantic features of the audio to be converted, a global timbre embedding associated with the prompt audio, and a local timbre embedding associated with the prompt audio can be determined. Then, the converted acoustic features can be generated based on the text embeddings, the semantic embeddings, the global timbre embedding, and the local timbre embedding.
[0045] Figure 3 A schematic diagram illustrating an example process 300 of implementing timbre conversion using a diffusion model based on self-attention according to some embodiments of the present disclosure is shown. As Figure 3As shown, the process 300 can extract the original text 302 (i.e., the content spoken by the speaker in the audio to be converted, also referred to as the text to be synthesized) and the original semantic features 310 from the audio to be converted, and extract the prompt text 304 (i.e., the content spoken by the speaker in the prompt audio) and the prompt semantic features 312 from the prompt audio. In addition, the process 300 can also extract the prompt acoustic features 318 from the prompt audio. In some related technologies, the audio with the specified timbre can be generated based on the prompt semantic features 312 and the original semantic features 310 only, however, the audio generated in this way has low similarity to the timbre of the prompt audio, and the pronunciation accuracy is also low. For this reason, in the process 300, the prompt semantic features 312 and the original semantic features 310 can be taken as a skeleton, and they are fused with information of multiple modalities (i.e., text and timbre) and multiple scales of one modality (i.e., global timbre information and local timbre information) to generate the timbre-converted audio, so as to improve the timbre similarity and pronunciation accuracy.
[0046] As shown in FIG. 3, the process 300 can include the following steps. Figure 3 As shown in the process 300, the text encoder 306 can generate the text embedding 308 based on the original text 302 and the prompt text 304. The size of the text embedding 308 is [T1+T2, C], where T1 represents the length of the prompt text, T2 represents the length of the original text, and C represents a specific vector dimension. [T1+T2, C] can represent T1+T2 vectors with a dimension of C. The model structure of the text encoder 306 can be a convolutional neural network with padding. The padding can adjust the output size of the convolutional layer, and can avoid information loss. In addition, the text encoder 306 can also be a Transformer. In this way, the text information can be provided for generating the timbre-converted audio, and the text information can improve the pronunciation accuracy of the generated audio.
[0047] As shown in FIG. 3, the process 300 can include the following steps. Figure 3 As shown, the global timbre encoder 320 can generate the global timbre embedding 322 based on the global information of the prompt acoustic features 318. The global information refers to the overall information of the acoustic features. In some embodiments, when generating the global timbre embedding 322, the prompt acoustic features 318 can be taken as a whole in the time dimension to generate the global timbre embedding 322. The size of the generated global timbre embedding 322 is [1, C]. In this way, the global-scale timbre information can be provided for generating the timbre-converted audio, so as to enrich the scale of the timbre-related information, and improve the authenticity and timbre similarity of the generated audio.
[0048] The input to the global timbre encoder 320 is a segment of acoustic features, and the output is a vector without a time dimension (which can also be understood as having a time dimension of 1). In some embodiments, the global timbre encoder 320 can be implemented using an ECAPA-TDNN structure. ECAPA-TDNN is a neural network structure that adds an attention mechanism to a Time Delay Neural Network (TDNN). Using this structure to implement the global timbre encoder 320 allows it to effectively learn the feature dependencies in the time dimension and can improve the feature representation ability by dynamically adjusting the feature importance of different channels, thereby capturing speech features at different time scales and enhancing the encoder's ability to recognize speech patterns.
[0049] like Figure 3 As shown, process 300 can also concatenate the cue acoustic feature 318 with the acoustic feature 326 to be synthesized to generate acoustic feature 328. At this time, since the acoustic feature 326 to be synthesized is unknown (i.e., the part the model needs to predict), it can be initialized using specific initial values (i.e., placeholders). Then, the local timbre encoder 330 can generate a local timbre embedding 332 based on the local information of the acoustic feature 328. Local information refers to information about a portion of the acoustic features, such as an acoustic feature corresponding to one or a portion of all audio frames. In some embodiments, when generating the local timbre embedding 332, the acoustic feature 328 can be segmented into multiple local acoustic features along the time dimension, and then the local timbre embedding 332 can be generated based on these multiple local acoustic features. The generated local timbre embedding 332 has a size of [T3+T4, C], where T3 is the length of the cue acoustic feature 318 of the cue audio, and T4 is the length of the acoustic feature 326 to be synthesized. The local timbre encoder 330 may, for example, include one or more fully connected layers to make the dimension of the generated embedding C.
[0050] In this way, acoustic feature 328 is divided into T3+T4 local acoustic features, and corresponding timbre embeddings are generated for each local acoustic feature and combined into local timbre embedding 332. This can provide local scale timbre information for generating audio with timbre conversion, thereby enriching the scale of timbre information during the generation process and improving the realism and timbre similarity of the generated audio.
[0051] like Figure 3As shown, in the process 300, the semantic encoder 314 can generate a semantic embedding 316 based on the original semantic features 310 and the prompt semantic features 312. The semantic embedding 316 has a size of [T3+T4, C], i.e., the same size as the local timbre embedding 332. The input of the semantic encoder is the concatenated semantic features generated by concatenating the original semantic features 310 and the prompt semantic features 312, which can have a different length than the acoustic features (i.e., the prompt acoustic features 318 and the acoustic features 328) (e.g., the semantic features can be 20 samples per second, and the acoustic features can be 40 samples per second). In the case where the lengths are different, the semantic encoder can up-sample (e.g., with a deconvolutional layer), down-sample (e.g., with a convolutional layer), or a combination of up-sampling and down-sampling, the semantic features to make the frequency (i.e., the number of samples per second) of the output semantic embedding 316 consistent with the frequency of the acoustic features. In this way, the size of the generated semantic embedding 316 can be aligned with the size of the local timbre embedding 332 to facilitate the subsequent information fusion.
[0052] As described above, the global timbre embedding 322 has a size of [1, C], and for information fusion, the process 300 can repeat the global timbre embedding 322 in the time dimension T3+T4 times to generate a repeated global timbre embedding 324 that also has a size of [T3+T4, C]. In one aspect, the repeated global timbre embedding 324 includes T3+T4 repeated vectors of size C, where each vector is generated based on the entire information of the prompt acoustic features 318. In another aspect, the local timbre embedding 332 includes T3+T4 different vectors of size C, where each vector is generated based on the information of one audio frame (or one sample point) in the acoustic features 328. In this way, timbre information at both global and local scales can be provided.
[0053] To recover the predicted acoustic features from the noise, the process 300 can generate a noisy acoustic feature 334 and convert the noisy acoustic feature 334 into a noisy acoustic embedding 338 using a noisy acoustic feature encoder 336, which also has a size of [T3+T4, C]. Then, the process 300 can add the semantic embedding 316, the repeated global timbre embedding 324, the local timbre embedding 332, and the noisy acoustic embedding 338, all of which have a size of [T3+T4, C], to generate a fused acoustic embedding that also has a size of [T3+T4, C]. Then, the process 300 can concatenate the generated fused acoustic embedding with the text embedding 308 in the time dimension, which has a size of [T1+T2, C], to generate a fused multi-modal embedding 339 that has a size of [T1+T2+T3+T4, C].
[0054] As described above, the global timbre embedding 322 has a size of [1, C], and for information fusion, the process 300 can repeat the global timbre embedding 322 in the time dimension T3+T4 times to generate a repeated global timbre embedding 324 that also has a size of [T3+T4, C]. In one aspect, the repeated global timbre embedding 324 includes T3+T4 repeated vectors of size C, where each vector is generated based on the entire information of the prompt acoustic features 318. In another aspect, the local timbre embedding 332 includes T3+T4 different vectors of size C, where each vector is generated based on the information of one audio frame (or one sample point) in the acoustic features 328. In this way, timbre information at both global and local scales can be provided. Figure 3As shown, the self-attention-based diffusion model 340 can generate predicted acoustic features 342 based on the fusion of multimodal embeddings 339. The predicted acoustic features 342 include predicted acoustic features 344 corresponding to cue acoustic features 318 and predicted acoustic features 346 corresponding to the acoustic features 326 to be synthesized. In process 300, the predicted acoustic features 344 corresponding to cue acoustic features 318 can be discarded, retaining only the predicted acoustic features 346 corresponding to the acoustic features 326 to be synthesized.
[0055] During training, the noise-added acoustic feature 334 can be generated by adding noise to the true value of the timbre-converted acoustic feature. Then, the loss between the predicted acoustic feature 346 and the true value can be calculated, and the self-attention-based diffusion model 340, text encoder 306, semantic encoder 314, global timbre encoder 320, and local timbre encoder 330 can be jointly trained based on this loss.
[0056] During the inference process, the acoustic feature 334 can be generated based on random noise. The predicted acoustic feature 346 is a timbre-converted acoustic feature generated by performing multiple noise reduction operations on the random noise (e.g., it could be...). Figure 1 The converted acoustic features 118 are then used. The timbre-converted audio can then be generated based on the predicted acoustic features 346 and using a vocoder.
[0057] In this way, the self-attention-based diffusion model 340 can perform timbre conversion based on information from multiple modalities (i.e., text and timbre) and information at multiple scales within a single modality (i.e., global timbre information and local timbre information). Thus, text information helps the model improve the pronunciation accuracy of the converted audio, while multi-scale timbre information helps the model improve timbre similarity.
[0058] Figure 4 A schematic diagram of an example architecture 400 of a self-attention-based diffusion model according to some embodiments of the present disclosure is shown. Figure 4 As shown, architecture 400 includes self-attention blocks 402, 404, 406, 408, 410, and 412, each of which can have a Transformer architecture. In architecture 400, these self-attention blocks are connected in series, meaning that the output of each upper self-attention block is used as at least a portion of the input of its lower adjacent self-attention block. Architecture 400 also includes multiple skip connections; for example, the output of self-attention block 402 is connected to self-attention block 412 via skip connection 414, and the output of self-attention block 404 is connected to self-attention block 410 via skip connection 416.
[0059] In architecture 400, the self-attention-based diffusion model receives input 418 and generates output 420. Architecture 400 feeds input 418 into self-attention blocks 402, each of which can independently process the input data and uses the Transformer architecture to extract and learn high-level features. Because the self-attention mechanism can process global information across the entire input sequence, the model can better understand and represent the complex patterns and relationships in input 418. The serial connections between self-attention blocks allow information to flow top-down through the model, with each block further processing and refining features based on the previous block, thus progressively improving the data representation capabilities.
[0060] Skip connections in architecture 400 prevent information from previous blocks from being forgotten by subsequent blocks, thus mitigating the vanishing gradient problem in the network architecture. Furthermore, these skip connections enable rapid feature propagation, facilitating the direct transfer of key information between blocks, thereby improving model efficiency and stability.
[0061] In the training and inference process of the timbre conversion model, the semantic features of the audio to be converted are used to generate the skeleton of the converted audio. However, in the semantic features of the audio to be converted (e.g., Figure 3 The original semantic features (310) inevitably contain some timbre information, which affects the timbre similarity between the converted audio and the prompt audio. In some related technologies, engineers need to carefully select semantic features containing less timbre information and generate audio based on the selected semantic features. However, these selected semantic features still contain timbre information, and the quality of the generated audio will be reduced due to the limitation of available semantic features.
[0062] Therefore, in some embodiments, a two-stage training process can be employed to train the timbre conversion model. In some embodiments, semantic features of the training audio can be determined. In a first training stage, an untrained self-attention-based diffusion model can be pre-trained based on the training audio and its semantic features. After the first training stage, semantic features of timbre changes can be generated using the pre-trained self-attention-based diffusion model based on the training audio and random audio, wherein the timbre of the random audio differs from that of the training audio. In a second training stage, the pre-trained self-attention-based diffusion model can be trained based on the training audio and the semantic features of timbre changes.
[0063] Figures 5A to 5C A schematic diagram of an example process 500 for training a self-attention-based diffusion model in two stages according to some embodiments of the present disclosure is shown. Figure 5A This illustrates the process of pre-training the model in the first stage.Figure 5B The process of generating a semantic feature of a timbre change using a pre-trained model is shown, Figure 5C The process of further training the pre-trained model using the generated semantic feature of the timbre change is shown.
[0064] As Figure 5A shown, in the first training stage, the process 500 can obtain an audio to be converted 502 as a training sample, which can be a segment of speech of a certain speaker. The process 500 can extract a text 504 and a semantic feature 506 (i.e., a semantic feature to be converted) from the audio 502, which can be any semantic feature extracted from the audio 502 and can contain timbre information of the audio 502. Then, the process 500 can randomly cut a partial audio 508 from the audio 502 as a prompt audio. Then, the self-attention-based diffusion model 510 can generate a predicted acoustic feature 512 (also referred to as a first predicted acoustic feature) based on the text 504, the semantic feature 506, and the partial audio 508, which can refer to the process 300 as Figure 3 shown.
[0065] In the process 500, a real acoustic feature 514 can also be extracted from the audio 502 as a ground truth. Then, the process 500 can pre-train the self-attention-based diffusion model 510 by calculating a loss 516 between the predicted acoustic feature 512 and the real acoustic feature 514, and minimizing the value of the loss 516. In this way, the first training stage is completed, and a pre-trained self-attention-based diffusion model 520 (as Figure 5B shown) can be obtained. At this time, since the semantic feature 506 contains the timbre information of the audio 502, the predicted acoustic feature 512 also contains some timbre information of the audio 502, which will result in a decrease in the timbre similarity of the converted audio.
[0066] As Figure 5B shown, after the completion of the first training stage, the process 500 enters the data production stage. In the data production stage, the process 500 can obtain a random audio 518, the timbre (i.e., the speaker) of which is different from that of the audio 502. The pre-trained self-attention-based diffusion model 520 can generate a timbre-changed acoustic feature 522 based on the text 504 and the semantic feature 506 of the audio 502, and the random audio 518 (i.e., as a prompt audio), which can refer to the process 300 as Figure 3The process 300 is shown. At this point, the timbre information in the timbre-changed acoustic feature 522 is fused with the timbre information of the semantic feature 506 and the timbre information of the random audio 518. Therefore, the timbre of the timbre-changed acoustic feature 522 is altered and differs from the timbre of the audio 502. It should be understood that the data involved in this disclosure (including but not limited to the data itself, the acquisition or use of the data) shall comply with the requirements of relevant laws, regulations, and provisions.
[0067] like Figure 5B As shown, process 500 can utilize vocoder 524 and generate timbre-changed audio 526 based on timbre-changed acoustic features 522. Then, process 500 can utilize semantic feature extractor 528 and generate timbre-changed semantic features 530 based on the timbre-changed audio 526. In this way, timbre-changed semantic features 530 can be generated based on semantic features 506, and timbre-changed semantic features 530 can have timbre information different from semantic features 506.
[0068] After the data production phase is completed, process 500 can proceed to the second training phase. For example... Figure 5C As shown, compared to Figure 5A In the first training phase, the pre-trained self-attention-based diffusion model can generate predicted acoustic features 532 based on semantic features 530 (instead of semantic features 506) of timbre changes, text 504 of audio 502, and a portion of audio 508. Then, process 500 can further train the pre-trained self-attention-based diffusion model 520 by calculating the loss 534 between the predicted acoustic features 532 and the true acoustic features 514 and minimizing the value of the loss 534.
[0069] By training the model in this way, since the semantic feature 530 of timbre change no longer contains the timbre information of audio 502, the trained self-attention-based diffusion model can ignore the original timbre in the audio to be converted when generating the timbre-converted audio, thereby improving the timbre similarity between the converted audio and the prompt audio.
[0070] Figure 6 A block diagram of a timbre conversion apparatus 600 according to some embodiments of the present disclosure is shown. Figure 6As shown, device 600 includes a semantic feature determination module 602, configured to determine the semantic features of the audio to be converted, which has an original timbre. Device 600 also includes a cue audio acquisition module 604, configured to acquire cue audio, which has a target timbre different from the original timbre. Device 600 further includes an acoustic feature generation module 606, configured to generate converted acoustic features based on the semantic features of the audio to be converted and the cue audio, using a self-attention-based diffusion model. Furthermore, device 600 includes a converted audio generation module 608, configured to generate converted audio based on the converted acoustic features; the converted audio is the audio whose timbre has been converted from the original audio to the target timbre.
[0071] It is understood that by utilizing the apparatus 600 of this disclosure, at least one of the many advantages achievable by the methods or processes described above can be realized. For example, apparatus 600 can improve the timbre similarity between the converted audio and the prompt audio, and can also improve the pronunciation accuracy of the converted audio. Furthermore, apparatus 600 can achieve timbre conversion by providing only a prompt audio without pre-training a model for the timbre of the prompt audio, thereby reducing the time and cost for training the model and enabling users to easily perform timbre conversion, thus improving the user experience.
[0072] Figure 7 Block diagrams of electronic devices 700 according to some embodiments of the present disclosure are shown. Device 700 may be the device or apparatus described in the embodiments of the present disclosure. Figure 7 As shown, device 700 includes a central processing unit (CPU) and / or a graphics processing unit (GPU) 701, which can perform various appropriate actions and processes according to computer program instructions stored in read-only memory (ROM) 702 or loaded from storage unit 708 into random access memory (RM) 703. RM 703 can also store various programs and data required for the operation of device 700. CPU / GPU 701, ROM 702, and RM 703 are interconnected via bus 704. Input / output (I / O) interface 705 is also connected to bus 704. Although not shown in... Figure 7 As shown, device 700 may also include a coprocessor.
[0073] A number of the components in device 700 are connected to the I / O interface 705, including an input unit 706, such as a keyboard, mouse, etc.; an output unit 707, such as various types of displays, speakers, etc.; a storage unit 708, such as a disk, a CD, etc.; and a communication unit 709, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 709 allows device 700 to exchange information / data with other devices over a computer network, such as the Internet, and / or various telecommunication networks.
[0074] The various methods or processes described above can be performed by the CPU / GPU 701. For example, in some embodiments, the methods can be implemented as a computer software program tangibly embodied in a machine readable medium, such as the storage unit 708. In some embodiments, portions or all of the computer program can be loaded and / or installed onto device 700 via the ROM 702 and / or the communication unit 709. When the computer program is loaded onto the RM 703 and executed by the CPU / GPU 701, one or more of the steps or actions of the methods or processes described above can be performed.
[0075] In some embodiments, the methods and processes described above can be implemented as a computer program product. The computer program product can include a computer readable storage medium having computer readable program instructions embodied therewith.
[0076] The computer readable storage medium can be a tangible device that can retain and store instructions for use by an instruction execution device. The computer readable storage medium can be, for example, but is not limited to, an electronic storage device, a magnetic storage device, an optical storage device, an electromagnetic storage device, a semiconductor storage device, or any suitable combination of the foregoing. More specific examples (a non-exhaustive list) of the computer readable storage medium include the following: a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), a static random access memory (SRAM), a portable compact disc read-only memory (CD-ROM), a digital versatile disk (DVD), a memory stick, a floppy disk, a mechanically encoded device such as punch-cards or punched tape, a
[0077] Computer readable program instructions described herein can be downloaded to respective computing / processing devices from a computer readable storage medium or to an external computer or external storage device via a network, for example, the Internet, a local area network, a wide area network and / or a wireless network. The network can comprise copper transmission cables, optical transmission fibers, wireless transmission, routers, firewalls, switches, gateway computers and / or edge servers. A network adapter card or network interface in each computing / processing device receives computer readable program instructions from the network and forwards the computer readable program instructions for storage in a computer readable storage medium within the respective computing / processing device.
[0078] Computer readable program instructions for carrying out operations of the present disclosure can be assembler instructions, instruction-set-architecture (ISA) instructions, machine instructions, machine dependent instructions, microcode, firmware instructions, state-setting data, or either source code or object code written in any combination of one or more programming languages, including an object oriented programming language and conventional procedural programming languages. The computer readable program instructions can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider). In some embodiments, electronic circuitry including, for example, programmable logic circuitry, field-programmable gate arrays (FPGA), or programmable logic arrays (PLA) can execute the computer readable program instructions by utilizing state information of the computer readable program instructions to personalize the electronic circuitry, in order to perform aspects of the present disclosure.
[0079] These computer readable program instructions can be provided to a processor of a general purpose computer, special purpose computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions / acts specified in the flowchart and / or block diagram block or blocks. These computer readable program instructions can also be stored in a computer readable storage medium that can include a non-transitory computer readable storage medium that can direct a computer, a programmable data processing apparatus, and / or other devices to function in a particular manner, such that the computer readable storage medium having instructions for causing an apparatus to implement various aspects of the functions / acts specified in the flowchart and / or block diagram block or blocks is produced.
[0080] The computer readable program instructions can also be loaded onto a computer, other programmable data processing apparatus, or other device to cause a series of operational steps to be performed on the computer, other programmable data processing apparatus, or other device to produce a computer implemented process such that the instructions which execute on the computer, other programmable data processing apparatus, or other device implement the functions / acts specified in the flowchart and / or block diagram block or blocks.
[0081] The flow diagrams and the block diagrams in the drawings are presented to illustrate the architecture, functionality, and operation of possible implementations of devices, methods and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flow diagrams and the block diagrams can represent a module, segment, or portion of instructions, which comprises one or more executable instructions for implementing the specified logic functions. In some alternative implementations, the functions noted in the blocks can occur out of the order noted in the figures. For example, two blocks shown in succession may, in fact, be executed substantially concurrently, or the blocks may sometimes be executed in the reverse order, depending upon the functionality involved. It will also be noted that each block of the block diagrams and / or flow diagrams, and combinations thereof, can be implemented by special purpose hardware-based systems that perform the specified functions or acts, or combinations of special purpose hardware and
[0082] Embodiments of the present disclosure have been described above, and the description is intended to be illustrative of the embodiments and not restrictive. Many modifications and variations of the described embodiments are possible and are apparent to those of ordinary skill in the art. The selection of terms is intended to best describe the principles of the embodiments, the practical application, or technical improvements over the existing technology, or to enable others skilled in the art to understand the embodiments disclosed herein.
[0083] Some example implementations of the present disclosure are listed below.
[0084] Example 1. A method for timbre conversion, comprising:
[0085] determining semantic features of audio to be converted, the audio to be converted having an original timbre;
[0086] obtaining prompt audio, the prompt audio having a target timbre different from the original timbre;
[0087] generating converted acoustic features based on the semantic features of the audio to be converted and the prompt audio, using a diffusion model based on self-attention; and
[0088] generate, based on the converted acoustic feature, a converted audio, the converted audio being an audio that converts a timbre of the audio to be converted to the target timbre.
[0089] Example 2. The method of example 1, wherein the semantic feature of the audio to be converted is an original semantic feature, and generating, based on the semantic feature of the audio to be converted and the prompt audio, the converted acoustic feature with the self-attention-based diffusion model comprises:
[0090] determining a text embedding associated with a prompt text of the prompt audio and an original text of the audio to be converted;
[0091] determining a semantic embedding associated with a prompt semantic feature of the prompt audio and the original semantic feature of the audio to be converted;
[0092] determining a global timbre embedding associated with the prompt audio;
[0093] determining a local timbre embedding associated with the prompt audio; and
[0094] generating the converted acoustic feature based on the text embedding, the semantic embedding, the global timbre embedding, and the local timbre embedding.
[0095] Example 3. The method of examples 1-2, wherein determining the text embedding associated with the prompt text of the prompt audio and the original text of the audio to be converted comprises:
[0096] generating, based on the prompt text and the original text, the text embedding with a text encoder.
[0097] Example 4. The method of examples 1-3, wherein determining the semantic embedding associated with the prompt semantic feature of the prompt audio and the original semantic feature of the audio to be converted comprises:
[0098] generating, based on the prompt semantic feature and the original semantic feature, the semantic embedding with a semantic encoder.
[0099] Example 5. The method of examples 1-4, wherein determining the global timbre embedding associated with the prompt audio comprises:
[0100] determining a prompt acoustic feature of the prompt audio;
[0101] generating, based on the prompt acoustic feature, the global timbre embedding with a global timbre encoder and treating the prompt acoustic feature as a whole in a temporal dimension.
[0102] Example 6. The method of examples 1-5, wherein determining the local timbre embedding associated with the prompt audio comprises:
[0103] determining a prompt acoustic feature of the prompt audio;
[0104] segmenting the prompt acoustic feature into a plurality of local acoustic features in a time dimension; and
[0105] generating the local timbre embedding with a local timbre encoder based on the plurality of local acoustic features.
[0106] Example 7. The method of examples 1-6, wherein generating the converted acoustic feature based on the text embedding, the semantic feature embedding, and the timbre embedding comprises:
[0107] generating random noise;
[0108] generating a noise-added acoustic embedding with a noise-added acoustic feature encoder based on the random noise;
[0109] generating a fusion embedding based on the text embedding, the semantic embedding, the global timbre embedding, the local timbre embedding, and the noise-added acoustic embedding;
[0110] generating the converted acoustic feature with the self-attention-based diffusion model based on the fusion embedding.
[0111] Example 8. The method of examples 1-7, wherein the training process of the self- attention-based diffusion model comprises:
[0112] determining a semantic feature of a training audio;
[0113] in a first training stage, pre-training an untrained self-attention-based diffusion model based on the training audio and the semantic feature of the training audio;
[0114] after the first training stage, generating a timbre-altered semantic feature with the pre-trained self-attention-based diffusion model based on the training audio and a random audio, a timbre of the random audio being different from a timbre of the training audio;
[0115] in a second training stage, training the pre-trained self-attention-based diffusion model based on the training audio and the timbre-altered semantic feature.
[0116] Example 9. The method of examples 1-8, wherein in the first training stage, pre-training the self-attention-based diffusion model based on the training audio and the semantic feature of the training audio comprises:
[0117] extracting a partial audio from the training audio;
[0118] generating, based on the semantic feature of the training audio and the partial audio, a first predicted acoustic feature using the untrained self-attention based diffusion model;
[0119] determining an acoustic feature of the training audio; and
[0120] pre-training the untrained self-attention based diffusion model by computing a loss between the first predicted acoustic feature and the acoustic feature of the training audio.
[0121] Example 10. The method of any of examples 1-9, wherein, after the first training phase, generating, based on the training audio and the random audio, the timbre-altered semantic feature using the pre-trained self-attention based diffusion model comprises:
[0122] generating, based on the semantic feature of the training audio and the random audio, a timbre-altered acoustic feature using the pre-trained self-attention based diffusion model; and
[0123] generating the timbre-altered semantic feature based on the timbre-altered acoustic feature.
[0124] Example 11. The method of any of examples 1-10, wherein generating the timbre-altered semantic feature based on the timbre-altered acoustic feature comprises:
[0125] generating a timbre-altered audio based on the timbre-altered acoustic feature; and
[0126] generating the timbre-altered semantic feature based on the timbre-altered audio.
[0127] Example 12. The method of any of examples 1-11, wherein, in the second training phase, training the pre-trained self-attention based diffusion model based on the training audio and the timbre-altered semantic feature comprises:
[0128] extracting a partial audio from the training audio;
[0129] generating, based on the timbre-altered semantic feature and the partial audio, a second predicted acoustic feature using the pre-trained self-attention based diffusion model;
[0130] determining an acoustic feature of the training audio; and
[0131] training the pre-trained self-attention based diffusion model by computing a loss between the second predicted acoustic feature and an acoustic feature of the training audio.
[0132] Example 13. An apparatus for timbre conversion, the apparatus comprising:
[0133] a semantic feature determination module configured to determine a semantic feature of a to-be-converted audio, the to-be-converted audio having an original timbre;
[0134] a prompt audio obtaining module configured to obtain a prompt audio, the prompt audio having a target timbre different from the original timbre;
[0135] an acoustic feature generation module configured to generate a converted acoustic feature based on the semantic feature of the to-be-converted audio and the prompt audio, using a self-attention based diffusion model; and
[0136] a converted audio generation module configured to generate a converted audio based on the converted acoustic feature, the converted audio being an audio whose timbre of the to-be-converted audio is converted to the target timbre.
[0137] Example 14. The apparatus of example 13, wherein the acoustic feature generation module comprises:
[0138] a text embedding determination module configured to determine a text embedding associated with a prompt text of the prompt audio and an original text of the to-be-converted audio;
[0139] a semantic embedding determination module configured to determine a semantic embedding associated with a prompt semantic feature of the prompt audio and the original semantic feature of the to-be-converted audio;
[0140] a global timbre embedding determination module configured to determine a global timbre embedding associated with the prompt audio;
[0141] a local timbre embedding determination module configured to determine a local timbre embedding associated with the prompt audio; and
[0142] a multi-modal embedding usage module configured to generate the converted acoustic feature based on the text embedding, the semantic embedding, the global timbre embedding, and the local timbre embedding.
[0143] Example 15. The apparatus of examples 13-14, wherein the text embedding determination module comprises:
[0144] a text encoder usage module configured to generate the text embedding based on the prompt text and the original text, using a text encoder.
[0145] Example 16. The apparatus of examples 13-15, wherein the semantic embedding determination module comprises:
[0146] a semantic encoder usage module configured to generate the semantic embedding with a semantic encoder based on the prompt semantic feature and the original semantic feature.
[0147] Example 17. The apparatus of examples 13-16, wherein the global timbre embedding determination module comprises:
[0148] a first prompt acoustic feature determination module configured to determine a prompt acoustic feature of the prompt audio;
[0149] a prompt acoustic feature usage module configured to generate the global timbre embedding with a global timbre encoder based on the prompt acoustic feature and as a whole in a time dimension.
[0150] Example 18. The apparatus of examples 13-17, wherein the local timbre embedding determination module comprises:
[0151] a second prompt acoustic feature determination module configured to determine a prompt acoustic feature of the prompt audio;
[0152] a local acoustic feature generation module configured to split the prompt acoustic feature into a plurality of local acoustic features in a time dimension; and
[0153] a local acoustic feature usage module configured to generate the local timbre embedding with a local timbre encoder based on the plurality of local acoustic features.
[0154] Example 19. The apparatus of examples 13-18, wherein the multi-modal embedding usage module comprises:
[0155] a random noise generation module configured to generate a random noise;
[0156] a random noise usage module configured to generate a noisy acoustic embedding with a noisy acoustic feature encoder based on the random noise;
[0157] a fusion embedding generation module configured to generate a fusion embedding based on the text embedding, the semantic embedding, the global timbre embedding, the local timbre embedding, and the noisy acoustic embedding;
[0158] a fusion embedding usage module configured to generate the converted acoustic feature with the self-attention based diffusion model based on the fusion embedding.
[0159] Example 20. The apparatus according to any of examples 13-19, wherein the training process of the self-attention based diffusion model comprises:
[0160] a training semantic feature determination module configured to determine semantic features of the training audio;
[0161] a first training module configured to pre-train an untrained self-attention based diffusion model based on the training audio and the semantic features of the training audio in a first training stage;
[0162] a data production module configured to generate timbre-altered semantic features based on the training audio and random audio with different timbres from the training audio using the pre-trained self-attention based diffusion model after the first training stage;
[0163] a second training module configured to train the pre-trained self-attention based diffusion model based on the training audio and the timbre-altered semantic features in a second training stage.
[0164] Example 21. The apparatus according to any of examples 13-20, wherein the first training module comprises:
[0165] a first partial audio extraction module configured to extract partial audio from the training audio;
[0166] a first predicted acoustic feature generation module configured to generate first predicted acoustic features based on the semantic features of the training audio and the partial audio using the untrained self-attention based diffusion model;
[0167] a first training acoustic feature determination module configured to determine acoustic features of the training audio; and
[0168] a first predicted acoustic feature usage module configured to pre-train the untrained self-attention based diffusion model by calculating a loss between the first predicted acoustic features and the acoustic features of the training audio.
[0169] Example 22. The apparatus according to any of examples 13-21, wherein the data production module comprises:
[0170] a timbre-altered acoustic feature generation module configured to generate timbre-altered acoustic features based on the semantic features of the training audio and the random audio using the pre-trained self-attention based diffusion model; and
[0171] a timbre-altered semantic feature generation module configured to generate the timbre-altered semantic features based on the timbre-altered acoustic features.
[0172] Example 23. The apparatus of any of examples 13-22, wherein the timbre-altered semantic feature generation module comprises:
[0173] a timbre-altered audio generation module configured to generate timbre-altered audio based on the timbre-altered acoustic features; and
[0174] a timbre-altered audio usage module configured to generate the timbre-altered semantic features based on the timbre-altered audio.
[0175] Example 24. The apparatus of any of examples 13-23, wherein the second training module comprises:
[0176] a second partial audio extraction module configured to extract partial audio from the training audio;
[0177] a second predicted acoustic feature generation module configured to generate second predicted acoustic features based on the timbre-altered semantic features and the partial audio using the pre-trained self-attention based diffusion model;
[0178] a second training acoustic feature determination module configured to determine acoustic features of the training audio; and
[0179] a second predicted acoustic feature usage module configured to train the pre-trained self-attention based diffusion model by computing a loss between the second predicted acoustic features and the acoustic features of the training audio.
[0180] Example 25. An electronic device, comprising:
[0181] a processor; and
[0182] a memory coupled with the processor, the memory having instructions stored therein that when executed by the processor cause the electronic device to perform actions comprising:
[0183] determining semantic features of audio to be converted, the audio to be converted having an original timbre;
[0184] obtaining prompt audio, the prompt audio having a target timbre different from the original timbre;
[0185] generating converted acoustic features based on the semantic features of the audio to be converted and the prompt audio using a self-attention based diffusion model; and
[0186] generating converted audio based on the converted acoustic features, the converted audio being audio having the original timbre of the audio to be converted converted to the target timbre.
[0187] Example 26. The method of example 25, wherein the semantic feature of the audio to be converted is an original semantic feature, and generating the converted acoustic feature with the self-attention based diffusion model based on the semantic feature of the audio to be converted and the prompt audio comprises:
[0188] determining a text embedding associated with a prompt text of the prompt audio and an original text of the audio to be converted;
[0189] determining a semantic embedding associated with a prompt semantic feature of the prompt audio and the original semantic feature of the audio to be converted;
[0190] determining a global timbre embedding associated with the prompt audio;
[0191] determining a local timbre embedding associated with the prompt audio; and
[0192] generating the converted acoustic feature based on the text embedding, the semantic embedding, the global timbre embedding, and the local timbre embedding.
[0193] Example 27. The method of examples 25-26, wherein determining the text embedding associated with the prompt text of the prompt audio and the original text of the audio to be converted comprises:
[0194] generating the text embedding with a text encoder based on the prompt text and the original text.
[0195] Example 28. The method of examples 25-27, wherein determining the semantic embedding associated with the prompt semantic feature of the prompt audio and the original semantic feature of the audio to be converted comprises:
[0196] generating the semantic embedding with a semantic encoder based on the prompt semantic feature and the original semantic feature.
[0197] Example 29. The method of examples 25-28, wherein determining the global timbre embedding associated with the prompt audio comprises:
[0198] determining a prompt acoustic feature of the prompt audio;
[0199] generating the global timbre embedding with a global timbre encoder based on the prompt acoustic feature and as a whole in a time dimension.
[0200] Example 30. The method of examples 25-29, wherein determining the local timbre embedding associated with the prompt audio comprises:
[0201] determining a prompt acoustic feature of the prompt audio;
[0202] segmenting the prompt acoustic feature into a plurality of local acoustic features in a time dimension; and
[0203] generating, based on the plurality of local acoustic features, the local timbre embedding using a local timbre encoder.
[0204] Example 31. The method of any of examples 25-30, wherein generating the converted acoustic feature based on the text embedding, the semantic feature embedding, and the timbre embedding comprises:
[0205] generating random noise;
[0206] generating, based on the random noise, a noisy acoustic embedding using a noisy acoustic feature encoder;
[0207] generating a fusion embedding based on the text embedding, the semantic embedding, the global timbre embedding, the local timbre embedding, and the noisy acoustic embedding;
[0208] generating, based on the fusion embedding, the converted acoustic feature using the self-attention based diffusion model.
[0209] Example 32. The method of any of examples 25-31, wherein the training process of the self-attention based diffusion model comprises:
[0210] determining semantic features of training audio;
[0211] in a first training phase, pre-training an untrained self-attention based diffusion model based on the training audio and the semantic features of the training audio;
[0212] after the first training phase, generating, based on the training audio and random audio, timbre-altered semantic features using the pre-trained self-attention based diffusion model, the random audio having a different timbre than the training audio;
[0213] in a second training phase, training the pre-trained self-attention based diffusion model based on the training audio and the timbre-altered semantic features.
[0214] Example 33. The method of any of examples 25-32, wherein in the first training phase, pre-training the self-attention based diffusion model based on the training audio and the semantic features of the training audio comprises:
[0215] extracting partial audio from the training audio;
[0216] generate, based on the semantic features of the training audio and the random audio, timbre-altered acoustic features using the pre-trained self-attention based diffusion model; and
[0217] determine acoustic features of the training audio; and
[0218] pre-train the untrained self-attention based diffusion model by computing a loss between the first predicted acoustic features and the acoustic features of the training audio.
[0219] Example 34. The method of examples 25-33, wherein, after the first training phase, generating, based on the training audio and the random audio, timbre-altered semantic features using the pre-trained self-attention based diffusion model comprises:
[0220] generate, based on the semantic features of the training audio and the random audio, timbre-altered acoustic features using the pre-trained self-attention based diffusion model; and
[0221] generate the timbre-altered semantic features based on the timbre-altered acoustic features.
[0222] Example 35. The method of examples 25-34, wherein generating the timbre-altered semantic features based on the timbre-altered acoustic features comprises:
[0223] generate timbre-altered audio based on the timbre-altered acoustic features; and
[0224] generate the timbre-altered semantic features based on the timbre-altered audio.
[0225] Example 36. The method of examples 25-35, wherein, in the second training phase, training the pre-trained self-attention based diffusion model based on the training audio and the timbre-altered semantic features comprises:
[0226] extract partial audio from the training audio;
[0227] generate, based on the timbre-altered semantic features and the partial audio, second predicted acoustic features using the pre-trained self-attention based diffusion model;
[0228] determine acoustic features of the training audio; and
[0229] train the pre-trained self-attention based diffusion model by computing a loss between the second predicted acoustic features and the acoustic features of the training audio.
[0230] Although the present disclosure has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. A method for timbre conversion, comprising: Determine the semantic features of the audio to be converted, wherein the audio to be converted has the original timbre; Obtain a prompt audio, wherein the prompt audio has a target timbre that is different from the original timbre; Based on the semantic features of the audio to be converted and the cue audio, a self-attention-based diffusion model is used to generate the converted acoustic features. as well as The converted audio is generated based on the converted acoustic features. The converted audio is an audio in which the timbre of the audio to be converted is converted into the target timbre.
2. The method according to claim 1, wherein the semantic features of the audio to be converted are original semantic features, and generating the converted acoustic features using the self-attention-based diffusion model based on the semantic features of the audio to be converted and the cue audio comprises: Determine the text embedding associated with the prompt text of the prompt audio and the original text of the audio to be converted; Determine the semantic embedding associated with the prompt semantic features of the prompt audio and the original semantic features of the audio to be converted; Determine the global timbre embedding associated with the cue audio; Determine the local timbre embedding associated with the cue audio; and The transformed acoustic features are generated based on the text embedding, the semantic embedding, the global timbre embedding, and the local timbre embedding.
3. The method of claim 2, wherein determining the text embedding associated with the prompt text of the prompt audio and the original text of the audio to be converted comprises: The text embedding is generated using a text encoder based on the prompt text and the original text.
4. The method of claim 2, wherein determining the semantic embedding associated with the prompt semantic features of the prompt audio and the original semantic features of the audio to be converted comprises: Based on the aforementioned prompt semantic features and the original semantic features, a semantic encoder is used to generate the semantic embedding.
5. The method of claim 2, wherein determining the global timbre embedding associated with the cue audio comprises: Determine the acoustic characteristics of the prompt audio; Based on the aforementioned acoustic features, a global timbre encoder is used, and the acoustic features are treated as a whole in the time dimension to generate the global timbre embedding.
6. The method of claim 2, wherein determining the local timbre embedding associated with the cue audio comprises: Determine the acoustic characteristics of the prompt audio; The acoustic features of the prompt are divided into multiple local acoustic features according to the time dimension; as well as Based on the aforementioned multiple local acoustic features, a local timbre encoder is used to generate the local timbre embedding.
7. The method of claim 2, wherein generating the transformed acoustic features based on the text embedding, the semantic embedding, and the timbre embedding comprises: Generate random noise; Based on the random noise, a noise-added feature encoder is used to generate a noise-added embedding; A fusion embedding is generated based on the text embedding, semantic embedding, global timbre embedding, local timbre embedding, and noise-added semantic embedding. Based on the fusion embedding, the transformed acoustic features are generated using the self-attention-based diffusion model.
8. The method according to claim 1, wherein the training process of the self-attention-based diffusion model includes: Determine the semantic features of the training audio; In the first training phase, the untrained self-attention-based diffusion model is pre-trained based on the training audio and its semantic features. After the first training phase, based on the training audio and random audio, a pre-trained self-attention-based diffusion model is used to generate semantic features of timbre changes, wherein the timbre of the random audio is different from that of the training audio. In the second training phase, the pre-trained self-attention-based diffusion model is trained based on the training audio and the semantic features of the timbre changes.
9. The method of claim 8, wherein in the first training phase, pre-training the self-attention-based diffusion model based on the training audio and the semantic features of the training audio comprises: Extract a portion of the audio from the training audio; Based on the semantic features of the training audio and the partial audio, the first predicted acoustic features are generated using the untrained self-attention-based diffusion model. Determine the acoustic features of the training audio; as well as The untrained self-attention-based diffusion model is pre-trained by calculating the loss between the first predicted acoustic features and the acoustic features of the training audio.
10. The method of claim 8, wherein after the first training phase, generating semantic features of the timbre change using the pre-trained self-attention-based diffusion model based on the training audio and the random audio comprises: Based on the semantic features of the training audio and the random audio, the pre-trained self-attention-based diffusion model is used to generate acoustic features with timbre changes. as well as The semantic features of the timbre change are generated based on the acoustic features of the timbre change.
11. The method of claim 10, wherein generating semantic features of the timbre change based on the acoustic features of the timbre change comprises: Audio with altered timbre is generated based on the acoustic features of the timbre change. as well as The semantic features of the timbre change are generated based on the audio with the changed timbre.
12. The method of claim 8, wherein in the second training phase, training the pre-trained self-attention-based diffusion model based on the training audio and the semantic features of the timbre change comprises: Extract a portion of the audio from the training audio; Based on the semantic features of the timbre change and the partial audio, the pre-trained self-attention-based diffusion model is used to generate a second predicted acoustic feature. Determine the acoustic features of the training audio; as well as The pre-trained self-attention-based diffusion model is trained by calculating the loss between the second predicted acoustic features and the acoustic features of the training audio.
13. An apparatus for editing audio content, the apparatus comprising: A semantic feature determination module is configured to determine the semantic features of the audio to be converted, the audio having the original timbre; The prompt audio acquisition module is configured to acquire prompt audio, wherein the prompt audio has a target timbre that is different from the original timbre; The acoustic feature generation module is configured to generate converted acoustic features based on the semantic features of the audio to be converted and the cue audio, using a self-attention-based diffusion model. as well as The audio conversion generation module is configured to generate converted audio based on the converted acoustic features, wherein the converted audio is an audio in which the timbre of the audio to be converted is converted into the target timbre.
14. An electronic device comprising: processor; as well as A memory coupled to the processor, the memory having instructions stored therein, which, when executed by the processor, cause the electronic device to perform the method according to any one of claims 1 to 12.
15. A computer program product comprising computer-executable instructions, wherein the computer-executable instructions are executed by a processor to implement the method according to any one of claims 1 to 12.