Controllable diffusion-based speech generation model

By introducing a prosody conversion engine and decoder, the problem of uncontrollable conversion of speech prosody features in existing technologies is solved, enabling independent adjustment of intonation, stress, and speech rate, and generating speech waveforms that match the target speech.

CN121753097APending Publication Date: 2026-03-27QUALCOMM INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-08-29
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing diffusion-based speech conversion technologies struggle to achieve efficient and controllable conversion of speech prosodic features, especially when maintaining the speech content unchanged, making it difficult to independently adjust details such as intonation, stress, and speech rate.

Method used

A prosody conversion engine and decoder are introduced. By extracting the prosodic features of the source and target speech, content embedding and speaker embedding are generated. A diffusion decoder is used to generate controllable speech prosodic data. Frame-level intonation control is achieved by combining the speech rate control component.

Benefits of technology

It achieves efficient and controllable conversion of speech prosodic features, and can independently adjust intonation, stress and speech rate to generate speech waveforms that match the target speech.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121753097A_ABST
    Figure CN121753097A_ABST
Patent Text Reader

Abstract

Systems and techniques described herein relate to a diffusion-based model for generating converted speech from source speech based on target speech. For example, a device may extract first rhythm data from input data, and may generate content embedding based on the input data. The device may extract second rhythm data from the target speech, generate a speaker insert from the target speech, and generate a rhythm insert from the second rhythm data. The device may generate converted rhythm data based on the first rhythm data and the rhythm embedding. The device may then generate a converted spectrogram based on the converted rhythm data, the speaker embedding, and the content embedding.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present disclosure generally relates to processing speech signals. For example, aspects of the present disclosure relate to a diffusion-based model for generating a transformed speech from a source speech based on a target speech (e.g., the transformed speech has prosodic characteristics of the target speech but maintains the same content as the source speech). BACKGROUND

[0002] Diffusion-based speech transformation is a technique that includes an encoder and decoder structure, where a source speech is provided to a mean speech encoder to generate a content embedding. The source speech and a target speech are provided to a speaker encoder to generate a speaker embedding. The content embedding and the speaker embedding are provided to a diffusion decoder that synthesizes a spectrogram depending on a condition vector associated with the content embedding and the speaker embedding. This approach depends on general speaker characteristics and utilizes a single embedding vector for speech transformation. SUMMARY

[0003] Systems and techniques are described herein for providing controllable diffusion-based speech generation models that introduce a transformation process that provides additional controllability to prosodic features of speech. According to some aspects, an apparatus for generating output speech from input data is provided. The apparatus includes one or more memories configured to store the input data and one or more processors coupled to the one or more memories and configured to: extract first prosodic data from the input data; generate a content embedding based on the input data; extract second prosodic data from a target speech; generate a speaker embedding from the target speech; generate a prosodic embedding from the second prosodic data; and generate transformed prosodic data based on the first prosodic data and the prosodic embedding.

[0004] In some aspects, a method for generating output speech from input data is provided. The method includes: extracting first prosodic data from the input data; generating a content embedding based on the input data; extracting second prosodic data from a target speech; generating a speaker embedding from the target speech; generating a prosodic embedding from the second prosodic data; generating transformed prosodic data based on the first prosodic data and the prosodic embedding; and generating a transformed spectrogram based on the transformed prosodic data, the speaker embedding, and the content embedding.

[0005] In some aspects, a non-transitory computer-readable medium having instructions stored thereon that, when executed by one or more processors, cause the one or more processors to be configured to: extract first prosody data from input data; generate a content embedding based on the input data; extract second prosody data from a target speech; generate a speaker embedding from the target speech; and generate a prosody embedding from the second prosody data; generate converted prosody data based on the first prosody data and the prosody embedding.

[0006] In some aspects, an apparatus is provided that includes means for extracting first prosody data from input data; means for generating a content embedding based on the input data; means for extracting second prosody data from a target speech; means for generating a speaker embedding from the target speech; means for generating a prosody embedding from the second prosody data; means for generating converted prosody data based on the first prosody data and the prosody embedding; and means for generating converted spectrograms based on the converted prosody data, the speaker embedding, and the content embedding.

[0007] In some aspects, one or more of the apparatuses described herein is, includes, and / or operates as an extended reality (XR) device or system (e.g., a virtual reality (VR) device, an augmented reality (AR) device, or a mixed reality (MR) device), a mobile device or wireless communication device (e.g., a mobile telephone or other mobile device), a wearable device (e.g., a network-enabled watch or other wearable device), a camera, a personal computer, a laptop computer, a vehicle or computing device or component of a vehicle, a server computer or server device (e.g., an edge- or cloud-based server, a personal computer acting as a server device, a mobile device such as a mobile phone acting as a server device, an XR device acting as a server device, a vehicle acting as a server device, a network router, or other device acting as a server device), another device, or a combination thereof. In some aspects, the apparatus includes one or more cameras for capturing one or more images. In some aspects, the apparatus also includes a display for displaying one or more images, notifications, and / or other displayable data. In some aspects, the apparatuses described above can include one or more sensors (e.g., one or more inertial measurement units (IMUs), such as one or more gyroscopes, one or more gyro tests, one or more accelerometers, any combination thereof, and / or other sensors).

[0008] This Summary is intended to identify key or essential features of the claimed subject matter, but it is not intended to identify key or essential features of the claimed subject matter. This subject matter should be understood from reading the entire specification of this patent, including any claims, the proper scope of which should be determined with reference to the entire specification, any or all drawings, and the claims.

[0009] The foregoing and other features and aspects will become more apparent from the following description, claims and accompanying drawings. Attached Figure Description

[0010] The exemplary aspects of this application are described in detail below with reference to the following figures:

[0011] Figure 1 This is a conceptual diagram illustrating various speech generation models according to various aspects of this disclosure;

[0012] Figure 2 This is a conceptual diagram illustrating examples of diffusion models according to various aspects of this disclosure;

[0013] Figure 3A Baseline speech conversion systems according to various aspects of this disclosure are illustrated;

[0014] Figure 3B It is a block diagram of the diffusion decoder training process according to various aspects of this disclosure;

[0015] Figure 4A This is a conceptual diagram illustrating an overall system for generating converted speech according to various aspects of this disclosure;

[0016] Figure 4B A conceptual diagram illustrating a Hubert model for generating conversion rates according to various aspects of this disclosure is shown;

[0017] Figure 5A Examples of training and inference schemes for highly controllable diffusion-based speech generation models are illustrated according to various aspects of this disclosure;

[0018] Figure 5B Examples of training and inference schemes for highly controllable diffusion-based speech generation models are illustrated according to various aspects of this disclosure;

[0019] Figure 6 Examples of training and inference schemes for highly controllable diffusion-based speech generation models are illustrated according to various aspects of this disclosure;

[0020] Figure 7 An example process is illustrated using a controllable diffusion-based speech generation model according to various aspects of this disclosure;

[0021] Figure 8 These are illustrations of example system architectures for implementing certain aspects described herein, based on various aspects of this disclosure; and

[0022] Figure 9 Example neural networks according to various aspects of this disclosure are illustrated. Detailed Implementation

[0023] Certain aspects of this disclosure are provided below. Some of these aspects may be applied independently, and some may be applied in combination, as will be apparent to those skilled in the art. Specific details are set forth in the following description for purposes of explanation in order to provide a thorough understanding of the various aspects of this application. However, it will be apparent that various aspects may be practiced without these specific details. The accompanying drawings and descriptions are not intended to be limiting.

[0024] The following description provides only exemplary aspects and is not intended to limit the scope, applicability, or configuration of this disclosure. Rather, the following description of the exemplary aspects will provide those skilled in the art with a description that can be used to implement the exemplary aspects. It should be understood that various changes can be made to the function and arrangement of the elements without departing from the scope of this application as set forth in the appended claims.

[0025] Figure 1 This is a conceptual diagram illustrating various speech generation models 100 according to some aspects of this disclosure. A speech generation model can represent a model that generates speech from text, speech input, or other types of input. Generally, speech generation model 100 generates desired speech through appropriate adjustments. For example, test sequence 102 can be provided to text-to-speech (TTS) model 104, which analyzes the text input and generates a speech waveform 106. Speech conversion (VC) model 112 can receive speech 108 (e.g., a speech waveform) and can receive speaker information 110 (e.g., data about the characteristics of human speech) and can convert speech 108 into speech waveform 106. In one example, speaker information may include, for example, prosodic data, which includes the characteristics of the speaker (e.g., speaker information 110). VC model 112 can adjust or modify the characteristics of speech 108 to match speaker information 110 by adjusting the prosody, velocity, pitch, or any other characteristics of speech 108. Typically, the content spoken (e.g., the actual words or sentences in speech 108) does not change, so that the same content (e.g., the words or concepts in the content) is provided in speech waveform 106, but with different characteristics.

[0026] In the example system, speech 108 may also be provided to a style / emotion transfer model 116, which receives style vectors / emotion recognition 114. In some aspects, the generated speech waveform 106 alters the style of speech 108, such as from happy to sad, or from a normal state to an angry or surprised state. Style vectors / emotion recognition 114 can change the style or emotion from a first state to a second state. This disclosure provides various methods for transferring speech from a first state to a second state. Part of this disclosure includes the ability to use a highly controllable transfer engine or module. For example, the transfer engine or module can provide frame-level intonation control and can utilize prosodic features such as fundamental frequency f0, energy, and speech-associated features. In one aspect, the transfer engine or module can provide speech rate control without a conventional automatic speech recognition model.

[0027] Figure 2 This is a conceptual diagram illustrating an example of a diffusion model 200 according to some aspects of this disclosure. Diffusion models are typically used in computer vision contexts where high-quality images can be generated from text prompts. Diffusion models learn a progressive generation process rather than a one-step generation process. In the forward diffusion process 202, for each step or timestamp associated with image 206 from time 0 to time T, Figure 2 Various mathematical operations 204 are illustrated. In the forward diffusion process 202, model 200 artificially generates noise to produce more noise samples, as shown in the process from timestamp 0 to timestamp T. The diffusion model 200 artificially adds noise to the original image 206 at each step. In the backward diffusion process 212, mathematical operations 208 utilize UNet 210 (e.g., a convolutional neural network with a "U"-shaped architecture) to allow the model to estimate the noise components at each step from timestamp T (the noisy image) to refine and reduce the noise at each step, thereby generating the original image 206 shown at timestamp 0. UNet is a deep learning architecture for semantic segmentation. In one example, UNet 210 may include, for example, Figure 2The shrinking and expanding paths are shown. The shrinking path follows a typical architecture of a convolutional network. The shrinking path may consist of repeated applications of two 3×3 convolutions (without padding), each followed by a Rectified Linear Unit (ReLU) and a 2×2 max-pooling operation that downsamples and / or compresses the information using a stride of 2. "Stride of 2" indicates that after each operation of the kernel or filter (e.g., a max-pooling kernel / filter), the kernel or filter is moved two positions. At each downsampling step, the structure doubles the number of feature channels. Each step in the expanding path consists of upsampling of the feature map, followed by a 2×2 convolution that halves the number of feature channels ("up convolution"), a concatenation with the corresponding cropped feature map from the shrinking path, and two 3×3 convolutions, each followed by a ReLU. Cropping may be performed because boundary pixels are lost in each convolution. At the final layer, 1×1 convolutions are used to map each 64-component feature vector to the desired number of classes. Overall, a sample Unet network has 23 convolutional layers. Using a diffusion model process, a realistic and easily controllable generative model can be obtained.

[0028] Recently, diffusion model 200 has been used for generative image modeling. Several text-to-image generation services also exist, such as Dall-E or Midjourney. These services, for example, generate images from text descriptions, much like the image version of ChatGPT. The core process of diffusion model 200 is a multi-step generation process, which includes, for example... Figure 2 The forward diffusion process 202 and the reverse diffusion process 212 are shown.

[0029] The diffusion model 200 can also be successfully applied to speech generation models. Figure 3AA baseline speech conversion system 300 is illustrated. In some cases, the baseline speech conversion system 300 may be a diffusion speech conversion (DiffVC) system. The baseline speech conversion system 300 includes an average speech encoder 306 and a diffusion decoder 322. The average speech encoder 306 (also referred to as a content encoder) takes an average spectrum 304 from the source speech 302 as input, which is different from text or content-related features. The speaker encoder 316 receives the target speech 314 and the source speech 302 and generates a speaker embedding 318. The embeddings disclosed herein generally represent the transformation of data from one context to another for further processing. Embeddings are typically dense digital representations of real-world objects (e.g., such as audio) and relationships and may be represented as vectors. The diffusion decoder 322 generates an output spectrogram 324 from the average spectrogram 312 or content embedding 320 and speaker information or speaker embedding 318. In natural language processing, word embeddings are representations of words. Embeddings can be used for text analysis. Typically, this representation is a real-valued vector that encodes the meaning of a word so that words that are expected to be closer in the vector space are similar in meaning.

[0030] Figure 3A Some of the flow lines in the diagram relate to the training process. For example, flow line 308 illustrates the data flow during the training phase, and flow line 310 relates to the flow during the inference or prediction phase when the baseline speech conversion system 300 generates the converted speech.

[0031] Figure 3B The training process of the diffusion decoder is illustrated in 350. During training, the source speech is fed to the average speech encoder 306. The model attempts to reconstruct itself from content-only features or content embeddings 320 and speaker embeddings 318, which is consistent with... Figure 3A The process shown in the transformation process is the same. Therefore, the process of learning reconstruction is similar to the process of learning transformation, which enables the model to handle unseen speaker information. In some respects, training in any context can be performed in real time on the device.

[0032] The diffusion decoder can be, for example Figure 3A The diffusion decoder 322. Speaker embedding 352, first spectrogram 354, and second spectrogram 356 can be provided to cascade module 358. Value 370 can also be provided to cascade module 358 for cascading with other data to provide output data to U-Net 360, which can be a noise estimator. A noise image 362 can be generated. Noise scheduler 368 generates noise 366 from the values ​​and determines a mean squared error loss 364 value between noise 366 and noise image 362. The mean squared error loss 364 is used to train diffusion decoder 322.

[0033] Figure 3A and Figure 3B The method shown has limitations. One limitation is that, apart from the content embedding 320, the single speaker embedding 318 is the only factor controlling the output or converted speech 326. The (single) speaker embedding 318 spans the entire time and frequency domains for adjustment. The speaker encoder 316 can capture global speaker or speech characteristics. However, another limitation is that... Figure 3A In the structure shown, it is difficult to further control the speech conversion process, such as intonation, stress, and speech rate.

[0034] Figure 4A An illustrated method for providing a highly controllable diffusion-based speech generation model is presented. A system 400 for generating converted speech is disclosed, comprising a prosodic transformation engine 422 and a decoder 430. Source speech 402 is provided to a first prosodic extractor 408 and a content encoder 406. Source speech 402 is also provided to a global speech rate predictor 412. Target speech 404 is provided to the global speech rate predictor 412, a second prosodic extractor 416, and a speaker encoder 420. Content encoder 406 generates content embedding 432. The first prosodic extractor 408 generates first output data 410, which may be one or more of the original prosodic features, such as a fundamental frequency f0, an energy value (e.g., a logarithmic energy value loge), and a velocity. The global speech rate predictor 412 receives source speech 402 and target speech 404 and generates a reference speech rate or R. SR 414. A second prosodic extractor 416 generates second output data 418, which may be one or more of the original prosodic features, such as a fundamental frequency f0, an energy value (e.g., a logarithmic energy value loge), and velocity. A speaker encoder 420 generates a speaker embedding 436 from the target speech 404. A prosodic transformation engine 422 includes a prosodic encoder 428 that generates the prosodic embedding 426 from the second output data 418. The prosodic encoder 428 captures global prosodic features and may receive the original prosodic features and low-frequency band spectral information (second output data 418) as input.

[0035] Reference speech rate or R can be generated using HuberT-based unit and duration prediction components. SR A value of 414. For example, for speech rate control, such as... Figure 4BAs shown, the system can use the HuBERT model (a self-supervised model) to obtain the unit values ​​and duration values ​​of the source and target speech. Details about the HuBERT model can be found in "HuBERT: Self-Supervised Speech Representation Learning by Masking Predicted Hidden Units" by Hsu et al., June 14, 2021 (arXiv: 2106.07447v1), which is incorporated herein by reference. "Unit" means speech lexical index, which is similarly used as a phoneme, and duration is its duration based on the number of speech frames. Once the unit-duration pairs are obtained, the speech rate can be calculated using the average duration. Figure 4B An example of the process is shown.

[0036] The prosody conversion engine 422 also includes a prosody conversion model 424 that receives first output data 410 (e.g., speech) and a prosodic embedding 426, and generates third output data 434, which may be, for example, a corrected fundamental frequency f'0 and a corrected energy value logE'. The prosodic embedding 426 may also be characterized as a global prosodic embedding. The second output data 418 may include the original prosodic features and a low-frequency mel spectrogram. The mel spectrogram has two important changes compared to a conventional spectrogram that plots frequency versus time. The mel spectrogram uses the Mel scale (or melody scale) instead of frequencies on the y-axis, and when using color, the mel spectrogram uses a decibel scale instead of amplitude to indicate color. The use of the mel spectrogram is to adjust the data to be more consistent with how humans perceive sound, since most of the sounds humans can hear are concentrated in a narrow frequency range.

[0037] Next, decoder 430 may include diffusion decoder 438, which receives content embedding 432 from content encoder 406, third output data 434 from prosodic transformation model 424, and speaker embedding 436 from speaker encoder 420. Diffusion decoder 438 generates a transformed spectrogram 440, which may be provided to speech rate control component 442. Speech rate control component 442 may also receive R... SR 414 (e.g., a conversion rate between source speech 402 and target speech 404), which generates a rate-controlled spectrogram that can be provided to a vocoder 444 (e.g., a neural vocoder) capable of generating converted speech 446. The converted speech 446 represents a synthesized waveform of the speech spectrum from the converted spectrogram 440, at a speech control value generated by the speech rate control component 442. The speech rate control component 442 is based on the conversion rate or R... SR414 is used to manipulate the speech rate of the converted spectrogram 440. The prosodic transformation model 424 uses prosodic embeddings 426 from the prosodic encoder 428 to convert the raw prosodic features from the source speech 402 into prosodic features of the target speech 404. The diffusion decoder 438 may also be explicitly a non-diffusion decoder. The decoder is not limited to the diffusion decoder 438 herein, but may include other types of decoders.

[0038] Figure 4A The main difference between the proposed method and the baseline model lies in the system 400 used to generate the transformed speech, or speech generation model, which uses frame-level prosodic features as an additional control factor in the diffusion model within the diffusion decoder 438. This model comprises two parts: a prosodic transformation engine 422 and a decoder 430. In the prosodic transformation engine 422, the original prosodic features of the source speech 402 are extracted, and those features are transformed to match the features of the target speech 404.

[0039] Figure 4B Figure 450 illustrates the operation of a HuBERT-based cell and direction prediction model 452. The input wave value “wav” can have a frequency f. wav The HuBERT-based unit and direction prediction model 452 generates predictions related to speech rate control using source and target speaker data. The goal is to obtain the speech rate conversion rate or R... SR 414. The output of the HuBERT-based cell and orientation prediction model 452 is the frequency value f. unit For example, 6,6,1,1,5,5,5. This value can be converted into a mel spectrogram f. mel The frequency values ​​in the data are, for example, 6,6,1,1,1,5,5,5,5. The data can then be divided into unit values ​​u (6,1,5) and direction d. u 2,3,4. R SR 414 can be equal to E[d] u_source ] / E[d u_target The output of the diffusion decoder 438 can utilize R. SR 414 Resampling. Therefore, based on predictions of the speech rates of both the source and target speakers, the global speech rate predictor 412 can provide R... used by the speech rate control component 442. SR 414.

[0040] Figure 5A Examples are shown for use Figure 4A A block diagram of the training and inference scheme of the prosodic encoder 428 500. Figure 5A and Figure 5BThe training and inference schemes of the proposed model are illustrated. In one aspect, the method uses pre-trained modules for the speaker encoder and content encoder. The prosodic encoder can be trained in a self-supervised manner, similar to the process associated with autoencoders. This process attempts to extract intermediate embeddings while reconstructing itself. Prosodic features and low-frequency band mel spectrograms are used to focus the model on the prosodic aspects of the data. Except that the model's input and output are prosodic features instead of mel spectrograms, the prosodic conversion model is trained similarly to other speech conversion systems.

[0041] The example process provides prosodic transformation training for the prosodic encoder 428. Figure 5A A source speech 402 or target speech 404 is shown provided to a first prosodic extractor 408, which generates output data 502, such as one or more of a fundamental frequency f0, energy, and velocity. The data may also include a low-frequency band mel spectrogram. The data is provided to a prosodic encoder 428 and a loss engine 506. The prosodic encoder 428 generates a prosodic embedding 526, which is provided to a decoder 438. The prosodic embedding 526 may include frame-level and sentence-level data. The decoder 438 generates output data 504, which includes one or more of a fundamental frequency f0, energy, and velocity. The output data 504 may also include other information, such as a low-frequency band mel spectrogram. The output data 504 may be provided to the loss engine 506. The loss engine 506 can generate or determine a loss (e.g., a loss value) between the output data 502 and the output data 504. This loss can then be used to train the prosody transition engine 422 (e.g., by performing backpropagation to tune the parameters of the prosody transition engine 422, such as the weights).

[0042] Figure 5B Examples are shown for use Figure 4A Another block diagram of the training and inference scheme of the prosodic encoder 428 520. Figure 5BA source speech 402 or a target speech 404 is shown provided to a prosodic extractor 408, which generates output data 522, such as a fundamental frequency f0, energy, and velocity. The output data 522 may also include a low-frequency band mel spectrogram. The output data 522 is provided to a prosodic encoder 428 and a loss engine 528. The prosodic encoder 428 generates a prosodic embedding 526, which is provided to a prosodic transformation model 424. The prosodic embedding 526 may be a frame-by-frame representation of the target speech 404 or the source speech 402. In some cases, the target speech 404 may be test data, and the source speech 402 may be training data. In some aspects, data may be provided to the prosodic encoder 428, which generates a sentence-level (and / or frame-level) representation 524. The sentence-level (and / or frame-level) representation 524 may also be provided to the prosodic transformation model 424. Prosody transformation model 424 receives prosody embeddings 526 (e.g., frame-by-frame representations) and, in some cases, sentence-level (and / or frame-level) representations 524. Based on the prosody embeddings 526 (e.g., frame-by-frame representations) and, in some cases, the sentence-level (and / or frame-level) representations 524, the prosody transformation model 424 can generate third output data 534, which may include one or more of a fundamental frequency f0, energy, and velocity, as well as the transformed speech. The output data 534 may be provided to a loss engine 528. The loss engine 528 can generate or determine a loss (e.g., a loss value) between the output data 502 and the output data 504. The loss generated by the loss engine 528 can be used to train the prosody transformation engine 422 (e.g., by performing backpropagation to tune the parameters of the prosody transformation engine 422, such as weights).

[0043] Figure 6A highly controllable diffusion-based speech generation model training and inference scheme 600 is illustrated. For decoder training, in one aspect, the method may include freezing the encoder weights during decoder training. Different utterances by the same speaker can be used as reference speech. Thus, source speech 402 is provided to content encoder 406, which produces content embedding 432. Source speech 402 is provided to first prosodic extractor 408, which produces second output data 418 (e.g., fundamental frequency f0 and energy value). Source speech 402 is provided to speaker encoder 420, which produces speaker embedding 436. Content embedding 432, second output data 418, and speaker embedding 436, together with noise scheduler t 602, are provided to diffusion decoder 438. Diffusion decoder 438 produces estimated noise 604, which is provided to mean squared error (MSE) loss engine 606. The source speech 402 is also provided to the MSE loss engine 606 as a noise scheduler t 608, which can be live noise 610. The output of the MSE loss engine 606 can be used to train the diffusion decoder 438. The source speech 402 is converted into content embedding 432 via the content encoder 406. The source speech 402 is converted into speaker embedding 436 via the speaker encoder 420. The source speech 402 is converted into prosody or second output data 418 (e.g., f0, LogE) via the first prosodic extractor 408, ultimately resulting in reconstructed speech. The MSE loss engine 606 can compute the loss between (N(source, t), N(reconstructed, t)).

[0044] In reasoning, such as Figure 4A As shown, target speech 404 from the target speaker is given as reference speech. Source speech 402 is used to generate content embedding 432. Target speech 404 is used to generate speaker embedding 436 and prosodic embedding 426. The transformed spectrogram 440 from the diffusion decoder 438 can be a synthesized spectrogram using content embedding 432, speaker embedding 436, and prosodic embedding 426. Also as... Figure 4A As shown, the speech rate control component 442 and the vocoder 444 (e.g., a neural vocoder) are used to acquire speech output or converted speech 446.

[0045] Figure 7 This is a flowchart illustrating an example process 700 for generating converted speech from input data (such as target speech 402, 404). Process 700 may include any one or more steps disclosed herein. Process 700 may be performed using a system, apparatus, or computing device (or components thereof, such as chipsets, one or more processors, etc.) (collectively, the system). In some aspects, the system may include Figure 4AThe system comprises a content encoder 406, one or more prosodic extractors 408, 416, a global speech rate predictor 412, a prosodic conversion engine 422 having a prosodic encoder 428 and a prosodic conversion model 424, a speaker encoder 420, a decoder 430 having a diffusion decoder 438 (or a decoder of another type), a speech rate control component 442 and a vocoder 444, a computing system 800, or a combination thereof.

[0046] At operation 702, the system (or a component thereof) can extract first prosodic data from the input data. In some aspects, the input data may include one or more of speech data, text data, or other types of data. The first prosodic data may include one or more of a fundamental frequency, energy value, and velocity value.

[0047] At operation 704, the system (or its components) can generate content embeddings based on input data.

[0048] At operation 706, the system (or its components) can extract second prosodic data from the target speech. In some aspects, the second prosodic data may include one or more of a fundamental frequency, energy value, and velocity value.

[0049] At operation 708, the system (or a component thereof) can generate speaker embeddings from target speech.

[0050] At operation 710, the system (or a component thereof) can generate a prosodic embedding from the second prosodic data.

[0051] At operation 712, the system (or its components) may generate converted prosodic data based on the first prosodic data and the prosodic embedding. In some aspects, the input data includes speech data. In such aspects, the system may include one or more microphones configured to capture speech data. In some cases, the system may include one or more speakers configured to output speech data including the converted prosodic data.

[0052] In some respects, the system (or its components) can generate transformed spectrograms based on transformed prosodic data, speaker embeddings, and content embeddings.

[0053] In some respects, the system (or its components) can be accessed via a decoder (e.g., Figure 4A The diffusion decoder 438 generates a converted spectrogram based on the converted prosodic data, speaker embeddings, and content embeddings. The decoder may include a diffusion decoder, a non-diffusion decoder, or a different type of diffusion decoder.

[0054] In some aspects, the system (or its components) may generate a predicted global speech rate based on transformed prosodic data, and generate the speech rate of the transformed spectrogram via a rate control engine (e.g., speech rate control component 442); and via a vocoder (e.g., Figure 4A The vocoder (444) generates converted speech based on the input data. In some aspects, the rate control engine can manipulate the speech rate based on the predicted speed. In some cases, the vocoder (e.g., Figure 4A The vocoder (444) may include a neural vocoder or some other type of vocoder. The vocoder analyzes and synthesizes human speech signals. In some respects, the vocoder examines speech by measuring how its spectral characteristics change over time. The vocoder generates a series of signals representing those frequencies at any given time when the user speaks, or based on the input data. This signal can be divided into multiple frequency bands, and the signal level at each band gives an instantaneous representation of the spectral energy content. To recreate speech, the vocoder reverses the process of passing data through stages that filter the frequency content based on a series of digital steps from the original recording to a broadband noise source.

[0055] In some aspects, the system (or its components) may extract first prosodic data from input data via a first prosodic extractor engine. The system (or its components) may generate content embeddings based on the input data via a content encoder. The system (or its components) may further extract second prosodic data from the target speech via a second prosodic extractor engine. The system (or its components) may generate speaker embeddings from the target speech via a speaker encoder.

[0056] In some aspects, the system (or its components) may generate prosodic embeddings from second prosodic data via a prosodic encoder (e.g., prosodic encoder 428). The system (or its components) may generate transformed prosodic data based on the first prosodic data and the prosodic embeddings via a prosodic transformation engine (e.g., prosodic transformation engine 422). The system (or its components) may generate a transformed spectrogram based on the transformed prosodic data, speaker embeddings, and content embeddings via a decoder (e.g., diffusion decoder 438). In some cases, the prosodic encoder may generate prosodic embeddings at one or more of the frame and / or sentence levels, or at different granularities, thereby enhancing the controllability of prosodic features.

[0057] In some aspects, the system (or its components) may be or may include a decoder (e.g., a diffusion decoder 438). In such aspects, the decoder may be configured to synthesize a speech spectrum conditioned on content embeddings, speaker embeddings, and transformed prosodic data.

[0058] In some aspects, the system (or a component thereof) may be or may include a prosodic encoder (e.g., prosodic encoder 428). In such aspects, the prosodic encoder may be configured to generate prosodic embeddings at the frame level to enable frame-level intonation control.

[0059] In some respects, the system (or its components) can generate the speech rate of a converted spectrogram based on converted prosodic data via a rate control engine, independent of an automatic speech recognition model.

[0060] In some respects, a non-transitory computer-readable medium (e.g., Figure 8 The memory 815, ROM 820, RAM 825, or cache 811 has instructions stored thereon that, when executed by one or more processors (e.g., processor 812), configure the one or more processors to: extract first prosodic data from input data; generate a content embedding based on the input data; extract second prosodic data from target speech; generate a speaker embedding from the target speech; and generate a prosodic embedding from the second prosodic data; generate converted prosodic data based on the first prosodic data and the prosodic embedding; and generate a converted spectrogram based on the converted prosodic data, the speaker embedding, and the content embedding.

[0061] In some aspects, an apparatus may include: components for extracting first prosodic data from input data; components for generating content embeddings based on the input data; components for extracting second prosodic data from target speech; components for generating speaker embeddings from the target speech; components for generating prosodic embeddings from the second prosodic data; components for generating converted prosodic data based on the first prosodic data and the prosodic embeddings; and components for generating converted spectrograms based on the converted prosodic data, the speaker embeddings, and the content embeddings. Components for performing any of the above functions may include... Figure 4A The system 400 for generating converted speech includes a content encoder 406, one or more prosodic extractors 408, 416, a global speech rate predictor 412, a prosodic conversion engine 422 with a prosodic encoder 428 and a prosodic conversion model 424, a speaker encoder 420, a decoder 430 with a diffusion decoder 438, a speech rate control component 442 and a vocoder 444, a computing system 800, or a combination thereof.

[0062] The system, apparatus, or computing device configured to perform process 700 may include any suitable device, such as a mobile device (e.g., a mobile phone), a desktop computing device, a tablet computing device, an XR device (e.g., a VR headset, an AR headset, AR glasses, etc.), a wearable device (e.g., a connected watch or smartwatch, or another wearable device), a server computer, a vehicle (e.g., an autonomous vehicle) or a vehicle's computing device, a robotic device, a laptop computer, a smart TV, a camera, and / or any other computing device with the resource capability to perform the processes described herein (including process 700 and / or any other processes described herein). In some cases, the computing device or apparatus may include various components, such as one or more input devices, one or more output devices, one or more processors, one or more microprocessors, one or more microcomputers, one or more cameras, one or more sensors, and / or other components configured to perform the steps of the processes described herein. In some examples, the computing device may include a display, a network interface configured to communicate and / or receive data, any combination thereof, and / or other components. The network interface may be configured to communicate and / or receive Internet Protocol (IP)-based data or other types of data.

[0063] A component capable of implementing a computing device in a circuit. For example, the component may include electronic circuitry or other electronic hardware, and / or may be implemented using electronic circuitry or other electronic hardware, which may include one or more programmable electronic circuits (e.g., a microprocessor, graphics processing unit (GPU), digital signal processor (DSP), central processing unit (CPU), and / or other suitable electronic circuitry), and / or may include computer software, firmware, or any combination thereof for performing the various operations described herein, and / or may be implemented using computer software, firmware, or any combination thereof for performing the various operations described herein.

[0064] Process 700 is illustrated as a logic flowchart, the operations of which represent a sequence of operations that can be implemented in hardware, computer instructions, or combinations thereof. In the context of computer instructions, each operation represents a computer-executable instruction stored on one or more computer-readable storage media that, when executed by one or more processors, performs the described operation. Generally, computer-executable instructions include routines, programs, objects, components, data structures, etc., that perform a specific function or implement a specific data type. The order in which the operations are described is not intended to be construed as limiting, and any number of described operations can be combined in any order and / or in parallel to implement the process.

[0065] Additionally, process 700 and / or any other process described herein may be executed under the control of one or more computer systems configured with executable instructions, and may be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) that executes jointly on one or more processors, by hardware, or a combination thereof. As noted above, the code may be stored on a computer-readable or machine-readable storage medium, for example, in the form of a computer program comprising multiple instructions executable by one or more processors. The computer-readable or machine-readable storage medium may be non-transitory.

[0066] Figure 8 These are illustrations illustrating examples of systems used to implement certain aspects of this disclosure. Specifically, Figure 8 An example of a computing system 800 is illustrated. This computing system can be any computing device, such as a computing system, a camera system, or any component thereof, wherein the components of the system communicate with each other using connection 805. Connection 805 can be a physical connection using a bus, or a direct connection to processor 812, such as in a chipset architecture. Connection 805 can also be a virtual connection, a networking connection, or a logical connection.

[0067] In some examples, the computing system 800 is a distributed system, wherein the functionality described herein may be distributed across a data center, multiple data centers, a peer-to-peer network, etc. In some examples, one or more of the described system components represent a plurality of such components, each performing some or all of the functionality targeted by the described component. In some examples, the components may be physical or virtual devices.

[0068] Example system 800 includes at least one processing unit (CPU or processor) 812 and a connection 805 that couples various system components, including system memory 815 (such as read-only memory (ROM) 820 and random access memory (RAM) 825), to processor 812. Computing system 800 may include a cache 811 of high-speed memory that is directly connected to, closely proximate to, or integrated into processor 812.

[0069] Processor 812 may include any general-purpose processor and hardware or software services (such as services 832, 834, and 836 stored in storage device 830 and configured to control processor 812), as well as dedicated processors in which software instructions are incorporated into the actual processor design. Processor 812 may be a substantially completely independent computing system containing multiple cores or processors, buses, memory controllers, caches, etc. Multi-core processors may be symmetric or asymmetric.

[0070] To enable user interaction, the computing system 800 includes an input device 845 that can represent any number of input mechanisms, such as a microphone for voice, a touch-sensitive screen for gesture or graphical input, a keyboard, a mouse, motion input, voice input, etc. The computing system 800 may also include an output device 835 that can be one or more of multiple output mechanisms. In some instances, a multi-mode system allows the user to provide multiple types of input / output to communicate with the computing system 800. The computing system 800 may include a communication interface 840, which typically governs and manages user input and system output.

[0071] The communication interface can perform or facilitate the receiving and / or transmitting of wired or wireless communications using wired and / or wireless transceivers, including utilizing audio jacks / plugs, microphone jacks / plugs, Universal Serial Bus (USB) ports / plugs, Apple... ® Lightning ® Ports / plugs, Ethernet ports / plugs, fiber optic ports / plugs, dedicated wired ports / plugs, Bluetooth ® Wireless signal transmission, Bluetooth ® Low-power (BLE) wireless signal transmission, IBEACON ® Wireless signal transmission, radio frequency identification (RFID) wireless signal transmission, near field communication (NFC) wireless signal transmission, dedicated short range communication (DSRC) wireless signal transmission, 802.11 Wi-Fi wireless signal transmission, wireless local area network (WLAN) signal transmission, visible light communication (VLC), microwave access global interoperability (WiMAX), infrared (IR) wireless signal transmission, public switched telephone network (PSTN) signal transmission, integrated services digital network (ISDN) signal transmission, 3G / 4G / 5G / LTE cellular data network wireless signal transmission, ad hoc network signal transmission, radio wave signal transmission, microwave signal transmission, infrared signal transmission, visible light signal transmission, ultraviolet light signal transmission, wireless signal transmission along the electromagnetic spectrum, or those communications in some combination thereof.

[0072] The communication interface 840 may also include one or more Global Navigation Satellite System (GNSS) receivers or transceivers for determining the location of the computing system 800 based on one or more signals received from one or more satellites associated with one or more GNSS systems. GNSS systems include, but are not limited to, the U.S. Global Positioning System (GPS), Russia's Global Navigation Satellite System (GLONASS), China's BeiDou Navigation Satellite System (BDS), and Europe's Galileo GNSS. There are no limitations on operation on any particular hardware arrangement, and therefore the basic features here can be easily replaced to obtain improved hardware or firmware arrangements as they are developed.

[0073] Storage device 830 may be a non-volatile and / or non-transitory and / or computer-readable storage device, and may be a hard disk or other type of computer-readable medium capable of storing data accessible by a computer, such as magnetic tape, flash memory cards, solid-state storage devices, digital multifunction disks, cartridges, floppy disks, hard disks, magnetic tapes, magnetic stripes, any other magnetic storage media, flash memory, memristor memory, any other solid-state storage, CD-ROM discs, rewritable CD discs, DVD discs, Blu-ray discs (BDD discs), holographic discs, another optical medium, secure digital (SD) cards, microSD cards, Memory Sticks. ® Cards, smart card chips, EMV chips, Subscriber Identity Module (SIM) cards, mini / micro / nano / micro SIM cards, another integrated circuit (IC) chip / card, random access memory (RAM), static RAM (SRAM), dynamic RAM (DRAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash EPROM, cache memory (L1 / L2 / L3 / L4 / L5 / L#), resistive random access memory (RRAM / ReRAM), phase change memory (PCM), spin-transfer torque RAM (STT-RAM), another memory chip or cassette and / or combinations thereof.

[0074] Storage device 830 may include software services, servers, services, etc., which enable the system to perform functions when the code defining such software is executed by processor 812. In some examples, hardware services performing specific functions may include software components for performing functions stored in a computer-readable medium connected to necessary hardware components such as processor 812, connection 805, output device 835, etc. The term "computer-readable medium" includes, but is not limited to, portable or non-portable storage devices, optical storage devices, and various other media capable of storing, containing, or carrying instructions and / or data. Computer-readable media may include non-transitory media in which data can be stored and which does not include carrier waves and / or transient electronic signals propagating wirelessly or over a wired connection. Examples of non-transitory media may include, but are not limited to, magnetic disks or magnetic tapes, optical storage media such as compact discs (CDs) or digital versatile discs (DVDs), flash memory, memory, or memory devices. Computer-readable media may store code and / or machine-executable instructions thereon, which may represent procedures, functions, subroutines, programs, routines, subroutines, modules, software packages, classes, or any combination of instructions, data structures, or program statements. Code segments may be coupled to other code segments or hardware circuitry by passing and / or receiving information, data, arguments, parameters, or memory contents. Information, arguments, parameters, data, etc., may be passed, forwarded, or transmitted via any suitable means, including memory sharing, message passing, token passing, network transmission, etc.

[0075] In some respects, computer-readable storage devices, media, and memories may include cables or wireless signals containing bit streams, etc. However, when referred to, non-transitory computer-readable storage media explicitly exclude media such as energy, carrier signals, electromagnetic waves, and the signals themselves.

[0076] As described in this article, Figure 9 The neural network 900 can be implemented using one or more neural networks. Figure 9 It is possible to be Figure 9The neural network 900 is an exemplary example of a deep learning neural network 900. An input layer 920 includes input data. In one exemplary example, the input layer 920 may include data representing pixels of an input video frame. The neural network 900 includes multiple hidden layers 922a, 922b through a final hidden layer 922n. Hidden layers 922a, 922b through 922n comprise “n” hidden layers, where “n” is an integer greater than or equal to one. The multiple hidden layers can include as many layers as needed for a given application. The neural network 900 further includes an output layer 924 that provides the output produced by the processing performed by the hidden layers 922a, 922b through the final hidden layer 922n. In one exemplary example, the output layer 924 may provide a classification of objects in the input video frame. The classification may include a category identifying the type of object (e.g., person, dog, cat, or other object).

[0077] Neural network 900 is a multi-layered neural network composed of interconnected nodes. Each node can represent a piece of information. The information associated with these nodes is shared between different layers, and each layer retains information while processing it. In some cases, neural network 900 may include a feedforward network, in which case there are no feedback connections where the network's output is fed back into itself. In some cases, neural network 900 may include a recurrent neural network, which may have loops that allow information to be carried across nodes as input is read.

[0078] Information can be exchanged between nodes via node-to-node interconnects between layers. Nodes in input layer 920 can activate a set of nodes in the first hidden layer 922a. For example, as shown, each input node in input layer 920 is connected to each node in the first hidden layer 922a. Nodes in hidden layers 922a, 922b, up to the last hidden layer 922n, can transform the information of each input node by applying activation functions. The information derived from this transformation can then be passed to nodes in the next hidden layer 922b, activating those nodes, which can then perform their own specified functions. Example functions include convolution, upsampling, data transformation, and / or any other suitable functions. The output of hidden layer 922b can then activate nodes in the next hidden layer, and so on. The output of the last hidden layer 922n can activate one or more nodes in output layer 924, providing the output at those nodes. In some cases, although a node in neural network 900 (e.g., node 926) is shown as having multiple output lines, the node has a single output and is shown as all lines output from the node representing the same output value.

[0079] In some cases, each node or the interconnection between nodes may have weights, which are a set of parameters derived from the training of the neural network 900. Once the neural network 900 has been trained, it can be referred to as a trained neural network and can be used to classify one or more objects. For example, the interconnection between nodes may represent a piece of information learned about the interconnected nodes. This interconnection may have tunable numerical weights that can be tuned (e.g., based on the training dataset), allowing the neural network 900 to adapt to the input and learn as more and more data is processed.

[0080] The neural network 900 is pre-trained to process features from the data in the input layer 920 using different hidden layers 922a, 922b up to the final hidden layer 922n, in order to provide an output through the output layer 924. In an example where the neural network 900 is used to identify objects in an image, the neural network 900 can be trained using training data that includes both images and labels. For example, training images can be input into the network, where each training image has a label indicating the category of one or more objects in each image (basically, indicating to the network what the objects are and what features they have). In an exemplary example, the training images could include images of the number 2, in which case the image label could be [0 0 1 0 0 0 00 0 0].

[0081] In some cases, the neural network 900 can use a training process called backpropagation to adjust the weights of its nodes. Backpropagation includes forward pass, loss function, back pass, and weight update. For each training iteration, forward pass, loss function, back pass, and parameter update are performed. This process can be repeated a certain number of times for each set of training images until the neural network 900 is trained well enough that the weights of each layer are accurately tuned.

[0082] For an example of identifying objects in an image, the forward pass may include passing a training image through a neural network 900. The weights are initially randomized before training the neural network 900. The image may include, for example, a numerical array representing pixels of the image. Each number in the array may include a value from 0 to 255 describing the intensity of the pixel at that location in the array. In one example, the array may include a 28×28×3 numerical array with 28 rows and 28 columns of pixels and 3 color components (such as red, green, and blue, or lightness and two chroma components, etc.).

[0083] For the first training iteration of a neural network 900, the output will likely include values ​​due to the weights being randomly selected during initialization without prioritizing any particular class. For example, if the output is a vector with probabilities that an object includes different classes, the probability values ​​for each class may be equal or at least very similar (e.g., 0.1 for each of ten possible classes). Using the initial weights, the neural network 900 cannot determine low-level features and therefore cannot make an accurate determination of what the object's classification might be. A loss function can be used to analyze the error in the output. Any suitable loss function can be defined. An example of a loss function is Mean Squared Error (MSE). MSE is defined as... It calculates the sum of half the square of the base truth output (e.g., the actual answer) minus the square of the predicted output (e.g., the predicted answer). The loss can be set to equal to The value of .

[0084] For the first training image, the loss (or error) will be high because the actual value will be significantly different from the predicted output. The goal of training is to minimize the loss so that the predicted output matches the training label. The Neural Network 900 performs backpropagation by determining which inputs (weights) contribute most to the network's loss and can adjust the weights to reduce and eventually minimize this loss.

[0085] The derivative of the loss with respect to the weights (denoted as dL / dW, where W is the weight at a specific layer) can be calculated to determine the weights that contribute the most to the network's loss. After calculating the derivative, a weight update can be performed by updating all the weights of the filter. For example, the weights can be updated so that they change in the opposite direction of the gradient. A weight update can be expressed as... Where w represents the weight, w i Let represent the initial weights, and η represent the learning rate. The learning rate can be set to any suitable value, where a high learning rate includes larger weight updates, while a lower value indicates smaller weight updates.

[0086] In some cases, self-supervised learning can be used to train neural networks.

[0087] Neural Network 900 can include any suitable deep network. An example includes a Convolutional Neural Network (CNN), which includes an input layer and an output layer, with multiple hidden layers between them. The hidden layers of a CNN consist of a series of convolutional, non-linear, and / or pooling (for downsampling) layers, and may include one or more fully connected layers. Neural Network 900 can include any other deep network besides CNNs, such as autoencoders, deep belief networks (DBNs), recurrent neural networks (RNNs), and so on.

[0088] Specific details are provided in the foregoing description to provide a thorough understanding of the aspects and examples presented herein. However, those skilled in the art will understand that these aspects can be practiced without these specific details. For clarity, in some instances, the technology may be presented as comprising individual functional blocks, including devices, device components, steps, or routines in a method embodied in software or a combination of hardware and software. Additional components may be used in addition to those shown in the figures and / or described herein. For example, circuits, systems, networks, processes, and other components may be shown as components in block diagram form to avoid obscuring these aspects with unnecessary details. In other instances, well-known circuits, processes, algorithms, structures, and techniques may be shown without unnecessary details to avoid obscuring the aspects.

[0089] Various aspects described above can be presented as processes or methods, depicted as flowcharts, diagrams, data flow graphs, structure diagrams, or block diagrams. While flowcharts may describe operations as sequential processes, many operations within an operation can be executed in parallel or concurrently. Furthermore, the order of operations can be rearranged. A process terminates when its operations are completed, but it may have additional steps not included in the accompanying diagrams. A process can correspond to a method, function, procedure, subroutine, subroutine, etc. When a process corresponds to a function, its termination may correspond to the function returning to its calling function or the main function.

[0090] The processes and methods described in the examples above can be implemented using stored computer-executable instructions or computer-executable instructions otherwise obtainable from a computer-readable medium. Such instructions may include, for example, instructions and data that configure, or otherwise configure, a general-purpose computer, special-purpose computer, or processing device to perform a function or group of functions. The portion may be accessible via a network of the computer resources used. The computer-executable instructions may be, for example, binary, intermediate format instructions such as assembly language, firmware, or source code. Examples of computer-readable media that can be used to store the instructions, the information used, and / or information created during the methods according to the described examples include disks or optical discs, flash memory, USB devices with non-volatile memory, networked storage devices, etc.

[0091] Devices implementing the processes and methods according to these disclosures may include hardware, software, firmware, middleware, microcode, hardware description languages, or any combination thereof, and may take any of a variety of form factors. When implemented as software, firmware, middleware, or microcode, program code or code segments (e.g., computer program products) for performing necessary tasks may be stored in a computer-readable or machine-readable medium. A processor performs the necessary tasks. Typical examples of form factors include laptops, smartphones, mobile phones, tablet devices, or other small form factor personal computers, personal digital assistants, rack-mounted devices, standalone devices, etc. The functionality described herein may also be embodied in peripheral devices or interlocking cards. By further example, such functionality may also be implemented on circuit boards of different chips or different processes executed on a single device.

[0092] Instructions, media for delivering such instructions, computing resources for executing them, and other structures for supporting such computing resources are example components for providing the functionality described in this disclosure.

[0093] In the foregoing description, aspects of this application have been described with reference to their specific aspects, but those skilled in the art will recognize that this application is not limited thereto. Therefore, although illustrative aspects of this application have been described in detail herein, it is to be understood that the inventive concepts can be implemented and employed in various other ways, and the appended claims are not intended to be construed as including these variations unless limited by prior art. The various features and aspects of the applications described above can be used individually or in combination. Furthermore, aspects can be utilized in any number of environments and applications beyond those described herein without departing from the broader spirit and scope of this specification. Therefore, the specification and drawings should be considered illustrative rather than restrictive. For illustrative purposes, the methods are described in a particular order. It should be understood that, in alternative aspects, the methods may be performed in a different order than described.

[0094] Those skilled in the art will understand that, without departing from the scope of this description, the less than ("<") and greater than (">") symbols or terms used herein may be replaced with less than or equal to ("<"), respectively. ") and greater than or equal to (" The symbol ) is used instead.

[0095] When a component is described as being “configured” to perform certain operations, such configuration can be achieved, for example, by designing electronic circuits or other hardware to perform the operations, by programming programmable electronic circuits (e.g., microprocessors or other suitable electronic circuits) to perform the operations, or any combination thereof.

[0096] The phrase “coupled to” means any component that is physically connected directly or indirectly to another component, and / or any component that communicates directly or indirectly with another component (e.g., connected to another component via a wired or wireless connection and / or other suitable communication interface).

[0097] Claim language or other languages ​​that state "at least one of" and / or "one or more of" in a set indicate that one member of the set or multiple members of the set (in any combination) satisfy the claim. For example, claim language stating "at least one of A and B" or "at least one of A or B" means A, B, or A and B. In another example, claim language stating "at least one of A, B, and C" or "at least one of A, B, or C" means A, B, C, or A and B, or A and C, or B and C, A and B and C, or any repetition is information or data (e.g., A and A, B and B, C and C, A and A and B, etc.), or any other ordering, repetition, or combination of A, B, and C. The language "at least one of" and / or "one or more of" in a set does not limit the set to the items listed in the set. For example, the language of a claim stating "at least one of A and B" or "at least one of A or B" may refer to A, B, or A and B, and may additionally include items not listed in the set of A and B. The phrases "at least one" and "one or more" are used interchangeably herein.

[0098] Claims using phrases such as "at least one processor, the at least one processor being configured to," "at least one processor being configured to," "one or more processors, the one or more processors being configured to," or "one or more processors being configured to," or other languages, indicate that one or more processors (in any combination) are capable of performing associated operations. For example, a claim using the phrase "at least one processor, the at least one processor being configured to: X, Y, and Z" means that a single processor can be used to perform operations X, Y, and Z; or that multiple processors are each assigned a specific subset of tasks to perform operations X, Y, and Z, such that the multiple processors together perform X, Y, and Z; or that a group of multiple processors work together to perform operations X, Y, and Z. In another example, a claim using the phrase "at least one processor, the at least one processor being configured to: X, Y, and Z" could mean that any single processor can perform only at least one subset of operations X, Y, and Z.

[0099] When referring to one or more elements that perform functions (e.g., steps of a method), one element may perform all functions, or more than one element may jointly perform these functions. When more than one element jointly performs these functions, each function does not need to be performed by every single element (e.g., different functions may be performed by different elements), and / or each function does not need to be performed by only one element as a whole (e.g., different elements may perform different sub-functions of a function). Similarly, when referring to one or more elements configured to cause another element (e.g., a device) to perform functions, one element may be configured to cause another element to perform all functions, or more than one element may be jointly configured to cause another element to perform these functions.

[0100] When referring to an entity that performs or is configured to perform functions (e.g., steps of a method) (e.g., any entity or device described herein), the entity may be configured to cause one or more elements (individually or collectively) to perform those functions. One or more components of the entity may include at least one memory, at least one processor, at least one communication interface, another component configured to perform one or more of those functions, and / or any combination thereof. When referring to an entity that performs functions, the entity may be configured to cause one component to perform all functions, or to cause more than one component to perform those functions collectively. When the entity is configured to cause more than one component to perform those functions collectively, each function does not need to be performed by every single component (e.g., different functions may be performed by different components), and / or each function does not need to be performed by only one component as a whole (e.g., different components may perform different sub-functions of a function).

[0101] The various exemplary logic blocks, modules, circuits, and algorithm steps described in conjunction with the examples disclosed herein can be implemented as electronic hardware, computer software, firmware, or combinations thereof. To clearly illustrate this interchangeability between hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described in general terms of their functionality. Whether such functionality is implemented as hardware or software depends on the specific application and the design constraints imposed on the overall system. Those skilled in the art may implement the described functionality in different ways for each specific application, but such implementation decisions should not be construed as departing from the scope of this application.

[0102] The techniques described herein can also be implemented in electronic hardware, computer software, firmware, or any combination thereof. Such techniques can be implemented in any of a variety of devices, such as general-purpose computers, wireless communication devices (mobile phones), or integrated circuit devices with multiple uses, including applications in wireless communication devices (mobile phones) and other devices. Any feature described as a module or component can be implemented together in an integrated logic device or separately as discrete but interoperable logic devices. If implemented in software, the techniques can be implemented at least in part by a computer-readable data storage medium comprising program code including instructions that, when executed, perform one or more of the methods, algorithms, and / or operations described above. The computer-readable data storage medium can form part of a computer program product, which may include packaging material. The computer-readable medium may include memory or data storage media, such as random access memory (RAM) (such as synchronous dynamic random access memory (SDRAM)), read-only memory (ROM), non-volatile random access memory (NVRAM), electrically erasable programmable read-only memory (EEPROM), flash memory, magnetic or optical data storage media, etc. Additionally or alternatively, the technology may be implemented at least in part by a computer-readable communication medium that carries or conveys program code in the form of instructions or data structures that can be accessed, read and / or executed by a computer, such as propagated signals or waves.

[0103] The program code can be executed by a processor, which may include one or more processors, such as one or more digital signal processors (DSPs), general-purpose microprocessors, application-specific integrated circuits (ASICs), field-programmable arrays (FPGAs), or other equivalent integrated or discrete logic circuits. Such processors can be configured to perform any of the techniques described in this disclosure. A general-purpose processor may be a microprocessor; however, in alternatives, the processor may be any conventional processor, controller, microcontroller, or state machine. The processor may also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. Therefore, as used herein, the term "processor" may refer to any of the foregoing structures, any combination of the foregoing structures, or any other structure or means suitable for implementing the techniques described herein.

[0104] The exemplary aspects of this disclosure include:

[0105] Aspect 1. An apparatus for generating output speech from input data, the apparatus comprising: one or more memories configured to store the input data; and one or more processors coupled to the one or more memories and configured to: extract first prosodic data from the input data; generate content embeddings based on the input data; extract second prosodic data from target speech; generate speaker embeddings from the target speech; generate prosodic embeddings from the second prosodic data; and generate transformed prosodic data based on the first prosodic data and the prosodic embeddings.

[0106] Aspect 2. The apparatus according to aspect 1, wherein the input data includes one or more of voice data or text data.

[0107] Aspect 3. The apparatus according to aspect 2, wherein the input data includes one of voice data and text data.

[0108] Aspect 4. The apparatus according to any one of Aspects 1 to 3, wherein the first rhythm data includes one or more of a fundamental frequency, an energy value, and a velocity value.

[0109] Aspect 5. The apparatus according to any one of Aspects 1 to 4, wherein the second rhythm data includes one or more of a fundamental frequency, an energy value, and a velocity value.

[0110] Aspect 6. The apparatus according to any one of Aspects 1 to 5, wherein the one or more processors are configured to: generate a converted spectrogram based on the converted prosodic data, the speaker embedding, and the content embedding; and generate the converted spectrogram via a decoder based on the converted prosodic data, the speaker embedding, and the content embedding, the decoder comprising a diffuse decoder or a non-diffuse decoder.

[0111] Aspect 7. The apparatus according to aspect 6, wherein the one or more processors are configured to: generate a predicted global speech rate based on the converted prosodic data, and generate the speech rate of the converted spectrogram via a rate control engine; and generate the converted speech based on the input data via a vocoder.

[0112] Aspect 8. The apparatus according to aspect 7, wherein the vocoder includes a neural vocoder.

[0113] Aspect 9. The apparatus according to any one of Aspects 6 to 8, wherein the one or more processors are configured to: extract first prosodic data from the input data via a first prosodic extractor engine; generate the content embedding based on the input data via a content encoder; extract the second prosodic data from target speech via a second prosodic extractor engine; generate the speaker embedding from the target speech via a speaker encoder; generate the prosodic embedding from the second prosodic data via a prosodic encoder; generate converted prosodic data based on the first prosodic data and the prosodic embedding via a prosodic conversion engine; and generate the converted spectrogram based on the converted prosodic data, the speaker embedding, and the content embedding via a decoder.

[0114] Aspect 10. The apparatus according to aspect 9, wherein the apparatus includes the decoder, and wherein the decoder is configured to synthesize a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosodic data.

[0115] Aspect 11. The apparatus according to any one of Aspects 7 to 10, wherein the rate control engine is configured to manipulate speech rate depending on the predicted speed.

[0116] Aspect 12. The apparatus according to any one of aspects 9 to 11, wherein the prosodic encoder is configured to generate the prosodic embedding at one or more of the frame level or the sentence level.

[0117] Aspect 13. The apparatus according to aspect 12, wherein the apparatus includes the prosody encoder, and wherein the prosody encoder is configured to generate the prosody embedding at the frame level to enable frame-level intonation control.

[0118] Aspect 14. The apparatus according to any one of Aspects 7 to 13, wherein the one or more processors are configured to: generate the speech rate of the converted spectrogram independently of an automatic speech recognition model based on the converted prosodic data via the rate control engine.

[0119] Aspect 15. A method for generating output speech from input data, the method comprising: extracting first prosodic data from the input data; generating a content embedding based on the input data; extracting second prosodic data from the target speech; generating a speaker embedding from the target speech; generating a prosodic embedding from the second prosodic data; and generating transformed prosodic data based on the first prosodic data and the prosodic embedding.

[0120] Aspect 16. The method according to aspect 15, wherein the input data includes one or more of voice data or text data.

[0121] Aspect 17. The method according to aspect 16, wherein the input data includes one of voice data and text data.

[0122] Aspect 18. The method according to any one of Aspects 15 to 17, wherein the first prosodic data includes one or more of a fundamental frequency, an energy value, and a velocity value.

[0123] Aspect 19. The method according to any one of Aspects 15 to 18, wherein the second prosodic data includes one or more of a fundamental frequency, an energy value, and a velocity value.

[0124] Aspect 20. The method according to any one of Aspects 15 to 19, the method further comprising: generating a converted spectrogram based on the converted prosodic data, the speaker embedding, and the content embedding; and generating the converted spectrogram via a decoder based on the converted prosodic data, the speaker embedding, and the content embedding, the decoder comprising a diffuse decoder or a non-diffuse decoder.

[0125] Aspect 21. The method according to aspect 20, the method further comprising: generating a predicted global speech rate based on the converted prosodic data, and generating the speech rate of the converted spectrogram via a rate control engine; and generating the converted speech via a vocoder based on the input data.

[0126] Aspect 22. The method according to aspect 21, wherein the vocoder includes a neural vocoder.

[0127] Aspect 23. The method according to any one of Aspects 20 to 22, the method further comprising: extracting first prosodic data from the input data via a first prosodic extractor engine; generating the content embedding based on the input data via a content encoder; extracting the second prosodic data from target speech via a second prosodic extractor engine; generating the speaker embedding from the target speech via a speaker encoder; generating the prosodic embedding from the second prosodic data via a prosodic encoder; generating converted prosodic data based on the first prosodic data and the prosodic embedding via a prosodic conversion engine; and generating the converted spectrogram based on the converted prosodic data, the speaker embedding, and the content embedding via a decoder.

[0128] Aspect 24. The method according to aspect 23, wherein the method is performed by a decoder, and wherein the decoder is configured to synthesize a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosodic data.

[0129] Aspect 25. The method according to any one of aspects 21 to 24, wherein the rate control engine is configured to manipulate speech rate depending on the predicted rate.

[0130] Aspect 26. The method according to any one of aspects 23 to 25, wherein the prosodic encoder is configured to generate the prosodic embedding at one or more of the frame level or sentence level.

[0131] Aspect 27. The method according to aspect 26, wherein the method is performed by a prosodic encoder, and wherein the prosodic encoder is configured to generate the prosodic embedding at the frame level to enable frame-level intonation control.

[0132] Aspect 28. The method according to any one of Aspects 21 to 27, the method further comprising: generating the speech rate of the converted spectrogram independently of an automatic speech recognition model based on the converted prosodic data via the rate control engine.

[0133] Aspect 29. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, configuring the one or more processors to: extract first prosodic data from input data; generate a content embedding based on the input data; extract second prosodic data from target speech; generate a speaker embedding from the target speech; generate a prosodic embedding from the second prosodic data; and generate transformed prosodic data based on the first prosodic data and the prosodic embedding.

[0134] Aspect 30. An apparatus comprising: means for extracting first prosodic data from input data; means for generating content embeddings based on the input data; means for extracting second prosodic data from target speech; means for generating speaker embeddings from the target speech; means for generating prosodic embeddings from the second prosodic data; and means for generating transformed prosodic data based on the first prosodic data and the prosodic embeddings.

[0135] Aspect 31. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, configuring the one or more processors to perform any one of aspects 15 to 28.

[0136] Aspect 32. A non-transitory computer-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, configuring the one or more processors to perform any one of aspects 15 to 28.

Claims

1. An apparatus to generate output speech from input data, the apparatus comprising: one or more memories configured to store the input data; and one or more processors coupled to the one or more memories and configured to: extract first prosody data from the input data; generate a content embedding based on the input data; extract second prosody data from a target speech; generate a speaker embedding from the target speech; generate a prosody embedding from the second prosody data; and generate converted prosody data based on the first prosody data and the prosody embedding.

2. The apparatus of claim 1, wherein the input data comprises one or more of speech data or text data.

3. The apparatus of claim 2, wherein the input data comprises one of speech data and text data.

4. The apparatus of claim 1, wherein the first prosody data comprises one or more of a fundamental frequency, an energy value, and a velocity value.

5. The apparatus of claim 1, wherein the second prosody data comprises one or more of a fundamental frequency, an energy value, and a velocity value.

6. The apparatus of claim 1, wherein the one or more processors are configured to: generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding; and generate, via a decoder, the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding, the decoder comprising a diffused decoder or a non-diffused decoder.

7. The apparatus of claim 6, wherein the one or more processors are configured to: generate a predicted global speaking rate based on the converted prosody data, and generate a speaking rate of the converted spectrogram via a rate control engine; and generate, via a vocoder, converted speech based on the input data.

8. The apparatus of claim 7, wherein the vocoder comprises a neural vocoder.

9. The apparatus of claim 6, wherein the one or more processors are configured to: extract the first prosody data from the input data via a first prosody extractor engine; generate the content embedding based on the input data via a content encoder; extract the second prosody data from a target speech via a second prosody extractor engine; generate the speaker embedding from the target speech via a speaker encoder; generate the prosody embedding from the second prosody data via a prosody encoder; generate converted prosody data based on the first prosody data and the prosody embedding via a prosody conversion engine; generate, via a decoder, the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding.

10. The apparatus of claim 9, wherein the apparatus comprises the decoder, and wherein the decoder is configured to synthesize a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosody data. ​ ​ ​ 11. The apparatus of claim 7, wherein the rate control engine is configured to manipulate a speech rate depending on a predicted velocity.

12. The apparatus of claim 9, the prosody encoder is configured to generate the prosody embedding at one or more of a frame level or a sentence level.

13. The apparatus of claim 12, wherein the apparatus comprises the prosody encoder, and wherein the prosody encoder is configured to generate the prosody embedding at the frame level to enable frame level prosody control.

14. The apparatus of claim 7, wherein the one or more processors are configured to: generate, via the rate control engine, the speech rate of the converted spectrogram independent of an automatic speech recognition model based on the converted prosody data.

15. The apparatus of claim 1, wherein the input data comprises speech data, the apparatus further comprising one or more microphones configured to capture the speech data.

16. The apparatus of claim 1, the apparatus further comprising one or more speakers configured to output speech data comprising the converted prosody data.

17. A method of generating output speech from input, the method comprising: extracting first prosody data from input data; generating a content embedding based on the input data; extracting second prosody data from a target speech; generating a speaker embedding from the target speech; generating a prosody embedding from the second prosody data; and generating converted prosody data based on the first prosody data and the prosody embedding.

18. The method of claim 17, wherein the input data comprises one or more of speech data or text data.

19. The method of claim 18, wherein the input data comprises one of speech data and text data.

20. The method of claim 17, wherein the first prosody data comprises one or more of a fundamental frequency, an energy value, and a velocity value.

21. The method of claim 17, wherein the second prosody data comprises one or more of a fundamental frequency, an energy value, and a velocity value.

22. The method of claim 17, the method further comprising: generating a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding; and generating, via a decoder, the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding, the decoder comprising a diffused decoder or a non-diffused decoder.

23. The method of claim 22, the method further comprising: generating a predicted global speech rate based on the converted prosody data, and generating a speech rate of the converted spectrogram via a rate control engine; and generating converted speech based on the input data via a vocoder.

24. The method of claim 23, wherein the vocoder comprises a neural vocoder.

25. The method of claim 22, the method further comprising: extracting, via a first prosody extractor engine, the first prosody data from the input data; generating, via a content encoder, the content embedding based on the input data; extracting, via a second prosody extractor engine, the second prosody data from a target speech; generating, via a speaker encoder, the speaker embedding from the target speech; generating, via a prosody encoder, the prosody embedding from the second prosody data; generating, via a prosody conversion engine, converted prosody data based on the first prosody data and the prosody embedding; and generating, via a decoder, the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding.

26. The method of claim 25, wherein the method is performed by a decoder, and wherein the decoder is configured to synthesize speech spectrograms conditioned on the content embedding, the speaker embedding, and the converted prosody data.

27. The method of claim 23, wherein the rate control engine is configured to manipulate speech rate in dependence on a predicted velocity.

28. The method of claim 25, the prosody encoder is configured to generate the prosody embedding at one or more of a frame level or a sentence level.

29. The method of claim 28, wherein the method is performed by a prosody encoder, and wherein the prosody encoder is configured to generate the prosody embedding at the frame level to enable frame-level prosody control.

30. The method of claim 23, the method further comprising: generating, via the rate control engine, the speech rate of the converted spectrogram independent of an automatic speech recognition model based on the converted prosody data.