Voice conversion method using diffusion model and system for executing same

WO2026106243A1PCT designated stage Publication Date: 2026-05-21INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
INDUSTRY UNIVERSITY COOPERATION FOUNDATION HANYANG UNIVERSITY
Filing Date
2025-11-07
Publication Date
2026-05-21

AI Technical Summary

Technical Problem

Existing speech conversion systems using diffusion models struggle to achieve naturalness and timbre fidelity in voice conversion due to the intertwining of speaker's voice and speech content, leading to unnatural conversion results.

Method used

A speech conversion method using a diffusion model that employs multiple automatic speech recognition models and information distortion techniques to precisely extract timbre and separate content data, pitch data, and style data, generating a Mel spectrogram that reflects the target speaker's voice.

Benefits of technology

The method enhances the naturalness and quality of converted speech by accurately preserving the target speaker's timbre and linguistic content, while improving computational efficiency by using a single diffusion model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025018339_21052026_PF_FP_ABST
    Figure KR2025018339_21052026_PF_FP_ABST
Patent Text Reader

Abstract

Disclosed are a voice conversion method using a diffusion model and a system for executing same. A voice conversion method using a diffusion model according to one embodiment disclosed herein is performed by a system including at least one processor and comprises: a data extraction process of extracting style data from a target speaker voice on the basis of original voice data; a source filter encoding process of synthesizing content data, the style data, and pitch data on the basis of source-filter theory to generate a mel spectrogram reflecting the timbre of the speaker voice; and a data conversion process of conditionally providing the style data and the pitch data using a pre-trained diffusion model, and restoring the speaker voice from the mel spectrogram on the basis of the condition.
Need to check novelty before this filing date? Find Prior Art

Description

Speech conversion method using a diffusion model and a system for executing the same

[0001] The disclosed invention relates to a speech conversion method using a diffusion model that can improve the naturalness and timbre fidelity of the converted speech by utilizing a diffusion model and information distortion, and a system for executing the same.

[0002] With the recent advancement of artificial intelligence technology, many companies are introducing services such as voice conversion and voice synthesis. Voice conversion refers to transforming an input voice into a target voice while preserving the content of the speech. When a specific speech signal is provided as input to a model, voice conversion is effectively performed by obtaining a spectrum in the frequency / time domain through the Short-Time Fourier Transform (STFT) and then applying it to the conversion model. The reason for using a spectrum is that the spectrum provides information for conversion in a Mel filter bank (mel). Voice conversion refers to converting the speech of a source speaker into the speech of a target speaker. The information contained in speech can include linguistic content and the speaker's vocal characteristics (rhythm, pitch range, timbre, etc.); voice conversion is a technology that modifies or replaces the speaker's vocal characteristics without altering the aforementioned linguistic content.

[0003] Since voice data consists of the speaker's speech content and voice characteristics (voice traits) organically intertwined, it is difficult to obtain a natural voice conversion result by separating the speaker's voice and speech content during the voice conversion process.

[0004] With the recent advancement of artificial intelligence technology, AI technology is being actively applied to speech conversion. Generally, two digital voice files are converted into small-dimensional vectors through a neural network (e.g., an encoder). The encoded information is a latent vector, and after modifying the latent vector, it is decoded through a neural network (e.g., a decoder) to generate voice information converted from the latent vector.

[0005] However, in the process of extracting features from voice data to input voice files into a neural network, the speaker's utterance content and voice features are intertwined, and unless the utterance content and voice features are perfectly separated, the voice conversion result becomes unnatural.

[0006] Although speech conversion systems applying diffusion models have recently been developed, existing systems perform speech conversion using only a single diffusion model. General diffusion model-based speech conversion systems are primarily used as a post-processing step to achieve higher-quality speech conversion, and training is conducted conditionally, mainly using style information, to ensure that style information is more heavily reflected during speech generation. In this process, it is important to generate speech that does not compromise the naturalness and quality of the voice while separating and then recombining pronunciation components, such as pitch and timbre, during conversion to the speaker's voice.

[0007] In conventional diffusion model-based speech conversion systems, while the diffusion model can help improve speech quality, it has the problem of having limitations in reflecting the speaker's timbre to be changed.

[0008] The disclosed embodiment for solving these problems relates to a speech conversion method using a diffusion model and a system for executing the same, which can improve the naturalness, quality, and speaker similarity of the converted speech by using a diffusion model and multiple automatic speech recognition models during speech conversion to improve speech quality and using information distortion techniques to precisely extract timbre to further enhance speaker similarity.

[0009] A speech conversion method using a diffusion model according to a disclosed embodiment is a speech conversion method using a diffusion model performed by a system including at least one processor, comprising: a data extraction process for separating content data and pitch data based on original speech data and extracting style data from a target speaker's voice; a source filter encoding process for synthesizing the content data, style data, and pitch data based on source filter theory to generate a Mel spectrogram reflecting the timbre of the speaker's voice; and a data conversion process for conditionally providing the style data and pitch data using a pre-trained diffusion model and restoring the speaker's voice from the Mel spectrogram based on the condition.

[0010] Alternatively, the data extraction process described above extracts content data using the architecture of a pre-trained Automatic Speech Recognition (ASR) model.

[0011] Alternatively, the data extraction process described above extracts the content data and style data using an information distortion function that utilizes content-preserved perturbation (f(·)) information and pitch perturbation (g(·)) information.

[0012] Alternatively, the information distortion function equation for the above-mentioned content preservation variation is represented by the following mathematical formula 1, and the information distortion function equation for the above-mentioned pitch variation information is represented by the following mathematical formula 2.

[0013] Alternatively, the source filter encoding process comprises a source Mel spectrogram (Z) based on the content data (c) and style data (s). src ) generating a step of generating a filter Mel spectrogram (Z) based on the pitch data (F0) and style data(s). ftr A step of generating ); and the source Mel spectrogram (Z src ) and filter Mel spectrogram (Z ftr Target Mel spectrogram (X) through element-wise addition operations between ) mel Mel spectrogram approx. )(Z mel It includes the step of generating ).

[0014] Alternatively, the above method further includes a training process that trains the diffusion model using objective functions of reconstruction loss and diffusion loss.

[0015] Alternatively, the above learning process uses the following mathematical formula to the reconstruction loss (L recon Calculate ) and the above reconstruction loss (L recon Target Mel Spectrogram (X) through ) mel ) and Mel spectrogram (Z mel The learning process is to make ) identical to each other.

[0016] Alternatively, the above learning process uses the following mathematical formula to the above diffusion loss (L diff Calculate ) and the above diffusion loss (L diff Using ), the above diffusion model scores (s θ It is to make it predict ).

[0017] A system according to a disclosed embodiment is a system for performing speech conversion using a diffusion model, comprising: a data extraction module that separates content data and pitch data based on original speech data and extracts style data from a target speaker's voice; a source filter encoding module that synthesizes the content data, style data, and pitch data based on source filter theory to generate a Mel spectrogram reflecting the timbre of the speaker's voice; and a data conversion module that conditionally provides the style data and pitch data using a pre-trained diffusion model and restores the speaker's voice from the Mel spectrogram based on the condition.

[0018] Alternatively, the data extraction module extracts content data from the original voice data using the architecture of an automatic speech recognition (ASR) model.

[0019] Alternatively, the data extraction module further includes a style encoder that extracts style data(s) from a reference voice that convey non-verbal attributes not included in the text of the original voice data.

[0020] Alternatively, the data extraction module further includes a pitch extractor that extracts pitch data (F0) of the voice from the original voice data.

[0021] Alternatively, the data extraction module extracts the content data and style data using an information distortion function that utilizes content-preserved perturbation (f(·)) information and pitch perturbation (g(·)) information.

[0022] Alternatively, the information distortion function equation for the above-mentioned content preservation variation is represented by the following mathematical formula 1, and the information distortion function equation for the above-mentioned pitch variation information is represented by the following mathematical formula 2.

[0023] Alternatively, the source filter encoding module generates a source Mel spectrogram (Z based on the pitch data (F0) and style data(s) src A source encoder that generates ); a filtered Mel spectrogram (Z) based on the content data (c) and style data (s). ftr A filter encoder that generates ); and the source Mel spectrogram (Z src ) and filter Mel spectrogram (Z ftr Target Mel spectrogram (X) through element-wise addition operations between ) mel Mel spectrogram approx. )(Z mel It includes an adder that outputs )

[0024] The voice conversion method using a diffusion model and the system for executing the same disclosed in the embodiment can convert the original speaker's voice into the voice of a target speaker while preserving linguistic content by using original voice data that does not rely on text, and at this time, by using a diffusion model and information distortion technology, it is possible to generate natural, high-quality voice while preserving the timbre of the target speaker's voice well.

[0025] In addition, in the disclosed embodiment, the speech conversion method using a diffusion model and the system for executing the same can secure computational efficiency by using only one diffusion model, and can better extract timbre features from the original speech data to be converted using information distortion techniques, as well as improve the quality of the converted speech by using multiple automatic speech recognition models.

[0026] Figure 1 is a control block diagram of a system that executes a speech conversion method using a diffusion model.

[0027] Figure 2 is a diagram illustrating a speech conversion system using a disclosed diffusion model.

[0028] Figure 3 is a diagram illustrating the configuration of the disclosed source filter encoding module and data conversion module.

[0029] Figure 4 is a flowchart illustrating a speech conversion method using a disclosed diffusion model.

[0030] Figure 5 is a flowchart illustrating a method for training a disclosed diffusion model.

[0031] Figure 6 is a diagram showing the results of comparing the similarity between a speaker's voice using a disclosed diffusion model and a conventional voice conversion method.

[0032] Figure 7 is a diagram showing the results of objective and subjective evaluations of a speech conversion method using a disclosed diffusion model.

[0033] Throughout the specification, the same reference numerals refer to the same components. This specification does not describe all elements of the embodiments, and general content in the art to which the invention pertains or content that overlaps between embodiments is omitted.

[0034] Throughout the specification, when a part is described as being "connected" to another part, this includes not only cases where they are directly connected but also cases where they are indirectly connected, and indirect connections include connections made via a wireless communication network.

[0035] Furthermore, when it is stated that a part "includes" a certain component, this means that, unless specifically stated otherwise, it does not exclude other components but may include additional components.

[0036] Singular expressions include plural expressions unless there is an obvious exception in the context.

[0037] In addition, terms such as "~part," "~unit," "~block," "~part," and "~module" may refer to a unit that processes at least one function or operation. For example, the above terms may refer to at least one piece of hardware such as an FPGA (field-programmable gate array) or an ASIC (application specific integrated circuit), at least one piece of software stored in memory, or at least one process processed by a processor.

[0038] The symbols attached to each step are used to identify each step and do not indicate the order of the steps relative to one another; the steps may be performed differently from the specified order unless a specific order is clearly indicated in the context.

[0039] Hereinafter, with reference to the attached drawings, embodiments relating to a speech conversion method using a diffusion model according to the disclosed embodiment and a system for executing the same will be described in detail.

[0040] Figure 1 is a control block diagram of a system for executing a method for training a speech recognition model based on the disclosed connectionist time classification.

[0041] A system (1) that executes a voice conversion method using a diffusion model can be implemented as a computer or portable terminal (hereinafter user terminal, 3) that can connect to a communication network such as the internet. Here, the computer includes, for example, a laptop, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser, and the portable terminal can be implemented as, for example, any type of handheld-based wireless communication device such as a smartphone, etc., as a wireless communication device that ensures portability and mobility.

[0042] Referring to FIG. 1, a user terminal (3) may include a communication interface (11) for receiving various data, such as an artificial intelligence model or training data required for voice conversion, from an external server (2); a memory (12) for storing the received training data and artificial intelligence model; an input unit (13) for receiving text to be converted into voice or other various user input commands; a processor (10) for controlling the overall system; and an output unit (14) for outputting the synthesized voice performed by the processor (10) as sound. However, since FIG. 1 is merely an example, the user terminal (3) may include other configurations for implementing a computing environment. Additionally, only some of the disclosed configurations may be included in the user terminal (3).

[0043] Specifically, the communication interface (11) may include one or more components that enable communication with an external communication network, and may include, for example, at least one of a short-range communication module, a wired communication module, and a wireless communication module.

[0044] The short-range communication module may include various short-range communication modules that transmit and receive signals using a wireless communication network at short range, such as a Bluetooth module, an infrared communication module, an RFID (Radio Frequency Identification) communication module, a WLAN (Wireless Local Access Network) communication module, an NFC communication module, and a Zigbee communication module.

[0045] Wired communication modules may include various wired communication modules such as Local Area Network (LAN) modules, Wide Area Network (WAN) modules, or Value Added Network (VAN) modules, as well as various cable communication modules such as USB (Universal Serial Bus), HDMI (High Definition Multimedia Interface), DVI (Digital Visual Interface), RS-232 (recommended standard 232), power line communication, or POTS (plain old telephone service).

[0046] In addition to Wi-Fi modules and WiBro (Wireless broadband) modules, the wireless communication module may include wireless communication modules that support various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), and LTE (Long Term Evolution).

[0047] The memory (12) may store an artificial intelligence model required to implement a voice conversion method using the disclosed diffusion model, an algorithm required for the operation of the processor (10), or a program for implementing the algorithm. To this end, the memory (12) may be implemented as at least one of a non-volatile memory device such as a cache, ROM (Read Only Memory), PROM (Programmable ROM), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), and flash memory, a volatile memory device such as RAM (Random Access Memory), or a storage medium such as a hard disk drive (HDD) and a CD-ROM, but is not limited thereto.

[0048] The input unit (13) can receive various input commands. Specifically, the input unit (13) may include a microphone that receives voice that speaks into a video containing text as subtitles, in addition to receiving input text. Furthermore, the input unit (13) may include hardware devices such as various buttons, switches, pedals, keyboards, mice, trackballs, various levers, handles, or sticks for receiving execution commands required for the operation of the processor (10). Additionally, the input unit (13) may include a GUI (Graphical User Interface), i.e., a software device such as a touch pad, for user input commands. The touch pad may be implemented as a touch screen panel (TSP) and form a layered structure with the display of the output unit (14).

[0049] The output unit (14) may include a speaker that outputs the converted voice as sound, as well as a display for outputting the learning results or inference results of an artificial intelligence model. The display may be provided as a Cathode Ray Tube (CRT), Digital Light Processing (DLP) panel, Plasma Display Panel, Liquid Crystal Display (LCD) panel, Electro Luminescence (EL) panel, Electrophoretic Display (EPD) panel, Electrochromic Display (ECD) panel, Light Emitting Diode (LED) panel, or Organic Light Emitting Diode (OLED) panel, but is not limited thereto.

[0050] The processor (10) can fine-tune or train an artificial intelligence model stored in memory (12) and perform speech recognition functions through the trained artificial intelligence model for input text.

[0051] Diffusion models fundamentally model the speech conversion process probabilistically and operate based on the basic principles of forward and reverse processes. The forward process (or propagation process) starts with the original speech data and gradually adds noise; this noise addition process involves multiple stages, at which point the speech moves slightly further away from the original, eventually reaching a state of nearly pure noise. The reverse process (or inversion process) is the opposite of the forward process, starting from a state of pure noise and returning to the original speech data. The reverse process gradually removes noise based on specific conditions, such as style and pitch data, and this is the part that the model must learn.

[0052] Specifically, the processor (10) enables the learning process of the diffusion model using an objective function. To this end, the processor (10) can perform learning so that the target Mel spectrogram for the speaker pitch to be changed based on reconstruction loss and the Mel spectrogram output through the source filter encoding process become identical to each other, and the diffusion model can use diffusion loss to effectively approximate the conditional score function so that the diffusion model scores (s θ Training is performed to predict ).

[0053] A specific method for the processor (10) to perform speech conversion by a learned diffusion model or to learn a diffusion model or other artificial intelligence model will be described later through other drawings below.

[0054] Meanwhile, the processor (10) may refer to a data processing device embedded in hardware having a physically structured circuit to perform a function expressed by code or instructions included in a program. Examples of such data processing devices embedded in hardware may include, but are not limited to, processing devices such as a microprocessor, a central processing unit (CPU), a processor core, a multiprocessor, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a Graphics Processing Unit (GPU), and a neural Processing Unit (NPU). The processor (10) may be provided in multiple units.

[0055] The hardware configurations provided in the user terminal (10) can exchange data and signals via Network Termination (NT) of a digital network such as ISDN (Integrated Services Digital Network).

[0056] FIG. 2 is a diagram illustrating a speech conversion system using a disclosed diffusion model, and FIG. 3 is a diagram illustrating the configuration of a disclosed source filter encoding module and a data conversion module.

[0057] Referring to FIG. 2, the system (100) includes, but is not limited to, a data extraction module (110), a source filter encoding module (120), and a data conversion module (130).

[0058] Since the original voice data is largely composed of content, speaker, and prosody, the data extraction module (110) extracts content data corresponding to the content, style data corresponding to the speaker's voice, and pitch data corresponding to the prosody, respectively.

[0059] The data extraction module (110) separates and extracts rich data including content data, style vectors, timbre, and pitch based on the original voice data. This data extraction module (110) includes XLS-R (111), Whisper (112), style encoder (113), and pitch extractor (114). Here, XLS-R (111) and Whisper (112) are the architectures of an automatic speech recognition model. Architecture represents the design or structure of a deep learning network, and the model can be described as the result learned according to that design.

[0060] Automatic speech recognition (ASR) models can recognize speech and convert it into corresponding text sequences (words or subwords). ASR models are either end-to-end models or hybrid HMM-DNN models.

[0061] XLS-R (111) is a large-scale model for learning cross-language speech representations based on wav2vec 2.0. XLS-R has been trained on nearly 500,000 hours of speech audio released in 128 languages ​​and achieves state-of-the-art results across a wide range of tasks, domains, data systems, and languages. XLS-R (111) utilizes large-scale multilingual data augmentation and contrastive learning techniques to learn universal speech representations that can be transmitted between languages ​​and domains.

[0062] Whisper (112) is a general-purpose model designed for speech recognition in noisy or resource-constrained environments and can perform various speech-related tasks. Whisper (112) uses a minimal approach to weak supervision and data preprocessing, which achieves state-of-the-art results demonstrating the potential to use advanced machine learning techniques for speech processing. Whisper (112), trained on large datasets of various audio, can perform multilingual speech recognition, speech translation, and language identification.

[0063] The style encoder (113) extracts tone data or style data(s) from the reference voice that convey non-verbal attributes not included in the text of the original voice data.

[0064] The pitch extractor (114) extracts the pitch of the voice from the original voice data. Here, pitch is a frequency characteristic of the voice signal that determines the pitch of the voice and reflects the emotional expression or nuance of pronunciation. Pitch information can mainly be represented by the F0 (Fundamental Frequency) value.

[0065] In this way, when the data extraction module (110) extracts content data, style data, and pitch data from the original voice data, it can extract content data and style data by applying an information distortion technique using content-preserved perturbation (f(·)) information and pitch perturbation (g(·)) information.

[0066] At this time, the information distortion function equation for content preservation variation can be expressed by the following mathematical formula 1, and the information distortion function equation for pitch variation information can be expressed by the following mathematical formula 2.

[0067]

[0068]

[0069] In the above mathematical formulas 1 and 2, h fs , h pr , h peq , h pm represents formant shifting, pitch randomization, parametric equalizer function, and pitch normalization function, respectively, and each information distortion function receives voice data as input. The data extraction module (110) can effectively extract content data and style data from voice data through the corresponding functions.

[0070] Here, h fs ε is the ratio for how much formant shift to perform, and is randomly sampled from a uniform distribution U(1, 1.4). In other words, after randomly sampling the shift ratio, it is randomly decided once again whether to apply that shift ratio. h pr The pitch range ratio and pitch shift ratio are randomly sampled from U(1, 2) and U(1, 1.5), respectively, and then randomly determined whether to apply the corresponding ratio values. Here, the randomly sampled random frequency is a randomly set frequency shift or range from a uniform distribution and applied randomly.

[0071] The source filter encoding module (120) synthesizes content data (c), style data (s), and prosodic data (F0) based on source filter theory to create a target Mel spectrogram (Z mel Creates ).

[0072] The source filter encoding module (120) includes a source encoder (121), a filter encoder (122), and a summer (123), as shown in FIG. 3.

[0073] The source encoder (121) receives pitch data (F0), which is the pitch of the sound, and style data (s) as input, and the source Mel spectrogram (Z src ) generates, and the filter encoder (122) receives content data (c) and style data (s) as input and filters the Mel spectrogram (Z ftr It generates ). The summer (123) is Z src Wow Z ftr Target Mel spectrogram X through element-wise addition operations between mel Z approximated by mel Creates.

[0074] The source filter encoding module (120) is a source Mel spectrogram (Z) as shown in Equation 3 below. src It can display the filter MEL spectrogram (Zftr) and target MEL spectrogram.

[0075]

[0076]

[0077]

[0078] The data conversion module (130) uses a pre-trained diffusion model to restore a Mel spectrogram that reflects the timbre of the speaker's voice based on the target Mel spectrogram. The data conversion module (130) generates a conditional part of the diffusion model, allowing pitch data (F0) and style data(s) to be input into the diffusion model as a conditional part. Accordingly, the diffusion model can be trained such that both pitch data and style data are reflected when generating speech, with pitch data and style data input as a conditional part. As a result, the present invention can generate a converted speech that accurately reflects both pitch data (pitching) and style data (timbre) after precisely extracting the timbre.

[0079] Mel spectrogram (Z mel ) is X in the diffusion model mel,T ~ N(Zmel,I It is used as ) and reverse process training is performed to generate MEL. Gaussian noise (X mel,T Original Mel X from ) mel,0 When learning the reverse process to, pitch data (F0) and style data (s) are provided as inputs conditionally so that the Mel generated from Gaussian noise can better reflect the pitch data (F0) and style data (s). That is, the diffusion model learns a distribution as shown in Equation 4 below.

[0080]

[0081] Figure 4 is a flowchart illustrating a speech conversion method using the disclosed diffusion model. To avoid redundant explanations, they are described together below.

[0082] A voice conversion method performed by a system (1) that executes a voice conversion method using a disclosed diffusion model extracts content data (c) from original voice data using an ASR model, extracts pitch data (F0) using a pitch extractor (114), and extracts style data (s) from a target speaker voice using a style encoder (113) (S110).

[0083] The source encoder (121) generates a source Mel spectrogram (Z based on pitch data (F0) and style data(s) src ) generates, and the filter encoder (122) filters the Mel spectrogram (Z) based on the content data (c) and style data(s). ftr After generating ), the summer (123) is Z src Wow Z ftr Target Mel spectrogram X through element-wise addition operations between mel Z approximated by mel Prints (S120).

[0084] The diffusion model uses Z as prior data. melWhen using and conditionally inputting pitch data (F0) and style data(s), speaker speech is restored from the Mel spectrogram to match the condition, thereby providing speech with high quality and naturalness (S130). At this time, the reverse process of the diffusion model to obtain the converted speech can be defined as shown in Equation 5 below.

[0085]

[0086] In the above mathematical formula 5, X t is the Mel spectrogram with noise added by a t-step, ZZ mel is the intermediate Mel spectrogram obtained through the source-filter encoder, s θ ε is the score predicted by the diffusion model, β t is a noise scheduling function, Each represents reverse Brownian motion.

[0087] Figure 5 is a flowchart illustrating a method for training a disclosed diffusion model.

[0088] The diffusion model is trained using the reconstruction loss and diffusion loss objective functions.

[0089] First, the processor (10) can calculate the reconstruction loss using the following mathematical formula 6 (S210).

[0090]

[0091] In the above mathematical formula 6, X mel is the target Mel spectrogram, and Z mel The Mel spectrogram is output from the source filter encoding module (120), and the processor (10) performs training so that the target Mel spectrogram and the Mel spectrogram become identical through reconstruction loss (S220).

[0092] Diffusion models smooth and transform the original data by adding noise during the forward process, and then reverse this process to generate new data from the noise during the reverse process. Each denoising step in the reverse process generally requires estimating a score function.

[0093] The processor (10) has pitch data (F0), style data(s), original voice data (X0), diffusion timestep (t), and t-step diffusion data (X t ), neural network predicted score (socre)(S θ ), log-density slope of t-step noise data( )Prior data(Z mel ) defines. At this time, the diffusion module defines the input data (X) obtained from the original voice data through the mathematical formula of the forward process in Equation 7 below. t ) and conditional(F0, s), Z m (or Z mel from s θ is predicted, and the original voice data (X0) is restored through the reverse process of the diffusion model. The processor (10) calculates the diffusion loss of the following mathematical formula 8 (S230), and, so that the diffusion model can effectively approximate the conditional score function, the score (s θ Training is performed to predict ) (S240).

[0094]

[0095]

[0096] Figure 6 is a diagram showing the results of comparing the similarity between a speaker's voice and a voice conversion method using the disclosed diffusion model and a conventional voice conversion method, and Figure 7 is a diagram showing the results of objective and subjective evaluations of the voice conversion method using the disclosed diffusion model.

[0097] In order to conduct comparative experiments on the speech conversion technology applied to the proposed model and existing models, training, validation, and testing were performed using “train-clean-100,” a sub-dataset of the LibriTTS dataset used by many researchers worldwide, and the generalization performance of the system on the new dataset was tested using another dataset, VCTK. The LibriTTS “train-clean-100” dataset consists of 58.78 hours of speech from 247 speakers, and the VCTK dataset consists of 46 hours of speech from 109 speakers. All speech was downsampled to 16 kHz and converted into 80 log-scale Mel spectrograms using the Short Time Fourier Transform (STFT) and Mel filter, with the window size set to 1,280 and the window movement width set to 320. A style encoder was used to extract timbre, and a 128-dimensional vector was used. For the diffusion model, a Unet-based architecture with 64 initial channels was used, and the noise schedule parameters β0 and β1 were set to 0.05 and 20, respectively. 1e -5The speech conversion system was optimized for 150 epochs with a batch size of 64 using the AdamW optimizer (β1= 0.8, β2= 0.99) with a learning rate, and the vocoder synthesized speech from Mel spectrograms using a pre-trained HiFi-GAN. All experiments were conducted in the PyTorch toolkit, and the training process was accelerated using a single NVIDIA GeForce RTX 30990 GPU with 24GB of memory.

[0098] Figure 6 shows a comparison of voice quality and similarity to the target speaker's voice for an unseen speaker during training, using the existing voice conversion method and the proposed voice conversion method. Ground Truth refers to the actual voice, and Ground Truth (Mel + Vocoder) refers to the voice generated by converting the actual voice into a Mel spectrogram and feeding it as the input to a pre-trained vocoder. AutoVC and Diff-HierVC are control models, and ProposedVC refers to the model proposed in this invention. SECS stands for speaker encoder cosine similarity and is an objective indicator evaluating whether the converted voice closely mimics the target speaker's voice. WER and CER are the word error rate and character error rate, respectively, which are objective indicators obtained by evaluating the converted voice using a speech recognition system. nMOS and sMOS stand for naturalness mean opinion score and similarity mean opinion score, respectively, and are subjective evaluation indicators based on an audience.

[0099] Figure 7 shows the objective and subjective evaluation of the methods of the proposed model, including the method with Multi Content (MC) added, the method with Pitch Perturbation (PP) added.

[0100] As shown in FIGS. 6 and 7, it was confirmed that the performance of the proposed model is superior to that of comparison models, and it can be seen that the method of applying the information distortion function proposed in the present invention is highly effective in improving the quality of the converted speech and speaker similarity.

[0101] Although FIGS. 3 and 4 describe each process as being executed sequentially, this is merely an illustrative explanation of the technical concept of one embodiment of the present invention. In other words, a person skilled in the art to which one embodiment of the present invention belongs can modify and adapt the process in various ways, such as changing the order described in each figure or executing one or more of the processes in parallel, without departing from the essential characteristics of one embodiment of the present invention; therefore, FIGS. 3 and 4 are not limited to a chronological order.

[0102] Meanwhile, the processes illustrated in FIGS. 3 and 4 can be implemented as computer-readable code on a computer-readable recording medium. A computer-readable recording medium includes all types of recording devices in which data that can be read by a computer system is stored. That is, a computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, floppy disk, hard disk, etc.) and optical reading media (e.g., CD-ROM, DVD, etc.). In addition, the computer-readable recording medium can be distributed across networked computer systems, allowing computer-readable code to be stored and executed in a distributed manner.

Claims

1. A speech conversion method using a diffusion model, performed by a system comprising at least one processor, wherein A data extraction process that separates content data and pitch data based on original voice data, and extracts style data from a target speaker's voice; A source filter encoding process that synthesizes the content data, style data, and pitch data based on source filter theory to generate a Mel spectrogram reflecting the timbre of the speaker's voice; and A data conversion process that uses a pre-trained diffusion model to conditionally provide the style data and pitch data, and restores the speaker's voice from the Mel spectrogram based on the condition; A method that includes 2. In Paragraph 1, The above data extraction process is, A method for extracting content data using the architecture of a pre-trained Automatic Speech Recognition (ASR) model.

3. In Paragraph 1, The above data extraction process is, A method for extracting content data and style data using an information distortion function that utilizes content-preserved perturbation (f(·)) information and pitch perturbation (g(·) ) information.

4. In Paragraph 3, The information distortion function equation for the above-mentioned content preservation variation is represented by the following Equation 1, and the information distortion function equation for the above-mentioned pitch variation information is represented by the following Equation 2. [Mathematical Formula 1] [Mathematical Formula 2] h fs : formant shifting, h pr : Pitch randomization, h peq : Parametric equalizer function, h pm : Each representing a pitch normalization function, method.

5. In Paragraph 1, The above source filter encoding process is, Based on the above pitch data (F0) and style data(s), the source Mel spectrogram (Z src Step of generating ); Based on the above content data (c) and style data(s), filter Mel spectrogram (Z ftr Step of generating ); and The above source mel spectrogram (Z src ) and filter Mel spectrogram (Z ftr Target Mel spectrogram (X) through element-wise addition operations between ) mel Mel spectrogram (Z) approximated by ) mel Step of generating ); A method that includes 6. In Paragraph 1, The learning process further includes training the above diffusion model using objective functions of reconstruction loss and diffusion loss. method.

7. In Paragraph 6, The above learning process is, Using the following mathematical formula, the above reconstruction loss (L recon Calculate ) and the above reconstruction loss (L recon Target Mel Spectrogram (X) through ) mel ) and Mel spectrogram (Z mel It is proceeding with learning so that ) become identical to each other, [Mathematical Formula] method.

8. In Paragraph 6, The above learning process is, Using the following mathematical formula, the above diffusion loss (L diff Calculate ) and the above diffusion loss (L diff Using ), the above diffusion model scores (s θ ) which makes it predict, [Mathematical Formula] F0: Pitch data, s : style data, X0: Original voice data, t : diffusion timestep, X t : t-step diffusion data, : Representing the log-density slopes of the t-step noise data, method.

9. As a system that performs speech conversion using a diffusion model, A data extraction module that separates content data and pitch data based on original voice data, and extracts style data from a target speaker's voice; A source filter encoding module that synthesizes the content data, style data, and pitch data based on source filter theory to generate a Mel spectrogram reflecting the timbre of the speaker's voice; and A data conversion module that uses a pre-trained diffusion model to conditionally provide the style data and pitch data, and restores the speaker's voice from the Mel spectrogram based on the condition; including, System.

10. In Paragraph 9, The above data extraction module is, Extracting content data from the original voice data using the architecture of an automatic speech recognition (ASR) model, System.

11. In Paragraph 9, The above data extraction module is, Further including a style encoder that extracts style data(s) from a reference voice that convey non-verbal attributes not included in the text of the original voice data, System.

12. In Paragraph 9, The above data extraction module is, The method further includes a pitch extractor that extracts pitch data (F0) of the voice from the original voice data. System.

13. In Paragraph 9, The above data extraction module is, Extracting the above content data and style data using an information distortion function that utilizes content-preserved perturbation (f(·)) information and pitch perturbation (g(·) ) information, System.

14. In Paragraph 13, The information distortion function equation for the above-mentioned content preservation variation is represented by the following Equation 1, and the information distortion function equation for the above-mentioned pitch variation information is represented by the following Equation 2. [Mathematical Formula 1] [Mathematical Formula 2] h fs : formant shifting, h pr : Pitch randomization, h peq : Parametric equalizer function, h pm : Each representing a pitch normalization function, System.

15. In Paragraph 9, The above source filter encoding module is, Based on the above pitch data (F0) and style data(s), the source mel spectrogram (Z src Source encoder that generates ); Based on the above content data (c) and style data(s), filter Mel spectrogram (Z ftr A filter encoder that generates ); and The above source mel spectrogram (Z src ) and filter Mel spectrogram (Z ftr Target Mel spectrogram (X) through element-wise addition operations between ) mel Mel spectrogram approx. )(Z mel Summer that outputs ); including, System.

16. A computer program stored on a computer-readable storage medium, wherein the computer program performs operations for speech conversion using a diffusion model when executed on one or more processors, and The above operations are, An operation to separate content data and pitch data based on original voice data, and to extract style data from the target speaker's voice; Based on source filter theory, the operation of synthesizing the content data, style data, and pitch data to generate a Mel spectrogram reflecting the timbre of the speaker's voice; and An operation of conditionally providing style data and pitch data using a pre-trained diffusion model, and restoring speaker speech from the Mel spectrogram based on the condition; including, Computer program.