Text-to-speech conversion method based on few sampling steps and system thereof
Patent Information
- Application Number
- PCT/KR2026/004502
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2026-03-17
- Filing Date
- 2026-03-20
- Publication Date
- 2026-09-24
Smart Images

Figure KR2026004502_24092026_PF_FP_ABST
Abstract
Description
A small number of sampling steps-based text-to-speech conversion method and system
[0001] The present disclosure relates to a text-to-speech conversion method and system based on a small number of sampling steps, and more specifically, to a method and system that generates a synthetic speech reflecting speaker characteristics extracted from a speaker's speech sample using a flow matching-based decoder, performs inference with fewer sampling steps than during pre-training, and further learns the generated output through adversarial learning with a discriminator, thereby outputting high-quality synthetic speech in real time with only a small number of steps.
[0002] Text-to-Speech (TTS) technology converts input text into natural-sounding speech and is widely used in various fields, including video conferencing assistants, smart speakers, voice assistants, and accessibility tools. In particular, it has recently emerged as a core component for implementing real-time conversational AI agents when combined with large-scale language models.
[0003] Conventional TTS systems have largely evolved into statistical-based and deep learning-based methods. Although statistical-based methods modeled speech using Hidden Markov Models (HMMs), they had limitations in that the naturalness and speaker similarity of the generated speech did not reach the level of human speech. Subsequently, with the advancement of deep learning technology, neural network-based models such as WaveNet, Tacotron, and FastSpeech emerged, significantly improving sound quality and naturalness.
[0004] However, among deep learning-based TTS models, approaches based on diffusion models, in particular, achieve high-quality speech synthesis but have a fundamental problem in that they require dozens of iterative sampling steps during the inference process. For example, modern models such as F5-TTS and E2-TTS typically require more than 32 sampling steps for high-quality speech synthesis, which results in a degraded Real-Time Factor (RTF), making it difficult to apply to real-time applications that are sensitive to latency.
[0005] To address these issues, a flow matching-based generative model has been proposed. Flow matching-based models can estimate target distributions with fewer sampling steps than diffusion models by estimating paths that are close to a straight line between data distributions. In particular, Optimal Transport (OT)-based Conditional Flow Matching (CFM) can further reduce the number of sampling steps by optimizing probabilistic flow paths to learn paths that are even closer to a straight line.
[0006] However, even with flow matching-based models, reducing the number of inference steps to a minimum of 2 to 4 results in a problem where speech quality deteriorates rapidly. In other words, additional technical means are required to compensate for the quality degradation caused by the reduction in the number of steps. Conventional knowledge distillation methods involve a complex process of training a student model from a teacher model, and have limitations in that it is difficult to achieve sufficient performance in terms of speaker similarity.
[0007] Meanwhile, when training Speech-to-Text (STT) models using voice data collected from specific acoustic environments, such as conference rooms, the lack of training data for diverse speakers and acoustic conditions becomes a major factor limiting model performance. To address this, data augmentation methods utilizing speech synthesis technology are gaining attention; however, there is a problem in that TTS systems lacking real-time capability are difficult to apply to large-scale data augmentation pipelines.
[0008] Therefore, a new TTS technology is required that can generate high-quality synthesized speech in real time, satisfying both speaker similarity and sound quality with only a very small number of sampling steps.
[0009] One embodiment of the present disclosure is devised to solve the problems of the prior art as described above, and aims to provide a method and system capable of synthesizing high-quality speech in real time with only a few sampling steps by applying adversarial post-training to a flow matching-based speech synthesis model.
[0010] In addition, one embodiment of the present disclosure aims to provide a zero-shot TTS method and system that generates a synthesized speech corresponding to an input text while preserving the speech characteristics and acoustic environment of a target speaker using only a short speech sample of the target speaker.
[0011] Furthermore, one embodiment of the present disclosure aims to provide a method and system for augmenting data usable for training a speech recognition model by rapidly generating a large amount of synthesized speech that reflects the acoustic characteristics of a specific acoustic environment using a small amount of speech samples collected in a specific acoustic environment.
[0012] However, the technical problems that the various embodiments of the present disclosure aim to solve are not limited to those described above, and there may be other technical problems that can be achieved through the technical means described in this specification.
[0013] One embodiment is,
[0014] A method executed by a computer comprises: receiving text; receiving a voice sample containing the voice characteristics of a speaker; generating a synthetic voice corresponding to the text, which reflects the speaker characteristics of the voice sample, using at least one artificial intelligence model, wherein the at least one artificial intelligence model is further trained through adversarial learning with a discriminator on an output generated through inference executed with fewer sampling steps than during prior training; and inputting the synthetic voice to at least one subsequent processing component.
[0015] In another aspect, the above method may further include the step of generating the synthesized speech in real time and outputting it through a user interface.
[0016] In another aspect, the method may further include the step of displaying at least one of the speaker similarity, sound quality, and synthesis speed of the synthesized speech through a user interface.
[0017] In another aspect, the voice sample is spoken in a first language and the text is written in a second language different from the first language, and the method may further include the step of at least one processor using the artificial intelligence model to generate a synthetic voice corresponding to the second language while maintaining the speaker characteristics of the voice sample.
[0018] In another aspect, the method may further include the steps of: receiving voice samples of each of a plurality of speakers; extracting voice characteristics of the plurality of speakers and storing them in at least one memory; and the step of the at least one processor generating a plurality of synthesized voices in real time by applying the voice characteristics of each of the plurality of speakers to at least a portion of the text.
[0019] In another aspect, the text includes a prompt text corresponding to the voice sample and a text to be synthesized, and the synthesized voice may be a method corresponding to the text to be synthesized.
[0020] In another aspect, the step of generating a synthetic voice corresponding to the text may include: encoding the voice sample into a latent expression; generating a synthetic latent expression corresponding to the target text for synthesis based on a flow matching method with a prompt extracted from the latent expression as a condition; and decoding the synthetic latent expression into a synthetic voice.
[0021] In another aspect, the flow matching method may include an Optimal Transport-based conditional flow matching method that generates the synthetic latent expression by approximating the path from the noise to the synthetic latent expression as a straight line.
[0022] In another aspect, the step of generating the synthetic latent expression may include: concatenating zero padding of a length corresponding to the synthesis target text to the latent expression of the voice sample; and generating the synthetic latent expression by infilling the segment corresponding to the zero padding.
[0023] In another aspect, the discriminator may be a method for simultaneously generating an unconditional output that determines the sound quality for the output of the at least one artificial intelligence model and a conditional output that determines the speaker similarity with the voice sample.
[0024] In another aspect, the adversarial learning may be a method based on a combination of a reconstruction loss that minimizes the difference between the output of the at least one artificial intelligence model and the latent representation of the speech sample, an adversarial loss that induces the artificial intelligence model to deceive the discriminator, and a feature matching loss that minimizes the difference between the intermediate features of the discriminator and the intermediate features of the output of the at least one artificial intelligence model.
[0025] In another aspect, the step of generating the synthesized speech may include: determining the utterance length for each phoneme of the text to be synthesized and adjusting the utterance speed of the synthesized speech according to the determined utterance length; and generating the synthesized speech that reflects the adjusted utterance speed.
[0026] In another aspect, the step of generating the synthesized speech may include: extracting fundamental frequency and voiced / unvoiced sound information from the speech sample; and generating the synthesized speech that reflects the fundamental frequency and voiced / unvoiced sound information.
[0027] In another aspect, the subsequent processing component may include a speech recognition model, the speech sample is obtained in a predetermined acoustic environment, and the synthesized speech may include a plurality of synthesized speeches reflecting the acoustic environment and be utilized as training data for the speech recognition model.
[0028] In another aspect, the step of acquiring the text comprises: a step of extracting speech from a video file or an audio file; a step of acquiring text of a first language by inputting the extracted speech into a speech recognition model; and a step of acquiring text of a second language by translating the text of the first language into a second language; and the step of acquiring the speech sample comprises: a step of acquiring a newly input speech sample or a speech sample of a selected speaker among a plurality of pre-stored speaker speeches; and the step of generating the synthesized speech may include a step of generating a synthesized speech corresponding to the text of the second language that reflects the speaker characteristics of the speech sample.
[0029] In another aspect, the step of generating the synthesized speech may include: determining the utterance length for each phoneme of the text of the second language; and adjusting the determined utterance length to be synchronized with the utterance timing of each segment of the video file to generate the synthesized speech corresponding to the text of the second language.
[0030] One embodiment is,
[0031] A system is provided comprising: at least one memory; and at least one processor that reads at least one instruction stored in the at least one memory and performs a text-to-speech conversion method based on a small number of sampling steps, wherein the at least one instruction includes the steps of: receiving text; receiving a voice sample including the voice characteristics of a speaker; and the step of the at least one processor using at least one artificial intelligence model to generate a synthetic voice corresponding to the text, which reflects the speaker characteristics of the voice sample; wherein the at least one artificial intelligence model is further trained through adversarial learning with a discriminator on an output generated through inference executed with fewer sampling steps than during prior training; and the step of inputting the synthetic voice to at least one subsequent processing component.
[0032] In another aspect, the system comprises: a plurality of neurons configured in an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons; wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights, and may further comprise a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network.
[0033] In another aspect, the system comprises: a plurality of neurons organized into an array including at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; wherein each of the plurality of neurons may further comprise an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network that is connected to at least one other neuron through any one of the plurality of synapse circuits.
[0034] A text-to-speech conversion method and system based on a small number of sampling steps according to one embodiment of the present disclosure can generate high-quality synthesized speech without degradation of sound quality due to the reduction in the number of steps by performing adversarial post-training with a discriminator while fixing the flow matching-based decoder to a significantly smaller number of sampling steps (e.g., 4 steps) than during pre-training, and can achieve a real-time processing rate that is about 4 times faster than F5-TTS (32 steps), thereby having the effect of being applicable to real-time applications sensitive to latency.
[0035] In addition, one embodiment of the present disclosure utilizes a Joint Conditional-Unconditional (JCU) discriminator that simultaneously generates an unconditional output (sound quality determination) and a conditional output (speaker similarity determination), thereby enabling adversarial learning that simultaneously improves sound quality and speaker similarity. This has the effect of generating synthetic speech that more accurately reproduces speaker characteristics compared to conventional GAN-based methods that simply optimize sound quality.
[0036] In addition, one embodiment of the present disclosure employs an optimal transport (OT)-based conditional flow matching method to approximate the path from noise to a synthetic potential representation as close to a straight line, thereby enabling the generation of a synthetic potential representation close to a target distribution with a small number of sampling steps, and thus provides a basis for high-quality speech synthesis with only a few steps.
[0037] Furthermore, one embodiment of the present disclosure employs an infilling method that encodes a speaker's voice sample into a latent expression and utilizes it as a prompt, thereby enabling precise control that explicitly separates the speaker's prompt text and the text to be synthesized, allowing only the voice to be synthesized to be generated.
[0038] In addition, one embodiment of the present disclosure has the effect of enabling the construction of a data augmentation pipeline that improves the performance of a speech recognition model under various acoustic conditions by using a speaker voice sample obtained in a specific acoustic environment (e.g., a conference room) to generate a plurality of synthesized voices reflecting the characteristics of the said acoustic environment and augmenting them as training data for a speech recognition model.
[0039] However, the effects obtainable from the present disclosure are not limited to those mentioned above, and there may be other effects that can be clearly understood by a person skilled in the art to which the present invention pertains through the configurations described above.
[0040] FIG. 1 illustrates an example of a block diagram of a computing system implementing a text-to-speech conversion method based on a small number of sampling steps according to one embodiment of the present disclosure.
[0041] FIG. 2 briefly illustrates the structure of a neuromorphic circuit that may be included in a processor according to one embodiment.
[0042] FIG. 3 illustrates an example of a block diagram of a computing device implementing a text-to-speech conversion method based on a small number of sampling steps according to one embodiment of the present disclosure.
[0043] FIG. 4 is a block diagram illustrating the internal architecture and data processing pipeline of an artificial intelligence model according to one embodiment of the present disclosure.
[0044] FIG. 5 illustrates an example of a block diagram in another aspect of a computing device implementing a text-to-speech conversion method based on a few sampling steps according to one embodiment of the present disclosure.
[0045] FIG. 6 is a block diagram illustrating the data flow and system interaction of a small number of sampling steps-based text-to-speech conversion service application process according to one embodiment of the present disclosure.
[0046] FIG. 7 illustrates the architecture of a Universal Dynamic Multi-Agent System according to one embodiment of the present disclosure.
[0047] FIG. 8 is a flowchart of a text-to-speech conversion method based on a small number of sampling steps according to one embodiment of the present disclosure.
[0048] FIG. 9 is a detailed flowchart of the steps for generating synthetic latent expressions based on latent expression encoding and flow matching according to one embodiment of the present disclosure.
[0049] FIG. 10 is an overall block diagram of a text-to-speech system based on a small number of sampling steps according to one embodiment of the present disclosure.
[0050] FIG. 11 is a block diagram illustrating the learning structure of a text-to-speech conversion model according to one embodiment of the present disclosure.
[0051] FIG. 12 is a block diagram illustrating the inference structure of a text-to-speech conversion model according to one embodiment of the present disclosure.
[0052] FIG. 13 is a block diagram illustrating the structure of a combined conditional-unconditional (JCU) discriminator according to one embodiment of the present disclosure.
[0053] FIG. 14 illustrates an exemplary user interface for displaying speaker similarity, sound quality, and synthesis speed of a synthesized voice according to one embodiment of the present disclosure.
[0054] FIG. 15 illustrates an exemplary user interface that supports multilingual speaker characteristic preservation synthesis according to one embodiment of the present disclosure.
[0055] FIG. 16 illustrates an exemplary user interface that supports real-time synthetic voice switching of multiple speakers according to one embodiment of the present disclosure.
[0056] FIG. 17 illustrates steps 1 to 3 of a synthetic speech-based data augmentation pipeline for training a speech recognition model according to one embodiment of the present disclosure.
[0057] FIG. 18 illustrates, exemplarily, four steps of a synthetic speech-based data augmentation pipeline for training a speech recognition model according to one embodiment of the present disclosure.
[0058] FIG. 19 illustrates an exemplary user interface that supports automatic dubbing according to one embodiment of the present disclosure.
[0059] As the embodiments of the present disclosure are subject to various modifications and may have various embodiments, specific embodiments are illustrated in the drawings and described in detail in the detailed description. The effects and features of the present disclosure and the methods for achieving them will become clear by referring to the embodiments described below in detail together with the drawings. However, the present disclosure is not limited to the embodiments disclosed below but may be implemented in various forms. In the following embodiments, terms such as "first," "second," etc., are used not in a limiting sense but for the purpose of distinguishing one component from another. Also, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, terms such as "include" or "have" mean that the features or components described in the specification exist, and do not preclude the possibility that one or more other features or components may be added. Additionally, in the drawings, the size of components may be exaggerated or reduced for convenience of explanation. For example, the size and thickness of each component shown in the drawings are arbitrarily depicted for convenience of explanation, so the embodiments of the present disclosure are not necessarily limited to those depicted.
[0060] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings. When describing with reference to the drawings, identical or corresponding components are given the same reference numerals, and redundant descriptions thereof will be omitted.
[0061]
[0062] - A text-to-speech system based on a small number of sampling steps (1000)
[0063] A system (1000) according to one embodiment extracts speaker characteristics from a speaker's voice sample and can generate synthetic speech corresponding to an input text in real time using a flow matching-based decoder with only a small number of sampling steps. By solving the problem of multiple sampling steps associated with conventional diffusion model-based TTS through adversarial post-training (AP) in a fixed state with fewer steps than during pre-training, the system (1000) achieves a real-time throughput (RTF) level applicable even to latency-sensitive real-time applications.
[0064] In addition, the present system (1000) can simultaneously optimize sound quality and speaker similarity through a combined conditional-unconditional (JCU) discriminator, thereby achieving a speaker similarity score (SMOS) and word error rate (WER) that are equivalent to or superior to existing models with only a few steps. As such, the present system (1000) can be implemented not as a simple set of software algorithms, but as a technical system in which specialized artificial intelligence models such as VAE, flow matching decoders, and JCU discriminators are organically combined with physical computational resources that control them.
[0065] FIG. 1 illustrates an example of a block diagram of a computing system (1000) implementing a small number of sampling steps-based text-to-speech conversion method according to one embodiment of the present disclosure.
[0066] Referring to FIG. 1, a computing system (1000) implementing a text-to-speech conversion service based on a small number of sampling steps according to one embodiment includes a user computing device (110), a server computing system (130), and a training computing system (150), and the devices can communicate through a network (170).
[0067]
[0068] A text-to-speech conversion method based on a small number of sampling steps according to one embodiment of the present disclosure may be implemented and provided locally by a user computing device (110), implemented and provided in the form of a web service by a server computing system (130) communicating with the user computing device (110), or implemented and provided by the user computing device (110) and the server computing system (130) in conjunction with each other.
[0069] In this embodiment, the user computing device (110) and / or server computing system (130) can train a machine learning model (120 and / or 140) through interaction with a training computing system (150) that is communicatedly connected via a network (170).
[0070] Additionally, in embodiments of the present disclosure, machine learning models (120, 140) may include various forms of AI agents or agentic architectures beyond simple prediction models.
[0071] In one embodiment, the machine learning model may be an AI agent having a structure that receives system prompts and user prompts, autonomously plans and executes tasks through core components such as planning, memory, reasoning, and tools, and improves itself through feedback.
[0072] In another embodiment, the machine learning model may include a Search Augmented Generative (RAG) architecture that retrieves relevant information from an external database and generates a response based thereon to provide an accurate answer based on the latest information or expertise.
[0073] In another embodiment, the machine learning model may be a multi-agent system in which multiple AI agents cooperate to achieve a specific goal. The multi-agent system may have a supervisory pattern in which a central supervisor agent distributes tasks to subordinate expert agents and aggregates the results. Alternatively, it may have a hierarchical pattern in which a meta-agent acts as an intermediary manager to control and coordinate subordinate agents. Furthermore, it is possible to include a multi-agent debate pattern in which multiple agents present different opinions and derive an optimal conclusion through discussion and evaluation.
[0074] The server computing system (130) can host AI agents such as those mentioned above, particularly multi-agent systems requiring complex computations or large-scale long-term memory. Additionally, the server computing system (130) includes a Multi-Channel Processing (MCP) server to manage integration with various external tools, and can relay communication with cloud APIs, payment services, search engines, etc.
[0075] The training computing system (150) can generate a ToolFormer model that learns how to use a specific tool, or perform iterative learning that gradually improves the performance of the agent through a self-reflection mechanism in which another LLM evaluates and modifies the results generated by the agent.
[0076] The training computing system (150) may be separate from the server computing system (130) or part of the server computing system (130). Additionally, in some embodiments, the training computing system (150) may be separate from the user computing device (110) or part of the user computing device (110).
[0077] And at this time, the artificial intelligence model can be 1) trained directly locally by a user computing device (110), 2) trained by the server computing system (130) and the user computing device (110) interacting with each other through a network (170), and 3) trained by a separate training computing system (150) using various training and learning techniques. It may also be implemented by transmitting the artificial intelligence model trained by the training computing system (150) to the user computing device (110) and / or the server computing system (130) through the network (170) to provide / update it.
[0078]
[0079] - User Computing Device (110: User Computing Device)
[0080] The user computing device (110) may include all other types of computing devices, such as a smartphone, a mobile phone, a digital broadcasting device, a PDA (personal digital assistants), a PMP (portable multimedia player), a desktop, a wearable device, an embedded computing device and / or a tablet PC.
[0081] The user computing device (110) may include at least one processor (111) and memory (112). Here, the processor (111) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.
[0082] In particular, according to the embodiment, this processor (111) may be configured based on a Field Programmable Gate Array (FPGA) implementation and / or an Application Specific Integrated Circuit (ASIC), which is a hardware technology for implementing a certain digital circuit.
[0083] Here, a field programmable gate array (FPGA) can refer to a flexible digital circuit that is programmable according to user needs.
[0084] As an example, a field programmable gate array implementation may include a register that temporarily stores data and controls the flow and timing of signals to maintain intermediate results or state information of operations to support synchronized operation of the FPGA, programmable logic that programs operations within the FPGA to perform specific functions or operations as logic circuits configurable according to user needs, and an input interface that receives signals from external devices or sensors and transmits them to internal circuits as a channel for receiving data from outside the FPGA.
[0085] Through the combination of the above components, a field-programmable gate array implementation can provide flexible and various types of digital circuits.
[0086] As an example, the application-dedicated integrated circuit may include a register, which is a small memory device for temporarily storing and managing data and supports the rapid processing of ASIC operations by storing intermediate calculation results or state information; a microprocessor, which is a central processing unit that performs control and operations within the ASIC and coordinates the operation of the entire system by performing various operations or generating control signals when necessary; and an input block, which is an interface for receiving data from the outside, which receives data to be processed by the ASIC and transmits it internally, and receives various input data through connections with sensors or external devices.
[0087] Through the combination of the components mentioned above, an application-specific integrated circuit can perform specific purpose tasks in an optimized manner.
[0088] For example, ASICs can have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits.
[0089] FIG. 2 briefly illustrates the structure of a neuromorphic circuit (600) that may be included in a processor (111, 131, 151) according to one embodiment.
[0090] Referring to FIG. 2, for example, a neuromorphic circuit (600) may include a plurality of presynaptic neuron circuits (610), a plurality of presynaptic lines (611) extending laterally from the plurality of presynaptic neuron circuits (610), a plurality of postsynaptic neuron circuits (620), a plurality of postsynaptic lines (621) extending longitudinally from the plurality of postsynaptic neuron circuits (620), and a plurality of synaptic circuits (630) provided at the intersection of the plurality of presynaptic lines (611) and the plurality of postsynaptic lines (621).
[0091] A plurality of free synaptic neuron circuits (610) can transmit signals input from the outside in the form of electrical signals to a plurality of synaptic circuits (630) through a plurality of free synaptic lines (611).
[0092] Additionally, a plurality of post-synaptic neuron circuits (620) can receive electrical signals from a plurality of synaptic circuits (630) through a plurality of post-synaptic lines (621).
[0093] Furthermore, multiple post-synaptic neuron circuits (620) may transmit electrical signals to multiple synaptic circuits (630) through multiple post-synaptic lines (621).
[0094] A plurality of synapse circuits (630) can store weights included in layers constituting a neural network system implemented by a neuromorphic circuit (600) and perform a predetermined operation based on the weights and input data.
[0095] For example, each of the plurality of synaptic circuits (630) may include a resistive memory cell having a variable resistance. In this case, the resistance value of the plurality of synaptic circuits (630) changes by a voltage applied through the plurality of presynaptic neuron circuits (610) or the plurality of postsynaptic neuron circuits (620), and can store weight data according to this resistance change.
[0096] The neuromorphic circuit (600) is formed by mimicking the structure of neurons and synapses, which are essential elements of the human brain. When a deep neural network (DNN) is realized using the neuromorphic circuit (600), the data processing speed can be improved and power consumption can be reduced compared to when the existing von Neumann structure is utilized.
[0097] The memory (112) of the user computing device (110) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc., and combinations thereof, and may include web storage of a server that performs memory storage functions on the internet. This memory (112) may store data (113) and instructions (114) necessary for the at least one processor (111) to perform functional operations such as training an artificial intelligence model or performing data filtering through an artificial intelligence model.
[0098] In one embodiment, the user computing device (110) may store at least one machine learning model (120). For example, the user computing device (110) may be various machine learning models, such as a plurality of neural networks (e.g., deep neural networks) that perform a text-to-speech conversion method based on a few sampling steps, or other types of machine learning models including non-linear models and / or linear models, and may be composed of a combination thereof.
[0099] For example, the machine learning model (120) may include linear regression, decision tree, random forest, gradient boosting or / and a few sampling steps-based text-to-speech model. And the neural network may include at least one of feed-forward neural networks, recurrent neural networks (e.g., long short-term memory recurrent neural networks), convolutional neural networks, Transformer or / and other forms of neural networks.
[0100] Additionally, according to an embodiment, the user computing device (110) may store a model to be used in each process and a prompt template that serves as the basis for input to the model in order to perform at least part of the process for a small number of sampling steps-based text-to-speech conversion method through a large language model (LLM).
[0101] In one embodiment, a user computing device (110) receives at least one machine learning model (120) from a server computing system (130) through a network (170), stores it in memory (112), and then executes the stored machine learning model (120) by a processor (111) to perform a text-to-speech conversion method based on a few sampling steps, etc.
[0102] In another embodiment, the user computing device (110) can provide a text-to-speech conversion service based on a small number of sampling steps to the user by performing operations through a machine learning model (140) including at least one machine learning model (140) in conjunction with a server computing system (130) and communicating related data externally.
[0103] For example, a user computing device (110) can provide a text-to-speech conversion service based on a small number of sampling steps in which a server computing system (130) provides an output for the user's input using a machine learning model (140) via the web.
[0104] Additionally, the artificial intelligence model can be implemented in such a way that at least some of the machine learning models (120 and / or 140) are executed on a user computing device (110) and the rest are executed on a server computing system (130).
[0105] Additionally, the user computing device (110) may include at least one input component (121) that detects user input.
[0106] For example, the user input component (121) may include a touch sensor (e.g., a touch screen and / or a touch pad, etc.) that detects a touch of the user's input medium (e.g., a finger or a stylus), an image sensor that detects the user's motion input, a microphone that detects the user's voice input, a button, a mouse and / or a keyboard, etc.
[0107] Here, the image sensor may include an image processing module. Specifically, the image sensor may process still images or video obtained by an image sensor device (e.g., CMOS or CCD).
[0108] In addition, the image sensor can process a still image or video acquired through the image sensor device using an image recognition process (e.g., OCR, etc.) and / or an image processing module to extract necessary information and transmit the extracted information to a processor.
[0109] Additionally, the input component (121) can receive input from an external controller (e.g., mouse, keyboard, etc.) based on an interface module, and in this case, may include an external output device (e.g., speaker).
[0110] At this time, the interface module may be configured to include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module, an audio I / O (Input / Output) port, a video I / O (Input / Output) port, an earphone port, a power amplifier, an RF circuit, a transceiver, and other communication circuits.
[0111] In addition, the external output device may include a display system that outputs various information related to a text-to-speech conversion service based on a small number of sampling steps as a graphic image.
[0112] Such a display system may be implemented by including at least one of a liquid crystal display (LCD), a thin film transistor-liquid crystal display (TFT LCD), an organic light-emitting diode (OLED), a flexible display, a 3D display, and an e-ink display.
[0113] Meanwhile, the user computing device (110) including the above-described components may further perform at least some of the functional operations performed by the server computing system (130) described later.
[0114]
[0115] -Server Computing System (130: Server Computing System)
[0116] The server computing system (130) can perform a series of processes to provide a few sampling steps-based text-to-speech conversion services.
[0117] In detail, in an embodiment, the server computing system (130) can provide a small number of sampling steps-based text-to-speech conversion service by exchanging data necessary to drive a small number of sampling steps-based text-to-speech conversion service process with an external device, such as a user computing device (110).
[0118] More specifically, in an embodiment, the server computing system (130) can provide an environment in which an application for providing a few sampling steps-based text-to-speech conversion service on a user computing device (110) can operate.
[0119] To this end, the server computing system (130) may include an application program, data and / or instructions, etc. for the application to operate, and may transmit and receive various data based thereon with the external device.
[0120] A server computing system (130) may include at least one processor (131) and memory (132). Here, the processor (131) of the server computing system (130) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or electrical units for performing other functions, or a plurality of electrically connected processors.
[0121] For example, ASICs may have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits (see Fig. 2).
[0122] And the memory (132) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (132) may store data (133) and instructions (134) necessary for the processor (131) to perform functional operations, such as training an artificial intelligence model or executing a text-to-speech conversion method based on a small number of sampling steps.
[0123] In one embodiment, the server computing system (130) may be implemented to include at least one computing device. For example, the server computing system (130) may be implemented to operate a plurality of computing devices according to a sequential computing architecture, a parallel computing architecture, or a combination thereof. Additionally, the server computing system (130) may include a plurality of computing devices connected to a network (170).
[0124] Additionally, the server computing system (130) may store at least one machine learning model (140). For example, the server computing system (130) may include a neural network and / or other multi-layer non-linear model as the machine learning model (140). Exemplary neural networks may include a feed-forward neural network, a deep neural network, a recurrent neural network, and a convolutional neural network.
[0125] In an embodiment, the server computing system (130) may further include a data store computing system (hereinafter, data store) which is a storage for continuously storing and managing raw data that forms the basis of a small number of sampling steps-based text-to-speech conversion service.
[0126] Such data stores may include various forms of data storage, ranging from file systems to cloud storage. For example, a data store may include at least one database among a relational database that uses a structured query language (SQL) to define and manipulate data, a NoSQL database designed for flexibility and scalability to process unstructured and semi-structured data, a data warehouse optimized for querying and analysis by centralizing large volumes of data from multiple sources as a system used for reporting and data analysis, a data warehouse that stores large volumes of raw data in basic formats such as structured data, semi-structured data, and unstructured data, and a local storage device or Network Attached Storage (NAS) that stores data in files in a format generally accessible by a computer operating system.
[0127]
[0128] - Training Computing System (150: Training Computing System)
[0129] The training computing system (150) may include at least one processor (151) and memory (152). Here, the processor (151) of the training computing system (150) may be composed of at least one of a central processing unit (CPU), a graphics processing unit (GPU), ASICs (application specific integrated circuits), DSPs (digital signal processors), DSPDs (digital signal processing devices), PLDs (programmable logic devices), FPGAs (field programmable gate arrays), controllers, microcontrollers, microprocessors, and / or other electrical units for performing functions, or a plurality of electrically connected processors.
[0130] For example, ASICs may have a structure of a neuromorphic circuit in the form of an array containing multiple neuron circuits (see Fig. 2).
[0131] And the memory (152) may include one or more non-transient / transient computer-readable storage media such as RAM, ROM, EEPROM, EPROM, flash memory device, magnetic disk, etc. and combinations thereof. This memory (152) may store data (153) and instructions (154) necessary for the processor (151) to perform learning of an artificial intelligence model, etc.
[0132] For example, the training computing system (150) may include a model trainer (160) that trains a machine learning model (120 and / or 140) stored in a user computing device (110) and / or a server computing system (130) using various training or learning techniques, such as back propagation of error.
[0133] For example, such a model trainer (160) can perform updates to one or more parameters of a machine learning model (120 and / or 140) for a small number of sampling steps-based text-to-speech conversion service in a backpropagation manner based on a defined loss function.
[0134] In some embodiments, performing backpropagation of the error may include performing truncated backpropagation through time. The model trainer (160) may perform a number of generalization techniques (e.g., weight devaluation, dropout and / or knowledge distillation, etc.) to improve the generalization ability of the machine learning model (120 and / or 140) being trained.
[0135] Additionally, the model trainer (160) can train a machine learning model (120 and / or 140) based on a series of training data (161). Here, the training data (161) may include data of different forms, such as images, audio samples and / or text, for example. Examples of image types that may be used may include video frames, LiDAR point clouds, X-ray images, computed tomography scans, hyperspectral images and / or various other forms of images.
[0136] These training data (161) may be provided by a user computing device (110) and / or a server computing system (130). When the training computing device trains a machine learning model (120 and / or 140) on specific data of the user computing device (110), the machine learning model (120 and / or 140) may be characterized as a personalized model.
[0137] And the model trainer (160) includes computer logic that is utilized to provide the desired function.
[0138] Additionally, the model trainer (160) may be implemented as hardware, firmware, and / or software that controls a general-purpose processor. In one embodiment, the model trainer (160) may include a program file stored in a storage device, be loaded into memory (152), and be executed by one or more processors (151). In another embodiment, the model trainer (160) includes one or more sets of computer-executable data (153) and instructions (154) stored in a tangible computer-readable storage medium, such as a RAM hard disk or an optical or magnetic medium.
[0139] In this system (1000), communication between the user computing device (110) and the server computing system (130) can be established via a wired / wireless network (170).
[0140] Network (170) includes, but is not limited to, 3GPP (3rd Generation Partnership Project) network, LTE (Long Term Evolution) network, WIMAX (World Interoperability for Microwave Access) network, Internet, LAN (Local Area Network), Wireless LAN (Wireless Local Area Network), WAN (Wide Area Network), PAN (Personal Area Network), Bluetooth network, satellite broadcasting network, analog broadcasting network and / or DMB (Digital Multimedia Broadcasting) network.
[0141] Generally, communication through the network (170) can be performed using any type of wired and / or wireless connection through various communication protocols (e.g., TCP / IP, HTTP, SMTP and / or FTP, etc.), encodings or formats (e.g., HTML and / or XML, etc.), and / or protection schemes (e.g., VPN, Secure HTTP and / or SSL, etc.).
[0142] The server computing system (130) may further include a plurality of specialized engines and repositories that are logically and physically separated to perform a few sampling step-based text-to-speech conversion methods.
[0143] In one embodiment, the server computing system (130) may include at least one engine among a reasoning engine that processes user natural language input or system events to establish an action plan, a membership management engine, and a supervision signal generation engine.
[0144] Here, the term engine may include not only a set of instructions executed by a processor (131) to perform specific logic, but also dedicated hardware circuits to accelerate said logic.
[0145] Such engines may run on hardware accelerators optimized to handle the computational load of large-scale generative artificial intelligence models. The hardware accelerators are processors specialized for matrix operations and vector processing and may include at least one of the aforementioned Tensor Processing Unit (TPU), Graphics Processing Unit (GPU), Field-Programmable Gate Array (FPGA), or Application-Specific Integrated Circuit (ASIC). These hardware accelerators can provide technical improvements that distribute the computational load of large-scale language models (LLM) and / or diffusion models and enable real-time inference.
[0146] Additionally, the data (133) may include speaker embedding vectors, text-speech pair training data, and learned parameters of a flow matching model. The server computing system (130) stores and manages multiple speaker profiles by mapping them to speaker identifiers, thereby enabling the system to quickly retrieve the embeddings of the corresponding speaker upon a real-time synthesis request and provide them as condition inputs to the flow matching decoder (32).
[0147] Additionally, the user computing device (110) may include a trigger event detector that provides an interface for interaction with an agent and initiates the operation of the agent. The trigger event detector can detect not only user input but also the arrival of a specific time, a change in the state of an external system, etc., and transmit a processing request to a server computing system (130).
[0148] Additionally, the model trainer (160) of the training computing system (150) may include a Supervision Signal Engine. The Supervision Signal Engine may compare the output generated by the agent (e.g., Raw Output) with a verified result (e.g., Grounded Output) obtained through an external tool (e.g., search engine, API) to calculate a difference value, and execute a reinforcement learning process to update the Reward Model or fine-tune the agent model based on this.
[0149] As such, in one embodiment, the system (1000) of the present disclosure may be implemented not as a simple set of software algorithms, but as a technical system in which specialized hardware accelerators, vectorized data storage, and physical engines controlling them are organically combined.
[0150] FIG. 3 illustrates an example of a block diagram of a computing device (100) implementing a text-to-speech conversion method based on a small number of sampling steps according to one embodiment of the present disclosure.
[0151] Including FIG. 3, the computing device (100) included in the user computing device (110), server computing system (130), and training computing system (150) may include a plurality of applications (e.g., Application 1 to Application N). Each application may include a machine learning library and one or more machine learning models. For example, the applications may include an image processing application (e.g., Detection, Classification and / or Segmentation, etc.), a text messaging application, an email application, a dictation application, a virtual keyboard application, a browser application and / or a chat-bot application, etc.
[0152] Additionally, for example, the computing device (100) may include a model that performs a text-to-speech conversion method based on a few sampling steps, and an application that provides related services to the user, i.e., a specialized application for text-to-speech conversion based on a few sampling steps.
[0153] In an embodiment, the computing device (100) may include a model trainer (160) for training an artificial intelligence model, and by storing and operating the trained artificial intelligence model, it may provide output data according to a predetermined input data (e.g., natural language query).
[0154] Each application of the computing device (100) can communicate with a number of other components of the computing device (100), such as, for example, at least one sensor, a context manager, a device state component, and / or additional components. In one embodiment, each application can communicate with each device component using an API (e.g., a public API). In one embodiment, the API used by each application may be specific to that application.
[0155] FIG. 4 is a block diagram illustrating the internal architecture and data processing pipeline of an artificial intelligence model according to one embodiment of the present disclosure.
[0156] Referring to FIG. 4, a computing device may have a pipeline structure that receives input data (202) and generates output data (212) through a series of conversion processes to perform a text-to-speech conversion method based on a few sampling steps. This process is performed through a preprocessing module (204), an encoder / embedding model (206), a neural network layer (208), and a decoder / generation head (210).
[0157] First, the preprocessing module (204) receives input data (202) (e.g., text prompt, image, or multimodal signal) from a user or system. The preprocessing module (204) performs tokenization and normalization on the input data to generate a sequence of tokens, which are the smallest units that the model can process.
[0158] Next, the encoder / embedding model (206) receives the generated token as input and converts it into a vector / embedding mapped to a number in a high-dimensional vector space. At this stage, the discrete information of the input data is converted into a continuous numeric matrix, making it a form that can be computed by the machine learning model.
[0159] Next, the neural network layer (208) receives the vector / embedding and performs deep computation. The neural network layer (208) may have a structure in which a plurality of sub-layers (e.g., Layer 1 to Layer N) are stacked. Each layer abstracts and refines input features through an attention mechanism or convolution operation, etc.
[0160] In particular, the final output of the neural network layer (208) is defined as a latent representation. This latent representation has a structure different from the original input data (202) and may correspond to an intermediate representation in which the semantic features of the data are highly compressed and abstracted. This implies that it is not a simple transmission of data, but a technical data structure that is valid only within the system.
[0161] Finally, the decoder / generation head (210) receives the potential representation as a conditioning input. The decoder / generation head (210) interprets the compressed potential representation and reconstructs or generates output data (212) in a form recognizable by the user (e.g., natural language text, image pixels, control codes, etc.).
[0162] This stepwise data transformation structure (token → vector → latent representation → output) can clearly demonstrate that it functions not as a simple sequence of operations, but as a concrete device that technically processes input data to generate useful information.
[0163] FIG. 5 illustrates an example of a block diagram in another aspect of a computing device (200) implementing a text-to-speech conversion method based on a few sampling steps according to one embodiment of the present disclosure.
[0164] Referring to FIG. 5, the computing device (200) includes a plurality of applications (e.g., Application 1 to Application N). Each application can communicate with a central intelligence layer. For example, applications may include an image processing application, a text messaging application, an email application, a dictation application, a virtual keyboard application and / or a browser application. In one embodiment, each application can communicate with the central intelligence layer (and a model stored therein) using an API (e.g., a common API across all applications).
[0165] Additionally, in one embodiment of the present disclosure, the application may include a text-to-speech application based on a few sampling steps, a speech recognition (STT) application, a multi-speaker real-time synthesis application, a conference assistance application and / or a logging and analysis application, etc.
[0166] The central intelligence layer may include a number of machine learning models. For example, as illustrated in FIG. 5, at least some of the machine learning models may be provided for each application and managed by the central intelligence layer. In another embodiment, two or more applications may share a single machine learning model. For example, in some embodiment, the central intelligence layer may provide a single model for all applications. In some embodiment, the central intelligence layer may be included within the operating system of the computing device (200) or otherwise implemented.
[0167] In one embodiment, the central intelligence layer may be integrated as part of the operating system or implemented as a separate logical layer, and may perform the role of transmitting input time series data to the corresponding model to return a prediction result.
[0168] The central intelligence layer can communicate with the central device data layer. The central device data layer may be a centralized data store for the computing device (200).
[0169] For example, the central device data layer can integrate and store sensor data, device status information, external environment information, etc., stored within the computing device (200), and provide this as input data required for a text-to-speech conversion service based on a small number of sampling steps. Each device component (e.g., sensor, state manager, etc.) can communicate with the corresponding data layer through a private API, etc.
[0170] As illustrated in FIG. 5, the central device data layer can communicate with a number of other components of the computing device (200), such as, for example, one or more sensors, a context manager, a device state component, and / or additional components. In some embodiments, the central device data layer can communicate with each device component using an API (e.g., a private API).
[0171] The technology described herein may refer to servers, databases, software applications, and other computer-based systems, as well as actions taken and information transmitted to or from said systems. It will be recognized that the inherent flexibility of computer-based systems allows for a wide range of possible configurations, combinations, division of tasks, and functionality between and from components. For example, the processes described herein may be implemented using a single device or component or multiple devices or components operating in combination. Databases and applications may be implemented in a single system or in a distributed system across multiple systems. Distributed components may operate sequentially or in parallel.
[0172] FIG. 6 is a block diagram illustrating the data flow and system interaction of a small number of sampling steps-based text-to-speech conversion service application process according to one embodiment of the present disclosure.
[0173] FIG. 6 is a block diagram illustrating the data flow and system interaction of a small number of sampling steps-based text-to-speech conversion service application process according to one embodiment of the present disclosure.
[0174] Referring to FIG. 6, the computing system may be configured as an organic data pipeline between a trigger event detector (310), an inference engine (330), and a mobile device screen (350).
[0175] First, the trigger event detector (310) is configured to monitor and detect a signal initiating the operation of the system. The trigger event detector (310) receives at least one of i) a time event indicating the arrival of a specific point in time, ii) a user action resulting from a user's physical input, or iii) a system state indicating a change in internal system data, and transmits an activation signal to the inference engine (330). This means that the service can be actively initiated depending on the situation without an explicit request from the user.
[0176] The reasoning engine (330) performs a multi-stage operation process that transforms raw data into a final result in response to the activation signal. This process is implemented as a series of logically connected prompt chains.
[0177] For example, the first prompt, contextualization (332), allows the inference engine (330) to receive unstructured raw data (e.g., user logs, channel metadata), analyze and refine it, and generate compressed summary information.
[0178] And the second prompt, Content Generation (334), can generate multiple candidate results that match the user's intent or situation using a generative AI model based on the summary information generated above.
[0179] In addition, the third prompt, Verification (336), performs filtering and verification by comparing the generated candidate results with a predefined policy (e.g., safety guidelines, format rules) to derive a reliable final result.
[0180] Finally, the final result generated by the inference engine (330) can be transmitted to and implemented on the mobile device screen (350) through an auto pre-filling (354) operation. Specifically, the system changes the state of the interface by directly writing the final result to a memory address of a specific target field (352) (e.g., text input field, setting value) within the mobile device screen (350). Here, the mobile device screen (350) may be an example of a user computing device (110).
[0181] Such a configuration can go beyond the simple display of information and provide specific technical means for data generated by an external trigger to physically control and complete the input interface of the user terminal.
[0182] FIG. 7 illustrates the architecture of a multi-agent system (Universal Dynamic Multi-Agent System, 400) according to one embodiment of the present disclosure.
[0183] Referring to FIG. 7, the system (400) may be configured around an orchestration engine (410) that determines and executes an optimal agent collaboration structure in real time according to the nature of the user's request or task. The system (400) may operate in an organically combined manner, including a complexity analyzer (405), an agent pool (420), a shared memory fabric (430), and a tool execution interface (440).
[0184] 1. Input Analysis and Dynamic Topology
[0185] The complexity analyzer (405) evaluates the complexity of the user query and the required domain expertise. Based on this evaluation result, the orchestration engine (410) dynamically configures an optimal agent collaboration topology for task resolution.
[0186] For example, in the case of a simple question, the engine (410) activates a single agent mode.
[0187] If a complex plan is required, the engine (410) configures a Supervisor pattern and instantiates one supervisor agent to control subordinate agents.
[0188] When high accuracy is required, the engine (410) configures a debate pattern and sets a control path so that multiple agents perform mutual criticism.
[0189] 2. Agent Pool and Instantiation (Agent Pool)
[0190] The agent pool (420) is a repository of template agents that have prompts and tool sets specialized for specific functions (e.g., web search, code generation, data analysis). The orchestration engine (410) can select and activate the necessary agents at runtime according to the determined topology.
[0191] 3. Shared Memory Fabric
[0192] The shared memory fabric (430) is a data pipeline that synchronizes the state and context between multiple collaborating agents. By mediating the short-term memory of individual agents and the knowledge base of the entire system, it can ensure that the output of Agent A is transferred to the input of Agent B without loss.
[0193] 4. Tool Execution & Feedback Loop
[0194] Each agent communicates with an external API (search engine, calculator, AWS, etc.) through a tool execution interface (440). The execution result generated at this time is fed back to the orchestration engine (410), and the engine can determine the consistency of the result and perform self-correction logic to instruct the agent to rework or proceed to the next step.
[0195] This structure enables the implementation of an adaptive artificial intelligence system that flexibly changes the system's processing structure according to the nature of the input problem, rather than a fixed (static) algorithm.
[0196] Meanwhile, the orchestration engine (410) analyzes the characteristics of the task received from the complexity analyzer (405) (e.g., creativity, logic, whether coding is required) and selects specialized agents waiting in the agent pool (420) to form a dynamic collaboration topology. The topology can be reconfigured into various operation modes as follows depending on the type of task.
[0197] First Operation Mode: Sequential Pattern
[0198] For tasks such as a single-flow creation or report writing, the orchestration engine (410) connects multiple agents in series. For example, it forms a pipeline in the order of user agent -> writer agent -> style agent to control the output of the previous stage to be passed as input to the next stage.
[0199] Second Operation Mode: Supervisor Pattern
[0200] In cases where complex sub-tasks are mixed, the engine (410) forms a centralized star topology in which one agent is designated as a supervisor and the remaining agents (e.g., research agent, math agent) are assigned as workers. The supervisor agent distributes tasks to sub-agents and aggregates the results.
[0201] Third Operation Mode: Hierarchical Pattern
[0202] In cases of high complexity, such as large-scale project management, the engine (410) establishes a command system by forming a tree structure topology in which a meta-agent is placed as the upper manager and multiple specialized agent groups are placed below it.
[0203] 4th Operation Mode: Debate Pattern
[0204] In cases where the correct answer is unclear or high reliability is required (e.g., social issue analysis), the engine (410) constructs a competitive topology in which multiple agents present different perspectives on the same topic and perform mutual criticism and voting to derive the optimal conclusion.
[0205] Fifth Operation Mode: Mixture-of-Agents Pattern
[0206] When parallel processing is required, the engine (410) forms a parallel processing structure that places multiple agents in multiple layers to perform tasks simultaneously and finally integrates the results through an aggregator agent.
[0207] Additionally, the agent pool (420) includes various template agents that can be deployed into the topology. Each agent is instantiated to perform the following core methodologies under the control of the orchestration engine (410).
[0208] For example, at least one agent can execute a loop that repeats reasoning and action to perform reasoning and action (ReAct Paradigm) that refines answers based on results from external tools (e.g., Google Search).
[0209] In addition, at least one agent can perform code-based behavior (CodeAct Paradigm) that involves complex calculations or data analysis by generating and executing executable code (e.g., Python) instead of natural language.
[0210] In addition, at least one agent can perform self-reflection to improve the quality of the output by carrying out a metacognitive process of self-evaluating (Critique) and modifying the generated output.
[0211] Furthermore, the system provides a shared memory fabric (430) to enable individual agents to cooperate organically without disconnection. This prevents context loss by managing the conversation history (Short-Term Memory) between agents and an external knowledge base in an integrated manner. Additionally, the tool execution interface (440) securely connects the agents with external APIs, cloud services, and databases through protocols such as Multi-Channel Processing (MCP).
[0212] In this way, the system (400) of the present disclosure implements an adaptive artificial intelligence platform that is not a fixed single model, but rather dynamically changes the system's processing structure and behavior according to the nature of the input problem.
[0213] Specifically, one embodiment of the present disclosure may be implemented as an Agentic RAG (Retrieval-Augmented Generation) system that supports advanced question-answering by including agent functions. This system is a search-based generation system that retrieves information from external data sources and generates answers based thereon.
[0214] The system's data processing pipeline can consist of a data extraction phase and an Agentic RAG pipeline phase. In the data extraction phase, content in various formats, such as text and images, is collected from designated websites. Subsequently, text and metadata are extracted from the collected content, the text is chunked into small units, and each chunk is vectorized using an embedding model and stored in a vector database.
[0215] The Agentic RAG pipeline can be divided into search and generation phases. In the search phase, when a user query is input, query rewriting and embedding are performed, and highly relevant documents are identified through similarity searches in a vector database. Subsequently, the retrieved results can be ranked based on relevance to construct context. In the generation phase, the user query and the retrieved context are combined to construct a prompt, which is then input into a large language model to generate the final response. During this process, the agent can respond to complex queries by utilizing agentic elements such as memory storage, calling external tools, and planning.
[0216] An AI agent according to one embodiment of the present disclosure can support complex decision-making and task execution through a hierarchical memory structure similar to human memory. The agent's memory can be broadly divided into short-term memory and long-term memory.
[0217] Short-term memory is a temporary memory space focused on the currently ongoing workflow and may include working memory, which manages workflow-specific reasoning and task flows, and cache memory, which provides quick access to frequently used data or results.
[0218] Long-term memory is a memory space based on knowledge and experience that is continuously preserved, and may include episodic memory, which records events or incidents manually saved in a specific workflow; semantic memory, which stores conceptual or factual knowledge; and procedural memory, which stores knowledge of how to perform specific tasks or procedural knowledge.
[0219] This memory structure can be integrated with the language model framework through a central memory controller. Additionally, it can leverage external knowledge or support real-time integration by connecting with external vector databases, semantic databases, or third-party APIs via the MCP server. Through this, the agent can generate user-customized responses that comprehensively reflect past experiences and current context.
[0220] One embodiment of the present disclosure may include various agentic workflows to solve complex problems. These workflows may be designed to suit specific business purposes.
[0221] For example, the system may include a Plan and Execute workflow. In this workflow, a Planner breaks down a single top-level task into multiple sub-tasks, specialized agents process each sub-task, and then integrates the results. This can be utilized for business process automation or data pipeline orchestration.
[0222] As another example, the system may include an Orchestrator-Worker workflow. In this structure, a central orchestrator language model breaks down tasks, distributes them to multiple worker language models for processing, and then integrates the results. This can be used in the implementation of Agentic RAGs or coding agents.
[0223] As another example, the system may include a routing workflow. This workflow is structured to analyze input tasks, classify them into the most suitable ones among several predefined subtasks, and forward them to a specialized language model or path capable of handling the task. This can be applied to customer support agents or multi-agent discussion systems.
[0224] One embodiment of the present disclosure may include a protocol for efficient and secure communication between a plurality of agents or between an agent and an external tool.
[0225] In one embodiment, an Agent2Agent (A2A) protocol may be used for communication between agents. The A2A protocol can enhance security by enabling each agent to communicate without directly sharing their internal data. Through this protocol, multiple agents can share tasks and negotiate, and each agent can operate independently using its own language model, framework, and database.
[0226] In another embodiment, the Multi-Channel Processing (MCP) protocol may be used for communication between an agent and an external function server. MCP has a structure that separates each external function, such as file access, search, and cloud API calls, into a separate server for communication. For example, one agent may communicate with a local file system or a search engine via the MCP protocol, while another agent may communicate with a cloud provider such as AWS or a communication tool such as Slack via the same protocol.
[0227]
[0228] Hereinafter, a text-to-speech conversion method in which a computing system (1000) according to the present disclosure extracts speaker characteristics from a speaker's voice sample using an artificial intelligence model and generates high-quality synthetic speech in real time with only a few sampling steps will be described in detail with reference to FIGS. 8 to 18.
[0229] In one embodiment, the text-to-speech conversion method may be performed by a processor (131) included in a server computing system (130). However, it is not limited thereto, and at least a part of the text-to-speech conversion method may be performed by a processor (111) of a user computing device (110) or a processor (151) of a training computing system (150), and another part may be performed by a processor (131) included in a server computing system (130).
[0230] First, text can be received through the user computing device (110) of the computing system (1000) (S101).
[0231] For example, a user computing device (110) can receive text from a user. In one embodiment, the user computing device (110) can receive text directly entered by the user through a text input window included in the user interface. In this case, the user computing device (110) can automatically detect the language of the entered text and select an artificial intelligence model (30) corresponding to the language, or perform text preprocessing to normalize special characters, numbers, abbreviations, etc. into a form suitable for speech synthesis.
[0232] In another embodiment, the user computing device (110) can extract text from a document file uploaded by the user. For example, when a document file such as .txt, .docx, or .pdf is input, the server computing system (130) can extract body text using a parser corresponding to the format and perform preprocessing to filter out elements unnecessary for speech synthesis, such as headers, footers, tables, and comments.
[0233] In another embodiment, the user computing device (110) may receive an image file containing text and extract the text. In this case, the server computing system (130) may use an Optical Character Recognition (OCR) model to detect text regions within the image and perform preprocessing to recognize characters. In one embodiment, the server computing system (130) may further perform preprocessing, such as tilt correction, noise removal, and resolution enhancement of the image, before performing OCR to improve the accuracy of character recognition.
[0234] In another embodiment, the user computing device (110) may receive text from the user's voice input. In this case, the user computing device (110) may input the user's utterance received through the microphone into a speech-to-text (STT) model to convert it into text, and utilize the converted text as input for text to be synthesized. This embodiment can be usefully utilized in scenarios where the voice spoken by the user is converted into the voice of another speaker, or where content spoken in a first language is output in real time as synthesized speech in a second language.
[0235] In another embodiment, the user computing device (110) may receive text from an external system via a network. For example, text generated or processed by an external component, such as a conversational AI agent, a translation system, or a document management system, may be received via an Application Programming Interface (API) and used as input for a speech synthesis task. This embodiment may be utilized to automatically generate synthesized speech for a large amount of text generated by an external system by linking with a synthetic speech-based data augmentation pipeline for training a speech recognition model according to one embodiment of the present disclosure.
[0236] In another embodiment, the user computing device (110) can receive text through a clipboard or a share sheet of the operating system. For example, the user can paste text copied from another application via the clipboard or receive text from an external application using the share sheet of the mobile operating system. This embodiment enhances user convenience in that text already written or viewed by the user can be immediately utilized as input for speech synthesis without a separate text input process.
[0237] In another embodiment, the user computing device (110) can extract text from a URL entered by the user. In this case, the server computing system (130) can crawl the webpage corresponding to the entered URL to extract the body text and perform preprocessing to filter out elements unnecessary for speech synthesis, such as advertisements, menus, and navigation elements. For example, text content on the web, such as news articles or blog posts, can be immediately converted into speech simply by entering a URL, which can be usefully utilized in terms of accessibility, such as reducing visual fatigue and supporting multitasking.
[0238] As such, the user computing device (110) can receive text in various ways, such as direct input by the user, uploading document files, image-based OCR extraction, speech recognition conversion, integration with external systems, clipboard sharing, and URL-based web crawling. The text-to-speech conversion method according to the present disclosure is not limited to the receiving methods described above, and any method by which the user computing device (110) can acquire text may be applied. The server computing system (130) processes the received text as input for a speech synthesis task regardless of the input method of the text, and can transmit it to the artificial intelligence model (30) after performing preprocessing corresponding to each input method.
[0239] Next, the computing system (1000) can receive a voice sample containing the voice characteristics of the speaker (S103).
[0240] In an embodiment, the user computing device (110) can receive voice samples from the user. In one embodiment, the user computing device (110) can receive voice samples by recording the user's speech in real time through a microphone. In this case, the user computing device (110) can convert the recorded voice samples into a form suitable for extracting speaker characteristics by performing preprocessing such as sampling rate normalization, removal of silent intervals, and removal of background noise. In one embodiment, if the length of the received voice sample does not meet a preset minimum length, the server computing system (130) can provide a notification to the user requesting additional recording to prevent degradation of the quality of speaker characteristic extraction.
[0241] In another embodiment, the user computing device (110) may receive an audio file uploaded by the user as a voice sample. For example, when a file of various audio formats such as .wav, .mp3, .flac is input, the server computing system (130) may extract audio data using a decoder corresponding to the format and perform preprocessing to normalize the sampling rate and the number of channels to match the input specifications of the artificial intelligence model (30).
[0242] In another embodiment, the user computing device (110) can receive an audio track separated from a video file uploaded by the user as a voice sample. For example, when a video file such as .mp4, .avi, or .mov is input, the server computing system (130) can extract an audio track from the video and perform source separation preprocessing to separate sound elements other than the speaker's voice, such as background music or sound effects. Through this, the user can utilize the speaker's voice included in the video as a voice sample without a separate audio file.
[0243] In another embodiment, the user computing device (110) may receive voice samples from a pre-stored speaker profile. In this case, the server computing system (130) may directly load the speaker embedding vector included in the pre-registered speaker profile from memory and use it as a conditional input for the artificial intelligence model (30). In this embodiment, the encoding process of the voice samples may be omitted, so processing delay can be minimized when synthesizing the voice of the same speaker repeatedly.
[0244] In another embodiment, the user computing device (110) may receive voice samples from an external system via a network. For example, voice samples may be received via an API from a speaker database, a voice storage system, or an external recording device and used as input for a voice synthesis task. This embodiment can be usefully utilized in a scenario where voice samples are received in real time from a separate sound acquisition device, such as a microphone array installed in a conference room, and automatically linked to a synthesized voice-based data augmentation pipeline.
[0245] As such, the user computing device (110) can receive voice samples in various ways, such as real-time microphone recording, audio file uploading, audio track extraction within a video file, loading a pre-stored speaker profile, and linking with an external system. The text-to-speech conversion method according to the present disclosure is not limited to the receiving methods described above, and any method by which the user computing device (110) can acquire voice samples may be applied. The server computing system (130) processes the received voice samples as input for a speech synthesis task regardless of the input method of the received voice samples, and can transmit them to an artificial intelligence model (30) after performing preprocessing corresponding to each input method.
[0246] Next, the server computing system (130) can generate a synthetic voice corresponding to the text that reflects the speaker characteristics of the voice sample using at least one artificial intelligence model (30) (S105).
[0247] Here, at least one artificial intelligence model (30) can be further trained through adversarial learning with a discriminator (34) on the output generated by inference executed with fewer sampling steps than during prior training. A detailed explanation of this will be provided later with reference to FIGS. 11 to 13.
[0248] Referring to FIGS. 9 and FIGS. 10, step (S105) may include a step of encoding a voice sample into a latent expression (S1051), a step of generating a synthetic latent expression based on a flow matching method with a prompt extracted from the latent expression as a condition (S1053), and a step of decoding the synthetic latent expression into synthetic speech (S1055).
[0249] First, the server computing system (130) can encode a voice sample received through the voice sample upload unit (20) into a potential representation through the VAE encoder (31) (S1051).
[0250] Specifically, the VAE encoder (31) can receive a Mel-spectrogram converted from the waveform of a speech sample and generate a latent representation (z₁) in a latent space. In this process, the VAE encoder (31) can compressively store key speech characteristics, such as the speaker's timbre, intonation, and speech rhythm, embedded in the speech signal in a low-dimensional latent representation.
[0251] Meanwhile, the text received through the text input unit (10) passes through a Style Encoder and a Text Encoder to obtain text condition embeddings (c textIt can be converted into ). Then, a duration predictor predicts the utterance length for each phoneme of the text, and a context embedding (c) expanded into a phoneme-unit sequence can be generated through an aligner and a length regulator.
[0252] Next, the server computing system (130) can generate a synthetic latent expression based on a flow matching method using a prompt extracted from the latent expression as a condition (S1053).
[0253] Specifically, the flow matching decoder (32) applies masking to the encoded latent expression (z₁) and constructs the prompt portion (p=(1-m)⊙z₁) and the zero padding corresponding to the synthesis target section by concatenating them. Here, the masking variable (m) is a binary mask that specifies the synthesis target section, where the prompt portion is set to m=0 and the synthesis target section is set to m=1. That is, the prompt (p) is the unmasked portion of the latent expression corresponding to the speech sample and contains the speech characteristics of the speaker, and the synthesis target section is initialized with zero padding and then filled through infilling.
[0254] Next, the flow matching decoder (32) can generate a synthetic latent representation by performing infilling using the Optimal Transport (OT) based Conditional Flow Matching (OT-CFM) method with the prompt (p) and context embedding (c) as conditional inputs.
[0255] The OT-CFM method approximates the path from the Gaussian noise distribution to the target latent representation distribution as a straight line based on optimal transport theory, thereby enabling the generation of a synthetic latent representation close to the target distribution with only a few steps, without tens of iteration steps. In one embodiment, the flow matching decoder (32) can perform 4-step inference by applying fixed time steps [0, 0.08, 0.29, 0.62] of the Euler integration method during inference.
[0256] Through this, the flow matching decoder (32) can sufficiently reflect condition information consisting of a prompt (p) and a context embedding (c) in the vector field estimation during the initial integration step starting from Gaussian noise by setting the initial time step interval (0 → 0.08) densely. That is, since the path to converge to the target latent expression distribution can be stably estimated through subsequent steps (0.08 → 0.29 → 0.62) while the initial conditions are well reflected, a synthetic latent expression in which both speaker characteristics and voice quality are preserved can be generated despite the small number of steps.
[0257] Next, the server computing system (130) can decode the synthetic potential expression into synthetic speech (S1055).
[0258] Specifically, the VAE decoder (33) receives the synthetic latent expression generated by the flow matching decoder (32), restores the Mel spectrogram, and converts it back into a speech waveform to output the final synthesized speech. At this time, the section corresponding to the prompt (p) is discarded, and only the waveform corresponding to the synthesis target section is transmitted to the subsequent processing component through the synthesized speech output unit (50).
[0259] Hereinafter, with reference to FIGS. 11 to 13, the learning structure and inference process of the artificial intelligence model (30) will be described in detail. The artificial intelligence model (30) acquires the ability to generate high-quality synthetic speech with only a few sampling steps through a three-stage learning process, and during inference after the completion of learning, it utilizes the VAE encoder (31), flow matching decoder (32), and VAE decoder (33) without a discriminator (34) to generate synthetic speech reflecting the speaker's voice characteristics in real time using only the received voice samples and text.
[0260]
[0261] Step 1: VAE Pre-training
[0262] Referring to FIG. 11, in the first step, the VAE encoder (31) and VAE decoder (33) are pre-trained using speech waveforms and Mel spectrograms.
[0263] Here, a Mel-spectrogram is a two-dimensional representation of a speech waveform transformed into the time-frequency domain, created by applying the Mel scale, which mimics human auditory characteristics, to transform the frequency axis. Specifically, it is generated by applying the Short-Time Fourier Transform (STFT) to the speech waveform to create a spectrogram, and then converting it to the Mel scale through a Mel filterbank. Mel-spectrograms are widely used as input for speech synthesis models because they can compressively represent the acoustic characteristics of a speech signal in a form that is easy for artificial intelligence models to process.
[0264] For example, the VAE encoder (31) receives the Mel spectrogram of a voice sample as input and the latent representation (VAE Latent ) is generated, and the VAE decoder (33) is trained to reconstruct the original speech waveform from the potential representation.
[0265] In this process, the pitch extractor extracts the fundamental frequency (F0) and voiced / unvoiced flags from the original speech waveform and inputs them together to the VAE decoder (33). Through this, the VAE decoder (33) is trained to accurately reproduce even fine pitch characteristics that are difficult to restore using only latent representations. The loss function of the VAE pre-training can be composed of the sum of the reconstruction loss and the KL divergence loss. Once the first stage of training is completed, all parameters of the VAE encoder (31) and VAE decoder (33) are frozen and are not updated in subsequent training stages.
[0266]
[0267] Phase 2: Flow Matching Learning
[0268] Referring to FIG. 11, in the second step, the flow matching decoder (32) can be trained using a latent representation (z₁) generated by a fixed VAE encoder (31). Specifically, a masking variable (m) is applied to the latent representation (z₁) to form a training target (m⊙z₁) and a prompt (Prompt p: (1-m)⊙z₁), and the flow matching decoder (32) can be trained to infill the masked segments based on the prompt (p) and context embedding (c).
[0269] At this time, in the second stage, the flow matching decoder (32) can be freely trained to minimize flow matching loss by using sufficient sampling steps without limiting the number of steps. That is, in this stage, the number of inference steps is not fixed, and the flow matching decoder (32) focuses on accurately modeling the target latent representation distribution, and optimization for a small number of steps of inference is performed separately in the third stage described later.
[0270] For example, assuming the text is "Hello, Good luck!" and the voice sample is an utterance corresponding to "Hello," the VAE encoder (31) generates a latent expression (z₁) corresponding to the entire "Hello, Good luck!", and the masking variable (m) sets the prompt interval corresponding to "Hello," to m=0 and the synthesis target interval corresponding to "Good luck!" to m=1. Accordingly, the prompt (p=(1-m)⊙z₁) is in a state where the latent expression of "Hello," is maintained, and the learning target (m⊙z₁) is a latent expression corresponding to "Good luck!".
[0271] A flow matching decoder (32) can be trained to receive a prompt (p) which is a potential representation of "Hello," and a context embedding (c) containing text condition information of "Good luck!" as inputs, and to infill a "Good luck!" section initialized with zero padding to generate a potential representation of that section.
[0272] That is, the flow matching decoder (32) can generate a synthetic voice such as "Good luck!" spoken by the same speaker by using a prompt (p) containing the speaker's vocal characteristics as a clue to generate a latent expression that naturally follows the text to be synthesized.
[0273] In another example, even when the voice sample is an utterance corresponding to the Korean "Hello" and the target text for synthesis is the English "Good luck!", the flow matching decoder (32) can generate English synthesized speech that reflects the phonetic characteristics of a Korean speaker by infilling a potential expression corresponding to the English text based on a prompt (p) extracted from the potential expression of the Korean voice sample. In this way, the infilling method can be usefully utilized in multilingual synthesis scenarios, as it enables the generation of synthesized speech that maintains speaker characteristics not only when the prompt text and the target text for synthesis are of the same language but also when they are of different languages.
[0274] Meanwhile, the server computing system (130) inputs the Mel spectrogram into a Style Encoder to extract speaker style information, and inputs the extracted speaker style information and text (x) together into a Text Encoder to obtain text condition embeddings (c text Can generate ).
[0275] The server computing system (130) generates text condition embeddings (c text ) can be passed to a Duration Predictor and an Aligner to determine the utterance length (Duration, D) of each phoneme, and a context embedding (c) expanded on a frame-by-frame basis according to the utterance length can be generated through a Length Regulator.
[0276] For example, if the text (x) is "How are you?", the Duration Predictor predicts the utterance length (D) for each phoneme, such as "H", "ow", "a", "r", "e", "y", "ou", etc. At this time, the Aligner aligns each phoneme embedding by repeating it for the corresponding length based on the predicted utterance length to form a frame-unit sequence, and the Length Regulator can generate a context embedding (c) by extending it to a time resolution suitable for processing by the flow matching decoder (32). By explicitly predicting and controlling the utterance length at the phoneme level in this way, the utterance speed and rhythm of the synthesized speech can be naturally controlled.
[0277] Here, the Duration Predictor is a neural network module composed of multiple 1D Convolution layers, a ReLU activation function, and Layer Normalization, which receives each phoneme embedding as input and can predict the utterance length of the corresponding phoneme as a scalar value. In one embodiment, the Duration Predictor is trained to minimize the Duration Loss by using the actual utterance length extracted from the Aligner as the correct answer during training, and can automatically predict the utterance length of each phoneme using only text during inference.
[0278] The Aligner is a neural network module that learns the alignment relationship between actual speech and text during training, and can estimate which frame segment of speech each phoneme corresponds to based on monotonic alignment. The alignment information estimated by the Aligner is converted into the actual utterance length (D) of each phoneme and used as a training target for the Duration Predictor. During inference, the utterance length can be determined solely by the prediction value of the Duration Predictor without the Aligner.
[0279] The Length Regulator is a deterministic module that does not have separate learnable parameters and performs the role of expanding each phoneme embedding into a frame-unit sequence by repeatedly duplicating each phoneme embedding a corresponding number of times according to the utterance length (D) predicted by the Duration Predictor. Through this, the phoneme-unit text embedding can be converted into a context embedding (c) having a temporal resolution suitable for processing by the flow matching decoder (32).
[0280] In addition, the Pitch Predictor can predict the pitch of the target synthesis section by receiving the lower 20 bins of the Mel spectrogram as input. Here, a bin refers to an individual frequency band unit that constitutes the frequency axis of the Mel spectrogram, and is obtained by dividing the entire frequency range into a certain number of sections through the Mel filterbank.
[0281] In one embodiment, when the entire Mel spectrogram consists of 80 bins, the lower 20 bins correspond to the low-frequency band and contain information closely related to pitch characteristics, such as the fundamental frequency (F0) and formant of the speech signal. By selectively receiving only the lower 20 bins, where pitch-related information is concentrated, instead of the entire Mel spectrogram, the pitch predictor can reduce learning noise caused by unnecessary high-frequency information and increase the accuracy of pitch prediction.
[0282] At this time, the lower 20 bins of the input Mel spectrogram are in a form with a masking variable (m) applied, and the prompt intervals are maintained as original values while the synthesis target intervals are masked and can be input into the pitch predictor.
[0283] The pitch predictor predicts the pitch of the masked target segment by referring to the pitch pattern of the prompt segment, and during training, it can be trained in a direction that minimizes pitch loss by using the actual fundamental frequency (F0) and voiced / unvoiced flags extracted from the original speech waveform by the pitch extractor in the first stage VAE pre-training as the correct answer.
[0284] The loss function of the second stage may consist of the sum of the flow matching loss (FM Loss), duration loss, and pitch loss. The flow matching loss (FM Loss) is a loss that minimizes the difference between the vector field predicted by the flow matching decoder (32) and the actual latent representation (z₁); the duration loss is a loss that minimizes the difference between the utterance length predicted by the duration predictor and the actual utterance length estimated by the aligner; and the pitch loss is a loss that minimizes the difference between the pitch predicted by the pitch predictor and the actual pitch extracted by the pitch extractor.
[0285] When the second stage of learning is completed, the flow matching decoder (32) has the ability to generate high-quality potential representations under sufficient sampling steps. However, since the number of inference steps must be significantly reduced for real-time speech synthesis, in the third stage described later, the inference steps of the flow matching decoder (32) are fixed at 4 steps, and adversarial learning with the discriminator (34) is performed to compensate for the degradation of voice quality due to the reduction in the number of steps.
[0286]
[0287] Phase 3: Adversarial Post-Training (AP)
[0288] Referring to FIGS. 11 and 13, in the third step, the flow matching decoder (32) can perform adversarial learning with the discriminator (34) while fixing the inference step to 4 steps. This is a key learning step to compensate for the degradation of voice quality caused by a small number of sampling steps. GAN-based speech synthesis models had a problem of reduced speaker similarity using a discriminator that optimizes only voice quality, but the discriminator (34) according to one embodiment of the present disclosure can solve this problem by adopting a combined conditional-unconditional structure that evaluates voice quality and speaker similarity simultaneously.
[0289] Referring to FIG. 13, the discriminator (34) may be configured as a Joint Conditional-Unconditional (JCU) discriminator. The JCU discriminator (34) has a shared encoder path that receives the entire latent representation (Latent z₁) as input and processes it through a plurality of convolutional (Conv) layers, from which it branches into two independent output paths.
[0290] The unconditional output path processes the features of the shared encoder through an additional convolutional layer to determine whether the sound quality of the synthesized speech itself is indistinguishable from the actual speech, and evaluates only acoustic naturalness regardless of speaker information. For example, if artifacts or unnatural acoustic patterns occur in the synthesized speech due to fractional sampling of 4 steps, the unconditional output path distinguishes them from the actual speech and induces the flow matching decoder (32) to generate a synthetic potential expression of more natural sound quality.
[0291] The Conditional Output path generates a speaker condition vector through an Avg. Pooling and Projection layer for the prompt (p: (1-m)⊙z₁), combines it with the intermediate features of the shared encoder, and processes it through an additional convolution layer to determine how similar the synthesized speech is to the original speaker's speech sample. For example, if the synthesized latent expression generated by the flow matching decoder (32) differs from the original speaker's timbre or intonation, even if the sound quality is natural, the Conditional Output path detects this and induces the flow matching decoder (32) to reproduce the speaker characteristics more precisely.
[0292] In this way, the JCU discriminator (34) simultaneously produces an unconditional output and a conditional output, thereby enabling adversarial learning that optimizes both voice quality and speaker similarity simultaneously using only a single discriminator. This has higher learning efficiency compared to the method of operating separate discriminators for voice quality optimization and speaker similarity optimization, and the two goals work complementarily to generate a synthesized speech that satisfies both voice quality and speaker similarity even in a small number of sampling steps.
[0293] Loss function of adversarial post-learning (L AP ) is selectively applied only to the synthesis target section through the masking variable (m) as in Equation (1) below, and the reconstruction loss (L recon ), JCU discriminator loss(L JCU ) and feature matching loss (λ fm ·L fm It can be composed of a weighted sum of ).
[0294]
[0295] Equation (1):
[0296]
[0297] Reconstruction loss (L recon) minimizes the difference between the output of the flow matching decoder (32) and the actual potential representation (z₁), and the JCU discriminator loss (L JCU ) is an adversarial loss that induces the flow matching decoder (32) to deceive the discriminator (34), and a feature matching loss (L fm ) minimizes the difference between the intermediate layer features of the discriminator (34) and the intermediate layer features of the generated output, thereby inducing a match even with fine characteristics of the sound quality. λ fm is a weight hyperparameter that controls the contribution of the feature matching loss.
[0298] Here, the 4 steps refer to the number of sampling steps performed by the flow matching decoder (32) to move from Gaussian noise to the target latent representation during the inference process of generating synthetic speech. That is, adversarial post-learning is a process of repeatedly performing adversarial learning with the discriminator (34) on the output generated by the inference method fixed to 4 steps in this way, and in one embodiment, the repeated learning is approximately 200,000 steps, batch size 5, and learning rate 1× It can be performed under the condition of.
[0299]
[0300] reasoning process
[0301] Referring to FIG. 12, during inference, a speech sample (Speech Prompt yp) is simultaneously input to a VAE encoder (31) and a style encoder. The VAE encoder (31) generates a latent prompt (Latent Prompt zp) from the speech sample, and the style encoder extracts speaker style information from the speech sample and transmits it to a text encoder.
[0302] The Text Encoder receives the entire text (xp, xt) formed by concatenating the prompt text (xp) and the target text for synthesis, and the text condition embedding (c textGenerates ). The Duration Predictor generates text condition embeddings (c text Based on ), the utterance length (D) of each phoneme is predicted, and a frame-by-frame expanded context embedding (c) is generated through an aligner and a length regulator.
[0303] Zero-padded text corresponding to the length of the text to be synthesized is concatenated to the potential prompt (zp) and provided as input to the flow matching decoder (32). Additionally, the pitch predictor receives the zero-padded Mel spectrogram as input and predicts the pitch of the section to be synthesized; among the predicted pitches, only the part corresponding to the section to be synthesized is transmitted to the VAE decoder (33), and the part corresponding to the prompt section is discarded.
[0304] The flow matching decoder (32) generates a synthetic latent expression by performing 4-step inference with fixed time steps [0, 0.08, 0.29, 0.62], and among the generated synthetic latent expressions, the part corresponding to the prompt section is discarded, and only the section to be synthesized is transmitted to the VAE decoder (33). Finally, the VAE decoder (33) receives the synthetic latent expression and pitch information and outputs a synthetic speech waveform.
[0305] The synthesized voice waveform generated in this way is transmitted to at least one subsequent processing component through the synthesized voice output unit (50) and can be utilized in various embodiments.
[0306] For example, the generated synthesized speech can be played back in real time through a user interface or provided as a visualization along with quality indicators such as speaker similarity and sound quality.
[0307] In addition, a system according to one embodiment of the present disclosure can be widely applied to a multilingual synthesis scenario that generates synthesized speech of a second language from a speech sample of a first language, a multi-speaker synthesis scenario that switches speakers in real time using a plurality of speaker profiles, a data augmentation scenario that generates a large amount of synthesized speech based on speech samples collected in a specific acoustic environment and utilizes it as training data for a speech recognition model, or an automatic dubbing scenario that generates speech optimized for lip movements by precisely tracking the speech timing and pitch of the original speaker in a video. Hereinafter, each embodiment will be described in detail with reference to FIGS. 14 to 19.
[0308] FIG. 14 illustrates an exemplary user interface for displaying speaker similarity, sound quality, and synthesis speed of a synthesized voice according to one embodiment of the present disclosure.
[0309] Referring to FIG. 14, the user interface may display a voice synthesis monitoring dashboard (500) including a speaker similarity display unit (510), a sound quality display unit (520), a synthesis speed display unit (530), a synthesized voice waveform output unit (540), and a synthesis history unit (550).
[0310] The speaker similarity display unit (510) is a component that displays the similarity between the synthesized voice and the original speaker voice sample as numerical and visual indicators. In one embodiment, the speaker similarity is calculated and displayed as a Speaker Mean Opinion Score (SMOS), and can also display whether there is an increase or decrease compared to a reference model. For example, as shown in FIG. 14, the speaker similarity display unit (510) can display SMOS 3.65 / 4.0 and also display that it has improved by +0.05 compared to the reference model.
[0311] The sound quality display unit (520) is a component that displays the naturalness and acoustic quality of the synthesized voice as numerical and visual indicators. In one embodiment, the sound quality is calculated and displayed as a Mean Opinion Score (MOS), and can also display whether there is an increase or decrease compared to a reference model. For example, as shown in FIG. 14, the sound quality display unit (520) can display an MOS of 3.58 / 5.0 and also display that it has improved by +0.12 compared to the reference model.
[0312] The synthesis speed display unit (530) is a component that displays the generation speed of the synthesized speech as a Real-Time Factor (RTF). RTF is a value obtained by dividing the time taken for synthesis by the playback time of the generated speech. An RTF value of less than 1 indicates that real-time synthesis is possible, and a lower value indicates that faster synthesis is possible. For example, as shown in FIG. 14, the synthesis speed display unit (530) can display an RTF of 0.052 and also indicate that real-time synthesis is possible with 4-step inference.
[0313] The synthesized voice waveform output unit (540) is a component that visualizes and displays the waveform of the generated synthesized voice in real time and also displays numerical values for each detailed item of speaker similarity. In one embodiment, the synthesized voice waveform output unit (540) can visually indicate the real-time synthesis progress by displaying the waveform of the section where synthesis is completed as a solid line and the section where synthesis is in progress as a dotted line. In addition, the synthesized voice waveform output unit (540) can display the similarity for each of the detailed items of speaker similarity—timbre, intonation, and speech speed—as a bar graph in the form of a percentage (%). For example, as shown in FIG. 14, by displaying the speaker similarity for each item in a subdivided manner, such as timbre 88%, intonation 82%, and speech speed 91%, the user can evaluate the quality of speaker characteristic reproduction of the synthesized voice from various angles.
[0314] The synthesis history section (550) is a component that displays the history of past synthesis requests along with time, speaker, SMOS, MOS, RTF, and status information. In one embodiment, the synthesis history section (550) displays synthesis results for multiple speakers in a time series, allowing the user to check and compare trends in synthesis quality by speaker. For example, as shown in FIG. 14, the synthesis history section (550) displays the synthesis time, SMOS, MOS, RTF, and completion status for Speaker A, Speaker B, and Speaker C, respectively, in a table format, allowing the user to compare the synthesis quality for each speaker at a glance.
[0315] In another embodiment, the user interface (500) may additionally provide a function to save the synthesized voice file to a local storage or transmit it to an external system. For example, the user may select a desired item from the past synthesis results displayed in the synthesis history section (550) to download the corresponding synthesized voice file or transmit it to an external voice recognition system or data storage via an API.
[0316] In another embodiment, the user interface (500) may additionally include a parameter setting interface that allows the user to directly adjust the trade-off between speaker similarity, sound quality, and synthesis speed. For example, the user may directly set parameters through the interface, such as reducing the number of sampling steps to increase synthesis speed or adjusting the weight of speaker embeddings to increase speaker similarity, and changes in quality indicators resulting from the setting changes may be reflected in real time on the speaker similarity display unit (510), sound quality display unit (520), and synthesis speed display unit (530).
[0317] FIG. 15 illustrates an exemplary user interface supporting multilingual speech synthesis according to one embodiment of the present disclosure.
[0318] Referring to FIG. 15, the multilingual speech synthesis interface (700) may include an input setting unit (710) and a synthesis result output unit (720).
[0319] The input setting section (710) may include an input language selection area, an output language selection area, a voice sample input area, a text to be synthesized input area, and a synthesis execution button. The input language selection area is an interface component that sets the first language in which the speaker's voice sample was spoken, and is implemented in the form of a drop-down menu to select one of a plurality of languages.
[0320] The output language selection area is an interface component for setting a second language in which the text to be synthesized is written, and is similarly implemented in the form of a drop-down menu. In one embodiment, a two-way switching button () is displayed between the input language and the output language, so that the user can easily switch between the input language and the output language. For example, as shown in FIG. 15, when Korean (KR) is set as the input language and English (US English) is set as the output language, synthesized speech corresponding to English text can be generated while maintaining the speaker characteristics of the speech sample spoken in Korean.
[0321] The voice sample input area is an interface component for uploading a voice sample file of a speaker speaking in a first language. In one embodiment, the voice sample input area may display metadata such as the filename, playback length, sampling rate, and speaker identifier of the uploaded voice sample file.
[0322] For example, as illustrated in FIG. 15, when the "speaker_sample_KR.wav" file is uploaded, metadata such as playback length 3.2 seconds, sampling rate 22 kHz, and speaker A may be displayed together. The server computing system (130) inputs the uploaded voice sample into a VAE encoder (31) and a Style Encoder to extract language-agnostic speaker characteristics, thereby enabling the application of speaker-specific characteristics, such as timbre, intonation, and speech style, which are not dependent on the phonological system of the first language, to the synthesis of the second language.
[0323] The synthesis target text input area is an interface component that receives synthesis target text written in a second language, and is implemented in the form of a text input window. In one embodiment, the synthesis target text input area may display the currently set output language as a label to guide the user to input text in the correct language.
[0324] For example, as illustrated in FIG. 15, if the output language is set to English, an "EN" label is displayed, and English text such as "Good morning. Today, I will present the quarterly results for our team." can be entered. The synthesis execution button is a button that allows the user to start synthesis after the input settings are completed. When the button is clicked, the server computing system (130) transmits the set speaker characteristics and the text to be synthesized to the artificial intelligence model (30) to execute the speech synthesis task.
[0325] The synthesis result output section (720) may include a synthesis completion status display area, a synthesized speech waveform display area, a quality indicator display area, and a speaker characteristic retention rate display area. The synthesis completion status display area is a component that displays the current output language and the synthesis progress status, and when synthesis is completed, it displays the completion status along with the output language code.
[0326] The synthesized speech waveform display area is a component that visualizes and displays the waveform of the generated synthesized speech, and can display the total playback time and current playback position of the synthesized speech together. For example, as shown in FIG. 15, the total playback time of the synthesized speech can be displayed as 0:05.4 seconds. The quality indicator display area is a component that displays the speaker similarity, sound quality (MOS), and synthesis speed (RTF) of the generated synthesized speech as numerical values. For example, as shown in FIG. 15, when English synthesized speech is generated from a Korean speech sample, a speaker similarity of 3.62, an MOS of 3.55, and an RTF of 0.054 can be displayed.
[0327] The speaker characteristic retention rate display area is a component that displays, in the form of a bar graph of percentages (%), how much of the speaker characteristics of the first language voice sample are retained in the second language synthesized voice, for detailed items such as timbre, intonation, and speech rate. For example, as shown in FIG. 15, in the cross-language synthesis from Korean (KO) to English (EN), speaker characteristic retention rates of 86% for timbre, 79% for intonation, and 88% for speech rate may be displayed. Through this, the user can intuitively check how faithfully the speaker characteristics of the first language voice sample are reflected in the second language synthesized voice item by item.
[0328] FIG. 16 illustrates an exemplary user interface that supports real-time speech synthesis switching of multiple speakers according to one embodiment of the present disclosure.
[0329] Referring to FIG. 16, the multi-speaker real-time speech synthesis interface (800) may include a speaker registration and management unit (810) and a speaker assignment unit (820) for each segment.
[0330] The speaker registration and management unit (810) is an interface component that registers, displays, and manages multiple speaker profiles. Each speaker profile may be displayed including a speaker identifier, a voice sample filename, a playback length, and a voice waveform preview.
[0331] For example, as illustrated in FIG. 16, the speaker registration and management unit (810) has three speaker profiles registered, including speaker A (sample_A.wav · 3.2s), speaker B (sample_B.wav · 4.1s), and speaker C (sample_C.wav · 2.8s), and the voice waveform of the corresponding speaker is visualized and displayed at the bottom of each profile.
[0332] The user can add a new speaker profile by uploading a new speaker's voice sample through the "+ Add Speaker" button at the bottom of the speaker registration and management section (810). The server computing system (130) inputs the registered voice sample into the VAE encoder (31) and the Style Encoder to extract a speaker embedding vector, maps it to a speaker identifier, and stores it in memory. This allows the stored speaker embedding to be immediately retrieved without delay, without re-encoding the voice sample every time a synthesis request is made, and used as a condition input for the flow matching decoder (32).
[0333] The section-by-section speaker assignment unit (820) is an interface component that divides the text to be synthesized into sections and assigns a speaker to each section. The section-by-section speaker assignment unit (820) displays the section number, assigned speaker, text to be synthesized, playback length of the synthesized voice, and synthesis status in a table format.
[0334] In one embodiment, the user may select and assign one of the speaker profiles registered in the speaker registration and management unit (810) to each section, and the same speaker may be assigned to multiple sections. For example, as shown in FIG. 16, Speaker A is assigned to section 01 and the text "Hello, I will start today's meeting. First, let's review the performance of the last quarter." (3.8s) is synthesized and completed; Speaker B is assigned to section 02 and the text "Yes, sales increased by 12% compared to the last quarter. Performance in the new product line was outstanding." (4.2s) is synthesized and completed; Speaker C is assigned to section 03 and the text "Additionally, I will also share the results of the customer satisfaction survey." (2.6s) is currently being generated and can be seen as Speaker A and Speaker B are assigned to sections 04 and 05, respectively, and are waiting.
[0335] In this way, the speaker assignment unit (820) for each section enables the sequential generation of synthesized voices that reflect the voice characteristics of the speaker corresponding to each speech section in various scenarios such as meetings, conversations, and narrations where multiple speakers take turns speaking.
[0336] At the bottom of the section-specific speaker assignment section (820), there is a full synthesis execution button, a real-time synthesis progress status display area, and a synthesis speed display area. The full synthesis execution button is a button that sequentially starts synthesis for all registered sections. When the button is clicked, the server computing system (130) provides the embedding vectors of the speakers assigned to each section in order from section 01 as condition inputs to the flow matching decoder (32) to generate synthesized speech.
[0337] The real-time synthesis progress display area displays the number of sections synthesized so far relative to the total sections in the form of a progress bar and text. For example, as shown in FIG. 16, the state where synthesis of 2 sections out of a total of 5 sections is completed can be displayed as "2 / 5 sections completed". The synthesis speed display area displays the RTF and the number of sampling steps currently being applied to the synthesis, and, for example as shown in FIG. 16, can be displayed as "RTF 0.052 · 4-step" to indicate that real-time synthesis is possible with 4-step inference.
[0338] FIG. 17 exemplarily illustrates steps 1 to 3 of a synthetic speech-based data augmentation platform for training a speech recognition model according to one embodiment of the present disclosure. FIG. 18 exemplarily illustrates step 4 of the platform.
[0339] Referring to FIGS. 17 and 18, the STT learning data augmentation platform (900) includes an acoustic environment setting unit (910), a synthesis setting unit (920), a synthesis execution status unit (930), and an STT learning data output unit (940). By utilizing each component sequentially, the user can build a synthesized speech dataset for training a speech recognition model through a four-step process of acoustic environment setting, synthesis parameter configuration, real-time synthesis execution, and training data export.
[0340] The acoustic environment setting unit (910) is an interface component that registers voice samples to define the characteristics of the acoustic environment to be reflected in the synthesized voice and displays the acoustic parameters of the environment. In one embodiment, a user can upload a voice sample file recorded in a specific acoustic environment through the acoustic environment setting unit (910), and the server computing system (130) automatically analyzes and displays the acoustic characteristics of the environment from the uploaded voice sample.
[0341] For example, as illustrated in FIG. 17, when the file “conference_room_A.wav (Conference Room A · 12.4s · 3 people · 44kHz) is registered, the acoustic environment setting unit (910) can analyze and display “Conference Room A” as the acoustic environment, “0.42s” as the reverberation time (RT60), “38dB” as the background noise, and “3 people” as the number of sample speakers. Here, the reverberation time (RT60) is the time it takes for the sound pressure level to decrease by 60dB after the sound source is stopped, and is a representative indicator representing the acoustic characteristics of the space. The server computing system (130) reflects the acoustic environment characteristics analyzed in this way into the speaker embedding through the VAE encoder (31) and the Style Encoder, thereby ensuring that the reverberation and background noise characteristics of the corresponding acoustic environment are automatically included in the synthesized speech generated thereafter.
[0342] The synthesis setting section (920) is an interface component for setting detailed parameters for building a synthetic speech dataset. The synthesis setting section (920) may include a synthetic text selection area, a sampling step selection area, a speaker selection area, and a total planned sentence number display area. The synthetic text selection area is a drop-down menu for selecting a text corpus file to be used for generating synthetic speech, and allows selecting a text corpus containing various domains and sentence types suitable for training a speech recognition model.
[0343] For example, as shown in FIG. 17, when "STT_corpus_v2.txt (1,200 sentences)" is selected, synthesized speech is generated for 1,200 sentences included in the text corpus. The sampling step selection area is a drop-down menu for selecting the number of inference steps of the flow matching decoder (32) to be applied to the synthesis; for example, as shown in FIG. 17, when "4-step (AP applied)" is selected, synthesis is performed in a 4-step inference mode with adversarial post-learning applied. The speaker selection area is a component for selecting a speaker profile to be used for synthesis, and allows selecting a speaker to participate in the synthesis from among multiple speakers registered in the acoustic environment setting unit (910). For example, as shown in FIG. 17, when Speaker A, Speaker B, and Speaker C are all selected, 1,200 sentences × 3 people = 3,600 sentences are displayed in the total sentences to be synthesized display area, allowing the user to check the total scale of synthesis in advance.
[0344] The synthesis execution status section (930) is an interface component that displays the overall progress rate and the progress status by speaker in real time during synthesis execution. The synthesis execution status section (930) includes an overall progress display area and a speaker-specific progress status table. The overall progress display area displays the number of sentences synthesized so far relative to the total number of sentences scheduled for synthesis in the form of a progress bar and text.
[0345] For example, as illustrated in FIG. 17, the state in which 1,247 sentences out of a total of 3,600 sentences have been synthesized may be indicated as "1,247 / 3,600 sentences completed". A progress table by speaker displays the synthesis progress rate, the number of completed sentences, RTF, and status for each speaker. For example, as illustrated in FIG. 17, the status may be displayed as follows: Speaker A has completed 696 sentences out of 1,200 (RTF 0.052, in progress), Speaker B has completed 551 sentences out of 1,200 (RTF 0.051, in progress), and Speaker C has completed 0 sentences out of 1,200 (waiting). In one embodiment, the server computing system (130) can improve the processing speed of the entire data augmentation pipeline by performing synthesis for multiple speakers in parallel. The artificial intelligence model (30) of the present disclosure achieves an RTF of approximately 0.052 with 4-step inference, so a synthetic speech dataset for 3,600 sentences can be built at a speed approximately 19 times faster than real-time.
[0346] Referring to FIG. 18, the STT training data output unit (940) is an interface component that displays summary information of synthesized voice data and exports the generated synthesized voice and the corresponding transcribed text as training data for a speech recognition model. The STT training data output unit (940) may include an area displaying the number of generated voices, an area displaying the total voice length, an area displaying the acoustic environment, a generated data preview area, and an STT training data export button.
[0347] The area displaying the number of completed voices is a component that displays the total number of voice files that have been synthesized. For example, as shown in FIG. 18, it may be displayed that there are 1,247 voice files that have been synthesized. The area displaying the total voice length is a component that displays the total playback time of the generated synthesized voice, and for example, as shown in FIG. 18, it may be displayed that a total of 18.4 hours of synthesized voice data have been generated. The area displaying the acoustic environment is a component that displays the acoustic environment reflected in the generated synthesized voice, and for example, as shown in FIG. 18, it may be displayed that the synthesized voice reflects the acoustic environment characteristics of "Conference Room A".
[0348] The generated data preview area is a component that previews the generated synthesized speech and its corresponding transcribed text in a list format. In one embodiment, each item is displayed including a sequence number, a voice waveform preview, transcribed text, and speaker information. For example, as illustrated in FIG. 18, item 001 may display "There are a total of three agenda items for today's meeting" in the voice of Speaker A, item 002 may display "I will review the sales report for the last quarter" in the voice of Speaker B, and item 003 may display "Please check the new product launch schedule" in the voice of Speaker A. This allows the user to check the quality and content of the generated synthesized speech data in advance before exporting.
[0349] The STT training data export button is a button that pairs a generated synthetic speech file (WAV) with its corresponding transcribed text file (TXT) and exports them as training data for a speech recognition model. In one embodiment, the exported data can be directly input into the training pipeline of the speech recognition model and can be used to expand the scale and diversity of the training dataset by mixing it with actual recording data. For example, in cases where actual recording data for a specific acoustic environment is scarce, the data augmentation pipeline of the present disclosure can be used to rapidly construct a large-scale synthetic speech dataset reflecting the acoustic characteristics of that environment, thereby effectively improving the robustness and recognition accuracy of the speech recognition model for that environment.
[0350] As such, the STT learning data augmentation platform (900) according to one embodiment of the present disclosure can rapidly construct a large-scale multi-speaker synthetic speech dataset that reflects the acoustic characteristics of an environment using only a small number of speaker voice samples collected in a specific acoustic environment, and can significantly improve the processing efficiency of the data augmentation pipeline through high-speed synthesis based on 4-step inference. Through this, the problem of insufficient training data for speech recognition models can be effectively resolved, and the robustness and recognition accuracy of the speech recognition model for various acoustic environments and speakers can be improved.
[0351] FIG. 19 illustrates an exemplary user interface that supports automatic dubbing according to one embodiment of the present disclosure.
[0352] Referring to FIG. 19, the user interface may display an automatic dubbing interface (850) including a source video, a voice actor setting section (851), and a section-by-section dubbing result display section (852).
[0353] The source video and voice actor setting unit (851) is a component that receives a source video or audio file to be dubbed and configures a voice sample and language settings of a speaker to be used for dubbing. In one embodiment, the source video and voice actor setting unit (851) receives a video file or audio file, extracts the original voice contained in the file, and inputs the extracted voice into a voice recognition model to obtain text of the original language.
[0354] For example, as illustrated in FIG. 19, the source video and voice actor setting unit (851) can receive an original English video file (source_video_EN.mp4) and extract the voice contained in the video into English text through STT (Speech-to-Text) conversion.
[0355] The source video and voice actor setting unit (851) can support both scenarios: using the text obtained through STT conversion as is without translation, or using it after translating it from the first language into the second language. In the case where translation is omitted, the text of the original language obtained through STT conversion is used as is as input for the generation of synthesized speech.
[0356] For example, by converting the utterance of foreign speaker A into STT and generating English synthesized speech that reflects the voice characteristics of native speaker B without translation, this can be applied to a voice replacement dubbing scenario where only the voice is replaced while maintaining the language of the original utterance.
[0357] When performing translation, the text of the first language (e.g., English, EN) obtained through STT conversion is translated into the second language (e.g., Korean, KO), and the translated text is subsequently used as input for synthetic speech generation.
[0358] For example, as illustrated in FIG. 19, this can be applied to a multilingual dubbing scenario in which the audio of an original English video is converted into STT and translated into Korean, and then a Korean synthesized voice reflecting the voice characteristics of Korean voice actor A is generated. At this time, the source video and voice actor setting unit (851) visually display the three-stage pipeline flow leading to STT conversion, translation, and TTS synthesis, thereby allowing the user to intuitively understand the overall progress of the automatic dubbing process.
[0359] The source video and voice actor setting unit (851) provide a function to set a voice sample of a speaker to be used for dubbing. In one embodiment, the voice sample of the speaker may be a voice sample newly entered by the user or a voice of a speaker selected from a plurality of pre-stored speaker voices.
[0360] For example, as illustrated in FIG. 19, when a voice sample (voice_KO_A.wav, 4.2s) of Korean voice actor A is registered, the artificial intelligence model extracts speaker characteristics from the sample and generates a synthetic voice corresponding to the translated Korean text by reflecting them. Through this, the original video spoken in English can be automatically dubbed into a Korean synthetic voice that preserves the voice characteristics of Korean voice actor A.
[0361] The source video and voice actor setting section (851) may additionally provide a function to set dubbing parameters including the original language, translation language, speech timing synchronization and inference method.
[0362] In one embodiment, when the speech timing synchronization function is enabled (ON), the speech length and speech speed of the synthesized speech are automatically adjusted to match the speech timing of each segment of the original video using a duration predictor and a length regulator included in the artificial intelligence model, thereby allowing the synthesized speech to be naturally synchronized with the mouth shape and speech timing of the original video.
[0363] In particular, in multilingual dubbing scenarios where translation is performed, the utterance length of the synthesized speech may differ from the original utterance length due to the difference in text length between the first language and the second language. In this case, the length regulator adjusts the duration assigned to each phoneme of the text to be synthesized so that the total utterance length of the synthesized speech matches the utterance length of the original segment, and the utterance speed is automatically adjusted accordingly, thereby maintaining timing synchronization with the original video. For example, as shown in FIG. 19, the original language can be set to EN, the translation language to KO, the utterance timing synchronization to ON, and the inference method to 4-step AP.
[0364] The section-by-section dubbing result display unit (852) is a component that displays the original text, translated text, original speech length, synthesized speech length, and dubbing progress status for each section of the source video. In one embodiment, the section-by-section dubbing result display unit (852) displays the original language (EN) text extracted through STT and the second language (KO) text translated therefrom in correspondence for each section, and displays the original speech length and synthesized speech length of each section together so that the user can check whether the speech timing is synchronized.
[0365] For example, as shown in FIG. 19, for section 01, it is indicated that the synthesized utterance length is adjusted to 3.1s for the same original utterance length of 3.1s (3.1s), and for section 02, it is indicated that the synthesized utterance length is synchronized to 4.8s for the original utterance length of 4.8s. This indicates that the timing of the original utterance is accurately maintained through the adjustment of the utterance speed, even when the text length changes after translation.
[0366] The section-by-section dubbing result display unit (852) displays the dubbing progress status of each section as completed, in progress, and waiting. In one embodiment, the completed state indicates a section where the synthesis of the voice and speech timing synchronization have been successfully completed, the in progress state indicates a section where synthesis is currently in progress, and the waiting state indicates a section where synthesis has not yet started. For example, as shown in FIG. 19, sections 01 and 02 are displayed as completed, section 03 as in progress, and sections 04 and 05 as waiting, and the progress bar at the bottom displays that 2 sections out of a total of 5 sections have been completed (2 / 5 sections completed).
[0367] At the bottom of the user interface (850), a dubbing execution button, a real-time dubbing progress bar, and an RTF indicator are displayed. In one embodiment, the Real-Time Factor (RTF) indicator is a value obtained by dividing the time taken for synthesis by the playback time of the generated voice, and if the value is less than 1, it means that real-time dubbing is possible. For example, as shown in FIG. 19, an RTF of 0.052 and 4-step inference are displayed, indicating that the automatic dubbing system according to the present embodiment achieves a synthesis speed at which real-time dubbing is possible with only 4 fractional sampling steps.
[0368] In another embodiment, the automatic dubbing interface (850) may additionally provide a function to replace the original audio of the source video with the dubbed synthesized audio and output it as a final dubbed video file. For example, after all sections of dubbing are completed, the user can generate a dubbed video file in which the audio track of the original video is replaced with the synthesized audio via the final output button, and save it to a local storage or transmit it to an external system.
[0369] In another embodiment, the automatic dubbing interface (850) may also be applied to a scenario in which different speakers are automatically identified in segments through a speaker diarization function for a video featuring multiple speakers, and a voice sample of a target speaker corresponding to each identified speaker is applied to perform multi-speaker automatic dubbing. In this case, a multi-speaker synthesis structure that stores multiple speaker embeddings in memory and rapidly switches between segments according to an embodiment of the present disclosure may be utilized together.
[0370] Meanwhile, the embodiments according to the present disclosure described above may be implemented in the form of program instructions that can be executed through various computer components and recorded on a computer-readable recording medium. The computer-readable recording medium may include program instructions, data files, data structures, etc., either alone or in combination. The program instructions recorded on the computer-readable recording medium may be those specifically designed and configured for the present disclosure or those known and available to those skilled in the art of computer software. Examples of computer-readable recording media include magnetic media such as hard disks, floppy disks, and magnetic tapes; optical recording media such as CD-ROMs and DVDs; magneto-optical media such as floptical disks; and hardware devices specifically configured to store and execute program instructions, such as ROM, RAM, and flash memory. Examples of program instructions include machine code, such as that generated by a compiler, as well as high-level language code that can be executed by a computer using an interpreter, etc. Hardware devices may be modified into one or more software modules to perform processing according to the present disclosure, and vice versa.
[0371] The specific embodiments described in this disclosure are examples and do not limit the scope of this disclosure in any way. For the sake of brevity of the specification, descriptions of prior electronic configurations, control systems, software, and other functional aspects of said systems may be omitted. Additionally, the connections of lines or connecting members between components shown in the drawings are illustrative of functional connections and / or physical or circuit connections, and may be replaced or additionally represented as various functional connections, physical connections, or circuit connections in actual devices. Furthermore, unless specifically stated as “essential,” “importantly,” etc., a component may not be strictly necessary for the application of the embodiments of this disclosure.
[0372] Furthermore, although the detailed description of the present disclosure has been explained with reference to preferred embodiments of the present disclosure, those skilled in the art or those with ordinary knowledge in the art will understand that various modifications and changes can be made to the embodiments of the present disclosure without departing from the spirit and technical scope of the present disclosure as set forth in the claims below. Accordingly, the technical scope of the present disclosure should not be limited to the contents described in the detailed description of the specification but should be determined by the claims.
[0373] Various embodiments of the present disclosure have industrial applicability in that they can be widely applied to actual service environments requiring both high synthesis quality and real-time processing performance in various voice service fields, such as voice assistants, multilingual automatic dubbing, and augmentation of training data for speech recognition models, by applying adversarial post-learning to a flow matching-based decoder to generate high-quality synthesized speech reflecting speaker characteristics in real time with only a small number of sampling steps.
Claims
1. As a method executed by a computer, Step to acquire text; A step of acquiring a voice sample containing the voice characteristics of the speaker; A step in which at least one processor uses at least one artificial intelligence model to generate a synthetic speech corresponding to the text, which reflects the speaker characteristics of the speech sample; wherein the at least one artificial intelligence model is further trained through adversarial learning with a discriminator on an output generated through inference executed with fewer sampling steps than during prior training, and A method comprising the step of inputting the synthesized speech into at least one subsequent processing component.
2. In Paragraph 1, A method further comprising the step of generating the synthesized speech in real time and outputting it through a user interface.
3. In Paragraph 1, A method further comprising the step of displaying at least one of the speaker similarity, sound quality, and synthesis speed of the synthesized speech through a user interface.
4. In Paragraph 1, The above voice sample is spoken in a first language, and the above text is written in a second language different from the first language, and A method further comprising the step of: at least one processor using the artificial intelligence model to generate a synthetic speech corresponding to the second language while maintaining the speaker characteristics of the speech sample.
5. In Paragraph 1, A step of acquiring voice samples of each of multiple speakers; A step of extracting voice characteristics of the plurality of speakers and storing them in at least one memory; and A method further comprising the step of generating a plurality of synthesized speeches in real time by applying speaker characteristics of each of the plurality of speakers to at least a portion of the text by the at least one processor.
6. In Paragraph 1, A method in which the above text includes a prompt text corresponding to the above voice sample and a text to be synthesized, and the synthesized voice corresponds to the text to be synthesized.
7. In Paragraph 6, The step of generating a synthesized speech corresponding to the above text is: A step of encoding the above voice sample into a latent representation; A step of generating a synthetic latent expression corresponding to the synthesis target text based on a flow matching method using a prompt extracted from the above latent expression as a condition; and A method comprising the step of decoding the above synthetic latent expression into synthetic speech.
8. In Paragraph 7, The above flow matching method is an Optimal Transport-based conditional flow matching method, comprising a method for generating the synthetic latent expression by approximating the path from the noise to the synthetic latent expression as a straight line.
9. In Paragraph 7, The step of generating the above synthetic latent expression is, A step of concatenating zero padding of a length corresponding to the synthesis target text to the potential representation of the voice sample; and A method comprising the step of generating the synthetic latent expression by infilling the section corresponding to the zero padding.
10. In Paragraph 1, A method in which the discriminator simultaneously generates an unconditional output for determining sound quality and a conditional output for determining speaker similarity with the voice sample for the output of at least one artificial intelligence model.
11. In Paragraph 1, A method based on a combination of adversarial learning, wherein the adversarial learning is based on a reconstruction loss that minimizes the difference between the output of at least one artificial intelligence model and the latent representation of the speech sample, an adversarial loss that induces the artificial intelligence model to deceive the discriminator, and a feature matching loss that minimizes the difference between the intermediate features of the discriminator and the intermediate features of the output of at least one artificial intelligence model.
12. In Paragraph 6, The step of generating the above-mentioned synthesized voice is, A step of determining the utterance length for each phoneme of the synthesized target text and adjusting the utterance speed of the synthesized speech according to the determined utterance length; and A method comprising the step of generating a synthesized voice reflecting the above-mentioned controlled speech rate.
13. In Paragraph 6, The step of generating the above-mentioned synthesized voice is, A step of extracting fundamental frequency and voiced / unvoiced sound information from the above voice sample; and A method comprising the step of generating a synthesized speech reflecting the above fundamental frequency and voiced / unvoiced sound information.
14. In Paragraph 1, The above subsequent processing component includes a speech recognition model, and A method in which the above-mentioned voice sample is obtained in a predetermined acoustic environment, and the above-mentioned synthesized voice includes a plurality of synthesized voices reflecting the acoustic environment and is utilized as training data for the above-mentioned voice recognition model.
15. In Paragraph 1, The step of obtaining the above text is, A step of extracting voice from a video file or audio file; A step of inputting the extracted voice into a voice recognition model to obtain text of the first language; and The method includes the step of translating the text of the first language into the second language to obtain the text of the second language; and The step of acquiring the above voice sample is, The method includes the step of acquiring a newly input voice sample or a voice sample of a selected speaker among a plurality of pre-stored speaker voices; The step of generating the above-mentioned synthesized voice is, A method comprising the step of generating a synthesized speech corresponding to the text of the second language, which reflects the speaker characteristics of the voice sample.
16. In Paragraph 15, The step of generating the above-mentioned synthesized voice is, A step of determining the utterance length for each phoneme of the text of the second language; and A method comprising the step of generating a synthesized speech corresponding to the text of the second language by adjusting the determined utterance length to be synchronized with the utterance timing of each segment of the video file.
17. At least one memory; and At least one processor that performs a method of reading at least one instruction stored in at least one memory and generating a response to a user's target prediction request; The above at least one instruction is, Step of acquiring the text to be synthesized; A step of acquiring a voice sample containing the voice characteristics of the speaker; A step in which at least one processor uses at least one artificial intelligence model to generate a synthetic speech corresponding to the target text for synthesis, the synthetic speech reflecting the speaker characteristics of the voice sample; wherein the at least one artificial intelligence model includes a generative model additionally trained through adversarial learning with a discriminator on an output executed with fewer sampling steps than during prior training; and A system comprising a command to perform the step of inputting the synthesized speech into at least one subsequent processing component.
18. In Paragraph 17, A plurality of neurons comprising an array including at least one register, at least one programmable logic, and at least one input interface; a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; and at least one routing network that controls the data flow between the plurality of neurons; comprising A system further comprising a Field Programmable Gate Array (FPGA) implementation for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through the routing network to establish a transmission path for the weights.
19. In Paragraph 17, A plurality of neurons organized into an array comprising at least one register, at least one microprocessor, and at least one input; and a plurality of synapse circuits storing synapse weights that regulate the connection strength between the plurality of neurons; comprising A system comprising an Application Specific Integrated Circuit (ASIC) for a predetermined artificial neural network, wherein each of the plurality of neurons is connected to at least one other neuron through any one of the plurality of synaptic circuits.