System and method for fine-tuning text-based voice by using user interface

The method addresses the limitations of existing text-to-speech technologies by using a user interface to fine-tune emotional expression through emotion embedding vectors and feedback, enhancing emotional control and user experience.

WO2026054579A1PCT designated stage Publication Date: 2026-03-12NEOSAPIENCE INC
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-05
Publication Date
2026-03-12

AI Technical Summary

Technical Problem

Existing text-to-speech technologies struggle to precisely control emotional expression and require repetitive user input for subtle emotional adjustments, leading to inefficiencies and user fatigue.

Method used

A method and system for fine-tuning text-based speech through a user interface, utilizing emotion embedding vectors and feedback mechanisms to intuitively adjust emotional nuances, allowing users to select and modify emotional information visually and through natural language inputs.

Benefits of technology

Enables precise and intuitive emotional control in synthetic voices, improving naturalness and accuracy of emotional expression, making it accessible even for non-experts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025013790_12032026_PF_FP_ABST
    Figure KR2025013790_12032026_PF_FP_ABST
Patent Text Reader

Abstract

The present disclosure relates to a method, performed by at least one processor, for fine-tuning a text-based voice by using a user interface (UI), the method comprising the steps of: acquiring a text and a voice; acquiring feedback on the acquired voice by using a UI; and fine-tuning the acquired voice on the basis of the acquired feedback.
Need to check novelty before this filing date? Find Prior Art

Description

Text-based speech fine-tuning system and method via user interface

[0001] The present disclosure relates to a text-based speech fine-tuning system and method through a user interface, and more specifically, to a system and method for selectively fine-tuning only certain segments included in speech.

[0002] Emotional synthesis voice technology is being utilized as a primary means to enhance user experience in the field of Text-to-Speech (TTS). While existing TTS technologies have focused on naturally converting sentences into speech, recent advancements have moved beyond simple speech accuracy to express the speaker's intent or mood, including emotions, intonation, and speaking style. In particular, demand for emotion-reflective voice generation is increasing across various application fields, such as movie dubbing, character voices, navigation, and educational content.

[0003] To address these demands, various methods for emotional regulation have been proposed, but each of these conventional technologies has clear limitations.

[0004] First, the emotion label-based method generates voice by manually selecting a predefined emotion label (e.g., joy, sadness, anger, etc.) for each sentence, which has limitations in that the range of emotional expression is limited and it is difficult to control subtle emotional differences. In addition, since the user must set the emotion directly for every sentence, there is a lot of repetitive work and user fatigue increases.

[0005] Next, the clustering-based method selects similar emotional expressions as clusters in the emotion embedding space. However, the adjustable range is determined by the number of clusters defined in advance, and it is difficult to clearly know what emotion each cluster represents, making it difficult for users to intuitively select the desired emotion.

[0006] In addition, the natural language prompt-based method sets emotions based on the description of the emotion input by the user in natural language, which provides a high degree of freedom of expression, but has limitations in that the input intent may not be accurately reflected or the prediction results may not be consistent due to the ambiguity of natural language.

[0007] Finally, retake-based methods rely on the model's randomness to generate multiple voices for the same text input and select the one closest to the desired emotion. This method has low controllability and requires many repeated attempts to obtain the desired voice, which is inefficient.

[0008] As a result, existing technologies have limitations in precisely controlling emotional expression or in fine-tuning it based on the user's intuitive feedback. In applications that seek to precisely reflect emotional nuances, a more flexible and intuitive emotional control method is required.

[0009] The background technology described above is something that the inventor possessed or acquired in the process of deriving the disclosure of the present application, and cannot necessarily be said to be a publicly known technology disclosed to the general public prior to the present application.

[0010] The present disclosure provides a method for fine-tuning text-based voice through a user interface, a computer program stored on a recording medium, and a device (system) to solve the above-described problems.

[0011] The present disclosure can be implemented in various ways, including as a method, a system (device), or a computer program stored on a readable storage medium.

[0012] According to one embodiment of the present disclosure, a method for fine-tuning text-based speech through a user interface performed by at least one processor may include the steps of acquiring text and speech, acquiring feedback on the acquired speech through a user interface (UI), and fine-tuning the acquired speech based on the acquired feedback.

[0013] Additionally, the step of obtaining feedback may include a step of obtaining, as feedback, an emotion information modification request that requests that initial emotion information corresponding to a voice corresponding to at least a portion of the acquired voice be modified to target emotion information.

[0014] Additionally, the step of obtaining a request for modifying emotional information may include a step of providing a plurality of emotional information corresponding to different types of emotions through a user interface and a step of selecting one of the provided plurality of emotional information as target emotional information.

[0015] Additionally, the step of providing multiple emotional information includes the step of visualizing multiple emotional information and providing the visualized multiple emotional information through a user interface, and the visualized multiple emotional information is a result obtained by reducing multiple emotional embedding vectors extracted from each of the multiple emotional information to two dimensions or three dimensions through a dimensionality reduction algorithm, and may be a two-dimensional plane or a three-dimensional space in which multiple points corresponding to each of the multiple emotional embedding vectors are displayed.

[0016] Additionally, the step of fine-tuning the acquired voice may include, if there exists an emotion embedding vector corresponding to any one of the emotion information among the previously generated multiple emotion embedding vectors, a step of regenerating voice corresponding to at least a portion of the acquired voice using an emotion embedding vector corresponding to any one of the emotion information.

[0017] Additionally, the step of fine-tuning the acquired voice may include, in cases where there is no emotion embedding vector corresponding to any one of the emotion information among the previously generated multiple emotion embedding vectors, a step of generating an emotion embedding vector corresponding to any one of the emotion information using a previously trained emotion prediction model or emotion search model, and a step of regenerating voice corresponding to at least a portion of the acquired voice using the generated emotion embedding vector.

[0018] In addition, the step of providing multiple emotional information includes a step of generating multiple example voices corresponding to each of the multiple emotional information using speaker information corresponding to the acquired voice and each of the multiple emotional information; and

[0019] When specific emotional information is selected from among multiple emotional information from a user, a step of providing an example voice corresponding to the specific emotional information from among multiple generated example voices may be included.

[0020] Additionally, the step of providing multiple emotional information may include, when one or more search terms related to emotions are entered by a user through a user interface, a step of searching for emotional information corresponding to one or more entered search terms among the multiple emotional information, and a step of providing the emotional information searched through the user interface.

[0021] Additionally, the step of fine-tuning the acquired voice may include the step of calculating a difference between a first embedding vector corresponding to initial emotional information and a second embedding vector corresponding to target emotional information, and the step of regenerating a voice corresponding to at least a portion of the acquired voice based on the calculated difference and a preset scale factor.

[0022] Additionally, the step of obtaining feedback may include a step of obtaining, as feedback, a step of obtaining a speech style modification request that requests that an initial speech style corresponding to a speech corresponding to at least a portion of the acquired speech be modified to a target speech style.

[0023] Additionally, the step of obtaining feedback may include a step of obtaining text-based feedback for a voice corresponding to at least a portion of the acquired voice through a user interface, and the step of modifying the acquired voice may include a step of obtaining modifications to the voice corresponding to at least a portion of the acquired voice by analyzing the obtained text-based feedback, and a step of regenerating the voice corresponding to at least a portion of the acquired voice based on the obtained modifications.

[0024] In addition, the step of obtaining feedback includes the step of providing an emotion regulation interface through a user interface and the step of obtaining a request for modification of a voice corresponding to at least a portion of a voice obtained as feedback through the provided emotion regulation interface, wherein the provided emotion regulation interface includes a plurality of attribute setting elements that individually set each of a plurality of parameters related to an attribute of an emotion, and the obtained modification request may be a parameter adjustment value obtained through at least one attribute setting element among the plurality of attribute setting elements.

[0025] In addition, the step of acquiring voice includes a step of generating speaker and emotion condition information based on speaker information of a specific speaker and initial emotion information for a specific emotion, and a step of generating a synthetic voice including a voice spoken by a specific speaker with a specific emotion by converting a plurality of sentences into voices based on the generated speaker and emotion condition information through a pre-trained text-to-speech conversion model, and the initial emotion information may be a plurality of emotion information derived corresponding to each of a plurality of sentences by analyzing a document including a plurality of sentences through a pre-trained emotion prediction model or emotion retrieval model.

[0026] In addition, the step of generating a synthetic voice including a voice spoken by a specific speaker with a specific emotion may include, when one document is a script, a step of obtaining additional information including at least one of genre information of the script, situational and location information for a scene corresponding to the script, and a step of generating a synthetic voice by converting a plurality of sentences into voices based on the generated speaker and emotion condition information and the obtained additional information.

[0027] A computer program stored on a computer-readable recording medium may be provided to execute the method for fine-tuning text-based voice through the user interface described above on a computer.

[0028] According to one embodiment of the present disclosure, an information processing system includes a memory and at least one processor coupled to the memory and configured to execute at least one computer-readable program contained in the memory, wherein the at least one program may include instructions for acquiring text and speech, obtaining feedback on the acquired speech through a user interface (UI), and fine-tuning the acquired speech based on the obtained feedback.

[0029] According to some embodiments of the present disclosure, by fine-tuning the synthetic voice according to feedback obtained through a user interface (UI), the emotion of the synthetic voice can be finely adjusted based on the intuitive feedback of the user, thereby improving the naturalness and accuracy of emotional expression.

[0030] According to some embodiments of the present disclosure, there is an advantage in that even non-experts can easily generate a voice with a desired emotional style by providing various forms of user interfaces (UIs) that can fine-tune a synthetic voice.

[0031] The effects of the present disclosure are not limited to the effects mentioned above, and other effects not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs (referred to as “one skilled in the art”) from the description of the claims.

[0032] Embodiments of the present disclosure will be described below with reference to the accompanying drawings, wherein like reference numerals represent similar elements, but are not limited thereto.

[0033] FIG. 1 is a drawing showing an example of fine-tuning a synthesized voice according to one embodiment of the present disclosure.

[0034] FIG. 2 is a schematic diagram showing a configuration in which an information processing system is connected to enable communication with a plurality of user terminals in order to provide a text-based voice fine-tuning service through a user interface in one embodiment of the present disclosure.

[0035] FIG. 3 is a block diagram showing the internal configuration of a user terminal and an information processing system according to one embodiment of the present disclosure.

[0036] FIG. 4 is a flowchart of a text-based voice fine-tuning method through a user interface according to one embodiment of the present disclosure.

[0037] FIG. 5 is a flowchart of a method for obtaining feedback using a plurality of emotional information provided through a user interface according to one embodiment of the present disclosure.

[0038] FIG. 6 is a drawing exemplarily illustrating a plurality of emotional information visualized on a two-dimensional plane according to one embodiment of the present disclosure.

[0039] FIG. 7 is a diagram exemplarily illustrating a plurality of emotional information visualized in a three-dimensional space according to one embodiment of the present disclosure.

[0040] FIG. 8 is a flowchart of a method for fine-tuning a synthetic speech based on text-based feedback obtained through a user interface according to one embodiment of the present disclosure.

[0041] FIG. 9 is a flowchart of a method for fine-tuning a synthetic voice based on feedback obtained through an emotion regulation interface provided through a user interface according to one embodiment of the present disclosure.

[0042] FIG. 10 is a drawing illustrating an emotion control interface according to one embodiment of the present disclosure in an exemplary manner.

[0043] FIG. 11 is a flowchart of a method for generating an emotion prediction model according to one embodiment of the present disclosure.

[0044] FIG. 12 is a flowchart of a method for generating an emotion search model according to one embodiment of the present disclosure.

[0045] FIG. 13 is a flowchart of a synthetic speech generation method according to one embodiment of the present disclosure.

[0046] FIG. 14 is a flowchart of a method for generating speaker and emotion condition information according to one embodiment of the present disclosure.

[0047] FIG. 15 is a flowchart of a method for generating a condition generation model according to one embodiment of the present disclosure.

[0048] Hereinafter, specific details for implementing the present disclosure will be described in detail with reference to the attached drawings. However, in the following description, specific descriptions regarding widely known functions or configurations will be omitted if there is a risk that the gist of the present disclosure may be unnecessarily obscured.

[0049] In the attached drawings, identical or corresponding components are assigned the same reference numerals. Furthermore, in the description of the embodiments below, duplicate descriptions of identical or corresponding components may be omitted. However, even if a description of a component is omitted, it is not intended that such component is not included in any embodiment.

[0050] The advantages and features of the disclosed embodiments, and methods for achieving them, will become clearer with reference to the embodiments described below, along with the accompanying drawings. However, the present disclosure is not limited to the embodiments disclosed below and may be implemented in various different forms. These embodiments are provided solely to ensure the completeness of the disclosure and to fully inform those skilled in the art of the scope of the invention.

[0051] The terms used in this specification will be briefly explained, followed by a detailed description of the disclosed embodiments. The terms used in this specification have been selected from widely used, current terms, taking into account the functions of the present disclosure. However, these terms may vary depending on the intentions of engineers working in the relevant field, precedents, the emergence of new technologies, etc. Furthermore, in certain cases, terms may be arbitrarily selected by the applicant, and in such cases, their meanings will be described in detail in the relevant description of the invention. Therefore, the terms used in this disclosure should not be defined simply as names of terms, but rather based on their meanings and the overall content of the present disclosure.

[0052] In this specification, singular expressions include plural expressions unless the context clearly indicates otherwise. Furthermore, plural expressions include singular expressions unless the context clearly indicates otherwise. When a part of the specification is said to include a component, this does not exclude other components, but rather implies that other components may be included, unless otherwise specifically stated.

[0053] Also, the term 'module' or 'part' used in the specification means a software or hardware component, and the 'module' or 'part' performs certain roles. However, the 'module' or 'part' is not limited to software or hardware. The 'module' or 'part' may be configured to reside on an addressable storage medium and may be configured to execute one or more processors. Thus, as an example, the 'module' or 'part' may include at least one of components such as software components, object-oriented software components, class components, and task components, processes, functions, attributes, procedures, subroutines, segments of program code, drivers, firmware, microcode, circuitry, data, databases, data structures, tables, arrays, or variables. The functionality provided within the components and 'modules' or 'parts' may be combined into a smaller number of components and 'modules' or 'parts', or further separated into additional components and 'modules' or 'parts'.

[0054] In one embodiment of the present disclosure, a 'module' or 'part' may be implemented as a processor and memory. The term 'processor' should be broadly interpreted to include a general-purpose processor, a central processing unit (CPU), a microprocessor, a digital signal processor (DSP), a controller, a microcontroller, a state machine, etc. In some contexts, the term 'processor' may refer to an application-specific integrated circuit (ASIC), a programmable logic device (PLD), a field programmable gate array (FPGA), etc. The term 'processor' may also refer to a combination of processing devices, such as, for example, a combination of a DSP and a microprocessor, a combination of multiple microprocessors, a combination of one or more microprocessors combined with a DSP core, or any other combination of such configurations. Additionally, the term 'memory' should be broadly interpreted to include any electronic component capable of storing electronic information. 'Memory' may refer to various types of processor-readable media, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, magnetic or optical data storage, registers, etc. Memory is said to be in electronic communication with the processor if the processor can read information from, and / or write information to, the memory. Memory integrated in a processor is in electronic communication with the processor.

[0055] In the present disclosure, the "system" may include, but is not limited to, at least one of a server device and a cloud device. For example, the system may be comprised of one or more server devices. As another example, the system may be comprised of one or more cloud devices. As yet another example, the system may be configured and operated by a combination of a server device and a cloud device.

[0056] FIG. 1 is a diagram illustrating an example of fine-tuning a synthesized voice according to one embodiment of the present disclosure. As illustrated in FIG. 1, an information processing system (100) can receive a voice (110) and feedback (120) corresponding to the voice (110) to generate a fine-tuned voice (130).

[0057] The information processing system (100) can acquire voice (110) and text (140). In one embodiment, the voice (110) may be acquired in correspondence with the text. For example, the voice (110) may be a synthetic voice generated in correspondence with the text based on text-to-speech conversion. For example, the voice (110) may be a synthetic voice generated by converting text into speech through a pre-trained Text-to-Speech (TTS) model. Additionally, or alternatively, the voice (110) may be a voice recording of a real person's voice, for example, acquired in correspondence with the text (e.g., the result of recording a voice reading a pre-prepared script).

[0058] Additionally or otherwise, the information processing system (100) can obtain voice (110) and text (140) by recognizing text from voice (110) through a pre-learned speech-to-text (STT) model. Additionally or otherwise, the information processing system (100) can receive input for text corresponding to voice (110). As another example, voice (110) and text (140) can be obtained by inputting text (140) corresponding to voice (110) by a user.

[0059] In FIG. 1, voice (110) and text (140) are shown as being input together into the information processing system (100), but are not limited thereto. As described above, the information processing system (100) may receive only voice (110) to obtain text (140) from voice (110), and conversely, may receive text (140) to obtain voice from text (140).

[0060] In addition, the voice (110) is a voice generated by converting text into voice based on speaker and condition information generated based on speaker information for a specific speaker and emotion information for a specific emotion, and may be a voice uttered by a specific speaker with a specific emotion.

[0061] In one embodiment, feedback (120) may be user input obtained from a user through a user interface (UI) for fine-tuning the voice (110). For example, feedback (120) may be a request to modify emotional information or a request to modify a speaking style obtained through the user interface (UI).

[0062] Here, the request for modifying emotional information may be a request to modify the initial emotional information corresponding to the voice corresponding to at least a portion of the voice (110) to target emotional information.

[0063] Additionally, the request for modifying the speech style here may be a request to modify the initial speech style corresponding to the speech corresponding to at least a portion of the speech (110) to a target speech style. Here, the speech style may be, but is not limited to, a dialect, an accent, etc.

[0064] In one embodiment, the information processing system (100) can generate a finely tuned voice (130) by fine-tuning the voice (110) based on feedback (120).

[0065] Here, the fine-tuned voice (130) may be the result of fine-tuning a voice corresponding to a portion of the voice (110) according to target emotion information included in the feedback (120), or the result of fine-tuning a voice corresponding to a portion of the voice (110) according to target speech style included in the feedback (120). Here, fine-tuning may involve modifying the voice corresponding to a portion of the portion, but in some cases, it may involve regenerating the voice corresponding to a portion of the portion.

[0066] FIG. 2 is a schematic diagram showing a configuration in which an information processing system (230) is connected to communicate with a plurality of user terminals (210_1, 210_2, 210_3) to provide a text-based voice fine-tuning service through a user interface in one embodiment of the present disclosure. As illustrated, the plurality of user terminals (210_1, 210_2, 210_3) may be connected to an information processing system (230) capable of performing text-based synthesized voice fine-tuning through a user interface via a network (220). Here, the plurality of user terminals (210_1, 210_2, 210_3) may include a user terminal that receives text-based voice fine-tuning or related services through a user interface. In one embodiment, the information processing system (230) may include one or more server devices and / or databases capable of storing, providing, and executing computer-executable programs and data related to fine-tuning text-based synthetic speech through a user interface, or one or more distributed computing devices and / or distributed databases based on cloud computing services.

[0067] The text-based voice fine-tuning service through a user interface provided by the information processing system (230) can be provided to the user through an application or web browser installed on each of the multiple user terminals (210_1, 210_2, 210_3). For example, the information processing system (230) can provide a text-based voice fine-tuning service through a user interface received from the user terminals (210_1, 210_2, 210_3) through an application, etc., or perform corresponding processing.

[0068] Multiple user terminals (210_1, 210_2, 210_3) can communicate with an information processing system (230) through a network (220). The network (220) can be configured to enable communication between the multiple user terminals (210_1, 210_2, 210_3) and the information processing system (230). Depending on the installation environment, the network (220) may be configured as a wired network such as a TCP / IP-based network, Ethernet, Power Line Communication, telephone line communication device and RS-serial communication, a mobile communication network, a Wireless LAN (WLAN), Wi-Fi, Bluetooth and ZigBee, or a combination thereof. The communication method is not limited, and may include not only a communication method utilizing a communication network (e.g., a mobile communication network, wired Internet, wireless Internet, broadcasting network, satellite network, etc.) that the network (220) may include, but also short-range wireless communication between user terminals (210_1, 210_2, 210_3).

[0069] In FIG. 2, a mobile phone terminal (210_1), a tablet terminal (210_2), and a PC terminal (210_3) are illustrated as examples of user terminals, but are not limited thereto, and the user terminals (210_1, 210_2, 210_3) may be any computing device capable of wired and / or wireless communication and capable of installing and executing applications or web browsers. For example, the user terminals may include AI speakers, smartphones, mobile phones, navigation systems, computers, laptops, digital broadcasting terminals, PDAs (Personal Digital Assistants), PMPs (Portable Multimedia Players), tablet PCs, game consoles, wearable devices, IoT (Internet of Things) devices, VR (virtual reality) devices, AR (augmented reality) devices, MR (Mixed reality) devices, set-top boxes, etc. Additionally, FIG. 2 illustrates three user terminals (210_1, 210_2, 210_3) communicating with an information processing system (230) through a network (220), but is not limited thereto, and may be configured so that a different number of user terminals communicate with an information processing system (230) through a network (220).

[0070] FIG. 3 is a block diagram showing the internal configuration of a user terminal (210) and an information processing system (230) according to one embodiment of the present disclosure. The user terminal (210) may refer to any computing device capable of executing an application or a web browser and capable of wired / wireless communication, and may include, for example, a mobile phone terminal (210_1), a tablet terminal (210_2), a PC terminal (210_3) of FIG. 2. As illustrated, the user terminal (210) may include a memory (312), a processor (314), a communication module (316), and an input / output interface (318). Similarly, the information processing system (230) may include a memory (332), a processor (334), a communication module (336), and an input / output interface (338). As illustrated in FIG. 3, the user terminal (210) and the information processing system (230) may be configured to communicate information and / or data through the network (220) using their respective communication modules (316, 336). Additionally, the input / output device (320) may be configured to input information and / or data to the user terminal (210) or output information and / or data generated from the user terminal (210) through the input / output interface (318).

[0071] The memory (312, 332) may include any non-transitory computer-readable recording medium. According to one embodiment, the memory (312, 332) may include a permanent mass storage device such as a read-only memory (ROM), a disk drive, a solid state drive (SSD), a flash memory, etc. As another example, a permanent mass storage device such as a ROM, an SSD, a flash memory, a disk drive, etc. may be included in the user terminal (210) or the information processing system (230) as a separate permanent storage device distinct from the memory. In addition, the memory (312, 332) may store an operating system and at least one program code (e.g., code for an application installed in the user terminal (210).

[0072] These software components may be loaded from a computer-readable recording medium separate from the memory (312, 332). This separate computer-readable recording medium may include a recording medium directly connectable to the user terminal (210) and the information processing system (230), and may include, for example, a computer-readable recording medium such as a floppy drive, a disk, a tape, a DVD / CD-ROM drive, a memory card, etc. As another example, the software components may be loaded into the memory (312, 332) through a communication module other than a computer-readable recording medium. For example, at least one program may be loaded into the memory (312, 332) based on a computer program that is installed by files provided by developers or a file distribution system that distributes installation files of applications through a network (220).

[0073] The processor (314, 334) may be configured to process instructions of a computer program by performing basic arithmetic, logic, and input / output operations. Instructions may be provided to the processor (314, 334) by a memory (312, 332) or a communication module (316, 336). For example, the processor (314, 334) may be configured to execute instructions received according to program code stored in a storage device such as the memory (312, 332).

[0074] The communication module (316, 336) may provide a configuration or function for the user terminal (210) and the information processing system (230) to communicate with each other via the network (220), and may provide a configuration or function for the user terminal (210) and / or the information processing system (230) to communicate with another user terminal or another system (e.g., a separate cloud system). For example, a request or data (e.g., a text-based voice fine-tuning request via a user interface) generated by the processor (314) of the user terminal (210) according to program code stored in a recording device such as memory (312) may be transmitted to the information processing system (230) via the network (220) under the control of the communication module (316). Conversely, control signals or commands provided under the control of the processor (334) of the information processing system (230) can be received by the user terminal (210) through the communication module (336) and the network (220) via the communication module (316) of the user terminal (210). For example, the user terminal (210) can receive fine-tuning results of text-based voice through a user interface from the information processing system (230) via the communication module (316).

[0075] The input / output interface (318) may be a means for interfacing with an input / output device (320). As an example, the input device may include a device such as a camera, keyboard, microphone, mouse, etc., including an audio sensor and / or an image sensor, and the output device may include a device such as a display, a speaker, a haptic feedback device, etc. As another example, the input / output interface (318) may be a means for interfacing with a device that has a configuration or function integrated into one for performing input and output, such as a touch screen. For example, when the processor (314) of the user terminal (210) processes a command of a computer program loaded into the memory (312), a service screen configured using information and / or data provided by the information processing system (230) or another user terminal may be displayed on the display through the input / output interface (318). In FIG. 3, the input / output device (320) is illustrated as not being included in the user terminal (210), but is not limited thereto, and may be configured as a single device with the user terminal (210). In addition, the input / output interface (338) of the information processing system (230) may be a means for interfacing with a device (not shown) for input or output that is connected to the information processing system (230) or that the information processing system (230) may include. In FIG. 3, the input / output interfaces (318, 338) are illustrated as elements configured separately from the processors (314, 334), but are not limited thereto, and the input / output interfaces (318, 338) may be configured to be included in the processors (314, 334).

[0076] The user terminal (210) and the information processing system (230) may include more components than those shown in FIG. 3. However, it is not necessary to explicitly illustrate most of the conventional components. According to one embodiment, the user terminal (210) may be implemented to include at least some of the input / output devices (320) described above. In addition, the user terminal (210) may further include other components, such as a transceiver, a Global Positioning System (GPS) module, a camera, various sensors, a database, etc. For example, if the user terminal (210) is a smartphone, it may include components that a smartphone generally includes, and various components, such as an acceleration sensor, a gyro sensor, a camera module, various physical buttons, buttons using a touch panel, input / output ports, and a vibrator for vibration, may be implemented to be further included in the user terminal (210). According to one embodiment, the processor (314) of the user terminal (210) may be configured to operate an application, etc. At this time, code associated with the application and / or program may be loaded into the memory (312) of the user terminal (210).

[0077] While a program for an application, etc. is running, the processor (314) may receive text, images, videos, voices and / or actions, etc. input or selected through input devices such as a camera, microphone, including a touch screen, keyboard, audio sensor and / or image sensor connected to an input / output interface (318), and may store the received text, images, videos, voices and / or actions, etc. in the memory (312) or provide them to the information processing system (230) through the communication module (316) and the network (220). For example, the processor (314) may receive data related to the generation of speaker and emotional condition information for text-to-speech conversion by a user and provide the data to the information processing system (230) through the communication module (316) and the network (220).

[0078] The processor (314) of the user terminal (210) may be configured to manage, process, and / or store information and / or data received from an input device (320), another user terminal, an information processing system (230), and / or multiple external systems. The information and / or data processed by the processor (314) may be provided to the information processing system (230) via a communication module (316) and a network (220). The processor (314) of the user terminal (210) may transmit information and / or data to an input / output device (320) via an input / output interface (318) and output the information and / or data.

[0079] The processor (334) of the information processing system (230) may be configured to manage, process, and / or store information and / or data received from multiple user terminals (210) and / or multiple external systems. Information and / or data processed by the processor (334) may be provided to the user terminal (210) via a communication module (336) and a network (220).

[0080] The processor (334) of the information processing system (230) may be configured to output processed information and / or data through an output device (320) such as a display output capable device (e.g., a touch screen, a display, etc.) or a voice output capable device (e.g., a speaker) of the user terminal (210).

[0081] FIG. 4 is a flowchart of a method for fine-tuning text-based speech through a user interface according to one embodiment of the present disclosure. The method (400) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of the method (400) may be divided and performed by at least one processor of the user terminal and at least one processor of the information processing system.

[0082] In method (400), first, the processor can obtain text and voice (S410).

[0083] In one embodiment, the processor can generate synthetic speech corresponding to text based on text-to-speech conversion. For example, the processor can generate synthetic speech by converting text into speech based on a pre-trained text-to-speech conversion model. In this case, the processor can generate synthetic speech including speech uttered by a specific speaker with a specific emotion by converting text into speech based on speaker and emotion condition information generated from speaker information and emotion information. However, the processor is not limited thereto and may receive voice data corresponding to the speech from a user or receive voice data directly through a separate voice input device (e.g., a microphone). For example, the speech may be a recording of a real person's voice, or it may be acquired in correspondence with the text (e.g., the result of recording a voice reading a pre-prepared script).

[0084] Additionally or alternatively, the processor can acquire speech and text by recognizing text through a speech-to-text (STT) model trained on speech. Additionally or alternatively, the processor can receive input for text corresponding to the speech. As another example, the speech and text can be acquired by a user inputting text corresponding to the speech. Thereafter, the processor can obtain feedback on the speech through a user interface (UI) (S420). For example, the processor can provide a UI, and by providing a synthetic speech through the UI, it can obtain feedback instructing fine-tuning of the synthetic speech.

[0085] Here, the feedback may be, but is not limited to, an emotion information modification request that requests that the initial emotion information corresponding to the voice corresponding to at least a portion of the synthetic voice be modified to target emotion information, or a speech style modification request that requests that the initial speech style corresponding to the voice corresponding to at least a portion of the synthetic voice be modified to the target speech style.

[0086] Afterwards, the processor can fine-tune the voice based on the feedback (S430).

[0087] In one embodiment, the processor may regenerate speech corresponding to at least a portion of the synthetic speech based on the target emotional information upon receiving an emotional information modification request, which requests that initial emotional information corresponding to speech corresponding to at least a portion of the synthetic speech be modified to target emotional information as feedback. For example, the processor may calculate a difference between a first embedding vector corresponding to the initial emotional information and a second embedding vector corresponding to the target emotional information, and may regenerate speech corresponding to at least a portion of the synthetic speech based on the calculated difference.

[0088] At this time, the processor can regenerate a voice corresponding to at least a portion of the synthetic voice by reflecting a scale factor of a predetermined size to the difference between the first embedding vector and the second embedding vector. For example, the processor can reflect a scale factor of 1 when it is desired to change the initial emotional information to the same degree as the feeling corresponding to the target emotional information, can reflect a scale factor of less than 1 when it is desired to change the feeling corresponding to the target emotional information less, and can reflect a scale factor of greater than 1 when it is desired to change the feeling corresponding to the target emotional information more.

[0089] In one embodiment, the processor can regenerate speech corresponding to at least a portion of the synthetic speech based on the target speech style upon obtaining a speech style modification request that requests, as feedback, to modify an initial speech style corresponding to the speech corresponding to at least a portion of the synthetic speech to a target speech style.

[0090] FIG. 5 is a flowchart of a method for obtaining feedback using multiple pieces of emotional information provided through a user interface according to one embodiment of the present disclosure. Method (500) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of method (500) may be performed separately by at least one processor of the user terminal and at least one processor of the information processing system.

[0091] In method (500), first, the processor can provide multiple emotional information via a user interface (UI) (S510). The processor can provide multiple emotional information corresponding to different types of emotions, allowing the user to select target emotional information via the user interface (UI).

[0092] In one embodiment, the processor may visualize a plurality of emotional information and provide the visualized plurality of emotional information through a user interface (UI). Here, the visualized plurality of information may be a result obtained by reducing a plurality of emotional embedding vectors extracted from each of the plurality of emotional information into two dimensions or three dimensions through a dimensionality reduction algorithm. For example, the visualized plurality of emotional information may be a two-dimensional plane (600) in which a plurality of two-dimensional points corresponding to each of the plurality of emotional embedding vectors are displayed, as shown in FIG. 6, or a three-dimensional space (700) in which a plurality of three-dimensional points corresponding to each of the plurality of emotional embedding vectors are displayed, as shown in FIG. 7.

[0093] For example, the processor can generate multiple visualized emotion information by mapping multiple emotion embedding vectors into a 2D plane or 3D space based on dimensionality reduction algorithms such as Principal Component Analysis (PCA), t-SNE (t-distributed Stochastic Neighbor Embedding), and UMAP (Uniform Manifold Approximation and Projection).

[0094] In one embodiment, the processor may generate multiple example voices corresponding to each of the multiple emotional information and may provide the multiple emotional information and multiple example voices together. For example, the processor may generate multiple example voices including voices in which the same speaker as the synthesized voice utters an emotion corresponding to each of the multiple emotional information, based on a pre-trained text-to-speech conversion model and using speaker information corresponding to the synthesized voice and each of the multiple emotional information. Subsequently, the processor may provide the multiple emotional information and the multiple example voices corresponding to each of the multiple emotional information together through a user interface (UI), or, in some cases, when a specific emotional information among the multiple emotional information is selected by the user, the processor may provide an example voice corresponding to the specific emotional information among the multiple example voices.

[0095] In one embodiment, when a processor receives one or more search terms related to emotions from a user through a user interface (UI), it can search for emotional information corresponding to one or more search terms among a plurality of emotional information and provide the searched emotional information through the user interface (UI). For example, when the processor obtains the search term 'sadness' from a user, it can select only emotional information related to 'sadness' among a plurality of emotional information and provide it to the user. For instance, the processor may display only points for emotional information corresponding to the user's search term on a two-dimensional plane or a three-dimensional space output through the user interface (UI), but is not limited thereto.

[0096] Subsequently, the processor can select one of the multiple emotional information provided through the user interface (UI) as the target emotional information (S520). For example, if user input is obtained from the user through one of the multiple emotional information, the processor can select one of the emotional information as the target emotional information.

[0097] FIG. 8 is a flowchart of a method for fine-tuning synthesized speech according to text-based feedback obtained through a user interface according to one embodiment of the present disclosure. The method (800) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of the method (800) may be divided and performed by at least one processor of the user terminal and at least one processor of the information processing system.

[0098] Referring to method (800), first, the processor can obtain text-based feedback for speech corresponding to at least a portion of the speech through a user interface (UI) (S810). For example, the processor can obtain feedback in the form of words, phrases, clauses, or sentences from the user, such as “at the end, with a angrier voice,” but is not limited thereto.

[0099] Subsequently, the processor can obtain modifications to the voice corresponding to at least a portion of the voice as it analyzes text-based feedback (S820).

[0100] For example, the processor can perform natural language analysis on text-based feedback to extract one or more keywords, and based on the extracted one or more keywords, extract modifications including target sections for fine-tuning (e.g., the end) and requested modifications (e.g., a more angry voice).

[0101] As another example, the processor can receive text-based feedback in natural language form, analyze the importance of each word and phrase in the feedback through an attention mechanism, and extract corrections based on this, including sections to be fine-tuned (e.g., the end) and correction requests (e.g., a more angry voice).

[0102] Thereafter, the processor can regenerate a voice corresponding to at least a portion of the speech based on the modifications obtained from the text-based feedback (S830). For example, if the processor wishes to modify a voice corresponding to at least a portion of the synthetic speech based on the modifications, the processor can modify an emotion embedding vector corresponding to the voice corresponding to at least a portion of the speech based on the modifications, and regenerate a voice corresponding to at least a portion of the speech using the modified emotion embedding vector.

[0103] FIG. 9 is a flowchart illustrating a method for fine-tuning a synthetic voice based on feedback obtained through an emotion regulation interface provided through a user interface according to an embodiment of the present disclosure, and FIG. 10 is an exemplary diagram illustrating an emotion regulation interface according to an embodiment of the present disclosure. The method (900) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of the method (900) may be performed separately by at least one processor of the user terminal and at least one processor of the information processing system.

[0104] In the method (900), first, the processor may provide an emotion control interface (1000) through a user interface (UI) (S910). Here, the emotion control interface (1000) may include a plurality of attribute setting elements (1010) that individually set each of a plurality of parameters related to the attributes of the emotion. For example, the plurality of attribute setting elements (1010) may include a plurality of attribute setting elements (1010) that set parameters related to the degree of positive-negative emotion, parameters related to the intensity or energy of the emotion, parameters related to the pitch of the tone, parameters related to the speed of speech, parameters related to intonation, etc.

[0105] Subsequently, the processor can obtain a request for modification of the voice corresponding to at least a portion of the voice as feedback through the emotion control interface (1000) (S920). Here, the modification request obtained through the emotion control interface (1000) may be a parameter adjustment value obtained through at least one of the multiple attribute setting elements (1010).

[0106] Subsequently, the processor can fine-tune the voice based on a modification request obtained through the emotion control interface (1000) (S930). For example, the processor can generate a second embedding vector corresponding to target emotion information by modifying a first embedding vector corresponding to initial emotion information based on parameter adjustment values ​​obtained through the emotion control interface (1000), and can regenerate voice corresponding to at least a portion of the synthesized voice using the second embedding vector.

[0107] FIG. 11 is a flowchart of a method for generating an emotion prediction model according to one embodiment of the present disclosure. The method (1100) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of the method (1100) may be divided and performed by at least one processor of the user terminal and at least one processor of the information processing system.

[0108] In method (1100), first, the processor can generate input data using multiple learning sentences written in text-based form and included in one document (S1110).

[0109] For example, the processor can extract sentence embedding vectors from each of multiple training sentences and extract speaker embedding vectors from information regarding the speaker who uttered the speech corresponding to each of the multiple training sentences, and can generate a combined embedding vector as input data by concatenating the sentence embedding vectors and the speaker embedding vectors. This combined embedding vector can be utilized as input data when generating a transformer-based sentiment prediction model.

[0110] As another example, the processor can tokenize multiple training sentences, insert a sentence-end indicator (e.g., a sentence end token) at the end of each tokenized training sentence, and then concatenate them to generate a single-sentence token sequence as input data. This single-sentence token sequence can be used as input data when creating a sentiment prediction model using the GPT2 architecture.

[0111] Subsequently, the processor can generate correct answer data using information regarding the emotion contained in the pitch corresponding to each of the multiple training sentences (S1120). For example, the processor can extract an emotion embedding vector as correct answer data from the information regarding the emotion contained in the pitch corresponding to each of the multiple training sentences.

[0112] Thereafter, the processor can create a hypothesis prediction model by learning using the above input data and the above correct answer data (S1130).

[0113] In one embodiment, the processor can create an emotion prediction model by learning to minimize the error between the output data and the correct data derived by inputting input data based on a regression-based loss function.

[0114] For example, the processor can set the emotion embedding vector generated as the correct answer data as the target embedding vector, and train based on the MSE loss function so that the error between the embedding vector derived by inputting input data and the target embedding vector is minimized.

[0115] Here, the processor can be set to not inflict loss on tokens (time-steps) without emotion when the emotion prediction model has the GPT2 structure, so that learning can be focused only on locations where emotion expression is required.

[0116] FIG. 12 is a flowchart of a method for generating an emotion search model according to one embodiment of the present disclosure. The method (1200) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of the method (1200) may be divided and performed by at least one processor of the user terminal and at least one processor of the information processing system.

[0117] In the method (1200), first, the processor can generate text embedding vectors by analyzing sentences using a text encoder (S1210). For example, the processor can extract multiple text embedding vectors corresponding to each of multiple training sentences written based on text using a text encoder.

[0118] In one embodiment, the processor may extract a text embedding vector for each of a plurality of training sentences by considering adjacent training sentences.

[0119] For example, the processor may use a text encoder to extract a first text embedding vector from a target training sentence among a plurality of training sentences written based on text, extract a second text embedding vector from training sentences before and after the target training sentence among the plurality of training sentences, and generate a target text embedding vector by reflecting the second text embedding vector on the first text embedding vector. For example, the processor may extract a text embedding vector that reflects information about surrounding sentences by reflecting the second text embedding vector on the first text embedding vector through cross-attention.

[0120] As another example, the processor can generate a target text embedding vector by using a text encoder to extract a first text embedding vector from a target training sentence among multiple training sentences written based on text, extract a second text embedding vector from training sentences before and after the target training sentence among multiple training sentences, extract a third text embedding vector from multiple training sentences, and reflect the second text embedding vector and the third text embedding vector into the first text embedding vector. For instance, the processor can extract a text embedding vector in which information regarding surrounding sentences and information regarding the entire document are appropriately reflected by reflecting the second text embedding vector and the third text embedding vector into the first text embedding vector through cross attention.

[0121] Thereafter, the processor may extract an audio embedding vector by analyzing the speech (840) using an audio encoder (S1220). For example, the processor may extract a target audio embedding vector from the speech corresponding to the target learning sentence using an audio encoder.

[0122] Subsequently, the processor can generate an emotion search model by training it using the target text embedding vector extracted through the text encoder and the target audio embedding vector extracted through the audio encoder as training data (S1230). For example, the processor can generate an emotion search model by training it such that the similarity between the target text embedding vector and the target audio embedding vector is maximized.

[0123] Figure 13 is a flowchart of a method for generating synthetic voice according to one embodiment of the present disclosure. Method (1300) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of method (1300) may be performed separately by at least one processor of the user terminal and at least one processor of the information processing system.

[0124] In method (1300), first, the processor can use speaker information and emotion information as conditions to generate speaker and emotion condition information in a combined form (S1310). Here, the speaker and emotion condition information may be an integrated embedding vector (global condition embedding) that simultaneously reflects information on two conditions: speaker and emotion.

[0125] In one embodiment, the processor can obtain speaker and emotion condition information as result data by inputting speaker information and emotion information into a pre-learned condition generation model.

[0126] Here, the initial emotional information may be multiple emotional information pieces derived corresponding to each of the multiple sentences by analyzing a single document containing multiple sentences using a pre-trained emotional prediction model or emotional retrieval model. For example, if the text for which a synthetic voice is to be generated is a sentence contained within a single script, the emotional information pieces may be multiple emotional information pieces derived by analyzing multiple sentences contained within the single script using a pre-trained emotional prediction model or emotional retrieval model. By utilizing the multiple emotional information pieces derived in this way to generate speaker and emotional condition information, a synthetic voice can be generated that reflects context-sensitive emotional information.

[0127] Subsequently, the processor can generate a synthesized speech containing the voice of a specific speaker speaking each of the multiple sentences with a specific emotion by converting multiple sentences into speech based on speaker and emotion condition information through a pre-trained text-to-speech conversion model (S1320).

[0128] In one embodiment, when a document is a script, the processor may obtain additional information related to the script and, based on the additional information, generate synthesized speech by converting a plurality of sentences into speech. Here, the additional information may include at least one of genre information of the script and situation and location information regarding a scene corresponding to the script, and by further utilizing such additional information in generating synthesized speech, more sophisticated synthesized speech can be generated.

[0129] FIG. 14 is a flowchart of a method for generating speaker and emotion condition information according to one embodiment of the present disclosure. The method (1400) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of the method (1400) may be divided and performed by at least one processor of the user terminal and at least one processor of the information processing system.

[0130] In the method (1000), first, the processor can generate a condition generation model based on a machine learning-based learning method (S1410).

[0131] In one embodiment, the processor can generate a condition generation model that receives speaker and emotion information as input data and derives speaker and emotion condition information by learning using learning data that inputs information related to the speaker and information about the emotion and the corresponding speaker and emotion condition information as correct answer data.

[0132] Afterwards, the processor can obtain speaker information (S1420).

[0133] In one embodiment, speaker information may be information obtained from speech uttered by a specific speaker. For example, speaker information may include at least one of an embedding vector extracted from speech uttered by the speaker, a speaker identification index (label), or the speaker's speech signal (waveform). In particular, the speaker embedding vector is extracted through a pretrained model such as Pyannote, ECAPA-TDNN, or TitaNet, and may refer to a high-dimensional vector that concisely represents the speech characteristics of a specific speaker. Additionally, speaker information may be a speaker speech sample (reference speech) containing the voice of a specific speaker. That is, speaker information may be not only a fixed vector but also the raw speech signal itself.

[0134] In one embodiment, the processor may obtain at least one of a speaker embedding vector and a speaker reference speech sample from speech uttered by a specific speaker as speaker information.

[0135] Thereafter, the processor can obtain emotional information corresponding to the emotional voice (S1430).

[0136] In one embodiment, the emotion information may be information obtained through speech or emotion labels reflecting a specific emotion. For example, the emotion information may include at least one of an emotion embedding vector, an emotion identification index (label), or a speech signal (waveform) containing an emotion. In particular, the emotion embedding vector is a vector representing an emotional state (e.g., happiness, anger, sadness, etc.) and may be extracted by passing an emotion prediction model (emotion search model) or an emotion index through an embedding layer, but is not limited thereto. Additionally, the emotion information may be an emotion speech sample (reference speech) in which a specific emotion is expressed. That is, the emotion information may be not only a fixed vector but also the raw speech signal itself.

[0137] In one embodiment, the processor may obtain at least one of an emotion embedding vector and an emotion reference speech sample from a speech containing a specific emotion as emotion information.

[0138] Thereafter, the processor can generate speaker and emotion condition information (global condition embedding) as result data by inputting speaker information and emotion information as input data into a condition generation model (S1440).

[0139] Figure 15 is a flowchart of a method for generating a condition generation model according to one embodiment of the present disclosure. Method (1500) may be performed by at least one processor of a user terminal or an information processing system. Alternatively, the steps of method (1500) may be performed separately by at least one processor of the user terminal and at least one processor of the information processing system.

[0140] In method (1500), first, the processor may generate first input data and second input data using information about the speaker and emotion (S1510). Here, the information about the speaker used to generate the first input data and the second input data may include, but is not limited to, a voice signal corresponding to a voice spoken by the speaker, a voice embedding vector extracted from the voice spoken by the speaker, and a speaker index (or speaker label) indicating the speaker.

[0141] Additionally, here, the information about the emotion used to generate the first input data and the second input data may include, but is not limited to, a voice signal corresponding to a short emotional voice, a voice embedding vector extracted from the emotional voice, and an emotion index (or emotion label) indicating the emotion.

[0142] That is, the processor can generate first input data and second input data by combining information about the speaker and information about emotions as described above.

[0143] For example, the processor may generate first input data using a voice signal corresponding to a voice spoken by a first speaker with a first emotion, and may generate second input data using a voice signal corresponding to a voice spoken by a second speaker with a second emotion.

[0144] As another example, the processor may generate first input data using a first speech embedding vector extracted from speech uttered by a first speaker with a second emotion, and may generate second input data using a second speech embedding vector generated by passing an emotion index indicating the first emotion through an embedding layer.

[0145] As another example, the processor may generate a first voice uttered by a first speaker with a second emotion by modulating a voice uttered by a first speaker with a first emotion, and generate first input data using the first voice, and may generate a second voice uttered by a second speaker with a first emotion by modulating a voice uttered by the first speaker with the first emotion, and generate second input data using the second voice.

[0146] As another example, the processor may generate first input data by removing information about a first emotion from a first embedding vector extracted from a first voice, and may generate second input data by removing information about the first emotion from a second embedding vector extracted from a second voice.

[0147] Thereafter, the processor can generate correct answer data based on speaker and emotional condition information (S1520). For example, the processor can generate an embedding vector extracted from a speech uttered by a specific speaker with a specific emotion as correct answer data.

[0148] Thereafter, the processor can generate a condition generation model using the first input data, the second input data, and the correct answer data (S1530). For example, the processor can train the model to minimize the error between the result data and the correct answer data derived by inputting the first and second input data.

[0149] For example, if the condition generation model is a multi-layer perceptron (MLP) model, the processor can be trained to minimize the error between the result data and the correct answer data derived by inputting the first input data and the second input data based on a regression-based loss function (e.g., MSE loss, cosine-distance loss).

[0150] As another example, if the condition generation model is a transformer-based diffusion model, the processor can be trained to minimize the error between the result data and the correct data derived by inputting the first input data and the second input data based on a general loss function (e.g., MSE loss, etc.) used in the diffusion model.

[0151] The above-described method may be provided as a computer program stored on a computer-readable recording medium for execution on a computer. The medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording means or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program instructions, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0152] The methods, operations, or techniques of the present disclosure may be implemented by various means. For example, these techniques may be implemented in hardware, firmware, software, or a combination thereof. Those skilled in the art will appreciate that the various exemplary logical blocks, modules, circuits, and algorithm steps described in connection with the disclosure herein may be implemented as electronic hardware, computer software, or a combination of both. To clearly illustrate this interchangeability of hardware and software, various exemplary components, blocks, modules, circuits, and steps have been described above generally in terms of their functionality. Whether such functionality is implemented as hardware or software will depend on the particular application and the design requirements imposed on the overall system. Those skilled in the art may implement the described functionality in various ways for each particular application, but such implementations should not be construed as departing from the scope of the present disclosure.

[0153] In a hardware implementation, the processing units used to perform the techniques may be implemented within one or more ASICs, DSPs, digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), processors, controllers, microcontrollers, microprocessors, electronic devices, other electronic units designed to perform the functions described herein, a computer, or a combination thereof.

[0154] Accordingly, the various exemplary logical blocks, modules, and circuits described in connection with the present disclosure may be implemented or performed by any combination of a general-purpose processor, a DSP, an ASIC, an FPGA or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or those designed to perform the functions described herein. A general-purpose processor may be a microprocessor, but in the alternative, the processor may be any conventional processor, controller, microcontroller, or state machine. A processor may also be implemented as a combination of computing devices, e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in conjunction with a DSP core, or any other such configuration.

[0155] In a firmware and / or software implementation, the techniques may be implemented as instructions stored on a computer-readable medium, such as random access memory (RAM), read-only memory (ROM), non-volatile random access memory (NVRAM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable PROM (EEPROM), flash memory, a compact disc (CD), a magnetic or optical data storage device, etc. The instructions may be executable by one or more processors and may cause the processor(s) to perform certain aspects of the functionality described herein.

[0156] When implemented in software, the techniques may be stored on or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media includes both computer storage media and communication media, including any medium that facilitates transfer of a computer program from one place to another. Storage media may be any available media that can be accessed by a computer. By way of example, and not limitation, such computer-readable media can include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and that can be accessed by a computer. Also, any connection is properly termed a computer-readable medium.

[0157] For example, if the software is transmitted from a website, server, or other remote source using coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technologies such as infrared, radio, and microwave, then the coaxial cable, fiber optic cable, twisted pair, digital subscriber line, or wireless technologies such as infrared, radio, and microwave are included within the definition of media. Disk and disc, as used herein, includes compact discs, laser discs, optical discs, digital versatile discs (DVDs), floppy disks, and Blu-ray discs, where disks usually reproduce data magnetically, whereas discs reproduce data optically using lasers. Combinations of the above should also be included within the scope of computer-readable media.

[0158] The software module may reside in RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, registers, a hard disk, a removable disk, a CD-ROM, or any other known form of storage medium. An exemplary storage medium may be connected to a processor so that the processor can read information from the storage medium or write information to the storage medium. Alternatively, the storage medium may be integrated into the processor. The processor and the storage medium may exist within an ASIC. The ASIC may exist within a user terminal. Alternatively, the processor and the storage medium may exist as separate components within the user terminal.

[0159] While the embodiments described above have been described as utilizing aspects of the presently disclosed subject matter in one or more standalone computer systems, the present disclosure is not limited thereto and may be implemented in conjunction with any computing environment, such as a network or distributed computing environment. Furthermore, aspects of the present disclosure may be implemented in multiple processing chips or devices, and storage may be similarly affected across multiple devices. Such devices may include personal computers, network servers, and portable devices.

[0160] Although the present disclosure has been described in relation to some embodiments, various modifications and changes may be made without departing from the scope of the present disclosure as understood by a person skilled in the art to which the invention of the present disclosure pertains. Furthermore, such modifications and changes should be considered to fall within the scope of the claims appended to this specification.

Claims

1. A method for fine-tuning text-based speech through a user interface performed by at least one processor, Steps for acquiring text and voice; A step of obtaining feedback on the acquired voice through a user interface (UI); and A step of fine-tuning the acquired voice based on the acquired feedback. A method for fine-tuning text-based speech through a user interface, comprising:

2. In paragraph 1, The steps for obtaining the above feedback are: As feedback, a step of obtaining an emotion information modification request that requests that initial emotion information corresponding to a voice corresponding to at least a portion of the acquired voice be modified to target emotion information. A method for fine-tuning text-based speech through a user interface, comprising:

3. In paragraph 2, The step of obtaining the above emotional information modification request is: A step of providing multiple emotional information corresponding to different types of emotions through the user interface; and A step of selecting one of the multiple emotional information provided above as target emotional information. A method for fine-tuning text-based speech through a user interface, comprising:

4. In paragraph 3, The step of providing the above multiple emotional information is: A step of visualizing the plurality of emotional information and providing the visualized plurality of emotional information through the user interface. Includes, The above visualized multiple emotional information is, A method for fine-tuning text-based speech through a user interface, wherein the result is a two-dimensional plane or three-dimensional space in which a plurality of points corresponding to each of the plurality of emotion embedding vectors are displayed, by reducing a plurality of emotion embedding vectors extracted from each of the plurality of emotion information into two or three dimensions through a dimensionality reduction algorithm.

5. In paragraph 3, The step of fine-tuning the acquired voice is as follows: A step of regenerating a voice corresponding to at least a portion of the acquired voice using an emotion embedding vector corresponding to one of the emotion information, if an emotion embedding vector corresponding to one of the emotion information exists among a plurality of previously generated emotion embedding vectors. A method for fine-tuning text-based speech through a user interface, comprising:

6. In Paragraph 3, The step of providing the above multiple emotional information is: A step of generating a plurality of example voices corresponding to each of the plurality of emotional information by using speaker information corresponding to the acquired voice and each of the plurality of emotional information; and When specific emotional information is selected from the plurality of emotional information from the user, a step of providing an example voice corresponding to the specific emotional information from among the plurality of generated example voices A method for fine-tuning text-based speech through a user interface, comprising:

7. In Paragraph 3, The step of providing the above multiple emotional information is: When one or more search words related to emotions are input from a user through the user interface, a step of searching for emotional information corresponding to one or more of the input search words among the plurality of emotional information; and Step of providing the searched emotion information through the above user interface A method for fine-tuning text-based speech through a user interface, comprising:

8. In Paragraph 2, The step of fine-tuning the acquired voice above is, A step of calculating the difference between a first embedding vector corresponding to the initial emotion information and a second embedding vector corresponding to the target emotion information; and A step of regenerating speech corresponding to at least a portion of the acquired speech based on the difference calculated above and a pre-set scale factor. A method for fine-tuning text-based speech through a user interface, comprising:

9. In paragraph 1, The steps for obtaining the above feedback are: As feedback, a step of obtaining a speech style modification request that requests that an initial speech style corresponding to a speech corresponding to at least a portion of the acquired speech be modified to a target speech style. A method for fine-tuning text-based speech through a user interface, comprising:

10. In paragraph 1, The steps for obtaining the above feedback are: A step of obtaining text-based feedback for a voice corresponding to at least a portion of the acquired voice through the user interface. Includes, The step of fine-tuning the acquired voice is as follows: A step of obtaining corrections to the voice corresponding to at least a portion of the obtained voice by analyzing the obtained text-based feedback; and A step of regenerating a voice corresponding to at least a portion of the acquired voice based on the acquired modification content. A method for fine-tuning text-based speech through a user interface, comprising:

11. In paragraph 1, The steps for obtaining the above feedback are: A step of providing an emotion regulation interface through the above user interface; and A step of obtaining a request for modification of a voice corresponding to at least a portion of the acquired voice as feedback through the emotion regulation interface provided above, The emotional regulation interface provided above is: Contains multiple attribute setting elements that individually set each of multiple parameters related to the attribute of the emotion, The above obtained modification request is, A method for fine-tuning text-based speech through a user interface, wherein the parameter adjustment value is obtained through at least one attribute setting element among the above multiple attribute setting elements.

12. In paragraph 1, The steps of acquiring the above voice are: A step of generating speaker and emotion condition information based on speaker information of a specific speaker and initial emotion information for a specific emotion; and A step of generating a synthetic voice that includes a voice spoken by a specific speaker with a specific emotion by converting multiple sentences into voice based on the generated speaker and emotion condition information through a pre-trained text-to-speech conversion model. Includes, The above initial emotional information is, A text-based speech fine-tuning method through a user interface, wherein multiple emotion information is derived corresponding to each of the multiple sentences by analyzing a single document containing the multiple sentences through a previously learned emotion prediction model or emotion search model.

13. In paragraph 12, The step of generating a synthesized voice including a voice uttered by the aforementioned specific speaker with a specific emotion is: If the above document is a script, a step of obtaining additional information including at least one of genre information of the script, situation and location information for a scene corresponding to the script; and A step of generating synthesized speech by converting the plurality of sentences into speech based on the speaker and emotion condition information generated above and the additional information obtained above. A method for fine-tuning text-based speech through a user interface, comprising:

14. A computer program stored on a computer-readable recording medium for executing a method according to any one of paragraphs 1 through 13 on a computer.

15. As an information processing system, memory; and At least one processor connected to the memory and configured to execute at least one computer-readable program contained in the memory. Including, At least one program above, Acquire text and voice, Feedback on the acquired voice is obtained through a user interface (UI), and An information processing system comprising commands for fine-tuning the acquired voice based on the acquired feedback.

Citation Information

Patent Citations

  • Text voice converting device

    JP1998011083A

  • Speech synthesizer, speech synthesizing method, and speech synthesizing program

    JP2012108378A

  • Method and system of synthesizing emotional speech based on personal prosody model and recording medium

    KR1020120117041A

  • Composite panel with improved vibration characteristics and manufacturing method therefor

    KR1020230126226A

  • Plasma generating device of both sides discharge and system using it

    KR102279607B1