Dubbing generation system and method to which voice technology based on linguistics and cognitive science is applied using artificial intelligence
The dubbing generation system leverages AI and voice technology to convert audio signals into text, enabling the creation of culturally and linguistically nuanced dubbing that sounds natural and reduces production time.
Patent Information
- Application Number
- PCT/KR2024/017832
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-24
- Filing Date
- 2024-11-12
- Publication Date
- 2025-05-30
AI Technical Summary
Current dubbing creation methods are time-consuming and struggle to convey the linguistic and cultural nuances of the original content, often resulting in unnatural-sounding dubbing.
A dubbing generation system and method utilizing artificial intelligence and voice technology, which includes a server and terminal setup, uses an AI model to convert audio signals into text composed of international phonetic symbols, allowing for the generation of dubbing information that reflects regional, social, cultural, and linguistic differences.
The system enables the rapid generation of natural-sounding dubbing that mimics the pronunciation of a native speaker, effectively addressing the challenges of conveying cultural and linguistic nuances while reducing production time.
Smart Images

Figure KR2024017832_30052025_PF_FP_ABST
Abstract
Description
A dubbing generation system and method using artificial intelligence and voice technology based on linguistics and cognitive science
[0001] The present disclosure relates to a dubbing generation system and method using artificial intelligence-based voice technology.
[0002] The video streaming market, related to the media content market, is growing rapidly. Furthermore, as the video streaming market expands beyond the domestic market to the global market, the need for dubbing generation technology is increasing.
[0003] Specifically, in the growing media content market, individual creators are providing high-quality videos through dubbing to enhance their competitiveness. Consequently, the dubbing creation service market is expanding.
[0004] Currently, dubbing is created by recording the results of translating the original language, so not only does it take a long time to produce, but the dubbing's ability to convey meaning can vary depending on the producer's level of linguistic and cultural understanding.
[0005] Additionally, in the case of dubbing production, there is a problem that it is difficult to convey the language and emotions of a native speaker, which may cause the video itself to become unnatural.
[0006] Accordingly, there is a need for technology that can create dubbing that reflects the linguistic and cultural differences of the country where the original work was produced, while also being able to insert it into a video without any unnaturalness.
[0007] The purpose of this disclosure is to provide a system and method that can generate natural dubbing that reflects language-specific characteristics, cultural differences between countries, and the speaker's emotions through artificial intelligence.
[0008] In order to achieve the above-described purpose, a dubbing generation system according to the present disclosure can be provided, which includes a server and a terminal, wherein the server includes a communication unit configured to receive image information from the terminal, a processor including an artificial intelligence model that converts an audio signal included in the image information into text and generates dubbing information based on the converted text, and wherein the text is composed of international phonetic symbols.
[0009] In one embodiment, the processor may generate dubbing information for the voice information based on the translated text and the translation result for the voice signal.
[0010] In one embodiment, the processor may convert the translation result into information composed of international phonetic symbols based on the converted text, and generate dubbing information based on the converted information.
[0011] In one embodiment, the processor can remove the audio signal from the image information and combine the dubbing information.
[0012] In one embodiment, the artificial intelligence model may be a transformer.
[0013] In addition, a dubbing generation method of a dubbing generation system including a server and a terminal according to the present disclosure includes a step in which the server receives video information from the terminal; a step in which the server converts a voice signal included in the video information into text; and a step in which the server generates dubbing information based on the converted text, wherein the text may be composed of international phonetic symbols.
[0014] According to the present disclosure, dubbing can be generated that reflects regional, social, cultural, and linguistic differences. Accordingly, the dubbing generation system according to the present disclosure can generate natural dubbing that sounds like it was spoken by a native speaker.
[0015] Additionally, according to the present disclosure, it becomes possible to produce dubbing at a high speed without requiring the speaker to record his / her voice for dubbing production.
[0016] Figure 1 is an overall system diagram of the present disclosure.
[0017] FIG. 2 is a block diagram of a server included in the dubbing generation system of the present disclosure.
[0018] Figure 3 is a block diagram of a terminal included in the dubbing generation system of the present disclosure.
[0019] FIG. 4 is a block diagram of a processor included in the dubbing generation system of the present disclosure.
[0020] Figure 5 is a conceptual diagram showing a subtitle and dubbing service.
[0021] FIG. 6 is a flowchart of a method for generating dubbing using international phonetic symbols according to the present disclosure.
[0022] FIG. 7 is a conceptual diagram illustrating an embodiment of a dubbing generation system according to the present disclosure that converts a voice signal into text composed of international phonetic symbols.
[0023] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and any content that is common in the technical field to which this disclosure pertains or that overlaps between embodiments is omitted. The terms "part, module, element, block" used in the specification may be implemented in software or hardware, and depending on the embodiments, multiple "parts, modules, elements, blocks" may be implemented as a single component, or a single "part, module, element, block" may include multiple components.
[0024] Throughout the specification, when a part is said to be "connected" to another part, this includes not only direct connection but also indirect connection, and indirect connection includes connection via a wireless communication network.
[0025] Additionally, when a part is said to "include" a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise specifically stated.
[0026] Throughout the specification, when we say that an element is "on" another element, this includes not only cases where the element is in contact with the other element, but also cases where another element exists between the two elements.
[0027] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.
[0028] Singular expressions include plural expressions unless the context clearly indicates otherwise.
[0029] The identification codes for each step are used for convenience of explanation and do not describe the order of each step. Each step may be performed in a different order than specified unless the context clearly indicates a specific order.
[0030] The operating principle and embodiments of the present disclosure are described below with reference to the attached drawings.
[0031] As used herein, the term "system according to the present disclosure" encompasses various devices capable of performing computational processing and providing results to a user. For example, the system according to the present disclosure may include a computer, a server device, and a portable terminal, or may be any one of them.
[0032] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.
[0033] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
[0034] The above portable terminal may include, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, a smart phone, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted device (HMD).
[0035] Below, a dubbing generation system according to the present disclosure is described.
[0036] Referring to FIG. 1, the dubbing generation system according to the present disclosure may include at least one of a server (10) and a terminal (20). Specifically, the dubbing generation system according to the present disclosure may be implemented solely by the server (10) or the terminal (20), or may be implemented in the form of a system including at least one of the server (10) and the terminal (20). The description of the dubbing generation system described below may be applied to both cases where the dubbing generation system according to the present disclosure is implemented solely by the server (10) or the terminal (20), or may be implemented in the form of a system including at least one of the server (10) and the terminal (20).
[0037] The server (10) is connected to at least one terminal (20) via a network, transmits information to each of a plurality of terminals, and generates data necessary for dubbing generation learning based on information received from at least one of the terminals (20).
[0038] Meanwhile, it is obvious to those skilled in the art that the terminal (20) is not limited to the above-described portable terminal, and may include a processor-equipped notebook, desktop, laptop, tablet PC, slate PC, etc.
[0039] Below, each of the server (10) and the terminal (20) for implementing the dubbing generation system according to the present disclosure will be described.
[0040] FIG. 2 is a block diagram of a server included in the dubbing generation system of the present disclosure.
[0041] A server (100) according to the present disclosure may include at least one of a communication unit (110), a storage unit (120), and a processor (130).
[0042] The communication unit (110) can communicate with at least one of a terminal, an external storage (e.g., a database (140)), an external server, and a cloud server.
[0043] Meanwhile, an external server or cloud server may be configured to perform at least a portion of the role of the processor (130). That is, data processing or data operations, etc. may be performed on an external server or cloud server, and the present invention does not impose any particular limitations on this method.
[0044] Meanwhile, the communication unit (110) can support various communication methods according to the communication standards of the communication target (e.g., electronic device, external server, device, etc.).
[0045] For example, the communication unit (110) may be configured to communicate with a communication target using at least one of WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.
[0046] Next, the storage unit (120) may be configured to store various information related to the present invention. In the present invention, the storage unit (120) may be provided in the device according to the present invention itself. Alternatively, at least a portion of the storage unit (120) may refer to at least one of a database (DB, 140) and a cloud storage (or cloud server). That is, the storage unit (120) may be sufficient as long as it is a space where information required for the device and method according to the present invention is stored, and it can be understood that there are no restrictions on physical space. Accordingly, hereinafter, the storage unit (120), the database (140), the external storage, and the cloud storage (or cloud server) will not be separately distinguished, and will all be referred to as the storage unit (120).
[0047] Next, the processor (130) may be configured to control the overall operation of the device related to the present invention. The processor (130) may process signals, data, information, etc. input or output through the components discussed above, or provide or process appropriate information or functions to the user.
[0048] The processor (130) includes at least one CPU (Central Processing Unit) and can perform functions according to the present invention.
[0049] At least one component may be added or deleted to correspond to the performance of the components illustrated in FIG. 2. Furthermore, it will be readily apparent to those skilled in the art that the relative positions of the components may be altered to correspond to the performance or structure of the device.
[0050] Hereinafter, a terminal included in the dubbing generation system of the present disclosure will be described in detail.
[0051] Figure 3 is a block diagram of a terminal included in the dubbing generation system of the present disclosure.
[0052] Referring to FIG. 3, a terminal (200) according to the present disclosure may include a communication unit (210), an input unit (220), a display unit (230), a processor (240), etc. The components illustrated in FIG. 3 are not essential for implementing a dubbing generation system according to the present disclosure, and thus, the terminal described in this specification may have more or fewer components than the components listed above.
[0053] Among the above components, the communication unit (210) may include one or more components that enable communication with an external device, and may include, for example, at least one of a broadcast reception module, a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
[0054] The wired communication module may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), RS-1302 (recommended standard 1302), power line communication, or plain old telephone service (POTS).
[0055] The wireless communication module may include a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G, in addition to a WiFi module and a Wireless Broadband module.
[0056] The input unit (220) is for inputting video information (or signals), audio information (or signals), data, or information input from a user, and may include at least one camera, at least one microphone, and at least one user input unit. Voice data or image data collected by the input unit may be analyzed and processed into a user control command.
[0057] The display unit (230) is intended to generate output related to visual, auditory, or tactile sensations, and may include at least one of a display unit, an audio output unit, a haptic module, and an optical output unit. The display unit may be formed as a touch screen by forming a mutual layer structure with a touch sensor or by forming an integral structure with the touch sensor. Such a touch screen may function as a user input unit that provides an input interface between the device and a user, and at the same time, provide an output interface between the device and the user.
[0058] The display unit displays (outputs) information processed by this device. For example, the display unit may display execution screen information of an application program (e.g., an application) running on this device, or UI (User Interface) or GUI (Graphical User Interface) information based on such execution screen information.
[0059] In addition to the above-described components, the above-described terminal may further include an interface unit and a memory.
[0060] The interface unit serves as a passageway for various types of external devices connected to the device. The interface unit may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module (SIM), an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. The device may perform appropriate control related to the external device connected to the interface unit.
[0061] The memory can store data supporting various functions of the device, programs for the operation of the processor, input / output data (e.g., music files, still images, moving images, etc.), and a plurality of application programs (or applications) running on the device, data for the operation of the device, and commands. At least some of these application programs can be downloaded from an external server via wireless communication.
[0062] Such memory may include at least one type of storage medium among flash memory type, hard disk type, SSD (Solid State Disk type), SDD (Silicon Disk Drive type), multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. In addition, the memory may be a database that is separate from the device but connected by wire or wirelessly.
[0063] Meanwhile, the terminal described above includes a processor (240). The processor may be implemented as a memory storing data regarding an algorithm for controlling the operation of components within the device or a program reproducing the algorithm, and at least one processor (not shown) that performs the aforementioned operations using the data stored in the memory. In this case, the memory and the processor may be implemented as separate chips. Alternatively, the memory and the processor may be implemented as a single chip.
[0064] Meanwhile, the processor may control any one or a combination of the components described above to implement various embodiments of the present disclosure described in the drawings below on the device.
[0065] Meanwhile, at least one component may be added or deleted in accordance with the performance of the components illustrated in Figures 1 to 3. Furthermore, it will be readily apparent to those skilled in the art that the relative positions of the components may be altered in accordance with the performance or structure of the device.
[0066] Meanwhile, as illustrated in FIG. 4, a processor included in at least one of the server and the terminal may include multiple modules for implementing the dubbing generation system described below. Specifically, the processor (300) may include a voice recognition module (310) and an artificial intelligence module (320). While the dubbing generation method described below is described as being implemented by the operations of the modules, the performance of each step described below need not necessarily be performed by the modules.
[0067] Below, the artificial intelligence described in the present invention is described in detail.
[0068] The artificial intelligence-related functions according to the present disclosure are operated through the processor and memory installed in the above-described server and terminal. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU or a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. One or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in the memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0069] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is learned by a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed in the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0070] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0071] According to an exemplary embodiment of the present disclosure, a processor can implement artificial intelligence. Artificial intelligence refers to a machine learning method based on an artificial neural network that imitates human neurons (biological neurons) to enable machines to learn. Artificial intelligence methodologies can be categorized into supervised learning, in which input data and output data are provided together as training data depending on the learning method, so that the solution (output data) to the problem (input data) is determined; unsupervised learning, in which only input data is provided without output data, so that the solution (output data) to the problem (input data) is not determined; and reinforcement learning, in which a reward (Reward) is provided from an external environment whenever an action (Action) is taken in the current state (State), and learning is performed in a direction to maximize this reward. In addition, artificial intelligence methodologies can be categorized according to the architecture of the learning model. The architectures of widely used deep learning technologies can be categorized into convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformers, and generative adversarial networks (GANs).
[0072] The present device and system may include an artificial intelligence model. The artificial intelligence model may be a single artificial intelligence model or may be implemented as multiple artificial intelligence models. The artificial intelligence model may be composed of a neural network (or artificial neural network) and may include statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network may refer to a model in general that has problem-solving capabilities by changing the binding strength of synapses through learning, formed by artificial neurons (nodes) that form a network by combining synapses. The neurons of the neural network may include a combination of weights or biases. The neural network may include one or more layers composed of one or more neurons or nodes. For example, the device may include an input layer, a hidden layer, and an output layer. The neural network constituting the device can infer a desired result (output) from an arbitrary input (input) by changing the weights of neurons through learning.
[0073] The processor can create a neural network, train (or learn) a neural network, perform a calculation based on received input data, generate an information signal based on the calculation result, or retrain the neural network. The models of the neural network can include various types of models such as CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restrcted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Classification Network, etc., but are not limited thereto. The processor can include one or more processors for performing calculations according to the models of the neural network. For example, the neural network can be a deep neural network. It may include a deep neural network.
[0074] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Network), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Depp Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), Generative Adversarial Network (GAN), Liquid State Machine (LSM), Extreme Learning Machine (ELM), It will be understood by those skilled in the art that any neural network may be included, including but not limited to ESN (Echo State Network), DRN (Deep Residual Network), DNC (Differentiable Neural Computer), NTM (Neural Turning Machine), CN (Capsule Network), KN (Kohonen Network), and AN (Attention Network).
[0075] According to an exemplary embodiment of the present disclosure, the processor may be configured to perform a process for generating a CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, Region with Convolution Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restrcted Boltzman Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA for natural language processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics for vision processing, Visual Understanding, Video Synthesis, ResNet for data intelligence, Anomaly Detection, Prediction, Time-Series Forecasting, Various artificial intelligence structures and algorithms, including optimization, recommendation, and data creation, can be utilized, but are not limited thereto. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.
[0076] Below, a subtitle and dubbing provision service provided by a dubbing generation system and method according to the present disclosure is described.
[0077] Figure 5 is a conceptual diagram showing a subtitle and dubbing service.
[0078] Referring to FIG. 5, a step of receiving audio information is performed. Here, the step of receiving audio information may be a step of receiving video information. The Jamik dubbing generation system according to the present disclosure receives video information and utilizes the audio information contained in the video information for dubbing generation. Alternatively, the step of receiving audio information may be a step of receiving audio information itself, rather than video information.
[0079] In this specification, audio information may include all data related to sound included in an image. In one embodiment, audio information may include audio signals and background sounds included in the image information.
[0080] The processor receives image information and processes audio information included in the image information.
[0081] Next, a step is performed to separate the voice signal and background sound information from the voice information.
[0082] The processor can separate the voice signal and background sound information from the voice information and create separate files.
[0083] The audio signal is the speaker's voice included in the video information and is the information that is the target of dubbing. Later, the audio signal separated from the audio information is generated as at least one dubbing signal.
[0084] The separation of the above speech signal and background sound information can be performed through a separate artificial intelligence model trained to separate the speech signal and background sound information from the speech information.
[0085] Next, the step of converting the separated voice signal into text is performed.
[0086] The above voice signal can be converted into text through the voice recognition module (310). Here, the text can be formed in the language that is the basis for producing the video from which the voice signal is separated.
[0087] In this specification, the language that is the basis of video production is called the “source language,” and the language in which dubbing is to be created is called the “target language.”
[0088] Next, the step of translating the converted text composed in the original language into the target language is performed.
[0089] The above translation result is in the form of text in the target language and can be provided as subtitles for the original video.
[0090] The type of the above target language can be input by the video producer or the user viewing the video.
[0091] In one embodiment, a video producer can generate subtitles for at least one target language from the video production stage by specifying the target language in advance.
[0092] In one embodiment, a video producer may generate subtitles by allowing viewers to specify a target language of their choice, rather than creating subtitles themselves. Subtitles created by a specific viewer can then be provided to viewers who request subtitles in that target language.
[0093] The above translation results can then be used to create dubbing, as described later.
[0094] Meanwhile, the above translation can be performed using a separate AI model trained to translate text in the source language into the target language. The type of AI model used for translation is not specifically limited.
[0095] Meanwhile, the processor generates the subtitles and then synthesizes them into the original video. At this time, the processor can set a time interval for inserting the subtitles by considering the time interval in which the audio signal is generated within the video information. Specifically, the processor can match the time interval in which the audio signal is generated within the video information at the syllable, word, or sentence level. The time information matched in this manner is also matched to the converted text after the audio signal is converted into text. Thereafter, the time information can also be matched to the result of translating the converted text. The processor can insert the subtitles into the video information based on the time information.
[0096] Next, a step is performed to generate a speech signal (hereinafter, dubbing information) in a target language based on at least one of the separated speech signal, the converted text, and the translation result.
[0097] The dubbing generation method according to the present disclosure can be implemented in the step of generating the above-described voice signal.
[0098] Hereinafter, with reference to FIGS. 6 and 7, a method for generating dubbing information in a target language will be described in detail.
[0099] Referring to Fig. 6, a step is performed in which an artificial intelligence module (320) receives a voice signal (S210).
[0100] Here, the information input to the artificial intelligence module (320) may be audio information included in the image information, or an audio signal in which background sound information is separated from the audio information.
[0101] Next, a step of converting the separated voice signal into text is performed (S221).
[0102] The above voice signal can be converted into text through the voice recognition module (310). Here, the text can be in the original language.
[0103] Next, a step is performed to translate the converted text into a target language (S222).
[0104] The above converted text composed of the original language is converted into a text composed of the target language through translation.
[0105] The above translation may be performed by a known artificial intelligence model or by a separately trained artificial intelligence model.
[0106] The above steps S221 and S222 can be omitted by utilizing the translation results described in the subtitle and dubbing provision service described in FIG. 5.
[0107] Separately from the above text translation, the next step is to have the artificial intelligence model convert the voice signal into text composed of the International Phonetic Alphabet (IPA) (S223).
[0108] The above artificial intelligence module (320) may include an artificial intelligence model trained to convert voice information into the International Phonetic Association (IPA).
[0109] In one embodiment, the artificial intelligence model may be, but is not limited to, a transformer.
[0110] In one embodiment, referring to FIG. 7, the artificial intelligence model can receive voice information such as “Could you lend me ten thousand won?” and convert it into the following international phonetic symbols.
[0111]
[0112] Specifically, the artificial intelligence model can be trained such that 'src' is a speech signal composed of Korean, 'tgt' is an English IPA embedding, and 'output' is an English IPA embedding.
[0113] Thereafter, the artificial intelligence model generates dubbing information based on the translation result and text composed of international phonetic symbols (S230).
[0114] Specifically, the artificial intelligence model can generate dubbing information for the voice information based on the translated text and the translation result for the voice signal.
[0115] Based on the converted text, the AI model converts the translation result into information structured in the International Phonetic Alphabet (IPA). To this end, the AI model can be trained to receive text structured in the IPA, generated from a speech signal in the source language, and the translation result of the speech signal in the source language. The model then generates text structured in the IPA and translated into the target language. Utilizing the IPA during the training process allows for the generation of dubbing information that reflects regional, social, cultural, and linguistic differences between the source and target languages.
[0116] Additionally, as described above, when utilizing the International Phonetic Alphabet when training an artificial intelligence model to generate dubbing information, it becomes possible to generate dubbing information that pronounces the target language naturally.
[0117] Again, referring to FIG. 5, a step of synthesizing the generated voice signal (dubbing information) and the separated background sound information is performed.
[0118] The processor can use the translation result to generate a subtitle for the image information, and use the subtitle generation result to synthesize the generated voice signal and the separated background sound information.
[0119] Specifically, the processor can search for time information at which the subtitle is output from the image information, and synthesize the generated voice signal and the separated background sound information by utilizing the searched time information.
[0120] The processor can match the time intervals in which audio signals are generated within the video information, at the syllable, word, or sentence level. The time information matched in this manner is then matched to the converted text after the audio signal is converted into text. The time information can then be matched to the translated text. Based on this time information, the processor can insert subtitles into the video information. Furthermore, the processor can utilize the aforementioned time information to generate dubbing.
[0121] Meanwhile, the processor may utilize the result of separating the voice signal and the background sound information from the voice information included in the image information to synthesize the generated voice signal and the separated background sound information. Specifically, the processor may utilize the time information at which the voice signal was separated from the voice information to synthesize the generated voice signal and the separated background sound information.
[0122] When separating a voice signal from background sound information, the processor can match the background sound information with information defining the time intervals from which the voice signal was separated. This time interval information can be defined at the syllable, word, or sentence level and matched to the background sound information. This time information can be utilized when synthesizing dubbing information with the background sound information.
[0123] As described above, according to the dubbing generation system according to the present disclosure, since background sound separated from voice information is synthesized after dubbing is generated, dubbing can be generated without loss of sound source.
[0124] In addition, according to the present disclosure, dubbing can be generated at a faster speed compared to conventional methods of manually generating dubbing.
[0125] Additionally, according to the present disclosure, it becomes possible to create dubbing that reflects regional, social, cultural, and linguistic differences.
[0126] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.
[0127] Computer-readable storage media include all types of storage media that store instructions that can be deciphered by a computer. Examples include read-only memory (ROM), random access memory (RAM), magnetic tape, magnetic disks, flash memory, and optical data storage devices.
[0128] The disclosed embodiments have been described with reference to the attached drawings as described above. Those skilled in the art will understand that the present disclosure can be implemented in forms other than the disclosed embodiments without altering the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be construed as limiting.
Claims
1. In a dubbing generation system including a server and a terminal, The above server, A communication unit configured to receive image information from the terminal; A processor including an artificial intelligence model that converts a voice signal included in the above image information into text and generates dubbing information based on the converted text, A dubbing generation system characterized in that the above text is composed of international phonetic symbols.
2. In paragraph 1, The above processor, A dubbing generation system characterized by generating dubbing information for the voice information based on the translated text and the translation result for the voice signal.
3. In paragraph 2, The above processor, A dubbing generation system characterized in that, based on the converted text, the processor converts the translation result into information composed of international phonetic symbols and generates dubbing information based on the converted information.
4. In paragraph 3, The above processor, A dubbing generation system characterized by removing the audio signal from the video information and combining the dubbing information.
5. In paragraph 4, A dubbing generation system characterized in that the above artificial intelligence model is a transformer.
6. In a dubbing generation method of a dubbing generation system including a server and a terminal, A step in which the server receives image information from the terminal; A step in which the server converts a voice signal included in the image information into text; and The above server comprises a step of generating dubbing information based on the converted text, A dubbing generation method characterized in that the above text is composed of international phonetic symbols.
7. In paragraph 6, The steps for generating the above dubbing information are: A dubbing generation method characterized by generating dubbing information for the voice information based on the converted text and the translation result for the voice signal.
8. In paragraph 7, The steps for generating the above dubbing information are: A dubbing generation method, characterized in that the processor converts the translation result into information composed of an international phonetic symbol based on the converted text and generates dubbing information based on the converted information.
9. In paragraph 8, A dubbing generation method characterized by further comprising the step of removing the audio signal from the image information and combining the dubbing information.
10. In paragraph 9, A dubbing generation method characterized in that the above artificial intelligence model is a transformer.
Citation Information
Patent Citations
Synthesized speech generation method and text-to-speech synthesis device
JP2826215B2
Method, system and recording medium for converting grapheme to phoneme based on prosodic information
KR101735195B1
Method and apparatus for synthesizing singing voice with artificial neural network
KR102168529B1
Method and device for improving dysarthria
KR102499316B1
Translation and dubbing system for video contents
KR102546559B1