System and method for generating subtitles and voice dubbing to which artificial intelligence-based voice technology is applied
The AI-based subtitle and dubbing generation system addresses the inefficiencies and inaccuracies of manual methods by separating voice and background sounds, translating text, and synthesizing audio, resulting in culturally and linguistically accurate, natural, and efficient subtitle and dubbing content.
Patent Information
- Application Number
- PCT/KR2024/017829
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-22
- Filing Date
- 2024-11-12
- Publication Date
- 2025-05-30
AI Technical Summary
Current subtitle and dubbing creation methods are time-consuming and prone to errors due to manual processing, which can result in mistranslations and a lack of cultural and linguistic accuracy, leading to unnatural video content.
An AI-based subtitle and dubbing generation system that separates voice signals from background sounds, converts audio to text, translates text into target languages, and synthesizes the translated voice signals with background sounds to create natural subtitles and dubbing.
The system enables faster and more accurate generation of subtitles and dubbing that reflect regional, social, cultural, and linguistic differences, ensuring natural integration into video content without losing the original sound sources.
Smart Images

Figure KR2024017829_30052025_PF_FP_ABST
Abstract
Description
Subtitle and dubbing generation system and method using artificial intelligence-based voice technology
[0001] The present disclosure relates to a system and method for generating subtitles and dubbing using artificial intelligence-based voice technology.
[0002] The video streaming market, related to the media content market, is growing rapidly. Furthermore, as the video streaming market expands beyond the domestic market to the global market, the need for subtitle and dubbing technology is increasing.
[0003] Specifically, in the growing media content market, individual creators are providing high-quality videos through subtitle and dubbing to enhance their competitiveness. Consequently, the market for subtitle and dubbing services is expanding.
[0004] Currently, subtitles and dubbing are created manually. This not only takes a long time, but the meaning conveyed can vary depending on the creator's linguistic and cultural understanding. Furthermore, mistranslations during the translation process can result in the original dialogue being conveyed differently than intended.
[0005] Additionally, in the case of dubbing, there is a problem that it is difficult to convey the language and emotions of the native speaker, and the music, background music, and sound effects of the original content may be lost due to dubbing, making the video itself unnatural.
[0006] Accordingly, there is a need for technology that can generate subtitles and dubbing that reflect the linguistic and cultural differences of the country where the original work was produced, while remaining unnatural when inserted into a video.
[0007] The purpose of the present disclosure is to provide a system and method capable of generating natural subtitles and dubbing using artificial intelligence.
[0008] A subtitle and dubbing generation system for achieving the above-described purpose may include a server and a terminal, wherein the server comprises a communication unit configured to receive video information or audio information from the terminal, a processor configured to separate audio signals and background sound information from audio information included in the video information or the received audio information, convert the separated audio signals into text, translate the converted text into a target language, generate an audio signal composed of the target language based on at least one of the audio signal, the converted text, and the translation result, and synthesize the generated audio signal and the separated background sound information.
[0009] In one embodiment, the processor can use the translation result to generate a subtitle for the image information, and use the subtitle generation result to synthesize the generated voice signal and the separated background sound information.
[0010] In one embodiment, the processor can synthesize the generated audio signal and the separated background sound information by utilizing time information at which the subtitle is output from the image information.
[0011] In one embodiment, the processor may utilize the result of separating the voice signal and the background sound information from the voice information to synthesize the generated voice signal and the separated background sound information.
[0012] In one embodiment, the processor can synthesize the generated voice signal and the separated background sound information by utilizing time information from which the voice signal is separated from the voice information.
[0013] In addition, a method for generating subtitles and dubbing in a system including a server and a terminal according to the present disclosure may include: a step in which the server receives video information or audio information; a step in which the server separates a voice signal and background sound information from audio information included in the video information or the received audio information; a step in which the server converts the separated voice signal into text; a step in which the server translates the converted text into a target language; a step in which the server generates a voice signal composed of a target language based on at least one of the audio signal, the converted text, and the translation result; and a step in which the server synthesizes the generated voice signal and the separated background sound information.
[0014] According to the subtitle and dubbing generation system according to the present disclosure, since background sound separated from voice information is synthesized after dubbing is generated, dubbing can be generated without loss of sound source.
[0015] In addition, according to the present disclosure, subtitles and dubbing can be generated at a faster speed compared to conventional methods of manually generating subtitles and dubbing.
[0016] Additionally, according to the present disclosure, it becomes possible to create subtitles and dubbing that reflect regional, social, cultural, and linguistic differences.
[0017] Figure 1 is an overall system diagram of the present disclosure.
[0018] FIG. 2 is a block diagram of a server included in the subtitle and dubbing generation system of the present disclosure.
[0019] FIG. 3 is a block diagram of a terminal included in the subtitle and dubbing generation system of the present disclosure.
[0020] FIG. 4 is a block diagram of a processor included in the subtitle and dubbing generation system of the present disclosure.
[0021] Figures 5 and 6 are flowcharts of a method for generating subtitles and dubbing according to the present disclosure.
[0022] FIG. 7 is a flowchart of a method for generating dubbing using international phonetic symbols according to the present disclosure.
[0023] FIG. 8 is a conceptual diagram illustrating an embodiment of a dubbing generation system according to the present disclosure that converts a voice signal into text composed of international phonetic symbols.
[0024] Figure 9 is a conceptual diagram illustrating an artificial intelligence model included in a dubbing generation system according to the present disclosure.
[0025] Throughout this disclosure, the same reference numerals denote the same components. This disclosure does not describe all elements of the embodiments, and any content that is common in the technical field to which this disclosure pertains or that overlaps between embodiments is omitted. The terms "part, module, element, block" used in the specification may be implemented in software or hardware, and depending on the embodiments, multiple "parts, modules, elements, blocks" may be implemented as a single component, or a single "part, module, element, block" may include multiple components.
[0026] Throughout the specification, when a part is said to be "connected" to another part, this includes not only direct connection but also indirect connection, and indirect connection includes connection via a wireless communication network.
[0027] Additionally, when a part is said to "include" a component, this does not mean that it excludes other components, but rather that it may include other components, unless otherwise specifically stated.
[0028] Throughout the specification, when we say that an element is "on" another element, this includes not only cases where the element is in contact with the other element, but also cases where another element exists between the two elements.
[0029] The terms first, second, etc. are used to distinguish one component from another, and the components are not limited by the aforementioned terms.
[0030] Singular expressions include plural expressions unless the context clearly indicates otherwise.
[0031] The identification codes for each step are used for convenience of explanation and do not describe the order of each step. Each step may be performed in a different order than specified unless the context clearly indicates a specific order.
[0032] The operating principle and embodiments of the present disclosure are described below with reference to the attached drawings.
[0033] As used herein, the term "system according to the present disclosure" encompasses various devices capable of performing computational processing and providing results to a user. For example, the system according to the present disclosure may include a computer, a server device, and a portable terminal, or may be any one of them.
[0034] Here, the computer may include, for example, a notebook, desktop, laptop, tablet PC, slate PC, etc. equipped with a web browser.
[0035] The above server device is a server that processes information by communicating with an external device, and may include an application server, a computing server, a database server, a file server, a game server, a mail server, a proxy server, and a web server.
[0036] The above portable terminal may include, for example, a wireless communication device that ensures portability and mobility, and may include all kinds of handheld-based wireless communication devices such as a PCS (Personal Communication System), GSM (Global System for Mobile communications), PDC (Personal Digital Cellular), PHS (Personal Handyphone System), PDA (Personal Digital Assistant), IMT (International Mobile Telecommunication)-2000, CDMA (Code Division Multiple Access)-2000, W-CDMA (W-Code Division Multiple Access), WiBro (Wireless Broadband Internet) terminal, a smart phone, and a wearable device such as a watch, a ring, a bracelet, an anklet, a necklace, glasses, contact lenses, or a head-mounted device (HMD).
[0037] Below, a subtitle and dubbing generation system according to the present disclosure is described.
[0038] Referring to FIG. 1, the subtitle and dubbing generation system according to the present disclosure may include at least one of a server (10) and a terminal (20). Specifically, the subtitle and dubbing generation system according to the present disclosure may be implemented solely by the server (10) or the terminal (20), or may be implemented in the form of a system including at least one of the server (10) and the terminal (20). The description of the subtitle and dubbing generation system described below may be applied to both cases where the subtitle and dubbing generation system according to the present disclosure is implemented solely by the server (10) or the terminal (20), or may be implemented in the form of a system including at least one of the server (10) and the terminal (20).
[0039] The server (10) is connected to at least one terminal (20) via a network, transmits information to each of a plurality of terminals, and generates data necessary for subtitle and dubbing generation learning based on information received from at least one of the terminals (20).
[0040] Meanwhile, it is obvious to those skilled in the art that the terminal (20) is not limited to the above-described portable terminal, and may include a processor-equipped notebook, desktop, laptop, tablet PC, slate PC, etc.
[0041] Below, each of a server (10) and a terminal (20) for implementing a subtitle and dubbing generation system according to the present disclosure will be described.
[0042] FIG. 2 is a block diagram of a server included in the subtitle and dubbing generation system of the present disclosure.
[0043] A server (100) according to the present disclosure may include at least one of a communication unit (110), a storage unit (120), and a processor (130).
[0044] The communication unit (110) can communicate with at least one of a terminal, an external storage (e.g., a database (140)), an external server, and a cloud server.
[0045] Meanwhile, an external server or cloud server may be configured to perform at least a portion of the role of the processor (130). That is, data processing or data operations, etc. may be performed on an external server or cloud server, and the present invention does not impose any particular limitations on this method.
[0046] Meanwhile, the communication unit (110) can support various communication methods according to the communication standards of the communication target (e.g., electronic device, external server, device, etc.).
[0047] For example, the communication unit (110) may be configured to communicate with a communication target using at least one of WLAN (Wireless LAN), Wi-Fi (Wireless-Fidelity), Wi-Fi (Wireless Fidelity) Direct, DLNA (Digital Living Network Alliance), WiBro (Wireless Broadband), WiMAX (World Interoperability for Microwave Access), HSDPA (High Speed Downlink Packet Access), HSUPA (High Speed Uplink Packet Access), LTE (Long Term Evolution), LTE-A (Long Term Evolution-Advanced), 5G (5th Generation Mobile Telecommunication), Bluetooth™, RFID (Radio Frequency Identification), Infrared Data Association (IrDA), UWB (Ultra-Wideband), ZigBee, NFC (Near Field Communication), Wi-Fi Direct, and Wireless USB (Wireless Universal Serial Bus) technologies.
[0048] Next, the storage unit (120) may be configured to store various information related to the present invention. In the present invention, the storage unit (120) may be provided in the device according to the present invention itself. Alternatively, at least a portion of the storage unit (120) may refer to at least one of a database (DB, 140) and a cloud storage (or cloud server). That is, the storage unit (120) may be sufficient as long as it is a space where information required for the device and method according to the present invention is stored, and it can be understood that there are no restrictions on physical space. Accordingly, hereinafter, the storage unit (120), the database (140), the external storage, and the cloud storage (or cloud server) will not be separately distinguished, and will all be referred to as the storage unit (120).
[0049] Next, the processor (130) may be configured to control the overall operation of the device related to the present invention. The processor (130) may process signals, data, information, etc. input or output through the components discussed above, or provide or process appropriate information or functions to the user.
[0050] The processor (130) includes at least one CPU (Central Processing Unit) and can perform functions according to the present invention.
[0051] At least one component may be added or deleted to correspond to the performance of the components illustrated in FIG. 2. Furthermore, it will be readily apparent to those skilled in the art that the relative positions of the components may be altered to correspond to the performance or structure of the device.
[0052] Hereinafter, a terminal included in the subtitle and dubbing generation system of the present disclosure will be described in detail.
[0053] FIG. 3 is a block diagram of a terminal included in the subtitle and dubbing generation system of the present disclosure.
[0054] Referring to FIG. 3, a terminal (200) according to the present disclosure may include a communication unit (210), an input unit (220), a display unit (230), a processor (240), etc. The components illustrated in FIG. 3 are not essential for implementing a subtitle and dubbing generation system according to the present disclosure, and thus, the terminal described in this specification may have more or fewer components than the components listed above.
[0055] Among the above components, the communication unit (210) may include one or more components that enable communication with an external device, and may include, for example, at least one of a broadcast reception module, a wired communication module, a wireless communication module, a short-range communication module, and a location information module.
[0056] The wired communication module may include various wired communication modules such as a Local Area Network (LAN) module, a Wide Area Network (WAN) module, or a Value Added Network (VAN) module, as well as various cable communication modules such as a Universal Serial Bus (USB), a High Definition Multimedia Interface (HDMI), a Digital Visual Interface (DVI), RS-1302 (recommended standard 1302), power line communication, or plain old telephone service (POTS).
[0057] The wireless communication module may include a wireless communication module that supports various wireless communication methods such as GSM (global System for Mobile Communication), CDMA (Code Division Multiple Access), WCDMA (Wideband Code Division Multiple Access), UMTS (universal mobile telecommunications system), TDMA (Time Division Multiple Access), LTE (Long Term Evolution), 4G, 5G, and 6G, in addition to a WiFi module and a Wireless Broadband module.
[0058] The input unit (220) is for inputting video information (or signals), audio information (or signals), data, or information input from a user, and may include at least one camera, at least one microphone, and at least one user input unit. Voice data or image data collected by the input unit may be analyzed and processed into a user control command.
[0059] The display unit (230) is intended to generate output related to visual, auditory, or tactile sensations, and may include at least one of a display unit, an audio output unit, a haptic module, and an optical output unit. The display unit may be formed as a touch screen by forming a mutual layer structure with a touch sensor or by forming an integral structure with the touch sensor. Such a touch screen may function as a user input unit that provides an input interface between the device and a user, and at the same time, provide an output interface between the device and the user.
[0060] The display unit displays (outputs) information processed by this device. For example, the display unit may display execution screen information of an application program (e.g., an application) running on this device, or UI (User Interface) or GUI (Graphical User Interface) information based on such execution screen information.
[0061] In addition to the above-described components, the above-described terminal may further include an interface unit and a memory.
[0062] The interface unit serves as a passageway for various types of external devices connected to the device. The interface unit may include at least one of a wired / wireless headset port, an external charger port, a wired / wireless data port, a memory card port, a port for connecting a device equipped with an identification module (SIM), an audio I / O (Input / Output) port, a video I / O (Input / Output) port, and an earphone port. The device may perform appropriate control related to the external device connected to the interface unit.
[0063] The memory can store data supporting various functions of the device, programs for the operation of the processor, input / output data (e.g., music files, still images, moving images, etc.), and a plurality of application programs (or applications) running on the device, data for the operation of the device, and commands. At least some of these application programs can be downloaded from an external server via wireless communication.
[0064] Such memory may include at least one type of storage medium among flash memory type, hard disk type, SSD (Solid State Disk type), SDD (Silicon Disk Drive type), multimedia card micro type, card type memory (e.g., SD or XD memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, and optical disk. In addition, the memory may be a database that is separate from the device but connected by wire or wirelessly.
[0065] Meanwhile, the terminal described above includes a processor (240). The processor may be implemented as a memory storing data regarding an algorithm for controlling the operation of components within the device or a program reproducing the algorithm, and at least one processor (not shown) that performs the aforementioned operations using the data stored in the memory. In this case, the memory and the processor may be implemented as separate chips. Alternatively, the memory and the processor may be implemented as a single chip.
[0066] Meanwhile, the processor may control any one or a combination of the components described above to implement various embodiments of the present disclosure described in the drawings below on the device.
[0067] Meanwhile, at least one component may be added or deleted in accordance with the performance of the components illustrated in Figures 1 to 3. Furthermore, it will be readily apparent to those skilled in the art that the relative positions of the components may be altered in accordance with the performance or structure of the device.
[0068] Meanwhile, as illustrated in FIG. 4, a processor included in at least one of the server and the terminal may include multiple modules for implementing a subtitle and dubbing generation system, which will be described later. Specifically, the processor (300) may include a voice recognition module (310) and an artificial intelligence module (320). While the subtitle and dubbing generation method described below is described as being implemented by the operations of the modules, the performance of each step described below need not necessarily be performed by the modules.
[0069] Below, the artificial intelligence described in the present invention is described in detail.
[0070] The artificial intelligence-related functions according to the present disclosure are operated through the processor and memory installed in the above-described server and terminal. The processor may be composed of one or more processors. In this case, one or more processors may be a general-purpose processor such as a CPU, an AP, a DSP (Digital Signal Processor), a graphics-only processor such as a GPU or a VPU (Vision Processing Unit), or an artificial intelligence-only processor such as an NPU. One or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in the memory. Alternatively, if one or more processors are artificial intelligence-only processors, the artificial intelligence-only processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0071] The predefined operation rules or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that the basic artificial intelligence model is learned by a learning algorithm using a plurality of learning data, thereby creating a predefined operation rules or artificial intelligence model set to perform a desired characteristic (or purpose). This learning may be performed in the device itself on which the artificial intelligence according to the present disclosure is performed, or may be performed through a separate server and / or system. Examples of the learning algorithm include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.
[0072] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values, and performs neural network operations through operations between the operation results of the previous layer and the multiple weights. The multiple weights of the multiple neural network layers may be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model is reduced or minimized during the learning process. The artificial neural network may include a deep neural network (DNN), and examples thereof include, but are not limited to, a convolutional neural network (CNN), a deep neural network (DNN), a recurrent neural network (RNN), a restricted boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), or deep Q-networks.
[0073] According to an exemplary embodiment of the present disclosure, a processor can implement artificial intelligence. Artificial intelligence refers to a machine learning method based on an artificial neural network that imitates human neurons (biological neurons) to enable machines to learn. Artificial intelligence methodologies can be categorized into supervised learning, in which input data and output data are provided together as training data depending on the learning method, so that the solution (output data) to the problem (input data) is determined; unsupervised learning, in which only input data is provided without output data, so that the solution (output data) to the problem (input data) is not determined; and reinforcement learning, in which a reward (Reward) is provided from an external environment whenever an action (Action) is taken in the current state (State), and learning is performed in a direction to maximize this reward. In addition, artificial intelligence methodologies can be categorized according to the architecture of the learning model. The architectures of widely used deep learning technologies can be categorized into convolutional neural networks (CNNs), recurrent neural networks (RNNs), transformers, and generative adversarial networks (GANs).
[0074] The present device and system may include an artificial intelligence model. The artificial intelligence model may be a single artificial intelligence model or may be implemented as multiple artificial intelligence models. The artificial intelligence model may be composed of a neural network (or artificial neural network) and may include statistical learning algorithms that mimic biological neurons in machine learning and cognitive science. A neural network may refer to a model in general that has problem-solving capabilities by changing the binding strength of synapses through learning, formed by artificial neurons (nodes) that form a network by combining synapses. The neurons of the neural network may include a combination of weights or biases. The neural network may include one or more layers composed of one or more neurons or nodes. For example, the device may include an input layer, a hidden layer, and an output layer. The neural network constituting the device can infer a desired result (output) from an arbitrary input (input) by changing the weights of neurons through learning.
[0075] The processor can create a neural network, train (or learn) a neural network, perform a calculation based on received input data, generate an information signal based on the calculation result, or retrain the neural network. The models of the neural network can include various types of models such as CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, R-CNN (Region with Convolution Neural Network), RPN (Region Proposal Network), RNN (Recurrent Neural Network), S-DNN (Stacking-based deep Neural Network), S-SDNN (State-Space Dynamic Neural Network), Deconvolution Network, DBN (Deep Belief Network), RBM (Restrcted Boltzman Machine), Fully Convolutional Network, LSTM (Long Short-Term Memory) Network, Classification Network, etc., but are not limited thereto. The processor can include one or more processors for performing calculations according to the models of the neural network. For example, the neural network can be a deep neural network. It may include a deep neural network.
[0076] Neural networks include CNN (Convolutional Neural Network), RNN (Recurrent Neural Network), perceptron, multilayer perceptron, FF (Feed Forward), RBF (Radial Basis Network), DFF (Deep Feed Forward), LSTM (Long Short Term Memory), GRU (Gated Recurrent Unit), AE (Auto Encoder), VAE (Variational Auto) Encoder), DAE (Denoising Auto Encoder), SAE (Sparse Auto Encoder), MC (Markov Chain), HN (Hopfield Network), BM (Boltzmann Machine), RBM (Restricted Boltzmann Machine), DBN (Depp Belief Network), DCN (Deep Convolutional Network), DN (Deconvolutional Network), DCIGN (Deep Convolutional Inverse Graphics Network), Generative Adversarial Network (GAN), Liquid State Machine (LSM), Extreme Learning Machine (ELM), It will be understood by those skilled in the art that any neural network may be included, including but not limited to ESN (Echo State Network), DRN (Deep Residual Network), DNC (Differentiable Neural Computer), NTM (Neural Turning Machine), CN (Capsule Network), KN (Kohonen Network), and AN (Attention Network).
[0077] According to an exemplary embodiment of the present disclosure, the processor may be configured to perform a process for generating a CNN (Convolution Neural Network) such as GoogleNet, AlexNet, VGG Network, Region with Convolution Neural Network (R-CNN), Region Proposal Network (RPN), Recurrent Neural Network (RNN), Stacking-based deep Neural Network (S-DNN), State-Space Dynamic Neural Network (S-SDNN), Deconvolution Network, Deep Belief Network (DBN), Restrcted Boltzman Machine (RBM), Fully Convolutional Network, Long Short-Term Memory (LSTM) Network, Classification Network, Generative Modeling, eXplainable AI, Continual AI, Representation Learning, AI for Material Design, BERT, SP-BERT, MRC / QA for natural language processing, Text Analysis, Dialog System, GPT-3, GPT-4, Visual Analytics for vision processing, Visual Understanding, Video Synthesis, ResNet for data intelligence, Anomaly Detection, Prediction, Time-Series Forecasting, Various artificial intelligence structures and algorithms, including optimization, recommendation, and data creation, can be utilized, but are not limited thereto. Hereinafter, embodiments of the present disclosure will be described in detail with reference to the attached drawings.
[0078] Below, a method for generating subtitles and dubbing using the above-described components is described in detail.
[0079] Figures 5 and 6 are flowcharts of a method for generating subtitles and dubbing according to the present disclosure.
[0080] Referring to FIGS. 5 and 6, a step of receiving audio information is performed (S110). Here, the step of receiving audio information may be a step of receiving video information. The Jamik dubbing generation system according to the present disclosure receives video information and utilizes the audio information contained in the video information to generate subtitles and dubbing. Alternatively, the step of receiving audio information may be a step of receiving audio information itself, rather than video information.
[0081] In this specification, audio information may include all data related to sound included in an image. In one embodiment, audio information may include audio signals and background sounds included in the image information.
[0082] The processor receives image information and processes audio information included in the image information.
[0083] Next, a step of separating voice signal and background sound information from voice information is performed (S120).
[0084] The processor can separate the voice signal and background sound information from the voice information and create separate files.
[0085] The audio signal is the speaker's voice included in the video information, and is the information that is the subject of subtitle and dubbing generation. Later, the audio signal separated from the audio information is generated as at least one of subtitles and dubbing.
[0086] The separation of the above speech signal and background sound information can be performed through a separate artificial intelligence model trained to separate the speech signal and background sound information from the speech information.
[0087] Next, a step of converting the separated voice signal into text is performed (S130).
[0088] The above voice signal can be converted into text through the voice recognition module (310). Here, the text can be formed in the language that is the basis for producing the video from which the voice signal is separated.
[0089] In this specification, the language that is the basis of video production is called the “source language,” and the language in which subtitles and dubbing are to be created is called the “target language.”
[0090] Next, a step is performed to translate the converted text composed of the original language into the target language (S140).
[0091] The above translation result is in the form of text in the target language and can be provided as subtitles for the original video.
[0092] The type of the above target language can be input by the video producer or the user viewing the video.
[0093] In one embodiment, a video producer can generate subtitles for at least one target language from the video production stage by specifying the target language in advance.
[0094] In one embodiment, a video producer may generate subtitles by allowing viewers to specify a target language of their choice, rather than creating subtitles themselves. Subtitles created by a specific viewer can then be provided to viewers who request subtitles in that target language.
[0095] The above translation results can then be used to create dubbing, as described later.
[0096] Meanwhile, the above translation can be performed using a separate AI model trained to translate text in the source language into the target language. The type of AI model used for translation is not specifically limited.
[0097] Meanwhile, the processor generates the subtitles and then synthesizes them into the original video. At this time, the processor can set a time interval for inserting the subtitles by considering the time interval in which the audio signal is generated within the video information. Specifically, the processor can match the time interval in which the audio signal is generated within the video information at the syllable, word, or sentence level. The time information matched in this manner is also matched to the converted text after the audio signal is converted into text. Thereafter, the time information can also be matched to the result of translating the converted text. The processor can insert the subtitles into the video information based on the time information.
[0098] Next, a step is performed to generate a voice signal (hereinafter, dubbing information) in a target language based on at least one of a separated voice signal, a converted text, and a translation result (S150).
[0099] Hereinafter, with reference to FIGS. 7 to 9, a method for generating dubbing information in a target language will be described in detail.
[0100] Referring to Fig. 7, a step is performed in which an artificial intelligence module (320) receives a voice signal (S210).
[0101] Here, the information input to the artificial intelligence module (320) may be audio information included in the image information, or an audio signal in which background sound information is separated from the audio information.
[0102] The above artificial intelligence module (320) may include an artificial intelligence model trained to convert voice information into the International Phonetic Association (IPA).
[0103] Next, a step is performed in which the artificial intelligence model converts voice information into text composed of the International Phonetic Alphabet (IPA) (S220).
[0104] In one embodiment, the artificial intelligence model may be, but is not limited to, a transformer.
[0105] In one embodiment, referring to FIG. 8, the artificial intelligence model can receive the voice information “Could you lend me ten thousand won?” and convert it into the following international phonetic symbol.
[0106]
[0107] In one embodiment, referring to FIG. 9, the artificial intelligence model can be trained such that 'src' is a speech signal composed of Korean, 'tgt' is an English IPA embedding, and 'output' is an English IPA embedding.
[0108] Thereafter, the artificial intelligence model generates subtitle or dubbing information based on the converted text, i.e., the text composed of the international phonetic symbols (S230).
[0109] Specifically, the artificial intelligence model can generate dubbing information for the voice information based on the translated text and the translation result for the voice signal.
[0110] Here, the above translation result may be the translation result performed in S140.
[0111] Based on the converted text, the AI model converts the translation result into information structured in the International Phonetic Alphabet (IPA). To this end, the AI model can be trained to receive text structured in the IPA, generated from a speech signal in the source language, and the translation result of the speech signal in the source language. The model then generates text structured in the IPA and translated into the target language. Utilizing the IPA during the training process allows for the generation of dubbing information that reflects regional, social, cultural, and linguistic differences between the source and target languages.
[0112] Additionally, as described above, when utilizing the International Phonetic Alphabet when training an artificial intelligence model to generate dubbing information, it becomes possible to generate dubbing information that pronounces the target language naturally.
[0113] Again, referring to FIGS. 5 and 6, a step of synthesizing the generated voice signal (dubbing information) and the separated background sound information is performed (S160).
[0114] The processor can use the translation result generated in S140 to generate a subtitle for the image information, and use the subtitle generation result to synthesize the generated voice signal and the separated background sound information.
[0115] Specifically, the processor can search for time information at which the subtitle is output from the image information, and synthesize the generated voice signal and the separated background sound information by utilizing the searched time information.
[0116] The processor can match the time intervals in which audio signals are generated within the video information, at the syllable, word, or sentence level. The time information matched in this manner is then matched to the converted text after the audio signal is converted into text. The time information can then be matched to the translated text. Based on this time information, the processor can insert subtitles into the video information. Furthermore, the processor can utilize the aforementioned time information to generate dubbing.
[0117] Meanwhile, the processor may utilize the result of separating the voice signal and the background sound information from the voice information included in the image information to synthesize the generated voice signal and the separated background sound information. Specifically, the processor may utilize the time information at which the voice signal was separated from the voice information to synthesize the generated voice signal and the separated background sound information.
[0118] When separating a voice signal from background sound information, the processor can match the background sound information with information defining the time intervals from which the voice signal was separated. This time interval information can be defined at the syllable, word, or sentence level and matched to the background sound information. This time information can be utilized when synthesizing dubbing information with the background sound information.
[0119] As described above, according to the subtitle and dubbing generation system according to the present disclosure, since background sound separated from voice information is synthesized after dubbing is generated, dubbing can be generated without loss of sound source.
[0120] In addition, according to the present disclosure, subtitles and dubbing can be generated at a faster speed compared to conventional methods of manually generating subtitles and dubbing.
[0121] Additionally, according to the present disclosure, it becomes possible to create subtitles and dubbing that reflect regional, social, cultural, and linguistic differences.
[0122] Meanwhile, the disclosed embodiments may be implemented in the form of a recording medium storing computer-executable instructions. The instructions may be stored in the form of program code, and when executed by a processor, may generate program modules to perform the operations of the disclosed embodiments. The recording medium may be implemented as a computer-readable recording medium.
[0123] Computer-readable storage media include all types of storage media that store instructions that can be deciphered by a computer. Examples include read-only memory (ROM), random access memory (RAM), magnetic tape, magnetic disks, flash memory, and optical data storage devices.
[0124] The disclosed embodiments have been described with reference to the attached drawings as described above. Those skilled in the art will understand that the present disclosure can be implemented in forms other than the disclosed embodiments without altering the technical spirit or essential features of the present disclosure. The disclosed embodiments are illustrative and should not be construed as limiting.
Claims
1. In a subtitle and dubbing generation system including a server and a terminal, The above server, A communication unit configured to receive image information or voice information from the terminal; Separate voice signal and background sound information from the voice information included in the above image information or the received voice information, Convert the above separated voice signal into text, Translate the above converted text into the target language, Generating a speech signal in a target language based on at least one of the speech signal, the converted text and the translation result, A subtitle and dubbing generation system characterized by including a processor that synthesizes the generated voice signal and the separated background sound information.
2. In paragraph 1, The above processor, Using the above translation results, subtitles are created for the above video information, A subtitle and dubbing generation system characterized by synthesizing the generated voice signal and the separated background sound information by utilizing the subtitle generation result.
3. In paragraph 2, The above processor, Search for the time information at which the subtitle is output in the above video information, A subtitle and dubbing generation system characterized by synthesizing the generated voice signal and the separated background sound information by utilizing the above-mentioned searched time information.
4. In paragraph 3, The above processor, A subtitle and dubbing generation system characterized in that it synthesizes the generated voice signal and the separated background sound information by utilizing the result of separating the voice signal and the background sound information from the voice information.
5. In paragraph 4, The above processor, A subtitle and dubbing generation system characterized in that it synthesizes the generated voice signal and the separated background sound information by utilizing the time information from which the voice signal is separated from the above voice information.
6. A method for generating subtitles and dubbing in a system including a server and a terminal, A step in which the above server receives video information or audio information; A step in which the server separates voice information included in the video information or voice signal and background sound information from the received voice information; A step in which the server converts the separated voice signal into text; A step in which the server translates the converted text into a target language; A step in which the server generates a speech signal in a target language based on at least one of the speech signal, the converted text, and the translation result; A method for generating subtitles and dubbing, characterized in that the server comprises a step of synthesizing the generated voice signal and the separated background sound information.
7. In paragraph 6, The step of synthesizing the generated voice signal and the separated background sound information is: Using the above translation results, subtitles are created for the above video information, A subtitle and dubbing generation method characterized by synthesizing the generated voice signal and the separated background sound information by utilizing the above subtitle generation result.
8. In paragraph 2, The step of synthesizing the generated voice signal and the separated background sound information is: A method for generating subtitles and dubbing, characterized in that the generated voice signal and the separated background sound information are synthesized by utilizing the time information at which the subtitle is output from the video information.
9. In paragraph 8, The step of synthesizing the generated voice signal and the separated background sound information is: A method for generating subtitles and dubbing, characterized in that the generated voice signal and the separated background sound information are synthesized by utilizing the result of separating the voice signal and the background sound information from the voice information.
10. In paragraph 9, The step of synthesizing the generated voice signal and the separated background sound information is: A method for generating subtitles and dubbing, characterized in that the generated voice signal and the separated background sound information are synthesized by utilizing time information from which the voice signal is separated from the voice information.
Citation Information
Patent Citations
Server, control method and control program for server, information processing system, information processing method, portable terminal, control method and control program for portable terminal
JP2013198066A
System and method for performing automatic dubbing onan audio-visual stream
KR1020050118733A
Video Authoring System and Method
KR102343336B1
Automatic translation system of video contents for hearing-impaired and non-disabled
KR102463283B1
Translation and dubbing system for video contents
KR102546559B1