Artificial intelligence content production service device and driving method therefor
The AI content production service device addresses the challenges of consistency and differentiation in AI-generated content by using multiple AI models to convert text into general-purpose data and generate various content types, resulting in efficient and cost-effective production of high-quality AI secondary content.
Patent Information
- Application Number
- PCT/KR2024/013425
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-07
- Filing Date
- 2024-09-05
- Publication Date
- 2025-06-12
AI Technical Summary
Existing AI content generation technologies struggle to produce consistent and differentiated AI secondary content such as text-based audio, webtoons, and videos, while also being inefficient in processing large amounts of text and requiring manual expertise and high production costs.
An AI content production service device that utilizes multiple AI models to convert text into general-purpose data, generate prompts, and automatically produce various content types like audiobooks, webtoons, and videos, ensuring consistency and differentiation through advanced AI language models and content generation engines.
The solution enables efficient and cost-effective production of high-quality, consistent, and differentiated AI secondary content, addressing the limitations of existing technologies and meeting the demands of diverse content consumption trends.
Smart Images

Figure KR2024013425_12062025_PF_FP_ABST
Abstract
Description
Artificial intelligence content production service device and its operating method
[0001] The present invention relates to an artificial intelligence content production service device and a method for operating the device, and provides a service for producing AI secondary content such as text-based audio, webtoons, and videos that emphasize consistency within a work and differentiation between works, and relates to an artificial intelligence content production service device and a method for operating the device that converts text into multi-use data using an AI language model (e.g., GPT, LLaMa, other LLMs, etc.) and automatically generates various content such as audiobooks, webtoons, and videos at a commercial level by connecting scenarios and prompts generated using the data with different types of generation AI engines (text-to-contents).
[0002] With the diversification of content consumption preferences, personalized content is gaining traction, and consumers are increasingly using features like speed-up and video summaries to consume content, leading to explosive growth in content demand. However, content production (and supply) remains largely manual, relying on low labor costs, hindering significant innovation in content production speed and cost. Consequently, the content shortage is expected to worsen.
[0003] Meanwhile, the fragmentation of content demand is lowering the expected revenue generated by a single piece of content. The profitability of traditional mainstream media such as books, TV, and movies is deteriorating, and demand is diversifying across newer mediums like web novels, webtoons, YouTube, and OTT dramas. Even for a single IP, it's crucial to proactively diversify content production risk by securing diverse content types and supply channels. However, existing production methods are prohibitively expensive even for a single piece of content.
[0004] Furthermore, unlike the past, when Korean literature relied heavily on foreign literature, the Korean literature market has recently grown. In particular, the market for "web novels," a form of Korean genre literature with unique settings and cliches, is now surpassing the domestic novel market. Best-selling web novels are being adapted into secondary content such as dramas, movies, and webtoons, achieving box office success. Furthermore, Korean web novels and webtoons based on web novels are gaining popularity in Europe and the US, opening up new markets. Creating content utilizing web novel (Korean novel) IP is essential for penetrating overseas markets, such as the US, where audiobooks and webtoons are more popular than in Korea. To meet this demand, an AI content generation engine is needed to efficiently create and mass-produce secondary content based on web novels.
[0005] While conventional generative AI engines have been widely commercialized, they face significant limitations in content creation. First, while many generative AI engines exist, they focus on generating fragmented text, images, and videos. Text generation AI summarizes, analyzes, and generates short texts. Image generation AI takes a single sentence of text and generates a single image. Therefore, existing AI technologies alone have limitations in processing large amounts of text and immediately transforming it into content.
[0006] Furthermore, conventional generative AI engines require general users to learn how to use each AI model to create content, and manually connect and edit the output generated by each AI model. This production process offers little time or cost advantage over the traditional, full-process, manual content creation process. However, AI has the advantage of enabling work even without domain expertise. However, the quality of content produced by experts in traditional methods is significantly lower. Furthermore, excluding the use of current SaaS services, detailed configuration requires additional AI researchers and engineers. In the content industry, which is sustained by low labor costs, recruiting AI talent is a significant challenge. For AI audiobook production, a human must read each sentence, assign appropriate narrators and characters, and configure detailed options such as emotion, reading speed, and ending cadence. This same process is required even when using commercialized AI TTS SaaS services.
[0007] Conventional AI generators lack consistency in their results, making them unsuitable for creating long-form content or feature-length series. Due to the nature of AI technology, if detailed settings aren't provided, the resulting output will vary with each generation. This is the biggest challenge in AI content generation. For example, in audiobooks, the voice of the same character within a single novel changes slightly from sentence to sentence, and completely different voices are heard in the beginning and end of the story. Similarly, in webtoons and videos, the drawing style (sketch style) and colors vary across each scene, hindering the smooth flow of a character's narrative.
[0008] Moreover, generative AI engines lack differentiation in their output, which means that when content is mass-produced, the resulting output tends to be similar. This means that it's difficult to differentiate between works with similar settings. For example, web novels featuring a "blonde, blue-eyed noblewoman" as the main character are common. If actual webtoon artists draw webtoons, their distinctive art style allows users to identify the work just by looking at the characters. However, images generated by generative AI are the result of learning multiple art styles corresponding to "blonde" and "blue eyes," and each generative AI model generates images centered around a single stereotypical image. For this reason, when content is mass-produced, the resulting output tends to be similar across multiple works.
[0009] [Prior Art Literature]
[0010] [Patent Document]
[0011] (Patent Document 1) Korean Patent Publication No. 10-2313203 (October 8, 2021)
[0012] (Patent Document 2) Korean Patent Publication No. 10-2586799 (October 4, 2023)
[0013] (Patent Document 3) Korean Patent Publication No. 10-2022-0017068 (February 11, 2022)
[0014] The technical task to be achieved by the present invention is to provide a service for producing AI secondary content such as text-based audio, webtoons, and videos that maintain consistency within the work and differentiation between works, and to convert text into general-purpose data using an AI language model (e.g., GPT, LLaMa, or other LLM), and to automatically generate various content such as audiobooks, webtoons, and videos at a commercial level by connecting scenarios and prompts generated using this data with different types of generative AI engines.
[0015] The problems to be solved by the present invention are not limited to the problems mentioned above, and other problems not mentioned will be clearly understood by those skilled in the art from the description below.
[0016] An artificial intelligence content production device according to an embodiment of the present invention for achieving the above-described task includes a communication interface unit that receives text data for generating content using an artificial intelligence (AI) program from a user terminal device, and a control unit that converts the received text data into general data, converts the converted general data into text for prompts for each content, and generates and outputs content such as an audiobook, a webtoon, or a video using an AI model that generates each content based on the converted text for prompts for each content.
[0017] The above control unit can execute a first AI model to perform a conversion operation into the general data, execute a second AI model to perform a conversion operation into the text for the prompt, and execute a third AI model to perform a generation operation of the content.
[0018] The above control unit can operate multiple third AI models in parallel to generate each content based on the converted prompt text.
[0019] The above control unit can execute, as the plurality of third AI models, an audiobook model for generating the audiobook, a webtoon model for generating the webtoon, and a video model for generating the video.
[0020] The above control unit can execute the AI model to summarize only the portions to be made into content without performing a character analysis operation when the received text data is non-novel text, and separate the summarized sentences to generate a prompt.
[0021] In addition, a method for driving an artificial intelligence content production device according to an embodiment of the present invention for achieving the above-mentioned task includes a step of receiving text data for generating content using an artificial intelligence (AI) program from a user terminal device by a communication interface unit, and a step of converting the received text data into general data by a control unit, converting the converted general data into text for a prompt for each content, and generating and outputting content of an audiobook, webtoon, or video using an AI model for generating each content based on the text of the prompt for each content that has been converted.
[0022] The above generating and outputting step may include executing a first AI model to perform a conversion operation into the general data, executing a second AI model to perform a conversion operation into the text for the prompt, and executing a third AI model to perform a generation operation of the content.
[0023] The above generating and outputting step can operate multiple third AI models in parallel structure to generate each content based on the converted prompt text.
[0024] The above generating and outputting step may execute an audiobook model for generating the audiobook, a webtoon model for generating the webtoon, and a video model for generating the video as the plurality of third AI models.
[0025] The above generating and outputting step may be performed by executing the AI model to summarize only the portions to be made into content without performing character analysis if the received text data is non-novel text, and the summarized sentences may be separated to generate a prompt.
[0026] According to an embodiment of the present invention, a universal database (DB) that is primarily processed with a first model (e.g., gpt, etc.) that converts text into universal data can be constructed, and can have expandability to various contents, and can process large amounts of text while maintaining consistency and create contents that take advantage of differentiation for each work.
[0027] FIG. 1 is a diagram illustrating an artificial intelligence content production system according to an embodiment of the present invention.
[0028] Figure 2 is a diagram that schematically illustrates the process of generating data, scenarios, and final content using AI text analysis technology.
[0029] Figure 3 is a block diagram illustrating a detailed structure of the artificial intelligence content production device of Figure 1.
[0030] Figures 4 and 5 are drawings showing an artificial intelligence content production process according to an embodiment of the present invention.
[0031] Figures 6 and 7 are flowcharts showing the operation process of an AI model that performs text analysis.
[0032] Figure 8 is a flowchart illustrating the operation process of a webtoon creation model.
[0033] Figure 9 is a flowchart illustrating the driving process of an image generation model.
[0034] Figure 10 is a flowchart showing the operation process of the artificial intelligence content production device of Figure 1.
[0035] The embodiments of the present invention described below are provided to more clearly explain the present invention to a person having ordinary skill in the art, and the scope of the present invention is not limited by the following embodiments, and the following embodiments may be modified in various other forms.
[0036] The terminology used herein is used to describe particular embodiments and is not intended to limit the present invention. The singular forms used herein may include the plural forms unless the context clearly dictates otherwise. In addition, the terms "comprise" and / or "comprising" used herein specify the presence of a stated feature, step, number, operation, element, element, and / or group thereof, but do not exclude the presence or addition of one or more other features, steps, numbers, operations, elements, elements, and / or groups thereof. In addition, the term "connection" used herein not only means that certain elements are directly connected, but also includes a concept that indirectly connects elements by interposing another element between them.
[0037] In addition, when it is said in this specification that a certain element is located "on" another element, this includes not only cases where a certain element is in contact with another element, but also cases where another element exists between the two elements. The term "and / or" as used in this specification includes any one of the listed items and any and all combinations of one or more of them. In addition, terms of degree such as "about", "substantially", etc. as used in this specification are used to mean a range of or close to the numerical value or degree, taking into account inherent manufacturing and material tolerances, and are used to prevent infringers from unfairly using the disclosure that mentions exact or absolute numbers provided to help the understanding of this specification.
[0038] Hereinafter, embodiments of the present invention will be described in detail with reference to the attached drawings. The sizes and thicknesses of areas or parts depicted in the attached drawings may be somewhat exaggerated for clarity and convenience of explanation. Like reference numbers designate like components throughout the detailed description.
[0039] FIG. 1 is a diagram showing an artificial intelligence content production system according to an embodiment of the present invention, and FIG. 2 is a diagram briefly schematically showing the process of creating data, scenarios, and final content using AI text analysis technology.
[0040] As illustrated in FIG. 1, the artificial intelligence content production system (90) according to an embodiment of the present invention includes part or all of a user terminal device (100), a communication network (110), an artificial intelligence content production device (120), and a third-party device (130).
[0041] Here, “including some or all” means that some components, such as a third-party device (130), may be omitted to configure the artificial intelligence content production system (90), or some or all of the components constituting the artificial intelligence content production device (120) may be integrated into a network device (e.g., a wireless switching device, etc.) constituting a communication network (110), and is explained as including all in order to help a sufficient understanding of the invention.
[0042] The user terminal device (100) can access the service provided by the artificial intelligence content creation device (120) of FIG. 1 and use the service to create secondary content based on text directly input by the user or text such as an e-book. In other words, the creation of secondary content is based on text data provided by general users or publishers through the user terminal device (100), but content such as text, webtoons, and videos of new forms or contents can be provided. The user terminal device (100) can receive secondary content created by the artificial intelligence model installed in the artificial intelligence content creation device (120) by inputting text data in a designated area on the service screen provided to the artificial intelligence content creation device (120) of FIG. 1. Of course, the service screen can be provided in various layouts depending on the intention of the service provider, but by providing a text file written in the format of Hangul or MS Word .txt.
[0043] The user terminal device (100) may include various types of terminal devices. The user terminal device (100) may be any type of terminal device as long as it can access the AI content production device (120) of FIG. 1 and use the service according to an embodiment of the present invention. For example, the user terminal device (100) may include not only PC-based terminal devices such as desktop computers or laptop computers, but also mobile-based terminal devices such as smartphones, tablet PCs, and wearable devices worn by users on their wrists, etc. Furthermore, the user terminal device (100) may further include various devices such as smart TVs capable of wired and wireless Internet access. Of course, the user terminal device may further include artificial intelligence (AI) speakers that operate in conjunction with smart TVs, etc. For example, secondary content produced in the form of an audiobook may be output through an AI speaker so that users who have requested content production can listen to it.
[0044] When a user terminal device (100) provides text data to the artificial intelligence content creation device (120) of FIG. 1, of course, the text data here can be divided into novel and non-novel text data, and users of the user terminal device (100) can provide type information related to content creation, such as what type (or types) of content they want to create, while providing the text data. For example, they can provide information such as whether they want a new type of text for the input text data, whether they want to create a webtoon based on the text data, or whether they want to create a video, i.e., an image. For example, when user A's user terminal device (100) provides content type information for webtoon creation while providing text data, the device identification information of the user terminal device (100) is also provided, so that the artificial intelligence content creation device (120) of FIG. 1 can know which user used which content creation service.
[0045] The communication network (110) can be configured in various forms. The communication network (110) can include both wired and wireless communication networks. For example, a wired or wireless Internet network can be used or linked as the communication network (110). Here, the wired network includes an Internet network such as a cable network or a public switched telephone network (PSTN), and the wireless communication network includes a CDMA, WCDMA, GSM, EPC (Evolved Packet Core), LTE (Long Term Evolution), Wibro network, etc. Of course, the communication network (110) according to the embodiment of the present invention is not limited thereto, and can be used as an access network of a next-generation mobile communication system to be implemented in the future, for example, a cloud computing network in a cloud computing environment, a 5G network, a 6G network, etc. For example, if the communication network (110) is a wired communication network, an access point within the communication network can connect to a telephone exchange, etc., but if it is a wireless communication network, data can be processed by connecting to an SGSN or GGSN (Gateway GPRS Support Node) operated by a communication company, or data can be processed by connecting to various relays such as a BTS (Base Transceiver Station), NodeB, or e-NodeB.
[0046] The communication network (110) may include an access point. The access point may include a small base station, such as a femto or pico base station, which is often installed inside a building. Here, the femto or pico base station may be classified according to the maximum number of user terminal devices (100) or third-party devices (130) of FIG. 1 that can be connected to the small base station. Of course, the communication network (110) may include a short-range communication module for performing short-range communication, such as Zigbee or Wi-Fi, with the user terminal devices (100) or third-party devices (130). The access point may use TCP / IP or RTSP (Real-Time Streaming Protocol) for wireless communication. Here, short-range communication can be performed using various standards such as Bluetooth, Zigbee, infrared (IrDA), radio frequency (RF) such as ultra-high frequency (UHF) and very high frequency (VHF), and ultra-wideband (UWB) in addition to Wi-Fi. Accordingly, the access point can extract the location of the data packet, designate the best communication path for the extracted location, and transmit the data packet along the designated communication path to the next device, such as the artificial intelligence content production device (120). The access point can share multiple lines in a general network environment, and may include, for example, a router, a repeater, and a repeater.
[0047] The AI content production device (120) may include, for example, a cloud server, and may be configured to include a DB (120a) of FIG. 1 that is linked to the server. Since the server and the DB (120a) are connected via an internal dedicated network such as an intranet, data can be processed without a separate data compression operation during data processing, and data can be systematically and quickly classified and stored and managed for each user in the DB (120a). The AI content production device (120) according to an embodiment of the present invention can provide a service for producing AI secondary content such as text-based audio (books), webtoons, and videos that emphasize consistency and differentiation. To this end, the artificial intelligence content creation device (120) can be seen as implementing AI technology (or function) for creating audiobooks, webtoons, and videos from novel texts, AI technology for summarizing non-novel texts, extracting sentences, and then creating the above content, technology for converting novels into audiobooks, technology for determining the type of output content and increasing the range of output data using a new AI model, and technology for summarizing only the parts to be converted into content overall for non-novel texts without requiring character analysis, and generating prompts by separating (for example, summarized) sentences (or based on the separated sentences). The artificial intelligence content creation device (120) can be seen as providing a service according to an embodiment of the present invention by a hardware (H / W) module, a software (S / W) module, or a combination thereof for implementing the above technology.
[0048] The service provided by the AI content creation device (120) may be named "Oddnary" according to an embodiment of the present invention. The AI content creation device (120) may be equipped with and run a general-purpose TTC (text to contents) AI, such as a general-purpose AI for language, voice, image, and video, to solve the problems of existing generative AI models through the Oddnary service. More precisely, the AI content creation device (120) may use a first AI model, such as GPT, to convert pre-input text (data) into general-purpose data. Then, it uses a second AI model, such as GPT, llama, or other LLM, to convert the general-purpose data into text for each content, i.e., a prompt. Finally, it uses (multiple) third AI models that form a parallel structure to generate each content, to generate each content based on the prompt. This will be discussed in more detail later.
[0049] Alternatively, the AI content creation device (120) can perform automated operations without manual work using step-by-step text analysis and generation models. A text processing process that shortens the traditional work creation process is applied before the generation model is input. For example, audiobooks may require an editor's editing, scripts for dialogue and narration, and background sound / sound effects. Webtoons / videos require storyboards containing descriptions of backgrounds, character expressions, and movements for each cut (scene), scripts for creating dialogue (talking bags, voice scripts) for each scene, and background sound / sound effects. Audiobooks require a script for recording and tags for generating sound effects, while webtoons / videos may require scripts for scene-by-scene storyboarding, talk bags, and sound effect text generation. Furthermore, tags containing timing are required to synchronize separately generated scripts / storyboards and background sound / sound effect scripts. Therefore, the AI content creation device (120) can perform automated operations to perform the above operations according to a preset method, such as programming.
[0050] The AI content creation device (120) can also perform data post-processing operations to maintain consistency in content creation. This involves reintroducing key information from existing data into the same intellectual property (IP). For example, if an unidentified assassin appearing early in a novel is later revealed to be the protagonist's half-brother, this could involve maintaining consistency in voice casting (audiobook) and character depiction (webtoon / video). Another example would be a bright and healthy prince who participates in a war with a neighboring country and, three years later, transforms into a cold and pessimistic emperor. The prince and emperor should have the same basic voice (audiobook) and facial features and drawing style (webtoon / video). However, after three years, the prince's cynical personality should be reflected in a low-pitched voice, and his basic expression should be expressed in an annoyed and angry drawing style (webtoon / video). Thus, in embodiments of the present invention, data post-processing operations can maintain consistency in content creation within a single work.
[0051] Furthermore, the artificial intelligence content production device (120) according to an embodiment of the present invention can perform operations related to specifying and fixing model-specific option parameters to maintain differentiation. For example, in the case of specifying the drawing style, this may include Example 1) setting the blonde hair color as a combination of bright (#FFECD1), middle (#EFD679), and dark (#D3AD4D) colors, Example 2) specifying the eye width / height ratio: 1.6:1, 2:1, Example 3) specifying the height of the eye tail: Slanted eyes (sharp impression, eyes with upturned eye tails) - corner of the eye (located at the top 50% of the vertical length of the eye), tail of the eye (located at the top 20% of the vertical length of the eye), Kind eyes (kind impression, eyes with droopy eye tails) - corner of the eye (located at the top 50% of the vertical length of the eye), tail of the eye (located at the top 70% of the vertical length of the eye), etc. Also, in the case of guide voice and image input, examples include 1) zero-shot TTS: inputting and differentiating the main character's voice for a short amount of time, about 1 minute, 2) creating differentiated sound sources with the same tone of voice using technology (RVC, etc.) that changes the music AR voice to the singer of my choice, 3) inputting a guide image (cover image) - 1 picture drawn by an artist: learning the atmosphere of the background of the work and the character drawing style to differentiate the drawing style for each work, etc.
[0052] As described above, the artificial intelligence content production device (120) according to the embodiment of the present invention may be configured to include an audiobook model, a webtoon model, and a video model for generating each content of audio, webtoon, and video, respectively, and when text is input (for example, paid text is input after purchase), the input text data may be used to perform service operations such as (1) data generation, dialogue / narration separation, sentence separation, appearance and background description keyword extraction, (2) script and storyboard generation, (3) post-processing such as generation, synchronization, and synthesis in the form of content, and (4) content download / posting on a platform.
[0053] If the input text data is text data related to a novel, the AI content creation device (120) according to an embodiment of the present invention can perform operations for generating data, scenarios, and final content using AI text analysis technology for content creation, as shown in FIG. 2. As described above, when a desired service is selected through a service screen provided to the user terminal device (100) and text is input, the AI content creation device (120) can generate and provide various types of content, such as audiobooks, webtoons, animations / movies, and games, by changing data such as characters, titles, and personalities, as well as scenarios such as character casting and character-title matching, according to the service request. In this process, a TTS model can be executed in the case of audiobook creation. Additionally, an image creation model can be executed in the case of webtoon content, and a video creation model can be executed in the case of animation / movie content. In the case of games, a 3D object creation model can be executed.
[0054] In other words, the AI content production device (120) according to an embodiment of the present invention can include multiple AI models for generating various types of AI content and execute each AI model. Each AI model is an AI model, i.e., a program, specialized (or trained by learning different learning data) for generating different types of content, and each AI model may have different SW modules or algorithms for performing detailed operations, but above all, the type of learning data for generating each AI content may be different.
[0055] For example, the AI content creation device (120) according to an embodiment of the present invention can be equipped with an AI model such as a Large Language Model (LLM) and perform an operation to create an audiobook by utilizing the same. The AI content creation device (120) analyzes text data received from a third-party device (130) or pre-stored text data to identify who the speaker of the dialogue within the text is from a text (e.g., a set of sentences), and can also extract sound keywords (e.g., verbs, onomatopoeia, etc.) and background keywords (e.g., temporal background, spatial background) within the text, and can create an audiobook by combining the two. Of course, in this process, the AI content creation device (120) can analyze the text data by utilizing an AI module such as LLM and synthesize voice data based on the analysis results to create an audiobook. In this process, a developer or administrator sets a command through a prompt, and of course, in the embodiment of the present invention, the prompt can be automatically generated through an AI model for prompt generation, and data analysis or voice data synthesis can be performed accordingly.
[0056] Text data such as novels may include text for narration and dialogue-related lines. Of course, since text related to dialogue can be marked with double quotation marks, etc., it is entirely possible to distinguish lines through this. The AI content production device (120) first separates the text and dialogue to create an audiobook, and then confirms who the dialogue belongs to before synthesizing it into a specific voice. To this end, the AI content production device (120) can identify lines using an AI model that identifies relationships between words and sentences, or more precisely, relationships between objects and sentences within words. Alternatively, it is entirely possible to determine characters or characters based on their relationships with surrounding dialogue.
[0057] More specifically, the AI content production device (120) can extract character names from sentences, convert them into data, and manage them. In addition, for each sentence, the correlation between the "character name data" and the sentence can be inferred (or predicted), and if there is no valid character, the character name "Unknown" can be selected. In addition, the AI content production device (120) can learn in advance the target sentence, surrounding sentences (e.g., n sentences before and after the target sentence), and the character name spoken in the target sentence as training data for the AI model. Through this, the AI content production device (120) can generate an AI model that receives sentences and surrounding sentences as input and infers the character name, or execute the pre-generated model.
[0058] In addition, the AI content production device (120) can classify keywords that can make sounds (e.g., verbs: knock, open, hit, drop, etc. / onomatopoeia: whirrrik, swoosh, boom, tak, thump, etc.) within the text or dialogue that constitutes the text data through data analysis, and can specifically identify verbs and onomatopoeia related to sounds. To this end, the AI content production device (120) can infer the relationship between related sound keywords and sentences. The AI content production device (120) can perform an operation to extract background keywords, and such backgrounds can be used when synthesizing character voices or voices related to sound keywords. In order to extract background keywords from the text, the AI content production device (120) can operate with a concept such as a subset of the intersection of named entity recognition and semantic role determination methods in a broad sense. In other words, the AI content production device (120) according to an embodiment of the present invention can infer the relationship between a sentence and a background keyword with a single operation in this part. For this purpose, LLM can be utilized. Currently, many semantic role determination (SRL) and named entity recognition (NER) models capable of identifying words within sentences have been developed, and these can be utilized in embodiments of the present invention.
[0059] The AI content production device (120) according to an embodiment of the present invention can separate verbs, adverbs, and nouns from a sentence by utilizing a morphological analysis model for keyword identification. Many Korean morphological classification models capable of separating morphemes from sentences, as well as semantic role determination (SRL) and named entity recognition (NFR) models capable of identifying words within a sentence, have been developed and are widely available on the market. Therefore, it is entirely possible to utilize such analysis models in the embodiment of the present invention. The AI content production device (120) can use the model to determine and extract the sound relevance of separated verbs and adverbs, and can also determine the spatial background based on nouns. In this process, the AI content production device (120) can learn data in advance, such as target sentences, keywords of the target sentences, and keyword types, as training data, i.e., learning data. Through this, the AI content production device (120) can generate an AI model that receives a sentence as input and infers main keywords. Alternatively, the pre-generated AI model can be executed.
[0060] The third-party device (130) may refer to a computer of an administrator who operates the service of the AI content creation device (120) of FIG. 1, but may also include a server of a program development company or a computer used by a developer who develops and installs a program for providing services to the AI content creation device (120). For example, if the third-party device (130) is a server or computer of a developer or a developer, it may be connected to and performed by the AI content creation device (120) of FIG. 1. Representative examples thereof may include actions for generating prompts. In the case of a prompt, this may mean that the developer provides a kind of guideline or inputs setting information so that an action is performed accordingly. However, in the embodiment of the present invention, it is preferable to automatically generate and use prompts through an AI model. In this way, the third-party device (130) may essentially function as a device for generating audiobooks, webtoons, or videos. However, in the embodiment of the present invention, since the artificial intelligence content creation device (120) is an expensive device, it is preferable to perform various types of content creation operations according to the embodiment of the present invention in the artificial intelligence content creation device (120) under the assumption that it can be equipped with an artificial intelligence program such as LLM.
[0061] In addition to the above, the user terminal device (100), communication network (110), artificial intelligence content production device (120), and third-party device (130) of FIG. 1 can perform various operations, and since related contents will be continuously covered later, detailed contents will be replaced with those contents.
[0062] Figure 3 is a block diagram illustrating a detailed structure of the artificial intelligence content production device of Figure 1.
[0063] As illustrated in FIG. 3, the artificial intelligence content production device (120) of FIG. 1 according to an embodiment of the present invention includes part or all of a communication interface unit (300), a control unit (310), an artificial intelligence content production unit (320), and a storage unit (330).
[0064] Here, “including some or all” means that some components, such as the storage unit (330), may be omitted to configure the artificial intelligence content production device (120), or some components, such as the artificial intelligence content production unit (320), may be integrated into other components, such as the control unit (310), etc. In order to help a sufficient understanding of the invention, it is described as including all.
[0065] The communication interface unit (300) communicates with the user terminal device (100) and the third-party device (130) via the communication network (110) of Fig. 1. In the process of performing communication, the communication interface unit (300) can perform operations such as modulation / demodulation, muxing / demuxing, encoding / decoding, encryption / decryption, and scaling to convert resolution, which are obvious to those skilled in the art, and thus further description thereof will be omitted.
[0066] The communication interface unit (300) can provide a service screen for producing content generated by an artificial intelligence program under the control of the control unit (310) in response to a service request from a user terminal device (100). In addition, text data (e.g., novels and non-novel text data, etc.) provided through the service screen provided to the user terminal device (100), and type information related to the type of content to be generated using the text data, etc. can be received and transmitted to the control unit (310). Furthermore, the communication interface unit (300) can provide artificial intelligence content generated using text data provided by the user to the user terminal device (100).
[0067] The control unit (310) may include a processor such as a CPU, MPU, or GPU, and may perform overall control operations of the communication interface unit (300), the control unit (310), the artificial intelligence content production unit (320), and the storage unit (330) of FIG. 3. When a service request is made from a user terminal device (100), the control unit (310) may control the communication of the communication interface unit (300) to execute a graphic (GUI) program installed in the artificial intelligence content production unit (320) to provide a service screen of a specified layout. In addition, the control unit (310) may receive text data and type information of content creation type provided through the service screen from the communication interface unit (300), temporarily store them in the storage unit (330), and then retrieve them to provide them to the artificial intelligence content production unit (320). Thereafter, the control unit (310) can perform an operation to provide AI content, such as webtoons or video content, produced based on text data and type information related to content creation provided to the AI content production unit (320) and provide the content to the user terminal device (100) of FIG. 1.
[0068] Of course, the control unit (310) may perform an operation to build an artificial intelligence program, i.e., a model, in the artificial intelligence content production unit (320) in order to produce and provide services according to an embodiment of the present invention, i.e., artificial intelligence secondary content such as text data-based audio, webtoons, and videos, and the artificial intelligence model may be preceded by various operations such as designating and fixing model-specific optional parameters. For example, an operation to learn (or train) learning data for content production may be performed. Here, the model may be interpreted in various ways, but it can be understood as a software unit that is completed to perform a specific function among artificial intelligence programs. Of course, the embodiment of the present invention will not be particularly limited to such a concept.
[0069] The AI content creation unit (320) can perform operations to automatically generate input text data into content desired by the user, such as audio, webtoons, and videos. To this end, the AI content creation unit (320) according to an embodiment of the present invention can implement AI technology (or functions) for creating audiobooks, webtoons, and videos from novel texts, as well as AI technology for summarizing non-novel texts, extracting sentences, and then creating content. While the "AI" herein may refer to generative AI, as mentioned above, it can also be viewed as an improved AI capable of creating complete content by overcoming the limitations of general generative AI models. In an embodiment of the present invention, it can be viewed as executing an oddball creation AI that creates such complete content. The oddball creation AI (model) converts a text into general-purpose data and automatically generates various content, such as audiobooks, webtoons, and videos, using the converted general-purpose data. In this process, the input text data is analyzed to generate prompts, and based on the generated prompts, the AI model for each content creates content according to the user's request. For example, the AI content production department (320) analyzes text data and, if the data is not a novel, does not perform a separate character analysis operation. Instead, it summarizes only the parts to be made into content, and (based on the summary) separates sentences to perform a prompt generation operation (based on each sentence). Based on the generated prompts, content of each type can be produced.
[0070] The artificial intelligence content production unit (320) can be equipped with and execute multiple AI models. Of course, a software manager for controlling each AI model can execute a specific AI model at the request of the processor of the control unit (310). Of course, there is no particular limitation thereto. The AI model according to an embodiment of the present invention can be equipped with a first AI model (e.g., gpt) that converts input text data into general-purpose data, a second AI model (e.g., gpt, llama, other LLM, etc.) that converts (or generates) the converted general-purpose data into prompts (text) for each content, and a third AI model that generates each content according to the prompts (or commands of the prompts). Here, the third AI model can be configured with a plurality of third AI models that are configured in parallel and operate respectively. The third AI model can include an audiobook model that generates audio, a webtoon model that generates images such as webtoons, and a video model that generates videos.
[0071] The odd-even service according to an embodiment of the present invention improves upon existing technologies in terms of large-capacity text processing, novel content specialization, content production automation, and generative AI versatility compared to existing services, and automates the creation of final content so that it is easy for ordinary people, not just AI experts, to use.
[0072] The storage unit (330) can temporarily store various types of information or data processed under the control of the control unit (310). Here, the terms "information" and "data" are used interchangeably in practice, but can be distinguished, for example, as text data and content creation type information related to what type of content the text data will be created into. When a user of the user terminal device (100) requests content creation, the storage unit (330) can temporarily store text data and type information related to content creation and then provide them to the artificial intelligence content creation unit (320).
[0073] In addition to the above, the communication interface unit (300), control unit (310), artificial intelligence content production unit (320), and storage unit (330) of FIG. 3 can perform various operations, and other detailed information has been sufficiently explained above, so those contents will be used instead.
[0074] Meanwhile, the communication interface unit (300), control unit (310), artificial intelligence content production unit (320), and storage unit (330) of FIG. 3 according to an embodiment of the present invention are configured as physically separate hardware modules, but each module may store and execute software for performing the above operations. However, the software is a collection of software modules, and each module may be formed of hardware, so there is no particular limitation to the configuration, such as software or hardware. For example, the storage unit (330) may be hardware, such as storage or memory. However, since it is also possible to store information (repository) in software, there is no particular limitation to the above.
[0075] Meanwhile, as another embodiment of the present invention, the control unit (310) may include a CPU and a memory, and may be formed as a single chip. The CPU may include a control circuit, an operation unit (ALU), a command interpretation unit, and a registry, and the memory may include a RAM. The control circuit may perform a control operation, the operation unit may perform an operation of binary bit information, and the command interpretation unit may perform an operation of converting a high-level language into machine language and vice versa, including an interpreter or a compiler, and the registry may be involved in software data storage. According to the above configuration, for example, at the beginning of the operation of the artificial intelligence content creation device (120) of FIG. 1, a program stored in the artificial intelligence content creation unit (320) may be copied and loaded into memory, i.e., RAM, and then executed, thereby rapidly increasing the data operation processing speed. In the case of a deep learning model, it may be loaded into the GPU memory instead of the RAM and executed by accelerating the execution speed using the GPU.
[0076] FIG. 4 and FIG. 5 are diagrams showing an artificial intelligence content production process according to an embodiment of the present invention, FIG. 6 and FIG. 7 are flowcharts illustrating an operation process of an AI model performing text analysis, FIG. 8 is a flowchart illustrating an operation process of a webtoon creation model, and FIG. 9 is a flowchart illustrating an operation process of an image creation model.
[0077] For convenience of explanation, referring to FIGS. 4 to 9 together with FIG. 1, the AI content production device (120) according to an embodiment of the present invention can receive text and type information of the type of content to be produced (e.g., audiobook, webtoon, video, etc.) from the user terminal device (100) of FIG. 1. More precisely, the AI content production device (120) can input the received text data and type information into the AI model (S400, S501). Here, the text data can utilize user input text or web novel text (e.g., text sold within the platform). The user input can be input in the form of entering text on the web (e.g., on a web screen) or uploading a file such as a .txt file.
[0078] Next, the AI content production device (120) can analyze the input text and perform an operation to generate metadata (S410). A processing operation can be performed on the input text data. Here, processing can mean conversion from text to text data. Through processing, operations such as text analysis, character analysis, and data conversion can be performed. Here, text data can mean metadata for AI model prompts such as JSON rather than natural language. Text analysis can be performed differently for plot, context identification, and sentence-by-sentence division depending on ① whether it is a novel or ② whether it is a text other than a novel. For example, in the case of character analysis, ① for a novel, analysis can be performed on characters 1, 2, and 3, narrator, and ② for an explanation, analysis can be performed on the gender of the speaker and the age group of the explanation target (e.g., general current affairs knowledge - 20-40s, kiosk usage method for the elderly - 60s). In the case of data conversion operations, conversion can be done into data required for audiobooks / webtoons / videos, ① audiobooks can be converted into scripts composed of characters, text, and dialogues, and sound effect / background sound tags, ② webtoons can be converted into the form of full cut distribution, main character setting (head, face, costume description), scene-by-scene storyboards based on scene keywords for each cut, picture composition, and dialogue, ③ videos can be converted into the form of character, scene description (character expression, movement, background) scripts for 30-second units, background sound, and sound effects.
[0079] Furthermore, the AI content creation device (120) can perform a content conversion (text data to contents) operation (S420, S430, S440, S450). Here, the text data refers to metadata for prompts. The AI content creation device (120) can create audiobooks, webtoons, and video content through content conversion. Audiobooks can be created by automatically generating sound sources (TTS model) by character / narration, and then adding background music and sound effects to the existing sound sources. Webtoons can be created by creating characters and background images by scene (image creation AI model), and creating speech bubbles and synthesizing them. Videos can be created by creating videos by 30-second scene descriptions, and then generating audio sources using TTS to add them. In addition, the AI content creation device (120) can perform content provision and download operations, and can perform operations such as uploading created content by sellers, pricing by consumers, paid purchases by consumers, and possession within the platform.
[0080] For example, when an audiobook generation model is selected as shown in FIG. 4, the artificial intelligence content creation device (120) can perform detailed operations of S431 to S435. ① The reader purchases and selects (start / end designation) all / part of a romance web novel text (prior author consultation, platform profit sharing contract completed) for which the reader wants to modify the content. ② The selected text is entered. ③ The web novel character characteristics, narration, background music, sound effects, etc. are generated as data from the text. At this time, the collection and generation of additional character appearance description data for webtoon creation (hair color, eye color, clothing, accessories, posture, facial expression, etc.) are also included. ④ Based on the generated data, an appropriate AI VOICE is cast and assigned to the TTS model to generate a voice sound source. ⑤ Among the generated data, background music and sound effects are searched for in the DB or newly synthesized. ⑥ The sound sources of steps 4 and 5 are combined in consideration of synchronization to generate the final audiobook content.
[0081] In addition, when a webtoon generation model is selected, the artificial intelligence content creation device (120) can perform detailed operations of S441 to S445. ① The reader purchases and selects (start / end designation) all / part of a romance web novel text for which he or she wants to modify content (e.g., prior author consultation, completion of platform profit sharing contract). ② Enter the selected text. ③ Generate data such as web novel character traits, narration, background music, sound effects, etc. from the text → Collect and create additional character appearance description data for webtoon creation (hair color, eye color, clothing, accessories, posture, facial expression, etc.). ④ Create a prompt for each image generation model based on the generated data (the text description prompt style that generates images well differs for each model, such as midjourney, stable diffusion, etc.). ⑤ Create an image for each cut by assigning it to the corresponding image generation AI model. ⑥ Create a speech bubble dialogue and background dialogue image as a GIF using dialogue for each cut. ⑦ Create an integrated image with each layer of the images generated in 5 and 6. ⑧ By naturally connecting each cut, it creates a long webtoon image that can be scrolled.
[0082] Furthermore, the artificial intelligence content production device (120) can perform detailed operations of S451 to S455 when a video creation model is selected. ① The reader purchases and selects (start / end designation) all / part of the romance web novel text (prior author consultation, platform profit sharing contract completed) for which he / she wants to modify content. ② The selected text is entered. ③ The web novel character characteristics, narration, background music, sound effects, etc. are generated as data from the text. At this time, the collection and generation of additional character appearance description data for video creation (hair color, eye color, clothing, accessories, posture, facial expression, etc.) are also included. ④ The text that will compose a 30-second video - scene description (character, background, sound effect description) is generated. ⑤ The dialogue, background music, and sound effect sound sources suitable for the video are generated. ⑥ The video and sound sources generated in 4 and 5 are synchronized and integrated.
[0083] Figures 5 to 9 are flowcharts illustrating the process illustrated in Figure 4 in more detail. As seen in Figure 5, the artificial intelligence content production device (120) of Figure 1 can determine whether or not an input text is a novel by determining the type of the input text (S501, S502).
[0084] For example, in the case of a novel, AI models 1-1 and 1-2 can be executed to perform actions such as sentence separation, analysis, and prompt generation (S503), and in the case of a non-novel, AI models 1-1 (non-novel) and 1-2 can be executed to perform actions such as sentence separation, summary, and prompt generation (S504). Figure 6 shows in detail the actions of sentence separation, analysis, and prompt generation in AI model 1. In addition, Figure 7 shows in detail the process of analyzing text in AI model 1. AI model 1-1 (non-novel) can analyze input text data to summarize the text, generate summary data, separate sentences in the generated summary data, and write prompts (for example, corresponding to the separated sentences) to provide them to AI model 2.
[0085] In addition, the AI content creation device (120) can execute AI model 1-2 to perform operations such as tokenization, inference, and data creation on the analysis results generated in steps S503 and S504 (S505), and can also determine the type of output content (S506) and generate a prompt accordingly. A token may refer to a small unit that the AI learns, and a token may be a word, phrase, or sentence. The process of dividing speech into such small units may be referred to as tokenization. The AI content creation device (120) can determine the type information input together with the text data to determine the type of output content. For example, if a user requests the creation of an audiobook using text data, a prompt for using the audiobook creation model can be generated. For example, an audiobook script can be generated. Steps S507a to S507c can be viewed as a prompt creation process for applying to the AI model for creating each content.
[0086] As shown in S508a to S508c, depending on the type of content to be created, audiobook script data is generated for audiobooks, and webtoon storyboard data is generated for webtoons. Additionally, video scenario data can be generated for videos. In embodiments of the present invention, this may be included within the scope of the prompt generation process.
[0087] In addition, the artificial intelligence content production device (120) can determine whether audiobook script data, webtoon storyboard data, and video scenario data are customized, that is, whether they are personalized, and add custom data if personalization is necessary. In the process, data such as optional parameters, guide voices, and guide images can be utilized (S509, S511, S512, S513).
[0088] And, using the data generated in S508a to S508c above and the added custom data, each generation model, i.e., the audiobook generation model, the webtoon generation model, and the video generation model, can generate and output content based on the initially input text data (S510a, S510b, S510c).
[0089] Fig. 8 clearly shows the detailed operations of the webtoon generation model that leads to models 1, 2-2, and 3-2 (S800 to S812). As previously explained in Fig. 4, the webtoon generation model can output the final completed content through detailed operations such as webtoon storyboard generation, image generation for each cut, speech bags for each cut, other text generation, and image and text data integration. For example, in the case of novel text data, operations such as character profile generation, dialogue-character matching, dialogue emotional line analysis, sound effect keyword detection, and background description keyword detection can be performed, and respective data can be generated accordingly (S801a to S801e, S802a to S802e).
[0090] In addition, Fig. 9 clearly shows the detailed process of the video generation model that leads to Model 1, 2-3, and 3-3 (S900 to S911). As already explained in Fig. 4, in the case of video generation, the final completed video content can be created and output through actions such as creating a scenario for each video scene, creating multiple 30-second videos for each scene, creating multiple sound sources for each scene, and synchronizing 30-second videos and sound sources.
[0091] The detailed operations in FIGS. 6 to 9 are performed according to a process preset in the program, and since the detailed information has already been sufficiently explained above, we will replace it with that information.
[0092] Figure 10 is a flowchart showing the operation process of the artificial intelligence content production device of Figure 1.
[0093] For convenience of explanation, referring to FIG. 10 together with FIG. 1, an AI content production device (120) according to an embodiment of the present invention receives text data for generating content using an AI program from a user terminal device (100) (S1000). Furthermore, the AI content production device (120) can input the received text data into an AI model.
[0094] In addition, the AI content creation device (120) can convert the received text data into general data (e.g., converting data using a gpt model), convert the converted general data into text (or prompt) for each content-specific prompt (e.g., gpt, llama, other LLM, etc.), and generate and output content such as audiobooks, webtoons, or videos using an AI model that generates each content based on the converted prompt (text) for each content (S1010). This operation can be considered to correspond to the operation of FIG. 4 or FIG. 5.
[0095] Above all, the AI content creation device (120) according to an embodiment of the present invention can execute AI model 1-1 or AI model 1-1 (non-fiction) to analyze the input data depending on whether the input text data is a novel. Here, the two models corresponding to AI model 1-1 can be understood as AI models specialized in analyzing novel text and non-fiction text, respectively. In other words, the AI content creation device (120) can be equipped with a first AI model that converts text into general data, a second AI model that converts general data into text for prompts for each content, and a third AI model that generates each content based on the prompt, and executes them to create and output content desired by the user. A model that provides a service according to an embodiment of the present invention by combining these various types of models can be named an 'oddory creation AI model' in an embodiment of the present invention. In an embodiment of the present invention, an oddory service can be provided using the oddory creation AI model.
[0096] In addition to the above, the artificial intelligence content production device (120) of FIG. 1 can perform various operations, and other detailed information has been sufficiently explained above, so it will be replaced with that information.
[0097] Even though all components constituting the embodiments of the present invention have been described as being combined or operating in combination, the present invention is not necessarily limited to such embodiments. That is, within the scope of the present invention, all of the components may be selectively combined and operated one or more times. In addition, although all of the components may be implemented as individual hardware, some or all of the components may be selectively combined and implemented as a computer program having program modules that perform some or all of the functions of the combined hardware in one or more pieces. The codes and code segments constituting the computer program will be readily inferred by those skilled in the art. Such a computer program may be stored in a non-transitory computer-readable storage medium and read and executed by a computer, thereby implementing the embodiments of the present invention.
[0098] Here, the non-transitory readable storage medium refers to a medium that permanently stores data and can be read by a device, rather than a medium that stores data for a short period of time, such as a register, cache, or memory. Specifically, the above-described programs may be stored and provided on a non-transitory readable storage medium, such as a CD, DVD, hard disk, Blu-ray disc, USB, memory card, or ROM.
[0099] Although embodiments of the present invention have been described with reference to the attached drawings, those skilled in the art will appreciate that the present invention can be implemented in other specific forms without altering the technical spirit or essential characteristics of the present invention. Therefore, the embodiments described above should be understood to be illustrative in all respects and not restrictive.
[0100] [Explanation of symbols]
[0101] 100: User terminal device 110: Communication network
[0102] 120: AI content creation device 130: Third-party device
[0103] 300: Communication interface section 310: Control section
[0104] 320: AI Content Production Department 330: Storage Department
Claims
1. A communication interface unit that receives text data for generating content using an artificial intelligence (AI) program from a user terminal device; and A control unit that converts the received text data into multi-use data, converts the converted universal data into text for prompting each content, and generates and outputs content of an audiobook, webtoon, or video using an AI model that generates each content based on the converted text for prompting each content; including, The above control unit executes a first AI model to perform a conversion operation into the general data, executes a second AI model to perform a conversion operation into the text for the prompt, and executes a third AI model to perform a generation operation of the content. The above control unit connects and operates multiple third AI models in parallel to generate each content based on the converted text for the prophets. The above control unit is an artificial intelligence content production device that executes the AI model to perform a character analysis operation if the received text data is text related to a novel, and if the text is text other than a novel, does not perform the character analysis operation but summarizes only the portion to be made into content, and separates the summarized sentences to generate a prompt.
2. In paragraph 1, The above control unit is an artificial intelligence content production device that executes an audiobook model for generating the audiobook, a webtoon model for generating the webtoon, and a video model for generating the video as the plurality of third AI models.
3. A step in which the communication interface unit receives text data for generating content using an artificial intelligence (AI) program from a user terminal device; and A step of the control unit converting the received text data into multi-use data, converting the converted multi-use data into text for prompts for each content, and using an AI model that generates each content based on the text of the converted prompts for each content, and generating and outputting the content of an audiobook, webtoon, or video; including, The above generating and outputting steps are: The first AI model is executed to perform the conversion operation into the general data, the second AI model is executed to perform the conversion operation into the text for the prompt, and the third AI model is executed to perform the generation operation of the content. The above generating and outputting steps are: Multiple third AI models are connected in parallel and operated to generate each content based on the text for the above-mentioned converted prompt. The above generating and outputting steps are: A method for operating an artificial intelligence content production device, which executes the above AI model to perform character analysis if the received text data is text related to a novel, and does not perform character analysis if the text data is text other than a novel, but summarizes only the portion to be made into content, and generates a prompt by separating the summarized sentences.
4. In paragraph 3, The above generating and outputting steps are: A method for operating an artificial intelligence content production device, which executes an audiobook model for generating the audiobook, a webtoon model for generating the webtoon, and a video model for generating the video as the plurality of third AI models.
Citation Information
Patent Citations
Solder alloy, solder paste, solder ball, solder joint, and electronic device comprising the solder joint
KR1020240155028A
method and apparatus for automatically generating training dataset for FAQ and chatbot based on natural language processing using deep learning
KR102436549B1
Moving quarantine device
KR102464985B1
Apparatus for Providing Artificial Intelligence Content Production Service and Driving Method Thereof
KR102693273B1