Electronic apparatus and method for generating video content using a generative artificial intelligence model-based digital content

The use of a generative AI model in an electronic device simplifies the conversion of digital content into video by automating the process, reducing time and skill requirements.

JP2026060948APending Publication Date: 2026-04-08KAKAO ENTERTAINMENT CORP
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2025-09-29
Publication Date
2026-04-08

AI Technical Summary

Technical Problem

The process of converting digital content into video content is time-consuming and requires specialized skills, making it difficult for non-experts to produce video content efficiently.

Method used

An electronic device and method utilizing a generative artificial intelligence model to identify cut images and text from digital content, generate narration, select matching images, and incorporate video asset information to produce video content quickly and easily.

Benefits of technology

Reduces the time required to convert digital content into video and enables non-experts to create video content effortlessly.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026060948000001_ABST
    Figure 2026060948000001_ABST
Patent Text Reader

Abstract

To provide an electronic device and method for generating video content using digital content more quickly and easily. [Solution] The method involves inputting a first prompt, which includes multiple cut images identified using digital content, text corresponding to the multiple cut images, and analysis infrastructure information, into a generative artificial intelligence model to acquire analysis information about the digital content; inputting a second prompt, which includes video asset information including multiple cut images, text, and analysis information, and narration generation infrastructure information, into the generative artificial intelligence model to acquire narration information consisting of multiple sentences; inputting a third prompt, which includes video asset information and matching infrastructure information, into the generative artificial intelligence model to select at least one cut image from the multiple cut images that matches each sentence; and generating video content using the multiple cut images and narration information according to the selection results.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an electronic device and method for generating video content using a generative artificial intelligence model-based digital content.

Background Art

[0002] Marketing strategies for expanding the digital content market such as webtoons and web novels and attracting new readers have entered a new phase with the explosive growth of online video sharing platforms such as YouTube (registered trademark).

[0003] As one of the marketing strategies, the process of producing digital content into video content can generally be roughly divided into four stages: 1) synopsis and storyboard production, 2) production of image assets for video, 3) subtitle / narration work, and 4) video production work, etc.

[0004] To produce one video through such a process, workers have to read and understand digital content and need skills to handle editing programs, etc. The production hurdles are high and it takes at least two to three weeks, which is a difficult point.

[0005] On the other hand, generative artificial intelligence (GAI) is one of the artificial intelligence technologies that uses a deep learning model trained based on a large dataset to generate new content.

[0006] With the emergence of generative artificial intelligence models, new attempts to produce video content from digital content using these have become possible.

Summary of the Invention

Problems to be Solved by the Invention

[0007] The object of the present invention is to provide an electronic device and method for generating video content using digital content more quickly and easily. [Means for solving the problem]

[0008] An electronic device that generates video content using a generative artificial intelligence model-based digital content according to one embodiment of the present invention may include a processor that uses the digital content to identify a plurality of cut images and text corresponding to the plurality of cut images, inputs a first prompt including the plurality of cut images; the text; and analysis base information to a generative artificial intelligence model to obtain analysis information regarding the digital content, inputs a second prompt including video asset information including the plurality of cut images, the text; and narration generation base information to a generative artificial intelligence model to obtain narration information consisting of a plurality of sentences, inputs a third prompt including the video asset information; and matching base information to a generative artificial intelligence model to select at least one cut image from the plurality of cut images that matches each of the sentences, and generates video content using the plurality of cut images and the narration information according to the selection result.

[0009] The aforementioned analysis information may be analysis information relating to at least one character or scene from the cut images of the digital content.

[0010] The processor can input a fourth prompt, including the video asset information and the synopsis generation base information, into a generative artificial intelligence model to obtain synopsis information about the digital content.

[0011] The processor can input a fifth prompt, including the video asset information and character analysis infrastructure information, into a generative artificial intelligence model to acquire character information relating to digital content.

[0012] The processor can convert the narration information into audio content.

[0013] The matching base information may include at least one candidate cut image that matches each of the sentences and a score for the candidate cut image.

[0014] The processor can input a sixth prompt, including the video asset information and usage conditions for each image effect, into a generative artificial intelligence model to select an image effect corresponding to each of the multiple cut images.

[0015] The processor can identify at least one keyword for searching for background music for the video content based on the video asset information, and can select background music from a sound source database based on the at least one keyword.

[0016] A method for generating video content using a generative artificial intelligence model-based digital content performed by an electronic device according to one embodiment of the present invention includes the steps of: identifying a plurality of cut images and text corresponding to the plurality of cut images using the digital content; inputting a first prompt including the plurality of cut images, the text, and analysis base information into a generative artificial intelligence model to obtain analysis information relating to the digital content; inputting a second prompt including video asset information including the plurality of cut images, the text, and the analysis information, and narration generation base information into a generative artificial intelligence model to obtain narration information consisting of a plurality of sentences; inputting a third prompt including the video asset information and matching base information into a generative artificial intelligence model to select at least one cut image from the plurality of cut images that matches each of the sentences; and generating video content using the plurality of cut images and the narration information according to the selection result.

[0017] The method may further include a step of inputting a fourth prompt, which includes the video asset information and the synopsis generation base information, into a generative artificial intelligence model to obtain synopsis information relating to digital content.

[0018] The method may further include a step of inputting a fifth prompt, which includes the video asset information and character analysis base information, into a generative artificial intelligence model to acquire character information relating to digital content.

[0019] The step of generating the video content may include the step of converting the narration information into audio content.

[0020] The step of generating the video content may include inputting a sixth prompt, which includes the video asset information and usage conditions for each image effect, into a generative artificial intelligence model to select an image effect corresponding to each of the multiple cut images.

[0021] The step of generating the video content may include the step of identifying at least one keyword for searching for background music for the video content based on the video asset information; and the step of selecting background music from a sound source database based on the at least one keyword. [Effects of the Invention]

[0022] According to one embodiment of the present invention, the time required to convert digital content into video is reduced, and even non-experts can easily and quickly create video content. [Brief explanation of the drawing]

[0023] [Figure 1] This is a schematic diagram illustrating the operation of an electronic device according to one embodiment of the present invention. [Figure 2] This is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present invention. [Figure 3]The drawing illustrates an operation flowchart of an electronic device according to an embodiment of the present invention. [Figure 4] The drawing illustrates a state of generating video content using digital content according to an embodiment of the present invention.

Embodiments for Carrying Out the Invention

[0024] Hereinafter, preferred embodiments according to the present invention will be described in detail with reference to the accompanying drawings. The detailed description disclosed below together with the accompanying drawings is intended to explain exemplary embodiments of the present invention, and is not intended to show the only embodiments in which the present invention can be implemented. Parts not related to the description for clearly explaining the present invention in the drawings can be omitted, and the same reference numerals can be used for the same or similar components throughout the specification.

[0025] FIG. 1 is a schematic diagram illustrating the operation state of an electronic device according to an embodiment of the present invention.

[0026] Referring to FIG. 1, an electronic device 100 according to an embodiment of the present invention is a device that generates video content 30 using a generative artificial intelligence model 10 (hereinafter, also referred to as model 10) - based digital content 20, and can be implemented by a computer, a server, a smartphone, a tablet PC, a smart pad, a notebook computer, or the like.

[0027] The generative artificial intelligence model 10 can be a language model trained to provide a reply corresponding to an input query, and can include, for example, large language models (LLMs) and lightweight language models (sLLMs).

[0028] At this time, the electronic device 100 can construct and use a generative artificial intelligence model 10, or it can receive and store a pre-constructed generative artificial intelligence model 10 from an external source and use it. Alternatively, the electronic device 100 can use a pre-constructed generative artificial intelligence model 10 that provides cloud infrastructure services via a network. The electronic device 100 is not limited to any one of these methods of utilizing the generative artificial intelligence model 10. Furthermore, the electronic device 100 can utilize two or more generative artificial intelligence models 10 in the process of generating video content 30 using digital content 20.

[0029] The electronic device 100 can construct one or more programs consisting of one or more computer-executable instruction words to generate video content 30 from digital content 20 using a generative artificial intelligence model 10.

[0030] In this invention, the digital content 20 may be story-based content including images, such as a webtoon. Furthermore, the digital content 20 may also include story-based content such as a web novel or a novel.

[0031] In this invention, the video content 30 is a video generated using the digital content 20, and may consist of a combination of images, subtitles, narration, image effects, and background music. The video content 30 can be generated in various forms, such as short form and long form. The video content 30 can be generated for various purposes, such as publicity, preview, announcement, and summary of the digital content 20. In this case, the generation form and generation purpose of the video content 30 are not limited to just one.

[0032] This invention proposes a method for automating the process of creating basic data for visualizing digital content 20 using a generative artificial intelligence model 10, so that this data can be easily visualized.

[0033] The configuration and operation of an electronic device 100 according to one embodiment of the present invention will be described in detail below with reference to the drawings.

[0034] Figure 2 is a block diagram illustrating the configuration of an electronic device according to one embodiment of the present invention.

[0035] An electronic device 100 according to one embodiment of the present invention may include an input unit 110, a communication unit 120, a display unit 130, a storage unit 140, and a processor 150.

[0036] The input unit 110 generates input data in response to user input to the electronic device 100. For example, user input may be user input to start the operation of the electronic device 100, user input to generate and tune prompts, user input to check, correct and confirm results obtained from the generative artificial intelligence model 10, and is applicable without limitation to any other user input necessary to generate video content 30 using digital content 20.

[0037] The input unit 110 includes at least one input means. The input unit 110 may include a keyboard, keypad, dome switch, touch panel, touch key, mouse, menu button, etc.

[0038] The communication unit 120 can communicate with external devices such as servers to send and receive digital content 20, cut images, text, analysis base information, analysis information, narration generation base information, narration information, matching base information, selection results, generative artificial intelligence model 10, etc.

[0039] For this purpose, the communication unit 120 can perform wireless communication such as 5G (5th generation communication), LTE-A (Long Term Evolution-Advanced), LTE (Long Term Evolution), Wi-Fi (Wireless Fidelity), and Bluetooth, as well as wired communication such as LAN (Local Area Network), WAN (Wide Area Network), and power line communication.

[0040] The display unit 130 displays display data generated by the operation of the electronic device 100. The display unit 130 can display the entire or a part of the process of generating video content 30 using the digital content 20, such as a screen for separating cut images from the digital content 20 and extracting text, a screen for acquiring narration information and synopsis information using video asset information such as multiple cut images, text, and analysis information, and a screen for selecting image effects, background music, etc., to be applied to the video content 30.

[0041] The display unit 130 includes liquid crystal displays (LCDs), light-emitting diode (LED) displays, organic light-emitting diode (OLED) displays, micro-electro-mechanical systems (MEMS) displays, and electronic paper displays. The display unit 130 can be combined with the input unit 110 to be implemented as a touchscreen.

[0042] The storage unit 140 stores the operating program for the electronic device 100. The storage unit 140 includes non-volatile storage that can store data (information) regardless of whether power is supplied, and volatile memory where data is loaded for processing by the processor 150 and data cannot be stored unless power is supplied. Storage can be flash memory, HDD (Hard-Disc Drive), SSD (Solid-State Drive), ROM (Read Only Memory), etc., and memory can be buffer, RAM (Random Access Memory), etc.

[0043] The storage unit 140 can store digital content 20, cut images, text, analysis base information, analysis information, narration generation base information, narration information, matching base information, selection results, generative artificial intelligence model 10, etc., and can store calculation programs necessary in the process of performing multiple cut image and text identification, acquisition of analysis information related to digital content 20, acquisition of narration information, selection of cut images that match the narration information, and video content generation.

[0044] The processor 150 can execute software such as programs to control at least one other component of the electronic device 100 (e.g., hardware or software component), and can perform various data processing or calculations.

[0045] A processor 150 according to one embodiment of the present invention can use digital content 20 to identify a plurality of cut images and text corresponding to the plurality of cut images, input a first prompt including the plurality of cut images; the text; and analysis base information to a generative artificial intelligence model 10 to obtain analysis information regarding the digital content 20, input a second prompt including video asset information including the plurality of cut images; the text; and narration generation base information to the generative artificial intelligence model 10 to obtain narration information consisting of a plurality of sentences, input a third prompt including the video asset information; and matching base information to the generative artificial intelligence model 10 to select at least one cut image from the plurality of cut images that matches each sentence, and generate video content 30 using the plurality of cut images and the narration information according to the selection result.

[0046] At this time, the processor 150 can either construct and use a generative artificial intelligence model 10, or receive and store a pre-constructed generative artificial intelligence model 10 from an external source and use it. Alternatively, the processor 150 can use a pre-constructed generative artificial intelligence model 10 that provides cloud infrastructure services via a network. The method by which the processor 150 utilizes the generative artificial intelligence model 10 is not limited to just one of these methods. Furthermore, the processor 150 can utilize two or more generative artificial intelligence models 10 in the process of generating video content 30 using digital content 20.

[0047] On the other hand, the processor 150 can perform at least part of the data analysis, processing, and result information generation necessary to carry out the above operations using at least one of the following as rule-based or artificial intelligence algorithms: machine learning, neural network, or deep learning algorithm. Examples of neural networks include models such as CNN (Convolutional Neural Network), DNN (Deep Neural Network), RNN (Recurrent Neural Network), and transformer.

[0048] Figure 3 is a diagram illustrating the operation flowchart of an electronic device according to one embodiment of the present invention.

[0049] A processor 150 according to one embodiment of the present invention can use digital content 20 to identify multiple cut images and text corresponding to the multiple cut images (S10).

[0050] The digital content 20 may be story-based content that includes images, or story-based content that does not include images. The processor 150 can receive the necessary digital content 20 from a database that stores the digital content 20 via the communication unit 120, or it can obtain it from a storage unit 140 that includes the database.

[0051] First, if the digital content 20 is story-based content that includes images, especially if it is a long image that is displayed by scrolling, such as in a webtoon, it is necessary to separate the overall image from multiple cut images. A cut image is an image that is displayed in at least one frame of the video content 30. In this case, the size of the cut image can be set in various ways to match the size of the video, and the center and size of the cut image can be readjusted after the cut image is generated so that it can be displayed in the video content 30 through object recognition or the like.

[0052] The processor 150 can separate an image within the digital content 20 into multiple cut images. A variety of methods can be used to generate the cut images. For example, the processor 150 can generate cut images based on a heuristic-based algorithm that generates cut images based on background color, or based on an image processing model that has been trained to separate images through object recognition within the overall image.

[0053] The processor 150 can identify speech bubbles and text written in the background within multiple cut images. A variety of methods can be used to identify the text; for example, the processor 150 can identify text based on optical character recognition (OCR) technology.

[0054] On the other hand, if the digital content 20 is story-based content that does not contain images, it is possible to generate cut images. For example, cut images can be generated using the digital content 20 to be visualized, and the portion corresponding to the cut image can be identified in text. Here, the portion corresponding to the cut image may mean the portion of the story-based content that does not contain images and was involved in the generation of the cut image. Various methods can be used to generate the cut images.

[0055] The following explanation assumes that, regardless of the type of digital content (20 types), a cut image and its corresponding text are acquired.

[0056] A processor 150 according to one embodiment of the present invention can input a prompt (hereinafter referred to as the first prompt) containing multiple cut images, text, and analytical base information to a generative artificial intelligence model 10 to obtain analytical information regarding digital content 20 (S20).

[0057] A prompt is a query input to the generative artificial intelligence model 10, and can be pre-configured to enable the generative artificial intelligence model 10 to output results effectively based on the given prompt. Prompts can generally include task information, background information, example information, persona information, etc. Task information is information about the task that the model 10 must perform, and may be "Generate this," "Analyze this," etc. Background information is information that provides background to the task so that the model 10 can perform the requested task more accurately. Example information is information that describes examples of the results that the model 10 will output. Example information may include the output format of the output results. Persona information is information about a virtual person or role assigned to the model 10, and for example, in the case of digital content generation, it may be "You are a creative digital content worker." Prompts can be continuously tuned to obtain better results in the process of generating video content 30 from digital content 20.

[0058] On the other hand, this stage is in which the generative artificial intelligence model 10 is trained to understand the content of the digital content 20 and acquires basic information (analytical information) for obtaining other information described later (also called the image captioning stage). The analytical information may include analytical information about at least one character or scene from the cut images of the digital content 20.

[0059] The first prompt for obtaining analytical information about the digital content 20 may include analytical background information. The analytical background information may include operational information requesting Model 10 to analyze multiple cut images and text. In addition, the analytical background information may include the aforementioned background information, example information, and persona information. At this time, the analytical background information may include operational information requesting not only individual analysis of the cut images but also the ability to understand the context between the cut images and analyze the story of the digital content 20.

[0060] The following multiple cut images, text, and analytical information are fundamental information for generating video content 30, and when integrated, they are referred to as video asset information.

[0061] A processor 150 according to one embodiment of the present invention can input a prompt (hereinafter referred to as the second prompt) including video asset information and narration generation base information to a generative artificial intelligence model 10 to acquire narration information consisting of multiple sentences (S30).

[0062] A second prompt for obtaining narration information related to digital content 20 may include narration generation base information. The narration generation base information may include operational information requesting narration to be inserted into video content 30 using video asset information. Similarly, the narration generation base information may also include the aforementioned background information, example information, and persona information.

[0063] A processor 150 according to one embodiment of the present invention can input a prompt (hereinafter referred to as the third prompt) including video asset information and matching base information to a generative artificial intelligence model 10 and select at least one cut image from a plurality of cut images that matches each sentence (S40).

[0064] A third prompt for selecting at least one cut image from multiple cut images that matches each sentence may include matching base information. The matching base information may include work information requesting the selection of at least one cut image that matches each sentence of the narration information using video asset information. Similarly, the matching base information may also include the background information, example information, and persona information mentioned above.

[0065] On the other hand, the process of selecting a cut image can be done in one step based on matching base information, but it is not limited to this and may involve multiple steps.

[0066] For example, the processor 150 can, through the model 10, select one or more candidate cut images from among multiple cut images that best match each sentence of the narration information, and assign a score to each candidate cut image. The score for each candidate cut image means a score assigned based on the relationship between each sentence and the candidate cut image. In this case, the matching base information can include at least one candidate cut image that matches each sentence and a score for that candidate cut image. In this case, the processor 150 can input a third prompt including video asset information and matching base information to the generative artificial intelligence model 10 to select at least one final cut image from among multiple cut images that matches each sentence.

[0067] A processor 150 according to one embodiment of the present invention can generate video content 30 using multiple cut images and narration information based on the selection result (S50).

[0068] The selection results are based on the prior selection at stage S40, where at least one cut image was chosen to match each sentence of the narration information.

[0069] The processor 150 can convert narration information into audio content. The process of converting to audio content can be diverse. For example, the processor 150 can convert narration information into audio content based on a Text-to-Speech (TTS) algorithm that converts text to speech. The TTS algorithm may be an artificial intelligence model that has been trained to output speech corresponding to the input text, and the artificial intelligence model may have been trained based on the voice of a specific person. Alternatively, the processor 150 can use audio content in which actual voice actors or others have dubbed the narration information.

[0070] The processor 150 can combine multiple cut images and audio content to generate video content 30.

[0071] At this time, in addition to multiple cut images and audio content, image effects and background music may also be inserted. This will be described with reference to Figure 4.

[0072] According to one embodiment of the present invention, the entire process of generating video content from digital content may be performed automatically, or it may include a step of verifying the output of any one of the stages and regenerating it as necessary.

[0073] According to one embodiment of the present invention, the time required to convert digital content into video is reduced, and even non-experts can easily and quickly create video content.

[0074] Figure 4 is a diagram illustrating how video content is generated using digital content according to one embodiment of the present invention.

[0075] Regarding the operation of generating video content 500 from digital content 410, the content previously described with reference to Figure 3 applies to the above content, so explanations of redundant content will be omitted.

[0076] First, the electronic device 100 can acquire video asset information 420 from the digital content 410. The video asset information 420 may include multiple cut images 421, text 422, and analysis information 423.

[0077] The electronic device 100 can acquire narration information 430 using video asset information 420. In another embodiment, the electronic device 100 can acquire narration information 430 by also considering synopsis information 440, which will be described later, in addition to the video asset information 420.

[0078] The synopsis information 440 can be used not only for narration information 430, but also for image matching and background music extraction. The synopsis information 440 can be converted to TTS and updated with narration information, and each sentence can be segmented to an appropriate length and subtitle information can be added. The segmentation unit can be determined by an artificial intelligence model.

[0079] The plot summary information 440 can also be used for image matching. Image matching can be performed by an artificial intelligence model that grasps the episode number, mentioned characters, associated emotions, etc. of the plot summary. The electronic device 100 can use the video asset information 420 to obtain the plot summary information 440 and character information 450.

[0080] Specifically, the electronic device 100 can input a prompt (hereinafter referred to as the fourth prompt) containing video asset information 420 and plot generation base information to the generative artificial intelligence model 10 to obtain plot information 440 relating to the digital content 410.

[0081] The fourth prompt for obtaining synopsis information 440 regarding digital content 410 may include synopsis generation base information. The synopsis generation base information may include operational information requesting model 10 to use video asset information 420 to create a synopsis of the digital content 410. In addition, the synopsis generation base information may include the aforementioned background information, example information, and persona information. At this time, the synopsis generation base information may include operational information requesting that the synopsis of the digital content 410 be analyzed not only by analyzing individual cut images, but also by understanding the context between cut images.

[0082] In another embodiment, the electronic device 100 can obtain synopsis information 440 by additionally considering character information 450, which will be described later, in addition to the video asset information 420.

[0083] Character information 450 is used to understand who the protagonist is, what their name is, their condition, personality, and appearance, and can be used to generate synopsis information 440 and for image matching.

[0084] Character information 450 can be extracted from the video asset information 420, and the main events, scenes, and characters of the synopsis can be determined based on the character information 450. Synopsis information 440 can be extracted focusing on the protagonist, and settings can be made to exclude mentions of unnecessary characters.

[0085] It can also be used to accurately match the image of a character mentioned in the synopsis information of character information 450. In this case, the character's status value, which changes as the story progresses, may be reflected. The character's status value can include external factors such as age, clothing, hair length, and accessories.

[0086] The electronic device 100 can input prompts (hereinafter referred to as the fifth prompt) containing video asset information 420 and character analysis base information to the generative artificial intelligence model 10 to obtain character information 450 related to the digital content 410. The character information 450 may include information about the appearance and personality of the characters appearing in the digital content 410.

[0087] The fifth prompt for obtaining character information 450 related to digital content 410 may include character analysis base information. The character analysis base information may include work information requesting model 10 to analyze the characters appearing in digital content 410 using video asset information 420. In addition, the character analysis base information may include the background information, example information, and persona information mentioned above. At this time, the character analysis base information may include work information requesting that the character of digital content 410 be analyzed not only by analyzing individual cut images, but also by understanding the context before and after the cut images.

[0088] The electronic device 100 can input video asset information 420 and prompts (hereinafter referred to as the sixth prompt) including usage conditions for each image effect 460 to the generative artificial intelligence model 10 to select an image effect 460 corresponding to each of multiple cut images.

[0089] Image effect 460 refers to an effect used to display multiple cut images 421 in video content 500, which may include, for example, zoom in, zoom out, left in, right in, etc.

[0090] The sixth prompt for selecting image effects 460 related to digital content 410 may include usage conditions for each image effect. These usage conditions include mandatory and recommended conditions. Mandatory conditions are those that must be satisfied in order to use the image effect in question. For example, these might include the size and aspect ratio of the cut image. Recommended conditions are those that make the image effect suitable for use in the cut image. For example, zoom in is recommended when the cut image emphasizes or enlarges a single person on screen. In this case, the usage conditions may include not only an individual analysis of the cut image, but also work information that allows for the selection of image effects 460 for each cut image by understanding the image effects used before and after the cut images. In addition, the sixth prompt may include the aforementioned background information, example information, and persona information.

[0091] In another embodiment, the electronic device 100 can select an image effect 460 for each cut image by additionally considering at least one of the following: narration information 430, synopsis information 440, and character information 450, in addition to the video asset information 420.

[0092] The electronic device 100 can identify at least one keyword for searching for background music for video content based on the video asset information 420.

[0093] Keywords can be identified within a predefined keyword list, or from the video asset information 420. The predefined keyword list may be divided into categories such as theme, genre, and mood. For themes, examples include adventure, fantasy, summer, thriller, and romantic. For genres, examples include acoustic, blues, nursery rhyme, cinematic, classical, country, electronic, fantasy, folk, punk, hip hop, holiday, indie, jazz, pop, and retro. For moods, examples include epic, exciting, happy, and playful.

[0094] The electronic device 100 can select background music from a sound source database based on at least one keyword. The electronic device 100 can receive the necessary background music from the sound source database via the communication unit 120, or obtain it from the storage unit 140 which incorporates the sound source database.

[0095] The electronic device 100 can acquire audio content 480 using narration information 430. Furthermore, the electronic device 100 can acquire subtitle information 490 using narration information 430.

[0096] The electronic device 100 can generate video content 500 through a combination of multiple cut images 421, image effects 460, background music 470, audio content 480, and subtitle information 490 obtained through the above process.

[0097] The electronic device 100 can upload the generated video content 500 to a video service server such as YouTube (registered trademark) or transmit it to a user terminal to make use of the generated video content 500.

[0098] Alternatively, the electronic device 100 may directly play the shape content 500 in a form that can be viewed by the user through the display unit 130 and a speaker (not shown).

Claims

1. In an electronic device that generates video content using a generative artificial intelligence model-based digital content, Using digital content, identify multiple cut images and the text corresponding to the multiple cut images. The plurality of cut images; the text; and a first prompt including analytical infrastructure information are input to a generative artificial intelligence model to obtain analytical information regarding the digital content. The process involves inputting the aforementioned multiple cut images, video asset information including the text and analysis information, and a second prompt including narration generation base information into a generative artificial intelligence model to obtain narration information consisting of multiple sentences. The third prompt, including the aforementioned video asset information and matching infrastructure information, is input to a generative artificial intelligence model to select at least one cut image from the plurality of cut images that matches each of the sentences. An electronic device including a processor that generates video content using the multiple cut images and narration information according to the selection results.

2. The electronic device according to claim 1, characterized in that the analysis information is analysis information relating to at least one of the characters or scenes from the cut images of the digital content.

3. The aforementioned processor, The electronic device according to claim 1, which inputs a fourth prompt including the aforementioned video asset information and synopsis generation base information into a generative artificial intelligence model to obtain synopsis information relating to digital content.

4. The aforementioned processor, The electronic device according to claim 1, which inputs a fifth prompt including the aforementioned video asset information and character analysis base information into a generative artificial intelligence model to acquire character information relating to digital content.

5. The aforementioned processor, The electronic device according to claim 1, which converts the narration information into audio content.

6. The electronic device according to claim 1, wherein the matching base information includes at least one candidate cut image that matches each of the sentences and a score for the candidate cut image.

7. The aforementioned processor, The electronic device according to claim 1, wherein a sixth prompt including the aforementioned video asset information and usage conditions for each image effect is input to a generative artificial intelligence model to select an image effect corresponding to each of the plurality of cut images.

8. The aforementioned processor, Based on the aforementioned video asset information, identify at least one keyword for searching for background music for the video content, The electronic device according to claim 7, which selects background music from a sound source database based on at least one of the aforementioned keywords.

9. In a method for generating video content using a generative artificial intelligence model-based digital content performed by an electronic device, A step of identifying multiple cut images and the text corresponding to the multiple cut images using digital content; A step of inputting the aforementioned multiple cut images; the aforementioned text; and a first prompt including analytical infrastructure information into a generative artificial intelligence model to obtain analytical information regarding the digital content; A step of inputting video asset information including the aforementioned multiple cut images, the text and the analysis information; and a second prompt including narration generation base information into a generative artificial intelligence model to obtain narration information consisting of multiple sentences; A third prompt including the aforementioned video asset information and matching infrastructure information is input to a generative artificial intelligence model to select at least one cut image from the plurality of cut images that matches each of the sentences; A method comprising the step of generating video content using the multiple cut images and narration information according to the selection results.

10. The method according to claim 9, characterized in that the analysis information is analysis information relating to at least one of the characters or scenes from the cut images of the digital content.

11. The method according to claim 9, further comprising the step of inputting a fourth prompt, which includes the aforementioned video asset information and synopsis generation base information, into a generative artificial intelligence model to obtain synopsis information relating to digital content.

12. The method according to claim 9, further comprising the step of inputting a fifth prompt, which includes the aforementioned video asset information and character analysis base information, into a generative artificial intelligence model to obtain character information relating to digital content.

13. The step of generating the aforementioned video content is: The method according to claim 9, further comprising the step of converting the narration information into audio content.

14. The method according to claim 9, wherein the matching base information includes at least one candidate cut image that matches each of the sentences and a score for the candidate cut image.

15. The step of generating the aforementioned video content is: The method according to claim 9, further comprising the step of inputting a sixth prompt, including the aforementioned video asset information and usage conditions for each image effect, into a generative artificial intelligence model to select an image effect corresponding to each of the plurality of cut images.

16. The step of generating the aforementioned video content is: A step of identifying at least one keyword for searching for background music for video content based on the aforementioned video asset information; The method according to claim 15, further comprising the step of selecting background music from a sound source database based on at least one of the keywords.