Method for providing video, electronic device for supporting same, and storage medium

A generative AI model in electronic devices ensures consistent visualization of primary objects across scenes by analyzing and generating content, addressing the challenge of object consistency in video content.

WO2025193001A1PCT designated stage Publication Date: 2025-09-18SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/003358
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-04-12
Filing Date
2025-03-14
Publication Date
2025-09-18

AI Technical Summary

Technical Problem

Existing electronic devices lack efficient methods to consistently visualize primary objects across multiple scenes in videos based on user requests, leading to inconsistencies in image and video content generation.

Method used

Implementing a generative artificial intelligence model that identifies and maintains consistency of primary objects throughout a video by analyzing scenes and generating content based on input descriptions, using modules like script generation, main object analysis, and video generation to ensure coherent visualization of main objects.

Benefits of technology

Enables consistent visualization of primary objects across multiple scenes, enhancing the coherence and quality of video content generation, reducing the need for user intervention in maintaining object consistency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025003358_18092025_PF_FP_ABST
    Figure KR2025003358_18092025_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to one embodiment may comprise: a memory for storing instructions; and a processor. When executed by the processor, the instructions can instruct the electronic device to: identify first information associated with descriptions corresponding to a plurality of scenes; identify, on the basis of the first information, at least one main object included in at least some of the plurality of scenes; input descriptions corresponding to the at least one main object into a first model trained to output content on the basis of the reception of information associated with scenes, thereby acquiring, on the basis of information output from the first model, first content corresponding to the at least one main object; and, on the basis of the first information and the information about the first content, acquire a video including the plurality of scenes. The at least one main object can be consistently visualized in at least some of the plurality of scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Method for providing video, electronic devices supporting the same, and storage media

[0001] The present disclosure relates to a method for providing a video, an electronic device supporting the same, and a storage medium.

[0002] Thanks to the remarkable advancements in information and communication technology and semiconductor technology, the proliferation and use of various electronic devices is rapidly increasing. Electronic devices are being developed to enable users to carry and communicate with one another. An electronic device can refer to any device that performs a specific function based on its embedded software, such as a mobile communication terminal, tablet PC, audio / video device, desktop / laptop computer, or in-car navigation system.

[0003] Electronic devices may require technology to provide images or videos upon user request.

[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.

[0005] An electronic device according to one embodiment may include a memory storing instructions and a processor. The instructions, when executed by the processor, may cause the electronic device to identify first information associated with descriptions corresponding to a plurality of scenes. The instructions, when executed by the processor, may cause the electronic device to identify at least one main object included in at least some of the plurality of scenes based on the first information. The instructions, when executed by the processor, may cause the electronic device to input a description corresponding to the at least one main object to a first model trained to output content based on input of information associated with a scene, thereby obtaining first content corresponding to the at least one main object based on information output from the first model. The instructions, when executed by the processor, may cause the electronic device to obtain an image including the plurality of scenes based on the first information and information about the first content. The at least one primary object may be consistently visualized in at least some of the plurality of scenes.

[0006] A method according to one embodiment may include an operation of identifying first information associated with a description corresponding to a plurality of scenes. The method may include an operation of identifying a description corresponding to at least one main object included in at least some of the plurality of scenes based on the first information. The method may include an operation of obtaining first content corresponding to the at least one main object based on information output from the first model by inputting a description corresponding to the at least one main object into a first model configured to output content based on input of information associated with the scene. The method may include an operation of obtaining an image including the plurality of scenes based on the first information and information about the first content. The at least one main object may be consistently visualized in at least some of the plurality of scenes.

[0007] In one embodiment, a computer-readable medium having computer-executable instructions recorded thereon may cause the electronic device to receive a request associated with acquisition of an image including a plurality of scenes, when executed by a processor of the electronic device. The computer-executable instructions, when executed by the processor, may cause the electronic device to identify at least one primary object included in at least a portion of the plurality of scenes based on first information associated with a description associated with the plurality of scenes identified based on the request. The computer-executable instructions, when executed by the processor, may cause the electronic device to acquire first content corresponding to the at least one primary object based on information output from the first model by inputting a description corresponding to the at least one primary object into a first model configured to output content based on inputting information associated with the scene. The computer-executable instructions, when executed by the processor, may cause the electronic device to obtain an image including the plurality of scenes based on the first information and the information about the first content. The at least one primary object may be consistently visualized in at least some of the plurality of scenes.

[0008] In connection with the description of the drawings, the same or similar reference numerals may be used for the same or similar components.

[0009] FIG. 1 is a block diagram of an electronic device within a network environment, according to one embodiment.

[0010] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0011] FIG. 3 is a block diagram of a memory of an electronic device, according to one embodiment.

[0012] FIG. 4 is a flowchart illustrating a method for providing a video according to one embodiment.

[0013] FIG. 5a, FIG. 5b, and FIG. 5c are drawings for explaining a method of providing a video according to one embodiment.

[0014] FIG. 6 is a flowchart illustrating a method for obtaining a video according to one embodiment.

[0015] FIG. 7 is a signal flow diagram illustrating a method for acquiring a video according to one embodiment.

[0016] FIG. 8 is a flowchart illustrating a method for obtaining information associated with a description corresponding to a plurality of scenes, according to one embodiment.

[0017] FIG. 9 is a diagram illustrating a method for obtaining information associated with a description corresponding to a plurality of scenes according to one embodiment.

[0018] FIG. 10 is a flowchart illustrating a method for obtaining another image based on information corresponding to a stored primary object, according to one embodiment.

[0019] FIG. 11 is a signal flow diagram illustrating a method for acquiring another image based on information corresponding to a stored primary object, according to one embodiment.

[0020] FIG. 12 is a flowchart illustrating a method for providing content associated with a primary object, according to one embodiment.

[0021] FIG. 13 is a diagram illustrating a method for providing content associated with a main object according to one embodiment.

[0022] FIG. 14 is a flowchart illustrating a method of acquiring an image according to one embodiment.

[0023] Fig. 15 is a diagram for explaining a generative artificial intelligence model according to one embodiment.

[0024] FIG. 16 is a flowchart illustrating a method for acquiring an image according to one embodiment.

[0025] Hereinafter, embodiments of the present disclosure will be described in detail with reference to the drawings so that those skilled in the art can easily implement the present disclosure. However, the present disclosure may be implemented in various different forms and is not limited to the embodiments described herein. In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components. Furthermore, in the drawings and related descriptions, descriptions of well-known functions and configurations may be omitted for clarity and conciseness.

[0026] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100), according to one embodiment.

[0027] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0028] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor)) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0029] The auxiliary processor (123) may control at least a portion of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0030] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0031] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0032] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0033] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0034] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0035] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0036] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0037] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0038] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0039] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0040] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0041] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).

[0042] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0043] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0044] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) can support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0045] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0046] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0047] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0048] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0049] FIG. 2 is a block diagram of an electronic device according to one embodiment.

[0050] In one embodiment, the electronic device (201) may be the electronic device (101) of FIG. 1.

[0051] Referring to FIG. 2, in one embodiment, the electronic device (201) may include a communication circuit (210), a processor (220), a memory (230), and / or a display (240).

[0052] In one embodiment, the communication circuit (210) may be included in the communication module (190) of FIG. 1.

[0053] In one embodiment, the communication circuit (210) may transmit or receive signals for communication with an external electronic device (e.g., electronic device (102), electronic device (104), and / or server (108) of FIG. 1).

[0054] In one embodiment, the processor (220) may be included in the processor (120) of FIG. 1.

[0055] In one embodiment, the processor (220) may control the overall operation of providing a video. The processor (220) may include one or more processors for performing the operation of providing a video. For example, the processor (220) may correspond to multiple processors that collectively perform multiple operations by dividing them among the processors. The operation of providing a video performed by the processor (220) will be described in detail with reference to FIGS. 3 to 16 .

[0056] In one embodiment, the processor (220) may include a neural processing unit (NPU) for performing a video acquisition operation. For example, when the video acquisition operation is performed using an artificial intelligence model, the processor (220) may include, but is not limited to, an NPU capable of performing the video acquisition operation using an artificial intelligence model. For example, when the video acquisition operation is performed using a designated algorithm, the processor (220) may include a graphic processing unit (GPU) capable of performing the video acquisition operation using a designated algorithm.

[0057] In one embodiment, the memory (230) may be included in the memory (130) of FIG. 1.

[0058] In one embodiment, the memory (230) may store information for performing an operation to provide a video.

[0059] In one embodiment, the memory (230) may store a generative artificial intelligence (AI) model. The generative AI model may include an artificial intelligence model trained to output content similar to the input content, such as text, audio, and / or images, based on input content. The generative AI model may, for example, learn patterns of the content included in the training data. The type of content output by the generative AI model may include at least one of an image, audio, or text associated with a description of the output content. In one embodiment, the act of obtaining content using the generative model may be referred to as an "act of generating content." In one embodiment, the generative AI model may be stored in an external electronic device.

[0060] In one embodiment, the display (240) may be included in the display module (160) of FIG. 1.

[0061] In one embodiment, the display (240) can play a video stored in the memory (230). For example, the display (240) can display a video playback screen under the control of the processor (220). The screen displayed by the display (240) will be described in detail later with reference to FIGS. 3 to 16.

[0062] Although the electronic device (201) in FIG. 2 is illustrated as including a communication circuit (210), a processor (220), a memory (230), and / or a display (240), it is not limited thereto. In one embodiment, the electronic device (201) may further include at least one of the components included in the electronic device (101) of FIG. 1. For example, the processor (201) may further include at least one of an input module (150) (e.g., a microphone), an audio output module (155) (e.g., a speaker), or a camera module (180).

[0063] FIG. 3 is a block diagram of a memory of an electronic device, according to one embodiment.

[0064] In one embodiment, referring to FIG. 3, the memory (230) (e.g., the memory (230) of FIG. 2) may include a script generation module (310), a main object analysis module (320), a main object generation module (330), a main object editing module (340), a main object database (DB) (350), and a video generation module (360).

[0065] In one embodiment, modules (e.g., at least one of a script generation module (310), a main object analysis module (320), a main object generation module (330), a main object editing module (340), a main object DB (350), or a video generation module (360)) implemented in the electronic device (201) (or stored in the memory (230)) may be implemented in the form of a program, computer code, instructions, routines, processes, software, applications, firmware, or a combination of at least two or more thereof that can be executed by a processor (e.g., the processor (220) of FIG. 2). For example, when at least one of the modules is executed, the processor (220) may perform an operation corresponding to the at least one module. Hereinafter, the description that "a specific module performs an operation" may be understood as "as the specific module is executed, the processor (220) performs an operation corresponding to the specific module." In one embodiment, at least some of the modules may include multiple programs, but are not limited to what is described. Meanwhile, at least some of the modules may be implemented in hardware form. For example, at least some of the script generation module (310), the main object analysis module (320), the main object generation module (330), the main object editing module (340), the main object DB (350), or the video generation module (360) may be implemented as a processing circuit (not shown). In one embodiment, the modules may be implemented as a service or an application when running on an Android operating system.

[0066] In one embodiment, the script generation module (310) may include a generative AI model (e.g., a large language model (LLM)) trained to output scripts corresponding to multiple scenes based on input prompts. For simplicity, each script may be referred to as a description below. In one embodiment, the script generation module (310) may obtain a script corresponding to an input prompt by using an LLM stored in a different area from an area corresponding to the script generation module (310) in the memory (230). The prompt input to the script generation module (310) may include, for example, at least one of text information or image information. The script output by the script generation module (310) may include, for example, text information associated with a description corresponding to a scene.

[0067] In one embodiment, the information output by the script generation module (310) may vary depending on the type of the input prompt. For example, the script generation module (310) may output text information including a description associated with a plurality of scenes or a brief summary associated with a video, based on whether the type of the input prompt corresponds to a long type or a short type. In one embodiment, the information output by the script generation module (310) based on a long type of input prompt may be identical to the information output by the script generation module (310) based on a short type of input prompt.

[0068] In one embodiment, the script generation module (310) may output text information associated with descriptions corresponding to a plurality of scenes based on receiving a relatively long prompt. The relatively long input prompt may include a description associated with a video generated (or acquired) by the electronic device (201). The description associated with the video may include, for example, a relatively specific description associated with one or more scenes included in the video. The relatively long input prompt inputted to the script generation module (310) may be acquired based on a relatively short prompt by the electronic device (201) or may be acquired based on a user input, without limitation. The text information outputted by the script generation module (310) may, for example, correspond to each of the plurality of scenes. The output text information may include a script corresponding to a scene, such as, for example, a script corresponding to a first scene or a script corresponding to a second scene.

[0069] In one embodiment, the script generation module (310) may output a relatively long script based on a relatively short input prompt (e.g., a sentence). The relatively short input prompt may include, for example, a brief summary associated with a video generated (or acquired) by the electronic device (201). The script generation module (310) may also input a rule or condition along with the sentence. The rule or condition may include a command associated with the context of the script to be output by the script generation module (310). The script generation module (310) may output a script associated with at least one keyword based on at least one keyword included in the input sentence or at least one condition input together with the sentence. For example, the script output by the script generation module (310) may include text information associated with a story including a plurality of scenes. Typically, the story may include descriptions associated with the plurality of scenes. The script may include a plurality of paragraphs corresponding to each of the plurality of scenes. For example, a paragraph may include multiple sentences corresponding to a scene. In one embodiment, the operation of the script generation module (310) outputting a script corresponding to a scene may be referred to as an "operation of generating a script."

[0070] In one embodiment, the script generation module (310) can determine whether the prompt corresponds to a short type or a long type based on whether the information included in the prompt satisfies a condition associated with the type of the prompt. For example, the script generation module (310) can determine whether the prompt corresponds to a short type or a long type based on whether the number of sentences included in the prompt exceeds a threshold. The specific value of the threshold corresponding to the number of sentences may vary depending on the embodiment, and there is no limitation thereto. The script generation module (310) can determine whether the prompt corresponds to a short type or a long type based on whether the number of words included in the prompt exceeds a threshold. The specific value of the threshold corresponding to the number of words may vary depending on the embodiment, and there is no limitation thereto. In one embodiment, the threshold may be set to a specific value or a specific range.

[0071] In one embodiment, the main object analysis module (320) may output at least one script corresponding to at least one main object based on receiving a script output by the script generation module (310). In one embodiment, the main object may be an object that requires consistency to be maintained during the transition between scenes (or video frames) among a plurality of objects included in a video provided by the electronic device (201). The main object may include a main character and a main background. For example, if the type (or genre) of the description corresponding to the prompt or script corresponds to a novel, the main character may include a person, animal, or object that contributes relatively highly to the progress of the story. The main background may include a place that is central to the progress of the story. In one embodiment, the at least one main object may include at least one of a person or voice information corresponding to the person.

[0072] In one embodiment, the primary object analysis module (320) may include an artificial intelligence model trained to perform natural language processing (NLP) operations. For example, the primary object analysis module (320) may identify at least one primary object from input information based on a keyword detection algorithm. The primary object analysis module (320) may output a description corresponding to a primary object, such as a script, based on inputting descriptions corresponding to multiple scenes (e.g., including multiple scripts corresponding to the multiple scenes).

[0073] In one embodiment, the primary object analysis module (320) may identify primary objects based on at least one keyword (or word) identified from the input information, based on a relatively short prompt or script input. When a relatively short prompt is input, the primary object analysis module (320) may output (or generate) a relatively long script corresponding to the input prompt. The relatively long script may include scripts corresponding to each of a plurality of scenes. The relatively long script may also include the entire script corresponding to the story of the input prompt. The primary object analysis module (320) may identify at least one primary object based on at least one piece of information identified from the output script. The manner in which the primary object analysis module (320) identifies at least one primary object from a description (e.g., a relatively long script) corresponding to a plurality of scenes may be the same as or similar to the manner described below.

[0074] In one embodiment, the primary object analysis module (320) may identify at least one primary object based on at least one of the following: the frequency with which information corresponding to the object is identified within the input description, the relative position of the information corresponding to the object within the input description, the relative frequency of information corresponding to the object within a script similar to the input description, or the description level of the description associated with the object. For example, the primary object analysis module (320) may identify an object as a primary object if the frequency with which information corresponding to the object is identified within the input script is relatively high. In one embodiment, the primary object analysis module (320) may identify an object as a primary object if information corresponding to the object is repeatedly identified within a script corresponding to a plurality of scenes. The primary object analysis module (320) may determine that the closer the location at which information corresponding to the object is detected within the input script is to the starting point of the script, the higher the probability that the object corresponds to the primary object. The primary object analysis module (320) can determine that the probability that an object corresponds to a primary object is relatively high when the frequency with which information corresponding to an object is detected within similar descriptions that correspond to the same type as the input description is relatively high. The primary object analysis module (320) can determine that the probability that an object corresponds to a primary object is relatively high when the description level of the description associated with the object is high. The description level can be determined to be high as the description associated with the object is detailed. For example, when an object identified from an input prompt corresponds to a place or a person, the primary object can be determined to be an object associated with detailed descriptions as opposed to an object associated with a simple name, but there is no limitation thereto.

[0075] In one embodiment, although not illustrated in FIG. 3, the primary object analysis module (320) may receive user input associated with the designation of a primary object. Based on receiving the user input associated with the designation of a primary object, the primary object analysis module (320) may identify an object identified from the user input as a primary object.

[0076] In one embodiment, although not shown in FIG. 3, the primary object analysis module (320) may receive, from the primary object generation module (330), a script associated with an object acquired by the primary object generation module (330). The primary object analysis module (320) may identify a primary object based on at least one of a script associated with an object received from the primary object generation module (330) or a script corresponding to a scene received from the script generation module (310).

[0077] In one embodiment, the key object generation module (330) may include a generative AI model trained to output content associated with a key object based on inputting a description corresponding to the key object. The content associated with the key object may include, for example, at least one of text information associated with a description corresponding to the key object, image information corresponding to the key object, or audio information associated with the key object. Information that may be included in the content associated with the key object may be simply referred to as information about the content hereinafter. Text information associated with a description corresponding to the key object may include, for example, a "prompt corresponding to the key object." Image information corresponding to the key object may include a "preview image." In one embodiment, the prompt corresponding to the key object may be provided to the key object generation module (330) as input information for obtaining (or generating) image information corresponding to the key object. For example, the key object analysis module (320) may provide a prompt corresponding to the key object as input information to the key object generation module (330). The main object generation module (330) may output at least one of text associated with a description corresponding to the main object, image information corresponding to the main object, or audio information associated with the main object, based on a prompt corresponding to the input main object. In one embodiment, the text information associated with the description corresponding to the main object output by the main object generation module (330) may be different from the prompt input to the main object generation module (330). For example, the text information associated with the description corresponding to the main object output by the main object generation module (330) may include a user-friendly summary associated with the main object.In one embodiment, the electronic device (201) can improve the consistency of the shape corresponding to the visualized primary object based on user input related to the provision of content associated with the primary object and approval of the content associated with the primary object prior to acquisition of the video.

[0078] In one embodiment, the primary object editing module (340) may output a prompt corresponding to the primary object based on receiving content associated with the primary object. The primary object editing module (340) may transmit the changed prompt associated with the primary object to the primary object generation module (330), for example, based on confirming a user input associated with a change in text information associated with a description corresponding to the primary object. If the primary object editing module (340) does not confirm a user input associated with a change in text information associated with a description corresponding to the primary object, the primary object editing module (340) may transmit a prompt corresponding to the primary object to the primary object generation module (330) based on confirming a user input associated with a request for obtaining a preview image of the primary object.

[0079] In one embodiment, the primary object generation module (330) may store content associated with a primary object in the primary object DB (350), for example, based on user input associated with approval of content associated with the primary object. In one embodiment, user approval may not be required. The content stored in the primary object DB (350) may include, for example, an image associated with the primary object or a script corresponding to the primary object. The script corresponding to the primary object may include a prompt that is input to the primary object generation module (330) to generate (or obtain) an image or audio corresponding to the primary object. In one embodiment, the content stored in the primary object DB (350) may be used to generate an image associated with a scene that includes the primary object. The content stored in the primary object DB (350) may also be used to generate other images.

[0080] In one embodiment, the video generation module (360) may obtain a video including a plurality of scenes based on a description (e.g., which may include a plurality of scripts corresponding to a plurality of scenes) and information about a main object stored in the main object DB (350). For example, the plurality of scripts corresponding to a plurality of scenes input to the video generation module (360) may include information output by the script generation module (310). In one embodiment, the plurality of scripts corresponding to a plurality of scenes input to the video generation module (360) may include a prompt based on user input without an operation by the script generation module (310). In one embodiment, the video generation module (360) may also obtain a video including a plurality of scenes directly based on a user input (e.g., a prompt or a description) associated with video generation (or acquisition). For example, the user input associated with video generation may include information associated with a command, a request, or a summary. In one embodiment, a video including a plurality of scenes may be required to maintain consistency or connectivity of main objects included in the video. For example, both the first and last paragraphs of the prompt input to the script generation module (310) may include text information related to “boy’s behavior.” The key object analysis module (320) may automatically determine whether the text information related to the boy identified in different areas of the input prompt corresponds to the same object. If it is determined that the text information related to the boy corresponds to the same object, consistent visualization of the text information indicating “boy” may be required. The video generation module (360) may obtain descriptions corresponding to multiple scenes (e.g., multiple scripts corresponding to multiple scenes output from the script generation module (310).The video generation module (360) can obtain a script corresponding to a "boy" identified as a primary object from the primary object DB (350). This script can correspond to information about the aforementioned content, which can be output based on, for example, receiving input of a description corresponding to the primary object by a generative AI model trained to output content associated with the primary object.

[0081] The video generation module (360) can obtain a video including a plurality of scenes based on a script corresponding to a main object and a plurality of scripts (or information about content) corresponding to a plurality of scenes. The video generation module (360) can then obtain a video using the script associated with the main object (or information about the first content) together with the script corresponding to the scene, thereby consistently visualizing the main object. That is, since the main object is automatically identified in the description to be included in at least some of the plurality of scenes, and the script associated with the main object (or information about the first content) is available, for example, in a corresponding database, the video generation module (360) can consistently visualize the main object throughout the video, i.e., in all scenes in which the main object appears. In particular, user input may generally not be required to identify or approve the main object in other scenes. In one embodiment, the video generation module (360) can consistently visualize not only the main character but also the main background.

[0082] In one embodiment, the video generation module (360) may provide the acquired video through a display (e.g., the display of FIG. 2). The electronic device (201) may play the acquired video through the display. According to one embodiment, the electronic device (201) may execute the main object editing module (340) based on a user input for at least some frames of the video provided through the display. For example, the electronic device (201) may execute the main object editing module (340) based on a user input associated with the selection of a main object included in at least some frames. The electronic device (201) may also provide a message asking whether to change a prompt corresponding to the main object based on the user input associated with the selection of the main object. The electronic device (201) may provide a function for editing a shape corresponding to the main object based on executing the main object editing module (340).

[0083] In one embodiment, a plurality of modules (e.g., at least one of a script generation module (310), a key object analysis module (320), a key object generation module (330), a key object editing module (340), or a video generation module (360)) included in the electronic device (201) may be connected to an AI model. Each of the plurality of modules may perform at least one operation described in the present disclosure by using the AI ​​model. Each of the plurality of modules may perform at least one operation based on information output from the AI ​​model, for example, by inputting data into the AI ​​model.

[0084] In one embodiment, the modules of the present disclosure (e.g., at least one of the script generation module (310), the main object analysis module (320), the main object generation module (330), the main object editing module (340), or the video generation module (360)) may be implemented as on-device modules within the electronic device (201). AI models included in the modules of the present disclosure or AI models used by the modules may be implemented as on-device modules within the electronic device (201). In one embodiment, the AI ​​models may be implemented as a single integrated AI model. Each of the plurality of modules may be connected to the integrated AI model.

[0085] In one embodiment, the script generation module (310), the main object analysis module (320), and the main object generation module (330) may be implemented as on-device modules, and the video generation module (360) may be implemented as at least a part of an external electronic device (e.g., the server (108) of FIG. 1).

[0086] In one embodiment, the AI ​​models used by the modules of the present disclosure may be implemented as at least a part of an external electronic device (e.g., server (108) of FIG. 1). For example, the AI ​​models included in the external electronic device may be implemented as an integrated AI model, but are not limited thereto.

[0087] FIG. 4 is a flowchart (400) for explaining a method of providing a video according to one embodiment.

[0088] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0089] Referring to FIG. 4, according to one embodiment, in operation 401, the electronic device (201) (e.g., the processor (220) of FIG. 2) may identify first information associated with a description corresponding to a plurality of scenes. According to one embodiment, the electronic device (201) may obtain a prompt input through a display (e.g., the display (240) of FIG. 2). For example, the electronic device (201) may obtain text information associated with a description corresponding to the plurality of scenes based on a touch input on the display (e.g., a touch screen). In one embodiment, the first information may include at least one of text information, image information, or audio information.

[0090] In one embodiment, in operation 403, the electronic device (201) may identify at least one primary object. In particular, the electronic device (201) may identify information associated with a description corresponding to the at least one object. The manner in which the identification of the primary object may be performed has already been described above with reference to FIG. 3 and the primary object analysis module, and thus, overlapping descriptions may not be repeated. The electronic device (201) may identify information associated with a description corresponding to at least one primary object included in at least some of the plurality of scenes based on the first information. In one embodiment, the electronic device (201) may identify text information associated with the primary object from text information associated with the description corresponding to the plurality of scenes based on natural language processing.

[0091] In one embodiment, in operation 405, the electronic device (201) may obtain first content corresponding to at least one primary object based on information output from a first model. The electronic device (201) may obtain first content corresponding to the primary object based on information output from the first model by inputting a description corresponding to the at least one primary object (or information associated with the description) into a first model trained to output content based on inputting information associated with a scene. The first model may include a text-to-image model, and is not limited to a specific AI model if it is a generative AI model trained to output image information based on inputting text information. The first content may include, for example, a prompt corresponding to the primary object, a preview image of the primary object, or audio information associated with the primary object.

[0092] In one embodiment, in optional operation 407, the electronic device (201) may identify second information corresponding to at least one primary object. The electronic device (201) may identify second information corresponding to at least one primary object based on a user input for the first content. In one embodiment, the electronic device (201) may control the display to display the first content. For example, based on identifying a user input associated with a change in a prompt corresponding to the primary object, the electronic device (201) may identify the changed prompt and image information (or audio information) corresponding to the changed prompt as second information. Based on identifying a user input associated with approval of the first content, the electronic device (201) may identify the first content acquired by operation 405 as second information.

[0093] In one embodiment, at operation 409, the electronic device (201) may obtain an image including a plurality of scenes. The electronic device (201) may obtain the image including the plurality of scenes based on the first information and the information about the first content. According to one embodiment, the image may further be obtained based on the second information, if possible. The electronic device (201) may consistently visualize a main object included in at least some of the plurality of scenes by obtaining the image based on the first information associated with a description corresponding to the scene and the second information associated with a main object, particularly the information about the first content.

[0094] A generative AI model may include an AI model trained to output an image in a manner that removes noise from random noise. For example, the generative AI model may obtain an image based on an input prompt. The generative AI model may obtain a different image based on the same input prompt. A generative AI model trained based on a method of removing random noise may output a different image even when the text information included in the input prompt is the same. According to one embodiment, an electronic device (201) may consistently visualize a key object included in at least some of a plurality of scenes based on obtaining information about first content (and optionally, second information) associated with the key object prior to generating an image including a plurality of scenes.

[0095] In one embodiment, at least a portion of the operation of obtaining the first content corresponding to the at least one primary object and at least a portion of the operation of causing the image including the plurality of scenes to be obtained may be performed in parallel.

[0096] FIG. 5a, FIG. 5b, and FIG. 5c are drawings for explaining a method of providing a video according to one embodiment.

[0097] Referring to FIG. 5A, according to one embodiment, the electronic device (201) may display a window (510) associated with an input prompt and a window (520) associated with confirmation of a key object through the display (240) (e.g., the display (240) of FIG. 2). The electronic device (201) may confirm a relatively short input prompt, such as, for example, “After A meets B at street T, he delivers the luggage by motorcycle.” The input prompt may correspond to a user input (e.g., a touch input or an audio input) associated with video generation (or acquisition). The electronic device (201) may acquire a plurality of scripts corresponding to a plurality of scenes based on executing a script generation module (e.g., the script generation module (310) of FIG. 3). The plurality of scripts may be acquired, for example, by the script generation module, based on the input prompt. The electronic device (201) may display an input prompt on at least a portion (511) of a window (510) associated with the input prompt. The electronic device (201) may also display a plurality of scripts corresponding to a plurality of scenes acquired by the script generation module on the at least portion (511). The plurality of scripts corresponding to the plurality of scenes may be acquired, for example, as shown in Table 1.

[0098] Scene Number Script 1. A meets B as he leaves T Street with his luggage. 2. A loads his luggage onto his motorcycle and sets off at T Street. B stands nearby. 3. A delivers his luggage on his motorcycle to people at T Street. 4. A returns home, happy after completing his deliveries.

[0099] In one embodiment, based on user input, the electronic device (201) may obtain a relatively long input prompt, such as, for example, Table 2.

[0100] Input prompt: A meets B as he leaves T Street with his luggage. At T Street, A loads his luggage onto his motorcycle and sets off. B stands nearby. A delivers his luggage on his motorcycle to the people at T Street. A returns home and is happy after completing his deliveries.

[0101] In one embodiment, the electronic device (201) may convert a relatively long input prompt into a plurality of scripts corresponding to a plurality of scenes as shown in Table 1 based on executing a script generation module. In one embodiment, the electronic device (201) may display at least one main object identified from the plurality of scripts through the display (240). In one embodiment, the electronic device (201) may identify at least one main object based on at least one of information associated with a frequency in a description corresponding to the plurality of scenes or information associated with a correlation between the plurality of objects, among the plurality of objects included in the plurality of scenes. The electronic device (201) may display identification information of the identified at least one main object on a window (520) associated with the identification of the main object, for example. The identification information of the main object may include, for example, text information referring to the main object in the script. Table 3 includes at least one main object identified from scripts obtained based on the input prompt of Table 2.

[0102] Main ObjectsT Street, A, Luggage, B, Motorcycle

[0103] In one embodiment, the fonts corresponding to each of the displayed primary objects may be different. For example, each of the primary objects may be distinguished by a different color or font. Referring to FIG. 5B , the electronic device (201) may provide content associated with the primary objects through the display (240).

[0104] Referring to reference numerals 570a and 570b, the electronic device (201) may display a window (530) associated with identification information of a primary object through the display (240). Referring to reference numeral 570a, text information (571a) indicating a primary object 'A' may be displayed on the window (530) associated with identification information of the primary object. Referring to reference numeral 570b, text information (571b) indicating a primary object 'T distance' may be displayed on the window (530) associated with identification information of the primary object.

[0105] In one embodiment, the electronic device (201) may display a window (540) associated with a prompt of a primary object through the display (240). The electronic device (201) may display at least one display object (e.g., an icon or text) on at least a portion of the window (540) associated with the prompt of the primary object. The at least one display object may include, for example, a display object (541) associated with a refresh of the primary object, a display object (543) associated with a request to a previous step, or a display object (545) associated with an approval of a prompt of the primary object. The electronic device (201) may perform an action corresponding to the at least one display object based on a user input to the at least one display object. The user input to the at least one display object may be, but is not limited to, a touch input. The electronic device (201) may obtain a preview image of the primary object corresponding to a prompt (573a, 573b) of the primary object based on, for example, a user input for a display object (541) associated with refreshing the primary object. The electronic device (201) may confirm a modified prompt corresponding to the primary object based on a user input associated with modifying the prompt corresponding to the primary object. The electronic device (201) may obtain a preview image of the primary object obtained from the modified prompt based on the user input for the display object (541) associated with refreshing the primary object. The electronic device (201) may display the obtained preview image on a window (550) associated with the preview image of the primary object. The electronic device (201) may store a plurality of preview images and / or prompts corresponding to the primary object obtained based on a plurality of user inputs for the display object (541) associated with refreshing.The electronic device (201) can obtain multiple candidate images corresponding to a single primary object by storing multiple preview images. For example, the multiple candidate images corresponding to a single primary object can be stored as a type of library, but there is no limitation thereto. In one embodiment, the operation of the electronic device (201) obtaining a preview image of the primary object based on a user input associated with modifying a prompt corresponding to the primary object and a user input for a display object (541) associated with refreshing the primary object may be referred to as a "refresh operation of the preview image corresponding to the primary object." In one embodiment, the electronic device (201) can perform the refresh operation of the preview image corresponding to the primary object based on an on-device model. The electronic device (201) can obtain an image corresponding to the primary object based on communication with an AI model external to the electronic device (201) based on a user input for a display object (545) associated with approval of the primary object. For example, an image acquired by an external AI model may have a higher resolution than an image acquired by an on-device model, but there is no limitation thereto. The electronic device (201) may display a window (510) associated with an input prompt through the display (240) based on a user input to a display object (543) associated with a request to the previous step. The electronic device (201) may store a prompt for a key object in a key object DB (e.g., a key object DB (350) of FIG. 3) based on a user input to a display object (545) associated with an approval for a prompt for a key object.

[0106] In one embodiment, the electronic device (201) may display a window (550) associated with a preview image of a primary object through a display. Referring to reference numeral 570a, a preview image (575a) of a primary object 'A' obtained in response to a prompt (573a) of the primary object 'A' may be displayed on at least a portion of the window (550) associated with the preview image of the primary object. Referring to reference numeral 570b, a preview image (575b) of a primary object 'T distance' obtained in response to a prompt (573b) of the primary object 'T distance' may be displayed on at least a portion of the window (550) associated with the preview image of the primary object. In one embodiment, the primary object generation module may be configured to provide audio information corresponding to the primary object together with the preview image corresponding to the primary object. The electronic device (201) may display a preview image of the main object on at least a portion of a window (550) associated with the preview image of the main object based on the output information of the main object generation module, and output audio corresponding to the main object through an audio output module (e.g., the audio output module (155) of FIG. 1). For example, the audio corresponding to the main object may include a voice corresponding to a main character or a soundtrack corresponding to a main background. In one embodiment, the electronic device (201) may confirm a user input associated with a request to modify the audio corresponding to the main object. For example, the electronic device (201) may confirm a user input associated with a request to modify the audio corresponding to the main object based on receiving text information (e.g., a prompt) associated with a description of the audio corresponding to the main object. The electronic device (201) can obtain changed audio information corresponding to the main object based on a prompt associated with a description of audio corresponding to the input main object and / or a script associated with a description of audio corresponding to the main object output by a script generation module.According to one embodiment, the electronic device (201) may change audio information corresponding to the main object based on a user input associated with a request for modifying an image of the main object. The electronic device (201) may output audio corresponding to the main object through an audio output module based on the changed audio information. The electronic device (201) may store multiple audio files corresponding to one main object based on the change in audio corresponding to the main object.

[0107] In one embodiment, the electronic device (201) may display, through the display, at least one display object associated with a request for switching between primary objects or for storing information about primary objects. The electronic device (201) may display a screen associated with another primary object based on a user input for a display object (561) associated with a request for switching to a previous screen. For example, the electronic device (201) may change the screen displayed from the screen indicated by reference numeral 570b to the screen indicated by reference numeral 570a. The electronic device (201) may display a screen associated with another primary object based on a user input for a display object (563) associated with a request for switching to a next screen. For example, the electronic device (201) may change the screen displayed from the screen indicated by reference numeral 570a to the screen indicated by reference numeral 570b. The electronic device (201) may store information associated with a primary object in a memory (e.g., memory (230) of FIG. 2) based on a user input for a display object (565) associated with a storage request for information about the primary object. The information associated with the primary object may include, for example, a prompt for the primary object and a preview image of the primary object.

[0108] Referring to FIG. 5c, the electronic device (201) can display, through the display (240), a window (581) associated with a script corresponding to a scene and a window (583) associated with the display of the acquired image.

[0109] Referring to reference numeral 590a, in one embodiment, the electronic device (201) can obtain an image (593a) corresponding to the first scene based on a script (591a) corresponding to the first scene. The image (593a) corresponding to the first scene can include, for example, A (575a), T distance (575b), B (575c), and load (575d) corresponding to the main object. In one embodiment, the electronic device (201) can obtain a video based on sequentially connecting the first to fourth images.

[0110] Referring to reference numeral 590b, in one embodiment, the electronic device (201) can obtain an image (593b) corresponding to the second scene based on a script (591b) corresponding to the second scene.

[0111] Referring to reference numeral 590b, in one embodiment, the electronic device (201) may acquire an image (593b) corresponding to the second scene based on a script (591b) corresponding to the second scene. The image (593b) corresponding to the second scene may include, for example, objects A (575a), T street (575b), B (575c), luggage (575d), and a motorcycle (575e) corresponding to the main objects.

[0112] Referring to reference numeral 590c, in one embodiment, the electronic device (201) may acquire an image (593c) corresponding to the third scene based on a script (591c) corresponding to the third scene. The image (593c) corresponding to the third scene may include, for example, A (575a), T street (575b), luggage (575d), and a motorcycle (575e) corresponding to the main object.

[0113] Referring to reference numeral 590d, in one embodiment, the electronic device (201) may acquire an image (593d) corresponding to the fourth scene based on a script (591d) corresponding to the fourth scene. The image (593d) corresponding to the fourth scene may include, for example, A (575a) corresponding to a main object.

[0114] FIG. 6 is a flowchart (600) for explaining a method of obtaining a video according to one embodiment.

[0115] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0116] Referring to FIG. 6, according to one embodiment, in operation 601, the electronic device (201) (e.g., the processor (220) of FIG. 2) may obtain a first image including an image of at least one primary object based on first information and information about the first content (and optionally, second information). In one embodiment, the first image may include a video corresponding to a scene in which the primary object is first described within the input prompt. The first image may include a video corresponding to a scene in which a description level corresponding to the primary object is higher than a set level within the input prompt. Prior to obtaining a plurality of images corresponding to a plurality of scenes including the primary object, the electronic device (201) may obtain the first image including the image of at least one primary object based on the first information and information about the first content (and optionally, second information).

[0117] According to one embodiment, in operation 603, the electronic device (201) may provide a first image as an input for acquisition of at least one image. The electronic device (201) may provide the first image as an input for acquisition of at least one image corresponding to at least one scene including at least one main object. The electronic device (201) may obtain another scene including at least one object based on providing the first image including an image of at least one main object as an input together with first information and information about the first content (and optionally, second information). The electronic device (201) may consistently visualize the at least one object by obtaining another scene including the at least one object based on the first information, information about the first content (and optionally, second information), and the first image.

[0118] FIG. 7 is a signal flow diagram illustrating a method for acquiring a video according to one embodiment.

[0119] Referring to FIG. 7, in one embodiment, the electronic device (201) can obtain an image corresponding to a prompt or script based on the third model (710).

[0120] In one embodiment, the electronic device (201) may obtain a first image (731a) output from the third model (710) by inputting a script (721a) corresponding to a first scene and a prompt (723a) corresponding to a main object into the third model (710) based on executing a video generation module (e.g., the video generation module (360) of FIG. 3). The electronic device (201) may provide the first image (731a) as an input for obtaining a second image (731b), a third image (731c), or a fourth image (731d) based on the fact that the first image (731a) includes image information corresponding to the main object. In one embodiment, the third model (710) may include a generative AI model trained to output images corresponding to a plurality of scenes based on inputting a script and / or a prompt. The third model (710) may be implemented as part of the video generation module or may be implemented externally to the video generation module. If the third model (710) is implemented externally to the video generation module, the third model (710) may provide the video generation module with images corresponding to multiple scenes based on receiving commands from the video generation module.

[0121] In one embodiment, the electronic device (201) may obtain a second image (731b) output from the third model (710) by inputting a script (721b) corresponding to a second scene, a prompt (723b) corresponding to a main object, and a first image (731a) into the third model (710). The first image (731a) may be considered to represent information about the first content.

[0122] In one embodiment, the electronic device (201) can obtain a third image (731c) output from the third model (710) by inputting a script (721c) corresponding to a third scene, a prompt (723c) corresponding to a main object, and a first image (731a) into the third model (710).

[0123] In one embodiment, the electronic device (201) can obtain a fourth image (731d) output from the third model (710) by inputting a script (721d) corresponding to the fourth scene, a prompt (723d) corresponding to the main object, and a first image (731a) into the third model (710).

[0124] In one embodiment, a second image corresponding to a second scene, a third image corresponding to a third scene, and a fourth image corresponding to a fourth scene may be generated (or acquired) sequentially. In one embodiment, at least some of the second image, the third image, or the fourth image may be generated in parallel. In one embodiment, based on the fact that at least one of the second scene, the third scene, or the fourth scene does not contain a key object, a prompt corresponding to a key object may not be provided as input to the third model.

[0125] In one embodiment, the electronic device (201) can obtain an image (741) including a plurality of scenes based on sequentially connecting a first image (731a), a second image (731b), a third image (731c), and a fourth image (731d). In one embodiment, the electronic device (201) can also generate an entire video without sequentially connecting the plurality of images based on at least one of the first information (e.g., information associated with a description corresponding to the plurality of scenes), a prompt corresponding to a main object included in a main object DB (e.g., the main object DB (350) of FIG. 3), or an image corresponding to a main object included in the main object DB. The electronic device (201) can consistently visualize the main object by providing the first image (731a) as an input for obtaining the second image (731b), the third image (731c), or the fourth image (731d).

[0126] FIG. 8 is a flowchart (800) illustrating a method for obtaining information associated with a description corresponding to a plurality of scenes according to one embodiment.

[0127] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0128] Referring to FIG. 8, according to one embodiment, in operation 801, the electronic device (201) (e.g., the processor (220) of FIG. 2) may identify a user input associated with a first story. In one embodiment, the user input associated with the first story may include an input prompt of relatively short length. The user input associated with the first story may also include an input prompt of relatively long length that includes descriptions corresponding to a plurality of scenes. The type of the user input may include at least one of text, an image (e.g., a still image, a moving image), or audio (e.g., a voice).

[0129] According to one embodiment, in operation 803, the electronic device (201) may obtain information associated with a description corresponding to at least some of the plurality of scenes based on information output from the second model. The electronic device (201) may obtain text information associated with a description corresponding to at least some of the plurality of scenes based on the information output from the second model by inputting information corresponding to a user input into a second model trained to output text information associated with a description corresponding to the plurality of scenes based on inputting information associated with the story. The electronic device (201) may obtain a script of a relatively long length based on, for example, an input prompt of a relatively short length. The electronic device (201) may obtain a plurality of scripts corresponding to each of the plurality of scenes based on a relatively long input prompt including descriptions corresponding to the plurality of scenes. The electronic device (201) may obtain a script for identifying a key object based on an artificial intelligence model trained to identify first information associated with a description corresponding to the plurality of scenes based on inputting a user input.

[0130] FIG. 9 is a diagram illustrating a method for obtaining information associated with a description corresponding to a plurality of scenes according to one embodiment.

[0131] Referring to FIG. 9, according to one embodiment, the script generation module (310) may output a script (920) based on receiving a prompt (910).

[0132] In one embodiment, referring to reference numeral 911, the script generation module (310) may receive a relatively short prompt (913). Based on the prompt (913), the script generation module may output (915) a script (923) including a plurality of scenes.

[0133] In one embodiment, referring to reference numeral 921, the script (923) output by the script generation module (310) may include a plurality of scripts corresponding to each of a plurality of scenes. For example, the plurality of scripts may include a script (923a) corresponding to a first scene, a script (923b) corresponding to a second scene, a script (923c) corresponding to a third scene, and a script (923d) corresponding to a fourth scene.

[0134] FIG. 10 is a flowchart (1000) illustrating a method for obtaining another image based on information corresponding to a stored primary object, according to one embodiment.

[0135] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0136] Referring to FIG. 10, according to one embodiment, in operation 1001, the electronic device (201) (e.g., the processor (220) of FIG. 2) may store second information corresponding to at least one main object. The electronic device (201) may, for example, store the second information associated with the main object included in the video in a main object DB (e.g., the main object DB (350) of FIG. 3). In one embodiment, the electronic device (201) may store a video corresponding to a first story acquired by a video generation module (e.g., the video generation module (360) of FIG. 3) in a memory (e.g., the non-volatile memory (134) of FIG. 1).

[0137] In one embodiment, in operation 1003, the electronic device (201) may determine whether the second story is correlated with the first story. The electronic device (201) may determine whether there is a correlation between the first story and the second story based on confirming a request for generating an image associated with the second story. For example, the electronic device (201) may determine whether text information corresponding to the second story is correlated with text information corresponding to the first story, but is not limited thereto. In one embodiment, the electronic device (201) may determine whether the second story is correlated with the first story if the similarity between the scripts corresponding to the first story exceeds a set threshold based on comparing the similarity between the scripts corresponding to the first story and the scripts corresponding to the second story. The electronic device (201) may determine whether the second story is not correlated with the first story based on confirming that the similarity between the scripts is less than or equal to the set threshold. In one embodiment, the electronic device (201) may determine whether the second story is correlated with the first story based on comparing the similarity between information associated with at least one main object included in the first story and information associated with at least one main object included in the second story. For example, the electronic device (201) may determine that the second story is correlated with the first story based on determining that the similarity between the information associated with the main object exceeds a set threshold. The electronic device (201) may determine that the second story is not correlated with the first story based on determining that the similarity between the information associated with the main object is below the set threshold. In one embodiment, based on determining that the second story is not correlated with the first story (operation 1003-No), in operation 1007, the electronic device (201) may obtain an image associated with the second story based on text information corresponding to the second story.The electronic device (201) may not provide information associated with the first story as input for obtaining an image associated with the second story.

[0138] In one embodiment, prompts corresponding to the main characters and / or main backgrounds of the first story may be stored in the main object DB. Images corresponding to the main characters and / or main backgrounds of the first story may be stored in the main object DB. Based on confirmation that at least some of the plurality of main objects identified based on text information corresponding to the second story correspond to the information stored in the main object DB, the electronic device (201) may generate (or obtain) a video associated with the second story based on the information stored in the main object DB. For example, if the main characters or main backgrounds of the first story correspond to the main characters or main backgrounds of the second story, the second story may be confirmed to be related to the first story.

[0139] In one embodiment, if the main character of the second story corresponds to the main character of the first story, and the narrative time point of the second story is different from the narrative time point of the first story, the electronic device (201) may generate a video associated with the second story based on information stored in the main object DB. For example, if the narrative time point of the first story is associated with the main character's youth, and the narrative time point of the second story is associated with the main character's old age, the second story may be determined to be associated with the first story.

[0140] In one embodiment, if the narrative of the second story corresponds to the narrative of the first story and the protagonist of the second story is different from the protagonist of the first story, the electronic device (201) may generate a video associated with the second story based on information stored in the main object DB. The narrative may be associated, for example, with the main background and / or introduction of the story. Based on the fact that the narrative of the second story is similar to the first story and that the protagonist of the first story corresponds to a good role and the protagonist of the second story corresponds to an evil role, the electronic device (201) may generate a video associated with the second story using the information stored in the main object DB, but is not limited thereto. Based on the fact that the narrative of the second story is similar to the first story and that the ending of the second story is different from the ending of the first story, the electronic device (201) may also generate a video associated with the second story using information stored in the main object DB.

[0141] In one embodiment, based on the fact that the main object and / or narrative of the second story corresponds to the main object and / or narrative of the first story, and the genre of the second story is different from the genre of the first story, the electronic device (201) may generate a video associated with the second story based on information stored in the main object DB. For example, based on the fact that the main object and / or narrative of the second story corresponds to the main object and / or narrative of the first story, and the genre of the second story corresponds to an advertisement and the genre of the first story corresponds to a movie, the electronic device (201) may generate an advertisement video associated with the second story based on information stored in the main object DB.

[0142] In one embodiment, based on the fact that the main objects and / or narratives of the second story correspond to the main objects and / or narratives of the first story, and that the number of scenes included in the second story is less than the number of scenes included in the first story, the electronic device (201) may generate a video associated with the second story based on information stored in the main object DB. For example, based on the fact that the second story is a summary of the first story, the electronic device (201) may generate a video associated with the second story based on obtaining information associated with the main objects included in the first story from the main object DB.

[0143] In one embodiment, in operation 1005, the electronic device (201) may obtain an image associated with the second story using text information corresponding to the second story and second information. Based on confirming that the second story is related to the first story (operation 1003-Yes), the electronic device (201) may obtain an image associated with the second story using the second information. By obtaining another image based on information stored in the main object DB, the electronic device (201) may reduce a delay that may occur due to the acquisition of the image compared to a case where another image is obtained without the main object DB.

[0144] FIG. 11 is a signal flow diagram illustrating a method for acquiring another image based on information corresponding to a stored primary object, according to one embodiment.

[0145] Referring to FIG. 11, in one embodiment, the electronic device (201) can obtain a video based on different methods depending on whether text information corresponding to the second story is correlated with text information corresponding to the first story.

[0146] Referring to reference numeral 1110, the video generation module (360) can obtain a video based on the text information corresponding to the second story and the second information of the main object associated with the first story obtained from the main object DB (350) based on confirming that the text information corresponding to the second story is correlated with the text information corresponding to the first story.

[0147] Referring to reference numeral 1120, the video generation module (360) can obtain a video based on the text information corresponding to the second story and the information of the main object associated with the second story obtained from the main object DB (350) based on confirming that the text information corresponding to the second story is not related to the text information corresponding to the first story.

[0148] In one embodiment, the video generation module (360) can reduce the delay in obtaining a video by using information stored in the main object DB (350) to obtain another video when a correlation is confirmed.

[0149] FIG. 12 is a flowchart (1200) illustrating a method for providing content associated with a primary object according to one embodiment.

[0150] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0151] Referring to FIG. 12, according to one embodiment, in operation 1201, the electronic device (201) (e.g., the processor (220) of FIG. 2) may provide first content. The electronic device (201) may provide first content including text information associated with a description corresponding to at least one main object. In one embodiment, the electronic device (201) may provide first content including a preview image corresponding to at least one main object. The resolution of the main object included in the preview image may be lower than, for example, the resolution of the main object included in the acquired image. In one embodiment, the electronic device (201) may provide audio information as at least a portion of the first content corresponding to the at least one main object through an audio output module (e.g., the audio output module (155) of FIG. 1). A sampling rate corresponding to the audio information included in the first content may be relatively low. For example, a sampling rate corresponding to audio information included in the first content may be lower than a sampling rate corresponding to audio information included in an image obtained based on the first content. A duration corresponding to audio information included in the first content may be relatively short. For example, a playback time corresponding to audio information included in the first content may be shorter than a playback time corresponding to audio information included in an image obtained based on the first content, but there is no limitation thereto. An image obtained based on the first content may be generated based on a user input associated with a video generation request, for example. In order to reduce a delay caused by executing an on-device module, the electronic device (201) may set the resolution of a preview image output by the main object generation module to have a lower value than the resolution of an image output by the video generation module.The electronic device (201) can control a display (e.g., the display (240) of FIG. 2) to display first content acquired based on the execution of, for example, a main object generation module (e.g., the main object generation module (330) of FIG. 3).

[0152] In one embodiment, at operation 1203, the electronic device (201) may determine whether a user input associated with approval of the first content is confirmed. In one embodiment, the electronic device (201) may control the display (240) to display a display object associated with approval of the first content. The electronic device (201) may determine a user input associated with approval of the first content based on determining a user input for the display object associated with approval of the first content. The user input for the display object may include, but is not limited to, a touch input. In one embodiment, at operation 1205, the electronic device (201) may store second information based on determining a user input associated with approval of the first content (operation 1203 - Yes). The electronic device (201) may store the second information associated with at least one primary object in a primary object DB (e.g., primary object DB (350)).

[0153] In one embodiment, at operation 1207, the electronic device (201) may determine whether text information included in the first content has changed based on a user input associated with approval of the first content not being confirmed (operation 1203—No). Based on the determination of the change in the text information included in the first content, the electronic device (201) may determine second information corresponding to at least one primary object based on a user input to the first content. In one embodiment, the electronic device (201) may control the display to display a window associated with the input prompt. Based on the determination of the user input to the window associated with the input prompt, the electronic device (201) may determine that the text information included in the first content has changed. The user input to the window associated with the input prompt may include, but is not limited to, a touch input on at least a portion of the window and a text input associated with a change in a description corresponding to the primary object, for example. In one embodiment, in operation 1209, the electronic device (201) may obtain the first content based on the changed text information, based on confirmation that the text information included in the first content has changed (operation 1207 - Yes). The electronic device (201) may control the display to display the first content obtained based on the changed text information.

[0154] In one embodiment, the electronic device (201) may provide the first content based on confirming that the text information included in the first content has not changed (operation 1207—No). The electronic device (201) may obtain an image corresponding to the main object based on, for example, a script corresponding to the main object included in the first content provided by operation 1201. The electronic device (201) may obtain images of different main objects even when the prompt corresponding to the main object is the same based on executing the main object generation module (330). The electronic device (201) may repeat the operation of obtaining a preview image corresponding to the main object until a user input associated with approval of the first content is confirmed (operation 1203—Yes).

[0155] FIG. 13 is a diagram illustrating a method for providing content associated with a main object according to one embodiment.

[0156] In one embodiment, the electronic device (201) may display a window (1320) associated with a change in the shape of a primary object through the display (240). The electronic device may display the window (1320) associated with a change in the shape of a primary object based on a user input to a display object displayed on a window (e.g., window (540)) associated with a prompt of the primary object of FIG. 5B . The user input may include an input associated with selecting an option for changing the shape of the primary object or an input associated with editing a prompt associated with a description of the primary object. The electronic device (201) may display at least one display object on at least a portion of the window (1320) associated with a change in the shape of the primary object. The at least one display object may include, for example, a display object (1341) associated with a refresh of the primary object, a display object (1343) associated with a request to return to a previous step, or a display object (1345) associated with an acknowledgment of a prompt of the primary object. The electronic device (201) may perform an operation corresponding to at least one display object based on a user input for at least one display object. The user input for at least one display object may be a touch input, but is not limited thereto. For example, the electronic device (201) may obtain a preview image (575a) of the main object based on a user input for a display object (1341) associated with refreshing the main object. The electronic device (201) may display a preview image of the main object with a changed hairstyle or a changed outfit on a window (550) associated with the preview image of the main object based on a user input for display objects (1311a) associated with a change in the hairstyle of the main object or display objects (1311b) associated with a change in the outfit of the main object and a user input for display object (1341) associated with refreshing the main object.The electronic device (201) may display a window associated with an input prompt through the display (240) based on a user input to a display object (1343) associated with a request to a previous step. The electronic device (201) may store the prompt of the primary object in a primary object DB (e.g., the primary object DB (350) of FIG. 3) based on a user input to a display object (1345) associated with an approval of the prompt of the primary object.

[0157] Referring to reference numeral 1310a, the electronic device (201) may display, through the display (240), display objects (1311a) associated with a change in the hairstyle of the main object on at least a portion of a window (1320) associated with a change in the shape of the main object. The electronic device (201) may display a preview image of the main object with a changed hairstyle on a window (550) associated with the preview image of the main object, based on, for example, a user input associated with a selection of an option for the display objects (1311a) associated with a change in the hairstyle of the main object and a user input for a display object (1341) associated with a refresh of the main object.

[0158] Referring to reference numeral 1310b, the electronic device (201) may display display objects (1311b) associated with a change in the appearance of the main object on at least a portion of a window (1320) associated with a change in the appearance of the main object through the display (240). The electronic device (201) may display a preview image of the main object whose appearance has been changed on a window (550) associated with the preview image of the main object, based on, for example, a user input associated with a selection of an option for the display objects (1311b) associated with a change in the appearance of the main object and a user input for a display object (1341) associated with a refresh of the main object.

[0159] The electronic device (201) can store image information corresponding to the main object in the main object DB based on confirmation that the hairstyle and / or clothing of the main object has been determined.

[0160] FIG. 14 is a flowchart illustrating a method of acquiring an image according to one embodiment.

[0161] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0162] Referring to FIG. 14, according to one embodiment, in operation 1401, the electronic device (201) (e.g., the processor (220) of FIG. 2) may start generating an image including a plurality of scenes. In one embodiment, in operation 1403, the electronic device (201) may determine whether a key object is identified from a script corresponding to the scene. The electronic device (201) may determine whether at least one key object is identified based on executing a key object analysis module (e.g., the key object analysis module (320) of FIG. 3) within the script corresponding to the scene.

[0163] In one embodiment, based on determining that a primary object is not identified from a script corresponding to a scene (operation 1403-No), the electronic device (201) may, in operation 1407, generate an image based on the script corresponding to the scene. The electronic device (201) may generate an image based on the script corresponding to the scene without sending a query for information associated with the primary object to a primary object DB (e.g., primary object DB (350) of FIG. 3).

[0164] In one embodiment, based on confirming that a primary object is identified from a script corresponding to a scene (operation 1403 - Yes), the electronic device (201) can generate an image based on a prompt corresponding to the stored primary object and a script corresponding to the scene in operation 1405. The electronic device (201) can obtain a prompt corresponding to the primary object from a primary object DB. The electronic device (201) can consistently visualize the primary object by generating an image based on the obtained prompt corresponding to the primary object and the script corresponding to the scene. In one embodiment, the electronic device (201) can obtain an image including the plurality of scenes based on information output from the third model by inputting the first information and the second information into a third model trained to generate an image based on inputting information related to a description.

[0165] Fig. 15 is a diagram for explaining a generative artificial intelligence model according to one embodiment.

[0166] Referring to FIG. 15, a user query / response interface (1510), an application / service component (1530), a knowledge repository (1520), an AI framework (1540), and a generative AI model (1560) (e.g., the generative AI model (340) of FIG. 3) may be stored in a memory (230) (e.g., the memory (230) of FIG. 2) or stored in a separate server. At least some of the user query / response interface (1510), the application / service component (1530), the knowledge repository (1520), the AI ​​framework (1540), or the generative AI model (1560) may be implemented in software or hardware.

[0167] According to one embodiment, a user query / response interface (1510) may receive a user's input. The user's input may be in the form of natural language, images, and / or videos, but is not limited thereto. Furthermore, context information may also be transmitted when the user's input is transmitted. The context information may include various additional information at the time of the user's input. For example, the additional information may include information about the application currently being used by the user or information about the user's location. Furthermore, the user's input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Furthermore, the user's input may also be in a non-natural language form, such as selecting a menu. The user query / response interface (1510) may output the results of the generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user query / response interface (1510) may output the results of the generative artificial intelligence system to the user. The output can be in natural language form, in the form of specific content, or in the form of an action requested by the user.

[0168] The AI ​​framework (1540) can receive user input and coordinate and control each component necessary to perform the user's intention based on the user's query.

[0169] User input received from the user query / response interface (1510) can be transmitted to a prompt design component (1541). The prompt design component (1541) can be used to generate prompts suitable for inputting the user input into a large language model (LLM), a large vision model (LVM), or a large multimodal model (LMM). The prompt design component (1541) can be an AI component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design component (1541) can access a knowledge component including user preference data, a prompt library, and prompt examples based on the user input to generate prompts, and can transmit the generated prompts to the large language model (LLM) or the large multimodal model (LMM).

[0170] The API / Plug-in management component (1542) can communicate with external information when there is a request for additional information when passing user input as input to a generative model. The API / Plug-in management component (1542) can establish a channel for communicating with the outside of the AI ​​Interface through the API, and can enable access to various data sources (e.g., knowledge repositories (1520)) through the established channel. In addition, the API / Plug-in management component (1542) can request the application / service component (1530) through the API for an action that ultimately performs the user input, rather than an intermediate result, when the action needs to be performed in the application or service. Information obtained from the outside can be used to generate a prompt in the prompt design component (1541) together with the user input, or can be passed as input to the generative model.

[0171] The output modification component (also called a refiner component) (1543) can fine-tune the output from the generative model. For example, the output modification component (1543) can verify that the content generated through the LLM and / or LMM is not irrelevant, does not contain biased content, or does not contain harmful content. In addition, the output modification component (1543) can determine to what extent the content matches the user's desired result and, if necessary, can perform additional processing. The output modification component (1543) can additionally configure and provide the user with hints to avoid unwanted output.

[0172] A generative AI model (1560) can generally refer to an artificial intelligence neural network that creates new types of data based on user input information. A generative AI model (1560) can include an image-generating model and / or a language-generating model. Representative models for generating images include a generative adversarial network (GAN) and a variational auto encoder (VAE), and examples include a VAE and a Diffusion-based generative model that uses a Transformer structure. A language-generating model is a model trained to statistically output the most appropriate output based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there are also large multimodal models (LMMs) that can recognize various types of data input, such as text, images, and audio, and generate new data corresponding to them.

[0173] FIG. 16 is a flowchart illustrating a method for acquiring an image according to one embodiment.

[0174] In the following examples, the operations may be performed sequentially, but are not necessarily sequential. For example, the order of the operations may be changed, and at least two operations may be performed in parallel.

[0175] Referring to FIG. 16, according to one embodiment, in operation 1601, an electronic device (201) (e.g., processor (220) of FIG. 2) may transmit a request related to image generation to an external electronic device (1610) including a third model trained to generate an image based on inputting information associated with a description. The electronic device (201) may transmit a request related to image generation including first information and second information to the external electronic device (1610). The external electronic device (1610) may generate (or obtain) an image including a plurality of scenes based on the information included in the request and the third model. In one embodiment, in operation 1603, the electronic device (201) may receive, from the external electronic device (1610), an image including a plurality of scenes generated by the external electronic device (1610) in response to the request. The electronic device (201) can acquire an image including a plurality of scenes based on the first information and the second information. The electronic device (201) can provide (or reproduce) the acquired image through a display (e.g., the display (240) of FIG. 2). The electronic device (201) can acquire an image including a plurality of scenes through an external electronic device (1610) that provides a cloud service related to the generation of a video.

[0176] An electronic device (e.g., the electronic device 201 of FIG. 2) according to one embodiment may include a memory (e.g., the memory 230 of FIG. 2) that stores instructions, and a processor (e.g., the processor 220 of FIG. 2). The instructions, when executed by the processor 220, may cause the electronic device 201 to identify first information associated with a description corresponding to a plurality of scenes. The instructions, when executed by the processor 220, may cause the electronic device 201 to identify, based on the first information, information associated with a description corresponding to at least one main object included in at least some of the plurality of scenes. The instructions, when executed by the processor (220), may cause the electronic device (201) to obtain first content corresponding to the at least one main object based on information output from the first model by inputting information associated with a description corresponding to the at least one main object into a first model trained to output content based on inputting information associated with a scene. The instructions, when executed by the processor (220), may cause the electronic device (201) to identify second information corresponding to the at least one main object based on a user input for the first content. The instructions, when executed by the processor (220), may cause the electronic device (201) to obtain an image including the plurality of scenes based on the first information and the second information.

[0177] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to obtain a first image including an image of the at least one main object based on the first information and the second information. The instructions, when executed by the processor (220), may cause the electronic device (201) to provide the first image as an input for obtaining at least one image corresponding to at least one scene in which the at least one main object is included.

[0178] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to, at least as part of an operation of identifying information associated with a description corresponding to the at least one primary object included in at least some of the plurality of scenes based on the first information, identify the at least one primary object among the plurality of objects included in the plurality of scenes based on at least one of information associated with a frequency within the description corresponding to the plurality of scenes or information associated with a correlation between the plurality of objects.

[0179] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to, at least as part of an operation of identifying first information associated with a description corresponding to the plurality of scenes, identify a user input associated with a first story. The instructions, when executed by the processor (220), may cause the electronic device (201) to, at least as part of an operation of identifying first information associated with a description corresponding to the plurality of scenes, input information corresponding to the user input to a second model trained to output text information associated with a description corresponding to the plurality of scenes based on input of information associated with the story, thereby obtaining text information associated with a description corresponding to at least some of the plurality of scenes based on information output from the second model.

[0180] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to store second information corresponding to the at least one primary object. The instructions, when executed by the processor (220), may cause the electronic device (201) to determine whether there is a correlation between the first story and the second story based on confirming a request for generating an image associated with the second story. The instructions, when executed by the processor (220), may cause the electronic device (201) to obtain an image associated with the second story using the second information based on confirming that the second story is correlated with the first story.

[0181] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to provide the first content including text information associated with a description corresponding to the at least one primary object. The instructions, when executed by the processor (220), may cause the electronic device (201) to identify a change in the text information included in the first content as at least part of an operation of identifying the second information corresponding to the at least one primary object based on a user input for the first content.

[0182] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to provide the first content including a preview image corresponding to the at least one primary object.

[0183] In one embodiment, the resolution of the main object included in the preview image may be lower than the resolution of the main object included in the acquired image.

[0184] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to, at least as part of an operation of obtaining an image including the plurality of scenes based on the first information and the second information, input the first information and the second information into a third model trained to generate an image based on inputting information associated with a description, thereby obtaining an image including the plurality of scenes based on information output from the third model.

[0185] In one embodiment, the instructions, when executed by the processor (220), may cause the electronic device (201) to transmit, as at least part of an operation of obtaining an image including the plurality of scenes based on the first information and the second information, a request associated with generating an image including the first information and the second information to an external electronic device including a third model trained to generate an image based on inputting information associated with a description. The instructions, when executed by the processor (220), may cause the electronic device (201) to receive, from the external electronic device, an image including the plurality of scenes generated by the external electronic device in response to the request, as at least part of an operation of obtaining an image including the plurality of scenes based on the first information and the second information.

[0186] In one embodiment, the at least one primary object may include at least one of a person or voice information corresponding to the person.

[0187] In one embodiment, at least a portion of the operation of obtaining the first content corresponding to the at least one primary object and at least a portion of the operation of causing the image including the plurality of scenes to be obtained may be performed in parallel.

[0188] A method according to one embodiment may include an operation of identifying first information associated with a description corresponding to a plurality of scenes. The method may include an operation of identifying information associated with a description corresponding to at least one main object included in at least some of the plurality of scenes based on the first information. The method may include an operation of inputting information associated with a description corresponding to the at least one main object to a first model configured to output content based on input of information associated with a scene, thereby obtaining first content corresponding to the at least one main object based on information output from the first model. The method may include an operation of identifying second information corresponding to the at least one main object based on a user input for the first content. The method may include an operation of acquiring an image including the plurality of scenes based on the first information and the second information.

[0189] In one embodiment, the method may further include an operation of acquiring a first image including an image of the at least one main object based on the first information and the second information. The method may further include an operation of providing the first image as an input for acquiring at least one image corresponding to at least one scene including the at least one main object.

[0190] In one embodiment, the operation of identifying information associated with a description corresponding to the at least one main object included in at least some of the plurality of scenes based on the first information may include an operation of identifying the at least one main object based on at least one of information associated with a frequency in the description corresponding to the plurality of scenes or information associated with a correlation between the plurality of objects, among the plurality of objects included in the plurality of scenes.

[0191] In one embodiment, the operation of confirming first information associated with a description corresponding to the plurality of scenes may include an operation of confirming a user input associated with a first story. The operation of confirming first information associated with a description corresponding to the plurality of scenes may include an operation of inputting information corresponding to the user input into a second model trained to output text information associated with a description corresponding to the plurality of scenes based on inputting information associated with the story, thereby obtaining text information associated with a description corresponding to at least some of the plurality of scenes based on information output from the second model.

[0192] In one embodiment, the method may further include storing second information corresponding to the at least one primary object. The method may further include determining whether there is a correlation between the first story and the second story based on a confirmation of a request to generate an image associated with the second story. The method may further include obtaining an image associated with the second story using the second information based on a confirmation that the second story is correlated with the first story.

[0193] In one embodiment, the method may further include providing the first content including text information associated with a description corresponding to the at least one primary object. The operation of confirming the second information corresponding to the at least one primary object based on a user input for the first content may include confirming a change in the text information included in the first content.

[0194] In one embodiment, the method may further comprise providing the first content including a preview image corresponding to the at least one primary object.

[0195] In one embodiment, the operation of obtaining an image including the plurality of scenes based on the first information and the second information may include an operation of obtaining an image including the plurality of scenes based on information output from the third model by inputting the first information and the second information into a third model set to generate an image based on inputting information associated with a description.

[0196] In one embodiment, the operation of obtaining an image including the plurality of scenes based on the first information and the second information may include an operation of transmitting a request related to generation of an image including the first information and the second information to an external electronic device including a third model configured to generate an image based on inputting information associated with a description. The operation of obtaining an image including the plurality of scenes based on the first information and the second information may include an operation of receiving, from the external electronic device, an image including the plurality of scenes generated by the external electronic device in response to the request.

[0197] In one embodiment, a non-transitory computer-readable medium having recorded thereon computer-executable instructions, wherein the computer-executable instructions, when executed by a processor (220) of an electronic device (201), may cause the electronic device (201) to receive a request associated with acquisition of an image including a plurality of scenes. The computer-executable instructions, when executed by the processor (220) of the electronic device (201), may cause the electronic device (201) to acquire first content corresponding to at least one main object to be included in the image based on at least a portion of a description associated with the image identified based on the request. The computer-executable instructions, when executed by the processor (220) of the electronic device (201), may cause the electronic device (201) to provide the first content through a display. The computer-executable instructions, when executed by the processor (220) of the electronic device (201), may cause the electronic device (201) to identify information corresponding to the at least one main object based on a user input for the first content. The computer-executable instructions, when executed by the processor (220) of the electronic device (201), may cause the electronic device (201) to obtain an image including the plurality of scenes based on the description and the information corresponding to the at least one main object.

[0198] Additionally, the structure of the data used in the embodiments of the present document described above can be recorded on a computer-readable recording medium through various means. The computer-readable recording medium includes storage media such as magnetic storage media (e.g., ROM, floppy disk, hard disk) and optical reading media (e.g., CD-ROM, DVD).

[0199] An electronic device according to an embodiment disclosed in this document may take various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a home appliance. The electronic device according to an embodiment of this document is not limited to the aforementioned devices.

[0200] It should be understood that the embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0201] The term "module" used in one embodiment of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0202] An embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, "non-transitory" simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0203] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0204] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In the electronic device (201), Memory (230) for storing instructions; and A processor (220) is included, and the instructions, when executed by the processor (220), cause the electronic device (201) to: Check the first information associated with the description corresponding to multiple scenes, Based on the above first information, at least one main object included in at least some of the plurality of scenes is identified, By inputting a description corresponding to at least one main object into a first model trained to output content based on input information related to a scene, first content corresponding to at least one main object is obtained based on information output from the first model, An electronic device (201) that causes an image including the plurality of scenes to be acquired based on the first information and the information about the first content, wherein the at least one main object is consistently visualized in some of the plurality of scenes.

2. In paragraph 1, The above instructions, when executed by the processor (220), cause the electronic device (201) to: Based on the user input for the first content, causing the second information corresponding to the at least one main object to be identified, The above video is further obtained based on the second information, electronic device (201).

3. In paragraph 1 or 2, The above instructions, when executed by the processor (220), cause the electronic device (201) to: Based on the first information, the information about the first content, and the second information, a first image including an image of at least one main object is obtained, An electronic device (201) that causes the first image to be provided as an input for acquisition of at least one image corresponding to at least one scene including at least one main object.

4. In any one of paragraphs 1 to 3, The above instructions, when executed by the processor (220), cause the electronic device (201) to: An electronic device (201) that causes the at least one main object to be identified among a plurality of objects included in the plurality of scenes, based on at least one of a frequency of information corresponding to the main object identified in the first information, a relative position of information corresponding to the main object in information similar to the first information, a relative frequency of information corresponding to the main object in information, or a description level of a description associated with the main object in the first information, as at least part of an operation of identifying the at least one main object included in at least some of the plurality of scenes based on the first information.

5. In any one of paragraphs 1 to 4, The above instructions, when executed by the processor (220), cause the electronic device (201) to, at least as part of an operation of: identifying first information associated with a description corresponding to the plurality of scenes; Verify user input associated with a first story, where the story includes descriptions of multiple scenes, An electronic device (201) that inputs information corresponding to the user input to a second model trained to output text information associated with descriptions corresponding to a plurality of scenes based on inputting information related to the story, thereby causing text information associated with descriptions corresponding to at least some of the plurality of scenes to be obtained based on information output from the second model.

6. In any one of paragraphs 1 to 5, The above instructions, when executed by the processor (220), cause the electronic device (201) to: Store at least one of information about the first content or second information corresponding to the at least one main object, Based on the confirmation of the request for creation of a video related to the second story, whether there is a correlation between the first story and the second story is confirmed, An electronic device (201) that causes an image associated with the second story to be acquired using at least one of information about the first content or the second information, based on confirmation that the second story is related to the first story.

7. In any one of paragraphs 2 to 5, The above instructions, when executed by the processor (220), cause the electronic device (201) to: Providing the first content including text information associated with a description corresponding to at least one of the main objects, An electronic device (201) that causes at least one of: designation of at least one primary object, approval of the first content, or change of the text information included in the first content, to be confirmed as at least part of an operation of confirming the second information corresponding to the at least one primary object based on a user input for the first content.

8. In any one of paragraphs 1 to 6, The above instructions, when executed by the processor (220), cause the electronic device (201) to: Causing the first content to be provided, wherein the first content includes a preview image corresponding to at least one main object; An electronic device (201) wherein the resolution of the main object included in the above preview image may be lower than the resolution of the main object included in the above acquired image.

9. In any one of paragraphs 1 to 8, The above instructions, when executed by the processor (220), cause the electronic device (201) to: An electronic device (201) that causes an image including the plurality of scenes to be obtained based on information output from the third model by inputting the first information and information about the first content into a third model trained to generate an image based on inputting information related to a description, at least as a part of an operation of obtaining an image including the plurality of scenes based on information about the first information and the first content.

10. In any one of paragraphs 1 to 9, The above instructions, when executed by the processor (220), cause the electronic device (201) to obtain an image including the plurality of scenes based on the first information and the information about the first content, at least as part of: Transmitting a request associated with image generation including information about the first information and the first content to an external electronic device including a third model trained to generate an image based on inputting information associated with the description, An electronic device (201) that causes an image including the plurality of scenes generated by the external electronic device in response to the request to be received from the external electronic device.

11. In any one of paragraphs 1 to 10, An electronic device (201), wherein the at least one main object comprises at least one of a main background, a main character, or voice information corresponding to the main character.

12. In a method performed by an electronic device (201), An action of confirming (401) first information associated with a description corresponding to multiple scenes; An operation of identifying (403) at least one main object included in at least some of the plurality of scenes based on the first information; An operation of obtaining (405) first content corresponding to at least one main object based on information output from the first model by inputting a description corresponding to the at least one main object into a first model set to output content based on inputting information related to a scene; and A method comprising: obtaining (409) a video including the plurality of scenes based on the first information and the information about the first content, wherein the at least one main object is consistently visualized in some of the plurality of scenes.

13. In paragraph 12, Further comprising an action of confirming second information corresponding to at least one main object based on a user input for the first content, A method wherein the above video is further acquired based on the second information.

14. In paragraph 12 or 13, An operation of obtaining (601) a first image including an image of at least one main object based on the first information, information about the first content, and the second information; and A method further comprising the action of providing (603) the first image as an input for obtaining at least one image corresponding to at least one scene including at least one main object.

15. In a non-transitory computer-readable medium having recorded thereon computer-executable instructions, the computer-executable instructions, when executed by a processor (220) of an electronic device (201), cause the electronic device (201) to: Receive a request related to acquisition of an image containing multiple scenes, Based on the first information associated with the description associated with the plurality of scenes identified based on the above request, at least one main object included in at least some of the plurality of scenes is identified, By inputting a description corresponding to at least one main object into a first model set to output content based on inputting information related to a scene, first content corresponding to at least one main object is obtained based on information output from the first model, A computer-readable medium that causes an image including the plurality of scenes to be acquired based on the first information and the information about the first content, wherein the at least one main object is consistently visualized in at least some of the plurality of scenes.

Citation Information

Patent Citations

  • Apparatus and method for text-based performance pre-visualization

    KR1020160073750A

  • Method and Apparatus for converting text to scene

    KR1020160078703A

  • Method and system for training image segmentation model

    KR1020230155870A

  • Converting method food waste into compost

    KR102370596B1

  • Video generation method and server performing thereof

    KR102560609B1