Electronic device and method for generating story content data by using artificial intelligence model in electronic device

The electronic device generates personalized story content by processing user-selected data with AI models, addressing the lack of integrated text and image data creation in existing devices, resulting in engaging narratives through on-device and external AI model collaboration.

WO2026010489A1PCT designated stage Publication Date: 2026-01-08SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/099769
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-30
Filing Date
2025-03-13
Publication Date
2026-01-08

AI Technical Summary

Technical Problem

Existing electronic devices lack the capability to automatically generate personalized story content using AI models based on user-selected photos or videos, failing to effectively integrate text and image data to create engaging narratives.

Method used

An electronic device equipped with AI models processes user-selected content data to generate story content, including text and image data, by classifying and summarizing the content, and generating descriptive text adjacent to the images, utilizing both on-device and external AI models for enhanced functionality.

Benefits of technology

The solution enables the creation of personalized story content that integrates text and images, providing a user-friendly and engaging narrative experience, leveraging both on-device and external AI models for improved performance and adaptability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025099769_08012026_PF_FP_ABST
    Figure KR2025099769_08012026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device according to an embodiment may comprise a display, at least one processor, and a memory for storing instructions, wherein the instructions, when executed individually or collectively by the at least one processor, cause the electronic device to: generate a first text on the basis of content data included in a first input; extract, from the memory, at least one piece of first image data associated with the first text; generate a second text through a first AI model (231b in figure 2b) on the basis of the at least one piece of first image data; and generate and output first result data based on the at least one piece of first image data and the second text, wherein the first result data includes at least one of image data, text data, and sound data. Other embodiments may be included.
Need to check novelty before this filing date? Find Prior Art

Description

A method for generating story content data using artificial intelligence models in electronic devices and electronic devices.

[0001] The present disclosure relates to an electronic device and a method for generating story content data using an artificial intelligence model in the electronic device.

[0002] Electronic devices may provide features that automatically create stories by analyzing photos or videos stored on the electronic device.

[0003] An electronic device can classify photos or videos stored in the electronic device based on location, date, and time, and provide the classified photos or videos as a single story with music to a user of the electronic device.

[0004] In various applications, a topic script can be generated through an AI model based on content data of interest selected by the user, and a description can be generated through an AI model based on video data related to the topic searched on an electronic device, video data acquired from an external device, and video data generated through an AI model, and personalized story content can be provided based on the video data and description.

[0005] An electronic device according to an embodiment may include a display, at least one processor, and a memory storing instructions. The instructions according to an embodiment may be configured to cause the electronic device to display first content data through the display when individually or collectively executed by the at least one processor. The instructions according to an embodiment may be configured to cause the electronic device to obtain a first text based on at least a portion of the first content data selected by the first input based on confirming a first input for selecting at least a portion of the first content data to generate story content data. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to, based on identifying the first input for selecting at least a portion of the first content data to generate the story content data, obtain at least one first image data from the memory based on at least a portion of the first content data selected by the first input. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, and when the story content data including the first text and the at least one first image data is displayed, the first text may be displayed adjacent to the at least one first image data and used as a description of the at least one first image data.

[0006] An electronic device according to an embodiment may include a display, at least one processor, and a memory storing instructions. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a first text based on content data included in a first input. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to extract at least one first image data associated with the first text from the memory. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate a second text through a first AI model (231b of FIG. 2b) based on the at least one first image data. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, cause the electronic device to generate and output first result data based on the at least one first image data and the second text, wherein the first result data may include at least one of image data, text data, and audio data.

[0007] According to one embodiment, a method for generating story content data using an artificial intelligence model in an electronic device may include an operation of displaying first content data through a display of the electronic device. According to one embodiment, the method may include an operation of obtaining a first text based on at least a portion of the first content data selected by the first input, based on confirming a first input for selecting at least a portion of the first content data to generate the story content data. According to one embodiment, the method may include an operation of obtaining at least one first image data from a memory of the electronic device based on at least a portion of the first content data selected by the first input, based on confirming the first input for selecting at least a portion of the first content data to generate the story content data. According to one embodiment, the method includes an operation of generating story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, and when the story content data including the first text and the at least one first image data is displayed, the first text is displayed adjacent to the at least one first image data and can be used as a description of the at least one first image data.

[0008] In one embodiment, a non-volatile storage medium storing commands, wherein the commands are configured to cause the electronic device to perform at least one operation when executed by the electronic device, wherein the at least one operation may include an operation of displaying first content data through a display of the electronic device. In one embodiment, the at least one operation may include an operation of obtaining a first text based on at least a portion of the first content data selected by the first input, based on confirming a first input for selecting at least a portion of the first content data to generate story content data. In one embodiment, the at least one operation may include an operation of obtaining at least one first image data from a memory of the electronic device based on at least a portion of the first content data selected by the first input, based on confirming the first input for selecting at least a portion of the first content data to generate the story content data. In one embodiment, the at least one operation includes generating story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, and when the story content data including the first text and the at least one first image data is displayed, the first text may be displayed adjacent to the at least one first image data and used as a description of the at least one first image data.

[0009] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.

[0010] FIG. 2A is a block diagram of an electronic device according to one embodiment.

[0011] FIG. 2b is a block diagram illustrating the configuration of a processor and an AI model according to one embodiment.

[0012] FIG. 3 is a diagram for explaining an operation of generating story content data according to one embodiment.

[0013] FIGS. 4A, 4B, and 4C are drawings illustrating generation of a first text in an electronic device according to one embodiment.

[0014] FIGS. 5A and 5B are drawings for explaining a modification operation of a first text in an electronic device according to one embodiment.

[0015] FIGS. 6A, 6B, 6C, 6D, and 6E are drawings for explaining an operation of generating story content data in an electronic device according to one embodiment.

[0016] FIG. 7 is a diagram for explaining an operation of generating story content data in an electronic device according to one embodiment.

[0017] FIG. 8 is a diagram for explaining an operation of generating story content data in an electronic device according to one embodiment.

[0018] FIG. 9 is a flowchart illustrating an operation of generating story content data in an electronic device according to one embodiment.

[0019] FIG. 10 is a flowchart illustrating an operation of generating story content data in an electronic device according to one embodiment.

[0020] FIG. 1 is a block diagram of an electronic device (101) within a network environment (100) according to an embodiment. Referring to FIG. 1, in the network environment (100), the electronic device (101) may communicate with the electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (104) or the server (108) via a second network (199) (e.g., a long-range wireless communication network). According to an embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).

[0021] The processor (120) may control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing, for example, software (e.g., a program (140)), and may perform various data processing or calculations. According to one embodiment, as at least a part of the data processing or calculation, the processor (120) may store a command or data received from another component (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the command or data stored in the volatile memory (132), and store the resulting data in a non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or a secondary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together therewith. For example, if the electronic device (101) includes a main processor (121) and a secondary processor (123), the secondary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a specified function. The secondary processor (123) may be implemented separately from the main processor (121) or as a part thereof.

[0022] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0023] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto. The memory (130) can include volatile memory (132) or non-volatile memory (134).

[0024] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).

[0025] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0026] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0027] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0028] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).

[0029] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0030] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0031] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0032] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0033] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.

[0034] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).

[0035] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0036] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).

[0037] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.

[0038] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas, for example, by the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device via the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).

[0039] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0040] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0041] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server utilizing machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.

[0042] FIG. 2a is a block diagram of an electronic device according to an embodiment, and FIG. 2b is a block diagram for explaining the configuration of a processor and an AI model according to an embodiment.

[0043] Referring to the above FIGS. 2A and 2B, the electronic device (201) may include a processor (220), a memory (230), a display (260), and a communication circuit (290).

[0044] According to one embodiment, the processor (220) may perform overall control operations of the electronic device (201). The processor (220) according to one embodiment may execute software (e.g., the program (140) of FIG. 1) to control at least one other component (e.g., a hardware or software component) of the electronic device (201) connected to the processor (220), and may perform data processing or calculations based on instructions. An instruction according to one embodiment may include a command configured in a machine language that can be processed by the electronic device (201) or the processor (220). For example, the instruction may include a command corresponding to an operation instruction used in a program.

[0045] According to one embodiment, a processor (220) may generate a first text (e.g., topic script) based on at least one content data selected by a first input of a user.

[0046] At least one content data according to one embodiment may mean an object including at least one still image data (image data), video data, and / or link data displayed on a screen of an electronic device.

[0047] According to one embodiment, the processor (220) may, upon confirming selection of at least one content data included in the image data based on a first input of the user, generate a first text for generating story content data related to the at least one content data through the first AI model (231a).

[0048] According to one embodiment, when a processor (220) confirms a gesture (e.g., a circle gesture) for determining an area as a first input of a user, the processor (220) may generate a first text (e.g., a search word or a search sentence) for generating story content data related to some content data (e.g., at least a part of the first content) within an area confirmed by the gesture for determining an area in image data (e.g., first content data) through a first AI model (231a).

[0049] According to one embodiment, when a processor (220) confirms an operation (e.g., screen capture) that can confirm selection of the entire video data (e.g., first content data) as a first input of a user, the processor (220) can generate a first text (e.g., a search word or search sentence) for generating story content data related to the entire content data included in the video data (e.g., first content) through a first AI model (231a).

[0050] According to one embodiment, the processor (220) may display a user interface (e.g., a menu) for generating story content data upon confirming selection of at least one content data based on a first input of the user.

[0051] According to one embodiment, the processor (220) may, upon confirming the selection of at least one content data included in the image data based on the user's first input, transmit a prompt requesting the generation of content data and first text to the first AI model (231a).

[0052] According to one embodiment, a processor (220) may transmit at least one content data selected based on a first input of a user to a first AI model (231a).

[0053] According to one embodiment, a processor (220) may process at least one content data selected based on a user's first input into a form that can be transmitted to a first AI model (231a) (e.g., preprocessing such as normalization into a form suitable for model input, and / or LangChain form, etc.) and then transmit the processed data to the first AI model (231a).

[0054] According to one embodiment, a processor (220) can check at least one content data to be transmitted to the first AI model (231a) from a content DB of a memory (230) in which content data processed in a form that can be transmitted to the first AI model (231a) is stored.

[0055] According to one embodiment, a processor (220) may process at least one content data into a form that can be transmitted to a first AI model (231a) using a data conversion model included in an electronic device (201) or a server (e.g., preprocessing such as normalization into a form suitable for model input, and / or LangChain form, etc.) and then transmit the processed content data to the first AI model (231a).

[0056] According to one embodiment, the processor (220) may include, in the prompt transmitted to the first AI model (231a), a request for determining a main keyword from at least one input content data and generating a concise first text based on the determined main keyword.

[0057] According to one embodiment, the processor (220) may include, in a prompt transmitted to the first AI model (231a), a request for determining a main object among at least one input content data and converting the main object into a form usable by a search engine (e.g., vector form, language description form of content, main object tag information of content, etc.) in order to facilitate a search for image data based on the first text.

[0058] According to one embodiment, a processor (220) may transmit to a first AI model (231a) at least one content data selected by a user's first input, application information providing at least one content data, and information (e.g., name information, relationship information, location information, date information, and / or place name information) of text data included in at least one content data.

[0059] According to one embodiment, a first AI model (231a) can, when it is confirmed that at least one content data includes text and some image data among image data, summarize the text included in the at least one content data and generate a first text that describes some image data included in the content data as text.

[0060] According to one embodiment, the first text may be generated for use as a story subject for image search and / or generation of story content data.

[0061] According to one embodiment, the first AI model (231a) can generate a first text summarizing the text included in the content data when it is confirmed that at least one content data includes text.

[0062] According to one embodiment, the first AI model (231a) can generate a first text that describes some of the image data included in the content data in text when it is confirmed that at least one content data includes some of the image data.

[0063] The first AI model (231a) according to one embodiment may include a generative AI model, may include an LLM (Large Language Model), and may be an on-device model or an external AI model (281).

[0064] According to one embodiment, a processor (220) can confirm selection of content data included in image data based on a first input of a user, and generate text directly input by the user as first text.

[0065] According to one embodiment, a processor (220) can confirm a selection of content data included in image data based on a first input of a user, and generate the text selected by the user as a first text without processing the text.

[0066] According to one embodiment, the processor (220) can generate a first text based on tag information included in image data selected by a user among image data stored in the memory (230).

[0067] According to one embodiment, the processor (220) can generate a first text using a keyword that identifies a person included in image data selected by a user among image data stored in the memory (230).

[0068] According to one embodiment, the processor (220) can generate a first text in the form of writing a description generated through the first AI (231a) based on image data selected by the user from among image data stored in the memory (230).

[0069] According to one embodiment, a processor (220) may generate a first text in a manner of describing a description based on at least one image data related to image data selected by a user among image data stored in a memory (230).

[0070] According to one embodiment, the processor (220) can generate a first text based on voice or text input by a user through an AI Assistant.

[0071] According to one embodiment, the processor (220) can verify the validity of the first text when it confirms the creation of the first text.

[0072] According to one embodiment, the processor (220) may determine the validity of the first text based on whether the first text is related to user information stored in the electronic device (201).

[0073] According to one embodiment, the processor (220) may compare the main keyword among at least one keyword included in the first text with tag information stored in relation to image data (e.g., still image data and / or video data) stored in the electronic device (201).

[0074] According to one embodiment, the processor (220) may display a message indicating that story content data cannot be created with the first text if tag information related to the first text is not found among the tag information based on the comparison, or may provide a user interface for modifying the first text.

[0075] According to one embodiment, the processor (220) may perform an operation of analyzing and / or processing the first text upon confirming the creation of the first text.

[0076] According to one embodiment, a processor (220) may perform an operation of analyzing and / or processing a first text while identifying and deleting at least one keyword (e.g., a specific place name) that represents a narrow range among keywords included in the first text.

[0077] In one embodiment, the processor (220) may perform an operation of analyzing and / or processing the first text while changing at least one keyword included in the first text to a second keyword (e.g., campground) that represents a broader range within a similar (same) category as the first keyword (e.g., AAA campground).

[0078] According to one embodiment, a processor (220) may perform an operation of analyzing and / or processing the first text while selecting various additional keywords within a similar (identical) category to at least one keyword (e.g., campsite) included in the first text and providing them as candidate keywords (e.g., campsite, and / or glamping site, etc.).

[0079] According to one embodiment, a processor (220) may perform an operation of analyzing and / or processing a first text using user information.

[0080] According to one embodiment, the processor (220), when the user DB of the memory (230) includes information on a place visited by the user (e.g., BBB Campground), may perform an operation of extracting a third keyword (e.g., BBB Campground) included in a similar (same) category as the first keyword (e.g., AAA Campground) included in the first text from the place information (e.g., BBB Campground) and changing the first keyword (e.g., AAA Campground) to the third keyword (e.g., BBB Campground) while analyzing and / or processing the first text.

[0081] According to one embodiment, the processor (220) can output the first text when it confirms the creation of the first text, and output the first text modified through the user's second input.

[0082] According to one embodiment, the processor (220) may output a first text through a display (260) or an audio output device (not shown) included in the electronic device, and provide a user interface for modifying the first text.

[0083] According to one embodiment, the processor (220) may, upon confirming the modification of the first text based on the user's second input through the user interface for modifying the first text, output the modified first text through the display (260) or an audio output device (not shown) included in the electronic device (201).

[0084] According to one embodiment, the processor (220) can, upon confirming the creation of the first text, extract at least one first image data associated with the first text from the memory (230) or the user DB of the cloud.

[0085] Image data according to one embodiment may include still image data and / or moving image data.

[0086] According to one embodiment, the processor (220) can search and extract at least one first image data associated with at least one keyword included in the first text from the memory (230) or the user DB of the cloud.

[0087] According to one embodiment, the processor (220) can search for and extract at least one first image data associated with at least one keyword included in a first text by using a vector DB of a memory (230) in which image data is stored in a vector format or a text DB of a memory (230) in which image data is stored in a text format.

[0088] According to one embodiment, a processor (220) analyzes content data selected based on a first input of a user to identify and separate a main object from at least one object included in the content data, and searches for and extracts at least one first image data associated with the main object from a user DB of a memory (230) or a cloud.

[0089] According to one embodiment, if the processor (220) fails to obtain at least a portion of at least one first image data from the memory (230) or the user DB of the cloud, the processor (220) may obtain at least one second image data associated with some content data (e.g., at least a portion of the first content) that has confirmed the user's first input for generating story content from the image data (e.g., the first content data) from an external device (e.g., an external electronic device) of the electronic device (201) through crawling or RAG (Retrieval-Augmented Generation).

[0090] In an embodiment, when the processor (220) cannot extract image data associated with at least some keywords among at least one keyword included in the first text (e.g., search word or search sentence) from the user DB of the memory (230) or the cloud, the processor (220) can obtain at least one second image data associated with at least some keywords for which the associated image data was not extracted from the user DB of the memory or the cloud through crawling or RAG (Retrieval-Augmented Generation) from outside of the electronic device (201) (e.g., an external electronic device).

[0091] In an embodiment, when the processor (220) cannot extract image data associated with at least some keywords among at least one keyword included in a first text (e.g., a search word or a search sentence) from the user DB of the memory (230) or the cloud, the processor (220) determines whether external data can be utilized for at least some keywords for which associated image data was not extracted from the user DB of the memory (230) or the cloud, and then, using keywords determined to be capable of utilizing external data, obtains at least one second image data associated with at least some keywords for which associated image data was not extracted from the user DB of the memory (230) or the cloud through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device (201) (e.g., an external electronic device).

[0092] In an embodiment, when the processor (220) cannot extract (obtain) image data associated with at least some keywords among at least one keyword included in a first text (e.g., a search word or a search sentence) from the memory (230) or the user DB of the cloud, the processor (220) can obtain at least one second image data associated with at least some keywords for which the associated image data was not extracted from the memory (230) or the user DB of the cloud through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device (201) based on location information and date and time information included in metadata of at least one first image data extracted from the user DB of the memory (230) or the cloud.

[0093] According to an embodiment, the processor (220) may, when obtaining at least one second image data through crawling or RAG (Retrieval-Augmented Generation), determine whether to use the at least one second image data as image data of story content data associated with content data based on a correlation between the at least one first image data and the at least one second image data. For example, the processor (220) may determine to use the at least one second image data as image data of story content data associated with content data based on season information, place information, and / or person information included in the at least one first image data if the at least one first image data and the at least one second image data have the same season or place information.

[0094] According to one embodiment, if the processor (220) fails to obtain at least a portion of at least one first image data from the memory (230) or the user DB of the cloud, the processor (220) may obtain at least one second image data associated with some content data (e.g., at least a portion of the first content) that has confirmed the user's first input for generating story content from the image data (e.g., the first content data) through the second AI model (231b).

[0095] In an embodiment, the processor (220) may, when it is unable to extract image data associated with at least some of the keywords included in the first text from the memory (230) or the cloud DB, generate at least one third image data associated with at least some of the keywords for which the associated image data was not extracted from the memory (230) or the cloud user DB through the second AI model (231b).

[0096] In one embodiment, when the processor (220) cannot extract image data associated with at least some of the keywords included in the first text from the memory (230) or the user DB of the cloud, the processor (220) may transmit a prompt requesting the generation of at least some of the keywords of the first text, at least one first image data, at least one second image data, and / or at least one third image data associated with at least some of the keywords to the second AI model (231b).

[0097] According to one embodiment, the processor (220) may generate a description including specific text information about content data selected based on a first input of a user based on a first text, and transmit the description to the second AI model (231b), so that the second AI model (231b) generates at least one third image data similar to at least one first image data.

[0098] According to one embodiment, the processor (220) may generate a prompt to generate at least one third image data similar to at least one first image data through the second AI model (231b).

[0099] According to an embodiment, the processor (220) may generate a prompt based on at least one keyword included in the first text, the first text, information of searched image data, or personalized data acquired from a user database and / or memory, and input the generated prompt as an input value of the second AI model (231b) to generate at least one third image data similar to at least one first image data through the second AI model (231b). According to an embodiment, when the processor (220) generates at least one third image data through the second AI model (231b), the processor (220) may determine the suitability of the at least one third image data generated through the second AI model (231b) by using a similarity between the first text and the at least one third image data or a captioning function that describes the at least one third image data.

[0100] According to one embodiment, the processor (220) may request the re-generation of at least one third image data using the second AI model (231b) when the suitability of at least one third image data generated through the second AI model (231b) is below a standard for evaluation.

[0101] A second AI model (231b) according to one embodiment can generate third image data associated with at least some keywords by using at least some of the first text, at least some keywords of the first text, at least one first image data, and at least one second image data.

[0102] The second AI model (231b) according to one embodiment may include a generative AI model, may include an LLM (Large Language Model) and / or an Image representation AI model, and may be an on-device model or an external AI model (281).

[0103] The second AI model (231b) according to one embodiment may be the same AI model as the first AI model (231a), or may be a separate AI model.

[0104] According to an embodiment, the processor (220) may, if it is determined that at least one of the first image data extracted through a user DB of a memory (230) or a cloud based on the first text after generating the first text, at least one second image data acquired through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device (201), or at least one third image data generated through a second AI model (231b) needs to be modified, the processor (220) may notify the user of the modification of at least one image data for generating story content data.

[0105] According to one embodiment, the processor (220) can, when it confirms that at least one image data for generating story content data includes sensitive user information (e.g., resident registration number or passport number) using a function such as optical character recognition (OCR), delete the sensitive user information from the at least one image data for generating story content data, and notify the user of the deletion of the sensitive user information from the at least one image data for generating story content data.

[0106] According to one embodiment, the processor (220) may perform anonymization processing on at least one generated image data depending on the degree of a set security level (e.g., privacy level).

[0107] According to one embodiment, the processor (220) may perform anonymization processing operations, such as covering or deleting a portion of at least one face included in at least one image data, to protect personal information.

[0108] A processor (220) according to one embodiment may perform an unidentified processing operation on various objects such as text data, still image data (image data), and / or video data. A processor (220) according to one embodiment may perform an unidentified processing operation on data previously designated by a user for various objects such as text data, still image data (image data), and / or video data.

[0109] According to one embodiment, the processor (220) determines whether data requiring a de-identification operation is included in at least one image data, and then transmits information related to the data requiring a de-identification operation as input to an AI model to generate image data with the data obscured or deleted.

[0110] According to one embodiment, the processor (220), when it is confirmed that an object unrelated to the story content data is included among at least one piece of image data (e.g., still image and / or video data) for generating story content data, the processor (220) may use an in-painting technique that naturally fills in the missing portion to intentionally omit and then fill in the object unrelated to the story content data, thereby deleting the object unrelated to the story content data, and may notify the user of the deletion of the object unrelated to the story content data among at least one piece of image data for generating story content data.

[0111] According to one embodiment, a processor (220) may, when the facial expression of a person included in at least one image data for generating story content data is unnatural, change the facial expression of the person to a natural facial expression of the same person based on the image data using an in-painting technique, and notify the user of the change in facial expression of the person included in at least one image data for generating story content data.

[0112] According to an embodiment, when the processor (220) confirms that at least one piece of image data for generating story content data includes image data (e.g., image data of a mother smiling brightly after receiving flowers and a letter) that exactly matches the first text (e.g., a gift given on Mother's Day), the processor (220) may edit the image data that exactly matches the first text among the at least one piece of image data for generating story content data by using a technology such as Live Effect (e.g., a technology such as 3D Photography and / or Cinemagraph) or a technology for generating a moving image using a still image through an AI model to emphasize the image data that exactly matches the first text among the at least one piece of image data for generating story content data, and may notify the user that the image data that exactly matches the first text among the at least one piece of image data for generating story content data has been edited by emphasizing the image data.

[0113] According to an embodiment, the processor (220) verifies the inclusion of image data (e.g., image data of a user supporting the Leaning Tower of Pisa) that exactly matches the first text (e.g., a trip to Pisa with friends in midsummer) among at least one image data for generating story content data, but if it verifies the presence of people unrelated to the generation of the story content data in the image data that exactly matches the first text, it performs editing to delete people unrelated to the generation of the story content data or add people related to the generation of the story content data using an in-painting function, and notifies the user to perform editing to delete people unrelated to the generation of the story content data or add people related to the generation of the story content data.

[0114] According to one embodiment, when the processor (220) determines that modification of at least one image data for generating story content data is necessary, the processor (220) may edit at least one image data for generating story content data using various methods.

[0115] According to one embodiment, the processor (220) can generate a second text through a third AI model (231c) based on at least one first image data and user information.

[0116] According to one embodiment, the processor (220) can generate a second text through a third AI model (231c) based on at least one second image data and user information.

[0117] According to one embodiment, the processor (220) can generate a second text through a third AI model (231c) based on at least one first image data, at least one second image data, and at least some of the user information.

[0118] According to one embodiment, the processor (220) may transmit a prompt requesting the generation of a second text for at least one image data based on at least one of the first image data, at least one second image data, or at least one third image data and user information to the third AI model (231c).

[0119] According to one embodiment, the third AI model (231c) can generate personalized second text for at least one piece of image data using at least one piece of information from the user (e.g., name based on face, age and / or personality information, user's preferred places and / or objects, family relationships, etc.).

[0120] A third AI model (231c) according to one embodiment can generate one second text (e.g., a description in sentence form or a tag form) for at least one image data or at least one second text (e.g., a description in sentence form or a tag form) corresponding to at least one image data.

[0121] The third AI model (231c) according to one embodiment may include a generative AI model, may include an LLM (Large Language Model), and may be an on-device model or an external AI model (281).

[0122] The third AI model (231c) according to one embodiment may be the same AI model as at least one of the first AI model (231a) or the second AI model (231b), or may be a separate AI model.

[0123] A processor (220) according to one embodiment can generate first result data based on at least one first image data and a second text.

[0124] According to one embodiment, a processor (220) can generate first result data through a fourth AI model (231d) based on at least one first image data and a second text.

[0125] According to one embodiment, a processor (220) can generate first result data based on at least one second image data and a second text.

[0126] According to one embodiment, the processor (220) can generate first result data through the fourth AI model (231d) based on at least one second image data and a second text.

[0127] According to one embodiment, a processor (220) may generate first result data (e.g., story content data) based on at least one first image data, at least one second image data, at least one third image data, and at least a portion of second text.

[0128] According to one embodiment, a processor (220) may generate first result data (e.g., story content data) through a fourth AI model (231d) based on at least one first image data, at least one second image data, at least one third image data, and at least a portion of the second text.

[0129] A processor (220) according to one embodiment may output first result data including at least one of image data, text data, and audio data.

[0130] According to one embodiment, the processor (220) can output first result data including at least one of image data, text data, and audio data generated through the fourth AI model (231d).

[0131] According to one embodiment, a processor (220) may generate personalized first result data (e.g., story content data) in a result content style determined by the device or a result content style specified by a user (e.g., journal style) using a list of image data including at least one of at least one first image data, at least one second image data, or at least one third image data, and a second text for each image data.

[0132] According to one embodiment, the processor (220) may generate personalized first result data (e.g., story content data) in a result content style determined by the device or a result content style specified by the user (e.g., journal style) using a list of image data including at least one of at least one first image data, at least one second image data, or at least one third image data, and a second text for each image data, through the fourth AI model (231d).

[0133] According to one embodiment, a processor (220) can generate first result data (story content data) by arranging a list of image data including at least one of first image data, at least one second image data, or at least one third image data and second text in an associated manner.

[0134] According to one embodiment, the processor (220) can generate first result data (story content data) by arranging a list of image data including at least one of first image data, at least one second image data, or at least one third image data and second text in a related manner through the fourth AI model (231d).

[0135] According to an embodiment, the processor (220) may generate the first result data (e.g., story content data) such that, when the first result data (e.g., story content data) including the second text data and the list of image data is displayed on the display (260), the second text data is displayed adjacent to at least one of the first image data, the at least one second image data, or the at least one third image data included in the image data list, and thus the first result data (e.g., story content data) may be used as a description for the at least one image data. According to an embodiment, the processor (220) may generate the first result data (e.g., story content data) such that, when the first result data (e.g., story content data) including the second text data and the list of image data is displayed on the display (260), the second text data is displayed adjacent to at least one of the first image data, the at least one second image data, or the at least one third image data included in the image data list, and thus the first result data (e.g., story content data) may be used as a description for the at least one image data.

[0136] According to one embodiment, a processor (220) may generate a title of first result data (e.g., story content data) based on information of image data included in a first text and image data list and / or information of a user.

[0137] According to one embodiment, the processor (220) may use the first text as a title of the first result data (e.g., story content data).

[0138] A processor (220) according to one embodiment can generate first result data in which at least one image data and a second text included in an image data list are arranged based on a time order.

[0139] The fourth AI model (231d) according to one embodiment can generate first result data by arranging at least one image data and a second text included in the image data list based on the order of time.

[0140] A processor (220) according to an embodiment can generate first result data capable of changing the arrangement of at least one image data and a second text included in an image data list in a manner responsive to a user's interaction, regardless of the order of time or storage. A fourth AI model (231d) according to an embodiment can generate first result data capable of changing the arrangement of at least one image data and a second text included in an image data list in a manner responsive to a user's interaction, regardless of the order of time or storage.

[0141] A processor (220) according to one embodiment may group image data having similarity in an image data list and arrange a second text related to the grouped image data.

[0142] A fourth AI model (231d) according to one embodiment can group image data having similarity in an image data list and arrange one second text related to the grouped image data.

[0143] The fourth AI model (231d) according to one embodiment may include a generative AI model, may include an LLM (Large Language Model), and may be an on-device model or an external AI model (281).

[0144] The fourth AI model (231d) according to one embodiment may be the same AI model as at least one of the first AI model (231a), the second AI model (231b), or the third AI model (231c), or may be a separate AI model. The processor (220) according to one embodiment may use the fourth AI model (231d) to generate a portion of the first result data.

[0145] A processor (220) according to one embodiment can adjust the strength of the fictitiousness of the first result data based on a third input of the user, and the strength of the fictitiousness can be applied to at least one of the image data, text data, and audio data included in the first result data.

[0146] According to one embodiment, when generating first result data through the fourth AI model (231d), if a user inputs text corresponding to fiction, the processor (220) may generate first result data including a portion of the fiction in the input value.

[0147] According to one embodiment, the processor (220) can transmit text (e.g., numbers or language expressions) corresponding to fiction by the user to the fourth AI model (231d).

[0148] According to one embodiment, the processor (220) may input text (e.g., numbers or language expressions) corresponding to fiction by the user as variables of a generative AI model, such as temperature, top_k, and / or top_p, in a prompting operation within the fourth AI model (231d).

[0149] According to one embodiment, the processor (220) can control text (e.g., numbers or language expressions) corresponding to fiction to be processed by the user in a processing operation such as prompting engineering within the fourth AI model (231d).

[0150] According to one embodiment, when generating a first text (topic script) through a first AI model (231a), if an input of text corresponding to fiction is received by a user, the processor (220) may generate a first text that includes a portion of the fiction in the input value.

[0151] According to one embodiment, the processor (220) may transmit text (e.g., numbers or language expressions) corresponding to fictional content (fiction) by the user to the first AI model (231a). According to one embodiment, the processor (220) may generate a first text based on an input value by considering the characteristics of the first AI model (231c), or may generate a first text by synthesizing a certain portion of fictional content to the input value by adjusting a variable for generating the first text based on text (e.g., numbers or language expressions) corresponding to fictional content (fiction) by the user.

[0152] For example, when the processor (220) receives keywords (e.g., healing, night sky, and / or bonfire) in addition to keywords (e.g., double cherry blossoms, camping, and / or tent) detected based on content data selected by the user as a first input as text corresponding to fiction when generating a first text (topic script), the processor (220) may adjust variables for generating the first text to generate a first text that synthesizes a certain portion of the fiction into the input value.

[0153] According to one embodiment, the processor (220) may combine second result data generated by an external device with first result data to generate third result data. According to one embodiment, the processor (220) may confirm that the second result data includes image data captured by the external device.

[0154] According to one embodiment, the processor (220) may receive at least one of image data stored in an external device, image data captured by the external device, metadata of image data captured by the external device, and analysis data obtained by analyzing image data captured by the external device from the external device.

[0155] According to one embodiment, the processor (220) may generate third result data by combining the first result data with the second result data stored in the memory (230) of the electronic device (201).

[0156] According to one embodiment, a processor (220) may generate a prompt using second result data and first result data or external data, and transmit the prompt to an on-device AI model (231) or an external AI model (281), thereby generating third result data through the on-device AI model (231) or the external AI model (281).

[0157] According to one embodiment, a processor (220) may generate first result data (e.g., first story content data) using first text generated based on content data selected based on a first input of a user, and, when receiving second result data (e.g., second story content data) similar to the first result data from an external device, may compare and merge at least one image data and a second text included in the second result data with the first result data to generate third result data (e.g., third story content data) that is a merge of the first result data and the second result data.

[0158] For example, the processor (220) can confirm the time order by analyzing metadata and context of at least one image data included in the second result data and at least one image data included in the second result data, and can generate the third result data by simultaneously utilizing at least one image data included in each of the first result data and the second result data and merging them. The processor (220) can confirm metadata (e.g., EXIF) of at least one image data included in each of the first result data and the second result data, or can shorten the time for generating the third result data by checking the results analyzed in advance in the form of a pre-analyzed vector or vector DB.

[0159] According to one embodiment, the processor (220) may delete or emphasize at least one object included in the first result data based on the user's fourth input, and may apply the deletion or emphasis of at least one object to at least one of image data, text data, and audio data included in the first result data.

[0160] According to one embodiment, the processor (220) may, after generating first result data in various forms (e.g., a description, a journal, a slide show, and / or a video, etc.), edit (e.g., modify, delete, or emphasize) story content data representing the first result data based on a user's input.

[0161] For example, the processor (220) may modify the story content data representing the first result data, including modifying the first text (topic script), specifying required elements and removed elements of the story content data, adjusting the intensity of fictionality and limiting the first text (topic script), editing objects or characters, etc. in at least one image data for generating the story content data, adjusting the content and tone of the second text (description), and changing the form of the final story content data.

[0162] According to one embodiment, when a user's input for modifying story content data representing first result data includes a user's speech, the processor (220) can generate story content data that meets the user's intention by understanding the user's speech and generating a first text for modification for a portion that needs to be modified during the story content data generation sequence. In the process of generating story content data that meets the user's intention, the processor (220) can assist in the editing process for the story content data, including a process of evaluating whether the story content data has been generated to meet the user's intention based on an analysis of at least one image data included in the story content data.

[0163] For example, if the processor (220) wants to delete all characters not included in the character list registered by the user in the first story content data, the processor (220) may perform an operation of selecting characters not registered by the user from at least one image data included in the first story content data, performing masking on the characters not registered by the user to delete them, and then checking whether the characters not registered by the user are included. The processor (220) may perform an editing operation by completing the generation of the second text and the generation of the story content data through at least one image data that has been finally edited (changed).

[0164] According to one embodiment, the processor (220) may, if there is intermediate result data generated or collected before reaching the step of generating final story content data, provide intermediate result data in the operations performed at each step, and may generate story content data while editing it based on user feedback while providing the intermediate result data.

[0165] The memory (230) according to one embodiment may be implemented substantially identically or similarly to the memory (130) of FIG. 1.

[0166] According to one embodiment, the memory (230) may store an on-device AI model (231), and the on-device AI model (231) may include a first AI model (231a), a second AI model (231b), a third AI model (231c), and a fourth AI model (231d), or may include at least one AI model among the first AI model (231a), the second AI model (231b), the third AI model (231c), and the fourth AI model (231d), or may include at least one of the first AI model (231a), the second AI model (231b), the third AI model (231c), and the fourth AI model (231d) as one identical AI model, or may include the first AI model (231a), the second AI model (231b), the third AI model (231c), and the fourth AI model (231d) as one identical AI model.

[0167] At least some of the on-device AI models (231) according to one embodiment may be operated as external AI models.

[0168] According to one embodiment, the on-device AI model (231) can process sensitive information such as personal information, and the external AI model can process information excluding sensitive information such as personal information.

[0169] An on-device AI model (231) according to one embodiment is an artificial intelligence model installed in an electronic device (201) and can provide various functions without a network.

[0170] According to one embodiment, a plurality of artificial intelligence models may be stored in the memory (230).

[0171] According to one embodiment, each of the plurality of AI models may be AI models learned based on a designated type of learning algorithm, implemented to receive various types of data (or content) as input, perform operations, and output (or obtain) result data.

[0172] In one embodiment, the plurality of AI models may include a generative AI model.

[0173] According to one embodiment, a generative AI model may generate and output new content (e.g., text, images, and / or computer code, etc.) based on what has been learned in response to an input prompt. For example, in an electronic device (601), learning may be performed to output specific types of result data as output data by using data of specified types as input data based on a machine learning algorithm or a deep learning algorithm, thereby generating a plurality of AI models (e.g., machine learning models and deep learning models) and storing them in the electronic device (601), or AI models learned from an external electronic device (e.g., an external server) may be transmitted to and stored in the electronic device (601). For example, in the electronic device (601), input data (input data) may be output as output data of a model learned through artificial intelligence of specified types based on a machine learning algorithm or a deep learning algorithm. Machine learning algorithms include supervised learning algorithms such as linear regression and logistic regression, unsupervised learning algorithms such as clustering, visualization and dimensionality reduction, and association rule learning, and reinforcement learning algorithms, and deep learning algorithms may include artificial neural networks (ANNs), deep neural networks (DNNs), and convolution neural networks (CNNs), and may further include various learning algorithms without being limited to those described.The AI ​​model that has completed learning includes at least one computational operation (e.g., a convolutional layer or a pooling layer) for computing input data, and can be implemented to output result data by performing an operation on the input data based on the at least one computational operation.

[0174] According to one embodiment, an application that can be connected to an external AI model (281) may be stored in the memory (230).

[0175] According to one embodiment, the external AI model (281) may include a first AI model (231a), a second AI model (231b), a third AI model (231c), and a fourth AI model (231d), or may include at least one AI model among the first AI model (231a), the second AI model (231b), the third AI model (231c), and the fourth AI model (231d), or may include the first AI model (231a), the second AI model (231b), the third AI model (231c), and the fourth AI model (231d) as one and the same AI model.

[0176] A display (260) according to one embodiment may be implemented substantially identically or similarly to the display (160) of FIG. 1.

[0177] A display (260) according to one embodiment may display story content data including at least one image data and second text (e.g., a description) related to content data selected based on a first input of a user.

[0178] A communication circuit (290) according to an embodiment may form a communication connection with an external electronic device (e.g., another electronic device or a server) through various types of communication methods, and transmit and / or receive data. As described above, the communication method may include a communication method that establishes a direct communication connection such as Bluetooth and Wi-Fi direct, a communication method that uses an access point (AP) (e.g., Wi-Fi communication), or a communication method that uses cellular communication using a base station (e.g., 3G, 4G / LTE, 5G). Since the communication circuit (290) may be implemented like the communication module (190) described above in FIG. 1, a redundant description thereof will be omitted.

[0179] FIG. 3 is a diagram for explaining an operation of generating story content data according to one embodiment.

[0180] Referring to FIG. 3, according to an embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, upon confirming a first input (e.g., a circle gesture) (a1) of a user for content data of a first application in operation 311a, confirm at least one content data selected based on the first input of the user, or upon confirming a first input (e.g., a circle gesture) (a2) of a user for content data of a second application in operation 311b, confirm at least one content data selected based on the first input of the user.

[0181] According to one embodiment, in operation 313, the electronic device may generate a first text (e.g., topic script) including at least one keyword through a first AI model (e.g., the first AI model (231a) of FIG. 2b) based on at least one content data selected by a first input of the user.

[0182] According to one embodiment, in operation 315, the electronic device may extract at least one first image data (e.g., still image data (image data) and / or video data) associated with the first text from a memory of the electronic device or a user DB of a cloud.

[0183] According to one embodiment, the electronic device can extract at least one first image data related to at least one keyword included in the first text from the memory of the electronic device or a user DB of the cloud.

[0184] In one embodiment, when an electronic device cannot extract image data associated with at least some of at least one keyword included in a first text from a memory of the electronic device or a DB of a cloud, the electronic device can obtain at least one second image data associated with at least some of the keywords for which the associated image data was not extracted from the memory of the electronic device or the user DB of the cloud from an external device or generate at least one third image data through a second AI model (the second AI model (231b) of FIG. 2b).

[0185] According to one embodiment, the electronic device may, in operation 317, generate a second text (e.g., a description) through a third AI model (e.g., the third AI model (231c) of FIG. 2b) based on at least one first image data.

[0186] According to one embodiment, the electronic device may generate a second text describing and / or summarizing at least one first image data, a second text describing each of the at least one first image data, a second text grouping the at least one first image data according to characteristics and describing each group, or a second text describing a story based on the at least one first image data.

[0187] According to one embodiment, the electronic device may convert the second text into voice data and include at least one of the converted voice data and the first text in the first result data (e.g., story content data).

[0188] In one embodiment, the electronic device may convert the second text into subtitles and include them in the first result data (e.g., story content data).

[0189] According to one embodiment, the electronic device may, in operation 319, generate first result data (e.g., story content data) based on at least one image data and second text.

[0190] According to one embodiment, the electronic device may generate first result data (e.g., story content data) through a fourth AI model (e.g., the fourth AI model (231d) of FIG. 2B) based on at least one image data and a second text, and may output the first result data including at least one of the image data, the text data, and the audio data through a display or an audio output device included in the electronic device.

[0191] FIGS. 4A, 4B, and 4C are drawings illustrating generation of a first text in an electronic device according to one embodiment.

[0192] Referring to FIG. 4A according to an embodiment, according to an embodiment, when an electronic device (201) (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) confirms selection of conversation content with the other party through a first input (e.g., a circle gesture, etc.) (b1) of the user while executing a first application (e.g., a message application), the electronic device may transmit conversation content information included in the selected area based on the first input of the user as an input value of a first AI model (e.g., the first AI model (231a) of FIG. 2B) for generating a first text (topic script). According to one embodiment, the electronic device (201) may transmit at least one of the conversation content information (e.g., text information and / or image data information), information of a first application (e.g., a message application), name information of the other party, relationship information between the other party and me, and various other types of information related to the other party as input values ​​of a first AI model (e.g., the first AI model (231a) of FIG. 2b) for generating the first text, included in a selected area based on a first input of the user.

[0193] According to one embodiment, when the electronic device (201) transmits the other party's name information as an input value of the first AI model (e.g., the first AI model (231a) of FIG. 2b), the other party's name information may be included in the first text and generated.

[0194] Referring to FIG. 4B, according to one embodiment, when an electronic device (201) (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) confirms a first input (b2) (e.g., a circle gesture) of a user while displaying image data (411), the electronic device may determine whether the image data (411) is stored in a memory (e.g., the memory (230) of FIG. 2A) of the electronic device, and upon confirming that the image data (411) is stored in the memory of the electronic device, determine whether a person is included in content data within a selected area based on the first input (b2) of the user.

[0195] According to one embodiment, when the electronic device confirms that the content data within the selected area includes a person (e.g., the user's face and the user's family faces) based on the user's first input (b2), the electronic device may transmit the person information (e.g., the user's name, the user's family name information, and / or the relationship information with the user) and additionally the location information and date information stored as metadata of the image data (411) as input values ​​of the first AI model (e.g., the first AI model (231a) of FIG. 2b) for generating the first text (topic script).

[0196] Referring to FIG. 4C, according to one embodiment, when an electronic device (201) (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) confirms a first input (b3) (e.g., a circle gesture) of a user while displaying image data (411), the electronic device may display a user interface (413) (e.g., a menu) for generating story content data for content data within a selected area based on the first input (b3) of the user and / or content properties selected by the user.

[0197] According to one embodiment, when the electronic device (201) confirms a selection of a user interface (413) (e.g., a menu) by a user, the electronic device (201) may transmit content data within the selected area based on the user's first input (b3) as an input value of a first AI model (e.g., the first AI model (231a) of FIG. 2b) for generating a first text (topic script).

[0198] FIGS. 5A and 5B are drawings for explaining a modification operation of a first text in an electronic device according to one embodiment.

[0199] Referring to FIG. 5A, according to one embodiment, an electronic device (201) (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may generate a first text (511) (e.g., a topic script) through a first AI model (e.g., the first AI model (231a) of FIG. 2B) and then provide a user interface that allows editing of the first text (511) while displaying the first text (511) through a display of the electronic device (e.g., the display (260) of FIG. 2A).

[0200] Referring to FIG. 5B, according to an embodiment, when an electronic device (201) (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A-2B) confirms selection of a user interface for editing a first text (511) by a user, the electronic device (201) may display a keyboard (515) to enable the user to edit the first text. Upon confirming completion of the modification of the first text (e.g., deleting the keywords “valley” and “Wangpicheon” and adding the keyword “travel with family”) by the user, the electronic device may display the modified first text (513) through a display of the electronic device (e.g., the display (260) of FIG. 2A).

[0201] FIGS. 6A, 6B, 6C, 6D, and 6E are drawings for explaining the operation of generating story content data in an electronic device according to one embodiment. FIGS. 6B, 6C, 6D, and 6E are drawings that separately explain the operation flowchart of FIG. 6A, and the order included in the drawings may be changed or some operations may be changed or omitted when performed.

[0202] According to one embodiment, with reference to FIGS. 6A and 6B, according to one embodiment, when an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A and 2B) confirms a first input (c1) (e.g., a circle gesture) of a user while displaying image data (601a) (e.g., a blog) by executing an application (e.g., a browser application) in operation 601, the electronic device may transmit a prompt requesting generation of content data (e.g., text information and image data information) and / or first text within a selected area based on the first input (b2) of the user to a first AI model (e.g., the first AI model (231a) of FIG. 2B).

[0203] According to one embodiment, the first AI model (e.g., the first AI model (231a) of FIG. 2B) may, in operation 603, generate a first text including at least one keyword based on content data (e.g., text information and image data information) input as an input value.

[0204] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 605, confirm the generation of a first text (e.g., “Wangpicheon Campground with a valley and camping tents under double cherry blossoms”) generated through a first AI model (e.g., first AI model (231a) of FIG. 2B).

[0205] According to one embodiment, an electronic device (e.g., an electronic device (101) of FIG. 1 and / or an electronic device (201) of FIGS. 2A to 2B), in operation 607, a processor (220) according to one embodiment may perform an operation of analyzing content data selected based on a first input of a user to identify and separate a main object (607a) from at least one object included in the content data.

[0206] According to an embodiment, referring to FIGS. 6A and 6C , according to an embodiment, when an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B ) confirms the creation of a first text (e.g., "Wangpicheon Campground with a Valley and Camping Tents Under Double Cherry Blossoms") in operation 605, in operation 609, provides a user interface for confirming and / or modifying the first text, confirms completion of modification of the first text (e.g., deleting keywords "valley" and "Wangpicheon", adding keywords "travel with family") by the user through the user interface, and confirms the modified first text (e.g., "travel with family at a campground with a camping tent under double cherry blossoms") in operation 611.

[0207] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A-2B) may, upon confirming the creation of a first text (e.g., “Wangpicheon Campground with a valley and camping tents under double cherry blossoms”) in operation 605, perform an operation of analyzing and / or processing the first text in operation 611.

[0208] In one embodiment, the electronic device may perform an operation of analyzing and / or processing the first text while identifying and modifying and / or deleting at least one keyword (e.g., a specific place name such as “Wangpicheon”) that represents a narrow range of keywords included in the first text.

[0209] In one embodiment, the electronic device may perform an operation of analyzing and / or processing the first text while changing at least one keyword included in the first text to a second keyword (e.g., campsite) that represents a broader range within a similar (same) category as the first keyword (e.g., Wangpicheon Campsite).

[0210] In one embodiment, the electronic device may perform an operation of analyzing and / or processing the first text while selecting various additional keywords within a similar (identical) category to at least one keyword (e.g., campground) included in the first text and providing them as candidate keywords (e.g., campsite, and / or glamping site, etc.).

[0211] According to one embodiment, if the electronic device includes information on a place (e.g., Jeju Island Campground) visited by the user in the user DB of the memory of the electronic device (e.g., memory (230) of FIG. 2A), the electronic device may perform an operation of analyzing and / or processing the first text while extracting third keyword information (e.g., Jeju Island Campground) included in a similar (same) category as the first keyword (e.g., Wangpicheon Campground) included in the first text from the place information (e.g., Jeju Island Campground) and changing the first keyword (e.g., Wangpicheon Campground) to the third keyword information (e.g., Jeju Island Campground).

[0212] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may, in operation 613, extract at least one first image data associated with a first text from a user DB of a memory of the electronic device (e.g., the memory (230) of FIG. 2A) or a user DB of a cloud.

[0213] According to one embodiment, the electronic device can extract at least one first image data from a memory of the electronic device based on at least one keyword included in the first text.

[0214] According to one embodiment, the electronic device may extract at least one first image data associated with a key object (607) identified based on analysis of content data from a user DB of a memory of the electronic device (e.g., memory (230) of FIG. 2A) or a user DB of a cloud.

[0215] According to one embodiment, referring to FIGS. 6A and 6D , according to one embodiment, the electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B ) may, in operation 615, identify at least some keywords (e.g., camping tent "under" double cherry blossoms) from among at least one keyword included in the first text, for which image data cannot be extracted from the user DB of the memory of the electronic device (e.g., the memory (230) of FIG. 2A ).

[0216] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A-2B) may, in operation 617, generate a prompt requesting generation of third image data associated with at least some keywords, and transmit the prompt requesting generation of the first text, at least some keywords of the first text, at least one first image image, and / or third image data associated with at least some keywords to a second AI model (e.g., second AI model (231b) of FIG. 2B).

[0217] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 619, acquire at least one second image data associated with at least some keywords through crawling or Retrieval-Augmented Generation (RAG) from outside the electronic device, and transmit the acquired at least one second image data to a second AI model.

[0218] According to one embodiment, the AI ​​model (e.g., the second AI model (231b) of FIG. 2b) may, in operation 621, generate at least one third image data associated with at least some keywords (e.g., camping tent “under” double cherry blossoms) for which image data cannot be extracted from a user DB of a memory of the electronic device (e.g., memory (230) of FIG. 2a) based on a prompt requesting generation of third image data associated with at least some keywords of the first text, at least some keywords of the first text, at least one first image image, at least one second image image, and / or at least some keywords. For example, an AI model (e.g., a second AI model (231b) of FIG. 2b) can generate at least one third image data (633) associated with at least some keywords (e.g., camping tent "under" double cherry blossoms) by using image data (631a) including "double cherry blossoms" and image data (631b) including "camping tent" from at least one of the first image data or the second image data.

[0219] According to one embodiment, referring to FIGS. 6A and 6E, according to one embodiment, an electronic device (e.g., an electronic device (101) of FIG. 1 and / or an electronic device (201) of FIGS. 2A to 2B) may transmit a prompt requesting generation of at least one first image data (613a) extracted from a user DB of a memory of the electronic device (e.g., a memory (230) of FIG. 2A), at least one second image data (635) acquired from the outside of the electronic device through crawling or RAG (Retrieval-Augmented Generation), or at least one third image data (633) generated through a second AI model (e.g., a second AI model (231b) of FIG. 2B), user information, and / or second text (e.g., a description) for the at least one image data, to a third AI model (e.g., a third AI model (231b) of FIG. 2B).

[0220] According to one embodiment, the third AI model (e.g., the third AI model (231b) of FIG. 2b) may, in operation 623, generate personalized second text for at least one piece of image data using information of the user (e.g., name, age and / or personality information based on face, preferred places and / or objects of the user, family relationships, etc.).

[0221] According to one embodiment, the third AI model (e.g., the third AI model (231c) of FIG. 2b) can generate one second text (e.g., a description in sentence form or a tag form) or each second text (e.g., a description in sentence form or a tag form) for at least one image data.

[0222] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate first result data (627) (e.g., story content data) based on at least one first image data, at least one second image data, at least one third image data, and / or at least a portion of second text in operation 625, and the result data (627) (e.g., story content data) may include at least one of image data, text data, and audio data. The first result data (627) may be generated using a fourth AI model.

[0223] According to one embodiment, the processor (220) or the fourth AI model (e.g., the fourth AI model (231d) of FIG. 2b) may generate first result data in which at least one image data included in the image data list and a second text are arranged based on the order of time.

[0224] According to one embodiment, the processor (220) or the fourth AI model (e.g., the fourth AI model (231d) of FIG. 2b) may generate first result data capable of changing the arrangement of at least one image data and second text included in the image data list in a manner responsive to a user interaction, regardless of the order of time or storage.

[0225] According to one embodiment, the processor (220) or the fourth AI model (e.g., the fourth AI model (231d) of FIG. 2b) may group image data having similarity in an image data list including at least one of at least one first image data, at least one second image data, and at least one third image data, and place one second text related to the grouped image data.

[0226] FIG. 7 is a diagram for explaining an operation of generating story content data in an electronic device according to an embodiment. FIG. 7 can explain an operation of confirming content data for generating story content data based on a user's first input for generating story content data, 3D modeling at least one image data related to the content data, and then generating interactive story content data that can be played back according to a form factor (e.g., VR environment, mobile device, and / or PC environment, etc.).

[0227] Referring to FIG. 7, according to an embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B), in operation 701, when confirming a first input (e.g., a circle gesture) of a user while displaying image data (701a) (e.g., a storybook cover) through application execution, analyzes content data (e.g., text information (rabbit and turtle) and image data information) within a selected area based on the first input of the user, and if it is determined that interactive story content data can be generated based on the analysis result, transmits a prompt requesting generation of a first text corresponding to the content data (e.g., text information (rabbit and turtle) and image data information) and a story script to a first AI model (e.g., the first AI model (231a) of FIG. 2B).

[0228] According to one embodiment, the first AI model (e.g., the first AI model (231a) of FIG. 2b) may generate the first text (e.g., a story about a rabbit and a turtle running a race, and then the rabbit resting and the turtle crawling to the end) based on a prompt requesting the generation of the first text corresponding to the content data (e.g., text information (rabbit and turtle) and image data information) and the story script through the first AI model (e.g., the first AI model (231a) of FIG. 2b) in operation 703.

[0229] According to one embodiment, the first AI model (e.g., the first AI model (231a) of FIG. 2b) can describe in detail the situation and relationships between characters, dialogue, etc. in the first text corresponding to the story script.

[0230] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 705, verify a first text generated by a first AI model (e.g., a story about a rabbit and a turtle running a race, with the rabbit resting and the turtle crawling to the end).

[0231] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may, in operation 707, extract at least one first image data (e.g., rabbit and turtle) associated with a first text from a user DB of a memory of the electronic device. According to one embodiment, when at least one first image data associated with the first text does not exist in the user DB of the memory of the electronic device, the electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may extract similar image data (e.g., image data of a rabbit or turtle dressed up in a school play).

[0232] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A-2B) may, in operation 709, determine whether there is an interactive 3D model associated with at least one first image data or similar image data.

[0233] According to one embodiment, the electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A-2B), in operation 711, if there is no interactive 3D model associated with at least one first image data or similar image data, may generate at least one first image data or similar image data into interactive 3D models (e.g., a 3D person dressed up as a rabbit, a 3D person dressed up as a turtle, etc.) through a fifth AI model (e.g., an on-device AI model (231) of FIG. 2B or an external AI model (281) of FIG. 2B).

[0234] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 713, generate second text describing each scene using a 3D model through a third AI model (e.g., third AI model (231c) of FIG. 2B).

[0235] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 715, generate first result data (717) (e.g., story content data composed of a 3D model) based on each scene and second text using a 3D model through a fourth AI model (e.g., fourth AI model (231d) of FIG. 2B).

[0236] According to one embodiment, an electronic device (e.g., an electronic device (101) of FIG. 1 and / or an electronic device (201) of FIGS. 2A to 2B) may analyze, using a fourth AI model (e.g., a fourth AI model (231d) of FIG. 2B), components existing in the environment or background (e.g., grass, large trees, etc.) or sounds that may occur when interacting with a user (e.g., a rabbit jumping sound and / or a turtle walking sound), from content data stored in the electronic device, and then generate some of the sounds by reflecting them in first result data (717) (e.g., story content data composed of a 3D model).

[0237] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may synthesize at least one image data taken at a place (e.g., a park) frequently visited by a user through a fourth AI model (e.g., the fourth AI model (231d) of FIG. 2B) to generate a background model (e.g., a 3D model of the park) by reflecting the synthesized data into first result data (717) (e.g., story content data composed of a 3D model).

[0238] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A-2B) may generate a 3D model capable of interacting as required in a target form factor (e.g., a VR environment) or extract the same from a memory of the electronic device, and then input the first text (e.g., a story script) into an engine for generating story content data capable of interacting based on the first text and executing the first text.

[0239] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate story content data by taking into account various directing elements such as a movement method of an interactive 3D model and / or an ending branch according to the flow of the story, taking into account an output method of the story content data and / or a form factor characteristic of the electronic device.

[0240] FIG. 8 is a diagram for explaining an operation of generating story content data in an electronic device according to one embodiment.

[0241] According to one embodiment, when an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) confirms selection of conversation content with a counterpart through a first input of a user while executing a first application (e.g., a message application) in operation 801, the electronic device may transmit a prompt requesting generation of a first text based on conversation content information (e.g., schedule of a Japanese travel plan, reservation information, etc.) included in a selected area based on the user's first input, image data (e.g., still image data and / or video data) received from the counterpart through the first application (e.g., a message application) selected by the user, and / or conversation content information (e.g., schedule of a Japanese travel plan, reservation information, etc.) and image data (e.g., still image data and / or video data) as an input value of a first AI model (e.g., the first AI model (231a) of FIG. 2B).

[0242] According to one embodiment, the first AI model (e.g., the first AI model (231a) of FIG. 2B) may, in operation 803, generate a first text including at least one keyword based on input conversation content information (e.g., schedule of a Japan travel plan, reservation information, etc.) and image data (e.g., still image data and / or video data).

[0243] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 805, confirm the generation of a first text (e.g., “Osaka cherry blossoms, restaurants, streets, etc. itinerary, experience, etc. at the second day’s lodging”) generated through a first AI model (e.g., first AI model (231a) of FIG. 2B).

[0244] According to one embodiment, the electronic device may verify the existence of previously generated second result data (831) (e.g., second story content data) based on the first text, and determine that supplementation of the second result data (831) (e.g., second story content data) is necessary based on conversation content information (e.g., schedule of a Japan travel plan, reservation information, etc.) and image data (e.g., still image data and / or video data).

[0245] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may, in operation 807, extract at least one first image data associated with a first text from a user DB of a memory of the electronic device (e.g., the memory (230) of FIG. 2A) or a user DB of a cloud.

[0246] According to one embodiment, the electronic device can extract at least one first image data from a memory of the electronic device based on at least one keyword included in the first text.

[0247] According to one embodiment, the electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may, in operation 811, identify at least some keywords (e.g., “ramen ordered by a friend,” “Osaka Castle taken from various angles”) from among at least one keyword included in the first text, for which image data cannot be extracted from the user DB of the memory (e.g., “memory (230) of FIG. 2A”) of the electronic device.

[0248] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A-2B) may, in operation 813, generate a prompt requesting generation of third image data associated with at least some keywords, and transmit the prompt requesting generation of the first text, at least some keywords of the first text, at least one first image image, and third image data associated with at least some keywords to a second AI model (e.g., second AI model (231b) of FIG. 2B).

[0249] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may, in operation 815, acquire at least one second image data associated with at least some keywords through crawling or Retrieval-Augmented Generation (RAG) from outside the electronic device, and transmit the acquired at least one second image data to a second AI model.

[0250] According to one embodiment, the AI ​​model (e.g., the second AI model (231b) of FIG. 2b) may, in operation 817, generate at least one third image data associated with at least some keywords (e.g., “ramen ordered by a friend,” “Osaka Castle taken from various angles”) for which image data cannot be extracted from a user DB of a memory of the electronic device (e.g., the memory (230) of FIG. 2a) based on a prompt requesting generation of third image data associated with the first text, at least some keywords of the first text, at least one first image image, at least one second image image, and / or at least some keywords.

[0251] According to one embodiment, an electronic device (e.g., an electronic device (101) of FIG. 1 and / or an electronic device (201) of FIGS. 2A to 2B) may transmit a prompt requesting generation of at least one first image data extracted from a user DB of a memory of the electronic device (e.g., a memory (230) of FIG. 2A), at least one second image data acquired from the outside of the electronic device through crawling or Retrieval-Augmented Generation (RAG), or at least one third image data generated through a second AI model, and / or a second text (e.g., a description) for the at least one image data to a third AI model (e.g., a third AI model (231c) of FIG. 2B).

[0252] According to one embodiment, the third AI model (e.g., the third AI model (231c) of FIG. 2b) may, in operation 819, generate one second text (e.g., a description in sentence form or a tag form) or each second text (e.g., a description in sentence form or a tag form) for at least one image data.

[0253] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate new first result data (833) (e.g., first story content data) through a fourth AI model (e.g., fourth AI model (231d) of FIG. 2B) based on at least one first image data, at least one second image data, at least one third image data, and second text in operation 819, and the new first result data (833) (e.g., first story content data) may be output to include at least one of image data, text data, and audio data.

[0254] According to one embodiment, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may compare previously generated second result data (831) (e.g., second story content data) with newly generated first result data (833) (e.g., first story content data), and if the first result data (833) (e.g., first story content data) includes more content data (e.g., text data and video data) than the first result data (831) (e.g., first story content data), update the second result data (831) (e.g., second story content data) with the first result data (833) (e.g., first story content data) and store it.

[0255] According to one embodiment, an electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) can generate the first result data (story content data) in various forms, not limited to a format, such as a story in a journal format, a short-form format, a vlog format, etc., and can also update the story content data by generating and then replacing or inserting image data (e.g., still image data or video data) into the previously generated first result data (e.g., story content data) through a generative AI model.

[0256] An electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) according to an embodiment may include a display (e.g., display module (160) of FIG. 1 and / or display (260) of FIG. 2A), at least one processor (e.g., processor (120) of FIG. 1 and / or processor (220) of FIGS. 2A to 2B), and a memory (e.g., memory (130) of FIG. 1 and / or memory (230) of FIG. 2A) for storing instructions. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to display first content data through the display. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to, based on a first input for selecting at least a portion of the first content data to generate story content data, obtain first text data based on at least a portion of the first content data selected by the first input. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to, based on a first input for selecting at least a portion of the first content data to generate the story content data, obtain at least one first image data from the memory based on at least a portion of the first content data selected by the first input.In one embodiment, the instructions, when individually or collectively executed by the at least one processor, cause the electronic device to generate story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, and when the story content data including the first text data and the at least one first image data is displayed, the first text data may be displayed adjacent to the at least one first image data and used as a description of the at least one first image data.

[0257] In one embodiment, the commands, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to, if it fails to acquire at least a portion of the at least one first image data from the memory, acquire at least one second image data associated with at least a portion of the first content data from an external electronic device through crawling or Retrieval-Augmented Generation (RAG). In one embodiment, the commands, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate story content data including at least a portion of the first text, the at least one first image data, and the at least one second image data.

[0258] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate, through an AI model, at least one third image data associated with at least a portion of the first content data if at least a portion of the at least one first image data is not obtained from the memory. In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate story content data including at least a portion of the first text, the at least one first image data, and the at least one third image data.

[0259] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to generate a search term based on at least a portion of the first content data. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the at least one first image data based on at least a portion of the first content data from the memory using the search term, wherein the search term may include at least one keyword.

[0260] In one embodiment, the commands, when individually or collectively executed by the at least one processor, may cause the electronic device to compare the at least one keyword included in a search term generated based on at least a portion of the first content data with tag information of image data stored in the memory. In one embodiment, the commands, when individually or collectively executed by the at least one processor, may cause the electronic device to determine the search term as an invalid search term if the at least one keyword is not included in tag information of the image data stored in the memory. In one embodiment, the commands, when individually or collectively executed by the at least one processor, may cause the electronic device to display a message indicating that generation of story content data is impossible or a UI for modifying the search term if the search term is determined to be an invalid first text.

[0261] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to convert the first text describing the at least one first image data into voice data. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to convert the first text describing the at least one first image data into voice data.

[0262] An electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) according to an embodiment may include a display (e.g., display module (160) of FIG. 1 and / or display (260) of FIG. 2A), at least one processor (e.g., processor (120) of FIG. 1 and / or processor (220) of FIGS. 2A to 2B), and a memory (e.g., memory (130) of FIG. 1 and / or memory (230) of FIG. 2A) for storing instructions. The instructions according to an embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate a first text based on content data included in a first input. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to extract from the memory at least one first image data associated with the first text. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate a second text through a first AI model (e.g., the second AI model (231b) of FIG. 2b) based on the at least one first image data. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate and output first result data based on the at least one first image data and the second text, wherein the first result data may include at least one of image data, text data, and audio data.

[0263] According to one embodiment, the first input includes image data input by a user, and according to one embodiment, the commands, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate the first text through a second AI model (e.g., the first AI model (231a) of FIG. 2b) based on content data included in the image data input by the user.

[0264] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to output the first text through the display or an audio output device included in the electronic device. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to, after outputting the first text, output a first text modified based on a second input through the display or an audio output device included in the electronic device. The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to extract from the memory at least one first image data associated with the modified first text.

[0265] The commands according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to identify at least one keyword included in the first text. The commands according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to change the first keyword to the second keyword when identifying a second keyword included in the same category as the first keyword among the at least one keyword in the location information stored in the memory.

[0266] In one embodiment, the commands, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to acquire at least one second image data associated with the first text from an external source of the electronic device through crawling or Retrieval-Augmented Generation (RAG). In one embodiment, the commands, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to generate the second text through the first AI model based on the at least one first image data and the at least one second image data. In one embodiment, the commands, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to output the first result data based on the at least one first image data, the at least one second image data, and the second text.

[0267] In one embodiment, the first text includes at least one keyword, and the instructions in one embodiment, when individually or collectively executed by the at least one processor, may be configured to cause the electronic device to, when the electronic device cannot extract image data associated with at least some of the at least one keyword from the memory, acquire at least one second image data associated with the at least some of the keywords through crawling or Retrieval-Augmented Generation (RAG) from outside the electronic device.

[0268] At least a portion of the first result data according to one embodiment comprises a story in chronological order, and the instructions according to one embodiment, when individually or collectively executed by the at least one processor, cause the electronic device to adjust the strength of the fiction of the data, wherein the strength of the fiction can be applied to at least one of image data, text data and audio data included in the first result data.

[0269] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, cause the electronic device to delete or emphasize at least one object included in the first result data according to a fourth input received from a user, and the deletion or emphasis of the at least one object may be applied to at least one of image data, text data and audio data included in the first result data.

[0270] The instructions according to one embodiment, when individually or collectively executed by the at least one processor, may cause the electronic device to combine second result data generated by an external device with the first result data to generate third result data.

[0271] The second result data according to one embodiment includes image data captured by the external device, and the commands according to one embodiment, when individually or collectively executed by the at least one processor, can cause the electronic device to receive, from the external device, at least one of image data captured by the external device, metadata of the image data captured by the external device, and analysis data obtained by analyzing the image data captured by the external device.

[0272] FIG. 9 is a flowchart illustrating operations for generating story content data in an electronic device according to an embodiment. The operations for generating story content data may include operations 901 to 909. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, at least two operations may be performed in parallel, or other operations may be added.

[0273] In operation 901, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) can confirm at least one content data selected by a first input of a user.

[0274] In one embodiment, an electronic device may identify a gesture for determining an area (e.g., a circle gesture and / or a screen capture) as a first user input for generating story content data.

[0275] The first input according to one embodiment may include various user inputs, such as gestures, and / or multimodal inputs.

[0276] An electronic device according to one embodiment may identify an action (e.g., screen capture) that allows selection of the entire image data as a first user input for generating story content data.

[0277] In one embodiment, an electronic device may display a user interface (e.g., a menu) for generating story content data upon confirming selection of a portion of content data from at least one content data based on a first input of a user.

[0278] In operation 903, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate a first text based on at least one piece of content data.

[0279] According to one embodiment, when an electronic device confirms a selection of content data included in image data based on a first input of a user, the electronic device may transmit a prompt requesting generation of content data and first text (e.g., topic script, search word, and search sentence) to a first AI model (e.g., the first AI model (231a) of FIG. 2b).

[0280] An electronic device according to one embodiment can transmit at least one content data selected based on a first input of a user to a first AI model (231a).

[0281] According to one embodiment, an electronic device may process content data selected based on a user's first input into a form that can be transmitted to a first AI model (e.g., preprocessing such as normalization into a form suitable for model input, and / or LangChain form, etc.) and then transmit the processed content data to the first AI model.

[0282] An electronic device according to one embodiment can check content data to be transmitted to the first AI model from a content DB of a memory (e.g., memory (230) of FIG. 2A) in which content data processed in a form that can be transmitted to the first AI model is stored.

[0283] According to one embodiment, an electronic device may process content data into a form that can be transmitted to a first AI model (e.g., preprocessing such as normalization into a form suitable for model input, and / or LangChain form, etc.) using a data conversion model included in the electronic device or server, and then transmit the processed content data to the first AI model.

[0284] In one embodiment, the electronic device may include, in a prompt transmitted to the first AI model, a request for determining a main keyword from the input content data and generating a concise first text based on the determined main keyword.

[0285] An electronic device according to one embodiment may include, in a prompt transmitted to a first AI model, a request for determining a main object among input content data and converting the main object into a form usable by a search engine (e.g., vector form, language description form of the content, and / or main object tag information of the content, etc.) to facilitate a search for image data based on the first text.

[0286] An electronic device according to one embodiment may transmit at least one content data selected by a first input of a user to a first AI model, application information providing the at least one content data, and information (e.g., name information, relationship information, location information, date information, and / or place name information) of text data included in the at least one content data.

[0287] An electronic device according to one embodiment may, upon confirming the creation of a first text, output the first text and output a modified first text through a second input from the user.

[0288] An electronic device according to one embodiment may output a first text through a display (e.g., a display (260) of FIG. 2A) or an audio output device (not shown) included in the electronic device, and provide a user interface for modifying the first text.

[0289] An electronic device according to one embodiment may output the modified first text through a display or an audio output device (not shown) included in the electronic device upon confirming modification of the first text based on a second input of the user through a user interface for modification of the first text.

[0290] An electronic device according to one embodiment may perform an operation of analyzing and / or processing the first text upon confirming the creation of the first text.

[0291] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text while identifying and deleting at least one keyword (e.g., a specific place name) that represents a narrow range among keywords included in the first text.

[0292] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text while changing at least one keyword included in the first text to a second keyword that represents a broader range within a similar category as the first keyword.

[0293] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text while selecting various additional keywords within a similar category to at least one keyword included in the first text and providing the selected keywords as candidate keywords.

[0294] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text using user information.

[0295] An electronic device according to one embodiment may, when information on places visited by a user is included in a user DB of a memory (e.g., memory (230) of FIG. 2A), perform an operation of extracting third keyword information similar to a first keyword included in a first text from the place information and changing the first keyword into the third keyword information, and analyzing and / or processing the first text.

[0296] In operation 905, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may extract at least one first image data associated with a first text.

[0297] According to one embodiment, when the electronic device confirms the creation of the first text, it can extract at least one first image data associated with the first text from a memory (e.g., memory (230) of FIG. 2A) or a user DB in the cloud.

[0298] An electronic device according to one embodiment can search and extract at least one first image data associated with at least one keyword included in a first text from a user DB of a memory or a cloud.

[0299] An electronic device according to one embodiment can search and extract at least one first image data associated with at least one keyword included in a first text by using a vector DB of a memory in which image data is stored in a vector form or a text DB of a memory in which image data is stored in a text form.

[0300] According to one embodiment, an electronic device analyzes at least one content data selected based on a first input of a user to identify and separate a main object from at least one object included in the content data, and searches for and extracts at least one first image data associated with the main object from a user DB in a memory or a cloud.

[0301] In one embodiment, an electronic device may generate at least one first image data associated with a primary object through a second AI model (e.g., the second AI model (231b) of FIG. 2b) when it is unable to extract at least one first image data associated with a primary object from a user DB of a memory or a cloud.

[0302] In operation 907, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate a second text based on at least one first image data and user information.

[0303] An electronic device according to one embodiment can generate a second text through a third AI model (e.g., the third AI model (231c) of FIG. 2b) based on at least one first image data and user information.

[0304] An electronic device according to one embodiment may transmit a prompt to a third AI model requesting generation of a second text for at least one image based on at least one first image data and information of a user.

[0305] A third AI model according to one embodiment may generate personalized second text for at least one piece of image data using information about the user (e.g., name, age and / or personality information based on face, preferred locations and / or objects of the user, family relationships, etc.).

[0306] A third AI model according to one embodiment can generate one second text (e.g., a description in sentence form or a tag form) or each second text (e.g., a description in sentence form or a tag form) for at least one image data.

[0307] In operation 909, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate first result data based on at least one first image data and a second text.

[0308] An electronic device according to one embodiment may generate first result data (e.g., story content data) through a fourth AI model (e.g., the fourth AI model (231d)) based on at least one first image data and at least a part of a second text.

[0309] An electronic device according to one embodiment may output first result data including at least one of image data, text data, and audio data generated through a fourth AI model.

[0310] An electronic device according to one embodiment may generate personalized first result data (e.g., story content data) in a journal style using at least one first image data and at least a portion of a second text for the at least one first image data through a fourth AI model.

[0311] An electronic device according to one embodiment may generate a title of first result data (e.g., story content data) based on at least some of a first text, at least one first image data, and / or user information.

[0312] According to one embodiment, the processor (220) may use the first text as a title of the first result data (e.g., story content data).

[0313] A fourth AI model according to one embodiment can generate first result data in which at least one first image data and a second text are arranged based on a time sequence.

[0314] A fourth AI model according to one embodiment can generate first result data that can change the arrangement of at least one first image data and second text in a manner responsive to a user interaction, regardless of the time order or storage order.

[0315] A fourth AI model (231d) according to one embodiment can group first image data having similarity in at least one first image data and place one second text related to the grouped first image data.

[0316] FIG. 10 is a flowchart illustrating operations for generating story content data in an electronic device according to an embodiment. The operations for generating story content data may include operations 1001 to 1015. In the following embodiments, the operations may be performed sequentially, but are not necessarily performed sequentially. For example, the order of the operations may be changed, at least two operations may be performed in parallel, or other operations may be added.

[0317] In operation 1001, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) can verify at least one content data selected by a first input of a user.

[0318] An electronic device according to one embodiment may identify a gesture for determining an area (e.g., a circle gesture) as a first user input for generating story content data.

[0319] An electronic device according to one embodiment may identify an action (e.g., screen capture) that allows selection of the entire image data as a first user input for generating story content data.

[0320] In one embodiment, the electronic device may display a user interface (e.g., a menu) for generating story content data when the electronic device confirms selection of a portion of content data from at least one content data based on a first input of the user.

[0321] In operation 1003, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate a first text based on content data.

[0322] According to one embodiment, when an electronic device confirms a selection of content data included in image data based on a first input of a user, the electronic device may transmit a prompt requesting generation of content data and first text (e.g., topic script, search word, or search sentence) to a first AI model (e.g., the first AI model (231a) of FIG. 2b).

[0323] An electronic device according to one embodiment may, upon confirming the creation of a first text, output the first text and output a modified first text through a second input from the user.

[0324] An electronic device according to one embodiment may output a first text through a display (e.g., a display (260) of FIG. 2A) or an audio output device (not shown) included in the electronic device, and provide a user interface for modifying the first text.

[0325] An electronic device according to one embodiment may output the modified first text through a display or an audio output device (not shown) included in the electronic device upon confirming modification of the first text based on a second input of the user through a user interface for modification of the first text.

[0326] An electronic device according to one embodiment may perform an operation of analyzing and / or processing the first text upon confirming the creation of the first text.

[0327] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text while identifying and deleting at least one keyword (e.g., a specific place name) that represents a narrow range among keywords included in the first text.

[0328] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text while changing at least one keyword included in the first text to a second keyword that represents a broader range within a similar category as the first keyword.

[0329] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text while selecting various additional keywords within a similar (identical) category to at least one keyword included in the first text and providing the keywords as candidate keywords.

[0330] An electronic device according to one embodiment may perform an operation of analyzing and / or processing a first text using user information.

[0331] In one embodiment, an electronic device may, when information on places visited by a user is included in a user DB of a memory (e.g., memory (230) of FIG. 2A), perform an operation of extracting third keyword information similar to (identical to) a first keyword included in a first text from the place information and analyzing and / or processing the first text while changing the first keyword to the third keyword information.

[0332] In operation 1005, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may extract at least one first image data associated with a first text.

[0333] According to one embodiment, when the electronic device confirms the creation of the first text, it can extract at least one first image data associated with the first text from a memory (e.g., memory (230) of FIG. 2A) or a user DB in the cloud.

[0334] An electronic device according to one embodiment can search and extract at least one first image data associated with at least one keyword included in a first text from a user DB of a memory or a cloud.

[0335] An electronic device according to one embodiment can search and extract at least one first image data associated with at least one keyword included in a first text by using a vector DB of a memory in which image data is stored in a vector form or a text DB of a memory in which image data is stored in a text form.

[0336] According to one embodiment, an electronic device analyzes content data selected based on a first input of a user to identify and separate a main object from at least one object included in the content data, and searches for and extracts at least one first image data associated with the main object from a user DB in a memory or a cloud.

[0337] In operation 1007, the electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may check whether at least some keywords, among at least one keyword included in the first text, exist for which associated image data has not been extracted from a user DB of a memory or a cloud. In operation 1007, if the electronic device checks whether at least some keywords, among at least one keyword included in the first text, exist for which associated image data has not been extracted from a user DB of a memory or a cloud, the electronic device may acquire at least one second image data from outside the electronic device and generate at least one third image data through an AI model in operation 1009.

[0338] In one embodiment, when an electronic device cannot extract image data associated with at least some keywords among at least one keyword (search word) included in a first text from a user DB of a memory or a cloud, the electronic device can obtain at least one second image data associated with at least some keywords for which the associated image data was not extracted from the user DB of the memory or the cloud through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device.

[0339] In one embodiment, when an electronic device cannot extract image data associated with at least some keywords among at least one keyword (search word) included in a first text from a user DB of a memory or a cloud, the electronic device can obtain at least one second image data associated with at least some keywords for which the associated image data was not extracted from the user DB of the memory or the cloud through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device based on location information and date and time information included in metadata of at least one first image data extracted from the user DB of the memory or the cloud.

[0340] According to one embodiment, an electronic device may, upon acquiring at least one second image data acquired through crawling or RAG (Retrieval-Augmented Generation), determine whether to use the at least one second image data as image data of story content data associated with content data based on a correlation between the at least one first image data and the at least one second image data.

[0341] In one embodiment, when an electronic device cannot extract image data associated with at least some keywords among at least one keyword included in a first text from a DB of a memory or a cloud, the electronic device can generate at least one third image data associated with at least some keywords for which the associated image data was not extracted from the user DB of the memory or the cloud through a second AI model (e.g., the second AI model (231b) of FIG. 2b).

[0342] In one embodiment, when an electronic device cannot extract image data associated with at least some keywords among at least one keyword included in a first text from a user DB of a memory or a cloud, the electronic device may transmit a prompt to a second AI model requesting generation of the first text, at least some keywords of the first text, at least one first image data, at least one second image data, and at least one third image data associated with at least some keywords.

[0343] An electronic device according to one embodiment may generate a description including specific text information about content data selected based on a first input of a user based on a first text, and transmit the description to the second AI model, such that the second AI model generates at least one third image data similar to at least one first image data.

[0344] An electronic device according to one embodiment may, when generating at least one third image data through a second AI model, determine whether to use at least one third image data generated through the second AI model by using a similarity between a first text and the at least one third image data or a captioning function that describes the at least one third image data.

[0345] An electronic device according to one embodiment may determine to use at least one third image data when the similarity of at least one third image data generated through the second AI model is greater than or equal to a criterion (threshold value) for evaluation.

[0346] An electronic device according to one embodiment may request re-generation of at least one third image data using the second AI model when the similarity of at least one third image data generated through the second AI model is below a criterion (threshold value) for evaluation.

[0347] According to one embodiment, a second AI model may generate at least one third image data associated with at least some of the keywords by using at least some of the first text, at least some of the keywords of the first text, at least one first image data, and at least one second image data.

[0348] In operation 1011, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate a second text based on at least one first image data, at least one second image data, at least one third image data, and user information.

[0349] An electronic device according to one embodiment may generate a second text through a third AI model (e.g., the third AI model (231c) of FIG. 2b) based on at least one first image data, at least one second image data, at least one third image data, and user information.

[0350] An electronic device according to one embodiment may transmit a prompt to a third AI model requesting generation of a second text for at least one image data based on at least one first image data, at least one second image data, at least one third image data, and at least some of the user's information.

[0351] A third AI model according to one embodiment may generate personalized second text for at least one piece of image data using information about the user (e.g., name, age and / or personality information based on face, preferred locations and / or objects of the user, family relationships, etc.).

[0352] A third AI model according to one embodiment can generate one second text (e.g., a description in sentence form or a tag form) or each second text (e.g., a description in sentence form or a tag form) for at least one image data.

[0353] In operation 1013, an electronic device (e.g., electronic device (101) of FIG. 1 and / or electronic device (201) of FIGS. 2A to 2B) may generate first result data based on at least one first image data, at least one second image data, at least one third image data, and second text.

[0354] An electronic device according to one embodiment may generate first result data (e.g., story content data) through a fourth AI model (e.g., the fourth AI model (231d)) based on at least one first image data, at least one second image data, at least one third image data, and at least a portion of a second text.

[0355] An electronic device according to one embodiment may output first result data including at least one of image data, text data, and audio data generated through a fourth AI model.

[0356] An electronic device according to one embodiment may generate personalized first result data (e.g., story content data) in a journal style through a fourth AI model based on a list of image data including at least one of at least one first image data, at least one second image data, and at least one third image data, and a second text for each image data. An electronic device according to one embodiment may generate a title of the first result data (e.g., story content data) based on the first text, the at least one first image data, and / or user information.

[0357] According to one embodiment, the processor (220) may use the first text as a title of the first result data (e.g., story content data).

[0358] A fourth AI model according to one embodiment can generate first result data in which at least one first image data, at least one second image data, at least one third image data, and a second text are arranged based on a time sequence.

[0359] A fourth AI model according to one embodiment can generate first result data that can change the arrangement of at least one first image data, at least one second image data, at least one third image data, and at least some of the second text in a manner responsive to a user interaction, regardless of the order of time or storage.

[0360] A fourth AI model according to one embodiment may group image data having similarities among at least one first image data, at least one second image data, and at least one third image data, and place one second text related to the grouped image data.

[0361] In operation 1007, if the electronic device does not confirm the existence of at least some keywords among at least one keyword included in the first text and for which associated image data is not extracted from the user DB of the memory or cloud, in operation 1015, the electronic device (e.g., the electronic device (101) of FIG. 1 and / or the electronic device (201) of FIGS. 2A to 2B) may generate first result data based on at least one first image data and at least some of the second text.

[0362] According to an embodiment, a method for generating story content data using an artificial intelligence model in an electronic device (e.g., the electronic device 101 of FIG. 1 and / or the electronic device 201 of FIGS. 2A to 2B) may include an operation of displaying first content data through a display of the electronic device. According to an embodiment, the method may include an operation of obtaining a first text based on at least a portion of the first content data selected by the first input, based on confirming a first input for selecting at least a portion of the first content data to generate the story content data. According to an embodiment, the method may include an operation of obtaining at least one first image data from a memory of the electronic device based on at least a portion of the first content data selected by the first input, based on confirming the first input for selecting at least a portion of the first content data to generate the story content data. According to one embodiment, the method includes an operation of generating story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, and when the story content data including the first text and the at least one first image data is displayed, the first text is displayed adjacent to the at least one first image data and can be used as a description of the at least one first image data.

[0363] According to one embodiment, the method may include an operation of acquiring at least one second image data associated with at least one part of the first content data from an external electronic device through crawling or Retrieval-Augmented Generation (RAG), if at least a part of the at least one first image data is not acquired from the memory of the electronic device. According to one embodiment, the method may further include an operation of generating story content data including at least a part of the first text, the at least one first image data, and the at least one second image data.

[0364] According to one embodiment, the method may include an operation of generating at least one third image data associated with at least a portion of the first content data through an AI model if at least a portion of the at least one first image data is not obtained from the memory of the electronic device. According to one embodiment, the method may further include an operation of generating story content data including at least a portion of the first text, the at least one first image data, and the at least one third image data.

[0365] The method according to one embodiment further includes an operation of obtaining at least one first image data based on at least a portion of the first content data from the memory using the search word, wherein the search word may include at least one keyword.

[0366] Electronic devices according to embodiments disclosed herein may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to embodiments disclosed herein are not limited to the aforementioned devices.

[0367] The embodiments of this document and the terminology used herein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.

[0368] The term "module" used in one embodiment of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0369] An embodiment of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101) or an electronic device (301)). For example, a processor (e.g., a processor (520)) of the machine (e.g., an electronic device (301)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.

[0370] According to one embodiment, the method according to one embodiment disclosed in the present document may be provided as a computer program product. The computer program product may be traded between sellers and buyers as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or may be provided through an application store (e.g., Play Store). TM ) or directly between two user devices (e.g., smart phones), online distribution (e.g., downloading or uploading). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily created in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.

[0371] According to one embodiment, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to one embodiment, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to one embodiment, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device (101 in FIG. 1; 201 in FIG. 2a to FIG. 2b), Display (160 in Fig. 1; 260 in Fig. 2a); At least one processor (120 in FIG. 1; 220 in FIGS. 2A to 2B); and It includes a memory (130 in Fig. 1; 230 in Fig. 2a) that stores commands, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Through the above display, the first content data is displayed, Based on identifying a first input for selecting at least a portion of the first content data to generate story content data, obtaining a first text based on at least a portion of the first content data selected by the first input, Based on the first input for selecting at least a portion of the first content data to generate the story content data, obtaining at least one first image data from the memory based on at least a portion of the first content data selected by the first input, Generate story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, An electronic device in which, when the story content data including the first text and the at least one first image data is displayed, the first text is displayed adjacent to the at least one first image data and is used as a description of the at least one first image data.

2. In paragraph 1, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: If at least a portion of the at least one first image data is not obtained from the memory, at least one second image data associated with at least a portion of the first content data is obtained from an external electronic device through crawling or RAG (Retrieval-Augmented Generation), An electronic device configured to generate story content data including at least a portion of the first text, the at least one first image data, and the at least one second image data.

3. In paragraph 1 or 2, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: If at least a portion of the at least one first image data is not obtained from the memory, at least one third image data associated with at least a portion of the first content data is generated through the AI ​​model, An electronic device configured to generate story content data including at least a portion of the first text, the at least one first image data, and the at least one third image data.

4. In any one of paragraphs 1 to 3, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Generate a search term based on at least a portion of the first content data, Obtaining at least one first image data based on at least a part of the first content data from the memory using the search word, An electronic device wherein the search term includes at least one keyword.

5. In any one of paragraphs 1 to 4, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Comparing the at least one keyword included in the search term generated based on at least a portion of the first content data with the tag information of the image data stored in the memory, If at least one keyword is not included in the tag information of the image data stored in the memory, the search term is determined to be an invalid search term, An electronic device that displays a message indicating that generation of story content data is impossible or a UI for modifying the search term when the search term is determined to be an invalid first text.

6. In any one of paragraphs 1 to 5, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Converting the first text describing the at least one first image data into voice data, An electronic device wherein the above story content data includes at least one of the voice data and the first text.

7. A method for generating story content using an AI model in an electronic device, An action of displaying first content data through a display of the electronic device; An operation of obtaining a first text based on at least a portion of the first content data selected by the first input, based on identifying a first input for selecting at least a portion of the first content data to generate story content data; An operation of obtaining at least one first image data from a memory of the electronic device based on at least a portion of the first content data selected by the first input, based on confirming the first input for selecting at least a portion of the first content data to generate the story content data; and An operation of generating story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, A method in which, when the story content data including the first text and the at least one first image data is displayed, the first text is displayed adjacent to the at least one first image data and is used as a description of the at least one first image data.

8. In a non-volatile storage medium storing commands, the commands are configured to cause the electronic device to perform at least one operation when executed by the electronic device, wherein the at least one operation is: An action of displaying first content data through a display of the electronic device; An operation of obtaining first text data based on at least a portion of the first content data selected by the first input, based on identifying a first input for selecting at least a portion of the first content data to generate story content data; An operation of obtaining at least one first image data from a memory of the electronic device based on at least a portion of the first content data selected by the first input, based on confirming the first input for selecting at least a portion of the first content data to generate the story content data; and An operation of generating story content data including the first text and the at least one first image data obtained based on at least a portion of the first content data selected by the first input, A storage medium in which, when the story content data including the first text data and the at least one first image data is displayed, the first text data is displayed adjacent to the at least one first image data and is used as a description of the at least one first image data.

9. In the electronic device (101 in Fig. 1; 201 in Figs. 2a to 2b), Display (160 in Fig. 1; 260 in Fig. 2a); At least one processor (120 in FIG. 1; 220 in FIGS. 2A to 2B); and It includes a memory (130 in Fig. 1; 230 in Fig. 2a) that stores commands, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Generate a first text based on content data included in the first input, Extracting at least one first image data associated with the first text from the memory, Generating a second text through a first AI model (231b of FIG. 2b) based on at least one first image data, Generating and outputting first result data based on at least one first image data and the second text, An electronic device wherein the first result data includes at least one of image data, text data, and audio data.

10. In paragraph 9, The first input includes image data input by the user, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device configured to generate the first text through a second AI model (231a in FIG. 2b) based on content data included in image data input by the user.

11. In paragraph 9 or 10, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Outputting the above first text through the display or an audio output device included in the electronic device, After the first text is output, the first text modified based on the second input is output through the display or an audio output device included in the electronic device, An electronic device configured to extract at least one first image data associated with the modified first text from the memory.

12. In paragraph 9 or paragraph 11, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Check at least one keyword contained in the first text above, If a second keyword is identified that is included in the same category as the first keyword among the at least one keyword in the location information stored in the above memory, An electronic device configured to change the first keyword to the second keyword.

13. In any one of paragraphs 9 to 12, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Obtaining at least one second image data associated with the first text through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device, Generating the second text through the first AI model based on the at least one first image data and the at least one second image data, An electronic device configured to output the first result data based on the at least one first image data, the at least one second image data, and the second text.

14. In any one of paragraphs 9 to 13, The first text above contains at least one keyword, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: An electronic device configured to obtain at least one second image data associated with at least some of the keywords through crawling or RAG (Retrieval-Augmented Generation) from outside the electronic device when image data associated with at least some of the keywords cannot be extracted from the memory.

15. In any one of paragraphs 9 to 14, At least some of the above first result data comprises a story in chronological order, The above instructions, when individually or collectively executed by the at least one processor, cause the electronic device to: Adjust the strength of the fictitiousness of the first result data according to the third input received from the user, An electronic device in which the above fictional strength is applied to at least one of the image data, text data, and audio data included in the first result data.

Citation Information

Patent Citations

  • Picture book creation system, picture book creation program, and picture book creation method

    JP7462991B1

  • Prefabricated hoist

    KR1020210151563A

  • Method for generating story-board based on sound source

    KR102643498B1

  • Method and apparatus for generating storyboards for video production using artificial intelligence

    KR102678148B1

  • KR20240039777A