Method for generating images and electronic device performing same
By classifying and scoring images based on user input and metadata, the electronic device generates composite images that align with user preferences, improving the relevance and quality of output.
Patent Information
- Application Number
- PCT/KR2025/006339
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-08-13
- Filing Date
- 2025-05-12
- Publication Date
- 2026-01-02
AI Technical Summary
Existing generative models struggle to generate outputs tailored to a user's intent and personal preferences based on direct user inputs, leading to suboptimal results.
An electronic device determines candidate images based on user input and metadata, classifies them into sets, assigns importance and preference scores, and generates a composite image using these scores to align with user preferences.
The method effectively generates images that better match user intent and preferences, enhancing the relevance and quality of output.
Smart Images

Figure KR2025006339_02012026_PF_FP_ABST
Abstract
Description
Image generation method and electronic device performing the same
[0001] One embodiment disclosed in this document relates to an electronic device for generating an image, a method thereof, and a storage medium, and hereinafter, a technology for generating an image based on a user input is disclosed.
[0002] Various features can be supported to enhance the convenience of using electronic devices. For example, personalized services based on the user's electronic device usage are actively being developed. AI-based assist functions can monitor the overall usage status of electronic devices and provide customized assistance functions.
[0003] Generative models, which can generate media such as text and images in response to prompts, can be used in a variety of applications, including user experience, art, and development. The quality of the output from a generative model can vary depending on the prompts inputted. A method is needed to generate output tailored to the user's intent and personal preferences for prompts directly entered by the user.
[0004] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above-described matters constitute prior art related to the present disclosure.
[0005] In one embodiment, a method performed by an electronic device may include, in response to obtaining a user input, determining a plurality of candidate images from among pre-stored images based on the user input and metadata of images pre-stored in the electronic device. The method performed by the electronic device may include an operation of dividing the plurality of candidate images into one or more image sets, each image set including one or more images that satisfy a predetermined condition of mutual similarity. The method performed by the electronic device may include an operation of determining an importance score for each of the one or more image sets based on personal data of a user of the electronic device, the plurality of candidate images, and metadata of the plurality of candidate images. The method performed by the electronic device may include an operation of determining a target image set satisfying a predetermined condition from among the one or more image sets based on the importance score. The method performed by the electronic device may include an operation of determining an evaluation score and a preference score for each of the one or more images included in the target image set based on personal data of the user, the target image set, and metadata of the target image set. The method performed by the electronic device may include an operation of generating a composite image using the one or more images included in the target image set based on the evaluation score and the preference score.
[0006] According to one embodiment, a non-transitory computer-readable recording medium can store one or more programs including commands for executing the above-described method of operation.
[0007] According to an embodiment, an electronic device may include at least one processor including processing circuitry. The electronic device may include a memory including one or more storage media storing instructions. When the instructions are individually or collectively executed by the at least one processor, the electronic device may: in response to obtaining a user input, determine a plurality of candidate images from among the pre-stored images based on the user input and metadata of the images pre-stored in the electronic device. When the instructions are individually or collectively executed by the at least one processor, the electronic device may: classify the plurality of candidate images into one or more image sets, each image set including one or more images satisfying a condition for which mutual similarity is determined. When the instructions are individually or collectively executed by the at least one processor, the electronic device may: determine an importance score for each of the one or more image sets based on personal data of a user of the electronic device, the plurality of candidate images, and metadata of the plurality of candidate images. When the above commands are individually or collectively executed by the at least one processor, the electronic device may be caused to: determine a target image set satisfying a predetermined condition among the one or more image sets based on the importance score. When the above commands are individually or collectively executed by the at least one processor, the electronic device may be caused to: determine an evaluation score and a preference score for each of the one or more images included in the target image set based on the user's personal data, the target image set, and metadata of the target image set.When the above instructions are individually or collectively executed by the at least one processor, the electronic device may be configured to: generate a composite image using the one or more images included in the target image set based on the evaluation score and the preference score.
[0008] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0009] FIG. 2 is a diagram illustrating an artificial intelligence system according to one embodiment.
[0010] Figure 3 is a configuration diagram of an artificial intelligence system according to one embodiment.
[0011] Figure 4 is a drawing illustrating an image generation method according to an example.
[0012] Figure 5 is a flowchart of an image generation method according to one embodiment.
[0013] Figure 6 is a flowchart of a method for determining an importance score according to an example.
[0014] Figure 7 is a flowchart of a method for determining an evaluation score according to an example.
[0015] Figure 8 is a flowchart of a method for determining a preference score according to an example.
[0016] Figure 9 is a flowchart of an image generation method according to one embodiment.
[0017] Figure 10 is a flowchart of a method for determining an intimacy score according to an example.
[0018] Figure 11 is a flowchart of an image generation method according to one embodiment.
[0019] Figure 12 is a flowchart of a method for generating a panoramic image according to one embodiment.
[0020] Figure 13 is a block diagram of an artificial intelligence system according to one embodiment.
[0021] Figure 14 is a drawing illustrating an image generation method according to an example.
[0022] Figure 15 is a drawing illustrating an image generation method according to an example.
[0023] Hereinafter, an embodiment of the present document may be described with reference to the attached drawings.
[0024] FIG. 1 is a block diagram of an electronic device within a network environment according to one embodiment.
[0025] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) via a first network (198) (e.g., a short-range wireless communication network), or may communicate with an electronic device (104) or a server (108) via a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) via the server (108). According to one embodiment, the electronic device (101) may include a processor (120), a memory (130), an input module (150), an audio output module (155), a display module (160), an audio module (170), a sensor module (176), an interface (177), a connection terminal (178), a haptic module (179), a camera module (180), a power management module (188), a battery (189), a communication module (190), a subscriber identification module (196), or an antenna module (197). In some embodiments, the electronic device (101) may omit at least one of these components (e.g., the connection terminal (178)), or may have one or more other components added. In some embodiments, some of these components (e.g., the sensor module (176), the camera module (180), or the antenna module (197)) may be integrated into one component (e.g., the display module (160)).
[0026] The processor (120) may, for example, execute software (e.g., a program (140)) to control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (120) may store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in a volatile memory (132), process the commands or data stored in the volatile memory (132), and store result data in a non-volatile memory (134).
[0027] According to one embodiment, the processor (120) may be implemented as a circuit (e.g., a processing circuit) such as a system on chip (SoC) or an integrated circuit (IC). The processor (120) may include one or more processors. For example, the processor (120) may include a combination of one or more processors such as a CPU, a GPU, an MPU, an AP, and a CP.
[0028] According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (121). For example, when the electronic device (101) includes the main processor (121) and the auxiliary processor (123), the auxiliary processor (123) may be configured to use less power than the main processor (121) or to be specialized for a given function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as a part thereof.
[0029] The auxiliary processor (123) may control at least a part of functions or states associated with at least one component (e.g., a display module (160), a sensor module (176), or a communication module (190)) of the electronic device (101), for example, on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (123) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (180) or a communication module (190)). In one embodiment, the auxiliary processor (123) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, in the electronic device (101) itself where artificial intelligence is performed, or can be performed through a separate server (e.g., server (108)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0030] The memory (130) can store various data used by at least one component (e.g., processor (120) or sensor module (176)) of the electronic device (101). The data can include, for example, software (e.g., program (140)) and input data or output data for commands related thereto.
[0031] According to one embodiment, the memory (130) may include one or more memories. The instructions stored in the memory (130) may be stored in a single memory. The instructions stored in the memory (130) may be divided and stored in a plurality of memories. The instructions stored in the memory (130) may be individually or collectively executed by the processor (120) to cause the electronic device (101) (e.g., the electronic device (201) of FIG. 2) to perform and / or control the prompting described with reference to FIGS. 2 to 8. The instructions stored in the memory (130) may be individually or collectively executed by a plurality of processors to cause the electronic device (101) (e.g., the electronic device (201) of FIG. 2) to perform and / or control the prompting described with reference to FIGS. 2 to 8. According to one embodiment, the memory (130) may include volatile memory (132) or non-volatile memory (134).
[0032] The program (140) may be stored as software in the memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0033] The input module (150) can receive commands or data to be used in a component of the electronic device (101) (e.g., a processor (120)) from an external source (e.g., a user) of the electronic device (101). The input module (150) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0034] The audio output module (155) can output audio signals to the outside of the electronic device (101). The audio output module (155) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. According to one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0035] The display module (160) can visually provide information to an external party (e.g., a user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0036] The audio module (170) can convert sound into an electrical signal, or vice versa, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150), output sound through the sound output module (155), or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (101).
[0037] The sensor module (176) can detect the operating status (e.g., power or temperature) of the electronic device (101) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (176) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0038] The interface (177) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (101) with an external electronic device (e.g., the electronic device (102)). In one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0039] The connection terminal (178) may include a connector through which the electronic device (101) may be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0040] A haptic module (179) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. According to one embodiment, the haptic module (179) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0041] The camera module (180) can capture still images and videos. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0042] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented as, for example, at least a part of a power management integrated circuit (PMIC).
[0043] A battery (189) may power at least one component of the electronic device (101). In one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0044] The communication module (190) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may operate independently from the processor (120) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (194) (e.g., a local area network (LAN) communication module, or a power line communication module). Among these communication modules, the corresponding communication module can communicate with an external electronic device (104) via a first network (198) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (199) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules can be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can verify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) by using subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (196).
[0045] The wireless communication module (192) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). The NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimization of terminal power and connection of multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (192) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (192) can support various technologies for securing performance in a high-frequency band, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), an external electronic device (e.g., the electronic device (104)), or a network system (e.g., the second network (199)). According to one embodiment, the wireless communication module (192) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0046] The antenna module (197) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (198) or the second network (199), may be selected from the plurality of antennas by, for example, the communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (197).
[0047] In one embodiment, the antenna module (197) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high-frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high-frequency band.
[0048] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0049] According to one embodiment, commands or data may be transmitted or received between the electronic device (101) and an external electronic device (104) via a server (108) connected to a second network (199). Each of the external electronic devices (102 or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations executed in the electronic device (101) may be executed in one or more of the external electronic devices (102, 104, or 108). For example, when the electronic device (101) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (101) may, instead of or in addition to executing the function or service by itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (101) may provide an ultra-low latency service by using distributed computing or mobile edge computing, for example. In one embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or a neural network. According to one embodiment, the external electronic device (104) or the server (108) may be included in the second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0050] FIG. 2 is a diagram illustrating an artificial intelligence system according to one embodiment.
[0051] The artificial intelligence system (200) may include a user query / response interface (210), an AI framework (220), a knowledge component (230), an application / service component (240), and / or a generative model (250).
[0052] In an artificial intelligence system (hereinafter, referred to as the system) (200), a user query / response interface (210) can receive an input. The input can include a user input and / or data obtained or generated by an electronic device (e.g., the electronic device (101) or the electronic device (301) described above). The data can include images, videos, and / or sensor data generated by at least one processor (e.g., at least one processor (120) or the processor (310)) of the electronic device (e.g., illuminance data around the electronic device obtained from a sensor or sensor hub (e.g., a coprocessor (123), posture data (or orientation data) of the electronic device, a temperature inside the electronic device (e.g., a temperature of the display module (160) or a temperature of the at least one processor (120)), size information of a display area of the display module (160), and / or an image obtained through an image sensor (e.g., included in a camera module (180)) of the electronic device). For example, the user input may be a type of input such as natural language, touch data obtained through a touch circuit included in the display module (160) (e.g., used to identify input from a finger and / or stylus), an image, audio, and / or video. Additionally, when the user input is transmitted, context information may also be transmitted. The context information may include various side information related to the time when the user input is input into the system (200). For example, there may be application information currently being used by the user or location information of the user.Additionally, user input may be a mixed type of input, including natural language, images, audio, video, and / or contextual information, as described above. Furthermore, user input may include non-natural language input, such as selecting a menu.
[0053] The user query / response interface (210) may provide the user with output from the generative artificial intelligence system. The output may include results (or result information) generated or obtained by the system (200) based at least in part on the input. The output may include a natural language-based response and / or specific content. The output may also include an action requested by the user. For example, the output may have a format based on user settings of the electronic device.
[0054] The AI framework (220) can receive user input. Based on the user input (e.g., the user's query), the AI framework (220) can coordinate and control one or more components necessary to perform an action corresponding to the user's intent.
[0055] User input received from the user query / response interface (210) can be transmitted to a prompt design component (221). The prompt design component (221) can be used to generate a prompt suitable as input to a generative model (250) based on the user input.
[0056] The prompt design component (221) may be an AI component that uses a machine learning algorithm or a neural network. The prompt design component (221) may generate improved prompts through learning over time. The prompt design component (221) may access a knowledge repository (230) to generate prompts based on user input. The knowledge repository (230) may include user preference data, a prompt library, and / or prompt examples. The prompt design component (223) may provide the generated prompts to a generative model (e.g., LLM, LVM, and / or LMM).
[0057] The APIs / Plugins management component (223) can communicate with an external information source based on a request for additional information when user input is transmitted to the generative model (250).
[0058] The APIs / Plugins management component (223) can establish a communication channel for communication with the outside of the system (200) via the API. The APIs / Plugins management component (223) can enable access to various data sources via the communication channel. For example, the APIs / Plugins management component (223) can be used to request another component (e.g., an application / service component (240)) to perform feedback (or response) according to the prompt. The acquired information can be used to generate a prompt by the prompt design component (221) together with user input, or can be used as input to the generative model (250).
[0059] The APIs / Plugins management component (223) can request a final action via an API when a final action in response to user input, rather than an intermediate action, must be performed by an application or service.
[0060] The refiner component (225) can at least partially tune (or adjust) (or change) the results (e.g., content) obtained (or output) from the generative model (250). For example, the refiner component (225) can determine the relevance (e.g., score) between the output (e.g., content) of the generative model and the user input. For example, the refiner component (225) can determine whether the output contains biased information (e.g., selective information). For example, the refiner component (225) can determine whether the output contains harmful information (e.g., violent content or profanity).
[0061] The refinement component (225) can determine the degree of matching (e.g., score) between the output of the generative model (250) and the user input (e.g., the intent of the user input). If the refinement component (225) determines that the output of the generative model (250) does not correspond to the user input, the refinement component (225) can modify the output so that it corresponds to the user input.
[0062] The refinement component (225) can provide hints (e.g., hints for prompt generation) to the user so that the user can obtain information that matches the user's intention from the generative model (250).
[0063] In one embodiment, the generative model (250) may form at least a portion of an artificial intelligence neural network. The generative model (250) may include a model that generates images or a model that generates language. The image generation model may include, for example, a generative adversarial network (GAN), a variational autoencoder (VAE), or a diffusion-based model using a VAE and a transformer. The language generation model may include, for example, a large language model (LLM), a large multimodal model (LMM), a large vision model (LVM), a large vision language model (LVLM), or a large action model (LAM). The LAM may automatically generate actions for an environment (e.g., a robot, a car, an electronic device (101), or a program (140)). Additionally, for at least some AI models (e.g. LLM), there may be a LoRA (low-rank adaptation) adapter that is fine-tuned for a specific task or specific situation, for example.
[0064] Figure 3 is a configuration diagram of an artificial intelligence system according to one embodiment.
[0065] According to one embodiment, the artificial intelligence system (300) (hereinafter, the system) may include an electronic device (301) (e.g., the electronic device (101) of FIG. 1). The electronic device (301) may be a device such as a mobile terminal (e.g., a smartphone, a tablet, a laptop) or a fixed terminal (e.g., a personal computer (PC)).
[0066] According to one embodiment, the electronic device (301) may include at least some of the components of the electronic device (101) of FIG. 1. The electronic device (301) may include a processor (310) including processing circuitry (e.g., processor (120) of FIG. 1). The processor (310) may include at least one processor. The electronic device (301) may include a memory (320) including one or more storage media for storing instructions (e.g., memory (130) of FIG. 1).
[0067] In the system (300), the electronic device (301) can generate an image using an artificial intelligence model (330).
[0068] According to one embodiment, the artificial intelligence model (or artificial intelligence neural network) (330) may include various foundation models such as a language model, a code model, an image model, and / or other artificial intelligence neural network models. The artificial intelligence model (330) may include an LLM, an LVM, and / or an LVLM (e.g., a generative model (250) of FIG. 2). For convenience of explanation, the present disclosure will describe LLM and / or LVM as examples.
[0069] The artificial intelligence model (330) that can be used in the present disclosure may include an LLM, a language model based on an artificial intelligence neural network that has learned a large amount of text data through pre-training. The LLM may include a relatively larger number of parameters (e.g., approximately 10 billion or more) than conventional general language models. The LLM may utilize a transformer artificial intelligence neural network structure based on an attention mechanism.
[0070] In one embodiment, the training of the LLM may include pre-training and / or fine-tuning. Pre-training may involve training the LLM to acquire general language knowledge using a large amount of text data. For example, pre-training may involve self-supervised learning, which predicts the next word in a text string using a previous word string. Fine-tuning may involve training the LLM to be suitable for a specific domain (e.g., chatbot, AI assistant, translation, summary generation, question answering) and / or task. Fine-tuning may involve further training (e.g., supervised learning, adaptive learning) the LLM using a dataset corresponding to the specific domain and / or task based on the pre-trained model. The LLM may perform a task based on text input containing natural language, referred to as a prompt.
[0071] In one embodiment, fine-tuning can be omitted in LLM learning. Users can control the prompts provided to the LLM to improve performance on a desired task. For example, users can control whether the prompts provide additional examples of tasks and / or guidance for performing the task, such as in-context learning, zero-shot learning, and / or few-shot learning. Publicly available LLMs include Bidirectional Encoder Representations from Transformer (BERT) and generative pre-trained transformer (GPT).
[0072] The term "LLM" can refer to the language neural network model itself, but it can also refer to the model of an LLM-based application (e.g., chatbot, AI assistant, translation, summary generation, text classification, sentence generation). For example, an LLM-based chatbot like ChatGPT or an LLM-based translator can also be referred to as "LLM."
[0073] "LLM" may include an inference engine utilizing the LLM neural network model. For example, "inputting an input prompt to the LLM" may mean "inputting the input prompt to an inference engine based on the LLM." For example, "the output of the LLM for the input prompt" may mean the output information of the last neural network layer of the LLM obtained when the input prompt is input to the LLM-based inference engine, and / or the output information modified through additional processing.
[0074] The attention mechanism is a technology that enables an artificial intelligence model (330) to focus (attention) on important parts within input data. The attention mechanism can be utilized to predict output data by predicting the degree to which some of the time-series input data (e.g., time-series input data such as voice or video, or input data of some layers of a neural network) contributes to the output of the intermediate layer and / or the final output of the neural network. A recurrent neural network (RNN) structure that sequentially processes each element of a sequence may have poor prediction performance when there is information dependency between long time-series distances, but the attention mechanism can consider the information dependency between long time-series distances by controlling the degree of weight concentration (attention level) within the entire and / or part of the context of the input data. The transformer can be configured as an encoder-decoder structure. The encoder can process the input data and output compressed information (e.g., contextual representation). The decoder can process compressed information and output data in token units. Each encoder and decoder can include an independent attention network, and may further include a cross-attention network connecting the encoder and decoder.
[0075] According to one embodiment, an external server (e.g., server (108) of FIG. 1) or an external electronic device (e.g., electronic device (102) or electronic device (104) of FIG. 1) may include (or store) an artificial intelligence model (330). The artificial intelligence model (330) may be embedded (or installed or deployed) in the external server or the external electronic device. The electronic device (301) may offload at least a portion of the work associated with generating an image generated by the electronic device (301) to the external server or the external electronic device. The external server or the external electronic device may perform at least a portion of the work associated with generating an image generated by the electronic device (301) using the artificial intelligence model (330). The external server or the external electronic device may transmit the result of the work (or a portion of the work) performed using the artificial intelligence model (330) to the electronic device (301). The electronic device (301) can receive the results of a task related to image generation from an external server or an external electronic device. The electronic device (301) can provide an image based on the results of the task received from the external server or the external electronic device.
[0076] In the system (300), the method by which the electronic device (301) generates an image is not limited to the illustrated example. According to one embodiment, the electronic device (301) may include (or store) an artificial intelligence model (330). The artificial intelligence model (330) may be built into (or installed or distributed) the electronic device (301). The electronic device (301) may generate an image by on-device artificial intelligence computing using the artificial intelligence model (330) included in the electronic device (301). The electronic device (301) may provide an image generated using the artificial intelligence model (330) (e.g., LLM, LVM, LVLM).
[0077] According to one embodiment, the system (300) may be a system based on the artificial intelligence system (200) of FIG. 2. The system (300) may include at least some components of the artificial intelligence system (200). For example, the electronic device (301) may be implemented as a user query / response interface (210).
[0078] In the system (300), the electronic device (301) can obtain data such as images, audio, video, voice, text, call records, messages, web page visit records, notes, or sensor information. The electronic device (301) can obtain data based on user input such as taking pictures, recording, typing, drawing, or operating an application. The electronic device (301) can receive data from an external server (e.g., server (108) of FIG. 1) or an external electronic device (e.g., electronic device (102) or electronic device (104) of FIG. 1, or external electronic device (302)).
[0079] An electronic device (301) can receive (or acquire) user input for the electronic device (301). The user input may include voice data corresponding to a user's speech, text data by a user's typing, or touch input data for icons, buttons, images, text, or various indicators acquired through a display of the electronic device.
[0080] The electronic device (301) can generate a prompt based at least in part on user input to the electronic device (301). The prompt can include natural language, text, photos, videos, audio, or any combination thereof. In one embodiment, the electronic device (301) can generate a prompt according to a predetermined template based at least in part on user input. In one embodiment, the electronic device (301) can generate a prompt based at least on user input using an artificial intelligence model (330).
[0081] The artificial intelligence model (330) can generate output data based on a prompt. The output data may be a multimedia file containing elements such as text, images (e.g., photos, videos), or graphics. The types and meaning of the output data generated by the artificial intelligence model (330) are not limited to the examples described, and for example, the artificial intelligence model (330) can classify images or output selected images based on a prompt containing instructions (or commands) composed of text and images.
[0082] FIG. 4 is a drawing (400) illustrating an image generation method according to an example.
[0083] According to one embodiment, in an artificial intelligence system (300) (hereinafter, system), an electronic device (e.g., electronic device (101) of FIG. 1 or electronic device (301) of FIG. 3) can obtain user input.
[0084] An electronic device can obtain user input regarding image generation. The electronic device can receive user input regarding image generation, which may or may not include instructions such as "the best picture" or similar meanings (e.g., "the best shot (or image)", "a picture of my favorite style"). For example, the electronic device can obtain user input including instructions or commands for image generation, such as "the best shot I took in Jeju Island" or "make a family photo of our graduation." For example, the electronic device can obtain user input including instructions or commands for generating a panoramic image using multiple images, such as "make the best panorama shot (or image) I took recently." A panoramic image can represent an image that is a combination of multiple images.
[0085] The electronic device can generate an optimal image that is expected to be preferred by the user based on at least one of the user input, images previously stored on the electronic device, metadata of the previously stored images, and personal data of the user of the electronic device, based on the user input regarding image generation.
[0086] In response to receiving a user input, the electronic device may determine a plurality of candidate images (410) from among images stored in the electronic device that are associated with (or similar to) the user input.
[0087] An electronic device may classify a plurality of candidate images (410) into one or more image sets (420) (e.g., a first image set, a second image set, a third image set, a fourth image set). Each of the one or more image sets (420) may include one or more images that are similar to each other (or that satisfy a condition of mutual similarity).
[0088] An electronic device can determine an importance score for each of one or more sets of images (420). The importance score is an indicator of a user's interest in an image, based on factors such as whether the user of the electronic device selected the image, performed additional operations on it, exported it externally, or captured it under special circumstances. A method for determining the importance score is described in detail with reference to FIG. 6.
[0089] The electronic device can determine a target image set (430) (e.g., a first image set, a second image set, a third image set) that satisfies a predetermined condition among one or more image sets (420) based on the importance score.
[0090] The electronic device can determine an evaluation score and a preference score for each of one or more images included in the target image set (430). A method for determining the evaluation score and the preference score is described in detail with reference to FIGS. 7 and 8, respectively.
[0091] The electronic device can generate a composite image (440) using one or more images included in the target image set (430) based on the evaluation score and the preference score.
[0092] The target image set (430) may include a plurality of target image sets (e.g., a first image set, a second image set, a third image set). The electronic device may generate a plurality of composite images (440) corresponding to each of the plurality of target image sets.
[0093] The electronic device may select a first target image (41) having the highest evaluation score among one or more images included in the first image set. The electronic device may select a second target image (43) having the highest preference score among one or more images included in the first image set. The electronic device may generate a composite image using the first target image (41) and the second target image (43).
[0094] The electronic device may generate a composite image using the image (43) in a composition of images deemed to be preferred by the user, if the second image set includes one image (43).
[0095] The electronic device may select a third target image (44) with the highest evaluation score among one or more images included in the third image set. The electronic device may select a fourth target image (45) with the highest preference score among one or more images included in the third image set. The electronic device may generate a composite image using the third target image (44) and the fourth target image (45).
[0096] The electronic device can generate a panoramic composite image (450) by connecting multiple composite images (440).
[0097] As a non-limiting example, the electronic device may calculate a significance score for each of the plurality of composite images (440). A method for determining the significance score is described in detail with reference to FIG. 6. The electronic device may adjust the plurality of composite images (440) by cropping, reducing, or enlarging the plurality of composite images (440) to a size or ratio proportional to the significance score of each of the plurality of composite images (440). The electronic device may generate a panoramic composite image (450) by connecting the adjusted plurality of composite images.
[0098] According to one embodiment, the electronic device may generate a composite image by using a target image (e.g., a first target image (41), a third target image (44)) having the highest evaluation score from each of a plurality of target image sets (e.g., a first image set, a second image set, a third image set). The electronic device may select the first target image (41) having the highest evaluation score from among one or more images included in the first image set. The electronic device may select the third target image (44) having the highest evaluation score from among one or more images included in the third image set. The electronic device may generate the composite image by using the first target image (41) and the third target image (44).
[0099] Figure 5 is a flowchart of an image generation method according to one embodiment.
[0100] According to one embodiment, the operations 510 to 560 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 3) and a memory (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 3).
[0101] An electronic device can receive (or acquire) user input for the electronic device. The electronic device can acquire user input related to image generation. The user input may include voice data corresponding to the user's speech, text data typed by the user, or touch input data for icons, buttons, images, text, or various indicators acquired through the display of the electronic device.
[0102] User input can be text data or voice data representing instructions or commands to generate an image, such as 'Best shot taken in Jeju Island' or 'Make a family photo for graduation.'
[0103] User input can be obtained based on a predetermined template. For example, the electronic device can display a predetermined template including elements such as 'person,' 'location,' 'event,' 'date (or date range),' 'animal,' 'object,' 'weather,' 'place,' 'mood,' 'composition,' and 'source image' on the display of the electronic device. The user can select options for at least some of the elements included in the predetermined template, input text, or select an image that serves as the basis for generating a composite image through the display of the electronic device. The electronic device can obtain user input based on the predetermined template.
[0104] In operation 510, the electronic device, in response to obtaining a user input, may determine a plurality of candidate images from among the previously stored images based on the user input and / or metadata of the images previously stored in the electronic device.
[0105] The electronic device may determine a plurality of candidate images from among the images selected by the user, if the user input includes images selected by the user from among images previously stored on the electronic device. For example, the user input may include an input for selecting a plurality of images directly from a gallery application.
[0106] The images previously stored on the electronic device may be images stored in the electronic device's internal or external storage. These images may include images stored in the cloud associated with the electronic device's user account. These images may be images selected by the user.
[0107] Image metadata may include at least one of the following: location, date, or time the image was taken; objects included in the image (e.g., people, animals, objects, background); resolution, lighting, composition, orientation, color space (e.g., RGB, YCbCr, grayscale); whether a flash was used; exposure time; or lens focal length; information about the camera that took the image (e.g., make, model); favorite information about the image; or event information about the image (e.g., graduation, wedding, party).
[0108] The multiple candidate images may represent images related to (or similar to) the user input among previously stored images. The electronic device may classify multiple candidate images among the previously stored images that are highly related to the user input. For example, the electronic device may determine multiple candidate images among the previously stored images whose similarity to the user input exceeds a predetermined threshold.
[0109] According to one embodiment, the electronic device may determine a plurality of candidate images from among pre-stored images using a trained artificial intelligence model (e.g., a classifier) (e.g., the artificial intelligence model (330) of FIG. 3). In response to providing metadata of pre-stored images and / or at least a portion of user input as input to the trained artificial intelligence model, the electronic device may obtain a plurality of candidate images output by the trained artificial intelligence model.
[0110] According to one embodiment, the electronic device can extract at least one keyword included in a user input. The electronic device can determine a plurality of candidate images corresponding to at least one keyword included in the user input from among previously stored images.
[0111] For example, the user input may represent an instruction or command to generate an image, such as "Best shot taken in Jeju Island" or "Make a family photo of a graduation ceremony." The electronic device may determine a plurality of candidate images, such as images taken in Jeju Island corresponding to the keyword included in the user input ("Jeju Island") or images containing family member objects taken at a graduation ceremony corresponding to the keywords ("graduation ceremony" and "family").
[0112] In one embodiment, the electronic device may receive a user input specifying a range of images. For example, the user input may include an instruction or command specifying a range of images, such as "Make the best panorama with photos from last year" or "Make the best images by season." The electronic device may then determine a plurality of candidate images from among images taken or stored last year or images taken or stored by season (or quarter).
[0113] In operation 520, the electronic device may divide the plurality of candidate images into one or more image sets, each image set including one or more images that satisfy a mutual similarity condition.
[0114] According to one embodiment, the electronic device may determine one or more image sets, each of which includes one or more images that satisfy a predetermined condition (e.g., a predetermined threshold or greater) with respect to mutual similarity with respect to factors such as overall color, blurriness, mood, an object or object characteristics (e.g., an identifier, a relationship or intimacy between a user of the electronic device and a person, a background, a number of objects), or composition, among a plurality of candidate images. Each image set may include one or more images that are similar to each other (or that satisfy the predetermined condition of mutual similarity).
[0115] According to one embodiment, the electronic device may use a trained artificial intelligence model (e.g., a classifier) (e.g., the artificial intelligence model (330) of FIG. 3) to classify a plurality of candidate images into one or more image sets, each of which includes one or more images that satisfy a condition of mutual similarity.
[0116] In one embodiment, the mutual similarity of one or more images satisfying a given condition may include the similarity of all pairs of the images satisfying the given condition. For example, the mutual similarity of the first image, the second image, and the third image satisfying the given condition may include the similarity of each of the first image-second image pair, the first image-third image pair, and the second image-third image pair satisfying the given condition.
[0117] According to one embodiment, the mutual similarity of one or more images satisfying a predetermined condition may include the similarity between any one of the images and the similarity between the image and other images satisfying a predetermined condition. For example, the mutual similarity between the first image, the second image, and the third image satisfying a predetermined condition may include the similarity between the first image and the second image, and the similarity between the first image and the third image satisfying a predetermined condition.
[0118] According to one embodiment, for any image among a plurality of candidate images, if there is no other image that satisfies a condition of similarity with the image, the image set may include only the image.
[0119] In operation 530, the electronic device may determine an importance score for each of one or more sets of images based on the user's personal data, the plurality of candidate images, and metadata of the plurality of candidate images.
[0120] The importance score is an indicator of a user's interest in an image, based on factors such as whether the user of the electronic device selected the image, performed additional actions on it, exported it externally, or captured it in a special situation.
[0121] The user's personal data may include at least one of contacts stored on the electronic device (e.g., phone book, email address), images, information about the user's relationship, call logs (e.g., received calls, outgoing calls, missed calls), messages (e.g., received messages, outgoing messages), email, audio, schedules (e.g., calendar, notifications), application data (e.g., settings, login information, user data, custom tags, files created within the application, files stored through the application), memos, notes, web browsing data (e.g., bookmarks, history, cache of a web browser), location history, or preference data.
[0122] The method for determining the importance score is described in detail with reference to FIG. 6.
[0123] In operation 540, the electronic device may determine a target image set (e.g., target image set (430) of FIG. 4) that satisfies a predetermined condition (e.g., above or exceeding a predetermined threshold) among one or more image sets based on the importance score.
[0124] According to one embodiment, the electronic device may determine an image set having an importance score equal to or greater than a predetermined threshold among one or more image sets as a target image set.
[0125] In operation 550, the electronic device may determine an evaluation score and a preference score for each of one or more images included in the target image set based on the user's personal data, the target image set, and metadata of the target image set.
[0126] The evaluation score is an indicator that evaluates the quality (or completeness) of an image based on factors such as the depiction of objects within the image or the brightness of the image. The method for determining the evaluation score is described in detail with reference to Figure 7.
[0127] A preference score is an indicator of a user's preference for an image, based on factors such as whether the user of the electronic device selected the image, performed additional actions, or exported it. The method for determining the preference score is described in detail with reference to FIG. 8.
[0128] In operation 560, the electronic device may generate a composite image using one or more images included in the target image set based on the evaluation score and the preference score. Specifically, the electronic device may generate a composite image using at least some of the one or more images included in the target image set based on the evaluation score and the preference score.
[0129] According to one embodiment, the electronic device can generate a synthetic image using an artificial intelligence model (e.g., the artificial intelligence model (330) of FIG. 3). A method for generating a synthetic image is described in detail with reference to FIGS. 9, 11, and 12.
[0130] According to one embodiment, an electronic device can determine preference scores for images previously stored in the electronic device. The electronic device can determine a target composition based on the preference scores for the previously stored images. The target composition may represent a composition of images deemed preferred by the user.
[0131] An electronic device can generate a composite image based on a target composition. The electronic device can generate a composite image expressed in the target composition using one or more images included in a set of target images.
[0132] For example, when generating an image containing a single object (or a person object) based on user input, the target composition may be a composition in which the object is centered. For example, when generating an image containing multiple objects (or person objects) based on user input, the target composition may be a composition in which the objects are arranged in three equal parts.
[0133] Figure 6 is a flowchart of a method for determining an importance score according to an example.
[0134] According to one embodiment, the operations 610 to 660 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0135] According to one embodiment, operation 530 of determining an importance score for each of one or more image sets of FIG. 5 may include operations 610 to 660.
[0136] Referring to FIG. 6, operations 610 to 660 are depicted as being performed sequentially, but each operation may be performed in parallel or in any order. Depending on the embodiment, some of the operations may be omitted, or additional operations may be performed based on importance conditions.
[0137] According to one embodiment, the electronic device may determine an importance score for each of one or more sets of images based on personal data of the user, a plurality of candidate images, and metadata of the plurality of candidate images.
[0138] According to one embodiment, the electronic device can determine an importance score of the first image set based on whether each of one or more images included in the first image set among the one or more image sets satisfies a predetermined importance condition.
[0139] The given importance condition may include at least one of a condition for displaying the image as a favorite, a condition for the image to be associated with a social media application, a condition for the image to be one of multiple images captured within a given area, a condition for the image to be captured outside the user's usual activity area, a condition for the image to correspond to a given event, or a condition for the image to be an edited image.
[0140] In operation 610, the electronic device may increase the importance score of the first image set if an image included in the first image set satisfies a favorite display condition.
[0141] Users can mark previously saved images as favorites (or "likes") through the gallery application on their electronic devices. The electronic device may store information about the user's favorites as metadata for the image, or information about user-defined tags as personal data for the user.
[0142] In operation 620, the electronic device may increase the importance score of the first image set if the image included in the first image set satisfies a condition for being associated with a social media application. The social media application may include an application (e.g., a messaging application, a phone application, or a social media service (SNS) application) that allows users to interact with external users, such as by sending and receiving messages, chatting, making calls, posting posts, or leaving comments.
[0143] For example, if an image included in the first image set is an image that was transmitted to an external user through a social media application or received and stored from an external user, the image may satisfy a condition associated with the social media application.
[0144] In operation 630, the electronic device may increase the importance score of the first image set if the image included in the first image set satisfies a condition that the image is one of multiple images acquired within a certain area.
[0145] The electronic device can determine whether an image included in the first image set satisfies a condition of being one of multiple images acquired within a given area based on the location, date, or time at which the image was captured.
[0146] In operation 640, the electronic device may increase the importance score of the first image set if the images included in the first image set satisfy a condition that they are acquired outside the user's usual activity area.
[0147] The electronic device may determine whether an image included in the first set of images satisfies the condition of being acquired outside the user's usual area of activity based on the location history contained in the user's personal data, based on whether the image was acquired at a location that differs somewhat from the user's usual location history.
[0148] In operation 650, the electronic device may increase an importance score of the first image set if an image included in the first image set satisfies a condition corresponding to a given event.
[0149] A given event may include, for example, a graduation ceremony, a wedding, a party, or a trip. The electronic device may determine whether an image included in the first image set satisfies a condition corresponding to the given event based on the image metadata.
[0150] In operation 660, the electronic device may increase the importance score of the first image set if the image included in the first image set satisfies a condition that the image is an edited image.
[0151] The electronic device may increase the importance score of the first image set if the images included in the first image set are images edited in a photo editing application or a gallery application.
[0152] The electronic device may repeat the operations of FIG. 6 for each of one or more sets of images.
[0153] According to one embodiment, the electronic device may determine an importance score for a composite image generated using one or more images (or at least some of the one or more images included in the target image set) included in a target image set (e.g., the first image set, the second image set, or the third image set of FIG. 4 ). For example, the electronic device may determine an importance score for each of a plurality of composite images (e.g., the plurality of composite images (440) of FIG. 4 ). The electronic device may determine the importance score for the composite image based on personal data of the user, one or more images used to generate the composite image, and metadata of the one or more images used to generate the composite image.
[0154] According to one embodiment, the electronic device may perform operations 610 to 660 for each of one or more images used to generate the composite image. For example, if one or more images used to generate the composite image satisfy a condition for displaying favorites, the electronic device may increase the importance score of the composite image. If one or more images used to generate the composite image satisfy a condition for being associated with a social media application, the electronic device may increase the importance score of the composite image. If one or more images used to generate the composite image satisfy a condition for being one of multiple images acquired within a certain area, the electronic device may increase the importance score of the composite image. If one or more images used to generate the composite image satisfy a condition for being acquired outside of a user's usual activity area, the electronic device may increase the importance score of the composite image. If one or more images used to generate the composite image satisfy a condition corresponding to a predetermined event, the electronic device may increase the importance score of the composite image. The electronic device may increase the importance score of the composite image if one or more images used to generate the composite image satisfy a condition that the images are edited images.
[0155] Figure 7 is a flowchart of a method for determining an evaluation score according to an example.
[0156] According to one embodiment, the operations 710 to 740 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0157] According to one embodiment, operation 550 of determining an evaluation score and a preference score for each of one or more images included in the target image set of FIG. 5 may include operations 710 to 740.
[0158] Referring to FIG. 7, operations 710 to 740 are depicted as being performed sequentially, but each operation may be performed in parallel or in any order. Depending on the embodiment, some of the operations may be omitted, or additional operations may be performed based on evaluation conditions.
[0159] According to one embodiment, the electronic device may determine an evaluation score and a preference score for each of one or more images included in the target image set based on the user's personal data, the target image set, and metadata of the target image set. A method for determining the evaluation score is described below.
[0160] According to one embodiment, the electronic device can determine an evaluation score of a first image based on whether the first image included in the target image set satisfies a predetermined evaluation condition.
[0161] The given evaluation conditions may include at least one of the following: a condition in which the main object included in the image is unobstructed (or non-occluded), a condition in which the main object included in the image is in focus, a condition in which the brightness of the image is appropriate, or a condition in which the sharpness of the image is appropriate.
[0162] In operation 710, the electronic device may increase the evaluation score of the first image if an object included in the first image does not obscure the main object.
[0163] In operation 720, the electronic device may increase the evaluation score of the first image if a primary object included in the first image is in focus.
[0164] The primary object in an image can be determined in various ways. For example, the primary object in an image may be a person corresponding to the user of the electronic device. The primary object in an image may be an object selected by the user. The primary object in an image may be an object located at or near the center of the image. The primary object in an image may be a person whose intimacy with the user of the electronic device exceeds a threshold.
[0165] The primary object contained in an image can be determined based on user input. For example, if the user input includes instructions for a specific object (e.g., "Create a panoramic image of a family photo," "Take my best shot in Jeju Island"), the electronic device can determine the object indicated by the user input (e.g., a person object corresponding to the family, a person object corresponding to the user of the electronic device) as the primary object.
[0166] Electronic devices can determine different evaluation scores based on the focusing method of an image. For example, an electronic device may increase the evaluation score more when the focus is on a key object in an out-of-focus image, compared to an image captured in a pan-focus manner, where all parts are sharp, and an image captured in an out-of-focus manner, where only a portion is in focus.
[0167] In operation 730, the electronic device may increase the evaluation score of the first image if the brightness of the first image is at an appropriate level (e.g., the brightness is not higher or lower than a predetermined level compared to an average of a set of target images including the first image).
[0168] In operation 740, the electronic device may increase the evaluation score of the first image if the sharpness (or saturation) of the first image is at an appropriate level (e.g., within a predetermined range).
[0169] In operations 730 and 740, the standard for an appropriate degree (e.g., a predetermined degree, a predetermined range) of brightness or sharpness (or, luminosity or saturation) of an image is not fixed or limited to a specified value or range, and may vary or change depending on the settings of the electronic device, the preferences of the user of the electronic device, an average of at least some of the images previously stored in the electronic device, and is not limited to the present disclosure.
[0170] The electronic device can repeat the operations of FIG. 7 for each of one or more images included in the target image set.
[0171] Figure 8 is a flowchart of a method for determining a preference score according to an example.
[0172] According to one embodiment, the operations 810 to 830 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0173] According to one embodiment, operation 550 of determining an evaluation score and a preference score for each of one or more images included in the target image set of FIG. 5 may include operations 810 to 830.
[0174] Referring to FIG. 7, operations 810 to 830 are depicted as being performed sequentially, but each operation may be performed in parallel or in any order. Depending on the embodiment, some of the operations may be omitted, or additional operations may be performed based on preference conditions.
[0175] According to one embodiment, an electronic device may determine an evaluation score and a preference score for each of one or more images included in a target image set based on a user's personal data, a target image set, and metadata of the target image set. A method for determining a preference score is described below.
[0176] According to one embodiment, the electronic device can determine a preference score of a first image based on whether the first image included in the target image set satisfies a predetermined preference condition.
[0177] The established preference conditions may include at least one of a condition for displaying the image as a favorite, a condition for the image to be associated with a social media application, or a condition for the image to be an edited image.
[0178] In operation 810, the electronic device may increase a preference score of the first image if the first image satisfies a favorite display condition.
[0179] In operation 820, the electronic device may increase a preference score of the first image if the first image satisfies a condition of being associated with a social media application.
[0180] In operation 830, the electronic device may increase the preference score of the first image if the first image satisfies a condition that the first image is an edited image.
[0181] The electronic device can repeat the operations of FIG. 8 for each of one or more images included in the target image set.
[0182] Figure 9 is a flowchart of an image generation method according to one embodiment.
[0183] According to one embodiment, the operations 910 to 930 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0184] According to one embodiment, operation 560 of generating the composite image of FIG. 5 may include operations 910 to 930.
[0185] In operation 910, the electronic device may select a first target image with the highest evaluation score among one or more images included in the target image set. Any description of the evaluation score that overlaps with the description provided above with reference to FIG. 7 will be omitted.
[0186] In operation 920, the electronic device may select a second target image with the highest preference score among one or more images included in the target image set. Any description of the preference score that overlaps with the description provided above with reference to FIG. 8 will be omitted.
[0187] In operation 930, the electronic device can generate a composite image using the first target image and the second target image.
[0188] In one embodiment, the electronic device can generate a composite image using a background of a first target image and an object of a second target image.
[0189] The electronic device can generate a prompt based on a first target image with the highest evaluation score in a set of target images and a second target image with the highest preference score in the set of target images. The electronic device can generate text for generating a composite image based on the first target image and the second target image using a language model (e.g., the generative model (250) of FIG. 2 or the artificial intelligence model (330) of FIG. 3). For example, the electronic device can generate text for generating a composite image, such as "Create an image that synthesizes the background of [the first target image] and the first object, the second object, and the third object of [the second target image]" using a language model. The prompt can be in a multi-modal format including the text for generating the composite image and the first target image and the second target image.
[0190] The electronic device can generate a synthetic image based on a prompt using a vision model (e.g., a generative model (250) of FIG. 2 or an artificial intelligence model (330) of FIG. 3) (or a vision language model). The electronic device can generate a synthetic image in which the background of a first target image and an object of a second target image are synthesized based on the prompt using the vision model.
[0191] Figure 10 is a flowchart of a method for determining an intimacy score according to an example.
[0192] According to one embodiment, the operations 1010 to 1030 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0193] In one embodiment, operations 1010 to 1030 may be performed in advance of the operations of FIG. 5. In one embodiment, operations 1010 to 1030 may be performed in parallel with at least some of the operations of FIG. 5.
[0194] Referring to FIG. 10, operations 1010 to 1030 are depicted as being performed sequentially, but each operation may be performed in parallel or in any order. Depending on the embodiment, some of the operations may be omitted, or additional operations related to the intimacy score may be performed.
[0195] According to one embodiment, an electronic device can identify a person object included in at least some of the images previously stored on the electronic device. The electronic device can determine an intimacy score for the identified person object based on the user's personal data and the images previously stored on the electronic device.
[0196] An intimacy score is an indicator of a relationship between a non-user of an electronic device (or a person object corresponding to another person) and the user of the electronic device (or the person object corresponding to the user), such as an acquaintance, family member, close acquaintance, or non-acquaintance. An intimacy score can also indicate relationships between people other than the user of the electronic device.
[0197] The electronic device can determine an intimacy score for a person object identified in a pre-stored image based on a predetermined intimacy evaluation factor for at least some of the images pre-stored in the electronic device.
[0198] The established intimacy evaluation factors may include at least one of the frequency with which the person object is included in images previously stored on the electronic device, the distance between the user (or the person object corresponding to the user) and the person object in images containing the user and the person object, or the frequency with which an external user corresponding to the person object exchanged messages and phone calls with the user of the electronic device.
[0199] In operation 1010, the electronic device may increase an intimacy score for the person object based on the frequency with which the person object is included in images previously stored on the electronic device.
[0200] In operation 1020, the electronic device may increase an intimacy score for a person object based on a distance between the user and the person object in an image including the user and the person object together.
[0201] In operation 1030, the electronic device may increase an intimacy score for the person object based on the frequency with which an external user corresponding to the person object exchanges messages and phone calls with the user of the electronic device.
[0202] The electronic device may increase an intimacy score for a person object based on the frequency with which an external user corresponding to the person object interacts with the user, such as sending or receiving messages through a social media application of the electronic device, chatting, calling, or commenting.
[0203] The electronic device may store the aforementioned intimacy score as user relationship information indicating a relationship between the user of the electronic device and the person object(s) included in at least some of the images previously stored in the electronic device. For example, the user relationship information may include information that the intimacy score for the user of the electronic device with respect to a first person (or a person object corresponding to the first person) is N (or N%). The relationship information may further include a relationship type determined based on the user's personal information. For example, the relationship information may include information that the user of the electronic device has an intimacy score for respect to a second person of M, and that the second person is a family member of the user. It may be determined that the second person is a family member of the user based on the user's contact information.
[0204] Figure 11 is a flowchart of an image generation method according to one embodiment.
[0205] According to one embodiment, the operations 1110 to 1140 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0206] According to one embodiment, operation 560 of generating the composite image of FIG. 5 may include operations 1110 to 1140.
[0207] In operation 1110, the electronic device may select a first target image with the highest evaluation score among one or more images included in the target image set. Any description of the evaluation score that overlaps with the description provided above with reference to FIG. 7 will be omitted.
[0208] In operation 1120, the electronic device can identify a person object associated with the user of the electronic device based on an intimacy score in the target image set. That is, the electronic device can identify a person object associated with the user in the target image set based on an intimacy score for the person object included in at least some of the previously stored images. Any description of the intimacy score that overlaps with the above description with reference to FIG. 10 will be omitted.
[0209] The electronic device can identify a human object contained in a set of target images.
[0210] The electronic device may identify a person object associated with the user of the electronic device among the person object(s) included in the target image set based on the intimacy score (or relationship information of the user). For example, the electronic device may identify a person object associated with the user of the electronic device among the person objects included in the target image set whose intimacy score satisfies a predetermined condition (e.g., an intimacy score equal to or exceeding a predetermined threshold).
[0211] An electronic device may identify a person object associated with the user of the electronic device among the person objects included in the target image set based on the user's personal data. For example, the electronic device may identify a person object corresponding to a contact stored in the electronic device among the person objects included in the target image set, or a person object corresponding to a user of an external electronic device with which the user frequently exchanges messages, as the person object associated with the user of the electronic device.
[0212] In operation 1130, the electronic device may select a second target image with the highest preference score from among one or more images containing a person object associated with the user in the target image set. Any description of the preference score that overlaps with the description provided above with reference to FIG. 8 will be omitted.
[0213] In operation 1140, the electronic device can generate a composite image using the first target image and the second target image.
[0214] In one embodiment, the electronic device can generate a composite image using a background of a first target image and a human object of a second target image.
[0215] An electronic device can generate a prompt based on a first target image and a second target image. The electronic device can generate text for generating a composite image based on the first target image and the second target image using a language model (e.g., a generative model (250) of FIG. 2 or an artificial intelligence model (330) of FIG. 3). For example, the electronic device can generate text for generating a composite image, such as "Create an image that synthesizes the background of [the first target image] and the person object of [the second target image]," using a language model. The prompt can be in a multi-modal format including text for generating a composite image and the first target image and the second target image.
[0216] The electronic device can generate a synthetic image based on a prompt using a vision model (e.g., a generative model (250) of FIG. 2 or an artificial intelligence model (330) of FIG. 3) (or a vision language model). The electronic device can generate a synthetic image based on a prompt using the vision model, in which the background of the first target image and the person object of the second target image are synthesized.
[0217] According to one embodiment, the electronic device can identify a plurality of person objects (e.g., a first person object, a second person object) associated with a user of the electronic device among person objects included in a target image set based on an intimacy score.
[0218] The electronic device may select a second target image having a highest preference score from among one or more images containing a first human object associated with the user of the electronic device in the set of target images. The electronic device may select a third target image having a highest preference score from among one or more images containing a second human object associated with the user of the electronic device in the set of target images.
[0219] The electronic device can generate a composite image using a first target image, a second target image, and a third target image. For example, the electronic device can generate a composite image using the background of the first target image, the first person object of the second target image, and the second person object of the third target image.
[0220] According to one embodiment, the electronic device may select one or more images, for each of a plurality of target image sets (e.g., the first image set, the second image set, and the third image set of FIG. 4), whose evaluation scores satisfy a predetermined condition (e.g., being equal to or greater than a predetermined threshold), from among the plurality of target image sets. The electronic device may select the first target image having the highest preference score from among the selected one or more images.
[0221] The electronic device may, for each of a plurality of target image sets, select one or more images whose affinity scores for the identified person object satisfy a predetermined condition (e.g., a predetermined threshold or higher). The electronic device may then select a second target image with the highest affinity score from among the selected one or more images.
[0222] An electronic device can generate a composite image using a first target image and a second target image. The electronic device can generate the composite image using a background of the first target image and a person object of the second target image.
[0223] Figure 12 is a flowchart of a method for generating a panoramic image according to one embodiment.
[0224] According to one embodiment, the operations 1210 and 1220 below may be performed by an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 3). The electronic device may include at least some of the components of the electronic device (101) described in FIG. 1. For example, the electronic device may include at least one processor including a processing circuit (e.g., the processor (120) of FIG. 1), a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1), and a communication unit (e.g., the communication module (190) of FIG. 1).
[0225] According to one embodiment, operation 560 of generating the composite image of FIG. 5 may include operations 1210 and 1220.
[0226] According to one embodiment, the target image set determined in operation 540 of FIG. 5 may include a plurality of target image sets.
[0227] In operation 1210, the electronic device can generate a plurality of composite images, each corresponding to a plurality of target image sets. A description of the method for generating the composite images, which overlaps with the description provided above with reference to FIGS. 9 and 11, will be omitted.
[0228] In operation 1220, the electronic device can generate a panoramic composite image by connecting a plurality of composite images generated respectively corresponding to a plurality of target image sets. The electronic device can generate the panoramic composite image by connecting the plurality of composite images using a vision model (e.g., the generative model (250) of FIG. 2 or the artificial intelligence model (330) of FIG. 3) (or a vision language model).
[0229] The electronic device can calculate a significance score for each of the multiple composite images, as described above with reference to FIG. 6. The electronic device can adjust the multiple composite images by cropping, reducing, or enlarging them to a size or ratio proportional to the significance score of each of the multiple composite images. The electronic device can then generate a panoramic composite image by connecting the adjusted multiple composite images.
[0230] An electronic device can determine the connection order of multiple composite images based on user input. The electronic device can generate a panoramic composite image by connecting the multiple composite images according to the determined connection order.
[0231] When an electronic device receives a user input including a temporal range of images, the electronic device can generate a panoramic composite image by connecting multiple composite images in the order in which the composite images were captured or stored. For example, the electronic device can generate a panoramic composite image by connecting multiple composite images in seasonal order (spring-summer-fall-winter) based on a user input such as "Create a panorama shot of the best images of each season." The electronic device can generate a panoramic composite image by connecting multiple composite images in the shooting order based on a user input such as "Create a best panorama with last year's photos."
[0232] When an electronic device receives user input that includes the designation of a person object, the electronic device can determine the connection order based on the intimacy score between the user and the person object included in each of the multiple composite images. For example, the electronic device can generate a panoramic composite image by designating an image containing a person object with a high intimacy score as the center order.
[0233] Figure 13 is a block diagram of an artificial intelligence system according to one embodiment.
[0234] According to one embodiment, the artificial intelligence system (300) (hereinafter, the system) may include an electronic device (301) (e.g., the electronic device (101) of FIG. 1). The electronic device (301) may be a device such as a mobile terminal (e.g., a smartphone, a tablet, a laptop) or a fixed terminal (e.g., a personal computer (PC)).
[0235] According to one embodiment, the electronic device (301) may include at least a portion of the configuration of the electronic device (101) of FIG. 1. The electronic device (301) may include a processor (310) including processing circuitry (e.g., the processor (120) of FIG. 1). The processor (310) may include at least one processor. The electronic device (301) may include a memory (320) including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1). The memory (320) may include an information collection module (321), a photo classification module (322), an image selection module (323), an object detection module (324), and an image generation module (325). The information collection module (321), the photo classification module (322), the image selection module (323), the object detection module (324), and the image generation module (325) may be software (e.g., the program (140) of FIG. 1 ), commands, or a set of commands executable by the processor (310) stored in the memory (320). When the commands stored in the memory (320) (e.g., the information collection module (321), the photo classification module (322), the image selection module (323), the object detection module (324), or the image generation module (325)) are individually or collectively executed by at least one processor of the processor (310), the electronic device (301) may be caused to perform at least a part of an operation, function, or process of the image generation method of the present disclosure.
[0236] The electronic device (301) can acquire data such as images, audio, video, voice, text, call records, messages, web page visit records, notes, or sensor information through the information collection module (321).
[0237] The electronic device (301) can obtain metadata of an image (or an image previously stored in the electronic device (301)) through the information collection module (321). The metadata of the image may include at least one of the location, date or time at which the image was taken, objects included in the image (e.g., people, animals, objects, background), resolution, lighting, composition, direction, color space (e.g., RGB, YCbCr, grayscale), whether a flash was used, exposure time, or lens focal length, information about the camera that took the image (e.g., manufacturer, model), favorite information about the image, or event information about the image (e.g., graduation, wedding, party).
[0238] The electronic device (301) can obtain personal data of a user of the electronic device (301) through the information collection module (321). The personal data of the user may include at least one of contacts (e.g., phone book, email address), images, relationship information of the user, call records (e.g., received calls, outgoing calls, missed calls), messages (e.g., received messages, outgoing messages), emails, audio, schedules (e.g., calendar, notifications), application data (e.g., settings, login information, user data, custom tags, files generated within the application, files stored through the application), memos, notes, web browsing data (e.g., bookmarks, history, cache of a web browser), location history, or preference data stored in the electronic device.
[0239] In response to obtaining a user input, the electronic device (301) can determine a plurality of candidate images from among images previously stored in the electronic device (301) through the photo classification module (322). The electronic device (301) can classify a plurality of candidate images highly related to the user input from among the previously stored images through the photo classification module (322).
[0240] The electronic device (301) can classify a plurality of candidate images into one or more image sets, each of which includes one or more images that satisfy a condition of mutual similarity, through the photo classification module (322).
[0241] The electronic device (301) can determine an importance score for each of one or more image sets through the image selection module (323). The electronic device (301) can determine a target image set that satisfies a predetermined condition among one or more image sets based on the importance score through the image selection module (323).
[0242] The electronic device (301) can determine an evaluation score and a preference score for each of one or more images included in a target image set through the image selection module (323). The electronic device (301) can select a first target image having the highest evaluation score among one or more images included in the target image set through the image selection module (323). The electronic device (301) can select a second target image having the highest preference score among one or more images included in the target image set through the image selection module (323). The electronic device (301) can select a second target image having the highest preference score among one or more images including a person object associated with a user in the target image set through the image selection module (323).
[0243] The electronic device (301) can identify an object (or a person object) included in at least some of the images previously stored in the electronic device (301) through the object detection module (324). The electronic device (301) can determine an evaluation score for the image based on the object identified through the object detection module (324). The electronic device (301) can identify a person object associated with the user based on an intimacy score in the target image set.
[0244] The electronic device (301) can generate a composite image using one or more images included in a target image set based on an evaluation score and a preference score through the image generation module (325). The electronic device (301) can generate a composite image using a first target image having the highest evaluation score and a second target image having the highest preference score through the image generation module (325). The electronic device (301) can generate a composite image using a background of the first target image and an object (or a person object) of the second target image through the image generation module (325).
[0245] In the system (300), the electronic device (301) can perform at least a part of the operations, functions, or processes for generating the aforementioned image using an artificial intelligence model (330). The artificial intelligence model (330) may be an on-device model included (or stored) in the electronic device (301). The artificial intelligence model (330) may be a model included (or stored) in an external server (e.g., server (108) of FIG. 1) or an external electronic device (e.g., electronic device (102) or electronic device (104) of FIG. 1).
[0246] The electronic device (301) can determine a plurality of candidate images from among images previously stored in the electronic device (301) using an artificial intelligence model (330). In response to providing metadata of previously stored images and at least a portion of user input as input to the pre-trained artificial intelligence model (330), the electronic device (301) can obtain a plurality of candidate images output by the pre-trained artificial intelligence model (330).
[0247] For example, the user input may represent an instruction or command to generate an image, such as "Best shot taken in Jeju Island." In response to providing "Jeju Island" included in the user input and metadata of previously stored images as input to the artificial intelligence model (330), the electronic device may obtain a plurality of candidate images taken in Jeju Island, which are output by the artificial intelligence model (330).
[0248] The electronic device (301) may provide the user's relationship information (or, intimacy score) as input to a pre-trained artificial intelligence model (330). For example, the user input may represent an instruction or command to create an image, such as "Make a family photo of the graduation ceremony." In response to providing "graduation ceremony," "family," the user's relationship information, and metadata of previously stored images included in the user input as input to the artificial intelligence model (330), the electronic device may obtain a plurality of candidate images corresponding to the "graduation ceremony" event, which include person object(s) corresponding to the family output by the artificial intelligence model (330).
[0249] The electronic device (301) can use an artificial intelligence model (330) to classify a plurality of candidate images into one or more image sets, each of which includes one or more images that satisfy a condition of mutual similarity. In response to providing a plurality of candidate images as input to a pre-trained artificial intelligence model (330), the electronic device (301) can obtain one or more distinguished image sets (or a plurality of candidate images, each labeled for each image set) output by the pre-trained artificial intelligence model (330).
[0250] The electronic device (301) can determine a target image set using an artificial intelligence model (330). The electronic device (301) can generate a prompt based on a predetermined importance condition, the user's personal data, a plurality of candidate images, and metadata of the plurality of candidate images. For example, the electronic device can generate text for image classification (or selection) using a language model, such as, "Output a [target image set] that satisfies a [determined condition (e.g., above or above a predetermined threshold)] based on a [determined importance condition] among [one or more image sets] using [the user's personal data], [the plurality of candidate images], and [the metadata of the plurality of candidate images]." The prompt can be in a multi-modal format including the text for image classification and at least a portion of the aforementioned plurality of candidate images, the user's personal data, the predetermined importance condition, and metadata of the plurality of candidate images. In response to providing a prompt as input to an artificial intelligence model (330), the electronic device (301) can obtain a set of target images (or labeled images among a plurality of candidate images) output by the artificial intelligence model (330).
[0251] The electronic device (301) may select a first target image with the highest evaluation score among one or more images included in a target image set using an artificial intelligence model (330). The electronic device (301) may generate a prompt based on a predetermined evaluation condition, the user's personal data, the target image set, and metadata of the target image set. For example, the electronic device (301) may generate text for image classification, such as 'Output an image based on [the predetermined evaluation condition] among the [target image set] using [the user's personal data], [the target image set], and [the metadata of the target image set]' using a language model. The prompt may be in a multi-modal format including the text for image classification and at least a portion of the aforementioned target image set, the user's personal data, the predetermined evaluation condition, and the metadata of the target image set. In response to providing a prompt as input to the artificial intelligence model (330), the electronic device (301) can obtain a first target image (or a labeled image among one or more images included in the target image set) output by the artificial intelligence model (330).
[0252] The electronic device (301) can use an artificial intelligence model (330) to select a second target image with the highest preference score among one or more images included in a target image set. The method for selecting the first target image described above can be similarly applied to a method for selecting the second target image, with some necessary modifications, such as replacing "predetermined evaluation conditions" with "predetermined preference conditions."
[0253] The electronic device (301) may use an artificial intelligence model (330) to select a second target image having the highest preference score from among one or more images containing a person object associated with the user in a target image set. The electronic device (301) may generate a prompt based on a predetermined preference condition, a predetermined intimacy evaluation factor (or, user's relationship information), the user's personal data, a target image set, and metadata of the target image set. For example, the electronic device (301) may use a language model to generate text for image classification, such as, 'Output images containing [the first person] and [the second person] based on the [determined preference condition] among the [target image set] using [the determined intimacy evaluation factor (or, user's relationship information)], [the user's personal data], [the target image set], and [the metadata of the target image set].' The prompt may be in a multi-modal format, including text for image classification, the aforementioned target image set, the user's personal data, a predetermined preference condition, a predetermined intimacy evaluation factor (or, the user's relationship information), and at least a portion of the metadata of the target image set. In response to providing the prompt as an input to the artificial intelligence model (330), the electronic device (301) may obtain a second target image (or, a labeled image among one or more images included in the target image set) output by the artificial intelligence model (330).
[0254] The electronic device (301) can generate a synthetic image using an artificial intelligence model (330). The electronic device (301) can generate a prompt based on a predetermined evaluation condition, a predetermined preference condition, a predetermined intimacy evaluation factor (or, the user's relationship information), the user's personal data, a plurality of candidate images, and the metadata of the plurality of candidate images. For example, the electronic device (301) can generate text for image generation, such as 'Synthesize an image suitable for [user input] using [a predetermined evaluation condition], [a predetermined preference condition], [a predetermined intimacy evaluation factor (or, the user's relationship information)], [the user's personal data], [a plurality of candidate images], and [the metadata of the plurality of candidate images]', using a language model. The prompt can be in a multi-modal format including the text for image generation and at least a portion of the aforementioned predetermined evaluation condition, predetermined preference condition, predetermined intimacy evaluation factor (or, the user's relationship information), the user's personal data, the plurality of candidate images, and the metadata of the plurality of candidate images. The electronic device (301) can obtain a synthetic image output by the artificial intelligence model (330) in response to providing a prompt as input to the artificial intelligence model (330).
[0255] Figure 14 is a drawing illustrating an image generation method according to an example.
[0256] According to one embodiment, in an artificial intelligence system (300) (hereinafter, system), an electronic device (e.g., electronic device (101) of FIG. 1 or electronic device (301) of FIGS. 3 and 13) can obtain user input.
[0257] In response to receiving a user input, the electronic device can determine a plurality of candidate images associated with (or similar to) the user input from among images previously stored in the electronic device.
[0258] An electronic device may classify a plurality of candidate images into one or more image sets. Each of the one or more image sets may include one or more images that are similar to each other (or that satisfy a predetermined condition of mutual similarity).
[0259] The electronic device can determine a significance score for each of one or more sets of images. Any description of the significance score that overlaps with the description provided above with reference to FIG. 6 will be omitted.
[0260] An electronic device may determine a target image set (1410) that satisfies a predetermined condition among one or more image sets based on an importance score. The target image set (1410) may include a plurality of target image sets (e.g., a first image set, a second image set, a third image set).
[0261] According to one embodiment, the electronic device can generate a plurality of composite images (1421-1423), each corresponding to a plurality of target image sets. The electronic device can display the plurality of composite images (1421-1423) through a display of the electronic device. The electronic device can store an image selected by the user from among the plurality of composite images (1421-1423).
[0262] The electronic device can determine an evaluation score and a preference score for each of one or more images (1401 to 1411) included in the target image set (1410). Any description of the evaluation score and preference score that overlaps with the description provided above with reference to FIGS. 7 and 8 will be omitted.
[0263] The electronic device may select a first target image having the highest evaluation score among one or more images (1401 to 1403) included in a first image set. The electronic device may identify a person object (A) associated with the user from the first image set based on the intimacy score described above with reference to FIG. 10. The electronic device may select a second target image having the highest preference score from one or more images (1401 to 1403) including a person object (A) associated with the user from the first image set. The electronic device may generate a composite image (1421) using the first target image and the second target image. The electronic device may generate the composite image (1421) using the background of the first target image and the person object (A) of the second target image. The electronic device may generate the composite image (1421) with a target composition determined based on the preference scores of pre-stored images.
[0264] The electronic device can generate a composite image (1422) with a target composition using the image (1404), if the second image set includes one image (1404).
[0265] The electronic device can identify a person object (A) associated with the user and a person object (C) not associated with the user based on the intimacy score in the second image set. The electronic device can generate a composite image (1422) by removing the person object (C) not associated with the user.
[0266] The electronic device may select a first target image with the highest evaluation score from among one or more images (1405 to 1411) included in the third image set. The electronic device may identify a person object (A, B) associated with the user based on the intimacy score in the third image set.
[0267] The electronic device may select a second target image having the highest preference score from among one or more images (1405 to 1407) containing a human object (A) associated with the user in the third image set.
[0268] The electronic device may select a third target image having the highest preference score from among one or more images (1405 to 1407, 1411) containing a person object (B) associated with the user in the third image set.
[0269] The electronic device can generate a composite image (1423) using a first target image, a second target image, and a third target image. The electronic device can generate the composite image (1423) using the background of the first target image, the person object (A) of the second target image, and the person object (B) of the third target image. The electronic device can generate the composite image (1423) with a target composition determined based on the preference scores of the previously stored images.
[0270] According to one embodiment, the electronic device can generate a panoramic composite image (1440) by connecting a plurality of composite images (1421 to 1423). The electronic device can display the panoramic composite image (1440) through a display of the electronic device.
[0271] According to one embodiment, the electronic device may generate a composite image (1430) corresponding to the target image set (1410). For example, if the user input indicates an instruction or command to generate one image, such as 'Give me the best shot (or one best shot) taken in Jeju Island', the electronic device may generate one composite image (1430). The electronic device may select a first target image having the highest evaluation score among one or more images (1401 to 1411) included in the target image set (1410), a second target image having the highest preference score among one or more images (1401 to 1407) including a person object (A) associated with the user, and a third target image having the highest preference score among one or more images (1405 to 1407, 1411) including a person object (B) associated with the user. An electronic device can generate a composite image (1430) using the background of a first target image, a person object (A) of a second target image, and a person object (B) of a third target image. The electronic device can display the composite image (1430) through a display of the electronic device.
[0272] Figure 15 is a drawing illustrating an image generation method according to an example.
[0273] Screens (1510, 1520, 1530) are examples of screens displayed through a display of an electronic device (e.g., the electronic device (101) of FIG. 1 or the electronic device (301) of FIGS. 3 and 13). The electronic device may include at least a portion of the configuration of the electronic device (101) of FIG. 1 or the electronic device (301) of FIG. 13. The electronic device may include a processor including a processing circuit (e.g., the processor (120) of FIG. 1 or the processor (310) of FIG. 13). The electronic device may include a memory including one or more storage media for storing instructions (e.g., the memory (130) of FIG. 1 or the memory (320) of FIG. 13).
[0274] On the screen (1510), the electronic device can provide a panoramic composite image (1501). For example, at the end of the year, at a specific time of the year, or on an anniversary such as the birthday of the user of the electronic device, the panoramic composite image (1501) can be automatically generated and displayed in the gallery application.
[0275] For example, a user can check the entire image by scrolling the panoramic composite image (1501) left and right.
[0276] In one embodiment, on screen (1510), the electronic device can directly start generating a panoramic composite image.
[0277] On screen (1520), the electronic device may display a screen for obtaining user input regarding image generation.
[0278] According to one embodiment, an electronic device may provide (or recommend) a theme (or topic) for an image to be generated based on at least one of images previously stored in the electronic device, metadata of the previously stored images, and personal data of a user of the electronic device. The user may select some of the recommended themes. Based on a user input selecting a specific theme, the electronic device may generate a panoramic composite image corresponding to the theme.
[0279] According to one embodiment, the electronic device may provide options that are referenced in the creation of a panoramic composite image (e.g., size, time period or location at which the image was taken or stored, people (or people objects) to be included in the image, keywords).
[0280] For example, if a device has multiple images taken at a specific location among its stored images, it can recommend that location. The device can suggest effective keywords based on the detailed address where the images were taken. For example, the device might suggest "Seongsu-dong" as a recommended location for images taken throughout Seongsu-dong. If multiple images were taken at a specific cafe in Seongsu-dong on the same day or at different times, the device might suggest "Seongsu-dong Cafe" as a recommended location.
[0281] For example, an electronic device may suggest that a person object included in many of the previously stored images be included in the creation of a panoramic composite image. The electronic device may recommend the person object based on the user's relationship information. The electronic device may also present identifying information (e.g., name, relationship, phone number) of the person object along with an image of the person object.
[0282] For example, the electronic device can present keywords associated with an image, such as event information of the image (e.g., graduation, wedding, party, birthday), and various information extracted through object recognition of the image (e.g., clothing color, pose, background, atmosphere).
[0283] Users can select options such as the size of the panoramic composite image to be created, the time period or location when the previously stored images to be used to create the panoramic composite image were taken, the people to be included in the image, and keywords associated with the image.
[0284] The electronic device can select a plurality of candidate images from among images corresponding to the selected options based on a user input selecting options to be referenced in generating a panoramic composite image, and generate a panoramic composite image in a manner as described with reference to FIG. 12.
[0285] On the screen (1530), the electronic device can display the generated panoramic composite images (1502, 1503). For example, the electronic device can display at least one panoramic composite image (1503) along with a representative panoramic composite image (1502) that best suits the user's preference. For example, the user can check the entire image by scrolling the representative panoramic composite image (1502) left and right or pinching. The user can check the entire image of the representative panoramic composite image (1502) by rotating the electronic device. The user can select a portion of at least one panoramic composite image (1503) to check the entire image.
[0286] In one embodiment, the electronic device may regenerate a panoramic composite image. For example, the electronic device may generate a new panoramic composite image based on user input for regenerating the panoramic composite image or user input for modifying an option selected on the screen (1520).
[0287] The technical problems to be achieved in the present disclosure are not limited to the technical problems mentioned above, and other technical problems not mentioned will be clearly understood by a person having ordinary knowledge in the technical field to which the present disclosure pertains.
[0288] In one embodiment, a method performed by an electronic device (101, 301) may include, in response to obtaining a user input, an operation of determining a plurality of candidate images from among images previously stored in the electronic device (101, 301) based on the user input and metadata of the images previously stored in the electronic device (101, 301); an operation of dividing the plurality of candidate images into one or more image sets, each image set including one or more images that satisfy a predetermined condition of mutual similarity; an operation of determining an importance score for each of the one or more image sets based on personal data of a user of the electronic device (101, 301), the plurality of candidate images, and metadata of the plurality of candidate images; an operation of determining a target image set satisfying a predetermined condition from among the one or more image sets based on the importance score; an operation of determining an evaluation score and a preference score for each of the one or more images included in the target image set based on the personal data of the user, the target image set, and metadata of the target image set; and an operation of generating a composite image using the one or more images included in the target image set based on the evaluation score and the preference score.
[0289] According to one embodiment, the operation of determining an importance score for each of one or more sets of images based on the user's personal data, the plurality of candidate images, and metadata of the plurality of candidate images may include the operation of determining an importance score of the first set of images based on whether each of one or more images included in the first set of images from among the one or more sets of images satisfies a predetermined importance condition. The predetermined importance condition may include at least one of a favorite display condition, a condition of being associated with a social media application, or a condition of being one of multiple images acquired within a certain area.
[0290] According to one embodiment, the operation of determining an evaluation score and a preference score for each of one or more images included in the target image set based on the user's personal data, the target image set, and the metadata of the target image set may include the operation of determining the evaluation score of the first image based on whether the first image included in the target image set satisfies a predetermined evaluation condition; and the operation of determining the preference score of the first image based on whether the first image included in the target image set satisfies a predetermined preference condition.
[0291] According to one embodiment, the operation of generating a composite image using one or more images included in a target image set based on an evaluation score and a preference score may include: selecting a first target image having a highest evaluation score among the one or more images included in the target image set; selecting a second target image having a highest preference score among the one or more images included in the target image set; and generating a composite image using the first target image and the second target image.
[0292] In one embodiment, the operation of generating a composite image using the first target image and the second target image may include: generating a prompt based on the first target image having the highest evaluation score in the set of target images and the second target image having the highest preference score in the set of target images; and generating a composite image based on the prompt using the vision model.
[0293] According to one embodiment, a method performed by an electronic device (101, 301) may further include: identifying a person object included in at least some of the previously stored images; and determining an intimacy score for the identified person object based on the user's personal data and the previously stored images.
[0294] According to one embodiment, the operation of generating a composite image using one or more images included in a target image set based on an evaluation score and a preference score may include: selecting a first target image having a highest evaluation score among one or more images included in the target image set; identifying a person object associated with the user in the target image set based on an intimacy score for a person object included in at least some of the previously stored images; selecting a second target image having a highest preference score among one or more images including a person object associated with the user in the target image set; and generating a composite image using the first target image and the second target image.
[0295] According to one embodiment, a method performed by an electronic device (101, 301) further includes an operation of determining preference scores of images previously stored in the electronic device (101, 301); and an operation of determining a target composition based on the preference scores of the previously stored images, and the operation of generating a composite image using one or more images included in a target image set based on the evaluation scores and the preference scores may include an operation of generating a composite image based on the target composition.
[0296] In one embodiment, the target image set may include a plurality of target image sets. The operation of generating a composite image may include an operation of generating a plurality of composite images, each corresponding to the plurality of target image sets.
[0297] According to one embodiment, the target image set may include a plurality of target image sets. The method performed by the electronic device (101, 301) may further include an operation of generating a panoramic composite image by connecting a plurality of composite images generated respectively corresponding to the plurality of target image sets.
[0298] According to one embodiment, a non-transitory computer-readable recording medium can store a program for executing a method performed by an electronic device (101, 301) in combination with hardware.
[0299] According to one embodiment, an electronic device (101, 301) comprises at least one processor (120, 310) comprising processing circuitry; And a memory (120, 320) including one or more storage media storing instructions, and when the instructions are individually or collectively executed by at least one processor (120, 310), causing the electronic device (101, 301) to: in response to obtaining a user input, determine a plurality of candidate images from among the images previously stored in the electronic device (101, 301) based on the user input and metadata of the images previously stored in the electronic device (101, 301), divide the plurality of candidate images into one or more image sets each including one or more images that satisfy a predetermined condition of mutual similarity, determine an importance score for each of the one or more image sets based on personal data of the user of the electronic device (101, 301), the plurality of candidate images, and metadata of the plurality of candidate images, determine a target image set satisfying a predetermined condition from among the one or more image sets based on the importance score, determine an evaluation score and a preference score for each of the one or more images included in the target image set based on the personal data of the user, the target image set, and metadata of the target image set, and determine an evaluation score and a preference score for each of the one or more images included in the target image set based on the evaluation score and a synthetic image can be generated using one or more images included in the target image set based on the preference score.
[0300] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be caused to: determine an importance score of a first image set based on whether each of one or more images included in the first image set from among one or more image sets satisfies a defined importance condition. The defined importance condition may include at least one of a favorite display condition, a condition of being associated with a social media application, or a condition of being one of multiple images acquired within a certain area.
[0301] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be configured to: determine an evaluation score of a first image included in a target image set according to whether the first image satisfies a given evaluation condition; and determine a preference score of the first image included in the target image set according to whether the first image satisfies a given preference condition.
[0302] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be configured to: select a first target image having a highest evaluation score from among one or more images included in a set of target images; select a second target image having a highest preference score from among one or more images included in the set of target images; and generate a composite image using the first target image and the second target image.
[0303] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be configured to: generate a prompt based on a first target image having a highest evaluation score in a set of target images and a second target image having a highest preference score in the set of target images, and, using a vision model, generate a synthetic image based on the prompt.
[0304] According to one embodiment, the instructions, when individually or collectively executed by at least one processor (120, 310), may cause the electronic device (101, 301) to: identify a person object included in at least some of the stored images, and determine an intimacy score for the identified person object based on the user's personal data and the stored images.
[0305] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be configured to: select a first target image having a highest evaluation score from among one or more images included in a set of target images; identify a person object associated with the user from the set of target images based on an intimacy score for a person object included in at least some of the previously stored images; select a second target image having a highest affinity score from among one or more images including a person object associated with the user from the set of target images; and generate a composite image using the first target image and the second target image.
[0306] According to one embodiment, when the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be caused to: determine preference scores of images previously stored in the electronic device (101, 301), determine a target composition based on the preference scores of the previously stored images, and generate a composite image based on the target composition.
[0307] According to one embodiment, the target image set may include a plurality of target image sets. When the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may: generate a plurality of composite images, each corresponding to the plurality of target image sets.
[0308] According to one embodiment, the target image set may include a plurality of target image sets. When the instructions are individually or collectively executed by at least one processor (120, 310), the electronic device (101, 301) may be caused to: generate a panoramic composite image by concatenating a plurality of composite images generated respectively corresponding to the plurality of target image sets.
[0309] According to one embodiment, a non-transitory computer-readable recording medium stores one or more programs including instructions, which, when individually or collectively executed by at least one processor of an electronic device, cause the electronic device to: in response to obtaining a user input, determine a plurality of candidate images from among previously stored images based on the user input and metadata of images previously stored in the electronic device (101, 301); divide the plurality of candidate images into one or more image sets, each image including one or more images satisfying a predetermined condition of mutual similarity; determine an importance score for each of the one or more image sets based on personal data of a user of the electronic device (101, 301), the plurality of candidate images, and metadata of the plurality of candidate images; determine a target image set satisfying a predetermined condition from among the one or more image sets based on the importance score; determine an evaluation score and a preference score for each of the one or more images included in the target image set based on personal data of the user, the target image set, and metadata of the target image set; And an operation of generating a synthetic image using one or more images included in the target image set based on the evaluation score and preference score can be performed.
[0310] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned will be clearly understood by a person having ordinary skill in the art to which the present disclosure pertains.
[0311] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0312] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0313] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0314] Various embodiments of the present document may be implemented as software (e.g., a program (140)) including one or more instructions stored in a storage medium (e.g., an internal memory (136) or an external memory (138)) readable by a machine (e.g., an electronic device (101)). For example, a processor (e.g., a processor (120)) of the machine (e.g., an electronic device (101)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0315] According to one embodiment, the method according to various embodiments disclosed in this document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0316] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separately arranged in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0317] The embodiments described above may be implemented using hardware components, software components, and / or a combination of hardware components and software components. For example, the devices, methods, and components described in the embodiments may be implemented using a general-purpose computer or a special-purpose computer, such as, for example, a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing unit may execute an operating system (OS) and software applications running on the operating system. Furthermore, the processing unit may access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing unit is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing unit may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.
[0318] Software may include a computer program, code, instructions, or a combination of one or more of these, and may configure a processing device to perform a desired operation or, independently or collectively, command the processing device. The software and / or data may be permanently or temporarily embodied in any type of machine, component, physical device, virtual equipment, computer storage medium or device, or transmitted signal wave, for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on a computer-readable recording medium.
[0319] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. The computer-readable medium may include program commands, data files, data structures, etc., alone or in combination, and the program commands recorded on the medium may be those specially designed and configured for the embodiment or may be known and available to those skilled in the art of computer software. Examples of the computer-readable recording medium include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and hardware devices specially configured to store and execute program commands such as ROMs, RAMs, and flash memories. Examples of program commands include not only machine language codes such as those generated by a compiler, but also high-level language codes that can be executed by a computer using an interpreter, etc.
[0320] The hardware device described above may be configured to operate as one or more software modules to perform the operations of the embodiment, and vice versa.
[0321] Although the embodiments have been described with limited drawings, those skilled in the art will appreciate that various technical modifications and variations can be applied based on the described embodiments. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.
[0322] Therefore, other implementations, embodiments and equivalents to the claims also fall within the scope of the claims described below.
Claims
1. In the electronic device (101, 301), At least one processor (120, 310) comprising processing circuitry; and A memory (120, 320) comprising one or more storage media for storing instructions, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: In response to obtaining a user input, a plurality of candidate images are determined among the pre-stored images based on the user input and metadata of the images pre-stored in the electronic device (101, 301), Divide the above plurality of candidate images into one or more image sets, each of which includes one or more images that satisfy a condition of mutual similarity, Determine an importance score for each of the one or more image sets based on personal data of a user of the electronic device (101, 301), the plurality of candidate images, and metadata of the plurality of candidate images; Based on the importance score, a target image set satisfying a predetermined condition is determined among the one or more image sets, Determine an evaluation score and a preference score for each of one or more images included in the target image set based on the personal data of the user, the target image set, and the metadata of the target image set; Generating a synthetic image using one or more images included in the target image set based on the evaluation score and the preference score To do, Electronic devices (101, 301).
2. In paragraph 1, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Determine the importance score of the first image set based on whether each of one or more images included in the first image set among the one or more image sets satisfies a defined importance condition; The above-mentioned importance conditions are: Including at least one of the following conditions: a favorite display condition, a condition associated with a social media application, or a condition of being one of multiple images acquired within a certain area; Electronic devices (101, 301).
3. In either of paragraphs 1 and 2, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Determine the evaluation score of the first image based on whether the first image included in the target image set satisfies a predetermined evaluation condition, Determine the preference score of the first image based on whether the first image included in the target image set satisfies a predetermined preference condition. Electronic devices (101, 301).
4. In any one of paragraphs 1 to 3, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Selecting a first target image having the highest evaluation score among the one or more images included in the target image set; Selecting a second target image having the highest preference score among the one or more images included in the target image set; Generating the composite image using the first target image and the second target image To do, Electronic devices (101, 301).
5. In any one of paragraphs 1 to 4, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Generate a prompt based on the first target image with the highest evaluation score in the target image set and the second target image with the highest preference score in the target image set, Using the vision model, generate the synthetic image based on the prompt. To do, Electronic devices (101, 301).
6. In any one of paragraphs 1 to 5, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Identifying a person object contained in at least some of the above stored images, Determining an intimacy score for the identified person object based on the personal data of the user and the previously stored images To do, Electronic devices (101, 301).
7. In any one of paragraphs 1 to 6, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Selecting a first target image having the highest evaluation score among the one or more images included in the target image set; Identifying a person object associated with the user in the target image set based on an intimacy score for the person object included in at least some of the stored images; Selecting a second target image having the highest preference score among one or more images containing the person object associated with the user from the set of target images; Generating the composite image using the first target image and the second target image To do, Electronic devices (101, 301).
8. In any one of paragraphs 1 to 7, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Determine the preference scores of images previously stored in the above electronic device (101, 301), Determine the target composition based on the preference scores of the above-mentioned stored images, Generate the composite image based on the target composition above. To do, Electronic devices (101, 301).
9. In any one of paragraphs 1 to 8, The above target image set includes a plurality of target image sets, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: Generating multiple synthetic images corresponding to each of the multiple target image sets above. To do, Electronic devices (101, 301).
10. In any one of paragraphs 1 to 9, The above target image set includes a plurality of target image sets, When the above instructions are individually or collectively executed by the at least one processor (120, 310), the electronic device (101, 301) causes: A panoramic composite image is created by connecting multiple composite images generated corresponding to each of the multiple target image sets. To do, Electronic devices (101, 301).
11. In a method performed by an electronic device (101, 301), In response to obtaining a user input, an operation of determining a plurality of candidate images among the pre-stored images based on the user input and metadata of the images pre-stored in the electronic device (101, 301); An operation of dividing the above plurality of candidate images into one or more image sets, each of which includes one or more images that satisfy a condition of mutual similarity; An operation of determining an importance score for each of the one or more image sets based on personal data of a user of the electronic device (101, 301), the plurality of candidate images, and metadata of the plurality of candidate images; An operation of determining a target image set that satisfies a predetermined condition among the one or more image sets based on the importance score; An operation of determining an evaluation score and a preference score for each of one or more images included in the target image set based on the personal data of the user, the target image set, and metadata of the target image set; and An operation of generating a synthetic image using one or more images included in the target image set based on the evaluation score and the preference score. including, method.
12. In paragraph 11, An operation of generating the composite image using the one or more images included in the target image set based on the evaluation score and the preference score is as follows: An operation of selecting a first target image having the highest evaluation score among the one or more images included in the target image set; An operation of selecting a second target image having the highest preference score among the one or more images included in the target image set; and An operation of generating the composite image using the first target image and the second target image. including, method.
13. In any one of paragraphs 11 and 12, The operation of generating the above composite image is: An operation of generating a prompt based on a first target image having the highest evaluation score in the target image set and a second target image having the highest preference score in the target image set; and An operation of generating the synthetic image based on the prompt using the vision model. including, method.
14. In any one of paragraphs 11 to 13, The above target image set includes a plurality of target image sets, The operation of generating the above composite image is: An operation of generating a plurality of synthetic images corresponding to each of the plurality of target image sets. Includes, An operation of creating a panoramic composite image by connecting the above multiple composite images. including more, method.
15. A non-transitory computer-readable recording medium storing a program for executing a method according to any one of claims 11 to 14 in combination with hardware.
Citation Information
Patent Citations
Deodorizing filter coating composition for air cleaners and deodorizing filter using the same
KR1020210147750A
High-speed Wireless transmitting system with 60GHz frequency band using drones
KR1020240151404A
Sintering furnace for electrode active materail
KR1020250001321A
Operation part for surgical instrument and surgical instrument for electrocautery equipped with the operation part
KR1020250157048A
KR20200001034A