Electronic device, and method for generating text describing image of electronic device

WO2026205733A1PCT designated stage Publication Date: 2026-10-01SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2026/001512
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2025-05-27
Filing Date
2026-01-26
Publication Date
2026-10-01

Smart Images

  • Figure KR2026001512_01102026_PF_FP_ABST
    Figure KR2026001512_01102026_PF_FP_ABST
Patent Text Reader

Abstract

In an electronic device and an operating method of the electronic device, according to one embodiment, the electronic device may comprise at least one processor, and a memory for storing at least one computer program including instructions. When executed individually or collectively by the at least one processor, the at least one computer program can cause the electronic device to: generate a first text describing a first image; identify a first word included in the first text describing the first image; identify a second image related to the first word from among images related to a user of the electronic device; generate, on the basis of the second image, a second text including information of the user included in the second image; and output all of the first image, the second image and the second text.
Need to check novelty before this filing date? Find Prior Art

Description

How to generate text describing electronic devices and images of electronic devices

[0001] The present disclosure relates to an electronic device and a method of operating an electronic device, and to a technique for generating text describing an image of an electronic device and an electronic device.

[0002] An electronic device can generate text describing an image based on the image. For example, the electronic device may input an image and a command instructing the generation of text describing the image into an artificial intelligence model, and generate text describing the image using the result output by the artificial intelligence model. The text describing the image may include information indicating features of at least some regions included in the image.

[0003] The information described above may be provided as related art for the purpose of aiding understanding of this document. None of the above is to be claimed as prior art related to this document, nor can it be used to determine prior art.

[0004] When an electronic device generates text describing an image, it may not reflect the user's information. When an electronic device generates text that does not reflect the user's information, it may generate text that does not correspond to the user's intent.

[0005] The technical problems to be solved in this document are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art to which this invention belongs from the description below.

[0006] In an electronic device according to one embodiment, the electronic device may include at least one processor. The electronic device may include a memory that stores at least one computer program containing instructions. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may generate a first text describing a first image. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may identify a first word included in the first text describing the first image. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may identify a second image related to the first word among images related to the user of the electronic device. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may generate a second text containing user information included in the second image based on the second image. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may output all of the first image, the second image, and the second text.

[0007] In a method of operating an electronic device according to one embodiment, the method of operating the electronic device may include an operation of generating a first text describing a first image. The method of operation may include an operation of verifying a first word included in the first text describing the first image. The method of operation may include an operation of verifying a second image related to the first word among images related to a user of the electronic device. The method of operation may include an operation of generating a second text including user information included in the second image based on the second image. The method of operation may include an operation of outputting the first image, the second image, and the second text.

[0008] When an electronic device receives user input instructing it to output text describing a first image, it may generate a second text containing user information based on a second image containing metadata containing user information. By including user information in the second text describing the first image, the electronic device may improve the user's understanding of the second text describing the first image.

[0009] The electronic device may provide the user with the second image used to generate the second text by outputting the second image together with the second text and the first image. By providing the user with the second image used to generate the second text, the electronic device may improve the user's understanding of the second text.

[0010] The effects obtainable from the present disclosure are not limited to those mentioned above, and other unmentioned effects will be clearly understood by those skilled in the art to which the present disclosure belongs from the description below.

[0011] FIG. 1 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.

[0012] Figure 2 is a schematic diagram of an exemplary artificial intelligence system.

[0013] FIG. 3 is a block diagram of an electronic device according to one embodiment.

[0014] FIG. 4 is a drawing illustrating an image input to an electronic device and text generated by the electronic device according to one embodiment.

[0015] FIG. 5 is a block diagram of an electronic device according to one embodiment.

[0016] FIG. 6 is a diagram illustrating the operation of an electronic device for verifying a word included in a first text describing a first image according to one embodiment.

[0017] FIG. 7 is a diagram illustrating the operation of an electronic device that verifies a second image based on a first word according to one embodiment.

[0018] FIG. 8 is a diagram illustrating an operation to generate a second text containing user information related to a first word based on a second image of an electronic device according to one embodiment.

[0019] FIGS. 9a and 9b are drawings for explaining the operation of outputting a second text of an electronic device according to one embodiment.

[0020] FIG. 10 is a diagram illustrating the operation of outputting a second text of an electronic device according to one embodiment.

[0021] FIG. 11 is a flowchart of the operation of an electronic device according to one embodiment.

[0022] FIG. 1 is a block diagram of an exemplary electronic device (100) capable of performing the operations described in this document.

[0023] Referring to FIG. 1, the electronic device (100) may be one of various forms of electronic devices, such as a notebook (190), smartphones (191) having various form factors (e.g., a bar-type smartphone (191-1), a foldable-type smartphone (191-2), or a sliderable (or rollable)-type smartphone (191-3)), a tablet (192), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 1 are illustrative only and are not intended to limit the implementations described or claimed herein. The electronic device (100) may be referred to as a mobile device, a user device, a multifunction device, a portable device, or a server.

[0024] The electronic device (100) may include components comprising at least one processor (110) (hereinafter referred to as processor (110)), at least one memory (120) (hereinafter referred to as memory (120)), at least one display (140) (hereinafter referred to as display (140)), at least one image sensor (150) (hereinafter referred to as image sensor (150)), at least one communication circuit (160) (hereinafter referred to as communication circuit (160)), and / or at least one sensor (170) (hereinafter referred to as sensor (170)). The components are merely exemplary. For example, the electronic device (100) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuit, antenna, rechargeable battery, or input / output interface). For example, some components may be omitted from the electronic device (100). For example, some components may be integrated into a single component.

[0025] The processor (110) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing operations. The processor (110) may include at least one electrical circuit and may process instructions (or programs, data, etc.) stored in memory (120) individually or collectively in a distributed manner. The processor (110) may include a processor assembly comprising one or more processing circuits. The processor (110) may include any processing circuit that is operative to control the performance and operations of one or more components of the electronic device (100) (e.g., memory (120), display (140), image sensor (150), communication circuit (160), and / or sensor (170)). For example, the processor (110) (e.g., application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (110) may be implemented with a plurality of cores (or at least one core circuit), a plurality of chips, or a plurality of chipsets. For example, the processor (110) may include one or more processing circuits. For example, the processor (110) may include one or more processing circuits configured to perform the various functions of the present disclosure individually and / or collectively. As an example without limitation, at least a portion of the processor (110) may be included in a first chip of the electronic device (100), and at least another portion of the processor (110) may be included in a second chip of the electronic device (100) different from the first chip of the electronic device (100).

[0026] For example, the processor (110) may include a central processing unit (111), a graphics processing unit (112), a neural processing unit (113), an image signal processor (114), a display controller (115), a memory controller (116), a storage controller (117), a communication processor (118), and / or a sensor interface (119). These components of the processor (110) are merely exemplary. For example, the processor (110) may include other components. For example, some components of the processor (110) may be omitted from the processor (110). For example, some components of the processor (110) may be included as separate components of the electronic device (100) outside of the processor (110). For example, some components of the processor (110) (e.g., memory controller (116)) may be included in other components (e.g., at least part of memory (120), an interface (e.g. available for connection to at least one component of the electronic device (100)), a display (140) and / or an image sensor (150)).

[0027] The processor (110) may cause other components of the electronic device (100) to perform various operations by executing instructions stored in memory (120). The CPU (111) (or central processing circuit) may be configured to control the components of the processor (110) based on the execution of instructions stored in memory (120) (e.g., volatile memory (121) and / or non-volatile memory (122)). The GPU (112) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (113) (or neural processing circuit, or AI (artificial intelligence) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). An ISP (114) (or image signal processing circuit) may be configured to process a raw image acquired through an image sensor (150) into a format suitable for a component within the electronic device (100) or a component of the processor (110). A display controller (115) (or display control circuit, or DPU (display processing unit)) may be configured to process an image acquired from a CPU (111), GPU (112), ISP (114), or memory (120) (e.g., volatile memory (121)) into a format suitable for a display (140). A memory controller (116) (or memory control circuit) may be configured to control reading data from the volatile memory (121) and writing data to the volatile memory (121). A storage controller (117) (or storage control circuit) may be configured to control reading data from the non-volatile memory (122) and writing data to the non-volatile memory (122).The CP (118) (communication processing circuit) may be configured to process data obtained from a component of the processor (110) into a format suitable for transmitting to another electronic device via the communication circuit (160), or to process data obtained from another electronic device via the communication circuit (160) into a format suitable for processing by the component of the processor (110). For example, the communication circuit (160) may include one or more communication circuits. The sensor interface (119) (or sensing data processing circuit, sensor hub) may be configured to process data regarding the state of the electronic device (100) and / or the state around the electronic device (100), obtained through the sensor (170), into a format suitable for the component of the processor (110).

[0028] Memory (120) may include one or more storage media (or one or more storage devices). For example, memory (120) may include a memory assembly comprising one or more storage media. For example, the one or more storage media may include a hard drive, a permanent memory such as flash memory, read-only memory (ROM) (e.g., non-volatile memory (122)), a semi-permanent memory such as random access memory (RAM) (e.g., volatile memory (121)), any other suitable type of storage (or storage assembly), or any combination thereof. Memory (120) may include a cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (100). As an example not limited to, the cache memory may be included within the processor (110). The memory (120) may be fixedly embedded within the electronic device (100) or incorporated into one or more suitable types of components (e.g., a SIM (subscriber identity module) card and / or an SD (secure digital) card) that can be repeatedly inserted into and removed from the electronic device (100).

[0029] For example, memory (120) may store one or more software applications, such as operating system (or system) software applications, firmware software applications, driver software applications, plugin (e.g., add-in, add-on, and / or applet) software applications, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (110). For example, memory (120) may store instructions that can be called by an application programming interface (API). For example, memory (120) may store instructions within a library.

[0030] Figure 2 is a schematic diagram of an exemplary artificial intelligence system.

[0031] Referring to FIG. 2, an artificial intelligence system according to one embodiment may include a user interface (210), a database (250), an application and service component (260), an AI framework (220), and a generative AI model (230).

[0032] A user interface (210) may receive a user query. The input may include user input and / or data obtained or generated by an electronic device (e.g., the electronic device (100) described above, the electronic device (300) of FIG. 3, or the electronic device (500) of FIG. 5). The data may include images, videos, and / or sensor data generated by at least one processor of the electronic device (e.g., at least one processor (110) or the processor (510) of FIG. 5) (e.g., illuminance data around the electronic device obtained from a sensor (170) or sensor hub, attitude data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., temperature of the display (140) or temperature of at least one processor (110)), size information of the display area of ​​the display (140), and / or images obtained through an image sensor (150) of the electronic device. For example, the user query may be in the form of natural language, touch data obtained through a touch circuit included in the display (140) (e.g., used to identify input from a finger and / or stylus), an image, audio, and / or video. Additionally, context information may be transmitted along with the user query. The context information may include various side information related to the time when the user query is input into the artificial intelligence system. Examples include application information currently being used by the user or location information of the user. As another example, the user query may also be a non-natural language input that does not generate natural language, such as a design request or modification.In addition, a mixed form of the natural language, images, sounds, and context information described above is also possible. Furthermore, the user interface (210) can output results of the artificial intelligence system to the user. The output may include results (or result information) generated or obtained by the artificial intelligence system based on at least part of the input. The output may be in the form of natural language or specific content, and may also be provided in a form such as an action requested by the user. For example, the output may have a format according to the user settings of the electronic device.

[0033] The AI ​​framework (220) can receive a user query and coordinate and control each component necessary to perform the user's intent. The AI ​​framework (220) may include a prompt design component (221), an APIs / Plugins Management component (223), and an output modification component (225).

[0034] User queries or actions entered in the user interface (210) can be transmitted to a prompt design component (221). The prompt design component (221) can be used to generate prompts suitable for input into a large language model (LLM), a large vision model (LVM), or large multimodal models. The prompt design component (221) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (221) can generate prompts by accessing a database (250) (e.g., a knowledge component) containing user preference data, a prompt library, and prompt examples, and can transmit them to the large language model (LLM), large vision model (LVM), and / or large multimodal model (LMM).

[0035] The application and plugin management component (223) can perform the role of communicating with external information when there is a request for additional information when user input is transmitted as input to the generative AI model (230). The application and plugin management component (223) establishes a channel to communicate with the outside of the artificial intelligence system through an application programming interface (API), thereby enabling access to various data sources. For example, the application and plugin management component (223) can be used to request other components (e.g., application and service components (260)) that perform feedback (or response) according to the prompt. The acquired information can be used to generate a prompt by the prompt design component (221) together with the user input, or can be used as input to the generative AI model (230). Additionally, the application and plugin management component (223) can request the action through the API if the application or service needs to perform an action that ultimately executes the user query rather than an intermediate result. Information obtained from an external source can be transmitted as input to a generative AI model (230) along with user input.

[0036] The output modification component (225) can finely tune (or adjust) (or change) the output output from the generative AI model (230). For example, the output modification component (225) can determine the relevance (e.g., score) between the output (e.g., content) of the generative AI model and the user input. For example, the output modification component (225) can verify whether the content generated through the Large Language Model (LLM), Large Vision Model (LVM), or Large Multimodal Model (LMM) contains the aforementioned relevance, biased information (e.g., selective information), or harmful information (e.g., violent content or profanity). Additionally, the output modification component (225) can determine the extent to which the output matches the desired result and, if additional processing is required, proceed with that process. Additionally, the output modification component (225) can configure and provide to the user a hint to avoid unwanted output.

[0037] A generative AI model (230) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (230) may include an image generation model and / or a language generation model. An image generation model may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). An example of an image generation model is a diffusion-based generative model that uses the structures of a VAE and a transformer. Additionally, a language generation model is a model trained to output the statistically most appropriate output value based on input values, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. It may also include large multimodal models (LMMs) capable of recognizing various forms of data input, such as text, images, and voice, and generating new data corresponding to them.

[0038] In one embodiment, the AI ​​framework (220) and / or generative artificial intelligence model (230) may be included within an AI module (e.g., including a processing circuit) (e.g., the AI ​​module (323) of FIG. 3) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor of the electronic device (e.g., at least one processor (110) or processor (510)). For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0039] FIG. 3 is a block diagram of an electronic device according to one embodiment.

[0040] An electronic device (300) (e.g., the electronic device (100) of FIG. 1) may include an application layer (310), a framework layer (320), and / or a hardware layer (330).

[0041] The application layer (310) may include at least one component that performs a function according to the control of the user of the electronic device (300). For example, the application layer (310) may include a video output module (311), a voice output module (312), and / or an image output module (313).

[0042] The video output module (311) can output a video stored in the memory (332) of the electronic device (300) or on an external server. For example, the video output module (311) can convert a specific area (e.g., the area representing the object represented by the first word described in FIG. 9) included in an image (e.g., the first image (610) described in FIG. 6) output by the electronic device (300) into a dynamic form while executing a specific function (e.g., image to text).

[0043] The voice output module (312) can output sound stored in the memory (332) of the electronic device (300) or on an external server. The voice output module (312) can output sound generated by the electronic device (300) as well as sound generated by the electronic device (300) (e.g., sound generated based on text generated by the image to text function (e.g., the second text (801) described in FIG. 8).

[0044] The image output module (313) can output an image stored in the memory (332) of the electronic device (300) or on an external server (e.g., an image related to the user described in FIG. 7). The image output module (313) can provide various functions (e.g., a zoom function) for converting the output image (e.g., a first image (610) described in FIG. 6 or a second image (720, 740) described in FIG. 7). The image output module (313) can manage (e.g., create and edit) metadata included in the image (e.g., metadata described in FIG. 5). According to one embodiment, the image output module (313) can execute specific functions related to the image (e.g., image to text, image to speech).

[0045] The framework layer (320) may include at least one component connecting the application layer (310) and the hardware layer (330). For example, the framework layer (320) may include a pen module (321), a capture module (322), and / or an AI module (323). Some modules included in the framework layer (320) (e.g., the pen module (321)) may be omitted.

[0046] The pen module (321) can process user input received by a stylus pen. The pen module (321) can identify a predetermined gesture indicated by the user input received by the stylus pen and transmit the predetermined gesture to the application layer (310) or the hardware layer (330).

[0047] The AI ​​module (323) may include an image generation module, a large vision model (LVM) and / or a large language model (LLM). The image generation module may generate various images based on a specific image (e.g., an image stored in an electronic device (300) or an image included in an internet search result) (e.g., a first image (610) described in FIG. 6, a second image (720, 740) described in FIG. 7). The image generation module may crop or resize at least a portion of the specific image and apply graphic effects to at least a portion of the specific image (e.g., a first image (610) described in FIG. 6, a second image (720, 740) described in FIG. 7). LVM may include an artificial intelligence model (e.g., the generative artificial intelligence model (230) of FIG. 2) configured to analyze specific images (e.g., the first image (610) described in FIG. 6, the second images (720, 740) described in FIG. 7). LVM can generate information describing a specific image based on the results of analyzing a specific image (e.g., the result of executing an image-to-text or image-to-speech function (e.g., a first text (620) describing the first image in FIG. 6, text describing the second image (720, 740) described in FIG. 8, and a second text (801)). LLM may include an artificial intelligence model (e.g., the generative artificial intelligence model (230) of FIG. 2) configured to analyze specific text (e.g., the result of executing an image-to-text or image-to-speech function (e.g., a first text (620) describing the first image in FIG. 6, text describing the second image (720, 740) described in FIG. 8, and a second text (801))). For example, LLM can change the style of specific text, summarize specific text, or identify the sentiment indicated by specific text.

[0048] The hardware layer (330) may include at least one component constituting the hardware of the electronic device (300). The hardware layer (330) may include a processor (331) (e.g., processor (110) of FIG. 1, processor (510) described in FIG. 5)) and / or memory (332) (memory (120) of FIG. 1, memory (520) described in FIG. 5). The components included in the hardware layer (330) will be described in detail in FIG. 5.

[0049] FIG. 4 is a drawing illustrating an image input to an electronic device and text generated by the electronic device according to one embodiment.

[0050] An electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (300) of FIG. 3) can generate text (420) describing an image based on an image (410). For example, the electronic device (300) can input a command (or prompt) instructing the generation of text (420) describing an image and the image (410) into a generative artificial intelligence model (230), and generate text (420) describing an image using the result output by the generative artificial intelligence model (230).

[0051] The electronic device (300) can generate text (420) describing an image based on objects included in the image (410) (e.g., airplane, sea, sandy beach, sky, tourist, resort) (or areas representing objects included in the image (410)). For example, the electronic device (300) can identify objects included in the image (410) and generate information representing objects included in the image (410). The electronic device (300) can generate text (420) describing the image by combining information representing objects included in the image (410).

[0052] Information representing an object included in the image (410) may include information representing the characteristics of the object included in the image (410) (e.g., color, number, location, pose, size). For example, information representing a sandy beach may include information representing the color of the sandy beach (e.g., white). Information representing the sea may include information representing the color of the sea (e.g., turquoise). Information representing tourists may include information representing the number of tourists (e.g., many). Information representing a resort may include information representing the relative location of the resort (e.g., behind the beach).

[0053] The electronic device (300) can generate text (420) describing the image based on the image (410), thereby allowing a user who cannot view the image (410) to view information representing the image (410) in the form of text.

[0054] When the electronic device (300) generates text (420) describing an image that does not reflect the user's information (or information associated with the user) of the electronic device (300), it may also generate text describing an image that does not match the user's intention. The electronic device (300) described in the drawings below (e.g., the electronic device (500) of FIG. 5) may provide the user with text that reflects the user's information (or information associated with the user) by generating text (e.g., the second text (801) of FIG. 8) describing an image (410) (e.g., the first image (610) described in FIG. 6) that includes the user's information (or information associated with the user).

[0055] FIG. 5 is a block diagram of an electronic device according to one embodiment.

[0056] An electronic device (500) (e.g., electronic device (100) of FIG. 1, electronic device (300) of FIG. 3) may include a processor (510) (e.g., processor (110) of FIG. 1, processor (310) of FIG. 3), memory (520) (e.g., memory (120) of FIG. 1, memory (320) of FIG. 3)) and / or a display (530) (e.g., display (140) of FIG. 1). Various embodiments of this document may be implemented even if some of the illustrated configurations are omitted or replaced with other configurations. In addition to the illustrated configurations, the electronic device (500) may further include at least some of the configurations and / or functions of the electronic device (100) of FIG. 1 and the electronic device (300) of FIG. 3. At least some of the components of each electronic device (100, 300) shown (or not shown) may be operatively, functionally, and / or electrically connected.

[0057] The processor (510) may include at least one processing circuitry, and the processor (510) may include at least one processor. The operations of the electronic device (500) or the processor (510) described in FIGS. 1 to 10 may be performed individually or collectively by at least one processor included in the processor (510).

[0058] The memory (520) can store at least one computer program, and at least one computer program may include instructions that can be executed by a processor (510). The operation of the electronic device (500) or processor (510) described in FIGS. 1 to 10 can be performed according to the execution of instructions contained in the memory (420).

[0059] The processor (510) may receive user input instructing it to generate text describing a first image. The first image may include an image stored in memory (520) and / or an image included in the result of execution of a function executable by an electronic device (e.g., internet search). The processor (510) may display an area (or graphic object, button) for receiving user input instructing it to generate the first image and / or text describing the first image while running various applications (e.g., internet browser, gallery application) capable of providing the first image (or enabling the display of the first image). The processor (510) may receive user input while displaying the first image.

[0060] The processor (510) can perform the operation of generating a first text based on user input instructing to generate text describing a first image. The processor (510) can generate a first text describing a first image based on the first image. For example, the processor (510) can input a command (or prompt) instructing to generate text describing a first image and the first image into a generative artificial intelligence model (230), and generate a first text describing the first image using the result output by the generative artificial intelligence model (230).

[0061] According to one embodiment, the processor (510) may generate text describing the first image based on the first image and generate a first text describing the first image by summarizing the text describing the first image. The processor (510) may input a command (or prompt) instructing to summarize the text describing the first image and the text describing the first image into a generative artificial intelligence model (230) (e.g., an LLM included in the AI ​​module (323) of FIG. 2) and generate a first text summarizing the text describing the first image using the result output by the generative artificial intelligence model (230).

[0062] The first text is text that is output by applying the first image to the generative artificial intelligence model (230), and can be generated based on data used to train the generative artificial intelligence model (230). According to one example, the generative artificial intelligence model (230) can be trained based on data related to various users, rather than data related to the user of the electronic device (500), and the first text may not contain information related to the user of the electronic device (500).

[0063] The first text may be generated based on data (or information) included in a command (or prompt) that instructs the generation of the first text. According to one example, a generative artificial intelligence model (230) may be configured to generate text containing data (or information) included in the command (or prompt). The generative artificial intelligence model (230) may generate a first text that does not contain information associated with the user based on a command (or prompt) that does not contain information associated with the user.

[0064] A processor (510) can identify at least one first word included in a first text. For example, the processor (510) can control a generative artificial intelligence model (230) (e.g., an LLM included in the AI ​​module (323) of FIG. 2) to identify at least one first word among the words included in the first text that satisfies a predetermined condition. The processor (510) can input a command (or prompt) instructing the generative artificial intelligence model (230) to identify the first text and at least one first word satisfying a predetermined condition, and can identify at least one first word output by the generative artificial intelligence model (230). The predetermined condition may include a condition that the word among the words included in the first text has the highest importance (or has an importance level greater than or equal to a threshold). According to one embodiment, the first word may be composed of a plurality of words.

[0065] The first word may refer to a word that has higher importance than words other than the first word included in the first text. Higher importance may include having a larger area representing the object represented by the word included in the first image. For example, the first word may have a larger area representing the object represented by the first word than the area representing the object represented by words other than the first word. Alternatively, higher importance may include having more information included in the metadata of the first image. For example, the first word may include more information included in the metadata of the first image than words other than the first word.

[0066] The processor (510) can determine (or verify) a second image associated with a first word. For example, the processor (510) can verify a second image associated with a first word among images associated with a user of the electronic device (500).

[0067] An image associated with a user of the electronic device (500) may include an image stored in a memory (420) included in the electronic device (500) and / or an image with a history confirmed by a user of the electronic device (500). An image with a history confirmed by a user of the electronic device (500) may include an image with a history processed by various functions executable by the electronic device (500) (e.g., content output, internet search, message transmission and reception) (or an image with a history displayed on a display (140) included in the electronic device (500)).

[0068] Each image associated with a user of the electronic device (500) may include metadata of the image associated with the user of the electronic device (500). The metadata may include information of the user (or information associated with the user) represented by the image associated with the user of the electronic device (500) (e.g., information representing an object included in the image, information representing the time (and date) when the image was taken, information representing the location where the image was taken, and / or information representing an event associated with the image). The object included in the image may include objects and / or people. The information representing a person included in the image may include information representing the relationship between the person included in the image and the user (e.g., friends, family). The event associated with the image may include the user's schedule related to the time and / or location when the image was taken. For example, when the processor (510) acquires an image, it may store information representing an event associated with the image as metadata based on information representing the user's schedule stored in the electronic device (500) (or memory (420)).

[0069] The images of the present disclosure (e.g., an image related to a user, a first image, a second image) may be a main image included in an image file for each of the images. Each image file may include a main image, an image (or at least a part of an image) set with a setting value different from that of the main image (e.g., a setting value for resolution, illuminance, saturation, or color information), and / or metadata.

[0070] According to one embodiment, the second image associated with the first word may refer to an image including an area representing an object substantially identical to the object indicated by the first word.

[0071] The processor (510) can identify a second image among images related to a user of the electronic device (500) that includes a first word (or at least a part of the first word) in the metadata. Including the first word (or at least a part of the first word) in the metadata may include including in the metadata a word substantially identical to the first word (e.g., a word representing an object substantially identical to the object represented by the first word) or information having a correlation with the first word greater than or equal to a threshold value.

[0072] According to one embodiment, the processor (510) may determine a second image using a personal data core (PDC) included in the electronic device (500) that stores and retrieves images related to the user. The personal data core may store the user's personal data (e.g., the user's schedule, the user's events and / or the user's images). The personal data core may retrieve personal data (e.g., a second image) that has a relevance to a specific word (e.g., a first word) greater than or equal to a threshold value.

[0073] According to one embodiment, if the processor (510) does not identify information among images related to the user of the electronic device (500) where the relevance is greater than (or exceeds) a threshold, it may input a first word to an external server and verify a second image received from the external server. For example, the processor (510) may determine (or verify) an image included in the internet search results of the first word as the second image related to the first word.

[0074] According to one embodiment, the processor (510) can identify a second image associated with a specific first word based on a specific first word included in a first text describing a first image. The processor (510) can identify a second image that includes at least a portion of the specific first word as metadata among images stored in memory (420) (or images with a history of being identified by a user). For example, the processor (510) can determine the image with the highest correlation to the specific first word as the second image based on the metadata of an image associated with a user of the electronic device (500). If there are multiple images that include at least a portion of the specific first word as metadata, the processor (510) can determine the image that includes the largest portion of the specific first word as metadata as the second image associated with the first word. Alternatively, if there are multiple images that include at least a portion of the specific first word as metadata, the processor (510) can determine the image most recently identified (or captured) by the user as the second image associated with the first word.

[0075] According to one embodiment, the processor (510) can identify a second image associated with a specific first word based on a specific first word included in a first text describing a first image.

[0076] According to one embodiment, if an image containing information in the metadata that has a relevance to a specific first word greater than or equal to a threshold value is not found, the processor (510) inputs the specific first word to an external server and can check a second image received from the external server. For example, the processor (510) may determine (or confirm) an image included in the internet search results of the specific first word as a second image related to the first word.

[0077] If the processor (510) identifies (or determines) a second image associated with a first word, it may generate text describing the second image. The processor (510) may generate text describing the second image based on an object included in the second image. If there are multiple first words included in the first text describing the first image, the processor (510) may generate text describing each second image associated with each first word.

[0078] According to one embodiment, the processor (510) may generate text describing the second image using metadata of the second image. For example, the processor (510) may input a command (or prompt) instructing to generate text describing the second image, the second image and / or metadata of the second image into a generative artificial intelligence model (230), and generate text describing the second image using the result output by the generative artificial intelligence model (230). The processor (510) may generate text describing the second image, which includes at least some of information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event related to the image.

[0079] The processor (510) can control the generative artificial intelligence model (230) to identify at least one word among the words included in the text describing the second image that satisfies a predetermined condition. The processor (510) can input the text describing the second image into the generative artificial intelligence model (230) and identify at least one word output by the generative artificial intelligence model (230). The predetermined condition may include a condition that the word among the words included in the second text has the highest importance (or has an importance level greater than or equal to a threshold). The at least one word satisfying the predetermined condition may refer to a word that has a higher importance than words other than the second word included in the second text. Having a higher importance may include having a larger area of ​​the object represented by the word included in the second image. Having a higher importance may include having more information included in the metadata of the second image.

[0080] The processor (510) can identify a second word associated with a first word (or a second word corresponding to a first word) among at least one word satisfying a predetermined condition included in the text describing the second image. The second word associated with the first word may include a word representing an object substantially identical to the object represented by the first word.

[0081] The processor (510) can identify an area in the second image that represents an object represented by the second word and / or an area in the first image that represents an object represented by the first word. The area in the second image that represents an object represented by the second word and the area in the first image that represents an object represented by the first word may represent substantially the same object.

[0082] The processor (510) can generate text that complements a first text describing a first image. The text that complements the first text may include text that complements a first word included in the first text.

[0083] The processor (510) can generate text that complements the first word based on user information (or information associated with the user) included in the metadata of the second image associated with the first word. The metadata of the second image may include user information (or information associated with the user) (e.g., information indicating an object included in the second image, information indicating the time (and date) when the second image was taken, information indicating the location where the second image was taken, and / or information indicating an event associated with the second image). The object included in the second image may include objects and / or people. The information indicating a person included in the second image may include information indicating the relationship between the person included in the image and the user (e.g., friends, family). The event associated with the second image may include the user's schedule related to the time and / or location when the image was taken.

[0084] According to one embodiment, text that complements the first word may refer to text that complements (or modifies) the first word to include user information (or information associated with the user). For example, the processor (510) may generate a prompt instructing the generation of text that complements the first word based on the first word, the second word, the first image (or the area representing the object represented by the first word), the second image (or the area representing the object represented by the second word), and / or at least part of the metadata of the second image (e.g., user information). The processor (510) may input the first word, the second word, the first image (or the area representing the object represented by the first word), the second image (or the area representing the object represented by the second word), and / or at least part of the metadata of the second image (e.g., user information), and / or the prompt instructing the generation of text that complements the first word into a generative artificial intelligence model (230), and generate text that complements the first word based on the result output by the generative artificial intelligence model (230).

[0085] According to one embodiment, the processor (510) may generate text that complements the first word based on an area representing an object represented by the second word and an area representing an object represented by the first word. For example, the processor (510) may compare an area representing an object represented by the second word and an area representing an object represented by the first word, and generate text that complements the first word to include the result of the comparison. For example, the processor (510) may generate text that complements the first word to include the result of comparing the characteristics (e.g., color, size, shape, and / or ratio) of each of the area representing the second word included in the second image and the area representing the first word included in the first image.

[0086] The processor (510) may generate a second text containing user information (or information associated with the user) included in a second image (or metadata of the second image) based on text that complements the first word and / or the first text. For example, the processor (510) may generate the second text by replacing the first word included in the first text with text that complements the first word. The processor (510) may generate the second text by inserting (or placing) text that complements the first word in an adjacent part of the first word included in the first text.

[0087] The second text may include text (or at least a part of text) that complements at least one first word. The processor (510) inputs a prompt generated based on an image related to the user (e.g., the second image) and user information (or information associated with the user) included in the metadata of the image related to the user (e.g., the second image) into a generative artificial intelligence model (230), and may generate text that complements the first word based on the result output by the generative artificial intelligence model (230). The generative artificial intelligence model (230) may generate text that complements the first word based on user information included in the prompt. Unlike the first text, the second text may include user information (or information associated with the user).

[0088] The processor (510) can output a second text and a first image containing user information (or information associated with the user) on a display (530) included in the electronic device (500). When the processor (510) outputs the first image and the second text, it can output the second image together.

[0089] When the processor (510) outputs the second text, it may apply a first graphic effect to the text that complements the first word included in the second text. Alternatively, the processor (510) may output the first graphic effect on the area containing the second text. The first graphic effect applied to the text that complements the first word may be an effect that distinguishes the text that complements the first word from the rest included in the second text. For example, the processor (510) may make the text that complements the first word sparkle, change it to a different color, or highlight it (e.g., highlight it in bold).

[0090] When the processor (510) outputs the first image, it may apply a second graphic effect to the area representing the object represented by the first word. The second graphic effect may be an effect that distinguishes the area representing the object represented by the first word from other areas included in the first image. For example, the processor (510) may make the area representing the object represented by the first word sparkle or emphasize the boundary line of the area representing the object represented by the first word. As another example, when the processor (510) outputs the first image, it may separate the area representing the object represented by the first word. The processor (510) may output the separated area on the display as a separate layer from the first image. According to one embodiment, when the processor (510) applies the first graphic effect to text that complements a specific first word, it may also apply the second graphic effect to the area representing the object represented by the specific first word.

[0091] According to one embodiment, the processor (510) may be configured to output both the first image and the second text on the display (530). The processor (510) may be configured to output a second image associated with a specific first word in response to receiving user input selecting an area representing an object represented by a specific first word. The processor (510) may output a second image associated with another first word in response to receiving user input selecting an area representing another first word. When a user of the electronic device (500) selects a specific area included in the first image, the user may see a second image associated with the specific area (or the first word indicated by the specific area).

[0092] According to one embodiment, when the processor (510) outputs the second image, it may output information included in the metadata of the second image based on user input (e.g., a long press) received in the area where the second image is displayed. The metadata of the second image may include user information (or information associated with the user) (e.g., information indicating an object (person) included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event (e.g., a schedule) associated with the image).

[0093] According to one embodiment, the processor (510) may output text describing the second image based on user input (e.g., long press) received in the area where the second image is displayed when the second image is output.

[0094] According to one embodiment, the operation of the electronic device (500) of the present disclosure generating text (e.g., a second text) containing user information (or information associated with the user) describing a first image may be applied in the same way to a second image. For example, a processor (510) may identify a third word among the words included in the text describing the second image, identify a third image associated with the third word, and generate a third text containing user information (or information associated with the user) included in the third image based on the third image. The third text may include text that complements the third word based on the third image.

[0095] When the electronic device (500) (or processor (510)) receives user input instructing it to output text describing a first image, it may generate a second text containing user information (or information associated with the user) based on a second image containing metadata containing user information (or information associated with the user). The electronic device (500) (or processor (510)) may improve the user's understanding of the second text describing the first image by including user information (or information associated with the user) in the second text describing the first image.

[0096] The electronic device (500) (or processor (510)) can provide the user with the second image used to generate the second text by outputting the second image together with the second text and the first image. The electronic device (500) (or processor (510)) can improve the user's understanding of the second text by providing the user with the second image used to generate the second text.

[0097] FIG. 6 is a diagram for explaining the operation of an electronic device that identifies a word included in a first text describing a first image according to one embodiment.

[0098] The electronic device (500) may receive user input instructing it to generate text describing the first image (610). The first image (610) may include an image stored in memory (520) and / or an image included in the result of execution of a function (e.g., internet search) that can be executed by the electronic device (500). The electronic device (500) may display an area (or graphic object, button) for receiving user input instructing it to generate the first image (610) and / or text describing the first image (610) while running various applications (e.g., internet browser, gallery application) that can provide the first image (610) (or enable the display of the first image (610)). The electronic device (500) may receive user input while displaying the first image (610).

[0099] The electronic device (500) can perform the operation of generating a first text (620) based on user input instructing it to generate text describing a first image (610). The electronic device (500) can generate a first text (620) describing a first image based on the first image (610).

[0100] The operation of generating text (420) describing an image based on the image (410) of FIG. 4 can be applied to the operation of generating first text (620) describing a first image based on the first image (610) in FIG. 6.

[0101] According to one embodiment, the electronic device (500) may generate text (not shown) describing the first image based on the first image (610), and may generate a first text (620) describing the first image by summarizing the text (not shown) describing the first image. The electronic device (500) may input a command (or prompt) instructing to summarize the text describing the first image and the text describing the first image into a generative artificial intelligence model (e.g., the generative artificial intelligence model (230) of FIG. 2), and generate a first text (620) summarizing the text (not shown) describing the first image using the result output by the generative artificial intelligence model (230). For example, the electronic device (500) may generate a first text (620) summarizing the text describing the first image of [Table 1] and the text describing the first image of [Table 2].

[0102] This photo captures a stunning scene of the famous Maho Beach on Saint Martin Island. A white Delty Airways passenger plane is flying low just above the beach against a backdrop of blue skies and white clouds. Below, the clear turquoise waters of the Caribbean Sea and white sandy beaches stretch out, with many tourists gathered on the beach to watch the plane land. Behind the beach, modern white resort buildings are visible, and a crane under construction can be seen on one of them. Because this location is very close to Princess Juliana International Airport, it is a world-famous tourist spot where visitors can witness the special sight of planes flying right over the beach. The aircraft appears to be a Delty Airways Boeing 757, featuring the airline's traditional red and blue livery.

[0103] This is a scene of a Delty aircraft landing over Maho Beach in Saint Martin. Many tourists are watching the landing against the backdrop of turquoise waters, white sandy beaches, and a blue sky. Modern resort buildings are visible behind the beach, and this location is a famous tourist spot known for its low-altitude landings due to its proximity to the airport.

[0104] The first text (620) is text output by applying the first image (610) to the generative artificial intelligence model (230), and can be generated based on data used to train the generative artificial intelligence model (230). According to one example, the generative artificial intelligence model (230) may be trained based on data related to various users, rather than data related to the user of the electronic device (500), and the first text (620) may not contain information related to the user of the electronic device (500). The first text (620) may be generated based on data (or information) included in a command (or prompt) that instructs the generation of the first text (620). According to one example, the generative artificial intelligence model (230) may be configured to generate text containing data (or information) included in the command (or prompt). A generative artificial intelligence model (230) can generate a first text (620) that does not contain information associated with the user based on a command (or prompt) that does not contain information associated with the user. An electronic device (500) can identify at least one first word (621, 622, 623, 624) included in the first text (620) (e.g., "Delti aircraft landing," "turquoise sea and white sandy beach and blue sky," "many tourists watching the plane landing," "this place is very close to the airport and is a famous tourist spot where you can see the plane's low landing"). For example, the electronic device (500) can control a generative artificial intelligence model (e.g., the generative artificial intelligence model (230) of FIG. 2) to identify at least one first word (621, 622, 623, 624) among the words included in the first text (620) that satisfies a predetermined condition.The electronic device (500) inputs a command (or prompt) instructing to check a word satisfying a predetermined condition and a first text (620) into a generative artificial intelligence model, and can check at least one first word (621, 622, 623, 624) (e.g., "Delti aircraft landing," "turquoise sea and white sandy beach and blue sky," "many tourists watching the plane landing," "this place is very close to the airport and a famous tourist spot where you can see the plane's low landing") output by the generative artificial intelligence model (230). The predetermined condition may include a condition that the word with the highest importance among the words included in the first text (620) (or whose importance is above a threshold).

[0105] According to one embodiment, the first word (621, 622, 623, 624) may be composed of multiple words. For example, at least one first word (621, 622, 623, 624) (e.g., "Delty aircraft landing," "turquoise sea and white sandy beach and blue sky," "many tourists watching the plane landing," "this place is very close to the airport and is a famous tourist spot where you can see the low landing of the plane") may each be composed of two or more words.

[0106] FIG. 7 is a diagram illustrating the operation of an electronic device that verifies a second image based on a first word according to one embodiment.

[0107] The description of some of the first words (621, 622, 623, 624) in Fig. 7 can be applied equally to the rest.

[0108] The electronic device (500) can determine (or identify) a second image (720, 740) associated with a first word (622, 624). For example, the electronic device (500) can identify a second image (720, 740) associated with a specific first word (622, 624) (e.g., "turquoise sea and white sandy beach and blue sky," "a famous tourist spot very close to the airport where you can see the low landing of airplanes") among images associated with the user of the electronic device (500).

[0109] An image associated with a user of the electronic device (500) may include an image stored in a memory included in the electronic device (500) and / or an image with a history confirmed by a user of the electronic device (500). An image with a history confirmed by a user of the electronic device (500) may include an image with a history processed by various operations of the electronic device (500) (e.g., content output, internet search, message transmission / reception) (or an image with a history displayed on a display included in the electronic device (500)).

[0110] An image associated with a user of the electronic device (500) may include metadata. The metadata may include information about the user (or information associated with the user) (e.g., information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event associated with the image). An object included in the image may include objects and / or people. Information indicating a person included in the image may include information indicating the relationship between the person included in the image and the user (e.g., friends, family). An event associated with the image may include the user's schedule related to the time and / or location when the image was taken. When the electronic device (500) acquires an image, it may store information indicating an event associated with the image as metadata for the image based on information indicating the user's schedule stored in the electronic device (500) (or memory).

[0111] The electronic device (500) may determine an image among images related to the user of the electronic device (500) that includes a first word (622, 624) (or at least a part of the first word) in the metadata as a second image (720, 740). Including the first word (622, 624) (or at least a part of the first word) in the metadata may include including in the metadata information that is substantially identical to the first word (622, 624) (e.g., a word representing an object substantially identical to the object represented by the first word) or that has a correlation with the first word (622, 624) greater than a threshold value.

[0112] For example, the electronic device (500) can identify a second image (720) that includes at least a portion (e.g., "sea") of a specific first word (622) (e.g., "turquoise sea and white sandy beach and blue sky") among images related to the user of the electronic device (500) in its metadata. For example, the electronic device (500) can identify a second image (740) that represents substantially the same object (e.g., "airplane") as the object (e.g., airplane) represented by a specific first word (624) ("This place is very close to the airport and is a famous tourist spot where you can see the low landing of airplanes") among images related to the user of the electronic device (500).

[0113] According to one embodiment, the electronic device (500) can determine a second image (720, 740) using a personal data core (PDC) included in the electronic device (500) that stores and retrieves information related to the user. The electronic device can verify necessary information using the personal data. For example, the electronic device can retrieve and verify necessary information using the personal data core.

[0114] The personalized data core can collect and manage users' personalized data. For example, the personalized data core may include a database, an activity engine, a context engine, a collection module, and / or a search module. The collection module can collect users' personalized data. The database may be implemented in a portion of the electronic device's memory and can classify and store users' personalized data by type. The search module can identify and output users' personalized data that matches a search term. The context engine can identify (or infer) the context of the personalized data and include it in the personalized data. The personalized data core can collect and manage users' personalized data using a large language model (LLM).

[0115] The personalized data core can store the user's personalized data (e.g., text history, web page surfing history, analysis results of images (e.g., tags), schedules, routes, and / or images of the user). The personalized data core can search for personalized data (e.g., second images (720, 740)) whose relevance to a specific word (e.g., first word (622, 624)) is greater than or equal to a threshold. The personalized data core can store not only the user's personalized data but also information analyzing the user's personalized data, the time (e.g., date) when the user's personalized data was acquired, and information indicating other personalized data associated with each personalized data. According to one embodiment, if the electronic device (500) does not identify an image among the images associated with the user that includes at least a portion of the first word (621) (e.g., "Delty aircraft landing") in the metadata, it may input the first word (621) (e.g., "Delty aircraft landing") into an external server and identify (or determine) the image received from the external server as the second image. For example, the electronic device (500) may determine an image included in the internet search results for the first word (621) (e.g., "Delty aircraft landing") as a second image related to the first word (621) (e.g., "Delty aircraft landing"). Alternatively, the electronic device (500) may determine an image included in the internet search results for the first word (623) (e.g., "many tourists watching the plane landing") as a second image related to the first word (623) (e.g., "many tourists watching the plane landing") if no image containing information in the metadata with a relevance level above a threshold for the first word (623) is found.

[0116] The electronic device (500) can determine the image with the highest correlation to the first word (622, 624) as the second image (720, 740) based on metadata of the image related to the user of the electronic device (500).

[0117] For example, if there are multiple images that include at least a portion of the first word (622) (e.g., "turquoise sea and white sandy beach and blue sky") as metadata, the electronic device (500) can determine (or identify) the image that includes the largest portion (e.g., sea, sandy beach, sky) of the first word (622) (e.g., "turquoise sea and white sandy beach and blue sky") as metadata as the second image (720) associated with the first word (622) (e.g., "turquoise sea and white sandy beach and blue sky"). Alternatively, if there are multiple images containing at least a portion of the first word (624) (e.g., "This place is very close to the airport and is a famous tourist spot where you can see the low landing of an airplane") as metadata, the electronic device (500) may determine the most recently identified (or captured) image by the user as the second image (740) associated with the first word (624) (e.g., "This place is very close to the airport and is a famous tourist spot where you can see the low landing of an airplane").

[0118] FIG. 8 is a diagram illustrating an operation to generate a second text containing user information (or information associated with the user) related to a first word based on a second image of an electronic device according to one embodiment.

[0119] The description of some of the first words (621, 622, 623, 624) in Fig. 8 may also be applied to the rest.

[0120] The electronic device (500) can generate text describing the second image (720, 740) when it identifies (or determines) the second image (720, 740) associated with the first word (622, 624). The electronic device (500) can generate text describing the second image (720, 740) based on an object included in the second image (720, 740). The electronic device (500) can generate text describing each of the second images (720, 740) associated with each first word (622, 624) when there are multiple first words (622, 624) included in the first text (e.g., the first text (620) of FIG. 6) describing the first image (e.g., the first image (610) of FIG. 6). The operation of generating text (620) describing an image based on the image (610) of FIG. 4 can be applied to the operation of generating text describing a second image based on the second image (720, 740) in FIG. 8.

[0121] According to one embodiment, the electronic device (500) may generate text describing the second image (720, 740) using metadata of the second image (720, 740). For example, the electronic device (500) may input a command (or prompt) instructing to generate text describing the second image (720, 740), the second image (720, 740), and / or metadata of the second image (720, 740) into a generative artificial intelligence model (e.g., the generative artificial intelligence model (230) of FIG. 2), and generate text describing the second image (720, 740) using the result output by the generative artificial intelligence (230). The electronic device (500) can generate text describing the second image (720, 740), including at least some of information indicating an object included in the second image (720, 740), information indicating the time (and date) when the second image (720, 740) was taken, information indicating the location where the second image (720, 740) was taken, and / or information indicating an event related to the second image (720, 740).

[0122] The electronic device (500) can control a generative artificial intelligence (230) model to identify at least one word among the words included in the text describing the second image (720, 740) that satisfies a predetermined condition. The electronic device (500) can input the text describing the second image (720, 740) into the generative artificial intelligence model (230) and identify at least one word output by the generative artificial intelligence model (230). The predetermined condition may include a condition that the word among the words included in the text describing the second image has the highest importance (or has an importance level greater than or equal to a threshold).

[0123] The electronic device (500) can identify a second word (or a second word corresponding to the first word (622, 624)) among at least one word satisfying a predetermined condition included in the text describing the second image (720, 740). The second word associated with the first word (622, 624) may include a word representing an object substantially identical to the object represented by the first word (622, 624). For example, the area (741) representing the object represented by the second word included in the second image (740) and the area representing the object represented by the first word (624) included in the first image (610) may represent substantially the same object.

[0124] The electronic device (500) can generate text that complements the first text (620) describing the first image (610). The text that complements the first text (620) may include text (820) that complements the first word (622) (e.g., "turquoise sea and white sandy beach and blue sky") included in the first text (620) (e.g., "The beach here is the same color as Cancun, where I traveled last time").

[0125] The electronic device (500) can generate a prompt instructing the generation of text that complements the first word (622) based on at least part of the metadata of the first word (622), the second word, the first image (610) (or the area representing the object represented by the first word (622)), the second image (720) (or the area representing the object represented by the second word), and / or the second image (720) (e.g., user information). The electronic device (500) inputs a prompt to a generative artificial intelligence model (230) instructing the generation of text that complements the first word (622), at least part of the metadata of the first word (622), the second word, the first image (610) (or the area representing the object represented by the first word (622)), the second image (720) (or the area representing the object represented by the second word), and / or the second image (720) (e.g., user information) and / or the first word (622), and can generate text (820) that complements the first word (622) based on the result output by the generative artificial intelligence model (230).

[0126] The electronic device (500) can generate text (820) (e.g., "The beach here is the same color as Cancun, where I traveled last time") that complements the first word (622) (e.g., "turquoise sea and white sandy beach and blue sky") based on user information (or information associated with the user) included in the metadata of the second image (720). The metadata of the second image (720) may include user information (or information associated with the user) (e.g., information indicating an object included in the second image (720) (e.g., "beach"), information indicating the time (and date) when the second image (720) was taken (e.g., "last summer"), information indicating the location where the second image (720) was taken (e.g., "Cancun"), and / or information indicating an event associated with the second image (720) (e.g., "family trip"). The objects included in the second image (720) may include objects (e.g., "beach") and / or people.

[0127] According to one embodiment, the electronic device (500) can generate text (840) that complements the first word (624) (e.g., "Do you remember taking a picture with an airplane on the beach when you went to Jeju last time? It is five times closer than that place.") based on an area (741) representing an object represented by a second word (e.g., airplane) included in a second image (740) and an area representing an object represented by a first word (624) included in a first image (610) (e.g., "This place is very close to the airport and is a famous tourist spot where you can see the airplane's low landing"). For example, the electronic device (500) can compare an area (741) representing an object represented by a second word (e.g., airplane) and an area representing an object represented by a first word, and generate text (840) that complements the first word (624) to include the result of the comparison (e.g., "The area (710) representing an airplane in the second image (740) is five times closer to the shooting location than the area (610) representing an airplane in the first image") (e.g., "Do you remember taking a picture with an airplane on the beach when you went to Jeju last time? It is five times closer than that place."). For example, the electronic device (500) can generate text (840) that complements the first word, which includes the result of comparing the characteristics (e.g., color, size, shape and / or ratio) of each of the area (741) representing the object represented by the second word (e.g., airplane) included in the second image (840) and the area representing the object represented by the first word (624) included in the first image (610).

[0128] The electronic device (500) can generate a second text (801) containing user information (or information associated with the user) included in a second image (720, 740) based on the text (820, 840) that complements the first word and / or the first text (620). For example, the electronic device (500) can generate the second text (801) by replacing the first word (622) (e.g., "turquoise sea and white sandy beach, and blue sky") included in the first text (620) with the text (820) that complements the first word (e.g., "The beach here is the same color as Cancun, where I traveled last time"). The electronic device (500) can generate a second text (801) by inserting (or placing) text (840) that complements the first word (624) (e.g., "This place is very close to the airport and a famous tourist spot where you can see the low landing of airplanes") (e.g., "Do you remember taking a picture with an airplane on the beach when you went to Jeju last time? It is five times closer than that place") into an adjacent part of the first text (620).

[0129] The second text (801) may include text (820, 840) (or at least a part of the text) that complements at least one first word. The electronic device (500) inputs a prompt generated based on user information (or information associated with the user) included in the metadata of an image associated with the user (e.g., the second image (720, 740)) and user information included in the metadata of the image associated with the user (e.g., the second image (720, 740)) into a generative artificial intelligence model (230), and can generate text (820, 840) that complements the first word based on the result output by the generative artificial intelligence model (230). The generative artificial intelligence model (230) can generate text (820, 840) that complements the first word based on user information included in the prompt. Unlike the first text (620), the second text (801) may include user information (or information associated with the user) (e.g., user information included in the metadata of the second image (720, 740)).

[0130] FIGS. 9a and 9b are drawings for explaining the operation of outputting a second text of an electronic device according to one embodiment.

[0131] Referring to FIGS. 9a and 9b, the electronic device (500) may output a second text (801) (or a sound corresponding to the second text (801)) containing user information (or information associated with the user) and a first image (610). When the electronic device (500) outputs the first image (610) and the second text (801) (or a sound corresponding to the second text (801)), it may output the second image (720, 740) together.

[0132] Referring to FIG. 9a, when the electronic device (500) outputs the second text (801), it may apply a first graphic effect (901) to text (840) that complements the first word included in the second text (801) (e.g., "Do you remember taking a picture with an airplane on the beach when you went to Jeju last time? It is about five times closer than that place."). The first graphic effect (901) applied to the text (840) that complements the first word (624) may be an effect that distinguishes the text that complements the first word (624) from the rest included in the second text (801). For example, the electronic device (500) may make the text (840) that complements the first word sparkle, change it to a different color, or highlight it (e.g., highlight it with bold text).

[0133] Referring to FIGS. 9a and 9b, when the electronic device (500) outputs the first image (610), it may apply a second graphic effect (902) to an area (614) representing an object (e.g., airplane) indicated by the first word (624) ("This place is very close to the airport and is a famous tourist spot where you can see the low landing of airplanes"). The second graphic effect (902) may be an effect that distinguishes the area (614) representing the object indicated by the first word (e.g., airplane) from other areas included in the first image (610). For example, the electronic device (500) may make the area (614) representing the object (e.g., airplane) represented by the first word (624) ("This place is very close to the airport and is a famous tourist spot where you can see the low landing of airplanes") flash, or highlight the boundary line of the area (614) representing the object (e.g., airplane) represented by the first word (624) ("This place is very close to the airport and is a famous tourist spot where you can see the low landing of airplanes"). When the electronic device (500) applies the first graphic effect (901) to the text (840) that complements the specific first word (624), it may also apply the second graphic effect (902) together to the area (614) representing the object represented by the specific first word (624).

[0134] Referring to FIG. 9a, the electronic device (500) may be configured to output both the first image (610) and the second text (801) on the display (530). The electronic device (500) may be configured to output the second image (740) associated with the first word (624) in response to receiving user input selecting an area (614) representing an object (e.g., airplane) represented by the first word (624) (e.g., "This place is very close to the airport and is a famous tourist spot where you can see the low landing of airplanes"). The electronic device (500) can output a second image (720) associated with another first word (622) (e.g., "turquoise sea and white sandy beach and blue sky") in response to receiving user input selecting an area (611) representing an object (e.g., sea and sky) represented by another first word (622) (e.g., "turquoise sea and white sandy beach and blue sky").

[0135] Referring to FIG. 9b, the electronic device (500) can output the second text (801) in the form of sound. For example, the electronic device (500) can perform text-to-speech (TTS) of the second text (801) and, using the result of performing TTS, output a sound corresponding to the second text (801). The electronic device (500) can output the sound corresponding to the second text (801) and the first image (610) together. For example, the electronic device can output the sound corresponding to the second text (801) through a sound output component (e.g., a speaker) and output the first image (610) through a display (530).

[0136] Referring to FIGS. 9a and 9b, when the electronic device (500) outputs the second image (720, 740), it may output information included in the metadata of the second image (720, 740) based on user input (e.g., long press) received in the area where the second image (720, 740) is displayed. The metadata of the second image (720, 740) may include user information (or information associated with the user) (e.g., information indicating an object (person) included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event (e.g., a schedule) associated with the image).

[0137] According to one embodiment, when the electronic device (500) outputs the second image (720, 740), it may output text describing the second image (e.g., text describing the second image described in FIG. 8) based on user input (e.g., long press) received in the area where the second image (720, 740) is displayed. It should be understood that the description of the first text (620) and the second text (801) describing the first image (610) may be applied in the same way to the text describing the second image (720, 740). For example, the operation of the electronic device (500) generating text (second text (801)) containing user information (or information associated with the user) describing the first image (610) may be applied in the same way to the second image (720, 740). The electronic device (500) can identify a third word among the words included in the text describing the second image (720, 740), identify a third image related to the third word, and generate a third text including user information (or information associated with the user) included in the third image based on the third image. The third text may also include text that complements the third word based on the third image.

[0138] When the electronic device (500) of the present disclosure receives user input instructing it to output text describing a first image (610), it may generate a second text (801) containing user information (or user-related information) based on a second image (720, 740) containing metadata containing user information (or user-related information). By including user information (or user-related information) in the second text (801) describing the first image (610), the electronic device (500) may improve the user's understanding of the text describing the first image (610).

[0139] Additionally, the electronic device (500) can provide the user with the second image (720, 740) used to generate the second text (801) by outputting the second image (720, 740) together with the second text (801) and the first image (610). By providing the user with the second image (720, 740) used to generate the second text (801), the electronic device (500) can improve the user's understanding of the second text (801).

[0140] FIG. 10 is a diagram illustrating the operation of outputting a second text of an electronic device according to one embodiment.

[0141] According to one embodiment, the electronic device (500) may identify a second image (741) that includes at least a portion of a specific first word among images with a history confirmed by a user (e.g., a scene from a movie with a history output by a media application) as metadata. The electronic device (500) may output a second text (1001) (or a sound corresponding to the second text (1001)) generated based on the second image (741), which is an image with a history confirmed by a user, together with the first image (610). The second text (1001) generated based on the second image (741), which is an image with a history confirmed by a user, may include information different from the text generated by the image obtained by the user (e.g., the second image (740) of FIG. 9a) (e.g., the second text (801) of FIG. 9a). For example, the second text (1001) may include information related to a history confirmed by the user of the second image (741) (e.g., watching movie A last week) (e.g., "It is a location close to a scene from movie A watched last week") (1002). As another example, the second text (1001) may include information related to the user of the other image (721) (e.g., watching movie B that the user likes) (e.g., "It looks like the same sea color as movie B that the user likes") (1003).

[0142] The electronic device (500) can improve the user's understanding of the second text (1001) describing the first image (610) by including user information (e.g., information related to the second image (741), which is an image that has been verified by the user) in the second text (1001) describing the first image (610).

[0143] FIG. 11 is a flowchart of the operation of an electronic device according to one embodiment.

[0144] An electronic device (e.g., the electronic device (100) of FIG. 1, the electronic device (300) of FIG. 3, the electronic device (500) of FIG. 5)) can, in operation 1110, generate a first text (e.g., the first text (620) of FIG. 6) describing a first image (e.g., the first image (610) of FIG. 6).

[0145] For example, the electronic device (500) may input a command (or prompt) instructing the generation of text describing the first image (610) and the first image (610) into a generative artificial intelligence model (230), and generate a first text (620) describing the first image (610) using the result output by the generative artificial intelligence model (230).

[0146] According to one embodiment, the electronic device (500) may generate text describing the first image (610) based on the first image (610) and generate a first text (620) describing the first image (610) by summarizing the text describing the first image (610). The electronic device (500) may input a command (or prompt) instructing to summarize the text describing the first image (610) and the text describing the first image (610) into a generative artificial intelligence model (230), and generate a first text (620) summarizing the text describing the first image (610) using the result output by the generative artificial intelligence model (230).

[0147] The electronic device (500) can identify the first word (e.g., the first words (621, 622, 623, 624) of FIG. 6) included in the first text (620) describing the first image (610) in operation 1120.

[0148] The electronic device (500) can identify at least one first word (621, 622, 623, 624) included in the first text (620). For example, the electronic device (500) can control a generative artificial intelligence model (230) to identify at least one first word (621, 622, 623, 624) among the words included in the first text (620) that satisfies a predetermined condition. The electronic device (500) can input the first text (620) into the generative artificial intelligence model (230) and identify at least one first word (621, 622, 623, 624) output by the generative artificial intelligence model (230). The predetermined condition may include a condition that the word among the words included in the first text (620) has the highest importance (or has an importance level greater than or equal to a threshold). According to one embodiment, the first word (621, 622, 623, 624) may be composed of a plurality of words.

[0149] In operation 1130, the electronic device (500) can identify a second image (e.g., the second image (720, 740) of FIG. 7) associated with the first word (621, 622, 623, 624) among images associated with the user of the electronic device (500).

[0150] An image associated with a user of the electronic device (500) may include an image stored in a memory included in the electronic device (500) (e.g., memory (520) of FIG. 5, memory (120) of FIG. 1) and / or an image with a history confirmed by a user of the electronic device (500). An image with a history confirmed by a user of the electronic device (500) may include an image with a history processed by various functions executed by the electronic device (500) (e.g., content output, internet search, message transmission and reception) (or an image with a history output on a display (530) included in the electronic device (500).

[0151] An image associated with a user of the electronic device (500) may include metadata containing information about the user (or information associated with the user) included in the image associated with the user of the electronic device (500). The information about the user (or information associated with the user) may include information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event associated with the image. The object included in the image may include objects and / or people. The information indicating a person included in the image may include information indicating the relationship between the person included in the image and the user (e.g., friends, family). The event associated with the image may include the user's schedule related to the time and / or location when the image was taken. When the electronic device (500) acquires an image, it may store information indicating an event associated with the image as metadata of the image based on information indicating the user's schedule stored in the electronic device (500) (or memory (530)).

[0152] The electronic device (500) may determine as a second image (720, 740) an image containing a first word (621, 622, 623, 624) (or at least a part of the first word (621, 622, 623, 624)) in the metadata among images related to the user of the electronic device (500). Including the first word (621, 622, 623, 624) (or at least a part of the first word (621, 622, 623, 624)) in the metadata may include including in the metadata information that is substantially identical to the first word (621, 622, 623, 624) (e.g., a word representing an object substantially identical to the object represented by the first word (621, 622, 623, 624)) or that has a correlation with the first word (621, 622, 623, 624) above a threshold value.

[0153] According to one embodiment, the electronic device (500) can determine a second image (720, 740) using a personal data core (PDC) included in the electronic device (500) that stores and retrieves images related to the user. The personal data core can store the user's personal data (e.g., the user's schedule, the user's events and / or the user's images). The personal data core can retrieve personal data (e.g., the second image (720, 740)) that has a relevance to a specific word (e.g., the first word (621, 622, 623, 624)) above a threshold value.

[0154] According to one embodiment, if information having a relevance level greater than a threshold is not found among images related to the user of the electronic device (500), the electronic device (500) inputs the first word (621, 622, 623, 624) to an external server and can receive the second image (720, 740) from the external server (or determine the image received from the external server as the second image (720, 740)). For example, the electronic device (500) may determine an image included in the internet search results of the first word (621, 622, 623, 624) as the second image (720, 740) related to the first word (621, 622, 623, 624).

[0155] According to one embodiment, the electronic device (500) can identify a second image (720, 740) associated with a specific first word (621, 622, 623, 624) based on a specific first word (621, 622, 623, 624) included in a first text (620) describing a first image (610). The electronic device (500) can identify a second image (720, 740) that includes at least a portion of the specific first word (621, 622, 623, 624) among an image stored in memory (or an image with a history of being identified by a user) as metadata. For example, the electronic device (500) may determine the image with the highest degree of correlation to a specific first word (621, 622, 623, 624) as the second image (720, 740) based on the metadata of an image related to the user of the electronic device (500). If there are multiple images that include at least a portion of the specific first word (621, 622, 623, 624) as metadata, the electronic device (500) may determine the image that includes the largest portion of the specific first word (621, 622, 623, 624) as metadata as the second image (720, 740) related to the first word (621, 622, 623, 624). Alternatively, if there are multiple images containing at least a portion of a specific first word (621, 622, 623, 624) as metadata, the electronic device (500) may determine the most recently identified (or captured) image by the user as the second image (720, 740) associated with the first word (621, 622, 623, 624).

[0156] According to one embodiment, the electronic device (500) can identify a second image (720, 740) associated with a specific first word (621, 622, 623, 624) based on a specific first word (621, 622, 623, 624) included in a first text (620) describing a first image (610).

[0157] According to one embodiment, if an image containing information in metadata that has a correlation with a specific first word (621, 622, 623, 624) greater than or equal to a threshold value is not found, the electronic device (500) inputs the specific first word (621, 622, 623, 624) to an external server and checks the second image (720, 740) received from the external server (or determines the image received from the external server as the second image (720, 740)). For example, the electronic device (500) may determine an image included in the internet search results of the specific first word (621, 622, 623, 624) as the second image (720, 740) associated with the first word (621, 622, 623, 624).

[0158] The electronic device (500) can, in operation 1140, generate a second text (801) containing user information (or information associated with the user) included in the second image (720, 740) associated with the first word (621, 622, 623, 624), based on the second image (720, 740).

[0159] The electronic device (500) can generate text describing the second image (720, 740) when it identifies (or determines) the second image (720, 740) associated with the first word (621, 622, 623, 624). The electronic device (500) can generate text describing the second image (720, 740) based on an object included in the second image (720, 740). If there are multiple first words (621, 622, 623, 624) included in the first text (620) describing the first image (610), the electronic device (500) can generate text describing each of the second images (720, 740) associated with each first word (621, 622, 623, 624).

[0160] According to one embodiment, the electronic device (500) may generate text describing the second image (720, 740) using metadata of the second image (720, 740). For example, the electronic device (500) may input a command (or prompt) instructing to generate text describing the second image (720, 740), the second image (720, 740), and / or metadata of the second image (720, 740) into a generative artificial intelligence model (230), and generate text describing the second image (720, 740) using the result output by the generative artificial intelligence model (230). The electronic device (500) may generate text describing the second image (720, 740) including at least some of information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event related to the image.

[0161] The electronic device (500) can control a generative artificial intelligence model (230) to identify at least one word among the words included in the text describing the second image (720, 740) that satisfies a predetermined condition. The electronic device (500) can input the text describing the second image (720, 740) into the generative artificial intelligence model (230) and identify at least one word output by the generative artificial intelligence model (230). The predetermined condition may include a condition that the word among the words included in the second text (801) has the highest importance (or has an importance level greater than or equal to a threshold).

[0162] The electronic device (500) can identify a second word (or a second word corresponding to the first word (621, 622, 623, 624)) among at least one word satisfying a predetermined condition included in the text describing the second image (720, 740). The second word associated with the first word (621, 622, 623, 624) may include a word representing an object substantially identical to the object represented by the first word (621, 622, 623, 624).

[0163] The electronic device (500) can identify an area representing an object represented by the second word among the areas included in the second image (720, 740) and / or an area representing an object represented by the first word (621, 622, 623, 624) among the areas included in the first image (610). The area representing an object represented by the second word included in the second image (720, 740) and the area representing an object represented by the first word (621, 622, 623, 624) included in the first image (610) may represent substantially the same object.

[0164] The electronic device (500) can generate text that complements the first text (620) describing the first image (610). The text that complements the first text (620) may include text that complements the first words (621, 622, 623, 624) included in the first text (620).

[0165] The electronic device (500) can generate text that complements the first word (621, 622, 623, 624) based on user information (or information associated with the user) included in the metadata of the second image (720, 740) associated with the first word (621, 622, 623, 624). The metadata of the second image (720, 740) may include user information (or information associated with the user) (e.g., information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event associated with the image). The object included in the image may include objects and / or people.

[0166] According to one embodiment, the electronic device (500) can generate text that complements the first word (621, 622, 623, 624) based on an area representing an object represented by the second word and an area representing an object represented by the first word (621, 622, 623, 624). For example, the electronic device (500) can compare an area representing an object represented by the second word and an area representing an object represented by the first word (621, 622, 623, 624), and generate text that complements the first word (621, 622, 623, 624) to include the result of the comparison. For example, the electronic device (500) can generate text that complements the first word (621, 622, 623, 624) including the result of comparing the characteristics (e.g., color, size, shape and / or ratio) of each of the area represented by the second word included in the second image (720, 740) and the area represented by the first word (621, 622, 623, 624) included in the first image (610).

[0167] The electronic device (500) can generate a second text (801) containing user information (or information associated with the user) included in a second image (720, 740) based on text that complements the first words (621, 622, 623, 624) and / or the first text (620). For example, the electronic device (500) can generate the second text (801) by replacing the first words (621, 622, 623, 624) included in the first text (620) with text that complements the first words (621, 622, 623, 624). The electronic device (500) can generate a second text (801) by inserting (or placing) text that complements the first word (621, 622, 623, 624) in an adjacent part of the first word (621, 622, 623, 624) included in the first text (620).

[0168] The electronic device (500) can output a first image (610), a second image (720, 740), and a second text (801) in operation 1150.

[0169] The electronic device (500) can output both a second text (801) containing user information (or information associated with the user) and a first image (610) on a display (530) included in the electronic device (500). When the electronic device (500) outputs the first image (610) and the second text (801), it can output the second image (720, 740) together. Outputting the second text (801) of the electronic device (500) of the present disclosure may include not only outputting the second text (801) on the display (530) but also outputting a sound corresponding to the second text (801) (e.g., the result of the second text (801)) through an audio output configuration (e.g., a speaker).

[0170] When the electronic device (500) outputs the second text (801), it may apply a first graphic effect to a part corresponding to the first word (621, 622, 623, 624) included in the second text (801) (or text that complements the first word (621, 622, 623, 624) included in the second text (801), user information (or information associated with the user) included in the second text (801). The first graphic effect applied to the text that complements the first word (621, 622, 623, 624) may be an effect that distinguishes the text that complements the first word (621, 622, 623, 624) from the remainder included in the second text (801). For example, the electronic device (500) can make the text complementing the first word (621, 622, 623, 624) sparkle, change to a different color, or highlight (e.g., highlight in bold).

[0171] When the electronic device (500) outputs the first image (610), it may apply a second graphic effect to the area representing the object represented by the first word (621, 622, 623, 624). The second graphic effect may be an effect that distinguishes the area representing the object represented by the first word (621, 622, 623, 624) from other areas included in the first image (610). For example, the electronic device (500) may make the area representing the object represented by the first word (621, 622, 623, 624) sparkle or emphasize the boundary line of the area representing the object represented by the first word (621, 622, 623, 624).

[0172] According to one embodiment, the electronic device (500) may be configured to output a first image (610) and a second text (801) on a display (530). The electronic device (500) may be configured to output a second image (720, 740) associated with a specific first word (621, 622, 623, 624) in response to receiving user input selecting an area representing an object represented by a specific first word (621, 622, 623, 624). The electronic device (500) may output a second image (720, 740) associated with another first word (621, 622, 623, 624) in response to receiving user input selecting an area representing another first word (621, 622, 623, 624).

[0173] According to one embodiment, when the electronic device (500) outputs the second image (720, 740), it may output information included in the metadata of the second image (720, 740) based on user input (e.g., long press) received in the area where the second image (720, 740) is displayed. The metadata of the second image (720, 740) may include user information (or information associated with the user) (e.g., information indicating an object (person) included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event (e.g., a schedule) associated with the image).

[0174] According to one embodiment, when the electronic device (500) outputs the second image (720, 740), it may output text describing the second image (720, 740) based on user input (e.g., long press) received in the area where the second image (720, 740) is displayed.

[0175] According to one embodiment, the operation of the electronic device (500) of the present disclosure generating and outputting text (second text (801)) containing user information (or information associated with the user) describing the first image (610) may be applied in the same way to the second image (720, 740). For example, the electronic device (500) may identify a third word among the words included in the text describing the second image (720, 740), identify a third image associated with the third word, and generate a third text containing user information (or information associated with the user) included in the third image based on the third image. The electronic device (500) may output the second image (720, 740), the third image, and / or the third text. The third text may include text that complements the third word based on the third image.

[0176] In an electronic device according to one embodiment, the electronic device may include at least one processor. The electronic device may include a memory that stores at least one computer program containing instructions. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may generate a first text describing a first image. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may identify a first word included in the first text describing the first image. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may identify a second image related to the first word among images related to the user of the electronic device. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may generate a second text containing user information included in the second image based on the second image. When the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may output all of the first image, the second image, and the second text.

[0177] In an electronic device according to one embodiment, each of the images associated with a user of the electronic device may include metadata comprising user information represented by each of the images associated with the user of the electronic device. The user information may include information representing an object included in the image, information representing the time (and date) when the image was taken, information representing the location where the image was taken, and / or information representing an event associated with the image.

[0178] In an electronic device according to one embodiment, the second at least one computer program may be executed individually or collectively by the at least one processor to enable the electronic device to identify a second image containing metadata including at least a portion of the first word among images associated with the user of the electronic device.

[0179] In an electronic device according to one embodiment, images associated with a user of the electronic device may include an image stored in the memory and / or an image with a history of being verified by the user. An image with a history of being verified by the user may include an image with a history of being processed by a function executed by the electronic device and / or an image with a history of being displayed on the display of the electronic device.

[0180] In an electronic device according to one embodiment, the at least one computer program may, when executed individually or collectively by the at least one processor, cause the electronic device to input the first word to an external server if the second image is not identified based on images associated with the user of the electronic device. The at least one computer program may, when executed individually or collectively by the at least one processor, cause the electronic device to identify the second image received from the external server.

[0181] In an electronic device according to one embodiment, the at least one computer program may, when executed individually or collectively by the at least one processor, cause the electronic device to generate text describing the second image when generating the second text. The at least one computer program may, when executed individually or collectively by the at least one processor, cause the electronic device to identify a second word related to the first word included in the text describing the second image. The at least one computer program may, when executed individually or collectively by the at least one processor, cause the electronic device to generate text supplementing the first word based on user information included in the metadata of the second image. The at least one computer program may, when executed individually or collectively by the at least one processor, cause the electronic device to generate a second text based on the text supplementing the first word and the first text.

[0182] In an electronic device according to one embodiment, the at least one computer program may be executed individually or collectively by the at least one processor to enable the electronic device to generate text that complements the first word, the text comprising the result of comparing an area representing an object represented by the second word included in the second image and an area representing an object represented by the first word included in the first image.

[0183] In an electronic device according to one embodiment, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device may output a first graphic effect on the portion corresponding to the first word included in the second text, which distinguishes the portion corresponding to the first word included in the second text from the remainder of the portion included in the second text.

[0184] In an electronic device according to one embodiment, when the electronic device outputs the first image, the at least one computer program may be configured to apply a second graphic effect to the area representing the object represented by the first word, which distinguishes the area representing the object represented by the first word from other areas included in the first image when the electronic device is executed individually or collectively by the at least one processor.

[0185] In an electronic device according to one embodiment, the first aspect of claim 1, wherein the at least one computer program, when executed individually or collectively by the at least one processor, the electronic device may output information included in the metadata of the second image and / or text describing the second image based on user input received in an area where the second image is displayed.

[0186] In an electronic device according to one embodiment, the first aspect of claim, wherein the at least one computer program may cause the electronic device to generate text describing the second image when executed individually or collectively by the at least one processor. The at least one computer program may cause the electronic device to generate text describing the second image when executed individually or collectively by the at least one processor.

[0187] The third word can be identified among the words included in the text describing the second image. The at least one computer program can be executed individually or collectively by the at least one processor to enable the electronic device to identify a third image associated with the third word. The at least one computer program can be executed individually or collectively by the at least one processor to enable the electronic device to generate a third text containing user information included in the third image based on the third image. The at least one computer program can be executed individually or collectively by the at least one processor to enable the electronic device to output the second image, the third image, and / or the third text.

[0188] In an electronic device according to one embodiment, the at least one computer program may be configured to control a generative artificial intelligence model to identify at least one word among the words included in the first text that satisfies a predetermined condition when the electronic device is executed individually or collectively by the at least one processor, when the first word included in the first text is identified. The predetermined condition may include a condition in which the importance of the words included in the first text is greater than or equal to a threshold value.

[0189] In a method of operating an electronic device according to one embodiment, the method of operating the electronic device may include an operation of generating a first text describing a first image. The method of operation may include an operation of verifying a first word included in the first text describing the first image. The method of operation may include an operation of verifying a second image related to the first word among images related to a user of the electronic device. The method of operation may include an operation of generating a second text including user information included in the second image based on the second image. The method of operation may include an operation of outputting the first image, the second image, and the second text.

[0190] In a method of operating an electronic device according to one embodiment, each of the images associated with a user of the electronic device may include metadata including user information represented by each of the images associated with the user of the electronic device.

[0191] The user information above may include information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event related to the image.

[0192] In a method of operating an electronic device according to one embodiment, the method of operating the electronic device may include an operation of checking a second image containing metadata including at least a portion of the first word among images related to a user of the electronic device.

[0193] In a method of operating an electronic device according to one embodiment, images related to a user of the electronic device may include an image stored in the memory of the electronic device and / or an image with a history of being verified by the user. The image with a history of being verified by the user may include an image with a history of being processed by a function executed by the electronic device and / or an image with a history of being displayed on the display of the electronic device.

[0194] In a method of operating an electronic device according to one embodiment, the method of operating the electronic device may include, when a second image is not confirmed based on images related to a user of the electronic device, inputting the first word to an external server and confirming the second image received from the external server.

[0195] In a method of operation of an electronic device according to one embodiment, the method of operation of the electronic device may include, when generating the second text: generating text describing the second image and identifying a second word related to the first word included in the text describing the second image. The method of operation may include generating text that complements the first word based on user information included in the metadata of the second image. The method of operation may include generating a second text based on the text that complements the first word and the first text.

[0196] In a method of operating an electronic device according to one embodiment, the method of operating the electronic device may include an operation of generating text that complements the first word, the result of comparing an area representing an object represented by the second word included in the second image and an area representing an object represented by the first word included in the first image.

[0197] In a method of operating an electronic device according to one embodiment, the method of operating the electronic device may include, when outputting the second text, an operation of outputting a first graphic effect on a portion corresponding to the first word included in the second text to distinguish the portion corresponding to the first word included in the second text from the remaining portion included in the second text.

[0198] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.

[0199] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any possible combination of items listed together in the corresponding phrase. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.

[0200] As used in this document, the term "module" may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be a component formed as a whole, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).

[0201] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored from the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' merely means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and this term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.

[0202] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0203] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to the integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In an electronic device, At least one processor; Memory for storing at least one computer program including instructions; When the above at least one computer program is executed individually or collectively by the above at least one processor, the electronic device, Generate a first text describing a first image, and Identify the first word included in the first text describing the above first image. Identify the second image related to the first word among the images related to the user of the electronic device, and Based on the second image above, a second text including user information included in the second image is generated, and An electronic device that outputs the first image, the second image, and the second text.

2. In claim 1, each of the images related to the user of the electronic device is It includes metadata containing user information represented by each of the images related to the user of the electronic device, and An electronic device wherein the user information includes information indicating an object included in the image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event related to the image.

3. In paragraph 2, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, An electronic device that enables the identification of a second image containing metadata including at least a portion of the first word among images related to the user of the electronic device.

4. In paragraph 1, the images related to the user of the electronic device are, It includes images stored in the memory and / or images with a history verified by the user, Images with a history confirmed by the above user, An electronic device comprising an image with a history processed by a function executed by the electronic device and / or an image with a history output on the display of said electronic device.

5. In claim 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, An electronic device that inputs the first word to an external server and checks the second image received from the external server when the second image is not confirmed based on images related to the user of the electronic device.

6. In claim 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, When generating the above second text: Generate text describing the above second image, and Identifying the second word related to the first word included in the text describing the second image above, and Based on user information included in the metadata of the second image above, text that complements the first word is generated, and An electronic device that generates a second text based on the first text and a text that complements the first word.

7. In paragraph 6, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, An electronic device that generates text that complements the first word, comprising a result of comparing an area representing an object represented by the second word included in the second image and an area representing an object represented by the first word included in the first image.

8. In paragraph 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, An electronic device that, when outputting the second text, outputs a first graphic effect on the portion corresponding to the first word included in the second text to distinguish the portion corresponding to the first word included in the second text from the remainder of the portion included in the second text.

9. In paragraph 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, An electronic device that, when outputting the first image, applies a second graphic effect to the area representing the object represented by the first word to distinguish the area representing the object represented by the first word from other areas included in the first image.

10. In claim 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, An electronic device that outputs information included in the metadata of the second image and / or text describing the second image based on user input received in an area where the second image is displayed.

11. In claim 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, Generate text describing the above second image, and Identify the third word among the words included in the text describing the second image above, and Check the third image related to the above third word, and Based on the third image above, a third text including user information included in the third image is generated, and An electronic device capable of outputting the second image, the third image and / or the third text.

12. In claim 1, when the at least one computer program is executed individually or collectively by the at least one processor, the electronic device, When checking a first word included in a first text, the generative artificial intelligence model is controlled to check at least one word among the words included in the first text that satisfies a predetermined condition, and The above-determined condition is an electronic device that includes a condition in which the importance of a word included in the first text is greater than or equal to a threshold value.

13. In a method of operating an electronic device, the method of operating the electronic device is, The action of generating a first text describing a first image, The operation of verifying the first word included in the first text describing the first image above, The operation of identifying a second image related to the first word among images related to a user of an electronic device, An operation to generate a second text including user information included in the second image based on the second image, and A method of operating an electronic device comprising the operation of outputting the first image, the second image, and the second text.

14. In claim 13, each of the images associated with the user of the electronic device is It includes metadata containing user information represented by each of the images related to the user of the electronic device, and A method of operating an electronic device, wherein the user information includes information indicating an object included in an image, information indicating the time (and date) when the image was taken, information indicating the location where the image was taken, and / or information indicating an event related to the image.

15. In claim 14, the method of operating the electronic device is, A method of operating an electronic device comprising the operation of identifying a second image containing metadata including at least a portion of the first word among images related to the user of the electronic device.