Method for providing response on basis of modality determination and electronic device therefor
By using AI models to analyze both text and image data within a conversation, the electronic device generates response messages that are highly relevant to the conversation topic, addressing the challenge of relevance in existing systems.
Patent Information
- Application Number
- PCT/KR2025/006845
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-09-19
- Filing Date
- 2025-05-20
- Publication Date
- 2026-01-15
AI Technical Summary
Electronic devices struggle to suggest response messages highly relevant to the topic of a conversation due to their inability to consider the relationship between the conversation topic and objects in the exchanged images.
An electronic device equipped with processors, displays, and communication circuits uses AI models to identify and analyze both text and image data within a conversation, generating a response message that considers the relationship between the conversation topic and image objects.
The device can suggest response messages that are highly relevant to the conversation topic by integrating AI models to analyze image data and text, enhancing the relevance and accuracy of responses.
Smart Images

Figure KR2025006845_15012026_PF_FP_ABST
Abstract
Description
Method for providing response based on modality determination and electronic device therefor
[0001] Embodiments disclosed in this document relate to a method for providing a response based on determination of modality and an electronic device therefor.
[0002] A variety of services are being developed using generative AI. Early services supported generating text-based responses or generating related images based on images. With the advancement of generative AI, services supporting various input formats are being introduced. For example, a service is being developed that suggests response messages to received messages using conversational messages exchanged with a specified user.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art in connection with the present disclosure.
[0004] Typically, electronic devices may analyze not only the conversational form of messages exchanged with a given user, but also the content transmitted during the exchange of messages, in order to suggest a response message.
[0005] For example, if an electronic device receives an image while exchanging messages with a designated user, it can identify objects contained in the image and associate them with the topic of the conversation to suggest an appropriate response message.
[0006] However, electronic devices have difficulty suggesting response messages that are highly relevant to the topic of the conversation because they do not consider the relationship between the topic of the conversation and the objects included in the image when identifying the objects included in the image.
[0007] The technical problems to be achieved in this document are not limited to the technical problems mentioned above, and other technical problems not mentioned can be clearly understood by a person having ordinary skill in the technical field to which the present invention belongs from the description below.
[0008] According to various embodiments, an electronic device includes at least one processor, a display, a communication circuit, and a memory operatively connected to the at least one processor, the display, and the communication circuit and storing at least one instruction, wherein the at least one instruction, when individually or collectively executed by the at least one processor, causes the electronic device to: output, through the display, a conversation screen in which a received message including at least one text and at least one image received from a specified user through the communication circuit is displayed; identify image data for the at least one image displayed as at least a part of the received message on the conversation screen; obtain text data describing an object included in the image data based on the identified image data and a specified first artificial intelligence (AI) model; and generate a response message for the received message based on the at least one text included in the received message and the obtained text data based on the image data and a specified second AI model.
[0009] An operating method of an electronic device according to various embodiments may include an operation of displaying a conversation screen in which a received message including at least one text and at least one image received from a specified user is displayed, an operation of identifying image data for the at least one image displayed as at least a part of the received message on the conversation screen, an operation of obtaining text data describing an object included in the image data based on the identified image data and a specified first artificial intelligence (AI) model, and an operation of generating a response message for the received message based on the obtained text data based on the at least one text and the image data included in the received message and a specified second AI model.
[0010] A computer-readable storage medium according to various embodiments may store instructions that, when executed by a processor of an electronic device, cause the electronic device to: output a conversation screen on which a received message including at least one text and at least one image received from a designated user is displayed, through a display; identify image data for the at least one image displayed as at least a part of the received message on the conversation screen; obtain text data describing an object included in the image data based on the identified image data and a designated first artificial intelligence (AI) model; and generate a response message for the received message based on the obtained text data based on the at least one text included in the received message and the image data and a designated second AI model.
[0011] An electronic device according to various embodiments disclosed in this document can suggest a response message highly relevant to the topic of a conversation by identifying an object included in an image by considering the relationship between the topic of the conversation and the object included in the image.
[0012] The effects that can be obtained from the present disclosure are not limited to the effects mentioned above, and other effects that are not mentioned can be clearly understood by a person having ordinary skill in the art to which the present disclosure belongs from the description below.
[0013] FIG. 1 illustrates a block diagram of an electronic device according to one embodiment.
[0014] FIG. 2 is a drawing for explaining a message function of an electronic device according to various embodiments.
[0015] FIG. 3 is a drawing for explaining the dialogue screen output operation of an electronic device according to various embodiments.
[0016] Figure 4 illustrates a block diagram of an artificial intelligence system according to various embodiments.
[0017] Figure 5 illustrates the structure of an artificial intelligence model according to one embodiment.
[0018] Figure 6 illustrates another block diagram of an artificial intelligence system according to various embodiments.
[0019] FIG. 7 is a diagram for explaining a response message generation operation of an electronic device according to various embodiments.
[0020] FIG. 8 is a diagram for explaining an operation of identifying conversation information and image data on a conversation screen according to various embodiments.
[0021] FIG. 9 is a diagram illustrating an operation of creating a view hierarchy relationship for a conversation screen in an electronic device according to various embodiments.
[0022] FIG. 10 is a diagram for explaining an operation of obtaining description information for image data in an electronic device according to various embodiments.
[0023] Figure 11 is a drawing for explaining description information according to various embodiments.
[0024] FIG. 12 is another diagram illustrating a response message generation operation of an electronic device according to various embodiments.
[0025] FIGS. 13A and 13B are diagrams illustrating a series of operations for providing a response message in an electronic device according to various embodiments.
[0026] FIGS. 14A and 14B are diagrams illustrating an electronic device supporting a response message generation function according to various embodiments.
[0027] FIG. 15 is another diagram for explaining a response message generation operation of an electronic device according to various embodiments.
[0028] FIG. 16a and FIG. 16b are diagrams for explaining response messages generated according to conversation styles according to various embodiments.
[0029] Figure 17 is a diagram illustrating the configuration of a dialogue screen according to various embodiments.
[0030] FIG. 18 is another diagram for explaining a response message generation operation of an electronic device according to various embodiments.
[0031] FIG. 19 is another diagram for explaining a response message generation operation of an electronic device according to various embodiments.
[0032] FIG. 20 is a diagram for explaining a response message generation procedure according to various embodiments.
[0033] FIG. 21 is another diagram illustrating a response message generation operation of an electronic device according to various embodiments.
[0034] FIG. 22 is a flowchart illustrating the operation of an electronic device according to various embodiments.
[0035] FIG. 23 is a flowchart illustrating the operation of an electronic device according to various embodiments.
[0036] FIG. 24 is a flowchart illustrating the operation of an electronic device according to various embodiments.
[0037] FIG. 25 is a block diagram of an exemplary electronic device capable of performing the operations described in this document.
[0038] Hereinafter, various embodiments of this document are described with reference to the attached drawings. However, this is not intended to limit the technology described in this document to specific embodiments, and it should be understood that various modifications, equivalents, and / or alternatives of the embodiments of this document are included. In connection with the description of the drawings, similar reference numerals may be used for similar components.
[0039]
[0040] FIG. 1 illustrates a block diagram of an electronic device according to one embodiment.
[0041] Referring to FIG. 1, an electronic device (100) according to various embodiments may include a processor (120), a memory (130), a display (140), a communication circuit (150), and an interface (160). Depending on the embodiment, the electronic device (100) may be implemented to have more or fewer components than the aforementioned components. For example, the electronic device (100) may include a configuration similar to the electronic device (2500) described below with reference to FIG. 25. For example, each of the processor (120), the memory (130), the display (140), and the communication circuit (150) illustrated in FIG. 1 may correspond to at least one processor (2510), at least one memory (2520), at least one display (2540), and at least one communication circuit (2560) of FIG. 25.
[0042] According to various embodiments, the processor (120) may be communicatively, electrically, operatively, or functionally connected to a memory (130), a display (140), a communication circuit (150), and / or an interface (160).
[0043] In various embodiments of the present disclosure, when one component is "operably" connected to another component, it can mean that the component is connected so that it can operate the other component. For example, the component can operate the other component by transmitting a control signal to the other component, either directly or via another component. In various embodiments of the present disclosure, when one component is "functionally" connected to another component, it can mean that the component is connected so that it can execute a function of the other component. For example, the component can execute a function of the other component by transmitting a control signal to the other component, either directly or via another component.
[0044] According to one embodiment, the processor (120) may include at least one processor. For example, the processor (120) may include an application processor (AP), a central processing unit (CPU), an image signal processor (ISP), a graphics processing unit (GPU), a neural processing unit (NPU), a tensor processing unit (TPU), and / or a communication processor (CP). The processor (120) may include at least one chip or one chipset. In the present disclosure, the processor (120) may be referred to as a hardware component having an architecture by at least one processing circuit. For example, the processor (120) may be mounted on a substrate (e.g., a printed circuit board) located inside the electronic device (100) and may communicate with other components of the electronic device (100) through at least one conductive path formed on the substrate.
[0045] According to various embodiments, the memory (130) may store instructions. When executed by the processor (120), the instructions may cause the electronic device (100) to perform various operations. For example, the instructions may be individually or collectively executed by at least one processor to cause the electronic device (100) to perform various operations. In various embodiments of the present disclosure, the operation of the electronic device (100) may be referred to as an operation performed by the processor (120) by executing instructions stored in the memory (130). The memory (130) may be referred to as a hardware component for data storage.
[0046] According to various embodiments, the display (140) may be configured to provide visual information (e.g., text, images, videos, icons, or symbols, etc.) to a user and receive user input (e.g., touch input). According to one embodiment, the display (140) may include multiple displays. For example, the display (140) may include a front display and / or a rear display.
[0047] According to various embodiments, the communication circuit (150) may support performing communication with an external device. According to one embodiment, the communication circuit (150) may be a device including hardware and software for transmitting and receiving signals (e.g., commands or data) between the electronic device (100) and the external device. For example, the communication circuit (150) may communicate with the external device via a first network (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)).
[0048] According to various embodiments, the interface (160) may include at least one device configured to receive input. In one embodiment, the interface (160) may include a touch circuit (161) configured to receive a touch input (e.g., a touch screen display). In one embodiment, the interface (180) may include at least one microphone (162) configured to receive a voice input. In one embodiment, the interface (160) may include at least one device for output. For example, the interface (160) may include a haptic module for tactile output, at least one speaker (163) for sound output, and / or an indicator. In one embodiment, the interface (180) may include a button (185). In this regard, the processor (120) may be configured to receive an input using the interface (160) and process the received input.
[0049] According to various embodiments, the electronic device (100) may provide a message function that outputs messages sent and received with a designated user in chronological order. This will be described in detail with reference to FIGS. 2 to 25 below. Furthermore, at least one of the various embodiments described with reference to FIGS. 2 to 25 below may be combined with other embodiments.
[0050]
[0051] FIG. 2 is a diagram illustrating a message function of an electronic device according to various embodiments. FIG. 3 is a diagram illustrating a dialogue screen output operation of an electronic device according to various embodiments.
[0052] As illustrated in 200 of FIG. 2, when a message function is executed, an electronic device (100) according to various embodiments can output a conversation screen (210) composed of a conversation area (211) and an input area (213).
[0053] According to one embodiment, the conversation area (211) may be an area where received messages (211-1) and transmitted messages (211-2) are output. For example, the received messages (211-1) and transmitted messages (211-2) output in the conversation area (211) may be output in chronological order in the form of a conversation.
[0054] According to one embodiment, the input area (213) refers to an area for entering a message to be sent to a designated user. For example, a virtual keyboard may be output in the input area (213). In addition, characters and / or combinations of characters selected from the virtual keyboard may be output in the input area (213).
[0055] According to an embodiment, the received message (211-1) may include a text message (211-1a), as illustrated. According to an embodiment, the received message (211-1) may also include various types of content, such as image data (211-1b), video data, and music data, as illustrated. However, this is merely an example, and various embodiments are not limited thereto. For example, the transmitted message (211-2) may also include various types of content, such as image data, video data, and music data, in addition to a text message.
[0056] The conversation screen (210) of the aforementioned message function may be output by an application framework (320) stored in an electronic device (100) (e.g., memory (130)), as illustrated in FIG. 3. For example, the application framework (320) may be a collection of classes or libraries usable in an application, and may be composed of software, firmware, hardware, or a combination of at least two or more of these.
[0057] Referring to FIG. 3, the application framework (320) can provide screen information (321) related to the conversation screen (210) based on an activity (311) transmitted from an application (310) (e.g., a message application).
[0058] According to one embodiment, the activity (311) transmitted from the application (310) may include at least one of execution of the application (310), reception of a message, transmission of a message, or termination of execution of the application (310).
[0059] For example, when an activity related to execution is received from an application (310), the application framework (320) can generate screen information (321) related to a conversation area (211) and an input area (213). For example, the application framework (320) can generate screen information (321) that defines the arrangement (e.g., size, position, and color) of the conversation area (211) and the input area (213).
[0060] In addition, when an activity related to receiving a message (or sending a message) is received from the application (310), the application framework (320) can generate screen information (321) related to the received message (211-1) (or the sent message (211-2)). For example, the application framework (320) can generate screen information (321) defining the arrangement of the received message (211-1) (or the sent message (211-2)) within the conversation area (211).
[0061] In addition, when an activity related to the termination of application execution is received from the application (310), the application framework (320) can generate screen information (321) that terminates the output of the dialogue area (211) and the input area (213).
[0062] Additionally or optionally, the electronic device (100) according to various embodiments may provide a response message generation function. The response message generation function may be a function in which the electronic device (100) automatically generates an appropriate response message (221) for a received message (e.g., the most recently received message) (211-1) and recommends it to the user, as illustrated in 220 of FIG. 2.
[0063] For example, the electronic device (100) may recognize image data (211-1b) received from a designated user as a single received message and recommend a response message (221) therefor. Accordingly, the user may select or ignore the recommended response message (221), and the electronic device (100) may transmit the response message (221) as a transmission message in response to detecting an input for selecting the recommended response message (221).
[0064] In this regard, an electronic device (100) according to various embodiments may generate a response message (221) using an artificial intelligence system. This will be described in detail with reference to FIGS. 4 and 5 below.
[0065]
[0066] Figure 4 illustrates a block diagram of an artificial intelligence system according to various embodiments.
[0067] Referring to FIGS. 1 and 4, according to one embodiment, an artificial intelligence system (400) may be configured to process received input using an artificial intelligence model and provide a result generated using the artificial intelligence model. In one example, the artificial intelligence system (400) may be implemented by an electronic device (100). For example, components of the artificial intelligence system (400) may be software modules (e.g., threads, functions, databases, and / or programs) that are implemented by executing instructions stored in the electronic device (100) (e.g., memory (130)).
[0068] In the example of FIG. 4, the artificial intelligence system (400) may include an I / O (input / output) interface (410), an artificial intelligence framework (420), a knowledge DB (460), an application (470), an artificial intelligence model DB (480), and / or a response message generation module (490).
[0069] According to various embodiments, the I / O (input / output) interface (410) may provide user input and / or context information to the artificial intelligence framework (420) and / or the response message generation module (490).
[0070] The user input may include a voice input (e.g., natural language input) obtained using a microphone (162). The user input may include a verbal input and / or a non-verbal input (e.g., an input for selecting a menu). The user input may include a text input obtained by performing a speech-to-text (STT) on the voice input, an input entered through a button (164), and / or a text input obtained through an interface such as a virtual keyboard. The user input may include an input obtained from an external device communicatively connected to the electronic device (100). The user input may include information obtained using a camera (not shown). The user input may include each of the above-described pieces of information or any combination of the above-described pieces of information.
[0071] Context information may include information related to the environment of the electronic device (100). For example, the context information may include a battery SoC (state of charge) of the electronic device (100), background sound acquired using a microphone (162), movement information of the electronic device (100) acquired using a motion sensor (not shown), and proximity information between the electronic device (100) and an external object. The context information may include information on an application currently running on the electronic device (100) and / or location information of the electronic device (100). The context information may include each of the above-described pieces of information or any combination of the above-described pieces of information.
[0072] According to various embodiments, the I / O (input / output) interface (410) may be configured to output results generated by an artificial intelligence model. For example, the I / O interface (410) may receive results from the AI framework (420). The I / O interface (410) may output the results in the form of natural language or content in any form. For example, the I / O interface (410) may visually output the results using the display (140). For example, the I / O interface (410) may audibly output the results using the speaker (163). For example, the I / O interface (410) may transmit the results to an external device using the communication circuit (150), thereby causing the external device to output the results.
[0073] According to various embodiments, the artificial intelligence framework (420) may be configured to receive user input and / or context information from the I / O interface (410), and control components to perform an action corresponding to an intention (e.g., performing a task) corresponding to a user query based on a user query included in the user input. For example, the artificial intelligence framework (420) may include a prompt manager (430), an application manager (440), and / or an output manager (450).
[0074] According to various embodiments, the prompt manager (430) may generate prompts for input into an artificial intelligence model from user input and / or context information. For example, the prompt manager (430) may process the user input and / or context information into prompts in a form for input into a large language model (LLM), a large multi-modal model (LMM), and / or a large vision model (LVM). For example, if the input is a voice input, text conversion and natural language understanding may be performed on the voice input. For example, if the input includes an image, feature extraction or object identification may be performed on the image. The prompt manager (430) may generate prompts using text information, feature information, and / or object information extracted from the voice input. In one example, the prompt manager (430) may utilize a trained machine learning algorithm or an artificial intelligence neural network to generate prompts.
[0075] According to various embodiments, the knowledge DB (460) may store history information related to the electronic device (100). For example, the knowledge DB (460) may store user preference data, a prompt library, and / or prompt example data generated based on user input. The prompt manager (430) may generate prompts from user input and / or context information using data stored in the knowledge DB (460). The prompt manager (430) may transmit the generated prompts to an artificial intelligence model (e.g., LLM, LVM, or LMM).
[0076] According to various embodiments, the application manager (440) may provide an interface between external components (e.g., knowledge DB (460), application (470), artificial intelligence model DB (480), and / or response message generation module (490)) and the artificial intelligence framework (420). For example, the application manager (440) may establish a channel between the external components and the artificial intelligence framework (420) using an application programming interface (API) or a plug-in.
[0077] For example, when user input is passed to an artificial intelligence model or when a prompt is generated, additional information may be requested. The application manager (440) can obtain the additional information by communicating with the application (470). The application manager (440) can provide a notification requesting additional information using the application (470) and can receive the additional information from the application (470). The application manager (440) can obtain the additional information by accessing an external database (e.g., a knowledge database (460). The application manager (440) can transmit the additional information to the prompt manager (430) and / or the artificial intelligence model.
[0078] In one example, a user input may include an intent to perform a specific task. In this case, the output generated by the AI model may include an action or a sequence of actions for performing the task corresponding to the intent. The application manager (440) may execute the action or sequence of actions included in the output using at least one application or service associated with the task.
[0079] According to various embodiments, the output manager (450) may perform fine-tuning on the output from the AI model. For example, fine-tuning may include tuning the output based on content policies (e.g., user terms of use, harmfulness policies, or prohibited content policies) and / or relevance (e.g., relevance between user input and the output).
[0080] For example, the output manager (450) may determine whether the output contains prohibited content (e.g., racially or politically biased content). For example, the output manager (450) may determine whether the output contains harmful content (e.g., content requiring age verification or content that violates public order and morals). For example, the output manager (450) may determine whether content that violates the Terms of Service exists. If prohibited content, harmful content, and / or content that violates the Terms of Service exists, the output manager (450) may exclude the content from the output or replace the content with other content. In one example, if prohibited content, harmful content, and / or content that violates the Terms of Service exists, the output manager (450) may provide a notification informing the user that the content cannot be provided. In one example, the output manager (450) may provide guidance information to the user to prevent prohibited content, harmful content, and / or content that violates the Terms of Service from being output.
[0081] For example, the output manager (450) can verify the correlation between the output and the user input (e.g., the intent of the user input). For example, the output manager (450) can utilize an artificial intelligence model to extract information about the output. The output manager (450) can identify the degree of correlation between the extracted information and the user input by comparing the extracted information with the user input. If the degree of correlation is below a threshold, the output manager (450) can perform additional actions. For example, the output manager (450) can modify the user input and provide the modified user input to the prompt manager (430). The artificial intelligence system (400) can process the prompt generated based on the modified user input using the artificial intelligence model, thereby generating an output that matches the intent of the user input.
[0082] According to various embodiments, the application (470) may include any application (e.g., application (310)) installed on the electronic device (100). The application (470) may act as a user interface (e.g., front-end) and / or a client-end for the artificial intelligence framework (420). The application (470) may provide a user interface for obtaining input. In one example, the application (470) may be configured to process operations directed by the application manager (440).
[0083] According to various embodiments, the artificial intelligence model DB (480) may store at least one artificial intelligence model. The at least one artificial intelligence model may include an LMM, an LVM, and / or an LMM. For example, the artificial intelligence model may include at least one generative artificial intelligence model. The artificial intelligence model may include an artificial intelligence model trained to generate images and / or language. For example, the artificial intelligence model may include a generative adversarial network (GAN), a variational autoencoder (VAE), and / or a diffusion-based generative model for generating images. The diffusion-based generative model may utilize a VAE and a transformer structure. For example, the artificial intelligence model may include a model (e.g., CHAT-GPT 3, CHAT-GPT 4) trained to generate statistical output values based on input values for generating text data.
[0084] According to various embodiments, the response message generation module (490) enables the electronic device (100) to generate an appropriate response message (221) for a received message (211-1).
[0085] According to one embodiment, the response message generation module (490) can generate a response message (221) using text messages (221-1a and 211-2) and image data (211-1b) exchanged with a designated user. For example, the response message generation module (490) can analyze text messages (221-1a and 211-2) exchanged with a designated user to determine the topic of the conversation and interpret the image data (211-1b) according to the topic of the conversation. In addition, the response message generation module (490) can input the topic of the conversation and the interpretation of the image data (211-1b) into at least one artificial intelligence model to generate a response message (221). This will be described in more detail with reference to FIGS. 7 to 21 below.
[0086]
[0087] Figure 5 illustrates the structure of an artificial intelligence model according to one embodiment.
[0088] Referring to FIG. 5, according to one embodiment, the artificial intelligence model DB (480) of FIG. 4 may include at least one artificial neural network model (500). The artificial neural network model (500) may include a plurality of hidden layers (520). The hidden layers (520) may be positioned between an input layer (510) and an output layer (530), and may include at least one layer learned while transmitting data (x1, x2, x3, ..., xn) (n is an integer greater than or equal to 4) transmitted from the input layer (510) to the output layer (530). For example, the hidden layers (520) may include a first hidden layer (520-1), a second hidden layer (520-2), and a third and Mth hidden layers (520-M) (M is an integer greater than or equal to 3). The number of hidden layers illustrated in FIG. 5 is merely an example, and embodiments of the present disclosure are not limited thereto. For example, the artificial neural network model (500) may include more hidden layers than the number of hidden layers illustrated in FIG. 5, or may include fewer hidden layers than the number of hidden layers illustrated in FIG. 5. In one example, the artificial neural network model (500) may correspond to a multi-layer perception (MLP) model. Each of the hidden layers (520) may include a plurality of nodes (N1, N2, ..., Nk). The weight values of each node may be learned using input data.
[0089] For example, the electronic device (100) of FIG. 1 can input the generated prompt as input data into an artificial neural network model (500). The input data is calculated by the artificial neural network model (500), and the artificial neural network model (500) can output output data (Y).
[0090] The artificial neural network model (500) is an example of an artificial intelligence model, and the embodiments of the present disclosure are not limited thereto. In the present disclosure, the term "artificial intelligence model" may include a generative artificial intelligence model. For example, the artificial intelligence model may include a large language model (LLM), a large multi-modal model (LMM), and / or a large vision model (LVM).
[0091] As described above with reference to FIGS. 4 and 5, the artificial intelligence system (400) according to various embodiments may be implemented by an electronic device (100). However, this is merely exemplary, and the various embodiments are not limited thereto.
[0092] According to an embodiment, a part of the artificial intelligence system (400) may be implemented by an external device (e.g., a server) (600), as illustrated in FIG. 6. In this case, the instruction database (460) and / or the artificial intelligence model database (480) may be implemented by the external device (600). In this case, the electronic device (100) may transmit and receive information by communicating with the external device (600) using the communication circuit (150).
[0093] As described above, the electronic device (100) according to various embodiments can process the received message (211-1) and the transmitted message (211-2) as inputs of an artificial intelligence model and obtain a response message (221) as a result. The obtaining (or generation) of the response message (221) will be described in more detail with reference to FIGS. 7 to 21 below. In addition, at least one of the various embodiments described with reference to FIGS. 7 to 21 below may be combined with other embodiments.
[0094]
[0095] FIG. 7 is a diagram illustrating an operation of generating a response message in an electronic device according to various embodiments. FIG. 8 is a diagram illustrating an operation of identifying conversation information and image data in a conversation screen according to various embodiments. FIG. 9 is a diagram illustrating an operation of generating a view hierarchy relationship (e.g., a view hierarchy diagram) for a conversation screen in an electronic device according to various embodiments.
[0096] Referring to FIG. 7, according to various embodiments, the application framework (320) (or electronic device (100)) may request the generation of a response message (730) (e.g., the response message (221) of FIG. 2) to the response message generation module (490).
[0097] According to one embodiment, the application framework (320) may provide, as part of such a request, conversation information (710) (e.g., text-based received message (211-1) and transmitted message (211-2)) and image data (720) (e.g., image data (211-1b)) included in the conversation screen (210) of the message function to the response message generation module (490). For example, the application framework (320) may select conversation information (710) and image data (720) to be provided to the response message generation module (490) based on the transmission and reception time of the message. For example, a certain number of conversation information (710) and image data (720) that have been recently transmitted and received in chronological order may be provided to the response message generation module (490).
[0098] In this regard, the application framework (320) can identify (or extract) conversation information (710) and image data (720) from the conversation screen (210) of the message function.
[0099] According to one embodiment, as illustrated in FIG. 8, when an activity (810) (e.g., message reception) is transmitted from an application (310) (e.g., message application), the application framework (320) can obtain (820) a conversation screen (210) (or a received message) of the message function and perform view analysis (830) based on the same. Through this view analysis, the application framework (320) can identify a view in the conversation screen (210) and, based on the view, identify conversation information (710) and image data (720) in the conversation screen (210).
[0100] For example, a view may include a text view (831) that displays text and an image view that displays an image. According to an embodiment, the application framework (320) may identify conversation information (710) based on the text view (831). For example, at least a portion of the text displayed in the text view (831) may be identified as conversation information (710). In addition, the framework (320) may identify image data (720) based on the image view (833). For example, at least a portion of the image displayed in the image view (833) may be identified as image data (720).
[0101] Additionally or optionally, the application framework (320) may provide image data (720) in string form to the response message generation module (490). In this regard, the application framework (320) may perform an operation of converting the identified image data (720) into a string (e.g., base64 conversion) (835).
[0102] However, this is merely an example, and various embodiments are not limited thereto. For example, the application framework (320) may analyze various views, such as a button view that includes input functions for event processing.
[0103] According to one embodiment, the application framework (320) may, as part of an operation of identifying conversation information (710) and image data (720), create a view hierarchy relationship (or screen hierarchy relationship) for the screen (210) of the message function, and based on this, identify the conversation information (710) and image data (720) in the screen (210) of the message function.
[0104] For example, the view hierarchy relationship may be information that expresses the hierarchical characteristics of the parent-child relationship by making each view included in the conversation screen (210) a node.
[0105] According to one embodiment, the conversation screen (210) may include a first view (910) (e.g., one view including a conversation area (211) and an input area (213), as shown in 900 of FIG. 9. In addition, the conversation screen (210) may include a second view (911) (e.g., a view including a conversation area (211)) and a third view (913) (e.g., a view including an input area (213)) displayed within the first view (910), and may include fourth views (911-1) to sixth views (911-3) (e.g., views including respective transmitted and received messages) displayed within the second view (911).
[0106] In this regard, the electronic device (100) can create a view hierarchy relationship (screen hierarchy relationship) for the conversation screen (210), as shown in 930 of FIG. 9.
[0107] For example, a hierarchical relationship can be created indicating that a first view (910) (e.g., a first node) includes a second view (911) (e.g., a second node) and a third view (913) (e.g., a third node) as children, and that the second view (911) includes (e.g., is dependent on) a fourth view (911-1) (e.g., a fourth node) to a sixth view (911-3) (e.g., a sixth node).
[0108] According to one embodiment, the electronic device (100) may generate view information (e.g., first view information to sixth view information (923)) for each view (e.g., first view (911) to sixth view (911-3)) included in the generated hierarchical relationship. However, this is merely exemplary, and various embodiments are not limited thereto. For example, as illustrated, the electronic device (100) may generate view information (e.g., fourth view information (921), fifth view information (922), sixth view information (923)) for views belonging to the lowest layer (e.g., fourth view (911-1), fifth view (911-2), and sixth view (911-3)).
[0109] According to one embodiment, the view information (e.g., fourth view information (921), fifth view information (922), sixth view information (923)) may include identification information and arrangement (e.g., size, position, and color) information for each view (e.g., fourth view (911-1), fifth view (911-2), and sixth view (911-3)).
[0110] For example, the identification information may include identification information for each view, identification information for a view in a higher layer that each view depends on, and identification information indicating the type of each view. For example, the identification information related to the fourth view (911-1) may include the ID of the fourth view (911-1), the ID of the second view (911) that the fourth view (911-1) depends on, and an ID indicating that the fourth view (911-1) is a text view. Similarly, the identification information related to the sixth view (911-3) may include the ID of the sixth view (911-3), the ID of the second view (911) that the sixth view (911-3) depends on, and an ID indicating that the sixth view (911-3) is an image view.
[0111] For example, the array information may include coordinate information indicating the location of each view, size information indicating the length and width of each view, and color information indicating the background color of each view.
[0112] According to one embodiment, the application framework (320) can identify (or extract) conversation information (710) and image data (720) from the conversation screen (210) based on at least view information for each view.
[0113]
[0114] As described above with reference to FIG. 8, the application framework (320) according to various embodiments can identify conversation information (710) and image data (720) at the time when the activity (810) is transmitted from the application (310). However, this is merely exemplary, and various embodiments are not limited thereto. For example, the application framework (320) can identify conversation information (710) and image data (720) at specified intervals regardless of the transmission of the activity (810).
[0115] According to various embodiments, the response message generation module (490) may generate a response message (730) based on conversation information (710) and image data (720) provided from the application framework (320) and provide the same to the application framework (320).
[0116] According to one embodiment, the response message generation module (490) may analyze conversation information (710) to confirm the topic (or content) of the conversation when generating a response message, and may analyze image data (720) to identify an object included in the image data (720). For example, the response message generation module (490) may identify the type of object included in the image data (720) and associate it with the topic of the conversation to suggest an appropriate response message (730).
[0117] As described above, when an electronic device (100) according to various embodiments receives image data (720) in the process of exchanging messages with a designated user, it can suggest an appropriate response message (730) by associating an object included in the image data (720) with the topic of the conversation.
[0118] Additionally or optionally, the electronic device (100) according to various embodiments may not simply identify the type of object included in the image data (720), but may also obtain descriptive information describing the situation, behavior, psychology, etc. of the object included in the image data (720) based on the topic of the conversation. This descriptive information may be utilized in generating a response message (730), thereby suggesting a response message with a higher relevance to the topic of the conversation. This will be described in more detail with reference to FIGS. 10 to 13 below.
[0119]
[0120] FIG. 10 is a diagram illustrating an operation of obtaining description information for image data in an electronic device according to various embodiments. FIG. 11 is a diagram illustrating description information according to various embodiments.
[0121] Referring to FIG. 10, the response message generation module (490) (or the electronic device (100)) may obtain description information (1020) regarding image data (720) using a first artificial intelligence (AI) model and use the same to generate a response message. According to one embodiment, the first AI model (1000) may be a large multi-modal model (LMM). In addition, the description information (1020) may be text data describing an object included in the image data (720). For example, the description information (1020) may be information describing a situation, behavior, psychology, etc., regarding an object included in the image data (720).
[0122] In this regard, the response message generation module (490) may provide the image data (720) identified in the dialogue screen (210) of the message function as an input to the first AI model (1000). For example, the identified image data (720) and a prompt (1010) requesting description information (1020) therefor may be provided as an input to the first AI model (1000). Depending on the embodiment, the response message generation module (490) may provide a prompt requesting simple description information (1020) that briefly describes the image data (720), detailed description information (1020) that more specifically describes the image data (720), or both, as an input to the first AI model (1000).
[0123] According to various embodiments, the first AI model (1000) may output description information (1020) about image data (720) based on input provided from the response message generation module (490). For example, the first AI model (1000) may analyze the image data (720) provided from the response message generation module (490) to generate description information (1020) that describes the external characteristics of an object included in the image data (720).
[0124] According to one embodiment, for the image data (720) (e.g., an image of a donut and coffee) shown in 1100 of FIG. 11, the first AI model (1000) may generate description information (1020) (or text data) expressing the external features of the donut and coffee included in the image data (720) in text format. In this regard, the first AI model (1000) may generate description information (1020) including at least one of simple description information (e.g., description information describing the donut and coffee included in the image data (720) as a dessert) (1111) that briefly describes the image data (720), as shown in 1110 of FIG. 11, and detailed description information (e.g., description information describing the donut and coffee included in the image data (720) as food arranged on a wooden tray) (1113) that more specifically describes the image data (720).
[0125] Additionally or optionally, the first AI model (1000) according to various embodiments may generate description information (1020) that interprets image data (720) to fit the topic of the conversation, rather than simply generating description information (1020) that describes the external features of an object.
[0126] In this regard, the response message generation module (490) according to various embodiments may also provide the conversation information (710) identified in the conversation screen (210) as input to the first AI model (1000) as part of the operation of requesting description information (1020). For example, messages transmitted and received by the electronic device (100) over a certain period of time may be input to the AI model (1000). For example, a certain number of messages recently transmitted and received in chronological order may be input to the AI model (1000).
[0127] Accordingly, the first AI model (1000) can identify the topic of the conversation based on the conversation information (710) provided by the response message generation module (490) and use it to generate description information (1020) for the image data (720).
[0128] For example, with respect to the conversation information (710) illustrated in 1100 of FIG. 11, the first AI model (1000) can identify that a conversation related to fatigue is taking place. In this regard, the AI model (1000) determines that the topic of the conversation is related to fatigue, and thus, as illustrated in 1120 of FIG. 11, can generate description information (e.g., description information describing donuts and coffee included in the image data (720) as foods that can help relieve fatigue) (1020) that interprets the objects included in the image data (720) to fit the topic of the conversation.
[0129] As described above, the response message generation module (490) according to various embodiments can obtain description information (1020) regarding image data (720) using the first AI model (1000) and use this to generate a response message. This will be described in more detail with reference to FIG. 12 below.
[0130]
[0131] FIG. 12 is another diagram illustrating a response message generation operation of an electronic device according to various embodiments.
[0132] Referring to FIG. 12, according to various embodiments, the application framework (320) (or electronic device (100)) may request the generation of a response message (730) (e.g., the response message (221) of FIG. 2) to the response message generation module (490).
[0133] According to one embodiment, as part of such a request, the application framework (320) may provide description information (1020) regarding the conversation information (710) and image data (720) included in the conversation screen (210) of the message function to the response message generation module (490). For example, the application framework (320) may obtain the description information (1020) through the first AI model.
[0134] According to various embodiments, the response message generation module (490) may request a response message from the second AI model (1200) based on a request from the application framework (320). The second AI model (1200) may be a large language model (LLM). However, this is merely an example, and various embodiments are not limited thereto. For example, the second AI model (1200) may be a large multi-modal model (LMM).
[0135] According to one embodiment, the response message generation module (490) may provide conversation information (710) and description information (1020) as inputs to the second AI model (1200). Accordingly, the second AI model (1200) may generate a response message (1210) based on the conversation information (710) and description information (1020) and provide the same to the application framework (320) through the response message generation module (490). According to an embodiment, the second AI model (1200) may be the same AI model as the first AI model (1000) described above with reference to FIG. 10. According to an embodiment, the second AI model (1200) may be a different AI model from the first AI model (1000) described above with reference to FIG. 10.
[0136]
[0137] FIGS. 13A and 13B are diagrams illustrating a series of operations for providing a response message in an electronic device according to various embodiments.
[0138] Referring to FIGS. 13A and 13B , an electronic device (100) according to various embodiments may transmit and receive messages with a designated user. In this regard, the electronic device (100) may output a conversation screen (210) consisting of a conversation area (211) and an input area (213), as illustrated in 1300 of FIG. 13A .
[0139] According to various embodiments, the electronic device (100) may generate a response message based on the text-type transmitted and received messages and image data (720) displayed in the conversation area (211). According to one embodiment, the electronic device (100) may automatically generate an appropriate response message (1311) for a received message (e.g., the most recently received message) and recommend it to the user, as illustrated in 1310 of FIG. 13A.
[0140] According to various embodiments, the electronic device (100) may identify a response message to be transmitted to a designated user among the recommended response messages (1311). According to one embodiment, the electronic device (100) may detect an input for selecting one of the recommended response messages (1311), as illustrated in 1320 of FIG. 13b.
[0141] According to various embodiments, the electronic device (100) may, in response to detecting an input for selecting a recommended response message (221), transmit the selected response message (221) as a transmission message, as illustrated in 1330 of FIG. 13b.
[0142]
[0143] FIGS. 14A and 14B are diagrams illustrating an electronic device supporting a response message generation function according to various embodiments.
[0144] Referring to FIGS. 14a and 14b, the response message generation function according to various embodiments may be provided by various types of electronic devices.
[0145] According to one embodiment, the response message generation function may be provided by a watch-type wearable device, as illustrated in 1410 of FIG. 14A.
[0146] According to one embodiment, the response message generation function may be provided by a smartphone having various form factors, as illustrated in 1420 of FIG. 14b.
[0147] As described above, the electronic device (100) according to various embodiments may generate a response message (1210) based on the description information (1020) and the conversation information (710) regarding the image data (720). Additionally or alternatively, the electronic device (100) according to various embodiments may also generate the response message (1210) by further considering the conversation style between users. This will be described in more detail with reference to FIG. 15 below.
[0148]
[0149] FIG. 15 is another diagram illustrating a response message generation operation of an electronic device according to various embodiments. FIG. 16a and FIG. 16b are diagrams illustrating a response message generated according to a conversation style between users according to various embodiments.
[0150] Referring to FIG. 15, according to various embodiments, the application framework (320) may request the generation of a response message (1210) (e.g., the response message (221) of FIG. 2) to the response message generation module (490).
[0151] According to one embodiment, as part of such a request, the application framework (320) may provide description information (1020) about the conversation information (710) and image data (720) included in the conversation screen (210) of the message function to the response message generation module (490).
[0152] Additionally, the application framework (320) may provide conversation style information indicating conversation styles between users to the response message generation module (490). According to one embodiment, the conversation style may be associated with a social relationship between users. The social relationship may include a first relationship that uses informal speech toward a designated user (e.g., a counterpart user) and a second relationship that uses formal speech toward the designated user.
[0153] In this regard, a conversation analysis module (1500) configured to analyze a conversation screen (210) output by an application framework (320) to identify social relationships between users and provide the same to the application framework (320) may be provided as a component of the electronic device (100). However, this is merely exemplary, and various embodiments are not limited thereto. For example, the conversation style may be related to whether it is a daily conversation or a business conversation, and depending on the embodiment, the conversation analysis module (1500) may also provide the tone of the conversation (e.g., serious, sad, happy, excited) as conversation style information.
[0154] According to various embodiments, the response message generation module (490) may request a response message from the second AI model (1200) based on a request from the application framework (320).
[0155] According to one embodiment, the response message generation module (490) may provide conversation information (710), description information (1020), and conversation style information (1501) as inputs to the second AI model (1200). Accordingly, the second AI model (1200) may generate a response message (1210) based on the conversation information (710), description information (1020), and conversation style information (1501), and provide the same to the application framework (320) through the response message generation module (490).
[0156] For example, as illustrated in 1610 of FIG. 16A, when a transmitted and received message of a conversation screen (210) includes an expression of polite speech, conversation style information (1501) indicating a first relationship between users that uses an expression of polite speech may be provided as an input to a second AI model (1200). Accordingly, the second AI model (1200) according to various embodiments may generate a response message (1210) corresponding to the first relationship (e.g., a response message expressed in polite speech), as illustrated in 1611 of FIG. 16A.
[0157] As another example, as illustrated in 1620 of FIG. 16b, when a transmitted and received message of a conversation screen (210) includes an expression of polite speech, conversation style information (1501) indicating a second relationship between users using an expression of polite speech may be provided as an input to a second AI model (1200). Accordingly, the second AI model (1200) according to various embodiments may generate a response message (1210) corresponding to the second relationship (e.g., a response message expressed in polite speech), as illustrated in 1621 of FIG. 16b.
[0158] As described above, the electronic device (100) according to various embodiments may utilize description information (1020), conversation information (710), and conversation style information (1501) regarding image data (720) to generate a response message (1210). Additionally or alternatively, the electronic device (100) according to various embodiments may also generate the response message (1210) by additionally considering information related to a designated user (e.g., a counterpart user). This will be described in more detail with reference to FIG. 17 below.
[0159]
[0160] Figure 17 is a diagram illustrating the configuration of a dialogue screen according to various embodiments.
[0161] Referring to FIG. 17, a conversation screen (210) according to various embodiments may be composed of a conversation area (211) and an input area (213). In addition, the conversation screen (210) may include identification information (1710) about the other user with whom messages are sent and received.
[0162] According to one embodiment, the identification information (1710) for the counterpart user may include the counterpart user's phone number, name or initials stored in the electronic device (100).
[0163] In this regard, the electronic device (100) according to various embodiments can utilize identification information (1710) about the other user to generate a response message (1210).
[0164] According to an embodiment, the electronic device (100) can identify a social relationship between users based on identification information (1710) about the other user, and can generate a response message (1210) corresponding to the identified social relationship.
[0165] According to an embodiment, the electronic device (100) can determine the intimacy between users based on identification information (1710) about the other user, and can generate a response message (1210) corresponding to the determined intimacy.
[0166] As described above, the electronic device (100) according to various embodiments may provide a content-type message (e.g., image data (720)) included in a transmitted and received message as input to an AI model (e.g., a first AI model (1000) and a second AI model (1200)) to generate a response message. Additionally or alternatively, the electronic device (100) according to various embodiments may convert a content-type message into a text format and then provide the converted message as input to the AI model. This will be described in more detail with reference to FIG. 18 below.
[0167]
[0168] FIG. 18 is another diagram for explaining a response message generation operation of an electronic device according to various embodiments.
[0169] Referring to FIG. 18, according to various embodiments, the application framework (320) (or electronic device (100)) may request the generation of a response message (730) (e.g., the response message (221) of FIG. 2) to the response message generation module (490).
[0170] According to one embodiment, the application framework (320) may, as part of such a request, provide a message (1810) in the form of content included in a conversation screen (210) of a message function to the response message generation module (490). For example, the message (1810) in the form of content may include at least one of image data (720) output through a conversation area (211) of the conversation screen (210), uniform resource locator (URL) information, or a music file attached to the conversation area (211).
[0171] According to various embodiments, the response message generation module (490) may obtain a text message (1803) for a content message (1810). According to one embodiment, the response message generation module (490) may provide the content message (1810) to the first AI model (1000) and obtain a text message (1803) in response thereto. For example, the first AI model (1000) may generate the text message (1803) by converting the content message (1810) into a string (e.g., base64 conversion). According to an embodiment, the text message (1803) may include description information (1020) about the content message (1810).
[0172] According to various embodiments, the response message generation module (490) may provide the acquired text message (1803) to the second AI model (1200), and obtain a response message (1805) in response thereto. According to one embodiment, the response message generation module (490) may provide the text message (1803) to the second AI model (1200) together with a prompt requesting generation of the response message. Depending on the embodiment, the prompt and the text message (1803) may be provided as one combined.
[0173] For example, if the content-type message (1803) is image data (720), the response message generation module (490) may provide the text-type message (1803) along with a prompt requesting generation of a response message (1805) for the image data (720).
[0174] For example, if the content-type message (1803) is a music file, the response message generation module (490) may provide a text-type message (1803) along with a prompt requesting generation of a response message (1805) related to the music file.
[0175] For example, if the content-type message (1803) is URL information, the response message generation module (490) may provide a text-type message (1803) along with a prompt requesting the generation of a response message (1805) related to the URL information. In this regard, the response message generation module (490) may collect additional information based on the URL information and utilize it to generate the response message (1805). This will be described in more detail with reference to FIGS. 19 and 20 below.
[0176]
[0177] FIG. 19 is another diagram illustrating a response message generation operation of an electronic device according to various embodiments. FIG. 20 is a diagram illustrating a response message generation procedure according to various embodiments.
[0178] Referring to FIG. 19, an electronic device (100) according to various embodiments may transmit and receive messages with a designated user. In this regard, the electronic device (100) may output a conversation screen (210) consisting of a conversation area (211) and an input area (213), as illustrated in 1910 of FIG. 19. According to one embodiment, the conversation area may include URL information (1911).
[0179] In this regard, the electronic device (100) according to various embodiments can perform view analysis on the conversation screen (210).
[0180] According to one embodiment, as illustrated in FIG. 20, the electronic device (100) can identify a text view (2001) displaying text through view analysis, and then prepare (2011) for identifying URL information included in the text view (2001).
[0181] According to one embodiment, when identification of URL information is prepared (2011), the electronic device (100) can identify the URL information included in the text view (2001) through text analysis (2003) and then prepare for information retrieval (2013).
[0182] According to one embodiment, when information retrieval is prepared (2013), the electronic device (100) can obtain information necessary for generating a response message through information retrieval (2005) (e.g., URL access).
[0183] According to one embodiment, when information acquisition is complete, the electronic device (100) may provide (2015) a prompt to the AI model (2007) requesting the generation of the retrieved information (URL content) and a response message (1805) therefor.
[0184] Accordingly, the AI model (2007) generates summary information (2017) on the searched information, and the electronic device (100) can output it (2019) through a dialogue screen (e.g., text view).
[0185] As described above, the electronic device (100) according to various embodiments may generate a response message (1805) using multiple different AI models. Additionally or alternatively, the electronic device (100) according to various embodiments may also generate a response message (1805) using a single AI model. This will be described in more detail with reference to FIG. 21 below.
[0186]
[0187] FIG. 21 is another diagram illustrating a response message generation operation of an electronic device according to various embodiments.
[0188] Referring to FIG. 21, according to various embodiments, the application framework (320) (or electronic device (100)) may request the generation of a response message (730) (e.g., the response message (221) of FIG. 2) to the response message generation module (490).
[0189] According to one embodiment, the application framework (320) may, as part of such a request, provide a message (1810) in the form of content included in a conversation screen (210) of a message function to the response message generation module (490). For example, the message (1810) in the form of content may include at least one of image data (720) output through a conversation area (211) of the conversation screen (210), uniform resource locator (URL) information, or a music file attached to the conversation area (211).
[0190] According to various embodiments, the response message generation module (490) may obtain a text message (1803) for a content message (1810). According to one embodiment, the response message generation module (490) may generate a text message (2103) by converting the content message (1810) into a string (e.g., base64 conversion) (2101). According to an embodiment, the text message (2103) may include description information (1020) for the content message (1810).
[0191] According to various embodiments, the response message generation module (490) may provide the acquired text-based message (2103) to the AI model (2100), and obtain a response message (1805) in response thereto. According to one embodiment, the response message generation module (490) may provide the text-based message (2103) to the AI model (2100) along with a prompt requesting generation of the response message (1805). Accordingly, the AI model (2100) may analyze the text-based message (2103) to generate the response message (1805). For example, the AI model (2100) may include a large vision model (LVM).
[0192]
[0193] An electronic device (100) according to various embodiments may include at least one processor (120), a display (140), a communication circuit (150), and a memory (130) operatively connected to the at least one processor (120), the display (140), and the communication circuit (150) and storing at least one command. According to one embodiment, the at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: output, through the display (140), a conversation screen (210) on which a received message (211-1) including at least one text (211-1a) and at least one image (211-1b) received from a designated user through the communication circuit (150) is displayed, identify image data (720) for the at least one image (211-b) displayed as at least a part of the received message (211-1) on the conversation screen (210), obtain text data (1020) describing an object included in the image data (720) based on the identified image data (720) and a designated first artificial intelligence (AI) model (1000), and And, it may be configured to generate a response message (730, 1210) for the received message (211-1) based on the acquired text data (1020) based on the image data (720) and the designated second AI model (1200).
[0194] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: display a transmission message transmitted to the designated user via the communication circuit (150) through the conversation screen (210), and obtain the text data (1020) based on at least a portion of the transmission message, the identified image data (720), and the designated first AI model (1000).
[0195] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: identify a subject of a message to be transmitted or received with the designated user based on at least a portion of the transmitted message, and obtain the text data (1020) corresponding to the identified subject.
[0196] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: display a transmission message transmitted to the designated user via the communication circuit (150) via the conversation screen (210), and generate a response message (730, 1210) for the received message (211-1) based on at least a portion of the transmission message, the acquired text data (1020), and the designated second AI model (1200).
[0197] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: display a transmission message transmitted to the designated user through the communication circuit (150) through the conversation screen (210), analyze a conversation style based on the transmission message (211-2) and the received message (211-1) displayed on the conversation screen (210), and generate a response message (730, 1210) for the received message (211-1) based on information (1501) related to the conversation style, at least a portion of the transmission message (211-2), the acquired text data (1020), and the designated second AI model (1200).
[0198] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: if the conversational style is associated with the expression of honorifics, express the response message (730, 1210) in honorifics; and if the conversational style is associated with the expression of informal speech, express the response message (730, 1210) in informal speech.
[0199] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: identify at least one view constituting the conversation screen, and identify the image data (720) based on the identified at least one view.
[0200] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: identify the at least one view based on a screen hierarchy relationship with respect to the conversation screen.
[0201] According to various embodiments, the at least one instruction, when individually or collectively executed by the at least one processor (120), may be configured to cause the electronic device (100) to: transmit the generated response message (730, 1210) to the designated user.
[0202] According to various embodiments, the first AI model (1000) specified above may include a large multi-modal model (LMM), and the second AI model (1200) specified above may include a large language model (LLM).
[0203]
[0204] Figure 22 is a flowchart illustrating the operation of an electronic device according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0205] Referring to FIG. 22, an electronic device (100) according to various embodiments may, in operation 2210, output a conversation screen (210) including transmitted and received messages (211-1, 211-2) and image data (720). In this regard, the electronic device (100) may output the transmitted and received messages (211-1, 211-2) in a conversation format based on the transmission and reception time.
[0206] According to various embodiments, an electronic device (100) may, in operation 2220, obtain description information (1020) regarding image data (720) using transmitted and received messages (211-1, 211-2), image data (720), and an AI model (e.g., a first AI model (1000)). The description information (1020) may be information describing a situation, behavior, psychology, etc. regarding an object included in the image data (720). In this regard, reference may be made to the description related to FIG. 10 described above.
[0207] According to various embodiments, the electronic device (100) may, in operation 2230, generate a response message for image data based on the description information (1020). According to one embodiment, the electronic device (100) may recognize image data received from a designated user as a single received message and generate a response message therefor. In this regard, reference may be made to the description related to FIG. 12 described above.
[0208]
[0209] Figure 23 is a flowchart illustrating the operation of an electronic device according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0210] Referring to FIG. 23, an electronic device (100) according to various embodiments may, in operation 2310, obtain hierarchical information regarding a conversation screen (210). According to one embodiment, the hierarchical information may be information expressing the hierarchical characteristics of a parent-child relationship by making each view included in the conversation screen a node. In this regard, reference may be made to the descriptions related to FIGS. 8 and 9 described above.
[0211] According to various embodiments, the electronic device (100) may, in operation 2320, obtain image data from a conversation screen (210) based on layer information. In this regard, the electronic device (100) may identify a text view (831) displaying text and an image view displaying an image on the conversation screen (210). According to one embodiment, the electronic device (100) may obtain at least a portion of an image displayed in the image view (833) as image data (720).
[0212] An electronic device (100) according to various embodiments may, in operation 2330, obtain description information (1020) regarding image data using the image data and an AI model. The description information (1020) may be information describing the situation, behavior, psychology, etc. of an object included in the image data (720). In this regard, reference may be made to the description related to FIG. 10 described above.
[0213] According to various embodiments, an electronic device (100) may, in operation 2340, identify a conversation style (1501) using transmitted and received messages and an AI model. According to one embodiment, the conversation style (1501) may be associated with social relationships between users. For example, the electronic device (100) may determine whether a conversation between users uses polite or polite language.
[0214] An electronic device (100) according to various embodiments may, in operation 2350, generate a response message (1210) based on description information (1020) and a conversation style (1501). In this regard, reference may be made to the description related to FIG. 15 described above.
[0215]
[0216] Figure 24 is a flowchart illustrating the operation of an electronic device according to various embodiments. While the operations in the following embodiments may be performed sequentially, they are not necessarily performed sequentially. For example, the order of the operations may be changed, and at least two operations may be performed in parallel. Furthermore, at least one of the aforementioned operations may be omitted depending on the embodiment.
[0217] Referring to FIG. 24, an electronic device (100) according to various embodiments may output a conversation screen (210) including transmitted and received messages and image data (720) in operation 2410. In this regard, the electronic device (100) may output transmitted and received messages in a conversation format based on the transmission and reception time.
[0218] An electronic device (100) according to various embodiments may, in operation 2420, obtain description information (1020) regarding image data (720) using transmitted and received messages, image data (720), and an AI model (e.g., a first AI model (1000)). The description information (1020) may be information describing a situation, behavior, psychology, etc. regarding an object included in the image data (720). In this regard, reference may be made to the description related to FIG. 10 described above.
[0219] According to various embodiments, the electronic device (100) can, in operation 2430, identify a conversation style (1501) using transmitted and received messages and an AI model. According to one embodiment, the conversation style (1501) may be associated with social relationships between users. For example, the electronic device (100) can determine whether a conversation between users uses polite or polite language.
[0220] According to various embodiments, the electronic device (100) may, in operation 2440, generate a response image (1210) for image data based on description information (1020) and conversation style information (1501). According to one embodiment, the electronic device (100) may generate a response message in the form of an image instead of a response message in the form of text. In this regard, the electronic device (100) may provide the description information (1020) and conversation style information (1501) as inputs to a generative adversarial network (GAN), and obtain a response image as an output.
[0221]
[0222] FIG. 25 is a block diagram of an exemplary electronic device (2500) capable of performing the operations described in this document.
[0223] Referring to FIG. 25, the electronic device (2500) may be one of various forms of electronic devices, such as a notebook (2590), smartphones (2591) having various form factors (e.g., a bar-type smartphone (2591-1), a foldable-type smartphone (2591-2), or a sliderable (or rollable) type smartphone (2591-3)), a tablet (2592), a cellular phone (not shown), and other similar computing devices (not shown). The components, their relationships, and their functions illustrated in FIG. 1 are exemplary only and do not limit the implementations described or claimed in this document. The electronic device (2500) may be referred to as a mobile device, a user device, a multi-function device, a portable device, or a server.
[0224] The electronic device (2500) may include components including at least one processor (2510) (hereinafter referred to as processor (2510)), at least one memory (2520) (hereinafter referred to as memory (2520)), at least one display (2540) (hereinafter referred to as display (2540)), at least one image sensor (2550) (hereinafter referred to as image sensor (2550)), at least one communication circuit (2560) (hereinafter referred to as communication circuit (2560)), and / or at least one sensor (2570) (hereinafter referred to as sensor (2570)). The above components are merely exemplary. For example, the electronic device (2500) may include other components (e.g., power management integrated circuitry (PMIC), audio processing circuitry, an antenna, a rechargeable battery, or an input / output interface). For example, some components may be omitted from the electronic device (2500). For example, several components can be combined into one component.
[0225] The processor (2510) may be implemented as one or more IC (integrated circuit (or circuitry)) chips and may perform various data processing. The processor (2510) may include at least one electrical circuit and may individually or collectively perform distributed processing of instructions (or programs, data, etc.) stored in the memory (2520). The processor (2510) may include a processor assembly including one or more processing circuits. The processor (2510) may include any processing circuit operative to control the performance and operations of one or more components of the electronic device (2500) (e.g., the memory (2520), the display (2540), the image sensor (2550), the communication circuit (2560), and / or the sensor (2570)). For example, the processor (2510) (e.g., an application processor (AP)) may be implemented as a system on chip (SoC) (e.g., a single chip or chipset). For example, the processor (2510) may be implemented as multiple cores (or at least one core circuit), multiple chips, or multiple chipsets. For example, the processor (2510) may include one or more processing circuits. For example, the processor (2510) may include one or more processing circuits configured to individually and / or collectively perform various functions of the present disclosure. As a non-limiting example, at least a portion of the processor (2510) may be included in a first chip of the electronic device (2500), and at least another portion of the processor (2510) may be included in a second chip of the electronic device (2500) that is different from the first chip of the electronic device (2500).
[0226] For example, the processor (2510) may include a central processing unit (CPU) (2511), a graphics processing unit (GPU) (2512), a neural processing unit (NPU) (2513), an image signal processor (ISP) (2514), a display controller (2515), a memory controller (2516), a storage controller (2517), a communication processor (CP) (2518), and / or a sensor interface (2519). These components of the processor (2510) are merely exemplary. For example, the processor (2510) may further include other components. For example, some components of the processor (2510) may be omitted from the processor (2510). For example, some components of the processor (2510) may be included as separate components of the electronic device (2500) outside the processor (2510). For example, some components of the processor (2510) (e.g., memory controller (2516)) may be included within other components (e.g., at least a portion of memory (2520), an interface (e.g., available for connection to at least one component of the electronic device (2500)), a display (2540) and / or an image sensor (2550)).
[0227] The processor (2510) may cause other components of the electronic device (2500) to perform various operations by executing instructions stored in the memory (2520). The CPU (2511) (or central processing circuit) may be configured to control components of the processor (2510) based on the execution of instructions stored in the memory (2520) (e.g., volatile memory (2521) and / or non-volatile memory (2522)). The GPU (2512) (or graphics processing circuit) may be configured to execute parallel operations (e.g., rendering). The NPU (2513) (or neural processing circuit, or artificial intelligence (AI) chip) may be configured to execute operations for an artificial intelligence model (e.g., convolution computation). The ISP (2514) (or image signal processing circuit) may be configured to process a raw image acquired through the image sensor (2550) into a format suitable for a component within the electronic device (2500) or a component of the processor (2510). The display controller (2515) (or display control circuit, or display processing unit (DPU)) may be configured to process an image acquired from the CPU (2511), the GPU (2512), the ISP (2514), or the memory (2520) (e.g., the volatile memory (2521)) into a format suitable for the display (2540). The memory controller (2516) (or memory control circuit) may be configured to control reading data from the volatile memory (2521) and writing data to the volatile memory (2521). The storage controller (2517) (or storage control circuit) may be configured to control reading data from and writing data to the nonvolatile memory (2522).The CP (2518) (communication processing circuit) may be configured to process data obtained from a component of the processor (2510) into a format suitable for transmission to another electronic device via the communication circuit (2560), or to process data obtained from another electronic device via the communication circuit (2560) into a format suitable for processing by the component of the processor (2510). For example, the communication circuit (2560) may include one or more communication circuits. The sensor interface (2519) (or sensing data processing circuit, sensor hub) may be configured to process data on the state of the electronic device (2500) and / or the state of the surroundings of the electronic device (2500), obtained via the sensor (2570), into a format suitable for the component of the processor (2510).
[0228] The memory (2520) may include one or more storage media (or one or more storage devices). For example, the memory (2520) may include a memory assembly including one or more storage media. For example, the one or more storage media may include permanent memory (e.g., non-volatile memory (2522)) such as a hard drive, flash memory, read-only memory (ROM), semi-permanent memory (e.g., volatile memory (2521)) such as random access memory (RAM), any other suitable type of storage (or storage assembly), or any combination thereof. The memory (2520) may include cache memory, which is one or more different types of memory used to temporarily store data for a function or feature of the electronic device (2500). As a non-limiting example, the cache memory may be included within the processor (2510). The memory (2520) may be fixedly embedded within the electronic device (2500) or incorporated into one or more suitable types of components (e.g., a subscriber identity module (SIM) card and / or a secure digital (SD) card) that may be repeatedly inserted into and removed from the electronic device (2500).
[0229] For example, the memory (2520) may store one or more software applications, such as an operating system (or system) software application, a firmware software application, a driver software application, a plug-in (e.g., add-in, add-on, and / or applet) software application, and / or any other suitable software applications. For example, the one or more software applications may include instructions executable by the processor (2510). For example, the memory (2520) may store instructions callable by an application programming interface (API). For example, the memory (2520) may store instructions within a library.
[0230]
[0231] An operating method of an electronic device (100) according to various embodiments includes an operation of outputting, through a display (140), a conversation screen (210) in which a received message (211-1) including at least one text (211-1a) and at least one image (211-1b) received from a designated user is displayed, an operation of identifying image data (720) for the at least one image (211-1b) displayed as at least a part of the received message (211-1) on the conversation screen (210), an operation of obtaining text data (1020) describing an object included in the image data (720) based on the identified image data (720) and a designated first artificial intelligence (AI) model (1000), and an operation of obtaining text data (1020) based on the at least one text included in the received message (211-1) and the image data (720) and a designated second AI model (1200). It may include an action of generating a response message (730, 1210) to a message (211-1).
[0232] According to various embodiments, the operating method of the electronic device (100) may include an operation of displaying a transmission message transmitted to the designated user through the conversation screen (210) and an operation of obtaining the text data (1020) based on at least a portion of the transmission message, the identified image data (720) and the designated first AI model (1000).
[0233] According to various embodiments, the operating method of the electronic device (100) may include an operation of identifying a subject of a message to be transmitted and received with the designated user based on at least a portion of the transmitted message, and an operation of obtaining the text data (1020) corresponding to the identified subject.
[0234] According to various embodiments, the method of operating the electronic device (100) may include an operation of displaying a transmission message transmitted to the designated user through the conversation screen (210), and an operation of generating a response message (730, 1210) for the received message (211-1) based on at least a portion of the transmission message, the acquired text data (1020), and the designated second AI model (1200).
[0235] According to various embodiments, the operating method of the electronic device (100) may include an operation of displaying a transmission message transmitted to the designated user through the conversation screen (210), an operation of analyzing a conversation style based on the transmission message (211-2) and the reception message (211-1) displayed on the conversation screen (210), and an operation of generating a response message (730, 1210) for the reception message (211-1) based on information (1501) related to the conversation style, at least a portion of the transmission message (211-2), the acquired text data (1020), and the designated second AI model (1200).
[0236] According to various embodiments, the method of operating the electronic device (100) may include, if the conversation style is associated with the expression of honorifics, an operation of expressing the response message (730, 1210) in honorifics, and if the conversation style is associated with the expression of informal speech, an operation of expressing the response message (730, 1210) in informal speech.
[0237] According to various embodiments, the operating method of the electronic device (100) may include an operation of identifying at least one view constituting the conversation screen and an operation of identifying the image data (720) based on the identified at least one view.
[0238] According to various embodiments, the method of operating the electronic device (100) may include an operation of identifying at least one view based on a screen hierarchy relationship for the conversation screen.
[0239] According to various embodiments, the method of operating the electronic device (100) may include an operation of transmitting the generated response message (730, 1210) to the designated user.
[0240] According to various embodiments, the first AI model (1000) specified above may include a large multi-modal model (LMM), and the second AI model (1200) specified above may include a large language model (LLM).
[0241]
[0242] According to various embodiments, a computer-readable storage medium, when executed by a processor (120) of an electronic device (100), causes the electronic device (100) to: output a conversation screen through a display (140) in which a received message (211-1) including at least one text (211-1a) and at least one image (211-1b) received from a designated user is displayed; identify image data (720) for the at least one image (211-1b) displayed as at least a part of the received message (211-1) on the conversation screen; acquire text data describing an object included in the image data (720) based on the identified image data (720) and a designated first artificial intelligence (AI) model; and generate a response message for the received message (211-1) based on the acquired text data based on the at least one text included in the received message (211-1) and the image data (720) and a designated second AI model. You can store the instructions that you want to do.
Claims
1. In an electronic device (100), At least one processor (120); display (140); Communication circuit (150); and A memory (130) operatively connected to at least one processor (120), the display (140) and the communication circuit (150) and storing at least one command, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: A dialogue screen (210) is displayed through the display (140) on which a received message (211-1) including at least one text (211-1a) and at least one image (211-1b) received from a user specified through the communication circuit (150) is displayed, Identifying image data (720) for at least one image (211-1b) displayed as at least a part of the received message (211-1) in the above conversation screen (210), Obtain text data (1020) describing an object included in the image data (720) based on the identified image data and the designated first artificial intelligence (AI) model (1000), An electronic device configured to generate a response message (730, 1210) for the received message (211-1) based on the acquired text data (1020) based on the at least one text (211-1a) included in the received message (211) and the image data (720) and a designated second AI model (1200).
2. In paragraph 1, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: The transmission message transmitted to the designated user through the above communication circuit (150) is displayed through the above conversation screen (210), An electronic device configured to obtain the text data (1020) based on at least a portion of the transmitted message, the identified image data (720) and the designated first AI model (1000).
3. In paragraph 2, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: Identifying the subject of a message to be sent or received from the designated user based on at least a portion of the above-mentioned transmitted message; An electronic device configured to obtain the text data (1020) corresponding to the above-identified subject.
4. In paragraph 1, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: The transmission message transmitted to the designated user through the above communication circuit (150) is displayed through the above conversation screen (210), An electronic device configured to generate a response message (730, 1210) to the received message (211-1) based on at least a portion of the transmitted message, the acquired text data (1020) and the designated second AI model (1200).
5. In paragraph 1, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: The transmission message (211-2) transmitted to the designated user through the above communication circuit (150) is displayed through the above conversation screen (210), The conversation style is analyzed based on the transmitted message (211-2) and received message (211-1) displayed on the above conversation screen (210), An electronic device configured to generate a response message (730, 1210) to the received message (211-1) based on information (1501) related to the conversation style, at least a portion of the transmitted message (211-2), the acquired text data (1020), and the designated second AI model (1200).
6. In paragraph 5, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: If the above conversation style is associated with the expression of honorifics, the above response message (730, 1210) is expressed in the above honorifics, An electronic device configured to express the response message (730, 1210) in the informal language, if the above conversational style is associated with the informal language.
7. In paragraph 1, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: Identify at least one view that constitutes the above conversation screen, An electronic device configured to identify the image data (720) based on at least one view identified above.
8. In paragraph 7, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: An electronic device configured to identify at least one view based on a screen hierarchy relationship for the above conversation screen.
9. In paragraph 1, The at least one instruction, when individually or collectively executed by the at least one processor (120), causes the electronic device (100) to: An electronic device configured to transmit the generated response message (730, 1210) to the specified user.
10. In paragraph 1, The above-mentioned first AI model (1000) includes a large multi-modal model (LMM), The above-mentioned second AI model (1200) is an electronic device including a large language model (LLM).
11. In the operating method of an electronic device (100), An action of outputting a conversation screen (210) through a display (140) in which a received message (211-1) including at least one text (211-1a) and at least one image (211-1b) received from a designated user is displayed; An operation of identifying image data (720) for at least one image (211-1b) displayed as at least a part of the received message (211-1) in the above conversation screen (210); An operation of obtaining text data (1020) describing an object included in the image data (720) based on the identified image data (720) and a designated first artificial intelligence (AI) model (1000); and A method comprising an operation of generating a response message (730, 1210) for the received message (211-1) based on the acquired text data (1020) based on the at least one text (211-1a) included in the received message (211-1) and the image data (720) and a designated second AI model (1200).
12. In paragraph 11, An action of displaying a transmission message sent to the above-mentioned user through the above-mentioned conversation screen (210); and A method comprising obtaining the text data (1020) based on at least a portion of the transmitted message, the identified image data (720) and the designated first AI model (1000).
13. In paragraph 12, An operation of identifying a subject of a message to be sent or received from the designated user based on at least a portion of the transmitted message; and A method comprising an action of obtaining the text data (1020) corresponding to the above-identified subject.
14. In paragraph 11, An action of displaying a transmission message sent to the above-mentioned user through the above-mentioned conversation screen (210); A method comprising the action of generating a response message (730, 1210) to the received message (211-1) based on at least a portion of the transmitted message, the acquired text data (1020) and the designated second AI model (1200).
15. In paragraph 11, An action of displaying a transmission message (211-2) sent to the above-mentioned user through the above-mentioned conversation screen (210); An operation of analyzing a conversation style based on a transmitted message (211-2) and a received message (211-1) displayed on the above conversation screen (210); and A method comprising an operation of generating a response message (730, 1210) to the received message (211-1) based on information (1501) related to the conversation style, at least a portion of the transmitted message (211-2), the acquired text data (1020), and the designated second AI model (1200).
Citation Information
Patent Citations
User terminal apparatus for recommanding a reply message and method thereof
KR1020170054919A
Automatic suggested responses to images received in messages using language models.
KR102050334B1
Method of recommanding a reply message and apparatus thereof
KR102477272B1
Systems, apparatuses and methods for generating a user interface
US20160034441A1
KR20210007128A