Electronic device and control method therefor

WO2026160825A1PCT designated stage Publication Date: 2026-07-30SAMSUNG ELECTRONICS CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
SAMSUNG ELECTRONICS CO LTD
Filing Date
2026-01-20
Publication Date
2026-07-30

Smart Images

  • Figure KR2026001200_30072026_PF_FP_ABST
    Figure KR2026001200_30072026_PF_FP_ABST
Patent Text Reader

Abstract

An electronic device and a control method therefor are provided. This electronic device includes a display, a memory storing at least one instruction, and a processor, wherein the at least one instruction, when collectively or individually executed by the processor, causes the electronic device to: capture an image displayed on the display; obtain context information corresponding to the captured image by inputting the image to a generative artificial intelligence (AI) model; and store the context information in a long-term memory storage space or a short-term memory storage space on the basis of whether the context information is information related to a human state or corresponds to a pre-registered preferred keyword.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device and control method thereof

[0001] The present disclosure relates to an electronic device and a method for controlling the same.

[0002] Recently, users frequently use mobile devices. Users' mobile device usage records contain a significant amount of meaningful information about them. However, storing all of a user's mobile device usage records makes data management difficult and may provide users with unnecessary information.

[0003] Furthermore, the user's environment can continuously change over time. As the user's environment changes, the information meaningful to the user may also change. Therefore, there is a growing need for technology capable of selecting and storing information meaningful to the user by reflecting changes over time.

[0004] According to one embodiment of the present disclosure, an electronic device comprises a display, a memory in which at least one instruction is stored, and a processor. When the at least one instruction is executed collectively or individually by the processor, the electronic device captures an image displayed by the display, inputs the captured image into a generative AI model to obtain context information corresponding to the image, and stores the context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword.

[0005] The above context information may include a plurality of keywords representing the situation regarding the image and information regarding the relationship between the plurality of keywords.

[0006] If the above context information is information related to a person's state, the above context information can be stored in the above long-term memory storage space.

[0007] If the above context information is not information related to a person's state, the above context information may be stored in the short-term memory storage space, and the above context information may be stored in the long-term memory storage space based on whether the above context information stored in the short-term memory storage space corresponds to the preferred keyword.

[0008] If the context information stored in the short-term memory storage space corresponds to the preferred keyword and is information related to the completion of an action among the information related to a person's action, the context information can be stored in the long-term memory storage space.

[0009] It is possible to obtain a plurality of feature vectors corresponding to the plurality of keywords included in the context information, and to obtain the preferred keyword corresponding to at least one cluster identified based on the similarity between the plurality of feature vectors.

[0010] When an input corresponding to a user request is received, a response corresponding to the request or an operation can be performed based on the context information corresponding to the request stored in the short-term memory storage space or the long-term memory storage space.

[0011] If context information corresponding to the above request exists in the short-term memory storage space, a response corresponding to the above request is provided or an operation is performed based on the context information stored in the short-term memory storage space, and if context information corresponding to the above request does not exist in the short-term memory storage space but exists in the long-term memory storage space, a response corresponding to the above request is provided or an operation is performed based on the context information stored in the long-term memory storage space.

[0012] A control method for an electronic device including a display according to one embodiment of the present disclosure comprises: a step of capturing an image displayed by the display; a step of inputting the captured image into a generative AI model to obtain context information corresponding to the image; and a step of storing the context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword.

[0013] The above context information may include a plurality of keywords representing the situation regarding the image and information regarding the relationship between the plurality of keywords.

[0014] The step of storing the above context information in a long-term memory storage space or a short-term memory storage space may include the step of storing the context information in the long-term memory storage space if the context information is information related to a person's state.

[0015] The step of storing the above context information in a long-term memory storage space or a short-term memory storage space may include: a step of storing the above context information in the short-term memory storage space if the above context information is not information related to a person's state; and a step of storing the above context information in the long-term memory storage space based on whether the above context information stored in the short-term memory storage space corresponds to the preferred keyword.

[0016] The step of storing the above context information in a long-term memory storage space or a short-term memory storage space may include the step of storing the above context information in the long-term memory storage space if the context information stored in the short-term memory storage space corresponds to the preferred keyword and is information related to the completion of an action among information related to a person's action.

[0017] The above control method may include: a step of obtaining a plurality of feature vectors corresponding to the plurality of keywords included in the context information; and a step of obtaining the preferred keyword corresponding to at least one cluster identified based on the similarity between the plurality of feature vectors.

[0018] The above control method may include the step of, upon receiving an input corresponding to a user request, providing a response corresponding to the request or performing an operation based on the context information corresponding to the request stored in the short-term memory storage space or the long-term memory storage space.

[0019] The step of providing the above response or performing the action may include: a step of providing the response or performing the action corresponding to the request based on the context information stored in the short-term memory storage space if context information corresponding to the request exists in the short-term memory storage space; and a step of providing the response or performing the action corresponding to the request based on the context information stored in the long-term memory storage space if context information corresponding to the request does not exist in the short-term memory storage space but exists in the long-term memory storage space.

[0020] FIG. 1 is a schematic diagram showing the operation of an electronic device (100) according to one embodiment of the present disclosure.

[0021] FIG. 2 is a block diagram showing the configuration of an electronic device (100) according to one embodiment of the present disclosure.

[0022] FIG. 3 is a flowchart illustrating a method for an electronic device (100) according to one embodiment of the present disclosure to store context information in a storage space.

[0023] FIG. 4 is a diagram showing the acquisition of context information including a plurality of keywords from a captured image according to one embodiment of the present disclosure.

[0024] FIG. 5 is a diagram illustrating a method for obtaining a preferred keyword from a plurality of keywords according to one embodiment of the present disclosure.

[0025] FIG. 6 is a flowchart illustrating a method in which an electronic device provides a response or performs an operation corresponding to a user request when an input corresponding to a user request is received according to one embodiment of the present disclosure.

[0026] FIG. 7 is a flowchart illustrating a control method of an electronic device (100) according to one embodiment of the present disclosure.

[0027] The embodiments described herein are subject to various modifications and may have various forms; specific embodiments are illustrated in the drawings and described in detail in the detailed description. However, this is not intended to limit the scope of specific embodiments and should be understood to include various modifications, equivalents, and / or alternatives of the embodiments of the present disclosure. In relation to the description of the drawings, similar reference numerals may be used for similar components.

[0028] In describing the present disclosure, if it is determined that a detailed description of related known functions or configurations could unnecessarily obscure the essence of the present disclosure, such detailed description is omitted.

[0029] Additionally, the following embodiments may be modified in various other forms, and the scope of the technical concept of the present disclosure is not limited to the following embodiments. Rather, these embodiments are provided to make the present disclosure more faithful and complete and to fully convey the technical concept of the present disclosure to those skilled in the art.

[0030] The terms used in this disclosure are used merely to describe specific embodiments and are not intended to limit the scope of the rights. Singular expressions include plural expressions unless the context clearly indicates otherwise.

[0031] In the present disclosure, expressions such as “have,” “may have,” “include,” or “may include” indicate the presence of such features (e.g., numerical values, functions, actions, or components such as parts) and do not exclude the presence of additional features.

[0032] In the present disclosure, expressions such as “A or B,” “at least one of A or / and B,” or “one or more of A or / and B” may include all possible combinations of items listed together. For example, “A or B,” “at least one of A and B,” or “at least one of A or B” may refer to cases including (1) at least one A, (2) at least one B, or (3) both at least one A and at least one B.

[0033] Expressions such as "first," "second," "first," or "second" used in this disclosure may modify various components regardless of order and / or importance, and are used only to distinguish one component from another and do not limit said components.

[0034] Where it is stated that a certain component (e.g., a first component) is "(operatively or communicatively) coupled with / to" or "connected to" another component (e.g., a second component), it should be understood that the said certain component may be directly connected to the said other component or connected through another component (e.g., a third component).

[0035] On the other hand, when it is stated that a certain component (e.g., a first component) is "directly connected" or "directly coupled" to another component (e.g., a second component), it may be understood that no other component (e.g., a third component) exists between said certain component and said other component.

[0036] As used in this disclosure, the expression “configured to” may be replaced, depending on the context, with, for example, “suitable for,” “having the capacity to,” “designed to,” “adapted to,” “made to,” or “capable of.” The term “configured to” may not necessarily mean only “specifically designed to” in hardware.

[0037] Instead, in some situations, the expression “device configured to do something” may mean that the device is “capable of doing something” together with other devices or components. For example, the phrase “processor configured (or set) to perform A, B, and C” may mean a dedicated processor for performing those operations (e.g., an embedded processor), or a generic-purpose processor (e.g., a CPU or application processor) capable of performing those operations by executing one or more software programs stored in a memory device.

[0038] In the embodiments, a 'module' or 'part' performs at least one function or operation and may be implemented in hardware or software, or a combination of hardware and software. Additionally, a plurality of 'modules' or a plurality of 'parts' may be integrated into at least one module and implemented by at least one processor, except for a 'module' or 'part' that needs to be implemented in specific hardware.

[0039] An embodiment of the present disclosure will be described in more detail below with reference to the attached drawings.

[0040] Meanwhile, various elements and areas in the drawings are depicted schematically. Accordingly, the technical concept of the present invention is not limited by the relative sizes or spacing depicted in the attached drawings.

[0041] FIG. 1 is a schematic diagram showing the operation of an electronic device (100) according to one embodiment of the present disclosure.

[0042] An electronic device (100) can capture an image (10) displayed on a display. An electronic device (100) can capture an image (10) corresponding to a screen displayed to a user. For example, an electronic device (100) can capture an image (10) corresponding to a screen of an application being run or a search screen. An electronic device (100) can capture multiple images (10) corresponding to a screen that changes according to user input. Another example is that an electronic device (100) can capture an image (10) obtained through a camera. Referring to FIG. 1, an electronic device (100) displays a screen of an application for accommodation reservation and can capture an image (10) corresponding to the screen of the application.

[0043] The electronic device (100) can input a captured image (10) into a generative AI model to obtain context information corresponding to the image (10).

[0044] Context information may include multiple keywords representing the situation regarding the image (10) and information regarding the relationships between the multiple keywords. The multiple keywords may be keywords corresponding to text included in the image (10), or keywords corresponding to multiple objects included in the image (10). Information regarding the relationships between the multiple keywords may be information regarding the connection relationships between multiple keywords corresponding to the situation regarding the image (10). For example, in FIG. 1, when an electronic device (100) inputs an image (10) corresponding to the screen of an application related to accommodation reservation into a generative AI model, the electronic device (100) may obtain context information including information regarding the name of the accommodation, the rating of the accommodation, the location of the accommodation, the price of the accommodation, or the appearance of the accommodation included in the image (10). Alternatively, the context information may be cognitive information, screen understanding information, or screen context information.

[0045] The generative AI model may be an artificial intelligence model trained to output context information corresponding to an image (10) when an image (10) is input. The generative AI model may be an artificial intelligence model trained to output context information including multiple keywords representing the situation of the input image (10) and information about the relationships between the multiple keywords.

[0046] The electronic device (100) may store context information in a long-term memory storage space (30) or a short-term memory storage space (20) based on whether the context information corresponds to a type of context information or a previously registered preferred keyword. The type of context information may include whether the information is related to a person's state or whether it is related to a person's behavior.

[0047] Preferred keywords may be keywords that indicate a user's preferences. Specifically, preferred keywords may be keywords corresponding to fields or items that the user is interested in. For example, preferred keywords may include keywords corresponding to the food or movie genres that the user likes. Another example is that preferred keywords may include keywords corresponding to the user's hobbies.

[0048] The electronic device (100) may store context information in a long-term memory storage space (30) or a short-term memory storage space (20). The long-term memory storage space (30) may be a storage space capable of permanently storing data. In one embodiment, the electronic device (100) may store information related to a person's state or information corresponding to a preferred keyword among the context information in the long-term memory storage space (30). In another example, the electronic device (100) may store information related to the completion of an action in the long-term memory storage space (30).

[0049] On the other hand, the short-term memory storage space (20) is a space where data can be stored non-permanently. The electronic device (100) stores data in the short-term memory storage space (20) in the order in which it was acquired, and can delete the data from the short-term memory storage space (20) in the order in which it was stored after a preset time has passed. The electronic device (100) can store context information acquired from a captured image (10) in the short-term memory storage space (20).

[0050] When the electronic device (100) receives an input corresponding to a user request, it identifies whether context information corresponding to the user request exists in a long-term memory storage space (30) or a short-term memory storage space (20), and can provide a response corresponding to the user request or perform an action based on the context information stored in the long-term memory storage space (30) or the short-term memory storage space (20).

[0051] Referring to FIG. 1, an electronic device (100) is shown receiving input from a user requesting to know the place with the most reviews and the best rating among the accommodations viewed so far, and providing a response corresponding to the request. The electronic device (100) can obtain context information including information related to accommodations obtained from an image (10) corresponding to the screen of an application related to accommodation reservation. The electronic device (100) can store context information including information related to accommodations in a short-term memory storage space (20) or a long-term memory storage space (30). When the electronic device (100) receives input corresponding to a user request, it identifies whether context information corresponding to the user request exists in the short-term memory storage space (20) or the long-term memory storage space (30), and can provide a response based on the context information stored in the short-term memory storage space (20) or the long-term memory storage space (30).

[0052] FIG. 2 is a block diagram showing the configuration of an electronic device (100) according to one embodiment of the present disclosure.

[0053] Referring to FIG. 2, the electronic device (100) may include a display (110), memory (120), and a processor (130). Meanwhile, the configuration shown in FIG. 2 is merely an example of various embodiments, and some configurations may be omitted and new configurations may be added.

[0054] The display (110) can display an image under the control of the processor (130). The display (110) can be implemented as an LCD (Liquid Crystal Display Panel), OLED (Organic Light Emitting Diodes), etc., and the display (110) can also be implemented as a flexible display, transparent display, etc. depending on the case. However, the display (110) according to the present disclosure is not limited to a specific type. Although FIG. 2 is illustrated as if the electronic device (100) directly embeds the display (110), this can be interpreted to include not only cases where the display (110) is actually mounted on the electronic device (100) (e.g., TV, kiosk, smartphone, laptop PC, tablet PC, etc.), but also cases where it is connected to a separately provided display device (e.g., monitor, TV, beam projector, electronic whiteboard, etc.) via various wired or wireless communication methods (e.g., set-top box, PC, server, etc.).

[0055] At least one instruction regarding an electronic device (100) may be stored in the memory (120). Additionally, an operating system (O / S) for operating the electronic device (100) may be stored in the memory (120). Furthermore, various software programs or applications for operating the electronic device (100) may be stored in the memory (120) according to various embodiments of the present disclosure. Additionally, the memory (120) may be implemented as a volatile memory such as S-RAM (Static Random Access Memory) or D-RAM (Dynamic Random Access Memory), a non-volatile memory such as Flash Memory, ROM (Read Only Memory), EPROM (Erasable Programmable Read Only Memory), or EEPROM (Electrically Erasable Programmable Read Only Memory), a hard disk drive (HDD), or a solid-state drive (SSD).

[0056] Specifically, various software modules for operating an electronic device (100) according to various embodiments of the present disclosure may be stored in the memory (120), and the processor (130) may control the operation of the electronic device (100) by executing the various software modules stored in the memory (120). That is, the memory (120) is accessed by the processor (130), and reading / writing / modifying / deleting / updating of data by the processor (130) may be performed.

[0057] The memory (120) may be a configuration provided separately from the processor (130), may be an internal memory built into the processor (130), and may also be used to include a memory (120) card (not shown) (e.g., micro SD card, memory stick) or an external hard drive mounted on the electronic device (100).

[0058] The processor (130) controls the overall operation of the electronic device (100). Specifically, the processor (130) is connected to the configuration of the electronic device (100) including a display (110) and a memory (120), and can control the overall operation of the electronic device (100) by executing at least one instruction stored in the memory (120) as described above.

[0059] The processor (130) can be implemented in various ways. For example, the processor (130) may include or be defined by one or more of a central processing unit (CPU) that processes digital signals, a Micro Controller Unit (MCU), a micro processing unit (MPU), a controller, an application processor (AP), a communication processor (CP), or an ARM processor. Additionally, the processor (130) may be implemented as a System on Chip (SoC) or Large Scale Integration (LSI) with built-in processing algorithms, or as a Field Programmable Gate Array (FPGA). The processor (130) can perform various functions by executing computer executable instructions stored in memory (120).

[0060] The processor (130) can perform a method according to one or more embodiments of the present disclosure based on the execution of at least one instruction stored in memory (120).

[0061] According to one embodiment, the processor (130) can capture an image displayed by the display (110).

[0062] According to one embodiment, the processor (130) can input a captured image into a generative AI model to obtain context information corresponding to the image.

[0063] According to one embodiment, the processor (130) may store context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword.

[0064] Meanwhile, an artificial intelligence model refers to a model that inputs specific input values ​​into a specific function based on learned data to produce an output value. Artificial intelligence models can be referred to in various ways, such as neural network models, deep learning models, neural network models, and generative artificial intelligence models (generative AI models).

[0065] Meanwhile, the artificial intelligence-related function according to the present disclosure may be operated by having the artificial intelligence model included in an external server, transmitting input values ​​to an external device through a communication unit, and receiving output values.

[0066] Alternatively, an artificial intelligence model may be stored in the memory (120) of the electronic device (100) and operated through the processor (130) and the memory (120).

[0067] The processor (130) can be controlled to process input data according to predefined operation rules or artificial intelligence models stored in memory (120). Alternatively, if the processor (130) is an artificial intelligence-dedicated processor, the artificial intelligence-dedicated processor may be designed with a hardware structure specialized for processing a specific artificial intelligence model. The predefined operation rules or artificial intelligence models are characterized by being created through learning.

[0068] Here, "created through learning" means that a basic artificial intelligence model is trained using multiple learning data by a learning algorithm, thereby creating a predefined rule of operation or an artificial intelligence model configured to perform a desired characteristic (or objective). Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include, but are not limited to, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning.

[0069] An artificial intelligence model can be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values ​​and performs neural network operations through calculations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights can be updated during the learning process so that the loss or cost values ​​obtained by the artificial intelligence model are reduced or minimized.

[0070] Artificial neural networks may include deep neural networks (DNNs), such as, but are not limited to, Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), Generative Adversarial Networks (GANs), or Deep Q-Networks.

[0071] Hereinafter, the present disclosure will be described in more detail with reference to the drawings.

[0072] FIG. 3 is a flowchart illustrating a method for an electronic device (100) according to one embodiment of the present disclosure to store context information in a storage space.

[0073] The electronic device (100) can capture an image displayed by a display (S310). The image may be an image corresponding to a screen being displayed to a user using the electronic device (100). For example, the image may include a screen of an application running on the electronic device (100) or an image corresponding to a search screen of an internet browser. Alternatively, the image may include an image acquired through a camera and displayed on a display.

[0074] The electronic device (100) can capture multiple images that are displayed differently depending on user input. The electronic device (100) can capture multiple images corresponding to screens that are displayed differently depending on user input. Referring to FIG. 1, the electronic device (100) can capture multiple images corresponding to screens of an application that are displayed differently as the user searches for various accommodations.

[0075] The electronic device (100) can input a captured image into a generative AI model to obtain context information corresponding to the image (S320). The context information may include multiple keywords representing the situation regarding the image and information regarding the relationships between the multiple keywords. The multiple keywords may be keywords corresponding to text included in the image, or keywords corresponding to multiple objects included in the image. The information regarding the relationships between the multiple keywords may be information regarding the connection relationships between multiple keywords corresponding to the situation included in the image. For example, when the electronic device (100) inputs an image corresponding to the screen of an application related to accommodation reservation into a generative AI model, it can obtain context information including information about the name of the accommodation, the rating of the accommodation, the location of the accommodation, the price of the accommodation, or the exterior of the accommodation included in the image.

[0076] A generative AI model may be an artificial intelligence model trained to output contextual information corresponding to an image when an image is input. The generative AI model may be an artificial intelligence model trained to output contextual information including multiple keywords representing the situation regarding the input image and information about the relationships between the multiple keywords.

[0077] The electronic device (100) can obtain a preferred keyword based on a plurality of keywords included in context information (S330). The preferred keyword may be a keyword representing the user's preference. Specifically, the preferred keyword may be a keyword corresponding to a field or item that the user is interested in.

[0078] The electronic device (100) can acquire a plurality of feature vectors corresponding to a plurality of keywords and acquire a preferred keyword corresponding to at least one cluster identified based on the similarity between the plurality of feature vectors. A method for acquiring the preferred keyword is described in detail below in FIG. 5.

[0079] The electronic device (100) can store context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword.

[0080] Specifically, the electronic device (100) may store the context information in a long-term memory storage space (S380) if the context information is information related to a person's state (S340-Y). Information related to a person's state may be information about a user using the electronic device (100). Information related to a person's state may include information about the user's characteristics, physical state, and social state. Information about the user's characteristics may include the user's age, gender, occupation, place of residence, etc. Information about the physical state may refer to information about the user's health. Information about the social state may refer to information about the user's human relationships.

[0081] The electronic device (100) can permanently store information related to the user by storing context information in a long-term memory storage space, regardless of whether the context information corresponds to a preferred keyword, if the context information is information related to the state of the person. When the electronic device (100) receives an input corresponding to a user request, it can provide a user-customized response or perform an action based on the permanently stored information related to the user.

[0082] On the other hand, the electronic device (100) may store the context information in a short-term memory storage space (S350) if the context information is not information related to a person's state (S340-N). Specifically, the electronic device (100) may store the context information in a short-term memory storage space if the context information is information regarding a person's behavior or information regarding the state of an object. The electronic device (100) may store the acquired context information in a short-term memory storage space first if it is information regarding a person's behavior or information regarding the state of an object, and then store it in a long-term memory storage space based on whether it corresponds to a preferred keyword and whether it is information related to the completion of an action.

[0083] Meanwhile, the electronic device (100) can store context information in the order in which it was acquired in a short-term memory storage space. The electronic device (100) can delete the context information in the order in which it was stored after a preset time has passed. The electronic device (100) can increase the efficiency of the storage space by deleting unnecessary information over time instead of storing it permanently.

[0084] The electronic device (100) can store context information in a long-term memory storage space based on whether the context information stored in the short-term memory storage space corresponds to a preferred keyword and whether the information is related to a person's behavior.

[0085] The electronic device (100) can identify whether context information stored in a short-term memory storage space corresponds to a preferred keyword (S360). The electronic device (100) can identify whether the context information corresponds to the preferred keyword based on the similarity between the feature vector corresponding to the context information and the feature vector corresponding to the preferred keyword. For example, the electronic device (100) can identify that the context information corresponds to the preferred keyword if the distance between the feature vector corresponding to the context information and the feature vector corresponding to the preferred keyword is less than or equal to a preset value. For another example, the electronic device (100) can identify that the context information corresponds to the preferred keyword if the cosine similarity obtained based on the angle between the feature vector corresponding to the context information and the feature vector corresponding to the preferred keyword is greater than or equal to a preset value. However, the above-described examples are merely examples of obtaining similarity between feature vectors and are not limited thereto.

[0086] When the context information stored in the short-term memory storage space corresponds to a preferred keyword (S360-Y), the electronic device (100) can identify whether the context information is information related to the completion of an action among the information related to a person's action (S370). By storing the context information corresponding to the preferred keyword in the long-term memory storage space, the electronic device (100) can store information that reflects the user's preference.

[0087] Information related to a person's actions may be information related to an action that the user is performing through the electronic device (100). For example, information related to a person's actions may be information related to an action of searching for accommodations or information related to an action of booking accommodations. Information related to the completion of an action may be information related to the completion of an action that the user is performing through the electronic device (100). For example, information related to the completion of an action may be information related to an action of purchasing goods by proceeding with payment or information related to an action of booking accommodations. The electronic device (100) can identify that the context information is information related to the completion of an action if the context information includes keywords related to the completion of an action. For example, the electronic device (100) can identify that the context information obtained from the image is information related to the completion of an action if keywords such as 'purchase completed' or 'reservation completed' are identified from the image.

[0088] The electronic device (100) can store the context information in a long-term memory storage space if the context information is information related to the completion of an action among the information related to a person's action (S370-Y) (S380).

[0089] The electronic device (100) can store information necessary for the user after the completion of an action by storing context information related to the completion of the action in a long-term memory storage space.

[0090] The electronic device (100) may maintain the state in which context information is stored in short-term memory storage space if the context information does not correspond to a preferred keyword (S360-N) or is not information related to the completion of an action (S370-N). The electronic device (100) may delete the context information in the order in which it was stored after a preset time has elapsed since the context information was stored.

[0091] Meanwhile, in another embodiment of the present disclosure, the electronic device (100) may store context information in a long-term memory storage space or a short-term memory storage space based on clusters without acquiring preferred keywords. Specifically, the electronic device (100) may store context information in a long-term memory storage space or a short-term memory storage space depending on whether a feature vector corresponding to a keyword included in the context information is included in a specific cluster area within a vector space. This will be explained in detail below.

[0092] The electronic device (100) can store context information in a short-term memory storage space if the context information obtained from the captured image is not information related to a person's state.

[0093] Subsequently, the electronic device (100) can identify whether a feature vector corresponding to a keyword included in the context information is included in a cluster area within the vector space. Specifically, the electronic device (100) can identify at least one cluster (520) having high density within a vector space (520) where feature vectors for a plurality of keywords are mapped. The electronic device (100) can map a feature vector corresponding to a keyword included in the context information to the vector space (520) to identify whether the feature vector is included in a specific cluster area. If a feature vector corresponding to a keyword included in the context information is included in a specific cluster area within the vector space, the context information may be information with high user preference.

[0094] The electronic device (100) can identify whether the context information is information related to the completion of an action among information related to a person's action when a feature vector corresponding to a keyword included in the context information is included in a specific cluster area. If the context information is information related to the completion of an action among information related to a person's action, the electronic device (100) can store the context information in a long-term memory storage space. On the other hand, if the context information is not information related to the completion of a person's action, the context information can be maintained in a state where it is stored in a short-term memory storage space.

[0095] Meanwhile, the electronic device (100) can maintain the state in which the context information is stored in the short-term memory storage space if the feature vector corresponding to the keyword included in the context information is not included in any cluster area.

[0096] When the electronic device (100) receives an input corresponding to a user request, it may provide a response corresponding to the user request or perform an action based on context information stored in a short-term memory storage space or a long-term memory storage space. A detailed explanation is provided in FIG. 6.

[0097] FIGS. 4 and 5 are drawings illustrating a method for obtaining context information including a plurality of keywords from a captured image according to an embodiment of the present disclosure, and obtaining a preferred keyword based on the plurality of keywords.

[0098] Referring to FIG. 4, an image corresponding to the screen of an application used by a user and a plurality of keywords obtained from the image are shown.

[0099] In FIG. 4, a plurality of keywords included in context information obtained when an image corresponding to the screen of an application providing a movie is input into a generative AI model are illustrated. At this time, the plurality of keywords may include keywords corresponding to text included in the image. For example, in FIG. 4, the plurality of keywords may include the movie title, the name of the application, and the genre of the movie ('Youth Romance', 'Romantic Comedy' in FIG. 4), which are keywords corresponding to text included in the image. Alternatively, the plurality of keywords may include keywords corresponding to objects included in the image. For example, in FIG. 4, the plurality of keywords may include 'classic car,' which is a keyword corresponding to a picture of a 'car,' which is an object included in the image. In addition, the plurality of keywords may include keywords that can be inferred from keywords corresponding to text or objects included in the image. For example, in FIG. 4, the plurality of keywords may include keywords such as 'movie' and 'date,' which can be inferred from the movie title and the genre of the movie.

[0100] Additionally, when an image corresponding to the screen of an application that provides videos related to recipes is input into a generative AI model, multiple keywords included in the acquired context information are shown. Referring to FIG. 4, the multiple keywords may include keywords corresponding to text included in the image ('microwave recipe'), keywords corresponding to objects included in the image ('pumpkin', 'pumpkin recipe'), and keywords that can be inferred from keywords corresponding to text and objects included in the image ('healthy recipe', 'home cafe', 'easy cooking', etc.).

[0101] Although only multiple keywords are shown in FIG. 4 for convenience, the electronic device (100) is not limited thereto and can obtain context information including multiple keywords and relationships between multiple keywords through a generative AI model.

[0102] FIG. 5 is a diagram illustrating a method for obtaining a preferred keyword from a plurality of keywords according to one embodiment of the present disclosure.

[0103] The electronic device (100) can acquire multiple feature vectors corresponding to multiple keywords. The electronic device (100) can identify at least one cluster (520) based on the similarity between multiple feature vectors in a vector space (510) where multiple feature vectors are mapped. A cluster (520) may refer to a set of multiple feature vectors where the similarity is greater than or equal to a threshold value. Additionally, a cluster (520) may refer to a set of feature vectors included in an area where the density within the vector space is greater than or equal to a preset value. The electronic device (100) can identify a cluster (520) where the distance between multiple feature vectors is less than or equal to a preset value. Another example is that the electronic device (100) can identify a cluster (520) where the cosine similarity between multiple feature vectors is greater than or equal to a preset value. Referring to FIG. 5, the electronic device (100) can acquire feature vectors corresponding to each of a plurality of keywords and map them to a vector space (510), and identify a cluster (520), which is a set of feature vectors having a density greater than or equal to a preset value within the vector space (510).

[0104] The electronic device (100) can obtain a preferred keyword corresponding to at least one cluster. The electronic device (100) can obtain a feature vector corresponding to at least one cluster. In one embodiment, the electronic device (100) can obtain a feature vector corresponding to the center of at least one cluster. In another embodiment, the electronic device (100) can obtain a feature vector corresponding to at least one cluster using a density centroid or a Gaussian mean.

[0105] The electronic device (100) can obtain preferred keywords based on feature vectors corresponding to at least one cluster. Referring to FIG. 5, the electronic device (100) can obtain ‘romantic comedy’ and ‘healthy recipe’ as preferred keywords.

[0106] The electronic device (100) can acquire preferred keywords by accumulating multiple keywords included in context information obtained through an image. The electronic device (100) can delete keywords from the accumulated multiple keywords in the order they were stored after a preset time has elapsed. The electronic device (100) can acquire new preferred keywords based on the remaining multiple keywords. Through this, the electronic device (100) can acquire preferred keywords for fields or items of interest to the user by reflecting changes in the user's preferences over time.

[0107] FIG. 6 is a flowchart illustrating a method in which, when an input corresponding to a user request is received according to one embodiment of the present disclosure, an electronic device (100) provides a response corresponding to the user request or performs an operation.

[0108] The electronic device (100) can receive an input corresponding to a user request (S610). The input corresponding to the user request may be an input corresponding to a user's query or an input requesting the performance of an operation of the electronic device (100). The electronic device (100) can receive a user's voice or text input corresponding to the user request.

[0109] When the electronic device (100) receives an input corresponding to a user request, it can provide a response corresponding to the request or perform an action based on context information stored in a short-term memory storage space or a long-term memory storage space.

[0110] Specifically, the electronic device (100) can search whether context information corresponding to a user request exists in a short-term memory storage space (S620). The electronic device (100) can search whether context information corresponding to the user's intention included in the user request exists in a short-term memory storage space.

[0111] If context information corresponding to a user request exists in a short-term memory storage space (S630-Y), the electronic device (100) can provide a response corresponding to the request or perform an action based on the context information stored in the short-term memory storage space (S660). At this time, the electronic device (100) can obtain a response corresponding to the user request by inputting the user request and the context information stored in the short-term memory storage space together into an artificial intelligence model. Through this, the electronic device (100) can provide a more efficient and faster response by providing a response based on the information stored in the short-term memory storage space.

[0112] On the other hand, if the electronic device (100) does not have context information corresponding to a user request in a short-term memory storage space, it can search whether the context information exists in a long-term memory storage space (S640). The electronic device (100) can search whether the context information corresponding to the user's intention included in the user request exists in a long-term memory storage space.

[0113] If context information corresponding to a user request exists in a long-term memory storage space (S650-Y), the electronic device (100) can provide a response corresponding to the request or perform an action based on the context information stored in the long-term memory storage space (S670). At this time, the electronic device (100) can obtain a response corresponding to the user request by inputting the user request and the context information stored in the long-term memory storage space together into an artificial intelligence model. Through this, the electronic device (100) can provide a user-customized response using information related to the user's state or preference stored in the long-term memory storage space.

[0114] If the electronic device (100) does not have context information corresponding to the user request in the long-term memory storage space (S650-N), it can provide a response corresponding to the user request or perform an action (S680). At this time, the electronic device (100) can input the user request into an artificial intelligence model to obtain a response corresponding to the user request.

[0115] FIG. 7 is a flowchart illustrating a control method of an electronic device (100) according to one embodiment of the present disclosure.

[0116] The electronic device (100) can capture an image displayed by a display (S710). At this time, the image may be an image corresponding to the screen of an application being run or a captured image obtained through a camera.

[0117] The electronic device (100) can input a captured image into a generative AI model to obtain context information corresponding to the image (S720). At this time, the context information may include a plurality of keywords representing the situation regarding the image and information about the relationship between the plurality of keywords.

[0118] The electronic device (100) may store context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword (S730). At this time, the information related to a person's state may be information regarding a user using the electronic device (100). The information regarding a person's state may include information regarding the user's characteristics, physical state, and social state. Additionally, the preferred keyword may be a keyword indicating the user's preference. Specifically, the preferred keyword may be a keyword corresponding to a field or item that the user is interested in. The electronic device (100) may store information related to a person's state or information corresponding to a preferred keyword and related to the completion of an action among the context information in a long-term memory storage space.

[0119] The various embodiments described above may be implemented individually, but are not necessarily limited thereto, and may be implemented together in combination with at least one other embodiment, either partially or wholly.

[0120] The methods according to the various embodiments of the present disclosure described above can be implemented by software upgrade or hardware upgrade alone for an existing electronic device (100).

[0121] Meanwhile, according to a specific example of the present disclosure, the control method according to the various embodiments described above may be implemented as software comprising instructions stored on a non-transitory machine-readable storage media that can be read by various machines (e.g., computers), such as an electronic device (100).

[0122] Specifically, a program for performing a control method comprising the steps of: capturing an image displayed on a display; inputting the captured image into a generative AI model to obtain context information including context for the image; and storing the context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword, may be provided in a state stored on a non-transient computer-readable recording medium.

[0123] When stored software or instructions are executed by a processor, the processor may perform operations according to the various embodiments described above, either directly or by utilizing other components. Instructions may include code generated or executed by a compiler or an interpreter. Here, 'non-transient' means only that the storage medium does not contain a signal and is tangible, and does not distinguish whether data is stored semi-permanently or temporarily in the storage medium.

[0124] Additionally, according to one or more embodiments of the present disclosure, the method according to the various embodiments described above may be provided as included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed online through an online store, in addition to the non-transient readable recording media described above. In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created in a storage medium such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0125] Additionally, each component (e.g., module or program) according to the various embodiments described above may be composed of a single or multiple entities, and some of the aforementioned sub-components may be omitted, or other sub-components may be further included in the various embodiments. Generally or additionally, some components (e.g., module or program) may be integrated into a single entity to perform the functions performed by each of the respective components prior to integration in the same or similar manner. The operations performed by the module, program, or other components according to the various embodiments may be executed sequentially, in parallel, iteratively, or heuristically, or at least some operations may be executed in a different order, omitted, or other operations added.

[0126] Although preferred embodiments of the present disclosure have been illustrated and described above, the present disclosure is not limited to the specific embodiments described above. It is understood that various modifications can be made by those skilled in the art without departing from the essence of the present disclosure as claimed in the claims, and such modifications should not be understood individually from the technical spirit or perspective of the present disclosure.

Claims

1. In an electronic device, display, Memory in which at least one instruction is stored; and Includes a processor; When the above at least one instruction is executed collectively or individually by the processor, the electronic device, Capture the image displayed by the above display, and The above-mentioned captured image is input into a generative AI model to obtain context information corresponding to the image, and An electronic device that stores the context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword.

2. In Paragraph 1, The above context information comprises a plurality of keywords indicating the situation regarding the image and information regarding the relationship between the plurality of keywords, an electronic device.

3. In Paragraph 2, When the above at least one instruction is executed collectively or individually by the processor, the electronic device, An electronic device that stores the context information in the long-term memory storage space if the context information is information related to a person's state.

4. In Paragraph 3, When the above at least one instruction is executed collectively or individually by the processor, the electronic device, If the above context information is not information related to a person's state, the above context information is stored in the above short-term memory storage space, and An electronic device that stores context information in a long-term memory storage space based on whether the context information stored in the short-term memory storage space corresponds to the preferred keyword.

5. In Paragraph 4, When the above at least one instruction is executed collectively or individually by the processor, the electronic device, An electronic device that stores the context information stored in the short-term memory storage space in the long-term memory storage space if the context information corresponds to the preferred keyword and, among the information related to a person's behavior, is information related to the completion of the behavior.

6. In Paragraph 2, When the above at least one instruction is executed collectively or individually by the processor, the electronic device, A plurality of feature vectors corresponding to the plurality of keywords included in the above context information are obtained, and An electronic device for obtaining the preferred keyword corresponding to at least one cluster identified based on the similarity between the plurality of feature vectors.

7. In Paragraph 1, When the above at least one instruction is executed collectively or individually by the processor, the electronic device, An electronic device that, upon receiving input corresponding to a user request, provides a response corresponding to the request or performs an action based on context information corresponding to the request stored in the short-term memory storage space or the long-term memory storage space.

8. In Paragraph 7, When the above at least one instruction is executed collectively or individually by the processor, the electronic device, If context information corresponding to the above request exists in the short-term memory storage space, a response corresponding to the request is provided or an action is performed based on the context information stored in the short-term memory storage space, An electronic device that, if context information corresponding to the above request does not exist in the short-term memory storage space but exists in the long-term memory storage space, provides a response corresponding to the above request or performs an operation based on the context information stored in the long-term memory storage space.

9. A method for controlling an electronic device including a display, A step of capturing an image displayed by the above display; A step of inputting the above-mentioned captured image into a generative AI model to obtain context information corresponding to the image; and A control method comprising the step of storing the context information in a long-term memory storage space or a short-term memory storage space based on whether the context information is information related to a person's state or corresponds to a previously registered preferred keyword.

10. In Paragraph 9, A control method comprising the above context information including a plurality of keywords representing the situation regarding the image and information regarding the relationship between the plurality of keywords.

11. In Paragraph 10, The step of storing the above context information in a long-term memory storage space or a short-term memory storage space is, A control method comprising the step of storing the context information in the long-term memory storage space if the context information is information related to a person's state.

12. In Paragraph 11, The step of storing the above context information in a long-term memory storage space or a short-term memory storage space is, If the above context information is not information related to a person's state, the step of storing the above context information in the short-term memory storage space; and A control method comprising the step of storing the context information in the long-term memory storage space based on whether the context information stored in the short-term memory storage space corresponds to the preferred keyword.

13. In Paragraph 12, The step of storing the above context information in a long-term memory storage space or a short-term memory storage space is, A control method comprising the step of storing the context information in the short-term memory storage space in the long-term memory storage space if the context information stored in the short-term memory storage space corresponds to the preferred keyword and, among the information related to a person's behavior, is information related to the completion of a behavior.

14. In Paragraph 10, The above control method is, A step of obtaining a plurality of feature vectors corresponding to the plurality of keywords included in the context information; and A control method comprising the step of obtaining the preferred keyword corresponding to at least one cluster identified based on the similarity between the plurality of feature vectors.

15. In Paragraph 9, The above control method is, A control method comprising the step of, upon receiving an input corresponding to a user request, providing a response corresponding to the request or performing an action based on context information corresponding to the request stored in the short-term memory storage space or the long-term memory storage space.