Electronic device, method, and non-transitory computer-readable storage medium for acquiring information about screen

The electronic device captures and analyzes screenshots using a pre-trained model to efficiently manage context information, addressing memory and processing challenges in screen transitions and user interactions, thereby enhancing user interaction tracking.

WO2026034795A1PCT designated stage Publication Date: 2026-02-12SAMSUNG ELECTRONICS CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/KR2025/008722
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-09-24
Filing Date
2025-06-23
Publication Date
2026-02-12

AI Technical Summary

Technical Problem

Existing electronic devices face challenges in efficiently managing and analyzing context information related to screen transitions and user interactions, particularly in terms of memory requirements and processing complexity when capturing and analyzing screenshots at regular intervals.

Method used

The electronic device captures screenshots at regular intervals, stores them in a memory, and uses a pre-trained model to analyze the screenshots, obtaining context information about screen transitions and user interactions, including context information storage in a ring buffer and utilizing machine learning techniques to identify and classify objects and functions on the screen.

Benefits of technology

This approach allows for efficient management of context information, reducing memory requirements and processing complexity by analyzing screen transitions and user interactions, enabling effective tracking of user actions and screen changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure KR2025008722_12022026_PF_FP_ABST
    Figure KR2025008722_12022026_PF_FP_ABST
Patent Text Reader

Abstract

This electronic device may comprise: a memory storing instructions; a display; and at least one processor. The instructions may cause the electronic device to: display, on the display, a first screen including an executable object; receive a user input related to the executable object; display, on the display, a second screen corresponding to the executable object, on the basis of receiving the user input; obtain context information indicating that the screen displayed on the display is switched from the first screen to the second screen by the user input received through the first screen; and store the context information in the memory.
Need to check novelty before this filing date? Find Prior Art

Description

Electronic device, method, and non-transitory computer-readable storage medium for obtaining information about a screen

[0001] The present disclosure relates to an electronic device, a method, and a non-transitory computer-readable storage medium for obtaining information about a screen.

[0002] An electronic device may include a display. The electronic device may obtain data about a software application via a communication circuit. The electronic device may display a screen for the software application on the display. The electronic device may obtain data about the screen displayed on the display. The electronic device may store data about the screen displayed on the display in memory.

[0003] The above information may be provided as background art to aid in understanding the present disclosure.

[0004] No claim or determination is made as to whether any of the above is applicable as prior art to the present disclosure.

[0005] An electronic device is described. The electronic device may include a memory that stores instructions and includes one or more storage media. The electronic device may include a display. The electronic device may include at least one processor that includes a processing circuit. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a first screen including an executable object on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive a user input related to the executable object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a second screen corresponding to the executable object on the display based on receiving the user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain context information indicating that a screen displayed on the display has been switched from the first screen to the second screen based on the user input received through the first screen, based on receiving the user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store the context information in the memory.

[0006] A method is provided. The method can be executed in an electronic device having a display. The method can include an operation of displaying a first screen including an executable object on the display. The method can include an operation of receiving a user input related to the executable object. The method can include an operation of displaying a second screen corresponding to the executable object on the display based on receiving the user input. The method can include an operation of obtaining context information indicating that a screen displayed on the display is switched from the first screen to the second screen based on the user input received through the first screen, based on receiving the user input. The method can include an operation of storing the context information in the memory.

[0007] A non-transitory computer-readable storage medium is provided. The non-transitory computer-readable storage medium may store one or more programs. The one or more programs may include instructions that, when executed by an electronic device having a display, cause the electronic device to display a first screen including an executable object on the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive a user input related to the executable object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display a second screen corresponding to the executable object on the display based on receiving the user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain context information indicating that a screen displayed on the display has been switched from the first screen to the second screen based on the user input received through the first screen. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store the context information in the memory.

[0008] Figure 1 illustrates an example of an electronic device saving a screenshot of a screen.

[0009] Figure 2 is a simplified block diagram of an exemplary electronic device.

[0010] Figure 3 is a flowchart showing the operation of an electronic device that obtains context information.

[0011] Figure 4 illustrates the operation of an electronic device for obtaining context information.

[0012] Figure 5 illustrates the operation of an electronic device for obtaining context information using a screenshot stored in a ring buffer.

[0013] Figure 6 illustrates an example of obtaining context information using a portion of the screen.

[0014] Figure 7 illustrates an example in which two or more software applications are executed.

[0015] Figure 8 illustrates the operation of an electronic device for obtaining context information for each software application.

[0016] Figure 9 illustrates an example of user input for searching information.

[0017] Figure 10 illustrates the operation of an electronic device for obtaining context information about content.

[0018] Figure 11 illustrates the operation of an electronic device for obtaining context information for payment.

[0019] Figure 12 illustrates an exemplary operation of an electronic device that provides audio using context information.

[0020] Figure 13 illustrates the operation of an electronic device for obtaining context information about a photograph.

[0021] Figure 14a illustrates an example of context information provided based on an event.

[0022] Figure 14b illustrates an example of context information managed over time.

[0023] FIG. 15 is a block diagram of an electronic device within a network environment according to various embodiments.

[0024] FIG. 16 illustrates an example of a generative artificial intelligence system according to one embodiment.

[0025] Figure 1 illustrates an example of an electronic device saving a screenshot of a screen.

[0026] Referring to FIG. 1, the electronic device (100) can be used to obtain data regarding a screen. For example, the electronic device (100) can display a screen on a display (e.g., display (208) of FIG. 2). For example, the electronic device (100) can capture a screen displayed on the display. For example, the electronic device (100) can store the captured screen in a memory (e.g., memory (206) of FIG. 2). For example, the electronic device (100) can store the captured screen in a ring buffer (e.g., ring buffer (550) of FIG. 5). For example, the electronic device (100) can store a screenshot of the screen. For example, the screenshot can include a captured screen or a captured image.

[0027] The electronic device (100) can execute a software application. For example, the electronic device (100) can execute multiple software applications. For example, the electronic device (100) can display a screen (or execution screen) for a software application on the display. For example, screen (150) can be described as a screen for a software application for searching. For example, screen (160) can be described as a screen for a software application for payment. For example, screen (170) can be described as a screen for a software application for content (e.g., video content). For example, the electronic device (100) can display a screen differently based on user input. For example, the electronic device (100) can switch a screen displayed on a display (e.g., display (208) of FIG. 2) from screen (150) to screen (160) based on user input. For example, the electronic device (100) can switch a screen displayed on a display (e.g., display (208) of FIG. 2) from screen (160) to screen (170) based on a user input. For example, the electronic device (100) can switch a screen displayed on a display (e.g., display (208) of FIG. 2) from screen (170) to screen (150) based on a user input.

[0028] The electronic device (100) can periodically store data regarding a screen displayed on a display (e.g., the display (208) of FIG. 2). For example, the electronic device (100) can periodically store screenshots of the screen displayed on the display in the memory. For example, the electronic device (100) can store the screenshots at regular intervals (e.g., every 10 seconds). For example, the electronic device (100) can analyze the stored screenshots. For example, the electronic device (100) can analyze the screen by providing the stored screenshots to a pre-trained model (e.g., the pre-trained model (430) of FIG. 4). For example, the electronic device (100) can obtain information regarding objects included in the screen by analyzing the screen. For example, the electronic device (100) can obtain information regarding functions corresponding to each object included in the screen. Since the electronic device (100) acquires screenshots at regular intervals, the electronic device (100) may be required to include a large capacity memory.

[0029] The electronic device (100) may acquire a screenshot based on satisfying certain conditions to save memory for storing screenshots. For example, the electronic device (100) may store a screenshot of a screen displayed on the display based on receiving a user input. For example, the electronic device (100) may change the screen displayed on the display from a first screen to a second screen based on receiving a user input. For example, the electronic device (100) may acquire context information (e.g., context information (440) of FIG. 4) indicating that the screen has changed from the first screen to the second screen based on receiving a user input.

[0030] For example, the electronic device (100) may include hardware components used to perform or execute the above operations. The hardware components are described and exemplified with reference to FIG. 2.

[0031] Figure 2 is a simplified block diagram of an exemplary electronic device.

[0032] Referring to FIG. 2, the electronic device (100) may include at least one processor (207), a display (208), a memory (206), a speaker (210), and a camera (209).

[0033] At least one processor (207) may include a hardware component for processing data using instructions stored in the memory (206). The hardware component for processing data may include a central processing unit (CPU) (e.g., including processing circuitry). The hardware component for processing data may include a graphic processing unit (GPU) (e.g., including processing circuitry). The hardware component for processing data may include a display processing unit (DPU) (e.g., including processing circuitry). The hardware component for processing data may include a neural processing unit (NPU) (e.g., including processing circuitry).

[0034] At least one processor (207) may include one or more cores. For example, at least one processor (207) may have a multi-core processor structure such as a dual core, a quad core, or a hexa core.

[0035] The memory (206) may include hardware components for storing data and / or instructions input to and / or output from at least one processor (207). The memory (206) may include, for example, volatile memory such as random-access memory (RAM) and / or non-volatile memory such as read-only memory (ROM). The volatile memory may include, for example, at least one of dynamic RAM (DRAM), static RAM (SRAM), cache RAM, and pseudo SRAM (PSRAM). The non-volatile memory may include, for example, at least one of programmable ROM (PROM), erasable PROM (EPROM), electrically erasable PROM (EEPROM), flash memory, hard disk, compact disc, and embedded multimedia card (EMMC).

[0036] The display (208) can output visualized information. For example, the display (208) can output visualized information to the user under the control of at least one processor (207). The display (208) can include hardware components of the electronic device (100) used to display a screen. For example, the display (208) can include light-emitting elements and circuits (e.g., transistors) that control the light-emitting elements to emit light. For example, each of the light-emitting elements can include an organic light emitting diode (OLED) or a micro LED. However, the present invention is not limited thereto. For example, the display (208) can include a liquid crystal display (LCD).

[0037] The camera (209) may include one or more optical sensors (e.g., a charged coupled device (CCD) sensor, a complementary metal oxide semiconductor (CMOS) sensor) that generate electrical signals representing the color and / or brightness of light. For example, the camera (209) may be described as one or more image sensors. For example, the camera (209) may have a field of view (FOV) corresponding to the FOV of a user's eye. For example, the camera (209) may be used to acquire images of the eye of a user wearing the electronic device (100) (e.g., the user (1210) of FIG. 12). For example, the camera (209) may be included in a camera module (e.g., the camera module (1580) of FIG. 15).

[0038] The electronic device (100) may include a speaker (210). For example, the speaker (210) may be used to convert audio data into an audio signal. For example, the speaker (210) may be used to output audio representing context information.

[0039] At least one processor (207) can display a first screen (e.g., the first screen (410) of FIG. 4) including an executable object on the display (208). For example, the display (208) can be used to display a first screen (e.g., the first screen (410) of FIG. 4) or a second screen (e.g., the second screen (420) of FIG. 4). At least one processor (207) can receive a user input. For example, at least one processor (207) can obtain context information (e.g., the context information (440) of FIG. 4) indicating that a screen displayed on the display (208) is switched from the first screen to the second screen based on a user input. At least one processor (207) can obtain an image through the camera (209). At least one processor (207) can obtain an image through the camera (209) based on receipt of a user input. For example, a camera (209) may be used to acquire an image. At least one processor (207) may output audio through a speaker (210) using context information (e.g., context information (440) of FIG. 4). For example, the speaker (210) may be used to output audio representing the context information.

[0040] FIG. 3 is a flowchart illustrating the operation of an electronic device for acquiring context information. This method may be executed by the electronic device (100) illustrated in FIG. 3 or at least one processor (207) of the electronic device (100).

[0041] In the following examples, each operation may be performed sequentially, but is not necessarily performed sequentially. For example, the order of each operation may be changed, and at least two operations may be performed in parallel.

[0042] Referring to FIG. 3, in operation 310, at least one processor (207) may display a first screen (e.g., the first screen (410)) including an executable object (e.g., the executable object (930) of FIG. 9) on the display (208). For example, the first screen may include an executable object (or a UI (user interface) object). For example, the executable object (e.g., the executable object (930) of FIG. 9) may be described as an object for executing a pre-designated function within the first screen. For example, the executable object may include an executable object related to a designated software application executed within the electronic device (100). However, the present invention is not limited thereto. For example, the executable object may include an executable object related to an operating system (OS) executed within the electronic device (100). For example, the executable object may be an executable object included in a home screen. For example, the executable object may include an object for changing the settings of the OS. For example, the executable object may include an object for searching for a function within the OS.

[0043] In operation 320, at least one processor (207) may receive user input (e.g., user input (1050)) related to the executable object. For example, the user input may include touch input. For example, the user input may include voice input. For example, the user input may include input to an input device. For example, the user input may include gestures.

[0044] In operation 330, at least one processor (207) may display a second screen (e.g., the second screen (420) of FIG. 4) corresponding to an executable object (e.g., the executable object (930) of FIG. 9) on the display (208) based on receiving a user input (e.g., the user input (1050)). For example, at least one processor (207) may change a screen displayed on the display (208) from the first screen to the second screen based on receiving the user input. For example, at least one processor (207) may stop displaying the first screen and then display the second screen based on receiving the user input.

[0045] In operation 340, at least one processor (207) may obtain context information (e.g., context information (440) of FIG. 4) indicating that a screen displayed on the display (208) is switched from the first screen to the second screen based on a user input (e.g., user input (1050)) received through a first screen (e.g., the first screen (410) of FIG. 4)). For example, the at least one processor (207) may obtain context information (440) stored together with a timestamp indicating a time at which the user input was received. According to one embodiment, the context information (e.g., context information (440) of FIG. 4) may include content describing the first screen. For example, the context information (e.g., context information (440) of FIG. 4) may include content describing the content (e.g., image, text) of the first screen. In one embodiment, the context information (e.g., context information (440) of FIG. 4) may include content describing the second screen. For example, the context information (e.g., context information (440) of FIG. 4) may include content describing the content (e.g., image, text) of the second screen.

[0046] According to one embodiment, at least one processor (207) may display a first screen for searching (e.g., screen (150) of FIG. 1) on the display (208). For example, at least one processor (207) may receive a user input for entering a word. For example, at least one processor (207) may display a second screen including search results for the entered word on the display. For example, at least one processor (207) may obtain context information for the search. For example, the context information may include information indicating that a specific word was searched for at a specific time.

[0047] According to one embodiment, at least one processor (207) may display a first screen including a pop-up window. For example, at least one processor (207) may receive a user input for removing the pop-up window. For example, at least one processor (207) may display a second screen from which the pop-up window has been removed based on receiving the user input. For example, at least one processor (207) may obtain context information indicating that the user is not interested in the content included in the pop-up window. For example, at least one processor (207) may obtain context information indicating that the user of the electronic device (100) is not focused on the content included in the pop-up window based on the screen displayed on the display (208) switching from the first screen to the second screen. However, the present invention is not limited thereto. For example, at least one processor (207) may receive a user input for checking the content included in the pop-up window. For example, at least one processor (207) may obtain context information indicating interest in content included in the pop-up window.

[0048] According to one embodiment, at least one processor (207) may receive another user input regarding an executable object included in the second screen while displaying the second screen through the display (208). For example, based on receiving the other user input, the at least one processor (207) may switch from the second screen, from which the pop-up window is removed, to a third screen. For example, based on receiving the other user input, the at least one processor (207) may obtain context information indicating that the second screen is switched to the third screen. For example, the context information may not include information about the pop-up window. For example, the context information may not include information about the user input for removing the pop-up window.

[0049] According to one embodiment, at least one processor (207) may display a first screen including a photo library on the display (208). For example, at least one processor (207) may display a first screen including an executable object for executing a sharing function. For example, at least one processor (207) may receive a user input for sharing an image with a wife. For example, at least one processor (207) may execute a software application for sharing based on receiving the user input. For example, the sharing may be described as a function for transmitting data stored in the memory (206) of the electronic device (100) to another electronic device (not shown). For example, at least one processor (207) may obtain context information (e.g., context information (440) of FIG. 4) representing information about the image and information about another electronic device to which the image will be transmitted based on receiving the user input. For example, the information about the image may include information describing the image. For example, the information about the other electronic device may include information about a relationship (e.g., a couple) between a first user of the electronic device (100) and a second user of the other electronic device. For example, at least one processor (207) may obtain information describing the image by providing the image to a pre-trained model (e.g., the pre-trained model (430) of FIG. 4). For example, at least one processor (207) may obtain information about a relationship between the first user and the second user by providing communication information of the other electronic device (e.g., a phone number or a name for the other electronic device registered within the electronic device (100)) to a pre-trained model (e.g., the pre-trained model (430) of FIG. 4).For example, at least one processor (207) may obtain context information indicating that the image has been shared with the wife based on receiving a user input to share an image representing a child with the wife, who is a user of another electronic device.

[0050] For example, at least one processor (207) may store context information (e.g., context information (440) of FIG. 4) obtained based on a user input for sharing content together with the content. For example, the content may include context information (e.g., context information (440) of FIG. 4) obtained based on the user input. For example, context information (e.g., context information (440) of FIG. 4) obtained based on receiving a user input for sharing content may be stored as metadata for the content. For example, context information (e.g., context information (440) of FIG. 4) obtained based on receiving a user input may be stored in (or together with) the content in the form of an EXIF ​​(exchangeable image file format) file for the content. For example, the form in which context information (e.g., context information (440) of FIG. 4) obtained based on receiving the above user input is stored is not limited to EXIF.

[0051] For example, at least one processor (207) can use context information (e.g., context information (440) of FIG. 4) generated by executing a function for sharing content as a database. For example, at least one processor (207) can refer to context information (e.g., context information (440) of FIG. 4) while browsing content. For example, at least one processor (207) can refer to context information (e.g., context information (440) of FIG. 4) stored in the form of a database when executing a software application that browses content. For example, at least one processor (207) can store context information (e.g., context information (440) of FIG. 4) in the form of a database that can be referenced by a software application that browses content, in the memory (206) of the electronic device (100). For example, at least one processor (207) may store context information (e.g., context information (440) of FIG. 4) in the form of a database that can be referenced by a software application performing content browsing within an external electronic device (e.g., a server or a cloud server). For example, at least one processor (207) may cause the external electronic device to store the context information (e.g., context information (440) of FIG. 4) by transmitting the context information (e.g., context information (440) of FIG. 4) to the external electronic device via a communication circuit.

[0052] The acquisition of the above context information is described and illustrated in more detail with reference to FIG. 4.

[0053] Figure 4 illustrates the operation of an electronic device for obtaining context information.

[0054] Referring to FIG. 4, at least one processor (207) can obtain context information (440) using a pre-trained model (430). For example, at least one processor (207) can obtain context information (440) by providing data for a first screen (410) and data for a second screen (420) to the pre-trained model (430). For example, at least one processor (207) can provide a screenshot for the first screen (410) and a screenshot for the second screen (420) to the pre-trained model (430). For example, at least one processor (207) can provide an image captured for the first screen (410) and an image captured for the second screen (420) to the pre-trained model (430). For example, context information (440) may include information about a process of switching to a second screen (420) upon receiving a user input (e.g., user input (1050) of FIG. 10) from a first screen (410). For example, context information (440) may include information about a function that a user (e.g., user (1210)) attempts to execute by providing a user input for an executable object (e.g., executable object (930) of FIG. 9) included in the first screen (410).

[0055] The pre-trained model (430) can be described as a model trained using screen data. For example, state (450) can be described as a state in which the pre-trained model (430) is trained using screen data. For example, the pre-trained model (430) can be trained using multiple screens (e.g., screen (460-1), screen (460-2), and screen (460-N)). For example, the pre-trained model (430) can be trained using a machine learning technique.

[0056] For example, the pre-trained model (430) can analyze objects included in the screen. For example, the pre-trained model (430) can classify the types of objects included in the screen. For example, the pre-trained model (430) can identify text included in the screen. For example, the pre-trained model (430) can identify text included in the screen through an optical character recognition (OCR) technique. For example, the OCR technique can be referred to as optical character recognition (or digital character recognition). For example, the OCR technique can be described as a technique that identifies text using light reflected from text (or a technique that recognizes text in a digitized image). For example, the pre-trained model (430) can match text included in the screen with an executable object. For example, the pre-trained model (430) can determine or infer a function caused by a user input by learning that the user input transitions from a first screen to a second screen.

[0057] The pre-trained model (430) can analyze the screen using the OCR technique. However, the present invention is not limited thereto. For example, the pre-trained model (430) can analyze the screen using XML (extensible markup language). For example, at least one processor (207) can provide XML data for the first screen (410) to the pre-trained model (430). For example, at least one processor (207) can provide XML data for the second screen (420) to the pre-trained model (430). For example, at least one processor (207) can identify a function corresponding to an executable object included in the screen using the XML data. For example, at least one processor (207) can identify a function caused by a received user input using the XML data. For example, the pre-trained model (430) can be trained using XML data for screens (e.g., screen (460-1), screen (460-2), and screen (460-N)). At least one processor (207) can obtain context information (440) by providing XML data to the pre-trained model (430). However, the present invention is not limited thereto. For example, the at least one processor (207) can convert XML data into an image. For example, the at least one processor (207) can obtain an image corresponding to the XML data using the XML data, and then provide the obtained image to the pre-trained model (430) to obtain context information (440). For example, the pre-trained model (430) can be provided with data for a screen or a screenshot of the screen to analyze text and executable objects included in the screen. For example, a pre-trained model (430) can be provided with a screenshot of a screen and then describe or explain functions corresponding to text and executable objects included in the screen.

[0058] At least one processor (207) can obtain context information (440) using a pre-trained model (430). For example, at least one processor (207) can obtain context information (440) by providing a screenshot of the first screen (410), a screenshot of the second screen (420), and information about user input to the pre-trained model (430). For example, the information about the user input can include coordinate information about the user input on the display (208). For example, if XML data does not exist, at least one processor (207) can identify an executable object corresponding to the user input using the coordinate information about the user input received on the display (208).

[0059] For example, at least one processor (207) may display a first screen (410) for a software application on the display (208). For example, at least one processor (207) may use an OCR technique or XML data to analyze the first screen (410). For example, if at least one processor (207) is unable to obtain XML data, the OCR technique may be used.

[0060] The context information (440) may include information indicating that the screen has been switched from the first screen (410) to the second screen (420). For example, the context information (440) may include text data. For example, the context information (440) may include words or sentences expressing the process of switching from the first screen (410) to the second screen (420). However, the present invention is not limited thereto. For example, the context information (440) may include raw data for the first screen (410). For example, the context information (440) may include raw data for the second screen (420). For example, at least one processor (207) may determine the format of the context information (440) acquired through the pre-trained model (430). For example, at least one processor (207) may acquire the context information (440) in the form of raw data. For example, at least one processor (207) can obtain context information (440) in processed form.

[0061] At least one processor (207) can train a pre-trained model (430) using the first screen (410) and the second screen (420). For example, at least one processor (207) can train the pre-trained model (430) using the LoRA (low-rank adaptation) technique. For example, the LoRA technique can be described as a technique for inserting a trainable matrix into a model. For example, the LoRA technique can be described as a technique for fine-tuning the weights of matrices within a model by training the inserted matrix. For example, at least one processor (207) can obtain context information (440) using the pre-trained model (430). For example, at least one processor (207) can train a pre-trained model (430) using the acquired screenshots of the first screen (410) and the second screen (420) while displaying the first screen (410) or the second screen (420) on the display (208) via the LoRA technique. For example, at least one processor (207) can tune the weights of the pre-trained model (430) using the data for the first screen (410) and the data for the second screen (420). For example, at least one processor (207) can adjust the weight vector of the pre-trained model (430) using the data for the first screen (410) and / or the data for the second screen (420). For example, the electronic device (100) can acquire context information (440) more suitable for the situation via the LoRA technique. For example, the electronic device (100) can analyze the first screen (410) and the second screen (420) in more detail by adjusting the weight vector.

[0062] In one embodiment, the pre-trained model (430) can generate context information (440) using characteristics or functions of each of the software applications used for learning. For example, the pre-trained model (430) can be trained by grouping the functions or characteristics of each of a plurality of software applications. For example, the pre-trained model (430) can be classified into software applications for content playback, software applications for communication, or software applications for payment. The types of each of the plurality of software applications are described, but this is merely exemplary. For example, at least one processor (207) can obtain context information (440) using characteristics or functions of a software application corresponding to a screen being displayed through a display (208). For example, at least one processor (207) may obtain context information (440) using a pre-trained model (430) adjusted for a software application for payment based on identifying that the first screen displayed through the display (208) is a software application for payment. For example, at least one processor (207) may identify that the first screen is a screen for a software application for playing content. For example, at least one processor (207) may obtain context information (440) using a pre-trained model (430) whose weight vector is adjusted according to the software application for playing content based on identifying that the first screen is a screen for a software application for playing content.

[0063] At least one processor (207) may obtain context information (440) based on receipt of a user input (e.g., user input (1050)). For example, at least one processor (207) may transmit a first screenshot of the first screen (410) to a pre-trained model (430) from a ring buffer (e.g., ring buffer (550) of FIG. 5) based on the user input. The ring buffer is described and exemplified in more detail with reference to FIG. 5.

[0064] Figure 5 illustrates the operation of an electronic device for obtaining context information using a screenshot stored in a ring buffer.

[0065] Referring to FIG. 5, at least one processor (207) may store data regarding a screen in a ring buffer (550). For example, the ring buffer (550) may be described as a memory area for temporarily storing data. For example, the ring buffer (550) may include a First-In First-Out (FIFO) structure. For example, the ring buffer (550) may include a buffer for storing a specified number of images. For example, the ring buffer may be formed in a volatile memory. For example, at least one processor (207) may repeatedly store screenshots of a screen displayed on a display (208) in the ring buffer according to a specified period (e.g., 10 seconds). For example, at least one processor (207) may obtain a first screenshot of a first screen (410) that was last stored in the ring buffer (550) based on receiving a user input (e.g., user input (1050)). For example, at least one processor (207) can obtain context information (440) using a second screenshot for a second screen (420) displayed on a display (208) and a first screenshot for a first screen (410).

[0066] For example, at least one processor (207) can store a screenshot of the screen (510) and a screenshot of the screen (520) in the ring buffer (550). For example, at least one processor (207) can store a screenshot of the screen (510) and a screenshot of the screen (520) in the ring buffer (550) while displaying the screen (530) on the display (208). For example, at least one processor (207) can store a screenshot of the screen (530) being displayed on the display (208) in the ring buffer (550) as a specified period of time elapses. For example, at least one processor (207) can sequentially display the screens (510), (520), and (530) on the display (208). For example, at least one processor (207) can sequentially display screen (510), screen (520), and screen (530) on the display (208) based on user input or over time. For example, at least one processor (207) can sequentially store a screenshot of screen (510), a screenshot of screen (520), and a screenshot of screen (530) in a ring buffer (550).

[0067] For example, at least one processor (207) may obtain a screenshot of the screen (520) from the ring buffer (550) based on a user input (e.g., user input (1050)). For example, at least one processor (207) may obtain a screenshot of the screen (530) through the display (208) based on the user input. For example, at least one processor (207) may obtain context information (440) by providing the screenshot of the screen (520) and the screenshot of the screen (530) to a pre-trained model (430). However, the present invention is not limited thereto. For example, at least one processor (207) may obtain context information (440) by storing a screenshot of a screen (530) in a ring buffer (550) based on user input and then providing the screenshot of the screen (530) stored in the ring buffer (550) to a pre-trained model (430). The context information (440) is described and illustrated in more detail with reference to FIG. 6.

[0068] Figure 6 illustrates an example of obtaining context information using a portion of the screen.

[0069] Referring to FIG. 6, at least one processor (207) can obtain context information (440) using a portion of a first screen (410) and a portion of a second screen (420). For example, at least one processor (207) can display the first screen (410) on the display (208). For example, at least one processor (207) can change the screen displayed on the display (208) from the first screen (410) to the second screen (420) based on a user input (e.g., user input (1050) of FIG. 10). For example, at least one processor (207) can display an area (630) and an area (640) included in the first screen (410) on the display (208). For example, at least one processor (207) may display a second screen (420) including an area (650) and an area (660) on the display (208) based on the user input. For example, at least one processor (207) may maintain a visual object (670) in an area (630) on the second screen (420) while switching from the first screen (410) to the second screen (420) based on the user input. For example, at least one processor (207) may display a visual object (670) in an area (650) of the second screen (420) at the same location as the area (630) of the first screen (410). For example, at least one processor (207) may display a visual object (670) included in an area (630) of the first screen (410) on an area (650) of the second screen (420). For example, regions (630) and (650) may be referred to as fixed regions or invariant regions. For example, the fixed region or invariant region may be described as regions that are maintained independently of user input when the screen switches from the first screen to the second screen based on receiving user input.At least one processor (207) may stop displaying a visual object (680) on the display (208) based on the user input. For example, at least one processor (207) may refrain from or bypass displaying a visual object (680) included in a region (640) of the first screen (410) in a region (660) of the second screen (420). For example, at least one processor (207) may identify regions (630) and (640) as unchanging regions. For example, at least one processor (207) may identify regions (640) and (660) as changeable regions. For example, regions (640) and (660) may be referred to as variable regions. For example, the variable region may be described as a region that changes depending on the user input when the screen displayed through the display (208) switches from the first screen to the second screen based on receiving the user input. For example, at least one processor (207) may obtain context information (440) by providing a screenshot of the region (640) and a screenshot of the region (660) to the pre-trained model (430) based on the user input. For example, at least one processor (207) may refrain from or bypass providing the screenshot of the region (630) and the screenshot of the region (650) to the pre-trained model (430). For example, even though at least one processor (207) receives the user input, the region (630) and the region (650) may not be used to obtain context information (440) because they do not change. For example, at least one processor (207) can obtain context information (440) by providing a screenshot of a cropped area (640) of a first screen (410) and a screenshot of a cropped area (660) of a second screen (420) to a pre-trained model (430).

[0070] At least one processor (207) can obtain context information (440) using screens for each of the software applications. The context information (440) for each of the software applications is described and illustrated in more detail with reference to FIG. 7.

[0071] Figure 7 illustrates an example in which two or more software applications are executed.

[0072] Referring to FIG. 7, at least one processor (207) can display independent screens through the display (208). For example, at least one processor (207) can display an area (730) included in a first screen (410) and an area (740) included in the first screen (410) on the display (208). For example, at least one processor (207) can display an area (750) included in a second screen (420) and an area (760) included in the second screen (420) on the display (208) based on receiving a user input (e.g., user input (1050) of FIG. 10). For example, areas (730) and (750) can be referred to as first display areas. For example, areas (740) and (760) can be referred to as second display areas. At least one processor (207) can display independent screens using areas (730) and (740) while displaying the first screen (410) on the display (208). For example, at least one processor (207) can display a screen for searching on area (730). For example, at least one processor (207) can display a screen for video content on area (740). For example, at least one processor (207) can display a screen independent of the screen displayed on area (730) on area (740). For example, at least one processor (207) can change the screen displayed on area (730) by receiving a user input through area (730). For example, at least one processor (207) can maintain the screen displayed on area (740) while changing the screen displayed on area (730) when receiving a user input through area (730). For example, at least one processor (207) may maintain the screen displayed in the area (730) while changing the screen displayed in the area (740) when receiving user input through the area (740).For example, at least one processor (207) can simultaneously display a screen for a first software application and a screen for a second software application on the display (208). For example, at least one processor (207) can display a first screen (410) or a second screen (420) for the first software application in the first display area. For example, at least one processor (207) can display a third screen (e.g., a screen displayed in area (740)) or a fourth screen (e.g., a screen displayed in area (760)) for the second software application in the second display area. For example, at least one processor (207) can display the third screen including another executable object. For example, at least one processor (207) can receive another user input related to another executable object. For example, at least one processor (207) can display the fourth screen corresponding to the other executable object on the second display area based on receiving another user input. For example, at least one processor (207) may, based on receiving another user input, obtain other context information indicating that the screen displayed on the second display area has been switched from the third screen to the fourth screen by another user input received through the third screen. For example, at least one processor (207) may store the other context information in the memory (206).

[0073] According to one embodiment, at least one processor (207) may obtain context information (440) by cropping a portion of the first display area and then providing a screenshot of the cropped portion to a pre-trained model (430). For example, at least one processor (207) may display a first screen for a preset software application on the first display area. For example, at least one processor (207) may receive a user input while displaying the first screen. For example, at least one processor (207) may switch the screen displayed on the first display area from the first screen to a second screen based on receiving the user input. For example, at least one processor (207) may analyze the first portion and the second portion by providing a screenshot of each of a first portion cropped from the first screen and a second portion cropped from the second screen to a pre-trained model (430) based on receiving the user input. For example, at least one processor (207) can obtain context information (440) for the first part and the second part based on the provision.

[0074] According to one embodiment, a technique for obtaining context information (e.g., first context information) for a first software application and a technique for obtaining other context information (e.g., second context information) for a second software application may be different from each other. For example, at least one processor (207) may analyze the first screen (410) and the second screen (420) using an OCR technique to obtain context information (e.g., first context information) for the first software application. For example, at least one processor (207) may analyze the third screen and the fourth screen using XML data to obtain other context information (e.g., second context information) for the second software application.

[0075] At least one processor (207) can obtain context information (440) that hierarchically represents the screen. For example, at least one processor (207) can switch the screen displayed on the display (208) from the first screen (410) to the second screen (420) based on receiving a user input (e.g., user input (1050) of FIG. 10). For example, at least one processor (207) can switch from the first screen (410) to a third screen (not shown) based on receiving another user input. The first screen (410) can be connected to another screen based on the number of executable objects included in the first screen (410). For example, at least one processor (207) can switch from the first screen (410) to one of the second screen (420), the third screen (not shown), and the fourth screen (not shown) based on the user input. For example, at least one processor (207) may obtain context information (440) in which screens are hierarchically expressed based on the screens displayed on the display (208) being switched from the first screen (410) to one of the second screen, the third screen, and the fourth screen. For example, the context information (440) may include information about the second screen (420), the third screen, and the fourth screen, which are derived from the first screen (410).

[0076] According to one embodiment, at least one processor (207) may, while displaying a first screen (410) for a first software application on a display (208), display a second screen (420) for a second software application on the display (208) based on receiving a user input. For example, at least one processor (207) may, based on the user input, obtain context information (440) indicating that the software applications are switched (e.g., information indicating that a first screen of the first software application is switched to a second screen of the second software application).

[0077] According to one embodiment, at least one processor (207) may receive a user input for terminating the software application for the first screen (410) while displaying the first screen (410) for the first software application on the display (208). For example, based on receiving the user input, the at least one processor (207) may obtain context information (440) indicating termination of the first software application.

[0078] According to one embodiment, at least one processor (207) can obtain context information (440) by analyzing the screen for each of the regions. For example, at least one processor (207) can obtain context information (440) by providing a screenshot for a region (730) of a first screen (410) and a screenshot for a region (750) of a second screen (420) to a pre-trained model (430). For example, when a screen for a first software application (e.g., a software application for searching) and a screen for a second software application (e.g., a software application for video content) are simultaneously displayed on the display (208), at least one processor (207) can obtain context information (440) for each of the software applications. For example, at least one processor (207) can obtain context information (440) for the first software application and context information (440) for the second software application. Context information (440) for each of the above software applications is described and illustrated in more detail with reference to FIG. 8.

[0079] Figure 8 illustrates the operation of an electronic device for obtaining context information for each software application.

[0080] Referring to FIG. 8, at least one processor (207) may obtain context information (440) for each of the software applications. For example, at least one processor (207) may group context information (440) for each of the software applications. For example, at least one processor (207) may display a screen (810) for a first software application on a display (208). For example, at least one processor (207) may obtain context information (440-1) (e.g., first context information) for the first software application based on receiving a user input for an executable object included in the screen (810). For example, at least one processor (207) may display a screen (820) for a second software application on the display (208). For example, at least one processor (207) may obtain context information (440-2) (e.g., second context information) for a second software application based on receiving a user input for an executable object included in a screen (820). For example, at least one processor (207) may display a screen (830) for a third software application on the display (208). For example, at least one processor (207) may obtain context information (440-3) (e.g., third context information) for a third software application based on receiving a user input for an executable object included in a screen (830). For example, at least one processor (207) may obtain context information (440) in which screens are hierarchically expressed for each of the software applications.

[0081] According to one embodiment, at least one processor (207) may obtain data to be provided to a pre-trained model (430) of a software application when the software application is installed in the electronic device (100). For example, the data may include a screenshot of the software application. For example, the data may include XML data. For example, at least one processor (207) may analyze a screen that can be provided by the software application by providing the data to the pre-trained model (430) when the software application is installed. For example, at least one processor (207) may obtain information about a screen that can be displayed through the display (208) while the software application is running in the electronic device (100) by providing the data to the pre-trained model (430). For example, at least one processor (207) may obtain information about executable objects in a screen provided by the software application before the software application is executed by providing the data to a pre-trained model (430) when installing the software application. At least one processor (207) may update the installed software application. For example, at least one processor (207) may update information about a screen provided by the software application when updating the software application. For example, at least one processor (207) may additionally obtain data about a screen of the software application when updating the software application. For example, at least one processor (207) may update information about the screen by providing the data to a pre-trained model (430).For example, the information about the screen may include information about the user interface (UI) of the software application. For example, the information may include information about an executable object within the UI of the software application. For example, the information may include information about a function corresponding to the executable object within the UI.

[0082] At least one processor (207) may receive user input for retrieving information. For example, the user input may include scrolling. User input for retrieving information is described and exemplified in more detail with reference to FIG. 9.

[0083] Figure 9 illustrates an example of user input for searching information.

[0084] Referring to FIG. 9, at least one processor (207) may obtain context information (440) based on receiving a user input for scrolling. For example, the state (910) may be described as a search result for the Paris Olympics. For example, the first screen (410) may include an executable object (930) for scrolling. For example, at least one processor (207) may change a screen displayed on the display (208) from the first screen (410) to the second screen (420) based on receiving a user input for controlling the executable object (930). For example, at least one processor (207) may identify a start point and an end point of the executable object (930). For example, at least one processor (207) may obtain context information (440) indicating that information located between the start point and the end point is not of interest based on identifying the start point and the end point. For example, the state (920) can be described as a state of scrolling down from the first screen (410). For example, at least one processor (207) can obtain context information (440) indicating that the second screen (420) is displayed by scrolling from the first screen (410) based on receiving a user input for scrolling. For example, at least one processor (207) can obtain context information (440) indicating that the user input is an action for searching for content based on receiving a user input for scrolling on the first screen (410). For example, the context information (440) can include information indicating that the user input is scrolling to check the schedule of the Paris Olympics.

[0085] According to one embodiment, at least one processor (207) may obtain context information (440) based on receiving a user input for searching for information. Referring to FIG. 9, an operation for obtaining context information (440) for a software application for searching is illustrated, but the present disclosure is not limited thereto. For example, at least one processor (207) may obtain context information (440) based on receiving a user input for searching on a home screen of the electronic device (100). For example, the home screen may include a first screen displayed through the display (208) when the state of the electronic device (100) is switched from a power-off state to a power-on state. For example, at least one processor (207) may receive a user input for searching for information of the electronic device (100) within the home screen. For example, at least one processor (207) may obtain context information (440) related to the home screen based on receiving the user input.

[0086] At least one processor (207) may receive user input for controlling video content. Control of the video content is described and illustrated in more detail with reference to FIG. 10.

[0087] Figure 10 illustrates the operation of an electronic device for obtaining context information about content.

[0088] Referring to FIG. 10, at least one processor (207) may receive a user input for controlling video content. For example, a state (1010) may be described as a state in which a screen on which video content is played is displayed on a display (208). For example, a state (1020) may be described as a state in which a user input (1050) for controlling the playback speed of the video content is received. For example, at least one processor (207) may display the video content on an area (1030) included in a first screen (410). For example, the area (1030) may include an executable object (1040) for controlling the playback speed of the video content. For example, at least one processor (207) may obtain context information (440) based on receiving a user input (1050) for the executable object (1040). For example, at least one processor (207) may obtain context information (440) indicating that a specific video content is to be played at double speed. Referring to FIG. 10, an executable object for controlling the playback speed is illustrated, but this is merely exemplary. For example, the first screen (410) may include a time bar for determining the playback section of the video content.

[0089] According to one embodiment, at least one processor (207) may obtain context information (440) including text summarizing the video content based on receiving a user input for controlling the video content. For example, the user input may include a user input for playing the video content. For example, at least one processor (207) may obtain text representing the video content by analyzing a video frame included in the video content based on the video content being displayed on the display (208). For example, at least one processor (207) may obtain context information (440) summarizing the video content based on analyzing the video content. For example, at least one processor (207) may obtain context information (440) indicating that a video content of a dog being walked is being played.

[0090] At least one processor (207) may receive information about video content from an external electronic device (not shown) through a communication circuit (not shown). For example, when acquiring or generating context information (440), at least one processor (207) may receive information about video content from the external electronic device and then provide the information to a pre-trained model (430). For example, at least one processor (207) may use a uniform resource locator (URL) included in the first screen (410) to acquire information about the first screen (410) from the external electronic device (e.g., the creator of the video content, the playback time of the video content, or the content of the video content). For example, at least one processor (207) may analyze or acquire information about the URL by calling an application programming interface (API) corresponding to the URL based on identifying the URL. For example, at least one processor (207) may perform an OCR technique on the first screen (410), thereby providing the acquired data to a pre-trained model (430), and then obtaining context information (440). The data may include a title of the video content and / or the content of the video content.

[0091] At least one processor (207) can obtain context information (440) including the content of shopping by analyzing the screen for shopping. The context information (440) for shopping is described and exemplified in more detail with reference to FIG. 11.

[0092] Figure 11 illustrates the operation of an electronic device for obtaining context information for payment.

[0093] Referring to FIG. 11, at least one processor (207) may display a screen for payment on the display (208). For example, at least one processor (207) may generate or obtain contextual information (440) for payment based on receiving a user input. For example, state (1110) may be described as a state in which a screen including card information for payment is displayed on the display (208). For example, state (1120) may be described as a state in which a screen for payment details is displayed on the display (208). For example, at least one processor (207) may display a first screen (410) including an executable object (1130) on the display (208). For example, at least one processor (207) may display a second screen (420) based on a user input for the executable object (1130). For example, the second screen (420) may include details for payment. For example, the details may include one or more of the payment amount, payment location, and payment details. For example, at least one processor (207) may generate or obtain contextual information (440) indicating a purchase of a specific item at a specific location using a card corresponding to an executable object (1130) at a specific time based on the user input.

[0094] According to one embodiment, at least one processor (207) can obtain information regarding offline payment. For example, when a user wearing an external electronic device including a camera and a head mounted display (HMD) makes an offline payment using the electronic device (100), the electronic device (100) can obtain information to be provided to a pre-trained model (430) by identifying an item the user is purchasing through the camera. For example, the electronic device (100) can obtain context information (440) indicating the purchase of a specific item by providing the information obtained from the external electronic device to the pre-trained model (430).

[0095] Referring back to FIG. 3, at operation 350, at least one processor (207) may store context information (440) in memory (206). For example, at least one processor (207) may use the stored context information (440) to provide it to a user (e.g., user (1210) of FIG. 12). The provision of the context information (440) is described and exemplified in more detail with reference to FIG. 12.

[0096] Figure 12 illustrates an exemplary operation of an electronic device that provides audio using context information.

[0097] Referring to FIG. 12, at least one processor (207) may provide context information (440) based on other user input. For example, at least one processor (207) may identify an event for the context information (440). For example, the event may include receiving a user input. For example, a user (1210) may provide a user input to inquire about the context information (440) to at least one processor (207). For example, at least one processor (207) may receive a user input that causes the at least one processor (207) to output a result using the context information (440). For example, at least one processor (207) may output audio representing the context information (440) stored in the memory (206) through the speaker (210) based on identifying the event. For example, the electronic device (100) may include a microphone (not shown). For example, at least one processor (207) may receive a voice of a user (1210) through the microphone. For example, at least one processor (207) may output audio representing context information (440) related to the voice through the speaker (210) based on the reception of the voice. For example, at least one processor (207) may output audio including context information (440) about the voice through the speaker (210) based on the reception of the voice. For example, at least one processor (207) may output the audio using a voice assistant. For example, at least one processor (207) may determine context information (440) related to the voice by identifying the voice. For example, at least one processor (207) may determine a portion of stored context information (440) corresponding to the voice by analyzing the voice of the user (1210).

[0098] At least one processor (207) may identify an event of receiving the user's voice through the microphone. However, the present invention is not limited thereto. For example, based on identifying an event of receiving a touch input, the at least one processor (207) may output audio representing context information (440) stored in the memory (206) through the speaker (210). The at least one processor (207) may output audio representing context information (440) through the speaker (210). However, the present invention is not limited thereto. For example, the at least one processor (207) may display text representing context information (440) on the display (208).

[0099] At least one processor (207) may acquire contextual information (440) about the acquired image in the form of metadata while executing a software application for the camera. The form of the metadata for the image is described and exemplified in more detail with reference to FIG. 13.

[0100] Figure 13 illustrates the operation of an electronic device for obtaining context information about a photograph.

[0101] Referring to FIG. 13, at least one processor (207) can acquire an image through a camera (209). A state (1310) can be described as a state of preparing to acquire an image using a software application for the camera. For example, a state (1320) can be described as a state of acquiring an image based on receiving a user input. For example, at least one processor (207) can display a first screen (410) including a viewfinder area (1330) (or preview screen) (or preview image) on the display (208). For example, at least one processor (207) can display a first screen (410) including an executable object (1340) for acquiring an image including a visual object through the camera (209) on the display (208). For example, the viewfinder area (1330) in the first screen (410) may include an image to be acquired through the camera (209). For example, at least one processor (207) may acquire context information (440) based on receiving a user input for acquiring an image displayed in the viewfinder area (1330). For example, at least one processor (207) may provide data acquired by analyzing the image displayed in the viewfinder area (1330) to a pre-trained model (430) based on receiving the user input. For example, at least one processor (207) may acquire context information (440) describing the image (e.g., information indicating that athletes are performing a victory ceremony) by analyzing the image. For example, at least one processor (207) may acquire context information (440) describing a visual object included in the image by analyzing the image.For example, at least one processor (207) may store the acquired context information (440) together with the image acquired through the camera (209) in the memory (206). For example, at least one processor (207) may store the acquired context information (440) as metadata for the acquired image in the memory (206).

[0102] At least one processor (207) may determine the amount of context information (440) to be provided based on the type of event. For example, at least one processor (207) may provide a first portion of the context information (440) based on an event. For example, at least one processor (207) may provide a second portion of the context information (440) based on another event. Determining the amount of context information (440) is described and exemplified in more detail with reference to FIG. 14A.

[0103] Figure 14a illustrates an example of context information provided based on an event.

[0104] Referring to FIG. 14A, at least one processor (207) may identify a type of an event. For example, the at least one processor (207) may determine an amount of context information (440) to be provided from among the context information (440) stored in the memory (206) based on the type of the identified event. For example, the at least one processor (207) may identify a type of event for the context information (440). For example, the at least one processor (207) may determine a portion of the context information (440) to be provided to the user (1210) based on the type of the identified event. For example, the at least one processor (207) may provide a portion of the context information (440) based on the type of the identified event. For example, the at least one processor (207) may provide a first context information (1440-1) to an Nth context information (1440-N) based on the identification of a first event (e.g., an event in which a user requests a reservation at a restaurant). For example, at least one processor (207) may provide Nth context information (1440-N) based on identification of a second event (e.g., an event querying search history for one week). For example, at least one processor (207) may obtain second context information (1440-2) after obtaining first context information (1440-1). For example, at least one processor (207) may obtain third context information (1440-3) after obtaining second context information (1440-2). For example, at least one processor (207) may obtain Nth context information (1440-N) after obtaining third context information (1440-3). For example, at least one processor (207) may obtain first context information (1440-1) to Nth context information (1440-N) based on the passage of time.

[0105] Figure 14b illustrates an example of context information managed over time.

[0106] Referring to FIG. 14B, at least one processor (207) may summarize or delete context information (440) over time. For example, the at least one processor (207) may be required to summarize context information (440) because the capacity of the memory (206) is limited. For example, the at least one processor (207) may reduce the amount of context information (440) by summarizing context information (440) that has passed a specified period of time (e.g., one year) since it was acquired. For example, the at least one processor (207) may summarize context information (440) by gradually deleting information about details included in the context information (440). For example, the context information (440) may be classified into short-term context information (1465) or long-term context information (1460). For example, short-term context information (1465) may be described as context information that has been generated for a specified period of time (e.g., one year). For example, long-term context information (1460) may be described as context information that has not been generated for the specified period of time. For example, at least one processor (207) may manage the memory (206) by converting short-term context information (1465) into long-term context information (1460) over time. However, the present invention is not limited thereto. For example, the electronic device (100) may store short-term context information (1465) in the memory (206) of the electronic device (100). For example, the electronic device (100) may transmit long-term context information (1460) to an external electronic device (e.g., a server or a cloud server). For example, the electronic device (100) may generate content based on receiving long-term context information (1460) from the external electronic device after transmitting long-term context information (1460) to the external electronic device. For example, the long-term context information (1460) may be included in the external electronic device.

[0107] For example, the electronic device (100) may include a context information selection unit (1470) and a content generation unit (1480). For example, the context information selection unit (1470) may determine information necessary for providing content based on an input of a user (1210). For example, the context information selection unit (1470) may determine information necessary for providing content as short-term context information (1465) based on an input of a user (1210). For example, the context information selection unit (1470) may determine information necessary for providing content as long-term context information (1460) based on an input of a user (1210). For example, the context information selection unit (1470) may provide context information (e.g., long-term context information (1460) and short-term context information (1465)) used for providing content to a user (1210) to a software application related to the content. For example, when the electronic device (100) receives a user input for taking a photo, the context information selection unit (1470) can transmit short-term context information (1465) related to the photo to a software application for the camera. For example, the content generation unit (1480) can use the short-term context information (1465) related to the photo to generate metadata for the photo. For example, the metadata can include information about the location where the photo was taken and information about a person in the photo.

[0108] FIG. 15 is a block diagram of an electronic device (1501) within a network environment (1500) according to various embodiments.

[0109] Referring to FIG. 15, in a network environment (1500), an electronic device (1501) may communicate with an electronic device (1502) via a first network (1598) (e.g., a short-range wireless communication network), or may communicate with at least one of an electronic device (1504) or a server (1508) via a second network (1599) (e.g., a long-range wireless communication network). In one embodiment, the electronic device (1501) may communicate with the electronic device (1504) via the server (1508). According to one embodiment, the electronic device (1501) may include a processor (1520), a memory (1530), an input module (1550), an audio output module (1555), a display module (1560), an audio module (1570), a sensor module (1576), an interface (1577), a connection terminal (1578), a haptic module (1579), a camera module (1580), a power management module (1588), a battery (1589), a communication module (1590), a subscriber identification module (1596), or an antenna module (1597). In some embodiments, the electronic device (1501) may omit at least one of these components (e.g., the connection terminal (1578)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1576), camera module (1580), or antenna module (1597)) may be integrated into a single component (e.g., display module (1560)).

[0110] The processor (1520) may, for example, execute software (e.g., a program (1540)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1501) connected to the processor (1520) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1520) may store commands or data received from other components (e.g., a sensor module (1576) or a communication module (1590)) in a volatile memory (1532), process the commands or data stored in the volatile memory (1532), and store result data in a non-volatile memory (1534). According to one embodiment, the processor (1520) may include a main processor (1521) (e.g., a central processing unit or an application processor) or an auxiliary processor (1523) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1521). For example, when the electronic device (1501) includes the main processor (1521) and the auxiliary processor (1523), the auxiliary processor (1523) may be configured to use less power than the main processor (1521) or to be specialized for a given function. The auxiliary processor (1523) may be implemented separately from the main processor (1521) or as a part thereof.

[0111] The auxiliary processor (1523) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1560), the sensor module (1576), or the communication module (1590)) of the electronic device (1501), for example, on behalf of the main processor (1521) while the main processor (1521) is in an inactive (e.g., sleep) state, or together with the main processor (1521) while the main processor (1521) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1523) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1580) or a communication module (1590)). In one embodiment, the auxiliary processor (1523) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1501) where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1508)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.

[0112] The memory (1530) can store various data used by at least one component (e.g., the processor (1520) or the sensor module (1576)) of the electronic device (1501). The data can include, for example, software (e.g., the program (1540)) and input data or output data for commands related thereto. The memory (1530) can include volatile memory (1532) or non-volatile memory (1534).

[0113] The program (1540) may be stored as software in memory (1530) and may include, for example, an operating system (1542), middleware (1544), or an application (1546).

[0114] The input module (1550) can receive commands or data to be used in a component of the electronic device (1501) (e.g., a processor (1520)) from an external source (e.g., a user) of the electronic device (1501). The input module (1550) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).

[0115] The audio output module (1555) can output audio signals to the outside of the electronic device (1501). The audio output module (1555) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.

[0116] The display module (1560) can visually provide information to an external party (e.g., a user) of the electronic device (1501). The display module (1560) may include, for example, a display, a holographic device, or a projector, and a control circuit for controlling the device. In one embodiment, the display module (1560) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.

[0117] The audio module (1570) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (1570) can acquire sound through the input module (1550), output sound through the sound output module (1555), or an external electronic device (e.g., electronic device (1502)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1501).

[0118] The sensor module (1576) can detect the operating status (e.g., power or temperature) of the electronic device (1501) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1576) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.

[0119] The interface (1577) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1501) with an external electronic device (e.g., the electronic device (1502)). In one embodiment, the interface (1577) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.

[0120] The connection terminal (1578) may include a connector through which the electronic device (1501) may be physically connected to an external electronic device (e.g., the electronic device (1502)). In one embodiment, the connection terminal (1578) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).

[0121] The haptic module (1579) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1579) may include, for example, a motor, a piezoelectric element, or an electrical stimulation device.

[0122] The camera module (1580) can capture still images and videos. In one embodiment, the camera module (1580) may include one or more lenses, image sensors, image signal processors, or flashes.

[0123] The power management module (1588) can manage the power supplied to the electronic device (1501). According to one embodiment, the power management module (1588) can be implemented as at least a part of, for example, a power management integrated circuit (PMIC).

[0124] A battery (1589) may power at least one component of the electronic device (1501). In one embodiment, the battery (1589) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.

[0125] The communication module (1590) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1501) and an external electronic device (e.g., electronic device (1502), electronic device (1504), or server (1508)), and the performance of communication through the established communication channel. The communication module (1590) may operate independently from the processor (1520) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1590) may include a wireless communication module (1592) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1594) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1504) via a first network (1598) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1599) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1592) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1596) to identify or authenticate the electronic device (1501) within a communication network such as the first network (1598) or the second network (1599).

[0126] The wireless communication module (1592) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1592) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1592) can support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1592) can support various requirements specified in the electronic device (1501), an external electronic device (e.g., the electronic device (1504)), or a network system (e.g., the second network (1599)). According to one embodiment, the wireless communication module (1592) may support a peak data rate (e.g., 20 Gbps or more) for eMBB implementation, a loss coverage (e.g., 164 dB or less) for mMTC implementation, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC implementation.

[0127] The antenna module (1597) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1597) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1597) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1598) or the second network (1599), may be selected from the plurality of antennas by, for example, the communication module (1590). A signal or power may be transmitted or received between the communication module (1590) and an external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1597).

[0128] According to various embodiments, the antenna module (1597) may form a mmWave antenna module. In one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.

[0129] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).

[0130] According to one embodiment, commands or data may be transmitted or received between the electronic device (1501) and an external electronic device (1504) via a server (1508) connected to a second network (1599). Each of the external electronic devices (1502 or 1504) may be the same or a different type of device as the electronic device (1501). According to one embodiment, all or part of the operations executed in the electronic device (1501) may be executed in one or more of the external electronic devices (1502, 1504, or 1508). For example, when the electronic device (1501) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1501) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1501). The electronic device (1501) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1501) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1504) may include an Internet of Things (IoT) device. The server (1508) may be an intelligent server utilizing machine learning and / or a neural network.According to one embodiment, an external electronic device (1504) or server (1508) may be included within the second network (1599). The electronic device (1501) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.

[0131] Some of the operations described above may be executed (or performed) by an AI (artificial intelligence) system as described with reference to FIG. 16.

[0132] Figure 16 is a schematic diagram of an exemplary AI system.

[0133] Referring to FIG. 16, the AI ​​system (1600) may include an input / output interface (1610), an AI framework (1620), a generative AI model (1630), and / or a knowledge repository (1690).

[0134] The input / output interface (1610) can receive input. The input can include user input and / or data acquired or generated by an electronic device (e.g., the electronic device (100) or the electronic device (1501) described above). The data can include images, videos, and / or sensor data generated by at least one processor (e.g., at least one processor (207) or processor (1520)) of the electronic device (e.g., illuminance data around the electronic device acquired from a sensor or sensor hub (e.g., a coprocessor (1523), posture data (or orientation data) of the electronic device, temperature inside the electronic device (e.g., temperature of the display (208) or temperature of the at least one processor (207)), size information of a display area of ​​the display (208), and / or images acquired through an image sensor (e.g., included in a camera module (1580)) of the electronic device). The user input may include natural language, touch data obtained via touch circuitry included within the display panel (e.g., used to identify input from a finger and / or a stylus), images displayed (and / or to be displayed) on the display panel, and / or video. As a non-limiting example, the user input may be received by the input / output interface (1610) together with context information. The context information may be described as additional information obtained in relation to the user input. The context information may relate to a state when the user input is received (e.g., including a state of the electronic device and / or a state surrounding the electronic device (e.g., a user state)). For example, the context information may include information about one or more software applications running within the electronic device when the user input is received.For example, the contextual information may include information about the location of the electronic device (or the location of the user of the electronic device) at the time the user input is received. For example, the user input may be integrated with the contextual information. For example, the user input integrated with the contextual information may be received by the input / output interface (1610).

[0135] The input / output interface (1610) can transmit (or provide) output. The output may include a result (or result information) generated or obtained by the AI ​​system (1600) based at least in part on the input. The format of the output may vary. For example, the output may include natural language. For example, the output may include content (e.g., including media content and / or multimedia content). For example, the output may include an action related to a user of the electronic device. For example, the output may have a format according to a user setting of the electronic device.

[0136] The input / output interface (1610) can be described as a user query / response interface (1610).

[0137] The AI ​​framework (1620) can be used to obtain information (or data) about the input from the input / output interface (1610) and control one or more components related to the AI ​​system (1600) using the obtained information.

[0138] For example, the prompt design component (1621) within the AI ​​framework (1620) can use the acquired information to generate or obtain prompts for a generative AI model (1630) (e.g., including a large language model (LLM) or a large multimodal model (LMM)). For example, the prompt design component (1621) can be described as an AI component that utilizes a learning algorithm and / or a neural network to provide enhanced prompts over time. For example, the prompt design component (1621) can use the acquired information to access a knowledge component (e.g., a knowledge repository (1690)) that includes user preference data, a prompt library, and / or prompt examples to generate or obtain prompts. The generated prompts can be provided to the generative AI model (1630) (e.g., including an LLM or LMM).

[0139] For example, the API / plugin management component (1622) within the AI ​​framework (1620) may be utilized to facilitate communication for additional information requested (or induced) in connection with the prompt provided (or to be provided) to the generative AI model (1630). For example, the API / plugin management component (1622) may be utilized to create or establish channels for communication with various data sources (e.g., knowledge repositories (1690)). For example, the API / plugin management component (1622) may facilitate access to at least some of the data sources. For example, the API / plugin management component (1622) may be utilized to request another component (e.g., an application / service component (1680)) to perform feedback (or response) in response to the prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1622) may be provided to the prompt design component (1621) for generating a prompt. As a non-limiting example, information obtained (or generated) through the API / plugin management component (1622) may be provided to a generative AI model (1630).

[0140] For example, the improvement component (1623) within the AI ​​framework (1620) can at least partially tune (or adjust) (or change) the result (e.g., content) obtained (or output) from the generative AI model (1630). For example, the improvement component (1623) can determine or verify whether the content obtained from the generative AI model (1630) is relevant to the input. For example, the improvement component (1623) can determine or verify whether the content obtained from the generative AI model (1630) contains biased content. For example, the improvement component (1623) can determine or verify whether the content obtained from the generative AI model (1630) contains harmful content. For example, the improvement component (1623) can support or assist in performing additional processing to improve the content obtained from the generative AI model (1630). For example, the improvement component (1623) may support providing hints to the user to improve the content.

[0141] A generative AI model (1630) can be described as an artificial intelligence neural network that generates feedback in response to a prompt. For example, the feedback may include additional data and / or information related to the prompt, but relative to the prompt. For example, the feedback may include new content related to the prompt. For example, the generative AI model (1630) may include a model that generates images and / or a model that generates language. For example, the model that generates images may include a generative adversarial network (GAN) and / or a variational autoencoder (VAE). For example, the model that generates images may include a diffusion-based generative model (e.g., a transformer VAE). For example, the model that generates language may include CHAT-GPT 3 and / or CHAT-GPT 4. For example, a generative AI model (1630) may include an LMM that generates the feedback by recognizing text, images, and / or speech.

[0142] As a non-limiting example, the AI ​​framework (1620) and / or the generative AI model (1630) may be included within an AI module (e.g., including a processing circuit) within the electronic device. For example, the AI ​​module may be operatively coupled with at least one processor (e.g., at least one processor (207) or processor (1520)) of the electronic device. For example, the AI ​​module may be operatively coupled with a display driver circuit (e.g., a display driver circuit or a display driver integrated circuit (DDI)) of the electronic device. For example, the AI ​​module may be operatively coupled with a sensor hub of the electronic device for one or more sensors within the electronic device.

[0143] An electronic device (e.g., electronic device (100)) as described above may include a memory (e.g., memory (206)) that stores instructions. The electronic device may include a display (e.g., display (208)). The electronic device may include at least one processor (e.g., at least one processor (207)). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a first screen (e.g., first screen (410)) including an executable object (e.g., executable object (930)) on the display. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive a user input (e.g., user input (1050)) related to the executable object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a second screen (e.g., second screen (420)) corresponding to the executable object on the display based on receiving the user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain context information (e.g., context information (440)) indicating that a screen displayed on the display has been switched from the first screen to the second screen by the user input received through the first screen, based on receiving the user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store the context information in the memory.

[0144] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify text included in the first screen and the second screen by performing optical character recognition (OCR) on the first screen and the second screen based on receiving the user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the context information by providing the identified text to a pre-trained model (e.g., pre-trained model (430)).

[0145] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to repeatedly store screenshots of a screen displayed on the display according to a specified period in a ring buffer (e.g., ring buffer (550)) for storing a specified number of images based on a First-In First-Out (FIFO) basis. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain a first screenshot of the first screen, which was last stored in the ring buffer, based on receiving the user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the context information using a second screenshot of the second screen displayed on the display and the obtained first screenshot.

[0146] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain context information indicating that the user input is an action for retrieving content based on receiving the user input for scrolling on the first screen.

[0147] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain the context information including text describing the video content based on receiving the user input for controlling the video content.

[0148] According to one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain context information indicating that a portion of a screen displayed on the display has changed, using a first screenshot of the first area and a second screenshot of the second area, based on a determination that a visual object included in a first area within the first screen is not included in a second area corresponding to the first area within the second screen.

[0149] According to one embodiment, the context information may include context information for a first software application. The user input may include a first user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive a second user input for executing a second software application after receiving the first user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a third screen for the second software application, the third screen including another executable object on the display and different from the first software application, based on receiving the second user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive a third user input related to the other executable object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a fourth screen corresponding to the other executable object on the display based on receiving the third user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain another context information (e.g., third context information) indicating that the screen displayed on the display has been switched from the third screen to the fourth screen by the third user input received through the third screen, based on receiving the third user input.The above instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store the other context information in the memory.

[0150] According to one embodiment, the display may include a first display area on which the first screen or the second screen is displayed, and a second display area that is distinct and separate from the first display area. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a third screen including another executable object on the second display area. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to receive another user input related to the another executable object. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to display a fourth screen corresponding to the another executable object on the second display area, based on receiving the another user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to obtain other context information indicating that a screen displayed on the second display area has switched from the third screen to the fourth screen based on the other user input received through the third screen, based on receiving the other user input. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to store the other context information in the memory.

[0151] According to one embodiment, the electronic device may further include a speaker (e.g., speaker (210)). The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify an event requiring provision of the context information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to output audio representing the context information stored in the memory through the speaker, based on identifying the event.

[0152] In one embodiment, the instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to identify a type of event requiring provision of the context information. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to determine, based on the identified type, a portion of the context information to provide to a user of the electronic device. The instructions, when individually or collectively executed by the at least one processor, may cause the electronic device to provide, based on the identified type, the portion of the context information.

[0153] According to one embodiment, the electronic device may further include a camera (e.g., camera (209)). The executable object may include an executable object for obtaining an image including a visual object through the camera. The first screen may include an image to be obtained through the camera. The context information may describe the visual object included in the first screen.

[0154] As described above, a method performed by an electronic device (e.g., electronic device (100)) having a display (e.g., display (208)) may include an operation of displaying a first screen (e.g., first screen (410)) including an executable object (e.g., executable object (930)) on the display. The method may include an operation of receiving a user input (e.g., user input (1050)) related to the executable object. The method may include an operation of displaying a second screen (e.g., second screen (420)) corresponding to the executable object on the display based on receiving the user input. The method may include an operation of obtaining context information (e.g., context information (440)) indicating that a screen displayed on the display is switched from the first screen to the second screen by the user input received through the first screen based on receiving the user input. The method may include an operation of storing the context information in the memory.

[0155] In one embodiment, the method may include an operation of identifying text included in the first screen and the second screen by performing optical character recognition (OCR) on the first screen and the second screen based on receiving the user input. The method may include an operation of obtaining the context information by providing the identified text to a pre-trained model (e.g., pre-trained model (430)).

[0156] In one embodiment, the method may include an operation of repeatedly storing screenshots of a screen displayed on the display according to a specified period in a ring buffer (e.g., ring buffer (550)) for storing a specified number of images based on a First-In First-Out (FIFO) order. The method may include an operation of obtaining a first screenshot of the first screen, which was last stored in the ring buffer, based on receiving the user input. The method may include an operation of obtaining the context information using a second screenshot of the second screen displayed on the display and the obtained first screenshot.

[0157] According to one embodiment, the method may include an operation of obtaining context information indicating that the user input is an operation for searching content, based on receiving the user input for scrolling on the first screen.

[0158] According to one embodiment, the method may include an operation of obtaining context information including text describing the video content based on receiving the user input for controlling the video content.

[0159] According to one embodiment, the method may include obtaining context information indicating that a portion of a screen displayed on the display is changed using a first screenshot of the first area and a second screenshot of the second area, based on a determination that a visual object included in a first area within the first screen is not included in a second area corresponding to the first area within the second screen.

[0160] According to one embodiment, the context information may include context information for a first software application. The user input may include a first user input. The method may include, after receiving the first user input, receiving a second user input for executing a second software application. The method may include, based on receiving the second user input, displaying a third screen for the second software application, which includes another executable object on the display and is different from the first software application. The method may include, based on receiving the second user input, receiving a third user input related to the other executable object. The method may include, based on receiving the third user input, displaying a fourth screen corresponding to the other executable object on the display. The method may include, based on receiving the third user input, obtaining another context information indicating that a screen displayed on the display is switched from the third screen to the fourth screen by the third user input received through the third screen. The method may include an operation of storing the other context information in the memory.

[0161] According to one embodiment, the display may include a first display area on which the first screen or the second screen is displayed, and a second display area that is distinct and separated from the first display area. The method may include an operation of displaying a third screen including another executable object on the second display area. The method may include an operation of receiving another user input related to the other executable object. The method may include an operation of displaying a fourth screen corresponding to the other executable object on the second display area based on receiving the other user input. The method may include an operation of obtaining other context information indicating that a screen displayed on the second display area is switched from the third screen to the fourth screen by the other user input received through the third screen based on receiving the other user input. The method may include an operation of storing the other context information in the memory.

[0162] According to one embodiment, the electronic device may further include a speaker (e.g., speaker (210)). The method may include an operation of identifying an event requiring provision of the context information. The method may include an operation of outputting audio representing the context information stored in the memory through the speaker based on the identification of the event.

[0163] According to one embodiment, the method may include an operation of identifying a type of event requiring provision of the context information. The method may include an operation of determining a portion of the context information to be provided to a user of the electronic device based on the identified type. The method may include an operation of providing the portion of the context information based on the identified type.

[0164] According to one embodiment, the electronic device may further include a camera (e.g., camera (209)). The executable object may include an executable object for obtaining an image including a visual object through the camera. The first screen may include an image to be obtained through the camera. The context information may describe the visual object included in the first screen.

[0165] In a computer-readable storage medium having one or more programs stored thereon, as described above, the one or more programs may include instructions that, when executed by an electronic device (e.g., the electronic device (100)) having a display (e.g., the display (208)), cause the electronic device to display a first screen (e.g., the first screen (410)) including an executable object (e.g., the executable object (930)) on the display. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive a user input (e.g., the user input (1050)) related to the executable object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display a second screen (e.g., the second screen (420)) corresponding to the executable object on the display based on receiving the user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain context information (e.g., context information (440)) indicating that a screen displayed on the display has switched from the first screen to the second screen based on the user input received through the first screen, based on receiving the user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store the context information in the memory.

[0166] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify text included in the first screen and the second screen by performing optical character recognition (OCR) on the first screen and the second screen based on receiving the user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain the context information by providing the identified text to a pre-trained model (e.g., pre-trained model (430)).

[0167] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to repeatedly store screenshots of a screen displayed on the display according to a specified period in a ring buffer (e.g., ring buffer (550)) for storing a specified number of images based on a First-In First-Out (FIFO) order. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a first screenshot of the first screen, which was last stored in the ring buffer, based on receiving the user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain a second screenshot of the second screen displayed on the display and the context information using the obtained first screenshot.

[0168] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain context information indicating that the user input is an action for retrieving content based on receiving the user input for scrolling on the first screen.

[0169] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain context information including text describing the video content based on receiving the user input for controlling the video content.

[0170] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain context information indicating that a portion of a screen displayed on the display has changed, using a first screenshot of the first area and a second screenshot of the second area, based on a determination that a visual object included in a first area within the first screen is not included in a second area corresponding to the first area within the second screen.

[0171] According to one embodiment, the context information may include context information for a first software application. The user input may include a first user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive a second user input for executing a second software application after receiving the first user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display a third screen for the second software application, which includes another executable object on the display and is different from the first software application, based on receiving the second user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive a third user input related to the other executable object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display a fourth screen corresponding to the other executable object on the display based on receiving the third user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain another context information indicating that a screen displayed on the display has been switched from the third screen to the fourth screen based on the third user input received through the third screen.The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store the other context information in the memory.

[0172] According to one embodiment, the display may include a first display area on which the first screen or the second screen is displayed, and a second display area that is distinct and separate from the first display area. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display a third screen including another executable object on the second display area. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to receive another user input related to the another executable object. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to display a fourth screen corresponding to the another executable object on the second display area based on receiving the another user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to obtain other context information indicating that a screen displayed on the second display area has switched from the third screen to the fourth screen based on the other user input received through the third screen, based on receiving the other user input. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to store the other context information in the memory.

[0173] According to one embodiment, the electronic device may further include a speaker (e.g., speaker 210). The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify an event requiring provision of the context information. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to output audio representing the context information stored in the memory through the speaker based on the identification of the event.

[0174] According to one embodiment, the one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to identify a type of event requiring provision of the context information. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to determine, based on the identified type, a portion of the context information to provide to a user of the electronic device. The one or more programs may include instructions that, when executed by the electronic device, cause the electronic device to provide, based on the identified type, the portion of the context information.

[0175] According to one embodiment, the electronic device may further include a camera (e.g., camera (209)). The executable object may include an executable object for obtaining an image including a visual object through the camera. The first screen may include an image to be obtained through the camera. The context information may describe the visual object included in the first screen.

[0176] The devices described above may be implemented as hardware components, software components, and / or a combination of hardware components and software components. For example, the devices and components described in the embodiments may be implemented using one or more general-purpose computers or special-purpose computers, such as a processor, a controller, an arithmetic logic unit (ALU), a digital signal processor, a microcomputer, a field programmable gate array (FPGA), a programmable logic unit (PLU), a microprocessor, or any other device capable of executing instructions and responding to them. The processing device may execute an operating system (OS) and one or more software applications running on the operating system. The processing device may also access, store, manipulate, process, and generate data in response to the execution of the software. For ease of understanding, the processing device is sometimes described as being used alone; however, one of ordinary skill in the art will recognize that the processing device may include multiple processing elements and / or multiple types of processing elements. For example, a processing unit may include multiple processors, or a processor and a controller. Other processing configurations, such as parallel processors, are also possible.

[0177] Software may include a computer program, code, instructions, or a combination of one or more of these, which may configure a processing device to perform a desired operation or may independently or collectively command the processing device. The software and / or data may be embodied in any type of machine, component, physical device, computer storage medium, or device for interpretation by the processing device or for providing instructions or data to the processing device. The software may also be distributed over networked computer systems and stored or executed in a distributed manner. The software and data may be stored on one or more computer-readable recording media.

[0178] The method according to the embodiment may be implemented in the form of program commands that can be executed through various computer means and recorded on a computer-readable medium. In this case, the medium may be one that continuously stores a computer-executable program or one that temporarily stores it for execution or download. In addition, the medium may be various recording or storage means in the form of a single or multiple hardware combinations, and is not limited to a medium directly connected to a computer system, but may also be distributed over a network. Examples of the medium may include magnetic media such as hard disks, floppy disks, and magnetic tapes, optical recording media such as CD-ROMs and DVDs, magneto-optical media such as floptical disks, and those configured to store program commands, including ROM, RAM, and flash memory. In addition, examples of other media may include recording or storage media managed by app stores that distribute applications, sites that supply or distribute various software, servers, etc.

[0179] Although the embodiments described above have been described by way of limited examples and drawings, those skilled in the art will appreciate that various modifications and variations can be made based on the above teachings. For example, appropriate results can still be achieved even if the described techniques are performed in a different order than described, and / or components of the described systems, structures, devices, circuits, etc. are combined or combined in a different manner than described, or are replaced or substituted with other components or equivalents.

[0180] Therefore, other implementations, other embodiments, and equivalents of the claims are also within the scope of the claims described below. According to one embodiment, the method according to the various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a commodity between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be at least temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.

[0181] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.

Claims

1. In electronic devices, A memory comprising one or more storage media and storing instructions; display; and At least one processor comprising processing circuitry, The above instructions, when individually or collectively executed by the at least one processor, Displaying a first screen including an executable object on the display, Receive user input related to the above executable object, Based on receiving the above user input: Displaying a second screen corresponding to the executable object on the display, and By the user input received through the first screen, context information indicating that the screen displayed on the display is switched from the first screen to the second screen is obtained, and To store the context information in the above memory, causing the above electronic device, Electronic devices.

2. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Based on receiving the user input, performing optical character recognition (OCR) on the first screen and the second screen to identify text included in the first screen and the second screen, and By providing the identified text to a pre-trained model, the context information is obtained. causing the above electronic device, Electronic devices.

3. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Repeatedly saving screenshots of the screen displayed on the display according to a specified cycle within a ring buffer for storing a specified number of images based on FIFO (First-In First-Out), Based on receiving the user input, obtaining a first screenshot of the first screen, which was last stored in the ring buffer, and To obtain the context information using the second screenshot of the second screen displayed on the display and the first screenshot obtained, causing the above electronic device, Electronic devices.

4. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Based on receiving the user input for scrolling on the first screen, obtain the context information indicating that the user input is an action for searching content. causing the above electronic device, Electronic devices.

5. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Based on receiving the user input for controlling the video content, obtaining the context information including text describing the video content, causing the above electronic device, Electronic devices.

6. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Based on a determination that a visual object included in a first area within the first screen is not included in a second area corresponding to the first area within the second screen, using a first screenshot for the first area and a second screenshot for the second area, obtain context information indicating that a portion of a screen displayed on the display is changed. causing the above electronic device, Electronic devices.

7. In claim 1, the context information is: Contains contextual information about the first software application, The above user input is, The first user input is, The above instructions, when individually or collectively executed by the at least one processor, After receiving the first user input, receiving a second user input for executing a second software application, Based on receiving the second user input, displaying a third screen for the second software application, the third screen including another executable object on the display and different from the first software application; Receiving a third user input related to the above other executable object, Based on receiving the third user input above: Displaying a fourth screen corresponding to the above other executable object on the display, and By the third user input received through the third screen, another context information is obtained indicating that the screen displayed on the display is switched from the third screen to the fourth screen, and To store the other context information in the above memory, causing the above electronic device, Electronic devices.

8. In claim 1, the display, It includes a first display area in which the first screen or the second screen is displayed and a second display area that is distinct and separated from the first display area, The above instructions, when individually or collectively executed by the at least one processor, Display a third screen including another executable object on the second display area, Receive other user input related to the other executable object, Based on receiving the above other user input: Displaying a fourth screen corresponding to the other executable object on the second display area, and By the other user input received through the third screen, the screen displayed on the second display area acquires other context information indicating that it has switched from the third screen to the fourth screen, and To store the other context information in the above memory, causing the above electronic device, Electronic devices.

9. In claim 1, the electronic device, Including more speakers, The above instructions, when individually or collectively executed by the at least one processor, Identifying events that require provision of the above context information, and Based on identifying the above event, output audio representing the context information stored in the memory through the speaker. causing the above electronic device, Electronic devices.

10. In claim 1, The above instructions, when individually or collectively executed by the at least one processor, Identify the type of event for which provision of the above context information is required, Based on the types identified above: determining some of the context information to be provided to the user of the electronic device, and To provide some of the above context information, causing the above electronic device, Electronic devices.

11. In claim 1, the electronic device, Including more cameras, The above executable object is, An executable object for acquiring an image containing a visual object through the camera, The first screen above is, Including an image to be acquired through the above camera, and The above context information is, Describing the visual object included in the first screen, Electronic devices.

12. In a non-transitory computer-readable storage medium storing one or more programs, the one or more programs are: When executed by an electronic device having a display, Displaying a first screen including an executable object on the display, Receive user input related to the above executable object, Based on receiving the above user input: Displaying a second screen corresponding to the executable object on the display, and By the user input received through the first screen, context information indicating that the screen displayed on the display is switched from the first screen to the second screen is obtained, and To store the above context information in memory, comprising instructions causing the electronic device to operate; Non-transitory computer-readable storage medium.

13. In claim 12, The above one or more programs, when executed by the electronic device, Based on receiving the user input, performing optical character recognition (OCR) on the first screen and the second screen to identify text included in the first screen and the second screen, and By providing the identified text to a pre-trained model, the context information is obtained. comprising instructions causing the electronic device to operate; Non-transitory computer-readable storage medium.

14. In claim 12, The above one or more programs, when executed by the electronic device, Repeatedly saving screenshots of the screen displayed on the display according to a specified cycle within a ring buffer for storing a specified number of images based on FIFO (First-In First-Out), Based on receiving the user input, obtaining a first screenshot of the first screen, which was last stored in the ring buffer, and To obtain the context information using the second screenshot of the second screen displayed on the display and the first screenshot obtained, comprising instructions causing the electronic device to operate; Non-transitory computer-readable storage medium.

15. In claim 12, The above one or more programs, when executed by the electronic device, Based on receiving the user input for scrolling on the first screen, obtain the context information indicating that the user input is an action for searching content. comprising instructions causing the electronic device to operate; Non-transitory computer-readable storage medium.

Citation Information

Patent Citations

  • Generating action elements suggesting content for ongoing tasks

    EP4116843A1

  • Display apparatus and error detection method thereof

    KR1020030066889A

  • Controlling a mobile terminal with at least two display area

    KR1020100030387A

  • Method and device for providing display data

    KR1020150014139A

  • An intelligent recommendation system that integrates multi-attribute ratings and emotional ratings of online review data

    KR1020240058386A