Method for performing search on basis of captured screen and electronic device supporting same
Patent Information
- Application Number
- PCT/KR2026/001522
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2025-05-08
- Filing Date
- 2026-01-26
- Publication Date
- 2026-08-27
Smart Images

Figure KR2026001522_27082026_PF_FP_ABST
Abstract
Description
Method for performing a search based on a captured screen and an electronic device supporting the same
[0001] The present disclosure relates to a method for performing a search based on a captured screen and an electronic device supporting the same.
[0002] Various services and additional functions provided through electronic devices, such as portable electronic devices like smartphones, are gradually increasing. To enhance the utility value of these devices and satisfy the needs of diverse users, telecommunications service providers or electronic device manufacturers are competitively developing devices to offer various functions and differentiate themselves from competitors. Consequently, the various functions provided through electronic devices are also becoming increasingly sophisticated.
[0003] An electronic device may provide a search function using a captured image of at least a portion of a screen being displayed through a display. For example, while a screen is being displayed through the display, the electronic device may generate an image of a selected area (hereinafter also referred to as a "captured image") by capturing an area selected based on user input within the screen. The electronic device may detect text and / or objects within the captured image by analyzing the captured image. The electronic device may perform a search for the detected text and / or objects (and / or the captured image itself) using an artificial intelligence model and provide the results of the search.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] The aforementioned search function using captured images may find it difficult to provide the desired search results to the user because it performs the search by considering only the captured images. For example, when an area containing the text "Building A" is selected within a screen displayed via a display based on user input while a movie titled "Building A" is playing on a video application, an electronic device may provide information about the actual Building A (e.g., location of the actual Building A, photos of the actual Building A, history of the actual Building A) as a search result that is unrelated to the movie "Building A". On the other hand, the user may intend to receive information related to the movie "Building A" currently playing on the video application. Accordingly, when providing a search function using captured images, it may be necessary to perform the search by considering the application running on the electronic device (e.g., a video application) and the content provided through said application (e.g., the movie "Building A").
[0006] Various embodiments of the present disclosure relate to a method for performing a search based on a captured screen and an electronic device supporting the same, wherein the search is performed by considering an application running on an electronic device and content provided through said application when providing a search function using a captured image.
[0007] The technical problems that the present disclosure aims to solve are not limited to those mentioned above, and other unmentioned technical problems will be clearly understood by those skilled in the art from the description below.
[0008] An electronic device according to one embodiment may include a display, at least one processor including a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause an area including at least a portion of the screen to be captured based on user input while displaying the screen through the display. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to determine whether an application corresponding to the captured area is included in at least one designated application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause at least one action corresponding to information obtained from the captured area and set for the application based on whether the application is included in the at least one designated application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause at least one prompt to perform the at least one action on content being provided using the application. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to acquire data corresponding to the at least one prompt using an artificial intelligence model.
[0009] A method according to one embodiment may include an operation of capturing an area including at least a portion of the screen based on user input while displaying a screen through a display of an electronic device. The method may include an operation of determining whether an application corresponding to the captured area is included in at least one designated application. The method may include an operation of obtaining at least one action corresponding to information obtained from the captured area and set for the application based on whether the application is included in the at least one designated application. The method may include an operation of obtaining at least one prompt for performing the at least one action on content being provided using the application. The method may include an operation of obtaining data corresponding to the at least one prompt using an artificial intelligence model.
[0010] According to one embodiment, in a non-transient computer-readable storage medium storing computer-executable instructions, the computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to capture an area including at least a portion of the screen based on user input while displaying the screen through the display of the electronic device. The computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to determine whether an application corresponding to the captured area is included in at least one designated application. The computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to obtain at least one action set for the application and corresponding to information obtained from the captured area based on whether the application is included in the at least one designated application. The computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to obtain at least one prompt for performing the at least one action on content being provided using the application. When the above computer-executable instructions are executed individually or collectively by at least one processor, the electronic device may be caused to acquire data corresponding to the at least one prompt using an artificial intelligence model.
[0011] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0012] FIG. 2 is a diagram illustrating a generative artificial intelligence system according to one embodiment.
[0013] FIG. 3 is a block diagram of an electronic device according to one embodiment.
[0014] FIG. 4 is a flowchart illustrating a method for performing a search based on a captured screen according to one embodiment.
[0015] FIG. 5 is a diagram illustrating a method for capturing a selected area based on user input according to one embodiment.
[0016] FIG. 6 is a flowchart illustrating a method for obtaining at least one action based on the detection of at least one text or object from a captured area according to one embodiment.
[0017] FIG. 7 is a flowchart illustrating a method for obtaining at least one action based on the fact that the captured area is an area where content is provided, according to one embodiment.
[0018] FIG. 8 is a drawing for explaining a method of performing a search based on a captured screen according to one embodiment.
[0019] FIG. 9 is a drawing for explaining a method of performing a search based on a captured screen according to one embodiment.
[0020] FIG. 10a is a drawing for explaining a method of performing a search based on a captured screen according to one embodiment.
[0021] FIG. 10b is a drawing for explaining a method of performing a search based on a captured screen according to one embodiment.
[0022] FIG. 11 is a diagram illustrating a method for determining an application corresponding to a captured area based on the execution of a plurality of applications according to one embodiment.
[0023] FIG. 12 is a diagram illustrating a method for determining an application corresponding to a captured area based on the execution of a plurality of applications according to one embodiment.
[0024] FIG. 13 is a diagram illustrating a method for determining an application corresponding to a captured area based on the execution of a plurality of applications according to one embodiment.
[0025] FIG. 14 is a flowchart illustrating a method for performing a search based on capturing a video according to one embodiment.
[0026] FIG. 15 is a drawing for explaining a method of performing a search based on capturing a video according to one embodiment.
[0027] FIG. 1 is a block diagram of an electronic device (101) in a network environment (100) according to one embodiment.
[0028] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0029] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0030] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0031] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0032] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0033] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0034] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0035] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0036] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0037] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0038] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0039] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0040] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0041] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0042] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0043] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0044] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0045] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0046] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0047] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0048] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0049] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0050] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0051] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0052] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0053] Various embodiments of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0054] According to one embodiment, the method according to the various embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)) or an application store (e.g., Play Store). TM It can be distributed online (e.g., downloaded or uploaded) through ) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0055] According to various embodiments, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to various embodiments, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the multiple components in the same or similar manner as those performed by the corresponding component among the multiple components prior to integration. According to various embodiments, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0056] FIG. 2 is a diagram illustrating a generative artificial intelligence system according to one embodiment.
[0057] Referring to FIG. 2, the user query / response interface (210) can receive user input. The user input may be in the form of natural language, images, and / or videos, but there are no limitations. Additionally, context information may be transmitted along with the user input. The context information may include various additional information at the time of user input. For example, the additional information may include information about the application currently being used by the user or the user's location information. Furthermore, the user input may be in a mixed form of the aforementioned natural language, images, sounds, and context information. Additionally, the user input may be in a non-natural language form, such as selecting a menu. The user query / response interface (210) can output results from a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of actions requested by the user. The user query / response interface (210) can output results from a generative artificial intelligence system to the user. The output can be in the form of natural language or specific content, and it may also be provided in the form of actions requested by the user.
[0058] The AI framework (240) can receive input from the user and coordinate and control each component necessary to perform the user's intent based on the user's query.
[0059] User input received from the user query / response interface (210) can be transmitted to a prompt design component (241). The prompt design component (241) can be used to generate prompts suitable for inputting user input into a large language model (LLM) or a large multimodal model (LMM). The prompt design component (241) may be an AI component that uses machine learning algorithms or neural networks to develop better prompts over time. The prompt design component (241) can generate prompts by accessing a knowledge component containing user preference data, a prompt library, and prompt examples based on user input, and can transmit the generated prompts to the LLM or LMM.
[0060] The API / Plug-in management component (242) can perform the role of communicating with external information when there is a request for additional information when user input is passed as input to a generative model. The API / Plug-in management component (242) establishes a channel to communicate with the outside of the AI Interface via the API, and through the established channel, it can enable access to various data sources (e.g., knowledge repository (220)). Additionally, the API / Plug-in management component (242) can request the application / service component (230) via the API to perform an action that ultimately executes the user input, rather than an intermediate result, in the case where the application or service needs to perform such action. The information obtained from the outside may be used to generate a prompt in the prompt design component (241) along with the user input, or it may be passed as input to the generative model.
[0061] The output modification component (or refiner component) (243) can fine-tune the output of the generative model. For example, the output modification component (243) can verify whether the content generated through the LLM and / or LMM is irrelevant, contains biased content, or contains harmful content. Additionally, the output modification component (243) can determine the extent to which the output matches the desired result and, if necessary, proceed with the additional process. Furthermore, the output modification component (243) can configure and provide hints to the user to avoid unwanted output.
[0062] A generative AI model (260) generally refers to an artificial intelligence neural network that generates new forms of data based on user input information. A generative AI model (260) may include a model that generates images and / or a model that generates language. Models that generate images include, but are not limited to, GANs (generative adversarial networks) and VAEs (variational auto encoders), and examples include Diffusion-based generative models that use VAEs and Transformer structures. Models that generate language are models trained to output the most statistically appropriate output value based on input values, and examples include models such as CHAT-GPT 3 and CHAT-GPT 4. There are also LMMs (large multimodal models) that can recognize various forms of data input, such as text, images, and voice, and generate new data corresponding to them.
[0063] FIG. 3 is a block diagram of an electronic device (301) according to one embodiment.
[0064] Referring to FIG. 3, in one embodiment, the electronic device (301) may include a communication circuit (310), a display (320), a memory (330), and a processor (340).
[0065] In one embodiment, the electronic device (301) may be included in the electronic device (101) of FIG. 1.
[0066] In one embodiment, the communication circuit (310) may be included in the communication module (190) of FIG. 1.
[0067] In one embodiment, the display (320) may be included in the display module (160) of FIG. 1.
[0068] In one embodiment, the memory (330) may be included in the memory (130) of FIG. 1.
[0069] In one embodiment, the memory (330) may store information necessary to perform an operation of performing a search based on a captured screen.
[0070] In one embodiment, the memory (330) may store instructions so that the electronic device (301) performs an operation (e.g., operations performed by the electronic device (301) described later with reference to FIGS. 4 through 15) when executed individually or collectively by at least one processor (e.g., processor (340)).
[0071] In one embodiment, the memory (330) may store an artificial intelligence model. For example, the memory (330) may store an artificial intelligence model for performing at least part of the operation of performing a search based on a captured screen.
[0072] In one embodiment, the artificial intelligence model stored in memory (330) may include the generative AI model (260) shown in FIG. 2.
[0073] In one embodiment, the memory (330) can store the capture module.
[0074] In one embodiment, a capture module (hereinafter referred to as the "capture module") stored in memory (330) can capture an area including at least a portion of a screen displayed through the display (320). For example, the capture module can capture the entire area of the screen displayed through the display (320), a region of interest automatically selected within the entire area of the screen (e.g., an area including an object detected within the screen), or an area determined (e.g., selected) based on user input within the entire area of the screen (e.g., user input drawing a closed curve within the entire area of the screen using a user's finger or an electronic pen) (e.g., an area in the form of a bounding box surrounding the closed curve). In one embodiment, the operation of capturing an area including at least a portion of the screen displayed through the display (320) may include the operation of generating an image (hereinafter also referred to as the "capture image") of the area including at least a portion of the screen (hereinafter also referred to as the "captured area").
[0075] In one embodiment, the memory (330) may store various applications. For example, the memory (330) may store a video application, a music application, an image application (e.g., a gallery application), and / or an internet application (e.g., a web browser). However, the applications stored by the memory (330) are not limited to the aforementioned applications.
[0076] In one embodiment, the processor (340) may be included in the processor (120) of FIG. 1.
[0077] In one embodiment, the processor (340) can control the overall operation of performing a search based on a captured screen. The processor (340) may include at least one processor for performing the operation of performing a search based on a captured screen. The at least one processor may execute the instructions stored in memory (330) individually or collectively.
[0078] In FIG. 3, the electronic device (301) is illustrated as including a communication circuit (310), a display (320), a memory (330), and a processor (340), but is not limited thereto. For example, the electronic device (301) may further include at least one of the components of the electronic device (101) of FIG. 1 (e.g., a microphone included in the input module (150), a speaker included in the sound output module (155), a camera module (180), or a sensor module (176). For example, the electronic device (301) may not include a communication circuit (310).
[0079] FIG. 4 is a flowchart (400) for explaining a method of performing a search based on a captured screen according to one embodiment.
[0080] Referring to FIG. 4, in operation 401, in one embodiment, the processor (340) can capture an area including at least a portion of the screen based on user input while displaying the screen through the display (320).
[0081] In one embodiment, the processor (340) may display an execution screen of the application through the display (320) based on the execution of the application. Hereinafter, an application running (e.g., currently running) on the electronic device (301) will be referred to as the "first application." Additionally, a screen displayed (e.g., currently being displayed) through the display (320), including the execution screen of the first application, will also be referred to as the "first screen."
[0082] In one embodiment, the processor (340) may capture the entire area of the first screen while the first screen is displayed through the display (320) based on user input. For example, the processor (340) may capture the entire area of the first screen using a capture application (or capture module) based on user input to a hardware key of the electronic device (301) (e.g., user input pressing the power key together with the key to decrease the volume level, among the keys to increase the volume level and the keys to decrease the volume level), input input by moving the edge of the user's hand on the screen, or user input to a software key displayed through the display (320). However, the method of capturing the entire area of the first screen is not limited to the examples described above.
[0083] In one embodiment, the processor (340) can capture a region of interest that is automatically determined (e.g., selected) within the entire area of the first screen while the first screen is displayed through the display (320) based on user input. For example, the processor (340) can capture the entire area of the first screen based on user input as described above. The processor (340) can select a region of interest (e.g., an area containing an object) within the entire area of the captured first screen using an algorithm or artificial intelligence model for selecting a region of interest. The processor (340) can capture the selected region of interest.
[0084] In one embodiment, the processor (340) can capture an area determined (e.g., selected) based on user input within the entire area of the first screen. Hereinafter, with reference to FIG. 5, a method for capturing an area selected based on user input within the entire area of the first screen will be described.
[0085] FIG. 5 is a diagram illustrating a method for capturing a selected area based on user input according to one embodiment.
[0086] Referring to FIG. 5, in one embodiment, at reference numeral 501 of FIG. 5, the processor (340) can display the execution screen of the music application as a first screen (510) through the display (320) based on the music application being executed as a first application.
[0087] In one embodiment, the first screen (510) may include an album cover (521) of a song provided (e.g., currently provided) using a music application and objects (522, 523, 524) corresponding to the functions of the music application.
[0088] In one embodiment, the processor (340) may receive user input drawing a closed curve (520) while the first screen (510) is displayed. For example, the processor (340) may receive user input drawing a closed curve (520) using a user's finger or an electronic pen while the first screen (510) is displayed. In one embodiment, the processor (340) may capture an area determined by the closed curve (520) based on receiving user input drawing the closed curve (520) (e.g., an area corresponding to the inside of the closed curve (520) with the closed curve (520) as a boundary, hereinafter also referred to as the "inside area of the closed curve").
[0089] However, the method of capturing a selected area based on user input is not limited to the examples described above. For example, in reference numeral 502 of FIG. 5, the processor (340) may capture an area in the shape of a bounding box (530) containing the closed curve (520) (e.g., a rectangle formed by two straight lines adjacent to the closed curve (520) and parallel to the horizontal line of the screen (510) and two straight lines adjacent to the closed curve (520) and parallel to the vertical line of the screen (510)) instead of the operation of capturing an area inside the closed curve (520) based on receiving user input to draw the closed curve (520).
[0090] In the examples described above, the entire area of the first screen or a selected area within the first screen is illustrated as being captured based on user input for hardware keys (e.g., a key for increasing the volume level and a key for decreasing the volume level, and a power key), user input for software keys (e.g., icons or buttons displayed through the display (320)), or user input by moving the edge of the user's hand on the screen, but is not limited thereto. In one embodiment, the processor (340) may capture the entire area of the first screen or a selected area within the first screen based on user voice input (e.g., user voice input through a microphone). For example, the processor (340) may capture the entire area of the first screen based on user voice input through a microphone, such as "Capture and analyze the current screen." For example, the processor (340) can capture a selected area (e.g., the area of the window on the left among the multi-windows or the area of the image currently displayed in the pop-up window) within the entire area of the first screen based on user voice input through the microphone, such as “Capture and search the window on the left among the multi-windows” or “Search the image currently displayed in the pop-up window.”
[0091] Referring again to FIG. 4, for convenience of explanation, the area including at least a portion of the first screen, which is captured by one of the methods for capturing an area including at least a portion of the first screen described above, will be referred to as the "captured area."
[0092] In one embodiment, the operation of capturing an area including at least a portion of the first screen may include the operation of generating an image of the captured area. For example, the processor (340) may generate an image of the captured area (hereinafter also referred to as "captured image") by capturing an area including at least a portion of the first screen (e.g., by imaging the captured area using a capture module). However, it is not limited thereto. For example, the operation of capturing an area including at least a portion of the first screen may be an operation of checking information about the captured area (e.g., layout information about the captured area) without performing the operation of generating an image of the captured area.
[0093] In operation 403, in one embodiment, the processor (340) can determine whether the application corresponding to the captured area is included in at least one designated application.
[0094] In one embodiment, at least one designated application (hereinafter referred to as "at least one designated application") may include an application capable of playing at least one of video or audio. For example, at least one application may include an application capable of playing video or audio over time, such as a video application, a music application, a voice recording application, or a game application. However, at least one designated application is not limited to the examples described above.
[0095] In one embodiment, information about at least one designated application (e.g., the name, identifier, and / or type of at least one designated application) may be stored in memory (330).
[0096] In one embodiment, the processor (340) may determine the application as an application corresponding to the captured area based on the fact that the captured area includes at least a portion of the execution screen of the application. For example, the processor (340) may determine the first application as an application corresponding to the captured area based on the fact that at least a portion of the execution screen of the first application running (e.g., running) on the electronic device (301) is included in the captured area.
[0097] In one embodiment, the processor (340) may display a plurality of execution screens corresponding to each of the plurality of applications through the display (320) based on the execution of a plurality of applications in the electronic device (301). The processor (340) may determine an application among the plurality of applications that corresponds to a captured area based on the fact that the captured area includes a part of each of the plurality of execution screens. These operations will be described in more detail later with reference to FIGS. 11, 12, and 13.
[0098] In one embodiment, the processor (340) may perform a search for the captured area by analyzing only the captured area (e.g., the captured image) based on the fact that the application corresponding to the area captured in operation 403 is not included in at least one designated application. For example, the processor (340) may provide a search result by performing a search for text and / or objects detected from the captured image for the captured area, or the captured image itself, based on the fact that the application corresponding to the captured area (e.g., the first application) is not included in at least one designated application.
[0099] In operation 405, in one embodiment, the processor (340) may obtain at least one action (or command) corresponding to (e.g., mapped) information obtained from the captured area, which is set for the application (e.g., first application) corresponding to the captured area, based on confirming in operation 403 that the application is included in the at least one designated application. For example, the processor (340) may obtain at least one action mapped to information obtained from the captured area (e.g., at least one of text or object obtained from the captured area) associated with the application corresponding to the captured area, based on the application being included in the at least one designated application.
[0100] Hereinafter, with reference to FIGS. 6 and FIGS. 7, the operation of verifying at least one action of operation 405 will be described in more detail.
[0101] FIG. 6 is a flowchart (600) for explaining a method of obtaining at least one action based on the detection of at least one text or object from a captured area according to one embodiment.
[0102] Referring to FIG. 6, in one embodiment, operations 601 and 603 of FIG. 6 may be included in operation 405 of FIG. 4. However, this is not limited thereto, and part of operation 601 may be performed after operation 401 is performed and before operation 405 is performed.
[0103] In operation 601, in one embodiment, the processor (340) can determine whether at least one of text or an object is detected from the captured area based on the fact that the application corresponding to the captured area (e.g., the first application) is included in at least one designated application.
[0104] In one embodiment, the processor (340) can detect at least one of text or an object from a captured area by checking information for configuring a first screen (hereinafter referred to as "configuration information of the first screen") that is being displayed through a display (320), based on the fact that an application corresponding to the captured area (e.g., a first application) is included in at least one designated application. For example, the processor (340) can check the location of the captured area within the first screen based on the fact that an application corresponding to the captured area is included in at least one designated application. The processor (340) can check at least one of text or an object in at least a part of the captured area by checking the configuration information displayed at the confirmed location based on the configuration information of the first screen. For example, in reference numerals 501 and 502 of FIG. 5, the processor (340) can identify the inner area of a closed curve (520) or a bounding box (530) as a captured area. The processor (340) can identify the location of the captured area within the first screen (510). The processor (340) can identify a progress bar (524) as an object from the captured area by identifying configuration information corresponding to the location of the captured area based on configuration information of the first screen (510) (e.g., layout information of the first screen (510)). However, the processor (340) is not limited thereto, and by performing the aforementioned operations, the processor (340) can identify objects (522, 523) included in the captured area in addition to the progress bar (524).
[0105] However, it is not limited thereto. In one embodiment, the processor (340) can detect at least one of text or an object within the captured area (or captured image) by analyzing the captured image using an image analysis module for the captured area. For example, the processor (340) can detect at least one of text or an object within the captured area (or captured image) by using image analysis of the captured image for the captured area. For example, in reference numerals 501 and 502 of FIG. 5, the processor (340) can identify the inner area of the closed curve (520) or the bounding box (530) as the captured area. The processor (340) can detect objects (522, 523, 524) within the captured area (or captured image) by using image analysis of the captured image generated for the captured area.
[0106] In operation 603, in one embodiment, the processor (340) may obtain at least one action corresponding to at least one of the text or object that is set for an application corresponding to the captured area based on the detection of at least one of the text or object from the captured area.
[0107] In one embodiment, the processor (340) may set at least one action corresponding to at least one of text or object for at least one designated application (e.g., for each of at least one designated application). For example, the processor (340) may set at least one action that is at least partially different (or identical) depending on the application included in at least one designated application, even for at least one of the same text or the same object. For example, the processor (340) may set at least one action that is at least partially different (or identical) depending on at least one of the text or object for the same application.
[0108] In one embodiment, if at least one designated application includes a video application and a music application, the processor (340) may set at least one action corresponding to the same object (or text) differently for the video application and the music application.
[0109] In one embodiment, at least one action may include at least one command (or also referred to as "at least one function") to be performed using an artificial intelligence model.
[0110] In one embodiment, the processor (340) can set one or more actions corresponding to at least one of text or object for an application (e.g., an application included in at least one designated application).
[0111] For example, the processor (340) provides, with respect to a music application, one or more actions corresponding to a progress bar (e.g., progress bar (524)), information regarding a time interval corresponding to an area captured within the entire playback interval of a song provided by the music application (hereinafter referred to as the “entire playback interval”) (e.g., a time interval corresponding to positions (524-1, 524-2) on the progress bar (524) where the progress bar (524) and the closed curve (520) intersect in reference numeral 501, or a time interval corresponding to positions (524-3, 524-4) on the progress bar (524) where the progress bar (524) and the bounding box (530) intersect in reference numeral 502) (hereinafter referred to as the “time interval”), information regarding the current playback time within the entire playback time interval, and the lyrics of the song provided by the music application corresponding to the time interval. Using the music application, at least one of the following can be set: generating an audio portion corresponding to the time interval within the audio of a song being provided (e.g., an audio file), generating a preferred portion within the audio portion, searching for a music video corresponding to the determined time interval, or generating a clip of the music video. In generating a preferred portion within the audio portion, the preferred portion within the audio portion may be a portion that has been recommended by multiple users or played by multiple users within the audio portion. However, the method of setting at least one action corresponding to at least one of text or an object for the music application is not limited to the examples described above.
[0112] For example, the processor (340) may, for a video application, provide information about a time interval (hereinafter referred to as "time interval") corresponding to a captured area within an entire playback interval (hereinafter referred to as "entire playback interval") of a video (e.g., a movie) provided by the video application (e.g., a time interval corresponding to positions (1013-1, 1013-2) on the progress bar (1013) where the closed curve (1012) and the progress bar (1013) intersect in FIG. 10a), as one or more actions corresponding to a progress bar (e.g., the progress bar (1013) of FIG. 10a) (hereinafter referred to as "time interval"), provide information about the current playback time within the entire playback time interval, provide a summary of the content of the video provided corresponding to the time interval, create a video portion corresponding to the time interval within the video provided (e.g., a video file) (e.g., create a short form using the video portion corresponding to the time interval), or create a new one by editing the video portion. At least one of the video generation can be configured. In one embodiment, the generation of a new video by editing the video portion may include the generation of a desktop image based on the video portion or a change in the specified style of the video portion (e.g., converting the video portion to a realistic style or changing the style of the video portion to a cartoon style). However, the method of configuring at least one action corresponding to at least one of text or an object for a video application is not limited to the examples described above.
[0113] In one embodiment, the processor (340) can obtain the at least one action by identifying at least one action corresponding to at least one text or object detected from the captured area among one or more actions set for an application corresponding to the captured area.
[0114] FIG. 7 is a flowchart (700) for explaining a method of obtaining at least one action based on the fact that the captured area is an area where content is provided, according to one embodiment.
[0115] Referring to FIG. 7, in operation 701, in one embodiment, the processor (340) can determine whether the captured area corresponds to an area where content is provided within a screen (e.g., a first screen) based on the fact that the application corresponding to the captured area is included in at least one designated application.
[0116] In one embodiment, the processor (340) can determine whether the captured area corresponds to an area where content provided (e.g., currently provided) using the application within the first screen is provided, based on the fact that the application corresponding to the captured area is included in at least one designated application. For example, the processor (340) can determine whether the captured area includes at least a portion of an area where a movie is played as content provided using the video application within the first screen (e.g., area (1011) of FIG. 10a) as an execution screen of the video application, based on the fact that the video application corresponding to the captured area is included in at least one designated application.
[0117] In operation 703, in one embodiment, the processor (340) may obtain at least one action set for an application corresponding to a captured area based on the fact that the captured area corresponds to an area where content is provided within a screen (e.g., a first screen). For example, the processor (340) may obtain at least one action set for a video application based on the fact that the application corresponding to the captured area is a video application and the captured area corresponds to an area where content is provided within the first screen (e.g., providing summary information of a movie as content provided using the video application).
[0118] In one embodiment, in FIGS. 6 and 7, the processor (340) may obtain (e.g., generate) at least one prompt to be input to an artificial intelligence model by analyzing only the captured image of the captured area based on whether at least one of text or an object is not detected from the captured area or whether the captured area does not correspond to the area where content is provided, or obtain at least one prompt to perform a search on the captured image of the captured area itself.
[0119] Referring again to FIG. 4, in operation 407, the processor (340) can obtain at least one prompt to perform at least one action (e.g., at least one action described with reference to FIG. 6 and FIG. 7) on content provided (e.g., being provided) using an application corresponding to the captured area.
[0120] In one embodiment, the processor (340) may generate at least one prompt to perform at least one action corresponding to at least one of text or object detected from the captured area and set for the music application, based on the fact that the application corresponding to the captured area is a music application.
[0121] In one embodiment, the processor (340) may generate at least one prompt to perform at least one action corresponding to at least one of text or object detected from the captured area and set for the video application, based on the fact that the application corresponding to the captured area is a video application.
[0122] In one embodiment, the processor (340) may generate at least one prompt to perform at least one action set for the video application for the movie, based on the fact that the application corresponding to the captured area is a video application and the captured area is an area where a movie is played as content provided using the video application.
[0123] In operation 409, in one embodiment, the processor (340) can obtain data corresponding to at least one prompt obtained by performing operation 407 using an artificial intelligence model.
[0124] In one embodiment, the processor (340) may input at least one prompt obtained by performing operation 407 with an artificial intelligence model (e.g., a generative AI model (260)). In response to inputting at least one prompt obtained by performing operation 407 with an artificial intelligence model, the processor may obtain data corresponding to at least one prompt.
[0125] In one embodiment, the data corresponding to the at least one prompt may be data output by inputting the at least one prompt into an artificial intelligence model. For example, the data corresponding to the at least one prompt may include results obtained (e.g., generated) by performing a search related to the at least one prompt using an artificial intelligence model, and / or data obtained by performing an action indicated by the at least one prompt (e.g., video data, music data, text data). Examples of the data corresponding to the at least one prompt will be described in more detail later.
[0126] In one embodiment, although not illustrated in FIG. 4, the processor (340) may output the acquired data based on acquiring data corresponding to at least one prompt. For example, the processor (340) may display the acquired data through a display (320) or output audio corresponding to the acquired data through a speaker frame, depending on the acquired data (e.g., type of data).
[0127] In one embodiment, some of the operations described with reference to FIG. 4 may be performed on a server (e.g., server (108)) wirelessly connected to an electronic device (301). For example, operation 409 of FIG. 4 may be performed on a server (e.g., a generative artificial intelligence model stored on a server) wirelessly connected to an electronic device (301). When the server obtains said data by performing an operation to obtain said data corresponding to at least one prompt of operation 409, the processor (340) may receive said data from said server through a communication circuit (310).
[0128] FIG. 8 is a drawing (800) for explaining a method of performing a search based on a captured screen according to one embodiment.
[0129] Referring to FIG. 8, in one embodiment, the screen (810) may display a screen in which data corresponding to at least one prompt obtained by performing operations 403, 405, 407, and 409 of FIG. 4 is output, in the case where a captured area is determined based on a closed curve (520) (or bounding area (530)) drawn by a user as illustrated in reference numeral 501 (or reference numeral 502) of FIG. 5.
[0130] In one embodiment, at reference numeral 501 (or reference numeral 502) of FIG. 5, the processor (340) may determine the inner area of a closed curve (520) drawn by the user as a captured area while the first screen (510) is displayed through the display (320). The processor (340) may detect a progress bar (524) from the captured area based on the fact that a music application corresponding to the captured area is included in at least one designated application. The processor (340) can obtain (e.g., confirm) the provision of lyrics for a time interval (e.g., position (524-1, 524-2) corresponding to positions (524-1, 524-2) on the progress bar (524) during the total playback time of the content being provided using the music application, and the creation of audio data (e.g., audio file) corresponding to the time interval, as at least one action corresponding to the progress bar (524) for the music application. The processor (340) can obtain (e.g., create) a prompt including the provision of song lyrics for a time interval corresponding to positions (524-1, 524-2) on the progress bar (524) and the creation of audio data (e.g., audio file) for the song "ABC" as content being provided using the music application. The processor (340) can use an artificial intelligence model to obtain song lyrics (and song lyrics corresponding to the entire playback time interval of the song "ABC") and audio data corresponding to the time interval as data corresponding to the obtained prompt for the song "ABC" (and song lyrics corresponding to the entire playback time interval of the song "ABC"), and can display the obtained song lyrics (e.g., texts included in areas (821, 822, 823)) through a display (320) and output the obtained audio data through a speaker.For example, the processor (340) can display through the display (320) the lyrics of the time interval corresponding to positions (524-1, 524-2) on the progress bar (524) among the entire lyrics of the song “ABC” (e.g., lyrics included in areas (821, 822, 823)) (e.g., lyrics included in area (822)) so as to distinguish them from the lyrics of other time intervals (e.g., lyrics included in areas (821, 823)) (e.g., by highlighting the lyrics included in area (822).
[0131] In one embodiment, although not illustrated in FIG. 8, after data corresponding to at least one prompt is output using an artificial intelligence model, the processor (340) can perform additional searching using the artificial intelligence model based on user input regarding the output data.
[0132] FIG. 9 is a drawing (900) for explaining a method of performing a search based on a captured screen according to one embodiment.
[0133] Referring to FIG. 9, in one embodiment, as described above, the processor (340) can perform a search for a captured area by analyzing only the captured area (e.g., a captured image) based on the fact that the application corresponding to the area captured in operation 403 is not included in at least one designated application. For example, the processor (340) can provide a search result by performing a search for text and / or objects detected from the captured image of the captured area, or the captured image itself, based on the fact that the application corresponding to the captured area (e.g., a first application) is not included in at least one designated application.
[0134] In one embodiment, the processor (340) may display, through the display (320), an execution screen of a document application that includes an image representing a progress bar while the document application is running. The processor (340) may generate a captured image by capturing an area within the execution screen of the document application that includes an image representing a progress bar based on user input. The processor (340) may detect the progress bar as an object from the generated captured image based on confirming that the document application is not included in at least one designated application. The processor (340) may obtain a search result for the progress bar by performing a search for the detected progress bar. The processor (340) may display the obtained search result through the display (320). For example, as illustrated in FIG. 9, the processor (340) can display a screen (910) including an image (920) representing a progress bar as a search result and an image (930) representing a video player including a progress bar (931) through a display (320).
[0135] FIG. 10a is a drawing for explaining a method of performing a search based on a captured screen according to one embodiment.
[0136] FIG. 10b is a drawing for explaining a method of performing a search based on a captured screen according to one embodiment.
[0137] Referring to FIGS. 10a and 10b, in one embodiment, at reference numeral 1001 of FIG. 10a, the processor (340) may display a screen (1010) as an execution screen of a video application through a display (320) based on the execution of a video application. For example, the screen (1010) may include a screen (1011) in which the movie "ABC" is provided as content provided using the video application, a progress bar (1013), and text (1014) indicating the title of the movie "ABC". While the screen (1010) is displayed through the display (320), the processor (340) may determine the inner area of a closed curve (1012) as a captured area based on user input. The processor (340) can detect a progress bar (1013) and text (1014) from a captured area based on the fact that a movie application corresponding to the captured area is included in at least one designated application. The processor (340) can obtain (e.g., confirm) the provision of a summary of the content of the movie "ABC" at a time interval (e.g., location (1013-1) is the start point of the time interval and location (1013-2) is the end point of the time interval) corresponding to the positions (1013-1, 1013-2) on the progress bar (1013) during the total playback time of the movie "ABC" as content provided by the video application, as at least one action corresponding to the progress bar (1013) for the video application. The processor (340) can obtain at least one action corresponding to the text (1014) for the video application. The processor (340) can obtain (e.g., generate) prompts that perform at least one action corresponding to a progress bar (1013) and at least one action corresponding to text (1014) for the movie "ABC" as content provided using a video application.The processor (340) can generate data corresponding to the acquired prompts using an artificial intelligence model and output the generated data. For example, the processor (340) can display a screen (1020) through the display (320) containing a summary of the content of the movie "ABC" in a time interval corresponding to positions (1013-1, 1013-2) on the progress bar (524) for the movie "ABC", as illustrated in reference numeral 1002 of FIG. 10a. In one embodiment, although not illustrated in reference numeral 1002 of FIG. 10, the processor (340) can display through the display (320) data corresponding to a prompt for performing at least one action corresponding to the text (1014) for the movie "ABC", which is acquired using an artificial intelligence model.
[0138] In one embodiment, the processor (340) may output information representing the results before outputting the results, based on the acquisition of results (e.g., data corresponding to each of the multiple prompts) corresponding to each of the multiple prompts (e.g., multiple prompts corresponding to each of the multiple actions). The processor (340) may output a result selected based on user input within the output information.
[0139] In one embodiment, referring to FIGS. 10a and FIGS. 10b, the processor (340) may display a screen (1010) through a display (320) as shown in FIG. 10a. While the screen (1010) is being displayed, the processor (340) may determine the inner area of a closed curve (1012) as a captured area based on user input. The processor (340) may detect a progress bar (1013) and text (1014) from the captured area based on the fact that a movie application corresponding to the captured area is included in at least one designated application. The processor (340) may obtain (e.g., generate) a plurality of prompts that cause a plurality of actions corresponding to the progress bar (1013) and text (1014) to be performed for the movie "ABC" as content provided using a video application. The processor (340) may generate a plurality of results corresponding to the plurality of obtained prompts.
[0140] In one embodiment, the processor (340) may output information representing the plurality of results to allow a user to select a desired result among the plurality of results based on generating the plurality of results. For example, the processor (340) can use a window (1031) (e.g., a window displayed in the form of a pop-up) to display information (1041) that guides the user to select a desired result among the plurality of results, information (1051) that indicates a result corresponding to an action of "summarize in text" (e.g., an action that commands to provide a summary of the movie ABC as content as text), information (1052) that indicates a result corresponding to an action of "summarize in video" (e.g., an action that commands to provide a summary of the movie ABC as content as video), information (1053) that indicates a result corresponding to an action of "related shopping" (e.g., an action that commands to provide shopping information related to the movie ABC as content), information (1054) that indicates a result corresponding to an action of "video retouching" (e.g., an action that commands to provide a result of performing a retouching operation on the movie ABC as content), and information (1055) that recommends a result corresponding to an action of "summarize in text" through a display (320).
[0141] In one embodiment, the processor (340) can display a screen (1020) through the display (320) that provides a summary of the content of the movie "ABC" as text, as illustrated in reference numeral 1002 of FIG. 10a, based on selecting information (1051) that represents a result corresponding to the action of "summary in text" based on user input.
[0142] In the examples described above, the output of a result selected based on user input is illustrated as being obtained based on results corresponding to each of the multiple prompts, but is not limited thereto. For example, the processor (340) may output information representing the multiple actions obtained after multiple actions (or multiple prompts corresponding to the multiple actions) have been obtained. The processor (340) may search for and output a result corresponding to an action corresponding to the information selected based on user input among the information representing the multiple actions.
[0143] FIG. 11 is a diagram illustrating a method for determining an application corresponding to a captured area based on the execution of a plurality of applications according to one embodiment.
[0144] Referring to FIG. 11, in one embodiment, the processor (340) can display a plurality of execution screens corresponding to each of the plurality of applications through the display (320) based on the execution of a plurality of applications.
[0145] In one embodiment, at reference numeral 1101 of FIG. 11, the processor (340) may display a screen (1110) including a video application execution screen (1112) displayed in a pop-up form on a calendar application execution screen (1111) through a display (320). For example, while the calendar application execution screen (1111) is displayed through the display (320), the processor (340) may display the video application execution screen (1112) through the display (320) using a pop-up window.
[0146] In one embodiment, a video application is included in at least one designated application described with reference to operation 403 of FIG. 4, and a calendar application may not be included in the at least one designated application.
[0147] In one embodiment, as illustrated in reference numeral 1101, the processor (340) may determine, based on user input, an area (1113) including at least a portion of the execution screen (1112) of a video application and at least a portion of the execution screen (1111) of a calendar application as a captured area.
[0148] In one embodiment, the processor (340) can identify a layer corresponding to the layer arranged (or placed) at the top among the layers of the plurality of execution screens (also referred to as the "uppermost layer"), based on the fact that the captured area includes at least a portion of each of the plurality of execution screens corresponding to the plurality of applications. The processor (340) can determine the application corresponding to the identified uppermost layer as the application corresponding to the captured area. For example, in reference numeral 1102 of FIG. 11, layer (1121) represents a layer of the execution screen (1111) of a calendar application, layer (1122) represents a layer of the execution screen (1112) of a video application, and area (1123) may be the area where the area (1113) is captured.
[0149] In one embodiment, the processor (340) can determine that the captured area (1123) includes a part of the execution screen of a calendar application layer (1121) and a layer (1122) of the execution screen of a video application. Based on the fact that the captured area (1123) includes a part of the layer (1121) and the layer (1122), the processor (340) can determine the video application corresponding to the layer (1122) arranged at the top of the layer (1121) and the layer (1122) as the application corresponding to the captured area (1123).
[0150] However, the method of determining the application corresponding to the captured area based on the execution of multiple applications is not limited to the examples described above.
[0151] In one embodiment, the processor (340) can determine that an application running in the foreground among a plurality of applications is an application corresponding to the captured area, based on the fact that the captured area includes at least a portion of each of a plurality of execution screens (e.g., a plurality of execution screens corresponding to each of a plurality of applications).
[0152] In one embodiment, the processor (340) may determine the application having focus among the plurality of applications as the application corresponding to the captured area, based on the fact that the captured area includes at least a portion of each of the plurality of execution screens (e.g., a plurality of execution screens corresponding to each of the plurality of applications). For example, the processor (340) may determine the application corresponding to the execution screen where the last input (e.g., user input) was acquired among the plurality of execution screens as the application corresponding to the captured area, based on the fact that the captured area includes at least a portion of each of the plurality of execution screens (e.g., a plurality of execution screens corresponding to each of the plurality of applications).
[0153] FIG. 12 is a drawing (1200) for explaining a method for determining an application corresponding to a captured area based on the execution of a plurality of applications according to one embodiment.
[0154] Referring to FIG. 12, in one embodiment, the processor (340) can display a plurality of execution screens corresponding to each of the plurality of applications through the display (320) based on the execution of a plurality of applications.
[0155] In one embodiment, in FIG. 12, the processor (340) can display a screen (1210) including a calendar application execution screen (1220) and a video application execution screen (1230) through the display (320). For example, the processor (340) can simultaneously display the calendar application execution screen (1220) and the video application execution screen (1230) on the display (320) using a multi-window method.
[0156] In one embodiment, a video application is included in at least one designated application described with reference to operation 403 of FIG. 4, and a calendar application may not be included in the at least one designated application.
[0157] In one embodiment, the processor (340) may determine, based on user input, a captured area that includes both the execution screen (1220) of a calendar application and the execution screen (1230) of a video application.
[0158] In one embodiment, the processor (340) may determine the application that has focus among the plurality of applications as the application corresponding to the captured area, based on the fact that the captured area includes at least a portion of each of the plurality of execution screens (e.g., a plurality of execution screens corresponding to each of the plurality of applications). For example, the processor (340) may determine the application corresponding to the execution screen where the user input was last acquired among the execution screen (1220) and the execution screen (1230) as the application corresponding to the captured area, based on the fact that the execution screen (1220) of the calendar application and the execution screen (1230) of the video application are included.
[0159] FIG. 13 is a drawing (1300) for explaining a method for determining an application corresponding to a captured area based on the execution of a plurality of applications according to one embodiment.
[0160] Referring to FIG. 13, in one embodiment, the processor (340) may execute a second application linked to or included in the first application based on the execution of the first application. For example, the processor (340) may execute a video application using an address linked to the internet application (e.g., a URL (uniform resource locator)) based on the execution of the internet application.
[0161] In one embodiment, a video application is included in at least one designated application described with reference to operation 403 of FIG. 4, and an internet application may not be included in the at least one designated application.
[0162] In one embodiment, the processor (340) may determine, among the first application and the second application, the application corresponding to the execution screen including the execution screen of the first application and the execution screen of the second application that is captured as the application corresponding to the captured area.
[0163] In one embodiment, as illustrated in FIG. 13, the processor (340) may display, through the display (320), an execution screen (1330) of a video application provided using a URL linked to the internet application within an execution screen (1310) of an internet application. The processor (340) may determine, based on user input, that an area (1331) including at least a part of the execution screen (1330) is captured, and that the video application among the internet application and the video application is the application corresponding to the captured area. The processor (340) may determine, based on user input, that an area (1321) including at least a part of the screen portion (e.g., text (1320)) excluding the execution screen (1330) of the video application within the execution screen (1310) of the internet application is captured, and that the internet application among the internet application and the video application is the application corresponding to the captured area.
[0164] Although not previously described, in one embodiment, the processor (340) may determine an application corresponding to the captured area based on a specified criterion among the first application and the second application when the captured area includes at least a portion of the execution screen of the first application and at least a portion of the execution screen of the second application. For example, when the captured area includes at least a portion of the execution screen of the first application and at least a portion of the execution screen of the second application, the processor (340) may identify an execution screen among the execution screen of the first application and the execution screen of the second application that includes the center point (e.g., center of gravity) of the captured area. The processor (340) may determine an application corresponding to the captured area among the first application and the second application that corresponds to the identified execution screen. However, it is not limited thereto. For example, if the captured area includes at least a portion of the execution screen of the first application and at least a portion of the execution screen of the second application, the processor (340) may determine that a first area corresponding to the execution screen of the first application within the captured area corresponds to the first application, and a second area corresponding to the execution screen of the second application within the captured area corresponds to the first application.
[0165] FIG. 14 is a flowchart (1400) for explaining a method of performing a search based on capturing a video according to one embodiment.
[0166] FIG. 15 is a drawing for explaining a method of performing a search based on capturing a video according to one embodiment.
[0167] Referring to FIGS. 14 and 15, in operation 1401, in one embodiment, the processor (340) can create a new video by capturing a portion of the video based on user input while the video is being provided using a video application.
[0168] In one embodiment, as described above, in FIG. 10, the processor (340) can capture a portion of the video by selecting an area including a progress bar (1013) while the video is provided using the video through the display (320). For example, in FIG. 10, the processor (340) can identify a time interval corresponding to positions (1013-1, 1013-2) on the progress bar (1013) where the closed curve (1012) and the progress bar (1013) intersect, based on user input drawing a closed curve (1012). The processor (340) can obtain a file (e.g., a plurality of frames and audio data corresponding to the identified time interval) from among the video files.
[0169] In one embodiment, the processor (340) can generate a new video including the file corresponding to the identified time interval.
[0170] However, the method of creating a new video by capturing a portion of the video based on user input while the video is being provided using a video application is not limited to the examples described above. In one embodiment, the processor (340) may create a new video based on user input regarding an area where the video is provided using a video application. For example, in reference numeral 1501 of FIG. 15, the processor (340) may receive a touch input (e.g., a touch-down input) regarding an area (1510) where the video is provided by the user's finger (1521) while the video is being played. Based on the touch input, the processor (340) may start an operation to capture the video from a first point in time of the video corresponding to the time when the touch input was received (e.g., a point in time of the video corresponding to the time when the touch input was received within the total playback time of the video). In reference numeral 1502 of FIG. 15, the processor (340) can create a new video by capturing a portion of the video from a first point in time of the video up to a second point in time of the video corresponding to the point in time when the touch input is released after the touch input is maintained by the user's finger (1521) (e.g., a point in time of the video corresponding to the point in time when the touch input is released during the entire playback time of the video). In the example above, a new video is created by capturing a portion of the video during the time the touch input is maintained for the area (1510) where the video is provided, but it is not limited thereto.For example, the processor (340) can generate a new video by capturing a portion of the video corresponding to a time interval from the first point in time to the second point in time within the entire playback time of the video, based on the fact that a first touch input (e.g., touchdown input) for the video-providing area (1510) is received at the first point in time and a second touch input (e.g., touchdown input) for the video-providing area (1510) is received at the second point in time.
[0171] In one embodiment, the operation of generating a new video may include the operation of acquiring a plurality of image frames corresponding to the time interval (e.g., a time interval from a first point in time to a second point in time within the total playback time of the video) by continuously (e.g., sequentially) capturing the screen while the video is being played, and / or the operation of acquiring audio data of the video corresponding to the time interval.
[0172] In one embodiment, the operation of creating a new video may include the operation of obtaining a portion of the video corresponding to the time interval (e.g., the time interval from a first point in time to a second point in time within the total playback time of the video) from the entire file of the video.
[0173] In operation 1403, in one embodiment, the processor (340) can obtain at least one prompt corresponding to the newly generated video.
[0174] In one embodiment, the processor (340) may obtain at least one prompt corresponding to the generated new video (or text and / or objects detected within the generated new video) by performing at least some or all of operations 405 and 407 of FIG. 4 on the new video. For example, at least one prompt may be set including at least one of providing a summary of the content of the new video, generating a short form using the portion of the new video, or editing the portion of the new video. In one embodiment, editing the new video may include generating a desktop image based on the new video or changing the style of the new video to a specified style (e.g., making the new video realistic or changing the style of the new video to a cartoon style).
[0175] An electronic device according to one embodiment may include a display, at least one processor including a processing circuit, and a memory for storing instructions. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause an area including at least a portion of the screen to be captured based on user input while displaying the screen through the display. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to determine whether an application corresponding to the captured area is included in at least one designated application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause at least one action corresponding to information obtained from the captured area and set for the application based on whether the application is included in the at least one designated application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause at least one prompt to perform the at least one action on content being provided using the application. When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to acquire data corresponding to the at least one prompt using an artificial intelligence model.
[0176] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to determine whether at least one of text or an object is detected from the captured area based on the application being included in the at least one designated application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may cause to obtain at least one action set for the application and corresponding to at least one of the text or the object based on the detection of at least one of the text or the object from the captured area.
[0177] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to determine whether the at least one of the text or the object is detected from the captured area by checking configuration information corresponding to the captured area within the screen among the configuration information of the screen, based on the fact that the application is included in at least one designated application.
[0178] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to determine whether the captured area corresponds to an area where the content is provided within the screen, based on the application being included in the at least one designated application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to obtain at least one action set for the application, based on the fact that the captured area corresponds to an area where the content is provided within the screen.
[0179] In one embodiment, the captured area may include the entire area of the screen, a region of interest automatically selected within the entire area of the screen, or an area selected based on user input within the entire area of the screen.
[0180] In one embodiment, the specified at least one application may include an application capable of playing at least one of video or audio.
[0181] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may further cause at least one action to be set for the same object, at least some of which are different, according to an application included in the at least one designated application.
[0182] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may further cause at least one action to be set differently for at least one application included in the at least one designated application, depending on the object.
[0183] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may further cause the captured area to include at least a portion of each of the multiple execution screens corresponding to each of the multiple applications, and the top layer among the layers of the multiple execution screens. When the instructions are executed individually or collectively by the at least one processor, the electronic device may further cause the application corresponding to the top layer among the multiple applications to determine as the application corresponding to the captured area.
[0184] In one embodiment, when the instructions are executed individually or collectively by the at least one processor, the electronic device may further cause to generate a new video by capturing a portion of the video based on user input while the video is provided using a video application. When the instructions are executed individually or collectively by the at least one processor, the electronic device may further cause to obtain at least one prompt corresponding to the generated new video. It may further cause to obtain data corresponding to the at least one prompt using the artificial intelligence model.
[0185] A method according to one embodiment may include an operation of capturing an area including at least a portion of the screen based on user input while displaying a screen through a display of an electronic device. The method may include an operation of determining whether an application corresponding to the captured area is included in at least one designated application. The method may include an operation of obtaining at least one action corresponding to information obtained from the captured area and set for the application based on whether the application is included in the at least one designated application. The method may include an operation of obtaining at least one prompt for performing the at least one action on content being provided using the application. The method may include an operation of obtaining data corresponding to the at least one prompt using an artificial intelligence model.
[0186] In one embodiment, the operation of acquiring the at least one action may include an operation of determining whether at least one of text or an object is detected from the captured area based on the application being included in the at least one designated application. The operation of acquiring the at least one action may include an operation of acquiring at least one action that is set for the application and corresponds to at least one of the text or the object based on the detection of at least one of the text or the object from the captured area.
[0187] In one embodiment, the operation of determining whether at least one of the text or the object is detected from the captured area may include determining whether at least one of the text or the object is detected from the captured area by checking configuration information corresponding to the captured area within the screen among the configuration information of the screen, based on the application being included in at least one designated application.
[0188] In one embodiment, the operation of acquiring the at least one action may include an operation of determining whether the captured area corresponds to an area where the content is provided within the screen, based on the application being included in the at least one designated application. The operation of acquiring the at least one action may include an operation of acquiring at least one action set for the application, based on the fact that the captured area corresponds to an area where the content is provided within the screen.
[0189] In one embodiment, the captured area may include the entire area of the screen, an area of interest automatically selected within the entire area of the screen, or an area selected based on user input within the entire area of the screen.
[0190] In one embodiment, the specified at least one application may include an application capable of playing at least one of video or audio.
[0191] In one embodiment, the method may further include an operation of setting at least one action that is at least partially different for the same object, depending on an application included in the at least one designated application.
[0192] In one embodiment, the method may further include, depending on the object, an action of setting at least one action that is different in at least a part for the same application included in the at least one designated application.
[0193] In one embodiment, the method may further include an operation of identifying the top layer among the layers of the plurality of execution screens, based on the fact that the captured area includes at least a portion of each of the plurality of execution screens corresponding to the plurality of applications. The method may further include an operation of determining the application corresponding to the top layer among the plurality of applications as the application corresponding to the captured area.
[0194] According to one embodiment, in a non-transient computer-readable storage medium storing computer-executable instructions, the computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to capture an area including at least a portion of the screen based on user input while displaying the screen through the display of the electronic device. The computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to determine whether an application corresponding to the captured area is included in at least one designated application. The computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to obtain at least one action set for the application and corresponding to information obtained from the captured area based on whether the application is included in the at least one designated application. The computer-executable instructions may cause an electronic device, when executed individually or collectively by at least one processor, to obtain at least one prompt for performing the at least one action on content being provided using the application. When the above computer-executable instructions are executed individually or collectively by at least one processor, the electronic device may be caused to acquire data corresponding to the at least one prompt using an artificial intelligence model.
Claims
1. In the electronic device (301), Display (320); At least one processor (340) including processing circuitry; and It includes memory (330) for storing instructions, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: While displaying a screen through the above display, an area including at least a portion of the screen is captured based on user input, and Check whether the application corresponding to the above-mentioned captured area is included in at least one designated application, and Based on the fact that the above application is included in the at least one designated application, at least one action corresponding to information obtained from the captured area and set for the application is obtained, and Obtaining at least one prompt for performing at least one action on the content being provided using the above application, and An electronic device that causes data corresponding to at least one prompt to be acquired using an artificial intelligence model.
2. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on the fact that the above application is included in the at least one designated application, determining whether at least one of text or an object is detected from the captured area, and An electronic device that causes at least one action corresponding to at least one of the text or the object to be obtained for the application based on the detection of at least one of the text or the object from the captured area.
3. In Paragraph 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: An electronic device that causes to determine whether at least one of the text or the object is detected from the captured area by checking configuration information corresponding to the captured area within the screen among the configuration information of the screen, based on the fact that the above application is included in at least one designated application.
4. In any one of paragraphs 1 to 3, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on the fact that the above application is included in the at least one designated application, determining whether the captured area corresponds to an area where the content is provided within the screen, and An electronic device that causes at least one action set for the application to be obtained based on the fact that the above-mentioned captured area corresponds to the area where the content is provided within the above-mentioned screen.
5. In any one of paragraphs 1 to 4, The above-mentioned captured area is an electronic device comprising the entire area of the screen, a region of interest automatically selected within the entire area of the screen, or a region selected based on user input within the entire area of the screen.
6. In any one of paragraphs 1 to 5, The above-mentioned at least one application is an electronic device comprising an application capable of playing at least one of video or audio.
7. In any one of paragraphs 1 through 6, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: An electronic device that, depending on an application included in the above-mentioned at least one designated application, further causes at least one action to be set for the same object in which at least some parts are different.
8. In any one of paragraphs 1 through 7, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: An electronic device that, depending on the object, further causes at least one other action to be set for the same application included in at least one specified application.
9. In any one of paragraphs 1 through 8, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Based on the fact that the above-mentioned captured area includes at least a portion of each of a plurality of execution screens corresponding to a plurality of applications, the top layer among the layers of the plurality of execution screens is identified, and An electronic device that further causes the application corresponding to the top layer among the plurality of applications to be determined as the application corresponding to the captured area.
10. In any one of paragraphs 1 through 9, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: While a video is being provided using a video application, a new video is created by capturing a portion of the video based on user input, and Obtain at least one prompt corresponding to the new video generated above, and An electronic device that further causes to acquire data corresponding to at least one prompt using the above artificial intelligence model.
11. Regarding the method, An operation of capturing an area including at least a portion of the screen based on user input while displaying a screen through a display of an electronic device; An operation to determine whether the application corresponding to the above-mentioned captured area is included in at least one designated application; Based on the fact that the above application is included in the at least one designated application, an operation to obtain at least one action corresponding to information obtained from the captured area and set for the application; The operation of obtaining at least one prompt for performing the at least one action on the content being provided using the above application; and A method comprising the operation of acquiring data corresponding to at least one prompt using an artificial intelligence model.
12. In Paragraph 11, The operation of acquiring at least one of the above actions is: An operation to determine whether at least one of text or an object is detected from the captured area based on the fact that the above application is included in the at least one designated application; and A method comprising the operation of obtaining at least one action corresponding to at least one of the text or the object, which is set for the application and based on the detection of at least one of the text or the object from the captured area.
13. In Paragraph 12, The operation of determining whether at least one of the text or the object is detected from the above-mentioned captured area is: A method comprising checking whether at least one of the text or the object is detected from the captured area by checking configuration information corresponding to the captured area within the screen among the configuration information of the screen, based on the fact that the above application is included in at least one designated application.
14. In any one of paragraphs 11 through 13, The operation of acquiring at least one of the above actions is: An operation to determine whether the captured area corresponds to an area where the content is provided within the screen, based on the fact that the above application is included in the at least one designated application; and A method comprising obtaining at least one action set for the application based on the fact that the captured area corresponds to an area where the content is provided within the screen.
15. In a non-transient computer-readable storage medium storing computer-executable instructions, said computer-executable instructions, when executed individually or collectively by at least one processor, an electronic device: While displaying a screen through the display of the electronic device, an area including at least a portion of the screen is captured based on user input, and Check whether the application corresponding to the above-mentioned captured area is included in at least one designated application, and Based on the fact that the above application is included in the at least one designated application, at least one action corresponding to information obtained from the captured area and set for the application, and Obtaining at least one prompt for performing at least one action on the content being provided using the above application, and A computer-readable storage medium that causes data corresponding to at least one prompt to be obtained using an artificial intelligence model.