Electronic device using artificial intelligence model, operation method thereof, and storage medium
The electronic device uses AI models to process multiple images, generating association information and identifying target objects, addressing the limitations of existing AI technologies in dynamic environments for augmented and virtual reality systems.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- SAMSUNG ELECTRONICS CO LTD
- Filing Date
- 2025-10-28
- Publication Date
- 2026-06-04
AI Technical Summary
Existing AI technologies struggle to efficiently process and analyze visual data to identify and locate target objects across multiple images, particularly in dynamic environments, limiting their application in augmented and virtual reality systems.
An electronic device equipped with an artificial intelligence model, utilizing neural networks and large-scale language models, processes multiple images to generate association information between texts and images, enabling the identification and location of target objects through scene analysis and query verification.
Enhances the ability to accurately and efficiently identify and locate target objects within complex scenes, supporting advanced augmented and virtual reality applications by leveraging AI models for enhanced visual understanding and object recognition.
Smart Images

Figure KR2025017304_04062026_PF_FP_ABST
Abstract
Description
Electronic device using an artificial intelligence model, method of operation thereof, and storage medium
[0001] The present disclosure relates to an electronic device using an artificial intelligence model, a method of operation thereof, and a storage medium.
[0002] With the rapid advancement of AI (artificial intelligence) technology, various AI agent applications and services are being developed, and the related market is growing explosively. Beyond massive language models, AI is expected to expand infinitely in the scope of its application across diverse aspects of life; it will go beyond simply providing search results to creating new content by combining existing data or acting as a substitute for actual experts by being trained with specialized knowledge in specific fields. AI agent applications can provide commands that respond to the user's voice.
[0003] Recent advancements in artificial intelligence technology are enabling various AI applications based on large-scale model training. In particular, Large Language Models (LLMs) and Large Visual Models (LVMs) are demonstrating innovative performance in natural language processing and visual data processing, attracting attention across various industries. LLMs are models used to understand and generate human language, implemented by artificial neural networks containing tens of billions to trillions of parameters. These models learn from massive amounts of text data to acquire the ability to understand context and process complex sentence structures. Consequently, they are utilized in diverse application fields such as translation, summarization, question-answering systems, conversational interfaces, and content creation. Notably, they can provide performance optimized for specific domains through pretraining and fine-tuning techniques. LVMs are models used to process visual data, including images, videos, or 3D data. By training on large-scale image datasets, LVMs can perform various visual tasks such as object recognition, scene understanding, image generation, and video analysis. Furthermore, with the recent emergence of multimodal models that learn by combining text and images, new application technologies such as text-based image search, image description generation, and video captioning are also becoming possible.
[0004] The information described above may be provided as related art for the purpose of aiding understanding of the present disclosure. No claim or determination is made as to whether any of the foregoing may be applied as prior art related to the present disclosure.
[0005] The electronic device may include at least one processor and a memory for storing instructions.
[0006] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause each of the multiple texts according to scene analysis by an artificial intelligence model for multiple images.
[0007] Each of the above multiple texts can be associated with the scene analysis results for the objects included in each of the above multiple images.
[0008] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to generate and store multiple association information between each of the plurality of texts and each of the plurality of images.
[0009] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check a query for the location of a target object.
[0010] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify at least one first association information among the plurality of association information, which includes text corresponding to the target object.
[0011] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide a candidate location of the target object based on the at least one first association information.
[0012] A method of operation of an electronic device may include an operation of verifying each of a plurality of texts based on scene analysis by an artificial intelligence model for a plurality of images. Each of the plurality of texts may be associated with the scene analysis results for objects included in each of the plurality of images.
[0013] The method of operation of the above electronic device may include the operation of generating and storing multiple association information between each of the plurality of texts and each of the plurality of images.
[0014] The method of operation of the above electronic device may include an operation of checking a query for the location of a target object.
[0015] The method of operation of the electronic device may include an operation of identifying at least one first association information among the plurality of association information, which includes text corresponding to the target object.
[0016] The method of operation of the electronic device may include providing a candidate location of the target object based on at least one first association information.
[0017] A storage medium for storing instructions that can be read by a computer may be provided.
[0018] When the above instructions are executed individually or collectively by at least one processor of an electronic device, the electronic device may cause each of the multiple texts according to scene analysis by an artificial intelligence model for multiple images.
[0019] Each of the above multiple texts can be associated with the scene analysis results for the objects included in each of the above multiple images.
[0020] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to generate and store multiple association information between each of the plurality of texts and each of the plurality of images.
[0021] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check a query for the location of a target object.
[0022] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify at least one first association information among the plurality of association information, which includes text corresponding to the target object.
[0023] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide a candidate location of the target object based on the at least one first association information.
[0024] The electronic device may include at least one processor and a memory for storing instructions.
[0025] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check a query for the location of a target object.
[0026] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to include the target object.
[0027] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the first data for the at least one shooting location among the time series data for the location of the user of the electronic device.
[0028] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a candidate location identified by excluding the first data among the time series data.
[0029] The method of operation of the electronic device may include an operation to check a query for the location of a target object.
[0030] The method of operation of the electronic device may include an operation of checking data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to include the target object.
[0031] The method of operating the electronic device may include the operation of confirming a first data for at least one shooting location among time series data for the location of the user of the electronic device.
[0032] The method of operation of the above electronic device may include an operation of providing a candidate location identified by excluding the first data among the time series data.
[0033] A storage medium for storing instructions that can be read by a computer may be provided.
[0034] When the above instructions are executed individually or collectively by at least one processor of the electronic device, the electronic device may be caused to check a query for the location of a target object.
[0035] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to include the target object.
[0036] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the first data for the at least one shooting location among the time series data for the location of the user of the electronic device.
[0037] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a candidate location identified by excluding the first data among the time series data.
[0038] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0039] FIG. 2 is a perspective view for explaining the internal configuration of a wearable electronic device according to one embodiment of the present disclosure.
[0040] FIG. 3a is a drawing showing the front and rear of a wearable electronic device according to one embodiment.
[0041] FIG. 3b is a drawing showing the front and rear of a wearable electronic device according to one embodiment.
[0042] FIG. 4 is a drawing for explaining command processing according to one embodiment.
[0043] FIG. 5 is a diagram illustrating a method of operation of an electronic device according to one embodiment.
[0044] FIG. 6 is a diagram illustrating the process of providing candidate locations according to one embodiment.
[0045] FIG. 7 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0046] FIG. 8 is a diagram illustrating candidate location identification based on user movement information according to one embodiment.
[0047] FIG. 9a is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0048] FIGS. 9b, 9c, and 9d are drawings for explaining a target object selection process according to one embodiment.
[0049] FIG. 10 is a drawing for explaining the provision of candidate locations according to one embodiment.
[0050] FIG. 11 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0051] Functions related to artificial intelligence according to the present disclosure are operated through a processor and memory. The processor may be composed of one or more processors. In this case, the one or more processors may be general-purpose processors such as CPUs, APs, and DSPs (Digital Signal Processors), graphics-dedicated processors such as GPUs and VPUs (Vision Processing Units), or artificial intelligence-dedicated processors such as NPUs. The one or more processors control the processing of input data according to predefined operation rules or artificial intelligence models stored in memory. Alternatively, if the one or more processors are artificial intelligence-dedicated processors, the artificial intelligence-dedicated processors may be designed with a hardware structure specialized for processing a specific artificial intelligence model.
[0052] The predefined rules of operation or artificial intelligence models are characterized by being created through learning. Here, being created through learning means that a predefined rules of operation or artificial intelligence models configured to perform desired characteristics (or objectives) are created by a basic artificial intelligence model being trained using multiple learning data by a learning algorithm. Such learning may be performed on the device itself where the artificial intelligence according to the present disclosure is executed, or it may be performed through a separate server and / or system. Examples of learning algorithms include supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but are not limited to the examples described above.
[0053] An artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of previous layers and the multiple weights. The multiple weights possessed by the multiple neural network layers can be optimized based on the learning results of the artificial intelligence model. For example, the multiple weights may be updated so that the loss value or cost value obtained from the artificial intelligence model during the learning process is reduced or minimized. Artificial neural networks may include deep neural networks (DNNs), such as Convolutional Neural Networks (CNNs), Deep Neural Networks (DNNs), Recurrent Neural Networks (RNNs), Restricted Boltzmann Machines (RBMs), Deep Belief Networks (DBNs), Bidirectional Recurrent Deep Neural Networks (BRDNNs), or Deep Q-Networks, but are not limited to the examples mentioned above.
[0054] According to the present disclosure, image data can be used as input data for an artificial intelligence model to obtain output data that recognizes an image. The artificial intelligence model can be created through learning. Here, being created through learning means that a basic artificial intelligence model is trained using multiple training data by a learning algorithm, thereby creating a predefined operational rule or artificial intelligence model configured to perform a desired characteristic (or objective). The artificial intelligence model may be composed of multiple neural network layers. Each of the multiple neural network layers has multiple weight values and performs neural network operations through operations between the results of operations of the previous layer and the multiple weights.
[0055] Visual understanding is a technology that perceives and processes objects like human vision, and includes object recognition, object tracking, image retrieval, human recognition, scene recognition, 3D reconstruction / localization, and image enhancement.
[0056] FIG. 1 is a block diagram of an electronic device in a network environment according to one embodiment.
[0057] Referring to FIG. 1, in a network environment (100), an electronic device (101) may communicate with an electronic device (102) through a first network (198) (e.g., a short-range wireless communication network) or with at least one of an electronic device (104) or a server (108) through a second network (199) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (101) may communicate with the electronic device (104) through a server (108). According to one embodiment, the electronic device (101) may include a processor (120), memory (130), input module (150), sound output module (155), display module (160), audio module (170), sensor module (176), interface (177), connection terminal (178), haptic module (179), camera module (180), power management module (188), battery (189), communication module (190), subscriber identification module (196), or antenna module (197). In some embodiments, at least one of these components (e.g., connection terminal (178)) may be omitted from the electronic device (101), or one or more other components may be added. In some embodiments, some of these components (e.g., sensor module (176), camera module (180), or antenna module (197)) may be integrated into a single component (e.g., display module (160)).
[0058] The processor (120) can control at least one other component (e.g., a hardware or software component) of the electronic device (101) connected to the processor (120) by executing software (e.g., a program (140)), and can perform various data processing or operations. According to one embodiment, as at least part of the data processing or operations, the processor (120) can store commands or data received from other components (e.g., a sensor module (176) or a communication module (190)) in volatile memory (132), process the commands or data stored in volatile memory (132), and store the resulting data in non-volatile memory (134). According to one embodiment, the processor (120) may include a main processor (121) (e.g., a central processing unit or an application processor) or an auxiliary processor (123) that can operate independently or together with it (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor). For example, if the electronic device (101) includes a main processor (121) and an auxiliary processor (123), the auxiliary processor (123) may be configured to use lower power than the main processor (121) or to be specialized for a designated function. The auxiliary processor (123) may be implemented separately from the main processor (121) or as part thereof.
[0059] The auxiliary processor (123) may control at least some of the functions or states associated with at least one component of the electronic device (101) (e.g., display module (160), sensor module (176), or communication module (190)) on behalf of the main processor (121) while the main processor (121) is in an inactive (e.g., sleep) state, or together with the main processor (121) while the main processor (121) is in an active (e.g., application execution) state. According to one embodiment, the auxiliary processor (123) (e.g., image signal processor or communication processor) may be implemented as part of another functionally related component (e.g., camera module (180) or communication module (190)). According to one embodiment, the auxiliary processor (123) (e.g., neural network processing unit) may include a hardware structure specialized for processing an artificial intelligence model. The artificial intelligence model may be generated through machine learning. Such learning may be performed, for example, on the electronic device (101) itself where the artificial intelligence model is executed, or through a separate server (e.g., server (108)). The learning algorithm may include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model may include a plurality of artificial neural network layers.An artificial neural network may be a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to the hardware structure, the artificial intelligence model may include a software structure, either additionally or substantially.
[0060] The memory (130) can store various data used by at least one component of the electronic device (101) (e.g., processor (120) or sensor module (176)). The data may include, for example, input data or output data for software (e.g., program (140)) and related commands. The memory (130) may include volatile memory (132) or non-volatile memory (134).
[0061] The program (140) may be stored as software in memory (130) and may include, for example, an operating system (142), middleware (144), or an application (146).
[0062] The input module (150) can receive commands or data to be used for a component of the electronic device (101) (e.g., processor (120)) from outside the electronic device (101) (e.g., user). The input module (150) may include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0063] The sound output module (155) can output a sound signal to the outside of the electronic device (101). The sound output module (155) may include, for example, a speaker or a receiver. The speaker may be used for general purposes, such as multimedia playback or recording playback. The receiver may be used to receive incoming calls. According to one embodiment, the receiver may be implemented separately from the speaker or as part thereof.
[0064] The display module (160) can visually provide information to an external (e.g., user) of the electronic device (101). The display module (160) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling said device. According to one embodiment, the display module (160) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of the force generated by said touch.
[0065] The audio module (170) can convert sound into an electrical signal or, conversely, convert an electrical signal into sound. According to one embodiment, the audio module (170) can acquire sound through the input module (150) or output sound through the sound output module (155) or an external electronic device (e.g., electronic device (102)) (e.g., speaker or headphones) connected directly or wirelessly to the electronic device (101).
[0066] The sensor module (176) can detect the operating state of the electronic device (101) (e.g., power or temperature) or the external environmental state (e.g., user state) and generate an electrical signal or data value corresponding to the detected state. According to one embodiment, the sensor module (176) may include, for example, a gesture sensor, a gyroscope sensor, a barometric pressure sensor, a magnetic sensor, an accelerometer sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biosensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0067] The interface (177) may support one or more specified protocols that can be used for the electronic device (101) to be connected directly or wirelessly to an external electronic device (e.g., electronic device (102)). According to one embodiment, the interface (177) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0068] The connection terminal (178) may include a connector through which the electronic device (101) can be physically connected to an external electronic device (e.g., electronic device (102)). According to one embodiment, the connection terminal (178) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0069] The haptic module (179) can convert an electrical signal into a mechanical stimulus (e.g., vibration or movement) or an electrical stimulus that can be perceived by the user through tactile or kinesthetic senses. According to one embodiment, the haptic module (179) may include, for example, a motor, a piezoelectric element, or an electric stimulation device.
[0070] The camera module (180) can capture still images and video. According to one embodiment, the camera module (180) may include one or more lenses, image sensors, image signal processors, or flashes.
[0071] The power management module (188) can manage power supplied to the electronic device (101). According to one embodiment, the power management module (188) can be implemented, for example, as at least part of a power management integrated circuit (PMIC).
[0072] The battery (189) can supply power to at least one component of the electronic device (101). According to one embodiment, the battery (189) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0073] The communication module (190) can support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between an electronic device (101) and an external electronic device (e.g., electronic device (102), electronic device (104), or server (108)), and the performance of communication through the established communication channel. The communication module (190) may include one or more communication processors that operate independently of the processor (120) (e.g., application processor) and support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (190) may include a wireless communication module (192) (e.g., cellular communication module, short-range wireless communication module, or GNSS (global navigation satellite system) communication module) or a wired communication module (194) (e.g., LAN (local area network) communication module, or power line communication module). The corresponding communication module among these communication modules can communicate with an external electronic device (104) through a first network (198) (e.g., a short-range communication network such as Bluetooth, WiFi (wireless fidelity) direct, or IrDA (infrared data association)) or a second network (199) (e.g., a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a LAN or WAN)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (192) can identify or authenticate the electronic device (101) within a communication network such as the first network (198) or the second network (199) using subscriber information (e.g., International Mobile Subscriber Identifier (IMSI)) stored in the subscriber identification module (196).
[0074] The wireless communication module (192) can support 5G networks and next-generation communication technologies following 4G networks, for example, new radio access technology. NR access technology can support high-speed transmission of high-capacity data (enhanced mobile broadband (eMBB)), minimization of terminal power and connection of multiple terminals (massive machine type communications (mMTC)), or high reliability and low latency (ultra-reliable and low-latency communications (URLLC)). The wireless communication module (192) can support a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate, for example. The wireless communication module (192) can support various technologies for securing performance in the high-frequency band, such as beamforming, massive MIMO (multiple-input and multiple-output), full-dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large-scale antenna. The wireless communication module (192) can support various requirements specified in the electronic device (101), external electronic device (e.g., electronic device (104)), or network system (e.g., second network (199)). According to one embodiment, the wireless communication module (192) may support a Peak data rate (e.g., 20 Gbps or more) for eMBB realization, loss coverage (e.g., 164 dB or less) for mMTC realization, or U-plane latency (e.g., downlink (DL) and uplink (UL) each 0.5 ms or less, or round trip 1 ms or less) for URLLC realization.
[0075] An antenna module (197) can transmit a signal or power to or from an external source (e.g., an external electronic device). According to one embodiment, the antenna module (197) may include an antenna comprising a radiator made of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). According to one embodiment, the antenna module (197) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as a first network (198) or a second network (199), may be selected from the plurality of antennas, for example, by a communication module (190). A signal or power may be transmitted or received between the communication module (190) and an external electronic device through the selected at least one antenna. According to some embodiments, in addition to the radiator, other components (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as part of the antenna module (197).
[0076] According to various embodiments, the antenna module (197) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent to a first surface (e.g., bottom surface) of the printed circuit board and capable of supporting a specified high frequency band (e.g., mmWave band), and a plurality of antennas (e.g., array antennas) disposed on or adjacent to a second surface (e.g., top surface or side surface) of the printed circuit board and capable of transmitting or receiving a signal of the specified high frequency band.
[0077] At least some of the above components can be connected to each other via a communication method between peripheral devices (e.g., bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)) and exchange signals (e.g., commands or data) with each other.
[0078] According to one embodiment, commands or data may be transmitted or received between an electronic device (101) and an external electronic device (104) through a server (108) connected to a second network (199). Each of the external electronic devices (102, or 104) may be the same or a different type of device as the electronic device (101). According to one embodiment, all or part of the operations performed on the electronic device (101) may be performed on one or more of the external electronic devices (102, 104, or 108). For example, if the electronic device (101) needs to perform a function or service automatically or in response to a request from a user or another device, the electronic device (101) may request one or more external electronic devices to perform at least part of the function or service instead of performing the function or service itself or additionally. One or more external electronic devices that receive the above request may execute at least part of the requested function or service, or additional function or service related to the request, and transmit the result of the execution to the electronic device (101). The electronic device (101) may provide the result as is or additionally processed as at least part of the response to the request. For this purpose, for example, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used. The electronic device (101) may provide ultra-low latency services using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (104) may include an Internet of Things (IoT) device. The server (108) may be an intelligent server using machine learning and / or neural networks. According to one embodiment, the external electronic device (104) or the server (108) may be included within a second network (199).The electronic device (101) can be applied to intelligent services (e.g., smart home, smart city, smart car, or healthcare) based on 5G communication technology and IoT-related technology.
[0079] FIG. 2 is a perspective view for explaining the internal configuration of a wearable electronic device according to one embodiment of the present disclosure.
[0080] Referring to FIG. 2, a wearable electronic device (200) according to one embodiment of the present disclosure may include at least one of a light output module (211), a display member (201), and a camera module (250).
[0081] According to one embodiment of the present disclosure, the light output module (211) may include a light source capable of outputting an image and a lens that guides the image to a display member (201). According to one embodiment of the present disclosure, the light output module (211) may include at least one of a liquid crystal display (LCD), a digital mirror device (DMD), a liquid crystal on silicon (LCoS), an organic light emitting diode (OLED), an organic light emitting diode on silicon (OLEDoS), or a micro light emitting diode (micro LED).
[0082] According to one embodiment of the present disclosure, the display member (201) may include an optical waveguide (e.g., a waveguide). According to one embodiment of the present disclosure, an output image of an optical output module (211) incident on one end of the optical waveguide may propagate within the optical waveguide and be provided to a user. According to one embodiment of the present disclosure, the optical waveguide may include at least one diffractive element (e.g., a Diffractive Optical Element (DOE), a Holographic Optical Element (HOE)) or a reflective element (e.g., a reflective mirror). For example, the optical waveguide may guide the output image of the optical output module (211) to the user's eye using at least one diffractive element or reflective element.
[0083] According to one embodiment of the present disclosure, the camera module (250) can capture still images and / or video. According to one embodiment, the camera module (250) may be placed within a lens frame and around a display member (201).
[0084] According to one embodiment of the present disclosure, the first camera module (251) can capture and / or recognize the trajectory of the user's eye (e.g., pupil, iris) or gaze. According to one embodiment of the present disclosure, the first camera module (251) can periodically or non-periodically transmit information related to the trajectory of the user's eye or gaze (e.g., trajectory information) to a processor (e.g., processor (120) of FIG. 1).
[0085] According to one embodiment of the present disclosure, the second camera module (253) can capture an external image.
[0086] According to one embodiment of the present disclosure, a third camera module (255) may be used for hand detection and tracking and recognition of user gestures (e.g., hand movements). According to one embodiment of the present disclosure, a third camera module (255) may be used for 3 degrees of freedom (3DoF), 6DoF head tracking, location (space, environment) recognition, and / or movement recognition. According to one embodiment of the present disclosure, a second camera module (253) may be used for hand detection and tracking and recognition of user gestures. According to one embodiment of the present disclosure, at least one of the first camera module (251) to the third camera module (255) may be replaced with a sensor module (e.g., a LiDAR sensor). For example, the sensor module may include at least one of a vertical cavity surface emitting laser (VCSEL), an infrared sensor, and / or a photodiode.
[0087] FIG. 3a is a drawing showing the front and rear of a wearable electronic device according to one embodiment.
[0088] FIG. 3b is a drawing showing the front and rear of a wearable electronic device according to one embodiment.
[0089] Referring to FIG. 3a and FIG. 3b, in one embodiment, camera modules (311, 312, 313, 314, 315, 316) and / or a depth sensor (317) for acquiring information related to the surrounding environment of the wearable electronic device (300) may be disposed on the first surface (310) of the housing.
[0090] In one embodiment, camera modules (311, 312) can acquire images related to the surrounding environment of a wearable electronic device.
[0091] In one embodiment, camera modules (313, 314, 315, 316) can acquire images while the wearable electronic device is worn by a user. Camera modules (313, 314, 315, 316) can be used for hand detection, tracking, and user gesture (e.g., hand movements) recognition. Camera modules (313, 314, 315, 316) can be used for 3DoF, 6DoF head tracking, position (space, environment) recognition, and / or movement recognition. In one embodiment, camera modules (311, 312) may be used for hand detection and tracking and user gestures.
[0092] In one embodiment, the depth sensor (317) may be configured to transmit a signal and receive a signal reflected from an object, and may be used for purposes such as time of flight (TOF) to determine the distance to an object.
[0093] According to one embodiment, a face recognition camera module (325, 326,) and / or a display (321) (and / or a lens) may be disposed on the second surface (320) of the housing.
[0094] In one embodiment, a face recognition camera module (325, 326) adjacent to the display may be used to recognize the user's face or to recognize and / or track both of the user's eyes.
[0095] In one embodiment, the display (321) (and / or lens) may be disposed on a second surface (320) of the wearable electronic device (300). In one embodiment, the wearable electronic device (300) may not include camera modules (315, 316) among a plurality of camera modules (313, 314, 315, 316). Although not illustrated in FIG. 3a and 3b, the wearable electronic device (300) may further include at least one of the configurations illustrated in FIG. 2.
[0096] As described above, according to one embodiment, the wearable electronic device (300) may have a form factor for being worn on a user's head. The wearable electronic device (300) may further include a strap and / or a wearing member for being secured on a part of the user's body. The wearable electronic device (101) may provide a user experience based on augmented reality, virtual reality, and / or mixed reality while being worn on the user's head.
[0097] FIG. 4 is a drawing for explaining command processing according to one embodiment.
[0098] According to one embodiment, the target object location search module (410) may receive data from at least one wearable device (200, 300) and / or an electronic device (101). The target object location search module (410) may be included in (or executed by) a server (108), for example, but is not limited to. The server (108) may be, for example, an Internet of Things (IoT) server linked to a user account, but this is exemplary and there are no limitations on the implementation form, purpose, and / or number of the server (108). Those skilled in the art will understand that the target object location search module (410) may be included in (or executed by) the wearable device (200, 300), or the target object location search module (410) may be included in (or executed by) the electronic device (101). In one example, all modules included in the target object location search module (410) may be included in a single entity (or executed by a single entity). In one example, the modules included in the target object location search module (410) may be distributed among two or more entities (or executed by two or more entities).
[0099] The target object location search module (410) may provide information about candidate locations in response to a query regarding the location where the target object is located. The target object location search module (410) may include, but is not limited to, a device information management module (411), an AR data-to-text conversion module (412), a target data input-to-text conversion module (413), a device data-to-text conversion module (414), a candidate location verification module (415), and / or a UI creation module (416). The target object location search module (410) may be coupled with, for example, a large-scale language model (LLM) (420), but this is exemplary, and the target object location search module (410) may be implemented to include the LLM (420).
[0100] The device information management module (411) can register and / or manage devices registered by a user under the same account (e.g., electronic device (101), wearable device (200) and / or wearable device (300)). The device information management module (411) can perform registration and / or management based on device information (e.g., Device Type, Model Name, Serial Number, International Mobile Equipment Identification Number (IMEI), and / or MAC Address, but is not limited to). The device information management module (411) can manage data from devices for a single user account, but this is exemplary. For example, the device information management module (411) may manage data from devices registered under another account for which integrated management is allowed (e.g., an account of another user in a family relationship) in an integrated manner with data from devices registered under the user account.
[0101] The AR data-to-text conversion module (412) can convert an image captured by an AR device (e.g., a wearable device (200) and / or a wearable device (300)) into text using an LLM (420). The AR device (e.g., a wearable device (200) and / or a wearable device (300)) can, for example, periodically capture an image and provide it to the target object location search module (410). For example, the AR device (e.g., a wearable device (200) and / or a wearable device (300)) can provide the captured image to the target object location search module (410) based on the occurrence of an image transmission trigger. The image transmission trigger may be, for example, a change in the location where the AR device (e.g., a wearable device (200) and / or a wearable device (300)) is located and / or a change in the captured field of view (FoV), but is not limited thereto. For example, the user may move from a first location to a second location while wearing the wearable device (200). The wearable device (200) may provide at least some of the images taken at the second location to the target object location search module (410) based on the change of location to the second location. Accordingly, the target object location search module (410) can acquire an image of the scene that the user is looking at (or that is taken by the wearable device (200)) without the user entering a special command.
[0102] The AR data-to-text conversion module (412) can store and / or manage the converted text in association with location information and time information where the image was taken. For example, Table 1 is an example of prompting data for an LLM (420) generated by the AR data-to-text conversion module (412).
[0103] Please check the attached image. #Attachment: Image #1 Please analyze the scene in the attached image. For example, describe the objects included in the image in detail, including information that allows the location of the image to be inferred. You must describe it as realistically as possible without exaggeration.
[0104] For example, based on prompting data such as Table 1, LLM (420) can provide text corresponding to an image such as Table 2.
[0105] The scene analysis results for the attached image are as follows: There is a green tumbler on a desk in what appears to be a classroom. There is also a computer monitor on the desk. The green tumbler is located to the right of the computer monitor. The desk is brown.
[0106] The AR data-to-text conversion module (412) can store text corresponding to an image such as Table 2. The entire text provided by the LLM (420) such as Table 2 may be stored and / or managed, but this is exemplary, and those skilled in the art will understand that only information regarding the location of an object may be selected (or organized) and stored and / or managed. The converted data may be stored in a database defined by the name "DB_from_AR". The stored data may include an "AR" tag to indicate that it was generated from an AR device, but is not limited to this. Table 3 is an example of converted data (e.g., DB_from_AR).
[0107] There is a yellow tumbler on a desk in what appears to be a classroom. There is also a computer monitor on the desk. A green tumbler is to the right of the computer monitor. School 2024-08-25-10:00 ARV, there is a green tumbler on the table in the living room with the sofa. Home / Living Room 2024-08-26-00:20 AR There are a blanket and pillow on the bed in the room. Home / Room #1 2024-08-26-11:00 AR There is a green tumbler on the treadmill in the gym with exercise equipment. Gym 2024-08-26-13:00 AR There is a notebook, a ballpoint pen, and a green tumbler on the desk. Library 2024-08-27-19:00 AR It is the space with the bed. There is a green tumbler on the desk. Home / Room#1 2024-08-27-21:00AR It looks like a kitchen with a refrigerator and a sink. There is a green tumbler on the sink. Home 2024-08-27-21:06ARTV, it is a living room with a sofa Home / Living Room 2024-08-28-09:03AR There is a bed and a desk. Home 2024-08-29-23:00AR It is a kitchen with a sink and a refrigerator. Home 2024-08-30-11:00AR There is a desk, a chair, and school friends. School 2024-08-31-09:00AR
[0108] AR data may be provided from a wearable device (200) and / or a wearable device (300). Meanwhile, those skilled in the art will understand that text such as Table 3 may be acquired and stored based on at least one data including voice recordings as well as images from the AR device. Such wearable devices (200, 300) generate AR data based on the user's field of vision and may generate images by periodically capturing the surrounding environment while the user is wearing them. For example, a wearable device (200), such as AR smart glasses, or a wearable device (300), such as a VST, may capture the surrounding environment while the user moves or performs a specific task and transmit this data to an AR data-to-text conversion module (412). Location information provided by the wearable devices (200, 300) may be utilized to subdivide zones or determine the user's movement path based on wireless communication data with surrounding devices. When location information is stored, the specific zone may be defined by utilizing device information registered to an account in the vicinity. For example, speakers, air conditioners, air purifiers, and lighting can be registered as subdivided zones within the house structure, such as the living room and master bedroom. Information about surrounding devices can be searched via wireless communication while wearing an AR device, and accordingly, the user can check in detail which zone they are currently in. The target data input-text conversion module (413) can receive a query for a target object and convert it into descriptive information. For example, the user may input a query based on the electronic device (101). If the user performs user voice or text input such as "find a green tumbler," the query may include text. Meanwhile, in another example, a user specification for a green tumbler among the images provided by the electronic device (101) may be input.In this case, the electronic device (101) may provide information regarding user customization for an image rather than text to the target object location search module (410) as a query. In this case, the target data input-text conversion module (413) may extract a portion of the image corresponding to a tumbler based on the query. The extracted portion may be converted into descriptive information by the LLM (420). For example, the LLM (420) may provide target text of "green tumbler" corresponding to the extracted portion.
[0109] The device data-to-text conversion module (414) can collect images, video files, recorded voice files, and / or location information stored in a device registered to a user account, such as an electronic device (101), or a watch-type wearable device worn by the user, and convert them into descriptive information to store and / or manage them. In this process, device data may be provided from the electronic device (101), but there are no limitations. The electronic device (101) includes various types of electronic devices such as smartphones, tablets, laptops, desktop computers, etc., and such devices may provide data created or stored by the user, such as photo galleries, video libraries, and / or voice recording files. Image data may be converted into descriptive information and stored along with location and / or time information. Among the video files, a scene containing target data may be identified, and the scene may be saved as a screenshot and then converted into descriptive information. Recorded voice files may be converted into descriptive information to check if there is content related to target data, and if such content is found, the descriptive information may be stored along with location and / or time information. The converted data is stored in a database named "DB_from_Non-AR", and each data item is tagged with device information. The data provided from the electronic device (101) may include information related to the user's behavior. For example, a photo taken by the user at a specific location or a voice file recorded by the user may be used to provide context related to that location. The database may be created based on, for example, a relational database (RDBMS), a NoSQL database, and / or a vector database, but is not limited to one. Meanwhile, location and / or time information may be stored as DB_from_Non-AR without descriptive information.For example, Table 4 is an example of DB_from_Non-AR containing descriptive information, and Table 5 may be an example of DB_from_Non-AR not containing descriptive information.
[0110] Text Location Time Device The treadmill is operating at the location with various exercise equipment. Gym 2024-08-30- 15:00 smartphone
[0111] Location Time Device House 2024-08-29-16:00 Watch Gym 2024-08-30-15:00 Watch School 2024-08-31-09:00 Watch
[0112] The candidate location verification module (415) can apply the "Target" tag to data in the "DB_from_AR" database that contains the Target Text in the descriptive information. The candidate location verification module (415) can apply the "Non-Target" tag to data in the "DB_from_AR" database that does not contain the Target Text in the descriptive information. The candidate location verification module (415) can apply the "Latest_Date" tag to information where the Target Text was last recorded in a time series. Table 6 is an example of a database with Target / Non-Target and Latest_Date applied.
[0113] There is a yellow tumbler on a desk in what appears to be a classroom. There is also a computer monitor on the desk. The green tumbler is to the right of the computer monitor. School 2024-08-25- 10:00 ARTargetN / A There is a green tumbler on the table in the living room where the sofa is. Home / Living Room 2024-08-26- 0:20 ARTargetN / A There is a blanket and pillow on the bed in the room. Home / Room #1 2024-08-26- 11:00 ARNon-TargetN / A There is a green tumbler on the treadmill in the gym where the exercise equipment is. Gym 2024-08-26- 13:00 ARTargetN / A There is a notebook, a ballpoint pen, and a green tumbler on the desk. Library 2024-08-27- 19:00 ARTargetN / A It is the space where the bed is. There is a green tumbler on the desk. Home / Room#1 2024-08-27- 21:00 ARNon-TargetN / A It looks like a kitchen with a refrigerator and a sink. There is a green tumbler on the sink. Home 2024-08-27- 21:06 ARNon-TargetOTV, it is a living room with a sofa. Home / Living Room 2024-08-28- 09:03 ARNon-TargetN / A There is a bed and a desk. Home 2024-08-29- 23:00 ARNon-TargetN / A It is a kitchen with a sink and a refrigerator. Home 2024-08-30- 11:00 ARNon-TargetN / A There is a desk, a chair, and school friends. School 2024-08-31- 09:00 ARNon-TargetN / A
[0114] For example, as shown in Table 5, candidate locations may be determined based on data tagged as Latest_Date in DB_from_AR and / or data tagged as Target prior to that data (e.g., Latest_Date - 7). Data for candidate locations may be named the "Location_from_AR_Target" table. Meanwhile, location information where no target was recorded during the period from data tagged as "Latest_Date" to "Current_Date" may be named the "Location_from_AR_No-Target" table. Table 7 is an example of Location_from_AR_Target, and Table 8 is an example of Location_from_AR_No-Target.
[0115] Location Time Device TargetLatest_Date School 2024-08-25-10:00 ARTargetN / A Home / Living Room 2024-08-26-00:20 ARTargetN / A Gym 2024-08-26-13:00 ARTargetN / A Library 2024-08-27-19:00 ARTargetN / A Home / Room #1 2024-08-27-21:00 ARTargetN / A Home 2024-08-27-21:06 ARNon-TargetO
[0116] Location Time Device TargetLatest_Date Home / Living Room 2024-08-28-09:03 ARNon-TargetN / A Home 2024-08-29-23:00 ARNon-TargetN / A Home 2024-08-30-11:00 ARNon-TargetN / A School 2024-08-31-09:00 ARNon-TargetN / A
[0117] In one example, the candidate location verification module (415) may provide candidate locations based on "Location_from_AR_Target" and / or "Location_from_AR_No-Target". Location-based technologies such as Simultaneous Localization and Mapping (SLAM), WiFi, Bluetooth, Beacon, RFID, and / or UWB may be used, and user path analysis may be performed accordingly. The candidate location verification module (415) may also provide candidate locations by further using "Location_from_Non-AR". Table 9 is an example of Location_from_Non-AR.
[0118] Location Time Device Home 2024-08-29-16:00 Watch Gym 2024-08-30-15:00 Watch Gym 2024-08-30-15:00 Smartphone School 2024-08-31-09:00 Watch
[0119] Information regarding the user's movement path can be verified based on "Location_from_Non-AR". For example, "Location_from_AR-No-Target" may be a location where no target object is identified when an AR device is worn. In one example, the candidate location verification module (415) may verify candidate locations by excluding locations where no target object was identified based on "Location_from_AR-No-Target" from the total information of locations where the user was located. For example, "Location_from_AR-No-Target" information may be excluded from "Location_from_Non-AR", and this may be named the "Candidate_Location" table, which can be represented as Table 10.
[0120] Location Time Device Gym 2024-08-30-15:00 Watch Gym 2024-08-30-15:00 Smartphone
[0121] For example, there may be cases where a user visits a specific location but temporarily removes the AR device at that location, resulting in a specific object not being photographed, or where another user moves the user's object. Taking this into consideration, the candidate location verification module (415) may verify the Location_From_Non-AR table as is, as the Candidate_Location table. Meanwhile, as described above, those skilled in the art will understand that the Location_from_AR_Target table may be managed based on data from other user accounts that have visited the relevant location, in addition to the specific user account. Depending on the implementation, information for identifying the user account may be included in the “Device” field. The UI generation module (416) may provide the candidate location (e.g., the “Candidate_Location” table) verified by the candidate location verification module (415) as a UI that the user can recognize. The candidate location may be one, but may also be represented as multiple.
[0122] FIG. 5 is a diagram illustrating a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 5 is to be described with reference to FIG. 6. FIG. 6 is a diagram illustrating a process of providing candidate locations according to one embodiment. Hereinafter, the electronic device (101) is described as performing the operations of FIG. 5, but this is exemplary and there is no limitation on the entity performing the operations of FIG. 5. For example, the operations of FIG. 5 may be performed by a server (108). For example, those skilled in the art will understand that some of the operations of FIG. 5 may be executed by the electronic device (101), and the remaining parts may be executed by the server (108).
[0123] According to one embodiment, the electronic device (101) can identify each of the multiple texts (e.g., the first to sixth texts of FIG. 6) based on scene analysis by an artificial intelligence model for multiple images (621, 622, 623, 624, 625, 626) captured as in FIG. 6 in operation 501. For example, the electronic device (101) may request scene analysis for multiple images (621, 622, 623, 624, 625, 626) from an artificial intelligence model (e.g., LLM). The electronic device (101) may request scene analysis for multiple images (621, 622, 623, 624, 625, 626) from an artificial intelligence model (e.g., LLM) based on prompting data such as, for example, Table 1, but is not limited thereto. The electronic device (101) can identify each of the multiple texts (e.g., the first to sixth texts of FIG. 6) according to scene analysis from an artificial intelligence model (e.g., LLM). For example, text such as Table 2 can be identified.
[0124] According to one embodiment, the electronic device (101) may, in operation 503, generate and store a plurality of association information (631, 632, 633, 634, 635, 636) between each of a plurality of texts (e.g., the first to sixth texts of FIG. 6) and each of a plurality of images (621, 622, 623, 624, 625, 626). One association information (631) may include, for example, a first text corresponding to a first image (621), a shooting location (P1) of the first image (621), and / or the first image (621). For example, the first text may be stored and / or managed together with the shooting location (P1) of the first image (621). The first text (631) may also be stored and / or managed together with the time of capture (t1) of the first image (621). Accordingly, multiple associated information, such as Table 3, may be generated and stored and / or managed.
[0125] According to one embodiment, the electronic device (101) can check a query for the location of a target object (650) in operation 505. In one example, the electronic device (101) can check a query based on text entered by a user. The user may enter a query, for example, “Find my green tumbler,” based on the SIP of the electronic device (101). In one example, the electronic device (101) may check a query by analyzing the user’s voice. The user may utter a user voice, for example, “Find my green tumbler.” Based on the analysis of the user’s voice, the electronic device (101) can check a query where, for example, “green tumbler” is the target object (650). In one example, the electronic device (101) may check a query based on a user selection of a portion of an image. For example, the electronic device (101) can confirm that a specific part of an image has been selected by the user and can request the LLM to analyze the area. The target object (650) can be identified by the LLM, and accordingly, the electronic device (101) can also check the query. The methods for checking the query described above are exemplary and are not limited to the methods.
[0126] According to one embodiment, the electronic device (101) can identify at least one first association information (641) containing text corresponding to a target object (650) among a plurality of association information (631, 632, 633, 634, 635, 636) in operation 507. For example, association information (631, 632, 633) may be identified as at least one first association information (641) containing text corresponding to a target object (650) based on the target object (650) being included in the first text, second text, and third text. For example, even if the target object (650) is not exactly included in the first text, or even if a word similar to the target object (650) is included in the first text, it may be determined that the association information (631) corresponding to the first text corresponds to the target object (650).
[0127] According to one embodiment, the electronic device (101) may provide a candidate location of a target object based on at least one first association information (641) in operation 509. In one example, the electronic device (101) may provide a third association information (633), which is the most recent association information among at least one first association information (641), as a candidate location. In one example, the electronic device (101) may provide a candidate location based on the third association information (633), which is the most recent association information, and user movement information. For example, the electronic device (101) may check association information (634, 635, 636) after the third association information (633), which is the most recent association information. According to one embodiment, the electronic device (101) may provide a candidate location by excluding locations (e.g., P1, P4) included in the association information (634, 635, 636) among the user movement information. In one example, the electronic device (101) may identify an intermediate location between the location (P3) of the most recent associated information, the third associated information (633), and subsequent locations (e.g., P1, P4) based on user movement information, and may provide the intermediate location and / or the last location (P3) as a candidate location. There may be one candidate location or multiple candidate locations.
[0128] FIG. 7 is a diagram illustrating a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 7 will be described with reference to FIG. 8. FIG. 8 is a diagram illustrating candidate location identification based on user movement information according to one embodiment.
[0129] According to one embodiment, the electronic device (101) can identify each of a plurality of texts (e.g., first text to sixth texts) based on scene analysis by an artificial intelligence model for a plurality of images (e.g., images of FIG. 6 (621, 622, 623, 624, 625, 626)) captured in operation 701. As described above, the electronic device (101) can identify each of a plurality of texts (e.g., first text to sixth texts) based on input to an artificial intelligence model (e.g., LLM) of prompting data such as Table 1, but is not limited thereto. The electronic device (101) can generate and store multiple association information (631, 632, 633, 634, 635, 636) between each of the multiple texts (e.g., the first to sixth texts) and each of the multiple images (e.g., the images of FIG. 6 (621, 622, 623, 624, 625, 626)) in operation 703. The electronic device (101) can check a query regarding the location of a target object (e.g., 650 of FIG. 6) in operation 705. The method for checking the query has been described above and is not repeated here. In operation 707, the electronic device (101) can identify at least one first association information (641) containing text corresponding to a target object (e.g., 650 in FIG. 6) among a plurality of association information (631, 632, 633, 634, 635, 636). As described above, text containing the target object (e.g., 650 in FIG. 6)) and / or text containing words similar to the target object (e.g., 650 in FIG. 6)) may be identified as at least one first association information (641), but there are no limitations.
[0130] According to one embodiment, the electronic device (101) can check time series data (810) for the user's location in operation 709. In the time series data (810) for the user's location, for example, locations where the user stayed (811, 812, 813, 814, 815, 816, 817, 818, 819) may be managed in a time series order, but there are no restrictions on the method of representation. The time series data (810) may be checked, for example, based on the DB_from_Non-AR described above, but there are no restrictions. The time series data (810) may also be checked based on DB_from_Non-AR and / or DB_from_AR.
[0131] According to one embodiment, the electronic device (101) can identify a second association information (642) in which the target object is not identified based on the last association information (633) in the time series among the shooting locations corresponding to the first association information (641) in the 711 operation. For example, the electronic device (101) can identify association information (634, 635, 636) for time points after the last association information (633) in the time series among the shooting locations corresponding to the first association information (641) as the second association information (642), but there are no limitations.
[0132] The electronic device (101) can provide candidate locations of a confirmed target object by excluding data corresponding to the second association information (642) among the time-series data (810) regarding the user's location in operation 713. For example, the electronic device (101) can exclude data (815, 818, 819) corresponding to the second association information (642) among the locations (820) that are time-series after the last association information (633) among the shooting locations corresponding to the first association information (641). Accordingly, P3 corresponding to data (813) and P2 corresponding to data (814, 816, 817) can be provided as candidate locations. Although the AR device cannot confirm whether the target object is placed at P2, P2 may be provided as a candidate location based on the use of user movement information. Meanwhile, as this is exemplary, the electronic device (101) may provide all of the locations (820) that are chronologically after the last associated information (633) among the shooting locations corresponding to the first associated information (641) as candidate locations, or may provide only a part of them as candidate locations.
[0133] Meanwhile, in another example, the electronic device (101) may provide P2, which is an intermediate position between the last associated information (633) in time series and the data (815) where the target object is not identified, among the shooting positions corresponding to the first associated information (641), as at least some of the candidate positions.
[0134] FIG. 9a is a drawing for explaining a method of operation of an electronic device according to one embodiment. The embodiment of FIG. 9a will be explained with reference to FIG. 9b to 9d. FIG. 9b to 9d is a drawing for explaining a target object selection process according to one embodiment.
[0135] According to one embodiment, the electronic device (101) can identify each of the plurality of texts (e.g., the first text to the sixth text) based on scene analysis by an artificial intelligence model for a plurality of images captured in operation 901 (e.g., at least one image of FIG. 6 (e.g., 621, 622, 623, 624, 625, 626)). Meanwhile, the at least one image (e.g., 621, 622, 623, 624, 625, 626) is not limited to exemplary and may be applied to at least part of the present disclosure as well as to this embodiment. Furthermore, those skilled in the art will understand that the plurality of images may be replaced by a single image. As described above, the electronic device (101) can identify the plurality of texts (e.g., the first text to the sixth text) based on input to an artificial intelligence model (e.g., LLM) of prompting data such as Table 1 Each can be verified, but without limitation. In operation 903, the electronic device (101) may generate and store at least one association information (e.g., 631, 632, 633, 634, 635, 636) between the shooting positions of each of the plurality of texts (e.g., the first text through the sixth text) and each of the plurality of images (e.g., the images of FIG. 6 (621, 622, 623, 624, 625, 626)). As described above, those skilled in the art will understand that the at least one association information (e.g., 631, 632, 633, 634, 635, 636) is merely exemplary. In operation 905, the electronic device (101) may verify a user selection regarding the location of the target object. For example, as in FIG. 9b, the electronic device (101) [can verify] the first A screen (920) can be displayed. The first screen (920) may include an object representing at least one category (e.g., 921, 922, 923).When any one of the objects (921, 922, 923) is selected, the electronic device (101) may provide content within the selected category. For example, if an object (921) corresponding to a photograph is selected, the electronic device (101) may provide a stored image (930) as in FIG. 9c. The electronic device (101) may confirm a selection (932) for a first object (931) among the images (930). The selection (932) may be, for example, a touch on the first object (931), but there are no limitations. The electronic device (101) may generate target text corresponding to the user selection in the 907 operation. For example, the electronic device (101) may generate target text by performing recognition on the first object (931), but there are no limitations on the method of generation. The electronic device (101) can identify at least one first association information including target text among a plurality of association information in operation 909. The electronic device (101) can provide a candidate location of a target object based on at least one first association information in operation 911. As described above, in one example, the electronic device (101) can provide the most recent association information among at least one first association information as a candidate location. In one example, the electronic device (101) may provide a candidate location based on the most recent association information and user movement information. For example, the electronic device (101) can identify association information after the most recent association information. The electronic device (101) may provide a candidate location by excluding a location included in the association information after the most recent association information among the user movement information. In one example, the electronic device (101) may determine an intermediate location between the location of the most recent associated information and the location thereafter based on user movement information, and may provide the intermediate location and / or the last location as a candidate location.There may be one candidate location or multiple candidate locations.
[0136] In one example, as shown in FIG. 9a, an object (924) that causes a query to be entered directly may be provided. When the object (924) is selected, the electronic device (101) may provide a text input window (940) and a soft input panel (SIP) (942), as shown in FIG. 9d. A query text (941) corresponding to user input via the SIP (942) may be displayed in the input window (940). In one example, as shown in FIG. 9a, an object (925) that causes a query input based on user voice may be provided. When the object (925) is selected, the electronic device (101) may enter a listening mode and may verify the query based on the analysis of the user voice obtained in the listening mode.
[0137] FIG. 10 is a drawing for explaining the provision of candidate locations according to one embodiment.
[0138] According to one embodiment, the electronic device (101) may represent objects (1011, 1012, 1013, 1014) indicating a location and time of shooting where a target object (e.g., a green tumbler) is identified based on an image captured by an AR device (e.g., a wearable device (200, 300)). The electronic device (101) may provide objects (1021, 1022) for candidate locations identified based on time-series data associated with user movement and a location where the target object (e.g., a green tumbler) is identified not to be placed. The objects (1021, 1022) for candidate locations are illustrated as including information about the location, information about the data collection device, and information about the time of collection, but this is exemplary and not limited. As described above, the electronic device (101) can identify candidate locations by excluding locations among the locations corresponding to time-series data associated with user movement where a target object (e.g., a green tumbler) is not placed, but there are no limitations on the method of identification. For example, as described with reference to FIG. 8, objects (1011, 1012, 1013, 1014) may be provided according to at least one first association information (641) containing text corresponding to the target object (650), but there are no limitations. For example, as described with reference to FIG. 8, among the data corresponding to the last image and subsequent images in a time series among at least one image determined to contain a target object (e.g., data (820) of FIG. 8), by excluding the first data (e.g., 815, 818, 819) for at least one shooting location corresponding to at least one image determined not to contain a target object, candidate locations (e.g., P2, P3, which are locations corresponding to data (813, 814, 815, 817)) can be identified.For example, an object (1021, 1022) for a candidate location in FIG. 10 may be provided, such as a method for verifying a candidate location based on the process of FIG. 8 (e.g., P2, P3, which are locations corresponding to data (813, 814, 815, 817)), but there are no limitations.
[0139] FIG. 11 is a drawing for explaining the operation method of an electronic device according to one embodiment.
[0140] According to one embodiment, the electronic device (101) can check a query for the location of a target object in operation 1101. For example, as described with reference to FIGS. 9b through 9d, the query for the location of a target object can be checked based on the analysis results of text input, the analysis results of user voice, and / or the analysis results of a selection of a specific area of an image. In operation 1103, the electronic device (101) can check data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to contain the target object. For example, a plurality of images captured by the electronic device (101) and / or another external electronic device can be classified into at least one image determined to contain the target object and at least one image determined not to contain the target object. In operation 1105, the electronic device (101) can identify first data (e.g., 815, 818, 819) for at least one shooting location corresponding to at least one image determined not to contain a target object among time-series data for a user's location (e.g., time-series data for a user's location in FIG. 8 (810)). In operation 1107, the electronic device (101) can provide candidate locations identified by excluding the first data among the time-series data. In one example, the electronic device (101) may provide candidate locations identified by excluding the first data (e.g., 815, 818, 819) from data corresponding to the last image and subsequent images in time-series among at least one image determined to contain a target object (e.g., data (820) in FIG. 8).
[0141] Meanwhile, as this is exemplary, the electronic device (101) may provide a location corresponding to the last image in a chronological order among at least one image confirmed to contain the target object among a plurality of previously captured images as a candidate location. Alternatively, the electronic device (101) may provide an intermediate location between the location corresponding to the last image in a chronological order among at least one image confirmed to contain the target object and the location corresponding to an image confirmed not to contain the target object as a candidate location.
[0142] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a shooting location corresponding to the at least one first association information and a candidate location of the target object identified based on time-series data regarding the location of the user of the electronic device.
[0143] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a second association information in which the target object is not identified, based on the last association information in a time series among the shooting positions corresponding to the first association information.
[0144] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a candidate location of the target object identified by excluding data corresponding to the second association information from at least some of the time series data regarding the location of the user of the electronic device.
[0145] At least some of the time-series data regarding the location of the user of the electronic device may include the last associated information in the time series among the shooting locations corresponding to the first associated information and at least one data after the last associated information.
[0146] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a second association information in which the target object is not identified, based on the last association information in a time series among the shooting positions corresponding to the first association information.
[0147] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide an intermediate location between the location corresponding to the last associated information in the time series and the location of the second associated information as the candidate location.
[0148] A method of operation of an electronic device may include an operation of verifying each of a plurality of texts based on scene analysis by an artificial intelligence model for a plurality of images. Each of the plurality of texts may be associated with the scene analysis results for objects included in each of the plurality of images.
[0149] The method of operation of the above electronic device may include the operation of generating and storing multiple association information between each of the plurality of texts and each of the plurality of images.
[0150] The method of operation of the above electronic device may include an operation of checking a query for the location of a target object.
[0151] The method of operation of the electronic device may include an operation of identifying at least one first association information among the plurality of association information, which includes text corresponding to the target object.
[0152] The method of operation of the electronic device may include providing a candidate location of the target object based on at least one first association information.
[0153] The method of operation of the above electronic device may include the operation of generating multiple prompting data for scene analysis for each of the multiple images.
[0154] The method of operation of the above electronic device may include an operation of verifying each of the plurality of texts based on the response by the artificial intelligence model to the plurality of prompting data.
[0155] The method of operation of the electronic device may include providing a shooting location corresponding to the last associated information in a time series among the at least one first associated information as a candidate location of the target object.
[0156] The method of operation of the electronic device may include providing a shooting location corresponding to at least one first association information and a candidate location of the target object identified based on time series data regarding the location of the user of the electronic device.
[0157] The method of operation of the electronic device may include an operation of confirming a second association information in which the target object is not confirmed, based on the last association information in a time series among the shooting positions corresponding to the first association information.
[0158] The method of operation of the electronic device may include providing a candidate location of the target object identified by excluding data corresponding to the second association information from at least a portion of the time series data regarding the location of the user of the electronic device.
[0159] At least some of the time-series data regarding the location of the user of the electronic device may include the last associated information in the time series among the shooting locations corresponding to the first associated information and at least one data after the last associated information.
[0160] The method of operation of the electronic device may include an operation of confirming a second association information in which the target object is not confirmed, based on the last association information in a time series among the shooting positions corresponding to the first association information.
[0161] The method of operation of the electronic device may include providing an intermediate location between the location corresponding to the last associated information in the time series and the location of the second associated information as the candidate location.
[0162] A storage medium for storing instructions that can be read by a computer may be provided.
[0163] When the above instructions are executed individually or collectively by at least one processor of an electronic device, the electronic device may cause each of the multiple texts according to scene analysis by an artificial intelligence model for multiple images.
[0164] Each of the above multiple texts can be associated with the scene analysis results for the objects included in each of the above multiple images.
[0165] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to generate and store multiple association information between each of the plurality of texts and each of the plurality of images.
[0166] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check a query for the location of a target object.
[0167] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify at least one first association information among the plurality of association information, which includes text corresponding to the target object.
[0168] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause to provide a candidate location of the target object based on the at least one first association information.
[0169] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause the electronic device to generate multiple prompting data for scene analysis for each of the multiple images.
[0170] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may cause each of the plurality of texts to be verified based on the response by the artificial intelligence model to the plurality of prompting data.
[0171] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a shooting location corresponding to the last associated information in a time series among the at least one first associated information as a candidate location of the target object.
[0172] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a shooting location corresponding to the at least one first association information and a candidate location of the target object identified based on time-series data regarding the location of the user of the electronic device.
[0173] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a second association information in which the target object is not identified, based on the last association information in a time series among the shooting positions corresponding to the first association information.
[0174] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a candidate location of the target object identified by excluding data corresponding to the second association information from at least some of the time series data regarding the location of the user of the electronic device.
[0175] At least some of the time-series data regarding the location of the user of the electronic device may include the last associated information in the time series among the shooting locations corresponding to the first associated information and at least one data after the last associated information.
[0176] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify a second association information in which the target object is not identified, based on the last association information in a time series among the shooting positions corresponding to the first association information.
[0177] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide an intermediate location between the location corresponding to the last associated information in the time series and the location of the second associated information as the candidate location.
[0178] The electronic device may include at least one processor and a memory for storing instructions.
[0179] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check a query for the location of a target object.
[0180] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to include the target object.
[0181] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the first data for the at least one shooting location among the time series data for the location of the user of the electronic device.
[0182] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a candidate location identified by excluding the first data among the time series data.
[0183] The method of operation of the electronic device may include an operation to check a query for the location of a target object.
[0184] The method of operation of the electronic device may include an operation of checking data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to include the target object.
[0185] The method of operating the electronic device may include the operation of confirming a first data for at least one shooting location among time series data for the location of the user of the electronic device.
[0186] The method of operation of the above electronic device may include an operation of providing a candidate location identified by excluding the first data among the time series data.
[0187] A storage medium for storing instructions that can be read by a computer may be provided.
[0188] When the above instructions are executed individually or collectively by at least one processor of the electronic device, the electronic device may be caused to check a query for the location of a target object.
[0189] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to check data for at least one shooting location corresponding to at least one image among a plurality of previously captured images that is determined not to include the target object.
[0190] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to identify the first data for the at least one shooting location among the time series data for the location of the user of the electronic device.
[0191] When the above instructions are executed individually or collectively by the at least one processor, the electronic device may be caused to provide a candidate location identified by excluding the first data among the time series data.
[0192] The electronic device according to the various embodiments disclosed in this document may be of various forms. The electronic device may include, for example, a portable communication device (e.g., a smartphone), a computer device, a portable multimedia device, a portable medical device, a camera, a wearable device, or a consumer electronics device. The electronic device according to the embodiments of this document is not limited to the devices described above.
[0193] The various embodiments of this document and the terms used therein are not intended to limit the technical features described in this document to specific embodiments, and should be understood to include various modifications, equivalents, or substitutions of said embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of said items unless the relevant context clearly indicates otherwise. In this document, phrases such as "A or B," "at least one of A and B," "at least one of A or B," "A, B or C," "at least one of A, B and C," and "at least one of A, B, or C" may each include any one of the items listed together in the corresponding phrase, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used simply to distinguish said components from other said components and do not limit said components in any other aspect (e.g., importance or order). Where any (e.g., 1st) component is referred to as “coupled” or “connected” to another (e.g., 2nd) component, with or without the terms “functionally” or “communicationly,” it means that said any component may be connected to said other component directly (e.g., via a wire), wirelessly, or through a third component.
[0194] The term “module” as used in the various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit, for example. A module may be a component formed integrally, or a minimum unit of said component or a part thereof that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0195] One embodiment of the present document may be implemented as software (e.g., program (140)) comprising one or more instructions stored in a storage medium (e.g., internal memory (136) or external memory (138)) readable by a machine (e.g., electronic device (101)). For example, a processor (e.g., processor (120)) of the machine (e.g., electronic device (101)) may call at least one of the one or more instructions stored in the storage medium and execute it. This enables the machine to be operated to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code that can be executed by an interpreter. The storage medium readable by the machine may be provided in the form of a non-transitory storage medium. Here, 'non-temporary' simply means that the storage medium is a tangible device and does not contain a signal (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently and cases where it is stored temporarily.
[0196] According to one embodiment, the method according to the embodiments disclosed herein may be provided by being included in a computer program product. The computer program product may be traded between a seller and a buyer as a product. The computer program product may be distributed in the form of a device-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or distributed online (e.g., download or upload) through an application store (e.g., Play Store™) or directly between two user devices (e.g., smartphones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily created on a device-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or a relay server.
[0197] According to one embodiment, each component (e.g., module or program) of the components described above may include a singular or multiple entities, and some of the multiple entities may be separated and placed in other components. According to one embodiment, one or more of the components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Generally or additionally, multiple components (e.g., module or program) may be integrated into a single component. In this case, the integrated component may perform one or more functions of each of the components of the multiple components in the same or similar manner as those performed by the corresponding components among the multiple components prior to integration. According to one embodiment, operations performed by the module, program, or other components may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
Claims
1. In an electronic device, At least one processor; and The electronic device includes a memory for storing instructions, wherein the instructions are executed individually or collectively by the at least one processor, thereby: Each of the multiple texts is identified based on scene analysis by an artificial intelligence model for multiple images, and each of the multiple texts is associated with the scene analysis results for objects included in each of the multiple images, and Generating and storing multiple association information between each of the plurality of texts and each of the plurality of images, and Check the query for the location of the target object, and Identify at least one first association information among the plurality of association information above that includes text corresponding to the target object, and An electronic device that causes to provide a candidate location of the target object based on at least one first association information.
2. In Paragraph 1, When the above instructions are executed individually or collectively by the at least one processor, the electronic device: Generate multiple prompting data for scene analysis for each of the above multiple images, and An electronic device that causes each of the plurality of texts to be verified based on the response by the artificial intelligence model to the plurality of prompting data.
3. In any one of paragraphs 1 to 2, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that causes a shooting location corresponding to the last associated information in a time series among at least one first associated information to be provided as a candidate location of the target object.
4. In any one of paragraphs 1 to 3, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, An electronic device that causes to provide a shooting location corresponding to at least one first association information and a candidate location of the target object identified based on time-series data regarding the location of the user of the electronic device.
5. In any one of paragraphs 1 to 4, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on the last associated information in a time series among the shooting locations corresponding to the first associated information, a second associated information in which the target object is not identified is identified, and An electronic device that causes to provide a candidate location of the target object identified by excluding data corresponding to the second association information from at least a portion of the time series data regarding the location of the user of the electronic device.
6. In any one of paragraphs 1 through 5, An electronic device in which at least some of the time-series data regarding the location of a user of the electronic device comprises the last associated information in the time series among the shooting locations corresponding to the first associated information and at least one data after the last associated information.
7. In any one of paragraphs 1 through 6, When the above instructions are executed individually or collectively by the at least one processor, the electronic device, Based on the last associated information in a time series among the shooting locations corresponding to the first associated information, a second associated information in which the target object is not identified is identified, and An electronic device that causes an intermediate location between the location corresponding to the last associated information in the above time series and the location of the second associated information to be provided as the candidate location.
8. An operation of verifying each of a plurality of texts based on scene analysis by an artificial intelligence model for a plurality of images, wherein each of the plurality of texts is associated with a scene analysis result for objects included in each of the plurality of images; The operation of generating and storing multiple association information between each of the plurality of texts and each of the plurality of images; Action to check a query for the location of a target object; An operation to identify at least one first association information among the plurality of association information, including text corresponding to the target object; and An operation of providing a candidate location of the target object based on at least one first association information. A method of operation of an electronic device including 9. In Paragraph 8, The operation of generating multiple prompting data for scene analysis for each of the above multiple images; and An operation to verify each of the plurality of texts based on the response by the artificial intelligence model to the plurality of prompting data. A method of operation of an electronic device including 10. In any one of paragraphs 8 to 9, An operation of providing a shooting location corresponding to the last associated information in a time series among the at least one first associated information as a candidate location of the target object; A method of operation of an electronic device including 11. In any one of paragraphs 8 through 10, Operation of providing a shooting location corresponding to at least one first association information and a candidate location of the target object identified based on time-series data regarding the location of the user of the electronic device. A method of operation of an electronic device including 12. In any one of paragraphs 8 through 11, An operation to identify second association information in which the target object is not identified, based on the last association information in a time series among the shooting positions corresponding to the first association information; and An operation of providing a candidate location of the target object identified by excluding data corresponding to the second association information from at least a portion of the time-series data regarding the location of the user of the electronic device. A method of operation of an electronic device including 13. In any one of paragraphs 8 through 12, A method of operation of an electronic device wherein at least a portion of the time-series data regarding the location of a user of the electronic device comprises the last associated information in the time series among the shooting locations corresponding to the first associated information and at least one data after the last associated information.
14. In any one of paragraphs 8 through 13, An operation to identify second association information in which the target object is not identified, based on the last association information in a time series among the shooting positions corresponding to the first association information; and An operation of providing an intermediate location between the location corresponding to the last associated information in the above time series and the location of the second associated information as the candidate location. A method of operation of an electronic device including 15. In a storage medium storing computer-readable instructions, wherein the instructions are executed individually or collectively by at least one processor of an electronic device, the electronic device: Each of the multiple texts is identified based on scene analysis by an artificial intelligence model for multiple images, and each of the multiple texts is associated with the scene analysis results for objects included in each of the multiple images, and Generating and storing multiple association information between each of the plurality of texts and each of the plurality of images, and Check the query for the location of the target object, and Identify at least one first association information among the plurality of association information above that includes text corresponding to the target object, and A storage medium that causes to provide a candidate location of the target object based on at least one first association information.