Electronic device and method for operating electronic device
By determining an initial alignment relationship with external devices and adjusting the ROI based on posture changes, the electronic device accurately identifies and responds to user inputs, addressing the challenge of changing orientations and improving object recognition.
Patent Information
- Application Number
- PCT/KR2025/006361
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-11
- Filing Date
- 2025-05-12
- Publication Date
- 2025-12-26
AI Technical Summary
Electronic devices struggle to accurately identify the user's intended object within an image when the device's posture changes relative to the user's intended direction, leading to inaccurate responses.
The electronic device determines an initial alignment relationship with external devices worn by the user, adjusts the region of interest (ROI) based on posture changes, and uses AI models to identify and respond to user inputs, ensuring accurate object recognition and response generation.
This approach enhances the accuracy of object recognition and response generation by aligning the device's posture with the user's intended direction, even when the device's orientation changes, providing precise and relevant feedback.
Smart Images

Figure KR2025006361_26122025_PF_FP_ABST
Abstract
Description
Electronic devices and methods of operating electronic devices
[0001] Embodiments disclosed in this document relate to techniques for recognizing objects in an image in an electronic device and providing responses related to the objects.
[0002] Electronic devices (e.g., mobile terminals, wearable devices, or attachable devices) that include cameras can acquire external images using the cameras. When an image is acquired by an electronic device, the image may include multiple objects. The electronic device can recognize the multiple objects included in the image through image processing. A user can request the electronic device to perform a task or provide information related to an object in the acquired image, and the electronic device needs to identify the object intended by the user in the image to provide an appropriate response to the user's request. For example, the electronic device may set a portion of the image as the user's region of interest (ROI) to identify the object in the image.
[0003] The above information may be provided as background art to aid in understanding the present disclosure. No claim or determination is made as to whether any of the above is applicable as prior art related to the present disclosure.
[0004] An electronic device according to an embodiment disclosed in the present document includes a camera, a communication circuit, at least one sensor, a memory, and at least one processor, wherein the memory, when executed by the at least one processor, causes the electronic device to detect a degree of movement of the user using the at least one sensor while the electronic device is worn, attached, or fixed at a first location corresponding to a body part of the user, and, if the movement of the user is below a specified threshold value, detect a first initial posture of the electronic device using the at least one sensor, receive information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body part of the user through the communication circuit, determine an initial alignment relationship between the first initial posture of the electronic device and the second initial posture of the external electronic device, acquire an image including a plurality of objects outside the electronic device using the camera, detect the first posture of the electronic device using the at least one sensor, receive information related to the second posture of the external electronic device from the external electronic device, and determine the first initial posture, the second initial posture, the first posture, the second posture, or Instructions may be stored that determine a region of interest (ROI) of a user within the image based on at least some of the initial alignment relationships.
[0005] In addition, a method according to an embodiment disclosed in the present document may include an operation of detecting a degree of movement of a user using at least one sensor of an electronic device while the electronic device is worn, attached, or fixed at a first location corresponding to a part of a user's body, an operation of detecting a first initial posture of the electronic device using the at least one sensor when the movement of the user is less than or equal to a specified threshold, an operation of receiving information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another part of the user's body, an operation of determining an initial alignment relationship between the first initial posture of the electronic device and the second initial posture of the external electronic device, an operation of acquiring an image including a plurality of objects external to the electronic device using a camera of the electronic device, an operation of detecting the first posture of the electronic device using the at least one sensor, an operation of receiving information related to the second posture of the external electronic device from the external electronic device, and an operation of determining a region of interest (ROI) of the user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship. there is.
[0006] In addition, a storage medium according to an embodiment disclosed in the present document, when executed by at least one processor of an electronic device, detects a degree of movement of the user using at least one sensor of the electronic device while the electronic device is worn, attached, or fixed at a first location corresponding to a body part of the user, and, when the movement of the user is below a specified threshold value, detects a first initial posture of the electronic device using the at least one sensor, receives information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body part of the user, determines an initial alignment relationship between the first initial posture of the electronic device and the second initial posture of the external electronic device, acquires an image including a plurality of objects outside the electronic device using a camera of the electronic device, detects the first posture of the electronic device using the at least one sensor, receives information related to the second posture of the external electronic device from the external electronic device, and determines a region of interest of the user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship. Instructions and / or programs that determine return on investment (ROI) can be stored.
[0007] FIG. 1 is a diagram showing the usage status of an electronic device and an external electronic device according to one embodiment.
[0008] FIG. 2 is a block diagram of an electronic device according to one embodiment.
[0009] FIG. 3 is a drawing for explaining the operation of an electronic device according to one embodiment.
[0010] FIG. 4 is a drawing for explaining the operation of an electronic device according to one embodiment.
[0011] FIG. 5 is a diagram showing the configuration of an electronic device and an external electronic device according to one embodiment.
[0012] FIG. 6 is a diagram illustrating the structure of a multimodal model according to one embodiment.
[0013] FIG. 7 is a diagram showing a rotation matrix according to one embodiment.
[0014] FIG. 8 is a diagram illustrating an operation of an electronic device recognizing objects in an image according to one embodiment.
[0015] FIGS. 9A to 9C are drawings for explaining the operation of an electronic device according to one embodiment.
[0016] FIGS. 10A to 10C are drawings for explaining the operation of an electronic device according to one embodiment.
[0017] Fig. 11 is a flowchart of a method of operating an electronic device according to one embodiment.
[0018] Fig. 12 is a flowchart of a method of operating an electronic device according to one embodiment.
[0019] Fig. 13 is a flowchart of an operating method of an electronic device according to one embodiment.
[0020] Fig. 14 is a flowchart of a method of operating an electronic device according to one embodiment.
[0021] FIG. 15 illustrates an electronic device within a network environment according to various embodiments.
[0022] Figure 16 is a block diagram illustrating a program according to various embodiments.
[0023] FIG. 17 illustrates a generative artificial intelligence system according to one embodiment.
[0024] In connection with the description of the drawings, the same or similar reference numerals may be used for identical or similar components.
[0025] FIG. 1 is a diagram showing the usage status of an electronic device and an external electronic device according to one embodiment.
[0026] According to one embodiment, a user (1) may wear, attach, and / or secure an electronic device (111) (e.g., an electronic device (200) of FIG. 2, an electronic device (501) of FIG. 5, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) and at least one external electronic device (131, 133) (e.g., an external electronic device (503) of FIG. 5, an external electronic device (903) of FIGS. 9A to 9C, an external electronic device (1003) of FIGS. 10A to 10C, an electronic device (1501, 1502, 1504) of FIG. 15, or an electronic device (1700) of FIG. 17) at a location corresponding to a body part. According to one embodiment, the electronic device (111) may include various devices having a camera. For example, the electronic device (111) may include an attachable device that is attached to a location corresponding to a body part of the user (1) (e.g., clothes, a bag, or the body of the user (1)). The electronic device (111) may include a wearable device that is worn on the body of the user (1) (e.g., but not limited to, a smart watch, a smart ring, a smart bracelet, a smart band, a smart belt, smart shoes, smart glasses, or a smart necklace). The electronic device (111) may include a mobile terminal including a smartphone, and may include, for example, a foldable terminal that can be fixed in a folded state in a pocket of the clothes of the user (1).
[0027] For example, the electronic device (111) may include a camera. According to one embodiment, the electronic device (111) may be worn, attached, and / or fixed at a location corresponding to a body part of the user (1) so that the camera is positioned in an external direction away from the body of the user (1). For example, the electronic device (111) may use the camera to obtain an image including an object in an external direction (e.g., in a frontal direction and / or a side direction of the user (1)).
[0028] Hereinafter, an attachment device is described as an example of an electronic device (111), but the electronic device (111) is not limited to an attachment device.
[0029] According to one embodiment, at least one external electronic device (131, 133) may be worn, attached, and / or fixed at a location corresponding to a body part (e.g., head and / or ear) of the user (1) (e.g., a different location from the electronic device (111). For example, the at least one external electronic device (131, 133) may include, but is not limited to, a hearable device (131) (e.g., earphones), a headset, smart glasses (e.g., AR glasses), a video see through (VST) device, an optical see through (OST) device, or a head mounted device (HMD). For example, the at least one external electronic device (131, 133) may include a wireless earphone, and may include, for example, an earphone device including a true wireless stereo (TWS) device, an open wireless stereo (OWS) device, and / or a neck band.
[0030] For example, the external electronic device (111) may not include a camera, but is not limited thereto. For example, the external electronic device (111) may be worn on the body (e.g., head or ear) of the user (1) and may detect the direction of the user's (1) head and / or gaze direction using at least one sensor and / or microphone.
[0031] According to one embodiment, the electronic device (111) may determine an initial alignment relationship by comparing the postures with each of at least one external electronic device (131, 133). According to one embodiment, the electronic device (111) may determine a region of interest (ROI) within an image acquired through a camera based at least in part on the initial alignment relationship. For example, the electronic device (111) may detect a head or gaze direction of the user (1) based on posture information of at least one external electronic device (131, 133), and may detect a portion of the image corresponding to the head or gaze direction of the user (1) as an ROI. According to one embodiment, the electronic device (111) may compare the posture of the electronic device (111) and the posture of the at least one external electronic device (131, 133) based on the initial alignment relationship, and may recognize that the direction (e.g., camera direction) of the electronic device (111) has been unintentionally changed by the user (1), and may determine or adjust the ROI within the image based on the recognition result. According to one embodiment, the electronic device (111) can detect objects within the ROI, identify objects corresponding to user input among the detected objects, and generate and provide a response to the user input.
[0032] According to one embodiment, the electronic device (111) may additionally link with a plurality of external devices (151 to 158) in addition to at least one external electronic device (131, 133) to identify an object or generate a response to a user input. For example, the electronic device (111) may determine an ROI within an image, identify an object, and / or generate a response corresponding to a user input based on information (e.g., posture information, movement information, location information, gesture information, biometric information, and / or sensor information) received from the plurality of external devices (151 to 158). For example, the plurality of external devices may include a hearable device (131) (e.g., true wireless stereo (TWS) earphones or headset), smart glasses (133) (e.g., extended reality (XR), mixed reality (MR), augmented reality (AR), virtual reality (VR), video see through (VST), or optical see through (OST) device), a smart watch (151), a smart ring (152), a smart necklace (153), a smart band (154), a smart belt (155), a mobile terminal (156), smart shoes (157), or an attachable device (158) (e.g., a digital tattoo), but are not limited to those illustrated or described in FIG. 1.
[0033] Hereinafter, operations of an electronic device (111) determining an ROI within an image, detecting an object within the ROI, and generating and providing a response to a user input based on the detected object will be described in more detail according to various embodiments of the present disclosure.
[0034]
[0035] FIG. 2 is a block diagram of an electronic device according to one embodiment.
[0036] According to one embodiment, an electronic device (200) (e.g., electronic device (111) of FIG. 1, electronic device (501) of FIG. 5, electronic device (901) of FIGS. 9A to 9C, electronic device (1001) of FIGS. 10A to 10C, electronic device (1501) of FIG. 15, or electronic device (1700) of FIG. 17) may include a camera (210) (e.g., camera module (1580) of FIG. 15), a communication circuit (220) (e.g., communication module (1590) of FIG. 15), at least one sensor (230) (e.g., sensor module (1576) of FIG. 15), a memory (240) (e.g., memory (1530) of FIG. 15), and at least one processor (250) (e.g., processor (1520) of FIG. 15).
[0037] According to one embodiment, the camera (210) can acquire an image including a plurality of objects outside the electronic device (200). For example, the plurality of objects may include any object existing outside the electronic device (200). For example, when the electronic device (200) is worn, attached, or fixed at a first location corresponding to a part of the user's body, the camera (210) may be positioned facing the front direction and / or the side direction of the user. For example, the camera (210) may capture an image of the outside of the electronic device (200) including a direction other than the direction in which the user is (e.g., the front direction and / or the side direction of the user). For example, the camera (210) may capture a still image and / or a moving image. The camera (210) may capture an external image according to a specified cycle or continuously, or may capture an external image according to a specified condition (e.g., a user input).
[0038] According to one embodiment, the communication circuit (220) can transmit and receive information and / or data with at least one external electronic device (e.g., the external electronic device (131, 133) of FIG. 1, the external electronic device (503) of FIG. 5, the external electronic device (903) of FIGS. 9A to 9C, the external electronic device (1003) of FIGS. 10A to 10C, the electronic devices (1501, 1502, 1504) of FIG. 15, or the electronic device (1700) of FIG. 17). For example, the communication circuit (220) can receive information related to a posture of the external electronic device from the external electronic device. For example, the communication circuit (220) can receive information related to a movement of the user (e.g., a direction of the user's head) and / or a direction of the user's gaze from the external electronic device.
[0039] According to one embodiment, at least one sensor (230) can detect a posture of the electronic device (200). For example, the at least one sensor (230) can include, but is not limited to, an inertial sensor (230) (e.g., a gyroscope sensor (230), an acceleration sensor (230), and / or a geomagnetic sensor (230)), and / or a microphone. For example, the at least one sensor (230) can detect a posture in which the electronic device (200) is worn and / or attached to a user, and / or a posture in which the electronic device (200) is fixed at a position corresponding to a body part of the user.
[0040] According to one embodiment, the memory (240) may store instructions that control the operation of the electronic device (200) when executed individually or collectively by at least one processor (250). According to one embodiment, the instructions may be stored in one memory (240) or in a plurality of memories (240). The memory (240) may at least temporarily store information and / or data related to the operations of the electronic device (200). For example, the memory (240) may at least temporarily store an image acquired through the camera (210). The memory (240) may at least temporarily store information related to the posture of the electronic device (200) and / or an external electronic device. The memory (240) may store information regarding an initial alignment relationship between the electronic device (200) and the external electronic device. According to one embodiment, the memory (240) can store at least one artificial intelligence model (e.g., an artificial intelligence-based vision model (e.g., a large vision model (LVM)) and / or an artificial intelligence-based multimodal model (e.g., a large multimodal model (LMM)). According to various embodiments, the at least one artificial intelligence model may be implemented as a separate hardware module (e.g., an artificial intelligence processor (250), a chip, or a circuit).
[0041] According to one embodiment, at least one processor (250) can individually or collectively control the operations of the electronic device (200) by executing instructions stored in the memory (240). For example, operations described as being performed by the 'processor (250)' in the present disclosure can be understood as being performed by at least one processor (250) individually or collectively. For example, at least one processor (250) can independently or collectively control each of the operations of the electronic device (200) described below. According to one embodiment, at least one processor (250) can include a circuit such as a central processing unit (CPU), a microprocessor unit (MPU), an application processor (AP), a communication processor (CP), a system on chip (SoC), and / or an integrated circuit (IC).
[0042] According to one embodiment, the processor (250) may detect the degree of movement of the user using at least one sensor (230) while the electronic device (200) is worn, attached, or fixed at a first location corresponding to a part of the user's body. For example, the first location may be a part of the user's body, or an object corresponding to a part of the user's body (e.g., clothes worn by the user, a bag, an accessory, etc.).
[0043] In one embodiment, the processor (250) may perform actions related to initial posture alignment in response to recognizing a specified trigger condition. For example, the specified trigger condition may include recognizing that the electronic device (200) is connected to the external electronic device, that the user is wearing, attaching, or securing the electronic device (200) and / or the external electronic device, and / or receiving a specified user input.
[0044] According to one embodiment, the processor (250) may detect a first initial posture of the electronic device (200) using at least one sensor (230) when the user's movement is less than or equal to a specified threshold. For example, when the user's movement exceeds the threshold, it may be difficult to measure an accurate posture of the electronic device (200) and / or an external electronic device, or an error may occur in the posture measurement. The processor (250) may detect the first initial posture of the electronic device (200) when the user's movement is less than or equal to a specified threshold, thereby increasing the accuracy of the initial posture measurement. According to one embodiment, the processor (250) may recognize the user's ground contact time (GCT) using at least one sensor (230). The processor (250) may detect the first initial posture of the electronic device (200) based on the ground contact time. For example, when a user is walking, the user's movement may be relatively more stable and less variable during the time the user is in ground contact (e.g., the time the user's feet are in contact with the ground) than during the time the user is not in contact with the ground. The processor (250) may detect a first initial posture in the GCT to increase the accuracy of the initial posture detection.
[0045] According to one embodiment, the processor (250) may receive information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user. For example, the processor (250) may receive information related to a second initial posture of the external electronic device measured by at least one sensor (230) of the external electronic device from the external electronic device. According to one embodiment, the second location may be different from the first location. For example, the external processor (250) may be worn, attached, or fixed at a location capable of detecting the movement of the user's head. According to one embodiment, the processor (250) may receive information related to the second initial posture of the external electronic device based on the user's ground contact time (GCT) from the external electronic device. For example, the processor (250) may receive information related to the second initial posture of the external electronic device measured by the GCT by the external electronic device. According to one embodiment, the measurement times of the first initial posture and the second initial posture may correspond to each other.
[0046] According to one embodiment, the processor (250) may determine an initial alignment relationship between a first initial posture of the electronic device (200) and a second initial posture of the external electronic device. For example, the processor (250) may determine the initial alignment relationship based on a difference between the first initial posture and the second initial posture. For example, the initial alignment relationship may be information indicating a posture of the external electronic device based on the posture of the electronic device (200), or may be information indicating a posture of the electronic device (200) based on the posture of the external electronic device. According to one embodiment, the processor (250) may map and store the initial alignment relationship with the first initial posture and the second initial posture.
[0047] According to one embodiment, the processor (250) may determine an initial alignment relationship by comparing a first initial posture and a second initial posture corresponding to the same point in time. According to one embodiment, the processor (250) may acquire an external image corresponding to the point in time at which the first initial posture and the second initial posture are detected using the camera (210), and may map and store the first initial posture, the second initial posture, the initial alignment relationship, and / or the external image. For example, the processor (250) may identify an area of an external image captured by the camera (210) under the first initial posture, the second initial posture, and the initial alignment relationship.
[0048] According to one embodiment, the processor (250) may synchronize time information of the electronic device (200) and the external electronic device when determining the initial alignment relationship. For example, the processor (250) may recognize a delay time due to data (and / or information) transmission between the electronic device (200) and the external electronic device, and set a value for compensating for the delay time.
[0049] According to one embodiment, the processor (250) may acquire an image including a plurality of objects outside the electronic device (200) through the camera (210). For example, the plurality of objects may include any object outside the electronic device (200). For example, the processor (250) may acquire a plurality of images. For example, the electronic device may acquire a plurality of still images or a video (e.g., a plurality of video frames).
[0050] According to one embodiment, the processor (250) can detect the first posture of the electronic device (200) using at least one sensor (230). For example, the processor (250) can recognize whether the posture of the electronic device (200) has changed from the first initial posture. According to one embodiment, the processor (250) can detect the first posture of the electronic device (200) corresponding to the time at which an image is acquired through the camera (210).
[0051] According to one embodiment, the processor (250) may receive information related to a second posture of the external electronic device from the external electronic device. For example, the processor (250) may detect whether the posture of the external electronic device has changed from the second initial posture. According to one embodiment, the processor (250) may recognize a direction in which the user is looking (e.g., the direction of the user's head and / or the direction of the user's gaze) based on the information related to the second posture. According to one embodiment, the information related to the second posture may include information related to the direction in which the user is looking (e.g., the direction of the user's head and / or the direction of the user's gaze). According to one embodiment, the processor (250) may receive information related to a second posture of the external electronic device corresponding to a time point at which an image is acquired through the camera (210). For example, the time points at which the first posture and the second posture are measured may correspond to each other.
[0052] According to one embodiment, the processor (250) may determine a region of interest (ROI) of the user within an image based on at least a portion of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship. For example, the processor (250) may recognize the degree to which the relationship between the first posture and the second posture has changed compared to the initial alignment relationship. The processor (250) may recognize the degree to which the electronic device (200) and / or the external electronic device has changed from the first initial posture and / or the second initial posture based on the first posture and / or the second posture. The processor (250) may recognize an external region of current interest of the user based on at least a portion of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship. For example, the processor (250) may determine an ROI for each of the plurality of images. For example, the ROIs within the plurality of images may be determined to be the same region or different regions.
[0053] For example, assume that the electronic device (200) is an attachable device and the external electronic device is an earphone. When the user attaches the attachable device to, for example, clothing and uses the camera (210) of the attachable device to capture the front of the user, the first posture of the attachable device may change differently from the first initial posture depending on the movement of the user (e.g., the user's clothing). In this case, the second posture of the earphone worn on the user's ear may remain relatively unchanged from the second initial posture. Based on the initial alignment relationship and the second posture, the attachable device can detect that only the posture of the attachable device has changed to a posture different from the user's intention, and adjust the user's region of interest within the captured image. For example, when the user turns his head in a different direction, the second posture of the earphone may change from the second initial posture, and the first posture of the attachable device may remain relatively unchanged from the first initial posture. The processor (250) can recognize that only the direction of the user's head or gaze has moved based on the second posture and the initial alignment relationship of the earphone, and adjust the user's region of interest within the image to an area corresponding to the direction of the user's head or gaze. For example, if a user is moving while looking in the same direction, the first posture of the attached device and the second posture of the earphones may change relatively little or not at all from the first initial posture and the second initial posture, and the relationship between the first posture and the second posture may correspond to the initial alignment posture. In this case, even if the image acquired by the attached device changes according to the movement of the user, the user's region of interest in the newly acquired image may be determined (or maintained) as the region corresponding to the region of interest in the previously acquired image.
[0054] According to one embodiment, the processor (250) can identify an external object based on the determined ROI, and generate and provide a response corresponding to a user input including a query based on the identified object.
[0055] According to one embodiment, the processor (250) may identify at least one object included in a region of interest (ROI) within an image acquired through the camera (210). For example, the at least one object may include any object existing outside the electronic device (200). According to one embodiment, the processor (250) may identify the object within the ROI using an artificial intelligence-based vision model (e.g., a large vision model (LVM). For example, the processor (250) may identify at least one object included within a plurality of images. According to various embodiments, the processor (250) may identify an object outside the ROI based on a user input (e.g., a user utterance).
[0056] According to one embodiment, the processor (250) may determine the user's interest in each of the at least one object based on at least some of the following: a location of each of the at least one object within the region of interest (e.g., a distance from the center of the ROI), a user preference, a history of operation of the electronic device (200) and / or an external electronic device (e.g., a history of the user's use of functions of the electronic device (200), or a state of the electronic device (200) and / or the external electronic device. According to one embodiment, the processor (250) may estimate the user's interest in at least one object using an artificial intelligence-based vision model (e.g., a large vision model (LVM). For example, the processor (250) may provide the vision model with at least some of the positions of at least one object within the region of interest (e.g., the distance from the center of the ROI), user preferences, operation history of the electronic device (200) and / or an external electronic device (e.g., the user's history of using functions of the electronic device (200)), or states of the electronic device (200) and / or the external electronic device as inputs, and may obtain the interest in objects within the ROI as an output of the vision model. For example, the processor (250) may determine the interest in objects within a single image, or may determine the interest in objects within a plurality of images. According to one embodiment, when the processor (250) receives a user input, the processor (250) may determine and / or change the interest in objects based on the user input. For example, the processor (250) may extract user intention from the user input, and You can change the level of interest in objects based on intent.
[0057] According to one embodiment, the processor (250) may receive user input corresponding to a query related to an image. For example, the processor (250) may receive user input corresponding to the query related to the image through at least one of various input methods including voice, gesture, or touch from the user. For example, the processor (250) may recognize the user intent contained in the user input.
[0058] According to one embodiment, the processor (250) may specify an object related to a query among at least one object based on interest. For example, the processor (250) may specify at least one object in order of increasing interest. For example, the processor (250) may specify at least one object without considering the ROI and / or interest based on the user intent included in the user input. For example, if the processor (250) can specify at least one object through the user intent included in the user input, the processor (250) may specify at least one object without using the ROI and / or interest, or may specify at least one object using the ROI and / or interest. For example, the processor (250) may specify at least one object from a plurality of images. For example, the processor (250) may specify at least one object from at least one of the plurality of images acquired before the reception of the user input is completed. For example, when capturing multiple images (e.g., sequential images, videos) through a camera, the electronic device can acquire multiple images until the reception of user input is completed, and can also acquire multiple images while receiving user input. The processor (250) can identify at least one image from multiple images acquired during a specified period of time, rather than identifying an object from a single image at a specific point in time.
[0059] In one embodiment, the processor (250) may provide feedback prompting additional user input to identify an object when there are multiple objects with an interest level greater than or equal to a specified value, multiple objects with the same interest level, the interest level of an object cannot be estimated, or there are no objects with an interest level greater than or equal to the specified value. In response to the feedback, the processor (250) may receive additional user input. Based on the additional user input, the processor (250) may identify an object associated with the user input (e.g., a query).
[0060] According to one embodiment, the processor (250) may generate a response to a query based on a specified object and user input. For example, the processor (250) may generate a response including information related to the specified object. According to one embodiment, the processor (250) may generate a response to a query using an artificial intelligence-based multimodal model (e.g., a large multimodal model (LMM). For example, the processor (250) may provide information related to the specified object (e.g., a name of the specified object and / or an image of the specified object) and user input (e.g., query content) as inputs to the multimodal model, and may obtain a response to the query as an output of the multimodal model.
[0061] In one embodiment, the processor (250) may provide the generated response. For example, the processor (250) may display the response through a display of the electronic device (200) and / or an external electronic device, or output the response through a speaker of the electronic device (200) and / or an external electronic device.
[0062] According to various embodiments, the configuration of the electronic device (200) is not limited to that illustrated in FIG. 2, and at least some components may be omitted, or at least some components (e.g., at least one of the components of FIG. 5 or FIGS. 15 to 17) may be added.
[0063]
[0064] FIG. 3 is a drawing for explaining the operation of an electronic device according to one embodiment.
[0065] According to one embodiment, in operation 305, an electronic device (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (501) of FIG. 5, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) (e.g., an attachable device) may determine whether the electronic device is moving (e.g., a user's movement) and the user's walking situation based on sensor data of the electronic device. For example, the electronic device may detect whether the user is moving or walking using inertial sensor data. According to one embodiment, in order to increase the accuracy of posture detection, if the degree of the detected movement is below a threshold value or the user is not walking (including, for example, the user's ground contact time (GCT)), the electronic device may perform operation 310.
[0066] According to one embodiment, the electronic device may perform operation 305 when the electronic device and an external electronic device (e.g., the external electronic device (131, 133) of FIG. 1, the external electronic device (503) of FIG. 5, the external electronic device (903) of FIGS. 9A to 9C, the external electronic device (1003) of FIGS. 10A to 10C, the electronic devices (1501, 1502, 1504) of FIG. 15, or the electronic device (1700) of FIG. 17) are connected, when a specified user input is received, and / or when the electronic device detects that a user is attaching or wearing the electronic device and / or an external electronic device (e.g., earphones or smart glasses).
[0067] In one embodiment, in operation 310, the electronic device may detect a posture of the electronic device. For example, the posture of the electronic device may indicate a posture in which the electronic device is worn or attached to a location corresponding to a part of the user's body.
[0068] In one embodiment, in operation 315, the electronic device may detect the posture of the external electronic device. For example, the electronic device may receive information related to the posture of the external electronic device from the external electronic device. For example, the posture of the external electronic device may indicate a posture in which the external electronic device is worn or attached to a location corresponding to a part of the user's body.
[0069] In one embodiment, in operation 320, the electronic device may determine an initial alignment relationship between the electronic device and an external electronic device. For example, the initial alignment relationship may include information about a difference between a first initial posture of the electronic device and a second initial posture of the external electronic device.
[0070] According to one embodiment, in operation 325, the electronic device may determine a region of interest (ROI) of the user based on image data of the electronic device, a posture of the electronic device, a posture of an external electronic device, and / or an initial alignment relationship. For example, the electronic device may generate image data by capturing an exterior of the electronic device using a camera. For example, the image data may include at least one image. For example, the electronic device may recognize a direction in which the user is looking based on the posture of the electronic device, the posture of the external electronic device, and / or an initial alignment relationship, and determine an ROI corresponding to the direction in which the user is looking within the image data.
[0071] According to one embodiment, in operation 330, the electronic device may provide ROI-related information and image data as inputs to a vision AI (e.g., LVM) (hereinafter, referred to as an 'artificial intelligence-based vision model' in the present disclosure).
[0072] According to one embodiment, in operation 335, the electronic device may use vision AI to identify objects within a region of interest (ROI) and estimate interest in the objects. For example, the electronic device may obtain information about the identified objects within the ROI and / or interest in the identified objects as output from the vision AI.
[0073] According to one embodiment, in operation 340, the electronic device may provide information about the identified object, interest in the object, and a user query as inputs to a multimodal AI (e.g., LMM) (hereinafter referred to as an “AI-based multimodal model” in the present disclosure). For example, the user query may be obtained through various types of user input, including voice input, gesture input, and / or touch input.
[0074] In one embodiment, in operation 345, the electronic device may generate a response to a user query using multimodal AI. For example, the electronic device may obtain a response to a user query as an output of the multimodal AI.
[0075] In one embodiment, in operation 350, the electronic device may provide the user with visual, auditory, and / or tactile feedback corresponding to the response. For example, the feedback may include the generated response.
[0076] According to one embodiment, the electronic device can acquire an image in a direction the user is looking, or determine a region of interest corresponding to the direction the user is looking, even when the electronic device's posture (e.g., the direction of the electronic device's camera) is not facing the user's front direction. The electronic device can identify objects within the determined region of interest and generate a response to a user input (query) based on the identified objects, thereby improving the accuracy of identifying objects according to the user's intention and the accuracy of providing a response to a user input (query) that matches the user's intention.
[0077]
[0078] FIG. 4 is a diagram for explaining the operation of an electronic device according to one embodiment. For example, FIG. 4 illustrates an operation in which an electronic device (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (501) of FIG. 5, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) identifies an object and estimates a degree of interest using an artificial intelligence-based vision model. Hereinafter, descriptions overlapping with those of FIG. 3 will be briefly described or omitted.
[0079] According to one embodiment, in operation 425, the electronic device may acquire image data by capturing the outside of the electronic device using a camera. The electronic device may acquire sensor data (e.g., information related to the attitude of the electronic device) using at least one sensor. The electronic device may receive sensor data (e.g., information related to the attitude of the external electronic device) of the external electronic device from an external electronic device (e.g., the external electronic device (131, 133) of FIG. 1 , the external electronic device (503) of FIG. 5 , the external electronic device (903) of FIGS. 9A to 9C , the external electronic device (1003) of FIGS. 10A to 10C , the electronic devices (1501, 1502, 1504) of FIG. 15 , or the electronic device (1700) of FIG. 17 ). The electronic device may determine an initial alignment relationship between the electronic device and the external electronic device. The electronic device may determine a region of interest within the image data based on the image data, the sensor data of the electronic device, and the sensor data of the external electronic device.
[0080] In one embodiment, in operation 430, the electronic device may provide ROI-related information and image data as inputs to the vision AI. In one embodiment, the vision model may include a large-scale vision model including a vision transformer and ConvNeXt. In one embodiment, the vision AI may simultaneously perform object identification and object interest estimation through a multi-task learning method.
[0081] In one embodiment, in operation 431, an electronic device (e.g., vision AI) may perform annotation on image data. For example, the electronic device may extract a region of interest (ROI) from the image data, and bind and segment objects within the ROI.
[0082] In one embodiment, in operation 432, the electronic device (e.g., vision AI) may perform preprocessing. For example, the electronic device may scale and normalize the annotated image data.
[0083] In one embodiment, in operation 433, an electronic device (e.g., vision AI) may process data via a transformer. For example, the electronic device may perform operations to identify objects within a region of interest (ROI) and estimate their interest based on preprocessed image data.
[0084] In one embodiment, in operation 435, the electronic device may obtain the result of identifying an object within the ROI and the degree of interest in the identified object. For example, the electronic device may obtain the result of identifying an object within the ROI and the degree of interest in the identified object as an output of the vision AI.
[0085] According to various embodiments, at least some of the operations of FIGS. 3 and 4 may be omitted, or at least some operations may be added, and the operations of FIGS. 3 and 4 may be performed independently of each other, or at least some of the operations may be performed in an integrated manner.
[0086]
[0087] FIG. 5 is a diagram showing the configuration of an electronic device and an external electronic device according to one embodiment. For example, FIG. 5 shows the software configuration of an electronic device (501) (e.g., the electronic device (111) of FIG. 1 , the electronic device (200) of FIG. 2 , the electronic device (901) of FIGS. 9A to 9C , the electronic device (1001) of FIGS. 10A to 10C , the electronic device (1501) of FIG. 15 , or the electronic device (1700) of FIG. 17 ) and an external electronic device (503) (e.g., the external electronic device (131, 133) of FIG. 1 , the external electronic device (903) of FIGS. 9A to 9C , the external electronic device (1003) of FIGS. 10A to 10C , the electronic devices (1501, 1502, 1504) of FIG. 15 , or the electronic device (1700) of FIG. 17 ).
[0088] According to one embodiment, the electronic device (501) may include an application (510), a framework (520), and a library (530).
[0089] According to one embodiment, the application (510) may include at least one application (510). For example, FIG. 5 illustrates a health application (511) and a wearable application (513), but the present invention is not limited thereto. For example, the health application (511) may manage a user's biometric information and / or health information. For example, the wearable application (513) may control linkage with an external electronic device (503) (which is not limited to an external wearable device and may include an external attachable device or an external mobile terminal). For example, the wearable application (513) may control connection with the external electronic device (503) and control transmission and reception of data (or information) with the external electronic device (503).
[0090] According to one embodiment, the framework (520) may include a sensor manager (521), a location manager (523), a wearable manager (525), and a context manager (527).
[0091] According to one embodiment, the sensor manager (521) can control the operation of at least one sensor included in the electronic device (501) and manage measurement data through the sensor. For example, the sensor manager (521) can detect the posture of the electronic device (501) based on data measured using at least one sensor. For example, the sensor manager (521) can obtain external image data of the electronic device (501) through an image sensor (e.g., a camera).
[0092] According to one embodiment, the location manager (523) can measure and / or manage location information of the electronic device (501).
[0093] According to one embodiment, the wearable manager (525) can process and / or manage data of an external electronic device (503) acquired through a wearable application (513). For example, the wearable manager (525) can transmit information received from an external electronic device (503) acquired through a wearable application (513) to a context manager (527).
[0094] According to one embodiment, the context manager (527) may execute a context-aware algorithm based on information received through the wearable manager (525). For example, the context manager (527) may recognize the posture of the external electronic device (503), the direction of the user's head, the direction of the user's gaze, and / or the state of the external electronic device (503) based on information received from the external electronic device (503) through the context-aware algorithm.
[0095] According to one embodiment, the library (530) may store information related to services, functions, and / or operations of the electronic device (501) provided by the electronic device (501). According to one embodiment, the library (530) may include an AI library (531). For example, the AI library (531) may include at least one AI model (e.g., a vision model, a multimodal model, a language model, a machine learning model, and / or a deep learning model). For example, the at least one AI model may identify an object from image data and estimate a user's interest in the identified object. For example, the at least one AI model may generate a response to a user input (a user query).
[0096] According to one embodiment, the external electronic device (503) may include a network manager (540), a system manager (550), a service manager (560), a sensor algorithm (570), and an AI library (580).
[0097] In one embodiment, the network manager (540) can control data transmission and reception with the electronic device (501). For example, the network manager (540) can transmit sensor data of an external electronic device (503) (e.g., information related to the posture of the external electronic device (503)) to the electronic device (501).
[0098] According to one embodiment, the system manager (550) can control the overall operation of the external electronic device (503).
[0099] According to one embodiment, the service manager (560) can manage services provided by an external electronic device (503).
[0100] According to one embodiment, the sensor algorithm (570) can extract features to be provided as inputs to the AI library (580) from sensor data collected from the sensor. For example, the sensor algorithm (570) can provide the extracted features as inputs to at least one AI model included in the AI library (580). According to one embodiment, the sensor algorithm (570) can also provide sensor raw data as inputs to the AI library (580) (AI model). The sensor algorithm (570) can control the operation (e.g., sensor operation) of the electronic device (501) based on the output of the AI model received from the AI model.
[0101] In one embodiment, the AI library (580) may include at least one AI model. For example, the AI model may be trained online or offline, depending on the purpose. The output of the AI model may be transmitted to the sensor algorithm (570).
[0102]
[0103] FIG. 6 is a diagram illustrating the structure of a multimodal model according to one embodiment.
[0104] According to one embodiment, the multimodal model may include a feature extraction (620), an encoder (630) (e.g., a transformer), and a dense layer (640). Below, a description of the general operations and functions of the general components (respective layers) (620, 630, 640) that constitute the multimodal model that can be understood by those skilled in the art is omitted, and the operations of each component are not limited to what is described below.
[0105] According to one embodiment, various types of inputs (e.g., audio input (611), gesture input (613), visual input (615)) can be input to a multimodal model. For example, each of the feature extraction units (621, 623, 625) can extract features from each of the inputs (611, 613, 615) and provide the extracted features to an encoder (transformer) (630).
[0106] According to one embodiment, the encoder (630) can encode inputs (e.g., input data) (6111, 613, 615). For example, the encoder (630) can embed the inputs and perform positional encoding. For example, the encoder (630) can include positional information of the input values in the embedded values. The encoder (630) can evaluate the inputs. For example, the encoder (630) can evaluate the data values through attention (e.g., multi-head attention). The encoder (630) can perform summation and normalization based on the values on which attention has been performed and output the result value.
[0107] In one embodiment, a dense layer (640) (e.g., a fully connected layer) can connect both inputs and outputs of a multimodal model. For example, the dense layer (640) can connect each neuron (node) between the input and output of the multimodal model with a weight. For example, the dense layer (640) can connect the input and output through a linear or nonlinear function.
[0108] According to one embodiment, an electronic device (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (501) of FIG. 5, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) can receive a user input corresponding to a query related to an image. The electronic device can generate a response to the query using a multimodal model. For example, the electronic device can provide information about an object specified from a user input (query) and / or a ROI of an image as inputs to the multimodal model, and obtain a response to the query as an output (650) of the multimodal model.
[0109] According to various embodiments, the multimodal model that the electronic device uses to generate a response to user input is not limited to that illustrated in FIG. 6, and various artificial intelligence-based models may be used.
[0110]
[0111] FIG. 7 is a diagram showing a rotation matrix according to one embodiment.
[0112] According to one embodiment, the electronic device can detect the posture of the electronic device (e.g., the electronic device (111) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (501) of FIG. 5, the electronic device (901) of FIGS. 9A to 9C, the electronic device (1001) of FIGS. 10A to 10C, the electronic device (1501) of FIG. 15, or the electronic device (1700) of FIG. 17) and the posture of an external electronic device (e.g., the external electronic device (131, 133) of FIG. 1, the external electronic device (503) of FIG. 5, the external electronic device (903) of FIGS. 9A to 9C, the external electronic device (1003) of FIGS. 10A to 10C, the electronic devices (1501, 1502, 1504) of FIG. 15, or the electronic device (1700) of FIG. 17). The electronic device may generate a rotation matrix representing a difference value between the posture of the electronic device and the posture of the external electronic device based on the determinant illustrated in FIG. 7. For example, the rotation matrix may represent an initial alignment relationship between the electronic device and the external electronic device (i.e., a difference value between the initial posture of the electronic device and the initial posture of the external electronic device) and / or a difference value between the current posture of the electronic device and the current external electronic device. According to various embodiments, the electronic device may generate a posture of the external electronic device relative to the posture of the electronic device as a rotation matrix, or may generate a posture of the electronic device relative to the posture of the external electronic device as a rotation matrix.
[0113] For example, an electronic device can recognize the attitude of the electronic device and / or an external electronic device in three axes (e.g., x-axis, y-axis, z-axis). For example, the electronic device can calculate roll, pitch, and yaw for each axis using the angular values formed by the attitude of the electronic device and the attitude of the external electronic device with respect to three mutually orthogonal axes, and generate a rotation matrix based on the same.
[0114] For example, referring to FIG. 7, φ, θ, and ψ can represent angles between the electronic device and the external electronic device with respect to three orthogonal axes (e.g., x-axis, y-axis, z-axis), respectively. For example, Rx(φ) represents roll, Ry(θ) represents pitch, and Rz(ψ) represents yaw. The electronic device can generate a rotation matrix R using Rx(φ), Ry(θ), and Rz(ψ).
[0115] According to one embodiment, the electronic device can identify a difference in pose between the electronic device and an external electronic device using a rotation matrix, and determine or change a ROI within the image based on the difference in the identified pose, thereby increasing the accuracy of identifying and / or specifying an object corresponding to a user input (i.e., matching the user intent) within the image.
[0116]
[0117] FIG. 8 is a diagram for explaining an operation of an electronic device recognizing objects in an image according to one embodiment. For example, FIG. 8 shows a result of identifying objects (820) within an ROI determined by an electronic device (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (501) of FIG. 5, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17). FIG. 8 is merely an example, and the objects (820) identified by the electronic device are not limited to those illustrated in FIG. 8.
[0118] According to one embodiment, the electronic device can determine a region of interest (ROI) of a user within an image acquired through a camera. The electronic device can identify at least one object (820) included within the ROI. Referring to FIG. 8, the electronic device can identify a monitor (821), a desk (822), a trash can (823), a drawer (824), a board (825), a organizer (826), books (827), a window (828), and a wall (829) within the ROI of the image.
[0119] According to one embodiment, the electronic device may identify or provide information (e.g., a room) (810) related to a location (e.g., a ROI) appearing in an image.
[0120] According to one embodiment, the electronic device may estimate the user's interest in each of the identified objects (820). According to one embodiment, the electronic device may use an artificial intelligence-based vision model to identify objects within a ROI and / or estimate the interest in the identified objects (820). For example, the interest in the objects (820) may be determined based on the location of the objects (820) within the ROI (e.g., how close they are to the center of the ROI), the user's head direction, the user's gaze direction, the user's preference, and / or the user's device (or function) usage history.
[0121] For example, if the interest in monitor (821) is estimated to be 85%, the interest in desk (822) is estimated to be 75%, the interest in trash can (823) is estimated to be 15%, the interest in drawer (824) is estimated to be 12%, the interest in board (825) is estimated to be 5%, the interest in organizer (826) is estimated to be 3.5%, and the interest in books (827) is estimated to be 2.5%, the electronic device may provide information about the estimated interest in each object (820). In one embodiment, the electronic device may not provide information about the interest in window (828) and wall (829) if the interest in the corresponding object is estimated to be less than a specified value (e.g., 0%).
[0122] According to one embodiment, the electronic device may identify at least one object corresponding to the user's intent (e.g., a query included in a user input) based on the estimated interest level. For example, the electronic device may identify the monitor (821) with the highest interest level as the object corresponding to the user's intent (user query), or may identify the monitor (821) and the desk (822) with the highest interest level as the objects corresponding to the user's intent (user query).
[0123] According to one embodiment, the electronic device may provide information about the determined ROI, information about the identified objects (820) within the ROI (e.g., names of the objects (820)), and / or interest in each of the objects (820). For example, the electronic device may display information about the identified objects (820) and / or interest in each of the objects (820) in an image corresponding to the ROI.
[0124]
[0125] FIGS. 9A to 9C are diagrams for explaining the operation of an electronic device according to one embodiment. For example, FIGS. 9A to 9C each illustrate examples of an electronic device (901) (e.g., the electronic device (111) of FIG. 1 , the electronic device (200) of FIG. 2 , the electronic device (901) of FIGS. 9A to 9C , the electronic device (1001) of FIGS. 10A to 10C , the electronic device (1501) of FIG. 15 , or the electronic device (1700) of FIG. 17 ) adjusting an ROI.
[0126] According to one embodiment, FIG. 9A illustrates a case where the gaze of the user (1) moves while the user (1) is stationary. The electronic device (901) may acquire an external image (910) and determine an area of interest (o1) of the user (1) within the image (910) based on at least a part of a relationship (e.g., an initial alignment relationship) between the posture of the electronic device (901), the posture of the external electronic device (903), or the postures of the electronic device (901) and the external electronic device (903) (e.g., an initial alignment relationship). For example, the region of interest (o1) may include an area corresponding to the gaze direction (v1) of the user (1). The electronic device (901) may recognize objects within the region of interest (o1).
[0127] For example, when the gaze direction (v1, v2) of the user (1) changes at the bottom of FIG. 9a, the electronic device (901) can determine the area corresponding to the changed gaze direction (v2) of the user (1) as the area of interest (o2) and recognize objects within the area of interest (o2).
[0128] According to one embodiment, FIG. 9b illustrates a case where the posture of the electronic device (901) is changed while the user (1) is stationary.
[0129] For example, in the upper part of FIG. 9b, the electronic device (901) can determine the central area of the image (930) corresponding to the gaze direction (v3) of the user (1) in a stationary state as the area of interest (o3).
[0130] For example, in the case where only the posture of the electronic device (901) is changed without the movement of the user (1) in the lower part of FIG. 9b, the gaze direction (v3) of the user (1) is the same, but since the direction of the camera of the electronic device (901) is changed, the direction (v4) corresponding to the preset ROI of the image (940) captured by the camera may not match the gaze direction (v1) of the user (1). The electronic device (901) may determine the area corresponding to the gaze direction (v3) of the user (1) within the image (940) as the region of interest (o4) based on the posture of the electronic device (901), the posture of the external electronic device (903), and the initial alignment relationship between the electronic device (901) and the external electronic device (903). For example, an area corresponding to substantially the same external objects within the image may be determined as the region of interest (o4) even though the posture of the electronic device (901) has changed. For example, the region of interest (o3) may substantially correspond to the region of interest (o4). The electronic device (901) may identify an object within the determined region of interest (o4).
[0131] According to one embodiment, FIG. 9c illustrates a case where the position of the user (1) is moved (e.g., the user (1) moves while the posture of the electronic device (901) does not change and the gaze direction of the user (1) is also maintained).
[0132] For example, in the upper portion of FIG. 9c, the electronic device (901) can determine a region of interest (o5) within the image (950) based on the posture of the electronic device (901), the posture of the external electronic device (903), and the difference (e.g., initial alignment relationship) between the postures of the electronic device (901) and the external electronic device (903).
[0133] For example, in the lower part of FIG. 9c, the image (950) obtained may change depending on the movement of the user (1) (e.g., movement in the direction of the user's (1) gaze). For example, since the image (950) changes but the gaze direction (v5) of the user (1) and the posture of the electronic device (901) are maintained, the electronic device (901) may maintain the region of interest (o5, o6) within the image (950). For example, even if the region of interest (o5, o6) remains the same, objects identified within the region of interest (o5, o6) may change because the image (950) changes.
[0134] According to one embodiment, when the electronic device (901) receives a user query (e.g., “Read that sign”), it can identify an object corresponding to the user query among recognized objects, and generate and provide a response including information related to the identified object. For example, the electronic device (901) can determine a priority (e.g., the interest of the user (1)) for the recognized objects using an AI model. The electronic device (901) can identify an object with a high priority as the object corresponding to the user query, and generate a response to the query based on information related to the identified object. For example, if a user (1) makes a query saying “Read that signboard”, the ROI (o1, o2, o3, o4, o5, o6) within the image (910, 630, 640, 650) can be determined, the object (signboard) with the highest priority among the recognized objects within the ROI (o1, o2, o3, o4, o5, o6) can be specified, and a response (e.g., “That signboard is a Starbucks”) containing information related to the signboard (e.g., text included in the signboard, etc.) can be generated and provided. According to one embodiment, the electronic device (901) determines a region of interest of the user (1) (e.g., a region toward which the user's (1's) gaze is directed) within an image acquired based on a posture of the electronic device (901), a posture of an external electronic device (903), and a relationship (e.g., an initial alignment relationship) between the posture of the electronic device (901) and the posture of the external electronic device (903), and identifies an object within the determined region of interest to provide a response to the user's query, thereby increasing the accuracy of the response to the user's query.
[0135]
[0136] FIGS. 10A to 10C are drawings for explaining the operation of an electronic device according to one embodiment.
[0137] According to one embodiment, an electronic device (1001) (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) (e.g., an attachable device) is connected to an external electronic device (1003) (e.g., an external electronic device (131, 133) of FIG. 1, an external electronic device (903) of FIGS. 9A to 9C, an external electronic device (1003) of FIGS. 10A to 10C, an electronic device (1501, 1502, 1504) of FIG. 15, or an electronic device (1700) of FIG. 17)) (e.g., an earphone) of a user (1) and / or an external electronic device When the wearing of the device (1003) is detected, the image acquired through the camera, the initial posture of the electronic device (1001), and the initial posture of the external electronic device (1003) can be aligned. For example, the electronic device (1001) can measure the posture of the electronic device (1001) using at least one sensor. The electronic device (1001) can receive information about the posture of the external electronic device (1003) measured using the sensor of the external electronic device (1003) from the external electronic device (1003). For example, the electronic device (1001) can determine the initial alignment relationship based on the difference between the posture of the electronic device (1001) and the posture of the external electronic device (1003).
[0138] For example, the electronic device (1001) can adjust the area of an image captured by the camera based on the posture of the electronic device (1001) and the posture of the external electronic device (1003). For example, the electronic device (1001) can determine the external range captured by the camera based on the postures of the electronic device (1001) and the external electronic device (1003). The electronic device (1001) can map and store the shooting range of the camera (the area of an image captured by the camera) to the initial posture of the electronic device (1001), the initial posture of the external electronic device (1003), and / or the initial alignment relationship.
[0139] For example, referring to FIG. 10A, a user (1) is shown looking straight ahead while wearing an electronic device (1001) and an external electronic device (1003). In this case, the electronic device (1001) can acquire an image of the front direction of the user (1) using the camera of the electronic device (1001).
[0140] Referring to FIG. 10B, the user (1) can detect the head direction or gaze direction of the user (1) based on information received from the external electronic device (1003) (e.g., information related to the posture of the external electronic device (1003). For example, the electronic device (1001) can recognize that the posture of the external electronic device (1003) has changed from the initial posture (a state in which the user (1) is looking straight ahead) to the current posture (a state in which the user (1) is looking to the side), and can recognize that the user (1) is looking to the side. The electronic device (1001) can determine the ROI within the image based on the direction in which the user (1) is looking. According to one embodiment, the electronic device (1001) can obtain an image corresponding to the direction in which the user (1) is looking using a camera.
[0141] Referring to FIG. 10C, the electronic device (1001) may receive a user input including a query from the user (1). For example, the electronic device (1001) may receive a user input including a query related to an external object via voice, gesture, and / or touch. According to one embodiment, the electronic device (1001) may specify an object corresponding to the user input (query) among objects included in a region of interest (ROI). For example, the electronic device (1001) may identify objects within the ROI based on an artificial intelligence-based vision model, estimate the degree of interest in each object, and specify at least one object corresponding to the user input (query). According to one embodiment, the electronic device (1001) may obtain an image of a specified object using a camera. For example, the electronic device (1001) may extract an image of a specified object within a ROI of a captured image, or capture an image including the specified object using a camera. According to one embodiment, the electronic device (1001) may provide feedback prompting additional user (1) effort to specify an object when the electronic device (1001) cannot specify an object corresponding to the query or when more than a specified number of objects are specified, and may specify an object corresponding to the query based on the additional user input.
[0142] The electronic device (1001) can generate a response to a query based on information related to a specific object, user input, the posture of the electronic device (1001), the posture of the external electronic device (1003), and information about the initial alignment relationship. The electronic device (1001) can generate a response to a query using an artificial intelligence-based multimodal model. For example, if the user query is “Find where the ABC store sign is,” the electronic device (1001) can identify the ABC store sign among the identified objects and provide the user (1) with a response including information related to the ABC store.
[0143] In one embodiment, the electronic device (1001) may provide a response generated visually, audibly, and / or tactilely. For example, the electronic device (1001) may output the response via a display of the electronic device (1001) and / or an external electronic device (1003), or via a speaker.
[0144]
[0145] For example, when capturing an external image through an electronic device worn, attached, and / or fixed to a user, if the posture of the electronic device is not maintained and is distorted, it may be difficult to recognize an object within the image according to the user's intention. In this case, the accuracy may decrease when specifying an object corresponding to a user query and generating a response to the user query. According to embodiments of the present disclosure, an ROI that matches a user's intention within an image can be determined based on the posture of the electronic device, the posture of the external electronic device, the posture alignment relationship between the electronic device and the external electronic device, the user's movement, gaze, information related to the electronic device, and / or information related to the external electronic device, and an object that matches the user's intention can be specified. According to embodiments of the present disclosure, by improving the accuracy in specifying an object that matches a user query (user intent), a response suitable for the user's intent included in the user query can be generated and provided.
[0146] An electronic device according to one embodiment may include a camera, a communication circuit, at least one sensor, a memory, and at least one processor. The memory, when executed by the at least one processor, may cause the electronic device to detect the degree of movement of the user using the at least one sensor while the electronic device is worn, attached, or fixed to a first location corresponding to a part of the user's body.
[0147] The instructions, when executed by the at least one processor, may cause the electronic device to detect a first initial pose of the electronic device using the at least one sensor when the movement of the user is below a specified threshold.
[0148] The instructions, when executed by the at least one processor, may cause the electronic device to receive, through the communication circuit, information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user.
[0149] The instructions, when executed by the at least one processor, may cause the electronic device to determine an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device.
[0150] The instructions, when executed by the at least one processor, may cause the electronic device to obtain an image including a plurality of objects outside the electronic device using the camera.
[0151] The above instructions, when executed by the at least one processor, may cause the electronic device to detect a first posture of the electronic device using the at least one sensor.
[0152] The instructions, when executed by the at least one processor, may cause the electronic device to receive information related to a second posture of the external electronic device from the external electronic device.
[0153] The instructions, when executed by the at least one processor, may cause the electronic device to determine a region of interest (ROI) of the user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship.
[0154] The instructions, when executed by the at least one processor, may cause the electronic device to identify at least one object contained within the region of interest.
[0155] The instructions, when executed by the at least one processor, may cause the electronic device to determine a user's interest in each of the at least one object based at least in part on a location of each of the at least one object within the region of interest, a user preference for the at least one object, and a history of operation of the electronic device.
[0156] The above instructions, when executed by the at least one processor, may cause the electronic device to identify the at least one object using an artificial intelligence-based vision model.
[0157] The instructions, when executed by the at least one processor, may cause the electronic device to receive user input corresponding to a query related to the image.
[0158] The instructions, when executed by the at least one processor, may cause the electronic device to specify an object related to the query among the at least one object based on the interest.
[0159] The above instructions, when executed by the at least one processor, may cause the electronic device to generate a response to the query based on the specified object and the user input using an artificial intelligence-based multimodal model.
[0160] The above instructions, when executed by the at least one processor, may cause the electronic device to provide the generated response.
[0161] The instructions, when executed by the at least one processor, may cause the electronic device to provide feedback prompting additional user input to specify an object when there are multiple objects having a degree of interest greater than or equal to the specified value.
[0162] The instructions, when executed by the at least one processor, may cause the electronic device to receive the additional user input in response to the feedback.
[0163] The instructions, when executed by the at least one processor, may cause the electronic device to specify an object related to the query based on the additional user input.
[0164] The instructions, when executed by the at least one processor, may cause the electronic device to detect the first pose while acquiring the image using the camera.
[0165] The instructions, when executed by the at least one processor, may cause the electronic device to receive information related to the second posture of the external electronic device from the external electronic device while acquiring the image.
[0166] The instructions, when executed by the at least one processor, may cause the electronic device to recognize a ground contact time (GCT) of the user using the at least one sensor.
[0167] The instructions, when executed by the at least one processor, may cause the electronic device to detect the first initial pose of the electronic device at the time of ground contact.
[0168] The second initial posture of the external electronic device may correspond to the posture of the external electronic device at the time of ground contact.
[0169] The instructions, when executed by the at least one processor, may cause the electronic device to detect a degree of movement of the user to determine the initial alignment relationship based on the electronic device and the external electronic device being connected via the communication circuit, or receiving a specified user input.
[0170] The above instructions, when executed by the at least one processor, may cause the electronic device to synchronize time information of the electronic device and the external electronic device.
[0171] The above instructions, when executed by the at least one processor, may cause the electronic device to recognize and correct a delay time of information transmitted and received between the electronic device and the external electronic device based on the synchronized time information.
[0172] The second posture of the external electronic device can be changed according to the movement of the user.
[0173] The above instructions, when executed by the at least one processor, may cause the electronic device to recognize at least one of a head movement or a gaze direction of the user based on information related to the second posture.
[0174] The instructions, when executed by the at least one processor, may cause the electronic device to determine the ROI based at least in part on at least one of the user's head movement or gaze direction.
[0175]
[0176] Fig. 11 is a flowchart of a method of operating an electronic device according to one embodiment.
[0177] According to one embodiment, in operation 1110, an electronic device (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) may detect a degree of movement of the user using at least one sensor while the electronic device is worn, attached, or fixed at a first location corresponding to a body part of the user. For example, the first location may be a body part of the user, or an object corresponding to a body part of the user (e.g., clothes worn by the user, a bag, an accessory, etc.).
[0178] According to one embodiment, the electronic device may perform an action related to initial posture alignment (e.g., at least one of actions 1110 to 1140) in response to recognizing a specified trigger condition. For example, the specified trigger condition may include recognizing that the electronic device and an external electronic device (e.g., the external electronic device (131, 133) of FIG. 1 , the external electronic device (903) of FIGS. 9A to 9C , the external electronic device (1003) of FIGS. 10A to 10C , the electronic devices (1501, 1502, 1504) of FIG. 15 , or the electronic device (1700) of FIG. 17 ) are connected, that a user has worn, attached, or secured the electronic device and / or the external electronic device, and / or that a specified user input is received.
[0179] According to one embodiment, in operation 1120, the electronic device may detect a first initial posture of the electronic device using at least one sensor when the user's movement is less than or equal to a specified threshold. For example, when the user's movement exceeds the threshold, it may be difficult to measure the exact posture of the electronic device and / or an external electronic device, or errors may occur in the posture measurement. The electronic device may detect the first initial posture of the electronic device when the user's movement is less than or equal to the specified threshold, thereby increasing the accuracy of the initial posture measurement. According to one embodiment, the electronic device may recognize the user's ground contact time (GCT) using at least one sensor. The electronic device may detect the first initial posture of the electronic device at the ground contact time. For example, when the user is walking, the user's movement may be relatively more stable and less variable during the user's ground contact time (e.g., the time the user's foot is in contact with the ground) than during the time the user is not in contact with the ground. The electronic device may increase the accuracy of the initial posture detection by detecting the first initial posture using the GCT.
[0180] According to one embodiment, in operation 1130, the electronic device may receive information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user. For example, the electronic device may receive information related to a second initial posture of the external electronic device measured by at least one sensor of the external electronic device from the external electronic device. According to one embodiment, the second location may be different from the first location. For example, the external electronic device may be worn, attached, or fixed at a location capable of detecting head movement of the user. According to one embodiment, the electronic device may receive information related to the second initial posture of the external electronic device based on the ground contact time (GCT) of the user from the external electronic device. For example, the electronic device may receive information related to the second initial posture of the external electronic device measured by the GCT of the external electronic device.
[0181] According to one embodiment, in operation 1140, the electronic device may determine an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device. For example, the electronic device may determine the initial alignment relationship based on a difference between the first initial posture and the second initial posture. For example, the initial alignment relationship may be information indicating a posture of the external electronic device based on a posture of the electronic device, or may be information indicating a posture of the electronic device based on a posture of the external electronic device. According to one embodiment, the electronic device may store the initial alignment relationship by mapping it with the first initial posture and the second initial posture.
[0182] According to one embodiment, the electronic device can determine the initial alignment relationship by comparing the first initial pose and the second initial pose corresponding to the same point in time. According to one embodiment, the electronic device can use a camera to acquire an external image corresponding to the point in time at which the first initial pose and the second initial pose are detected, and map and store the first initial pose, the second initial pose, the initial alignment relationship, and / or the external image. For example, the electronic device can identify an area of the external image captured by the camera under the first initial pose, the second initial pose, and the initial alignment relationship.
[0183] In one embodiment, the electronic device may synchronize time information between the electronic device and an external electronic device when determining an initial alignment relationship. For example, the electronic device may recognize a delay in data (and / or information) transmission between the electronic device and the external electronic device and set a value to compensate for the delay.
[0184] According to one embodiment, in operation 1150, the electronic device may obtain an image including a plurality of objects external to the electronic device.
[0185] In one embodiment, at operation 1160, the electronic device may detect a first posture of the electronic device using at least one sensor. For example, the electronic device may recognize whether the posture of the electronic device has changed from the first initial posture. In one embodiment, the electronic device may detect the first posture of the electronic device corresponding to the time at which the image is acquired in operation 1150.
[0186] According to one embodiment, in operation 1170, the electronic device may receive information related to a second posture of the external electronic device from the external electronic device. For example, the electronic device may detect whether the posture of the external electronic device has changed from the second initial posture. According to one embodiment, the electronic device may recognize a direction in which the user is looking (e.g., the direction of the user's head and / or the direction of the user's gaze) based on the information related to the second posture. According to one embodiment, the information related to the second posture may include information related to the direction in which the user is looking (e.g., the direction of the user's head and / or the direction of the user's gaze). According to one embodiment, the electronic device may receive information related to the second posture of the external electronic device corresponding to the time at which the image is acquired in operation 1150.
[0187] According to one embodiment, at operation 1180, the electronic device may determine a region of interest (ROI) of the user within the image based on at least a portion of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship. For example, the electronic device may recognize a degree to which the relationship between the first posture and the second posture has changed compared to the initial alignment relationship. The electronic device may recognize a degree to which the electronic device and / or an external electronic device has changed from the first initial posture and / or the second initial posture based on the first posture and / or the second posture. The electronic device may recognize an external region of interest of the user currently based on at least a portion of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship.
[0188] For example, assume that the electronic device is an attachable device and the external electronic device is an earphone. If the user attaches the attachable device to, for example, their clothing and uses the attachable device's camera to capture the user's front view, the attachable device's first pose may change from the first initial pose depending on the movement of the user (e.g., the user's clothing). In this case, the second pose of the earphones worn on the user's ears may remain relatively unchanged from the second initial pose. Based on the initial alignment relationship and the second pose, the attachable device can detect that only the pose of the attachable device has changed to a pose different from the user's intention, and adjust the user's region of interest within the captured image. For example, if the user turns their head to a different direction, the second pose of the earphones may change from the second initial pose, and the first pose of the attachable device may remain relatively unchanged from the first initial pose. Based on the second pose and the initial alignment relationship of the earphones, the electronic device can recognize that only the direction of the user's head or gaze has moved, and adjust the user's region of interest within the image to an area corresponding to the direction of the user's head or gaze. For example, if a user is moving while looking in the same direction, the first posture of the attached device and the second posture of the earphones may change relatively little or not at all from the first initial posture and the second initial posture, and the relationship between the first posture and the second posture may correspond to the initial alignment posture. In this case, even if the image acquired by the attached device changes according to the movement of the user, the user's region of interest in the newly acquired image may be determined (or maintained) as the region corresponding to the region of interest in the previously acquired image.
[0189] According to one embodiment, the electronic device can identify an external object based on a determined ROI, and generate and provide a response corresponding to a user input including a query based on the identified object. The operations for providing a response based on the determined ROI are described in more detail in FIG. 12 .
[0190]
[0191] FIG. 12 is a flowchart of an operating method of an electronic device according to one embodiment. For example, the operations of FIG. 12 may be performed after the operations of FIG. 11, but are not limited thereto.
[0192] According to one embodiment, in operation 1210, an electronic device (e.g., the electronic device (111) of FIG. 1, the electronic device (200) of FIG. 2, the electronic device (901) of FIGS. 9A to 9C, the electronic device (1001) of FIGS. 10A to 10C, the electronic device (1501) of FIG. 15, or the electronic device (1700) of FIG. 17) may identify at least one object included in a region of interest (ROI) within an image acquired through a camera. For example, the at least one object may include any object existing outside the electronic device. According to one embodiment, the electronic device may identify the object within the ROI using an artificial intelligence-based vision model (e.g., a large vision model (LVM)).
[0193] According to one embodiment, an electronic device may identify at least one object based on a plurality of images acquired through a camera. For example, when the electronic device captures a plurality of images continuously or periodically, or captures a video (i.e., a plurality of consecutive image frames) in real time, the electronic device may identify at least one object included in the plurality of images and / or the video. For example, when acquiring a plurality of images, the electronic device may recognize objects included in each of the plurality of images and track the recognized objects. For example, the electronic device may determine the same or different ROI for each of the plurality of images. For example, in operation 1230, the electronic device may identify at least one object from at least one of the plurality of images acquired during a specified time period before receiving a user input. For example, the specified time period may be from the time when receiving a user input begins (e.g., the time when an utterance corresponding to the user input begins) to the time when receiving the user input ends (e.g., the time when an utterance corresponding to the user input ends), but is not limited thereto. The specified time may be set and / or changed based on the state of the electronic device and / or user input. According to one embodiment, in operation 1220, the electronic device may determine the user's interest in each of the at least one object based on at least some of the position of each of the at least one object within the region of interest (e.g., distance from the center of the ROI), user preference, operation history (e.g., history of use of functions of the electronic device) of the electronic device and / or an external electronic device (e.g., external electronic device 131, 133 of FIG. 1 , external electronic device 903 of FIGS. 9A to 9C , external electronic device 1003 of FIGS. 10A to 10C , electronic devices 1501, 1502, 1504 of FIG. 15 , or electronic device 1700 of FIG. 17 ), or state of the electronic device and / or the external electronic device.According to one embodiment, the electronic device can estimate the user's interest in at least one object using an artificial intelligence-based vision model (e.g., a large vision model (LVM). For example, the electronic device can provide the vision model with at least some of the following as inputs: a location of each of at least one object within a region of interest (e.g., a distance from a center of the ROI), a user preference, a history of operation of the electronic device and / or an external electronic device (e.g., a history of use of a function of the electronic device by the user), or a status of the electronic device and / or the external electronic device, and obtain the interest in objects within the ROI as an output of the vision model.
[0194] For example, an electronic device can recognize at least one object in at least one of a plurality of images. The electronic device can determine the user's interest in each of the objects included in at least one of the plurality of images. For example, if the electronic device acquires multiple images, the electronic device can determine the user's interest in the objects included in each image. For example, the electronic device can determine the interest in objects in a single image, or can determine the interest in objects in multiple images. For example, when determining the interest in objects recognized in multiple images, the electronic device can set a relatively high interest in an object on which the user gazed for a long time in the multiple images (for example, if the same object is included in the multiple images and the user gazed in the direction of the object for a relatively long time in each of the images (when each of the images was captured)), or can determine the user's interest in each of the objects included in the multiple images based on the movement of the user's face (for example, gaze and / or head direction) while acquiring the multiple images.
[0195] According to one embodiment, at operation 1230, the electronic device may receive user input corresponding to a query related to an image. For example, the electronic device may receive user input corresponding to the query related to the image through at least one of various input methods including voice, gesture, or touch from the user.
[0196] For example, an electronic device may receive multiple user inputs from a user. The electronic device may receive a first user input corresponding to a first query, such as, "What store is that red sign?" The electronic device may then receive a second user input corresponding to a second query, such as, "Then what is that?" For example, the second query (the second user input) may be related to the first query (the first user input).
[0197] According to one embodiment, in operation 1240, the electronic device may specify an object related to the query among at least one object based on a degree of interest. For example, the electronic device may specify at least one object in order of increasing interest.
[0198] According to one embodiment, the electronic device may provide feedback prompting additional user input to identify an object when there are multiple objects with a specified interest level or higher, multiple objects with the same interest level, the interest level of an object cannot be estimated, or there are no objects with a specified interest level or higher. In response to the feedback, the electronic device may receive additional user input. Based on the additional user input, the electronic device may identify an object related to the user input (e.g., a query).
[0199] For example, the electronic device can identify at least one object from at least one of a plurality of images acquired during a specified time period prior to receiving a user input. For example, if the time point at which the electronic device receives the user input is T, the electronic device can identify at least one object from at least one of the plurality of images acquired before T. For example, the electronic device can identify at least one object included in an image acquired at a time point when the user's movement (e.g., gaze and / or head movement) was relatively small among the plurality of images acquired before T.
[0200] For example, an electronic device can specify an object based on a user input. For example, when the electronic device receives a first user input (e.g., “What store is that red sign?”), the electronic device can recognize the user intent contained in the first user input. For example, the electronic device can recognize the meaning of “red sign” contained in the first user input, and specify at least one object corresponding to the red sign among a plurality of images acquired up to the time point (T1) of the first user input. For example, when the electronic device can specify an object based on a user input, the electronic device can specify the object without considering the ROI and / or the interest, or can specify the object based on the user input, the ROI, and / or the interest. For example, if an electronic device receives a second user input (e.g., “Then what is that?”) after receiving a first user input, the electronic device may identify at least one object from among a plurality of images acquired before the second user input time point (T2) based on the first user input, the second user input, the ROI of the plurality of images, and / or the interest of objects within the plurality of images. For example, if the electronic device receives a second user input after the first user input, the electronic device may identify the object based at least in part on the first user input (the first query). For example, if the first user input (e.g., the first query) includes an intent to ask about a “red sign,” the electronic device may determine that the second user input (e.g., the second query) also includes an intent to ask about a “signboard.” In this case, the electronic device may identify at least one object corresponding to the signboard included in the plurality of images.
[0201] According to one embodiment, in operation 1250, the electronic device may generate a response to a query based on a specified object and a user input. For example, the electronic device may generate a response including information related to the specified object. According to one embodiment, the electronic device may generate a response to the query using an artificial intelligence-based multimodal model (e.g., a large multimodal model (LMM)). For example, the electronic device may provide information related to the specified object (e.g., a name of the specified object and / or an image of the specified object) and a user input (e.g., a query content) as inputs to the multimodal model, and may obtain a response to the query as an output of the multimodal model.
[0202] For example, in response to a first user input (e.g., “What store is that red sign?”), the electronic device may obtain information related to at least one object corresponding to the red sign included in at least one of the plurality of images acquired up to the time of receiving the first user input. For example, the electronic device may recognize the shape of the sign included in the specific object, and the picture, figure, and / or text included in the sign. Based on the obtained information, the electronic device may generate a first response to the first user input (e.g., “This is XX Banjeom Chinese Restaurant”). For example, in response to a second user input (e.g., “Then what is that?”), the electronic device may obtain information related to the specific object (e.g., the sign related to the second user input). Based on the obtained information, the electronic device may generate a second response to the second user input (e.g., “This is an Italian restaurant”).
[0203] According to one embodiment, at operation 1260, the electronic device may provide the generated response. For example, the electronic device may display the response via a display of the electronic device and / or an external electronic device, or output the response via a speaker of the electronic device and / or an external electronic device.
[0204] According to various embodiments, the order of the operations of FIGS. 11 and 12 may be changed, at least some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 13 or FIG. 14) may be added.
[0205]
[0206] Fig. 13 is a flowchart of an operating method of an electronic device according to one embodiment. Below, descriptions that overlap with those of Figs. 11 and 12 are briefly described or omitted.
[0207] According to one embodiment, in operation 1310, an electronic device (e.g., an electronic device (111) of FIG. 1, an electronic device (200) of FIG. 2, an electronic device (901) of FIGS. 9A to 9C, an electronic device (1001) of FIGS. 10A to 10C, an electronic device (1501) of FIG. 15, or an electronic device (1700) of FIG. 17) may be connected to an external electronic device (e.g., an external electronic device (131, 133) of FIG. 1, an external electronic device (903) of FIGS. 9A to 9C, an external electronic device (1003) of FIGS. 10A to 10C, an electronic device (1501, 1502, 1504) of FIG. 15, or an electronic device (1700) of FIG. 17) through a designated communication method (e.g., Bluetooth, Wi-Fi, or UWB).
[0208] According to one embodiment, in operation 1320, the electronic device may acquire an external image through a camera.
[0209] According to one embodiment, in operation 1330, the electronic device may obtain measurement information through sensors of the electronic device and an external electronic device. For example, the electronic device may obtain information measured using an inertial sensor of the electronic device, and receive information measured using an inertial sensor of the external electronic device from the external electronic device.
[0210] According to one embodiment, in operation 1340, the electronic device may determine an initial alignment relationship between the electronic device and the external electronic device. The initial alignment relationship may include information for aligning the axes of the postures of the electronic device and the external electronic device to the same degree. For example, the electronic device may determine a rotation matrix representing a difference value of a posture of one of the electronic device and the external electronic device based on a reference posture (initial posture) of the other device as the initial alignment relationship. According to one embodiment, when determining the initial alignment relationship, the electronic device may synchronize time information of the electronic device and the external electronic device in order to compensate for a delay in data transmitted between the electronic device and the external electronic device.
[0211] In one embodiment, the electronic device can determine an initial alignment relationship when the user's movement falls below a specified threshold. For example, the electronic device can detect the user's ground contact time (GCT) and determine the initial alignment relationship based on the GCT-acquired posture information of the electronic device and an external electronic device.
[0212] In one embodiment, at operation 1350, the electronic device may determine a region of interest (ROI) based on the image and the initial alignment relationship. For example, the electronic device may determine the ROI within the image based on the direction of the user's head (or gaze) and the difference between the posture of the electronic device and an external electronic device and the initial alignment relationship.
[0213] According to one embodiment, at operation 1360, the electronic device may use vision AI (referred to herein as an “AI-based vision model”) to identify objects within a region of interest (ROI) and estimate a degree of interest for the objects. For example, the electronic device may determine the degree of interest of the user for each of the at least one object based on at least some of: a location of each of the at least one object within the region of interest (e.g., a distance from the center of the ROI), a user preference, a history of operation of the electronic device and / or an external electronic device (e.g., a history of use of functions of the electronic device by the user), or a state of the electronic device and / or the external electronic device. For example, the degree of interest may include a probability value that estimates the relevance of each object to the query.
[0214] According to one embodiment, at operation 1370, the electronic device may determine whether a user query has been input. For example, if a user query related to an image (or an object within the image) has been input, the electronic device may perform operation 1380, and if no user query has been input, the electronic device may repeatedly perform operations 1320 to 1360.
[0215] According to one embodiment, at operation 1380, the electronic device may generate a response to a user query based on interest using multimodal AI (referred to herein as an “AI-based multimodal model”). For example, the electronic device may specify at least one object with a high interest as an object corresponding to the query. The electronic device may generate a response to the user query based on information about the specified object. According to one embodiment, if there are multiple objects with an interest level greater than or equal to a specified value, the electronic device may prompt for additional user input to specify an object. The electronic device may specify an object corresponding to the query based on the additional user input, and generate a response to the query based on the specified object.
[0216] In one embodiment, at operation 1390, the electronic device may provide feedback based on the generated response. For example, the electronic device may provide information related to the generated response visually, audibly, and / or tactilely.
[0217] According to various embodiments, the order of the operations of FIG. 13 may be changed, at least some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 11, FIG. 12, or FIG. 14) may be added.
[0218]
[0219] FIG. 14 is a flowchart of an operating method of an electronic device according to one embodiment. For example, FIG. 14 illustrates operations for aligning an initial posture of an electronic device (e.g., an electronic device (111) of FIG. 1 , an electronic device (200) of FIG. 2 , an electronic device (901) of FIGS. 9A to 9C , an electronic device (1001) of FIGS. 10A to 10C , an electronic device (1501) of FIG. 15 , or an electronic device (1700) of FIG. 17 ) with an external electronic device (e.g., an external electronic device (131, 133) of FIG. 1 , an external electronic device (903) of FIGS. 9A to 9C , an external electronic device (1003) of FIGS. 10A to 10C , an electronic device (1501, 1502, 1504) of FIG. 15 , or an electronic device (1700) of FIG. 17 ). Descriptions that overlap with those of Figures 11 to 13 are briefly explained or omitted.
[0220] According to one embodiment, in operation 1410, the electronic device may obtain measurement information using sensors of the electronic device and an external electronic device. For example, the electronic device may obtain measurement information using inertial sensors of the electronic device and an external electronic device.
[0221] According to one embodiment, in operation 1420, the electronic device can recognize whether the electronic device and / or the external electronic device is worn or attached. For example, the electronic device can determine whether the electronic device and / or the external electronic device is worn or attached based on the information acquired in operation 1410. For example, the electronic device can perform operation 1430 if the electronic device and / or the external electronic device is worn or attached, and can perform operation 1410 if the electronic device or the external electronic device is not worn or attached.
[0222] According to one embodiment, at operation 1430, the electronic device may detect movement of the electronic device. For example, the electronic device may perform operation 1410 if movement of the electronic device is detected, and may perform operation 1440 if no movement is detected. For example, the electronic device may perform operations for initial posture alignment (operations 1440 and below) if movement of the electronic device is not detected.
[0223] According to one embodiment, at operation 1440, the electronic device may acquire an external image using a camera.
[0224] According to one embodiment, in operation 1450, the electronic device may obtain posture information of the electronic device and an external electronic device.
[0225] According to one embodiment, at operation 1460, the electronic device may generate a rotation matrix (see FIG. 7) between the electronic device and the external electronic device. For example, the rotation matrix may represent a difference value between an initial pose of the electronic device and an initial pose of the external electronic device.
[0226] According to one embodiment, at operation 1470, the electronic device may align the initial poses between the electronic device and the external electronic device based on a rotation matrix. For example, the electronic device may align the poses of the electronic device and the external electronic device to values relative to the same axis using the rotation matrix.
[0227] In one embodiment, at operation 1480, the electronic device may align the aligned initial pose and the image. For example, the electronic device may synchronize and store the aligned initial pose and the image acquired at the time of initial pose detection (e.g., a region of the captured image).
[0228] According to various embodiments, the order of the operations of FIG. 14 may be changed, at least some operations may be omitted, or at least one operation (e.g., at least one of the operations of FIG. 11, FIG. 12, or FIG. 13) may be added.
[0229]
[0230] According to one embodiment, a method of operating an electronic device may include an operation of detecting a degree of movement of a user using at least one sensor of the electronic device while the electronic device is worn, attached, or fixed at a first location corresponding to a body part of the user.
[0231] The method may include an operation of detecting a first initial posture of the electronic device using the at least one sensor when the movement of the user is below a specified threshold.
[0232] The method may include receiving information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user.
[0233] The method may include an operation of determining an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device.
[0234] The method may include an operation of obtaining an image including a plurality of objects outside the electronic device using a camera of the electronic device.
[0235] The method may include an operation of detecting a first posture of the electronic device using the at least one sensor.
[0236] The method may include an action of receiving information related to a second posture of the external electronic device from the external electronic device.
[0237] The method may include an operation of determining a region of interest (ROI) of the user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship.
[0238] The method may include an action of identifying at least one object contained within the region of interest.
[0239] The method may include determining a user's interest in each of the at least one object based at least in part on a location of each of the at least one object within the region of interest, a user preference for the at least one object, and a history of operation of the electronic device.
[0240] The operation of identifying at least one object may include an operation of identifying the at least one object using an artificial intelligence-based vision model.
[0241] The method may include receiving user input corresponding to a query related to the image.
[0242] The method may include an operation of specifying an object related to the query among the at least one object based on the interest.
[0243] The method may include an operation of generating a response to the query based on the specified object and the user input using an artificial intelligence-based multimodal model.
[0244] The method may include an action of providing the generated response.
[0245] The action of specifying the object may include providing feedback prompting additional user input to specify the object when there are multiple objects with the interest level greater than or equal to the specified value.
[0246] The method may include, in response to the feedback, receiving the additional user input.
[0247] The method may include an action of specifying an object related to the query based on the additional user input.
[0248] The operation of detecting the first posture may include an operation of detecting the first posture while acquiring the image using the camera.
[0249] The operation of receiving information related to the second posture may include an operation of receiving information related to the second posture of the external electronic device from the external electronic device while acquiring the image.
[0250] The operation of detecting the first initial posture may include an operation of recognizing the ground contact time (GCT) of the user using the at least one sensor.
[0251] The operation of detecting the first initial posture may include an operation of detecting the first initial posture of the electronic device at the ground contact time.
[0252] The second initial posture of the external electronic device may correspond to the posture of the external electronic device at the time of ground contact.
[0253] The operation of detecting the degree of movement of the user may include an operation of detecting the degree of movement of the user in response to the electronic device and the external electronic device being connected through the communication circuit or receiving a specified user input.
[0254] The operation of determining the ROI may include an operation of recognizing at least one of the user's head movement or gaze direction based on information related to the second posture.
[0255] The act of determining the ROI may include an act of determining the ROI based at least in part on at least one of the user's head movement or gaze direction.
[0256] A storage medium according to one embodiment may store instructions. The instructions, when executed by at least one processor of an electronic device, may cause the electronic device to detect a degree of movement of the user using at least one sensor of the electronic device while the electronic device is worn, attached, or fixed at a first location corresponding to a part of the user's body.
[0257] The instructions, when executed by at least one processor of the electronic device, may cause the electronic device to detect a first initial pose of the electronic device using the at least one sensor when a movement of the user is below a specified threshold.
[0258] The instructions, when executed by at least one processor of the electronic device, may cause the electronic device to receive information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user.
[0259] The instructions, when executed by at least one processor of the electronic device, may cause the electronic device to determine an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device.
[0260] The instructions, when executed by at least one processor of the electronic device, may cause the electronic device to obtain an image including a plurality of objects outside the electronic device using a camera of the electronic device.
[0261] The above instructions, when executed by at least one processor of the electronic device, may cause the electronic device to detect a first posture of the electronic device using the at least one sensor.
[0262] The instructions, when executed by at least one processor of the electronic device, may cause the electronic device to receive information related to a second posture of the external electronic device from the external electronic device.
[0263] The instructions, when executed by at least one processor of the electronic device, may cause the electronic device to determine a region of interest (ROI) of the user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship.
[0264] According to embodiments of the present disclosure, based on the posture of an electronic device, the posture of an external electronic device, the posture alignment relationship between the electronic device and the external electronic device, the user's movement, gaze, information related to the electronic device, and / or information related to the external electronic device, an ROI matching a user's intention within an image can be determined, and an object matching the user's intention can be specified. According to embodiments of the present disclosure, by improving the accuracy in specifying an object matching a user query (user intent), a response suitable for the user's intention included in the user query can be generated and provided.
[0265]
[0266] FIG. 15 is a block diagram of an electronic device (1501) within a network environment (1500) according to various embodiments. Referring to FIG. 15 , in the network environment (1500), the electronic device (1501) may communicate with the electronic device (1502) via a first network (1598) (e.g., a short-range wireless communication network), or may communicate with at least one of the electronic device (1504) or the server (1508) via a second network (1599) (e.g., a long-range wireless communication network). According to one embodiment, the electronic device (1501) may communicate with the electronic device (1504) via the server (1508). According to one embodiment, the electronic device (1501) may include a processor (1520), a memory (1530), an input module (1550), an audio output module (1555), a display module (1560), an audio module (1570), a sensor module (1576), an interface (1577), a connection terminal (1578), a haptic module (1579), a camera module (1580), a power management module (1588), a battery (1589), a communication module (1590), a subscriber identification module (1596), or an antenna module (1597). In some embodiments, the electronic device (1501) may omit at least one of these components (e.g., the connection terminal (1578)), or may have one or more other components added. In some embodiments, some of these components (e.g., sensor module (1576), camera module (1580), or antenna module (1597)) may be integrated into a single component (e.g., display module (1560)).
[0267] The processor (1520) may, for example, execute software (e.g., a program (1540)) to control at least one other component (e.g., a hardware or software component) of the electronic device (1501) connected to the processor (1520) and perform various data processing or operations. According to one embodiment, as at least a part of the data processing or operations, the processor (1520) may store commands or data received from other components (e.g., a sensor module (1576) or a communication module (1590)) in a volatile memory (1532), process the commands or data stored in the volatile memory (1532), and store result data in a non-volatile memory (1534). According to one embodiment, the processor (1520) may include a main processor (1521) (e.g., a central processing unit or an application processor) or a secondary processor (1523) (e.g., a graphics processing unit, a neural processing unit (NPU), an image signal processor, a sensor hub processor, or a communication processor) that can operate independently or together with the main processor (1521). For example, when the electronic device (1501) includes the main processor (1521) and the secondary processor (1523), the secondary processor (1523) may be configured to use less power than the main processor (1521) or to be specialized for a given function. The secondary processor (1523) may be implemented separately from the main processor (1521) or as a part thereof.
[0268] The auxiliary processor (1523) may control at least a portion of functions or states associated with at least one component (e.g., the display module (1560), the sensor module (1576), or the communication module (1590)) of the electronic device (1501), for example, on behalf of the main processor (1521) while the main processor (1521) is in an inactive (e.g., sleep) state, or together with the main processor (1521) while the main processor (1521) is in an active (e.g., application execution) state. In one embodiment, the auxiliary processor (1523) (e.g., an image signal processor or a communication processor) may be implemented as a part of another functionally related component (e.g., a camera module (1580) or a communication module (1590)). In one embodiment, the auxiliary processor (1523) (e.g., a neural network processing unit) may include a hardware structure specialized for processing artificial intelligence models. The artificial intelligence models may be generated through machine learning. This learning can be performed, for example, on the electronic device (1501) where the artificial intelligence model is executed, or can be performed through a separate server (e.g., server (1508)). The learning algorithm can include, for example, supervised learning, unsupervised learning, semi-supervised learning, or reinforcement learning, but is not limited to the examples described above. The artificial intelligence model can include multiple artificial neural network layers.The artificial neural network may be one of a deep neural network (DNN), a convolutional neural network (CNN), a recurrent neural network (RNN), a restricted Boltzmann machine (RBM), a deep belief network (DBN), a bidirectional recurrent deep neural network (BRDNN), a deep Q-network, or a combination of two or more of the above, but is not limited to the examples described above. In addition to, or alternatively to, a hardware structure, an artificial intelligence model may include a software structure.
[0269] The memory (1530) can store various data used by at least one component (e.g., the processor (1520) or the sensor module (1576)) of the electronic device (1501). The data can include, for example, software (e.g., the program (1540)) and input data or output data for commands related thereto. The memory (1530) can include volatile memory (1532) or non-volatile memory (1534).
[0270] The program (1540) may be stored as software in memory (1530) and may include, for example, an operating system (1542), middleware (1544), or an application (1546).
[0271] The input module (1550) can receive commands or data to be used in a component of the electronic device (1501) (e.g., a processor (1520)) from an external source (e.g., a user) of the electronic device (1501). The input module (1550) can include, for example, a microphone, a mouse, a keyboard, a key (e.g., a button), or a digital pen (e.g., a stylus pen).
[0272] The audio output module (1555) can output audio signals to the outside of the electronic device (1501). The audio output module (1555) can include, for example, a speaker or a receiver. The speaker can be used for general purposes, such as multimedia playback or recording playback. The receiver can be used to receive incoming calls. In one embodiment, the receiver can be implemented separately from the speaker or as part of the speaker.
[0273] The display module (1560) can visually provide information to an external party (e.g., a user) of the electronic device (1501). The display module (1560) may include, for example, a display, a holographic device, or a projector and a control circuit for controlling the device. In one embodiment, the display module (1560) may include a touch sensor configured to detect a touch, or a pressure sensor configured to measure the intensity of a force generated by the touch.
[0274] The audio module (1570) can convert sound into an electrical signal, or vice versa. According to one embodiment, the audio module (1570) can acquire sound through the input module (1550), output sound through the sound output module (1555), or an external electronic device (e.g., electronic device (1502)) (e.g., speaker or headphone) directly or wirelessly connected to the electronic device (1501).
[0275] The sensor module (1576) can detect the operating status (e.g., power or temperature) of the electronic device (1501) or the external environmental status (e.g., user status) and generate an electrical signal or data value corresponding to the detected status. According to one embodiment, the sensor module (1576) can include, for example, a gesture sensor, a gyro sensor, a barometric pressure sensor, a magnetic sensor, an acceleration sensor, a grip sensor, a proximity sensor, a color sensor, an IR (infrared) sensor, a biometric sensor, a temperature sensor, a humidity sensor, or an illuminance sensor.
[0276] The interface (1577) may support one or more designated protocols that may be used to directly or wirelessly connect the electronic device (1501) with an external electronic device (e.g., the electronic device (1502)). In one embodiment, the interface (1577) may include, for example, a high definition multimedia interface (HDMI), a universal serial bus (USB) interface, an SD card interface, or an audio interface.
[0277] The connection terminal (1578) may include a connector through which the electronic device (1501) may be physically connected to an external electronic device (e.g., the electronic device (1502)). According to one embodiment, the connection terminal (1578) may include, for example, an HDMI connector, a USB connector, an SD card connector, or an audio connector (e.g., a headphone connector).
[0278] The haptic module (1579) can convert electrical signals into mechanical stimuli (e.g., vibration or movement) or electrical stimuli that a user can perceive through tactile or kinesthetic sensations. In one embodiment, the haptic module (1579) can include, for example, a motor, a piezoelectric element, or an electrical stimulation device.
[0279] The camera module (1580) can capture still images and videos. According to one embodiment, the camera module (1580) may include one or more lenses, image sensors, image signal processors, or flashes.
[0280] The power management module (1588) can manage power supplied to the electronic device (1501). According to one embodiment, the power management module (1588) can be implemented, for example, as at least a part of a power management integrated circuit (PMIC).
[0281] A battery (1589) may power at least one component of the electronic device (1501). In one embodiment, the battery (1589) may include, for example, a non-rechargeable primary battery, a rechargeable secondary battery, or a fuel cell.
[0282] The communication module (1590) may support the establishment of a direct (e.g., wired) communication channel or a wireless communication channel between the electronic device (1501) and an external electronic device (e.g., electronic device (1502), electronic device (1504), or server (1508)), and the performance of communication through the established communication channel. The communication module (1590) may operate independently from the processor (1520) (e.g., application processor) and may include one or more communication processors that support direct (e.g., wired) communication or wireless communication. According to one embodiment, the communication module (1590) may include a wireless communication module (1592) (e.g., a cellular communication module, a short-range wireless communication module, or a global navigation satellite system (GNSS) communication module) or a wired communication module (1594) (e.g., a local area network (LAN) communication module, or a power line communication module). Any of these communication modules may communicate with an external electronic device (1504) via a first network (1598) (e.g., a short-range communication network such as Bluetooth, wireless fidelity (WiFi) direct, or infrared data association (IrDA)) or a second network (1599) (e.g., a long-range communication network such as a legacy cellular network, a 5G network, a next-generation communication network, the Internet, or a computer network (e.g., a local area network or a wide area network)). These various types of communication modules may be integrated into a single component (e.g., a single chip) or implemented as multiple separate components (e.g., multiple chips). The wireless communication module (1592) may use subscriber information (e.g., an international mobile subscriber identity (IMSI)) stored in the subscriber identification module (1596) to identify or authenticate the electronic device (1501) within a communication network such as the first network (1598) or the second network (1599).
[0283] The wireless communication module (1592) can support 5G networks and next-generation communication technologies following the 4G network, such as NR access technology (new radio access technology). NR access technology can support high-speed transmission of high-capacity data (eMBB (enhanced mobile broadband)), minimizing terminal power and connecting multiple terminals (mMTC (massive machine type communications)), or high reliability and low latency communications (URLLC (ultra-reliable and low-latency communications)). The wireless communication module (1592) can support, for example, a high-frequency band (e.g., mmWave band) to achieve a high data transmission rate. The wireless communication module (1592) can support various technologies for securing performance in high-frequency bands, such as beamforming, massive multiple-input and multiple-output (MIMO), full dimensional MIMO (FD-MIMO), array antenna, analog beam-forming, or large scale antenna. The wireless communication module (1592) can support various requirements specified in the electronic device (1501), an external electronic device (e.g., the electronic device (1504)), or a network system (e.g., the second network (1599)). According to one embodiment, the wireless communication module (1592) may support a peak data rate (e.g., 20 Gbps or more) for eMBB realization, a loss coverage (e.g., 164 dB or less) for mMTC realization, or a U-plane latency (e.g., 0.5 ms or less for downlink (DL) and uplink (UL), or 1 ms or less for round trip) for URLLC realization.
[0284] The antenna module (1597) can transmit or receive signals or power to or from an external device (e.g., an external electronic device). In one embodiment, the antenna module (1597) may include an antenna including a radiator formed of a conductor or a conductive pattern formed on a substrate (e.g., a PCB). In one embodiment, the antenna module (1597) may include a plurality of antennas (e.g., an array antenna). In this case, at least one antenna suitable for a communication method used in a communication network, such as the first network (1598) or the second network (1599), may be selected from the plurality of antennas by, for example, the communication module (1590). A signal or power may be transmitted or received between the communication module (1590) and the external electronic device via the at least one selected antenna. In some embodiments, in addition to the radiator, another component (e.g., a radio frequency integrated circuit (RFIC)) may be additionally formed as a part of the antenna module (1597).
[0285] According to various embodiments, the antenna module (1597) may form a mmWave antenna module. According to one embodiment, the mmWave antenna module may include a printed circuit board, an RFIC disposed on or adjacent a first side (e.g., a bottom side) of the printed circuit board and capable of supporting a designated high frequency band (e.g., a mmWave band), and a plurality of antennas (e.g., an array antenna) disposed on or adjacent a second side (e.g., a top side or a side side) of the printed circuit board and capable of transmitting or receiving signals in the designated high frequency band.
[0286] At least some of the above components can be interconnected and exchange signals (e.g., commands or data) with each other via a communication method between peripheral devices (e.g., a bus, GPIO (general purpose input and output), SPI (serial peripheral interface), or MIPI (mobile industry processor interface)).
[0287] According to one embodiment, commands or data may be transmitted or received between the electronic device (1501) and an external electronic device (1504) via a server (1508) connected to a second network (1599). Each of the external electronic devices (1502, or 104) may be the same or a different type of device as the electronic device (1501). According to one embodiment, all or part of the operations executed in the electronic device (1501) may be executed in one or more of the external electronic devices (1502, 104, or 108). For example, when the electronic device (1501) is to perform a certain function or service automatically or in response to a request from a user or another device, the electronic device (1501) may, instead of or in addition to executing the function or service itself, request one or more external electronic devices to perform the function or at least a part of the service. One or more external electronic devices that receive the request may execute at least a portion of the requested function or service, or an additional function or service related to the request, and transmit the result of the execution to the electronic device (1501). The electronic device (1501) may process the result as is or additionally and provide it as at least a portion of a response to the request. For this purpose, cloud computing, distributed computing, mobile edge computing (MEC), or client-server computing technology may be used, for example. The electronic device (1501) may provide an ultra-low latency service using, for example, distributed computing or mobile edge computing. In another embodiment, the external electronic device (1504) may include an Internet of Things (IoT) device. The server (1508) may be an intelligent server utilizing machine learning and / or a neural network.In one embodiment, an external electronic device (1504) or server (1508) may be included within the second network (1599). The electronic device (1501) may be applied to intelligent services (e.g., smart homes, smart cities, smart cars, or healthcare) based on 5G communication technology and IoT-related technology.
[0288]
[0289] FIG. 16 is a block diagram (1600) illustrating a program (1540) according to various embodiments. According to one embodiment, the program (1540) may include an operating system (1542), middleware (1544), or an application (1546) executable on the operating system (1542) for controlling one or more resources of the electronic device (1501). The operating system (1542) may include, for example, Android™, iOS™, Windows™, Symbian™, Tizen™, or Bada™. At least some of the programs (1540) may be preloaded onto the electronic device (1501), for example, during manufacturing, or may be downloaded or updated from an external electronic device (e.g., the electronic device (1502 or 1504), or a server (1508)) when used by a user.
[0290] The operating system (1542) may control the management (e.g., allocation or retrieval) of one or more system resources (e.g., processes, memory, or power) of the electronic device (1501). The operating system (1542) may additionally or alternatively include one or more driver programs for driving other hardware devices of the electronic device (1501), for example, an input module (1550), an audio output module (1555), a display module (1560), an audio module (1570), a sensor module (1576), an interface (1577), a haptic module (1579), a camera module (1580), a power management module (1588), a battery (1589), a communication module (1590), a subscriber identification module (1596), or an antenna module (1597).
[0291] Middleware (1544) may provide various functions to the application (1546) so that functions or information provided from one or more resources of the electronic device (1501) can be used by the application (1546). Middleware (1544) may include, for example, an application manager (1601), a window manager (1603), a multimedia manager (1605), a resource manager (1607), a power manager (1609), a database manager (1611), a package manager (1613), a connectivity manager (1615), a notification manager (1617), a location manager (1619), a graphics manager (1621), a security manager (1623), a call manager (1625), or a voice recognition manager (1627).
[0292] The application manager (1601) can, for example, manage the life cycle of the application (1546). The window manager (1603) can, for example, manage one or more GUI resources used on the screen. The multimedia manager (1605) can, for example, identify one or more formats required for playing media files, and perform encoding or decoding of a corresponding media file among the media files using a codec suitable for the corresponding format selected among the media files. The resource manager (1607) can, for example, manage the source code of the application (1546) or the memory space of the memory (1530). The power manager (1609) can, for example, manage the capacity, temperature, or power of the battery (1589), and use the corresponding information among these to determine or provide relevant information required for the operation of the electronic device (1501). According to one embodiment, the power manager (1609) may be interoperable with a basic input / output system (BIOS) (not shown) of the electronic device (1501).
[0293] The database manager (1611) may, for example, create, search, or modify a database to be used by the application (1546). The package manager (1613) may, for example, manage the installation or update of an application distributed in the form of a package file. The connectivity manager (1615) may, for example, manage a wireless connection or direct connection between the electronic device (1501) and an external electronic device. The notification manager (1617) may, for example, provide a function for notifying a user of the occurrence of a specified event (e.g., an incoming call, a message, or an alarm). The location manager (1619) may, for example, manage location information of the electronic device (1501). The graphics manager (1621) may, for example, manage one or more graphic effects to be provided to the user or a user interface related thereto.
[0294] The security manager (1623) may provide, for example, system security or user authentication. The telephony manager (1625) may manage, for example, a voice call function or a video call function provided by the electronic device (1501). The voice recognition manager (1627) may, for example, transmit the user's voice data to the server (1508) and receive, from the server (1508), a command corresponding to a function to be performed in the electronic device (1501) based at least in part on the voice data, or text data converted based at least in part on the voice data. In one embodiment, the middleware (1644) may dynamically delete some existing components or add new components. In one embodiment, at least a portion of the middleware (1544) may be included as part of the operating system (1542) or implemented as separate software different from the operating system (1542).
[0295] The application (1546) may include, for example, a home (1651), a dialer (1653), an SMS / MMS (1655), an instant message (IM) (1657), a browser (1659), a camera (1661), an alarm (1663), a contact (1665), a voice recognition (1667), an email (1669), a calendar (1671), a media player (1673), an album (1675), a watch (1677), a health (1679) (e.g., measuring biometric information such as the amount of exercise or blood sugar), or an environmental information (1681) (e.g., measuring barometric pressure, humidity, or temperature information) application. According to one embodiment, the application (1546) may further include an information exchange application (not shown) that can support information exchange between the electronic device (1501) and an external electronic device. The information exchange application may include, for example, a notification relay application configured to transmit designated information (e.g., a call, a message, or an alarm) to an external electronic device, or a device management application configured to manage an external electronic device. The notification relay application may, for example, transmit notification information corresponding to a designated event (e.g., receipt of an email) that occurred in another application (e.g., an email application (1669)) of the electronic device (1501) to the external electronic device. Additionally or alternatively, the notification relay application may receive notification information from the external electronic device and provide it to a user of the electronic device (1501).
[0296] A device management application may, for example, control the power (e.g., turning on or off) or the function (e.g., brightness, resolution, or focus) of an external electronic device or a component thereof (e.g., a display module or a camera module of the external electronic device) that communicates with the electronic device (1501). The device management application may additionally or alternatively support the installation, deletion, or update of an application running on the external electronic device.
[0297]
[0298] FIG. 17 illustrates a generative artificial intelligence system according to one embodiment.
[0299] Referring to FIG. 17, in a generative artificial intelligence system (1700), a user question / response interface (1710) can receive user input. The user input may be in the form of natural language, images, and / or videos. Additionally, contextual information may also be transmitted when the user input is transmitted. The contextual information may include various additional information at the time of the user input. For example, information on the application currently being used by the user or information on the user's location. Furthermore, the user input may also be in a form that combines the aforementioned natural language, images, sounds, and contextual information. Furthermore, the user input may also be in a form other than natural language, such as selecting a menu.
[0300] According to one embodiment, the user question / response interface (1710) may output the results of a generative artificial intelligence system to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user. The user question / response interface (1710) may output the results of a generative artificial intelligence system (1700) to the user. The output may be in the form of natural language or specific content, and may also be provided in the form of an action requested by the user.
[0301] According to one embodiment, the artificial intelligence framework (1720) can receive user input and coordinate and control each component or module necessary to perform the user's intention based on the user's query.
[0302] In one embodiment, user input received from the user question / response interface (1710) may be transmitted to a prompt design module (1721). The prompt design module (1721) may be used to generate prompts suitable for inputting the user input into a large language model (LMM) or a large multi-modal model (LMM). The prompt design module (1721) may be an artificial intelligence component that uses a machine learning algorithm or a neural network to develop better prompts over time. The prompt design module (1721) may access a database (1730) containing user preference data, a prompt library, and prompt examples based on the user input to generate prompts, and may transmit the generated prompts to the LLM or LMM.
[0303] According to one embodiment, the API / plug-in management module (1722) may communicate with external information when there is a request for additional information when passing user input as input to the generative artificial intelligence model (1750). The API / plug-in management module (1722) may establish a channel for communicating with the outside of the artificial intelligence interface through an application programming interface (API) and may enable access to various data sources through the established channel. In addition, the API / plug-in management module (1722) may request an action through the API that ultimately performs the user input, rather than an intermediate result, if the application / service module (1740) needs to perform the action. Information obtained from the outside may be used to generate a prompt in the prompt design module (1721) together with the user input or may be passed as an input to the generative artificial intelligence model (1750).
[0304] In one embodiment, the transformation module (1723) can fine-tune the output from the generative artificial intelligence model (1750). For example, the transformation module (1723) can verify whether the content generated through the LLM and / or LMM is irrelevant, biased, or harmful. In addition, the transformation module (1723) can determine to what extent the content matches the user's desired result and, if necessary, perform additional processing. The transformation module (1723) can additionally configure and provide the user with hints to avoid undesired output.
[0305] According to one embodiment, a generative artificial intelligence model (1750) may generally refer to an artificial intelligence neural network that creates new types of data based on user input information. The generative artificial intelligence model (1750) may include a model that generates images and / or a model that generates language. Representative models that generate images include a generative adversarial network (GAN) and a variational autoencoder (VAE), and examples include a diffusion-based generative model that uses a VAE and a transformer structure. A model that generates language is a model that is trained to statistically output the most appropriate output value based on an input value, and representative examples include models such as CHAT-GPT 3 and CHAT-GPT 4. In addition, there is also an LMM that can recognize various types of data input such as text, images, and voice and generate new data corresponding thereto.
[0306]
[0307] Electronic devices according to the various embodiments disclosed in this document may take various forms. Electronic devices may include, for example, portable communication devices (e.g., smartphones), computer devices, portable multimedia devices, portable medical devices, cameras, wearable devices, or home appliances. Electronic devices according to the embodiments of this document are not limited to the aforementioned devices.
[0308] The various embodiments of this document and the terminology used therein are not intended to limit the technical features described in this document to specific embodiments, but should be understood to include various modifications, equivalents, or substitutes of the embodiments. In connection with the description of the drawings, similar reference numerals may be used for similar or related components. The singular form of a noun corresponding to an item may include one or more of the items, unless the context clearly indicates otherwise. In this document, each of the phrases "A or B", "at least one of A and B", "at least one of A or B", "A, B, or C", "at least one of A, B, and C", and "at least one of A, B, or C" can include any one of the items listed together in the corresponding phrase among those phrases, or all possible combinations thereof. Terms such as "first," "second," or "first" or "second" may be used merely to distinguish one component from another, and do not limit the components in any other respect (e.g., importance or order). When a component (e.g., a first component) is referred to as "coupled" or "connected" to another component (e.g., a second component), with or without the terms "functionally" or "communicatively," it means that the component can be connected to the other component directly (e.g., wired), wirelessly, or through a third component.
[0309] The term "module" used in various embodiments of this document may include a unit implemented in hardware, software, or firmware, and may be used interchangeably with terms such as logic, logic block, component, or circuit. A module may be an integral component, or a minimum unit or part of such a component that performs one or more functions. For example, according to one embodiment, a module may be implemented in the form of an application-specific integrated circuit (ASIC).
[0310] Various embodiments of the present document may be implemented as software (e.g., a program (1540)) including one or more instructions stored in a storage medium (e.g., an internal memory (1536) or an external memory (1538)) readable by a machine (e.g., an electronic device (1501)). For example, a processor (e.g., a processor (1520)) of the machine (e.g., an electronic device (1501)) may call at least one instruction among the one or more instructions stored from the storage medium and execute it. This enables the machine to operate to perform at least one function according to the at least one called instruction. The one or more instructions may include code generated by a compiler or code executable by an interpreter. The machine-readable storage medium may be provided in the form of a non-transitory storage medium. Here, 'non-transitory' simply means that the storage medium is a tangible device and does not contain signals (e.g., electromagnetic waves), and the term does not distinguish between cases where data is stored semi-permanently or temporarily on the storage medium.
[0311] According to one embodiment, the method according to various embodiments disclosed in the present document may be provided as included in a computer program product. The computer program product may be traded as a product between a seller and a buyer. The computer program product may be distributed in the form of a machine-readable storage medium (e.g., compact disc read-only memory (CD-ROM)), or may be distributed online (e.g., downloaded or uploaded) via an application store (e.g., Play Store™) or directly between two user devices (e.g., smart phones). In the case of online distribution, at least a portion of the computer program product may be temporarily stored or temporarily generated in a machine-readable storage medium, such as the memory of a manufacturer's server, an application store's server, or an intermediary server.
[0312] According to various embodiments, each component (e.g., a module or a program) of the above-described components may include one or more entities, and some of the entities may be separated and placed in other components. According to various embodiments, one or more components or operations of the aforementioned components may be omitted, or one or more other components or operations may be added. Alternatively or additionally, a plurality of components (e.g., a module or a program) may be integrated into a single component. In such a case, the integrated component may perform one or more functions of each of the plurality of components identically or similarly to those performed by the corresponding component among the plurality of components prior to the integration. According to various embodiments, the operations performed by a module, program, or other component may be executed sequentially, in parallel, iteratively, or heuristically, or one or more of the operations may be executed in a different order, omitted, or one or more other operations may be added.
[0313]
[0314]
Claims
1. In electronic devices, camera; communication circuit; At least one sensor; memory; and Contains at least one processor, The above memory, when executed by the at least one processor, causes the electronic device to: The electronic device detects the degree of movement of the user using at least one sensor while being worn, attached, or fixed at a first location corresponding to a part of the user's body, If the movement of the user is below a specified threshold, detecting a first initial posture of the electronic device using the at least one sensor, Receive information related to a second initial posture of an external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user through the communication circuit; Determining an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device, Using the camera, an image including a plurality of objects outside the electronic device is obtained, Detecting a first posture of the electronic device using at least one sensor, Receive information related to a second posture of the external electronic device from the external electronic device, An electronic device storing instructions for determining a region of interest (ROI) of a user within the image based on at least some of the first initial pose, the second initial pose, the first pose, the second pose, or the initial alignment relationship.
2. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Identifying at least one object contained within the region of interest; An electronic device that determines a user's interest in each of the at least one object based at least in part on a location of each of the at least one object within the region of interest, a user preference for the at least one object, and a history of operation of the electronic device.
3. In claim 2, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device that identifies at least one object using an artificial intelligence-based vision model.
4. In claim 2, The above instructions, when executed by the at least one processor, cause the electronic device to: Receive user input corresponding to a query related to the above image, Based on the above interest, specify an object related to the query among the at least one object, Using an artificial intelligence-based multimodal model, a response to the query is generated based on the specified object and the user input, An electronic device configured to provide the above generated response.
5. In claim 4, The above instructions, when executed by the at least one processor, cause the electronic device to: If there are multiple objects with a level of interest greater than the specified value, provide feedback prompting additional user input to identify the objects. In response to the above feedback, receiving the additional user input, An electronic device that specifies an object related to the query based on the additional user input.
6. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Detecting the first pose while acquiring the image using the camera, An electronic device that receives information related to the second posture of the external electronic device from the external electronic device while acquiring the image.
7. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Recognize the user's ground contact time (GCT) using at least one sensor, To detect the first initial posture of the electronic device at the above ground contact time, The second initial posture of the external electronic device corresponds to the posture of the external electronic device at the time of ground contact.
8. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: An electronic device that detects the degree of movement of the user to determine the initial alignment relationship, based on the electronic device and the external electronic device being connected through the communication circuit or receiving a specified user input.
9. In claim 1, The above instructions, when executed by the at least one processor, cause the electronic device to: Synchronize the time information of the electronic device and the external electronic device, An electronic device that recognizes and corrects the delay time of information transmitted and received between the electronic device and the external electronic device based on the synchronized time information.
10. In claim 1, The second posture of the external electronic device changes according to the movement of the user, The above instructions, when executed by the at least one processor, cause the electronic device to: Based on the information related to the second posture, at least one of the user's head movement or gaze direction is recognized, An electronic device that determines the ROI based at least in part on at least one of the user's head movement or gaze direction.
11. In the method of operating an electronic device, An action of detecting the degree of movement of the user using at least one sensor of the electronic device while the electronic device is worn, attached, or fixed at a first location corresponding to a part of the user's body; An action of detecting a first initial posture of the electronic device using the at least one sensor when the movement of the user is below a specified threshold; An action of receiving information related to a second initial posture of an external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user; An operation of determining an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device; An action of obtaining an image including a plurality of objects outside the electronic device using a camera of the electronic device; An operation of detecting a first posture of the electronic device using at least one sensor; An operation of receiving information related to a second posture of the external electronic device from the external electronic device; and A method comprising an action of determining a region of interest (ROI) of a user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship.
12. In claim 11, An operation of identifying at least one object contained within the region of interest; and A method comprising determining a user's interest in each of said at least one object based at least in part on a location of each of said at least one object within said region of interest, a user preference for said at least one object, and a history of operation of said electronic device.
13. In claim 12, The action of identifying at least one object above, A method comprising an action of identifying at least one object using an artificial intelligence-based vision model.
14. In claim 12, An action for receiving user input corresponding to a query related to the image; An action of specifying an object related to the query among the at least one object based on the above interest; An operation of generating a response to the query based on the specified object and the user input using an artificial intelligence-based multimodal model; and A method comprising an action providing the generated response.
15. Instructions are stored in a non-transitory storage medium, The above instructions, when executed by at least one processor of an electronic device, cause the electronic device to: The electronic device is worn, attached, or fixed at a first location corresponding to a part of the user's body, and the degree of movement of the user is detected using at least one sensor of the electronic device, If the movement of the user is below a specified threshold, detecting a first initial posture of the electronic device using the at least one sensor, Receive information related to a second initial posture of the external electronic device from an external electronic device worn, attached, or fixed at a second location corresponding to another body of the user; Determining an initial alignment relationship between a first initial posture of the electronic device and a second initial posture of the external electronic device, Obtaining an image including a plurality of objects outside the electronic device using a camera of the electronic device, Detecting a first posture of the electronic device using at least one sensor, Receive information related to a second posture of the external electronic device from the external electronic device, A storage medium for determining a region of interest (ROI) of a user within the image based on at least some of the first initial posture, the second initial posture, the first posture, the second posture, or the initial alignment relationship.
Citation Information
Patent Citations
Method for determining region of interest of image and device for determining region of interest of image
KR1020160068447A
A composition for emitting glucose comprising hydrogel and epidermal growth factor receptor ligand as an active ingredient
KR1020210003052A
Apparatus that performs wireless communication, method of operating the apparatus, and wireless communication system including the same
KR1020250076351A
Organic solvent adsorption system having activated carbon fiber adsorption and desorption module
KR102723411B1
Method and system for controlling access to virtual and real-world environments for head mounted device
US20230298221A1