System and method for processing based on user query and gaze

By integrating cameras and eye-tracking technology into a head-mounted device and combining them with machine learning models to process user gaze and queries, the problem of low interaction accuracy in 3D environments of existing devices is solved, achieving more efficient operation and interaction capabilities.

CN121742708APending Publication Date: 2026-03-27APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511377449.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2025-09-16
Filing Date
2025-09-25
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

Existing electronic devices struggle to improve the accuracy and efficiency of operations when processing user queries and gaze information, especially when interacting in complex 3D environments.

Method used

It employs a head-mounted device equipped with a camera and audio/text input device, combining eye tracking and environmental image capture technology to determine the region of interest through user gaze and queries, and uses machine learning models to process images and voice commands to execute corresponding operations.

Benefits of technology

It improves the accuracy and efficiency of electronic devices in responding to user input in a three-dimensional environment, enhances their ability to interact with physical objects, and supports a variety of applications such as drawing, presentation, and word processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121742708A_ABST
    Figure CN121742708A_ABST
Patent Text Reader

Abstract

The invention relates to a system and method for processing based on user queries and gaze. In some examples, an electronic device in communication with one or more input devices detects an input and gaze direction of a user of the electronic device. In some examples, in response to an input, an electronic device captures one or more images. In some examples, the electronic device identifies at least a subset of the first image from the captured image using the detected gaze direction and a portion of the input. If certain criteria are met, the electronic device performs an operation using the processing circuitry based on the processing input, the captured image, and the identified subset of the first image.
Need to check novelty before this filing date? Find Prior Art

Description

Cross-references to related applications

[0001] This application claims the benefit of U.S. Provisional Application No. 63 / 699,659, filed September 26, 2024, and U.S. Patent Application No. 19 / 330,484, filed September 16, 2025, the entire disclosure of which is incorporated herein by reference for all purposes. Technical Field

[0002] This disclosure relates in its entirety to processing based on user queries and gaze, and more specifically, to performing actions based on processed images and subsets of images determined based on user queries and gaze. Background Technology

[0003] Electronic devices such as mobile phones and laptops may include digital assistants. A digital assistant in an electronic device can receive user queries in the form of natural language input and enable the electronic device to perform actions in response to the user query. Summary of the Invention

[0004] Electronic devices, such as head-mounted devices, are equipped with or communicate with one or more input devices. In some examples, the input device includes a camera for detecting the user's gaze and one or more cameras for detecting the environment. In some examples, the input device also includes one or more text or audio input components (e.g., a microphone, keyboard, touch sensor panel, etc.). In some examples, the electronic device uses one or more cameras to capture images of the environment and uses the user's gaze to capture a subset of the environmental images (e.g., a cropped version of the image). In effect, the gaze is used to capture the region of interest (ROI) to which the gaze is directed. The ROI may include one or more objects of interest. In some examples, one or more characteristics of the ROI are based on a user query (e.g., voice or text input). In some examples, the image, a subset of the image, and the user query are inputs from which actions can be determined. Using gaze in conjunction with a user query can improve the accuracy of actions performed by the electronic device in response to user input.

[0005] The accompanying drawings and detailed descriptions provide a comprehensive description of the examples, and it should be understood that the above-described invention does not limit the scope of this disclosure in any way. Attached Figure Description

[0006] To better understand the various examples described, reference should be made to the following detailed embodiments in conjunction with the accompanying drawings, in which similar reference numerals indicate corresponding parts throughout the drawings.

[0007] Figure 1Examples of electronic devices and handheld electronic devices that present a three-dimensional environment according to some examples of this disclosure are illustrated.

[0008] Figures 2A to 2B A block diagram illustrating an example architecture for an electronic device according to some examples of this disclosure is shown.

[0009] Figures 3A to 3R Examples of electronic devices according to this disclosure are illustrated in various ways when performing operations based on a combination of user queries, images of the three-dimensional environment, and cropped images of the three-dimensional environment.

[0010] Figure 4 Examples of methods for performing operations based on a combination of user queries with images of a 3D environment and cropped images, according to some examples of this disclosure, are illustrated. Detailed Implementation

[0011] The following description of the examples will be referenced to the accompanying drawings, which form part of the description and illustrate specific examples of alternative implementations by way of example. It should be understood that other examples and structural changes may be optionally used and optionally made without departing from the scope of the disclosed examples.

[0012] Electronic devices, such as head-mounted devices, are equipped with or communicate with one or more input devices. In some examples, the input device includes a camera for detecting the user's gaze and one or more cameras for detecting the environment. In some examples, the input device also includes one or more text or audio input components (e.g., a microphone, keyboard, touch sensor panel, etc.). In some examples, the electronic device uses one or more cameras to capture images of the environment and uses the user's gaze to capture a subset of the environmental images (e.g., a cropped version of the image). In effect, the gaze is used to capture the region of interest (ROI) to which the gaze is directed. The ROI may include one or more objects of interest. In some examples, one or more characteristics of the ROI are based on a user query (e.g., voice or text input). In some examples, the image, a subset of the image, and the user query are inputs from which actions can be determined. Using gaze in conjunction with a user query can improve the accuracy of actions performed by the electronic device in response to user input.

[0013] Although the following description uses the terms "first," "second," etc., to describe various elements, these elements should not be limited by the terms. These terms are only used to distinguish one element from another. For example, a first touch may be named a second touch, and similarly, a second touch may be named a first touch, without departing from the various examples described. Both a first touch and a second touch are touches, but they are not the same touch.

[0014] The terminology used in the description of the various described examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various described examples and the appended claims, the singular forms “an,” “a,” and “the” are intended to include the plural forms as well, unless the context expressly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the terms “comprising” and / or “including” as used in this specification specify the presence of the stated features, integers, steps, operations, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0015] Depending on the context, the term "if" may optionally be interpreted as meaning "when," "in," or "in response to determination" or "in response to detection." Similarly, depending on the context, the phrases "if it is determined..." or "if [the stated condition or event] is detected" may optionally be interpreted as meaning "in response to determination..." or "in response to detection of [the stated condition or event]."

[0016] Figure 1 An electronic device 101 is illustrated according to some examples of this disclosure, which presents an extended reality (XR) environment (e.g., a computer-generated environment that optionally includes representations of physical and / or virtual objects). In some examples, such as Figure 1 As shown, electronic device 101 is a head-mounted display or other head-mountable device configured to be worn on the head of a user of electronic device 101. See below for reference. Figure 2A An example of an architecture block diagram to describe electronic device 101. For example... Figure 1 As shown, electronic device 101 and various objects (discussed in further detail below) reside in a physical environment (denoted herein as 3D environment 130). 3D environment 130 may include physical features such as physical surfaces (e.g., floor, wall) or physical objects (e.g., table, lamp, etc.). In some examples, electronic device 101 may be configured to detect and / or capture images of the physical environment, including table 310 (refer to below). Figures 3A to 3R (Example in the field of view of the electronic device 101 under discussion).

[0017] In some examples, such as Figure 1 As shown, electronic device 101 includes one or more internal image sensors 114a oriented toward the user's face (e.g., referred to below). Figures 2A to 2B(The described eye-tracking camera). In some examples, an internal image sensor 114a is used for eye tracking (e.g., detecting the user's gaze). The internal image sensor 114a is optionally arranged on the left and right portions of the display 120 to enable eye tracking of the user's left and right eyes. In some examples, the electronic device 101 also includes external image sensors 114b and 114c facing outwards from the user to detect and / or capture the three-dimensional environment of the electronic device 101 and / or movement of the user's hands or other body parts.

[0018] In some examples, display 120 has a field of view visible to the user (e.g., it may correspond to or not correspond to the field of view of external image sensors 114b and 114c). Because display 120 is optionally part of a head-mounted device, the field of view of display 120 may be the same as or similar to the field of view of the user's eyes. In other examples, the field of view of display 120 may be smaller than the field of view of the user's eyes. In some examples, electronics 101 may be an optical pass-through device, through which display 120 is a transparent or translucent display through which parts of the three-dimensional environment can be viewed directly. In some examples, display 120 may be included within a transparent lens and may overlap with all or only a portion of the transparent lens. In other examples, electronics may be a video pass-through device, through which display 120 is an opaque display configured to display images of the three-dimensional environment captured by external image sensors 114b and 114c. Although a single display 120 is shown, it should be understood that display 120 may include a stereoscopic display pair. In some examples, the head-mounted device does not include a display 120 (e.g., optionally includes a transparent lens), and display functionality is achieved via electronic devices 160.

[0019] In some examples, electronic device 101 may be configured to communicate with a second electronic device, such as a companion device. For example, as Figure 1 As illustrated, electronic device 101 can communicate with handheld electronic device 160. In some examples, handheld electronic device 160 corresponds to mobile electronic device, such as a smartphone, tablet computer, smartwatch, or other electronic device. See below for reference. Figure 2B The following architectural block diagram is used to describe additional examples of the handheld electronic device 160. In some examples, electronic device 101 and handheld electronic device 160 are associated with the same user. For example, in Figure 1In this configuration, electronic device 101 may be positioned (e.g., mounted) on a user's head, and handheld electronic device 160 may be positioned near electronic device 101, such as in the user's hand 103 (e.g., hand 103 is holding handheld electronic device 160), and electronic device 101 and handheld electronic device 160 are associated with the same user account (e.g., the user is logged into a user account on both electronic device 101 and handheld electronic device 160). See below for further details. Figures 2A to 2B Additional details regarding the communication between electronic device 101 and handheld electronic device 160 are provided. Although primarily described herein as a handheld electronic device, it should be understood that handheld electronic device 160 may be a non-handheld device.

[0020] In some examples, when presenting a 3D environment that includes one or more physical objects, a user of the head-mounted device can initiate interaction with one or more physical objects in the 3D environment. In some examples, the interaction may include a user query. In some examples, the interaction may include additional input associated with other input devices. For example, the user's gaze may be tracked by an electronic device as input for identifying regions of interest corresponding to one or more physical objects associated with the user query. Additionally or alternatively, in some examples, hand-tracking input may be used to identify regions of interest corresponding to one or more physical objects.

[0021] In the following discussion, an electronic device communicating with a display generating component and / or one or more input devices is described. It should be understood that the electronic device may optionally communicate with one or more other physical user interface devices, such as a touch-sensitive surface, physical keyboard, mouse, joystick, hand-tracking device, eye-tracking device, stylus, etc. Furthermore, as described above, it should be understood that the described electronic device, display generating component, and touch-sensitive surface may optionally be distributed among two or more devices. It should be understood that in some examples, the electronic device does not include a display generating component or a display. Therefore, as used in this disclosure, information on or displayed by the electronic device may optionally be used to describe information output by the electronic device for display on a separate display device (touch-sensitive or non-touch-sensitive). Similarly, as used in this disclosure, input received on the electronic device (e.g., touch input received on a touch-sensitive surface of the electronic device, or touch input received on the surface of a stylus) may optionally be used to describe input received on a separate input device from which the electronic device receives input information.

[0022] The electronic devices described herein can support a variety of applications. For example, one or more input devices can be used to generate input for interacting with one or more applications, and / or one or more displays can be used to display the applications and associated user interfaces. One or more applications may include one or more of the following: drawing applications, presentation applications, word processing applications, website creation applications, disk editing applications, spreadsheet applications, game applications, telephone applications, video conferencing applications, email applications, instant messaging applications, fitness support applications, photo management applications, digital camera applications, digital video camera applications, web browsing applications, digital music player applications, TV channel browsing applications, and / or digital video player applications.

[0023] Figures 2A to 2B Block diagrams illustrating example architectures for electronic devices 201 and 260 according to some examples of this disclosure are shown. In some examples, electronic device 201 and / or electronic device 260 include one or more electronic devices. For example, electronic device 201 may be a portable device, an auxiliary device for communicating with another device, a head-mounted display, a head-mounted device, etc. In some examples, electronic device 201 corresponds to the above reference. Figure 1 The described electronic device 101. In some examples, electronic device 260 corresponds to the above reference. Figure 1 The handheld electronic device 160 is described.

[0024] like Figure 2A As illustrated, electronic device 201 may optionally include various sensors, such as one or more hand tracking sensors 202, one or more position sensors 204A, and one or more image sensors 206A (optionally corresponding to...). Figure 1 The internal image sensor 114a and / or external image sensors 114b and 114c, one or more touch-sensitive surfaces 209A, one or more motion and / or orientation sensors 210A, one or more eye-tracking sensors 212, one or more microphones 213A or other audio sensors, one or more body tracking sensors (e.g., torso and / or head tracking sensors), and one or more display generating components 214A (optionally corresponding to...) Figure 1 The electronic device 201 includes a display 120, one or more speakers 216A, one or more processors 218A, one or more memories 220A, and / or communication circuitry 222A. One or more communication buses 208A are optionally used for communication between the aforementioned components of the electronic device 201. Additionally, as... Figure 2BAs shown, electronic device 260 optionally includes one or more position sensors 204B, one or more image sensors 206B, one or more touch-sensitive surfaces 209B, one or more orientation sensors 210B, one or more microphones 213B, one or more display generating components 214B, one or more speakers 216B, one or more processors 218B, one or more memories 220B, and / or communication circuitry 222B. One or more communication buses 208B are optionally used for communication between the aforementioned components of electronic device 260. Electronic devices 201 and 260 are optionally configured to communicate via a wired or wireless connection between the two electronic devices (e.g., via communication circuitry 222A, 222B). For example, as Figure 2A As indicated, electronic device 260 can be used as an accessory device to electronic device 201.

[0025] Communication circuits 222A and 222B optionally include circuitry for communicating with electronic devices and networks such as the Internet, intranets, wired and / or wireless networks, cellular networks, and wireless local area networks (LANs). Communication circuits 222A and 222B optionally include circuitry for using near-field communication (NFC) and / or short-range communication such as Bluetooth. ® The circuit used for communication.

[0026] Processors 218A and 218B include one or more general-purpose processors, one or more graphics processors, and / or one or more digital signal processors. In some examples, memory 220A or 220B is a non-transitory computer-readable storage medium (e.g., flash memory, random access memory, or other volatile or non-volatile memory or storage device) storing a computer-readable program including instructions configured to be executed by processor 218A or 218B to perform the techniques, processes, and / or methods described below. In some examples, memory 220A and / or 220B may include more than one non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium may be any medium (e.g., excluding signals) that can tangibly contain or store computer-executable instructions for use by or in conjunction with an instruction execution system, apparatus, or device. In some examples, the storage medium is a transient computer-readable storage medium. In some examples, the storage medium is a non-transitory computer-readable storage medium. A non-transitory computer-readable storage medium may include, but is not limited to, magnetic storage devices, optical storage devices, and / or semiconductor storage devices. Examples of such storage devices include hard disks, optical discs based on compressed disc (CD), digital versatile optical disc (DVD), or Blu-ray technology, as well as persistent solid-state storage (such as flash memory, solid-state drives, etc.).

[0027] In some examples, display generating components 214A, 214B include a single display (e.g., a liquid crystal display (LCD), an organic light-emitting diode (OLED), or other types of display). In some examples, display generating components 214A, 214B include multiple displays. In some examples, display generating components 214A, 214B may include a touch-enabled display (e.g., a touchscreen), a projector, a holographic projector, a retinal projector, a transparent or translucent display, etc. In some examples, electronic devices 201 and 260 respectively include touch-sensitive surfaces 209A and 209B for receiving user input such as tap input and swipe input or other gestures. In some examples, display generating components 214A, 214B and touch-sensitive surfaces 209A, 209B form a touch-sensitive display (e.g., a touchscreen integrated with each of electronic devices 201 and 260 or a touchscreen external to each of electronic devices 201 and 260 that communicates with each of electronic devices 201 and 260).

[0028] Electronic devices 201 and 260 optionally include image sensor 206A and image sensor 206B, respectively. Image sensors 206A and 206B optionally include one or more visible light image sensors (such as charge-coupled device (CCD) sensors) and / or complementary metal-oxide-semiconductor (CMOS) sensors operable to acquire images of physical objects from a real-world environment. Image sensors 206A and 206B also optionally include one or more infrared (IR) sensors, such as passive or active IR sensors, for detecting infrared light from the real-world environment. For example, an active IR sensor includes an IR emitter for emitting infrared light into the real-world environment. Image sensors 206A and 206B also optionally include one or more cameras configured to capture movement of a physical object in the real-world environment. Image sensors 206A and 206B also optionally include one or more depth sensors configured to detect the distance between the physical object and electronic devices 201 and 260. In some examples, information from one or more depth sensors allows a device to identify objects in a real-world environment and distinguish them from other objects in the real-world environment. In some examples, one or more depth sensors allow a device to determine the texture and / or shape of objects in a real-world environment.

[0029] In some examples, electronic devices 201 and 260 combine a CCD sensor, an event camera, and a depth sensor to detect the three-dimensional environment surrounding them. In some examples, image sensors 206A and 206B include a first image sensor and a second image sensor. The first and second image sensors work cooperatively and are optionally configured to capture different information about physical objects in the real-world environment. In some examples, the first image sensor is a visible light image sensor, and the second image sensor is a depth sensor. In some examples, electronic devices 201 and 260 use image sensors 206A and 206B to detect the location and orientation of electronic devices 201 and 260 and / or display generation components 214A and 214B in the real-world environment. For example, electronic devices 201 and 260 use image sensors 206A and 206B to track the location and orientation of display generation components 214A and 214B relative to one or more fixed objects in the real-world environment.

[0030] In some examples, electronic devices 201 and 260 include microphones 213A and 213B or other audio sensors, respectively. Electronic devices 201 and 260 may optionally use microphones 213A and 213B to detect sound from a user and / or the user's real-world environment. In some examples, microphones 213A and 213B include arrays (multiple microphones) of microphones that optionally operate in cooperation, such as to identify ambient noise or locate sound sources in the space of a real-world environment.

[0031] In some examples, electronic devices 201 and 260 include position sensors 204A and 204B, respectively, for detecting the positions of electronic device 201A and / or display generating component 214A and the positions of electronic device 260 and / or display generating component 214B. For example, position sensors 204A and 204B may include Global Positioning System (GPS) receivers that receive data from one or more satellites and allow electronic devices 201 and 260 to determine the absolute location of the device in the physical world.

[0032] In some examples, electronic devices 201 and 260 include orientation sensors 210A and 210B, respectively, for detecting the orientation and / or movement of electronic device 201 and / or display generating component 214A, and the orientation and / or movement of electronic device 260 and / or display generating component 214B, respectively. For example, electronic devices 201 and 260 use orientation sensors 210A and 210B to track changes in the positioning and / or orientation of electronic devices 201 and 260 and / or display generating components 214A and 214B, such as changes relative to physical objects in a real-world environment. Orientation sensors 210A and 210B may optionally include one or more gyroscopes and / or one or more accelerometers.

[0033] In some examples, electronic device 201 includes a hand tracking sensor 202 and / or an eye tracking sensor 212 (and / or other body tracking sensors, such as a leg tracking sensor, a torso tracking sensor, and / or a head tracking sensor). The hand tracking sensor 202 is configured to track the localization / position of one or more portions of a user's hand, and / or the movement of one or more portions of the user's hand relative to the extended reality environment, relative to the display generating component 214A, and / or relative to another defined coordinate system. The eye tracking sensor 212 is configured to track the localization and movement of the user's gaze (more generally, the eyes, face, or head) relative to the real world or the extended reality environment and / or relative to the display generating component 214A. In some examples, the hand tracking sensor 202 and / or the eye tracking sensor 212 are implemented together with the display generating component 214A. In some examples, the hand tracking sensor 202 and / or the eye tracking sensor 212 are implemented separately from the display generating component 214A. In some examples, electronic device 201 optionally does not include the hand tracking sensor 202 and / or the eye tracking sensor 212. In some such examples, the generating component 214A is shown to be used by the electronic device 260 to provide an extended reality environment and utilizes input and other data collected via other sensors of the electronic device 201 (e.g., one or more position sensors 204A, one or more image sensors 206A, one or more touch-sensitive surfaces 209A, one or more motion and / or orientation sensors 210A, and / or one or more microphones 213A or other audio sensors) as input and data processed by the processor 218B of the electronic device 260. Additionally or alternatively, the electronic device 201 may optionally not include... Figure 2B Other components shown include position sensor 204B, image sensor 206B, touch-sensitive surface 209B, etc. In some such examples, it is shown that generating component 214A can be used by electronic device 260 to provide an extended reality environment, and electronic device 260 uses input and other data collected via one or more motion and / or orientation sensors 210A (and / or one or more microphones 213A) of electronic device 201 as input.

[0034] In some examples, the hand tracking sensor 202 (and / or other body tracking sensors, such as leg tracking sensors, torso tracking sensors, and / or head tracking sensors) may use image sensors 206 (e.g., one or more IR cameras, 3D cameras, depth cameras, etc.) that capture 3D information from the real world, including one or more body parts (e.g., a human user's hand, leg, or torso). In some examples, sufficient resolution is available to distinguish the hand to differentiate the fingers and their corresponding positions. In some examples, one or more image sensors 206A are positioned relative to the user to define the field of view and interaction space of the image sensors 206A, in which the finger / hand positions, orientations, and / or movements captured by the image sensors are used as input (e.g., to differentiate from the user's resting hand or other hands of other people in the real-world environment). Tracking fingers / hands to achieve input (e.g., gestures, touches, taps, etc.) may be advantageous because it does not require the user to touch, hold, or wear any type of beacon, sensor, or other marker.

[0035] In some examples, the eye-tracking sensor 212 includes at least one eye-tracking camera (e.g., an infrared (IR) camera) and / or an illumination source (e.g., an IR light source, such as an LED) that emits light toward the user's eyes. The eye-tracking camera may be pointed at the user's eyes to receive reflected IR light from the light source directly or indirectly from the eyes. In some examples, both eyes are tracked separately by the respective eye-tracking camera and illumination source, and focus / gaze can be determined by tracking both eyes. In some examples, one eye (e.g., the dominant eye) is tracked by one or more respective eye-tracking cameras / illumination sources.

[0036] Electronic devices 201 and 260 are not limited to Figures 2A to 2B The components and configurations may include fewer, additional, or supplementary components in various configurations. In some examples, electronic device 201 and / or electronic device 260 may each be implemented among multiple electronic devices (e.g., as a system). In some such examples, each(or multiple) electronic device may each include one or more of the same components discussed above, such as various sensors, one or more display generation components, one or more speakers, one or more processors, one or more memories, and / or communication circuitry. One or more persons using electronic device 201 and / or electronic device 260 may optionally be referred to herein as one or more users of the device. In some examples, electronic device 201 does not include a display, and electronic device 260 includes a display.

[0037] Now turn attention to interaction with one or more objects in the three-dimensional environment 130. One or more input devices of an electronic device (e.g., corresponding to electronic device 201) may be used to support the interaction. As described herein, the interaction may include user queries (e.g., natural language requests based on text or audio) and / or may include one or more images, which may optionally include one or more images captured by a camera and / or one or more subsets of images based on the user's gaze.

[0038] Figure 3A An example is illustrated by electronic device 101 presenting features corresponding to the physical environment (e.g., as referenced above). Figure 1 A three-dimensional environment 130 containing multiple physical objects within a discussed physical environment. In some examples, the multiple objects include a table 310 centrally positioned within the field of view of an electronic device 101 in the three-dimensional environment 130. In some examples, the table 310 optionally includes various cooking ingredients and / or cooking appliances. In some examples, the various cooking ingredients include apples 311, carrots 312, and pasta 313 on a first portion of the table 310. In some examples, such as... Figure 3A As shown, the first portion of the table 310 on which the aforementioned cooking ingredients are positioned corresponds to the top portion (e.g., surface) of the table 310, which is to the right of the center of the field of view of the electronic device 101. In some examples, the table 310 further includes a pot 314 on a second portion of the table 310 (e.g., different from the first portion). In some examples, such as Figure 3A As shown, pot 314 includes one or more ingredients suspended in a solution (e.g., chicken broth). In some examples, such as Figure 3A As shown, the second portion of the table 310 on which the pot 314 is positioned corresponds to the top portion (e.g., surface) of the table 310, which is on the left side of the field of view of the electronic device 101. The first and second portions of the table 310 are not necessarily limited to the right and left sides of the center of the field of view of the electronic device 101, respectively, and may optionally be displayed in various alternative combinations of positioning on the table 310 relative to the viewpoint of the electronic device 101.

[0039] In some examples, such as Figure 3A As shown, the three-dimensional environment 130 includes a hairpin 301 placed on the floor of the physical environment. In some examples, the hairpin 301 may optionally correspond to any one of a variety of small hair-related devices. For example, the hairpin 301 may optionally be a headband, a curved hair clip, a claw clip, etc. In some examples, such as Figure 3A As shown, electronic device 101 is displayed on hairpin 301 on the floor of three-dimensional environment 130. This placement optionally corresponds to the lower right portion of the field of view of electronic device 101.

[0040] In some examples, the 3D environment 130 includes multiple objects set on walls corresponding to the physical environment of the 3D environment 130. In some examples, such as Figure 3A As shown, the wall includes a poster 340 located at the upper left portion of the three-dimensional environment 130 relative to the viewpoint of the electronic device 101. The poster 340 may optionally include details corresponding to a concert (e.g., such as...). Figure 3A The images of the drummer and singer shown, and the website address 341 associated with the poster. In some examples, electronic device 101 responds to the reference below. Figure 3O Further details of user input (e.g., user gaze, user hand movement) are provided to perform the action associated with website address 341. In some examples, such as Figure 3A As shown, the wall of the physical environment includes a lower shelf 320 and an upper shelf 330 mounted on the right side of the wall relative to the viewpoint of the electronic device 101. Multiple books can be placed on each shelf of the respective unit. For example, as... Figure 3A As shown, the lower shelf 320 includes books 320a-320l, and the upper shelf 330 includes books 330a-330j. In some examples, the arrangement of the corresponding books on the corresponding shelves is not restrictive and can be arranged in any particular order relative to the corresponding groups of books on the corresponding shelves.

[0041] Figure 3B An example is illustrated where electronic device 101 detects a user gaze 360 ​​pointing towards pot 314, and performs cropping (e.g., cropping 350) of an image of the 3D environment 130 when the user gaze 360 ​​is pointing towards pot 314. In some examples, electronic device 101 detects the user gaze 360 ​​via one or more input devices. For example, one or more input devices correspond to... Figure 1 One or more internal image sensors 114a are used to detect the direction of the user's gaze 360. In some examples, when one or more internal image sensors 114a detect the direction of the user's gaze 360, the electronic device 101 correlates the direction of the gaze 360 ​​with physical objects in the three-dimensional environment 130 via external image sensors 114b and / or 114c. In some examples, the gaze 360 ​​points to one or more physical objects, as discussed in further detail below. In some examples, the gaze 360 ​​points to a single physical object, such as... Figure 3B As shown. In some examples, electronic device 101 crops a still image of the 3D environment 130 to produce crop 350. In some examples, electronic device 101 crops a live video feed of the 3D environment 130 to produce crop 350. In some examples, electronic device 101 performs cropping in response to any one of one or more internal image sensors 114a-114c detecting a user gaze 360. For example, as Figure 3BAs shown, one or more internal image sensors 114a-114c detect a gaze 360 ​​pointing towards the center point of the pot 314, while the electronics 101 delineates a sub-section of the three-dimensional environment 130 corresponding to the crop 350 (e.g., by...). Figure 3B (The dashed box shown is an example). In some examples, the electronic device 101 performs clipping after gaze 360 ​​is detected. In some examples, clipping 350 is generated concurrently with the detection of gaze 360. In some examples, the electronic device generates clipping 350 based on a predetermined radius and / or distance from gaze 360. It should be understood that clipping 350 can be various shapes (e.g., circles, squares, triangles, stars, etc.) corresponding to sub-parts of the three-dimensional environment 130, and is not necessarily limited to such shapes. Figure 3B The illustrated rectangular shape.

[0042] In some examples, electronic device 101 detects a physical object corresponding to the direction of gaze 360 ​​and determines a sub-section of the 3D environment 130 that encapsulates the entire physical object. In some examples, electronic device 101 performs clipping 350 based on user input, discussed in further detail below. In some examples, electronic device 101 processes an image of the 3D environment 130 and the direction of user gaze 360 ​​before determining one or more boundaries of clipping 350. At least one of the aforementioned inputs may optionally be processed by a large language learning model to determine the sub-section referred to below. Figure 3C The cropping 350 involves one or more boundaries, discussed in further detail. In some examples, electronic device 101 processes an image of the 3D environment 130 and determines one or more boundaries of the cropping 350 via a machine learning model (e.g., a neural network, deep learning, etc.) at electronic device 101. In some examples, electronic device 101 sends an image of the 3D environment 130 to an auxiliary electronic device (not shown), such as a server, desktop computer, and / or cloud-based electronic services. Optionally, a machine learning model, incorporating one or more features of the machine learning model discussed above, is stored at this auxiliary electronic device and configured to process the image of the 3D environment 130 and determine one or more boundaries of the cropping 350.

[0043] Figure 3C Examples of the detection of a user voice command 370 are illustrated according to some examples of this disclosure, when the user gazes 360 toward the pot 314 within the crop 350 in the three-dimensional environment 130, and when the handheld electronic device 160 processes the captured image of the three-dimensional environment 130, the crop 350, and the user voice command 370. In some examples, Figure 3B and Figure 3C They occur concurrently. In some examples, such as... Figure 3CAs shown, electronic device 101 detects voice command 370 simultaneously with detecting the user's gaze direction 360. In some examples, voice command 370 corresponds to a voice command spoken by the user of electronic device 101. In some examples, voice command 370 optionally corresponds to a voice command spoken by a secondary user different from the user of electronic device 101. In some examples, voice command 370 is derived from the above reference. Figures 2A to 2B The microphones 213A and 213B are discussed and detected, and sent as input data to the handheld electronic device 160.

[0044] In some examples, electronic device 101 sends data corresponding to the three-dimensional environment 130, including images, cropping 350, and voice commands 370, to handheld electronic device 160. In some examples, handheld electronic device 160 includes the data referenced above. Figure 3B The discussion concerns at least one or more characteristics of the auxiliary electronic device. In some examples, electronic device 101 processes the aforementioned inputs as inputs 160a-160c. In some examples, such as Figure 3C As shown, when the user gazes 360° towards the pot 314, the handheld electronic device 160 processes inputs 160a-160c. In some examples, the handheld electronic device 160 processes inputs 160a-160c via an internal machine learning model. Specifically, input 160c optionally corresponds to a voice command 370 detected by electronic device 101 and is optionally processed by a large language learning model at the handheld electronic device 160. In some examples, the handheld electronic device 160 processes inputs 160a-160c via a machine learning model stored at the handheld electronic device 160. In some examples, the handheld electronic device 160 sends one or more of the inputs 160a-160c to be processed at a third electronic device (not shown). The remaining inputs may optionally be processed at the handheld electronic device 160.

[0045] Figure 3D Examples of this disclosure illustrate the detection of a user voice command 370 paired with a hand press 380 (or other touch input) when the user gazes 360 toward a pot 314 within a cutout 350 in the three-dimensional environment 130, and when a handheld electronic device 160 processes data corresponding to the three-dimensional environment 130, the cutout 350, and a user voice command 370 paired with a hand press 380. In some examples, Figure 3D An alternative example procedure for capturing and cropping 350 is shown above (see reference above). Figure 3C As summarized above. (See reference above) Figure 3C As discussed, electronic device 101 optionally detects user voice command 370 and, in response, captures an image of the three-dimensional environment 130 (e.g., additionally exemplified by input 160a). In some examples, such as Figure 3D As illustrated, electronic device 101 may optionally require an additional hand press 380 as a trigger to capture an image of the three-dimensional environment 130. In some examples, electronic device 101 concurrently detects hand press 380 and user voice command 370. In some examples, electronic device 101 determines hand press 380 as valid input if input is detected within a threshold time (e.g., 0.1 seconds, 0.25 seconds, 0.5 seconds, 0.75 seconds, 1 second, etc.) of the detected user voice command 370 (e.g., before or after detection). This hand press 380 may optionally correspond to a user of electronic device 101, but is not limited to any particular user. For example, electronic device 101 may optionally detect hand press 380 performed by a user of a third electronic device.

[0046] Figure 3E Alternative example processes are illustrated, according to some examples of this disclosure, for detecting a user command (e.g., similar to a voice command 370) via text input 377 when the user gazes 360 toward a pot 314 enclosed by clipping 350 within a three-dimensional environment 130. In some examples, Figure 3E The alternative process shown includes, for example: Figure 3C and / or Figure 3D The process of detecting user commands is illustrated by one or more characteristics. In some examples, when the user gazes 360 towards the pot 314, the handheld electronic device 160 detects text input 377 directed to the keyboard, such as... Figure 3E As illustrated. In response to detecting this input, the handheld electronic device 160 optionally displays a visual representation 377a of the text input 377. For example, as Figure 3E As shown, the handheld electronic device 160 optionally detects text input 377 on its numeric keypad and, in response, displays a visual representation 377a of the text input 377 at the handheld electronic device 160 (e.g., "Set the timer to 1 hour and 25 minutes"). In some examples, the text input 377 includes one or more features of the user voice command 370 as discussed above. In some examples, the handheld electronic device 160 does not display a visual representation 3757 of the text input 377, but instead refers to the above... Figure 3D Input 160c is processed similarly to the corresponding command. For example, handheld electronic device 160 processes inputs 160a-160b as described above, and processes text input 377 as input 160c in a manner similar to the user voice command 370 illustrated above. As a result of the inputs and commands described above, handheld electronic device 160 may optionally perform operations associated with the inputs and commands, as discussed in further detail below.

[0047] Figure 3FExamples of this disclosure illustrate timer 162 associated with a dish 314 presented at a handheld electronic device 160 in response to user input (e.g., user voice command 370, text input 377) when the electronic device 101 presents a three-dimensional environment 130. In some examples, timer 162 automatically begins counting down from the time set by the input discussed above in response to the handheld electronic device 160 processing inputs 160a-160c. In some examples, such as Figure 3F As shown, timer 162 is displayed as a text box set in the upper portion of the handheld electronic device 160. In some examples, the handheld electronic device 160 activates timer 162 associated with dish 314 but does not display timer 162. In some examples, the handheld electronic device transmits timer 162 to electronic device 101 for storing and / or transmitting commands for performing an operation (e.g., running timer 162) at electronic device 101. In some examples, while the handheld electronic device 160 displays timer 162, the electronic device continues to detect the direction of the user's gaze (e.g., user gaze 360). In some examples, electronic device 101 and / or handheld electronic device 160 are configured to perform multiple operations in response to user voice command 370 or text input 377. For example, when electronic device 101 and / or handheld electronic device 160 runs timer 162, the electronic device may optionally detect user gaze 360, as discussed in further detail below. In another example, when electronic device 101 and / or handheld electronic device 160 run timer 162, the electronic device may optionally perform the actions described in the reference above. Figures 3A to 3F Any one of the operations in the overview.

[0048] Figure 3G Examples of some examples according to this disclosure illustrate an information graphical user interface 163 associated with a pot 314 within a cutout 350 in the three-dimensional environment 130 when the electronic device 101 presents the three-dimensional environment 130 and when a user gaze 360 ​​is detected. The information graphical user interface 163 is presented at the handheld electronic device 160 in response to a user voice command 371 (e.g., “What can I use this to cook?”). In some examples, such as Figure 3G As shown, electronic device 101 detects user voice command 371 and, in response, performs cropping 350 on the image of the 3D environment 130, similar to the cropping operation discussed above. In some examples, electronic device 101 sends data corresponding to the image of the 3D environment 130, user voice command 371, and cropping 350 to handheld electronic device 160 for use in conjunction with the above reference. Figure 3D Similar processing methods are discussed. In some examples, the image of the 3D environment 130, the user voice command 371, and the cropping 350 correspond to the above reference. Figure 3DThe inputs discussed are 160a-160c. In some examples, the handheld electronic device 160 processes the aforementioned inputs and, in response, presents the information graphical user interface 163 described below.

[0049] In some examples, such as Figure 3G As shown, the information graphical user interface 163 includes information related to the pot 314 and / or its contents (e.g., "recipe ideas including chicken soup"). For example, after determining that the user's voice command 371 is associated with the pot 314 within the cutter 350, the handheld electronic device 160 presents a list of recipes associated with chicken soup (e.g., "Recipe A", "Recipe B") within the information graphical user interface 163. Figure 3G As shown. In some examples, the information presented at the information graphical user interface 163 associated with pot 314 includes, but is not necessarily limited to, ingredients present within the three-dimensional environment 130. For example, electronic device 101 may optionally detect (e.g., via image sensors 114b and 114c) pasta 313, carrot 312, and / or apple 311 within the three-dimensional environment 130, and may optionally include these objects as ingredients in a suggested recipe displayed at the information graphical user interface 163. In some examples, the information graphical user interface includes one or more items that do not correspond to one or more physical objects (e.g., apple 311) within the three-dimensional environment 130. For example, handheld electronic device 160 may optionally determine the existence of a shopping list (not shown), which may optionally be stored at handheld electronic device 160 (e.g., stored in the memory of handheld electronic device 160 and / or associated with an application of handheld electronic device 160, such as a note-taking and / or text editing application or photo application), and may optionally present one or more recipes including one or more ingredients from the shopping list at the information graphical user interface 163.

[0050] Figure 3HExamples of the present disclosure illustrate an electronic device 101 that detects a user voice command 372 when a user gazes 360 toward multiple objects (e.g., spaghetti 313, carrot 312, apple 311) in a three-dimensional environment 130 and performs a cropping 351 of an image of the three-dimensional environment 130. A handheld electronic device 160 is also illustrated, processing an image of the three-dimensional environment 130, including cropping 350 of multiple objects, and the user voice command 372. In some examples, the handheld electronic device 160 uses a machine learning model as previously discussed above to process the voice command 372 (e.g., “What are these?”). Using this model, the electronic device 101 can optionally determine, using the direction of the user gaze 360, the cropping 351, and the voice command 372, that the user's detected voice command 372 most likely refers to spaghetti 313, carrot 312, and apple 311, and performs subsequent operations associated with the aforementioned items. In some examples, electronic device 101 performs clipping 351 based on a predetermined radius and / or a distance from the center of the direction of the user's gaze 360, which includes the spaghetti 313, carrot 312, and apple 311. In some examples, electronic device 101 determines the boundaries of clipping 351 based on the detection of one or more objects near the direction of the user's gaze 360. For example, as... Figure 3H As shown, electronic device 101 optionally detects spaghetti 313, carrot 312, and apple 311, and optionally determines the distance from each object to the direction of the user's gaze 360. If the corresponding object is within a threshold distance from the direction of the user's gaze 360 ​​and optionally within a threshold distance between each corresponding object, electronic device 101 optionally includes the identified object in crop 351. In some examples, electronic device 101 sends data corresponding to the image of the three-dimensional environment 130, crop 351, and user voice command 372 for processing as input 160a to 160c at handheld electronic device 160. In some examples, in response to processing input 160a to 160c, handheld electronic device 160 performs an operation associated with user voice command 372, as discussed in further detail below.

[0051] Figure 3I Alternative examples are illustrated where electronic device 101, according to some examples of this disclosure, presents a three-dimensional environment 130 while handheld electronic device 160 presents a list 164 of items associated with the multiple items discussed above. In some examples, such as Figure 3I As shown, project list 164 includes multiple projects (e.g., Figure 3H The representation of pasta 313, carrots 312, and apples 311, along with information associated with each corresponding item. For example, as... Figure 3HAs shown, the detection of user voice command 372 initiates operations describing multiple items. In response to processing inputs 160a to 160c, as described above, the handheld electronic device 160 optionally displays a representation of Apple 311 and optionally includes sub-sections of online websites associated with Apple 311 (e.g., online encyclopedias, FDA nutrition guides, etc.). In some examples, such as... Figure 3I As shown, item list 164 includes corresponding descriptions associated with each of the multiple items described above, in a non-specific order. In some examples, item list 164 includes one or more hyperlinks associated with the multiple items, the hyperlinks being configured to receive user input, such as selection of one or more hyperlinks.

[0052] Figures 3J to 3K Examples of the present disclosure are illustrated when a user gazes 360 toward objects (e.g., carrot 312 and hairpin 301) within cropped regions (e.g., cropped regions 352 and 353, respectively) in the three-dimensional environment 130, and examples of the detection of a user voice command 373 when a handheld electronic device 160 processes an image of the three-dimensional environment 130, a cropped image of the object (cropped region 352 or cropped region 353), and a user voice command 373.

[0053] In some examples, electronic device 101 detects the direction of the user's gaze 360 ​​as pointing in the direction previously referenced above. Figure 3H The discussion focuses on the region associated with cropping 351. In some examples, such as... Figure 3J The directions of the user's gaze at 360 degrees shown include, for example: Figure 3H One or more characteristics of the user's gaze direction at 360 degrees, as previously shown. In some examples, such as... Figure 3J As shown, the electronic device 101 determines via a machine learning model that the user's voice command 373 corresponds to a carrot 312 and generates a crop 352 from an image of the three-dimensional environment 130. For example, Figure 3H User voice command 372 may optionally include the phrase "these," while Figure 3J The user voice command 373 optionally includes the phrase "this". Through a combination of a machine learning model and the direction of the user's gaze 360, the electronic device 101 is configured to optionally detect the user's intention to select a set of objects, such as... Figure 3H As shown, or the intent to select an object from a set of objects, such as Figure 3JAs shown. In some examples, electronic device 101 sends data corresponding to the three-dimensional environment 130, including a crop 352 of carrot 312 and user voice commands 373, as inputs 160a to 160c to handheld electronic device 160 for processing and implementation of subsequent operations in a manner similar to that described above.

[0054] In some examples, such as Figure 3K As shown, the electronic device 101 determines the direction of the user's gaze 360 ​​as towards the hairpin 301 and performs cropping 353 on the image of the three-dimensional environment 130. For example, the electronic device 101 detects that the gaze 360 ​​has moved from pointing to the carrot 312 to pointing to the hairpin 301. In some examples, the electronic device 101 detects a portion of the user's voice command 373 (e.g., "this") and uses it in accordance with the above reference. Figure 3J The process outlined in the overview establishes an association with the card issuer 301 in a similar manner. In some examples, the electronic device 101 is associated with the referenced above. Figure 3J A similar manner is described, sending data corresponding to the 3D environment 130, cropping 353, and user voice commands 373 to the handheld electronic device 160. In some examples, the handheld electronic device 160 processes inputs 160a to 160c and performs operations associated with the user voice commands 373, such as... Figure 3L exemplified.

[0055] In some examples, such as Figure 3L As shown, electronic device 101 presents a three-dimensional environment 130, while handheld electronic device 160 presents an information graphical user interface 165 associated with the card issuer 301. In some examples, the information graphical user interface 165 includes the above-mentioned references. Figure 3G The discussion concerns one or more features of the information graphical user interface 163. For example, such as... Figure 3G and Figure 3L As shown, the handheld electronic device 160 may optionally display a representation of an object (e.g., a hair clip) within a correspondingly cropped image (e.g., crop 350, crop 353) and text information associated with the corresponding object (e.g., information identifying hairpin 301 as a hair clip). In some examples, such as Figure 3L As shown, electronic device 101 displays a three-dimensional environment 130, and handheld electronic device 160 concurrently displays an information graphical user interface 165, but may optionally display it during non-concurrent times.

[0056] Figure 3MAn electronic device 101, according to some examples of this disclosure, detects a user voice command 374 and performs cropping (e.g., cropping 354) of an image of a three-dimensional environment 130 based on the direction of the user's gaze 360 ​​corresponding to the poster 340, while a handheld electronic device 160 processes the image of the three-dimensional environment 130, cropping 354, and the user voice command 374. In some examples, such as Figure 3M As shown, the handheld electronic device 160 determines (optionally via any suitable machine learning algorithm) an association between at least a portion of the user's voice command 374 (e.g., "this") and text within a sub-part of a cropped section 354 (e.g., "November 25, 7pm-10pm"). In some examples, such as Figure 3M As shown, the electronic device 101 detects the direction of the user's gaze 360 ​​as pointing towards the poster 340 and generates a crop 354 encompassing the entire poster 340. In some examples, the electronic device 101 detects text in a sub-section of the crop 354 independently of the direction of the gaze 360. For example, as... Figure 3M As shown, electronic device 101 optionally detects the user's gaze direction 360 as pointing towards the band member illustrated by poster 340. In response, electronic device 101 optionally performs cropping 354 and optionally determines the text associated with the user's voice command 374 (e.g., "November 25th, 7pm-10pm"). Once the text is detected, electronic device 101 optionally sends data corresponding to a sub-part of cropping 354 as input 160b to handheld electronic device 160. In some examples, at least a portion of the user's voice command 374 corresponds to the following reference. Figure 4 Box 408 further discusses the “part of the input” in more detail. In some examples, the image of the 3D environment 130, crop 354, and user voice command 374 correspond to inputs 160a to 160c, and are processed in a similar manner to perform the actions described in the reference above. Figures 3A to 3L Similar operations are discussed below. In some examples, the handheld electronic device 160 performs operations associated with the poster 340, as discussed in further detail below.

[0057] Figure 3N Examples of operations associated with poster 340 performed at handheld electronic device 160 when electronic device 101 displays 3D environment 130, according to some examples of this disclosure, are illustrated. In some examples, due to receiving as referenced above... Figure 3M The discussion focuses on inputs 160a to 160c, where the handheld device adds and / or creates reminders 166 to the user's calendar. In some examples, this reminder is stored at the handheld electronic device 160, electronic device 101, and / or an auxiliary computer / service communicating with any of the electronic devices. In some examples, such as... Figure 3NAs shown, the handheld electronic device 160 displays the reminder 166 as a text entry, which includes the above reference. Figure 3M The text found within crop 354 is shown. In some examples, reminder 166 is added and stored at handheld electronic device 160, but not displayed. In some examples, in response to successful reminder creation, such as... Figure 3N As shown, the handheld electronic device 160 displays a notification (e.g., "Reminder Created") in the upper portion of its display. In some examples, the handheld electronic device 160 does not require user input to perform an operation associated with an object (e.g., poster 340), as discussed in further detail below.

[0058] Figure 3O An electronic device 101, according to some examples of this disclosure, detects that a user gaze 360 ​​is directed towards a website address 341 at a poster 340 within a cropped area (e.g., crop 355) in a three-dimensional environment 130, while a handheld electronic device 160 processes an image of the three-dimensional environment 130 and the website address 341. In some examples, in response to detecting that the user gaze 360 ​​is directed towards the website address 341, the handheld electronic device 160 automatically processes an image of the three-dimensional environment 130 (optionally corresponding to input 160a) and a cropped area 355 containing the website address 341 (optionally corresponding to input 160b), without requiring user input. In some examples, such as Figure 3P As shown, the handheld electronic device 160 displays a website 167 associated with website address 341 and / or poster 340. In some examples, such as Figure 3P As shown, website 167 includes interactive components (e.g., "Buy concert tickets here!") configured to initiate further actions associated with poster 340 (e.g., purchasing tickets). In some examples, when handheld electronic device 160 displays website 167, electronic device 101 performs actions related to... Figures 3A to 3N Any of the associated methods and / or operations.

[0059] Figure 3Q Examples of this disclosure illustrate how, when a user gazes 360 toward books 330a to 330j and performs a cropping (e.g., cropping 356) of a portion of the upper shelf 330 within an image of the 3D environment 130, and when the handheld electronic device 160 processes the image of the 3D environment 130, the book 330c within cropping 356, and the user's voice command 375, the electronic device 101 detects the user's voice command 375. In some examples, the foregoing items correspond to inputs 160a to 160c at the handheld electronic device 160. In some examples, the handheld electronic device 160 is referenced above... Figures 3A to 3OInputs 160a to 160c are processed in a similar manner as described. In some examples, the handheld electronic device 160 determines the corresponding book (e.g., book 330c) within the cutout 356 based on a combination of the user's voice command 375 and the direction of the user's gaze 360. In some examples, inputs 160a to 160c include... Figures 3A to 3O One or more characteristics of the inputs 160a to 160c associated with any of them. In some examples, in response to processing inputs 160a to 160c, the handheld electronic device 160 presents information associated with the book 330c, as discussed in further detail below.

[0060] Figure 3R An information graphical user interface 168 according to some examples of this disclosure is illustrated, which is associated with a book within a cropped area of ​​the three-dimensional environment 130 when the electronic device 101 presents the three-dimensional environment 130 and is presented at the handheld electronic device 160 in response to a user voice command. In some examples, such as Figure 3R As shown, the handheld electronic device 160 displays information associated with the author of the book 330c (e.g., in the information graphical user interface 168), while also providing information about the author related to previous operations performed by the handheld electronic device 160. For example, as Figure 3G As shown, the handheld electronic device 160 optionally presents an information graphical user interface 163, which optionally includes information associated with the pot 314. For example... Figure 3R As shown, the handheld electronic device 160 optionally presents an information graphical user interface 168, which optionally includes cooking-related information (e.g., recipe recommendations) associated with the author due to previous operations related to cooking. In some examples, in response to Figure 3Q Upon input 160a to 160c, the handheld electronic device 160 presents a webpage associated with the author of book 330c (e.g., "Jane Doe") at an information graphical user interface 168. In some examples, the handheld electronic device 160 presents information at an information graphical user interface 167 based on information stored on the device (e.g., the result of a previously performed related operation). Alternatively, in some examples, the handheld electronic device 160 responds only to receiving... Figure 3Q The information is presented in a graphical user interface 168 by inputting 160a to 160c.

[0061] Figure 4 This is a flowchart illustrating a method for performing operations based on a user query and a cropped image of an object, and an image of a 3D environment, according to some examples of this disclosure. This method may optionally be described in reference above. Figures 1 to 3RThe method is performed at the described electronic device (e.g., electronic device 101). Some operations in method 400 may be optionally combined, and / or the order of some operations may be optionally changed. In some examples, method 400 includes five steps (e.g., blocks 402 to 410).

[0062] In some examples, block 402 of method 400 involves detecting input according to some examples of this disclosure. In some examples, the input corresponds to the above reference. Figures 3A to 3R The described gaze 360°. In some examples, the detection step is facilitated by utilizing one or more internal image sensors 114a-114c positioned to capture input from the user. For example, one or more input devices may optionally correspond to the above reference. Figure 3B One or more internal image sensors 114a are discussed. In some examples, the input corresponds to the reference above. Figure 3D The discussion concerns the detection of hand pressure 380. This hand pressure 380 may optionally be performed by the hand 103 of a user, optionally corresponding to the electronic device 101. In some examples, the input corresponds to a sound detected by the electronic device 101. For example, such as... Figure 3C As shown, electronic device 101 detects a voice command 370 outlining an operation to be performed by electronic device 101 (e.g., "set the timer to 1 hour and 25 minutes"). In some examples, input detected by one or more input devices includes detecting one or more inputs performed by one or more different input devices. For example, electronic device 101 may optionally detect inputs as referenced above via one or more different input devices communicating with electronic device 101. Figure 3D The discussion covers hand pressure 380 and voice commands 370. This capture of input lays the foundation for subsequent steps of method 400 and the resulting various operations performed by electronic device 101 in box 410.

[0063] In some examples, according to method 400, box 404 relates to detecting a user's gaze direction according to some examples of this disclosure. In some examples, the gaze direction (e.g., user gaze 360) is detected by one or more input devices discussed above with reference to box 402. In some examples, the gaze direction includes one or more characteristics of user gaze 360 ​​as discussed above. In some examples, the gaze direction corresponds to one or more physical objects within the three-dimensional environment 130, as discussed above with reference to... Figure 3H As shown. In some examples, in response to the detection of gaze direction, electronic device 101 performs one or more operations as discussed in further detail below.

[0064] In some examples, according to method 400, block 406 relates to capturing one or more images of a three-dimensional environment 130 according to some examples of this disclosure. In some examples, electronic device 101 captures one or more images of the three-dimensional environment 130 via one or more input devices discussed in reference block 402 above. In some examples, the one or more images include one or more physical objects within the three-dimensional environment 130 discussed in reference block 404 above. In some examples, electronic device responds to as described in reference block 402 above. Figure 3D The discussion focuses on hand pressure 380° to capture one or more images. In some examples, the one or more images include at least the first image, which is discussed in further detail below.

[0065] In some examples, according to method 400, block 408 relates to identifying a subset of a first image in one or more images based on a detected gaze direction and detected input (e.g., cropping 350), according to some examples of this disclosure. In some examples, electronic device 101 identifies one or more physical objects within the subset based on a detected gaze direction (e.g., user gaze 360). In some examples, the detected input corresponds to any of the user voice commands 370 to 374 discussed above. In some examples, electronic device 101 identifies the gaze direction as associated with a region of the first image. For example, as... Figure 3Q As shown, the gaze direction may optionally point to the upper shelf 330, and in response, sub-sections (e.g., books 330a-330j) associated with the area including the upper shelf 330 are identified. In some examples, the electronic device also performs operations associated with sub-sections of the first image, as discussed in further detail in reference box 410 below.

[0066] In some examples, according to method 400, box 410 relates to performing an operation based on processed input (e.g., user voice command 372), one or more processed images (e.g., 3D environment 130), and a subset of the first image (e.g., cropping 350), according to some examples of this disclosure, based on determining that one or more criteria are met. In some examples, electronic device 101 performs an operation from at least a portion of the detected input (e.g., setting the above in...). Figure 3C Commands (such as the timer discussed above). For example, electronic device 101 may optionally detect user voice command 373 (e.g., “What is this?”) and correspond the command to perform an operation to card issuing 301, as referenced above. Figure 3KAs discussed, and in response, operations are performed at the handheld electronic device 160 to display information, such as using the information graphical user interface 165 associated with the card 301. In some examples, one or more criteria are met because the electronic device 101 successfully processes one or more inputs. In some examples, the electronic device 101 processes the input (e.g., user voice command 370) using a large language learning model as discussed above. In some examples, the electronic device processes the input, one or more images, and a sub-part of a first image in any order. In some examples, the electronic device processes the input, one or more images, and a sub-part of a first image concurrently. In some examples, the electronic device 101 and / or the handheld electronic device 160 determine that one or more criteria are met. In some examples, the handheld electronic device 160 determines that the processed input is considered valid input (e.g., a known command compared to a large language learning model) to meet one or more criteria. In some examples, the handheld electronic device 160 performs operations simultaneously while executing any of the boxes 402 to 408.

[0067] It should be understood that, Figure 4 The specific order in which the flowchart boxes are described is merely exemplary and is not intended to indicate that the described order is the only possible order in which these operations can be performed. Those skilled in the art will recognize that there are many ways to reorder the operations described herein.

[0068] In some examples, when an electronic device (e.g., electronic device 101) is connected to one or more input devices (e.g., Figure 1 When communicating with one or more internal image sensors 114a), the electronic device detects input via one or more input devices. In some examples, the electronic device detects the user's gaze direction via one or more input devices, such as those mentioned above. Figure 3B The discussion focuses on user gaze 360. In some examples, electronic devices capture one or more images via one or more input devices, such as inputs 160a and 160b, as... Figure 3C As shown. In some examples, electronic devices use the gaze direction and a portion of the input (e.g., Figure 3C The voice command 370 shown is used to identify at least a subset of the first image from one or more images, such as the one mentioned above. Figure 3C The discussion focuses on trimming 350. In some examples, it is determined, based on (e.g., by electronic device 101), that one or more criteria (e.g., detection) are met. Figure 3E The text input 377 shown above indicates that the electronic device performs operations, such as generating timer 162, based on the processing input, one or more images, and a subset of the first image via processing circuitry (e.g., processor 218A, processor 218B), as referenced above. Figure 3F The discussion.

[0069] In some examples, the electronic device crops a subset of the first image from the first image via processing circuitry (such as the reference above). Figure 3H The cropping discussed in section 351) identifies a subset of at least the first image (e.g., input 160a) from one or more images.

[0070] In some examples, electronic devices identify at least a subset of one or more images (e.g., crop 352) by identifying a predetermined region around the user's gaze direction (e.g., user gaze 360), as described above. Figure 3J As shown.

[0071] In some examples, the electronic device identifies a subset (e.g., crop 353) of at least one image (e.g., input 160a) from one or more images by identifying a region around the user's gaze direction, where the size of the region around the gaze direction is based on the distance of the user from one or more objects at the focal point of the user's gaze direction, such as crop 353 performed by the electronic device. Figure 3L As shown.

[0072] In some examples, the operation (e.g., generating timer 162 as discussed above) includes causing an auxiliary electronic device (e.g., a handheld electronic device 160) communicating with an electronic device (e.g., electronic device 101) via one or more output devices of the auxiliary electronic device (e.g., Figure 2B The display generation component 214B shown outputs information related to one or more objects included in a subset of the first image, such as information graphics 165, etc. Figure 3L As shown.

[0073] In some examples, the operation includes causing an auxiliary electronic device (e.g., a handheld device 160) communicating with an electronic device (e.g., electronic device 101) to launch an application based on one or more objects (e.g., poster 340) included in a subset of the first image, such as launching a display of website 167 at the application location. Figure 3P As shown.

[0074] In some examples, performing the operation includes scheduling future events or notifications corresponding to one or more objects included in a subset of the first image (e.g., poster 340), such as a handheld electronic device 160 displaying a reminder 166, as... Figure 3N As shown.

[0075] In some examples, the input includes language commands, such as Figure 3M The user voice command shown is 374.

[0076] In some examples, the language command corresponds to text input directed to an auxiliary electronic device (e.g., handheld electronic device 160) that communicates with the electronic device, such as text input 377, as referenced above. Figure 3E As shown and discussed.

[0077] In some examples, language commands (e.g., user voice command 374) correspond to signals generated by an audio sensor (such as...). Figure 2A and Figure 2B The audio input detected by the microphones 213A and 213B shown.

[0078] In some examples, the process of identifying a subset of the first image (e.g., 3D environment 130) (e.g., crop 354) is based on indicator pronouns in the audio input, such as electronic device 101 detecting "this" in user voice command 373, as... Figure 3K As shown.

[0079] In some examples, the first image is selected based on the time offset from when it detects an indicator pronoun (e.g., the user voice command 373 discussed above).

[0080] In some examples, the subset identifying at least the first image is based on the image segmentation model, the input, and the gaze orientation, such as Figure 3C The inputs 160a to 160c are shown in the table.

[0081] In some examples, one or more images are captured in response to an actuation (e.g., a hand pressing a button) or a touch sensor, such as... Figure 3D exemplified.

[0082] In some examples, capturing or selecting the first image (e.g., 160a) is based on detecting an indicator pronoun in the audio input, such as "this" in the user's voice command 371, as... Figure 3G As shown.

[0083] In some examples, an electronic device (e.g., a handheld electronic device 160) performs an operation by identifying a first subset of at least a first image (e.g., cropping 351), one or more first objects based on the first subset of at least the first image, and input (e.g., a user voice command 372), such as identifying spaghetti 313, carrots 312, and apples 311. Figure 3H and Figure 3I As shown. In some examples, the electronic device performs an operation, such as identifying a carrot 312, by performing a second operation different from the first operation based on a second subset (e.g., 352) that identifies at least a first image, one or more second objects that are different from one or more first objects based on the second subset of at least the first image, and input (e.g., user voice command 373). Figure 3J As shown.

[0084] In some examples, one or more images, a subset of the first images, and input are provided to accept one or more image inputs and one or more language inputs (such as...). Figure 3H The model shown has inputs 160a to 160c.

[0085] In some examples, the model is stored on an electronic device, such as... Figure 3H The electronic device 101 shown.

[0086] In some examples, the model is stored in an auxiliary electronic device (such as a handheld electronic device 160) that communicates with the electronic device. Figure 3H (as shown). In some examples, the electronic device sends the input, one or more images, and a subset of the first images to the auxiliary electronic device (e.g., handheld electronic device 160). In some examples, the electronic device (e.g., electronic device 101) receives the output of the model from the auxiliary electronic device.

[0087] In some examples, one or more criteria are included in processing the input, one or more images, and a subset of the first image (such as inputs 160a to 160c, etc.). Figure 3J (As shown) will provide the criteria that must be met when a request is made for one or more objects corresponding to a subset of the first image.

[0088] While the disclosed examples have been fully described with reference to the accompanying drawings, it should be noted that various changes and modifications will become apparent to those skilled in the art. It should be understood that such changes and modifications are considered to be included within the scope of the disclosed examples as defined by the appended claims.

[0089] For purposes of explanation, the foregoing description has been given by way of specific examples. However, the illustrative discussion above is not intended to be exhaustive or to limit the invention to the precise forms disclosed. Many modifications and variations are possible based on the teachings above. The examples were chosen and described in order to best elucidate the principles of the invention and its practical application, thereby enabling others skilled in the art to best utilize the invention with various modifications suitable for the particular intended use, as well as the various described examples.

Claims

1. A method, the method comprising: Electronic devices that communicate with one or more input devices: Input is detected via the one or more input devices; Detect the user's gaze direction via the one or more input devices; Capture one or more images via the one or more input devices; Use the gaze direction and a portion of the input to identify at least a subset of the first image from the one or more images; as well as Based on the determination that one or more criteria are met, operations are performed via processing circuitry based on the input, the one or more images, and the subset of the first image.

2. The method of claim 1, wherein identifying the subset of at least the first image among the one or more images comprises cropping the subset of the first image from the first image via the processing circuitry.

3. The method of claim 1, wherein identifying at least the subset of the first image among the one or more images includes identifying a predetermined region around the user's gaze direction.

4. The method of claim 1, wherein identifying at least a subset of the first image in the one or more images includes identifying a region around the user's gaze direction, wherein the size of the region around the gaze direction is based on the distance of the user from one or more objects at the focal point of the user's gaze direction.

5. The method of claim 1, wherein the operation includes causing an auxiliary electronic device communicating with the electronic device to output information relating to one or more objects included in the subset of the first image via one or more output devices of the auxiliary electronic device.

6. The method of claim 1, wherein the operation includes causing an auxiliary electronic device communicating with the electronic device to launch an application based on one or more objects included in the subset of the first image.

7. The method of claim 1, wherein performing the operation includes scheduling future events or notifications corresponding to one or more objects included in the subset of the first image.

8. The method of claim 1, wherein the input includes language commands.

9. The method of claim 8, wherein the subset identifying the first image is based on an indicator pronoun in the language command.

10. The method of claim 9, wherein the first image is selected based on the time offset from when the indicative pronoun is detected therefrom.

11. The method of claim 1, wherein identifying the subset of at least the first image is based on the image segmentation model, the input, and the gaze direction.

12. The method of claim 1, wherein the one or more images are captured in response to actuation of a button or touch sensor.

13. The method of claim 1, wherein capturing or selecting the first image is based on detecting an indicative pronoun in the input, wherein the input corresponds to an audio input.

14. The method of claim 1, wherein performing the operation comprises: A first operation is performed based on one or more first objects in the first subset of at least the first image and the input, according to the identification of a first subset of at least the first image; as well as Based on a second subset that identifies at least the first image and is different from the first subset, a second operation different from the first operation is performed based on one or more second objects in the second subset that are different from the one or more first objects and the input.

15. The method of claim 1, wherein the one or more images, the subset of the first images, and the input are provided to a model that accepts one or more image inputs and one or more language inputs.

16. The method of claim 15, wherein the model is stored at the electronic device.

17. The method of claim 15, wherein the model is stored at an auxiliary electronic device communicating with the electronic device, the method further comprising: The input, the one or more images, and the subset of the first image are sent to the auxiliary electronic device; as well as The output of the model is received from the auxiliary electronic device.

18. The method of claim 1, wherein the one or more criteria include criteria satisfied when processing the input, the one or more images, and the subset of the first image to provide a request for one or more objects corresponding to the subset of the first image.

19. An electronic device, the electronic device comprising: One or more processors; Memory; and One or more programs, the programs being stored in the memory and configured to be executed by the one or more processors, the programs including instructions for performing the method according to any one of claims 1 to 18.

20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions that, when executed by one or more processors of an electronic device, cause the electronic device to perform the method according to any one of claims 1 to 18.