Methods and devices for managing user-facing interactions with physical objects.

By adapting output modalities and mark parameters based on physical object interactions, the user interface systems optimize XR environment interactions, enhancing user experience and engagement.

JP2026076149APending Publication Date: 2026-05-11APPLE INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
APPLE INC
Filing Date
2025-12-18
Publication Date
2026-05-11

AI Technical Summary

Technical Problem

Existing user interface workflows fail to leverage input modalities effectively, missing opportunities to enhance user experience based on the type of input used.

Method used

Implement systems and methods that dynamically adjust output modalities and mark parameters based on physical object interactions, such as movement, pressure, and grip pose, within extended reality environments.

Benefits of technology

Enhances user experience by optimizing interactions with XR environments through adaptive output modalities and mark parameters, improving interaction efficiency and user engagement.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026076149000001_ABST
    Figure 2026076149000001_ABST
Patent Text Reader

Abstract

The present invention provides a system, method, and device for managing user-facing interactions with physical objects. [Solution] A first graphical element associated with a first set of output modalities is displayed in an XR environment. While the first graphical element is displayed, the movement of a physical object is detected. In response to the detection of this movement, if the movement causes the physical object to exceed a distance threshold for the first graphical element among the first graphical elements, the first output modality associated with the first graphical element is selected as the current output modality of the physical object. If the movement causes the physical object to exceed a distance threshold for the second graphical element among the first graphical elements, the second output modality associated with the second graphical element is selected as the current output modality of the physical object.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] This disclosure generally relates to systems, methods, and methods for interacting with and manipulating user interfaces, and more particularly for managing user interface-oriented interactions with physical objects. [Background technology]

[0002] Typically, users can interact with user interfaces through various input modalities such as touch input, voice input, and stylus / peripheral device input. However, the workflow for performing actions within the user interface may remain the same regardless of the input modality. This can lead to missed opportunities to enhance the user experience based on input modality and other factors.

[0003] This disclosure may have a more detailed description by reference to several exemplary implementations, some of which are shown in the accompanying drawings, as can be understood by those skilled in the art. [Brief explanation of the drawing]

[0004] [Figure 1] This is a block diagram of an exemplary operating architecture for several implementation configurations.

[0005] [Figure 2] This is a block diagram of an exemplary controller relating to several implementation configurations.

[0006] [Figure 3] This is a block diagram of an exemplary electronic device relating to several implementation configurations.

[0007] [Figure 4] This is a block diagram of an exemplary control device relating to several implementation configurations.

[0008] [Figure 5A] This is a block diagram of the first part of an exemplary content delivery architecture relating to several implementation forms.

[0009] [Figure 5B] The following are illustrative data structures related to several implementation configurations.

[0010] [Figure 5C] This is a block diagram of the second part of an exemplary content delivery architecture relating to several implementation forms.

[0011] [Figure 6A] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6B] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6C] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6D] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6E] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6F] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6G] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6H] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6I] This shows a sequence of instances for the first content delivery scenario, relating to several implementation forms. [Figure 6J]Shows the sequence of instances for a first content delivery scenario according to some implementation forms. [Figure 6K] Shows the sequence of instances for a first content delivery scenario according to some implementation forms. [Figure 6L] Shows the sequence of instances for a first content delivery scenario according to some implementation forms. [Figure 6M] Shows the sequence of instances for a first content delivery scenario according to some implementation forms. [Figure 6N] Shows the sequence of instances for a first content delivery scenario according to some implementation forms. [Figure 6O] Shows the sequence of instances for a first content delivery scenario according to some implementation forms. [Figure 6P] Shows the sequence of instances for a first content delivery scenario according to some implementation forms.

[0012] [Figure 7A] Shows the sequence of instances for a second content delivery scenario according to some implementation forms. [Figure 7B] Shows the sequence of instances for a second content delivery scenario according to some implementation forms. [Figure 7C] Shows the sequence of instances for a second content delivery scenario according to some implementation forms. [Figure 7D] Shows the sequence of instances for a second content delivery scenario according to some implementation forms. [Figure 7E] Shows the sequence of instances for a second content delivery scenario according to some implementation forms. [Figure 7F] Shows the sequence of instances for a second content delivery scenario according to some implementation forms. [Figure 7G]This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7H] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7I] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7J] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7K] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7L] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7M] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms. [Figure 7N] This shows a sequence of instances for a second content delivery scenario, relating to several implementation forms.

[0013] [Figure 8A] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8B] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8C] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8D] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8E] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8F]This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8G] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8H] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8I] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8J] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8K] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8L] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms. [Figure 8M] This shows a sequence of instances for a third content delivery scenario, relating to several implementation forms.

[0014] [Figure 9A] This document shows flowcharts illustrating how to select the output modality of a physical object when interacting with or manipulating an XR environment, across several implementation configurations. [Figure 9B] This document shows flowcharts illustrating how to select the output modality of a physical object when interacting with or manipulating an XR environment, across several implementation configurations. [Figure 9C] This document shows flowcharts illustrating how to select the output modality of a physical object when interacting with or manipulating an XR environment, across several implementation configurations.

[0015] [Figure 10A]This diagram shows a flowchart representing a method for changing the mark parameters based on a first input (pressure) value while directly marking on a physical surface, or based on a second input (pressure) value while indirectly marking, for several implementation configurations. [Figure 10B] This diagram shows a flowchart representing a method for changing the mark parameters based on a first input (pressure) value while directly marking on a physical surface, or based on a second input (pressure) value while indirectly marking, for several implementation configurations.

[0016] [Figure 11A] This shows flowcharts illustrating how to change the selected modality based on whether the user is currently holding a physical object, across several implementations. [Figure 11B] This shows flowcharts illustrating how to change the selected modality based on whether the user is currently holding a physical object, across several implementations. [Figure 11C] This shows flowcharts illustrating how to change the selected modality based on whether the user is currently holding a physical object, across several implementations. [Overview of the project]

[0017] By convention, various features shown in the drawings may not be depicted to scale. Therefore, the dimensions of various features may be arbitrarily enlarged or reduced for clarity. In addition, some drawings may not depict all components of a given system, method, or device. Finally, throughout this specification and the drawings, similar reference numerals may be used to indicate similar features.

[0018] The various implementations disclosed herein include devices, systems, and methods for selecting the output modality of a physical object when interacting with or manipulating an XR environment. According to some implementations, the method is performed in a computing system including non-temporary memory and one or more processors, the computing system being communicatively coupled to a display device and one or more input devices. The method includes: displaying a first set of graphical elements associated with a first set of output modalities in an augmented reality (XR) environment via a display device; detecting a first movement of a physical object while the first set of graphical elements are being displayed; selecting a first output modality associated with a first graphical element as the current output modality of the physical object, in response to the detection of the first movement of the physical object and determining that the first movement of the physical object has caused the physical object to exceed a distance threshold for a first graphical element among the first set of graphical elements; and selecting a second output modality associated with a second graphical element as the current output modality of the physical object, in determining that the first movement of the physical object has caused the physical object to exceed a distance threshold for a second graphical element among the first set of graphical elements.

[0019] Various implementations disclosed herein include devices, systems, and methods for modifying the parameters of a mark based on a first input (pressure) value while directly marking on a physical surface, or based on a second input (pressure) value while indirectly marking. According to some implementations, the method is performed in a computing system including non-temporary memory and one or more processors, the computing system being communicatively coupled to a display device and one or more input devices. The method includes displaying a user interface via the display device, detecting a marking input by a physical object while the user interface is being displayed, displaying a mark in the user interface via the display device based on the marking input, in accordance with the determination that the marking input is directed towards a physical surface, wherein the parameters of the mark displayed based on the marking input are determined based on how hard the physical object is pressed against the physical surface, and displaying a mark in the user interface via the display device based on the marking input, in accordance with the determination that the marking input is not directed towards a physical surface, wherein the parameters of the mark displayed based on the marking input are determined based on how hard the physical object is being held by the user.

[0020] Various implementations disclosed herein include devices, systems, and methods for changing the selection modality based on whether a user is currently holding a physical object. According to some implementations, the method is performed in a computing system including non-temporary memory and one or more processors, the computing system being communicatively coupled to a display device and one or more input devices. The method includes displaying content via the display device, detecting a selection input while the content is being displayed and while a physical object is being held by a user, and performing an action corresponding to the selection input in response to the detection of the selection input, the action including performing a selection action on a first portion of the content according to a determination that a grip pose associated with the manner in which the physical object is being held by the user corresponds to a first grip, wherein the first portion of the content is selected based on the direction in which a given portion of the physical object is facing, and performing a selection action on a second portion of the content different from the first portion of the content according to a determination that a grip pose associated with the manner in which the physical object is being held by the user does not correspond to a first grip, wherein the second portion of the content is selected based on the direction in which the user is looking.

[0021] In some implementations, the electronic device includes one or more displays, one or more processors, non-temporary memory, and one or more programs, the one or more programs being stored in the non-temporary memory and configured to be executed by one or more processors, and the one or more programs including instructions to perform or cause to perform any of the methods described herein. According to some implementations, the non-temporary computer-readable storage medium stores instructions internally, and when these instructions are executed by one or more processors of the device, the device performs or causes to perform any of the methods described herein based on these instructions. According to some implementations, the device includes one or more displays, one or more processors, non-temporary memory, and means to perform or cause to perform any of the methods described herein.

[0022] According to some implementations, a computing system includes one or more processors, non-temporary memory, an interface for communicating with a display device and one or more input devices, and one or more programs, the one or more programs being stored in the non-temporary memory and configured to be executed by one or more processors, and the one or more programs including instructions to perform or cause to perform any of the operations described herein. According to some implementations, a non-temporary computer-readable storage medium stores instructions internally, and when the instructions are executed by one or more processors of a computing system having an interface for communicating with a display device and one or more input devices, the computing system performs or causes to perform any of the operations described herein. According to some implementations, a computing system includes one or more processors, non-temporary memory, an interface for communicating with a display device and one or more input devices, and means to perform or cause to perform any of the operations described herein. [Modes for carrying out the invention]

[0023] Numerous details are provided to provide a full understanding of the exemplary implementations shown in the drawings. However, the drawings merely illustrate some exemplary embodiments of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective embodiments and / or variations do not include all of the specific details described herein. Furthermore, well-known systems, methods, components, devices and circuits are not described in exhaustive detail so as not to obscure more suitable embodiments of the exemplary implementations described herein.

[0024] The physical environment refers to the physical world that people can perceive and / or interact with without the aid of electronic devices. The physical environment may include physical features such as physical surfaces or physical objects. For example, the physical environment corresponds to a physical park that includes physical trees, physical buildings, and physical people. People can directly perceive and / or interact with the physical environment through their senses such as sight, touch, hearing, taste, and smell. In contrast, an extended reality (XR) environment refers to a fully or partially simulated environment that people perceive and / or interact with through electronic devices. For example, an XR environment may include augmented reality (AR) content, mixed reality (MR) content, and virtual reality (VR) content. In an XR system, a subset of a person's body movements or their representation is tracked, and in response, one or more properties of one or more virtual objects simulated within the XR environment are adjusted to behave according to at least one law of physics. As an example, an XR system can detect the rotation of a person's head and, in response, adjust the graphic content and sound field presented to that person in a manner similar to how such views and sounds would change in the physical environment. As another example, an XR system can detect the movement of an electronic device presenting an XR environment (e.g., a mobile phone, tablet, laptop) and, in response, adjust the graphic content and sound field presented to that person in a manner similar to how such views and sounds would change in the physical environment. In some situations (e.g., for accessibility reasons), an XR system can adjust the characteristics of the graphic content within the XR environment in response to a representation of bodily movement (e.g., a voice command).

[0025] The existence of a wide variety of electronic systems enables people to perceive and / or interact with various XR environments. Examples include head-mountable systems, projection-based systems, heads-up displays (HUDs), vehicle windshields with integrated display capabilities, windows with integrated display capabilities, displays formed as lenses designed to be positioned over a person's eyes (similar to contact lenses), headphones / earphones, speaker arrays, input systems (e.g., wearable or handheld controllers with or without haptic feedback), smartphones, tablets, and desktop / laptop computers. A head-mountable system may have one or more speakers and an integrated opaque display. Alternatively, a head-mountable system may be configured to accept an external opaque display (e.g., a smartphone). A head-mountable system may incorporate one or more imaging sensors for capturing images or videos of the physical environment and / or one or more microphones for capturing audio of the physical environment. A head-mountable system may have a transparent or translucent display instead of an opaque display. A transparent or translucent display may have a medium through which light representing an image is directed towards a person's eye. The display may utilize digital light projection, OLED, LED, μLED, liquid crystal on silicon, laser scanning light source, or any combination of these technologies. The medium may be an optical waveguide, a holographic medium, an optical coupler, an optical reflector, or any combination thereof. In some implementations, the transparent or translucent display may be configured to be selectively opaque. A projection-based system may employ retinal projection technology to project a graphical image onto a person's retina. The projection system may also be configured to project virtual objects into the physical environment, for example, as a hologram or onto a physical surface.

[0026] Figure 1 is a block diagram of an exemplary operating architecture 100 relating to several implementations. While relevant features are shown, those skilled in the art will understand from this disclosure that various other features are not shown in order to simplify the diagram so as not to obscure more appropriate embodiments of the exemplary implementations disclosed herein. For that purpose, as a non-limiting example, the operating architecture 100 includes an optional controller 110 and electronic devices 120 (e.g., a tablet, mobile phone, laptop, near-eye system, wearable computing device, etc.).

[0027] In some implementations, the controller 110 is configured to manage and coordinate the XR experience (which may be referred to herein as the “XR environment,” “virtual environment,” or “graphical environment”) for a user 149 having a left hand 150 and a right hand 152, as well as optionally for other users. In some implementations, the controller 110 includes a preferred combination of software, firmware, and / or hardware. The controller 110 is described in more detail below with reference to Figure 2. In some implementations, the controller 110 is a computing device that is local or remote to the physical environment 105. For example, the controller 110 is a local server located within the physical environment 105. In another embodiment, the controller 110 is a remote server located outside the physical environment 105 (e.g., a cloud server, a central server, etc.). In some implementations, the controller 110 is communicably coupled to the electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In some implementations, the functions of the controller 110 are provided by the electronic device 120. Therefore, in some implementations, the components of the controller 110 are integrated into the electronic device 120.

[0028] As shown in Figure 1, user 149 grasps the control device 130 with their right hand 152. As shown in Figure 1, the control device 130 includes a first end 176 and a second end 177. In various embodiments, the first end 176 corresponds to the tip of the control device 130 (e.g., the tip of a pencil), and the second end 177 corresponds to the opposite or lower end of the control device 130 (e.g., the eraser of a pencil). As shown in Figure 1, the control device 130 includes a touch-sensing surface 175 for receiving touch input from user 149. In some implementations, the control device 130 includes a preferred combination of software, firmware, and / or hardware. The control device 130 is described in more detail below with reference to Figure 4. In some implementations, the control device 130 corresponds to an electronic device having a wired or wireless communication channel to the controller 110. For example, the control device 130 corresponds to a stylus, a finger-worn device, a handheld device, etc. In some implementations, the controller 110 is coupled to the control device 130 in a communicative manner via one or more wired or wireless communication channels 146 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.).

[0029] In some implementations, the electronic device 120 is configured to present audio and / or video (A / V) content to the user 149. In some implementations, the electronic device 120 is configured to present a user interface (UI) and / or an XR environment 128 to the user 149. In some implementations, the electronic device 120 includes an appropriate combination of software, firmware, and / or hardware. The electronic device 120 is described in more detail below with reference to Figure 3.

[0030] In some implementations, the electronic device 120 presents the XR experience to the user 149 while the user 149 is physically present in a physical environment 105, including a table 107, within the field of view (FOV) 111 of the electronic device 120. Thus, in some implementations, the user 149 holds the electronic device 120 in their hand(s). In some implementations, while presenting the XR experience, the electronic device 120 presents XR content (sometimes referred to herein as “graphical content” or “virtual content”), including an XR cylinder 109, and is configured to allow video passthrough of the physical environment 105 (e.g., including a table 107 or its representation) on the display 122. For example, the XR environment 128, including the XR cylinder 109, is stereoscopic or three-dimensional (3D).

[0031] In one example, the XR cylinder 109 corresponds to display-locked content such that when the FOV 111 changes due to the translational and / or rotational movement of the electronic device 120, the XR cylinder 109 remains displayed at the same location on the display 122. In another example, the XR cylinder 109 corresponds to world-locked content such that even if the FOV 111 changes due to the translational and / or rotational movement of the electronic device 120, the XR cylinder 109 remains displayed at its original location. Therefore, in this example, if the FOV 111 does not include the origin location, the XR environment 128 does not include the XR cylinder 109. For example, the electronic device 120 corresponds to a near-eye system, mobile phone, tablet, laptop, wearable computing device, etc.

[0032] In some implementations, display 122 corresponds to an additional display that enables optical see-through of the physical environment 105, including the table 107. For example, display 122 corresponds to a transparent lens, and electronic device 120 corresponds to glasses worn by user 149. Thus, in some implementations, electronic device 120 presents a user interface by projecting XR content (e.g., XR cylinder 109) onto the additional display, and the user interface is then overlaid onto the physical environment 105 from the user 149's viewpoint.

[0033] In some implementations, the user 149 wears an electronic device 120, such as a near-eye system. Thus, the electronic device 120 includes one or more displays (e.g., a single display or one for each eye) provided for displaying XR content. For example, the electronic device 120 surrounds the user 149's field of view (FOV). In such implementations, the electronic device 120 presents the XR environment 128 by displaying data corresponding to the XR environment 128 on one or more displays, or by projecting data corresponding to the XR environment 128 onto the user 149's retina.

[0034] In some implementations, the electronic device 120 includes an integrated display (e.g., an internal display) that displays the XR environment 128. In some implementations, the electronic device 120 includes a head-mounted enclosure. In various implementations, the head-mounted enclosure includes a mounting area to which another device having a display can be mounted. For example, in some implementations, the electronic device 120 can be mounted on the head-mounted enclosure. In various implementations, the head-mounted enclosure is molded to form a receptacle for receiving another device (e.g., the electronic device 120) that includes a display. For example, in some implementations, the electronic device 120 is slid / snap-fitted into the head-mounted enclosure or mounted in another way. In some implementations, the display of the device mounted on the head-mounted enclosure presents (e.g., displays) the XR environment 128. In some implementations, the electronic device 120 is replaced by an XR chamber, enclosure, or room configured to present XR content to a user 149 without the user wearing or holding the electronic device 120.

[0035] In some implementations, the controller 110 and / or electronic device 120 move the user 149's XR representation within the XR environment 128 based on movement information (e.g., torso posture data, eye-tracking data, hand / limb / finger / limbs tracking data, etc.) from the electronic device 120 and / or an optional remote input device within the physical environment 105. In some implementations, the optional remote input device corresponds to a fixed or movable sensing device within the physical environment 105 (e.g., an image sensor, depth sensor, infrared (IR) sensor, event camera, microphone, etc.). In some implementations, each of the remote input devices is configured to collect / capture input data while the user 149 is physically present within the physical environment 105 and to provide the input data to the controller 110 and / or electronic device 120. In some implementations, the remote input device includes a microphone, and the input data includes audio data associated with the user 149 (e.g., speech samples). In some implementations, the remote input device includes an image sensor (e.g., a camera), and the input data includes an image of the user 149. In some implementations, the input data characterizes the torso posture of user 149 at different times. In some implementations, the input data characterizes the head posture of user 149 at different times. In some implementations, the input data characterizes hand tracking information associated with user 149's hands at different times. In some implementations, the input data characterizes the velocity and / or acceleration of parts of the user's torso, such as user 149's hands. In some implementations, the input data indicates the joint positions and / or joint orientations of user 149. In some implementations, the remote input device includes feedback devices such as speakers and lights.

[0036] Figure 2 is a block diagram of an example of the controller 110 relating to several implementation configurations. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the implementation configurations disclosed herein. Therefore, as a non-limiting example, in some implementations, the controller 110 includes one or more processing units 202 (e.g., a microprocessor, application-specific integrated circuit (ASIC), field-programmable gate array (FPGA), graphics processing unit (GPU), central processing unit (CPU), processing core, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), FireWire, Thunderbolt, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global Mobile Communication System (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZiGBEE, or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these and various other components.

[0037] In some implementations, one or more communication buses 204 include circuits that interconnect system components and control communication. In some implementations, one or more I / O devices 206 include at least one of the following: a keyboard, mouse, touchpad, touchscreen, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0038] Memory 220 includes high-speed random-access memory such as dynamic random-access memory (DRAM), static random-access memory (SRAM), double-data-rate random-access memory (DDRRAM), or other random-access solid-state memory devices. In some implementations, memory 220 includes non-volatile memory such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 220 optionally includes one or more storage devices located remotely from one or more processing units 202. Memory 220 includes a non-temporary computer-readable storage medium. In some implementations, memory 220 or the non-temporary computer-readable storage medium of memory 220 stores the following programs, modules, and data structures, or subsets thereof as described below with reference to Figure 2.

[0039] The operating system 230 includes procedures for handling various basic system services and procedures for performing hardware-dependent tasks.

[0040] In some implementations, the data acquisition unit 242 is configured to acquire data (e.g., captured image frames of the physical environment 105, presentation data, input data, user interaction data, camera pose tracking information, eye-tracking information, head / torso pose tracking information, hand / limb / finger / limb tracking information, sensor data, location data, etc.) from at least one of the controller 110's I / O devices 206, the electronic device 120's I / O devices and sensors 306, and an optional remote input device. For this purpose, in various implementations, the data acquisition unit 242 includes instructions and / or logic for this purpose, as well as heuristics and metadata for this purpose.

[0041] In some implementations, the mapper and locator engine 244 is configured to map the physical environment 105 and track the position / location of at least the electronic device 120 or the user 149 relative to the physical environment 105. For this purpose, in various implementations, the mapper and locator engine 244 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0042] In some implementations, the data transmission unit 246 is configured to transmit data (e.g., presentation data such as rendered image frames related to the XR environment, location data, etc.) to at least the electronic device 120 and optionally to one or more other devices. For this purpose, in various implementations, the data transmission unit 246 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0043] In some implementations, the privacy architecture 508 is configured to ingest data and filter user information and / or identifying information within the data based on one or more privacy filters. The privacy architecture 508 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the privacy architecture 508 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0044] In some implementations, the object tracking engine 510 is configured to determine / generate an object tracking vector 511 for tracking a physical object (e.g., a control device 130 or a proxy object) based on tracking data, and to update the object tracking vector 511 over time. For example, as shown in Figure 5B, the object tracking vector 511 includes translational values ​​572 of the physical object (e.g., associated with x, y, and z coordinates relative to the physical environment 105), rotational values ​​574 of the physical object (e.g., roll, pitch, and yaw), one or more pressure values ​​576 associated with the physical object, and optional touch input information 578 associated with the physical object. The object tracking engine 510 is described in more detail below with respect to Figure 5A. For this purpose, in various implementations, the object tracking engine 510 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0045] In some implementations, the eye-tracking engine 512 is configured to determine / generate an eye-tracking vector 513 (e.g., along with the gaze direction) based on input data, as shown in Figure 5B, and to update the eye-tracking vector 513 over time. For example, the gaze direction indicates a point, physical object, or region of interest (ROI) within the physical environment 105 that the user 149 is currently looking at (e.g., associated with x, y, and z coordinates relative to the physical environment 105 or the world as a whole). As another example, the gaze direction indicates a point, XR object, or region of interest (ROI) within the XR environment 128 that the user 149 is currently looking at (e.g., associated with x, y, and z coordinates relative to the XR environment 128). The eye-tracking engine 512 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the eye-tracking engine 512 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0046] In some implementations, the torso / head attitude tracking engine 514 is configured to determine / generate an attitude feature vector 515 based on input data and to update the attitude feature vector 515 over time. For example, as shown in Figure 5B, the attitude feature vector 515 includes a head attitude descriptor 592A (e.g., up, down, neutral), translation values ​​592B for the head attitude, rotation values ​​592C for the head attitude, a torso attitude descriptor 594A (e.g., upright, seated, prone), translation values ​​594B for the torso / limbs / arms / joints, rotation values ​​594C for the torso / limbs / arms / joints, etc. The torso / head attitude tracking engine 514 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the torso / head attitude tracking engine 514 includes instructions and / or logic for it, as well as heuristics and metadata for it. In some implementations, the object tracking engine 510, the gaze tracking engine 512, and the torso / head attitude tracking engine 514 may be located on an electronic device 120 in addition to, or instead of, the controller 110.

[0047] In some implementations, the content selection unit 542 is configured to select XR content (sometimes referred to herein as "graphical content" or "virtual content") from the content library 545 based on one or more user requests and / or inputs (e.g., voice commands, selections from the user interface (UI) menu of an XR content item). The content selection unit 542 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the content selection unit 542 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0048] In some implementations, the content library 545 contains multiple content items such as audio / visual (A / V) content, virtual agents (VA), and / or XR content, objects, items, and landscapes. For example, XR content may include 3D reconstructions of user-captured videos, movies, TV episodes, and / or other XR content. In some implementations, the content library 545 is pre-entered by the user 149 or created manually. In some implementations, the content library 545 is located locally with respect to the controller 110. In some implementations, the content library 545 is located remotely from the controller 110 (e.g., on a remote server, cloud server, etc.).

[0049] In some implementations, the input management unit 520 is configured to acquire and analyze input data from various input sensors. The input management unit 520 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the input management unit 520 includes instructions and / or logic, as well as heuristics and metadata. In some implementations, the input management unit 520 includes a data aggregation unit 521, a content selection engine 522, a grip attitude evaluation unit 524, an output modality selection unit 526, and a parameter adjustment unit 528.

[0050] In some implementations, the data aggregation unit 521 is configured to aggregate the object tracking vector 511, the gaze tracking vector 513, and the pose characterization vector 515, and to determine / generate a characterization vector 531 (shown in Figure 5A) based on them for subsequent downstream use. The data aggregation unit 521 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the data aggregation unit 521 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0051] In some implementations, the content selection engine 522 is configured to determine a selected content portion 523 (shown in Figure 5A) within the XR environment 128 based on a characterization vector 531 (or a portion thereof). The content selection engine 522 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the content selection engine 522 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0052] In some implementations, the grip posture evaluation unit 524 is configured to determine a grip posture 525 (as shown in Figure 5A) associated with the current state in which the physical object is held by the user 149, based on a characterization vector 531 (or a part thereof). For example, the grip posture 525 indicates the manner in which the user 149 grasps the physical object (e.g., a surrogate object, control device 130, etc.). For example, the grip posture 525 corresponds to one of the following: remote control grip, pointing / wand grip, writing grip, reverse writing grip, handle grip, thumb-top grip, level grip, gamepad grip, flute grip, or pyrotechnic grip. The grip posture evaluation unit 524 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the grip posture evaluation unit 524 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0053] In some implementations, the output modality selection unit 526 is configured to select the current output modality 527 (shown in Figure 5A) associated with how physical objects interact with or manipulate the XR environment 128. For example, a first output modality corresponds to selecting / manipulating objects / content within the XR environment 128, and a second output modality corresponds to sketching, drawing, writing, etc., within the XR environment 128. The output modality selection unit 526 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the output modality selection unit 526 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0054] In some implementations, the parameter adjustment unit 528 is configured to adjust parameter values ​​(e.g., thickness, brightness, color, texture, etc.) associated with a marking input directed to the XR environment 128 (as shown in Figure 5A) based on either a first input (pressure) value or a second input (pressure) value associated with a physical object. The parameter adjustment unit 528 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the parameter adjustment unit 528 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0055] In some implementations, the content management unit 530 is configured to manage and update the layout, settings, structure, etc., of the XR environment 128, which includes one or more of the following: VA, XR content, and one or more user interface (UI) elements associated with the XR content. The content management unit 530 is described in more detail below with reference to Figure 5C. For this purpose, in various implementations, the content management unit 530 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose. In some implementations, the content management unit 530 includes a buffer 534, a content update unit 536, and a feedback engine 538. In some implementations, the buffer 534 includes XR content, rendered image frames, etc., for one or more past instances and / or frames.

[0056] In some implementations, the content update unit 536 is configured to modify the XR environment 128 over time based on translational or rotational movements of electronic devices 120 or physical objects within the physical environment 105, user input (e.g., hand / limb tracking input, eye-tracking input, touch input, voice commands, operation input by physical objects, etc.). To this end, in various implementations, the content update unit 536 includes instructions and / or logic for this purpose, as well as heuristics and metadata for this purpose.

[0057] In some implementations, the feedback engine 538 is configured to generate sensory feedback associated with the XR environment 128 (e.g., visual feedback such as text or changes in lighting, audio feedback, haptic feedback, etc.). For this purpose, in various implementations, the feedback engine 538 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0058] In some implementations, the rendering engine 550 is configured to render the XR environment 128 (which may also be referred to herein as the “graphical environment” or “virtual environment”) or associated image frames, as well as VA, XR content, one or more UI elements associated with the XR content, etc. For this purpose, in various implementations, the rendering engine 550 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose. In some implementations, the rendering engine 550 includes a pose determination unit 552, a rendering unit 554, an optional image processing architecture 562, and an optional compositing unit 564. Those skilled in the art will understand that the optional image processing architecture 562 and the optional compositing unit 564 may be present for a video pass-through configuration but may be removed for a full VR or optical see-through configuration.

[0059] In some implementations, the attitude determination unit 552 is configured to determine the current camera attitude of the electronic device 120 and / or user 149 with respect to A / V content and / or XR content. The attitude determination unit 552 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the attitude determination unit 552 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0060] In some implementations, the rendering unit 554 is configured to render A / V content and / or XR content according to the current camera orientation relative to it. The rendering unit 554 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the rendering unit 554 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0061] In some implementations, the image processing architecture 562 is configured to acquire (e.g., receive, retrieve, or capture) an image stream containing one or more images of the physical environment 105 from the current camera orientation of the electronic device 120 and / or user 149. In some implementations, the image processing architecture 562 is also configured to perform one or more image processing operations on the image stream, such as warping, color correction, gamma correction, sharpening, noise reduction, and white balance. The image processing architecture 562 is described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the image processing architecture 562 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0062] In some implementations, the compositing unit 564 is configured to combine rendered A / V content and / or XR content with processed image streams from the physical environment 105 of the image processing architecture 562 to generate rendered image frames of the XR environment 128 for display. The compositing unit 564 will be described in more detail below with reference to Figure 5A. For this purpose, in various implementations, the compositing unit 564 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0063] Although the data acquisition unit 242, mapper and locator engine 244, data transmission unit 246, privacy architecture 508, object tracking engine 510, eye-tracking engine 512, torso / head pose tracking engine 514, content selection unit 542, content management unit 530, motion modality management unit 540, and rendering engine 550 are shown as existing on a single device (e.g., controller 110), it should be understood that in other implementations, any combination of the data acquisition unit 242, mapper and locator engine 244, data transmission unit 246, privacy architecture 508, object tracking engine 510, eye-tracking engine 512, torso / head pose tracking engine 514, content selection unit 542, content management unit 530, motion modality management unit 540, and rendering engine 550 may be located in separate computing devices.

[0064] In some implementations, the functions and / or components of the controller 110 are combined with, or provided by, the electronic device 120 shown below in Figure 3. Furthermore, Figure 2 is intended to illustrate the functions of various features present in a particular implementation, rather than a structural schematic of the implementations described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules shown separately in Figure 2 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various implementations. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation, and in some implementations, will partially depend on a specific combination of hardware, software, and / or firmware selected for that particular implementation.

[0065] Figure 3 is a block diagram of an example of an electronic device 120 (e.g., a mobile phone, tablet, laptop, near-eye system, wearable computing device, etc.) relating to several implementation forms. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the implementation forms disclosed herein. For that purpose, as an unrestricted example, in some implementations, the electronic device 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, infrared, Bluetooth, ZIGBEE, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more displays 312, image capture devices 370 (e.g., one or more optional in-facing and / or out-facing image sensors), memory 320, and one or more communication buses 304 for interconnecting these and various other components.

[0066] In some implementations, one or more communication buses 304 include circuits that interconnect system components and control communication. In some implementations, one or more I / O devices and sensors 306 include at least one of the following: an inertial measuring unit (IMU), an accelerometer, a gyroscope, a magnetometer, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen monitor, a blood glucose monitor, etc.), one or more microphones, one or more speakers, a tactile engine, a heating and / or cooling unit, a skin shear engine, one or more depth sensors (e.g., structured light, time of flight, LiDAR, etc.), a localization and mapping engine, an eye-tracking engine, a torso / head attitude tracking engine, a hand / limb / finger / limbs tracking engine, a camera attitude tracking engine, etc.

[0067] In some implementations, one or more displays 312 are configured to present the XR environment to the user. In some implementations, one or more displays 312 are also configured to present flat video content to the user (e.g., two-dimensional or "flat" files such as AVI, FLV, WMV, MOV, MP4 associated with a TV episode or movie, or live video passthrough of the physical environment 105). In some implementations, one or more displays 312 correspond to touchscreen displays. In some implementations, one or more displays 312 correspond to holographic, digital light processing (DLP), liquid crystal displays (LCD), reflective liquid crystals (LCoS), organic light-emitting field-effect transistors (OLET), organic light-emitting diodes (OLED), surface conduction electron emission displays (SED), field emission displays (FED), quantum dot light-emitting diodes (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some implementations, one or more displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holographic displays. For example, the electronic device 120 includes a single display. In another example, the electronic device 120 includes a display for each of the user's eyes. In some implementations, one or more displays 312 can present AR and VR content. In some implementations, one or more displays 312 can present AR or VR content.

[0068] In some implementations, the image capture device 370 includes one or more RGB cameras, IR image sensors, event-based cameras, etc. (e.g., with complementary metal-oxide-semiconductor (CMOS) image sensors or charge-coupled device (CCD) image sensors). In some implementations, the image capture device 370 includes a lens assembly, a photodiode, and a front-end architecture. In some implementations, the image capture device 370 includes outward-facing and / or inward-facing image sensors.

[0069] Memory 320 includes high-speed random-access memory, such as DRAM, SRAM, DDR RAM, or other random-access solid-state memory devices. In some implementations, memory 320 includes non-volatile memory, such as one or more magnetic disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-temporary computer-readable storage medium. In some implementations, memory 320, or the non-temporary computer-readable storage medium of memory 320, stores the following programs, modules, and data structures, or subsets thereof, including an optional operating system 330 and presentation engine 340.

[0070] The operating system 330 includes procedures for handling various basic system services and procedures for performing hardware-dependent tasks. In some implementations, the presentation engine 340 is configured to present media items and / or XR content to the user via one or more displays 312. For this purpose, in various implementations, the presentation engine 340 includes a data acquisition unit 342, a presentation unit 570, an interaction processing unit 540, and a data transmission unit 350.

[0071] In some implementations, the data acquisition unit 342 is configured to acquire data (e.g., presentation data such as rendered image frames related to a user interface or XR environment, input data, user interaction data, head tracking information, camera pose tracking information, gaze tracking information, hand / limb / finger / limbs tracking information, sensor data, location data, etc.) from at least one of the I / O devices and sensors 306 of the electronic device 120, the controller 110, and remote input devices. For this purpose, in various implementations, the data acquisition unit 342 includes instructions and / or logic for this purpose, as well as heuristics and metadata for this purpose.

[0072] In some implementations, the dialogue processing unit 540 is configured to detect user interactions with presented A / V content and / or XR content (e.g., gesture input detected via hand / limb tracking, eye-tracking input detected via eye-tracking, voice commands, etc.). To this end, in various implementations, the dialogue processing unit 540 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0073] In some implementations, the presentation unit 570 is configured to present and update A / V content and / or XR content (e.g., a user interface or rendered image frame associated with an XR environment 128, including VA, XR content, and one or more UI elements associated with the XR content) via one or more displays 312. For this purpose, in various implementations, the presentation unit 570 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0074] In some implementations, the data transmission unit 350 is configured to transmit data (e.g., presentation data, location data, user interaction data, head tracking information, camera posture tracking information, eye-tracking information, hand / limb / finger / limbs tracking information, etc.) to at least the controller 110. For this purpose, in various implementations, the data transmission unit 350 includes instructions and / or logic for that purpose, as well as heuristics and metadata for that purpose.

[0075] Although the data acquisition unit 342, the dialogue processing unit 540, the presentation unit 570, and the data transmission unit 350 are shown as existing on a single device (e.g., electronic device 120), it should be understood that in other implementations, any combination of the data acquisition unit 342, the dialogue processing unit 540, the presentation unit 570, and the data transmission unit 350 may be located in separate computing devices.

[0076] Furthermore, Figure 3 is intended to illustrate the functionality of various features that may be present in a particular implementation, rather than to show a structural schematic of the implementations described herein. As will be recognized by those skilled in the art, the separately shown items can be combined, and some items can be separated. For example, several functional modules separately shown in Figure 3 can be implemented within a single module, and the various functions of a single functional block can be implemented by one or more functional blocks in various implementations. The actual number of modules, as well as the division of certain functions and how functions are assigned between them, will vary depending on the implementation, and in some implementations, it will partially depend on a specific combination of hardware, software, and / or firmware selected for that particular implementation.

[0077] Figure 4 is a block diagram of an exemplary control device 130 relating to several implementation configurations. The control device 130 may also be simply called a stylus. The control device 130 includes non-temporary memory 402 (optionally including one or more computer-readable storage media), a memory controller 422, one or more processing units (CPUs) 420, a peripheral interface 418, an RF circuit 408, an input / output (I / O) subsystem 406, and other input or control devices 416. The control device 130 optionally includes an external port 424 and one or more optical sensors 464. The control device 130 optionally includes one or more contact intensity sensors 465 for detecting the intensity of contact of the control device 130 on an electronic device 100 (for example, when the control device 130 is used with a touch-sensing surface such as the display system 122 of the electronic device 120) or on another surface (for example, the surface of a desk). The control device 130 optionally includes one or more tactile output generators 463 for generating tactile outputs on the control device 130. These components optionally communicate via one or more communication buses or signal lines 403.

[0078] It should be understood that the control device 130 is merely an example of an electronic stylus, and that the control device 130 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of those components. The various components shown in Figure 4 are implemented in hardware, software, firmware, or a combination thereof, including one or more signal processing circuits and / or application-specific integrated circuits. In some implementations, some functions and / or operations of the control device 130 (e.g., the touch interpretation module 477) are provided by the controller 110 and / or electronic device 120. Therefore, in some implementations, some components of the control device 130 are integrated into the controller 110 and / or electronic device 120.

[0079] As shown in Figure 1, the control device 130 includes a first end 176 and a second end 177. In various embodiments, the first end 176 corresponds to the tip of the control device 130 (e.g., the tip of a pencil), and the second end 177 corresponds to the opposite or lower end of the control device 130 (e.g., the eraser of a pencil).

[0080] As shown in Figure 1, the control device 130 includes a touch-sensing surface 175 for receiving touch input from a user 149. In some implementations, the touch-sensing surface 175 corresponds to a capacitive touch element. The control device 130 includes a sensor or set of sensors that detect input from the user based on tactile and / or tactile contact with the touch-sensing surface 175. In some implementations, the control device 130 detects contact and its movement or disconnection using any of several currently known or future-developed touch-sensing technologies, including but not limited to capacitive, resistive, infrared and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements that determine one or more contact points with the touch-sensing surface 175. Because the control device 130 includes various sensors and various types of sensors, the control device 130 can detect a variety of different inputs from the user 149. In some implementations, one or more sensors can detect a single touch input or a series of touch inputs in response to the user tapping the touch-sensing surface 175 once or multiple times. In some implementations, one or more sensors can detect a swipe input on the control device 130 in response to a user stroking along the touch-sensing surface 175 with one or more fingers. In some implementations, if the speed at which the user strokes along the touch-sensing surface 175 exceeds a threshold, one or more sensors detect a flick input instead of a swipe input.

[0081] The control device 130 also includes one or more sensors that detect the orientation (e.g., angular position) and / or movement of the control device 130, such as one or more accelerometers 467, one or more gyroscopes 468, and one or more magnetometers 469. One or more sensors can detect various rotational movements of the control device 130 by the user, including the type and direction of rotation. For example, one or more sensors can detect that the user rolls and / or pivots the control device 130 and can detect the direction of the roll / pivot (e.g., clockwise or counterclockwise). In some implementations, the input to be detected depends on the angular positions of the first end 176 and the second end 177 of the control device 130 relative to the electronic device. For example, in some implementations, if the control device 130 is substantially perpendicular to the electronic device and the second end 177 (e.g., an eraser) is closer to the electronic device, an erasing operation is performed by bringing the surface of the electronic device into contact with the second end 177. On the other hand, if the control device 130 is substantially perpendicular to the electronic device and the first end 176 (e.g., tip) is closer to the electronic device, the marking operation is performed by bringing the surface of the electronic device into contact with the first end 176.

[0082] Memory 402 optionally includes high-speed random-access memory and optionally also includes non-volatile memory such as one or more flash memory devices or other non-volatile solid-state memory devices. Access to memory 402 by CPU(s) 420 and other components of control device 130, such as peripheral interface 418, is optionally controlled by memory controller 422.

[0083] The input and output peripherals of this stylus can be coupled to the CPU(s)420 and memory402 using the peripheral interface 418. One or more processors 420 operate or execute various software programs and / or instruction sets stored in memory402 to perform various functions for the control device 130 and to process data. In some implementations, the peripheral interface 418, CPU(s)420, and memory controller 422 are optionally implemented on a single chip, such as chip 404. In some other embodiments, they are optionally implemented on separate chips.

[0084] The RF (radio frequency) circuit 408 transmits and receives RF signals, also known as electromagnetic signals. The RF circuit 408 converts electrical signals to electromagnetic signals, or electromagnetic signals to electrical signals, and communicates with communication networks and / or other communication devices, such as the controller 110 and the electronic device 120, via electromagnetic signals. The RF circuit 408 optionally includes well-known circuits for performing these functions, which include, but are not limited to, antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, CODEC chipsets, subscriber identification module (SIM) cards, and memory. The RF circuit 408 optionally communicates wirelessly with networks such as the Internet, also known as the World Wide Web (WWW), intranets, and / or wireless networks such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs), and with other devices. Wireless communication may optionally include Global System for Mobile Communications (GSM) (registered trademark), Enhanced Data GSM Environment (EDGE) (registered trademark), High-Speed ​​Downlink Packet Access (HSDPA), High-Speed ​​Uplink Packet Access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Cell HSPA (DC-HSPA), Long-Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wireless Fidelity (Wi-Fi) (registered trademark) (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b Using any of several communication standards, protocols, and technologies, including, but not limited to, 802.11g and / or IEEE 802.11n, or any other suitable communication protocol, including a communication protocol not yet developed as of the filing date of this document.

[0085] The I / O subsystem 406 connects input / output peripherals on the control device 130, such as other input or control devices 416, to the peripheral device interface 418. The I / O subsystem 406 optionally includes one or more input controllers 460 for an optical sensor controller 458, an intensity sensor controller 459, a tactile feedback controller 461, and other input or control devices. One or more input controllers 460 receive electrical signals from and transmit electrical signals to the other input or control devices 416. The other input or control devices 416 optionally include physical buttons (e.g., push buttons, rocker buttons), dials, slider switches, joysticks, click wheels, etc. In some alternative embodiments, the input controller(s) 460 are optionally coupled to (or not coupled to) either an infrared port and / or a USB port.

[0086] The control device 130 also includes a power system 462 that supplies power to various components. The power system 462 optionally includes a power management system, one or more power sources (e.g., battery, alternating current (AC)), a recharge system, a power failure detection circuit, a power converter or inverter, a power status indicator (e.g., a light-emitting diode (LED)), and any other components associated with the generation, management, and distribution of power in portable devices and / or portable accessories.

[0087] The control device 130 also optionally includes one or more optical sensors 464. Figure 4 shows the optical sensors coupled to the optical sensor controller 458 in the I / O subsystem 406. The one or more optical sensors 464 optionally include charge-coupled devices (CCDs) or complementary metal-oxide-semiconductor (CMOS) phototransistors. The one or more optical sensors 464 receive light from the environment projected through one or more lenses and convert that light into data representing an image.

[0088] The control device 130 also optionally includes one or more contact strength sensors 465. Figure 4 shows a contact strength sensor coupled to a strength sensor controller 459 in the I / O subsystem 406. The contact strength sensor 465 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electric force sensors, pressure-power sensors, optical force sensors, capacitive touch-sensing surfaces, or other strength sensors (e.g., sensors used to measure the force (or pressure) of contact with a surface or with a user 149's grip). The contact strength sensor 465 receives contact strength information (e.g., pressure information, or a proxy for pressure information) from the environment. In some implementations, at least one contact strength sensor is positioned juxtaposed with or adjacent to the tip of the control device 130. In some implementations, at least one contact strength sensor is positioned juxtaposed with or adjacent to the body of the control device 130.

[0089] The control device 130 also optionally includes one or more proximity sensors 466. Figure 4 shows one or more proximity sensors 466 coupled to a peripheral interface 418. Alternatively, one or more proximity sensors 466 are optionally coupled to an input controller 460 in an I / O subsystem 406. In some implementations, one or more proximity sensors 466 determine the proximity of the control device 130 to an electronic device (e.g., electronic device 120).

[0090] The control device 130 also optionally includes one or more tactile output generators 463. Figure 4 shows a tactile output generator coupled to a tactile feedback controller 461 in the I / O subsystem 406. The one or more tactile output generators 463 optionally include one or more electroacoustic devices such as speakers or other audio components, and / or electromechanical devices that convert energy into linear motion, such as motors, solenoids, electroactive polymers, piezoelectric actuators, electrostatic actuators, or other tactile output generating components (e.g., components that convert electrical signals into tactile outputs on electronic devices). The one or more tactile output generators 463 receive a tactile feedback generation command from the tactile feedback module 433 and generate a tactile output on the control device 130 that can be sensed by the user of the control device 130. In some implementations, at least one tactile output generator is positioned alongside or adjacent to the length of the control device 130 (e.g., the main body or housing), and optionally generates a tactile output by moving the control device 130 vertically (e.g., parallel to the length of the control device 130) or laterally (e.g., normal to the length of the control device 130).

[0091] The control device 130 also optionally includes one or more accelerometers 467, one or more gyroscopes 468, and / or one or more magnetometers 469 (e.g., as part of an inertial measurement unit (IMU)) for obtaining information about the location and positional state of the control device 130. Figure 4 shows sensors 467, 468, and 469 coupled to the peripheral interface 418. Alternatively, sensors 467, 468, and 469 are optionally coupled to an input controller 460 in the I / O subsystem 406. The control device 130 optionally includes a GPS (or GLONASS or other global navigation system) receiver (not shown) for obtaining information about the location of the control device 130.

[0092] The control device 130 includes a touch sensing system 432. The touch sensing system 432 detects inputs received on the touch sensing surface 175. These inputs include those described herein with respect to the touch sensing surface 175 of the control device 130. For example, the touch sensing system 432 can detect tap inputs, swipe inputs, roll inputs, flick inputs, swipe inputs, and the like. The touch sensing system 432 works in conjunction with a touch interpretation module 477 to decode specific types of touch inputs (e.g., swipe / roll / flick / swipe) received on the touch sensing surface 175.

[0093] In some implementations, the software components stored in memory 402 include an operating system 426, a communication module (or instruction set) 428, a touch / motion module (or instruction set) 430, a position module (or instruction set) 431, and a Global Positioning System (GPS) module (or instruction set) 435. Furthermore, in some implementations, memory 402 stores device / global internal state 457, as shown in Figure 4. Additionally, memory 402 includes a touch interpretation module 477. The device / global internal state 457 includes one or more sensor states, including information obtained from various sensors of the stylus and other input or control devices 416, information regarding the position and / or orientation of the control device 130 (e.g., translation and / or rotation values), and location information regarding the location of the control device 130 (e.g., determined by the GPS module 435).

[0094] The operating system 426 (e.g., embedded operating systems such as iOS, Darwin®, RTXC®, LINUX®, UNIX®, OS X®, WINDOWS®, or VxWorks®) includes various software components and / or drivers for controlling and managing common system tasks (e.g., memory management, power management, etc.) and facilitating communication between various hardware and software components. The communication module 428 optionally facilitates communication with other devices via one or more external ports 424 and also includes various software components for processing data received by the RF circuit 408 and / or external ports 424. The external ports 424 (e.g., Universal Serial Bus (USB), FireWire, etc.) are adapted to connect to other devices directly or indirectly via a network (e.g., the Internet, Wi-Fi, etc.).

[0095] The contact / motion module 430 optionally detects contact with the control device 130 and other touch-sensitive devices of the control device 130 (e.g., buttons or other touch-sensitive components of the control device 130). The contact / motion module 430 includes software components for performing various operations related to contact detection (e.g., detection of the tip of a stylus on a touch-sensitive display such as the display 122 of the electronic device 120, or on another surface such as the surface of a desk), such as determining whether contact has occurred (e.g., detection of a touchdown event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a substitute for the force or pressure of the contact), determining whether there is movement of the contact (e.g., across the display 122 of the electronic device 120) and tracking of that movement, and determining whether the contact has been terminated (e.g., detection of a lift-off event or interruption of contact). In some implementations, the contact / motion module 430 receives contact data from the I / O subsystem 406. Determining the movement of a contact point, represented by a series of contact data, optionally includes determining the speed (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point. As described above, in some implementations, one or more of these operations related to contact detection are performed by an electronic device 120 or controller 110 (in addition to, or instead of, a stylus using the contact / motion module 430).

[0096] The contact / motion module 430 optionally detects gesture input from the control device 130. Different gestures from the control device 130 have different contact patterns (e.g., different motion, timing, and / or intensity of the detected contact). Thus, gestures are optionally detected by detecting a specific contact pattern. For example, detecting a single tap gesture involves detecting a touchdown event, followed by a lift-off event at the same position (or substantially the same position) as the touchdown event (e.g., at the icon's position). As another example, detecting a swipe gesture involves detecting a touchdown event, followed by one or more stylus drag events, and then a lift-off event. As described above, in some implementations, gesture detection is performed by an electronic device using the contact / motion module 430 (in addition to, or instead of, a stylus using the contact / motion module 430).

[0097] The position module 431, together with one or more accelerometers 467, one or more gyroscopes 468, and / or one or more magnetometers 469, optionally detects position information about the stylus, such as the attitude of the control device 130 in a specific reference frame (e.g., roll, pitch, and / or yaw). The position module 431, in conjunction with one or more accelerometers 467, one or more gyroscopes 468, and / or one or more magnetometers 469, optionally detects movement gestures of the control device 130, such as flicking, tapping, and rolling. The position module 431 includes software components for performing various operations related to detecting the position of the stylus in a specific coordinate system and detecting changes in the position of the stylus. In some implementations, the position module 431 detects the position state of the control device 130 relative to the physical environment 105 or the world as a whole, and detects changes in the position state of the control device 130.

[0098] The haptic feedback module 433 includes various software components that generate commands used by one or more haptic output generators 463 to generate haptic outputs at one or more locations on the control device 130 in response to user interaction with the control device 130. The GPS module 435 determines the location of the control device 130 and provides this information for use in various applications, such as applications that provide location-based services, such as applications for finding lost devices and / or accessories.

[0099] The touch interpretation module 477 works in conjunction with the touch sensing system 432 to determine (e.g., decode or identify) the type of touch input received on the touch sensing surface 175 of the control device 130. For example, the touch interpretation module 477 determines that a touch input corresponds to a swipe input (as opposed to a tap input) if the user strokes the touch sensing surface 175 of the control device 130 over a sufficient distance in a sufficiently short amount of time. As another example, the touch interpretation module 477 determines that a touch input corresponds to a flick input (as opposed to a swipe input) if the speed at which the user strokes the touch sensing surface 175 of the control device 130 is sufficiently faster than the speed at which a swipe input would be corresponding. The stroke threshold speed can be preset and modified. In various embodiments, the pressure and / or force received by the touch on the touch sensing surface determines the type of input. For example, a light touch may correspond to a first type of input, and a strong touch may correspond to a second type of input.

[0100] Each of the modules and applications identified above corresponds to one or more of the functions described above, as well as an executable instruction set for performing the methods described in this application (e.g., methods performed by a computer and other information processing methods described herein). These modules (i.e., instruction sets) do not need to be implemented as separate software programs, procedures, or modules; therefore, various subsets of these modules can be combined or otherwise rearranged in various embodiments, at their discretion. In some implementations, memory 402 optionally stores a subset of the modules and data structures identified above. Furthermore, memory 402 optionally stores additional modules and data structures not described above.

[0101] Figure 5A is a block diagram of the first part 500A of an exemplary content delivery architecture relating to several implementations. While relevant features are shown, those skilled in the art will understand from this disclosure that various other features are not shown in order to simplify so as not to obscure more appropriate embodiments of the exemplary implementations disclosed herein. For that purpose, as a non-limiting example, the content delivery architecture is rendered and presented by a computing system such as the controller 110 shown in Figures 1 and 2, the electronic devices 120 shown in Figures 1 and 3, and / or appropriate combinations thereof.

[0102] As shown in Figure 5A, one or more local sensors 502 of the controller 110, electronic device 120, and / or combinations thereof acquire local sensor data 503 associated with the physical environment 105. For example, the local sensor data 503 includes an image or stream thereof of the physical environment 105, simultaneous location and mapping (SLAM) information about the physical environment 105, the location of the electronic device 120 or user 149 relative to the physical environment 105, ambient lighting information about the physical environment 105, ambient sound information about the physical environment 105, acoustic information about the physical environment 105, dimensional information about the physical environment 105, and semantic labels for objects within the physical environment 105. In some implementations, the local sensor data 503 includes raw or post-processed information.

[0103] Similarly, as shown in Figure 5A, one or more remote sensors 504 associated with the physical environment 105, the control device 130, and / or any selected remote input device within the physical environment 105 acquire remote sensor data 505 associated with the physical environment 105. For example, the remote sensor data 505 includes an image or stream thereof of the physical environment 105, SLAM information of the physical environment 105 and the location of the electronic device 120 or user 149 relative to the physical environment 105, ambient lighting information of the physical environment 105, ambient sound information of the physical environment 105, acoustic information of the physical environment 105, dimensional information of the physical environment 105, semantic labels of objects within the physical environment 105, and / or similar. In some implementations, the remote sensor data 505 includes raw or post-processed information.

[0104] As shown in Figure 5A, tracking data 506 is acquired by at least one of the controller 110, the electronic device 120, or the control device 130 in order to locate and track the control device 130. For example, tracking data 506 includes an image or stream of the physical environment 105 captured by the outward-facing image sensor of the electronic device 120, which includes the control device 130. For another example, tracking data 506 corresponds to IMU information, accelerometer information, gyroscope information, magnetometer information, and / or similar from the integrated sensors of the control device 130.

[0105] In some implementations, the privacy architecture 508 captures local sensor data 503, remote sensor data 505, and tracking data 506. In some implementations, the privacy architecture 508 includes one or more privacy filters associated with user information and / or identification information. In some implementations, the privacy architecture 508 includes an opt-in function in which the electronic device 120 notifies the user 149 about which user information and / or identification information is being monitored and how the user information and / or identification information is being used. In some implementations, the privacy architecture 508 selectively prevents and / or restricts the content delivery architecture 500A / 500B or any part thereof from acquiring and / or transmitting user information. To this end, the privacy architecture 508 receives user preferences and / or selections from the user 149 in response to prompting the user 149 for user preferences and / or selections. In some implementations, the privacy architecture 508 prevents the content delivery architecture 500A / 500B from acquiring and / or transmitting user information unless and until informed consent is obtained from the user 149. In some implementations, the privacy architecture 508 anonymizes certain types of user information (e.g., scrambling, obfuscating, encrypting, etc.). For example, the privacy architecture 508 receives user input specifying which types of user information the privacy architecture 508 anonymizes. As another example, the privacy architecture 508 anonymizes certain types of user information that are likely to contain sensitive and / or identifying information, independently of user specification (e.g., automatically).

[0106] In some implementations, the object tracking engine 510 obtains tracking data 506 after being subjected to the privacy architecture 508. In some implementations, the object tracking engine 510 determines / generates an object tracking vector 511 for a physical object based on the tracking data 506 and updates the object tracking vector 511 over time. As an example, the physical object corresponds to a proxy object detected in a physical environment 105 that does not have a communication channel to a computing system (e.g., controller 110, electronic device 120, etc.), such as a pencil or pen. As another example, the physical object corresponds to an electronic device (e.g., control device 130) that has a wired or wireless communication channel to a computing system (e.g., controller 110, electronic device 120, etc.), such as a stylus, finger-worn device, or handheld device.

[0107] Figure 5B shows an exemplary data structure for an object tracking vector 511 relating to several implementations. As shown in Figure 5B, the object tracking vector 511 may correspond to an N-tuple characterization vector or characterization tensor containing a timestamp 571 (e.g., the most recent time the object tracking vector 511 was updated), one or more translational values ​​572 of the physical object (e.g., x, y, and z values ​​relative to the physical environment 105, the whole world, etc.), one or more rotational values ​​574 of the physical object (e.g., roll, pitch, and yaw values), one or more pressure values ​​576 associated with the physical object (e.g., a first input (pressure) value associated with contact between the end and surface of the control device 130, a second input (pressure) value associated with the amount of pressure applied to the body of the control device 130 while it is being gripped by the user 149, etc.), optional touch input information 578 (e.g., information associated with user touch input directed to the touch-sensing surface 175 of the control device 130), and / or miscellaneous information 579. Those skilled in the art will understand that the data structure for the object tracking vector 511 in Figure 5B is merely an example of how it may contain different information components in various other implementations, and how it can be structured in countless ways in various other implementations.

[0108] In some implementations, the eye-tracking engine 512 acquires local sensor data 503 and remote sensor data 505 after being subjected to the privacy architecture 508. In some implementations, the eye-tracking engine 512 determines / generates an eye-tracking vector 513 associated with the user's 149 gaze direction based on the input data and updates the eye-tracking vector 513 over time.

[0109] Figure 5B shows an exemplary data structure for the gaze tracking vector 513 in several implementations. As shown in Figure 5B, the gaze tracking vector 513 may correspond to an N-tuple characterization vector or characterization tensor containing a timestamp 581 (e.g., the most recent time the gaze tracking vector 513 was updated), one or more angular values ​​582 (e.g., roll, pitch, and yaw values) with respect to the user 149's current gaze direction, one or more translational values ​​584 (e.g., x, y, and z values ​​with respect to the physical environment 105, the whole world, etc.), and / or other information 586. Those skilled in the art will understand that the data structure of the gaze tracking vector 513 in Figure 5B is merely an example of how it may contain different information components in various other implementations, and how it can be structured in countless ways in various other implementations.

[0110] For example, the line of sight indicates a point, physical object, or region of interest (ROI) within the physical environment 105 that user 149 is currently looking at (e.g., associated with x, y, and z coordinates relative to the physical environment 105 or the world as a whole). As another example, the line of sight indicates a point, XR object, or region of interest (ROI) within the XR environment 128 that user 149 is currently looking at (e.g., associated with x, y, and z coordinates relative to the XR environment 128).

[0111] In some implementations, the torso / head attitude tracking engine 514 acquires local sensor data 503 and remote sensor data 505 after being subjected to the privacy architecture 508. In some implementations, the torso / head attitude tracking engine 514 determines / generates an attitude feature vector 515 based on the input data and updates the attitude feature vector 515 over time.

[0112] Figure 5B shows an exemplary data structure of the pose feature vector 515 for several implementations. As shown in Figure 5B, the pose feature vector 515 may correspond to an N-tuple feature vector or feature tensor containing a timestamp 591 (e.g., the most recent time the pose feature vector 515 was updated), a head pose descriptor 592A (e.g., up, down, neutral, etc.), a translation value of the head pose 592B, a rotation value of the head pose 592C, a torso pose descriptor 594A (e.g., standing, sitting, prone, etc.), a translation value of the torso / limbs / arms / joints 594B, a rotation value of the torso / limbs / arms / joints 594C, and / or miscellaneous information 596. In some implementations, the pose feature vector 515 also includes information associated with finger / hand / limb tracking. Those skilled in the art will understand that the data structure for the pose characterization vector 515 in Figure 5B is merely an example of how different information components may be included in various other implementations, and how it can be structured in countless ways in various other implementations.

[0113] In some implementations, the data aggregation unit 521 acquires an object tracking vector 511, a gaze tracking vector 513, and a pose characterization vector 515 (which may be collectively referred to as "input vectors 519" in this specification). In some implementations, the data aggregation unit 521 aggregates the object tracking vector 511, the gaze tracking vector 513, and the pose characterization vector 515 and determines / generates a characterization vector 531 based on them for subsequent downstream use.

[0114] In some implementations, the content selection engine 522 determines the selected content portion 523 within the XR environment 128 based on the characterization vector 531 (or a part thereof). For example, the content selection engine 522 determines the selected content portion 523 based on current context information, the user's gaze direction, body posture information associated with the user 149, head posture information associated with the user 149, hand / limb tracking information associated with the user 149, position information associated with the physical object, rotation information associated with the physical object, etc. As an example, the content selection engine 522, upon determining that the grip posture associated with the way the physical object is held by the user corresponds to a first grip (e.g., first grip = pointing / wand-shaped grip), performs a selection operation on a first portion of the content based on the direction pointed to by a given part of the physical object (e.g., outward-facing end) (and the rays projected from it). As another example, the content selection engine 522, upon determining that the grip posture associated with the way the physical object is held by the user does not correspond to a first grip, performs a selection operation on a second portion of the content based on the user's gaze direction.

[0115] In some implementations, the grip posture evaluation unit 524 determines the grip posture 525 associated with the current manner in which the physical object is held by the user 149, based on the characterization vector 531 (or a part thereof). For example, the grip posture evaluation unit 524 determines the grip posture 525 based on current context information, body posture information associated with the user 149, head posture information associated with the user 149, hand / limb tracking information associated with the user 149, position information associated with the physical object, rotation information associated with the physical object, etc. In some implementations, the grip posture 525 indicates the manner in which the user 149 grasps the physical object. For example, the grip posture 525 corresponds to one of the following: remote control grip, wand grip, writing grip, reverse writing grip, handle grip, thumb-top grip, level grip, gamepad grip, flute grip, or fire-starter grip.

[0116] In some implementations, the output modality selection unit 526 selects the current output modality 527 associated with the way in which the physical object interacts with or manipulates the XR environment 128. For example, a first output modality corresponds to selecting / manipulating objects / content within the XR environment 128, and a second output modality corresponds to sketching, drawing, writing, etc., within the XR environment 128. As an example, the output modality selection unit 526 selects the associated first output modality as the current output modality 527 of the physical object based on the determination that the physical object has exceeded a distance threshold for a first graphical element among a first set of graphical elements due to the movement of the physical object. As another example, the output modality selection unit 526 selects the second output modality as the current output modality 527 of the physical object based on the determination that the physical object has exceeded a distance threshold for a second graphical element among a first set of graphical elements due to the movement of the physical object.

[0117] In some implementations, the parameter adjustment unit 528 adjusts the parameter values ​​529 associated with the marking input directed to the XR environment 128 (e.g., mark thickness, brightness, color, texture, etc.) based on either a first input (pressure) value or a second input (pressure) value associated with the physical object. For example, the parameter adjustment unit 528 adjusts the parameter values ​​529 associated with the detected marking input directed to the XR environment 128 based on how strongly the physical object is pressed against the physical surface (e.g., based on the first input (pressure) value) if the marking input is directed to a physical surface. In another example, the parameter adjustment unit 528 adjusts the parameter values ​​529 associated with the detected marking input directed to the XR environment 128 based on how strongly the physical object is being held by the user 149 (e.g., based on the second input (pressure) value) if the marking input is not directed to a physical surface. In this example, the marking input is detected while the physical object or a predetermined portion of the physical object, such as the tip of the physical object, is not in contact with any physical surface in the physical environment 105.

[0118] Figure 5C is a block diagram of a second part 500B of an exemplary content delivery architecture relating to several implementations. While relevant features are shown, those skilled in the art will understand from this disclosure that various other features are not shown in order to simplify so as not to obscure more appropriate embodiments of the exemplary implementations disclosed herein. For that purpose, as a non-limiting example, the content delivery architecture is rendered and presented by a computing system such as the controller 110 shown in Figures 1 and 2, the electronic devices 120 shown in Figures 1 and 3, and / or appropriate combinations thereof. Figure 5C is similar to and modified from Figure 5A. Therefore, the same reference numerals are used in Figures 5A and 5C. Thus, for brevity, only the differences between Figure 5A and Figure 5C are described below.

[0119] In some implementations, the interaction processing unit 540 acquires (e.g., receives, retrieves, or detects) one or more user inputs 541 provided by the user 149 related to selecting A / V content, one or more VAs, and / or XR content for presentation. For example, one or more user inputs 541 correspond to gesture inputs that modify and / or manipulate XR content or VAs in the XR environment 128 detected via hand / limb tracking, gesture inputs that select XR content in the XR environment 128 or from a UI menu detected via hand / limb tracking, gaze inputs that select XR content in the XR environment 128 or from a UI menu detected via gaze tracking, voice commands that select XR content in the XR environment 128 or from a UI menu detected via a microphone, etc. In some implementations, the content selection unit 542 selects XR content 547 from the content library 545 based on one or more user inputs 541.

[0120] In various implementations, the content management unit 530 manages and updates the layout, setup, and structure of the XR environment 128, which includes one or more of the following: VA, XR content, and one or more UI elements associated with the XR content, based on the selected content portion 523, grip pose 525, output modality 527, parameter value 529, characterization vector 531, etc. For this purpose, the content management unit 530 includes a buffer 534, a content update unit 536, and a feedback engine 538.

[0121] In some implementations, the buffer 534 contains XR content, rendered image frames, etc., for one or more past instances and / or frames. In some implementations, the content update unit 536 modifies the XR environment 128 over time based on selected content portions 523, grip poses 525, output modalities 527, parameter values ​​529, characterization vectors 531, user inputs 541 associated with modifying and / or manipulating the XR content or VA, translational or rotational movements of objects in the physical environment 105, translational or rotational movements of electronic devices 120 (or user 149), etc. In some implementations, the feedback engine 538 generates sensory feedback associated with the XR environment 128 (e.g., visual feedback such as text or lighting changes, audio feedback, haptic feedback, etc.).

[0122] In some implementations, referring to the rendering engine 550 in Figure 5C, the attitude determination unit 552 determines the current camera attitude of the electronic device 120 and / or user 149 relative to the XR environment 128 and / or physical environment 105, at least partially based on the attitude characterization vector 515. In some implementations, the rendering unit 554 renders the VA, XR content 547, one or more UI elements associated with the XR content, etc., according to the current camera attitude relative to it.

[0123] In some implementations, an optional image processing architecture 562 acquires an image stream from an image capture device 370 containing one or more images of the physical environment 105 from the current camera orientation of the electronic device 120 and / or user 149. In some implementations, the image processing architecture 562 also performs one or more image processing operations on the image stream, such as warping, color correction, gamma correction, sharpening, noise reduction, and white balance. In some implementations, an optional compositing unit 564 combines the rendered XR content with the processed image stream of the physical environment 105 from the image processing architecture 562 to generate rendered image frames of the XR environment 128. In various implementations, a presentation unit 570 presents the rendered image frames of the XR environment 128 to the user 149 via one or more displays 312. Those skilled in the art will understand that the optional image processing architecture 562 and the optional compositing unit 564 may not be applicable to a fully virtual environment (or optical see-through scenario).

[0124] Figures 6A to 6P show sequences of instances 610 to 6160 for content delivery scenarios relating to several implementations. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the implementations disclosed herein. For that purpose, as non-limiting examples, sequences of instances 610 to 6160 are rendered and presented by computing systems such as the controller 110 shown in Figures 1 and 2, electronic devices 120 shown in Figures 1 and 3, and / or appropriate combinations thereof.

[0125] As shown in Figures 6A to 6P, the content delivery scenario includes a physical environment 105 and an XR environment 128 displayed on the display 122 of an electronic device 120 (for example, associated with user 149). The electronic device 120 presents the XR environment 128 to user 149 while user 149 is physically present in the physical environment 105, including the door 115, which is currently within the FOV 111 of the outward-facing image sensor of the electronic device 120. Thus, in some implementations, user 149 holds the electronic device 120 in their left hand 150, similar to the operating environment 100 in Figure 1.

[0126] In other words, in some implementations, the electronic device 120 is configured to present XR content and enable optical see-through or video pass-through of at least a portion of the physical environment 105 on the display 122 (e.g., a door 115 or its representation). For example, the electronic device 120 may be a mobile phone, tablet, laptop, near-eye system, wearable computing device, etc.

[0127] As shown in Figure 6A, during instance 610 of the content delivery scenario (for example, associated with time T1), the electronic device 120 presents an XR environment 128 including a representation 116 of the door 115 and a virtual agent (VA) 606. As shown in Figure 6A, the control device 130 is not currently being held by the user 149 and has not detected any input directed towards its touch-sensitive surface 175.

[0128] Figures 6B and 6C illustrate a sequence in which a first set of graphical elements associated with a first set of output modalities are displayed within the XR environment 128 in response to the detection of a touch input directed to the control device 130. As shown in Figure 6B, during an instance 620 of the content delivery scenario (e.g., associated with time T2), the control device 130 detects a swipe input 622 directed to the touch-sensitive surface 175. In some implementations, the control device 130 provides instructions for the swipe input 622 to the controller 110 and / or electronic device 120. In some implementations, the control device 130 communicates with the electronic device 120 and / or controller 130.

[0129] As shown in Figure 6C, during an instance 630 of the content delivery scenario (for example, associated with time T3), the electronic device 120 displays graphical elements 632A, 632B, 632C, and 632D (which may collectively be referred to herein as the first set of graphical elements 632) in response to receiving or detecting a swipe input 622 directed to the touch-sensitive surface 175 of the control device 130 in Figure 6B.

[0130] Furthermore, as shown in Figure 6C, the electronic device 120 displays a representation 153 of the user 149's right hand 152 holding a representation 131 of the control device 130. For example, the user 149's right hand 152 is currently holding the control device 130 in a pointing grip position. In some implementations, a first set of graphical elements 632 are functions of the current grip position. For example, graphical element 632A corresponds to an output modality associated with generating a pencil-like mark in the XR environment 128, graphical element 632B corresponds to an output modality associated with generating a pen-like mark in the XR environment 128, graphical element 632C corresponds to an output modality associated with generating a marker-like mark in the XR environment 128, and graphical element 632D corresponds to an output modality associated with generating an airbrush-like mark in the XR environment 128.

[0131] In Figure 6C, the spatial location of the representation 131 of the control device 130 lies outside the activation region 634 associated with the graphical element 632D. In some implementations, the activation region 634 corresponds to a predetermined distance threshold, such as a radius of X cm, surrounding the graphical element 632D. In some implementations, the activation region 634 corresponds to a deterministic distance threshold surrounding the graphical element 632D.

[0132] Figures 6D and 6E illustrate the sequence in which a first output modality (e.g., airbrush marking) is selected for the control device 130 upon determination that the movement of the control device 130 has caused the control device 130 (or its representation) to exceed the activation region 634 (e.g., distance threshold) for the graphical element 632D. As shown in Figure 6D, during instance 640 of the content delivery scenario (e.g., associated with time T4), the electronic device 120 detects that the movement of the control device 130 causes the spatial location of the representation 131 of the control device 130 to exceed (or enter) the activation region 634 (e.g., distance threshold) for the graphical element 632D. In response to the movement of the control device 130 causing the spatial location of the representation 131 of the control device 130 to exceed the activation region 634 for the graphical element 632D, the electronic device 120 modifies the appearance of the graphical element 632D to indicate its selection by displaying a boundary or frame 642 around the graphical element 632D. Those skilled in the art will understand that the appearance of the graphical element 632D may be otherwise modified to indicate its selection, such as by changing its brightness, color, texture, shape, size, glow, shadow, and / or equivalent.

[0133] As shown in Figure 6E, during an instance 650 of the content delivery scenario (for example, associated with time T5), the electronic device 120 stops displaying graphical elements 632A, 632B, and 632C in response to detecting that the spatial position of the control device 130's representation 131 has moved beyond the activation region 634 relative to the graphical element 632D in Figure 6D.

[0134] Furthermore, in Figure 6E, the electronic device 120, in response to detecting that the spatial location of the representation 131 of the control device 130 has moved beyond the activation region 634 relative to the graphical element 632D in Figure 6D, displays the graphical element 632D superimposed on the tip of the representation 131 of the control device 130 within the XR environment 128. In some implementations, in response to the selection of the graphical element 632D, the graphical element 632D remains fixed to the tip of the representation 131 of the control device 130, as shown in Figures 6E and 6F.

[0135] Figures 6E and 6F illustrate the sequence in which detection of a marking input causes one or more marks to be displayed in the XR environment 128 according to the currently selected first output modality (e.g., airbrush marks). As shown in Figure 6E, during instance 650 of the content delivery scenario (e.g., associated with time T5), the electronic device 120 detects a marking input 654 by the control device 130 via hand / limb tracking. As shown in Figure 6F, during instance 660 of the content delivery scenario (e.g., associated with time T6), the electronic device 120 displays an airbrush-like mark 662 in the XR environment 128 in response to detecting the marking input 654 in Figure 6E. For example, the shape, depth, length, angle, etc., of the airbrush-like mark 662 correspond to the spatial parameters of the marking input 654 (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc., associated with the marking input).

[0136] Figures 6G-6I show a sequence in which a second output modality (e.g., pen marking) is selected for the control device 130 upon determination that the control device 130 (or its representation) has crossed an activation region 634 (e.g., distance threshold) for the graphical element 632B, as shown in Figure 6G. During an instance 670 of the content delivery scenario (e.g., associated with time T7), the electronic device 120 displays graphical elements 632A, 632B, 632C, and 632D (which may collectively be referred to herein as the first set of graphical elements 632) in response to receiving or detecting a swipe input 622 directed to the touch-sensitive surface 175 of the control device 130 in Figure 6B.

[0137] As shown in Figure 6H, during an instance 680 of the content delivery scenario (e.g., associated with time T8), the electronic device 120 detects that the movement of the control device 130 causes the spatial location of the control device 130's representation 131 to exceed (or enter) the activation region 634 (e.g., distance threshold) for the graphical element 632B. In response to the movement of the control device 130 causing the spatial location of the control device 130's representation 131 to exceed the activation region 634 for the graphical element 632D, the electronic device 120 modifies the appearance of the graphical element 632B to indicate its selection by displaying a border or frame 642 around the graphical element 632B. Those skilled in the art will understand that the appearance of the graphical element 632B may be modified in other ways to indicate its selection, such as by changing its brightness, color, texture, shape, size, glow, shadow, and / or equivalent.

[0138] As shown in Figure 6I, during an instance 690 of the content delivery scenario (for example, associated with time T9), the electronic device 120 stops displaying graphical elements 632A, 632C, and 632D in response to detecting that the spatial location of the representation 131 of the control device 130 has moved beyond the activation region 634 for the graphical element 632B in Figure 6H due to the movement of the control device 130. In some implementations, in response to the selection of graphical element 632B, the graphical element 632B remains fixed to the tip of the representation 131 of the control device 130, as shown in Figures 6I to 6N.

[0139] Figures 6J and 6K illustrate the sequence in which the detection of a first marking input causes one or more marks to appear within the XR environment 128 according to the currently selected second output modality (e.g., pen marking) and the current measured value of the input (pressure). As shown in Figure 6J, the content delivery scenario (e.g., time T) 10 During instance 6100 (associated with), the electronic device 120 detects the marking input 6104 by the control device 130 via hand / limb tracking. While the marking input 6104 is detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of an input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149's right hand 152. For example, the input (pressure) value is detected by one or more pressure sensors integrated into the body of the control device 130. For another example, the input (pressure) value is detected by analyzing finger / skin deformation, etc., in one or more images captured by the outward-facing image sensor of the electronic device 120 using computer vision techniques. As shown in Figure 6J, the input (pressure) value indicator 6102 shows the current measured value 6103 of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149. In some implementations, the input (pressure) value indicator 6102 is a diagram for the reader to guide, which may or may not be displayed by the electronic device 120.

[0140] As shown in Figure 6K, the content delivery scenario (for example, time T) 11 During instance 6110 (associated with), the electronic device 120 displays a pen-shaped mark 6112 in the XR environment 128 in response to detecting the marking input 6104 in Figure 6J. For example, the shape, depth, length, angle, etc., of the pen-shaped mark 6112 correspond to the spatial parameters of the marking input 6104. Furthermore, in Figure 6K, the pen-shaped mark 6112 is associated with a first thickness value that corresponds to the current measurement value 6103 of the input (pressure) value in Figure 6J.

[0141] Figures 6L and 6M illustrate the sequence in which the detection of a second marking input causes one or more marks to appear within the XR environment 128 according to the currently selected second output modality (e.g., pen marking) and the current measured value of the input (pressure). As shown in Figure 6L, the content delivery scenario (e.g., time T) 12 During instance 6120 (associated with), the electronic device 120 detects the marking input 6122 by the control device 130 via hand / limb tracking. While the marking input 6122 is detected, the electronic device 120 also detects the input (pressure) value or obtains an indication of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149's right hand 152. As shown in Figure 6L, the input (pressure) value indicator 6102 shows the current measured value 6123 of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149. The current measured value 6123 of the input (pressure) value in Figure 6L is greater than the measured value 6103 of the input (pressure) value in Figure 6J.

[0142] As shown in Figure 6M, the content delivery scenario (for example, time T) 13During instance 6130 (associated with), the electronic device 120 displays a pen-shaped mark 6132 in the XR environment 128 in response to detecting the marking input 6122 in Figure 6L. For example, the shape, depth, length, angle, etc., of the pen-shaped mark 6132 correspond to the spatial parameters of the marking input 6122. Furthermore, in Figure 6M, the pen-shaped mark 6132 is associated with a second thickness value corresponding to the current measured value 6123 of the input (pressure) value in Figure 6L. The second thickness value associated with the pen-shaped mark 6132 in Figure 6M is greater than the first thickness value associated with the pen-shaped mark 6112 in Figure 6K.

[0143] Figures 6N and 6O show a sequence in which a second set of graphical elements associated with a second set of output modalities are displayed in the XR environment 128 in response to the detection of touch input directed to the control device 130. As shown in Figure 6N, the content delivery scenario (e.g., time T) 14 During instance 6140 (associated with), the control device 130 detects a swipe input 6142 directed to the touch-sensing surface 175. In some implementations, the control device 130 provides instructions for the swipe input 6142 to the controller 110 and / or electronic device 120.

[0144] As shown in Figure 6O, the content delivery scenario (for example, time T) 15 During instance 6150 (associated with), the electronic device 120 displays graphical elements 6152A, 6152B, 6152C, and 6152D (which may collectively be referred to herein as the second set of graphical elements 6152) in response to receiving an instruction for a swipe input 6142 directed to the touch-sensitive surface 175 of the control device 130 in Figure 6N, or detecting a swipe input 6142 directed to the touch-sensitive surface 175 of the control device 130 in Figure 6N.

[0145] Furthermore, as shown in Figure 6O, the electronic device 120 displays a representation 153 of the user 149's right hand 152 holding a representation 131 of the control device 130. For example, the user 149's right hand 152 is currently holding the control device 130 in a writing grip position with the first end 176 pointed downwards and the second end 177 pointed upwards. In some implementations, the second set of graphical elements 6152 are a function of the current grip position (e.g., writing grip position). For example, graphical element 6152A corresponds to an output modality associated with generating pencil-shaped marks within the XR environment 128, graphical element 6152B corresponds to an output modality associated with generating pen-shaped marks within the XR environment 128, graphical element 6152C corresponds to an output modality associated with generating thin brush-shaped marks within the XR environment 128, and graphical element 6152D corresponds to an output modality associated with generating thick brush-shaped marks within the XR environment 128.

[0146] Figures 6O and 6P show a sequence in which a third set of graphical elements associated with a third set of output modalities are displayed within the XR environment 128 in response to the detection of a change in the current grip posture of the control device 130. As shown in Figure 6O, the content delivery scenario (e.g., time T) 15 During instance 6150 (associated with), the electronic device 120 uses computer vision techniques to detect the current grip posture of the control device 130, thereby indicating that the user 149's right hand 152 is currently holding the control device 130 in a writing grip posture with the first end 176 pointing downwards and the second end 177 pointing upwards. However, between Figure 6O and Figure 6P, the electronic device 120 detects a change in the current grip posture of the control device 130 from the writing grip posture in Figure 6O to the reverse writing grip posture in Figure 6P.

[0147] As shown in Figure 6P, the content delivery scenario (for example, time T) 16During instance 6160 (associated with Figure 6O), the electronic device 120 uses computer vision techniques to detect the current grip orientation of the control device 130, thereby indicating that the user 149's right hand 152 is currently holding the control device 130 in a reverse writing grip position with the first end 176 pointing upward and the second end 177 pointing downward. Thus, the control device 130 is inverted 180 degrees with respect to the orientation of its ends between Figure 6O and Figure 6P. In Figure 6P, in response to detecting the change in the current grip orientation of the control device 130 from the writing grip position in Figure 6O to the reverse writing grip position in Figure 6P, the electronic device 120 displays graphical elements 6162A, 6162B, 6162C, and 6162D (which may collectively be referred to herein as a third set of graphical elements 6162).

[0148] In some implementations, a third set of graphical elements 6162 is a function of the current grip pose (e.g., reverse writing grip pose). For example, graphical element 6162A corresponds to an output modality associated with erasing or removing pixels in the XR environment 128 based on a first radius value, graphical element 6162B corresponds to an output modality associated with erasing or removing pixels in the XR environment 128 based on a second radius value greater than the first radius value, graphical element 6162C corresponds to an output modality associated with a measurement mark in the XR environment 128, and graphical element 6162D corresponds to an output modality associated with a cut mark in the XR environment 128.

[0149] Figures 7A to 7N show sequences of instances 710 to 7140 for content delivery scenarios relating to several implementations. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the implementations disclosed herein. For that purpose, as non-limiting examples, sequences of instances 710 to 7140 are rendered and presented by computing systems such as the controller 110 shown in Figures 1 and 2, electronic devices 120 shown in Figures 1 and 3, and / or appropriate combinations thereof.

[0150] As shown in Figures 7A to 7N, the content delivery scenario includes a physical environment 105 and an XR environment 128 displayed on the display 122 of an electronic device 120 (for example, associated with user 149). The electronic device 120 presents the XR environment 128 to user 149 while user 149 is physically present in the physical environment 105, which includes the position of table 107, currently within the FOV 111 of the outward-facing image sensor of the electronic device 120. Thus, in some implementations, user 149 holds the electronic device 120 in their left hand 150 or right hand 152.

[0151] In other words, in some implementations, the electronic device 120 is configured to present XR content and enable optical see-through or video pass-through of at least a portion of the physical environment 105 on the display 122 (e.g., table 107). For example, the electronic device 120 may be a mobile phone, tablet, laptop, near-eye system, wearable computing device, etc.

[0152] As shown in Figure 7A, during an instance 710 of the content delivery scenario (e.g., associated with time T1), the electronic device 120 presents an XR environment 128 that includes a portion of Table 107, a virtual agent (VA) 606, an XR substrate 718 (e.g., a 2D or 3D canvas), and a menu 712. As shown in Figure 7A, the electronic device 120 also displays a representation 151 of the left hand 150 of user 149 holding a representation 131 of the control device 130 within the XR environment 128. For example, user 149's left hand 150 is currently holding the control device 130 in a writing grip position. As shown in Figure 7A, the menu 712 includes several selectable options 714 associated with changing the appearance of a mark created within the XR environment (e.g., different color, texture, etc.). For example, option 714A is currently selected from among the several selectable options 714. In this example, option 714A corresponds to a first appearance of a mark (e.g., a black mark) created within the XR environment 128. As shown in Figure 7A, menu 712 also includes a slider 716 for adjusting the thickness of marks created within the XR environment 128.

[0153] Figures 7A and 7B illustrate the sequence by which the detection of a first marking input causes one or more marks to be displayed in the XR environment 128 according to a first measurement of the input (pressure) value. As shown in Figure 7A, during an instance 710 of the content delivery scenario (e.g., associated with time T1), the electronic device 120 detects a marking input 715 by the control device 130 via hand / limb tracking. While the marking input 715 is detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of an input (pressure) value associated with how tightly the control device 130 is being grasped by the user's left hand 150. As an example, the input (pressure) value is detected by one or more pressure sensors integrated into the body of the control device 130. As another example, the input (pressure) value is detected by analyzing finger / skin deformation, etc., in one or more images captured by the outward-facing image sensor of the electronic device 120 using computer vision techniques. As shown in Figure 7A, the input (pressure) value indicator 717 displays the current measured value 719 of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149. In some implementations, the input (pressure) value indicator 717 may or may not be displayed by the electronic device 120 and is a diagram for the reader to observe.

[0154] As shown in Figure 7B, during an instance 720 of the content delivery scenario (e.g., associated with time T2), the electronic device 120 displays a mark 722 on the XR substrate 718 in the XR environment 128 in response to detecting the marking input 715 in Figure 7A. For example, the shape, depth, length, angle, etc., of the mark 722 correspond to the spatial parameters of the marking input 715 (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc., associated with the marking input). Furthermore, in Figure 7B, the mark 722 is associated with a first thickness value corresponding to the current measurement value 719 of the input (pressure) value in Figure 7A.

[0155] Figures 7C and 7D illustrate the sequence in which the detection of a second marking input causes one or more marks to appear within the XR environment 128 according to a second measurement of the input (pressure) value. As shown in Figure 7C, during an instance 730 of the content delivery scenario (e.g., associated with time T3), the electronic device 120 detects a marking input 732 by the control device 130 via hand / limb tracking. While the marking input 732 is detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of an input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149's left hand 150. As shown in Figure 7C, the input (pressure) value indicator 717 shows the current measurement 739 of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149. For example, the current measurement 739 in Figure 7C is greater than the measurement 719 in Figure 7A.

[0156] As shown in Figure 7D, during an instance 740 of the content delivery scenario (for example, associated with time T4), the electronic device 120 displays a mark 742 on the XR substrate 718 in the XR environment 128 in response to detecting a marking input 732 in Figure 7C. For example, the shape, depth, length, angle, etc., of the mark 742 correspond to the spatial parameters of the marking input 732. Furthermore, in Figure 7D, the mark 742 is associated with a second thickness value corresponding to the current measurement value 739 of the input (pressure) value in Figure 7C. For example, the second thickness value associated with the mark 742 is greater than the first thickness value associated with the mark 722.

[0157] Figures 7E and 7F illustrate a sequence in which one or more marks are translated within the XR environment 128 by the detection of an operation input. As shown in Figure 7E, during an instance 750 of the content delivery scenario (e.g., associated with time T5), the electronic device 120 detects an operation input 752 by the control device 130 corresponding to translating the mark 742 within the XR environment 128. While the operation input 752 is detected, the electronic device 120 also detects a touch input 754 directed to the touch-sensitive surface 175 of the control device 130, or obtains instructions for a touch input 754 directed to the touch-sensitive surface 175 of the control device 130. As an example, the touch input 754 is detected by the touch sensor 175 of the control device 130. As another example, the touch input 754 is detected by analyzing one or more images captured by the outward-facing image sensor of the electronic device 120 using computer vision techniques.

[0158] As shown in Figure 7F, during instance 760 of the content delivery scenario (for example, associated with time T6), the electronic device 120 translates the mark 742 in the XR environment 128 in response to detecting the operation input 752 in Figure 7E, while also detecting the touch input 754 directed to the touch-sensitive surface 175 of the control device 130 in Figure 7E. In some implementations, the detection of the operation input 752 may be sufficient to cause the translation of the mark in the XR environment 128 without detecting the touch input 754 directed to the touch-sensitive surface 175 of the control device 130. In some implementations, the detection of the operation input 752 in conjunction with the detection of the touch input 754 directed to the touch-sensitive surface 175 of the control device 130 causes the translation of the mark in the XR environment 128. For example, the angle, direction, displacement, etc. of the translation of the mark 742 correspond to the spatial parameters of the operation input 752 in Figure 7E. In some implementations, the operation input 752 may also cause a rotational movement of the mark 742 based on the rotation parameter of the operation input 752.

[0159] Figures 7G and 7H illustrate a sequence in which the detection of a first marking input causes one or more marks to appear in the XR environment 128 according to a first measurement of the input (pressure) value. As shown in Figure 7G, during an instance 770 of the content delivery scenario (e.g., associated with time T7), the electronic device 120 detects a marking input 772 by a control device 130 directed to an input area 774 on the table 107 via hand / limb tracking. For example, the input area 774 corresponds to a portion of a plane associated with the surface of the table 107. In some implementations, the input area 774 is visualized, such as by an XR boundary. In some implementations, the input area 774 is not visualized within the XR environment 128.

[0160] While the marking input 772 is detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of an input (pressure) value associated with how hard the control device 130 is pressed against the table 107. As an example, the input (pressure) value is detected by one or more pressure sensors integrated into one of the tips of the control device 130. As another example, the input (pressure) value is detected by analyzing one or more images captured by the outward-facing image sensor of the electronic device 120 using computer vision techniques. As shown in Figure 7G, the input (pressure) value indicator 777 shows the current measured value 779 of the input (pressure) value associated with how hard the control device 130 is pressed against the table 107. According to some implementations, the input (pressure) value indicator 777 is a diagram for reader guidance that may or may not be displayed by the electronic device 120.

[0161] As shown in FIG. 7H, during an instance 780 of a content delivery scenario (e.g., associated with time T8), in response to detecting the marking input 715 in FIG. 7G, the electronic device 120 displays a mark 782B on the XR substrate 718 within the XR environment 128 and a mark 782 on the input region 774. For example, the shape, depth, length, angle, etc. of marks 782A and 782B correspond to the spatial parameters of the marking input 772 in FIG. 7G (e.g., position values, rotation values, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc. associated with the marking input). Further, in FIG. 7H, marks 782A and 782B are associated with a first thickness value corresponding to the current measured value 779 of the input (pressure) value in FIG. 7G.

[0162] FIGS. 7I and 7J show a sequence in which the detection of a second marking input causes one or more marks to be displayed within the XR environment 128 according to a second measured value of the input (pressure) value. As shown in FIG. 7I, during an instance 790 of a content delivery scenario (e.g., associated with time T9), the electronic device 120 detects a marking input 792 by the control device 130 directed at the input region 774 on the table 107 by hand / limb tracking. While the marking input 792 is being detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of the input (pressure) value associated with how strongly the control device 130 is pressed against the table 107. As shown in FIG. 7I, the input (pressure) value indicator 777 shows the current measured value 799 of the input (pressure) value associated with how strongly the control device 130 is pressed against the table 107. For example, the current measured value 799 in FIG. 7I is greater than the measured value 779 in FIG. 7G.

[0163] As shown in FIG. 7J, during an instance of a content delivery scenario (e.g., time T 10During instance 7100 (associated with), the electronic device 120, in response to detecting the marking input 792 in Figure 7I, displays mark 7102A on the XR substrate 718 in the XR environment 128 and mark 7102B on the input area 774. For example, the shape, depth, length, angle, etc., of marks 7102A and 7102B correspond to the spatial parameters of the marking input 792 in Figure 7I (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc., associated with the marking input). Furthermore, in Figure 7J, marks 7102A and 7102B are associated with a second thickness value corresponding to the current measurement value 799 of the input (pressure) value in Figure 7I. For example, the second thickness value associated with marks 7102A and 7102B is greater than the first thickness value associated with marks 782A and 782B.

[0164] Figures 7K and 7L illustrate the sequence in which the detection of the first content placement input causes the first XR content to be displayed within the XR environment 128 according to the current measurement of the input (pressure) value. As shown in Figure 7K, the content delivery scenario (e.g., time T) 11 During instance 7110 (associated with), the electronic device 120 presents an XR environment 128, which includes a portion of table 107, VA606, an XR substrate 7118 (e.g., a planar substrate), and a menu 7112. As shown in Figure 7K, the menu 7112 includes several selectable options 7114 associated with changing the appearance of the XR content placed within the XR environment (e.g., different shapes, colors, textures, etc.). For example, option 7114A is currently selected from among the several selectable options 7114. In this example, option 7114A corresponds to a first appearance of the XR content placed within the XR environment 128. As shown in Figure 7A, the menu 7112 also includes a slider 7116 for adjusting the size of the XR content placed within the XR environment 128.

[0165] As shown in Figure 7K, the content delivery scenario (for example, time T) 11During instance 7110 (associated with), the electronic device 120 detects a touch input 7111 directed to the touch-sensitive surface 175 of the control device 130, or receives instructions for a touch input 7111 directed to the touch-sensitive surface 175 of the control device 130, while the representation 131 of the control device 130 is at a distance 7115 above the XR substrate 7118. As an example, the touch input 7111 is detected by the touch-sensitive surface 175 of the control device 130. As another example, the touch input 7111 is detected by analyzing one or more images captured by the outward-facing image sensor of the electronic device 120 using computer vision techniques. For example, the touch input 7111 corresponds to placing XR content (e.g., a cube) within the XR environment 128. It will be understood by those skilled in the art that other XR content can similarly be placed within the XR environment 128.

[0166] While touch input 7111 is detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of an input (pressure) value associated with how tightly the control device 130 is being gripped by the user's left hand 150. As shown in Figure 7K, the input (pressure) value indicator 717 shows the current measured value 7119 of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149. According to some implementations, the input (pressure) value indicator 717 is a diagram for the reader to guide, which may or may not be displayed by the electronic device 120.

[0167] As shown in Figure 7L, the content delivery scenario (for example, time T) 12During instance 7120 (associated with), the electronic device 120 displays the first XR content 7122 at a distance 7115 above the XR substrate 7118 in the XR environment 128 in response to detecting the touch input 7111 in Figure 7K. As shown in Figure 7L, the electronic device 120 also displays a shadow 7124 associated with the first XR content 7122 on the XR substrate 7118. For example, the position and rotation values ​​of the first XR content 7122 and shadow 7124 correspond to the parameters (e.g., position value, rotation value, etc.) of the representation 131 of the control device 130 when the touch input 7111 is detected in Figure 7K. For example, the first XR content 7122 is associated with a first size value corresponding to the current measurement 7119 of the input (pressure) value in Figure 7K.

[0168] Figures 7M and 7N illustrate the sequence in which the detection of the second content placement input causes the second XR content to be displayed within the XR environment 128 according to the current measurement of the input (pressure) value. As shown in Figure 7M, the content delivery scenario (e.g., time T) 13 During instance 7130 (associated with), the electronic device 120 detects a touch input 7131 directed to the touch-sensing surface 175 of the control device 130, or obtains an instruction for a touch input 7131 directed to the touch-sensing surface 175 of the control device 130, while the representation 131 of the control device 130 is in contact with the XR substrate 7118.

[0169] While touch input 7131 is detected, the electronic device 120 also detects an input (pressure) value or obtains an indication of an input (pressure) value associated with how tightly the control device 130 is being gripped by the user's left hand 150. As shown in Figure 7M, the input (pressure) value indicator 717 shows the current measured value 7139 of the input (pressure) value associated with how tightly the control device 130 is being gripped by the user 149. According to some implementations, the input (pressure) value indicator 717 is a diagram for the reader to guide, which may or may not be displayed by the electronic device 120.

[0170] As shown in Figure 7N, during instance 7140 of the content delivery scenario (for example, associated with time T14), the electronic device 120 displays the second XR content 7142 on the XR substrate 7118 in the XR environment 128 in response to detecting the touch input 7131 in Figure 7M. As shown in Figure 7N, the electronic device 120 does not display the shadow associated with the second XR content 7142 on the XR substrate 7118. For example, the position and rotation values ​​of the second XR content 7142 correspond to the parameters (e.g., position value, rotation value, etc.) of the representation 131 of the control device 130 when the touch input 7131 was detected in Figure 7M. For example, the second XR content 7142 is associated with a second size value corresponding to the current measurement 7139 of the input (pressure) value in Figure 7K. For example, the second size value associated with the second XR content 7142 is greater than the first size value associated with the first XR content 7122.

[0171] Figures 8A to 8M show sequences of instances 810 to 8130 for content delivery scenarios relating to several implementations. While certain features are shown, those skilled in the art will understand from this disclosure that various other features have been omitted for brevity so as not to obscure more appropriate embodiments of the implementations disclosed herein. For that purpose, as non-limiting examples, sequences of instances 810 to 8130 are rendered and presented by computing systems such as the controller 110 shown in Figures 1 and 2, electronic devices 120 shown in Figures 1 and 3, and / or appropriate combinations thereof.

[0172] As shown in Figures 8A to 8M, the content delivery scenario includes a physical environment 105 and an XR environment 128 displayed on the display 122 of an electronic device 120 (for example, associated with user 149). The electronic device 120 presents the XR environment 128 to user 149 while user 149 is physically present in the physical environment 105, including a door 115, which is currently within the FOV 111 of the outward-facing image sensor of the electronic device 120. Thus, in some implementations, user 149 holds the electronic device 120 in their left hand 150 or right hand 152.

[0173] In other words, in some implementations, the electronic device 120 is configured to present XR content and enable optical see-through or video pass-through of at least a portion of the physical environment 105 on the display 122 (e.g., a door 115 or its representation). For example, the electronic device 120 may be a mobile phone, tablet, laptop, near-eye system, wearable computing device, etc.

[0174] As shown in Figure 8A, during an instance 810 of the content delivery scenario (e.g., associated with time T1), the electronic device 120 presents an XR environment 128 containing a representation 116 of a door 115 in the physical environment 105, a virtual agent (VA) 606, and XR content 802 (e.g., a cylinder). As shown in Figure 8A, the electronic device 120 also displays the gaze direction 806 associated with the focus of the user 149's eye in the XR environment 128, based on eye tracking. In some implementations, the gaze direction 806 is neither displayed nor visualized. As shown in Figure 8A, the electronic device 120 also displays a representation 153 of the user 149's right hand 152 holding a representation 805 of a proxy object 804 (e.g., a stick, ruler, or another physical object) in the XR environment 128. For example, the user 149's right hand 152 is currently holding the proxy object 804 in a pointing grip position.

[0175] Figures 8A and 8B illustrate a sequence in which the size of an indicator element changes based on distance. As shown in Figure 8A, the representation 805 of the proxy object 804 is at a first distance 814 from the XR content 802, and the electronic device 120 displays a first indicator element 812A on the XR content 802 having a first size corresponding to the point of agreement between the XR content 802 and the rays emanating from the leading edge / end of the representation 805 of the proxy object 804. In some implementations, the size of the first indicator element 812A is a function of the first distance 814. For example, the size of the indicator element increases as the distance decreases, and the size of the indicator element decreases as the distance increases.

[0176] As shown in Figure 8B, during an instance 820 of the content delivery scenario (for example, associated with time T2), the electronic device 120 displays a second indicator element 812B on the XR content 802, having a second size corresponding to the point of agreement between the XR content 802 and the rays emanating from the leading edge / end of the representation 805 of the proxy object 804. As shown in Figure 8B, the representation 805 of the proxy object 804 is at a second distance 824 from the XR content 802, which is smaller than the first distance 814 in Figure 8A. For example, the second size of the second indicator element 812B is larger than the first size of the first indicator element 812A.

[0177] Figures 8B to 8D show a sequence in which XR content is selected in response to detection that a proxy object is directed towards XR content, and the XR content is translated in response to translational movement of the proxy object. In some implementations, the electronic device 120 selects XR content 802 and modifies its appearance in response to detection that a proxy object 804 has been directed towards XR content 802 for at least a predetermined period of time. In some implementations, the electronic device 120 selects XR content 802 and modifies its appearance in response to detection that a proxy object 804 has been directed towards XR content 802 for at least a deterministic period of time.

[0178] As shown in Figure 8C, during an instance 830 of the content delivery scenario (for example, associated with time T3), the electronic device 120, in response to detecting that the proxy object 804 has been directed towards the XR content 802 for at least a predetermined or deterministic period of time in Figures 8A and 8B, changes the appearance of the XR content 802 to a cross-hatched appearance 802A to visually indicate its selection. As shown in Figure 8C, the electronic device 120 also detects a translational movement 832 of the proxy object 804 while the XR content 802A is selected and the representation 805 of the proxy object 804 remains directed towards the XR content 802A.

[0179] As shown in Figure 8D, during instance 840 of the content delivery scenario (for example, associated with time T4), the electronic device 120 translates the XR content 802 in the XR environment 128 in response to detecting the translational movement 832 of the proxy object 804 in Figure 8C. For example, the direction and displacement of the translational movement of the XR content 802 in the XR environment 128 correspond to the spatial parameters of the translational movement 832 in Figure 8C (e.g., change in position value, change in rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc.). It will be understood by those skilled in the art that the XR content 802 can also be rotated.

[0180] Figures 8E to 8G show a sequence in which XR content is selected in response to the detection of a gaze direction directed towards the XR content, and the XR content is transformed based on the translational movement of the gaze direction. In some implementations, the electronic device 120 selects XR content 802 and changes its appearance in response to the detection that the gaze direction 806 has been directed towards the XR content 802 for at least a predetermined period of time. In some implementations, the electronic device 120 selects XR content 802 and changes its appearance in response to the detection that the gaze direction 806 has been directed towards the XR content 802 for at least a deterministic period of time.

[0181] As shown in Figure 8E, during an instance 850 of the content delivery scenario (for example, associated with time T5), the electronic device 120 displays a gaze direction indicator element 852 on the XR content 802 associated with the focus of the user 149's eye in the XR environment 128, based on eye tracking. For example, the gaze direction indicator element 852 corresponds to the point of agreement between the XR content 802 and the ray emanating from the user 149's eye. As shown in Figure 8E, the electronic device 120 also displays a representation 153 of the user 149's right hand 152 holding a representation 131 of the control device 130 in the XR environment 128. For example, the user 149's right hand 152 is currently holding the control device 130 in a writing grip position that is not directed at any XR content in the XR environment 128.

[0182] As shown in Figure 8F, during an instance 860 of the content delivery scenario (for example, associated with time T6), the electronic device 120, in response to detecting a gaze direction 806 directed towards the XR content 802 for at least a predetermined or deterministic period in Figure 8E, changes the appearance of the XR content 802 to a cross-hatched appearance 802A to visually indicate its selection. As shown in Figure 8F, the electronic device 120 also detects a translational movement 862 of the gaze direction 806 while the XR content 802A is selected.

[0183] As shown in Figure 8G, during instance 870 of the content delivery scenario (for example, associated with time T7), the electronic device 120 translates the XR content 802 in the XR environment 128 in response to detecting the translational movement 862 in the line of sight direction 806 in Figure 8F. For example, the direction and displacement of the translational movement of the XR content 802 in the XR environment 128 correspond to the spatial parameters of the translational movement 862 in the line of sight direction 806 in Figure 8F (e.g., change in position value, change in rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc.).

[0184] Figures 8H to 8J illustrate another sequence in which XR content is selected in response to the detection of a gaze direction directed towards the XR content, and the XR content is translated based on the translational movement of the gaze direction. As shown in Figure 8H, during instance 880 of the content delivery scenario (e.g., associated with time T8), the electronic device 120 displays a gaze direction indicator element 852 on the XR content 802 associated with the focus in the XR environment 128 of the user 149's eye, based on gaze tracking. For example, the gaze direction indicator element 852 corresponds to a point of agreement between the XR content 802 and the ray emanating from the user 149's eye. As shown in Figure 8H, the electronic device 120 detects that neither the proxy object 804 nor the control device 130 is held by the user 149.

[0185] As shown in Figure 8I, during an instance 890 of the content delivery scenario (for example, associated with time T9), the electronic device 120, in response to detecting a gaze direction 806 directed towards the XR content 802 for at least a predetermined or deterministic period in Figure 8H, changes the appearance of the XR content 802 to a cross-hatched appearance 802A to visually indicate its selection. As shown in Figure 8I, the electronic device 120 also detects a translational movement 892 of the gaze direction 806 while the XR content 802A is selected.

[0186] As shown in Figure 8J, the content delivery scenario (for example, time T)10 During instance 8100 (associated with), the electronic device 120 translates the mark 802 in the XR environment 128 in response to detecting the translational movement 892 in the line of sight direction 806 in Figure 8I. For example, the direction and displacement of the translational movement of the XR content 802 in the XR environment 128 correspond to the spatial parameters of the translational movement 892 in the line of sight direction 806 in Figure 8I (e.g., change in position value, change in rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc.).

[0187] Figures 8K to 8M show a sequence in which XR content is selected in response to the detection of the gaze direction directed towards the XR content, and the XR content is translated based on hand / limb tracking input. As shown in Figure 8K, the content delivery scenario (e.g., time T) 11 During instance 8110 (associated with), the electronic device 120 displays a gaze direction indicator element 852 on the XR content 802 associated with the focus in the XR environment 128 of the user 149's eye, based on gaze tracking. For example, the gaze direction indicator element 852 corresponds to a point of agreement between the XR content 802 and the ray emanating from the user 149's eye. As shown in Figure 8K, the electronic device 120 detects that neither the proxy object 804 nor the control device 130 is held by the user 149.

[0188] As shown in Figure 8L, the content delivery scenario (for example, time T) 12 During instance 8120 (associated with), the electronic device 120, in response to detecting a gaze direction 806 directed towards the XR content 802 for at least a predetermined or deterministic period in Figure 8K, changes the appearance of the XR content 802 to a cross-hatched appearance 802A to visually indicate its selection. As shown in Figure 8L, the electronic device 120 displays a representation 153 of the user 149's right hand 152 near the XR content 802A, which is detected and tracked using hand / limb tracking. As shown in Figure 8L, the electronic device 120 also detects the translational movement 8122 of the user 149's right hand 152.

[0189] As shown in Figure 8M, the content delivery scenario (for example, time T) 13 During instance 8130 (associated with the XR environment 128), the electronic device 120 translates the XR content 802 in the XR environment 128 in response to detecting the translational movement 8122 in Figure 8L. For example, the direction and displacement of the translational movement of the XR content 802 in the XR environment 128 correspond to the spatial parameters of the translational movement 8122 in Figure 8L (e.g., change in position value, change in rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc.). It will be understood by those skilled in the art that other XR content can be rotated in the same way.

[0190] Figures 9A to 9C show flowchart representations of method 900 for selecting the output modality of a physical object when interacting with or manipulating an XR environment, relating to several implementations. In various implementations, method 900 is executed in a computing system including non-temporary memory and one or more processors, the computing system being communicatively coupled to a display device and (optional) one or more input devices (e.g., electronic device 120 shown in Figures 1 and 3, controller 110 in Figures 1 and 2, or an appropriate combination thereof). In some implementations, method 900 is executed by processing logic, which includes hardware, firmware, software, or a combination thereof. In some implementations, method 900 is executed by a processor that executes code stored in a non-temporary computer-readable medium (e.g., memory). In some implementations, the computing system corresponds to one of the following: a tablet, laptop, mobile phone, near-eye system, wearable computing device, etc.

[0191] Typically, users switch marking tools by selecting a new tool from a toolbar or menu. This can interrupt the user's current workflow and may lead them to search for a new tool for appropriate control. In contrast, the method described herein allows users to invoke a toolset display by swiping on a physical object (e.g., a proxy object such as a pencil, or an electronic device such as a stylus) and moving the physical object toward a graphical representation of one of the tools in the toolset. Thus, users can switch tools without interrupting their workflow.

[0192] As represented by block 902, method 900 includes displaying a first set of graphical elements associated with a first set of output modalities in an augmented reality (XR) environment via a display device. In some implementations, the output modalities cause changes within the UI or XR environment, such as adding, removing, or otherwise modifying pixels within the UI or XR environment. For example, the first set of graphical elements correspond to different tool types for creating, modifying, etc., marks within the XR environment, such as pencils, markers, paintbrushes, and erasers.

[0193] As an example, in Figure 6C, the electronic device 120 displays graphical elements 632A, 632B, 632C, and 632D (which may collectively be referred to as the first set of graphical elements 632 in this specification). In this example, graphical element 632A corresponds to an output modality associated with generating pencil-like marks in the XR environment 128, graphical element 632B corresponds to an output modality associated with generating pen-like marks in the XR environment 128, graphical element 632C corresponds to an output modality associated with generating marker-like marks in the XR environment 128, and graphical element 632D corresponds to an output modality associated with generating airbrush-like marks in the XR environment 128. As another example, in Figure 6O, the electronic device 120 displays graphical elements 6152A, 6152B, 6152C, and 6152D (which may collectively be referred to as the second set of graphical elements 6152 in this specification). In this example, graphical element 6152A corresponds to the output modality associated with generating pencil-shaped marks within the XR environment 128, graphical element 6152B corresponds to the output modality associated with generating pen-shaped marks within the XR environment 128, graphical element 6152C corresponds to the output modality associated with generating thin brush-shaped marks within the XR environment 128, and graphical element 6152D corresponds to the output modality associated with generating thick brush-shaped marks within the XR environment 128.

[0194] In some implementations, the display device corresponds to a transparent lens assembly, and the presentation of the XR environment is projected onto the transparent lens assembly. In some implementations, the display device corresponds to a near-eye system, and the presentation of the XR environment includes compositing the presentation of the XR environment with one or more images of the physical environment captured by an outward-facing image sensor.

[0195] In some implementations, Method 900 includes obtaining (e.g., receiving, retrieving, or detecting) a touch input directed to a physical object before displaying a first set of graphical elements, and displaying the first set of graphical elements in the XR environment includes displaying the first set of graphical elements in the XR environment in response to obtaining the touch input instruction. In some implementations, the physical object corresponds to a stylus having a touch-sensitive area capable of detecting touch input. For example, the stylus detects an upward or downward swipe gesture on its touch-sensitive surface, and the computing system obtains (e.g., receives or retrieves) a touch input instruction from the stylus. For example, Figures 6B and 6C show a sequence in which an electronic device 120 displays a first set of graphical elements 632 associated with a first set of output modalities in the XR environment 128 of Figure 6C in response to detecting a touch input 622 directed to a control device 130 in Figure 6B.

[0196] In some implementations, method 900 includes obtaining (e.g., receiving, taking, or determining) a grip pose associated with the current form in which a physical object is held by the user, before displaying a first set of graphical elements, wherein the first set of graphical elements is a function of the grip pose, and, in response to obtaining the grip pose, displaying the first set of graphical elements associated with the first set of output modalities in the XR environment via a display device, according to the determination that the grip pose corresponds to the first grip pose, and displaying the second set of graphical elements associated with the second set of output modalities in the XR environment via a display device, according to the determination that the grip pose corresponds to a second grip pose different from the first grip pose. For example, a pointing / wand grip corresponds to the first set of graphical elements associated with the first set of tools, and a writing grip corresponds to the second set of graphical elements associated with the second set of tools. In some implementations, the first and second set of output modalities include at least one overlapping output modality. In some implementations, the first and second multiple output modalities include mutually exclusive output modalities.

[0197] As an example, referring to Figure 6C, the electronic device 120 displays a graphical element 632 in the XR environment 128 based on the determination that the current grip position corresponds to a pointing grip position. As another example, referring to Figure 6O, the electronic device 120 displays a graphical element 6152 in the XR environment 128 based on the determination that the current grip position corresponds to a writing grip position in which the first end 176 is pointed downwards and the second end 177 is pointed upwards. As yet another example, referring to Figure 6P, the electronic device 120 displays a graphical element 6162 in the XR environment 128 based on the determination that the current grip position corresponds to a reverse writing grip position in which the first end 176 is pointed upwards and the second end 177 is pointed downwards.

[0198] In some implementations, method 900 includes, after displaying a first set of graphical elements associated with a first set of output modalities in the XR environment, detecting a change in grip posture from a first grip posture to a second grip posture, and, in response to detecting the change in grip posture, replacing the display of the first set of graphical elements in the XR environment with a second set of graphical elements associated with a second set of output modalities in the XR environment. In some implementations, the computing system also stops displaying the first set of graphical elements. For example, a pointing / wand grip corresponds to a first set of graphical elements associated with a first set of output modalities, and a writing grip corresponds to a second set of graphical elements associated with a second set of output modalities. In some implementations, the first and second sets of graphical elements include at least some overlapping output modalities. In some implementations, the first and second sets of graphical elements include mutually exclusive output modalities. For example, Figures 6O and 6P show a sequence in which the electronic device 120 replaces multiple graphical elements 6152 with multiple graphical elements 6162 in response to detecting a change in the current grip posture of the control device 130 (for example, a change from the writing grip posture in Figure 6O to the reverse writing grip posture in Figure 6P).

[0199] In some implementations, Method 900 includes obtaining (e.g., receiving, retrieving, or determining) information indicating whether a first or second end of a physical object is outward-facing (e.g., outward-facing toward a surface, user, computing system, etc.) before displaying a first set of graphical elements; displaying a first set of graphical elements associated with a first set of output modalities in an XR environment via a display device, in accordance with the determination that the first end of the physical object is outward-facing, based on the information obtained regarding whether the first or second end of the physical object is outward-facing; and displaying a second set of graphical elements associated with a second set of output modalities in an XR environment via a display device, in accordance with the determination that the second end of the physical object is outward-facing. For example, the outward-facing first end corresponds to a first set of graphical elements associated with a first set of output modalities (e.g., sketch and writing tools), and the outward-facing second end corresponds to a second set of graphical elements associated with a second set of output modalities (e.g., erase or edit tools). In some implementations, the first and second sets of output modalities include at least one overlapping set of output modalities. In some implementations, the first and second sets of output modalities include mutually exclusive set of output modalities.

[0200] As another example, referring to Figure 6O, the electronic device 120 displays a graphical element 6152 in the XR environment 128 based on the determination that the current grip position corresponds to a writing grip position in which the first end 176 is pointed downwards and the second end 177 is pointed upwards. As yet another example, referring to Figure 6P, the electronic device 120 displays a graphical element 6162 in the XR environment 128 based on the determination that the current grip position corresponds to a reverse writing grip position in which the first end 176 is pointed upwards and the second end 177 is pointed downwards.

[0201] In some implementations, Method 900 includes displaying a first set of graphical elements associated with a first set of output modalities in the XR environment, detecting a change from a first outward-facing end of a physical object to a second outward-facing end of a physical object, and, in response to detecting the change from the first outward-facing end of a physical object to the second outward-facing end of a physical object, displaying a second set of graphical elements associated with a second set of output modalities in the XR environment via a display device. In some implementations, the computing system also stops displaying the first set of graphical elements. For example, the first outward-facing end corresponds to a first set of graphical elements associated with a first set of output modality tools (e.g., a sketch tool and a writing tool), and the second outward-facing end corresponds to a second set of graphical elements associated with a second set of output modalities (e.g., an erase tool or an edit tool). For example, Figures 6O and 6P show a sequence in which the electronic device 120 replaces the display of multiple graphical elements 6152 with multiple graphical elements 6162 in response to detecting a change in the current grip posture of the control device 130 (for example, a change from a writing grip posture in Figure 6O to a reverse writing grip posture in Figure 6P).

[0202] As represented by block 904, method 900 includes detecting a first movement of a physical object while displaying a first set of graphical elements. In some implementations, the computing system obtains (e.g., receives, retrieves, or determines) translational and rotational values ​​of the physical object, and detecting the first movement corresponds to detecting a change in one of the translational or rotational values ​​of the physical object. For example, the computing system tracks the physical object via computer vision, magnetic sensors, positional information, etc. As an example, the physical object corresponds to a proxy object such as a pencil or pen that has no communication channel to the computing system. As another example, the physical object corresponds to an electronic device such as a stylus or finger-worn device that has a wired or wireless communication channel to a computing system including an IMU, accelerometer, gyroscope, magnetometer, etc. for 6-degree-of-freedom (6DOF) tracking.

[0203] In some implementations, the computing system maintains one or more N-tuple tracking vectors / tensors of the physical object (e.g., object tracking vectors 511 in Figures 5A and 5B) based on the tracking data 506. In some implementations, one or more N-tuple tracking vectors / tensors of the physical object (e.g., object tracking vectors 511 in Figures 5A and 5B) include translational values ​​of the physical object relative to the whole world or the current operating environment (e.g., x, y, and z), rotational values ​​of the physical object (e.g., roll, pitch, and yaw), grip attitude indications of the physical object (e.g., pointing, writing, erasing, painting, dictation, etc.), currently used tip / end indications (e.g., the physical object may have an asymmetric design with specific first and second tip ends, or a symmetric design with non-specific first and second tip ends), a first input (pressure) value relating to how strongly the physical object is pressed against the physical surface, a second input (pressure) value relating to how strongly the physical object is gripped by the user, touch input information, etc.

[0204] In some implementations, the tracking data 506 corresponds to one or more images of the physical environment, including the physical object, to enable 6DOF tracking via computer vision techniques. In some implementations, the tracking data 506 corresponds to data collected by various integrated sensors of the physical object, such as IMUs, accelerometers, gyroscopes, and magnetometers. For example, the tracking data 506 corresponds to raw or processed sensor data, such as translational values ​​associated with the physical object (relative to the physical environment or the world as a whole), rotational values ​​associated with the physical object (relative to gravity), velocity values ​​associated with the physical object, angular velocity values ​​associated with the physical object, acceleration values ​​associated with the physical object, angular acceleration values ​​associated with the physical object, a first input (pressure) value associated with how strongly the physical object is in contact with the physical surface, a second input (pressure) value associated with how strongly the physical object is being gripped or in contact with the physical surface by the user, and so on.

[0205] In some implementations, the computing system also acquires finger operation data detected by the physical object via a communication interface. For example, the finger operation data includes touch input or gestures directed towards the touch-sensitive area of ​​the physical object. For example, the finger operation data includes contact intensity data with respect to the body of the physical object. In some implementations, the physical object includes a touch-sensitive surface / area configured to detect touch input directed towards the physical object, such as a longitudinally extending touch-sensitive surface. In some implementations, the translational and rotational values ​​for the physical object include determining the translational and rotational values ​​for the physical object based on at least one of the following: IMU data from the physical object, one or more images of the physical environment 105 containing the physical object, magnetic tracking data, etc.

[0206] In some implementations, the computing system is further communicatively coupled to a physical object, and obtaining tracking data 506 associated with the physical object includes obtaining tracking data 506 from the physical object, where the tracking data corresponds to output data from one or more integrated sensors of the physical object. Figures 6C–6P show a user 149 holding a control device 130 used to communicate with an electronic device 120 and interact with an XR environment 128. For example, one or more integrated sensors include at least one of an IMU, accelerometer, gyroscope, GPS, magnetometer, one or more contact intensity sensors, touch-sensing surface, and / or similar. In some implementations, the tracking data 506 further indicates whether the tip of the physical object is in contact with a physical surface and the associated pressure value.

[0207] In some implementations, method 900 includes acquiring one or more images of the physical environment, recognizing physical objects using one or more images of the physical environment, and assigning physical objects (e.g., proxy objects) to function as focus selectors when interacting with the XR environment 128. Figures 8A to 8D show user 149 holding a proxy object 804 (e.g., a ruler, a stick, etc.) that is unable to communicate with the electronic device 120 and is used to interact with the XR environment 128. In some implementations, the computing system designates a physical object as a focus selector when it is held by the user. In some implementations, the computing system designates a physical object as a focus selector when it is held by the user and the physical object satisfies predefined constraints (e.g., maximum or minimum size, specific shape, digital rights management (DRM) non-qualifier, etc.). Thus, in some implementations, user 149 can use a home object to interact with the XR environment. In some implementations, the attitude and grip indicators may be fixed to the proxy object (or its representation) as the proxy object moves and / or the FOV moves.

[0208] As represented by block 906, in response to detecting a first movement of a physical object and determining that the first movement of the physical object has caused the physical object (e.g., a given part of the physical object, such as the tip of the physical object) to exceed a distance threshold for a first graphical element among a first set of graphical elements, method 900 includes selecting a first output modality associated with the first graphical element as the current output modality of the physical object. In some implementations, the distance threshold is either non-deterministic (i.e., a given X mm radius) or deterministic based on one or more factors such as user preference, tool usage history, depth of the graphical element relative to the scene, occlusion, current content, and current context.

[0209] In some implementations, a computing system or its components (e.g., the output modality selection unit 526 in Figure 5A) selects the associated first output modality as the current output modality 527 for a physical object upon determination that a first movement of the physical object has caused the physical object to exceed a distance threshold for a first graphical element among a first set of graphical elements. As an example, Figures 6D and 6E show a sequence in which an electronic device 120 selects a first output modality (e.g., airbrush marking) for a control device 130 upon determination that a movement of the control device 130 has caused the control device 130 (or its representation) to exceed an activation region 634 (e.g., a distance threshold) for a graphical element 632D.

[0210] In some implementations, upon determining that a first movement of a physical object causes the physical object to exceed a distance threshold relative to a first graphical element among a first set of graphical elements, method 900 includes maintaining the display of the first graphical element adjacent to the physical object and stopping the display of the remaining first set of graphical elements that do not contain the first graphical element. As an example, referring to Figures 6D and 6E, the electronic device 120 maintains the display of a first graphical element (e.g., graphical element 632D) among a first set of graphical elements (e.g., graphical element 632) superimposed on the leading edge of the representation 131 of the control device 130 in the XR environment 128, and removes the display of the remaining first set of graphical elements (e.g., graphical elements 632A, 632B, and 632C) from the XR environment 128.

[0211] In some implementations, method 900 includes detecting a second movement of a physical object after selecting a first output modality associated with a first graphical element as the current output modality of the physical object, and, in response to detecting the second movement of the physical object, moving the first graphical element based on the second movement of the physical object in order to maintain the display of the first graphical element adjacent to the physical object. In some implementations, the first graphical element is fixed to the outward-facing end / tip of the physical object. In some implementations, the first graphical element is presented from or offset to the side of the outward-facing end / tip of the physical object. In some implementations, the first graphical element "snaps" to the representation of the physical object. As an example, referring to Figures 6E and 6F, the electronic device 120 maintains the display of a graphical element 632D superimposed on the tip of the representation 131 of the control device 130 in the XR environment 128 after detecting the movement of the control device to perform a marking input 654.

[0212] In some implementations, method 900 includes obtaining (e.g., receiving, retrieving, or detecting) a touch input directed to a physical object after stopping the display of the remaining first plurality of graphical elements, and redisplaying the first plurality of graphical elements in the XR environment via a display device in response to obtaining the touch input instruction. In some implementations, the physical object corresponds to a stylus having a touch-sensitive area capable of detecting touch input. For example, the stylus detects an upward or downward swipe gesture on its touch-sensitive surface, and the computing system obtains (e.g., receives or retrieves) a touch input instruction from the stylus. As an example, referring to Figures 6N and 6O, the electronic device 120 displays a second plurality of graphical elements 6152 associated with a second plurality of output modalities in the XR environment 128 in response to detecting a touch input 6142 directed to a control device 130.

[0213] As represented by block 908, in response to detecting a first movement of a physical object and in accordance with the determination that the first movement of the physical object has caused the physical object to exceed a distance threshold for a second graphical element of a first set of graphical elements, method 900 includes selecting a second output modality associated with the second graphical element as the current output modality of the physical object.

[0214] In some implementations, a computing system or its components (e.g., the output modality selection unit 526 in Figure 5A) selects the associated second output modality as the current output modality 527 for a physical object upon determination that a first movement of the physical object has caused the physical object to exceed a distance threshold for a second graphical element among a first set of graphical elements. As an example, Figures 6G to 6I show a sequence in which an electronic device 120 selects a second output modality (e.g., pen marking) for a control device 130 upon determination that a movement of the control device 130 has caused the control device 130 (or its representation) to exceed an activation region 634 (e.g., a distance threshold) for a graphical element 632B.

[0215] In some implementations, upon determination that a first movement of a physical object causes the physical object to exceed a distance threshold relative to a second graphical element of a first set of graphical elements, method 900 includes maintaining the display of the second graphical element adjacent to the physical object and stopping the display of the remaining first set of graphical elements that do not include the second graphical element. As an example, referring to Figures 6H and 6I, the electronic device 120 maintains the display of the second graphical element (e.g., graphical element 632B) of a first set of graphical elements (e.g., graphical element 632) superimposed on the leading edge of the representation 131 of the control device 130 within the XR environment 128, and removes the display of the remaining first set of graphical elements (e.g., graphical elements 632A, 632C, and 632D) from the XR environment 128.

[0216] In some implementations, as represented by block 910, the first and second output modalities cause different visual changes within the XR environment. For example, the first output modality is associated with selecting / manipulating objects / content within the XR environment, while the second output modality is associated with sketching, drawing, writing, etc., within the XR environment. As an example, referring to Figure 6F, the electronic device 120 displays an airbrush-like mark 662 in the XR environment 128 in response to detecting a marking input 654 in Figure 6E, while the current output modality corresponds to a graphical element 632D. As another example, referring to Figure 6J, the electronic device 120 displays a pen-like mark 6112 in the XR environment 128 in response to detecting a marking input 6104 in Figure 6J, while the current output modality corresponds to a graphical element 632B.

[0217] In some implementations, upon determining that a first movement of a physical object does not cause the physical object to exceed a distance threshold relative to a first or second graphical element, method 900 includes maintaining the initial output modality as the current output modality of the physical object and maintaining the display of the first set of graphical elements. As an example, referring to Figure 6C, the electronic device 120 maintains the initial output modality as the current output modality of the control device 130 while the representation 131 of the control device 130 is outside the activation region 634. As another example, referring to Figure 6G, the electronic device 120 maintains the initial output modality as the current output modality of the control device 130 while the representation 131 of the control device 130 is outside the activation region 634.

[0218] In some implementations, as represented by block 912, method 900 includes detecting a subsequent marking input by a physical object after selecting a first output modality associated with a first graphical element as the current output modality of the physical object, and in response to detecting the subsequent marking input, displaying one or more marks in the XR environment via a display device based on the subsequent marking input (e.g., shape, displacement, etc. of the subsequent marking input) and the first output modality. According to some implementations, the computing system detects the subsequent marking input using the physical object by tracking the physical object in 3D using IMU data, computer vision, magnetic tracking, etc. In some implementations, one or more marks correspond to XR content displayed in the XR environment 128, such as sketches, handwritten text, or doodles. As an example, referring to Figure 6F, the electronic device 120 displays an airbrush-like mark 662 in the XR environment 128 in response to detecting a marking input 654 in Figure 6E, while the current output modality corresponds to graphical element 632D. For example, the shape, depth, length, and angle of the airbrush-like mark 662 correspond to the spatial parameters of the marking input 654 (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc., associated with the marking input).

[0219] In some implementations, as represented by block 914, in response to detecting a subsequent marking input, method 900 includes, according to a determination that an input associated with how hard a physical object is pressed against a physical surface corresponds to a first input value, displaying one or more marks having a first appearance in an XR environment via a display device based on a subsequent marking input (e.g., shape, displacement, etc. of the subsequent marking input) and a first output modality, wherein the first appearance is associated with parameters of one or more marks corresponding to the first input value; and, according to a determination that an input associated with how hard a physical object is pressed against a physical surface corresponds to a second input value, displaying one or more marks having a second appearance in an XR environment via a display device based on a subsequent marking input (e.g., shape, displacement, etc. of the subsequent marking input) and a first output modality, wherein the second appearance is associated with parameters of one or more marks corresponding to a second input value.

[0220] In some implementations, one or more marks correspond to XR content displayed within the XR environment 128, such as sketches, handwritten text, or doodles. In some implementations, a computing system obtains (e.g., receives, retrieves, or determines) first and second input (pressure) values ​​based on locally or remotely collected data. As an example, a physical object corresponds to an electronic device having a pressure sensor on one or both of its ends / tips to detect input (pressure) values ​​when pressed against a physical surface. In some implementations, as represented by block 916, the parameter corresponds to one of the radius, width, thickness, intensity, translucency, opacity, color, or texture of one or more marks in the XR environment.

[0221] For example, referring to Figure 7H, the electronic device 120, in response to detecting the marking input 715 in Figure 7G, displays mark 782A on the XR substrate 718 in the XR environment 128 and mark 782B on the input area 774. For example, the shape, depth, length, angle, etc., of marks 782A and 782B correspond to the spatial parameters of the marking input 772 in Figure 7G (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc., associated with the marking input). Furthermore, in Figure 7H, marks 782A and 782B are associated with a first thickness value corresponding to the current measurement value 779 of the input (pressure) value in Figure 7G.

[0222] In another example, referring to Figure 7J, the electronic device 120, in response to detecting the marking input 792 in Figure 7I, displays mark 7102A on the XR substrate 718 in the XR environment 128 and mark 7102B on the input area 774. For example, the shape, depth, length, angle, etc., of marks 7102A and 7102B correspond to the spatial parameters of the marking input 792 in Figure 7I (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc., associated with the marking input). Furthermore, in Figure 7J, marks 7102A and 7102B are associated with a second thickness value corresponding to the current measurement value 799 of the input (pressure) value in Figure 7I. For example, the second thickness value associated with marks 7102A and 7102B is greater than the first thickness value associated with marks 782A and 782B.

[0223] In some implementations, as represented by block 918, in response to detecting a subsequent marking input, method 900 includes, according to a determination that an input associated with how tightly a physical object is being gripped by a user corresponds to a first input value, displaying one or more marks having a first appearance in an XR environment via a display device based on a subsequent marking input (e.g., shape, displacement, etc. of the subsequent marking input) and a first output modality, wherein the first appearance is associated with parameters of one or more marks corresponding to the first input value; and, according to a determination that an input associated with how tightly a physical object is being gripped by a user corresponds to a second input value, displaying one or more marks having a second appearance in an XR environment via a display device based on a subsequent marking input (e.g., shape, displacement, etc. of the subsequent marking input) and a first output modality, wherein the second appearance is associated with parameters of one or more marks corresponding to a second input value.

[0224] In some implementations, one or more marks correspond to XR content displayed within the XR environment 128, such as sketches, handwritten text, or doodles. In some implementations, a computing system obtains (e.g., receives, retrieves, or determines) first and second input (pressure) values ​​based on locally or remotely collected data. As an example, a physical object corresponds to an electronic device with a built-in pressure sensor for detecting input (pressure) values ​​when grasped by a user. In some implementations, as represented by block 920, the parameters correspond to one of the radius, width, thickness, intensity, translucency, opacity, color, or texture of the mark in the XR environment.

[0225] For example, referring to Figure 6J, the electronic device 120 displays a pen-shaped mark 6112 in the XR environment 128 in response to detecting a marking input 6104 in Figure 6J, while the current output modality corresponds to graphical element 632B. For example, the shape, depth, length, angle, etc., of the pen-shaped mark 6112 correspond to the spatial parameters of the marking input 6104. Furthermore, in Figure 6K, the pen-shaped mark 6112 is associated with a first thickness value that corresponds to the current measured value 6103 of the input (pressure) value in Figure 6J.

[0226] In another example, referring to Figure 6M, the electronic device 120 displays a pen-shaped mark 6132 in the XR environment 128 in response to detecting a marking input 6122 in Figure 6L. For example, the shape, depth, length, angle, etc., of the pen-shaped mark 6132 correspond to the spatial parameters of the marking input 6122. Furthermore, in Figure 6M, the pen-shaped mark 6132 is associated with a second thickness value corresponding to the current measured value 6123 of the input (pressure) value in Figure 6L. The second thickness value associated with the pen-shaped mark 6132 in Figure 6M is greater than the first thickness value associated with the pen-shaped mark 6112 in Figure 6K.

[0227] Figures 10A and 10B show flowchart representations of method 1000, which modifies mark parameters based on a first input (pressure) value during direct marking on a physical surface, or based on a second input (pressure) value during indirect marking, relating to several implementations. In various implementations, method 1000 is executed in a computing system including non-temporary memory and one or more processors, the computing system being communicatively coupled to a display device and one or more (optional) input devices (e.g., electronic device 120 shown in Figures 1 and 3, controller 110 in Figures 1 and 2, or an appropriate combination thereof). In some implementations, method 1000 is executed by processing logic, which includes hardware, firmware, software, or a combination thereof. In some implementations, method 1000 is executed by a processor that executes code stored in a non-temporary computer-readable medium (e.g., memory). In some implementations, the computing system corresponds to one of the following: a tablet, laptop, mobile phone, near-eye system, wearable computing device, etc.

[0228] Typically, users can adjust marking parameters, such as line thickness, by moving sliders in a toolbar or control panel. This can interrupt the user's current workflow and lead them to search through various menus for appropriate control. In contrast, the method described herein adjusts the marking parameters based on a first input (pressure) value between the physical object (e.g., a proxy object or stylus) and the physical surface, or on a second input (pressure) value associated with the user's grip on the physical object, when the marking input is directed towards a physical surface. Thus, the user can adjust the marking parameters more quickly and efficiently.

[0229] As represented by block 1002, method 1000 includes displaying a user interface via a display device. In some implementations, the user interface includes a two-dimensional marking area where marks are displayed (e.g., a planar canvas). In some implementations, the user interface includes a three-dimensional marking area where marks are displayed (e.g., marks are associated with 3D painting or drawing) (1006). As an example, referring to Figure 7A, an electronic device 120 presents an XR environment 128 including an XR substrate 718 (e.g., a 2D or 3D canvas).

[0230] In some implementations, the display device corresponds to a transparent lens assembly, and the presentation of the user interface is projected onto the transparent lens assembly. In some implementations, the display device corresponds to a near-eye system, and the presentation of the user interface includes compositing the user interface presentation with one or more images of the physical environment captured by an outward-facing image sensor.

[0231] In some implementations, method 1000 includes displaying a user interface element (e.g., a toolbar, menu, etc.) having multiple different selectable tools associated with markings within the user interface via a display device. In some implementations, the user interface element is fixed to a point in space. For example, the user interface element can be moved to a new fixed point in space. In some implementations, as the user rotates their head, the user interface element remains anchored to a point in space and may move out of the field of view until the user completes a reverse head rotation (e.g., world / object lock). In some implementations, the user interface element is fixed to a point in the user's field of view of the computing system (e.g., head / body lock). For example, the user interface element can be moved to a new fixed point in the FOV. In some implementations, as the user rotates their head, the user interface element will remain fixed to a point in the FOV, such that a toolbar remains within the FOV.

[0232] As an example, referring to Figure 7A, the electronic device 120 presents an XR environment 128 that includes a menu 712. As shown in Figure 7A, the menu 712 includes several selectable options 714 associated with changing the appearance of marks created within the XR environment (e.g., different colors, textures, etc.). For example, option 714A is currently selected from among the several selectable options 714. In this example, option 714A corresponds to a first appearance of a mark (e.g., a black mark) created within the XR environment 128. As shown in Figure 7A, the menu 712 also includes a slider 716 for adjusting the thickness of marks created within the XR environment 128.

[0233] As represented by block 1008, while displaying a user interface, method 1000 includes detecting marking input by a physical object. For example, the marking input corresponds to the creation of 2D or 3D XR content such as sketches, handwritten text, or doodles. In some implementations, the computing system acquires (e.g., receives, retrieves, or determines) the translational and rotational values ​​of the physical object, and detecting a first movement corresponds to detecting a change in one of the translational or rotational values ​​of the physical object. For example, the computing system tracks the physical object via computer vision, magnetic sensors, etc. As an example, the physical object corresponds to a proxy object such as a pencil or pen that has no communication channel to the computing system. As another example, the physical object corresponds to an electronic device such as a stylus or finger-worn device that has a wired or wireless communication channel to the computing system, including an IMU, accelerometer, gyroscope, etc. for 6DOF tracking.

[0234] In some implementations, the computing system includes one or more N-tuple tracking vectors / tensors of a physical object (e.g., object tracking vectors 511 in Figures 5A and 5B), which include translational values ​​(e.g., x, y, and z) relative to the whole world or the current operating environment, rotational values ​​(e.g., roll, pitch, and yaw), grip attitude indications (e.g., pointing, writing, erasing, painting, dictation, etc.), currently used tip / end indications (e.g., the physical object may have an asymmetric design with specific first and second tip ends, or a symmetric design with non-specific first and second tip ends), a first input (pressure) value relating to how strongly the physical object is pressed against the physical surface, a second input (pressure) value relating to how strongly the physical object is gripped by the user, and so on.

[0235] In some implementations, a physical object includes a touch-sensitive surface / region configured to detect touch input directed at the physical object, such as a vertically extending touch-sensitive surface. In some implementations, obtaining the translational and rotational values ​​of a physical object involves determining the translational and rotational values ​​of the physical object based on at least one of the following: inertial measurement unit (IMU) data from the physical object, one or more images of the physical environment containing the physical object, magnetic tracking data, etc.

[0236] As represented by block 1010, in response to detecting a marking input and in accordance with the determination that the marking input is directed toward a physical surface (e.g., a tabletop, another plane, etc.), method 1000 includes displaying a mark in a user interface via a display device based on the marking input (e.g., shape, size, orientation, etc. of the marking input), and the parameters of the mark displayed based on the marking input are determined based on how hard the physical object is pressed against the physical surface. In some implementations, a computing system or its components (e.g., parameter adjustment unit 528 in Figure 5A) adjusts output parameters (e.g., mark thickness, brightness, color, texture, etc.) associated with the detected marking input directed toward the XR environment 128 based on how hard the physical object is pressed against the physical surface (e.g., a first input (pressure) value), in accordance with the determination that the marking input is directed toward a physical surface (e.g., a tabletop, another plane, etc.). In some implementations, the parameter corresponds to one of the following: radius, width, thickness, intensity, translucency, opacity, color, or texture of the mark in the user interface (1014).

[0237] In some implementations, the mark parameters are determined based on how strongly a predefined portion of a physical object, such as the tip of the physical object in contact with a physical surface in a 3D environment, is pressed against the physical surface. For example, the physical object corresponds to an electronic device having pressure sensors on one or both of its ends / tips to detect a first input (pressure) value when pressed against the physical surface. In some implementations, a computing system maps the marking input on the physical surface to a 3D marking area or 2D canvas in the XR environment. For example, the marking area and the physical surface correspond to vertical planes offset by Y cm.

[0238] As an example, Figures 7G and 7H show a sequence in which the detection of a marking input 772 causes marks 782A and 782B to appear in the XR environment 128 according to the measured value 779 of the input (pressure) value. For example, the shape, depth, length, angle, etc. of marks 782A and 782B correspond to the spatial parameters of the marking input 772 in Figure 7G (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc. associated with the marking input). Furthermore, in Figure 7H, marks 782A and 782B are associated with a first thickness value corresponding to the measured value 779 of the input (pressure) value in Figure 7G.

[0239] As another example, Figures 7I and 7J show a sequence in which the detection of a marking input 792 causes marks 7102A and 7102B to appear in the XR environment 128 according to the measured value 799 of the input (pressure) value. For example, the shape, depth, length, angle, etc. of marks 7102A and 7102B correspond to the spatial parameters of the marking input 792 in Figure 7I (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc. associated with the marking input). Furthermore, in Figure 7J, marks 7102A and 7102B are associated with a second thickness value corresponding to the current measured value 799 of the input (pressure) value in Figure 7I. For example, the second thickness value associated with marks 7102A and 7102B is greater than the first thickness value associated with marks 782A and 782B.

[0240] As represented by block 1012, in response to the detection of a marking input and in accordance with the determination that the marking input is not directed toward a physical surface, method 1000 includes displaying a mark in a user interface via a display device based on the marking input (e.g., shape, size, orientation, etc. of the marking input), and the parameters of the mark displayed based on the marking input are determined based on how tightly the physical object is being gripped by the user. In some implementations, a computing system or its components (e.g., parameter adjustment unit 528 in Figure 5A) adjusts output parameters (e.g., thickness, brightness, color, texture, etc. of the mark) associated with the detected marking input directed toward the XR environment 128 based on how tightly the physical object is being gripped by the user 149 (e.g., second input (pressure) value), in accordance with the determination that the marking input is not directed toward a physical surface. In some implementations, the parameters correspond to one of the radius, width, thickness, intensity, translucency, opacity, color, or texture of the mark in the user interface (1014).

[0241] In some implementations, the computing system detects a marking input while the physical object or a predefined portion of the physical object, such as its tip, is not in contact with any physical surface in the three-dimensional environment. For example, the physical object corresponds to an electronic device having a built-in pressure sensor for detecting a second input (pressure) value when grasped by a user.

[0242] As an example, Figures 7A and 7B show a sequence in which the detection of a marking input 715 causes a mark 722 to appear in the XR environment 128 according to the current measured value 719 of the input (pressure) value. For example, the shape, depth, length, angle, etc. of the mark 722 correspond to the spatial parameters of the marking input 715 (e.g., position value, rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc. associated with the marking input). Furthermore, in Figure 7B, the mark 722 is associated with a first thickness value corresponding to the current measured value 719 of the input (pressure) value in Figure 7A.

[0243] As another example, Figures 7C and 7D show a sequence in which the detection of a marking input 732 causes a mark 742 to appear in the XR environment 128 according to a measured value 739 of the input (pressure) value. For example, the shape, depth, length, angle, etc. of the mark 742 correspond to the spatial parameters of the marking input 732. Furthermore, in Figure 7D, the mark 742 is associated with a second thickness value corresponding to the current measured value 739 of the input (pressure) value in Figure 7C. For example, the second thickness value associated with the mark 742 is greater than the first thickness value associated with the mark 722.

[0244] In some implementations, as represented by block 1016, method 1000 includes, after displaying a mark in the user interface, detecting a subsequent input by a physical object associated with moving (e.g., translating and / or rotating) the mark in the user interface, and, in response to the detection of the subsequent input, moving the mark in the user interface based on the subsequent input. As an example, Figures 7E and 7F show a sequence in which an electronic device translates the mark 742 in the XR environment 128 in response to detecting an operation input 752. For example, the angle, direction, displacement, etc. of the translational movement of the mark 742 correspond to the spatial parameters of operation input 752 in Figure 7E. In some implementations, operation input 752 may also cause a rotational movement of the mark 742 based on the rotational parameters of operation input 752.

[0245] In some implementations, as represented by block 1018, detecting a subsequent input corresponds to obtaining an instruction that an affordance on a physical object has been activated and detecting at least one of rotational or translational movement of the physical object. For example, the activation of an affordance corresponds to the detection of a touch input directed toward the touch-sensing surface of the control device 130. As an example, Figures 7E and 7F show a sequence in which the electronic device 120 translates mark 742 in the XR environment 128 in response to the detection of an operation input 752, while also detecting a touch input 754 directed toward the touch-sensing surface 175 of the control device 130 in Figure 7E.

[0246] In some implementations, as represented by block 1020, detecting a subsequent input corresponds to obtaining an instruction that an input value associated with how tightly a physical object is being grasped by the user exceeds a threshold input value, and detecting at least one of the rotational or translational movement of the physical object. For example, an input (pressure) value corresponds to a selection of the subsequent input. In some implementations, the pressure threshold is either non-deterministic (i.e., a given pressure value) or deterministic based on one or more factors such as user preference, usage history, current content, or current context.

[0247] In some implementations, as represented by block 1022, in response to detecting subsequent input, method 1000 includes changing the appearance of at least some content within the user interface while moving a mark within the user interface. For example, the computing system increases the opacity, translucency, blur radius, etc., of at least some content, such as a 2D canvas or a 3D marking area.

[0248] In some implementations, as represented by block 1024, in response to the detection of a marking input and in accordance with the determination that the marking input is directed toward a physical surface, method 1000 includes displaying a simulated shadow in the XR environment corresponding to the distance between the physical surface and the physical object via a display device. In some implementations, the size, angle, etc., of the shadow change as the physical object approaches or moves away from the physical surface. For example, as the physical object moves away from the physical surface, the size of the simulated shadow increases and the associated opacity value decreases. Continuing this example, as the physical object approaches the physical surface, the size of the simulated shadow decreases and the associated opacity value increases. In some implementations, a shadow may also be shown when the marking input is not directed toward a physical surface.

[0249] As an example, Figures 7K and 7L illustrate a sequence in which the electronic device 120 displays a first XR content 7122 within the XR environment 128 according to the current measurement of the input (pressure) value 7119 in response to the detection of a first content placement input associated with a touch input 7111. As shown in Figure 7L, the electronic device 120 also displays a shadow 7124 associated with the first XR content 7122 on the XR substrate 7118. For example, the position and rotation values ​​of the first XR content 7122 and shadow 7124 correspond to the parameters (e.g., position value, rotation value, etc.) of the representation 131 of the control device 130 when the touch input 7111 is detected in Figure 7K. For example, the first XR content 7122 is associated with a first size value corresponding to the current measurement of the input (pressure) value 7119 in Figure 7K.

[0250] As another example, Figures 7M and 7N show a sequence in which the electronic device 120 displays a second XR content 7142 within the XR environment 128 according to a measured value 7139 of the input (pressure) value, in response to the detection of a second content placement input associated with a touch input 7131. As shown in Figure 7N, the electronic device 120 does not display a shadow associated with the second XR content 7142 on the XR substrate 7118. For example, the position and rotation values ​​of the second XR content 7142 correspond to the parameters (e.g., position value, rotation value, etc.) of the representation 131 of the control device 130 when the touch input 7131 is detected in Figure 7M. For example, the second XR content 7142 is associated with a second size value corresponding to the current measured value 7139 of the input (pressure) value in Figure 7K. For example, the second size value associated with the second XR content 7142 is greater than the first size value associated with the first XR content 7122.

[0251] FIG. 11 is a flowchart representation of a method 1100 for changing a selection modality based on whether a user is currently holding a physical object, according to some implementations. In various implementations, the method 1100 is executed in a computing system that includes a non - transient memory and one or more processors, and the computing system is communicatively coupled to a display device and (optionally) one or more input devices (e.g., the electronic device 120 shown in FIGS. 1 and 3, the controller 110 of FIGS. 1 and 2, or a suitable combination thereof). In some implementations, the method 1100 is executed by processing logic that includes hardware, firmware, software, or a combination thereof. In some implementations, the method 1100 is executed by a processor that executes code stored in a non - transient computer - readable medium (e.g., memory). In some implementations, the computing system corresponds to one of a tablet, a laptop, a mobile phone, a near - eye system, a wearable computing device, etc.

[0252] Typically, when a user navigates content within a user interface, the user is limited to one or more input modalities such as touch input, voice commands, etc. Further, the one or more input modalities may be applicable regardless of the current situation, such as while operating a vehicle, while moving, while having full hands, etc., which may pose concerns regarding usability and safety. In contrast, the methods described herein enable a user to select content based on the line of sight when not holding a physical object (e.g., a proxy object or a stylus) with a pointing grip, and also enable a user to select content based on the orientation of the physical object when holding the physical object with a pointing grip. Thus, the input modality for selecting content dynamically changes based on the current context.

[0253] As represented by block 1102, method 1100 includes displaying content via a display device. As an example, the content corresponds to volumetric content or 3D content within an XR environment. As another example, the content corresponds to flat or 2D content within a user interface (UI). For example, referring to FIGS. 8A-8D, electronic device 120 displays VA606 and XR content 802 within XR environment 128.

[0254] In some implementations, the display device corresponds to a transparent lens assembly and the presentation of the content is projected onto the transparent lens assembly. In some implementations, the display device corresponds to a near-eye system and presenting the content includes synthesizing the presentation of the content with one or more images of the physical environment captured by an outward-facing image sensor.

[0255] As represented by block 1104, while content is being displayed and a physical object is being held by a user, method 1100 includes detecting a selection input. As an example, the physical object corresponds to a proxy object detected in a physical environment without a communication channel to a computing system, such as a pencil or pen. Referring to Figures 8A–8D, the electronic device 120 displays a representation 153 of the user 149's right hand 152 holding a representation 805 of a proxy object 804 (e.g., a stick, ruler, or another physical object) in the XR environment 128. For example, the user 149's right hand 152 is currently holding the proxy object 804 in a pointing grip position. As another example, the physical object corresponds to an electronic device that has a wired or wireless communication channel to a computing system, such as a stylus, finger-worn device, or handheld device. Referring to Figures 8E to 8G, the electronic device 120 displays a representation 153 of the user 149's right hand 152, which is holding a representation 131 of the control device 130 in the XR environment 128. For example, the user 149's right hand 152 is currently holding the control device 130 in a writing grip position that is not directed towards any XR content in the XR environment 128.

[0256] As represented by block 1106, in response to detecting a selection input, method 1100 includes performing an action corresponding to the selection input. In some implementations, a computing system or its components (e.g., the content selection engine 522 in Figure 5A) determines the selected content portion 523 based on a characterization vector 531 (or a portion thereof). For example, the content selection engine 522 determines the selected content portion 523 based on current context information, the user's gaze direction, body posture information associated with the user 149, head posture information associated with the user 149, hand / limb tracking information associated with the user 149, position information associated with a physical object, rotation information associated with a physical object, etc.

[0257] As an example, the content selection engine 522, upon determining that the grip pose associated with the way the physical object is held by the user corresponds to a first grip (e.g., first grip = pointing / wand-shaped grip), performs a selection operation on a first part of the content based on the direction pointed to by a predetermined part of the physical object (e.g., an outward-facing end) (and the rays projected from it). As another example, the content selection engine 522, upon determining that the grip pose associated with the way the physical object is held by the user does not correspond to a first grip, performs a selection operation on a second part of the content based on the user's line of sight.

[0258] In some implementations, as represented by block 1108, method 1100 includes changing the appearance of a first or second portion of the content. For example, changing the appearance of a first or second portion of the content corresponds to changing the color, texture, brightness, etc., of the first or second portion of the content to indicate that it has been selected. For example, changing the appearance of a first or second portion of the content corresponds to displaying a bounding box, highlight, spotlight, etc., associated with the first or second portion of the content to indicate that it has been selected. For example, referring to Figure 8C, in response to detecting that proxy object 804 has been directed towards XR content 802 for at least a predetermined or deterministic period of time in Figures 8A and 8B, the electronic device 120 changes the appearance of XR content 802 to a cross-hatched appearance 802A to visually indicate its selection. For example, referring to Figure 6D, the electronic device 120, in response to detecting that the spatial location of the representation 131 of the control device 130 has moved beyond the activation region 634 relative to the graphical element 632D due to the movement of the control device 130, changes the appearance of the graphical element 632D by displaying a boundary or frame 642 around the graphical element 632D to indicate its selection.

[0259] As represented by block 1110, method 1100 includes performing a selection operation on a first portion of content, based on the determination that a grip pose associated with the manner in which a physical object is held by the user corresponds to a first grip (e.g., first grip = pointing / wand-shaped grip), the first portion of content is selected based on the direction (e.g., rays projected from) a given portion of the physical object (e.g., an outward-facing end) is pointing (e.g., rays projected from there) (e.g., regardless of the user's line of sight). In some implementations, a computing system or its components (e.g., content selection engine 522 in Figure 5A) performs a selection operation on a first portion of content based on the direction (e.g., rays projected from) a given portion of the physical object (e.g., an outward-facing end) is pointing, based on the determination that a grip pose associated with the manner in which a physical object is held by the user corresponds to a first grip (e.g., first grip = pointing / wand-shaped grip). As an example, Figures 8B and 8C show a sequence in which the electronic device 120 selects the XR content 802 in response to detecting that the proxy object 804 (or its representation 805) has been directed toward the XR content 802 for at least a predetermined or deterministic period of time, and according to the determination that the grip pose associated with the way the physical object 804 is held by the user 149 corresponds to a first grip (e.g., a pointing grip).

[0260] In some implementations, the computing system obtains (e.g., receives, retrieves, or determines) the translational and rotational values ​​of a physical object and obtains (e.g., receives, retrieves, or determines) the grip orientation associated with the current way the physical object is held by the user. For example, the computing system tracks the physical object via computer vision, magnetic sensors, etc. As an example, the physical object corresponds to a proxy object such as a pencil or pen that has no communication channel to the computing system. As another example, the physical object corresponds to an electronic device such as a stylus or finger-worn device that has a wired or wireless communication channel to the computing system, including an IMU, accelerometer, gyroscope, etc. for 6DOF tracking. In some implementations, the computing system includes one or more N-tuple tracking vectors / tensors of a physical object, which include translational values ​​(e.g., x, y, and z) relative to the world or the current operating environment, rotational values ​​(e.g., roll, pitch, and yaw), grip attitude indications (e.g., attitude, such as pointing, writing, erasing, painting, dictation, etc.), currently used tip / end indications (e.g., the physical object may have an asymmetric design with specific first and second tip ends, or a symmetric design with non-specific first and second tip ends), a first input (pressure) value relating to how hard the physical object is pressed against the physical surface, a second input (pressure) value relating to how hard the physical object is gripped by the user, and so on. In some implementations, the physical object includes a touch-sensitive surface / region configured to detect touch input directed at the physical object, such as a touch-sensitive surface extending in the longitudinal direction. In some implementations, obtaining the translational and rotational values ​​of a physical object involves determining the translational and rotational values ​​of the physical object based on at least one of the following: IMU data from the physical object, one or more images of the physical environment containing the physical object, magnetic tracking data, etc.

[0261] As represented by block 1112, following the determination that the grip pose associated with the manner in which a physical object is held by the user does not correspond to a first grip, method 1100 includes performing a selection operation on a second part of content different from the first part of content, the second part of content being selected based on the user's gaze direction (e.g., regardless of the direction of light rays projected from a given part of the physical object). In some implementations, a computing system or its components (e.g., the gaze tracking engine 512 in Figure 5A) determine and update a gaze tracking vector 513, which includes x and y coordinates, focal length, or focus, associated with the gaze direction to the whole world or the current operating environment. In some implementations, the computing system determines the gaze tracking vector 513 based on one or more images of the user's eyes from an inward-facing image sensor. In some implementations, the computing system determines a region of interest (ROI) (e.g., an N × M mm ROI) within the XR environment 128 based on the gaze direction.

[0262] In some implementations, the computing system or its components (e.g., the content selection engine 522 in Figure 5A) perform a selection operation on a second portion of the content based on the user's gaze direction, following a determination that the grip orientation associated with the way the physical object is held by the user does not correspond to a first grip. As an example, Figures 8E and 8F show a sequence in which the electronic device 120 selects the XR content 802 in response to detecting that the user's gaze direction 806 has been directed towards the XR content 802 for at least a predetermined or deterministic period, and following a determination that the grip orientation associated with the way the physical object 130 is held by the user 149 does not correspond to a first grip.

[0263] In some implementations, as represented by block 1114, following a determination that a grip pose corresponds to a first grip, method 1100 includes displaying a first graphical element via a display device that indicates the direction in which a given portion of a physical object is pointing toward the content. For example, the first graphical element is displayed at a coincidence point where a ray projected from a given portion of a physical object coincides with content, a 2D canvas, a 3D marking area, a backplane, etc., in the XR environment 128. As an example, referring to Figure 8A, the electronic device 120 displays a first indicator element 812A on the XR content 802 having a first size corresponding to a coincidence point between the XR content 802 and a ray emanating from the leading edge / end of a representation 805 of a proxy object 804.

[0264] In some implementations, the size parameter (e.g., radius) of the first graphical element is a function of the distance between the first part of the content and the physical object (1116). In some implementations, the size of the first indicator element increases as the distance between the first part of the content and the physical object decreases, and the size of the first indicator element decreases as the distance between the first part of the content and the physical object increases. As an example, referring to Figure 8A, the electronic device 120 displays a first indicator element 812A on the XR content 802 having a first size corresponding to the point of agreement between the XR content 802 and the rays emanating from the leading edge / end of the representation 805 of the proxy object 804. In this example, the size of the first indicator element 812A is a function of the first distance 814. As another example, referring to Figure 8B, the electronic device 120 displays a second indicator element 812B on the XR content 802, having a second size corresponding to the point of agreement between the XR content 802 and the rays emanating from the leading / ending of the representation 805 of the proxy object 804. As shown in Figure 8B, the representation 805 of the proxy object 804 is at a second distance 824 from the XR content 802, which is smaller than the first distance 814 in Figure 8A. For example, the second size of the second indicator element 812B is larger than the first size of the first indicator element 812A.

[0265] In some implementations, as represented by block 1118, following a determination that a grip posture does not correspond to a first grip, method 1100 includes displaying a second graphical element via a display device that indicates the user's gaze direction to the content. For example, the second graphical element is different from the first graphical element. For example, the second graphical element is displayed at matching points where rays projected from one or more of the user's eyes strike content in the XR environment, such as a 2D canvas, a 3D marking area, or a backplane. For example, referring to Figure 8E, the electronic device 120 displays a gaze direction indicator element 852 on the XR content 802 that is associated with the focus of the user's 149 eye in the XR environment 128, based on gaze tracking.

[0266] In some implementations, the size parameter (e.g., radius) of the second graphical element is a function of the distance between the second part of the content and one or more of the user's eyes (1120). In some implementations, the size of the second indicator element increases as the distance between the first part of the content and the physical object decreases, and the size of the second indicator element decreases as the distance between the first part of the content and the physical object increases.

[0267] In some implementations, as represented by block 1122, method 1100 includes detecting subsequent inputs by physical objects associated with the movement of content (e.g., translation and / or rotation) while the content is being displayed, and moving the content based on the subsequent inputs in response to the detection of these subsequent inputs. As an example, referring to Figure 8D, the electronic device 120 translates the XR content 802 in the XR environment 128 in response to detecting the translational movement 832 of the proxy object 804 in Figure 8C. For example, the direction and displacement of the translational movement of the XR content 802 in the XR environment 128 correspond to the spatial parameters of the translational movement 832 of the proxy object 804 in Figure 8C (e.g., change in position value, change in rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc.). It will be understood by those skilled in the art that the XR content 802 can also be rotated.

[0268] As another example, referring to Figure 8G, the electronic device 120 translates the XR content 802 within the XR environment 128 in response to detecting the translational movement 862 in the line of sight direction 806 in Figure 8F. For example, the direction and displacement of the translational movement of the XR content 802 within the XR environment 128 correspond to the spatial parameters of the translational movement 862 in the line of sight direction 806 in Figure 8F (e.g., changes in position value, changes in rotation value, displacement, spatial acceleration, spatial velocity, angular acceleration, angular velocity, etc.).

[0269] In some implementations, detecting a subsequent input corresponds to obtaining an instruction that an affordance on a physical object has been activated, and detecting at least one of the rotational or translational movement of the physical object (1124). For example, detecting the activation of an affordance corresponds to the selection portion of the subsequent input.

[0270] In some implementations, detecting a subsequent input corresponds to obtaining an instruction that an input value associated with how tightly a physical object is being grasped by the user exceeds a threshold input value, and detecting at least one of rotational or translational movement of the physical object (1126). For example, an input (pressure) value corresponds to a selection of the subsequent input. In some implementations, the pressure threshold is either non-deterministic (i.e., a given pressure value) or deterministic based on one or more factors such as user preference, usage history, current content, or current context.

[0271] In some implementations, the magnitude of the subsequent input is modified by an amplification factor to determine the magnitude of the content movement (1128). In some implementations, the amplification factor is non-deterministic (e.g., a predetermined value) or deterministic, based on one or more factors such as user preference, usage history, selected content, and current context.

[0272] While various embodiments of the implementations within the scope of the attached claims have been described above, it is clear that the various features of the above-described implementations may be embodied in a wide variety of forms, and the specific structures and / or functions described above are merely illustrative. Those skilled in the art will understand that the embodiments described herein may be implemented independently of any other embodiments, and two or more of these embodiments may be combined in various ways. For example, an apparatus may be implemented and / or a method may be performed using any number of embodiments described herein. In addition, such an apparatus may be implemented and / or a method may be performed using one or more of the embodiments described herein, or other structures and / or functions.

[0273] In this specification, terms such as “first,” “second,” etc., may be used to describe various elements, but it will be understood that these elements should not be limited by these terms. These terms are used solely to distinguish one element from another. For example, a “first media item” can be called a “second media item” without changing the meaning of the description, as long as the name is consistently changed for all occurrences of the “first media item” and the name is consistently changed for all occurrences of the “second media item.” Both the first and second media items are media items, but they are not the same media item.

[0274] The terms used herein are for the purpose of describing specific implementations and are not intended to limit the scope of the claims. When used in the descriptions of the implementations described and in the appended claims, the singular forms "a," "an," and "the" are intended to include the plural form unless the context explicitly indicates otherwise. Furthermore, when used herein, the term "and / or" should be understood to mean and include any and all possible combinations of one or more of the enumerated items relating to the description. When the terms "comprises" and / or "comprising" are used herein, they specify the presence of the described features, integers, steps, actions, elements, and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, actions, elements, components, and / or groups thereof.

[0275] When used herein, the term "if" may be interpreted, depending on the context, as meaning "at the time" or "on the occasion of" or "in accordance with the determination" that the previously stated condition is true, or "in accordance with the determination" or "in accordance with the determination" or "in accordance with the determination." Similarly, the phrases "[when it is determined that the previously stated condition is true]," "[when the previously stated condition is true]," or "[when the previously stated condition is true]" may be interpreted as meaning "in the event of the determination that the previously stated condition is true," "in accordance with the determination," "in accordance with the determination," "when detected," or "in accordance with the determination."

Claims

1. It is a method, In a computing system comprising non-temporary memory and one or more processors, and communicatively coupled to a display device and one or more input devices, The display device is used to display a first set of graphical elements associated with a first set of output modalities within an augmented reality (XR) environment. While the first set of graphical elements is being displayed, the first movement of a physical object is detected, In response to detecting the first movement of the physical object, In accordance with the determination that the physical object has exceeded a distance threshold for a first graphical element among the first plurality of graphical elements due to the first movement of the physical object, the first output modality associated with the first graphical element is selected as the current output modality of the physical object. In accordance with the determination that the physical object has exceeded the distance threshold for a second graphical element among the first plurality of graphical elements due to the first movement of the physical object, the second output modality associated with the second graphical element is selected as the current output modality of the physical object. Methods that include...

2. The method according to claim 1, wherein the first and second output modalities cause different visual changes within the XR environment.

3. According to the determination that the physical object has exceeded the distance threshold for the first graphical element among the first plurality of graphical elements as a result of the first movement of the physical object, Maintaining the display of the first graphical element adjacent to the physical object, To stop displaying the remaining first graphical elements that do not include the first graphical element, In accordance with the determination that the physical object has exceeded the distance threshold for the second graphical element among the first plurality of graphical elements due to the first movement of the physical object, Maintaining the display of the second graphical element adjacent to the physical object, To stop displaying the remaining first graphical elements that do not include the second graphical element, The method according to claim 1 or 2, further comprising:

4. After selecting the first output modality associated with the first graphical element as the current output modality of the physical object, the second movement of the physical object is detected. In response to detecting a second movement of the physical object, the first graphical element is moved based on the second movement of the physical object in order to maintain the display of the first graphical element adjacent to the physical object. The method according to claim 3, further comprising:

5. After stopping the display of the remaining first graphical elements, the system obtains a touch input instruction directed at the physical object. In response to receiving the aforementioned touch input instruction, the first set of graphical elements are redisplayed within the XR environment via the display device. The method according to claim 3 or 4, further comprising:

6. The process further includes obtaining touch input instructions directed to the physical object before displaying the first set of graphical elements, Displaying the first set of graphical elements within the XR environment is done in response to receiving the touch input instruction, The method according to any one of claims 1 to 5, including the method described in any one of claims 1 to 5.

7. In accordance with the determination that the physical object does not exceed the distance threshold with respect to the first graphical element or the second graphical element as a result of the first movement of the physical object, Maintaining the initial output modality as the current output modality of the physical object, Maintaining the display of the first set of graphical elements, The method according to any one of claims 1 to 6, further comprising:

8. After selecting the first output modality associated with the first graphical element as the current output modality of the physical object, the subsequent marking input by the physical object is detected. In response to detecting the subsequent marking input, one or more marks are displayed in the XR environment via the display device based on the subsequent marking input and the first output modality. The method according to any one of claims 1 to 7, further comprising:

9. In response to the detection of the subsequent marking input, Based on the determination that an input associated with how strongly the physical object is pressed against a physical surface corresponds to a first input value, the display device displays one or more marks having a first appearance, which is associated with the parameters of the one or more marks corresponding to the first input value, within the XR environment, based on the subsequent marking input and the first output modality. Based on the determination that the input associated with how strongly the physical object is pressed against the physical surface corresponds to a second input value, the display device displays one or more marks having a second appearance, which is associated with the parameters of the one or more marks corresponding to the second input value, within the XR environment, based on the subsequent marking input and the first output modality. The method according to claim 8, further comprising:

10. The method according to claim 9, wherein the parameter corresponds to one of the radius, width, thickness, intensity, translucency, opacity, color, or texture of the one or more marks in the XR environment.

11. In response to the detection of the subsequent marking input, In accordance with the determination that an input associated with how tightly the physical object is being grasped by the user corresponds to a first input value, the display device, based on the subsequent marking input and the first output modality, displays one or more marks having a first appearance, which are associated with the parameters of the one or more marks corresponding to the first input value, within the XR environment. In accordance with the determination that the input associated with how tightly the physical object is being held by the user corresponds to a second input value, the display device, based on the subsequent marking input and the first output modality, displays one or more marks having a second appearance, which is associated with the parameters of the one or more marks corresponding to the second input value, within the XR environment. The method according to claim 8, further comprising:

12. The method according to claim 11, wherein the parameter corresponds to one of the radius, width, thickness, intensity, translucency, opacity, color, or texture of the one or more marks in the XR environment.

13. Before displaying the first set of graphical elements, which are functions of the grip pose, the grip pose associated with the current manner in which the physical object is held by the user is obtained. In response to acquiring the aforementioned gripping position, In accordance with the determination that the grip posture corresponds to a first grip posture, the display device displays the first set of graphical elements associated with the first set of output modalities in the XR environment. In accordance with the determination that the grip posture corresponds to a second grip posture different from the first grip posture, the display device displays a second set of graphical elements associated with a second set of output modalities in the XR environment. The method according to any one of claims 1 to 12, further comprising:

14. After displaying the first set of graphical elements associated with the first set of output modalities within the XR environment, the change in grip posture from the first grip posture to the second grip posture is detected. In response to detecting the change in the grip posture, the display of the first set of graphical elements in the XR environment is replaced with the second set of graphical elements associated with the second set of output modalities in the XR environment. The method according to claim 13, further comprising:

15. Before displaying the first set of graphical elements, information is obtained indicating whether the first or second end of the physical object is facing outwards. In response to obtaining the information indicating whether the first end or the second end of the physical object is facing outward, In accordance with the determination that the first end of the physical object is facing outward, the display device displays the first set of graphical elements associated with the first set of output modalities in the XR environment. In accordance with the determination that the second end of the physical object is facing outward, the display device displays a second set of graphical elements associated with a second set of output modalities in the XR environment. The method according to any one of claims 1 to 14, further comprising:

16. After displaying the first set of graphical elements associated with the first set of output modalities within the XR environment, the change from the outward-facing first end of the physical object to the outward-facing second end of the physical object is detected. In response to detecting the change from the outward-facing first end of the physical object to the outward-facing second end of the physical object, the display device displays the second set of graphical elements associated with the second set of output modalities within the XR environment. The method according to claim 15, further comprising:

17. It is a device, One or more processors, Non-temporary memory and An interface for communicating with a display device and one or more input devices, One or more programs stored in the non-temporary memory, which, when executed by the one or more processors, cause the device to perform the method described in any one of claims 1 to 16, A device equipped with the following features.

18. Non-temporary memory for storing one or more programs, wherein when the one or more programs are executed by one or more processors of a device having interfaces for communicating with a display device and one or more input devices, the non-temporary memory causes the device to execute the method according to any one of claims 1 to 16.

19. It is a device, One or more processors, Non-temporary memory and An interface for communicating with a display device and one or more input devices, Means for causing the device to perform the method described in any one of claims 1 to 16, A device equipped with the following features.

20. It is a method, In a computing system comprising non-temporary memory and one or more processors, and communicatively coupled to a display device and one or more input devices, The display device is used to display a user interface, While the aforementioned user interface is displayed, the system detects marking input by a physical object, In response to the detection of the aforementioned marking input, In accordance with the determination that the marking input is directed toward a physical surface, a mark is displayed in the user interface based on the marking input via the display device, wherein the parameters of the mark displayed based on the marking input are determined based on how strongly the physical object is pressed against the physical surface. In accordance with the determination that the marking input is not directed toward the physical surface, the mark is displayed in the user interface via the display device based on the marking input, wherein the parameters of the mark displayed based on the marking input are determined based on how tightly the physical object is being grasped by the user. Methods that include...

21. The method according to claim 20, wherein the parameter corresponds to one of the radius, width, thickness, intensity, translucency, opacity, color, or texture of the mark in the user interface.

22. The method according to claim 20 or 21, wherein the user interface includes a two-dimensional marking area on which the mark is displayed.

23. The method according to claim 20 or 21, wherein the user interface includes a three-dimensional marking area on which the mark is displayed.

24. After displaying the mark within the user interface, the system detects subsequent input from the physical object associated with moving the mark within the user interface. In response to detecting the subsequent input, the mark is moved within the user interface based on the subsequent input, The method according to any one of claims 20 to 23, further comprising:

25. The subsequent input is detected. To obtain an instruction that the affordance on the aforementioned physical object has been activated, To detect at least one of the rotational or translational movement of the physical object, The method according to claim 24, which corresponds to the method of claim 24.

26. The subsequent input is detected. The process involves obtaining an instruction that an input value associated with how tightly the physical object is being held by the user exceeds a threshold input value, To detect at least one of the rotational or translational movement of the physical object, The method according to claim 24, which corresponds to the method of claim 24.

27. In response to detecting the subsequent input, the appearance of at least some of the content within the user interface is changed while moving the mark within the user interface. The method according to any one of claims 24 to 26, further comprising:

28. In response to the detection of the aforementioned marking input, In accordance with the determination that the marking input is directed toward the physical surface, the display device displays a simulated shadow in the XR environment corresponding to the distance between the physical surface and the physical object. The method according to any one of claims 24 to 27, further comprising:

29. The display device is used to display a user interface element having multiple different selectable tools associated with markings within the user interface. The method according to any one of claims 20 to 28, further comprising:

30. The method according to claim 29, wherein the user interface element is fixed to a point in space.

31. The method according to claim 29, wherein the user interface element is fixed to a point within the user's field of view of the computing system.

32. It is a device, One or more processors, Non-temporary memory and An interface for communicating with a display device and one or more input devices, One or more programs stored in the non-temporary memory, which, when executed by the one or more processors, cause the device to execute the method described in any one of claims 20 to 31, A device equipped with the following features.

33. Non-temporary memory for storing one or more programs, wherein when the one or more programs are executed by one or more processors of a device having interfaces for communicating with a display device and one or more input devices, the non-temporary memory causes the device to execute the method according to any one of claims 20 to 31.

34. It is a device, One or more processors, Non-temporary memory and An interface for communicating with a display device and one or more input devices, The device is provided with means for causing it to perform the method described in any one of claims 20 to 31, A device equipped with the following features.

35. It is a method, In a computing system comprising non-temporary memory and one or more processors, and communicatively coupled to a display device and one or more input devices, Displaying content via the aforementioned display device, Detecting a selection input while the aforementioned content is being displayed and while the physical object is being held by the user, The operation includes, in response to detecting the aforementioned selection input, executing an operation corresponding to the aforementioned selection input, wherein the operation is A selection operation is performed on a first portion of the content, based on a determination that the grip posture associated with the manner in which the physical object is held by the user corresponds to a first grip, wherein the first portion of the content is selected based on the direction in which a predetermined portion of the physical object is facing. The selection operation is performed on a second portion of the content that is different from the first portion of the content, based on the determination that the grip posture associated with the manner in which the physical object is held by the user does not correspond to the first grip, wherein the second portion of the content is selected based on the user's line of sight. Methods that include...

36. The method according to claim 35, wherein performing the selection operation includes changing the appearance of the first or second portion of the content.

37. In accordance with the determination that the grip posture corresponds to the first grip, a first graphical element is displayed via the display device indicating the direction in which the predetermined portion of the physical object is pointing toward the content. The method according to any one of claims 35 to 36, further comprising:

38. The method according to claim 37, wherein the size parameter of the first graphical element is a function of the distance between the first portion of the content and the physical object.

39. In accordance with the determination that the grip posture does not correspond to the first grip, a second graphical element indicating the user's gaze direction to the content is displayed via the display device. The method according to any one of claims 35 to 36, further comprising:

40. The method according to claim 39, wherein the size parameter of the second graphical element is a function of the distance between the second portion of the content and one or more of the user's eyes.

41. While the content is being displayed, subsequent input from the physical object associated with moving the content is detected. In response to detecting the subsequent input, the content is moved based on the subsequent input, The method according to any one of claims 35 to 40, further comprising:

42. The subsequent input is detected. To obtain an instruction that the affordance on the aforementioned physical object has been activated, To detect at least one of the rotational or translational movement of the physical object, The method according to claim 41, which corresponds to the present invention.

43. The subsequent input is detected. The process involves obtaining an instruction that an input value associated with how tightly the physical object is being held by the user exceeds a threshold input value, To detect at least one of the rotational or translational movement of the physical object, The method according to claim 41, which corresponds to the present invention.

44. The method according to any one of claims 41 to 43, wherein the magnitude of the subsequent input is modified by an amplification factor to determine the magnitude of the movement of the content.

45. It is a device, One or more processors, Non-temporary memory and An interface for communicating with a display device and one or more input devices, One or more programs stored in the non-temporary memory, which, when executed by the one or more processors, cause the device to execute the method described in any one of claims 35 to 44, A device equipped with the following features.

46. Non-temporary memory for storing one or more programs, wherein when the one or more programs are executed by one or more processors of a device having interfaces for communicating with a display device and one or more input devices, the non-temporary memory causes the device to execute the method according to any one of claims 35 to 44.

47. It is a device, One or more processors, Non-temporary memory and An interface for communicating with a display device and one or more input devices, Means for causing the device to perform the method described in any one of claims 35 to 44, A device equipped with the following features.