Method and apparatus for dynamically selecting an operational modality of an object

CN117616365BActive Publication Date: 2026-09-29APPLE INC
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202280047899.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-07-05
Filing Date
2022-07-01
Publication Date
2026-09-29
Estimated Expiration
2042-07-01

AI Technical Summary

Technical Problem

这可能是笨拙和繁琐的过程

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117616365B_ABST
    Figure CN117616365B_ABST
Patent Text Reader

Abstract

In one implementation, a method for dynamically selecting an operational modality of a physical object. The method includes obtaining a user input vector including at least one user input indicator value associated with one of a plurality of different input modalities; obtaining tracking data associated with the physical object; generating a first characterization vector for the physical object based on the user input vector and the tracking data, the first characterization vector including a pose value and a user grip value, wherein the pose value characterizes a spatial relationship between the physical object and a user of the computing system and the user grip value characterizes a manner in which the physical object is gripped by the user; and selecting a first operational modality as a current operational modality of the physical object based on the first characterization vector.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure generally relates to interacting with an environment having objects, and more specifically to systems, methods, and approaches for dynamically selecting operational modes of physical objects. Background Technology

[0002] To change the behavior of the stylus, users typically select different tools from a menu of available tools. This can be a clumsy and tedious process. Attached Figure Description

[0003] Therefore, this disclosure will be understood by those skilled in the art, and a more detailed description can be made with reference to some exemplary embodiments, some of which are shown in the accompanying drawings.

[0004] Figure 1 It is a block diagram of an exemplary operational architecture based on some specific implementations.

[0005] Figure 2 It is a block diagram of an exemplary controller based on some specific implementations.

[0006] Figure 3 It is a block diagram of an exemplary electronic device based on some specific implementations.

[0007] Figure 4 It is a block diagram of an exemplary control device based on some specific implementations.

[0008] Figure 5A This is a block diagram of the first part of an exemplary content delivery architecture based on some specific implementations.

[0009] Figure 5B Exemplary data structures according to some specific implementations are shown.

[0010] Figure 5C This is a block diagram of the second part of an exemplary content delivery architecture based on some specific implementations.

[0011] Figures 6A to 6R This shows a sequence of instances of content delivery scenarios based on some specific implementations.

[0012] Figure 7 It is a flowchart representation of a method for dynamically selecting the operation mode of a physical object based on some specific implementation.

[0013] As is customary, the various features shown in the accompanying drawings may not be drawn to scale. Therefore, for clarity, the dimensions of various features may be arbitrarily expanded or reduced. Additionally, some drawings may not depict all components of a given system, method, or apparatus. Finally, similar reference numerals may be used throughout the specification and drawings to denote similar features. Summary of the Invention

[0014] The various specific embodiments disclosed herein include devices, systems, and methods for dynamically selecting the operating mode of a physical object. According to some embodiments, the method is performed at a computing system including non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device. The method includes: obtaining a user input vector including at least one user input indicator value associated with one of a plurality of different input modes; obtaining tracking data associated with a physical object; generating a first representation vector of the physical object based on the user input vector and the tracking data, the first representation vector including an attitude value and a user grip value, wherein the attitude value represents the spatial relationship between the physical object and a user of the computing system, and the user grip value represents how the physical object is gripped by the user; and selecting a first operating mode as the current operating mode of the physical object based on the first representation vector.

[0015] According to some embodiments, an electronic device includes one or more displays, one or more processors, non-transitory memory, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium stores instructions that, when executed by one or more processors of the device, cause the device to perform or cause to perform any of the methods described herein. According to some embodiments, a device includes: one or more displays, one or more processors, non-transitory memory, and means for performing or causing to perform any of the methods described herein.

[0016] According to some embodiments, a computing system includes one or more processors, non-transitory memory, an interface for communicating with a display device and one or more input devices, and one or more programs; the one or more programs are stored in the non-transitory memory and configured to be executed by the one or more processors, and the one or more programs include instructions for performing or causing to perform operations of any of the methods described herein. According to some embodiments, a non-transitory computer-readable storage medium has instructions stored therein that, when executed by one or more processors of a computing system having an interface for communicating with a display device and one or more input devices, cause the computing system to perform or cause to perform operations of any of the methods described herein. According to some embodiments, a computing system includes one or more processors, non-transitory memory, an interface for communicating with a display device and one or more input devices, and means for performing or causing to perform operations of any of the methods described herein. Detailed Implementation

[0017] Numerous details have been described to provide a thorough understanding of the exemplary embodiments illustrated in the accompanying drawings. However, the drawings illustrate only some exemplary aspects of this disclosure and should not be considered limiting. Those skilled in the art will understand that other effective aspects and / or variations do not include all the specific details described herein. Furthermore, well-known systems, methods, components, devices, and circuits have not been described exhaustively so as not to obscure further relevant aspects of the exemplary embodiments described herein.

[0018] People can sense or interact with the physical environment or world without using electronic devices. Physical features such as physical objects or surfaces can be included within the physical environment. For example, a physical environment may correspond to a physical city with physical buildings, roads, and vehicles. People can directly perceive or interact with the physical environment through various means such as smell, sight, taste, hearing, and touch. This is in contrast to extended reality (XR) environments, which can refer to partially or fully simulated environments that people can sense or interact with using electronic devices. XR environments can include virtual reality (VR) content, mixed reality (MR) content, augmented reality (AR) content, and so on. Using an XR system, physical movement of a person or a portion of their representation can be tracked, and in response, the properties of virtual objects in the XR environment can be altered in a manner that conforms to at least one law of nature. For example, an XR system can detect head movement of a user and adjust the auditory and graphical content presented to the user in a way that simulates how sound and view would change in a physical environment. In other examples, an XR system can detect movement of electronic devices (e.g., laptops, tablets, mobile phones, etc.) that present the XR environment. Therefore, XR systems can simulate how sound and view will change in the physical environment to adjust the auditory and graphical content presented to the user. In some instances, other inputs, such as representations of body movement (e.g., voice commands), can enable XR systems to adjust the properties of the graphical content.

[0019] Numerous types of electronic systems allow users to sense or interact with their XR environment. An incomplete list of examples includes lenses with integrated display capabilities placed on the user's eyes (e.g., contact lenses), head-up displays (HUDs), projection-based systems, head-mounted systems, windows or windshields with integrated display technology, headsets / earpieces, input systems with or without haptic feedback (e.g., handheld or wearable controllers), smartphones, tablets, desktop / laptop computers, and speaker arrays. Head-mounted systems may include opaque displays and one or more speakers. Other head-mounted systems may be configured to receive opaque external displays, such as the opaque external display of a smartphone. Head-mounted systems may use one or more image sensors to capture images / video of the physical environment, or one or more microphones to capture audio of the physical environment. Some head-mounted systems may include transparent or semi-transparent displays instead of opaque displays. Transparent or semi-transparent displays may guide light representing an image to the user's eyes through media such as holographic media, optical waveguides, optical combiners, optical reflectors, other similar technologies, or combinations thereof. Various display technologies can be used, such as liquid crystal on silicon, LED, μLED, OLED, laser scanning light sources, digital light projection, or combinations thereof. In some examples, transparent or translucent displays can be selectively controlled to become opaque. Projection-based systems can utilize retinal projection technology, which projects images onto a user's retina, or can project virtual content into a physical environment, such as onto a physical surface or as a hologram.

[0020] Figure 1 This is a block diagram of an exemplary operating architecture 100 according to some specific implementations. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and so as not to obscure further relevant aspects of the exemplary specific implementations disclosed herein. Therefore, as a non-limiting example, operating architecture 100 includes optional processing device 110 and electronic device 120 (e.g., tablet computer, mobile phone, laptop computer, near-eye system, wearable computing device, etc.).

[0021] In some implementations, processing device 110 is configured to manage and coordinate the XR experience (sometimes referred to herein as an "XR environment," "virtual environment," or "graphical environment") of user 149 using left-hand 150 and right-hand 152, and optionally other users. In some implementations, processing device 110 includes a suitable combination of software, firmware, and / or hardware. References are provided below. Figure 2The processing device 110 is described in more detail. In some embodiments, the processing device 110 is a computing device located locally or remotely relative to a physical environment 105. For example, the processing device 110 is a local server located within the physical environment 105. In another example, the processing device 110 is a remote server (e.g., a cloud server, a central server, etc.) located outside the physical environment 105. In some embodiments, the processing device 110 is communicatively coupled to the electronic device 120 via one or more wired or wireless communication channels 144 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.). In some embodiments, the functionality of the processing device 110 is provided by the electronic device 120. Thus, in some embodiments, components of the processing device 110 are integrated into the electronic device 120.

[0022] like Figure 1 As shown, user 149 holds controller 130 in his / her right hand 152. Figure 1 As shown, the controller 130 includes a first end 176 and a second end 177. In various embodiments, the first end 176 corresponds to the tip of the controller 130 (e.g., the tip of a pencil), and the second end 177 corresponds to the opposite end or bottom end of the controller 130 (e.g., the eraser of a pencil). Figure 1 As shown, controller 130 includes a touch-sensitive surface 175 to receive touch input from user 149. In some specific implementations, controller 130 includes a suitable combination of software, firmware, and / or hardware. References below... Figure 4 The controller 130 is described in more detail. In some embodiments, the controller 130 corresponds to an electronic device having a wired or wireless communication channel to the processing device 110. For example, the controller 130 corresponds to a stylus, a wearable finger device, a handheld device, etc. In some embodiments, the processing device 110 is communicatively coupled to the controller 130 via one or more wired or wireless communication channels 146 (e.g., Bluetooth, IEEE 802.11x, IEEE 802.16x, IEEE 802.3x, etc.).

[0023] In some embodiments, electronic device 120 is configured to present audio and / or video (A / V) content to user 149. In some embodiments, electronic device 120 is configured to present a user interface (UI) and / or XR environment 128 to user 149. In some embodiments, electronic device 120 includes a suitable combination of software, firmware, and / or hardware. References are provided below. Figure 3 The electronic device 120 is described in more detail.

[0024] According to some embodiments, when user 149 is physically present within physical environment 105, electronic device 120 presents an XR experience to user 149, wherein physical environment 105 includes a table 107 located within the field of view (FOV) 111 of electronic device 120. Thus, in some embodiments, user 149 holds electronic device 120 in one or both of his / her hands. In some embodiments, when presenting the XR experience, electronic device 120 is configured to present XR content (sometimes referred to herein as “graphical content” or “virtual content”), including XR cylinder 109, and to enable video pass-through of physical environment 105 (e.g., including table 107 or a representation thereof) on display 122. For example, the XR environment 128 including XR cylinder 109 is stereoscopic or three-dimensional (3D).

[0025] In one example, XR cylinder 109 corresponds to display-locked content, such that when FOV 111 changes due to translational and / or rotational movement of electronic device 120, XR cylinder 109 remains displayed in the same position on display 122. As another example, XR cylinder 109 corresponds to world-locked content, such that when FOV 111 changes due to translational and / or rotational movement of electronic device 120, XR cylinder 109 remains displayed in its original position. Therefore, in this example, if FOV 111 does not include the original position, then XR environment 128 will not include XR cylinder 109. For example, electronic device 120 corresponds to near-eye systems, mobile phones, tablets, laptops, wearable computing devices, etc.

[0026] In some embodiments, display 122 corresponds to an additive display that enables optical transmission of the physical environment 105 (including table 107). For example, display 122 corresponds to a transparent lens, and electronic device 120 corresponds to a pair of glasses worn by user 149. Thus, in some embodiments, electronic device 120 presents a user interface by projecting XR content (e.g., XR cylinder 109) onto the additive display, which is then overlaid on the physical environment 105 from the perspective of user 149. In some embodiments, electronic device 120 presents a user interface by displaying XR content (e.g., XR cylinder 109) on the additive display, which is then overlaid on the physical environment 105 from the perspective of user 149.

[0027] In some embodiments, user 149 wears electronic device 120, such as a near-eye system. Therefore, electronic device 120 includes one or more displays (e.g., a single display or one display per eye) provided for displaying XR content. For example, electronic device 120 surrounds the field of view (FOV) of user 149. In such embodiments, electronic device 120 presents the XR environment 128 by displaying data corresponding to the XR environment 128 on one or more displays or by projecting data corresponding to the XR environment 128 onto the retina of user 149.

[0028] In some embodiments, electronic device 120 includes an integrated display (e.g., a built-in display) for displaying XR environment 128. In some embodiments, electronic device 120 includes a head-mountable housing. In various embodiments, the head-mountable housing includes an attachment area to which another device having a display can be attached. For example, in some embodiments, electronic device 120 can be attached to the head-mountable housing. In various embodiments, the head-mountable housing is shaped to form a receiver for receiving another device (e.g., electronic device 120) including a display. For example, in some embodiments, electronic device 120 slides / snapes into or otherwise attaches to the head-mountable housing. In some embodiments, the display of the device attached to the head-mountable housing presents (e.g., displays) XR environment 128. In some embodiments, electronic device 120 is replaced by an XR room, housing, or chamber configured to present XR content, in which user 149 does not wear electronic device 120.

[0029] In some embodiments, processing device 110 and / or electronic device 120 cause the XR representation of user 149 to move within XR environment 128 based on motion information (e.g., body posture data, eye tracking data, hand / limb / finger / end-point tracking data, etc.) from electronic device 120 and / or optional remote input devices within physical environment 105. In some embodiments, the optional remote input devices correspond to fixed or movable sensory devices (e.g., image sensors, depth sensors, infrared (IR) sensors, event cameras, microphones, etc.) within physical environment 105. In some embodiments, each remote input device is configured to collect / capture input data when user 149 is physically within physical environment 105 and provide the input data to processing device 110 and / or electronic device 120. In some embodiments, the remote input device includes a microphone, and the input data includes audio data (e.g., voice samples) associated with user 149. In some embodiments, the remote input device includes an image sensor (e.g., a camera), and the input data includes images of user 149. In some embodiments, the input data characterizes the user 149's body posture at different times. In some embodiments, the input data characterizes the user 149's head posture at different times. In some embodiments, the input data characterizes hand tracking information associated with the user 149's hand at different times. In some embodiments, the input data characterizes the velocity and / or acceleration of the user 149's body parts (such as his / her hand). In some embodiments, the input data indicates the user 149's joint positioning and / or joint orientation. In some embodiments, the remote input device includes feedback devices such as speakers, lights, etc.

[0030] Figure 2This is a block diagram of an example of a processing apparatus 110 according to some specific implementations. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, in some specific implementations, the processing device 110 includes one or more processing units 202 (e.g., microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), graphics processing units (GPUs), central processing units (CPUs), processing kernels, etc.), one or more input / output (I / O) devices 206, one or more communication interfaces 208 (e.g., Universal Serial Bus (USB), IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, Global System for Mobile Communications (GSM), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Global Positioning System (GPS), Infrared (IR), Bluetooth, ZigBee, and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 210, memory 220, and one or more communication buses 204 for interconnecting these components and various other components.

[0031] In some embodiments, the one or more communication buses 204 include circuitry for communication between interconnecting system components and control system components. In some embodiments, one or more I / O devices 206 include at least one of a keyboard, mouse, touchpad, touchscreen, joystick, one or more microphones, one or more speakers, one or more image sensors, one or more displays, etc.

[0032] Memory 220 includes high-speed random access memory, such as dynamic random access memory (DRAM), static random access memory (SRAM), double data rate random access memory (DDR RAM), or other random access solid-state memory devices. In some embodiments, memory 220 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 220 optionally includes one or more storage devices located remotely from the one or more processing units 202. Memory 220 includes a non-transitory computer-readable storage medium. In some embodiments, memory 220 or the non-transitory computer-readable storage medium of memory 220 stores the following reference. Figure 2 The following programs, modules, and data structures, or subsets thereof.

[0033] The operating system 230 includes processes for handling various basic system services and for performing hardware-related tasks.

[0034] In some implementations, the data acquisition unit 242 is configured to acquire data (e.g., captured image frames of the physical environment 105, presentation data, input data, user interaction data, camera pose tracking information, eye tracking information, head / body pose tracking information, hand / limb / finger / limb tracking information, sensor data, position data, etc.) from at least one of the I / O devices 206 of the processing device 110, the I / O devices and sensors 306 of the electronic device 120, and optional remote input devices. To this end, in various implementations, the data acquisition unit 242 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0035] In some implementations, the mapper and locator engine 244 is configured to map the physical environment 105 and at least track the location / position of the electronic device 120 or the user 149 relative to the physical environment 105. To this end, in various implementations, the mapper and locator engine 244 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0036] In some implementations, data transmitter 246 is configured to transmit data (e.g., rendering data, such as rendered image frames associated with an XR environment, location data, etc.) to at least electronic device 120 and optionally one or more other devices. To this end, in various implementations, data transmitter 246 includes instructions and / or logic for instructions, as well as heuristics and metadata for heuristics.

[0037] In some specific implementations, privacy architecture 508 is configured to ingest data and filter user information and / or identification information within that data based on one or more privacy filters. See below for reference. Figure 5A Privacy Architecture 508 is described in more detail. To this end, in various specific implementations, Privacy Architecture 508 includes instructions and / or logic for the instructions, as well as heuristics and metadata for the heuristics.

[0038] In some specific implementations, the object tracking engine 510 is configured to determine / generate object tracking vectors 511 for tracking physical objects (e.g., controller 130 or proxy objects) based on tracking data, and to update the pose representation vectors 515 and 511 over time. For example, as Figure 5B As shown, the object tracking vector 511 includes translation values ​​572 of the physical object (e.g., associated with x, y, and z coordinates relative to the physical environment 105 or the world), rotation values ​​574 of the physical object (e.g., roll, pitch, and yaw), one or more pressure values ​​576 associated with the physical object, optional touch input information 578 associated with the physical object, etc. See below for further details. Figure 5AThe object tracing engine 510 is described in more detail. To this end, in various specific implementations, the object tracing engine 510 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0039] In some specific implementations, the eye-tracking engine 512 is configured to determine / generate, such as, based on input data. Figure 5B The eye-tracking vector 513 shown is (e.g., having a gaze direction) and is updated over time. For example, the gaze direction indicates a point in the physical environment 105 that the user 149 is currently viewing (e.g., associated with x, y, and z coordinates relative to the physical environment 105 or the world), a physical object, or a region of interest (ROI). As another example, the gaze direction indicates a point in the XR environment 128 that the user 149 is currently viewing (e.g., associated with x, y, and z coordinates relative to the XR environment 128), an XR object, or a region of interest (ROI). See below for further details. Figure 5A The eye-tracking engine 512 is described in more detail. To this end, in various specific implementations, the eye-tracking engine 512 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0040] In some specific implementations, the body / head pose tracking engine 514 is configured to determine / generate a pose representation vector 515 based on input data and update the pose representation vector 515 over time. For example, as Figure 5B As shown, the posture representation vector 515 includes a head posture descriptor 592A (e.g., up, down, neutral, etc.), head posture translation values ​​592B, head posture rotation values ​​592C, body posture descriptor 594A (e.g., standing, sitting, prone, etc.), body part / limb / limb / joint translation values ​​594B, body part / limb / limb / joint rotation values ​​594C, etc. See below for reference. Figure 5A The body / head pose tracking engine 514 is described in more detail. Therefore, in various specific embodiments, the body / head pose tracking engine 514 includes instructions and / or logic components for those instructions, as well as heuristics and metadata for those heuristics. In some specific embodiments, as a supplement to or alternative to the processing device 110, the object tracking engine 510, the eye tracking engine 512, and the body / head pose tracking engine 514 may reside on the electronic device 120.

[0041] In some implementations, the content selector 522 is configured to select XR content (sometimes referred to herein as "graphic content" or "virtual content") from the content library 525 based on one or more user requests and / or inputs (e.g., voice commands, selections from the user interface (UI) menu of the XR content item, etc.). See below for reference. Figure 5A Content selector 522 is described in more detail. Therefore, in various specific implementations, content selector 522 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0042] In some implementations, content library 525 includes multiple content items, such as auditory / visual (A / V) content, virtual agents (VA) and / or XR content, objects, items, scenes, etc. As an example, XR content includes 3D reconstructions of user-captured videos, movies, TV series, and / or other XR content. In some implementations, content library 525 is pre-populated by user 149 or manually created. In some implementations, content library 525 is located locally relative to processing device 110. In some implementations, content library 525 is located remotely from processing device 110 (e.g., on a remote server, cloud server, etc.).

[0043] In some specific implementations, the operation mode manager 540 is configured to be based on Figure 5B The physical object shown (e.g., controller 130 or agent object) uses a representation vector 543, including pose values ​​and user grip values, to select the operating mode of the physical object. The representation vector 543 is a function of user input vectors (e.g., a combination of eye-tracking vector 513 and pose representation vector 515) and tracking data (e.g., object tracking vector 511). See below for further details. Figure 5A The operation modality manager 540 is described in more detail. To this end, in various specific implementations, the operation modality manager 540 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics. In some specific implementations, the operation modality manager 540 includes a representation engine 542 and an operation modality selector 544.

[0044] In some implementations, the representation engine 542 is configured to determine / generate a physical object representation vector 543 based on user input vectors (e.g., a combination of eye-tracking vector 513 and pose representation vector 515) and tracking data (e.g., object tracking vector 511). In some implementations, the representation engine 542 is also configured to update the pose representation vector 515 over time. Figure 5B As shown, the representation vector 543 includes user grip value 5102 and pose value 5104. See below for reference. Figure 5A The representation engine 542 is described in more detail. Therefore, in various specific implementations, the representation engine 542 includes instructions and / or logical components for those instructions, as well as heuristics and metadata for those heuristics.

[0045] In some implementations, the operation mode selector 544 is configured to select the current operation mode of the physical object (when interacting with the XR environment 128) based on the representation vector 543. For example, the operation mode may include dictation mode, digital assistant mode, navigation mode, manipulation mode, marking mode, erasing mode, pointing mode, embody mode, etc. See below for reference. Figure 5A The operation mode selector 544 is described in more detail. Therefore, in various specific implementations, the operation mode selector 544 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0046] In some implementations, the content manager 530 is configured to manage and update the layout, settings, structure, etc., of the XR environment 128, including one or more of the VA, XR content, and one or more user interface (UI) elements associated with the XR content. See below for reference. Figure 5C Content Manager 530 is described in more detail. To this end, in various embodiments, Content Manager 530 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics. In some embodiments, Content Manager 530 includes a buffer 534, a content updater 536, and a feedback engine 538. In some embodiments, buffer 534 includes XR content, rendered image frames, etc., for one or more past instances and / or frames.

[0047] In some implementations, the content updater 536 is configured to modify the XR environment 105 over time based on translational or rotational movement of physical objects within the electronic device 120 or physical environment 128, user input (e.g., hand / limb tracking input, eye tracking input, touch input, voice commands, manipulation input of physical objects, etc.). To this end, in various implementations, the content updater 536 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0048] In some implementations, the feedback engine 538 is configured to generate sensory feedback (e.g., visual feedback such as text or lighting changes, audio feedback, haptic feedback, etc.) associated with the XR environment 128. To this end, in various implementations, the feedback engine 538 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0049] In some implementations, rendering engine 550 is configured to render an XR environment 128 (sometimes referred to herein as a “graphics environment” or a “virtual environment”) or image frames associated with that XR environment, as well as VAs, XR content, one or more UI elements associated with the XR content, etc. To this end, in various implementations, rendering engine 550 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics. In some implementations, rendering engine 550 includes a pose determiner 552, a renderer 554, an optional image processing architecture 562, and an optional compositor 564. Those skilled in the art will understand that for a video passthrough configuration, the optional image processing architecture 562 and the optional compositor 564 may be present, but for a full VR or optical passthrough configuration, the optional image processing architecture and the optional compositor may be removed.

[0050] In some implementations, the pose determiner 552 is configured to determine the current camera pose of the electronic device 120 and / or the user 149 relative to the A / V content and / or XR content. See below for reference. Figure 5A The attitude determiner 552 is described in more detail. Therefore, in various specific implementations, the attitude determiner 552 includes instructions and / or logic for those instructions, as well as a heuristic and metadata for that heuristic.

[0051] In some implementations, renderer 554 is configured to render A / V content and / or XR content based on the current camera pose in relation to it. See below for reference. Figure 5A The renderer 554 is described in more detail. To this end, in various specific implementations, the renderer 554 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0052] In some implementations, the image processing architecture 562 is configured to acquire (e.g., receive, retrieve, or capture) an image stream including one or more images of the physical environment 105 from the current camera pose of the electronic device 120 and / or the user 149. In some implementations, the image processing architecture 562 is also configured to perform one or more image processing operations on the image stream, such as distortion, color correction, gamma correction, sharpening, noise reduction, white balance, etc. See below for reference. Figure 5A The image processing architecture 562 is described in more detail. To this end, in various specific implementations, the image processing architecture 562 includes instructions and / or logic components for those instructions, as well as heuristics and metadata for those heuristics.

[0053] In some implementations, compositor 564 is configured to composite rendered A / V content and / or XR content with a processed image stream from physical environment 105 of image processing architecture 562 to produce rendered image frames for display in XR environment 128. See below for reference. Figure 5A Synthesizer 564 is described in more detail. Therefore, in various specific implementations, synthesizer 564 includes instructions and / or logic components for those instructions, as well as heuristics and metadata for those heuristics.

[0054] Although the data acquirer 242, mapper and locator engine 244, data transmitter 246, privacy architecture 508, object tracking engine 510, eye tracking engine 512, body / head pose tracking engine 514, content selector 522, content manager 530, operation mode manager 540, and rendering engine 550 are shown residing on a single device (e.g., processing device 110), it should be understood that in other specific implementations, any combination of the data acquirer 242, mapper and locator engine 244, data transmitter 246, privacy architecture 508, object tracking engine 510, eye tracking engine 512, body / head pose tracking engine 514, content selector 522, content manager 530, operation mode manager 540, and rendering engine 550 may reside in a separate computing device.

[0055] In some specific implementations, the functions and / or components of the processing device 110 are the same as those described below. Figure 3 The electronic device 120 shown is combined with or provided by it. Furthermore, Figure 2 It is used more as a functional description of various features present in a specific implementation than as a structural diagram of the specific implementation described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 2 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0056] Figure 3This is a block diagram of an example of an electronic device 120 (e.g., a mobile phone, tablet computer, laptop computer, near-eye system, wearable computing device, etc.) according to some specific embodiments. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure more relevant aspects of the specific embodiments disclosed herein. For this purpose, as a non-limiting example, in some specific implementations, electronic device 120 includes one or more processing units 302 (e.g., microprocessors, ASICs, FPGAs, GPUs, CPUs, processing cores, etc.), one or more input / output (I / O) devices and sensors 306, one or more communication interfaces 308 (e.g., USB, IEEE 802.3x, IEEE 802.11x, IEEE 802.16x, GSM, CDMA, TDMA, GPS, IR, Bluetooth, ZIGBEE and / or similar types of interfaces), one or more programming (e.g., I / O) interfaces 310, one or more displays 312, image capture devices 370 (one or more optional internal and / or external image sensors), memory 320, and one or more communication buses 304 for interconnecting these components and various other components.

[0057] In some embodiments, one or more communication buses 304 include circuitry for interconnecting and communicating between control system components. In some embodiments, one or more I / O devices and sensors 306 include at least one of the following: an inertial measurement unit (IMU), an accelerometer, a gyroscope, a magnetometer, a thermometer, one or more physiological sensors (e.g., a blood pressure monitor, a heart rate monitor, a blood oxygen saturation monitor, a blood glucose monitor, etc.), one or more microphones, one or more speakers, a haptic engine, a heating and / or cooling unit, a skin shearing engine, one or more depth sensors (e.g., structured light, time-of-flight, LiDAR, etc.), a positioning and mapping engine, an eye-tracking engine, a body / head posture tracking engine, a hand / limb / finger / quadrant tracking engine, a camera posture tracking engine, etc.

[0058] In some embodiments, one or more displays 312 are configured to present an XR environment to a user. In some embodiments, one or more displays 312 are also configured to present planar video content to a user (e.g., two-dimensional or “planar” files such as AVI, FLV, WMV, MOV, MP4, etc., associated with a TV series or movie, or real-time video pass-through of the physical environment 105). In some embodiments, one or more displays 312 correspond to touchscreen displays. In some embodiments, one or more displays 312 correspond to holographic, digital light processing (DLP), liquid crystal display (LCD), liquid crystal on silicon (LCoS), organic light-emitting field-effect transistor (OLET), organic light-emitting diode (OLED), surface-conducting electron emitter display (SED), field emission display (FED), quantum dot light-emitting diode (QD-LED), microelectromechanical systems (MEMS), and / or similar display types. In some embodiments, one or more displays 312 correspond to waveguide displays such as diffraction, reflection, polarization, and holography. For example, electronic device 120 includes a single display. As another example, electronic device 120 includes displays for each of the user's eyes. In some implementations, one or more displays 312 are capable of displaying AR and VR content.

[0059] In some embodiments, the image capture device 370 corresponds to one or more RGB cameras (e.g., having a complementary metal-oxide-semiconductor (CMOS) image sensor or a charge-coupled device (CCD) image sensor), an IR image sensor, an event-based camera, etc. In some embodiments, the image capture device 370 includes a lens assembly, a photodiode, and a front-end architecture. In some embodiments, the image capture device 370 includes an externally oriented and / or internally oriented image sensor.

[0060] Memory 320 includes high-speed random access memory, such as DRAM, SRAM, DDR RAM, or other random access solid-state memory devices. In some embodiments, memory 320 includes non-volatile memory, such as one or more disk storage devices, optical disk storage devices, flash memory devices, or other non-volatile solid-state storage devices. Memory 320 optionally includes one or more storage devices located remotely from one or more processing units 302. Memory 320 includes a non-transitory computer-readable storage medium. In some embodiments, memory 320 or the non-transitory computer-readable storage medium of memory 320 stores programs, modules, and data structures, or subsets thereof, including optional operating system 330 and rendering engine 340.

[0061] Operating system 330 includes processes for handling various basic system services and for performing hardware-related tasks. In some implementations, rendering engine 340 is configured to present media items and / or XR content to a user via one or more displays 312. To this end, in various implementations, rendering engine 340 includes data acquirer 342, renderer 570, interactive handler 520, and data transfer device 350.

[0062] In some implementations, the data acquisition unit 342 is configured to acquire data (e.g., presentation data, such as rendered image frames associated with a user interface or XR environment, input data, user interaction data, head tracking information, camera pose tracking information, eye tracking information, hand / limb / finger / limb tracking information, sensor data, position data, etc.) from at least one of the I / O devices of the electronic device 120, sensor 306, processing device 110, and remote input device. To this end, in various implementations, the data acquisition unit 342 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0063] In some implementations, the interaction handler 520 is configured to detect user interactions with the presented A / V content and / or XR content (e.g., gesture input detected via hand / limb tracking, eye gaze input detected via eye tracking, voice commands, etc.). To this end, in various implementations, the interaction handler 520 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0064] In some implementations, renderer 570 is configured to render and update A / V content and / or XR content (e.g., rendered image frames associated with a user interface or XR environment 128, including VA, XR content, one or more UI elements associated with the XR content, etc.) via one or more displays 312. To this end, in various implementations, renderer 570 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0065] In some implementations, the data transmitter 350 is configured to transmit at least data to the processing device 110 (e.g., presentation data, location data, user interaction data, head tracking information, camera pose tracking information, eye tracking information, hand / limb / finger / limb tracking information, etc.). To this end, in various implementations, the data transmitter 350 includes instructions and / or logic for those instructions, as well as heuristics and metadata for those heuristics.

[0066] Although the data acquirer 342, the interaction handler 520, the presenter 570, and the data transmitter 350 are shown residing on a single device (e.g., electronic device 120), it should be understood that in other implementations, any combination of the data acquirer 342, the interaction handler 520, the presenter 570, and the data transmitter 350 may reside in a separate computing device.

[0067] also, Figure 3 It is used more as a functional description of various features present in a specific implementation than as a structural illustration of the specific implementation described herein. As those skilled in the art will recognize, items shown individually can be combined, and some items can be separated. For example, Figure 3 Some functional modules shown individually can be implemented in a single module, and the various functions of a single functional block can be implemented in various specific implementations through one or more functional blocks. The actual number of modules and the division of specific functions, as well as how features are allocated therein, will vary depending on the specific implementation, and in some specific implementations, it depends in part on the specific combination of hardware, software, and / or firmware selected for a particular implementation.

[0068] Figure 4 This is a block diagram of an exemplary controller 130 according to some specific implementation. The controller 130 is sometimes simply referred to as a stylus. The controller 130 includes a non-transitory memory 402 (which optionally includes one or more computer-readable storage media), a memory controller 422, one or more processing units (CPUs) 420, a peripheral interface 418, RF circuitry 408, an input / output (I / O) subsystem 406, and other input or control devices 416. The controller 130 optionally includes an external port 424 and one or more optical sensors 464. The controller 130 optionally includes one or more contact strength sensors 465 for detecting the strength of contact of the controller 130 on electronic device 100 (e.g., when the controller 130 is used with a touch-sensitive surface such as a display 122 of electronic device 120) or on other surfaces (e.g., a table surface). The controller 130 optionally includes one or more haptic output generators 463 for generating haptic output on the controller 130. These components optionally communicate via one or more communication buses or signal lines 403.

[0069] It should be understood that controller 130 is merely an example of an electronic stylus, and controller 130 may optionally have more or fewer components than those shown, may optionally combine two or more components, or may optionally have different configurations or arrangements of these components. Figure 4 The various components shown are implemented in hardware, software, firmware, or any combination thereof (including one or more signal processing circuits and / or application-specific integrated circuits).

[0070] like Figure 1 As shown, the controller 130 includes a first end 176 and a second end 177. In various embodiments, the first end 176 corresponds to the tip of the controller 130 (e.g., the tip of a pencil), and the second end 177 corresponds to the opposite end or bottom end of the controller 130 (e.g., the eraser of a pencil).

[0071] like Figure 1 As shown, controller 130 includes a touch-sensitive surface 175 to receive touch input from user 149. In some embodiments, touch-sensitive surface 175 corresponds to a capacitive touch element. Controller 130 includes sensors or a group of sensors that detect input from the user based on tactile and / or sensory contact with touch-sensitive surface 175. In some embodiments, controller 130 includes any of a variety of touch sensing technologies now known or to be developed hereafter, as well as other proximity sensor arrays or other elements for determining one or more points of contact with touch-sensitive surface 175, including but not limited to capacitive, resistive, infrared, and surface acoustic wave technologies. Because controller 130 includes a variety of sensors and sensor types, controller 130 can detect a variety of inputs from user 149. In some embodiments, one or more sensors may detect a single touch input or continuous touch input in response to a user tapping once or multiple times on touch-sensitive surface 175. In some embodiments, one or more sensors may detect a swipe input on controller 130 in response to a user swiping along touch-sensitive surface 175 with one or more fingers. In some implementations, if a user swipes along a touch-sensitive surface at a speed exceeding a threshold, one or more sensors detect a flick input instead of a swipe input.

[0072] The controller 130 also includes one or more sensors, such as one or more accelerometers 467, one or more gyroscopes 468, one or more magnetometers 469, etc., for detecting the orientation (e.g., angular position) and / or movement of the controller 130. These sensors may detect various rotational movements of the user on the controller 130, including the type and direction of rotation. For example, the sensors may detect the user rolling and / or rotating the controller 130, and may detect the direction of the roll / rotation (e.g., clockwise or counterclockwise). In some embodiments, the detected input depends on the angular position of the first end 176 and the second end 177 of the controller 130 relative to the electronic device 120, the world, a physical surface or object within the physical environment 105, a virtual surface or object within the XR environment 128, or the like. For example, in some embodiments, if the controller 130 is substantially perpendicular to the electronic device and the second end 177 (e.g., an eraser) is closer to the electronic device, contact between the surface of the electronic device and the second end 177 results in an erasing operation. On the other hand, if the controller 130 is substantially perpendicular to the electronic device and the first end 176 (e.g., the tip) is closer to the electronic device, then contact between the surface of the electronic device and the first end 176 results in a marking operation.

[0073] Memory 402 optionally includes high-speed random access memory and also optionally includes non-volatile memory, such as one or more flash memory devices or other non-volatile solid-state memory devices. Access to memory 402 by other components of controller 130 (such as CPU 420 and peripheral interface 418) is optionally controlled by memory controller 422.

[0074] Peripheral interface 418 can be used to couple input and output peripherals of the stylus to CPU 420 and memory 402. One or more processors 420 run or execute various software programs and / or instruction sets stored in memory 402 to perform various functions of controller 130 and process data. In some embodiments, peripheral interface 418, CPU 420, and memory controller 422 are optionally implemented on a single chip, such as chip 404. In some other embodiments, they are optionally implemented on separate chips.

[0075] RF (Radio Frequency) circuit 408 receives and transmits RF signals, also known as electromagnetic signals. RF circuit 408 converts electrical signals into electromagnetic signals / converts electromagnetic signals into electrical signals, and communicates with processing device 110, electronic device 120, communication networks, and / or other communication devices via electromagnetic signals. RF circuit 408 optionally includes well-known circuitry for performing these functions, including but not limited to antenna systems, RF transceivers, one or more amplifiers, tuners, one or more oscillators, digital signal processors, codec chipsets, Subscriber Identity Module (SIM) cards, memory, etc. RF circuit 408 optionally communicates wirelessly with networks and other devices, such as the Internet (also known as the World Wide Web (WWW)), intranets, and / or wireless networks (such as cellular telephone networks, wireless local area networks (LANs), and / or metropolitan area networks (MANs)). Wireless communication may optionally employ any of a variety of communication standards, protocols, and technologies, including but not limited to Global System for Mobile Communications (GSM), Enhanced Data GSM Environment (EDGE), High-Speed ​​Downlink Packet Access (HSDPA), High-Speed ​​Uplink Packet Access (HSUPA), Evolution, Data-Only (EV-DO), HSPA, HSPA+, Dual-Unit HSPA (DC-HSPA), Long Term Evolution (LTE), Near Field Communication (NFC), Wideband Code Division Multiple Access (W-CDMA), Code Division Multiple Access (CDMA), Time Division Multiple Access (TDMA), Bluetooth, Wi-Fi (e.g., IEEE 802.11a, IEEE 802.11ac, IEEE 802.11ax, IEEE 802.11b, IEEE 802.11g, and / or IEEE 802.11n), or any other suitable communication protocol, including those not yet developed as of the date of this filing.

[0076] I / O subsystem 406 couples input / output peripherals such as other input or control devices 416 on controller 130 to peripheral interface 418. I / O subsystem 406 optionally includes an optical sensor controller 458, an intensity sensor controller 459, a haptic feedback controller 461, and one or more input controllers 460 for other input or control devices. One or more input controllers 460 receive electrical signals from / send electrical signals to the other input or control device 416. Other input or control devices 416 optionally include physical buttons (e.g., push-buttons, rocker buttons, etc.), dial pads, slide switches, click wheels, etc. In some alternative embodiments, one or more input controllers 460 are optionally coupled to (or not coupled to) any of the following: an infrared port and / or a USB port.

[0077] The controller 130 also includes a power system 462 for supplying power to various components. The power system 462 optionally includes a power management system, one or more power sources (e.g., batteries, alternating current (AC)), a recharging system, power fault detection circuitry, a power converter or inverter, a power status indicator (e.g., light-emitting diodes (LEDs)), and any other components associated with the generation, management, and distribution of power in portable devices and / or portable accessories.

[0078] The controller 130 may optionally also include one or more optical sensors 464. Figure 4 An optical sensor coupled to an optical sensor controller 458 in an I / O subsystem 406 is shown. One or more optical sensors 464 optionally include charge-coupled devices (CCDs) or complementary metal-oxide-semiconductor (CMOS) phototransistors. The one or more optical sensors 464 receive light projected from the environment through one or more lenses and convert the light into data representing an image.

[0079] The controller 130 may optionally also include one or more contact strength sensors 465. Figure 4 A contact strength sensor coupled to a strength sensor controller 459 in I / O subsystem 406 is shown. The contact strength sensor 465 optionally includes one or more piezoresistive strain gauges, capacitive force sensors, electro-force sensors, piezoelectric sensors, optical force sensors, capacitive touch-sensitive surfaces, or other strength sensors (e.g., sensors for measuring force (or pressure) relative to a surface or relative to a user's grip 149). The contact strength sensor 465 receives contact strength information (e.g., pressure information or substitutes for pressure information) from the environment. In some embodiments, at least one contact strength sensor is arranged juxtaposed or adjacent to the tip of controller 130. In some embodiments, at least one contact strength sensor is arranged juxtaposed or adjacent to the body of controller 130.

[0080] The controller 130 may optionally also include one or more proximity sensors 466. Figure 4 One or more proximity sensors 466 coupled to peripheral device interface 418 are shown. Alternatively, one or more proximity sensors 466 are coupled to input controller 460 in I / O subsystem 406. In some embodiments, one or more proximity sensors 466 determine the proximity of controller 130 to electronic device (e.g., electronic device 120).

[0081] The controller 130 may optionally also include one or more haptic output generators 463. Figure 4A haptic output generator coupled to a haptic feedback controller 461 in an I / O subsystem 406 is shown. One or more haptic output generators 463 optionally include one or more electroacoustic devices (such as speakers or other audio components), and / or electromechanical devices that convert energy into linear motion (such as motors, solenoids, electroactive polymerizers, piezoelectric actuators, electrostatic actuators), or other haptic output generating components (e.g., components that convert electrical signals into haptic outputs on electronic devices). One or more haptic output generators 463 receive haptic feedback generation instructions from a haptic feedback module 433 and generate a haptic output on the controller 130 that can be felt by a user of the controller 130. In some embodiments, at least one haptic output generator is juxtaposed or adjacent to the length (e.g., body or housing) of the controller 130, and optionally, the haptic output is generated by moving the controller 130 vertically (e.g., in a direction parallel to the length of the controller 130) or laterally (e.g., in a direction perpendicular to the length of the controller 130).

[0082] The controller 130 may optionally also include one or more accelerometers 467, one or more gyroscopes 468 and / or one or more magnetometers 469 (e.g., as part of an inertial measurement unit (IMU)) for acquiring information about the position and position state of the controller 130. Figure 4 Sensors 467, 468, and 469, coupled to peripheral interface 418, are shown. Alternatively, sensors 467, 468, and 469 may be coupled to input controller 460 in I / O subsystem 406. Controller 130 may optionally include a GPS (or GLONASS or other global navigation system) receiver (not shown) for acquiring information about the location of controller 130.

[0083] Controller 130 includes a touch-sensitive system 432. Touch-sensitive system 432 detects input received at touch-sensitive surface 175. These inputs include those discussed herein with respect to touch-sensitive surface 175 of controller 130. For example, touch-sensitive system 432 can detect tap input, spin input, scroll input, flick input, swipe input, etc. Touch-sensitive system 432 cooperates with touch interpretation module 477 to interpret specific types of touch input (e.g., spin / scroll / flick / swipe / etc.) received at touch-sensitive surface 175.

[0084] In some implementations, the software components stored in memory 402 include an operating system 426, a communication module (or instruction set) 428, a contact / motion module (or instruction set) 430, a position module (or instruction set) 431, and a Global Positioning System (GPS) module (or instruction set) 435. Furthermore, in some implementations, memory 402 stores device / global internal state 457, such as... Figure 4 As shown. Furthermore, memory 402 includes a touch interpretation module 477. Device / global internal state 457 includes one or more of the following: sensor state, including information obtained from various sensors of the stylus and other input or control devices 416; position state, including information about the position and / or orientation of controller 130 (e.g., translation and / or rotation values) and position information about the location of controller 130 (e.g., determined by GPS module 435).

[0085] Operating system 426 (e.g., iOS, Darwin, RTXC, LINUX, UNIX, OS X, WINDOWS, or embedded operating systems such as VxWorks) includes various software components and / or drivers for controlling and managing general system tasks (e.g., memory management, power management, etc.) and facilitates communication between various hardware and software components. Communication module 428 optionally facilitates communication with other devices via one or more external ports 424 and also includes various software components for processing data received by RF circuitry 408 and / or external ports 424. External port 424 (e.g., Universal Serial Bus (USB), FireWire, etc.) is adapted to be directly coupled to other devices or indirectly coupled via a network (e.g., Internet, Wireless LAN, etc.). In some implementations, some functions or operations of controller 130 are provided by processing device 110 and / or electronic device 120.

[0086] The contact / motion module 430 optionally detects contact with the controller 130 and other touch-sensitive devices of the controller 130 (e.g., buttons or other touch-sensitive components of the controller 130). The contact / motion module 430 includes software components for performing various operations related to the detection of contact (e.g., detection of contact between the tip of a stylus and a touch-sensitive display such as the display 122 of electronic device 120 or with another surface such as a table surface). These various operations include determining whether a contact has occurred (e.g., detecting a touch press event), determining the intensity of the contact (e.g., the force or pressure of the contact, or a substitute for force or pressure), determining whether there is movement of the contact and tracking that movement (e.g., across the display 122 of electronic device 120), and determining whether the contact has stopped (e.g., detecting a lift-off event or a contact interruption). In some specific implementations, the contact / motion module 430 receives contact data from the I / O subsystem 406. Determining the movement of the contact point optionally includes determining the rate (magnitude), velocity (magnitude and direction), and / or acceleration (change in magnitude and / or direction) of the contact point, the movement of which is represented by a series of contact data. As described above, in some specific embodiments, one or more of these operations related to the detection of contact are performed by electronic device 120 or processing device 110 (as a supplement to or alternative to a stylus using contact / motion module 430).

[0087] The touch / motion module 430 optionally detects gesture inputs made by the controller 130. Different gestures made by the controller 130 have different contact patterns (e.g., different movements, timings, and / or intensities of the detected contact). Therefore, gestures are optionally detected by detecting specific contact patterns. For example, detecting a single tap gesture includes detecting a touch press event, followed by detecting a lift-off event at the same (or substantially the same) location as the touch press event (e.g., at the location of an icon). As another example, detecting a swipe gesture includes detecting a touch press event, followed by detecting one or more stylus drag events, and then detecting a lift-off event. As described above, in some specific implementations, gesture detection is performed by an electronic device using the touch / motion module 430 (as a supplement to or alternative to a stylus using the touch / motion module 430).

[0088] In conjunction with one or more accelerometers 467, one or more gyroscopes 468, and / or one or more magnetometers 469, the position module 431 optionally detects positional information about the stylus, such as the attitude of the controller 130 in a particular frame of reference (e.g., roll, pitch, and / or yaw). In conjunction with one or more accelerometers 467, one or more gyroscopes 468, and / or one or more magnetometers 469, the position module 431 optionally detects movement gestures, such as flicks, taps, and rotations of the controller 130. The position module 431 includes software components for performing various operations related to detecting the position of the stylus and detecting changes in the position of the stylus in a particular frame of reference. In some specific implementations, the position module 431 detects the positional state of the controller 130 relative to the physical environment 105 or the world at large, and detects changes in the positional state of the controller 130.

[0089] The haptic feedback module 433 includes various software components for generating instructions that are used by one or more haptic output generators 463 to produce haptic output at one or more locations on the controller 130 in response to user interaction with the controller 130. The GPS module 435 determines the location of the controller 130 and provides that information for use in various applications (e.g., applications that provide location-based services, such as applications for locating lost devices and / or accessories).

[0090] Touch interpretation module 477 cooperates with touch-sensitive system 432 to determine (e.g., interpret or recognize) the type of touch input received at touch-sensitive surface 175 of controller 130. For example, if a user swipes a sufficient distance on touch-sensitive surface 175 of controller 130 for a sufficiently short period of time, touch interpretation module 477 determines that the touch input corresponds to a swipe input (instead of a tap). Similarly, if the user swipes on touch-sensitive surface 175 of controller 130 quickly enough to correspond to a voice input that corresponds to a swipe input, touch interpretation module 477 determines that the touch input corresponds to a flick input (instead of a swipe). The threshold speed of the swipe can be preset and varied. In various embodiments, the pressure and / or force of the touch received at the touch-sensitive surface determines the type of input. For example, a light touch may correspond to a first type of input, while a stronger touch may correspond to a second type of input. In some specific embodiments, the functionality of touch interpretation module 477 is provided by processing device 110 and / or electronic device 120, such as touch input detection via data from a magnetic sensor or computer vision technology.

[0091] Each module and application identified above corresponds to a set of executable instructions for performing one or more of the functions described above, as well as the methods described in this application (e.g., computer-implemented methods and other information processing methods described herein). These modules (i.e., instruction sets) need not be implemented as separate software programs, processes, or modules; therefore, various subsets of these modules may optionally be combined or otherwise rearranged in various embodiments. In some specific embodiments, memory 402 optionally stores a subset of the modules and data structures described above. Furthermore, memory 402 optionally stores additional modules and data structures not described above.

[0092] Figure 5A This is a block diagram of the first part 500A of an exemplary content delivery architecture according to some specific implementations. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for brevity and to avoid obscuring further relevant aspects of the exemplary specific implementations disclosed herein. Therefore, as a non-limiting example, the content delivery architecture includes, in computing systems, such as… Figure 1 and Figure 2 The processing device 110 shown; Figure 1 and Figure 3 The electronic devices 120 shown; and / or suitable combinations thereof.

[0093] like Figure 5A As shown, one or more local sensors 502 of processing device 110, electronic device 120, and / or combinations thereof acquire local sensor data 503 associated with physical environment 105. For example, local sensor data 503 includes images or streams of physical environment 105, Simultaneous Localization and Mapping (SLAM) information of physical environment 105, the position of electronic device 120 or user 149 relative to physical environment 105, ambient lighting information of physical environment 105, ambient audio information of physical environment 105, acoustic information of physical environment 105, dimensional information of physical environment 105, semantic labels of objects within physical environment 105, etc. In some specific implementations, local sensor data 503 includes unprocessed or post-processed information.

[0094] Similarly, such as Figure 5AAs shown, one or more remote sensors 504 associated with an optional remote input device within physical environment 105 acquire remote sensor data 505 associated with physical environment 105. For example, remote sensor data 505 includes an image or stream of physical environment 105, SLAM information of physical environment 105, the position of electronic device 120 or user 149 relative to physical environment 105, ambient lighting information of physical environment 105, ambient audio information of physical environment 105, acoustic information of physical environment 105, dimensional information of physical environment 105, semantic tags of objects within physical environment 105, etc. In some embodiments, remote sensor data 505 includes unprocessed or post-processed information.

[0095] like Figure 5A As shown, tracking data 506 is obtained by at least one of processing device 110, electronic device 120, or controller 130 in order to locate and track controller 130. As an example, tracking data 506 includes images or streams of the physical environment 105 captured by an externally facing image sensor of electronic device 120, which includes controller 130. As another example, tracking data 506 corresponds to IMU information, accelerometer information, gyroscope information, magnetometer information, etc., from integrated sensors of controller 130.

[0096] According to some implementations, privacy architecture 508 acquires local sensor data 503, remote sensor data 505, and tracking data 506. In some implementations, privacy architecture 508 includes one or more privacy filters associated with user information and / or identification information. In some implementations, privacy architecture 508 includes an opt-in feature where electronic device 120 notifies user 149 which user information and / or identification information is being monitored and how such user information and / or identification information will be used. In some implementations, privacy architecture 508 selectively prevents and / or restricts content delivery architecture 500A / 500B or portions thereof from acquiring and / or transmitting user information. To this end, privacy architecture 508 receives user preferences and / or choices from user 149 in response to prompting user 149 to make user preferences and / or choices. In some implementations, privacy architecture 508 prevents content delivery architecture 500A / 500B from acquiring and / or transmitting user information unless and until privacy architecture 508 obtains informed consent from user 149. In some specific implementations, privacy architecture 508 anonymizes (e.g., scrambles, obfuscates, encrypts, etc.) certain types of user information. For example, privacy architecture 508 receives user input specifying which types of user information it will anonymize. As another example, privacy architecture 508 may anonymize certain types of user information independently of user specification (e.g., automatically) including sensitive and / or identifying information.

[0097] In some specific implementations, the object tracking engine 510 obtains the tracking data 506 after it has been processed by the privacy architecture 508. In some specific implementations, the object tracking engine 510 determines / generates an object tracking vector 511 based on the tracking data 506, and updates the object tracking vector 511 over time.

[0098] Figure 5B An exemplary data structure for object tracking vector 511, according to some specific implementation, is shown. For example... Figure 5B As shown, the object tracking vector 511 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 571 (e.g., the time when the object tracking vector 511 was last updated), one or more translation values ​​572 of the physical object (e.g., x, y, and z values ​​relative to the physical environment 105, the entire world, etc.), one or more rotation values ​​574 of the physical object (e.g., roll, pitch, and yaw values), one or more pressure values ​​576 associated with the physical object (e.g., a first pressure value associated with contact between the end and surface of the controller 130, a second pressure value associated with the amount of pressure applied to the body of the controller 130 when gripped by the user 149, etc.), optional touch input information 578 (e.g., information associated with user touch input pointing towards the touch-sensitive surface 175 of the controller 130), and / or miscellaneous information 579. Those skilled in the art will understand that... Figure 5B The data structure of the object tracking vector 511 in the example is merely one example, and it can include different information parts in various other implementations and can be constructed in various other implementations in a variety of ways.

[0099] In some specific implementations, the eye-tracking engine 512 acquires local sensor data 503 and remote sensor data 505 after processing by the privacy architecture 508. In some specific implementations, the eye-tracking engine 512 determines / generates an eye-tracking vector 513 based on the input data and updates the eye-tracking vector 513 over time.

[0100] Figure 5B An exemplary data structure for eye-tracking vector 513 according to some specific implementations is shown. Figure 5B As shown, the eye-tracking vector 513 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 581 (e.g., the time when the eye-tracking vector 513 was last updated), one or more angle values ​​582 for the current gaze direction (e.g., roll, pitch, and yaw values), one or more translation values ​​584 for the current gaze direction (e.g., x, y, and z values ​​relative to the physical environment 105, the entire world, etc.), and / or miscellaneous information 586. Those skilled in the art will understand that... Figure 5BThe data structure of the eye tracking vector 513 in the example is merely one example, and it can include different information parts in various other implementations and can be constructed in various other implementations in a variety of ways.

[0101] For example, gaze direction indicates a point in the physical environment 105 that user 149 is currently viewing (e.g., associated with x, y, and z coordinates relative to physical environment 105 or the world), a physical object, or a region of interest (ROI). As another example, gaze direction indicates a point in the XR environment 128 that user 149 is currently viewing (e.g., associated with x, y, and z coordinates relative to XR environment 128), an XR object, or a region of interest (ROI).

[0102] In some specific implementations, the body / head pose tracking engine 514 acquires local sensor data 503 and remote sensor data 505 after processing by the privacy architecture 508. In some specific implementations, the body / head pose tracking engine 514 determines / generates a pose representation vector 515 based on the input data and updates the pose representation vector 515 over time.

[0103] Figure 5B An exemplary data structure for the pose representation vector 515, according to some specific implementations, is shown. Figure 5B As shown, the pose representation vector 515 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 591 (e.g., the time when the pose representation vector 515 was last updated), a head pose descriptor 592A (e.g., up, down, neutral, etc.), a translation value 592B for the head pose, a rotation value 592C for the head pose, a body pose descriptor 594A (e.g., standing, sitting, prone, etc.), a translation value 594B for body parts / limbs / joints, a rotation value 594C for body parts / limbs / joints, and / or miscellaneous information 596. In some specific implementations, the pose representation vector 515 may also include information associated with finger / hand / limb tracking. Those skilled in the art will understand that... Figure 5B The data structure of the pose representation vector 515 in the example is merely one example; it can include different information components in various other implementations and can be constructed in various ways in various other implementations. According to some implementations, the eye tracking vector 513 and the pose representation vector 515 are collectively referred to as the user input vector 519.

[0104] In some specific implementations, the representation engine 542 acquires object tracking vector 511, eye tracking vector 513, and pose representation vector 515. In some specific implementations, the representation engine 542 determines / generates the physical object's representation vector 543 based on the object tracking vector 511, eye tracking vector 513, and pose representation vector 515.

[0105] Figure 5B Exemplary data structures for representing vector 543 are shown according to some specific implementations. For example... Figure 5B As shown, representation vector 543 may correspond to an N-tuple representation vector or representation tensor, which includes a timestamp 5101 (e.g., the time when representation vector 543 was last updated), a user grip value 5102 associated with the physical object, and a pose value 5104 associated with the physical object. In some embodiments, the user grip value 5102 indicates how the user 149 grips the physical object. For example, the user grip value 5102 corresponds to one of a remote control grip, a wand grip, a writing grip, a reverse writing grip, a handle grip, a thumb grip, a horizontal grip, a game controller grip, a flute grip, a lighter grip, etc. In some embodiments, the pose value 5104 indicates the orientation or position of the physical object relative to the user 149, a physical surface, or another object detectable by computer vision (CV). For example, the pose value 5104 corresponds to one of a neutral pose, a conductor / wand pose, a writing pose, a surface pose, a near-mouth pose, an aiming pose, etc.

[0106] As discussed above, in one example, the physical object corresponds to controller 130, which includes integrated sensors and is communicatively coupled to processing device 110. In this example, controller 130 corresponds to a stylus, a wearable finger device, a handheld device, etc. In another example, the physical object corresponds to a proxy object that is not communicatively coupled to processing device 110.

[0107] According to some specific implementations, the operation mode selector 544 obtains a representation vector 543 of the physical object and selects the current operation mode 545 of the physical object (when interacting with the XR environment 128) based on the representation vector 543. As an example, when the pose value corresponds to a neutral pose and the user grip value corresponds to a writing tool-like grip, the selected operation mode 545 may correspond to a marking mode. As another example, when the pose value corresponds to a neutral pose and the user grip value corresponds to a reverse writing tool-like grip, the selected operation mode 545 may correspond to an erasing mode. As yet another example, when the pose value corresponds to a near-mouth pose and the user grip value corresponds to a wand-like grip, the selected operation mode 545 may correspond to a dictation mode. As yet another example, when the pose value corresponds to a neutral pose and the user grip value corresponds to a wand-like grip, the selected operation mode 545 may correspond to a manipulation mode. As yet another example, when the pose value corresponds to an aiming pose and the user grip value corresponds to a wand-like grip, the selected operation mode 545 may correspond to a pointing mode.

[0108] Figure 5C This is a block diagram of the second part 500B of an exemplary content delivery architecture according to some specific implementations. Although relevant features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for brevity and to avoid obscuring further relevant aspects of the exemplary specific implementations disclosed herein. Therefore, as a non-limiting example, the content delivery architecture includes, in computing systems, such as… Figure 1 and Figure 2 The processing device 110 shown; Figure 1 and Figure 3 The electronic devices 120 shown; and / or suitable combinations thereof. Figure 5C Similar to and adapted from Figure 5A .therefore, Figure 5A and Figure 5C Similar reference numerals are used in [the document / section]. Therefore, for the sake of brevity, the following text will only describe [the specific reference numerals]. Figure 5A and Figure 5C The differences between them.

[0109] According to some specific implementations, the interaction handler 520 acquires (e.g., receives, retrieves, or detects) one or more user inputs 521 provided by the user 149, which are associated with selecting A / V content, one or more VA and / or XR content for presentation. For example, one or more user inputs 521 correspond to gesture input selecting XR content from a UI menu detected via hand / limb tracking, eye gaze input selecting XR content from a UI menu detected via eye tracking, voice command selecting XR content from a UI menu detected via a microphone, etc. In some specific implementations, the content selector 522 selects XR content 527 from the content library 525 based on one or more user inputs 521 (e.g., voice commands, selections from a menu of XR content items, etc.).

[0110] In various specific implementations, the content manager 530 manages and updates the layout, settings, and structure of the XR environment 128 based on object tracking vector 511, user input vector 519, selected operation mode 545, user input 521, etc. This XR environment includes one or more of the following: VA, XR content, and one or more UI elements associated with the XR content. To this end, the content manager 530 includes a buffer 534, a content updater 536, and a feedback engine 538.

[0111] In some implementations, buffer 534 includes XR content, rendered image frames, etc., for one or more past instances and / or frames. In some implementations, content updater 536 modifies XR environment 128 over time based on object tracking vector 511, user input vector 519, selected operating mode 545, user input 521 associated with modifying and / or manipulating XR content or VA, translation or rotation of objects within physical environment 105, translation or rotation of electronic device 120 (or user 149), etc. In some implementations, feedback engine 538 generates sensory feedback (e.g., visual feedback such as text or lighting changes), audio feedback, haptic feedback, etc.) associated with XR environment 128.

[0112] Based on some specific implementations, refer to Figure 5C The rendering engine 550 and pose determiner 552 determine the current camera pose of the electronic device 120 and / or user 149 relative to the XR environment 128 and / or physical environment 105, at least in part, based on the pose representation vector 515. In some implementations, renderer 554 renders the VA, XR content 527, one or more UI elements associated with the XR content, etc., according to the current camera pose relative to it.

[0113] According to some embodiments, optional image processing architecture 562 acquires an image stream from image capture device 370, which includes one or more images of the physical environment 105 from the current camera pose of electronic device 120 and / or user 149. In some embodiments, image processing architecture 562 also performs one or more image processing operations on the image stream, such as distortion, color correction, gamma correction, sharpening, noise reduction, white balance, etc. In some embodiments, optional compositor 564 composites the rendered XR content with the processed image stream from the physical environment 105 of image processing architecture 562 to produce rendered image frames of XR environment 128. In various embodiments, renderer 570 renders the rendered image frames of XR environment 128 to user 149 via one or more displays 312. Those skilled in the art will understand that optional image processing architecture 562 and optional compositor 564 may not be suitable for fully virtual environments (or optical pass-through scenes).

[0114] Figures 6A to 6N Sequences of instances 610 to 6140 according to some specific implementations of content delivery scenarios are shown. Although some specific features are shown, those skilled in the art will recognize from this disclosure that various other features are not shown for the sake of brevity and in order not to obscure further relevant aspects of the specific implementations disclosed herein. Therefore, as a non-limiting example, the sequence of instances 610 to 6140 is rendered and presented by a computing system such as... Figure 1 and Figure 2 The processing device 110 shown; Figure 1 and Figure 3 The electronic devices 120 shown; and / or suitable combinations thereof.

[0115] like Figures 6A to 6N As shown, the content delivery scenario includes a physical environment 105 and an XR environment 128 displayed on a display 122 of electronic device 120 (e.g., associated with user 149). When user 149 is physically present within physical environment 105, electronic device 120 presents XR environment 128 to user 149, which includes a door 115 located within the FOV 111 of the externally facing image sensor of electronic device 120. Therefore, in some specific implementations, user 149 holds electronic device 120 in his / her left hand 150, similar to... Figure 1 The operating environment is 100.

[0116] In other words, in some implementations, electronic device 120 is configured to present XR content and enable optical or video pass-through (e.g., door 115) of at least a portion of the physical environment 105 on display 122. For example, electronic device 120 corresponds to mobile phones, tablets, laptops, near-eye systems, wearable computing devices, etc.

[0117] like Figure 6A As shown, during instance 610 of a content delivery scenario (e.g., associated with time T1), electronic device 120 presents an XR environment 128, which includes a virtual agent (VA) 606 and a representation 116 of a gate 115 within a physical environment 105. Figure 6A As shown, the XR environment 128 also includes a visual indicator 612 of instance 610 having a current pose value (e.g., no current pose value because controller 130 is not currently being gripped by user 149) and a current user grip value (e.g., no current user grip value because controller 130 is not currently being gripped by user 149). Those skilled in the art will understand that the visual indicator 612 is merely an exemplary visualization with pose and user grip values, which can be modified or replaced in various other specific implementations. As an example, the visual indicator 612 could be a badge, icon, audio output, etc., instead of a text notification.

[0118] like Figure 6B As shown, during instance 620 of the content delivery scenario (e.g., associated with time T2), electronic device 120 presents XR environment 128, which includes VA 606, visual indicator 622 of instance 620 with current posture value (e.g., near mouth) and current user grip value (e.g., wand), and representation 131 of controller 130 held by representation 153 of user 149's right hand 152 (e.g., video passthrough of physical environment 105). Figure 6B An illustration of the current body posture 625 of user 149 in instance 620, where user 149 is grasping controller 130 near his / her mouth with his / her right hand 152.

[0119] In response to determining that the current posture value corresponds to "near mouth" and the current user grip value corresponds to "wand," the electronic device 120 selects dictation mode as the current operating mode of the controller 130 and also displays a notification 624 indicating "dictation mode activated." Thus, when the current operating mode of the controller 130 corresponds to dictation mode, the user 149 can indicate what information, notes, content, etc., should be displayed within the XR environment 128. Those skilled in the art will understand that the recording notification 624 is merely an exemplary visualization and can be modified or replaced in various other specific implementations.

[0120] like Figure 6C As shown, during instance 630 of the content delivery scenario (e.g., associated with time T3), electronic device 120 presents XR environment 128, which includes VA 606, visual indicator 632 of instance 630 with current posture value (e.g., neutral) and current user grip value (e.g., writing), and representation 131 of controller 130 held by representation 153 of user 149's right hand 152 (e.g., video passthrough of physical environment 105). Figures 6C to 6E Illustrations of the current body posture 635 of user 149, including instances 630 to 650, in which user 149 is gripping controller 130 with his / her right hand 152 in order to write / draw / paint / mark in space (e.g., three-dimensional space).

[0121] In response to determining that the current posture value corresponds to "neutral" and the current user grip value corresponds to "writing," the electronic device 120 selects marking mode as the current operating mode of the controller 130 and also displays a notification 634 indicating "marking mode is activated." Figure 6C As shown, relative to gravity, the first end 176 points downwards and the second end 177 points upwards. Thus, when the current operating mode of the controller 130 corresponds to the marking mode, the user 149 can write / draw / paint / mark in three-dimensional space, and the electronic device 120 displays the corresponding mark within the XR environment 128.

[0122] like Figure 6D As shown, during instance 640 of the content delivery scenario (e.g., associated with time T4), electronic device 120 detects a marker input 644 in space using controller 130 (e.g., based on tracking data 506 from controller 130, computer vision, etc.). Figure 6E As shown, during instance 650 of the content delivery scenario (e.g., associated with time T5), electronic device 120 responds to the detection Figure 6E The marker input 644 is displayed as marker 654 within the XR environment 128. For example, marker 654 corresponds to the shape, displacement, etc. of marker input 644. In another example, marker 654 corresponds to a function of the shape, displacement, etc. of marker input 644 and weighting coefficients (e.g., magnification, attenuation, etc.).

[0123] like Figure 6F As shown, during instance 660 of the content delivery scenario (e.g., associated with time T6), the electronic device 120 presents an XR environment 128, which includes a VA 606 and a visual indicator 662 of instance 660 having current pose values ​​(e.g., surface) and current user grip values ​​(e.g., reverse writing). Figures 6F to 6HThe illustration includes instances 660 to 680, showing the current body posture 665 of user 149, where user 149 is gripping controller 130 with his / her right hand 152 to write / draw / mark on display 122 of electronic device 120.

[0124] In response to determining that the current posture value corresponds to "surface" and the current user grip value corresponds to "reverse writing," the electronic device 120 selects the erase mode as the current operating mode of the controller 130 and also displays a notification 664 indicating "erasure mode activated." Figure 6F As shown, relative to gravity, the second end 177 points downwards and the first end 176 points upwards. Thus, when the current operating mode of the controller 130 corresponds to the erase mode, the user 149 can remove / erase pixel markers from the XR environment 128.

[0125] like Figure 6G As shown, during instance 670 of the content delivery scenario (e.g., associated with time T7), electronic device 120 detects an erase input 674 made by controller 130 on display 122 (e.g., touch-sensitive surface, etc.). Figure 6H As shown, during instance 680 of the content delivery scenario (e.g., associated with time T8), electronic device 120 responds to the detection Figure 6G The erase input 674 is used to remove the marker 654 within the XR environment 128.

[0126] like Figure 6I As shown, during instance 690 of the content delivery scenario (e.g., associated with time T9), electronic device 120 presents XR environment 128, which includes VA 606, visual indicator 692 of instance 690 with current posture value (e.g., aiming) and current user grip value (e.g., wand), and representation 696 of physical object 695 held by representation 153 of user 149's right hand 152 (e.g., video passthrough of physical environment 105). Figure 6I and Figure 6J An illustration of the current body posture 697 of user 149, including instances 690 and 6100, in which user 149 is grasping a physical object 695 (e.g., a ruler, a stick, a controller 130, etc.) with his / her right hand 152 in order to point in space.

[0127] In response to determining that the current posture value corresponds to "aiming" and the current user grip value corresponds to "wank," the electronic device 120 selects the pointing mode as the current operating mode of the physical object 695 and also displays a notification 694 indicating "pointing mode activated." Thus, when the current operating mode of the physical object 695 corresponds to the pointing mode, the user 149 can use the physical object 695 as a laser pointing type device within the XR environment 128. Figure 6I and Figure 6J As shown, physical object 695's representation 696 points to VA 606 within XR environment 128.

[0128] like Figure 6J As shown, during instance 6100 of the content delivery scenario (e.g., with time T) 10 (Associated), electronic device 120 responds to determining the representation 696 of physical object 695 pointing to Figure 6I VA 606 is located within the XR environment 128, and a crosshair 6102 (e.g., a focus selector) is presented alongside VA 606 within the XR environment 128.

[0129] like Figure 6K As shown, during instance 6110 of the content delivery scenario (e.g., with time T) 11 (Associated), the electronic device 120 presents an XR environment 128, which includes a VA 606, a first-view perspective 6116A of the XR content (e.g., the front of a cube, box, etc.), a visual indicator 6112 of an instance 6110 having a current pose value (e.g., neutral) and a current user grip value (e.g., a magic wand), and a representation 696 of a physical object 695 grasped by the user 149's right hand 152 (e.g., video passthrough of the physical environment 105). Figures 6K to 6M An illustration of the current body posture 6115 of user 149, including instances 6110 to 6130, in which user 149 is grasping a physical object 695 (e.g., a ruler, a stick, a controller 130, etc.) with his / her right hand 152 in order to manipulate XR content.

[0130] In response to determining that the current posture value corresponds to "neutral" and the current user grip value corresponds to "magic wand", the electronic device 120 selects a manipulation mode as the current operating mode of the physical object 695 and also displays a notification 6114 indicating that "manipulation mode is activated". Thus, when the current operating mode of the physical object 695 corresponds to the manipulation mode, the user 149 can use the physical object 695 to manipulate (e.g., translate, rotate, etc.) XR content within the XR environment 128.

[0131] like Figure 6L As shown, during instance 6120 of the content delivery scenario (e.g., with time T)12 (Associated), electronic device 120 detects manipulation input 6122 performed in space using physical object 695 (e.g., based on computer vision). For example, manipulation input 6122 corresponds to a 180° clockwise rotation input. Figure 6M As shown, during instance 6130 of the content delivery scenario (e.g., with time T) 13 (Associated), electronic device 120 responds to detecting Figure 6L The manipulation input 6122 rotates the XR content 180° clockwise to show a second view 6116B of the XR content (e.g., the back of a cube, box, etc.).

[0132] like Figure 6N As shown, during instance 6140 of the content delivery scenario (e.g., with time T) 14 (Associated), the electronic device 120 presents an XR environment 128, which includes a VA 606, a visual indicator 6142 of an instance 6140 having a current posture value (e.g., near mouth) and a current user grip value (e.g., wand), and a representation 696 of a physical object 695 held by a representation 153 of the user 149's right hand 152 (e.g., video passthrough of physical environment 105). Figures 6N to 6R Illustrations of the current body posture 6145 of user 149 in instances 6140 to 6180, where user 149 is grasping a physical object 695 (e.g., a ruler, stick, controller 130, etc.) near his / her mouth with his / her right hand 152.

[0133] In response to determining that the current posture value corresponds to "near mouth" and the current user grip value corresponds to "magic wand," and in response to detecting voice input 6146 from user 149 containing keywords or key phrases (e.g., "Hey, digital assistant," etc.), electronic device 120 selects digital assistant mode as the current operating mode of physical object 695, and also displays a notification 6144 indicating "digital assistant mode activated." Thus, when the current operating mode of physical object 695 corresponds to digital assistant mode, user 149 can provide audible commands, search strings, etc., for the digital assistant program to execute.

[0134] like Figure 6O As shown, during instance 6150 of the content delivery scenario (e.g., with time T) 15 (Associated), electronic device 120 displays a digital assistant (DA) indicator 6152A (e.g., icon, badge, image, text box, notification, etc.) within XR environment 128. Figure 6O As shown, the XR environment 128 also includes a visualization 6154 of the user 149's gaze direction, which is currently pointing to... Figure 6OThe DA indicator 6152A is shown in the image. Those skilled in the art will understand that the DA indicator 6152A is merely an exemplary visualization and can be modified or replaced in various other embodiments. Those skilled in the art will also understand that in various other embodiments, the visualization 6154 indicating the gaze direction may not be displayed.

[0135] like Figure 6P As shown, during instance 6160 of the content delivery scenario (e.g., with time T) 16 (Associatedly), electronic device 120 provides audio output 6162 (e.g., prompting user 149 to provide a search string or voice command) and in response to detecting that user 149's gaze has been directed toward DA indicator 6152A for at least Figure 6O The predefined time value is displayed as the DA indicator 6152B within the XR environment 128. For example... Figure 6P As shown, the DA indicator 6152B is compared with... Figure 6O The DA indicator 6152A in the image is associated with a large size. In some specific implementations, the electronic device 120 changes the appearance of the DA indicator (e.g., as shown) in response to detecting that the user 149's gaze has been directed at the DA indicator for at least a predefined amount of time. Figure 6O and Figure 6P (as shown in the image) to indicate that the DA is ready to receive search strings or voice commands.

[0136] like Figure 6Q As shown, during instance 6170 of the content delivery scenario (e.g., with time T) 17 (Associated), electronic device 120 detects voice input 6172 (e.g., a search string) from user 149. Figure 6R As shown, during instance 6180 of the content delivery scenario (e.g., with time T) 18 (Associated), the electronic device 120 presents a corresponding [symbol / relationship] within the XR environment. Figure 6Q The search results 6182 are associated with the search string detected by the voice input 6172. For example, the search results 6182 include text, audio content, video content, 3D content, XR content, etc.

[0137] Figure 7 This is a flowchart representation of a method 700 for dynamically selecting the operating mode of a physical object according to some specific implementation. In various specific implementations, method 700 is executed at a computing system including non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device and (optionally) one or more input devices (e.g., ...). Figure 1 and Figure 3 The electronic device 120 shown; Figure 1 and Figure 2The processing device 110; or a suitable combination thereof. In some embodiments, method 700 is performed by processing logic components (including hardware, firmware, software, or combinations thereof). In some embodiments, method 700 is performed by a processor executing code stored in a non-transitory computer-readable medium (e.g., memory). In some embodiments, the computing system corresponds to one of a tablet computer, laptop computer, mobile phone, near-eye system, wearable computing device, etc.

[0138] As discussed above, to change the behavior of a stylus, users typically select different tools from a menu of available tools. This can be a cumbersome and tedious process. In contrast, the innovation described in this paper enables users to dynamically change the operating modality of a physical object (e.g., a proxy object or electronic device, such as a stylus, a finger wearable device, a handheld device, etc.) based on its posture and grip. For example, a computing system (e.g., a presentation device) determines the posture and grip of the physical object based on tracking data associated with the physical object itself, as well as other inputs obtained by the computing system (such as computer vision, eye tracking, hand / limb tracking, voice input, etc.). In this way, users can seamlessly change the operating modality of a physical object (e.g., a handheld stylus) without disrupting their workflow.

[0139] As indicated by box 710, method 700 includes obtaining (e.g., receiving, retrieving, or detecting / collecting) a user input vector, which includes at least one user input indicator value associated with one of a plurality of different input modalities. For example, the user inputs the vector via one or more input devices of a computing system. Continuing this example, the one or more input devices may include an eye-tracking engine, a finger / hand / limb tracking engine, a head / body pose tracking engine, one or more microphones, an externally oriented image sensor, etc.

[0140] like Figure 5A As shown, the first part 500A of the content delivery architecture includes an eye-tracking engine 512, which determines / generates an eye-tracking vector 513 based on local sensor data 503 and remote sensor data 505, and updates the eye-tracking vector 513 over time. Furthermore, as... Figure 5A As shown, the first part 500A of the content delivery architecture includes a body / head pose tracking engine 514, which determines / generates a pose representation vector 515 based on local sensor data 503 and remote sensor data 505, and updates the pose representation vector 515 over time. According to some specific implementations, the eye tracking vector 513 and the pose representation vector 515 are collectively referred to as the user input vector 519. Figure 5B An exemplary data structure for eye tracking vector 513 and pose representation vector 515 is shown.

[0141] In some implementations, the computing system is also communicatively coupled to one or more externally oriented image sensors, and the acquisition of tracking data associated with a physical object includes analyzing image streams captured by one or more externally oriented image sensors to visually track the physical object.

[0142] In some implementations, the computing system is also communicatively coupled to a limb tracking subsystem that outputs one or more limb tracking indicator values, wherein the one or more limb tracking indicator values ​​are associated with a limb tracking modality among a plurality of input modalities, and wherein the user input vector includes the one or more limb tracking indicator values.

[0143] In some implementations, the computing system is also communicatively coupled to a head / body pose tracking subsystem that outputs one or more head / body pose tracking indicator values, wherein the one or more head / body pose tracking indicator values ​​are associated with a head / body pose tracking modality among a plurality of input modalities, and wherein the user input vector includes the one or more head / body pose tracking indicator values.

[0144] In some implementations, the computing system is also communicatively coupled to a speech detection subsystem that outputs one or more speech detection indicator values, wherein the one or more speech detection indicator values ​​are associated with a speech detection modality among a plurality of input modalities, and wherein the user input vector includes the one or more speech detection indicator values.

[0145] In some implementations, the computing system is also communicatively coupled to an eye-tracking subsystem that outputs one or more eye-tracking indicator values, wherein the one or more eye-tracking indicator values ​​are associated with an eye-tracking modality among a plurality of input modalities, and wherein the user input vector includes the one or more eye-tracking indicator values.

[0146] As indicated by box 720, method 700 includes obtaining (e.g., receiving, retrieving, or detecting / collecting) tracking data associated with a physical object. Figure 5A As shown, the first part 500A of the content delivery architecture includes an object tracking engine 510, which determines / generates an object tracking vector 511 based on tracking data 506 and updates the object tracking vector 511 over time. Figure 5B An exemplary data structure for object tracking vector 511 is shown.

[0147] As an example, a physical object corresponds to a proxy object detected within the physical environment that lacks a communication channel to the computing system, such as a pencil, pen, ruler, stick, etc. Figures 6I to 6RThe illustration shows a user 149 grasping a physical object 695 that cannot communicate with the electronic device 120 or be used to interact with the XR environment 128. As another example, the physical object corresponds to an electronic device with a wired or wireless communication channel to a computing system, such as a stylus, a finger-wearable device, a handheld device, etc. Figures 6A to 6H A user 149 is shown gripping a controller 130, which communicates with an electronic device 120 and is used to interact with an XR environment 128. Depending on some specific implementations, the computing system dynamically selects the operating mode of a physical object, regardless of whether the physical object corresponds to a proxy object or an electronic device.

[0148] In some implementations, one or more input devices include one or more externally oriented image sensors, and acquiring tracking data associated with a physical object includes analyzing an image stream captured by the one or more externally oriented image sensors to visually track the physical object. In some implementations, the tracking data corresponds to one or more images of the physical environment including the physical object, enabling the physical object to be tracked in six degrees of freedom (6DOF) via computer vision techniques. In some implementations, the tracking data corresponds to data collected by various integrated sensors of the physical object (such as GPS, IMU, accelerometer, gyroscope, magnetometer, etc.). For example, the tracking data corresponds to raw sensor data or processed data, such as translation values ​​associated with the physical object (relative to the physical environment or the world), rotation values ​​associated with the physical object (relative to gravity), velocity values ​​associated with the physical object, angular velocity values ​​associated with the physical object, acceleration values ​​associated with the physical object, angular acceleration values ​​associated with the physical object, a first pressure value associated with the force of the physical object contacting a physical surface, a second pressure value associated with the force of the user gripping the physical object, etc. In some implementations, the computing system also obtains finger manipulation data detected by the physical object via a communication interface. For example, finger manipulation data includes touch inputs or gestures pointing towards the touch-sensitive area of ​​the physical object. For example, finger manipulation data includes contact intensity data relative to the body of the physical object.

[0149] In some implementations, the computing system is also communicatively coupled to a physical object, and obtaining tracking data associated with the physical object includes obtaining tracking data from the physical object, wherein the tracking data corresponds to output data from one or more integrated sensors of the physical object. Therefore, according to some implementations, the tracking data associated with the physical object includes attitude values. Figures 6A to 6HA user 149 is shown gripping a controller 130, which communicates with electronics 120 and is used to interact with an XR environment 128. For example, the one or more integrated sensors include at least one of an IMU, accelerometer, gyroscope, GPS, magnetometer, one or more contact strength sensors, touch-sensitive surfaces, etc. In some implementations, the tracking data also indicates whether the tip of a physical object is in contact with a physical surface and the associated pressure value.

[0150] In some implementations, the computing system is also communicatively coupled to a physical object, and obtaining the user input vector includes obtaining the user input vector from the physical object, wherein the user input vector corresponds to output data from one or more integrated sensors of the physical object. Thus, according to some implementations, the user input vector includes a user grip value.

[0151] In some implementations, method 700 includes: obtaining one or more images of a physical environment; using one or more images of the physical environment to identify physical objects; and assigning physical objects (e.g., proxy objects) to act as focus selectors when interacting with a user interface. Figures 6I to 6R The illustration shows a user 149 grasping a physical object 695, which cannot communicate with the electronic device 120 or be used to interact with the XR environment 128. In some embodiments, the computing system designates the physical object as a focus selector when it is grasped by the user. In some embodiments, the computing system designates the physical object as a focus selector when it is grasped by the user and meets predefined constraints (e.g., maximum or minimum size, specific shape, Digital Rights Management (DRM) disqualifiers, etc.). Thus, in some embodiments, the user can use the physical object to interact with the XR environment. In some embodiments, gesture and grasp indicators can be anchored to the proxy object (or its representation) as the proxy object moves and / or the field of view (FOV) moves.

[0152] As shown in box 730, method 700 includes generating a first representation vector of a physical object based on a user input vector and tracking data. This first representation vector includes attitude values ​​and user grip values, where the attitude values ​​represent the spatial relationship between the physical object and the user of the computing system, and the user grip values ​​represent how the physical object is gripped by the user. Figure 5A As shown, the first part 500A of the content delivery architecture includes a representation engine 542 that determines / generates a physical object representation vector 543 based on user input vectors (e.g., a combination of eye tracking vector 513 and pose representation vector 515) and tracking data (e.g., object tracking vector 511) and updates the representation vector 543 over time. Figure 5B An exemplary data structure is shown, comprising a representation vector 543 including user grip value 5102 and posture value 5104.

[0153] In some implementations, user input vectors are used to fill in any occlusions or gaps in the tracking data and to disambiguate it. Therefore, in some implementations, if the pose values ​​and user grip values ​​determined / generated based on the tracking data fail to exceed predefined confidence levels, the computational system can use user input vectors to reduce resource consumption.

[0154] For example, pose values ​​correspond to one of the following: neutral pose, band conductor / wandistant pose, writing pose, surface pose, near-mouth pose, aiming pose, etc. In some implementations, pose values ​​indicate the orientation of a physical object, such as relative to a physical surface or another object detectable by computer vision techniques. In one example, the other object corresponds to a different physical or virtual object within the visible area of ​​the environment. In another example, the other object corresponds to a different physical or virtual object outside the visible area of ​​the environment but which can be inferred, such as a physical object moving near the user's mouth. For example, pose values ​​can indicate the proximity of a physical object relative to another object. For example, grip values ​​correspond to one of the following: remote control grip, wand grip, writing grip, reverse writing grip, gamepad grip, thumb-tip grip, horizontal grip, gamepad grip, flute grip, igniter grip, etc.

[0155] As indicated by box 740, method 700 includes selecting a first operating mode as the current operating mode of the physical object based on a first representation vector. As an example, see... Figure 6B In response to determining that the current posture value corresponds to "near mouth" and the current user grip value corresponds to "wand," the electronic device 120 selects dictation mode as the current operating mode of the controller 130 and also displays a notification 624 indicating "dictation mode activated." Thus, when the current operating mode of the controller 130 corresponds to dictation mode, the user 149 can indicate what information, notes, content, etc., should be displayed within the XR environment 128. In some embodiments where the physical object does not include one or more microphones, the microphone of the electronic device 120 can be used to sense the user's speech. In some embodiments where the physical object does not include one or more microphones, images captured by the image sensor of the electronic device 120 can be analyzed to read the user's lips and subsequently sense the user's speech. Those skilled in the art will understand that the recording notification 624 is merely an exemplary visualization and can be modified or replaced in various other embodiments.

[0156] As another example, see Figure 6CIn response to determining that the current posture value corresponds to "neutral" and the current user grip value corresponds to "writing," the electronic device 120 selects the marking mode as the current operating mode of the controller 130, and also displays a notification 634 indicating that "marking mode is activated." As yet another example, see... Figure 6F In response to determining that the current posture value corresponds to "surface" and the current user grip value corresponds to "reverse writing", the electronic device 120 selects the erase mode as the current operating mode of the controller 130 and also displays a notification 664 indicating that "erasure mode is activated".

[0157] As yet another example, see Figure 6I In response to determining that the current posture value corresponds to "aiming" and the current user grip value corresponds to "wank," the electronic device 120 selects the pointing mode as the current operating mode of the physical object 695 and also displays a notification 694 indicating that "pointing mode is activated." Thus, when the current operating mode of the physical object 695 corresponds to the pointing mode, the user 149 can use the physical object 695 as a laser pointing device within the XR environment 128. As another example, see... Figure 6I In response to determining that the current posture value corresponds to "neutral" and the current user grip value corresponds to "magic wand", the electronic device 120 selects a manipulation mode as the current operating mode of the physical object 695 and also displays a notification 6114 indicating that "manipulation mode is activated". Thus, when the current operating mode of the physical object 695 corresponds to the manipulation mode, the user 149 can use the physical object 695 to manipulate (e.g., translate, rotate, etc.) XR content within the XR environment 128.

[0158] In some specific implementations, as indicated by box 742, the first operating mode corresponds to one of the following: dictation mode, digital assistant mode, navigation mode, manipulation mode, marking mode, erasing mode, pointing mode, or representation mode. As an example, when the gesture value corresponds to a neutral gesture (or surface gesture) and the user grip value corresponds to a writing tool-like grip, the first operating mode corresponds to marking mode. As another example, when the gesture value corresponds to a neutral gesture (or surface gesture) and the user grip value corresponds to a reverse writing tool-like grip, the first operating mode corresponds to erasing mode. As yet another example, when the gesture value corresponds to a near-mouth gesture and the user grip value corresponds to a wand-like grip, the first operating mode corresponds to dictation mode. As yet another example, when the gesture value corresponds to a neutral gesture and the user grip value corresponds to a wand-like grip, the first operating mode corresponds to manipulation mode. As yet another example, when the gesture value corresponds to an aiming gesture and the user grip value corresponds to a wand-like grip, the first operating mode corresponds to pointing mode. As yet another example, when the posture value corresponds to a near-mouth posture and the user grip value corresponds to a wand-like grip, and the computing device detects voice input containing keywords or key phrases, the first operating mode corresponds to a digital assistant mode. Those skilled in the art will understand that different combinations of posture and user grip values ​​can trigger the aforementioned modes.

[0159] In some implementations, after selecting a first operating mode as the current output mode for the physical object, method 700 includes presenting an extended reality (XR) environment comprising one or more virtual objects via a display device, wherein the physical object is provided for interaction with one or more virtual objects within the XR environment according to the first operating mode. In some implementations, the one or more virtual objects are overlaid on the physical environment when displayed within the XR environment. As an example, in Figure 6A In the XR environment 128, there are a representation 116 of gate 115 and a virtual agent 606 that overlays the representation of the physical environment 105. As another example, in... Figure 6K In this context, the XR environment 128 includes a first view 6116A of XR content (e.g., the front of a cube, box, etc.) that is overlaid on a representation of the physical environment 105.

[0160] In some embodiments, the display device corresponds to a transparent lens assembly, and the presentation of the XR environment is projected onto the transparent lens assembly. In some embodiments, the display device corresponds to a near-eye system, and the presentation of the XR environment includes combining the presentation of the XR environment with one or more images of the physical environment captured by an externally facing image sensor.

[0161] In some implementations, the XR environment includes one or more visual indicators associated with at least one of pose values ​​and user grip values. As an example, in... Figure 6A In the XR environment 128, a visual indicator 612 of instance 610 includes an instance 610 having a current pose value (e.g., no current pose value because controller 130 is not currently being gripped by user 149) and a current user grip value (e.g., no user grip value because controller 130 is not currently being gripped by user 149). Those skilled in the art will understand that the visual indicator 612 is merely an exemplary visualization having pose and user grip values, and it can be modified or replaced in various other specific implementations. As an example, the visual indicator 612 could be a badge, icon, audio output, etc., instead of a text notification. As another example, in Figure 6B In the XR environment 128, there is a visual indicator 622 of instance 620 having a current posture value (e.g., near mouth) and a current user grip value (e.g., wand), and a representation 131 of controller 130 held by representation 153 of user 149's right hand 152 (e.g., video passthrough of physical environment 105).

[0162] In some implementations, one or more visual indicators correspond to icons, badges, text boxes, notifications, etc., to indicate the current pose and grip value to the user. In some implementations, one or more visual indicators are world-locked, head-locked, body-locked, etc. In one example, one or more visual indicators may be anchored to the physical object (or its representation) as the physical object moves and / or the field of view (FOV) moves. In some implementations, one or more visual indicators may change over time based on changes in tracking data and / or user input vectors. For example, visual indicators associated with pose values ​​include a representation of the physical object (e.g., an icon) and a representation of the associated pose.

[0163] In some implementations, the XR environment includes one or more visual indicators associated with the current operating modality. As an example, in... Figure 6B In this embodiment, XR environment 128 includes a notification 624 indicating that "dictation mode is activated." Those skilled in the art will understand that the recording notification 624 is merely an exemplary visualization and can be modified or replaced in various other embodiments. In some embodiments, one or more visual indicators correspond to icons, badges, text boxes, notifications, etc., to indicate the current operating mode to the user. In some embodiments, one or more visual indicators are world-locked, head-locked, body-locked, etc. In one example, one or more visual indicators may be anchored to a physical object (or its representation) as the physical object moves and / or the field of view (FOV) moves. In some embodiments, one or more visual indicators may change over time based on changes in tracking data and / or user input vectors.

[0164] In some specific implementations, after selecting a first operating mode as the current output mode of the physical object, method 700 includes: displaying a user interface via a display device; detecting modification input from the physical object to content within the user interface (e.g., tracking the physical object in 3D using IMU data, computer vision, magnetic tracking, etc.); and modifying the content within the user interface based on the modification input (e.g., modifying the size, shape, displacement, etc. of the input) and the first operating mode in response to detecting the modification input. As an example, Figures 6C to 6E This shows that the electronic device 120 responds to the detection Figure 6E The marker input 644 (e.g., modifying the input) is used to display a sequence of markers 654 within the XR environment 128. For example, marker 654 corresponds to the shape, displacement, etc., of marker input 644. In another example, marker 654 corresponds to a function of the shape, displacement, etc., of marker input 644 and weighting coefficients (e.g., magnification, attenuation, etc.). As another example, Figures 6K to 6M This shows that the electronic device 120 responds to the detection Figure 6L Manipulating input 6122 (e.g., modifying input) to transfer XR content from Figure 6K The first-person perspective 6116A (e.g., the front of a cube, box, etc.) is rotated 180° clockwise to Figure 6M The sequence of second perspective 6116B (e.g., the back of a cube, box, etc.).

[0165] In some implementations, method 700 includes: detecting a change in either a user input vector or tracking data associated with a physical object; determining a second representation vector of the physical object based on the change in either the user input vector or the tracking data associated with the physical object; and selecting a second operating mode as the current operating mode of the physical object based on the second representation vector, wherein the second operating mode is different from the first operating mode. In some implementations, method 700 includes: displaying a user interface via a display device; after selecting the second operating mode as the current output mode of the physical object, detecting a modification input to the content within the user interface (e.g., tracking the physical object in 3D using IMU data, computer vision, magnetic tracking, etc.); and modifying the content within the user interface based on the modification input (e.g., modifying the size, shape, displacement, etc. of the input) and the second operating mode in response to detecting the modification input.

[0166] In some implementations, method 700 includes: in response to determining that a first operating mode corresponds to invoking a digital assistant, displaying a digital assistant (DA) indicator within a user interface; and in response to determining that one or more eye-tracking indicator values ​​correspond to the DA indicator, changing the appearance of the DA indicator and enabling DA. In some implementations, changing the appearance of the DA indicator corresponds to at least one of scaling the size of the DA indicator, changing the color of the DA indicator, changing the brightness of the DA indicator, etc. As an example, Figures 6N to 6R This illustrates the sequence of current operating modes corresponding to Digital Assistant (DA) modes. Continuing the example, the appearance of the DA indicator changes in response to the detection that user 149's gaze has been directed at the DA indicator for at least a predefined amount of time (e.g., size changes from...). Figure 6O The DA indicator 6152A in the middle is increased to Figure 6P (DA indicator 6152B in the example). Continuing with this example, DA is based on... Figure 6Q User 149 performs a search operation by providing voice input 6172 (e.g., a search string), and electronic device 120 displays the search results 6182 within the XR environment, which correspond to... Figure 6Q The search string associated with the detected voice input 6172. For example, search results 6182 include text, audio content, video content, 3D content, XR content, etc.

[0167] In some implementations, as the physical object moves and / or the field of view (FOV) moves, the DA indicator is displayed near the physical object's location and remains anchored to the physical object (or its representation). In some implementations, the electronic device receives a search string and, in response, retrieves the results information from the DA to be displayed within the XR environment. In one example, the user can pan and / or rotate the results information within the XR environment. In another example, the user can pin the obtained information to a bulletin board or other storage within the XR environment.

[0168] While various aspects of specific embodiments within the scope of the appended claims have been described above, it should be apparent that the various features of the above-described embodiments can be embodied in a wide variety of forms, and any particular structure and / or function described above are merely illustrative. Based on this disclosure, those skilled in the art will understand that the aspects described herein can be implemented independently of any other aspects, and two or more of these aspects can be combined in various ways. For example, any number of aspects set forth herein can be used to implement an apparatus and / or practice a method. Furthermore, such an apparatus and / or such a method can be implemented using other structures and / or functions besides or different from one or more aspects set forth herein.

[0169] It will also be understood that while terms such as "first," "second," etc., may be used herein to describe various elements, these elements should not be limited by these terms. These terms are merely used to distinguish one element from another. For example, a first media item can be referred to as a second media item, and similarly, a second media item can be referred to as a first media item, which changes the meaning of the description, provided that any occurrence of "first media item" is consistently renamed and any occurrence of "second media item" is consistently renamed. The first media item and the second media item are both media items, but they are not the same media item.

[0170] The terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the claims. As used in the description of these embodiments and in the appended claims, the singular forms “a,” “an,” and “the” are intended to also cover the plural forms unless the context clearly indicates otherwise. It will also be understood that the term “and / or” as used herein refers to and covers any and all possible combinations of one or more of the associated listed items. It will also be understood that the term “comprising,” when used in this specification, specifies the presence of the stated features, integers, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.

[0171] As used herein, the term "if" can be interpreted as meaning "when the prerequisite is true" or "when the prerequisite is true" or "in response to determination" or "according to determination" or "in response to detection" that the prerequisite is true, depending on the context. Similarly, the phrases "if it is determined [the prerequisite is true]" or "if [the prerequisite is true]" or "when [the prerequisite is true]" are interpreted as meaning "when it is determined that the prerequisite is true" or "in response to determination" or "according to determination" that the prerequisite is true or "when the prerequisite is detected" or "in response to detection" that the prerequisite is true, depending on the context.

Claims

1. A method for dynamically selecting the operation mode of a physical object, comprising: In a computing system comprising non-transitory memory and one or more processors, wherein the computing system is communicatively coupled to a display device: Obtain a user input vector, the user input vector including at least one user input indicator value associated with one of a plurality of different input modalities; Obtain tracking data associated with the physical object; A first representation vector of the physical object is generated based on the user input vector and the tracking data associated with the physical object. The first representation vector includes both a pose value and a user grip value, wherein the pose value represents the position of the physical object relative to the user of the computing system, and the user grip value represents the way the physical object is gripped by the user. as well as The first operating mode is selected as the current operating mode of the physical object based on the first representation vector.

2. The method according to claim 1, wherein the first operating mode corresponds to one of the following: dictation mode, digital assistant mode, navigation mode, manipulation mode, marking mode, erasing mode, pointing mode, or display mode.

3. The method according to any one of claims 1 to 2, further comprising: After selecting the first operating mode as the current operating mode of the physical object: The user interface is displayed via the display device; as well as Detecting the physical object's input that modifies content within the user interface; and In response to detecting the modified input, the content within the user interface is modified based on the modified input and the first operation mode.

4. The method according to any one of claims 1 to 2, further comprising: Detect changes in either the user input vector or the tracking data associated with the physical object; as well as In response to detecting a change in either the user input vector or the tracking data associated with the physical object: A second representation vector of the physical object is determined based on the change in either the user input vector or the tracking data associated with the physical object; and Based on the second representation vector, a second operation mode is selected as the current operation mode of the physical object, wherein the second operation mode is different from the first operation mode.

5. The method according to claim 4, further comprising: The user interface is displayed via the display device; After selecting the second operation mode as the current operation mode of the physical object, the modification input of the physical object to the content within the user interface is detected; as well as In response to detecting the modified input, the content within the user interface is modified based on the modified input and the second operation mode.

6. The method according to claim 1, further comprising: After selecting the first operating mode as the current operating mode of the physical object, an extended reality XR environment including one or more virtual objects is presented via the display device, wherein the physical object is provided for interacting with the one or more virtual objects in the XR environment according to the first operating mode.

7. The method of claim 6, wherein the display device corresponds to a transparent lens assembly, and wherein the presentation of the XR environment is projected onto the transparent lens assembly.

8. The method of claim 6, wherein the display device corresponds to a near-eye system, and wherein presenting the XR environment comprises combining the presentation of the XR environment with one or more images of the physical environment captured by an externally oriented image sensor.

9. The method of any one of claims 6 to 8, wherein the XR environment includes one or more visual indicators associated with at least one of the pose value and the user grip value.

10. The method of any one of claims 6 to 8, wherein the XR environment includes one or more visual indicators associated with the current operating mode.

11. The method according to any one of claims 1 to 2, further comprising: Obtain one or more images of the physical environment; The physical object is identified using one or more images of the physical environment; as well as The physical object is assigned to act as a focus selector when interacting with the user interface.

12. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to the physical object, and wherein obtaining the tracking data associated with the physical object comprises obtaining the tracking data from the physical object, wherein the tracking data corresponds to output data from one or more integrated sensors of the physical object.

13. The method of claim 12, wherein the tracking data associated with the physical object includes the attitude value.

14. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to the physical object, and wherein obtaining the user input vector includes obtaining the user input vector from the physical object, wherein the user input vector corresponds to output data from one or more integrated sensors of the physical object.

15. The method of claim 14, wherein the user input vector includes the user grip value.

16. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to one or more externally oriented image sensors, and wherein obtaining the tracking data associated with the physical object comprises analyzing an image stream captured by the one or more externally oriented image sensors to visually track the physical object.

17. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to a limb tracking subsystem that outputs one or more limb tracking indicator values, wherein the one or more limb tracking indicator values ​​are associated with a limb tracking modality among the plurality of input modalities, and wherein the user input vector includes the one or more limb tracking indicator values.

18. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to a head / body pose tracking subsystem that outputs one or more head / body pose tracking indicator values, wherein the one or more head / body pose tracking indicator values ​​are associated with a head / body pose tracking modality among the plurality of input modalities, and wherein the user input vector includes the one or more head / body pose tracking indicator values.

19. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to a speech detection subsystem that outputs one or more speech detection indicator values, wherein the one or more speech detection indicator values ​​are associated with a speech detection modality among the plurality of input modalities, and wherein the user input vector includes the one or more speech detection indicator values.

20. The method of any one of claims 1 to 2, wherein the computing system is further communicatively coupled to an eye-tracking subsystem that outputs one or more eye-tracking indicator values, wherein the one or more eye-tracking indicator values ​​are associated with an eye-tracking modality among the plurality of input modalities, and wherein the user input vector includes the one or more eye-tracking indicator values.

21. The method of claim 20, further comprising: In response to determining that the first operating mode corresponds to the invocation of the digital assistant, the digital assistant DA indicator is displayed within the user interface; as well as In response to determining that the one or more eye-tracking indicator values ​​correspond to the DA indicator, the appearance of the DA indicator is changed and the DA is enabled.

22. The method of claim 21, wherein changing the appearance of the DA indicator corresponds to at least one of scaling the size of the DA indicator, changing the color of the DA indicator, or changing the brightness of the DA indicator.

23. A device for dynamically selecting the operating mode of a physical object, the device comprising: One or more processors; Non-transitory memory; An interface used for communicating with display devices; and One or more programs stored in the non-transitory memory, which, when executed by the one or more processors, cause the device to perform any one of the methods according to claims 1 to 22.

24. A non-transitory memory storing one or more programs, said one or more programs, when executed by one or more processors of a device having an interface for communicating with a display device, causing said device to perform any one of the methods according to claims 1 to 22.

25. A device for dynamically selecting the operating mode of a physical object, the device comprising: One or more processors; Non-transitory memory; An interface used for communicating with display devices; and An apparatus for causing the device to perform any one of the methods according to claims 1 to 22.

Citation Information

Patent Citations

  • Spatial, multi-modal control device for use with spatial operating system

    CN102460510A

  • Multi-device multi-user sensor correlation for pen and computing device interaction

    US20150363034A1