User interaction in extended reality

Extended reality technology creates a virtual desktop environment that addresses the mobility vs. screen size dilemma, enabling a flexible and productive workspace through wearable devices.

JP2025106348APending Publication Date: 2025-07-15サイトフル コンピューターズ リミテッド
View PDF 10 Cites 0 Cited by

Patent Information

Application Number
JP2025060396
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-02-07
Filing Date
2025-04-01
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Users face a productivity dilemma between mobility and screen size, as laptops lack the ability to provide a large, stationary workspace while maintaining portability.

Method used

Utilizing extended reality (XR) technology to create a virtual desktop environment that can be accessed via wearable devices, allowing users to interact with virtual objects and adjust display modes based on device movement.

Benefits of technology

Enables a mobile workspace with a large screen environment, enhancing user productivity by providing a flexible and adaptable virtual desktop experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025106348000001_ABST
    Figure 2025106348000001_ABST
Patent Text Reader

Abstract

To provide a screen in a virtual desktop shape.SOLUTION: A virtual object is allowed to be virtually presented in an environment through a wearable extended reality device operable in first and second display modes. A position of the virtual object, in a first display mode, is kept in the environment regardless of detected motion of the wearable extended reality device, and the virtual object, in a second display mode, moves in the environment according to the detected motion of the wearable extended device. The motion of the wearable extended reality device is allowed to be detected. Selection of the first or second display mode is allowed to be received. A display signal configured to present the virtual object in a mode corresponding to the selected display mode is allowed to be outputted for presentation through the wearable extended reality device according to the selected display mode.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] Cross - Reference to Related Applications This application claims the benefit of priority of U.S. Provisional Patent Application No. 63 / 147,051, filed on February 8, 2021; U.S. Provisional Patent Application No. 63 / 157,768, filed on March 7, 2021; U.S. Provisional Patent Application No. 63 / 173,095, filed on April 9, 2021; U.S. Provisional Patent Application No. 63 / 213,019, filed on June 21, 2021; U.S. Provisional Patent Application No. 63 / 215,500, filed on June 27, 2021; U.S. Provisional Patent Application No. 63 / 216,335, filed on June 29, 2021; U.S. Provisional Patent Application No. 63 / 226,977, filed on July 29, 2021; U.S. Provisional Patent Application No. 63 / 300,005, filed on January 16, 2022; U.S. Provisional Patent Application No. 63 / 307,207, filed on February 7, 2022; U.S. Provisional Patent Application No. 63 / 307,203, filed on February 7, 2022; and U.S. Provisional Patent Application No. 63 / 307,217, filed on February 7, 2022, and all of these U.S. provisional patent applications are hereby incorporated by reference in their entirety into this application.

[0002] The present disclosure generally relates to the field of extended reality. More specifically, the present disclosure relates to systems, methods, and devices for providing productivity applications using an extended reality environment.

Background Art

[0003] For years, PC users have faced a productivity dilemma of either sacrificing mobility (when choosing a desktop computer) or screen size (when choosing a laptop computer). One partial solution to this dilemma is to use a docking station. A docking station is an interface device for connecting a laptop computer to other devices. By plugging the laptop computer into the docking station, laptop users can enjoy the improved visibility provided by a larger monitor. However, since the large monitor is stationary, the user's mobility is improved but still limited. For example, even a laptop user with a docking station does not have the freedom to use two 32-inch screens at a desired location.

[0004] Some of the disclosed embodiments are directed to a new approach for solving the productivity dilemma, which provides a mobile environment using extended reality (XR) to provide a virtual desktop-like screen so that users can experience the comfort of a stationary workspace anywhere they desire. SUMMARY OF THE INVENTION

[0005] Embodiments consistent with the present disclosure provide systems, methods, and devices for providing and supporting productivity applications using an extended reality environment.

[0006] Some of the disclosed embodiments can include a system, method, and non-transitory computer-readable medium for controlling a viewpoint in an extended reality environment using a physical touch controller. Some of these embodiments include outputting a first display signal that reflects a first viewpoint of a scene for presentation via a wearable extended reality device, receiving a first input signal caused by a first multi-finger interaction with a touch sensor via the touch sensor, and in response to the first input signal, outputting a second display signal configured to present a second viewpoint of the scene via the wearable extended reality device by changing the first viewpoint of the scene for presentation via the wearable extended reality device, receiving a second input signal caused by a second multi-finger interaction with the touch sensor via the touch sensor, and in response to the second input signal, outputting a third display signal configured to present a third viewpoint of the scene via the wearable extended reality device by changing the second viewpoint of the scene for presentation via the wearable extended reality device.

[0007] Some of the disclosed embodiments can include a system, method, and non-transitory computer-readable medium for enabling gesture interaction with invisible virtual objects. Some of these embodiments include receiving image data captured by at least one image sensor of a wearable extended reality device, the image data including a display of a plurality of physical objects within a field of view associated with the at least one image sensor of the wearable extended reality device; displaying a plurality of virtual objects in a portion of the field of view, the portion of the field of view being associated with a display system of the wearable extended reality device; receiving a selection of a particular physical object from the plurality of physical objects; receiving a selection of a particular virtual object from the plurality of virtual objects for association with the particular physical object; docking the particular virtual object with the particular physical object; receiving a gesture input indicating that a hand is interacting with the particular virtual object when the particular physical object and the particular virtual object are outside of a portion of the field of view associated with the display system of the wearable extended reality device such that the particular virtual object is not visible to a user of the wearable extended reality device; and causing an output associated with the particular virtual object in response to the gesture input.

[0008] Some of the disclosed embodiments can include a system, a method, and a non-transitory computer-readable medium for performing an incremental convergence operation in an extended reality environment. Some of these embodiments can include displaying a plurality of distributed virtual objects across a plurality of virtual regions, the plurality of virtual regions including at least a first virtual region and a second virtual region different from the first virtual region; receiving an initial motion input that tends towards the first virtual region; highlighting a group of virtual objects within the first virtual region based on the initial motion input; receiving a refined motion input that tends towards a particular virtual object from within the highlighted group of virtual objects; and triggering a function associated with the particular virtual object based on the refined motion input.

[0009] Some of the disclosed embodiments can include a system, a method, and a non-transitory computer-readable medium for facilitating an environmentally adaptive extended reality display in a physical environment. Some of these embodiments can include virtually displaying content via a wearable extended reality device operating in the physical environment, wherein displaying the content via the wearable extended reality device is associated with at least one adjustable extended reality display parameter; obtaining image data from the wearable extended reality device; detecting a particular environmental change in the image data that is unrelated to the virtually displayed content; accessing a group of rules associating the environmental change with a change in the at least one adjustable extended reality display parameter; determining that the particular environmental change corresponds to a particular rule of the group of rules; and implementing the particular rule to adjust the at least one adjustable extended reality display parameter based on the particular environmental change.

[0010] Some of the disclosed embodiments can include a system, method, and non-transitory computer-readable medium for selectively controlling the display of virtual objects. Some of these embodiments include virtually presenting a plurality of virtual objects within an environment via a wearable extended reality device operable in a first display mode and a second display mode, wherein in the first display mode, the positions of the plurality of virtual objects are maintained within the environment regardless of the detected movement of the wearable extended reality device, and in the second display mode, the plurality of virtual objects move within the environment in response to the detected movement of the wearable extended reality device; detecting the movement of the wearable extended reality device; receiving a selection of the first display mode or the second display mode for use in virtually presenting the plurality of virtual objects while the wearable extended reality device is in motion; and outputting a display signal configured to present the plurality of virtual objects in a manner consistent with the selected display mode for presentation via the wearable extended reality device in response to the selected display mode.

[0011] Some of the disclosed embodiments can include a system, method, and non-transitory computer-readable medium for moving a virtual cursor along two transverse virtual planes. Some of these embodiments include generating a display via a wearable extended reality device, the display including the virtual cursor and a plurality of virtual objects located on a first virtual plane that intersects a second virtual plane on a physical surface; receiving a first two-dimensional input via a surface input device while the virtual cursor is displayed on the first virtual plane, the first two-dimensional input reflecting an intention to select a first virtual object located on the first virtual plane; causing a first cursor movement toward the first virtual object in response to the first two-dimensional input, the first cursor movement being along the first virtual plane; receiving a second two-dimensional input via the surface input device while the virtual cursor is displayed on the first virtual plane, the second two-dimensional input reflecting an intention to select a second virtual object that appears on the physical surface; and causing a second cursor movement toward the second virtual object in response to the second two-dimensional input, the second cursor movement including a partial movement along the first virtual plane and a partial movement along the second virtual plane.

[0012] Some of the disclosed embodiments can include a system, a method, and a non - transitory computer - readable medium for enabling cursor control in an extended reality space. Some of these embodiments include receiving, from an image sensor, first image data reflecting a first focus region of a user of a wearable extended reality device, where the first focus region is within an initial field of view in the extended reality space; causing a first presentation of a virtual cursor at the first focus region; receiving, from the image sensor, second image data reflecting a second focus region of the user outside the initial field of view in the extended reality space; receiving input data indicating a user desire to interact with the virtual cursor; and causing a second presentation of the virtual cursor at the second focus region in response to the input data.

[0013] Some of the disclosed embodiments can include a system, a method, and a non - transitory computer - readable medium for moving virtual content between virtual planes in a three - dimensional space. Some of these embodiments include using a wearable extended reality device to virtually present a plurality of virtual objects on a plurality of virtual planes, where the plurality of virtual planes includes a first virtual plane and a second virtual plane; outputting, for presentation via the wearable extended reality device, a first display signal reflecting a virtual object at a first position on the first virtual plane; receiving an in - plane input signal for moving the virtual object to a second position on the first virtual plane; causing the wearable extended reality device to virtually display an in - plane movement of the virtual object from the first position to the second position in response to the in - plane input signal; receiving an inter - plane input signal for moving the virtual object to a third position on the second virtual plane while the virtual object is at the second position; and causing the wearable extended reality device to virtually display an inter - plane movement of the virtual object from the second position to the third position in response to the inter - plane input signal.

[0014] Consistent with other disclosed embodiments, the non-transitory computer-readable storage medium can store program instructions that are executed by at least one processing device to execute any of the methods described herein.

[0015] The foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the claims.

[0016] The accompanying drawings, which are incorporated in and constitute a part of this disclosure, illustrate various disclosed embodiments.

Brief Description of the Drawings

[0017]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Figure 14

Figure 15

Figure 16

Figure 17

Figure 18

Figure 19

Figure 20

Figure 21

Figure 22

Figure 23

Figure 24

Figure 25

Figure 26

Figure 27

Figure 28

Figure 29A

Figure 29B

Figure 29C

Figure 30

Figure 31A

Figure 31B

Figure 32

Figure 33

Figure 34

Figure 35

Figure 36

Figure 37A

Figure 37B

Figure 38

Figure 39

Figure 40

Figure 41

Figure 42

Figure 43

Figure 44

Figure 45

Figure 46

Figure 47

Figure 48A

Figure 48B

Figure 49

Figure 50A

Figure 50B

Figure 50C

Figure 50D

Figure 51

Figure 52

Figure 53

Figure 54

Figure 55

Figure 56

Figure 57

Figure 58

[0018] The following detailed description refers to the accompanying drawings. Whenever possible, the same reference numbers are used in the drawings and the following description to refer to the same or similar parts. Although several exemplary embodiments are described herein, changes, adaptations, and other implementations are possible. For example, substitutions, additions, or changes can be made to the components shown in the drawings, and the exemplary methods described herein can be changed by substituting, rearranging, deleting, or adding steps to the disclosed methods. Accordingly, the following detailed description is not limited to specific embodiments and examples, but includes the general principles described herein and shown in the drawings, in addition to the general principles encompassed by the appended claims.

[0019] The present disclosure relates to systems and methods for providing an extended reality environment to a user. The term "extended reality environment", which may also be referred to as "extended reality", "extended reality space", or "extended reality environment", refers to all types of virtual reality composite environments and human-machine interactions that are at least partially generated by computer technology. The extended reality environment may be a completely simulated virtual environment, or a combination of a real environment and a virtual environment that a user can perceive from different perspectives. In some examples, the user can interact with elements of the extended reality environment. One non-limiting example of an extended reality environment may be a virtual reality environment, also known as "virtual reality" or "virtual environment". An immersive virtual reality environment may be a simulated non-physical environment that provides the user with the perception of being present in the virtual environment. Another non-limiting example of an extended reality environment may be an augmented reality environment, also known as "augmented reality" or "augmented environment". The augmented reality environment may include a direct or indirect view of the live physical real-world environment enhanced with virtual computer-generated perceptual information, such as virtual objects with which the user can interact. Another non-limiting example of an extended reality environment is a mixed reality environment, also known as "mixed reality" or "mixed environment". The mixed reality environment may be a hybrid of the physical real-world environment and the virtual environment, where physical objects and virtual objects coexist and can interact in real time. In some examples, both the augmented reality environment and the mixed reality environment can include a combination of the real world and the virtual world, real-time interaction, and accurate 3D registration of virtual objects and real objects. In some examples, both the augmented reality environment and the mixed reality environment can include constructive overlaid sensory information that can be added to the physical environment. In other examples, both the augmented reality environment and the mixed reality environment can include destructive virtual content that can mask at least a portion of the physical environment.

[0020] In some embodiments, the system and method can provide an extended reality environment using extended reality devices. The term "extended reality device" can include any type of device or system that enables a user to perceive and / or interact with an extended reality environment. An extended reality device can enable a user to perceive and / or interact with an extended reality environment through one or more sensory modalities. Some non-limiting examples of such sensory modalities can include vision, hearing, touch, proprioception, and smell. An example of an extended reality device is a virtual reality device that enables a user to perceive and / or interact with a virtual reality environment. Another example of an extended reality device is an augmented reality device that enables a user to perceive and / or interact with an augmented reality environment. Yet another example of an extended reality device is a mixed reality device that enables a user to perceive and / or interact with a mixed reality environment.

[0021] According to one aspect of the present disclosure, the extended reality device may be a wearable device such as a head-mounted device, for example, smart glasses, smart contact lenses, a headset, or any other device worn by a human for the purpose of presenting extended reality to the human. Other extended reality devices may include a holographic projector, or any other device or system capable of providing augmented reality (AR), virtual reality (VR), mixed reality (MR), or any immersive experience. Typical components of a wearable extended reality device may include a stereoscopic head-mounted display, a stereoscopic head-mounted sound system, head motion tracking sensors (e.g., gyroscopes, accelerometers, magnetometers, image sensors, structured light sensors, etc.), a head-mounted projector, eye tracking sensors, and at least one of the further components described below. According to another aspect of the present disclosure, the extended reality device may be a non-wearable extended reality device. Specifically, the non-wearable extended reality device may include a multi-projection environment device. In some embodiments, the extended reality device may be configured to change the viewing perspective of the extended reality environment in response to the movement of the user, particularly the movement of the user's head. In one example, a wearable extended reality device may change the viewing field of the extended reality environment in response to a change in the user's head posture, such as by changing the spatial direction without changing the user's spatial position in the extended reality environment. In another example, a non-wearable extended reality device may change the user's spatial position in the extended reality environment in response to a change in the user's position in the real world, such as by changing the user's spatial position in the extended reality environment without changing the direction of the viewing field with respect to the spatial position.

[0022] According to some embodiments, the extended reality device may include a digital communication device configured to perform at least one of receiving virtual content data configured to enable presentation of virtual content, transmitting virtual content for sharing with at least one external device, receiving context data from at least one external device, transmitting context data to at least one external device, transmitting usage data indicating usage of the extended reality device, and transmitting data based on information captured using at least one sensor included in the extended reality device. In a further embodiment, the extended reality device may include a memory for storing at least one of virtual data configured to enable presentation of virtual content, context data, usage data indicating usage of the extended reality device, sensor data based on information captured using at least one sensor included in the extended reality device, software instructions configured to cause a processing device to present virtual content, software instructions configured to cause the processing device to collect and analyze context data, software instructions configured to cause the processing device to collect and analyze usage data, and software instructions configured to cause the processing device to collect and analyze sensor data. In a further embodiment, the extended reality device may include a processing device configured to perform at least one of rendering virtual content, collecting and analyzing context data, collecting and analyzing usage data, and collecting and analyzing sensor data. In a further embodiment, the extended reality device may include one or more sensors.One or more sensors can include one or more image sensors (e.g., configured to capture images and / or videos of a user of the device and / or the user's environment), one or more motion sensors (e.g., accelerometers, gyroscopes, magnetometers, etc.), one or more positioning sensors (e.g., GPS, outdoor positioning sensors, indoor positioning sensors, etc.), one or more temperature sensors (e.g., configured to measure the temperature of at least a portion of the device and / or the environment), one or more contact sensors, one or more proximity sensors (e.g., configured to detect whether the device is currently being worn), one or more electrical impedance sensors (e.g., configured to measure the electrical impedance of a user), a gaze detector, an optical tracker, an electro-potential tracker (e.g., an electrooculogram (EOG) sensor), a video-based eye gaze tracker, an infrared / near-infrared sensor, a passive light sensor, or one or more eye gaze tracking sensors such as any other technique capable of determining where a person is looking or gazing.

[0023] In some embodiments, the system and method can use an input device to interact with an extended reality device. The term input device can include any physical device configured to receive input from a user or the user's environment and provide data to a computing device. The data provided to the computing device may be in digital and / or analog form. In one embodiment, the input device can store input received from the user in a memory device accessible by a processing device, and the processing device can access the data stored for analysis. In another embodiment, the input device can directly provide data to the processing device, for example via a bus or via another communication system configured to transfer data from the input device to the processing device. In some examples, the input received by the input device can include keystrokes, tactile input data, motion data, position data, gesture-based input data, direction data, or any other data for providing for computation. Some examples of input devices can include buttons, keys, keyboards, computer mice, touch pads, touch screens, joysticks, or any other mechanism capable of receiving input. Another example of an input device can include an integrated computing interface device that includes at least one physical component for receiving input from a user. The integrated computing interface device can include at least memory, a processing device, and at least one physical component for receiving input from a user. In one example, the integrated computing interface device may further include a digital network interface that enables digital communication with other computing devices. In one example, the integrated computing interface device may further include a physical component for outputting information to the user. In some examples, all components of the integrated computing interface device may be included in a single housing, and in other examples, the components may be distributed across two or more housings.Some non-limiting examples of physical components for receiving input from a user that can be included in an integrated computing interface device can include at least one of a button, a key, a keyboard, a touchpad, a touch screen, a joystick, or any other mechanism or sensor capable of receiving computing information. Some non-limiting examples of physical components for outputting information to a user can include at least one of an optical indicator (such as an LED indicator), a screen, a touch screen, a buzzer, an audio speaker, or any other audio, video, or tactile device capable of providing a human-perceivable output.

[0024] In some embodiments, one or more image sensors can be used to capture image data. In some examples, the image sensor may be included in an extended reality device, a wearable device, a wearable extended reality device, an input device, the user's environment, etc. In some examples, the image data may be read from memory, received from an external device, (e.g., generated using a generative model), etc. Some non-limiting examples of image data can include images, grayscale images, color images, 2D images, 3D images, videos, 2D videos, 3D videos, frames, images, data derived from other image data, etc. In some examples, the image data may be encoded in any analog or digital format. Some non-limiting examples of such formats can include raw formats, compressed formats, uncompressed formats, irreversible formats, reversible formats, JPEG, GIF, PNG, TIFF, BMP, NTSC, PAL, SECAM, MPEG, MPEG-4 Part14, MOV, WMV, FLV, AVI, AVCHD, WebM, MKV, etc.

[0025] In some embodiments, the extended reality device can receive a digital signal from, for example, an input device. The term digital signal refers to a discrete-time series of digital values. The digital signal can correspond to, for example, sensor data, text data, audio data, video data, virtual data, or any other form of data that provides perceptible information. Consistent with the present disclosure, the digital signal can be configured to cause the extended reality device to present virtual content. In one embodiment, the virtual content may be presented in a selected orientation. In this embodiment, the digital signal can indicate the position and angle of the viewpoint in an environment such as an extended reality environment. Specifically, the digital signal can include an encoding of the position and angle in six degrees of freedom coordinates (e.g., forward / backward, up / down, left / right, yaw, pitch, and roll). In another embodiment, the digital signal can include an encoding of the position as three-dimensional coordinates (e.g., x, y, and z), and an encoding of the angle as a vector resulting from the encoded position. Specifically, the digital signal can indicate the orientation and angle of the presented virtual content in the absolute coordinates of the environment, for example, by encoding the yaw, pitch, and roll of the virtual content relative to a standard default angle. In another embodiment, the digital signal can indicate the orientation and angle of the presented virtual content relative to the viewpoint of another object (e.g., a virtual object, a physical object, etc.), for example, by encoding the yaw, pitch, and roll of the virtual content relative to the direction corresponding to the viewpoint or the direction corresponding to another object. In another embodiment, such a digital signal can include one or more projections of the virtual content in a format ready for presentation (e.g., an image, a video, etc.). For example, each such projection can correspond to a specific orientation or a specific angle. In another embodiment, the digital signal can include a display of the virtual content, for example, by encoding an object in a three-dimensional array of voxels, a polygon mesh, or any other format capable of presenting the virtual content.

[0026] In some embodiments, the digital signal may be configured to cause an extended reality device to present virtual content. The term virtual content can include any type of data display that can be presented to a user by an extended reality device. Virtual content can include virtual objects, virtual content of inanimate objects, virtual content of living organisms configured to change over time or in response to a trigger, virtual two-dimensional content, virtual three-dimensional content, a virtual overlay on a portion of a physical environment or a physical object, a virtual addition to a physical environment or a physical object, virtual promotional content, a virtual representation of a physical object, a virtual representation of a physical environment, a virtual document, a virtual character or persona, a virtual computer screen, a virtual widget, or any other format for presenting information virtually. Consistent with the present disclosure, virtual content can include any visual presentation rendered by a computer or processing device. In one embodiment, the virtual content includes virtual objects that are rendered by a computer within a limited area and are visual presentations configured to represent a particular type of object (e.g., virtual objects of inanimate objects, virtual objects of living organisms, virtual furniture, virtual decorative objects, virtual widgets, or other virtual representations, etc.). The rendered visual presentation can change, for example, to reflect a change to a state object or a change in the viewing angle of the object so as to mimic a change in the appearance of a physical object. In another embodiment, the virtual content can include a virtual display (also referred to herein as a "virtual display screen" or "virtual screen") such as a virtual computer screen, a virtual tablet screen, or a virtual smartphone screen, configured to display information generated by an operating system, and the operating system can be configured to receive text data from a physical keyboard and / or a virtual keyboard and cause the display of text content on the virtual display screen. In one example, as shown in FIG. 1, the virtual content can include a virtual environment that includes a virtual computer screen and a plurality of virtual objects.In some examples, the virtual display may be a virtual object that mimics and / or extends the functionality of a physical display screen. For example, the virtual display may be presented in an extended reality environment (e.g., a mixed reality environment, an augmented reality environment, a virtual reality environment, etc.) using an extended reality device. In one example, the virtual display can present content generated by a normal operating system that can be equally presented on a physical display screen. In one example, text content input using a keyboard (e.g., using a physical keyboard, using a virtual keyboard, etc.) can be presented on the virtual display in real time when the text content is input. In one example, a virtual cursor may be presented on the virtual display, and the virtual cursor may be controlled by a pointing device (a physical pointing device, a virtual pointing device, a computer mouse, a joystick, a touchpad, a physical touch controller, etc.). In one example, one or more windows of a graphical user interface operating system can be presented on the virtual display. In another example, the content presented on the virtual display may be interactive, i.e., the reaction to the user's actions may change. In yet another example, the presentation of the virtual display may or may not include the presentation of a screen frame.

[0027] Some of the disclosed embodiments include a data structure or database and / or can access a data structure or database. The terms data structure and database consistent with this disclosure can include any collection of data values and the relationships between them. The data can be stored linearly, horizontally, hierarchically, relationally, non-relationally, unidimensionally, multi-dimensionally, operationally, in an ordered manner, in an unordered manner, object-orientedly, centrally, non-centrally, distributively, decentralizedly, customarily, or in any way that enables data access. By way of non-limiting example, data structures can include arrays, associative arrays, linked lists, binary trees, balanced trees, heaps, stacks, queues, sets, hash tables, records, tagged unions, entity-relationship models, graphs, hypergraphs, matrices, tensors, and the like. For example, data structures can include XML databases, RDBMS databases, SQL databases, or NoSQL alternatives for data storage / search, such as MongoDB, Redis, Couchbase, Datastax Enterprise Graph, Elastic Search, Splunk, Solr, Cassandra, Amazon DynamoDB, Scylla, HBase, and Neo4J. A data structure can be a component of the disclosed system or a remote computing component (e.g., a cloud-based data structure). The data within a data structure can be stored in contiguous or non-contiguous memory. Further, a data structure does not require that information be located in the same place. A data structure can be distributed across multiple servers that can be owned or operated by the same or different entities. Thus, the term data structure in the singular includes multiple data structures.

[0028] In some embodiments, the system can determine a reliability level of the received input or any determined value. The term reliability level refers to a numerical value or any other indication of a level (e.g., within a predetermined range) that indicates the amount of reliability the system has in the determined data. For example, the reliability level can have a value between 1 and 10. Alternatively, the reliability level can be displayed as a percentage or any other numerical or non-numerical indication. In some cases, the system can compare the reliability level to a threshold. The term threshold can refer to a reference value, level, point, or range of values. During operation, if the reliability level of the determined data exceeds the threshold (depending on a particular use case, or is below it), the system can follow a first course of action, and if the reliability level is below the threshold (depending on a particular use case, or exceeds the threshold), the system can follow a second course of action. The value of the threshold can be predetermined for each type of object being inspected or can be dynamically selected based on different considerations.

[0029] System Overview Referring now to FIG. 1, which shows a user using an exemplary extended reality system consistent with various embodiments of the present disclosure. FIG. 1 is an exemplary display of only one embodiment, and it should be understood that some of the illustrated elements may be omitted and other elements may be added within the scope of the present disclosure. As shown in the figure, user 100 is sitting behind table 102 and supporting keyboard 104 and mouse 106. Keyboard 104 is connected by wire 108 to a wearable extended reality device 110 that displays virtual content to user 100. Instead of or in addition to wire 108, keyboard 104 can be connected wirelessly to wearable extended reality device 110. For purposes of illustration, the wearable extended reality device is shown as a pair of smart glasses, but as previously described, wearable extended reality device 110 can be any type of head-mounted device used to present extended reality to user 100. The virtual content displayed by wearable extended reality device 110 includes a virtual screen 112 (also referred to herein as a "virtual display screen" or "virtual display") and a plurality of virtual widgets 114. Virtual widgets 114A - 114D are displayed adjacent to virtual screen 112, and virtual widget 114E is displayed on table 102. User 100 can use keyboard 104 to input text into document 116 displayed on virtual screen 112 and can use mouse 106 to control virtual cursor 118. In one example, virtual cursor 118 can move anywhere within virtual screen 112. In another example, virtual cursor 118 can move anywhere within virtual screen 112 and can move to any of virtual widgets 114A - 114D, but not to virtual widget 114E. In yet another example, virtual cursor 118 can move anywhere within virtual screen 112 and can move to any of virtual widgets 114A - 114E.In a further example, the virtual cursor 118 can move anywhere in the extended reality environment including the virtual screen 112 and the virtual widgets 114A-114E. In yet another example, the virtual cursor can move only on all available surfaces (i.e., virtual surfaces or physical surfaces), or on selected surfaces within the extended reality environment. Alternatively or in addition, the user 100 can use hand gestures recognized by the wearable extended reality device 110 to interact with any of the virtual widgets 114A-114E, or with a selected virtual widget. For example, the virtual widget 114E can be an interactive widget (e.g., a virtual slider controller) that can be manipulated with hand gestures.

[0030] FIG. 2 shows an example of a system 200 that provides an extended reality (XR) experience to a user such as user 100. FIG. 2 is an exemplary display of one embodiment, and it should be understood that within the scope of the present disclosure, some of the illustrated elements may be omitted and other elements may be added. System 200 may be computer-based and may include computer system components, wearable devices, workstations, tablets, handheld computing devices, memory devices, and / or an internal network connecting the components. System 200 may include or be connected to various network computing resources (e.g., servers, routers, switches, network connections, storage devices, etc.) to support the services provided by system 200. Consistent with the present disclosure, system 200 can include an input unit 202, an XR unit 204, a mobile communication device 206, and a remote processing unit 208. The remote processing unit 208 can include a server 210 coupled to one or more physical or virtual storage devices such as data structure 212. System 200 may also include or be connected to a communication network 214 that facilitates communication and data exchange between different system components and different entities associated with system 200.

[0031] In accordance with the present disclosure, the input unit 202 can include one or more devices capable of receiving input from the user 100. In one embodiment, the input unit 202 can include a text input device such as a keyboard 104. The text input device can include all possible types of devices and mechanisms for inputting text information into the system 200. Examples of text input devices can include mechanical keyboards, membrane keyboards, flexible keyboards, QWERTY keyboards, Dvorak keyboards, Colemak keyboards, coded keyboards, wireless keyboards, keypads, key-based control panels, or other arrangements of control keys, visual input devices, or any other mechanism for inputting text, whether the mechanism is provided in a physical form or presented virtually. In one embodiment, the input unit 202 can also include a pointing input device such as a mouse 106. The pointing input device can include all possible types of devices and mechanisms for inputting two-dimensional or three-dimensional information into the system 200. In one example, two-dimensional input from the pointing input device may be used to interact with virtual content presented via the XR unit 204. Examples of pointing input devices can include computer mice, trackballs, touchpads, trackpads, touchscreens, joysticks, pointing sticks, styli, light pens, or any other physical or virtual input mechanism. In one embodiment, the input unit 202 can also include a graphical input device such as a touchscreen configured to detect contact, movement, or interruption of movement. The graphical input device can use any of a plurality of touch sensitivity technologies including, but not limited to, capacitive, resistive, infrared, and surface acoustic wave technologies, as well as other proximity sensor arrays or other elements for determining one or more contact points. In one embodiment, the input unit 202 can also include one or more voice input devices such as a microphone.The voice input device can include all possible types of devices and mechanisms for inputting voice data to facilitate voice-enabled functions such as voice recognition, voice replication, digital recording, and telephone functions. In one embodiment, the input unit 202 may also include one or more image input devices, such as an image sensor, configured to capture image data. In one embodiment, the input unit 202 may also include one or more tactile gloves configured to capture hand movement and gesture data. In one embodiment, the input unit 202 can also include one or more proximity sensors configured to detect the presence and / or movement of objects within a selected area near the sensor.

[0032] According to some embodiments, the system can include at least one sensor configured to detect and / or measure characteristics related to the user, the user's actions, or the user's environment. An example of the at least one sensor is sensor 216 included in input unit 202. Sensor 216 can be an accelerometer, a touch sensor, a light sensor, an infrared sensor, a voice sensor, an image sensor, a proximity sensor, a positioning sensor, a gyroscope, a temperature sensor, a biometric sensor, or any other sensing device to facilitate related functions. Sensor 216 may be integrated with the input device, or connected to the input device, or separated from the input device. In one example, a thermometer can be included in mouse 106 to determine the body temperature of user 100. In another example, a positioning sensor may be integrated with keyboard 104 to determine the movement of user 100 relative to keyboard 104. Such a positioning sensor can be implemented using one of Global Positioning System (GPS), GLObal Navigation Satellite System (GLONASS), Galileo Global Navigation Satellite System, BeiDou Navigation Satellite System, other Global Navigation Satellite Systems (GNSS), Indian Regional Navigation Satellite System (IRNSS), Local Positioning System (LPS), Real-Time Location System (RTLS), Indoor Positioning System (IPS), Wi-Fi-based positioning system, cellular triangulation, image-based positioning technology, indoor positioning technology, outdoor positioning technology, or any other positioning technology.

[0033] According to some embodiments, the system can include one or more sensors for identifying the position and / or movement of physical devices (physical input devices, physical computing devices, keyboard 104, mouse 106, wearable extended reality device 110, etc.). The one or more sensors may be included in the physical device or may be external to the physical device. In some examples, an image sensor external to the physical device (e.g., an image sensor included in another physical device) can be used to capture image data of the physical device, and the image data can be analyzed to identify the position and / or movement of the physical device. For example, the image data may be analyzed using a visual object tracking algorithm to identify the movement of the physical device, and may be analyzed using a visual object detection algorithm to identify the position of the physical device (e.g., relative to the image sensor, in a global coordinate system, etc.). In some examples, an image sensor included in the physical device can be used to capture image data, and the image data can be analyzed to identify the position and / or movement of the physical device. For example, the image data may be analyzed using a visual odometry algorithm to identify the position of the physical device, and may be analyzed using an egomotion algorithm to identify the movement of the physical device, etc. In some examples, a positioning sensor such as an indoor positioning sensor or an outdoor positioning sensor may be included in the physical device and may be used to determine the position of the physical device. In some examples, a movement sensor such as an accelerometer or a gyroscope may be included in the physical device and may be used to determine the movement of the physical device. In some examples, a physical device such as a keyboard or a mouse may be configured to be placed on a physical surface. Such a physical device can include an optical mouse sensor (also known as a non-mechanical tracking engine) directed towards the physical surface, and the output of the optical mouse sensor can be analyzed to determine the movement of the physical device relative to the physical surface.

[0034] In accordance with the present disclosure, the XR unit 204 can include a wearable extended reality device configured to present virtual content to the user 100. An example of the wearable extended reality device is the wearable extended reality device 110. Further examples of the wearable extended reality device can include virtual reality (VR) devices, augmented reality (AR) devices, mixed reality (MR) devices, or any other device capable of generating extended reality content. Some non-limiting examples of such devices can include Nreal Light, Magic Leap One, Varjo, Quest 1 / 2, Vive, and the like. In some embodiments, the XR unit 204 can present virtual content to the user 100. Generally, extended reality devices can include all real and virtual composite environments generated by computer technology and wearability, as well as the interaction between humans and machines. As described above, the term "extended reality" (XR) refers to a superset that includes the entire spectrum from "complete reality" to "complete virtuality". It includes representative forms such as augmented reality (AR), mixed reality (MR), virtual reality (VR), and the interpolated regions therebetween. Therefore, it should be noted that the terms "XR device", "AR device", "VR device", and "MR device" can be used interchangeably herein and can refer to any of the various devices listed above.

[0035] In accordance with the present disclosure, the system can exchange data with various communication devices related to a user, such as mobile communication device 206. The term "communication device" is intended to include all possible types of devices that can exchange data using a digital communication network, an analog communication network, or any other communication network configured to transmit data. In some examples, the communication device can include smartphones, tablets, smartwatches, personal digital assistants, desktop computers, laptop computers, IoT devices, dedicated terminals, wearable communication devices, and any other device that enables data communication. In some cases, mobile communication device 206 can supplement or replace input unit 202. Specifically, mobile communication device 206 may be associated with a physical touch controller that can function as a pointing input device. Further, mobile communication device 206 may also be used, for example, to implement a virtual keyboard and replace a text input device. For example, when user 100 leaves table 102 and walks to the break room with their smart glasses, they can receive an email that requires a quick response. In this case, the user can choose to use their smartwatch as an input device and enter a response to the email while it is being virtually presented by the smart glasses.

[0036] In accordance with the present disclosure, embodiments of the system can include the use of a cloud server. The term "cloud server" refers to a computer platform that provides services over a network such as the Internet. In the exemplary embodiment shown in FIG. 2, the server 210 can use virtual machines that do not have to correspond to individual hardware. For example, computing and / or memory capabilities may be implemented by allocating an appropriate portion of the desired computing / memory power from a scalable repository such as a data center or a distributed computing environment. Specifically, in one embodiment, the remote processing unit 208 can be used with the XR unit 204 to provide virtual content to the user 100. In one configuration example, the server 210 may be a cloud server that functions as an operating system (OS) of the wearable extended reality device. In one example, the server 210 can implement the methods described herein using customized hardwired logic, one or more application specific integrated circuits (ASICs), field programmable gate arrays (FPGAs), firmware, and / or program logic that combines with the computer system to make the server 210 a dedicated machine.

[0037] In some embodiments, server 210 can access data structure 212 to determine virtual content for displaying, for example, user 100. Data structure 212 can utilize volatile or non-volatile, magnetic, semiconductor, tape, optical, removable, non-removable, other types of storage devices or tangible or non-transitory computer-readable media, or any medium or mechanism for storing information. Data structure 212 can be part of server 210 as shown, or can be separate from server 210. If data structure 212 is not part of server 210, server 210 can exchange data with data structure 212 via a communication link. Data structure 212 can include one or more memory devices that store data and instructions used to execute one or more features of the disclosed method. In one embodiment, data structure 212 can include any of a plurality of suitable data structures, from a small data structure hosted on a workstation to a large data structure distributed among data centers. Data structure 212 can also include any combination of one or more data structures controlled by a memory controller device (e.g., a server) or software.

[0038] In accordance with the present disclosure, a communication network can be any type of network (including infrastructure) that supports communication, exchanges information, and / or facilitates the exchange of information between components of a system. For example, communication network 214 within system 200 can include, for example, a telephone network, an extranet, an intranet, the Internet, satellite communication, offline communication, wireless communication, transponder communication, a local area network (LAN), a wireless network (e.g., a Wi-Fi / 802.11 network), a wide area network (WAN), a virtual private network (VPN), a digital communication network, an analog communication network, or any other mechanism or combination of mechanisms that enables data transmission.

[0039] The components and arrangement of the system 200 shown in FIG. 2 are for illustrative purposes only, as the system components used to implement the disclosed processes and features may vary, and are not intended to limit any embodiments.

[0040] FIG. 3 is a block diagram showing a configuration example of the input unit 202. It should be understood that FIG. 3 is an exemplary representation of only one embodiment, and within the scope of the present disclosure, some of the illustrated elements may be omitted and other elements may be added. In the embodiment of FIG. 3, the input unit 202 can access directly or indirectly a bus 300 (or other communication mechanism) that interconnects subsystems and components for transferring information within the input unit 202. For example, the bus 300 can interconnect a memory interface 310, a network interface 320, an input interface 330, a power supply 340, an output interface 350, a processing device 360, a sensor interface 370, and a database 380.

[0041] The memory interface 310 shown in FIG. 3 can be used to access software products and / or data stored in a non-transitory computer-readable medium. Generally, a non-transitory computer-readable storage medium refers to any type of physical memory that can store information or data readable by at least one processor. Examples include random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, CD ROMs, DVDs, flash drives, disks, any other optical data storage medium, any physical medium with a pattern of holes, PROM, EPROM, FLASH-EPROM or any other flash memory, NVRAM, caches, registers, any other memory chip or cartridge, and network versions thereof. The terms "memory" and "computer-readable storage medium" can refer to multiple structures such as multiple memories or computer-readable storage media located within the input unit or at a remote location. Further, one or more computer-readable storage media can be utilized when implementing a computer-implemented method. Thus, the term computer-readable storage medium should be understood to include tangible items and exclude carrier waves and transient signals. In the particular embodiment shown in FIG. 3, the memory interface 310 can be used to access software products and / or data stored in a memory device such as the memory device 311. The memory device 311 may include high-speed random access memory and / or non-volatile memory, such as one or more magnetic disk storage devices, one or more optical storage devices, and / or flash memory (e.g., NAND, NOR). Consistent with the present disclosure, the components of the memory device 311 may be distributed among multiple units of the system 200 and / or multiple memory devices.

[0042] The memory device 311 shown in FIG. 3 can include software modules for executing processes consistent with this disclosure. In particular, the memory device 311 can include an input determination module 312, an output determination module 313, a sensor communication module 314, a virtual content determination module 315, a virtual content communication module 316, and a database access module 317. Modules 312-317 can include software instructions for execution by at least one processor (e.g., processing device 360) associated with the input unit 202. The input determination module 312, the output determination module 313, the sensor communication module 314, the virtual content determination module 315, the virtual content communication module 316, and the database access module 317 can cooperate to perform various operations. For example, the input determination module 312 may determine text using data received from, for example, the keyboard 104. Thereafter, the output determination module 313 can cause a presentation of the most recently input text on, for example, a dedicated display 352 physically or wirelessly coupled to the keyboard 104. In this way, when the user 100 inputs, a preview of the input text can be viewed without constantly moving the head up and down to view the virtual screen 112. The sensor communication module 314 can receive data from different sensors to determine the state of the user 100. Thereafter, the virtual content determination module 315 can determine the virtual content to be displayed based on the received input and the determined status of the user 100. For example, the determined virtual content may be a virtual presentation of the most recently input text on a virtual screen disposed virtually adjacent to the keyboard 104. The virtual content communication module 316 may obtain virtual content not determined by the virtual content determination module 315 (e.g., an avatar of another user). The search for virtual content may be from the database 380, from the remote processing unit 208, or from any other source.

[0043] In some embodiments, the input determination module 312 can adjust the operation of the input interface 330 to receive pointer input 331, text input 332, voice input 333, and XR-related input 334. Details of pointer input, text input, and voice input have been described above. The term "XR-related input" can include any type of data that may cause a change in the virtual content displayed to the user 100. In one embodiment, the XR-related input 334 can include image data of the user 100, a wearable extended reality device (e.g., a detected hand gesture of the user 100). In another embodiment, the XR-related input 334 can include wireless communication indicating the presence of another user in proximity to the user 100. Consistent with the present disclosure, the input determination module 312 can receive different types of input data simultaneously. Thereafter, the input determination module 312 can further apply different rules based on the type of input detected. For example, pointer input may be prioritized over voice input.

[0044] In some embodiments, the output determination module 313 can adjust the operation of the output interface 350 to generate an output using the optical indicator 351, the display 352, and / or the speaker 353. Generally, the output generated by the output determination module 313 does not include virtual content presented by the wearable extended reality device. Instead, the output generated by the output determination module 313 includes various outputs related to the operation of the input unit 202 and / or the XR unit 204. In one embodiment, the optical indicator 351 can include an optical indicator indicating the state of the wearable extended reality device. For example, the optical indicator can display green light when the wearable extended reality device 110 is connected to the keyboard 104 and can blink when the battery of the wearable extended reality device 110 is low. In another embodiment, the display 352 can be used to display operation information. For example, the display can present an error message when the wearable extended reality device is inoperable. In another embodiment, the speaker 353 can be used to output sound, for example, when the user 100 wants to play music for another user.

[0045] In some embodiments, the sensor communication module 314 can adjust the operation of the sensor interface 370 to receive sensor data from one or more sensors integrated with or connected to an input device. The one or more sensors can include an audio sensor 371, an image sensor 372, a motion sensor 373, an environmental sensor 374 (e.g., a temperature sensor, an ambient light detector, etc.), and other sensors 375. In one embodiment, the data received from the sensor communication module 314 can be used to determine the physical orientation of the input device. The physical orientation of the input device can indicate the user's state and can be determined based on a combination of tilt movement, roll movement, and pan movement. Thereafter, the physical orientation of the input device can be used by the virtual content determination module 315 to change the display parameters of the virtual content to match the user's state (e.g., attentive, sleepy, active, sitting, standing, leaning back, leaning forward, walking, moving, riding, etc.).

[0046] In some embodiments, the virtual content determination module 315 can determine the virtual content to be displayed by the wearable extended reality device. The virtual content can be determined based on data from the input determination module 312, the sensor communication module 314, and other sources (e.g., the database 380). In some embodiments, determining the virtual content can include determining the distance, size, and orientation of virtual objects. The determination of the position of the virtual object may be determined based on the type of the virtual object. Specifically, with respect to the example shown in FIG. 1, since the virtual widget 114E is a virtual controller (e.g., a volume bar), the virtual content determination module 315 can determine to arrange four virtual widgets 114A-114D on both sides of the virtual screen 112 and arrange the virtual widget 114E on the table 102. The determination of the position of the virtual object may be further determined based on the user's preference. For example, in the case of a left-handed user, the virtual content determination module 315 may determine to arrange the virtual volume bar to the left of the keyboard 104, and in the case of a right-handed user, the virtual content determination module 315 may determine to arrange the virtual volume bar to the right of the keyboard 104.

[0047] In some embodiments, the virtual content communication module 316 may adjust the operation of the network interface 320 to obtain data from one or more sources to be presented to the user 100 as virtual content. The one or more sources can include other XR units 204, the user's mobile communication device 206, a remote processing unit 208, publicly available information, and the like. In one embodiment, the virtual content communication module 316 may communicate with the mobile communication device 206 to provide a virtual display of the mobile communication device 206. For example, the virtual display can enable the user 100 to read messages and interact with applications installed on the mobile communication device 206. The virtual content communication module 316 may also adjust the operation of the network interface 320 to share virtual content with other users. In one example, the virtual content communication module 316 may use data from the input determination module to identify a trigger (e.g., the trigger can include a user gesture) and transfer the content from the virtual display to a physical display (e.g., a TV) or a virtual display of a different user.

[0048] In some embodiments, the database access module 317 can cooperate with the database 380 to retrieve stored data. The retrieved data can include, for example, privacy levels associated with different virtual objects, relationships between virtual and physical objects, user preferences, the user's past behavior, and the like. As described above, the virtual content determination module 315 can use the data stored in the database 380 to determine virtual content. The database 380 can include separate databases, such as, for example, a vector database, a raster database, a tile database, a viewport database, and / or a user input database. The data stored in the database 380 may be received from the modules 314 - 317 or other components of the system 200. Further, the data stored in the database 380 may be provided as input using data input, data transfer, or data upload.

[0049] Modules 312-317 may be implemented in software, hardware, firmware, any combination thereof, etc. In some embodiments, any one or more of the data related to modules 312-317 and database 380 may be stored in XR unit 204, mobile communication device 206, or remote processing unit 208. The processing device of system 200 may be configured to execute the instructions of modules 312-317. In some embodiments, aspects of modules 312-317 may be executable by one or more processors, alone or in various combinations with each other, in hardware, software (including one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof. Specifically, modules 312-317 may be configured to interact with each other and / or with other modules of system 200 to perform functions consistent with some of the disclosed embodiments. For example, input unit 202 may execute instructions including an image processing algorithm on data from XR unit 204 to determine the movement of user 100's head. Further, throughout this specification, each function described with respect to input unit 202, or with respect to the components of input unit 202, may correspond to a set of instructions for performing said function. These instructions need not be implemented as separate software programs, procedures, or modules. Memory device 311 may include additional modules and instructions or fewer modules and instructions. For example, memory device 311 may store an operating system such as ANDROID, iOS, UNIX, OSX, WINDOWS, DARWIN, RTXC, LINUX, etc., or an embedded operating system such as VXWorkS. The operating system may process basic system services and include instructions for performing hardware-dependent tasks.

[0050] The network interface 320 shown in FIG. 3 can provide bidirectional data communication to a network such as the communication network 214. In one embodiment, the network interface 320 can include an Integrated Services Digital Network (ISDN) card, a cellular modem, a satellite modem, or a modem to provide a data communication connection via the Internet. As another example, the network interface 320 can include a Wireless Local Area Network (WLAN) card. In another embodiment, the network interface 320 can include an Ethernet port connected to a radio frequency receiver and transmitter and / or an optical (e.g., infrared) receiver and transmitter. The specific design and implementation of the network interface 320 can depend on the communication network in which the input unit 202 is intended to operate. For example, in some embodiments, the input unit 202 can include a network interface 320 designed to operate via a GSM network, a GPRS network, an EDGE network, a Wi-Fi or WiMax network, and a Bluetooth network. In any such implementation, the network interface 320 can be configured to transmit and receive electrical, electromagnetic, or optical signals that carry digital data streams or digital signals representing various types of information.

[0051] The input interface 330 shown in FIG. 3 can receive inputs from various input devices, such as a keyboard, a mouse, a touchpad, a touch screen, one or more buttons, a joystick, a microphone, an image sensor, and any other device configured to detect physical or virtual inputs. The received input may be in at least one form of text, voice, sound, hand gesture, body gesture, tactile information, and any other type of physical or virtual input generated by the user. In the illustrated embodiment, the input interface 330 can receive pointer input 331, text input 332, voice input 333, and XR-related input 334. In a further embodiment, the input interface 330 may be an integrated circuit that can function as a bridge between the processing device 360 and any of the above input devices.

[0052] The power supply 340 shown in FIG. 3 can supply electrical energy to the power input unit 202 and optionally can also supply power to the XR unit 204. In general, the power supply included in any device or system of the present disclosure can be any device that can repeatedly store, distribute, or transmit power, including but not limited to one or more batteries (e.g., lead-acid battery, lithium-ion battery, nickel-metal hydride battery, nickel-cadmium battery), one or more capacitors, one or more connections to an external power source, one or more power converters, or any combination thereof. Referring to the example shown in FIG. 3, the power supply may be portable, which means that the input unit 202 can be easily carried by hand (e.g., the total weight of the power supply 340 is,). The portability of the power supply enables the user 100 to use the input unit 202 in various situations. In other embodiments, the power supply 340 may be associated with a connection to an external power source (such as a power grid) that can be used to charge the power supply 340. Further, the power supply 340 may be configured to charge one or more batteries included in the XR unit 204. For example, a pair of extended reality glasses (e.g., wearable extended reality device 110) can be charged (e.g., wirelessly or non-wirelessly) when they are placed on or near the input unit 202.

[0053] The output interface 350 shown in FIG. 4 can cause output from various output devices, for example, using an optical indicator 351, a display 352, and / or a speaker 353. In one embodiment, the output interface 350 may be an integrated circuit that can function as a bridge between the processing device 360 and at least one of the above output devices. The optical indicator 351 can include one or more light sources, such as an LED array associated with different colors, for example. The display 352 can include a screen (e.g., an LCD or a dot matrix screen) or a touch screen. The speaker 353 can include audio headphones, a hearing aid type device, a speaker, bone conduction headphones, an interface that provides tactile cues, a vibration stimulation device, and the like.

[0054] The processing device 360 shown in FIG. 3 can include at least one processor configured to execute a computer program, application, method, process, or other software to implement the embodiments described in this disclosure. Generally, the processing device included in any device or system of this disclosure can include one or more integrated circuits, microchips, microcontrollers, microprocessors, central processing devices (CPUs), graphics processing devices (GPUs), digital signal processors (DSPs), all or part of a field programmable gate array (FPGA), or other circuits suitable for executing instructions or performing logical operations. The processing device can include at least one processor configured to implement the functions of the disclosed methods, such as a microprocessor manufactured by Intel (trademark). The processing device can include a single core or a multi-core processor that simultaneously executes parallel processing. In one example, the processing device may be a single-core processor configured with virtual processing technology. The processing device can implement virtual machine technology or other technologies to provide the ability to execute, control, execute, operate, store, etc. multiple software processes, applications, programs, etc. In another example, the processing device can include a multi-core processor configuration (e.g., dual, quad-core, etc.) configured to provide a parallel processing function that enables related devices of the processing device to execute multiple processes simultaneously. It should be understood that other types of processor configurations can be implemented to provide the functions disclosed herein.

[0055] The sensor interface 370 shown in FIG. 3 can obtain sensor data from various sensors, such as a voice sensor 371, an image sensor 372, a motion sensor 373, an environmental sensor 374, and other sensors 375. In one embodiment, the sensor interface 370 may be an integrated circuit that can function as a bridge between the processing device 360 and at least one of the above sensors.

[0056] The audio sensor 371 may include one or more audio sensors configured to capture audio by converting the audio into digital information. Some examples of audio sensors can include a microphone, a unidirectional microphone, a bidirectional microphone, a cardioid microphone, an omnidirectional microphone, an on-board microphone, a wired microphone, a wireless microphone, or any combination of the above. Consistent with the present disclosure, the processing device 360 can modify the presentation of virtual content based on data received from the audio sensor 371 (e.g., an audio command).

[0057] The image sensor 372 can include one or more image sensors configured to capture visual information by converting light into image data. Consistent with the present disclosure, the image sensor may be included in any device or system in the present disclosure and may be any device capable of detecting optical signals in the near-infrared, infrared, visible, and ultraviolet spectra and converting them into electrical signals. Examples of image sensors can include digital cameras, phone cameras, semiconductor charge-coupled devices (CCDs), complementary metal-oxide-semiconductor (CMOS), or N-type metal-oxide-semiconductor (NMOS, live MOS) active pixel sensors. The electrical signals can be used to generate image data. Consistent with the present disclosure, the image data can include pixel data streams, digital images, digital video streams, data derived from captured images, and data used to construct one or more 3D images, sequences of 3D images, 3D videos, or virtual 3D displays. The image data acquired by the image sensor 372 can be transmitted to any processing device of the system 200 by wired or wireless transmission. For example, the image data can be processed for detecting an object, detecting an event, detecting an action, detecting a face, detecting a person, recognizing a known person, or any other information that can be used by the system 200. Consistent with the present disclosure, the processing device 360 can change the presentation of virtual content based on the image data received from the image sensor 372.

[0058] The motion sensor 373 can include one or more motion sensors configured to measure the motion of the input unit 202 or the motion of an object within the environment of the input unit 202. Specifically, the motion sensor can perform at least one of detecting the motion of an object within the environment of the input unit 202, measuring the speed of an object within the environment of the input unit 202, measuring the acceleration of an object within the environment of the input unit 202, detecting the motion of the input unit 202, measuring the speed of the input unit 202, measuring the acceleration of the input unit 202, etc. In some embodiments, the motion sensor 373 can include one or more accelerometers configured to detect a change in appropriate acceleration of the input unit 202 and / or measure appropriate acceleration. In other embodiments, the motion sensor 373 can include one or more gyroscopes configured to detect a change in the orientation of the input unit 202 and / or measure information regarding the orientation of the input unit 202. In other embodiments, the motion sensor 373 can include one or more that use an image sensor, a LIDAR sensor, a radar sensor, or a proximity sensor. For example, by analyzing the captured images, the processing device can determine the motion of the input unit 202, for example, using an egomotion algorithm. Further, the processing device can determine the motion of an object within the environment of the input unit 202, for example, using an object tracking algorithm. Consistent with the present disclosure, the processing device 360 can change the presentation of the virtual content based on the determined motion of the input unit 202 or the determined motion of an object within the environment of the input unit 202. For example, make the virtual display follow the movement of the input unit 202.

[0059] The environmental sensor 374 can include one or more sensors of different types configured to capture data reflecting the environment of the input unit 202. In some embodiments, the environmental sensor 374 is configured to perform at least one of measuring chemical properties within the environment of the input unit 202, measuring changes in chemical properties within the environment of the input unit 202, detecting the presence of chemicals within the environment of the input unit 202, and measuring the concentration of chemicals within the environment of the input unit 202, and can include one or more chemical sensors. Examples of such chemical properties can include pH level, toxicity, and temperature. Examples of such chemicals can include electrolytes, specific enzymes, specific hormones, specific proteins, smoke, carbon dioxide, carbon monoxide, oxygen, ozone, hydrogen, and hydrogen sulfide. In other embodiments, the environmental sensor 374 can include one or more temperature sensors configured to detect changes in the temperature of the environment of the input unit 202 and / or measure the temperature of the environment of the input unit 202. In other embodiments, the environmental sensor 374 can include one or more barometers configured to detect changes in the air pressure within the environment of the input unit 202 and / or measure the air pressure within the environment of the input unit 202. In other embodiments, the environmental sensor 374 can include one or more light sensors configured to detect changes in ambient light within the environment of the input unit 202. Consistent with the present disclosure, the processing device 360 can modify the presentation of virtual content based on the input from the environmental sensor 374. For example, automatically reducing the brightness of virtual content when the environment of user 100 becomes dark.

[0060] The other sensor 375 can include a weight sensor, a light sensor, a resistance sensor, an ultrasonic sensor, a proximity sensor, a biometric sensor, or other detection devices for facilitating related functions. In certain embodiments, the other sensor 375 can include one or more positioning sensors configured to obtain positioning information of the input unit 202, detect a change in the position of the input unit 202, and / or measure the position of the input unit 202. Alternatively, the GPS software can enable the input unit 202 to access an external GPS receiver (e.g., connected via a serial port or Bluetooth). Consistent with the present disclosure, the processing device 360 can change the presentation of the virtual content based on the input from the other sensor 375. For example, personal information is presented only after identifying the user 100 using data from a biometric sensor.

[0061] The components and arrangements shown in FIG. 3 are not intended to limit any embodiments. As will be understood by those skilled in the art having the benefit of this disclosure, numerous modifications and / or changes can be made to the illustrated configuration of the input unit 202. For example, not all components are essential for the operation of the input unit. Any component can be arranged in any suitable part of the input unit, and the components can be rearranged in various configurations while providing the functions of the various embodiments. For example, some input units may not include all of the elements as shown in the input unit 202.

[0062] FIG. 4 is a block diagram showing a configuration example of the XR unit 204. FIG. 4 is an exemplary display of one embodiment, and it should be understood that within the scope of the present disclosure, some of the illustrated elements may be omitted and other elements may be added. In the embodiment of FIG. 4, the XR unit 204 may directly or indirectly access a bus 400 (or other communication mechanism) that interconnects subsystems and components for transferring information within the XR unit 204. For example, the bus 400 may interconnect a memory interface 410, a network interface 420, an input interface 430, a power supply 440, an output interface 450, a processing device 460, a sensor interface 470, and a database 480.

[0063] The memory interface 410 shown in FIG. 4 is assumed to have the same functions as those of the memory interface 310 described in detail above. The memory interface 410 can be used to access software products and / or data stored in a non-transitory computer-readable medium or a memory device such as the memory device 411. The memory device 411 can include software modules for executing processes consistent with the present disclosure. In particular, the memory device 411 can include an input determination module 412, an output determination module 413, a sensor communication module 414, a virtual content determination module 415, a virtual content communication module 416, and a database access module 417. Modules 412-417 may include software instructions for execution by at least one processor (e.g., the processing device 460) associated with the XR unit 204. The input determination module 412, the output determination module 413, the sensor communication module 414, the virtual content determination module 415, the virtual content communication module 416, and the database access module 417 can cooperate to perform various operations. For example, the input determination module 412 can determine a user interface (UI) input received from the input unit 202. At the same time, the sensor communication module 414 can receive data from different sensors to determine the state of the user 100. Based on the received input and the determined state of the user 100, the virtual content determination module 415 can determine the virtual content to be displayed. The virtual content communication module 416 can search for virtual content not determined by the virtual content determination module 415. The search for virtual content may be from the database 380, the database 480, the mobile communication device 206, or the remote processing unit 208. Based on the output of the virtual content determination module 415, the output determination module 413 can cause a change in the virtual content displayed to the user 100 by the projector 454.

[0064] In some embodiments, the input determination module 412 can adjust the operation of the input interface 430 to receive gesture input 431, virtual input 432, voice input 433, and UI input 434. Consistent with the present disclosure, the input determination module 412 can receive different types of input data simultaneously. In one embodiment, the input determination module 412 can apply different rules based on the type of detected input. For example, gesture input can be prioritized over virtual input. In some embodiments, the output determination module 413 can adjust the operation of the output interface 450 to generate an output using the optical indicator 451, the display 452, the speaker 453, and the projector 454. In one embodiment, the optical indicator 451 can include an optical indicator indicating the state of the wearable extended reality device. For example, the optical indicator can display green light when the wearable extended reality device 110 is connected to the input unit 202 and can blink when the battery of the wearable extended reality device 110 is low. In another embodiment, the display 452 can be used to display operation information. In another embodiment, the speaker 453 can include bone conduction headphones used to output sound to the user 100. In another embodiment, the projector 454 can present virtual content to the user 100.

[0065] The operations of the sensor communication module, the virtual content determination module, the virtual content communication module, and the database access module have been described above with reference to FIG. 3, and the details thereof will not be repeated here. The modules 412-417 may be implemented in software, hardware, firmware, a mixture of any of them, and the like.

[0066] The network interface 420 shown in FIG. 4 is assumed to have the same functions as the network interface 320 described in detail above. The specific design and implementation of the network interface 420 may depend on the communication network in which the XR unit 204 is intended to operate. For example, in some embodiments, the XR unit 204 is configured to be selectively connectable to the input unit 202 by wire. When connected by wire, the network interface 420 can enable communication with the input unit 202, and when not connected by wire, the network interface 420 can enable communication with the mobile communication device 206.

[0067] The input interface 430 shown in FIG. 4 is assumed to have the same functions as the input interface 330 described in detail above. In this case, the input interface 430 may communicate with an image sensor to obtain a gesture input 431 (e.g., the finger of user 100 pointing to a virtual object), communicate with another XR unit 204 to obtain a virtual input 432 (e.g., a virtual object shared with the XR unit 204, or a gesture of an avatar detected in the virtual environment), communicate with a microphone to obtain an audio input 433 (e.g., an audio command), and communicate with the input unit 202 to obtain a UI input 434 (e.g., virtual content determined by the virtual content determination module 315).

[0068] The power supply 440 shown in FIG. 4 is assumed to have the same functions as the power supply 340 described above, and supplies electrical energy to supply power to the XR unit 204. In some embodiments, the power supply 440 may be charged by the power supply 340. For example, the power supply 440 may be wirelessly changed when the XR unit 204 is placed on or near the input unit 202.

[0069] The output interface 450 shown in FIG. 4 is assumed to have the same functions as the output interface 350 described in detail above. In this case, the output interface 450 can cause outputs from the optical indicator 451, the display 452, the speaker 453, and the projector 454. The projector 454 may be any device, apparatus, instrument, etc. that can project (or direct) light to display virtual content on a surface. The surface may be part of the XR unit 204, part of the user 100's eye, or part of an object proximate to the user 100. In one embodiment, the projector 454 can include an illumination unit that focuses light within a limited solid angle by one or more mirrors and lenses and provides a high value of luminous intensity in a defined direction.

[0070] The processing device 460 shown in FIG. 4 is assumed to have the same functions as the processing device 360 described in detail above. When the XR unit 204 is connected to the input unit 202, the processing device 460 may cooperate with the processing device 360. Specifically, the processing device 460 can implement virtual machine technology or other technologies to provide the ability to execute, control, run, operate, store, etc. a plurality of software processes, applications, programs, etc. It should be understood that other types of processor configurations can be implemented to provide the capabilities disclosed herein.

[0071] The sensor interface 470 shown in FIG. 4 is assumed to have the same functions as the sensor interface 370 described in detail above. Specifically, the sensor interface 470 can communicate with the voice sensor 471, the image sensor 472, the motion sensor 473, the environment sensor 474, and other sensors 475. The operations of the voice sensor, the image sensor, the motion sensor, the environment sensor, and other sensors have been described above with reference to FIG. 3, and the details are not repeated here. It is understood that other types and combinations of sensors may be used to provide the capabilities disclosed herein.

[0072] The components and arrangements shown in FIG. 4 are not intended to limit any embodiment. As will be understood by those skilled in the art having the benefit of this disclosure, numerous variations and / or modifications can be made to the configuration of the illustrated XR unit 204. For example, not all components are necessarily essential for the operation of the XR unit 204 in all cases. Any component may be disposed in any suitable part of the system 200, and the components may be rearranged in various configurations while providing the functions of various embodiments. For example, some XR units may not include all of the elements within the XR unit 204 (e.g., the wearable extended reality device 110 may not have the optical indicator 451).

[0073] FIG. 5 is a block diagram showing a configuration example of the remote processing unit 208. It should be understood that FIG. 5 is an exemplary representation of one embodiment and that within the scope of this disclosure, some of the illustrated elements may be omitted and other elements may be added. In the embodiment of FIG. 5, the remote processing unit 208 can include a server 210 that directly or indirectly accesses a bus 500 (or other communication mechanism) that interconnects subsystems and components for transferring information within the server 210. For example, the bus 500 can interconnect a memory interface 510, a network interface 520, a power supply 540, a processing device 560, and a database 580. The remote processing unit 208 can also include one or more data structures. For example, data structures 212A, 212B, and 212C.

[0074] The memory interface 510 shown in FIG. 5 is assumed to have the same functions as the memory interface 310 described in detail above. The memory interface 510 can be used to access software products and / or data stored in a non-transitory computer-readable medium or in other memory devices such as the memory devices 311, 411, 511, or data structures 212A, 212B, and 212C. The memory device 511 can include software modules for executing processes consistent with the present disclosure. In particular, the memory device 511 can include a shared memory module 512, a node registration module 513, a load distribution module 514, one or more computing nodes 515, an internal communication module 516, an external communication module 517, and a database access module (not shown). The modules 512-517 can include software instructions for execution by at least one processor (e.g., the processing device 560) associated with the remote processing unit 208. The shared memory module 512, the node registration module 513, the load distribution module 514, the computing module 515, and the external communication module 517 can cooperate to perform various operations.

[0075] The shared memory module 512 can enable information sharing between the remote processing unit 208 and other components of the system 200. In some embodiments, the shared memory module 512 can be configured to allow the processing device 560 (and other processing devices within the system 200) to access, retrieve, and store data. For example, using the shared memory module 512, the processing device 560 can perform at least one of the steps of executing a software program stored in the memory device 511, the database 580, or the data structures 212A-C, storing information in the memory device 511, the database 580, or the data structures 212A-C, or retrieving information from the memory device 511, the database 580, or the data structures 212A-C.

[0076] The node registration module 513 can be configured to track the availability of one or more compute nodes 515. In some examples, the node registration module 513 may be implemented as a software program, such as a software program executed by one or more compute nodes 515, a hardware solution, or a software and hardware combined solution. In some implementations, the node registration module 513 can communicate with one or more compute nodes 515, for example, using the internal communication module 516. In some examples, one or more compute nodes 515 can notify their status to the node registration module 513 by sending messages, for example, at startup, shutdown, at regular intervals, at selected times, in response to queries received from the node registration module 513, or at any other determined time. In some examples, the node registration module 513 can query the status of one or more compute nodes 515 by sending messages, for example, at startup, at regular intervals, at selected times, or at any other determined time.

[0077] The load distribution module 514 can be configured to divide the workload among one or more computing nodes 515. In some examples, the load distribution module 514 may be implemented as a software program such as a software program to be executed by one or more of the computing nodes 515, a hardware solution, or a combined software and hardware solution. In some implementations, the load distribution module 514 can interact with the node registration module 513 to obtain information regarding the availability of one or more computing nodes 515. In some implementations, the load distribution module 514 can communicate with one or more computing nodes 515, for example, using the internal communication module 516. In some examples, one or more computing nodes 515 can notify the load distribution module 514 of their status by sending a message, for example, at startup, shutdown, at regular intervals, at a selected time, in response to a query received from the load distribution module 514, or at any other determined time. In some examples, the load distribution module 514 can query the status of one or more computing nodes 515 by sending a message, for example, at startup, at regular intervals, at a preselected time, or at any other determined time.

[0078] The internal communication module 516 may be configured to receive and / or transmit information from one or more components of the remote processing unit 208. For example, control signals and / or synchronization signals can be transmitted and / or received via the internal communication module 516. In one embodiment, input information of a computer program, output information of a computer program, and / or intermediate information of a computer program can be transmitted and / or received via the internal communication module 516. In another embodiment, the information received via the internal communication module 516 may be stored in the memory device 511, the database 580, the data structures 212A - C, or other memory devices within the system 200. For example, the information obtained from the data structure 212A may be transmitted using the internal communication module 516. In another example, input data can be received using the internal communication module 516 and stored in the data structure 212B.

[0079] The external communication module 517 may be configured to receive and / or transmit information from one or more components of the system 200. For example, control signals can be transmitted and / or received via the external communication module 517. In one embodiment, the information received via the external communication module 517 can be stored in the memory device 511, the database 580, the data structures 212A - C, and / or any memory device within the system 200. In another embodiment, the information retrieved from any of the data structures 212A - C may be transmitted to the XR unit 204 using the external communication module 517. In another embodiment, input data can be transmitted and / or received using the external communication module 517. Examples of such input data can include data received from the input unit 202, information captured from the environment of the user 100 using one or more sensors (e.g., audio sensor 471, image sensor 472, motion sensor 473, environmental sensor 474, other sensor 475), etc.

[0080] In some embodiments, aspects of modules 512 - 517 may be implemented in hardware, software (including one or more signal processing and / or application specific integrated circuits), firmware, or any combination thereof, executable by one or more processors, either alone or in various combinations with each other. Specifically, modules 512 - 517 may be configured to interact with each other and / or with other modules of system 200 to perform functions consistent with embodiments of the present disclosure. Memory device 511 can include additional modules and instructions or fewer modules and instructions.

[0081] The network interface 520, power supply 540, processing device 560, and database 580 shown in FIG. 5 are assumed to have functions similar to those of the similar elements described above with reference to FIGS. 4 and 5. The specific design and implementation of the foregoing components may vary based on the embodiment of system 200. Further, remote processing unit 208 can include more or fewer components. For example, remote processing unit 208 can include an input interface configured to receive input directly from one or more input devices.

[0082] In accordance with the present disclosure, a processing device of system 200 (e.g., a processor within mobile communication device 206, a processor within server 210, a processor within a wearable extended reality device such as wearable extended reality device 110, and / or a processor within an input device associated with wearable extended reality device 110 such as keyboard 104) can use machine learning algorithms to implement any of the methods disclosed herein. In some embodiments, a machine learning algorithm (also referred to as a machine learning model in the present disclosure) can be trained using training examples, as described, for example, below. Some non-limiting examples of such machine learning algorithms can include classification algorithms, data regression algorithms, image segmentation algorithms, visual detection algorithms (e.g., object detectors, face detectors, person detectors, motion detectors, edge detectors, etc.), visual recognition algorithms (e.g., face recognition, person recognition, object recognition, etc.), speech recognition algorithms, mathematical embedding algorithms, natural language processing algorithms, support vector machines, random forests, nearest neighbor algorithms, deep learning algorithms, artificial neural network algorithms, convolutional neural network algorithms, recurrent neural network algorithms, linear machine learning models, non-linear machine learning models, ensemble algorithms, and the like. For example, a trained machine learning algorithm can include an inference model such as a prediction model, a classification model, a data regression model, a clustering model, a segmentation model, an artificial neural network (e.g., a deep neural network, a convolutional neural network, a recurrent neural network, etc.), a random forest, a support vector machine, and the like. In some examples, a training example can include an input of the example and a desired output corresponding to the input of the example. Further, in some examples, a machine learning algorithm trained using training examples can generate a trained machine learning algorithm, and the trained machine learning algorithm can be used to estimate an output for an input not included in the training examples.In some examples, the engineers, scientists, processes, and machines that train machine learning algorithms can further use validation examples and / or test examples. For example, the validation examples and / or test examples can include exemplary inputs along with the desired outputs corresponding to the exemplary inputs, and the trained machine learning algorithm and / or the intermediate trained machine learning algorithm can be used to estimate the outputs of the exemplary inputs of the validation examples and / or test examples, the estimated outputs can be compared with the corresponding desired outputs, and the trained machine learning algorithm and / or the intermediate trained machine learning algorithm can be evaluated based on the results of the comparison. In some examples, the machine learning algorithm can have parameters and hyperparameters, the hyperparameters can be set manually by a person or automatically by a process external to the machine learning algorithm (such as a hyperparameter search algorithm), and the parameters of the machine learning algorithm can be set by the machine learning algorithm based on the training examples. In some implementations, the hyperparameters may be set based on the training examples and the validation examples, and the parameters may be set based on the training examples and the selected hyperparameters. For example, given the hyperparameters, the parameters may be conditionally independent from the validation examples.

[0083] In some embodiments, a trained machine learning algorithm (also referred to herein as a machine learning model and a trained machine learning model) can be used to analyze an input and generate an output, for example, as described below. In some examples, a trained machine learning algorithm can be used as an inference model that generates an inferred output when an input is provided. For example, a trained machine learning algorithm can include a classification algorithm, the input can include samples, and the inferred output can include the classification of the samples (e.g., a predicted label, a predicted tag, etc.). In another example, a trained machine learning algorithm can include a regression model, the input can include samples, and the inferred output can include an inferred value corresponding to the samples. In yet another example, a trained machine learning algorithm can include a clustering model, the input can include samples, and the inferred output can include the assignment of the samples to at least one cluster. In a further example, a trained machine learning algorithm can include a classification algorithm, the input can include an image, and the inferred output can include the classification of the item depicted in the image. In yet another example, a trained machine learning algorithm can include a regression model, the input can include an image, and the inferred output can include an inferred value corresponding to the item depicted in the image (e.g., an estimated characteristic of the item such as the size, volume, age of a person depicted in the image, the distance from an item depicted in the image, etc.). In a further example, a trained machine learning algorithm can include an image segmentation model, the input can include an image, and the inferred output can include the segmentation of the image. In yet another example, a trained machine learning algorithm can include an object detector, the input can include an image, and the inferred output can include one or more detected objects in the image and / or one or more positions of the objects in the image.In some examples, a trained machine learning algorithm can include one or more equations and / or one or more functions and / or one or more rules and / or one or more procedures, an input can be used as an input to the equation and / or function and / or rule and / or procedure, and an inferred output can be based on the output of the equation and / or function and / or rule and / or procedure (e.g., selecting one of the outputs of the equation and / or function and / or rule and / or procedure, using a statistical measure of the output of the equation and / or function and / or rule and / or procedure, etc.).

[0084] In accordance with the present disclosure, the processing device of the system 200 can analyze the image data captured by an image sensor (e.g., image sensor 372, image sensor 472, or any other image sensor) to implement any of the methods disclosed herein. In some embodiments, analyzing the image data can include analyzing the image data to obtain preprocessed image data and then analyzing the image data and / or the preprocessed image data to obtain a desired result. As will be appreciated by those skilled in the art, the following are examples, and the image data can be preprocessed using other types of preprocessing methods. In some examples, the image data can be preprocessed by using a conversion function to convert the image data to obtain converted image data, and the preprocessed image data can include the converted image data. For example, the converted image data can include one or more convolutions of the image data. For example, the conversion function can include one or more image filters such as a low-pass filter, a high-pass filter, a band-pass filter, an all-pass filter, etc. In some examples, the conversion function can include a non-linear function. In some examples, the image data can be preprocessed by smoothing at least a portion of the image data, for example, using Gaussian convolution, using a median filter, etc. In some examples, the image data can be preprocessed to obtain a different representation of the image data. For example, the preprocessed image data can include a representation of at least a portion of the image data in the frequency domain, a discrete Fourier transform of at least a portion of the image data, a discrete wavelet transform of at least a portion of the image data, a time / frequency representation of at least a portion of the image data, a low-dimensional representation of at least a portion of the image data, an irreversible representation of at least a portion of the image data, a lossless representation of at least a portion of the image data, a time series of any of the above, any combination of the above, etc. In some examples, the image data can be preprocessed to extract edges, and the preprocessed image data can include information based on and / or related to the extracted edges. In some examples, the image data can be preprocessed to extract image features from the image data.Some non-limiting examples of such image features can include information based on and / or related to edges, corners, blobs, ridges, Scale-Invariant Feature Transform (SIFT) features, temporal features, and the like. In some examples, analyzing the image data can include calculating at least one convolution of at least a portion of the image data and using the at least one calculated convolution to calculate at least one result value and / or to perform a determination, identification, recognition, classification, etc.

[0085] In accordance with another aspect of the present disclosure, the processing device of system 200 can analyze image data to implement any of the methods disclosed herein. In some embodiments, analyzing the image can include analyzing the image data and / or pre-processed image data using one or more rules, functions, procedures, artificial neural networks, object detection algorithms, face detection algorithms, visual event detection algorithms, behavior detection algorithms, motion detection algorithms, background subtraction algorithms, inference models, and the like. Some non-limiting examples of such inference models can include results for training examples of training algorithms such as manually pre-programmed inference models, classification models, regression models, machine learning algorithms, and / or deep learning algorithms, where the training examples can include examples of data instances, and in some cases, the data instances can be labeled with corresponding desired labels and / or results, etc. In some embodiments, analyzing the image data (e.g., by the methods, steps, and modules described herein) can include analyzing pixels, voxels, point clouds, distance data, etc. included in the image data.

[0086] Convolution can include convolutions of any dimension. A one-dimensional convolution is a function that transforms an original sequence of numbers into a transformed sequence of numbers. A one-dimensional convolution can be defined by a series of scalars. Each particular value within the transformed sequence can be determined by calculating a linear combination of the values within a subsequence of the original sequence that corresponds to the particular value. The resulting value of the calculated convolution can include any value within the transformed sequence. Similarly, an n-dimensional convolution is a function that transforms an original n-dimensional array into a transformed array. An n-dimensional convolution can be defined by an n-dimensional array of scalars (known as the kernel of the n-dimensional convolution). Each particular value within the transformed array can be determined by calculating a linear combination of the values within an n-dimensional region of the original array that corresponds to the particular value. The resulting value of the calculated convolution can include any value within the transformed array. In some examples, an image can include one or more components (e.g., color components, depth components, etc.), and each component can include a two-dimensional array of pixel values. In one example, calculating the convolution of an image can include calculating a two-dimensional convolution on one or more components of the image. In another example, calculating the convolution of an image can include stacking arrays from different components to create a three-dimensional array and calculating a three-dimensional convolution on the resulting three-dimensional array. In some examples, a video can include one or more components (e.g., color components, depth components, etc.), and each component can include a three-dimensional array of pixel values (having two spatial axes and one temporal axis). In one example, calculating the convolution of a video can include calculating a three-dimensional convolution on one or more components of the video. In another example, calculating the convolution of a video can include stacking arrays from different components to create a four-dimensional array and calculating a four-dimensional convolution on the resulting four-dimensional array.

[0087] When using a wearable extended reality device, there may be a desire to change the perspective view, such as by zooming in on an object or by changing the virtual direction of the object. It may be desirable to enable this type of control via a touch controller that interacts with the wearable extended reality device such that a pinch or rotation of a finger on the touch sensor changes the perspective view of the scene. The following disclosure describes various systems, methods, and non-transitory computer-readable media for controlling a viewpoint in an extended reality environment using a physical touch controller.

[0088] Some of the disclosed embodiments can include a system, a method, and a non-transitory computer-readable medium for controlling a viewpoint in an extended reality environment using a physical touch controller. The term non-transitory is intended to describe computer-readable storage media excluding propagating electromagnetic signals, but is not intended to limit the types of physical computer-readable storage devices included in the phrase computer-readable media. For example, the term non-transitory computer-readable media is intended to include storage devices of a type that do not necessarily store information persistently, such as random access memory (RAM). Program instructions and data stored in a tangible computer-accessible memory medium in non-transitory form can be further transmitted by a transmission medium or signal, such as an electrical signal, an electromagnetic signal, or a digital signal, which can be transmitted via a communication medium such as a network and / or a wireless link. The non-transitory computer-readable media can include any other tangible medium that can store data in a format readable by a magnetic disk, a card, a tape, a drum, a punch card, a paper tape, an optical disk, a barcode, and magnetic ink characters, or any other machine device.

[0089] The control of the viewpoint as used in the present disclosure can include, for example, changing the position of the scene at a fixed distance, changing the distance of the scene, changing the angle of the scene, changing the size of the scene, causing rotation of the scene, or any other operation of the scene that can potentially cause a change in the way the observer visualizes the scene. In one example, controlling the viewpoint can include changing the spatial transformation of the scene. Some non-limiting examples of spatial transformation can include translational transformation, rotational transformation, reflection transformation, dilation transformation, affine transformation, projective transformation, etc. In one example, controlling the viewpoint can include changing the first spatial transformation of the first part of the scene and changing the second spatial transformation of the second part of the scene (the second spatial transformation may be different from the first spatial transformation, the change to the second spatial transformation may be different from the change to the first spatial transformation, and / or the second part of the scene may be different from the first part of the scene). The scene can include, for example, a place, location, position, point, spot, area, arena, stage, set, trajectory, section, segment, part, clip, sequence, living or inanimate object, person, or any other component of the visual field. Changing the position of the scene at a fixed distance can include, for example, translating the scene horizontally, vertically, diagonally, rotationally, or in any other direction within the three-dimensional field. Changing the distance of the scene can include, for example, increasing or decreasing the amount of space between the scene and any reference point horizontally, vertically, obliquely, rotationally, or in any other direction within the three-dimensional field.

[0090] Changing the angle of a scene can include, for example, increasing or decreasing the gradient, slope, inclination, or any other tilt between two reference points. Changing the size of a scene can include, for example, enlarging, increasing, expanding, extending, stretching, increasing, swelling, elongating, making larger, bulging, widening, amplifying, spreading, stretching, extending, stretching, shrinking, decreasing, making smaller, reducing, narrowing, lowering, shortening, compressing, or contracting one or more dimensions of any component, a plurality of components, or a combination of components within the scene. Other operations on the scene can include any visual change to the scene relative to the original scene.

[0091] A physical touch controller can include a device that enables manipulation of information on a display through detection of finger contact with a surface or movement of a finger on the surface. Any reference to a finger herein can equally apply to any other finger, pen, other object such as a stylus, etc. Any reference to finger contact and / or movement herein can equally apply to contact and / or movement of any other finger, pen, stylus, multiple such objects (e.g., in a multi-touch surface), etc. The physical touch controller can be implemented via software, hardware, or any combination thereof. In some embodiments, the physical touch controller can include, for example, one or more computer chips, one or more circuit boards, one or more electronic circuits, or any combination thereof. In addition to or instead of this, the physical touch controller can include, for example, a storage medium having executable instructions adapted to communicate with a touch sensor. In certain embodiments, the physical touch controller may be connected to other components such as sensors and processors using a wired or wireless system. For example, the touch controller can be implemented through pressure sensing of a finger (or substitute for a finger) on a touch-sensitive surface or through an optical sensor that detects touch or movement of a finger. In some examples, the physical touch controller can enable manipulation of information on a display via at least one of a touch sensor, a touch pad, a track pad, a touch screen, etc.

[0092] Some disclosed embodiments of a computer-readable medium can include instructions that cause one or more functions to be performed by at least one processor when executed by the at least one processor. The instructions can include any sequence, command, instruction, program, code, or condition that indicates how something should be done, operated, or processed. The processor that executes the instructions can be a general-purpose computer, but for example, can utilize any of a variety of other technologies including a dedicated computer, a microcomputer, a minicomputer, a mainframe computer, a programmed microprocessor, a microcontroller, a peripheral integrated circuit element, a CSIC (customer specific integrated circuit), an ASIC (application specific integrated circuit), a logic circuit, a digital signal processor, an FPGA (field programmable gate array), a PLD (programmable logic device), a PLA (programmable logic array), or other programmable logic devices such as RFID integrated circuits, smart chips, or any other device or arrangement of devices capable of implementing the steps or processes of the present disclosure. In some embodiments, the at least one processor can comprise a plurality of processors that need not be physically in the same location. Each processor can be in a geographically separate location and can be connected to communicate with each other in any suitable way. In some embodiments, each of the plurality of processors can be composed of different physical parts of the device.

[0093] Instructions embodied on a non-transitory computer-readable medium can include various instructions that cause a processor to perform one or more particular tasks. A set of such instructions for performing a particular task can be characterized, for example, as a program, software program, software, engine, module, component, mechanism, unit, or tool. Some embodiments can include a plurality of software processing modules stored in memory and executed on at least one processor. The program modules can be in the form of any suitable programming language that is converted to machine language or object code to enable at least one processor to read the instructions. The programming language used can include assembly language, Ada, APL, Basic, C, C++, COBOL, dBase, Forth, FORTRAN, Java, Modula-2, Pascal, Prolog, REXX, Visual Basic, JavaScript, or any other set of character strings that produces machine code output.

[0094] Consistent with some of the disclosed embodiments, a non-transitory computer-readable medium can include instructions that, when executed by at least one processor, cause the at least one processor to output a first display signal that reflects a first perspective of a scene for presentation via a wearable extended reality device. The output can include one or more of generation, distribution, or supply of data. The output may be presented, for example, in the form of a graphic display that includes one or more still or moving images, text, icons, video, or any combination thereof. The graphic display may be two-dimensional, three-dimensional, holographic, or may include various other types of visual characteristics. The at least one processor can generate one or more analog or digital signals and transmit them to a display device for presenting a graphic display for the user to view. In some embodiments, the display device can include a wearable extended reality device. As described elsewhere in this disclosure, a wearable extended reality device can implement, for example, augmented reality technology, virtual reality technology, mixed reality technology, any combination thereof, or any other technology that can combine a real environment and a virtual environment. A wearable extended reality device can take the form of, for example, a hat, visor, helmet, goggles, glasses, or any other object that can be worn. The perspective of the scene used in this disclosure can include, for example, the orientation, angle, size, direction, position, aspect, spatial transformation, or any other virtual characteristic of the scene or a portion of the scene. The at least one processor can be configured to cause the presentation of a graphic display that includes the first perspective of the scene. By way of example, as shown in FIG. 7, the processor can be configured to cause the presentation of the first perspective of scene 710.

[0095] Some of the disclosed embodiments can include receiving, via a touch sensor, a first input signal caused by a first multi-finger interaction with the touch sensor. In addition to or instead of the foregoing, the touch sensor can include, for example, a piezoresistive sensor, a piezoelectric sensor, an optical sensor, a capacitance sensor, a strain gauge sensor, or any other type of sensor that determines information based on physical interaction. In some embodiments, the touch sensor may be incorporated into a keyboard, a mouse, an operating rod, an optical or trackball mouse, or any other type of tactile input device. In some embodiments, the touch sensor can include a touch screen used alone or in combination with other input devices. The touch screen can be used, for example, to detect touch operations that can be performed by a user by using a finger, a palm, a wrist, a knuckle, or any other body part of the user. In addition to or instead of this, the touch operation can be performed by a user, for example, by touching the touch screen with a touch pen, a stylus, or any other external input device, or by moving the device near the touch screen. Regardless of the type of touch sensor used, the touch sensors disclosed herein can be configured to convert physical touch information into an input signal.

[0096] The first input signal can include a signal that provides information regarding the user's touch based on the user's interaction with the touch sensor. The first input signal can include information regarding one or more parameters that can be determined based on the touch interaction with the touch sensor. For example, the first input signal can include information regarding pressure, force, strain, position, movement, velocity, acceleration, temperature, occupancy, or any other physical or mechanical property. The first multi-finger interaction can include any interaction with the touch sensor that includes the user contacting the touch sensor using two or more fingers (or, using any two or more objects of any type, such as a finger, a pen, a bar, etc., to contact the touch sensor). In some embodiments, the first multi-finger interaction can include, for example, the user pinching two or more fingers together, separating two or more fingers, or performing any other action with two or more fingers on the touch sensor. FIG. 6 shows some examples of various multi-finger interactions that can be used in connection with some embodiments of the present disclosure, such as pinch-in, pinch-out, rotation, and slide. The first multi-finger interaction 610 represents a pinch-in action where the user brings the thumb and index finger closer together. The second multi-finger interaction 612 represents a pinch-out action where the user moves the thumb and index finger further apart from each other. The third multi-finger interaction 614 represents a rotation action where the user moves the thumb and index finger together clockwise or counterclockwise. The fourth multi-finger interaction 616 represents a slide action where the user moves the thumb and index finger in opposite directions. In another example of a multi-finger interaction, the user can move two or more fingers in the same direction. Other multi-finger interactions can also be used.

[0097] Some of the disclosed embodiments can include outputting a second display signal configured to cause a second perspective of a scene to be presented via a wearable extended reality device by changing a first perspective of the scene in response to a first input signal for presentation via the wearable extended reality device. The second display signal can have characteristics similar to the first display signal as described above and can be configured to change the first perspective of the scene. Changing the perspective of the scene can include changing one or more of, for example, the orientation, angle, size, direction, position, aspect, spatial transformation parameters, spatial transformation, or any other visual characteristic of the scene or a portion of the scene. Accordingly, the second perspective of the presented scene can include, for example, a second orientation, a second angle, a second size, a second direction, a second position, a second aspect, a second spatial transformation, or one or more of any other changes in the characteristics of the presentation of the scene (or a portion of the scene) in an extended reality environment different from the first perspective of the scene.

[0098] In some examples, it is possible to receive an indication that a physical object is located at a particular position within the environment of the wearable extended reality device. In some examples, it is possible to receive image data captured using an image sensor included in a first wearable extended reality device. For example, the image data can be received from the image sensor, from the first wearable extended reality device, from an intermediate device external to the first wearable extended reality device, from a memory unit, etc. The image data can be analyzed to detect a physical object at a particular position within the environment. In another example, a radar, lidar, or sonar sensor can be used to detect the presence of a physical object at a particular position within the environment. In some examples, a second perspective of the scene can be selected based on a physical object located at a particular position. In one example, the second perspective of the scene can be selected such that a virtual object within the scene does not appear to collide with the physical object. In another example, the second perspective of the scene can be selected such that a particular virtual object is not (fully or partially) hidden from the user of the wearable extended reality device by the physical object (e.g., based on the position of the wearable extended reality device). In one example, a raycasting algorithm can be used to determine that a particular virtual object is hidden from the user of the wearable extended reality device by the physical object. Further, the second display signal can be based on the selection of the second perspective of the scene. In one example, if the physical object is not located at a particular position, in response to the first input signal, a first version of the second display signal can be output for presentation via the wearable extended reality device, the first version of the second display signal being configured to change the first perspective of the scene and thereby cause the second perspective of the scene to be presented via the wearable extended reality device, the second perspective being able to include the presentation of a virtual object at a particular position.Furthermore, when a physical object is placed at a specific location, in response to a first input signal, a second version of the second display signal can be output for presentation via a wearable extended reality device, and the second version of the second display signal can be configured to change a first perspective of the scene, thereby presenting an alternative perspective of the scene via the wearable extended reality device, where the alternative perspective may not include the presentation of virtual objects at the specific location and / or may include the presentation of virtual objects at an alternative location different from the specific location.

[0099] Some of the disclosed embodiments can further include receiving, via a touch sensor, a second input signal caused by a second multi-finger interaction with the touch sensor. The second input signal may be similar to the first input signal described above and may include a signal providing information regarding one or more parameters that can be determined based on a touch interaction with the touch sensor. The second multi-finger interaction can include any interaction with the touch sensor that includes the user contacting the touch sensor using two or more fingers of a hand. Thus, for example, the second multi-finger interaction can include a movement of two or more fingers of the hand similar to the movement described above with respect to the first multi-finger interaction.

[0100] In accordance with some of the disclosed embodiments, the first multi-finger interaction and the second multi-finger interaction can be of a common multi-finger interaction type. Common multi-finger interaction types can include interactions that share commonality with fingers of the same hand, fingers of the same type (e.g., index finger, thumb, middle finger, etc.), the same fingers of the opposite hand, the same type of movement, any combination thereof, or any other type of interaction between two fingers that share commonality with another type of interaction between two fingers. By way of example, both the first multi-finger interaction and the second multi-finger interaction can include a finger pinch.

[0101] Consistent with some of the disclosed embodiments, the first multi-finger interaction can be of a different type than the second multi-finger interaction type. For example, the first multi-finger interaction of a different type than the second multi-finger interaction type can include interactions involving different types of movements, different fingers or different hands, any combination thereof, or any other type of interaction between two fingers that is different from another type of interaction between two fingers. As an example, the first multi-finger interaction can be a user who increases the distance between the thumb and index finger on one hand, while the second multi-finger interaction can be a user who decreases the distance between the thumb and index finger on the other hand. In another example, the first multi-finger interaction can be a user who rotates the thumb and index finger clockwise with one hand, while the second multi-finger interaction can be a user who rotates the thumb and index finger counterclockwise with the other hand. In other examples, the first multi-finger interaction can be a user who increases the distance between the thumb and index finger clockwise on one hand, while the second multi-finger interaction can be a user who rotates the thumb and index finger clockwise on the other hand.

[0102] Some of the disclosed embodiments may further include outputting a third display signal configured to present a third perspective of a scene via a wearable extended reality device by changing a second perspective of the scene to be presented via the wearable extended reality device in response to a second input signal. The third display signal can have similar characteristics as the first display signal as described above and can be configured to change the second perspective of the scene. Changing the second perspective of the scene can include changes similar to those described above with respect to changing the first perspective of the scene. For example, changing the second perspective of the scene can also include changing one or more of the orientation, angle, size, direction, position, aspect, spatial transformation parameters, spatial transformation, or any other visual characteristic of the scene or a portion of the scene. Accordingly, the third perspective of the presented scene can include one or more of a third orientation, a third angle, a third size, a third direction, a third position, a third aspect, a third spatial transformation, or a change in any other characteristic of the presentation of the scene (or a portion of the scene) in an extended reality environment different from the second perspective of the scene.

[0103] As an example, FIG. 7 shows a set of exemplary changes in the perspective of a scene that is consistent with some embodiments of the present disclosure. For example, the processor may be configured to first output a first display signal that reflects a first perspective of scene 710 for presentation via a wearable extended reality device. The processor can receive a first input signal 712 caused by a first multi-finger interaction 620 (e.g., pinch) with the touch sensor via the touch sensor. In response to the first input signal 712, the processor may output a second display signal configured to present a second perspective of scene 714 via the wearable extended reality device by changing the first perspective of the scene for presentation via the wearable extended reality device. As shown in FIG. 7, the first perspective of scene 710 has been changed to the second perspective of scene 714. The processor may receive a second input signal 716 caused by a second multi-finger interaction 622 (e.g., slide) with the touch sensor via the touch sensor, for example, after the first perspective of scene 710 has been changed to the second perspective of scene 714. As shown in FIG. 7, the second multi-finger interaction may be an upward swipe or drag motion of the user's thumb and index finger. In response to the second input signal 716, the processor may output a third display signal configured to present a third perspective of scene 718 via the wearable extended reality device by changing the second perspective of scene 714 for presentation via the wearable extended reality device. As shown in FIG. 7, the second perspective of scene 714 has been changed to the third perspective of scene 718.

[0104] In accordance with some disclosed embodiments, a scene can include a plurality of content windows. A content window can be an area of an extended reality environment configured to display text, a still image or a moving image, or other visual content. For example, a content window can include a web browser window, a word processing window, an operating system window, or any other area on a screen that displays information for a particular program. In another example, a content window can include a virtual display (such as a virtual object that mimics a computer screen described herein). In some embodiments, the content windows can be configured such that, at a first viewpoint of the scene, all of the plurality of content windows can be displayed at the same virtual distance from the wearable extended reality device, and at a second viewpoint of the scene, at least one of the plurality of windows can be displayed at a virtual distance different from the distance of the other windows of the plurality of windows, and at a third viewpoint of the scene, at least one window can be displayed in an orientation different from the orientation of the other windows of the plurality of windows. In one example, a content window displayed at a distance different from the distance of the other content windows can include a content window that is displayed closer to the wearable extended reality device than the other content windows. As another example, a content window displayed at a distance different from the distance of the other content windows can include a content window that is displayed farther from the wearable extended reality device than the other content windows. Alternatively, a content window displayed at a distance different from the distance of the other content windows can include a content window that is displayed with a zoomed-in object, and the other content windows are displayed with a zoomed-out object.A further example of a content window that is displayed at a distance different from the distance of other content windows can include a content window that is displayed with a zoomed-out object, while other content windows are displayed with zoomed-in objects. A content window that is displayed in an orientation different from the orientation of other content windows can include any display that changes any aspect of the angle, tilt, inclination, direction, or relative physical position of the content window as compared to other content windows. For example, a content window that is displayed in an orientation different from the orientation of other content windows can include a content window that is rotated 90 degrees while other content windows remain at 0 degrees. As another example, a content window that is displayed in an orientation different from the orientation of other content windows can include a content window that faces the left side of a wearable extended reality device, while other content windows face the right side of the wearable extended reality device.

[0105] Further examples of the viewpoints of the scene are shown in FIG. 8. For example, as shown at 810, one viewpoint can include two content windows presented side by side. Another viewpoint can include two content windows that are tilted in the same direction, as shown at 812. In another viewpoint, as shown at 814, the two content windows may face inward toward each other. Alternatively, the example of 816 shows two content windows that face outward from each other and are tilted upward. Another example of 818 shows that the two content windows are curved toward each other. In yet another example, as shown at 820, the two content windows may be curved in the same direction, but one content window may be presented over the content of another window. These are just a few examples of the viewpoints of the scene, but any variation in the visual presentation of the scene is contemplated for use with the embodiments disclosed herein.

[0106] Consistent with some of the disclosed embodiments, a scene can include a plurality of virtual objects. A virtual object can include a display that does not exist in the real world, such as an icon, an image of something that may or may not exist in the real world, a two-dimensional virtual object, a three-dimensional virtual object, a virtual object of a living being, a virtual object of an inanimate object, or any other display that simulates a non-physical display, a physical thing in either reality or imagination. In some embodiments, one or more virtual objects can be configured such that, in a first perspective of the scene, one or more of the virtual objects can be displayed at the same distance or a variable virtual distance from a wearable extended reality device, and in a second perspective of the scene, at least one of the virtual objects is displayed in a size different from other virtual objects among the plurality of virtual objects, and in a third perspective of the scene, the size of at least one virtual object returns to the size of at least one virtual object before being presented in the second perspective of the scene. When one or more virtual objects are displayed at the same distance from the wearable extended reality device, a user of the wearable extended reality device can see one or more virtual objects with the same amount of space between the user and each of the one or more virtual objects. When one or more virtual objects are displayed at a variable virtual distance from the wearable extended reality device, a user of the wearable extended reality device can see each or a part of one or more virtual objects in different amounts of space away from the user. When at least one of the virtual objects is displayed in a size different from other ones among the plurality of virtual objects, a user of the wearable extended reality device can see at least one of the virtual objects as larger or smaller than other ones among the plurality of virtual objects. When the size of at least one virtual object returns to the size of at least one virtual object before being presented in the second perspective of the scene, a user of the wearable extended reality device can see at least one virtual object in a size smaller or larger than the size in the second perspective of the scene.

[0107] Consistent with some of the disclosed embodiments, the input from the touch sensor enables selective position control of a plurality of virtual objects and allows the plurality of virtual objects to be directed at variable virtual distances from the wearable extended reality device. The variable virtual distance can refer to the perceived distance of the virtual object from the appliance wearer. Just as the virtual object is not physical, the distance is not real either. Rather, the distance of the virtual object from the wearer is as virtual as the virtual object itself. In one example, the distance can appear as the size of the rendering of the virtual object, as the position of the virtual object when rendered from different viewpoints (e.g., from the two viewpoints of the observer's two eyes, from the moving viewpoint of the observer, etc.), as hiding or being hidden by other objects (physical or virtual) that overlap the virtual object within the viewer's field of view of the virtual object. The input from the touch sensor can change such a virtual distance (i.e., change the presentation to make the change in virtual distance perceptible to the wearer). Such user input can be configured to occur prior to the first input signal such that the presentation via the wearable extended reality device enables the user to view the touch sensor. In order for the user to be able to interact with the touch sensor, it may be desirable to enable the user to view the touch sensor when a plurality of virtual objects obscure the user's field of view so that the user cannot view the touch sensor. Selective position control can include the ability to change one or more sizes, orientations, distances, or any other spatial attributes of the virtual object compared to another virtual object. Thus, in some embodiments, the user can selectively change the distance of one or more virtual objects such that the virtual objects are disposed at variable distances from the wearable extended reality device.

[0108] Consistent with some of the disclosed embodiments, selective position control can include docking a plurality of virtual objects in physical space and enabling a wearer of a wearable extended reality device to walk within the physical space and separately adjust the position of each virtual object. Docking a plurality of virtual objects in physical space can include maintaining the size, orientation, distance, or any other spatial attribute of the plurality of virtual objects with reference to the physical space. This type of selective position control can enable a user to move within the physical space while maintaining the size, orientation, distance, or any other spatial attribute of the plurality of virtual objects that is maintained as the user moves. For example, a user can move toward and select a virtual object and adjust any one of the size, orientation, distance, or any other spatial attribute of that virtual object, while the size, orientation, distance, or any other spatial attribute of the other virtual objects can remain unchanged. The user can move to another virtual object within the physical space and adjust any one of the size, orientation, distance, or any other spatial attribute of that virtual object, while the size, orientation, distance, or any other spatial attribute of the other virtual objects can remain unchanged. The user can continue to do the same for some or all of the virtual objects as the user moves around the physical space. The user can adjust any one of the size, orientation, distance, or any other spatial attribute of each virtual object, while the other unadjusted characteristics or other spatial attributes of the other virtual objects can remain unchanged.

[0109] Consistent with some of the disclosed embodiments, a second perspective of the scene can be selected based on a first input signal and whether a housing including a touch sensor is placed on a physical surface during a first multi-finger interaction. For example, the position of the input device can help define the perspective of the presentation of the scene in an extended reality environment. Display characteristics can be different, for example, when the associated touch sensor is on a table rather than being held by the user. Alternatively, for example, the relative position (left or right) of the user with respect to the touch-sensitive display can affect the preferred perspective of the presentation of the scene in an extended reality environment. In some examples, at least one parameter of the second perspective may be determined based on the physical surface. For example, a virtual surface can be determined based on the physical surface (e.g., parallel to the physical surface, overlapping the physical surface, perpendicular to the physical surface, an orientation selected with respect to a virtual surface, etc.), and virtual objects can be moved on the virtual surface. In one example, the direction and / or distance of movement of a virtual object on the virtual surface can be selected based on the first input signal. For example, in response to the first input signal, the virtual object may be moved on the virtual surface in a first direction and a first distance, and in response to the second input signal, the virtual object may be moved on the virtual surface in a second direction and a second distance. In another example, the new position of the virtual object on the virtual surface may be selected based on the first input signal.

[0110] The viewpoint selection, which is partially based on whether the housing is disposed on the physical surface at that point in time, may also be desirable to enable a particular type of display depending on whether the user of the wearable extended reality device is mobile or stationary at a given point in time. For example, if the user is mobile, a particular type of scene viewpoint may not be desirable because it may interfere with the user's safe movement, for example, by blocking the user's field of view. When the user is moving, the user is likely to be holding the housing or otherwise attached to the housing rather than placing the housing on the physical surface during the first multi-finger interaction. When the user is stationary, the user can place the housing on the physical surface, thereby enabling the user to select other viewpoints that do not require the user to look at portions of the user's forward field of view. When the user is stationary, the user is likely to place the housing on the physical surface during the first multi-finger interaction rather than holding the housing or otherwise attached to the housing. In one example, the selection of the second viewpoint of the scene can be further based on the type of surface, for example, whether the surface is the top surface of a table. In another example, the third viewpoint of the scene may be selected based on the second input signal and whether the housing is disposed on the physical surface during the second multi-finger interaction.

[0111] In certain embodiments, the first input signal from the touch sensor can reflect two-dimensional coordinate input, and the second display signal is configured to change a first viewpoint of the scene to introduce a three-dimensional change to the scene. Two-dimensional coordinate input can include input that affects a geometric setting that lacks depth and has only two measurements such as length and width. The three-dimensional change to the scene can include a change to the scene having three measurements such as length, width, and depth. For example, the first viewpoint of the scene can include a cube. In this example, the first input signal from the touch sensor can reflect an input that includes the user moving the thumb and index finger in opposite directions in a sliding motion. The second display signal in this example may be configured to change the cube by increasing or decreasing the depth of the cube.

[0112] In some embodiments, the first input signal from the touch sensor can reflect Cartesian coordinate input, and the second display signal configured to change the viewpoint of the scene is in a spherical coordinate system. Cartesian coordinate input can include input that affects a geometric setting defined by three types of points, each type of point being on a separate orthogonal axis. The spherical coordinate system can include a three-dimensional space defined by three types of points, the types of points including two angles on the orthogonal axes and the distance from the origin of the coordinate system to the type of point created by the two angles on the orthogonal axes. For example, the first viewpoint of the scene can include a cube. In this example, the first input signal from the touch sensor can reflect an input that includes the user moving the thumb and index finger in opposite directions in a sliding motion. The second display signal in this example may be configured to change the cube by increasing or decreasing the depth of the cube. Cartesian input can be used to change the viewpoint of the scene in spherical coordinates by using any mathematical technique for conversion between coordinate systems to convert the Cartesian coordinates of the input to spherical coordinates for the change.

[0113] In some embodiments, each type of multi-finger interaction can control one of a plurality of actions for changing the perspective of a scene presented via a wearable extended reality device. For example, one or more actions can include zooming in, zooming out, rotating, and sliding, and a multi-finger interaction representing the user bringing the thumb and index finger closer together in a pinch-in motion can control the zoom-in action. As another example, a multi-finger interaction representing the user moving the thumb and index finger further apart from each other in a pinch-out motion can control the zoom-out action. As another example, in a rotation action, a multi-finger interaction representing the user moving their thumb and index finger together clockwise or counterclockwise can control the rotation action. As another example, a multi-finger interaction representing the user moving the thumb and index finger in opposite directions in a slide action can control the slide action.

[0114] In some embodiments, the plurality of actions for changing the perspective of a scene can include at least two of changing the perspective position with respect to the scene, changing the distance from the scene, changing the angle of the scene, and changing the size of the scene. However, these changes are not limited to the entire scene, and the changes in perspective position, distance, angle, and size with respect to the scene may be applied to one or more of the selected objects within the scene.

[0115] In some embodiments, the first input signal and the second input signal may be associated with a first operating mode of a touch sensor for controlling a viewpoint. The operating mode of the touch sensor can include, for example, the means, method, approach, mode of operation, procedure, process, system, technique, or any other functional state of the touch sensor. The first operating mode may be different from the second operating mode with respect to the means, method, approach, mode of operation, procedure, process, system, technique, or any other functional state of the touch sensor. For example, the first operating mode can include a touch sensor that generates an input signal representing a user who increases or decreases the distance between the thumb and the index finger, and the second operating mode can include a touch sensor that generates an input signal representing a user who moves the thumb and the index finger together in a clockwise or counterclockwise direction. As another example, the first operating mode can include a touch sensor that generates an input signal representing a user who performs an interaction of multiple fingers within a defined area of the touch sensor, and the second operating mode can include a touch sensor that generates an input signal representing a user who performs an interaction of multiple fingers at any location on the touch sensor.

[0116] In some embodiments, at least one processor may be configured to receive an initial input signal related to a second operating mode of a touch sensor for controlling a pointer presented via a wearable extended reality device via the touch sensor before receiving a first input signal. The initial input signal used in the present disclosure may provide information regarding a user's touch, for example, based on an interaction between the user and the touch sensor, and may include a signal for controlling the position of the pointer in an extended reality environment presented via the wearable extended reality device. The initial input signal may include information regarding pressure, force, strain, position, movement, acceleration, occupancy, and any other parameter that can be determined based on a touch interaction with the touch sensor. The pointer may include, for example, an arrow, a dial, a director, a guide, an indicator, a mark, a signal, or any other visual representation of a position in the content presented via the wearable extended reality device or the extended reality environment.

[0117] Some embodiments can include selecting a virtual object within a scene in response to an initial input signal. Selecting a virtual object within a scene can include selecting a virtual object within the scene for further manipulation. Selection of a virtual object within the scene can be achieved by using the initial input signal to control a pointer to perform a click, selection, highlighting, pinning, drawing around a box, or any other type of marking of the virtual object. For example, the virtual object can be a tree within a forest scene. The user can select a tree in the forest by guiding the position of the pointer to the tree by dragging a finger on a touch sensor and marking the selected tree by clicking on it. In some embodiments, at least one processor can be configured to cause a change in the display of the selected virtual object to be presented via a wearable extended reality device in response to a first input signal and a second input signal. The change in the display of the selected virtual object can include a change in the viewpoint position relative to the object, a change in the distance from the object, a change in the angle of the object, and a change in the size of the object. For example, the virtual object can be a tree within a forest scene. The user can increase the size of the tree by increasing the distance between the thumb and index finger and inputting a first input signal via the touch sensor. The user can then rotate the tree by moving the thumb and index finger together in a clockwise or counterclockwise direction and inputting a second input signal via the touch sensor.

[0118] In some embodiments, the instructions are configured such that a first display signal, a second display signal, and a third display signal are transmitted to the wearable extended reality device via a first communication channel, and a first input signal and a second input signal are received from a computing device other than the wearable extended reality device via a second communication channel. As used in this disclosure, the communication channel can include either a wired communication channel or a wireless communication channel. The wired communication channel can utilize a network cable, a coaxial cable, an Ethernet cable, or any other channel that transmits information via a wired connection. The wireless communication channel can utilize WiFi, Bluetooth™, NFC, or any other channel that transmits information without a wired connection. The computing device can be any device other than the wearable extended reality device, including a general computer such as a personal computer (PC), a UNIX workstation, a server, a mainframe computer, a personal digital assistant (PDA), a cellular phone, a smartphone, a tablet, any combination of these, or any other device for automatically performing calculations.

[0119] In some embodiments, the computing device can be a smartphone connectable to the wearable extended reality device, and the touch sensor can be the touch screen of the smartphone. As another example, the computing device can be a smartphone or a smartwatch connectable to the wearable extended reality device, and the touch sensor can be the touch screen of the smartwatch. In such embodiments, the first and second multi-finger interactions can occur on the touch screen of the smartphone or smartwatch.

[0120] In some embodiments, the computing device can include a keyboard connectable to a wearable extended reality device, and the touch sensor is associated with the computing device. The keyboard can be a standard typewriter-style keyboard (e.g., a QWERTY-style keyboard), or another suitable keyboard layout such as a Dvorak layout or a coded layout. In one embodiment, the keyboard can include at least 30 keys. The keyboard can include alphanumeric keys, function keys, modifier keys, cursor keys, system keys, multimedia control keys, or other physical keys that perform computer-specific functions when pressed. The keyboard can also include virtual keys or programmable keys, and the function of the keys can vary depending on the functions or applications executed on the integrated computing interface device. In some embodiments, the keyboard can include dedicated input keys for performing operations associated with the wearable extended reality device. In some embodiments, the keyboard can include dedicated input keys for changing the illumination of content projected via the wearable extended reality device (e.g., an extended reality environment, or virtual content, etc.). The touch sensor may be incorporated into the keyboard via a touchpad that is part of the keyboard. The touch sensor may be incorporated into the keyboard as one or more touch sensors disposed on one or more keys of the keyboard. In such a configuration, the touch sensor is not wearable and is not connected to the wearable extended reality device or configured to move with the wearable extended reality device.

[0121] Some embodiments can include obtaining default settings for displaying content using a wearable extended reality device. The default settings may be stored in memory and may be pre-set by a manufacturer, distributor, or user. The settings can include color, size, resolution, orientation, or any other condition related to the display. The default settings can include any pre-selected option for such conditions that is employed by a computer program or other mechanism when no alternative means are specified by the user or programmer during use. The default settings may be received from the user and may be determined based on the type, time, location, or any other type of heuristic data of the wearable extended reality device being used. Further embodiments of the non-transitory computer-readable medium include selecting a first perspective of a scene based on the default settings.

[0122] In some embodiments, the default settings are determined based on a previous perspective of the scene. The previous perspective of the scene can include any perspective prior to a given point in time for presenting the scene, such as a perspective from a previous use of the wearable extended reality device. For example, if the perspective of the scene was zoomed in by a certain amount in a previous use, the default settings can be set such that the default display of the scene is that perspective zoomed in by that certain amount. In other embodiments, the instructions can direct the processor to select a first perspective of the scene based on user preferences. The user preferences can include settings customized by the user of the wearable extended reality device, such as color, size, resolution, orientation, or any other condition related to the display that can be changed. For example, a user of a wearable extended reality device can set a user preference that directs that the first perspective be displayed in black and white only. In this example, the instructions can direct the processor to select a first perspective of the scene to be displayed in black and white only based on these user preferences.

[0123] In some embodiments, the default settings are associated with a keyboard paired with the wearable extended reality device. The paired keyboard may be the user's business keyboard, home keyboard, office keyboard, field keyboard, or any other type of peripheral device that enables the user to input text into any electronic device. The pairing can include either wired pairing or wireless pairing. Wired pairing can utilize a coaxial cable, Ethernet, or any other channel for transmitting information via a wired connection. Wireless pairing can utilize WiFi, Bluetooth™, or any other channel for transmitting information without a wired connection. For example, a user of a wearable extended reality device can input default settings into the wearable extended reality device by inputting information regarding the default settings via a keyboard wirelessly paired with the wearable extended reality device. In one example, the pairing can enable changes to virtual objects using the keyboard. For example, the virtual object can include the presentation of text input using the keyboard.

[0124] Other disclosed embodiments can include a method for controlling a viewpoint in an extended reality environment using a physical touch controller. By way of example, FIG. 9 shows a flowchart depicting an exemplary process 900 for changing the viewpoint of a scene, consistent with some embodiments of the present disclosure. The method includes, as shown in block 910, outputting a first display signal that reflects a first viewpoint of the scene for presentation via a wearable extended reality device. The method also includes, as shown in block 912, receiving, via a touch sensor, a first input signal caused by a first multi-finger interaction with the touch sensor. The method further includes outputting a second display signal configured to cause a second viewpoint of the scene to be presented via the wearable extended reality device by changing the first viewpoint of the scene in response to the first input signal for presentation via the wearable extended reality device. This step is shown in block 914. As shown in block 916, the method also includes receiving, via the touch sensor, a second input signal caused by a second multi-finger interaction with the touch sensor. Block 918 shows outputting a third display signal configured to cause a third viewpoint of the scene to be presented via the wearable extended reality device by changing the second viewpoint of the scene in response to the second input signal for presentation via the wearable extended reality device.

[0125] Some embodiments of the present disclosure may relate to enabling gesture interaction with invisible virtual objects. Virtual objects displayed by a wearable extended reality device can be docked to a physical object (e.g., such that the virtual object can have a fixed position relative to the physical object). If the docked position of the virtual object is outside the field of view portion associated with the display system of the wearable extended reality device such that the virtual object is not displayed by the display system, the user can interact with the virtual object by gesture interaction with the docking position. In another example, where the virtual object is docked in the field of view portion associated with the display system of the wearable extended reality device but is hidden by another object (physical or virtual) such that the virtual object is not displayed by the display system, the user can interact with the virtual object by gesture interaction with the docking position. Gesture interaction with the docking position can be captured (e.g., by an image sensor) and used to determine the user's interaction with the virtual object. Further details regarding enabling gesture interaction with invisible virtual objects are described below.

[0126] Some of the disclosed embodiments may relate to enabling gesture interaction with invisible virtual objects, and the embodiments are presented as a method, a system, an apparatus, and a non-transitory computer-readable medium. It should be understood that any of the presentation formats described below is equally applicable to the others. For example, one or more processes embodied in a non-transitory computer-readable medium may be executed as a method, a system, or an apparatus. Some aspects of such processes may be performed electronically via a network, which may be wired, wireless, or both. Other aspects of such processes may be performed using non-electronic means. In the broadest sense, a process is not limited to a particular physical or electronic means, but rather can be achieved using many different means. For example, some of the disclosed embodiments can include a method, a system, or an apparatus for enabling gesture interaction with invisible virtual objects. The method can include various processes described herein. The system can include at least one processor configured to execute various processes described herein. The apparatus can include at least one processor configured to execute various processes as described herein.

[0127] A non-transitory computer-readable medium can include instructions that, when executed by at least one processor, cause the at least one processor to perform various processes as described herein. The non-transitory computer-readable medium can include any type of physical memory where information or data readable by at least one processor can be stored. The non-transitory computer-readable medium can include, for example, random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, compact disc read-only memory (CD-ROM), digital versatile discs (DVD), flash drives, disks, any optical data storage medium, any physical medium having a pattern of holes, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), FLASH-EPROM or any other flash memory, non-volatile random access memory (NVRAM), caches, registers, any other memory chip or cartridge, or network versions thereof. The non-transitory computer-readable medium can refer to multiple structures such as multiple non-transitory computer-readable media located at a local location or a remote location. Further, one or more non-transitory computer-readable media can be utilized when implementing a computer-implemented method. Thus, the non-transitory computer-readable medium can include tangible items and can exclude carrier waves or transient signals.

[0128] The instructions included in the non-transitory computer-readable medium can include, for example, software instructions, computer programs, computer code, executable instructions, source code, machine instructions, machine language programs, or any other type of instruction for a computing device. The instructions included in the non-transitory computer-readable medium can be based on one or more of various types of desired programming languages and can include various processes for enabling gesture interaction with the invisible virtual objects described herein (e.g., instantiating).

[0129] At least one processor can execute instructions included in a non-transitory computer-readable medium to perform various operations for enabling gesture interaction with the invisible virtual objects described herein. The processor can include, for example, an integrated circuit, a microchip, a microcontroller, a microprocessor, a central processing device (CPU), a graphics processing device (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or other unit suitable for executing instructions or performing logical operations. The processor can include a single-core or multi-core processor. In some examples, the processor may be a single-core processor configured with virtualization technology. The processor can implement virtualization technology or other functions, for example, to provide the ability to execute, control, execute, operate, or store multiple software processes, applications, or programs. In another example, the processor can include a multi-core processor (e.g., dual-core, quad-core, or any desired number of cores) configured to provide a parallel processing function to enable a device associated with the processor to execute multiple processes simultaneously. Other types of processor configurations may be implemented to provide the functions described herein.

[0130] Some of the disclosed embodiments can relate to enabling gesture interactions with invisible virtual objects. Gestures can include any finger or hand movement, such as dragging, pinching, spreading, swiping, tapping, pointing, scrolling, rotating, flicking, touching, zooming in, zooming out, thumbing up, thumbing down, touch and hold, or any other movement of the hand. In some examples, gestures can include movements of the eyes, mouth, face, or other parts of a person's body. A virtual object can refer to a visual display rendered by a computing device and configured to represent an object. Virtual objects can include, for example, inanimate virtual objects, animate virtual objects, virtual furniture, virtual decorative objects, virtual widgets, virtual screens, or any other type of virtual display. A user can interact with a virtual object via a gesture. In some examples, as described in more detail herein, it is not necessary for the virtual object to be displayed to the user for the user to interact with the virtual object. For example, when a user docks a virtual object for volume control to a corner of a desk, each time the user touches the corner of the desk, a sensor (e.g., a camera) can detect the touch gesture and enable volume adjustment even if the virtual object for volume control is not being displayed at the time the interaction is initiated. Docking a virtual object to a physical object can, for example, fix the virtual object in a fixed position relative to the physical object. Further details regarding various processes for enabling gesture interactions with invisible virtual objects (including docking) are described below.

[0131] Some of the disclosed embodiments can include receiving image data captured by at least one image sensor of a wearable extended reality device (WER-device). The image data can include the display of a plurality of physical objects within the field of view associated with at least one image sensor of the wearable extended reality device. For example, a user can use a wearable extended reality device. The wearable extended reality device can include at least one image sensor. The at least one image sensor can include image sensor 372 or image sensor 472, or can be implemented similarly thereto. In some examples, the at least one image sensor can be configured to continuously or periodically capture image data of the location where the wearable extended reality device is present (e.g., an image of the physical space setting that the at least one image sensor may face). Capturing the image data can include, for example, the image sensor receiving an optical signal emitted from the environment in which the image sensor is present, such as visible light, infrared light, electromagnetic waves, or any other desired radiation. Capturing the image data can include, for example, converting the received optical signal into data that can be stored or processed by a computing device. At least one processor can receive the image data captured by at least one image sensor of the wearable extended reality device. In some examples, the at least one processor can include processing device 360, processing device 460, processing device 560, or any other processor associated with the wearable extended reality device, or can be implemented similarly thereto.

[0132] At least one image sensor of the wearable extended reality device can have a field of view. The field of view can refer to the spatial range that can be observed or detected at any given instant. For example, the field of view of the image sensor can include the solid angle to which the image sensor can be sensitive to radiation (e.g., visible light, infrared light, or other optical signals). In some examples, the field of view can include the angular field of view. The field of view can be measured horizontally, vertically, obliquely, or in any other desired manner.

[0133] The image data captured by at least one image sensor of the wearable extended reality device can include the display of a plurality of physical objects within the field of view associated with the at least one image sensor of the wearable extended reality device. The plurality of physical objects can include desks, tables, keyboards, mice, touch pads, lamps, cups, telephones, mobile devices, printers, screens, shelves, machines, vehicles, doors, windows, chairs, buttons, surfaces, or any other type of physical item that can be observed by the image sensor. The plurality of physical objects may be within the field of view of at least one image sensor of the wearable extended reality device, and the image data including the display of the plurality of physical objects may be captured by the at least one image sensor.

[0134] FIG. 10 is a flowchart showing an exemplary process 1000 for enabling gesture interaction with an invisible virtual object consistent with some embodiments of the present disclosure. Referring to FIG. 10, at step 1010, instructions included in a non-transitory computer-readable medium when executed by a processor can cause the processor to receive image data captured by at least one image sensor of a wearable extended reality device, the image data including the display of a plurality of physical objects within the field of view associated with the at least one image sensor of the wearable extended reality device.

[0135] FIG. 11 is a schematic diagram showing a field of view that coincides with the horizontal direction in some embodiments of the present disclosure. Referring to FIG. 11, the wearable extended reality device 1110 can include at least one image sensor. The field of view of the at least one image sensor can have a horizontal range 1112. FIG. 12 is a schematic diagram showing a field of view that coincides vertically with some embodiments of the present disclosure. Referring to FIG. 12, the field of view of the at least one image sensor of the wearable extended reality device 1110 can have a vertical range 1210.

[0136] FIG. 13 is a schematic diagram showing a user using an exemplary wearable extended reality device that coincides with some embodiments of the present disclosure. One or more elements as shown in FIG. 13 may be the same as or similar to one or more elements described in relation to FIG. 1. For example, referring to FIG. 13, the user 100 may be sitting behind the table 102. The user 100 can use a wearable extended reality device (for example, the wearable extended reality device 110). The keyboard 104 and the mouse 106 may be arranged on the table 102. The wearable extended reality device (for example, the wearable extended reality device 110) can display virtual content such as a virtual screen 112 and virtual widgets 114C, 114D, 114E, etc. to the user 100. The wearable extended reality device (for example, the wearable extended reality device 110) can include at least one image sensor. The at least one image sensor can have a field of view 1310 that can be represented by a solid angle from the point of the wearable extended reality device. The at least one image sensor may be configured to capture image data including the display of a plurality of physical objects within the field of view 1310. The plurality of physical objects can include, for example, the table 102, the keyboard 104, and the mouse 106.

[0137] Figures 14, 15, 16, 17, 18, and 19 show various usage snapshots of an exemplary system for enabling gesture interaction with invisible virtual objects that are consistent with some embodiments of the present disclosure. Figures 14, 15, 16, 17, 18, and 19 can show one or more elements as described in connection with FIG. 13 from another perspective (e.g., the perspective of user 100, the perspective of user 100's wearable extended reality device, or the perspective of at least one image sensor of user 100's wearable extended reality device). Referring to FIG. 14, a keyboard 104 and a mouse 106 may be placed on a table 102. The wearable extended reality device can display virtual content such as a virtual screen 112 and virtual widgets 114C, 114D, 114E, etc. to the user. The wearable extended reality device can include at least one image sensor. The at least one image sensor can have a field of view 1310 and be configured to capture image data including the display of a plurality of physical objects within the field of view 1310. The plurality of physical objects can include, for example, the table 102, the keyboard 104, and the mouse 106.

[0138] Some of the disclosed embodiments can include displaying a plurality of virtual objects in a portion of the field of view, where the portion of the field of view is associated with a display system of a wearable extended reality device. As described above, the virtual objects can include, for example, inanimate virtual objects, virtual objects of living organisms, virtual furniture, virtual decorative objects, virtual widgets, virtual screens, or any other type of virtual display. The wearable extended reality device can display the plurality of virtual objects to the user via the display system. The display system of the wearable extended reality device can include, for example, an optical head-mounted display, a monocular head-mounted display, a binocular head-mounted display, a see-through head-mounted display, a helmet-mounted display, or any other type of device configured to display an image to the user. In some examples, the display system of the wearable extended reality device can include a display configured to reflect a projected image and enable the user to view through the display. The display can be based on waveguide technology, diffractive optics, holographic optics, polarization optics, reflective optics, or other types of technology for combining an optical signal emitted from a projected image by a computing device and an optical signal emitted from a physical object.

[0139] The display system of the wearable extended reality device may be configured to display virtual content (e.g., one or more virtual objects) to a user in a part of the field of view of at least one image sensor of the wearable extended reality device. The part of the field of view may be associated with the display system of the wearable extended reality device. For example, the virtual content output by the display system of the wearable extended reality device may occupy a solid angle that may be within the solid angle representing the field of view of at least one image sensor of the wearable extended reality device. In some examples, the display system of the wearable extended reality device may be able to have a field of view in which virtual content can be displayed. The field of view of the display system may correspond to a part of the field of view of at least one image sensor. In some examples, the solid angle representing the field of view of the display system may have a vertex that is the same as or approximate to the vertex (e.g., the point from which an object can be viewed) of the solid angle representing the field of view of at least one image sensor.

[0140] Referring to FIG. 10, at step 1012, when executed by a processor, the instructions included in the non-transitory computer-readable medium can cause the processor to display a plurality of virtual objects in a part of the field of view, and the part of the field of view is associated with the display system of the wearable extended reality device.

[0141] For example, referring to FIG. 11, the wearable extended reality device 1110 may include at least one image sensor. The field of view of the at least one image sensor may have a horizontal range 1112. A part of the field of view of the at least one image sensor may have a horizontal range 1114. The horizontal range 1114 may be within the horizontal range 1112. In some examples, the vertex of the horizontal range 1112 may be the same as the vertex of the horizontal range 1114, or may be approximate (e.g., less than 1 mm from each other, less than 1 cm from each other, less than 10 cm from each other, etc.).

[0142] As another example, referring to FIG. 12, the field of view of at least one image sensor of the wearable extended reality device 1110 can have a vertical range 1210. A part of the field of view of at least one image sensor can have a vertical range 1212. The vertical range 1212 may be within the vertical range 1210. In some examples, the apex of the vertical range 1210 may be the same as the apex of the vertical range 1212, or (e.g., less than 1 mm from each other, less than 1 cm from each other, less than 10 cm from each other, etc.) may be approximated.

[0143] Referring to FIG. 13, a wearable extended reality device (e.g., the wearable extended reality device 110) can include at least one image sensor. The at least one image sensor can have a field of view 1310 that can be represented by a solid angle from a point of the wearable extended reality device. A part of the field of view 1310 may be 1312. The field of view portion 1312 can be represented by a solid angle within the solid angle representing the field of view 1310. The field of view portion 1312 may be associated with a display system of the wearable extended reality device. The wearable extended reality device can display a plurality of virtual objects (e.g., the virtual screen 112 and the virtual widgets 114C, 114D, 114E) in the field of view portion 1312. Continuing with the example, referring to FIG. 14, the field of view portion 1312 may be within the field of view 1310, and the wearable extended reality device may display a plurality of virtual objects (e.g., the virtual screen 112 and the virtual widgets 114C, 114D, 114E) in the field of view portion 1312.

[0144] In some examples, the field of view associated with at least one image sensor of the wearable extended reality device may be wider than a part of the field of view associated with the display system of the wearable extended reality device. In some examples, the horizontal range and / or vertical range of a part of the field of view associated with the display system may be 45 to 90 degrees. In other embodiments, a part of the field of view associated with the display system can be increased to the human field of view or the field of view. For example, a part of the field of view associated with the display system can be increased to about 200 degrees for the horizontal range and about 150 degrees for the vertical range. As another example, the field of view associated with at least one image sensor can be increased accordingly to be wider than a part of the field of view associated with the display system.

[0145] In some examples, the horizontal range of the field of view associated with at least one image sensor of the wearable extended reality device may exceed 120 degrees, and the horizontal range of a part of the field of view associated with the display system of the wearable extended reality device may be less than 120 degrees. In some examples, the vertical range of the field of view associated with at least one image sensor of the wearable extended reality device may exceed 120 degrees, and the vertical range of a part of the field of view associated with the display system of the wearable extended reality device may be less than 120 degrees.

[0146] In some examples, the horizontal range of the field of view associated with at least one image sensor of the wearable extended reality device may exceed 90 degrees, and the horizontal range of a part of the field of view associated with the display system of the wearable extended reality device may be less than 90 degrees. In some examples, the vertical range of the field of view associated with at least one image sensor of the wearable extended reality device may exceed 90 degrees, and the vertical range of a part of the field of view associated with the display system of the wearable extended reality device may be less than 90 degrees.

[0147] In some examples, the horizontal range of the field of view associated with at least one image sensor of the wearable extended reality device may exceed 150 degrees, and the horizontal range of a portion of the field of view associated with the display system of the wearable extended reality device may be less than 150 degrees. In some examples, the vertical range of the field of view associated with at least one image sensor of the wearable extended reality device may exceed 130 degrees, and the vertical range of a portion of the field of view associated with the display system of the wearable extended reality device may be less than 130 degrees.

[0148] In addition to or instead of this, the horizontal range and / or vertical range of a portion of the field of view associated with the display system of the wearable extended reality device may be configured to be any desired degree for displaying virtual content to the user. For example, the field of view associated with at least one image sensor of the wearable extended reality device may be configured to be wider than a portion of the field of view associated with the display system of the wearable extended reality device (e.g., with respect to the horizontal range and / or vertical range).

[0149] In some examples, the plurality of virtual objects displayed in a portion of the field of view associated with the display system of the wearable extended reality device can include a first virtual object that can be launched on a first operating system and a second virtual object that can be launched on a second operating system. An operating system can include software that controls the operation of a computer and instructs the processing of programs. A plurality of operating systems can be used to control and / or present virtual objects. For example, the first operating system may be implemented in relation to a keyboard (e.g., keyboard 104), and the second operating system may be implemented on a wearable extended reality device (e.g., wearable extended reality device 110). In addition to or instead of this, an input unit (e.g., input unit 202), a wearable extended reality device (e.g., extended reality unit 204), or a remote processing unit (e.g., remote processing unit 208) may each have an operating system. In another example, the first operating system can be launched on a remote cloud, and the second operating system can be executed on a local device (e.g., a wearable extended reality device, a keyboard, a smartphone in the vicinity of the wearable extended reality device, etc.). The first virtual object can be launched on one of the operating systems. The second virtual object may be launched on another one of the operating systems.

[0150] Some of the disclosed embodiments can include receiving a selection of a particular physical object from a plurality of physical objects. For example, at least one image sensor of a wearable extended reality device can continuously or periodically monitor a scene within the field of view of the at least one image sensor, and at least one processor can detect a display of a selection of a particular physical object from the plurality of physical objects based on the monitored scene. The display of the selection of the particular physical object can include, for example, a gesture by a user indicating the selection of the particular physical object (e.g., pointing at the particular physical object, touch-and-hold on the particular physical object, or tapping on the particular physical object). In some examples, at least one processor can determine a particular point or region on or near the selected particular physical object based on the gesture by the user. The particular point or region can be used to receive a docking virtual object. In some examples, the selected particular physical object can be part of a physical object (e.g., a particular point or region on a table). In some other examples, the selection of the particular physical object from the plurality of physical objects can be received from an external device, from a data structure specifying the selection of the physical object, from a memory unit, from the user, from an automated process (e.g., from an object detection algorithm used to analyze image data to detect and thereby select a particular physical object), etc.

[0151] In some examples, the selection of a particular physical object may be from a list or menu that includes a plurality of physical objects. For example, based on receiving image data captured by at least one image sensor that includes a display of a plurality of physical objects, at least one processor can analyze the image data to assign labels to each of the plurality of physical objects, and can generate a list or menu that includes labels that identify each of the physical objects. A user of the wearable extended reality device can call up the list or menu and select a particular physical object from the list or menu. The user can further indicate a particular point or region on or near the particular physical object selected to receive the docking virtual object.

[0152] A particular physical object can include any physical object selected from a plurality of physical objects. In some examples, a particular physical object can include a portion of a touchpad. In some examples, a particular physical object can include a non-control surface. A non-control surface can include surfaces such as a surface of a table, a surface of clothing, etc., that are not configured to provide control signals to a computing device. In some examples, a particular physical object can include a keyboard. In some examples, a particular physical object can include a mouse. In some examples, a particular physical object can include a table or a portion thereof. In some examples, a particular physical object can include a sensor.

[0153] Referring to FIG. 10, in step 1014, the instructions included in the non-transitory computer-readable medium when executed by the processor can cause the processor to receive a selection of a particular physical object from a plurality of physical objects. For example, referring to FIG. 15, the hand gesture 1510 can indicate a tap on the table 102. At least one processor can detect the hand gesture 1510 based on, for example, the image data captured by at least one image sensor having a field of view 1310, and based on the hand gesture 1510, can determine a display of the selection of a particular physical object (e.g., the table 102). In some examples, at least one processor can further determine a particular point or region on or near the selected particular physical object (e.g., the point or region of the table 102 tapped by the hand gesture 1510) for receiving a docking virtual object based on the hand gesture 1510. In some examples, the selected particular physical object can be a part of a physical object (e.g., a particular point or region on the table 102).

[0154] Some of the disclosed embodiments may include receiving a selection of a particular virtual object from a plurality of virtual objects for association with a particular physical object. For example, at least one image sensor of a wearable extended reality device can continuously or periodically monitor a scene within the field of view of the at least one image sensor. At least one processor can detect a display of a selection of a particular virtual object from a plurality of virtual objects based on the monitored scene. The display of the selection of the particular virtual object may include, for example, a gesture by a user indicating the selection of the particular virtual object (e.g., pointing at the particular virtual object, touch-and-hold on the particular virtual object, or tapping on the particular virtual object). In some examples, the selection of the particular virtual object may be from a list or menu that includes a plurality of virtual objects. For example, at least one processor can generate a list or menu that includes a plurality of virtual objects. A user of the wearable extended reality device can invoke the list or menu and select a particular virtual object from the list or menu. The particular virtual object may be selected for association with a particular physical object. In some other examples, the selection of the particular virtual object from a plurality of virtual objects for association with a particular physical object may be received from an external device, may be received from a data structure that associates a desired selection of virtual objects with different physical objects, may be received from a memory unit, may be received from a user, may be received from an automated process (such as from an algorithm that ranks the compatibility of virtual objects with physical objects), and so on.

[0155] Referring to FIG. 10, in step 1016, the instructions included in the non-transitory computer-readable medium when executed by the processor can cause the processor to receive a selection of a particular virtual object from a plurality of virtual objects for association with a particular physical object. Referring to FIG. 16, the hand gesture 1610 can indicate a tap on the virtual widget 114E. At least one processor can detect the hand gesture 1610 based on, for example, image data captured by at least one image sensor having a field of view 1310, and based on the hand gesture 1610, can determine a display of a selection of a particular virtual object (e.g., virtual widget 114E) for association with a particular physical object (e.g., table 102).

[0156] In some examples, both the selection of a particular physical object and the selection of a particular virtual object may be determined from the analysis of image data captured by at least one image sensor of a wearable extended reality device. The image data can include the display of one or more scenes within the field of view of the at least one image sensor. One or more gestures may enter the field of view of the at least one image sensor and be captured by the at least one image sensor. At least one processor can determine, based on the image data, that one or more gestures can indicate the selection of a particular physical object and / or the selection of a particular virtual object. In some examples, both the selection of a particular physical object and the selection of a particular virtual object may be determined from the detection of a single predetermined gesture. The single predetermined gesture can include any one of a variety of gestures, such as a tap, a touch-and-hold, a pointing, or any other desired gesture indicating the selection of an object. In one example, a data structure can associate a gesture with the selection of a pair of one physical object and one virtual object, and the data structure can be accessed based on a single predetermined gesture to select a particular physical object and a particular virtual object. In some other examples, the image data captured by at least one image sensor of a wearable extended reality may be analyzed using an object detection algorithm to detect a particular physical object of a particular category (and / or the location of a particular physical object), and the particular physical object may be selected to be the detected particular physical object. Further, a particular virtual object may be selected based on the particular physical object detected in the image data and / or based on the characteristics of the particular physical object determined by analyzing the image data. For example, the image data can be analyzed to determine the dimensions of a particular physical object, and a virtual object that fits the determined dimensions can be selected as the particular virtual object.

[0157] Some of the disclosed embodiments can include docking a particular virtual object to a particular physical object. Docking a particular virtual object to a particular physical object can include, for example, placing the particular virtual object at a position fixed relative to the particular physical object. In addition to or instead of this, docking a particular virtual object to a particular physical object can include attaching the particular virtual object to the particular physical object. The attachment can be done in any way that establishes an association. For example, depending on the choice of the system designer, physical commands (e.g., hard or soft button presses), voice commands, gestures, hovering, or any other predefined action to cause docking can be used. The particular virtual object can be attached at a particular point or region on or near the particular physical object. The particular point or region on or near the particular physical object can be specified, for example, by a user of a wearable extended reality device. Docking a particular virtual object to a particular physical object can be based on the selection of the particular physical object and the selection of the particular virtual object. In some examples, the selected particular physical object can be a part of the physical object (e.g., a particular point or region on a table), and the particular virtual object can be docked to the part of the physical object.

[0158] In some examples, docking a particular virtual object to a particular physical object may be based on receiving a user input indication for docking the particular virtual object to the particular physical object (e.g., a gesture of dragging the particular virtual object to the particular physical object). By docking the particular virtual object to the particular physical object, the particular virtual object can stay in a position that can be fixed relative to the particular physical object. Based on the docking, the particular virtual object can be displayed as being in a position that can be fixed relative to the particular physical object.

[0159] In some examples, docking a particular virtual object to a particular physical object may include connecting the position of the particular virtual object to the particular physical object. The particular virtual object that is docked can move with the particular physical object to which the particular virtual object is docked. The particular physical object can be associated with a dockable region where the virtual object can be arranged to be docked to the particular physical object. The dockable region can include, for example, a region (e.g., a surface) of the particular physical object, and / or a region surrounding the particular physical object. For example, the dockable region of a table can include the surface of the table. As another example, the dockable region of a keyboard can include a region surrounding the keyboard (e.g., when the keyboard is placed on a support surface), and / or a surface of the keyboard (e.g., a part thereof that does not have buttons or keys). In some examples, a particular virtual object can be dragged to a dockable region (e.g., a region of a particular physical object). In some examples, when a particular virtual object is moved, the dockable region is highlighted (e.g., using virtual content displayed by a wearable extended reality device), and the display of the dockable region is visually presented to the user and the particular virtual object can be placed and docked.

[0160] Referring to FIG. 10, at step 1018, when executed by a processor, the instructions included in the non - transitory computer - readable medium can cause the processor to dock a particular virtual object with a particular physical object.

[0161] Referring to FIG. 17, the hand gestures 1710, 1712, 1714 can indicate a drag of the virtual widget 114E to the table 102. At least one processor can detect the hand gestures 1710, 1712, 1714 based on, for example, image data captured by at least one image sensor having a field of view 1310, and based on the hand gestures 1710, 1712, 1714, can determine a display for docking a particular virtual object (e.g., virtual widget 114E) with a particular physical object (e.g., table 102). The particular virtual object may be docked to a particular point or region on the particular physical object, or may approximate the particular physical object. The particular point or region may be specified by the user. In some examples, the selected particular physical object may be a part of the physical object (e.g., a particular point or region on the table 102), and the particular virtual object may be docked to a part of the physical object. Referring to FIG. 18, a particular virtual object (e.g., virtual widget 114E) is docked to a particular physical object (e.g., table 102, or a particular point or region on the table 102).

[0162] In some examples, a particular physical object may be a keyboard, and at least some of the plurality of virtual objects may be associated with at least one rule that defines the behavior of the virtual objects when the wearable extended reality device is detached from the docking keyboard. Virtual object behavior can include, for example, undocking a docked virtual object, deactivating a docked virtual object, hiding a docked virtual object, rendering a non-dockable virtual object, or other desired actions. At least one rule that defines virtual object behavior may be configured to cause at least one processor to perform one or more of the virtual object behaviors on one or more corresponding virtual objects based on a wearable extended reality device that is a threshold distance (e.g., 1 meter, 2 meters, 3 meters, 5 meters, 10 meters, or any other desired length) away from the docked keyboard. In some examples, at least one rule may be based on at least one of the type of keyboard, time, or at least one of the other virtual objects docked to the keyboard. A rule can include any instruction that causes an effect based on an occurrence or a condition for the occurrence. For example, movement of the keyboard relative to a virtual object can potentially cause some change to the virtual object (e.g., the virtual object can move with the keyboard). Alternatively or in addition, the effect may be conditional. For example, at least one rule may be configured to cause at least one processor to perform a first virtual object action on a virtual object for a keyboard in an office and a second virtual object action on a virtual object for a keyboard in a home. As another example, at least one rule may be configured to cause at least one processor to perform a first virtual object behavior on a virtual object during non-business hours and a second virtual object behavior on the virtual object during business hours.As a further example, at least one rule may be configured to cause at least one processor to execute a first virtual object behavior of a virtual object when no other virtual object is docked to the keyboard, and to execute a second virtual object behavior of the virtual object when another virtual object is docked to the keyboard.

[0163] In some embodiments, when a particular physical object and a particular virtual object are outside of a portion of the field of view associated with the display system of the wearable extended reality device such that the particular virtual object is not visible to the user of the wearable extended reality device, these embodiments can include receiving a gesture input indicating that the hand is interacting with the particular virtual object. After the particular virtual object is docked to the particular physical object, the user can move the wearable extended reality device, and the position and / or orientation of the wearable extended reality device can change. Based on the change in position and / or orientation, a portion of the field of view associated with the display system of the wearable extended reality device may not cover the position where the particular virtual object is docked. Thus, the particular virtual object may not be displayed to the user and may become invisible to the user. When the position where the particular virtual object is docked is outside of a portion of the field of view associated with the display system, the user can perform a gesture to interact with the position where the particular virtual object is docked.

[0164] At least one image sensor of the wearable extended reality device can capture image data within the field of view of the at least one image sensor and can detect gesture inputs. Gesture inputs can include, for example, drag, pinch, spread, swipe, tap, pointing, grab, scroll, rotate, flick, touch, zoom in, zoom out, thumb up, thumb down, touch and hold, or any other movement of one or more fingers or hands. In some examples, the gesture input may be determined from the analysis of the image data captured by at least one image sensor of the wearable extended reality device. Based on the image data (e.g., based on the analysis of the image data using a gesture recognition algorithm), at least one processor can determine that the gesture input may indicate that a hand or a part thereof is interacting with a particular virtual object. The hand interacting with the particular virtual object can be determined based on the hand interacting with a particular physical object. In some examples, the hand interacting with the particular virtual object can be the hand of the user of the wearable extended reality device. In some examples, the hand interacting with the particular virtual object can be the hand of a person other than the user of the wearable extended reality device.

[0165] In some examples, the gesture input may be determined from the analysis of additional sensor data obtained by additional sensors associated with an input device connectable to the wearable extended reality device. The additional sensors may be disposed, connected, or associated with any other physical element. For example, the additional sensors may be associated with a keyboard. Thus, even if the user looks away from the keyboard that renders the keyboard outside the user's field of view, the additional sensors on the keyboard can sense gestures. The additional sensors are preferably image sensors, but do not necessarily have to be. Further, the keyboard-related sensors are just one example. The additional sensors may be associated with any item. For example, the additional sensors may include sensors on gloves such as haptic gloves, wired gloves, or data gloves. The additional sensors can include hand position sensors (e.g., image sensors or proximity sensors included in an article). In addition to or instead of this, the additional sensors can include the sensors described herein in relation to hand position sensors included in a keyboard. The sensors may be included in input devices other than the keyboard. The additional sensor data can include the position and / or orientation of a hand or part of a hand. In some examples, a particular physical object may include a part of a touchpad, and the gesture input may be determined based on the analysis of touch input received from the touchpad. In some examples, a particular physical object may include a sensor, the gesture input may be received from the sensor, and / or may be determined based on data received from the sensor.In some examples, although a particular physical object and a particular virtual object are within a portion of the field of view associated with the display system of the wearable extended reality device, they are hidden by another object (virtual or physical), so that when the particular virtual object is not visible to the user of the wearable extended reality device, some of the disclosed embodiments may include, for example, receiving a gesture input indicating that a hand is interacting with the particular virtual object from an analysis of further sensor data obtained by a further sensor associated with an input device connectable to the wearable extended reality device as described above.

[0166] Referring to FIG. 10, at step 1020, when executed by a processor, instructions included in a non-transitory computer-readable medium can cause the processor to receive a gesture input indicating that a hand is interacting with a particular virtual object such that when the particular physical object and the particular virtual object are outside a portion of the field of view associated with the display system of the wearable extended reality device, the particular virtual object is not visible to the user of the wearable extended reality device.

[0167] Referring to FIG. 19, the position and / or orientation of the wearable extended reality device can change, and the field of view 1310 and the field of view portion 1312 can change their respective coverage areas. A particular point or region where a particular virtual object (e.g., virtual widget 114E) is docked has moved outside the field of view portion 1312, and the virtual widget 114E is no longer visible to the user. A gesture 1910 can indicate an interaction with the point or region where the virtual widget 114E is docked. The gesture 1910 can occur within the field of view 1310 and can be captured by at least one image sensor of the wearable extended reality device. Based on the gesture 1910, at least one processor can determine that a hand may be interacting with the virtual widget 114E.

[0168] Some of the disclosed embodiments may include causing an output associated with a particular virtual object in response to a gesture input. The output associated with a particular virtual object can be based on the gesture input and can be generated as if the gesture input is interacting with the particular virtual object, even if the particular virtual object is not visible to the user of the wearable extended reality device. For example, if the gesture input is a tap, the output associated with the particular virtual object can include activating the particular virtual object. If the gesture input is a scroll up or a scroll down, the output associated with the particular virtual object can include, for example, moving a scroll bar up and down. In some examples, causing an output associated with a particular virtual object can include changing a previous output associated with the particular virtual object. Thus, even if a particular virtual object is outside the user's field of view, the user may be permitted to operate the virtual object. Such operations can include, for example, changing the volume (if the virtual object is a volume control), rotating or repositioning another virtual object associated with the particular virtual object (if the particular virtual object is a positioning control device), changing display parameters (if the particular virtual object is a display parameter control), moving the particular virtual object from outside the user's field of view into the user's field of view, and, depending on the choice of the system designer, activating a function, or any other operation.

[0169] Referring to FIG. 10, in step 1022, the instructions included in the non-transitory computer-readable medium when executed by the processor can cause the processor to perform an output related to a specific virtual object in response to a gesture input. Referring to FIG. 19, gesture 1910 may include, for example, a scroll up. In response to gesture 1910, at least one processor can cause an output associated with virtual widget 114E. For example, virtual widget 114E may be a volume bar, and the output associated with virtual widget 114E may increase the volume.

[0170] As previously suggested, in some examples, the specific virtual object may be a presentation control, and the output may include a change in presentation parameters related to the presentation control. The presentation control can include, for example, a start button, a next page button, a previous page button, an end button, a pause button, a last viewed button, a volume bar, a zoom in button, a zoom out button, a viewpoint selection, a security control, a privacy mode control, or any other function for controlling the presentation. In response to a gesture input indicating an interaction with a presentation control (which may not be visible to the user, for example), at least one processor can cause a change in the presentation parameters associated with the presentation control (such as a change in the page number of the presentation). In some examples, the presentation parameters can include at least one of volume, zoom level, viewpoint, security, or privacy mode. For example, the volume of the presentation may be changed. As another example, the zoom level of the presentation may be changed, or the viewpoint of the presentation may be changed. For example, by changing the presentation parameters, the security level of the presentation can be changed to modify the access of the viewer to the presentation based on the security authentication information or level associated with the viewer. As yet another example, by changing the presentation parameters, the privacy mode of the presentation can be changed to modify the access of the viewer to the presentation based on the trust level of the viewer.

[0171] In some examples, a particular virtual object may be an application icon, and the output may include virtual content associated with an application inside a portion of the field of view associated with the display system of the wearable extended reality device. For example, a particular virtual object may be an application icon such as a weather application, an email application, a web browser application, an instant messaging application, a shopping application, a navigation application, or any other computer program for performing one or more tasks. In response to a gesture input indicating an interaction with an application icon (which may not be visible to the user, for example), at least one processor can, for example, launch the application and display the virtual content of the application inside a portion of the field of view associated with the display system of the wearable extended reality device.

[0172] Some of the disclosed embodiments can include recommending default positions for a plurality of virtual objects based on associated functionality. The default position can be a specific point or area where the virtual object can be docked. At least one processor can determine the functionality associated with a plurality of virtual objects displayed in a portion of the field of view associated with the display system of the wearable extended reality device, for example, by accessing a data structure that associates the virtual objects with functionality. Based on the functionality, at least one processor can determine the default positions of the plurality of virtual objects. For example, if a virtual object is associated with the function of controlling volume, the default position of the virtual object can include, for example, the upper region of the keyboard, the corner of a table, or any other desired position for docking the virtual object. If a virtual object is associated with an application function (e.g., an email application), the default position of the virtual object can include, for example, the central region of a table or any other desired position for docking the virtual object. In another example, if a particular virtual object function is associated with one or more gestures that benefit from tactile feedback, the default position of the particular virtual object can be on a surface. At least one processor can provide a recommendation of the default position to the user, for example, when a virtual object is selected for docking. In some examples, the recommended default position may be displayed to the user as an overlay on top of a physical object in a portion of the field of view associated with the display system of the wearable extended reality device.

[0173] In some examples, a wearable extended reality device may be configured to pair with multiple different keyboards. Some of the disclosed embodiments can include changing the display of certain virtual objects based on a selected paired keyboard. Different keyboards may be specific physical objects for docking certain virtual objects. For example, a first keyboard may be in a home and may be associated with a first setting of a certain virtual object. A second keyboard may be in an office and may be associated with a second setting of a certain virtual object. The first setting may be different from the second setting. Different keyboards may be paired with a wearable extended reality device (e.g., can be used as an input device of a wearable extended reality device). A certain virtual object (e.g., a volume bar) may be docked to the first keyboard or may be docked to the second keyboard. The certain virtual object may be configured with different settings and may be displayed differently when docked with different keyboards. For example, a certain virtual object may be configured with a first setting when docked with the first keyboard at home and may be configured with a second setting when docked with the second keyboard in the office. As another example, a certain virtual object may be displayed to have a larger size when docked with the first keyboard at home than when docked with the second keyboard in the office.

[0174] Some disclosed embodiments can include virtually displaying a particular virtual object proximate to a particular physical object when it is determined that the particular physical object enters a portion of the field of view associated with the display system of the wearable extended reality device. The particular physical object (e.g., a particular point or region on a table) can enter a portion of the field of view associated with the display system of the wearable extended reality device. At least one processor can determine the entry of the particular physical object into a portion of the field of view based on, for example, image data captured by at least one image sensor of the wearable extended reality device. The newly captured image data can be compared with stored image data, and the physical object can be recognized based on the comparison. This can occur even if only a portion of the physical object enters the field of view since the comparison can be made based on only a portion of the physical object (e.g., the comparison can be performed on a set of image data of the physical object that is less than a complete set of the image data of the physical object). When the physical object or a portion thereof is recognized, at least one processor can cause the virtual presentation of a particular virtual object proximate to the particular physical object (e.g., even if the particular physical object does not fully enter a portion of the field of view). In some examples, the particular virtual object can be docked in a position proximate to the particular physical object (e.g., a keyboard). For example, a volume bar can be docked in a position adjacent to the keyboard. At least one processor can cause the virtual presentation of a particular virtual object proximate to the particular physical object when determining the entry of the particular physical object into a portion of the field of view.

[0175] Some of the disclosed embodiments can include determining whether the hand interacting with a particular virtual object is the hand of a user of a hand-worn extended reality device. (As used throughout, reference to a hand includes the complete hand or a part of a hand.) Hand source recognition can be performed in any one of several ways, or a combination of ways. For example, based on the relative position of the hand, the system can determine that the hand is that of the user of the hand-worn extended reality device. This can occur because the hand of the wearer can have an orientation within the field of view that is a telltale sign that the hand is the wearer's hand. This can be determined by performing image analysis on the current image of the hand and comparing it to stored images or image data associated with the orientation of the wearer's hand. Similarly, the hand of a person other than the wearer can have a different orientation, and the system can determine that the detected hand is that of a person other than the wearer, in the same way that the system determines that the hand is that of one of the wearers of the extended reality device.

[0176] As another example, the hands of different individuals can be different. The system can become capable of recognizing the hand of the extended reality device wearer by inspecting the unique characteristics of the skin or structure. As the wearer uses the system, the characteristics of the wearer's hand can be stored in a data structure, and image analysis can be performed to confirm that the current hand in the image is the wearer's hand. Similarly, the hands of individuals other than the wearer can be recognized, enabling the system to distinguish between multiple individuals based on the characteristics of those hands. This feature can enable unique system control. For example, when multiple individuals interact with the same virtual object within the same virtual space, the control and movement of the virtual object can vary based on the person interacting with the virtual object.

[0177] Accordingly, at least one processor can determine whether a hand interacting with a particular virtual object (e.g., gesture 1910) is the hand of a user of the wearable extended reality device, based on, for example, image data captured by at least one image sensor of the wearable extended reality device. The at least one processor can analyze the image data for hand identification, for example, by determining hand features based on the image data and comparing the determined hand features to stored features of the hand of the user of the wearable extended reality device. In some examples, the at least one processor can perform hand identification based on a particular object associated with the hand (e.g., a particular ring for identifying the user's hand). Some of the disclosed embodiments can include causing an output associated with a particular virtual object in response to a hand and gesture input that is the hand of the user of the wearable extended reality device interacting with the particular virtual object. For example, if the hand interacting with a particular virtual object (e.g., gesture 1910) is the hand of the user of the wearable extended reality device, the at least one processor can cause an output associated with the particular virtual object. Further disclosed embodiments can include refraining from causing an output associated with a particular virtual object in response to the hand interacting with the particular virtual object not being the hand of the user of the wearable extended reality device. For example, if the hand interacting with a particular virtual object (e.g., gesture 1910) is not the hand of the user of the wearable extended reality device, the at least one processor can refrain from causing an output associated with the particular virtual object.

[0178] When working in an extended reality environment, it can be difficult for a user to know whether the system has recognized the object intended for the kinematic selection (e.g., through a gesture or through a gaze), or whether the kinematic selection is accurate enough to indicate a single object. Thus, it may be useful to provide the user with intermediate feedback indicating that the kinematic selection is not accurate enough to indicate a single object. For example, at the start of a gesture, if the gesture is not very specific, a group of virtual objects in the direction of the gesture may be highlighted. As the user refines the selection, surrounding objects may stop being highlighted until only the selected object remains highlighted. This can show the user that the initial gesture is in the correct direction but needs to be further refined to indicate a single virtual object.

[0179] Various embodiments can be described with reference to systems, methods, apparatuses, and / or computer-readable media. One disclosure is not intended to be all disclosures. For example, as described herein, the disclosure of one or more processes implemented on a non-transitory computer-readable medium also constitutes the disclosure of a method implemented by a computer-readable medium and, for example, a system and / or apparatus for implementing a process implemented on a non-transitory computer-readable medium via at least one processor. Thus, in some embodiments, a non-transitory computer-readable medium can include instructions that, when executed by at least one processor, can cause the at least one processor to perform an incremental convergence operation in an extended environment. Some aspects of such processes can be performed electronically via a network that can be wired, wireless, or both. Other aspects of such processes may be performed using non-electronic means. In the broadest sense, the processes disclosed herein are not limited to specific physical or electronic means; rather, they can be achieved using any number of different means. As previously mentioned, the term non-transitory computer-readable medium should be construed expansively to encompass any medium that can store data in a manner readable by any computing device having at least one processor for executing operations, methods, or any other instructions stored in a memory.

[0180] A non-transitory computer-readable medium can include, for example, random access memory (RAM), read-only memory (ROM), volatile memory, non-volatile memory, hard drives, compact disc read-only memory (CD-ROM), digital versatile discs (DVD), flash drives, disks, any optical data storage media, any physical media having a pattern of holes, programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), FLASH-EPROM or any other flash memory, non-volatile random access memory (NVRAM), caches, registers, any other memory chips or cartridges, or network versions thereof. The non-transitory computer-readable medium can refer to multiple structures such as multiple non-transitory computer-readable media located at a local position or a remote position. Further, one or more non-transitory computer-readable media can be utilized when implementing a computer-implemented method. Thus, the non-transitory computer-readable medium can include tangible items and can exclude carrier waves or transient signals.

[0181] Instructions included in the non-transitory computer-readable medium can include, for example, software instructions, computer programs, computer code, executable instructions, source code, machine instructions, machine language programs, or any other type of instructions for a computing device. The instructions included in the non-transitory computer-readable medium can be based on one or more various types of desired programming languages and can include various processes for detecting, measuring, processing, and / or implementing incremental convergence in the extended reality environment described herein, or for providing feedback based thereon (e.g., for embodying).

[0182] At least one processor may be configured to execute instructions included in a non-transitory computer-readable medium to cause various processes to be executed to perform incremental convergence in the extended reality environment described herein. The processor can include, for example, an integrated circuit, a microchip, a microcontroller, a microprocessor, a central processing device (CPU), a graphics processing device (GPU), a digital signal processor (DSP), a field programmable gate array (FPGA), or other unit suitable for executing instructions or performing logical operations. The processor can include a single-core or multi-core processor. In some examples, the processor may be a single-core processor configured with virtualization technology. The processor can implement virtualization technology or other functions, for example, to provide the ability to execute, control, execute, operate, or store a plurality of software processes, applications, or programs. In another example, the processor can include a multi-core processor (e.g., dual-core, quad-core, or any desired number of cores) configured to provide a parallel processing function to enable a device associated with the processor to execute multiple processes simultaneously. Other types of processor configurations may be implemented to provide the functions described herein.

[0183] Some of the disclosed embodiments can relate to performing an incremental convergence operation in an extended reality environment (such as a virtual reality environment, an augmented reality environment, a mixed reality environment, etc.). The term incremental convergence can relate not only to the start and stop of a user's physical interaction with the extended reality environment, but also to any increase or decrease, or otherwise be represented. A user's physical interaction with the extended reality environment can function as a motion input that enables the user to select or otherwise interact with a virtual object within the coordinate system of the extended reality environment. The interaction with the extended reality environment can include hand gestures, body gestures, eye movements, and / or any combination of gesture-based interactions or gesture-based interactions. The incremental convergence operation can be related to a process for honing a selection or item. This can occur gradually over time such that the targeting becomes more acute or approaches a final selection or particular item. Stepwise convergence may occur based on feedback that helps to indicate a selection. For example, the feedback may be provided to the user during a kinematic selection process of virtual objects within the extended reality environment when the user approaches or otherwise engages with at least one of a plurality of virtual objects displayed on a user interface of the extended reality environment.

[0184] Some of the disclosed embodiments can include displaying a plurality of distributed virtual objects across a plurality of virtual regions. The virtual objects can be visually presented to a user in an extended reality environment via an extended reality device and can be related to virtual content including any type of data display with which the user can interact. The virtual objects can be considered to be distributed if they are not all placed at exactly the same location. For example, virtual objects can be considered to be distributed if they appear at different locations on a virtual display or in an extended reality environment. Objects that are located adjacent to each other are an example of distributed objects. The virtual display or extended reality environment can have regions that are close to each other or spaced apart from each other. In either case, the objects can be considered to be distributed across a plurality of regions.

[0185] The virtual objects can include any visual presentation rendered by a processing device such as a computer or a virtual object of an inanimate or living thing. In some cases, the virtual objects can be configured to change over time or in response to a trigger. In some embodiments, the virtual objects of a plurality of distributed virtual objects can include at least one virtual widget and / or option in a virtual menu. Virtual objects such as at least one virtual widget and / or option in a virtual menu can be used to issue commands, initiate a dialog sequence, change an interaction mode, and / or provide the user with quick access to information without fully opening the associated virtual object. In addition to or instead of this, the virtual objects can include at least one icon that can activate a script to cause an action associated with a particular virtual object associated with the icon, and / or can be linked to an associated program or application.

[0186] A plurality of distributed virtual objects may be associated with any number of virtual objects that can be visually displayed or presented to a user via a user interface of an extended reality environment. For example, the plurality of distributed virtual objects may be distributed across a virtual computer screen (also referred to herein as a virtual display) that displays virtual objects generated by an operating system. In some embodiments, the plurality of virtual regions are a plurality of virtual regions of the extended reality environment, and displaying the plurality of distributed virtual objects may include causing a wearable extended reality device to display the plurality of distributed virtual objects.

[0187] The user interface of an extended reality environment may be configured to display a plurality of virtual regions. The plurality of virtual regions may be related to at least two individual or connected subsets of space within the entire space of the user interface. Each of the plurality of virtual regions may be defined by at least one of its boundary, internal region, external region, or elements (such as virtual objects) included in the virtual region. The boundary may or may not be visually presented. For example, in a visual sense, the boundary of a given virtual region can be represented by any type of closed geometric shape that encompasses a subset of the entire space of the user interface and forms an interface between the internal region and the external region of the virtual region. The internal region of a virtual region can be related to a subset of the space inside the boundary of the virtual region and can include at least one of a plurality of dispersed virtual objects. Thus, a given virtual region can include a single virtual object or a number of virtual objects. The external region of a virtual region can be related to a subset of the space outside the boundary of the virtual region and can include a number of dispersed virtual objects excluding the virtual objects or virtual objects included inside the virtual region. A virtual object located in the external region of a virtual region may be located in the internal region of another virtual region among the plurality of virtual regions. In other embodiments, color or other visual characteristics can indicate different regions. Also, in other embodiments, regions can be launched together without a perceptible divider.

[0188] The plurality of virtual regions may include at least a first virtual region and a second virtual region different from the first virtual region. This term may indicate that the first virtual region includes at least one virtual object among a plurality of distributed virtual objects not included in the second virtual region, and the second virtual region may include at least one virtual object among a plurality of distributed virtual objects not included in the first virtual region. Thus, the plurality of virtual regions includes a given virtual object and at least one non-overlapping region. In addition to or instead of this, all of the virtual objects in the first virtual region may not be included in the second virtual region, and all of the virtual objects in the second virtual region may not be included in the first virtual region. In some embodiments, the first virtual region may be selected based on an initial motion input. As will be described in detail below, the initial motion input can include an interaction based on any gesture by which the user begins to show interest in or otherwise pay attention to a virtual region such as the first virtual region, for example, by first pointing to and / or looking at the first virtual region.

[0189] In some embodiments, the boundary of the first virtual region may be adaptable in real time in response to motion input. For example, the boundary of the first virtual region may be dynamically adaptable to include or exclude different numbers of a plurality of distributed virtual objects based on the user's motion input within the coordinate system of the extended reality environment. In one example, the boundary of the first virtual region may include a first number of a plurality of virtual objects based on a threshold confidence level at which the user may intend to make a kinematic selection of one of the first number of a plurality of virtual objects. In one non-limiting example, when at least one processor determines that the user may intend to select at least one of some adjacent virtual objects of a plurality of virtual objects based on the user pointing to the virtual object, the boundary may be defined to include each of the virtual objects that the user may intend to interact with beyond a threshold confidence level, for example, beyond a 50% confidence level. In another example, the boundary of the first virtual region may include a second number of a plurality of virtual objects that is less than the first number of a plurality of virtual objects based on a higher threshold confidence level at which the user may intend to make a kinematic selection of one of the second number of a plurality of virtual objects.

[0190] FIG. 20 is a flowchart showing an exemplary process 2000 for providing feedback based on incremental convergence in an extended reality environment, consistent with some embodiments of the present disclosure. Referring to step 2010 of FIG. 20, instructions included in a non-transitory computer-readable medium, when executed by at least one processor, may cause the at least one processor to display a plurality of distributed virtual objects across a plurality of virtual regions. As an example, FIG. 21 shows a non-limiting example of an extended reality environment that displays a plurality of distributed virtual objects across a plurality of virtual regions, consistent with some embodiments of the present disclosure.

[0191] Referring to FIG. 21, the user interface 2110 is presented within an exemplary extended reality environment from the perspective of a user of the extended reality environment. The user interface 2110 can display virtual content including a plurality of distributed virtual objects 2112 to the user such that the user can interact with the plurality of distributed virtual objects 2112 across the user interface 2110. The plurality of distributed virtual objects 2112 presented to the user in the extended reality environment can include at least one interactive virtual widget and / or option within a virtual menu. Within a given virtual object of the plurality of distributed virtual objects 2112, there can be additional virtual objects with which the user can interact. In some examples, the user interface 2110 can be associated with a virtual computer screen configured to display to the user a plurality of distributed virtual objects 2112 generated by an operating system such that the user can interact with the plurality of distributed virtual objects 2112.

[0192] A plurality of distributed virtual objects 2112 may be displayed across a plurality of virtual regions of the user interface 2110. For example, the user interface 2110 may include a first virtual region 2114 that includes a first subset of the plurality of distributed virtual objects 2112 and a second virtual region 2116 that includes a second subset of the plurality of distributed virtual objects 2112. The boundary of the first virtual region 2114 is depicted as a dashed line and is shown to encompass the first subset of the plurality of distributed virtual objects 2112, and the boundary of the second virtual region 2116 is depicted as a dashed line and is shown to encompass the second subset of the plurality of distributed virtual objects 2112 such that each of the plurality of distributed virtual objects 2112 is included in either the first virtual region 2114 or the second virtual region 2116. It should be understood that the boundaries of the virtual regions may not be visible to the user and are shown herein for illustrative purposes. Further, the boundaries of the first virtual region 2114 and the second virtual region 2116 in this figure are depicted to include arbitrarily selected first and second subsets of the plurality of distributed virtual objects 2112 for simplicity and may vary to include different numbers of the plurality of distributed virtual objects 2112 based on several factors including the user's past and / or current interactions with the extended reality environment.

[0193] Some of the disclosed embodiments can include receiving an initial motion input that tends towards a first virtual region. The term motion input can relate to a user's natural human movements and / or gestures. Such movements can occur in a real coordinate system that can be tracked and / or captured as input data that can be converted into virtual movements within the virtual coordinate system of an extended reality environment. A user's motion input can include hand gestures, body gestures (e.g., movements of the user's head or body), eye movements or other ocular movements, and / or any other gesture-based interactions, facial movements, or combinations of gesture-based interactions. In another example, a user's motion input can include the movement of a computer cursor in a virtual display and / or an extended reality environment. A user's motion input is detected and converted into a virtual interaction with the extended reality environment, thereby enabling the user to select a virtual object within the extended reality environment or otherwise interact with it.

[0194] In some embodiments of the present disclosure, the receiving of the motion input can be related to the capture and / or recognition of at least one motion input by components of a wearable extended reality device such as a stereoscopic head-mounted display, head motion tracking sensors (e.g., gyroscopes, accelerometers, magnetometers, image sensors, structured light sensors), eye tracking sensors, and / or any other component capable of recognizing at least one of the user's various sensory motion inputs and converting the sensed motion input into a readable electrical signal. In some examples, these components are utilized to detect or measure the speed and / or acceleration of a user's motion input and to detect or otherwise measure the time difference between specific motion inputs in order to distinguish the user's intent based on interactions based on the user's subtle gestures.

[0195] Furthermore, the term initial motion input can relate to any motion input where the user begins to show interest or otherwise pay attention to a virtual region such as a first virtual region that includes at least one virtual object. The user can indicate interest or otherwise pay attention to a first virtual region that includes at least one virtual object by generating a motion input that includes at least one slight movement towards the first virtual region that includes at least one virtual object. This slight movement towards the first virtual region can enable the user to receive a substantially seamless feedback of the motion input perceived within the extended reality environment and intuitively discover the result of the perceived motion input. For example, a user attempting to make a kinematic selection of a virtual object within the extended reality environment by way of an initial motion input such as pointing at and / or looking at a first virtual region that initially includes a group of virtual objects may be provided with initial feedback and / or intermediate feedback based on the initial motion input if the intention of the interaction is not specific enough to identify an intention to select a particular virtual object of interest.

[0196] In some examples, the first region and / or the second region may be pre-determined, for example, based on the system configuration. In some other examples, the first region and / or the second region may be created, changed, and / or discarded dynamically. In some examples, the first region and / or the second region may be determined based on a plurality of distributed virtual objects. In one example, a change to a plurality of distributed virtual objects (e.g., addition of a virtual object, deletion of a virtual object, change in position of a virtual object, change in state of a virtual object, etc.) can cause a change to the first region and / or the second region. Some non-limiting examples of such states of virtual objects can include virtual objects not associated with a notification, virtual objects associated with a notification, closed virtual objects, open virtual objects, active virtual objects, non-active virtual objects, etc.

[0197] In one example, at a given point in time (e.g., at system startup, at the time of change, etc.), a clustering algorithm can be used to cluster virtual objects into regions based on at least one of, for example, the position of the virtual object, the identity of the virtual object, or the state of the virtual object. In another example, at a given point in time (e.g., at system startup, at the time of change, etc.), a segmentation algorithm can be used to segment at least a part of the extended reality environment and / or at least a part of the virtual display into regions.

[0198] In some examples, the first region and / or the second region may be determined based on the dimensions of the virtual display and / or the extended reality environment. In one example, a change in the dimensions of the virtual display and / or the extended reality environment can cause a change in the first region and / or the second region. In one example, at a given point in time (e.g., at system startup, at the time of change, etc.), the virtual display and / or the extended reality environment can be divided, for example, into a grid of regions of a predetermined number of equal (or approximately equal) sizes, regions of a predetermined fixed size, and / or regions of various sizes.

[0199] In some examples, the first region and / or the second region may be determined based on conditions related to the extended reality device (such as ambient lighting conditions in the environment of the extended reality device, the movement of the extended reality device, the usage mode of the extended reality device, etc.). In one example, a change in the conditions related to the extended reality device can cause a change in the first region and / or the second region. In one example, brighter ambient lighting within the environment of the extended reality device can trigger the use of a larger region. In another example, a moving extended reality device can have a larger region than a non - moving extended reality device.

[0200] In some examples, the first region and / or the second region may be determined based on a motion input, e.g., based on at least one of a direction, a position, a pace, or an acceleration associated with the motion input. For example, a machine learning model can be trained using training examples to determine regions based on motion inputs. An example of such a training example can include a sample motion input and a sample distribution of the sample virtual object, along with a label indicating a desired region corresponding to the sample motion input and the sample distribution. In another example, a motion input associated with faster movement can cause the determination of a region that is farther from the position associated with the motion input than a motion input associated with slower movement.

[0201] Referring to step 2012 of FIG. 20, instructions included in a non-transitory computer-readable medium, when executed by at least one processor, may cause the at least one processor to receive an initial motion input that tends toward a first virtual region. In one example, the initial motion input may be received after the display of a plurality of distributed virtual objects over a plurality of virtual regions by step 2010. As an example, FIG. 22 shows a non-limiting example of an initial motion input of a user that tends toward a first virtual region of the extended reality environment of FIG. 21, consistent with some embodiments of the present disclosure.

[0202] Referring to FIG. 22, the user interface 2210 of the extended reality environment includes a first subset of a plurality of distributed virtual objects 2212 in a first virtual region 2214, grouped in a region different from a second subset of a plurality of distributed virtual objects 2216 in a second virtual region 2218. The first virtual region 2214 may be dynamically defined to include any number of distributed virtual objects, and the user 2220 may begin to show interest in or otherwise pay attention to the plurality of distributed virtual objects by generating an initial motion input including at least one gesture to the virtual objects. For example, the first subset of the plurality of distributed virtual objects 2212 to be included in the first virtual region 2214 may be determined based on the initial motion input from the user 2220. If it is determined that the user 2220 may want to interact with the first subset of the plurality of distributed virtual objects 2212, the first virtual region 2214 may be defined to include the virtual objects 2212. In response to any given initial motion input, the number of virtual objects 2212 included in the first virtual region 2214 may vary based on the specificity associated with the initial motion input of the user 2220 and / or the determined confidence level.

[0203] In the non - limiting example shown in FIG. 22, user 2220 is shown to express interest in at least one virtual object out of a first subset of a plurality of distributed virtual objects 2212 within a first virtual region 2214 by a hand gesture that points to the first virtual region 2214 that includes the first subset of the plurality of distributed virtual objects 2212. At the time shown in FIG. 22, an initial motion input of user 2220 attempting to make a kinematic selection of at least one out of the first subset of the plurality of distributed virtual objects 2212 may not be specific enough to clearly detect a particular one of the virtual objects 2212 that the user 2220 may wish to interact with. However, at least one processor can determine that the initial motion input of user 2220 is not intended to make a kinematic selection of at least one out of a second subset of a plurality of distributed virtual objects 2216 based on the user's hand gesture.

[0204] It should be understood that the hand of user 2220 shown in the figure does not necessarily represent a virtual image of the hand as seen by user 2220. Rather, in the illustrated example, the virtual hand of user 2220 shown is captured and transformed into a virtual interaction within the corresponding virtual coordinate system of the extended reality environment, thereby being a diagram of one type of motion input within the real - world coordinate system that enables the user to select or otherwise interact with virtual objects within the extended reality environment. However, in some embodiments, a virtual image of the hand of user 2220 can also be overlaid on the extended reality environment. Further, in some embodiments, the initial motion input of user 2220 may additionally or alternatively include an eye gesture of user 2220.

[0205] In some embodiments, the initial motion input can include a combination of gaze detection and gesture detection. The gaze detection and gesture detection of the initial motion input can be related to the capture and / or recognition of the user's hand and / or body gestures and / or eye movements. In a general sense, the gaze detection and gesture detection can be performed in a similar manner to the detection of the aforementioned user motion input. The gaze detection and gesture detection of the user's gestures and eye movements may be performed simultaneously, or both may be used to cause at least one processor to determine the intention of the user's combined initial motion input. In one example, when the detected gaze is directed towards a first virtual region and the detected gesture is directed towards a second virtual region, the non-transitory computer-readable medium can cause at least one processor to highlight a group of virtual objects within the first virtual region. In another example, when the detected gaze is directed towards a first virtual region and the detected gesture is directed towards a region without virtual objects, the non-transitory computer-readable medium can cause at least one processor to highlight a group of virtual objects within the first virtual region.

[0206] In some examples, the determination of the confidence level associated with a motion input, such as the initial motion input, can include determining the confidence level associated with the gaze detection and the confidence level associated with the gesture detection. In addition to or instead of this, the determination to prioritize the detected gaze over the detected gesture and vice versa can be informed or otherwise weighted based on an analysis of past and / or real-time motion inputs, including the user's previously memorized interactions with the extended reality environment, pre-selected and / or analyzed user preferences, artificial intelligence, machine learning, deep learning, and / or neural network processing techniques.

[0207] Referring to FIG. 22, user 2220 may simultaneously display a first interest in at least one virtual object of a first subset of a plurality of distributed virtual objects 2212 within the first virtual region 2214 by a first gesture towards the first virtual region 2214, and may display a second interest in at least one virtual object of a second subset of a plurality of distributed virtual objects 2216 within the second virtual region 2218 by a second gesture towards the second virtual region 2218. Alternatively, the second interest of the user may be directed by the second gesture to a region other than the first virtual region 2214 that does not include virtual objects.

[0208] For example, user 2220 may simultaneously display an interest in at least one virtual object within one virtual region via the line of sight, and may display a second interest in at least one other virtual object within a different virtual region or in a region that does not include virtual objects via a hand gesture. In some cases, when the gesture towards the first virtual region 2214 is related to the line of sight of user 2220 and the gesture towards the second virtual region 2218 is related to a hand gesture, the non-transitory computer-readable medium may cause at least one processor to highlight a group of virtual objects 2212 within the first virtual region 2214 based on the detected fixation and the detected gesture. In other examples, when the gesture towards the first virtual region 2214 is related to the line of sight of user 2220 and the hand gesture tends to be towards a region without virtual objects, the non-transitory computer-readable medium may cause at least one processor to highlight a group of virtual objects 2212 within the first virtual region 2214 based on the detected fixation and the detected gesture.

[0209] Some of the disclosed embodiments can include highlighting a group of virtual objects within a first virtual region based on an initial motion input. The group of virtual objects within the first virtual region that can be highlighted in response to the initial motion input can include any set of virtual objects within a region where the user can show interest or pay attention. This can occur, for example, by gesture and / or gaze in the direction of the group. Depending on any given initial motion input, the number of virtual objects included in the group of virtual objects within the first virtual region can depend on the specificity of the user's motion input and / or a threshold confidence level regarding the intent of the motion input. Further, the first virtual region can be dynamically defined to include groups of adjacent virtual objects where the user begins to show interest or otherwise starts to pay attention by generating an initial motion input to the group of virtual objects. In some embodiments, the highlighted group of virtual objects within the first virtual region can include at least two virtual objects out of a plurality of dispersed virtual objects. In one example, the highlighted group of virtual objects within the first virtual region can include all of the virtual objects within the first virtual region. In another example, the highlighted group of virtual objects within the first virtual region can include some but not all of the virtual objects within the first virtual region.

[0210] The group of virtual objects within the first virtual region may be highlighted in response to an initial motion input that tends toward the group of virtual objects within the first virtual region. In some embodiments, highlighting the group of virtual objects may be associated with the presentation of any visually perceptible indicator that can help emphasize or otherwise distinguish the group of objects within the first virtual region from virtual objects outside the first virtual region (e.g., virtual objects within a second region) for the user. In some embodiments, highlighting the group of virtual objects within the first virtual region based on the initial motion input can include changing the appearance of each virtual object within the group of virtual objects within the first virtual region. For example, the highlighting can include enlarging at least one virtual object within the group, changing the color of at least one virtual object within the group, causing or otherwise changing a shadow or glow around the frame of at least one virtual object within the group, and / or blinking at least one virtual object within the group, and / or moving at least one virtual object within the group relative to virtual objects within a second virtual region. In addition to or instead of this, highlighting the group of virtual objects within the first virtual region based on the initial motion input can include providing a visual representation of the group of virtual objects within the first virtual region based on the initial motion input.

[0211] In some embodiments, highlighting can include highlighting each of the virtual objects within a group of virtual objects in a first virtual region and / or highlighting any portion of the first virtual region that includes the group of virtual objects. The group of virtual objects may be equally highlighted or may be highlighted to various degrees depending on the degree of specificity with which the user can indicate interest in any given one or more of the virtual objects within the group of virtual objects. In some embodiments, highlighting a group of virtual objects includes causing the wearable extended reality device to highlight the group of virtual objects.

[0212] In other embodiments, highlighting a group of virtual objects in a first virtual region based on an initial motion input can include displaying or adding controls that enable interaction with at least one of the virtual objects within the group of virtual objects in the first virtual region. For example, the displayed controls can include buttons on virtual objects such as virtual widgets that can activate a script for a program or application or otherwise link to a program or application when selected by the user. In other embodiments, highlighting can be associated with any perceivable physical indicator that can help to highlight or otherwise distinguish the group of objects in the first virtual region, such as by providing physical feedback (e.g., haptic feedback) to the user in response to a motion input.

[0213] Referring to step 2014 of FIG. 20, when executed by at least one processor, the instructions included in the non-transitory computer-readable medium can cause the at least one processor to highlight a group of virtual objects within a first virtual region based on an initial motion input. As an example, FIG. 22 shows a non-limiting example of an instance in which a group of virtual objects within a first virtual region is highlighted based on a user's initial motion input, which is consistent with some embodiments of the present disclosure.

[0214] Referring to FIG. 22, when user 2220 shows interest in at least one virtual object of a group of virtual objects included in a first subset of a plurality of dispersed virtual objects 2212 within a first virtual region 2214, the group of virtual objects 2212 in the direction of the initial motion input may be highlighted in a visually perceivable manner by user 2220. When it is determined that user 2220 may potentially desire to kinematically select one of the virtual objects within the group of virtual objects 2212, the group of virtual objects 2212 is highlighted in response to the initial motion input of user 2220 that tends towards the first virtual region 2214. In the example shown in FIG. 22, the visual appearance of each virtual object of the group of virtual objects 2212 in the first virtual region 2214 is shown to be highlighted or otherwise distinguished from a second subset of the dispersed virtual objects 2212 in a second virtual region 2218 due to a shadow added around the frame of the virtual object 2216.

[0215] In addition to or instead of this, highlighting at least one virtual object of a group of virtual objects 2212 in the first virtual region 2214 can include displaying or adding controls associated with at least one virtual object of the group of virtual objects 2212 with which the user 2220 can interact. Referring to FIG. 23, user 2310 shows a controller 2318 included in a highlighted virtual object 2312 that enables user interaction with at least one virtual object 2312 within a group of virtual objects in the first virtual region 2314. Highlighting at least one virtual object of a group of virtual objects can provide the user with substantially seamless feedback within the extended reality environment based on the perceived motion input and can help the user intuitively discover the results of the perceived motion input.

[0216] Some of the disclosed embodiments can include receiving a refined motion input that tends to be directed towards a particular virtual object from among a group of highlighted virtual objects. In a general sense, receiving a refined motion input can be similar to receiving the motion input described above. A refined motion input can be associated with a more precise gesture, gaze, or other movement of the user. Such greater precision can enable a corresponding selection refinement. For example, in response to the perception of a group of highlighted virtual objects, the user can increase the accuracy or precision of the motion input relative to an initial motion input towards a single virtual object of interest within the group of highlighted virtual objects. A refined motion input can be recognized by reception, detection, or other means such that a confidence level associated with the refined motion input is determined by at least one processor to be greater than a determined confidence level associated with the initial motion input.

[0217] In some embodiments, receiving refined motion input can result in canceling the highlighting of at least some of the virtual objects within a group of virtual objects. For example, a user attempting to kinematically select a virtual object within an extended reality environment by an initial motion input may be provided with intermediate feedback, such as a highlighted group of virtual objects, based on the initial motion input when the intent of the interaction is not very specific. As the user reaches towards a more specific region of the first virtual region towards a particular virtual object that the user may intend to select, subsequent feedback can be provided to the user by de - highlighting the virtual objects surrounding the particular virtual object until only the single virtual object intended to be selected remains highlighted based on the user's refined movement. Alternatively, all of the virtual objects within a group of virtual objects may become un - highlighted in response to refined motion input. In addition to or instead of this, canceling the highlighting of at least some of the virtual objects within a group of virtual objects can include ceasing to provide the visual representation of at least some of the virtual objects within the group of virtual objects.

[0218] In some embodiments, the initial motion input may be associated with a first type of motion input, and the refined motion input may be associated with a second type of motion input that is different from the first type of motion input. For example, the first motion input may be a gaze, and the second motion input may be a gesture such as pointing. Alternatively, the first motion input may be a head rotating towards a specific region of the display, and the second motion input may be a focused gaze on a specific region or item. In a general sense, the first type of motion input and the second type of motion input can include any combination of interactions based on gestures such as hand gestures, body gestures (e.g., movements of the user's head or body), and eye movements.

[0219] Referring to step 2016 of FIG. 20, when instructions included in a non-transitory computer-readable medium are executed by at least one processor, the at least one processor can be made to receive a refined motion input that tends to be directed towards a particular virtual object from among a group of highlighted virtual objects. In one example, the refined motion input that tends to be directed towards a particular virtual object from among a group of highlighted virtual objects may be received after the highlighting of the group of virtual objects within the first virtual region by step 2014 and / or after the reception of the initial motion input by step 2012. As an example, FIG. 23 shows a non-limiting example of a user's refined motion input, consistent with some embodiments of the present disclosure, compared to the initial motion input shown in FIG. 22, where the refined motion input tends to be directed towards a particular virtual object from among a group of highlighted virtual objects. It should be understood that both the initial motion input and the refined motion input shown in FIGS. 22 and 23 represent the user's hand gestures, but each of the initial motion input and the refined motion input can include any combination of gesture-based interactions including hand gestures and / or eye movements.

[0220] Referring to FIG. 23, user 2310 is shown to display a particular interest in a particular virtual object 2312 from among a group of previously highlighted virtual objects within a first virtual region 2314 via refined motion input. At the point in time shown in FIG. 23, the refined motion input of user 2310 has a more distinct tendency to be directed towards a particular virtual object 2312 than towards other virtual objects 2316 within the first virtual region 2314. Thus, the confidence level associated with the refined motion input shown in FIG. 23 is higher than the confidence level associated with the initial motion input shown in FIG. 22, and the refined motion input more specifically indicates that user 2310 may wish to make a kinematic selection of a particular virtual object 2312. In response to the refined motion input, only the particular virtual object 2316 determined by at least one processor to be intended for selection beyond a predetermined confidence level remains highlighted based on the user's refined movement, and by de-emphasizing other virtual objects 2312 within the first virtual region 2314 surrounding the particular virtual object 2312, subsequent feedback can be provided to user 2310.

[0221] Some of the disclosed embodiments can include triggering functions associated with a particular virtual object based on a refined motion input. The term "triggering a function" associated with a particular virtual object can include any occurrence resulting from the identification of the particular virtual object. Such occurrences can include changing the appearance of the virtual object or initiating the operation of code associated with the virtual object. Thus, the selection of a virtual object can result in the presentation of a virtual menu, the presentation of additional content, or the triggering of an application. When a function associated with a particular virtual object is triggered, the virtual content and / or functions of the particular virtual object can change from a non-engaged state or otherwise highlighted state based on the initial motion input to an engaged state in response to the selection or engagement of the virtual object via the refined motion input. In some embodiments, the function associated with a particular virtual object includes causing the wearable extended reality device to initiate the display of particular virtual content.

[0222] The function associated with a particular virtual object can additionally include causing an action associated with the particular virtual object, such as being further highlighted. In addition to or instead of this, the function associated with a particular virtual object can include enlarging the particular virtual object, revealing additional virtual objects for selection within the particular virtual object or otherwise associated with the particular virtual object, displaying or hiding a further viewing area that includes virtual content associated with the particular virtual object, and / or initiating any virtual content or process that is functionally associated with the particular virtual object based on the refined motion input. In some embodiments, where the particular virtual object is an icon, the triggered function can include causing an action associated with the particular virtual object, including activating a script associated with the icon.

[0223] Referring to step 2018 of FIG. 20, when instructions included in a non-transitory computer-readable medium are executed by at least one processor, the at least one processor can be caused to trigger functions related to a specific virtual object based on refined motion input. As an example, FIG. 23 shows a non-limiting example of refined motion input provided by a user that triggers functions resulting from being associated with a specific virtual object within an extended reality environment, consistent with some embodiments of the present disclosure.

[0224] Referring to FIG. 23, in response to the refined motion input of user 2310, at least one function related to a specific virtual object 2312 intended to be selected can be triggered by at least one processor that causes an action related to the specific virtual object 2312. For example, one function related to the specific virtual object 2312 triggered in this example is to enlarge the specific virtual object in response to the refined motion input of user 2310. Another function related to the specific virtual object 2312 triggered in this example is to reveal a further virtual controller 2318 for making selections within the specific virtual object 2312. As shown in FIG. 24, when an action associated with a specific virtual object 2412 is triggered, such as revealing a further virtual object 2414 for selection within the specific virtual object 2412, user 2410 can interact with the further virtual object 2414.

[0225] Other embodiments can include tracking the elapsed time after receiving an initial motion input. The tracking of time can be used to determine whether a function should be triggered or a highlighting should be canceled based on the elapsed time from when the initial motion input was received until a refined motion input is received. For example, when a refined motion input is received during a predetermined period following the initial motion input, a function associated with a particular virtual object can be triggered. If a refined motion input is not received within a predetermined period, the highlighting of a group of virtual objects within a first virtual region may be canceled.

[0226] In some embodiments, the term "predetermined period following the initial motion input" can be related to the measurement of time after the initial motion input is received. For example, the predetermined period following the initial motion input can be related to the period that begins after the initial motion input is detected or otherwise recognized and / or converted into a virtual interaction with an extended reality environment and a group of virtual objects within a first virtual region is highlighted. In some cases, the tracking of the elapsed time can start when the user starts the initial motion input from a stop position or from another preliminary motion input. In other examples, the tracking of the elapsed time can start when the user ends the initial motion input or otherwise decelerates the initial motion input below a threshold level. The time at which the reception of the refined motion input is initiated can be based on a measurement of the distance of the user's initial motion input from a particular one of the group of virtual objects and / or the virtual region containing the virtual objects, and / or a measurement of the speed and / or acceleration associated with the user's initial motion input, or a change thereof.

[0227] Other embodiments can include receiving a preliminary movement input that tends to be directed towards a second virtual region before receiving an initial movement input, and highlighting a group of virtual objects within the second virtual region based on the preliminary movement input. In a general sense, the preliminary movement input may be similar to the movement input described above. The term preliminary movement input can relate to any movement input that allows a user to begin to show interest or otherwise pay attention to a region that includes at least one virtual object, before showing interest in another region that includes at least one different virtual object. In some examples, there may be at least one virtual object within a group of virtual objects from the second virtual region that is included within a group of virtual objects within the first virtual region upon receipt of the initial movement input that follows the preliminary movement input.

[0228] Upon receiving the initial movement input, the highlighting of the group of virtual objects within the second virtual region can be cancelled, and a group of virtual objects within the first virtual region can be highlighted. In one embodiment, cancelling the highlighting can be related to restoring or changing the group of virtual objects highlighted in response to the preliminary movement input to a non-highlighted state in response to the initial movement input. For example, the group of virtual objects within the second virtual region highlighted in response to the preliminary movement input may be restored to a non-highlighted state, and a new group of virtual objects within the first virtual region may transition from a non-highlighted state to a highlighted state in response to the initial movement input. In one example, cancelling the highlighting of the group of virtual objects within the second virtual region can occur when the group of virtual objects within the first virtual region is highlighted. In another example, cancelling the highlighting of the group of virtual objects within the second virtual region can begin to fade or otherwise transition to a non-highlighted state as the group of virtual objects within the first virtual region is highlighted or transitions to a highlighted state.

[0229] FIG. 25 is a flowchart showing an exemplary process 2500 for highlighting a group of virtual objects highlighted in response to a preparatory movement input and canceling the highlighting of the group of virtual objects in response to an initial movement input, consistent with some embodiments of the present disclosure. Referring to FIG. 25, at step 2510, instructions included in a non-transitory computer-readable medium when executed by at least one processor can cause the at least one processor to receive a preparatory movement input that tends to direct toward a second virtual region prior to receiving the initial movement input. At step 2512, at least one processor can be caused to highlight a group of virtual objects within the second virtual region based on the preparatory movement input. At step 2514, when the at least one processor receives the initial movement input, it can cancel the highlighting of the group of virtual objects within the second virtual region and highlight a group of virtual objects within the first virtual region.

[0230] Other embodiments can include determining a confidence level associated with an initial motion input and selecting visual characteristics of the highlighting based on the confidence level. The term confidence level can be related to a threshold confidence level that some of a plurality of distributed virtual objects are intended to be selected by a user based on a motion input such as an initial motion input. When the number of the plurality of distributed virtual objects is determined to fall within a threshold confidence level for the likelihood of being selected based on the initial motion input by at least one processor, a group of virtual objects can be included in the group of virtual objects to be highlighted based on the initial motion input. In one non-limiting example of determining a confidence level associated with the initial motion input, if it is determined that the user shows interest in or otherwise pays attention to the number of virtual objects having an evaluated confidence level, e.g., a confidence level of 50% or more, the interaction can cause a group of virtual objects to be highlighted. If it is determined that the user pays attention to some virtual objects having an evaluated confidence level, e.g., less than 50% confidence level, the interaction may not cause the virtual objects to be highlighted.

[0231] In some embodiments, the non-transitory computer-readable medium can cause at least one processor to analyze past virtual object selections to determine a user profile and, based on the user profile, determine a confidence level associated with a motion input, such as an initial motion input. Analyzing past virtual object selections can include any artificial intelligence, machine learning, deep learning, and / or neural network processing techniques. Such techniques can enable machine learning by consuming large amounts of unstructured data such as sensed motion inputs, pre-selection recognition processes, and / or video, as well as the user's preferences analyzed over a period of time. Any suitable computing system or group of computing systems can be used to perform an analysis of past and / or current virtual object selections, or any other interactions within an extended reality environment, using artificial intelligence. Further, data corresponding to the analyzed past and / or current virtual object selections, user profile, and confidence levels associated with the user's past and / or current motion inputs can be stored in and / or accessed from any suitable database and / or data structure including one or more memory devices. For example, at least one database can be configured and operable to record and / or reference the user's past interactions with the extended reality environment to identify confidence levels associated with new motion inputs.

[0232] Other embodiments can include determining a confidence level associated with refined motion input and withholding triggering a function associated with a particular virtual object when the confidence level is below a threshold. In a general sense, determining a confidence level associated with refined motion input can be performed in a similar manner as determining a confidence level associated with the initial motion input described above. If a particular virtual object is determined to fall within a threshold confidence level for the likelihood of being triggered based on refined motion input by at least one processor, triggering of the function associated with the particular virtual object may not be initiated, or may be aborted if already initiated. It should be understood that determining a confidence level associated with refined motion input, including gestures and eye movements, can include determining a confidence level associated with fixation detection and a confidence level associated with gesture detection. A determination of whether to prioritize detected fixation over detected gestures of the refined motion input, or vice versa, may be notified in a similar manner as the initial motion input described above, or may be weighted otherwise.

[0233] The disabling of a function associated with a particular virtual object in response to a refined motion input can be done in a manner similar to the cancellation of the highlighting of at least some of the virtual objects within a group of virtual objects in response to the motion input. The disabling of a function associated with a particular virtual object can be related to restoring the particular virtual object to a non - participating state or returning it to a highlighted state in response to a low confidence level associated with the refined motion input. In one non - limiting example of determining a confidence level associated with a refined motion input, if the confidence level evaluated for the user's expressed interest in a particular virtual object is determined to be less than, for example, a 90% confidence level, the interaction is found to include a non - specific intention of selection and the trigger for the function associated with the particular virtual object may not be initiated or the initiation may be stopped. If the confidence level evaluated for the user's expressed interest in a particular virtual object is determined to be, for example, 90% confidence level or higher, the interaction is found to include a specific intention of selection and intermediate actions associated with the particular virtual object can be generated, such as expanding the virtual object or presenting additional virtual objects for selection within the particular virtual object prior to the selection of the virtual object. If the user is determined to have shown interest in a particular virtual object within a nearly 100% confidence level, the interaction can trigger some action associated with the particular virtual object.

[0234] In some embodiments, the non-transitory computer-readable medium can cause at least one processor to access a user action history indicating at least one action performed prior to an initial motion input and determine a confidence level associated with a refined motion input based on the user action history. In a general sense, the action history can be stored, accessed, and / or analyzed in a manner similar to the previously described past virtual object selections, user profiles, and confidence levels associated with the user's previous motion inputs. Further, the user action history may be related to an action history including data from an ongoing session and / or a previous session, and may be related to interactions based on any type of gesture. In some examples, the user action history may be related to data including past and / or current patterns of interaction with an extended reality environment. The action history regarding an ongoing session and / or a previous session can be used by at least one processor to determine a confidence level associated with a newly initiated refined motion input.

[0235] In some embodiments, the threshold confidence level may be selected based on a particular virtual object. In one example, the previous selection history of the user associated with a particular virtual object, the user profile, and / or the user behavior history may make the threshold confidence level for the selection of the virtual object higher or lower than the threshold confidence level for the selection of another virtual object, such as a virtual object that is not used very frequently. In another example, if the selection of a particular virtual object causes user confusion (e.g., pausing or ending a data stream) or causes user data loss (e.g., closing an application or turning off the device power), the threshold can be selected to require a higher confidence level before triggering a function associated with the particular virtual object. In addition to or instead of this, the threshold confidence level for a particular virtual object can be increased and / or decreased based on an increased and / or decreased determination of the need for the virtual object based on at least one external factor, such as a notification and / or warning associated with the virtual object, or a program running in parallel that attempts to send and receive information with a program associated with the virtual object.

[0236] Some embodiments can include determining a first confidence level associated with an initial motion input, ceasing to highlight a group of virtual objects within a first virtual region if the first confidence level is lower than a first threshold, determining a second confidence level associated with a refined motion input, and ceasing to trigger a function associated with a particular virtual object if the second confidence level is lower than a second threshold. In a general sense, the first confidence level associated with the initial motion input and the second confidence level associated with the refined motion input may be similar to the confidence level associated with the initial motion input and the confidence level associated with the refined motion input, respectively, as described above. Further, in a general sense, ceasing to highlight a group of virtual objects and ceasing to trigger a function can be performed in a manner similar to ceasing to trigger a function as described above. In some embodiments, the second threshold for ceasing to trigger a function associated with a particular virtual object may be greater than the first threshold for hi...

Claims

1. A non-transitory computer-readable medium including instructions that cause at least one processor to perform operations for moving virtual content between virtual planes in a three-dimensional space when executed by the at least one processor, the operations comprising: Using a wearable extended reality device to virtually display a plurality of virtual objects on a plurality of virtual planes, the plurality of virtual planes including a first virtual plane and a second virtual plane; Outputting a first display signal that reflects a virtual object at a first position on the first virtual plane for presentation via the wearable extended reality device; Receiving an in-plane input signal for moving the virtual object to a second position on the first virtual plane; In response to the in-plane input signal, virtually displaying an in-plane movement of the virtual object from the first position to the second position on the wearable extended reality device; Receiving an inter-plane input signal for moving the virtual object to a third position on the second virtual plane while the virtual object is at the second position; In response to the inter-plane input signal, virtually displaying an inter-plane movement of the virtual object from the second position to the third position on the wearable extended reality device; A non-transitory computer-readable medium comprising the above.

2. The non-transitory computer-readable medium according to claim 1, wherein the first virtual plane is associated with a first distance from the wearable extended reality device, the second virtual plane is associated with a second distance from the wearable extended reality device, and the first distance associated with the first virtual plane is greater than the second distance associated with the second virtual plane.

3. A first curvature of the first virtual plane is substantially the same as a second curvature of the second virtual plane, and the operations further include changing an extended reality display of the virtual object to reflect a difference between the first distance and the second distance in response to the inter-plane input signal. The non-transitory computer-readable medium according to claim 2.

4. The first curvature of the first virtual plane is different from the second curvature of the second virtual plane, and the operation further includes changing the display of the virtual object to reflect a difference between the first distance and the second distance and a difference between the first curvature and the second curvature in response to the in-plane input signal. The non-transitory computer-readable medium according to claim 2.

5. The first virtual plane is associated with a first distance from a physical input device, the second virtual plane is associated with a second distance from the physical input device, and the first distance associated with the first virtual plane is greater than the second distance associated with the second virtual plane. The non-transitory computer-readable medium according to claim 1.

6. The first virtual plane is associated with a first distance from an edge of a physical surface, the second virtual plane is associated with a second distance from the edge of the physical surface, and the first distance associated with the first virtual plane is greater than the second distance associated with the second virtual plane. The non-transitory computer-readable medium according to claim 1.

7. The first virtual plane is curved, and the in-plane movement of the virtual object from the first position to the second position includes three-dimensional movement. The non-transitory computer-readable medium according to claim 1.

8. The first virtual plane and the second virtual plane are convex surfaces. The non-transitory computer-readable medium according to claim 1.

9. The in-plane movement includes moving the virtual object along two orthogonal axes, the movement along the first axis includes changing the dimensions of the virtual object, and the movement along the second axis excludes changing the dimensions of the virtual object. The non-transitory computer-readable medium according to claim 1.

10. The operation further includes changing the dimensions of the virtual object based on the curvature of the first virtual plane. The non-transitory computer-readable medium according to claim 9.

11. Both the in-plane input signal and the inter-plane input signal are received from a touch sensor, and the operation further includes identifying a first multi-finger interaction as a command for causing in-plane movement, and identifying a second multi-finger interaction different from the first multi-finger interaction as a command for causing inter-plane movement. The non-transitory computer-readable medium according to claim 1.

12. The operation further includes receiving a preliminary input signal for selecting a desired curvature in the first virtual plane, and the in-plane movement is determined based on the selected curvature of the first virtual plane. The non-transitory computer-readable medium according to claim 1.

13. Selecting the curvature in the first virtual plane affects the curvature of the second virtual plane. The non-transitory computer-readable medium according to claim 12.

14. The operation is virtually displaying a plurality of virtual objects on a plurality of virtual planes using a wearable extended reality device, the plurality of virtual planes including a first virtual plane and a second virtual plane, causing the wearable extended reality device to display in-plane movement of the virtual object and the other virtual object in response to the in-plane input signal, causing the wearable extended reality device to display inter-plane movement of the virtual object and the other virtual object in response to the inter-plane input signal, The non-transitory computer-readable medium according to claim 1 further includes.

15. The operation is receiving an input signal indicating that the virtual object is docked to a physical object, identifying a relative movement between the physical object and the wearable extended reality device, determining whether to move the virtual object to a different virtual plane based on the identified movement, The non-transitory computer-readable medium according to claim 1 further includes.

16. The operation further includes determining a relative movement caused by movement of the physical object, and causing the inter-plane movement of the virtual object to the different virtual plane in response to the determined relative movement. The non-transitory computer-readable medium according to claim 15.

17. The operation further includes determining the relative movement caused by the movement of the wearable extended reality device, and preventing the movement of the virtual object between the different virtual planes according to the determined relative movement. The non-transitory computer-readable medium according to claim 15.

18. The operation includes receiving an input signal indicating that the first virtual plane and the second virtual plane are docked to a physical object, identifying the movement of the physical object to a new position while the virtual object is in the second position and before receiving the input signal between the planes, determining new positions in the first virtual plane and the second virtual plane based on the new position of the physical object, causing the wearable extended reality device to display the movement of the virtual object from the second position on the first virtual plane at the original position of the first virtual plane to the second position on the first virtual plane at the new position of the first virtual plane according to the identified movement of the physical object, further including the display of the movement of the virtual object between the planes from the second position to the third position is the display of the movement of the virtual object from the second position on the first virtual plane at the new position of the first virtual plane to the third position on the second virtual plane at the new position of the second virtual plane, The non-transitory computer-readable medium according to claim 1.

19. A system for moving virtual content between planes in a three-dimensional space, the system comprising at least one processor, the processor using a wearable extended reality device to virtually present a plurality of virtual objects on a plurality of virtual planes, the plurality of virtual planes including a first virtual plane and a second virtual plane, outputting a first display signal reflecting a virtual object at a first position on the first virtual plane for presentation via the wearable extended reality device, receiving an in-plane input signal for moving the virtual object to a second position on the first virtual plane, ​ In response to the in-plane input signal, causing the wearable extended reality device to display an in-plane movement of the virtual object from the first position to the second position, while the virtual object is at the second position, receiving an inter-plane input signal for moving the virtual object to a third position on the second virtual plane, and in response to the inter-plane input signal, causing the wearable extended reality device to display an inter-plane movement of the virtual object from the second position to the third position, A system configured as described above.

20. A method for moving virtual content between planes in a three-dimensional space, the method comprising: using a wearable extended reality device to virtually present a plurality of virtual objects on a plurality of virtual planes, the plurality of virtual planes including a first virtual plane and a second virtual plane; outputting a first display signal reflecting a virtual object at a first position on the first virtual plane for presentation via the wearable extended reality device; receiving an in-plane input signal for moving the virtual object to a second position on the first virtual plane; in response to the in-plane input signal, causing the wearable extended reality device to display an in-plane movement of the virtual object from the first position to the second position; while the virtual object is at the second position, receiving an inter-plane input signal for moving the virtual object to a third position on the second virtual plane; and in response to the inter-plane input signal, causing the wearable extended reality device to display an inter-plane movement of the virtual object from the second position to the third position. A method comprising the steps of:

Citation Information

Patent Citations

  • Display device, display method and program

    JP2014044334A

  • Head mounted display device, method for controlling the same and computer program

    JP2016081209A

  • Controlling the brightness of the displayed image

    JP2016519322A

  • Information processing device, information processing method and program

    JP2019114078A

  • Augmented sense of reality system and color compensation method thereof

    JP2020017252A