Content providing system and method
Through parallel connections between mobile devices and edge and fog resource devices and optimized wireless connections, the problems of high latency and high computing load in augmented reality systems are solved, and low latency and efficient content provision are achieved, improving user experience.
Patent Information
- Application Number
- CN202080048293.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Priority Date
- 2019-05-01
- Filing Date
- 2020-05-01
- Publication Date
- 2025-08-05
- Estimated Expiration
- 2040-05-01
AI Technical Summary
In the prior art, augmented reality systems have problems such as high latency, large computing load, and unstable information acquisition in terms of content provision, especially in connections and data processing between mobile devices and multiple resource devices.
The parallel connection between mobile devices and edge resource devices and fog resource devices is adopted, through cellular towers and Wi-Fi connections, phased array antennas and radar hologram-type transmission connectors are used, combined with the arbitrator function and rendering engine, wireless connections and data processing are optimized to achieve low-latency content delivery.
It realizes low latency and efficient content delivery, improves user experience, and enhances robust connection and data processing capabilities between mobile devices and multiple resource devices.
Smart Images

Figure CN114127837B_ABST
Abstract
Description
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application claims priority to U.S. Provisional Patent Application No. 62 / 841,806, filed May 1, 2019, the entire contents of which are incorporated herein by reference in their entirety. Technical Field
[0003] The present invention relates to connected mobile computing systems, methods, and configurations, and more particularly to content provision systems, mobile computing systems, methods, and configurations featuring at least one wearable component that can be used for virtual and / or augmented reality operations. Background Art
[0004] Content delivery systems with one or more augmented reality systems are becoming popular for viewing the real world overlaid with digital content. The content delivery system may, for example, include a mobile device such as a head-mounted viewing assembly. The content delivery system may also include a resource device having a resource device dataset including content and a storage medium. The resource device transmits the content to the mobile device. The mobile device has a connected output device capable of providing output that can be sensed by the user. Summary of the Invention
[0005] The present invention provides a content providing system, comprising: a mobile device, which may have a mobile device processor; a mobile device communication interface, which is connected to the mobile device processor and a first resource device communication interface and receives first content sent by a first resource device transmitter under the control of the mobile device processor; and a mobile device output device, which is connected to the mobile device processor and can provide output that can be sensed by a user under the control of the mobile device processor.
[0006] The content providing system may further comprise a first resource device which may have a first resource device processor, a first resource device storage medium and a first resource device data set comprising first content on the first resource device storage medium, a first resource device communication interface which forms part of the first resource device and is connected to the first resource device processor and is under the control of the first resource device processor.
[0007] The content providing system may include a first resource device at a first location, wherein the mobile device communication interface establishes a first connection with the first resource device, and wherein the content is first content specific to a first geographic parameter of the first connection.
[0008] The content providing system may further comprise: a second resource device which may have a second resource device processor, a second resource device storage medium, a second resource device data set comprising second content on the second resource device storage medium, and a second resource device communication interface forming part of the second resource device and connected to the second resource device processor and under control of the second resource device processor, wherein the second resource device is at a second location, wherein the mobile device communication interface creates a second connection with the second resource device, and wherein the content is second content specific to a second geographic parameter of the second connection.
[0009] The content providing system may include: a mobile device including a head-mounted viewing assembly coupleable to a user's head, and first and second content that provides the user with at least one of additional content, augmented content, and information about a particular view of the world as seen by the user.
[0010] The content delivery system may also include a location island for the user to enter, wherein specific features have been preconfigured to be located and interpreted by the mobile device to determine geographic parameters relative to the world around the user.
[0011] The content providing system may include a particular feature that is a visually detectable feature.
[0012] The content providing system may include a specific feature that is a wireless connection related feature.
[0013] The content delivery system may also include a plurality of sensors coupled to the head-mounted viewing assembly, the plurality of sensors being used by the mobile device to determine geographic parameters relative to the world surrounding the user.
[0014] The content providing system may further include a user interface configured to allow a user to at least one of: ingest, utilize, view, and bypass certain information of the first or second content.
[0015] The content providing system may include a connection that is a wireless connection.
[0016] The content providing system may include: a first resource device at a first location, wherein the mobile device has a sensor to detect a first feature at the first location, and the first feature is used to determine a first geographic parameter associated with the first feature, and wherein the content is first content specific to the first geographic parameter.
[0017] The content providing system may include: a second resource device at a second location, wherein the mobile device has a sensor to detect a second feature at the second location, and the second feature is used to determine a second geographic parameter associated with the second feature, and wherein the first content is updated with second content specific to the second geographic parameter.
[0018] The content providing system may include: the mobile device including a head-mounted viewing assembly coupleable to a user's head, and the first and second content providing the user with at least one of additional content, augmented content, and information about a particular view of the world as seen by the user.
[0019] The content providing system may further include a spatial computing layer between the mobile device and a resource layer having a plurality of data sources, the spatial computing layer being programmed to receive data resources, integrate the data resources to determine an integration profile, and determine the first content based on the integration profile.
[0020] The content providing system may include: the spatial computing layer may include a spatial computing resource device, which may have a spatial computing resource device processor, a spatial computing resource device storage medium, and a spatial computing resource device data set on the spatial computing resource device storage medium, and may be executed by the processor to receive data resources, integrate the data resources to determine an integration profile, and determine the first content based on the integration profile.
[0021] The content providing system may further include an abstraction and arbitration layer that is inserted between the mobile device and the resource layer and is programmed to make workload decisions and distribute tasks based on the workload decisions.
[0022] The content providing system may also include a camera device that captures images of the physical world surrounding the mobile device, wherein the images are used to make workload decisions.
[0023] The content providing system may further comprise a camera device that captures images of the physical world surrounding the mobile device, wherein the images form one of the data resources.
[0024] The content providing system may include: the first resource device is an edge resource device, wherein the mobile device communication interface includes one or more mobile device receivers connected to the mobile device processor and connected to the second resource device communication interface in parallel with the connection of the first resource device to receive the second content.
[0025] The content providing system may include the second resource device being a fog resource device having a second delay that is slower than the first delay.
[0026] The content providing system may include: a mobile device communication interface including one or more mobile device receivers connected to the mobile device processor and connected to the mobile device processor and a third resource device communication interface in parallel with the connection to the second resource device to receive third content sent by a third resource device transmitter, wherein the third resource device is a cloud resource device having a third delay that is slower than the second delay.
[0027] The content providing system may include: connecting to the edge resource device via a cellular tower, and connecting to the fog resource device via a Wi-Fi connection device.
[0028] The content providing system may include: a cellular tower connected to a fog resource device.
[0029] The content providing system may include: a Wi-Fi connection device connected to a fog resource device.
[0030] The content providing system may also include at least one camera to capture at least first and second images, wherein the mobile device processor sends the first image to the edge resource device for faster processing and sends the second image to the fog resource device for slower processing.
[0031] The content providing system may include: at least one camera being a room camera that captures a first image of a user.
[0032] The content providing system may also include: a sensor that provides sensor input to the processor; a posture estimator that is executable by the processor to calculate the posture of the mobile device based on the sensor input, including at least one of the position and orientation of the mobile device; a steerable wireless connector that creates a steerable wireless connection between the mobile device and the edge resource device; and a steering system that is connected to the posture estimator and has an output that provides input to the steerable wireless connector to manipulate the steerable wireless connection to at least improve the connection.
[0033] The content providing system may include: the steerable wireless connector is a phased array antenna.
[0034] The content providing system may include: the steerable wireless connector being a radar hologram type transmission connector.
[0035] The content providing system may also include an arbitrator function that can be executed by the processor to determine how much edge and fog resources are available through the edge and fog resource devices, respectively, send processing tasks to the edge and fog resources based on the determination of available resources, and receive results returned from the edge and fog resources.
[0036] The content providing system may include: an arbitrator function executable by a processor to combine results from edge and fog resources.
[0037] The content providing system may also include a runtime controller function that can be executed by the processor to determine whether the process is a runtime process, if a determination is made that the task is a runtime process, then immediately execute the task without using the arbitrator function to make a determination, and if a determination is made that the task is not a runtime process, then use the arbitrator function to make a determination.
[0038] The content providing system may also include multiple edge resource devices, data is exchanged between the multiple edge resource devices and the fog resource device, and the data includes points in space captured by different sensors and sent to the edge resource devices; and a super point calculation function, which can be executed by the processor to determine a super point, which is a point selected from two or more points where data from the edge resource devices overlap.
[0039] The content providing system may further include a plurality of mobile devices, wherein each super point is used in each mobile device for location, orientation, or pose estimation of the corresponding mobile device.
[0040] The content providing system may also include a contextual trigger function executable at the processor to generate a contextual trigger for a set of superpoints and store the contextual trigger on a computer-readable medium.
[0041] The content providing system may further include a rendering engine executable by the mobile device processor, wherein the contextual trigger serves as a handle for rendering the object based on the first content.
[0042] The content providing system may also include a rendering function, executable by the mobile device processor, to connect the mobile device to a plurality of resource devices, wherein each resource device receives a respective rendering request, to receive a rendering from each of the remote devices based on the respective rendering request, to compare the renderings to determine a preferred rendering, and to select, using the mobile device processor, the preferred rendering as the first content sent by the first resource device transmitter.
[0043] The content providing system may include rendering a system having a polynomial prediction for rendering a frame into the future at which a mobile device is predicted to be positioned or viewed.
[0044] The present invention also provides a method for providing content, comprising: connecting a mobile device communication interface of a mobile device to a first resource device communication interface of a first resource device under the control of a mobile device processor of the mobile device, and receiving first content sent by a first resource device transmitter using the mobile device communication interface under the control of the mobile device processor.
[0045] The method may also include: under the control of the first resource device processor, storing a first resource device data set including the first content on a first resource device storage medium connected to the first resource device processor, and sending the first content using a first resource device communication interface connected to the first resource device processor and under the control of the first resource device processor.
[0046] The method may include a first resource device at a first location, wherein the mobile device communication interface creates a first connection with the first resource device, and wherein the content is first content specific to a first geographic parameter of the first connection.
[0047] The method may further include storing a second resource device data set including second content on a second resource device storage medium connected to the second resource device processor under control of the second resource device processor, and sending the second content using a second resource device communication interface connected to the second resource device processor and under control of the second resource device processor, wherein the second resource device is at a second location, wherein the mobile device communication interface creates a second connection with the second resource device, and wherein the content is second content specific to a second geographic parameter of the second connection.
[0048] The method may include the mobile device including a head-mounted viewing assembly coupleable to a user's head, and the first and second content providing the user with at least one of additional content, augmented content, and information regarding a particular view of the world as seen by the user.
[0049] The method may include the user entering a location island where certain features have been preconfigured to be located and interpreted by the mobile device to determine geographic parameters relative to the world around the user.
[0050] The method may include the particular feature being a visually detectable feature.
[0051] The method may include the specific characteristic being a wireless connection related characteristic.
[0052] The method may include connecting a plurality of sensors to the head mounted viewing assembly, the plurality of sensors being used by the mobile device to determine geographic parameters relative to the world around the user.
[0053] The method may further include receiving input from a user through a user interface to at least one of ingest, utilize, view, and bypass certain information of the first or second content.
[0054] The method may include the connection being a wireless connection.
[0055] The method may include a first resource device at a first location, wherein the mobile device has a sensor to detect a first feature at the first location, and the first feature is used to determine a first geographic parameter associated with the first feature, and wherein the content is first content specific to the first geographic parameter.
[0056] The method may include: the second resource device is at a second location, wherein the mobile device has a sensor to detect a second feature at the second location, and the second feature is used to determine a second geographic parameter associated with the second feature, and wherein the first content is updated with second content specific to the second geographic parameter.
[0057] The method may include the mobile device including a head-mounted viewing assembly coupleable to a user's head, and the first and second content providing the user with at least one of additional content, augmented content, and information regarding a particular view of the world as seen by the user.
[0058] The method may further include receiving data resources through a spatial computing layer between the mobile device and a resource layer having multiple data sources, integrating the data resources through the spatial computing layer to determine an integration profile, and determining first content based on the integration profile through the spatial computing layer.
[0059] The method may include: the spatial computing layer may include a spatial computing resource device, which has a spatial computing resource device processor; a spatial computing resource device storage medium, and a spatial computing resource device data set on the spatial computing resource device storage medium, and can be executed by the processor to receive data resources, integrate the data resources to determine an integration profile, and determine the first content based on the integration profile.
[0060] The method may further include making workload decisions using an abstraction and arbitration layer interposed between the mobile device and the resource layer, and distributing tasks based on the workload using the abstraction and arbitration layer.
[0061] The method may further include capturing, with the camera device, an image of the physical world surrounding the mobile device, wherein the image is used to make the workload decision.
[0062] The method may further comprise capturing, with the camera device, an image of the physical world surrounding the mobile device, wherein the image forms one of the data resources.
[0063] The method may include: the first resource device is an edge resource device, and also includes: under the control of the mobile device processor of the mobile device and in parallel with the connection of the first resource device, connecting the mobile device communication interface of the mobile device to the second resource device communication interface of the second resource device, and receiving the second content sent by the second resource device transmitter using the mobile device communication interface under the control of the mobile device processor.
[0064] The method may include the second resource device being a fog resource device having a second delay that is slower than the first delay.
[0065] The method may also include: connecting the mobile device communication interface of the mobile device to a third resource device communication interface of a third resource device under the control of the mobile device processor of the mobile device and in parallel with the connection of the second resource device, wherein the third resource device is a cloud resource device having a third delay that is slower than the second delay; and receiving third content sent by the third resource device transmitter using the mobile device communication interface under the control of the mobile device processor.
[0066] The method may include connecting to the edge resource device via a cellular tower, and connecting to the fog resource device via a Wi-Fi connection device.
[0067] The method may include connecting a cellular tower to a fog resource device.
[0068] The method may include connecting a Wi-Fi connected device to a fog resource device.
[0069] The method may further include capturing at least first and second images with the at least one camera, wherein the mobile device processor sends the first image to the edge resource device and sends the second image to the fog resource device.
[0070] The method may include at least one camera being a room camera that captures a first image of the user.
[0071] The method may also include: receiving sensor input through a processor; determining, using the processor, a posture of the mobile device based on the sensor input, including at least one of a position and an orientation of the mobile device; and manipulating, using the processor, a manipulatable wireless connector that creates a wireless connection between the mobile device and the edge resource device based on the posture to at least improve the connection.
[0072] The method may include the steerable wireless connector being a phased array antenna.
[0073] The method may include the steerable wireless connector being a radar hologram type transmission connector.
[0074] The method may also include: determining, using an arbitrator function executed by the processor, how much edge and fog resources are available through the edge and fog resource devices, respectively; sending processing tasks to the edge and fog resources using the arbitrator function based on the determination of available resources; and receiving results from the edge and fog resources using the arbitrator function.
[0075] The method may further include combining results from the edge and fog resources using an arbitrator function.
[0076] The method may further include determining, by the mobile device processor, whether the process is a runtime process, immediately executing the task without making a determination using an arbitrator function if a determination is made that the task is a runtime process, and making a determination using an arbitrator function if a determination is made that the task is not a runtime process.
[0077] The method may also include: exchanging data between multiple edge resource devices and fog resource devices, the data including points in space captured by different sensors and sent to the edge resource devices; and determining a super point, which is a point selected from two or more points where data from the edge resource devices overlap.
[0078] The method may further include using each super point in the plurality of mobile devices for location, orientation, or pose estimation of the corresponding mobile device.
[0079] The method may also include generating, with the processor, contextual triggers for a set of superpoints, and storing, with the processor, the contextual triggers on a computer-readable medium.
[0080] The method may further include using the contextual trigger as a handle for rendering the object based on the first content.
[0081] The method may also include: connecting the mobile device to multiple resource devices under the control of the mobile device processor; sending one or more rendering requests through the mobile device processor, wherein each resource device receives a corresponding rendering request; receiving a rendering from each of the remote devices based on the corresponding rendering request using the mobile device processor; comparing the renderings to determine a preferred rendering using the mobile device processor; and selecting the preferred rendering as the first content sent by the first resource device transmitter using the mobile device communication interface under the control of the mobile device processor.
[0082] The method may include rendering a system having a polynomial prediction for rendering a frame into the future that the mobile device is predicted to be positioned or viewed. BRIEF DESCRIPTION OF THE DRAWINGS
[0083] The present invention is further described by way of example with reference to the accompanying drawings, in which:
[0084] Figure 1 is a perspective view of an augmented reality system, a mobile computing system, a wearable computing system, and a content providing system according to an embodiment of the present invention;
[0085] Figures 2 to 5 is a mobile computing system (such as Figure 1 An overhead view of a moving scene as a user of a wearable computing system (XR) operates in the world;
[0086] Figures 6 to 8 It can be formed Figure 1 A block diagram of a wireless device that is part of a system;
[0087] Figure 9 is the view of the ArUco marker;
[0088] Figure 10 is a flow chart of a user wearing an augmented reality system navigating the world using a "location island";
[0089] Figure 11 Further details are shown Figure 1 A perspective of the system;
[0090] Figure 12 is a flow chart of a user wearing an augmented reality system navigating the world using connected resources for positioning;
[0091] Figure 13 is a flow chart of a user wearing an augmented reality system navigating the world using geometric shapes for positioning;
[0092] Figure 14 is a diagram illustrating the concept of "spatial computing";
[0093] Figure 15 is a diagram showing another way of representing the relationship between a user and the physical world using a spatial computing system;
[0094] Figure 16 It is a block diagram that depicts the hierarchical connection elements of a spatial computing environment.
[0095] Figure 17 A block diagram of the fundamental principles of how humans process and store information within and generally within a spatial computing architecture;
[0096] Figure 18 It is a block diagram of a human-centric spatial computing layer and information coupling with many different data sources;
[0097] Figure 19A and Figure 19B is a block diagram of a configuration in which a user wears a Figure 1 A system of systems as depicted, wherein “edge” computing and / or storage resources are typically located closer to users than “fog” computing and / or storage resources, which are closer than the typically more powerful and more distant “cloud” resources;
[0098] Figure 20A and Figure 20B is similar to Figure 19A and Figure 19B A block diagram of a user of a connected computing system where computation is distributed among edge, fog, and cloud computing resources based on latency and computational requirements is shown;
[0099] Figure 21 It is a block diagram of a human-centric spatial computing layer and information coupling with many different data sources;
[0100] Figure 22A and Figure 22B is a block diagram of a configuration in which a room with multiple cameras surrounding a user is utilized, and images from the cameras can be separated and directed to different computing resources for various reasons;
[0101] Figure 23A and Figure 23B A block diagram of various "Internet of Things" resources connected to a user's local computing resources via edge computing resources;
[0102] Figure 24A and Figure 24B is a block diagram of the types of wearable technology that can be connected to edge computing resources;
[0103] Figure 25A and Figure 25E is a block diagram of a configuration that allows a user to couple their local system to external resources for additional computing, storage, and / or power using a wired connection, such as via direct coupling to one or more antennas, computing workstations, laptops, mobile computing devices such as smartphones and / or tablets, edge computing resources, and power sources for charging their local computing system power supply (i.e., battery). Figure 25A ), interconnected auxiliary computing components ( Figure 25B ), wireless coupling with other computing resources ( Figure 25C ), coupled to the car ( Figure 25D ), and with additional computing and / or storage resources ( Figure 25E );
[0104] Figure 26A and Figure 26C is a perspective view featuring manipulable connections and focused or focused connections toward one or more specific mobile computing devices;
[0105] Figure 27 is a perspective diagram showing fog computing (which may also be referred to as "ambient computing") with different computing "rings" corresponding to latency levels relative to user devices;
[0106] Figures 28A to 28C is a block diagram of a system in which there may be a communication layer including various forms of connectivity (including fiber, coaxial cable, twisted pair cable satellite, various other wireless connectivity modalities) between the edge, fog, and cloud layers;
[0107] Figure 29A and Figure 29Bis a block diagram of various types of connectivity resources using hardware-based connectivity as well as various wireless connectivity paradigms;
[0108] Figure 30A and Figure 30B is a hardware coupled to a belt-pack computing component ( Figure 30A ) or a block diagram of a configuration of a head wearable assembly with tablet type interconnection ( FIG. 30 );
[0109] Figure 31 is a flow chart of an example paradigm for arbitration and allocation with respect to external resources, such as edge computing, fog computing, and cloud computing resources;
[0110] Figure 32 is a diagram illustrating the concept of a human-centered integrated spatial computing ("MagicVerse") general operation content provision system;
[0111] Figure 33 is a schematic diagram illustrating the connection of multiple overlapping edge computing nodes within a larger fog computing node, wherein seamless handoff or transfer is enabled between the edge computing devices;
[0112] Figure 34 is a block diagram of components for a general startup / bootstrap procedure that may have interconnected distributed resources;
[0113] Figure 35 is a schematic diagram of a massively multiplayer online (MMO) configuration, showing the generalization of computational requirements relative to the size of individual XR user nodes;
[0114] Figure 36 Yes Figure 35 Block diagram of various computing stacks for human-centric integrated spatial computing shown;
[0115] Figure 37 is a schematic diagram of a configuration for discovering, switching, and controlling elements within a direct radius of a mobile user;
[0116] Figure 38 is a block diagram of a superpoint-based simultaneous localization and mapping ("SLAM") system;
[0117] Figure 39 is a schematic diagram illustrating further details of the connection of multiple overlapping edge computing nodes within a larger fog computing node, wherein seamless handoff or transfer is enabled between the edge computing devices;
[0118] Figure 40is a schematic diagram of an edge node that may include sensors capable of creating a depth map of the world, for example, this may include a pair of stereo cameras, an RGB-D camera, a LiDAR device, and / or a structured light sensor, each of which may also include an IMU, a microphone array, and / or a speaker and / or function as a Wi-Fi or 5G antenna;
[0119] Figure 41 A schematic diagram of a "traversable world" system, where each online XR creates a part of the aggregate model of the environment;
[0120] Figure 42 It is a block diagram of the system that reproduces the digital twin of the world;
[0121] Figure 43 is a block diagram of a system for filtering spatial information;
[0122] Figure 44 is a schematic diagram showing a classic implementation of pose determination during the world reconstruction phase for manipulation, anchor points, or superpoints;
[0123] Figure 45 is a timeline for one embodiment of pose estimation using an anchor map;
[0124] Figure 46 is a timeline of a system that uses adaptive computing power edge / fog / cloud resources to render parallel frames as predictions and selects the frame closest to the actual value at the last moment;
[0125] Figure 47 is a simplified flow chart of the physical world, where we use the framework described above at different levels for different processes in spatial computing; and
[0126] Figures 48 to 66 is an illustration of various exemplary embodiments featuring various XR devices for use in various scenarios using converged spatial computing. DETAILED DESCRIPTION
[0127] Figure 1A content delivery system is shown that features an augmented reality system having a head-mounted viewing component (2), a handheld controller component (4), and an interconnected auxiliary computing or controller component (6) that can be configured to be worn on a user as a belt pack, etc. Each of these components can be connected to each other (10, 12, 14, 16, 17, 18) and to other connected resources (8), such as cloud computing or cloud storage resources, via wired or wireless communication configurations, such as those specified by IEEE 802.11, Bluetooth (RTM), and other connection standards and configurations. As described, for example, in U.S. patent application Ser. Nos. 14 / 555,585, 14 / 690,401, 14 / 331,218, 15 / 481,255, and 62 / 518,539, each of which is incorporated herein by reference in its entirety, aspects of such assemblies, such as the depicted two optical elements (20), and various embodiments of a visual assembly through which a user can see the world around them; the visual assembly can be generated by associated system components for an augmented reality experience. A need exists for systems and components that optimize compact and durable connectivity in wearable computing systems.
[0128] Figure 1 The content providing system is an example of a content providing system, which includes: a mobile device (head-mounted viewing component (2)) having a mobile device processor; a mobile device communication interface, which is connected to the mobile device processor and the first resource device communication interface and receives first content sent by the first resource device transmitter under the control of the mobile device processor; and a mobile device output device, which is connected to the mobile device processor and is capable of providing output that can be sensed by a user under the control of the mobile device processor. The content providing system also includes: a first resource device (connected resource (8)), which has a first resource device processor, a first resource device storage medium and a first resource device data set including the first content on the first resource device storage medium, the first resource device communication interface, which forms a part of the first resource device and is connected to the first resource device processor, and is under the control of the first resource device processor.
[0129] refer to Figure 2 , depicting a mobile computing system (such as reference Figure 1 A travel scene (160) in which a user of the described wearable computing system operates in a world. Figure 2A user's home (22) is shown featuring at least one wireless device (40) configured to connect to the user's wearable computing system. As the user navigates the world around him, here in one illustrative example, the user travels (30) from home (22, point A-80) to work (24, point B-82, point C-84, point D-86, point E-88), then from work (24) to a park (26) for a walk (28, point K-100, point L-102, point M-104) before returning (34, point N-106, point O-108) to rest at home (22) - along the way, his mobile computing system communicates with various wireless devices (40, 42, 44, 46, 48, 50, 52, 54, and 56). Figure 3 and Figure 4 Preferably, the mobile computing system is configured to utilize the various wireless devices and information exchanged therewith to provide a relatively low latency and robust connection experience to the user, generally subject to user preferences selectable by the user.
[0130] The mobile computing system can be configured to allow the user to select certain aspects of his computing experience for the day. For example, through a graphical user interface, voice control and / or gestures, the user can input to the mobile computing system that he will have a typical workday, a usual route, and stop for a walk in the park on the way home. The mobile computing system has an "artificial intelligence" aspect so that it uses integration with the user's electronic calendar to temporarily understand his schedule, which is subject to rapid confirmation. For example, when he leaves for work, the system can be configured to say or display: "Going to work, usual route and usual computing configuration", and this usual route can be obtained from previous GPS and / or mobile triangulation data by his mobile computing system. The "usual computing configuration" can be customized by the user and subject to regulations, for example, the system can be configured to present only certain non-obstructed visuals, no advertisements, and no shopping or other information not relevant to driving while the user is driving, and provide the audio version of a news program or a current favorite audio book while the user is driving to work. As the user navigates driving on his way to work, he can leave the connection with his home wireless device (40) and enter or maintain connections with other wireless devices (42, 44, 46, 48). Each of these wireless devices may be configured to provide information about the user's experience to the user's mobile computing system with relatively low latency (ie, by locally storing certain information that may be about the user at that location). Figure 6 and Figure 7 shows certain aspects of a wireless device that may be used as described herein, Figure 8 and Figure 9Embodiments featuring non-storage beacon and / or tag configurations may also be used to connect directly to locally relevant cloud-based information without the benefit of local storage.
[0131] For example, as a user travels from point A (80) to point B (82) to point C (84), local wireless devices (44) around point C (84) may be configured to transmit geometric information to the user's mobile system, which geometric information may be used on the user's mobile computing system to highlight where a trench is created at such a location so that the user can clearly see and / or understand the hazard when driving by, and this geometric information (e.g., which may feature a highlighted outline of the trench, or may feature one or more photographs or other non-geometric information) may be stored locally on the local wireless device (44) so that it does not need to be retrieved from more remote resources, which may involve greater latency in getting the information to the driver. In addition to reducing latency, local storage may also be used to reduce the overall computational load on the user's mobile computing system because the mobile system may receive information that it would otherwise have to generate or construct based on sensors (e.g., which may include part of the local mobile hardware).
[0132] Once the user arrives at their work parking lot (24), for example, the system can be configured to detect walking speed and, by the user, to review their day's schedule with the user as they walk to the office via integration with their computerized calendar system. Certain additional information not resident on their local mobile computing system can be retrieved from a local source (e.g., 48, 50), which can feature a particular storage capacity, again to facilitate less mobile overhead and lower latency than a direct cloud connection.
[0133] refer to Figure 4Once in the office (24), the user may connect with various wireless devices (50, 60, 62, 64, 66, 68, 70, 72, 74), each of which may be configured to provide location-based information. For example, when at point F (90), the user's mobile computing system may be configured to detect the location (such as by GPS, computer vision, marker or beacon identification, and / or triangulation of the wireless devices (60, 62, 64)), and then quickly upload information about the location (such as a dense triangular mesh of the room's geometry), or some information about whose office the room is, information about the person, or other information that may be considered relevant, from local storage (i.e., from the wireless devices 60, 62, 64) to his mobile computing system, such as by an artificial intelligence agent operating automatically on the user's mobile computing system. Various other wireless devices (50, 66, 68, 70, 72, 74) may be located elsewhere in the office and configured to feature other location-based information, again providing low-latency and robust mobile computing capabilities to local users without requiring everything (such as determining room geometry) to be done from scratch in real time by sensor infrastructure local to the mobile computing system.
[0134] refer to Figure 3 , similar wireless device resources (40, 56, 58) may be used in the home (22) to assist with location-based information as the user navigates (P-110, Q-112, R-114, S-116, T-118, U-120) home using their mobile computing system. In an office (24) or home (22) environment, the mobile computing system may be configured to utilize external resources that are completely different from driving. For example, the artificial intelligence component of the user's mobile computing system may know that the user likes to view the previous week's evening news highlights (perhaps in a display format that is generally unacceptable when driving but acceptable when walking, or automatically expanded when the user stops walking and sits or stands still) when the walking speed is detected, so when the walking speed is detected, the system may be configured to deliver such highlights from local storage between these two hours, while also collecting other location-based information, such as the location of various objects or structures within the house at relevant locations (i.e., to reduce the computer vision processing load).
[0135] Similarly, when the user navigates walking (28) through the park (26), as shown in Figure 5 As shown in the enlarged view in , local wireless device resources (54) can be used to provide location-based information, such as contextual information related to the sculpture garden that the user may be observing as he walks, and such information can be displayed or reproduced as audio as the user moves around in a customized and / or customizable manner in a strolling scene in the park (i.e., as opposed to driving or walking around at home or work).
[0136] refer to Figure 6 , one or more of the aforementioned wireless devices (40, 42, 44, 46, 48, 50, 52, 54 and Figure 3 and Figure 4 Other wireless devices shown in the enlarged view of Figure 6 The system shown includes a local controller (134), such as a processor, connected (138) to a power source (132), such as a battery; a transceiver (130), such as transmit and receive antennas configured to wirelessly communicate with the mobile computing system and other computing systems and resources, such as by using mobile telecommunications (i.e., GSM, EDGE, HSPA / +, 3G, 4G, 5G), Wi-Fi (i.e., IEEE 802.11 standards, such as 802.11a, 802.11b, 802.11g, 802.11n, Wi-Fi 6 - also known as IEEE 802.11 AX, IEEE 802.11AY, IEEE 802.11AX-Halo, which are relatively low power variations that may be most useful for devices relatively close to a user), WiMax, and / or Bluetooth (RTM, i.e., 1.x, 2.x, 3.x, 4.x) configurations; and a local storage device (136), such as a mass storage device or memory device. The storage device (136) can be connected (140) to an external storage resource (146), such as a cloud storage resource, the local power source (132) can be connected (142) to an external power source (148), such as for long-term charging or replenishment, and the transceiver (130) can be connected (144) to an external connectivity resource (150) to provide, for example, access to an Internet backbone. All of these local and connected resources can be configured based on the location of such devices to provide local context-specific information to local users, whether such information is related to transportation, shopping, weather, architecture, culture, etc. Figure 7 Shows something like Figure 6 an embodiment of an embodiment of, but without a local storage facility - its components are connected (141) to a remote storage resource (146), such as a cloud resource, Figure 7 Such an embodiment in can be used in various configurations in place of, for example, Figure 6 The embodiments in do not benefit from direct local storage (as described above, such local storage may be beneficial in reducing delays in providing information to mobile systems in the area). Figure 8In a further scenario where there is no local storage capability, a transmitter beacon (41) type device, e.g., a close-up transmitter (131) only, not a bidirectional transceiver, such as a transmitting antenna configured to communicate wirelessly with a mobile computing system and other computing systems and resources, such as by using mobile telecommunications (i.e., GSM, EDGE, HSPA / +, 3G, 4G, 5G), Wi-Fi (i.e., 802.11 standards such as 802.11a, 802.11b, 802.11g, 802.11n), WiMax, and / or Bluetooth (RTM, i.e., 1.x, 2.x, 3.x, 4.x) configurations); and a relatively long-life battery (132) can be used to connect to a locally located mobile computing device to share location or beacon identification information that serves as a pointer to connect the mobile computing system with relevant cloud resources (i.e., bypassing local storage, but providing information similar to: you are here + a pointer to the relevant cloud resource). Reference Figure 9 ,In a very basic scenario, non-electronic markers (43), such as ArUco markers, can be used to also serve as pointers to connect the mobile computing system with the relevant cloud resources (i.e., bypassing local storage, but providing information similar to : you are here + a pointer to the relevant cloud resource).
[0137] As described above, in order to reduce latency and generally increase useful access to relevant location-based information, wireless devices with localized storage resources (such as Figure 6 The depicted wireless devices may be located inside a structure, such as a residence, a business, etc., as well as outside a city district, a mall or store, etc. Similarly, wireless devices that do not have localized storage capabilities but are connected to or pointed to remote storage resources may also be located inside a structure, such as a residence, a business, etc., as well as outside a city district, a mall or store, etc.
[0138] The mobile computing system may be customized by the user to present information filtered on a temporal basis, such as by how old or "out-of-date" such information is. For example, a user may be able to configure the system to only present traffic information, etc., that is 10 minutes old or newer while he is driving (i.e., the temporal aspect may be customizable / configurable). Alternatively, a user may be able to configure the system to only present architectural information (i.e., the location of walls within a building), etc., that is 1 year old or newer (i.e., the temporal aspect may be customizable / configurable).
[0139] refer to Figures 10 to 13 It is often desirable to configure the system such that the user's position and / or orientation (i.e., via determining the position and / or orientation of a coupled component, such as a head-mounted viewing component 2 that may be coupled to the user's head) can be used to provide the user with additional and / or enhanced content and / or information about the user's particular view of the world as the user navigates the world.
[0140] For example, Figure 10 As shown in the example of Figure 1 The depicted augmented reality system (200) navigates a world. A user may enter an area (such as a walkable area, or a functional volume inside or outside a building) where certain features (such as intentionally visually detectable features) and wireless connectivity-related features have been preconfigured to be located and interpreted by the user's augmented reality system, such that the system is configured to determine the user's position and / or orientation relative to the world immediately surrounding the user. Such relatively information-rich areas may be referred to as "location islands." For example, certain connectivity resources (8) may include wireless connectivity devices, such as 802.11 devices, which may broadcast information such as an SSID and / or IP address, and whose relative signal strength may be determined and may be correlated with proximity. Further detectable features may, for example, include Bluetooth, audio, and / or infrared beacons with known locations, and / or posters or other visual features with known locations. Combined detection and analysis of these inputs, such as by a plurality of sensors (which may include components such as a monochrome camera, a color camera, a Bluetooth detector, a microphone, a depth camera, a stereo camera, etc.) connected to the head wearable component (2) of the subject system, may be used to determine the position and / or orientation of the user (202) based on analysis of information about predetermined or known locations of such items, which information about predetermined or known locations of such items may, for example, be contained on a connected resource (8), such as a cloud storage system, like, for example, a reference Figure 6 The cloud storage system described.
[0141] Reference again Figure 10Once the user's initial position and / or orientation is determined, the sensors of the user's augmented reality system, along with the specific features of the location island, may be used to maintain an updated determination of the user's position and / or orientation within the area or volume (204). As the user views and / or navigates around the venue, given the updated determination of the user's position and / or orientation relative to the venue's coordinate system, certain specific content and information may be presented to the user through the user's augmented reality system, including but not limited to content and information about other remote venues via a "traversable world" configuration (such as, for example, the content and information described in U.S. patent application Ser. No. 13 / 663,466, which is incorporated herein by reference in its entirety). The "traversable world" configuration may be configured to, for example, allow other users and objects to virtually "teleport" to different locations to see images about the venue and / or communicate with other people, whether real or virtual (206). The user interface of the user's augmented reality system may be configured to allow the user to ingest, utilize, view, and / or circumvent certain information presented through the user's augmented reality system. For example, if a user is walking through a shopping area that is particularly identifiable, feature-rich (i.e., such as a "location island") and content-rich, but does not want to see any virtual presentation of information about the shopping at that time, the user can configure his or her system to not display such information, and instead display only information that has been selected for display, such as urgent personal message information.
[0142] According to the reference Figure 10 In additional details, the first resource device is at a first location, wherein the mobile device communication interface establishes a first connection with the first resource device, and wherein the content is first content specific to a first geographic parameter of the first connection. The content providing system further comprises: a second resource device having a second resource device processor, a second resource device storage medium, a second resource device dataset including second content on the second resource device storage medium, and a second resource device communication interface forming part of the second resource device and connected to the second resource device processor and under control of the second resource device processor, wherein the second resource device is at a second location, wherein the mobile device communication interface establishes a second connection with the second resource device, and wherein the content is second content specific to a second geographic parameter of the second connection.
[0143] refer to Figure 11 , showing something similar to Figure 1 , but also illustrates a system of systems that highlight a plurality of wireless connection resources that may be used to assist a user in positioning (i.e., in determining the position and / or orientation of a component, such as a head-coupled component 2 that is operatively coupleable to a user's head). For example, with reference to Figure 11In addition to the main system components (2, 4, 6) being connected to each other and to a connection resource (8) such as cloud storage or cloud computing resources, these system components can be wirelessly coupled to devices that can assist in locating the user, for example, a Bluetooth device (222), such as a transmitter beacon with a known identity and / or location; an 802.11 device (218), such as a Wi-Fi router with a specific SSID, IP address identifier and / or signal strength or proximity sensing and / or transmission capability; a vehicle or component thereof (220), which can be configured to transmit information regarding speed, location and / or orientation (e.g., certain speedometer systems within certain motor vehicles can be configured to transmit instantaneous speed and approximate GPS location by coupling with a component having GPS tracking capability, such as an onboard GPS tracking device). Such speed, location and / or orientation information relative to the vehicle in which the user is located can be utilized to reduce display "jitter" and also to help present information to the user regarding real-world features that can be seen through the vehicle window (such as a label of a hilltop being passed, or a vehicle window). a display image of a user's system (e.g., a display image of a user's system or other features exterior to a vehicle); in certain embodiments involving vehicles or other structures having viewing entrances exterior to such vehicles or structures, information regarding the geometry of such vehicles, structures, and / or entrances, such as from a connected resource 8 cloud repository, may be utilized to appropriately place virtual content for each user relative to the vehicle or structure); a mobile connectivity network transceiver (210), such as one configured for LTE connectivity that can not only connect to the user's system but also provide triangulated position and / or orientation integration and integrated GPS information; a GPS transmitter and / or transceiver (212) configured to provide location information to connected devices; an audio transmitter or transceiver beacon (214, 216), such as an audio transmitter or transceiver beacon configured to help locate or guide nearby systems by using frequencies that are not normally audible (e.g., in various embodiments, an audio transmitter or transceiver can be used to help mobile systems, such as augmented reality systems, "honing" with minimal intrusion. in upon”) or positioning (i.e., similar to how a first person in the dark can whistle to a second person in the dark to help the second person find the first person) not only an audio transmitter or transceiver, but also another adjacent or co-located positioning asset such as a light, infrared, RF or other beacon, transmitter and / or transceiver (i.e., through the sensor suite available on the augmented reality system (such as Figure 1 and Figure 11The sensor suite (close-up in the figure) automatically, or in other embodiments manually or semi-automatically, causes the audio emitter and / or transceiver to be directionally represented in a user interface for the user, such as via a visual indicator, such as an arrow in the user interface, and / or an audio indicator through an integrated speaker in the head-mounted assembly), and / or an infrared beacon detectable by the user's augmented reality system to similarly attract and / or identify information about position and / or orientation.
[0144] refer to Figure 12 , showing the Figure 11 An embodiment of the operation of the system of the system shown. A user wears an augmented reality system to navigate (200). Within the range of various wireless connectivity resources, such as a mobile telecommunications transceiver (such as LTE), a GPS device, an 802.11 device, and various types of beacons (such as Bluetooth RF, audio, and / or infrared beacons), the user's augmented reality system can be configured to determine the user's position and / or orientation relative to the world immediately surrounding the user (224). Once the user's initial position and / or orientation is determined, the sensors of the user's augmented reality system, along with the particular wireless connectivity resources, can be used to maintain an updated determination of the user's position and / or orientation within an area or volume (226). As the user views and / or navigates around the place, given the updated determination of the user's position and / or orientation relative to the place's coordinate system, certain content and information may be presented to the user through the user's augmented reality system, including but not limited to content and information about other remote places via a "traversable world" configuration (such as, for example, the content and information described in U.S. patent application Ser. No. 13 / 663,466, which is incorporated herein by reference in its entirety), the "traversable world" configuration being configured to, for example, allow other users and objects to be virtually "teleported" to different locations to see images about the place, and / or communicate with other people, real or virtual, (206). The user interface of the user's augmented reality system may be configured to allow the user to ingest, utilize, view, and / or circumvent certain information presented through the user's augmented reality system (208), as described above with respect to Figure 10 Described by example.
[0145] refer to Figure 13 In another embodiment, other detectable resources such as buildings, skylines, horizons, and / or different geometries of panoramas may be analyzed, such as via computer vision and / or image or feature processing techniques, utilizing connected systems and resources such as Figure 1 and Figure 11The user wears an augmented reality system to navigate a world (200). Within the vicinity of various structures or other detectable resources (such as one or more buildings, a skyline, a horizon line, and / or different geometric shapes of a panorama), the user's augmented reality system may be configured to determine the user's position and / or orientation relative to the world immediately surrounding the user by processing, thresholding, and / or comparing aspects of such images with known images of such scenes or resources (228). Once the user's initial position and / or orientation is determined, sensors of the user's augmented reality system (such as color, monochrome, and / or infrared cameras) may be used to maintain an updated determination of the user's position and / or orientation within the area or volume (230). As the user views and / or navigates around the place, given an updated determination of the user's position and / or orientation relative to the place's coordinate system, certain content and information may be presented to the user through the user's augmented reality system, including, but not limited to, content and information about other remote places configured via a "navigable world" (such as, for example, the content and information described in U.S. patent application Ser. No. 13 / 663,466, which is incorporated herein by reference in its entirety), which "navigable world" may be configured to, for example, allow other users and objects to virtually "teleport" to different locations to see images about the place, and / or communicate with other people, real or virtual, (206). The user interface of the user's augmented reality system may be configured to allow the user to ingest, utilize, view, and / or circumvent certain information presented through the user's augmented reality system, such as, for example, the content and information described above with respect to Figure 10 described.
[0146] therefore, Figure 13 Additional details are described wherein the first resource device is at a first location, wherein the mobile device has a sensor that detects a first feature at the first location, and the first feature is used to determine a first geographic parameter associated with the first feature, and wherein the content is first content specific to the first geographic parameter.
[0147] refer to Figures 14 to 18 , presents a paradigm for interconnected or integrated computing, which may be referred to as "spatial computing". Figures 1 to 13 As described, connected and portable personal computing systems (such as Figure 1Such systems (such as those shown) can be integrated into the user's immediate world, enabling the user to be present in the space around them while also interacting with the computing system in complex ways. This is due in part to the portable nature of such systems, but also due to the connectivity of various resources as described above, and also due to the fact that various embodiments are designed to facilitate normal activities while also operating the computing system. In other words, in various instances, such computing systems can be worn and operated by a person as they move about and work through their daily living spaces, indoors or outdoors, and the computing system can be configured and adapted to provide specialized and customized functionality to the user based on or in response to certain inputs provided by the spatial environment surrounding the user.
[0148] refer to Figure 14 The concept of "spatial computing" can be defined as being associated with multiple attributes, including but not limited to presence, persistence, scale, perceptibility, interactivity, respect, and feeling. For example, in terms of presence, the subject-integrated system components can be configured to amplify the user's capabilities without affecting the user's presence in the physical world. In terms of persistence, the subject-integrated system components can be configured to facilitate the life cycle of the user's attention, focus, and interaction with the physical world around them and any "digital residents" associated therewith (i.e., visitors such as digital or virtual characters or digital representations) to grow in context and depth over time, where digital content and residents can be configured to continue their life paths even if the subject computing system is not actively utilized by the user. In terms of scale, the subject-integrated system components can be configured to deliver anything from very large scale to relatively small scale images or content to the user. In terms of perceptibility, the subject-integrated system components can be configured to utilize information, signals, connected devices, and other inputs to provide the user with an enhanced understanding of the physical and digital world around them. In terms of interactivity, the subject-integrated system components may be configured to present digital content that responds to natural human signals or inputs (such as head, eyes, hands, voice, face, and other inputs) and inputs associated with various tools. In terms of respect, the subject-integrated system components may be configured to provide images, content, and digital behaviors that are integrated into the world around the user, such as by providing integrated light and shadow projections to match the room, by adding reverberation to the audio to match the sound physics of the room, or by appropriately occluding various digital / virtual objects relative to each other as perceived in reality. In terms of sensation, the subject-integrated system components may be configured to provide an integrated level of awareness of who the user is and what the world around the user is like, enabling the system to understand the world in a comprehensive way as the user does, and to use this comprehensive intelligence to provide a personalized and subjective experience to the user.
[0149] refer to Figure 15, illustrates yet another way of representing the relationship between a user and the physical world in terms of a spatial computing system. With breakthroughs in computing hardware, such as in central processing units ("CPUs"), graphics processing units ("GPUs"), along with ubiquitous high-bandwidth data transfer capabilities and relatively inexpensive and high-speed storage devices, there is a convergence of important factors for spatial computing configurations. At the center of the depicted spatial computing paradigm is the user, who has a highly evolved and capable human computing system (i.e., a brain). Immediately adjacent to the depicted user are intuitive interface layers, including such Figure 1 The system shown, wherein a user wears a wearable computing system and interacts with it using voice commands, gestures, handheld components, etc., as described in the description incorporated by reference above. The human-machine interface device (250) may also include computing systems such as tablet computers, smart phones, and laptop computers, vehicles (e.g., autonomous vehicles), various types of wearable electronic devices, drones, robots, and various other systems that provide access to computing and connectivity resources to human users, for example, in Figure 16 , which shows yet another hierarchical description of the connected elements of a spatial computing environment.
[0150] The next layer depicted is a security / encryption layer (252). This layer is configured to isolate and protect the user from other systems or users that may want to gain access to the user's data or metadata, which may include, for example, what the user likes, what the user does, where the user is located. Technologies such as blockchain can be used to assist in securing and configuring such a security layer, as well as interacting with digital identities that can be configured to securely represent who the user is in the digital world; and can be facilitated by biometric authentication technologies.
[0151] The next adjacent positioning layer is the human-centric integrated spatial computing layer (254), which may also be referred to under the trade name "MagicVerse". This spatial computing layer is also Figure 16 As shown in the figure, it can be configured to receive data from multiple sources, including but not limited to external developers, private developers, government sources, artificial intelligence sources, pre-trained models, deep learning models, psychological databases, current events databases, data sources about individual operations of one or more users (which may also be referred to as the user's "life stream"), device data sources, authoritative data sources, corporate data sources, learning data sources, mesh data sources, contextual data sources, government services, public sources, competitor sources, communication sources, emerging data sources, device data sources, grid or mapping operations, contextualized information, device data sources and management services.
[0152] Forward Reference Figure 17 When assembling spatial computing architectures, it is useful to understand some basic principles of how humans process and store information. Figure 17As shown, the human brain is able to receive a large amount of information due to the operation of the eyes, ears, and other human senses, and this information can enter the sensory memory buffer. It turns out that much of this information is not retained, perhaps because it is as useful as the other retained parts. Components of useful information can be moved to the working memory buffer and may eventually be moved to long-term memory, where there is potentially unlimited storage and unlimited retention. In various embodiments, the spatial computing system can be configured to increase the proportion of information pushed into working memory that is actually useful to the user. In other words, the preference is to add value to the user as much as possible, especially when the cognitive load on the user's brain is increased.
[0153] Return Reference Figure 15 , a human-centric spatial computing layer or "MagicVerse" that interfaces with everything around it. For example, Figure 18 In various embodiments, a human-centric spatial computing layer is shown in coupling information with many different data sources in a configuration that is as parallel as possible. The human-centric spatial computing layer can be understood as integration with all systems around the user, from smartphones to IoT-capable devices, to virtual reality / augmented reality / mixed reality (so-called "XR") devices, to various vehicles, databases, networks, and systems.
[0154] Figure 15 The next layer shown is an abstraction and mediation layer (256) inserted between the user, human-machine device interface, human-centric spatial computing layer and the computing resource layer such as edge computing (258), fog computing (260) and cloud computing resources (262).
[0155] Figures 15 to 18 Additional details of a spatial computing layer between a mobile device and a resource layer are described, the resource layer having multiple data sources and being programmed to receive data resources, integrate the data resources to determine an integration profile, and determine first content based on the integration profile. The spatial computing layer includes a spatial computing resource device having a spatial computing resource device processor, a spatial computing resource device storage medium, and a spatial computing resource device data set on the spatial computing resource device storage medium, and the spatial computing resource device data set is executable by the processor to receive data resources, integrate the data resources to determine an integration profile, and determine first content based on the integration profile. The content provision system also includes an abstraction and arbitration layer that is inserted between the mobile device and the resource layer and is programmed to make workload decisions and distribute tasks based on the workload decisions.
[0156] A portion of the three outermost layers is depicted as missing to represent the fact that the real physical world (264) is part of the integrated system. In other words, the world can be used to assist in calculations, make various decisions, and identify various items. The content providing system includes a camera device that captures images of the physical world surrounding the mobile device. The images can be used to make workload decisions. The images can form one of the data resources.
[0157] refer to Figures 19A to 26C , shows various embodiments of connection alternatives suitable for use with various subject space computing configurations. For example, referring to Figure 19A , shows a configuration that utilizes relatively high bandwidth mobile telecommunication connections (such as those available using 3G, 4G, 5G, LTE, and other telecommunication technologies) and local network connections (such as endpoints connected via Wi-Fi and cable), where such connection improvements can be used to connect to remote (i.e., not directly onboard the user) computing resources with relatively low latency.
[0158] Figure 19A Shows a user wearing Figure 1 The depicted system of systems configuration includes, for example, a head wearable component (2) and an interconnected auxiliary computing or controller component (6, also referred to as a "belt pack" due to the fact that in some variations such a computing component may be configured to be attached to a user's belt or waist region). Figure 19B Shows something like Figure 19A The configuration of the configuration is a configuration of the configuration of the configuration, except that the user wears a single component configuration (i.e., it only has a head wearable component 2, where the functionality of the interconnect auxiliary computing or controller component 6 is off-person in the form of fog, edge and / or cloud resources). Return to reference Figure 19A , so-called “edge” computing and / or storage resources are typically located closer to users than “fog” computing and / or storage resources, which are closer than the typically more powerful and more distant “cloud” resources. Figure 19A and Figure 19B As shown, edge resources, which are typically closer (and less powerful in terms of raw computing resources), will be available to users' local interconnected computing resources (2, 6) at relatively lower latency than fog resources (which will typically have an intermediate level of raw computing resources) and cloud resources located further away. Fog resources are generally defined as having intermediate latency, and cloud resources have the most latency and typically the most raw computing power. Thus, with such a configuration, there is more latency and more raw computing resource power as resources are located further away from the user.
[0159] Figure 19A and Figure 19BAdditional details are described, wherein the mobile device communication interface includes one or more mobile device receivers connected to the mobile device processor and connected to the second resource device communication interface in parallel with the connection to the first resource device to receive second content. In the given example, the second resource device is a fog resource device having a second latency that is slower than the first latency. The mobile device communication interface includes one or more mobile device receivers connected to the mobile device processor and connected to the third resource device communication interface in parallel with the connection to the second resource device to receive third content sent by a third resource device transmitter, wherein the third resource device is a cloud resource device having a third latency that is slower than the second latency. The connection to the edge resource device is through a cellular tower, and the connection to the fog resource device is through a Wi-Fi connected device. The cellular tower is connected to the fog resource device. The Wi-Fi connected device is connected to the fog resource device.
[0160] In various embodiments, such configuration is controlled to distribute workload to various resources depending on the need for computing workload with respect to latency. For example, referring to Figure 20A and Figure 20B , which shows the user and something like Figure 19A and Figure 19B The connected computing system shown can distribute computing among edge, fog and cloud computing resources based on latency and computing requirements. Figure 19A and Figure 19B , such computing needs can be directed from the user's local resources through application programming interfaces ("APIs") configured to perform computational abstraction, artificial intelligence ("AI") abstraction, network abstraction, and arbitrate computation to distribute computational tasks to various edge, fog, and cloud computing resources based on latency and computational demand. In other words, no matter which XR computing device the user has locally (i.e., such as an XR computing device) that will interface with the human-centric integrated spatial computing system (again, also referred to as the "MagicVerse"), Figure 1The wearable system shown, or tablet computer, etc.), can separate workloads and based on what format that workload appears in, where it needs to be used for external compute and / or storage resources, and the size of the file types involved - the compute abstraction node can be configured to manipulate such processing and guidance, such as by reading specific file formats, caching packets of specific sizes, moving certain documents or files as needed, etc. For example, the AI abstraction node can be configured to receive a workload and check what kind of processing model needs to be utilized and run, and fetch specific workload elements from memory so that it can run faster when more data is received. The network abstraction node can be configured to constantly analyze connected network resources, analyze their connection quality, signal strength and the ability of various types of processes that will be encountered, so that it can help to guide workloads to various network resources as best as possible. The arbitration node can be configured to perform when dividing various tasks and subtasks and sending them to various resources. Reference Figure 21 , the computational load can be distributed in multiple ways and directions so that the net result at the endpoint is optimized in terms of performance and latency. For example, head pose determination (i.e., when wearing a headset such as Figure 1 When the head-worn assembly 2 is depicted, the orientation of the user's head relative to the user's surroundings will generally affect the user's ability to navigate in situations such as Figure 1Perception in the augmented reality configuration shown), where relatively low latency may be most important, can be run at least primarily using edge resources, while semantic tagging services where ultra-low latency is less important (i.e., labels are not absolutely necessary at runtime) can be run using data center cloud resources. As described above, edge resources are typically configured to have lower latency and lower raw computing power, and they may be located at the "edge" where the computing is needed, hence the name of the resource. There are many different ways to configure edge computing resources. For example, commercially available edge-type resources may include resources sold by Intel, Inc. under the trademark "Movidius Neural Computer Stick" (TM) or Nvidia, Inc. under "Jetson Nano" (TM) or "Jetson TX2" (TM), or resources sold by Nvidia, Inc. under "AGX Xavier" (TM). Each of these typically includes CPU and graphics GPU resources, which can be coupled to camera devices, microphone devices, memory devices, etc., and connected to user endpoints, for example, with a high-speed bandwidth connection. Such resources can also be configured to include custom processors, such as application-specific integrated circuits ("ASICs"), which can be specifically used for deep learning or neural network tasks. Edge computing nodes can also be configured to be aggregated to increase the computing resources of any room a user is in (i.e., a room with 5 edge nodes can be configured to functionally have 5 times the computing power of a room with no edge nodes), abstraction layers and arbitration layers (such as Figure 19A and Figure 19B The abstraction layer and arbitration layer shown) coordinate such activities).
[0161] refer to Figure 22A and Figure 22B , shows a configuration utilizing a room with multiple cameras around a user. For various reasons, the images from the cameras can be separated and directed to different computing resources. For example, in a scenario where the head pose determination process will be repeated at a relatively high frequency, frames can be sent in parallel to edge nodes for relatively low latency processing, which can be referred to as a dynamic resolution computing scheme (in one variant, it can be configured for tensor training decomposition to perform dynamic degradation of resolution, which can be simplified to taking tensors and combining lower rank features so that the system only computes fragments or portions). Frames can also be sent to fog and cloud resources for further contextualization and understanding relative to the user's XR computing system ( Figure 22A Shown with Figure 1 A similar user computing configuration is shown in the , Figure 19A 、 20A , 23A and 24A user calculation configurations are also the same), Figure 22B Shown as Figure 19B 、 20B, 23B and 24B, there is no interconnected auxiliary computing component (6), since the functionality of such resources is instead located within the aggregation of edge, fog and cloud resources.
[0162] Figure 22A and 22B Additional details of at least one camera that captures at least first and second images are shown, wherein the mobile device processor sends the first image to an edge resource device for faster processing and sends the second image to a fog resource device for slower processing. The at least one camera is a room camera that captures the first image of the user.
[0163] refer to Figure 23A and 23B , various "Internet of Things" (i.e., configured to readily interface with network control, typically through an Internet connection such as an IEEE 802.11-based wireless network) resources can be connected to a user's local computing resources via edge computing resources, and thus, processing and control with respect to each of these resources can be at least partially accomplished outside of the wearable computing system.
[0164] refer to Figure 24A and 24B Many types of wearable technologies can connect to edge computing resources, such as via available network connection modalities, such as IEEE 802.11 Wi-Fi modality, WiFi-6 modality, and / or Bluetooth (TM) modality.
[0165] refer to Figure 25A , for something like Figure 1 The system of systems depicted may be configured to allow a user to couple their local system to external resources for additional computing, storage, and / or power using a wired connection, such as via direct coupling to one or more antennas, computing workstations, laptops, mobile computing devices such as smartphones and / or tablets, edge computing resources, and power supplies for powering the local computing system power source (i.e., batteries). Such connectivity may be facilitated by an API configured to operate as a layer above each connected resource operating system, such as Android(TM), iOS(TM), Windows(TM), Linux(TM), etc. Preferably, such resources may be added or disconnected by the user during operation, depending on the user's ability to maintain close proximity and desire to utilize such resources. In another embodiment, multiple such resources may be bridged together in an interconnected manner with the user's local computing system. Reference Figure 25B , showing something similar to Figure 25AEmbodiments of the embodiment of the embodiment, except that the interconnection auxiliary computing component (6) can be coupled to the user and wirelessly coupled to the auditory wearable component (2) such as via Bluetooth, 802.11, WiFi-6, Wi-Fi Halo, etc. Such a configuration can be used in the above context, such as in Figure 19A 、 20A , 22A, 23A, and 24A (in other words, with such a "hybrid" configuration, the user still has the interconnected auxiliary computing component 6 on their body, but the component is wirelessly connected to the head wearable component 2 rather than being connected through a tethered coupling). Figure 25C , the user is shown with a head-mounted assembly such as Figure 1 The head-mounted assembly (element 2) shown in FIG, which is wirelessly coupled to other computing resources, such as Figure 19B 、 20B In the embodiments of 22B, 23B and 24B, an XR device on the user's personal device is also shown, which can also be coupled to the depicted remote resources, such as through a wired or wireless connection. Figure 25D It is shown that such remote resources can be coupled not only to each other, but also to various local computing resources, such as resources that can be coupled to, for example, a car in which a user may be sitting. Figure 25E , showing something similar to Figure 25C Configuration of the configuration, but also remind, such as reference Figure 19B 、 20B , 21, 22B, 23B and 24B, additional computing and / or storage resources may also be coupled to each other, such as fog and / or cloud resources.
[0166] refer to Figures 26A to 26C ,As relatively high-bandwidth, low-latency mobile connections become more common, instead of or in addition to placing mobile devices in a cloud of valid signals from a variety of sources, e.g. Figure 26A As shown, beamforming configurations such as Figure 26B The beamforming configuration of the phased array antenna configuration shown in close-up can be used to effectively "steer" and concentrate or focus connections toward one or more specific mobile computing devices, such as Figure 26C With enhanced knowledge or understanding of the location and / or orientation of a particular mobile computing device (e.g., in various embodiments, due to gesture determination techniques, such as techniques involving a camera positioned on or coupled to such a mobile computing device), connection resources can be provided in at least a more directional and conservative manner (i.e., gestures can be mapped and fed back to beamforming antennas to more precisely guide connections). In one variation, as Figure 26CThe configuration shown can be implemented, for example, using a radar hologram type transmission connector, so that focused signal coherence is used to form an effective point-to-point communication link.
[0167] Figures 26A to 26C Additional features of the content provision system are described, further comprising: a sensor (350) that provides sensor input to a processor (352); a pose estimator (352) executable by the processor (352) to calculate a pose of the mobile device (head-mounted viewing assembly (2)) based on the sensor input, including at least one of a position and an orientation of the mobile device; a steerable wireless connector (358) that creates a steerable wireless connection (360) between the mobile device and an edge resource device (362); and a steering system connected to the pose estimator and having an output that provides input to the steerable wireless connector to steer the steerable wireless connection to at least improve the connection. Although all of these features are not present in 26A to 26C, they may be inferred from 26A to 26C or other figures and associated descriptions herein.
[0168] refer to Figure 27 , shows a depiction of fog computing (which may also be referred to as "ambient computing") with different computing "rings" corresponding to latency levels relative to user devices, such as various XR devices, robots, cars, and other devices shown in the outer rings. The adjacent rings inward show edge computing resources, which may include various edge node hardware, such as the edge node hardware described above, which may include various forms of proprietary 5G antennas and connectivity modalities. The center ring represents cloud resources.
[0169] refer to Figures 28A to 28C ,Between the edge, fog and cloud layers, there can be a communication layer that includes various forms of connectivity, including fiber, coaxial cable, twisted pair cable satellite, and various other wireless connectivity modalities. Figure 28A shows users wirelessly connected to edge resources that can sequentially connect to fog and cloud resources, Figure 28B shows a configuration of user hardware connections (i.e., via cables, optical fibers, or other non-wireless connection modalities) to edge resources that can sequentially connect to fog and cloud resources, Figure 28C The user is shown to be both wirelessly and hardwired to all resources in parallel, and therefore various such connection permutations are contemplated. Figure 21One challenge with managing such resources is how to move and distribute the computational load up and down the hierarchy between edge, fog, and cloud resources. In various embodiments, a deep arbitration layer based on reinforcement learning can be utilized (i.e., in one variation, function requests can be sent to edge resources, certain tasks can be sent to fog resources, and the "reward" of the reinforcement learning paradigm might be whatever compute node configuration finishes fastest, and thus a feedback loop optimizes the responsive arbitration configuration, where the overall functionality is one that improves as the resources "learn" how to optimize).
[0170] Return Reference Figure 15 , depicting the physical world impacting the edge, cloud, and fog layers, which means that the farther away something is, the closer it is functionally to a cloud data center, and those resources can be accessed in parallel.
[0171] refer to Figure 29A and Figure 29B , showing various types of connection resources for XR devices, such as for Figure 1 The head-wearable component (2) of the system of systems shown, as described above, such components can be connected to other resources using hardware-based connections as well as various wireless connection paradigms. Figure 30A , having a head wearable component hardware coupled to a belt-pack computing component similar to Figure 1 The configuration of the configuration can be connected to the reference as above for example Figure 19A 、 20A , 22A, 24A and 25A, the various resources of leaving people, Figure 30B It shows that one or more XR devices (such as head wearable AR components, tablet computers or smart phones) can be connected to the Figure 19B 、 20B , 22B, 24B and 25B, 25C, 25D and 25E.
[0172] refer to Figure 31 , illustrates a paradigm (400 to 428) for arbitrating and allocating resources relative to external resources, such as edge computing (416A), fog computing (416B), and cloud computing (416C). For example, Figure 31As shown, data can be received (400) from a user's local XR device (i.e., such as a wearable computing component, tablet computer, smartphone), a simple logical format check (402) can be performed, and the system can be configured to verify that the correct choice was made in the format check using simple error checking (404). The data can then be converted to a preferred format (406) for use in an optimized computing interface. Alternatively, APIs for various connected resources can be utilized. Additional error checking (408) can be performed to confirm the appropriate conversion. A decision (410) is then presented as to whether the relevant process is a runtime process. If so, the system can be configured to immediately process the issue and return it (412) to the connected device. If not, batch processing of external resources can be considered, in which case an arbitrator function (414) is configured to analyze how many cloud, edge, and / or fog resource (416A to 416C) instances are available and will be utilized, and send out processing tasks. Once the data is processed, it can be combined (422) back into a single model, which can be returned (424) to the device. A copy of the combined model (418) can be taken along with the relevant resource configuration details and fed back (420) to the arbitrator (414) and run in the background with different scenarios (and will continue to do so as it "learns" and optimizes the configuration to apply to that type of data input, with the error function eventually moving towards zero).
[0173] Figure 31 A content provisioning system (400 to 428) is described as including an arbitrator function executable by a processor to determine how many edge and fog resources are available through edge and fog resource devices, respectively, send processing tasks to the edge and fog resources based on the determination of available resources, and receive returned results from the edge and fog resources. The arbitrator function is executable by the processor to combine the results from the edge and fog resources. The content provisioning system also includes a runtime controller function executable by the processor to determine whether a process is a runtime process, if a determination is made that the task is a runtime process, then immediately execute the task without making a determination using the arbitrator function, and if a determination is made that the task is not a runtime process, then making a determination using the arbitrator function.
[0174] refer to Figure 32, shows a variation of the general operational diagram for human-centric integrated spatial computing ("MagicVerse"). The basic operational elements for creating spatial computing type experiences are shown in a relatively simplified form. For example, referring to the bottom-most depicted layer, the physical world can be at least partially reconstructed geometrically via simultaneous localization and mapping ("SLAM") and visual odometry techniques. The system can also be connected to information related to environmental factors such as weather, power configuration, nearby IoT devices, and connectivity factors. Such data can be used as input to create what can be called a "digital twin" of the environment (or a digitally reconstructed model of the physical world, such as from Figure 32 The system is preferably configured to leverage the presence and data capture of other devices within the same or nearby environments so that the aggregation of these data sets can be presented and utilized to assist in interaction, mapping, positioning, orientation, interaction, and analysis relative to such environments. The depicted “rendering layer” can be configured to aggregate all information being rendered by local devices and other nearby devices that may be interconnected by resources such as edge, fog, and / or cloud. This includes creating regions for starting and stopping predictions, enabling more complex and higher dimensional predictions at runtime. Also depicted is a “semantic layer” that can be configured to segment and label objects in the physical world as represented in a digital model of the world to be displayed to one or more users. This layer can generally be considered the foundation for natural language processing and other recurrent neural networks (“RNNs”), convolutional neural networks (“CNNs”), and deep learning methods. A “context layer” can be configured to provide results returned from the semantic layer in human-relevant terms based on inferred actions and configurations. This layer and the semantic layer can be configured to provide key inputs to “artificial intelligence” functionality. The “experience layer” and “application layer” are exposed directly to the user in the form of executable user interfaces, for example, using various XR devices such as Figure 1 These upper layers are configured to leverage information from all of the depicted layers below them to create a multi-user, multi-device, multi-location, multi-application experience, preferably in a manner that is as usable as possible to a human operator.
[0175] refer to Figure 33 , shows a view showing the connections of multiple overlapping edge computing nodes within a larger fog computing node, enabling seamless handoff or transfer between edge computing devices. In a system designed for a massively multiplayer online (MMO)-like system, various components can be configured to scan the nearby physical world environment in real time or near real time to continuously build more refined mesh and map construction information. In parallel, feature identifiers used within these meshes and maps can be used to create very reliable and high-confidence points, which can be referred to as "superpoints." These superpoints can be contextualized to represent physical objects in the world and used to overlap multiple devices in the same space. When the same application is running on multiple devices, these superpoints can be used to create anchors referenced to a common coordinate system, which is used to project the content's transformation matrix to each user on each associated XR device. In other words, each of these users can utilize superpoints to assist with pose determination and positioning (i.e., position and orientation determination in the local coordinate system relative to the superpoints), effectively becoming reliable waypoints for XR users to navigate in real or virtual space.
[0176] As in an MMO-like system, content is stored based on the coordinate system within which the users are interacting. This can be done in a database where a particular piece of content can have JSON (JavaScript Object Notation) that describes the location, orientation, state, timestamp, owner, etc. This stored information is typically not replicated locally on the user's XR device until the user crosses a location-based boundary and begins downloading content in the predicted future. To predict the future, machine learning and / or deep learning algorithms can be configured to use information such as speed, acceleration, distance, direction, behavior, and other factors. The content can be configured to always be in the same location unless interacted with by one or more users to be relocated, either manually (i.e., such as drag and drop using the XR interface) or through programmatic manipulation. Once new content is downloaded based on the aforementioned factors, it can be saved to the local computing device, such as with a reference Figure 1 The local computing device shown in the form of a computing package (6) is provided. Depending on contextual factors and available resources, such downloads may be performed continuously, periodically, or on demand. In various embodiments, such as Figure 1 In the depicted embodiments, where no significant connected computing resources are available, all rendering (such as for head poses and other mixed reality general operation runtime components) may be done on local computing resources (i.e., such as a CPU and GPU, which may be located locally to the user).
[0177] Figure 33A content provision system is described, including a plurality of edge resource devices, wherein data is exchanged between the plurality of edge resource devices and fog resource devices, the data including spatial points captured by different sensors and transmitted to the edge resource devices; and a superpoint calculation function, executable by a processor, for determining a superpoint, which is a point selected from two or more points where data from the edge resource devices overlap. The content provision system also includes a plurality of mobile devices, wherein each superpoint is used in each mobile device to estimate the location, orientation, or pose of the corresponding mobile device.
[0178] refer to Figure 34 , shows a general stratup / bootup procedure for a configuration that can have interconnected distributed resources (i.e., such as edge, fog and / or cloud resources) and end-to-end encryption between them to provide connectivity, for example, through wireless or wired connections, such as the aforementioned WiFi, mobile wireless, and other hardware-based and wireless-based connection paradigms.
[0179] refer to Figure 35 , shows another MMO-like configuration, which shows the generalization of computational requirements relative to the size of individual XR user nodes. In other words, in various embodiments, it takes a relatively large amount of external resources to support each individual XR user node, as shown in FIG. Figure 16 As shown, many of these resources are driven toward each instance of human-centric integrated spatial computing.
[0180] Figure 35 A content provision system is described that includes a contextual trigger function that is executed by a processor to generate a contextual trigger for a set of superpoints and store the contextual trigger on a computer-readable medium.
[0181] refer to Figure 36 , shows the use of Figure 35Various computing stacks for an embodiment of human-centric integrated spatial computing are shown. The physical world is shared by all users, and based on the parameters of each specific XR device used by each user, the system is configured to abstract interactions with each such device. For example, an Android™ phone will use the Android™ operating system and can use AR core™ for perception and anchoring. This allows the system to work with an already developed perception stack that can be optimized for interaction with many specific devices in a human-centric integrated spatial computing world. To enable use with many different devices, a software development kit ("SDK") can be distributed. The term "spatial atlas" can be used to refer to a grid ownership and operational configuration, where various objects, locations, and map features can be labeled, owned, sold, excluded, protected, and managed using a central authoritative database. A "context trigger" is a description of one or more superpoints and associated metadata that can be used to describe what various things are in the real and / or digital world and how they can be utilized in various ways. For example, superpoints on the four corners of a phone structure can be designated to cause the phone structure to operate as a hilt in a digital environment. At the top of the illustrated stack, application client and backend information is shown, which may represent an XR and / or server-level interaction layer for a user to interact with the human-centric integrated spatial computing configuration.
[0182] refer to Figure 37 , showing the configuration for discovering, switching, and controlling elements within the direct radius of the XR user. The rectangular stack refers to application management tools suitable for spatial computing, and the circular stack belongs to the integration of digital models of the physical world (i.e., "digital twins"), third-party developer content via published SDKs and APIs, and XR configuration as a service, where users of XR devices can effectively log in to a curated environment to experience augmented or virtual reality images, sounds, etc.
[0183] What might be called "spatial understanding" is an important skill for machines that must interact with the physical 3D world either directly (e.g., a robot that walks and picks up thrashes) or indirectly (e.g., mixed reality glasses that create high-quality 3D graphics that respect the geometry of the 3D world). Humans generally already have excellent spatial reasoning skills, but even "smart glasses" worn by a user must perform their own version of spatial reasoning. Computer scientists have developed certain algorithms suitable for extracting three-dimensional data and spatially reasoning about the world. The concept of "superpoints" has been described above and in the associated incorporated references, and in this paper we describe a superpoint-based spatial reasoning formulation that relies heavily on the input image stream.
[0184] As mentioned above, a specific class of 3D construction algorithms associated with spatial computing configurations is often referred to as simultaneous localization and mapping ("SLAM"). SLAM systems can be configured to take as input a series of images (color or depth images) and other sensor readings (such as an inertial measurement unit ("IMU")) and provide real-time localization of the current device pose. The pose is typically a rotation matrix R and a translation vector t relative to the camera coordinate system. SLAM algorithms can be configured to generate poses because they interleave two core operations behind the scene: mapping and localization with respect to a map.
[0185] The term "visual SLAM" can be used to refer to variants of SLAM that rely heavily on camera images. In contrast to 3D scanners (such as LIDAR) developed by the robotics community to help perform SLAM for industrial applications, cameras are significantly smaller, more ubiquitous, cheaper, and easier to work with. Modern off-the-shelf camera modules are very small, and using camera images, it is possible to build very small and effective visual SLAM systems.
[0186] What can be called "monocular visual SLAM" is a variation of visual SLAM that uses a single camera. The main benefit of using a single camera is the reduced client form factor, as opposed to using two or more cameras. With two cameras, great care must be taken to keep the multi-camera assembly rigid.
[0187] The two main disadvantages of monocular vision SLAM are as follows:
[0188] 1. Monocular vision SLAM is generally unable to recover 3D structure when there is no parallax, such as when the camera system rotates about the camera center. This is particularly problematic when the system initially starts without 3D structure and when there is insufficient parallax motion during the initialization of the algorithm's internal 3D map.
[0189] 2. The second challenge is that monocular visual SLAM is usually not able to recover the absolute scale of the world. For any given monocular trajectory with recovered point depth Zi, it is possible to multiply all point 3D coordinates by a scalar alpha and also scale the translation by alpha.
[0190] However, a small amount of additional information, such as odometry readings from a local IMU or depth data from an RGBD sensor, is enough to prevent monocular algorithms from degrading overall.
[0191] Visual SLAM systems operate on images and produce poses and 3D maps. The entire algorithm can be broken down into two stages: the front end and the back end. The goal of the front end is to extract salient 2D image features and describe them so that the original RGB image is no longer required. The task of the front end is often called "data abstraction" because the high-dimensional image is reduced to 2D points and descriptors, whose nearest neighbor relationships are preserved. For properly trained descriptors, we can take the Euclidean distance between them to determine whether they correspond to the same physical 3D point - but Euclidean distance on the original image is meaningless. The goal of the back end is to take the abstractions from the front end, namely the extracted 2D points and descriptors, and stitch them together to create a 3D map.
[0192] As described above and in the incorporated references, superpoint is a term that can be used for convolutional deep learning-based front-ends designed for monocular visual SLAM. Traditional computer vision front-ends for visual SLAM can consist of hand-crafted 2D keypoint detectors and descriptors. Traditional methods typically follow the following steps: 1.) extract 2D keypoints, 2.) crop patches from the image around the extracted 2D keypoints, 3.) compute descriptors for each patch. The "deep learning" configuration allows us to train a single multi-head convolutional neural network that jointly performs interest point and descriptor computation.
[0193] Given a set of 2D points tracked in an image, "bundle adjustment" is the term used for algorithms that can be used to jointly optimize the 3D structure and camera pose that best explain the 2D point observations. The algorithm can be configured to minimize the reprojection error (or rectification error) of the 3D points using a nonlinear least squares formula.
[0194] SuperPoint, a deep learning formulation for feature extraction, can be configured as a network with very little manual engineering (i.e., manual input). While a network can be designed to take an image as input and provide 2D point locations and associated descriptors, it typically does not do so until it has first been trained on an appropriate dataset. How these 2D point knowledge is extracted is never explicitly described. The network is trained using backpropagation on a labeled dataset.
[0195] We have previously described how to train the first part of a superpoint: keypoint localization of the head. This can be achieved by creating a large synthetic dataset of corners and the resulting network (which is just like a superpoint, but does not contain descriptors), which can be called a "magicpoint".
[0196] Once we have the singularity, we have a way to extract 2D keypoints from arbitrary images. Typically, we still need to 1.) improve the performance of keypoint detection on real-world images and 2.) add descriptors to the network. Improving performance on real-world images means we must train on real-world images. Adding descriptors means we typically have to train on image pairs, as we must provide the algorithm with positive and negative keypoint pairs for it to learn its descriptor insertions.
[0197] So-called singularity configurations can be run on real images using synthetic homographies with a procedure called “homography adaptation.” This provides better labeling on those images than running the singularity only once per image. The synthetic homography can be used to take an input image I and create two warped versions I’ and I”. Since a combination of homographies is still a homography, one can train a pair of images (I’, I”) with the homography between them.
[0198] The resulting system can be referred to as superpoint-v1 (superpoint_v1) and is the result of running Singularity on random real-world images using homography adaptation.
[0199] At this point, we have what we can call "superpoint-v1", which provides all the convolutional front-ends needed for a bare-bones visual odometry or visual SLAM system. However, superpoint_v1 is trained using random (non-time-sequential) images, and all image-to-image variation is due to synthetic homographies and synthetic noise. To make a better superpoint system, we can retrain "superpoint_v2" on real-world sequences using the output of SLAM.
[0200] refer to Figure 38 ,Superpoint-based SLAM consists of two key components, a deep learning-based front end (see Box 2) and a bundle adjustment-based back end (see Boxes 3 and 4). For the front end or "feature extraction" stage of the pipeline, one can use a superpoint network as described above, which produces 2d point locations and real-valued descriptors for each point.
[0201] The backend can be configured to perform two tasks: provide the current measured pose (positioning, see Figure 1 2 in ), and integrating the current measurements into the current 3D map (map update, see Figure 1The system can utilize auxiliary input in the form of depth images (such as obtained from a multi-view stereo system, a depth sensor, or a deep regression network based on deep learning) and auxiliary poses (such as obtained from existing augmented reality frameworks such as Apple's (TM) ARKit (TM) and Google's ARCore (TM) in consumer smartphones). The main system can be configured to not require auxiliary input (see Figure 38 1, bottom), but they can be used to improve localization and map building.
[0202] Positioning module (see Figure 38 Block 3 in
[15] can be configured to take as input the points and descriptors and the current 3D map (which is a collection of 3D points and their associated descriptors) and produce a current pose estimate. This module can leverage information from the depth map by associating effective real-world depth values with observed 2D points. Auxiliary pose information from block 1 can also be fed as input to the localization module. In the absence of auxiliary information such as depth or pose, the module can be configured to use a perspective-n-point ("PnP") algorithm to estimate the transformation between 3D points in the map and 2D point observations from the current image.
[0203] The map update module can be configured to take as input the current map and the current image observation with the estimated pose, and produce a new, updated map. The map can be updated by minimizing the reprojection error of 3D points in a large number of keyframes. This can be achieved by setting up a bundle adjustment optimization problem, as is commonly done in the known "Structure-from-Motion" computer vision literature. The bundle adjustment problem is a nonlinear least-squares optimization problem and can be efficiently solved using a second-order Levenberg–Marquardt algorithm.
[0204] Typically, a system using a superpoint can be configured so that it requires certain design decisions regarding the placement of computations. At one extreme, all computations involved in processing the camera sensor and forming a well-performing intensity image must occur locally (see Figure 38 Box 1 in the figure). Below is a series of four distributed SLAM systems that perform a subset of the necessary computations locally, with the remaining operations performed in the cloud. The four systems are as follows: 1. Local on-device, 2. Cloud map building, 3. Local superpoints, and 4. All cloud.
[0205] SLAM on a local device (100% local). In one extreme case, it is possible to take all the information required for localization and mapping and put it directly on the device with the camera. In this scenario, Figure 38 Boxes 1, 2, 3, and 4 in the example can all be executed locally.
[0206] Local superpoint, local localization, cloud map construction SLAM (66.7% local). Another client variant is one that performs superpoint extraction locally and localization against a known map. With this configuration, the only part running on the cloud is the map update operation (box 4). The cloud component will update the map (using a potentially larger set of computing resources) and send a version of the map down to the client. This version allows for more seamless tracking in the presence of communication channel interruptions.
[0207] Local superpoint, cloud positioning, cloud mapping SLAM (33.3% local). In this embodiment, although camera capture must be performed locally on the device (i.e., a camera next to a server rack in a data center will not help with SLAM), it is possible to perform only a subset of the computations on the local device and the rest in the cloud. In this version of the client, boxes 1 and 2 (see Figure 38 ) is performed locally, while blocks 3 and 4 are performed in the cloud. To achieve this, the local system typically has to send points and descriptors to the cloud for further processing.
[0208] Cloud-based SLAM (0% local). At the other extreme, it is possible to perform all SLAM computations in the cloud (i.e., edge, fog, cloud resources as described above, for brevity, in this section we just refer to "the cloud"), where the device only provides images and communication channels to the cloud computing resources. With such a thin client configuration, boxes 2, 3, and 4 can be performed in the cloud. Box 1 (image formation and capture) is still performed locally. In order to make such a system real-time (i.e., 30+ frames per second, "fps"), we may need to encode the image quickly and send it to the remote computing resource. The time required to encode the image into a suitable payload and send the image over the network must be minimized.
[0209] Comparison and bandwidth requirements. In each of the configurations discussed above that involve a cloud component, some information from the local device typically has to be sent to the cloud. In the case of cloud SLAM, one has to encode the image and send it to the SLAM system in the cloud. In the case of local super point processing, one may need to send points and descriptors to the cloud. In the case of a hybrid system with cloud map building and on-device localization, one may need to send points, descriptors, and estimated poses to the cloud, but it does not need to be done at 30+ fps. If map management happens in the cloud, but there is a local localization module, then information about the current map may be sent from the cloud to the client periodically.
[0210]
[0211] Table 1. Local vs. Cloud SLAM Computation Resource Allocation. Map* indicates that updated maps do not have to be sent from the cloud very quickly.
[0212] Assistance from other sensors and computing resources. The output of other sensors (such as IMUs and depth sensors) can be used to supplement super points, which typically only process raw images (color or grayscale). These additional sensors must usually be located on the same device as the physical image sensor. One can call these extra bits of information auxiliary inputs because they are not a hard requirement of our working method. From a computing resource perspective, one can also add additional computing units (such as more CPUs or more GPUs). For example, the additional computing resources may be placed away from the local device or right in a cloud data center. In this way, when the load is low, the additional computing resources can be used for other tasks.
[0213] Superpoints on head-mounted displays, smartphones, and other clients. The subject superpoint-based SLAM framework is designed to work across a wide spectrum of devices, which we call clients. At one extreme, a client can be a barebones image sensor with a Wi-Fi module and just enough computation to encode and send images over the network. At the other extreme, as described above, a client can include multiple cameras, a head-mounted display, additional locally coupled computational resources such as edge annotation, and so on.
[0214] Image-based localization and relocalization over time. By focusing on machine learning-based visual information extraction and summarization, the subject method is designed to be more robust to lighting changes and environmental changes that typically span 1-2 days in any given environment. This facilitates SLAM sessions lasting multiple days. Using classic image feature extraction procedures, only a small subset of the extracted 2D features are matchable in the tracking scene. Due to the extensive use of RANSAC and other outlier rejection mechanisms in traditional SLAM methods, these 2D features are typically not very robust to the task of relocalization over large temporal variations.
[0215] Targeting across devices and retargeting across time. By focusing on imagery, it may be easier to build a map using one client and then use that map within another.
[0216] Multi-user positioning and mapping. By performing cloud update operations in the cloud, it may be relatively easy to enable multiple clients to share and update a single 3D map.
[0217] The foregoing configuration facilitates the development and use of spatial computing systems with highly distributed resource pools, such as those described below and above with reference to various edge, fog, and cloud resource integrations.
[0218] refer to Figure 39 With the advent of deep learning configurations, increased bandwidth communications, and advanced communications chipsets, it is possible to configure systems with low latency and relatively high computational power. Described below are further details about various configurations where certain services are moved from local computational hardware on the user to a more distributed location. This effectively transforms the XR device associated with the user into a thin-client configuration, where the local hardware is relaying computational information rather than doing all the "heavy lifting" itself.
[0219] Thus, in these related embodiments, which may also be referred to as variations on the theme of "adaptive neural computing," one can draw from many of the same services and sources as in a less connected configuration, and the challenge becomes more focused on connectivity rather than locally hosting all the necessary computer hardware. Leveraging the speeds and bandwidth achievable with IEEE 802.11ax / ay (i.e., Wi-Fi 6) and 5G, one can rethink the way various tasks related to spatial computing can be performed.
[0220] One of the challenges with adaptive neural computing configurations in general is relocating the computational load of the services required to operate XR devices. As described above, in various embodiments, this can be facilitated by creating edge node computing devices and sensors, as well as optimized fog nodes and cloud computing infrastructure.
[0221] Adaptive neural computing edge nodes can be integrated IoT-type devices placed at the point of need or computing "edge" (i.e., easily connected using conventional network infrastructure). Such devices may include small computers capable of high-bandwidth connections, high-speed memory, GPUs, CPUs, and camera interfaces. As described above, suitable edge nodes include, but are not limited to, those sold by Nvidia (TM) and Intel (TM).
[0222] refer to Figure 40 Suitable edge nodes may also include sensors capable of creating a depth map of the world. For example, this may include a pair of stereo cameras, an RGB-D camera, a LiDAR device, and / or a structured light sensor, each of which may also include an IMU, a microphone array, and / or a speaker. Preferred edge nodes may also serve as Wi-Fi or 5G antennas.
[0223] Such edge computing devices can be configured as master devices for low-latency operation to facilitate integration with spatial computing systems. For example, Figures 25A to 25C An embodiment of progressive enhancement of capabilities for users to interface with data and the digital world is shown.
[0224] The integrated edge nodes can be configured to scale from portable and small to full on-premises server capabilities. As mentioned above, the purpose of using this distributed computing across the cloud system is to perform low-latency operations at the point of need, not on a small computer tethered to the wearable device.
[0225] The world reconstruction (i.e., meshing, SLAM) described above refers to an MMO-like configuration, where all or most of the computational resources are on the user's person, using a continuous operation configuration based on local services. One can use an "absolute" coordinate system as the world model (the geometry of the digital twin). This model can be fed through a series of transformations until it is placed on the physical environment in a true-to-scale configuration. These transformations are essentially calibration files that take into account factors intrinsic and extrinsic to the user's system, the type of device the user is using, or the user's current state.
[0226] refer to Figure 41 , through a “traversable world” system where each online XR device creates part of an aggregate model for the environment, additional users depend on scans of other devices to create such an aggregate. In contrast, with an absolute world model, scans can be captured or created in advance of runtime use by a specific user, so that when a user with an XR device enters a given room, there is at least a baseline model of such a room that can be used, while new / additional scans can be performed to improve such a model. The so-called “baseline” model becomes raw data that can be contextualized so that when a user’s XR device enters that space, there is less latency and the experience is more natural (i.e., the user’s XR device already knows that something is a “chair” or is able to quickly determine this based on available information).
[0227] Once we have the world model in the authoritative server, we can upload and download it from the digital twin server to the edge / fog nodes in the local environment on a schedule determined to be most useful. This ensures that the devices are as up-to-date as required by the specific application.
[0228] Once a coordinate system is established, it can be populated with data. There are various ways to achieve this. For example, data from different users' XR devices can be pipelined or included in the local storage of XR devices that are new to the locale. In another embodiment, various technologies and hardware (such as robots, sensor-integrated backpacks, etc.) can be used to provide periodically updated meshing / scanning of various environments. What these data sources do is capture data from the world and, using accessible world object identifiers, temporal features, contrast features, and supervised feature definitions, we can align maps, and through sensor fusion techniques such as the Extended Kalman-Bucy Filter (EKF) or Feedback Particle Filter (FPF), we can perform continuous-time nonlinear filtering and align these maps at runtime if necessary. This combined map may become the basis for how to perform tasks in spatial computation. The main goal in this embodiment is efficiency, which means that we do not need to constantly rebuild the world, we only need to capture differences or deltas in the mapped world. In other words, if the mesh of a particular environment is stable, there is no need to replace the entire mesh, only to add deltas or changes.
[0229] With all the raw map data in authoritative grid storage (such as a cloud server), it’s now possible to take that data and create intelligently smoothed and textured views of the world that can be expressed in much greater detail. This is because you might want to use that map not only to align digital content with it, but also to recreate the world for virtual reality through XR technology. Previously, you could map the world up and down to pre-identify key features in the raw geometry before consuming the data.
[0230] refer to Figure 42 , one may need to use raw grid data such as those from outdoor sources (e.g. Figure 42 The digital twin of the world (shown on the far left) that recreates the world may utilize many methods and combine them into a mesh database. For example, methods for recreating the external world may include stereo camera reconstruction, structure from motion, multi-view stereo reconstruction, image-based modeling, LiDar-based modeling, inverse process modeling, facade decomposition, facade modeling, ground reconstruction, aerial reconstruction, and a large number of model reconstructions. These methods can all contribute information to the mesh database. The raw information can then be contextualized. One way to do this is called "brand recognition" through multiple layers of language processing, RNNs, CNNs, supervised learning methods, and other methods. One can initially contextualize the information through many layers, first, one can use techniques such as "edge shrinkage" to identify and simplify the point cloud and geometric information. For example, one can use a persistent homology framework to analyze these changes. Refer again to Figure 42, one can identify various methods for taking raw point clouds or geometric information and contextualizing it for consumption by connected XR users, thinking that one goal of such a configuration is to unify users and the physical world into an ever-growing and enriching digital world. Employing such a process and configuration makes the world contextual, providing the ability to scale to more algorithms and models. Reference Figure 43 , one can take this authoritative queryable mesh and extract features that can be identified as permanent, semi-permanent, and temporary. We can determine this using learning algorithms, such as automatic parametric model reconstruction from point clouds. This is another technique that can be used to fill in gaps or scanning artifacts in a room interior and create a 3D reconstruction of the interior room using common features such as walls and doorways.
[0231] The next level might be to identify semi-permanent structures or objects in the room. One implementation utilizes "hierarchical fine-grained segmentation" (i.e., segmenting into smaller and smaller features) to achieve this, which involves semantically labeling features on an object until the features are no longer distinguishable from each other. An example of this would be to continually segment and then label segments of a chair.
[0232] After understanding the objects in the room, semi-permanent objects can be identified as those that can move but will likely move infrequently and / or require significant effort to move, such as a large dining table. One can then identify dynamically moving objects, such as a laptop or clouds in the sky. This can be achieved, for example, through object-level scene reconstruction using active object analysis. By continuously scanning the room and looking for dynamic objects, the system can be configured to increase the probability of proper segmentation.
[0233] Return Reference Figure 43 , all this algorithmic rigor facilitates filtering of spatial information so that users don’t need to use it unless it is important, or services that enable XR experiences require less computation because they only track clusters of points representing segmented objects.
[0234] Once the room is segmented, labeled, and contextualized, people may turn to connecting the room to digital “outlets” - a metaphor that can be used to form an intersection between the digital world and the physical world. A digital outlet can be an object, device, robot, or anything that plugs the digital world into the physical world. The intention to interact with the underlying digital world and how it affects the physical world can be communicated through these outlets, such as through metadata to each specific outlet. For example, if the physical world is too dark, the user interface of the XR device using such an outlet can be paired with an Internet of Things (“IoT”) controller that is integrated with an application and user interface (“UI”) that changes the light settings of the device. As another example, when a particular remotely located XR device user is considering which port or outlet to connect to when he wants to view a particular building, he can select an outlet that, based on metadata, will take him virtually directly into a very gridded room with a large amount of pre-existing full-color data about all the features of the room. This interaction may require many processes, and such processes are combined into tools that users can use to change their understanding of the perceived world.
[0235] It is now possible to have a fully contextualized and robust digital twin of the physical world. Figure 44 , for many XR devices (such as Figure 1 One of the most important processes for runtime comfort in systems that feature head-wearable display components is relatively low-latency determination of head pose or whatever relative to the real or virtual world of the system in question (i.e., not all systems have head-wearable XR configurations for users). Figure 44 As shown, head poses can occur on devices with limited computational resources, but because the relevant space is already predefined or prepopulated in terms of world reconstruction, one can use identified permanent, semi-permanent, and temporary features and identify anchor points within that classification, such as superpoints, as described above. Superpoints are computational graphs that can be used to describe mathematical relationships between specific regions whose characteristics are related to other features that increase the probability of identification of those regions. Using superpoints to lock poses is different from conventional pose determination methods. Again referring to the head pose with head wearable XR user configuration Figure 44 , the system can be configured to optimize the features tracked for head pose in 6DOF so that errors propagated from other spaces or residual fitting errors to the world map are spatially filtered out of the equation. The system can be configured to perform pose calculations on a local processor (such as a custom ASIC) packaged with the sensor (i.e., camera).
[0236] Pose-related data can be streamed from the user's XR device or from other XR devices to available edge nodes, fog nodes, or the cloud. As mentioned above, such distribution can be configured to require only a small amount of information, and when the computational load is low, the device can stream images back to the edge node to ensure that the positioning has not drifted.
[0237] exist Figure 44 In
[15] , a classic implementation of posture determination is shown. The user's XR device, which can be configured to continuously map the environment, builds or assists in building a world map. Such a world map is the content of the posture reference. Figure 44 A one-dimensional line representing the sparse points found during the meshing process is shown. This line is then fitted with the model to simplify computations, and this allows the residuals of the fit and error propagation to play an important role in the registration. Since the digital object collides with the physical world, these errors can lead to misregistration, scale ambiguity, and jitter that destroy the interaction paradigm. The fitted model can now be decomposed into reasonable fragments for transmission and memory management based on the local scans from the device in each space. Due to the aforementioned residual errors, these decomposed models are by definition misregistered. The amount of misregistration depends on the global error, as well as the quality of the scans that acquired the geometry and point cloud information.
[0238] Figure 44 The second row in
[15] shows how anchor points or super points are defined during the world reconstruction phase. These anchor points become the positioning tool used to locate all devices across a range of XR devices to an absolute coordinate system. Figure 45 An embodiment for pose estimation using anchor maps is shown.
[0239] The next element in this spatial computing implementation is the route taken to render objects to the XR device. Since these devices can have a variety of operating systems, one can employ streaming or edge rendering protocols configured to exploit the spatial nature of the digital and physical worlds, with distributed computing resources and varying levels of latency, as described above. That said, based on the gesture approach described above, the device's location in the room is generally known. Thanks to the reconstructed world, the room is also known.
[0240] Edge rendering can be facilitated by modern connectivity availability and should become increasingly common as more XR systems become used. Remote rendering can be implemented through many variations of the classic rendering pipeline, but will generally provide more latency than is acceptable in most spatial computing configurations to efficiently deliver inferred data.
[0241] One approach for implementing remote rendering for spatial computing systems is to leverage conventional rendering and streaming technology used by companies that stream movies or TV shows, such as Netflix™. In other words, the system can be used to implement the current rendering pipeline on distributed edge / fog / cloud resources, and then take that output, convert the 3D content into 2D video, and stream it to the device.
[0242] Figure 46 Another embodiment is shown where one can leverage the adaptive compute power of edge / fog / cloud resources to render parallel frames as predictions, and select the frame that is closest to the actual value at the last moment. Due to the nature of spatial computing (i.e., the proximity of the pose to the rendered content can be detected), the system can be configured so that for a given number of frames ahead of real time (such as, for example, four frames), the system can render multiple copies based on head pose measurements and content placement in the absolute world. This process can be repeated for a given number of frames, and before rendering directly onto the device, the last pose value can be taken from the device, and its associated frame can be sent to the device to best match the model. This allows the system to have a polynomial prediction for rendering frames into the future that the XR device is predicted to be posed or "looking at."
[0243] As mentioned above, for the configurations described herein, one may take advantage of the latest developments in wireless connectivity, including but not limited to WiFi-6, which may also be referred to as the compatible IEEE 802.11ax standard; or any successor that will be able to efficiently send signals to the device.
[0244] One overall pattern for rendering on a specific XR device may include computing all relevant processes in a distributed cloud, streaming the results to the XR device—directly to the device’s frame buffer—and producing an image for the user on a display.
[0245] The aforementioned superpoint technique can be used to aggregate meshes or form geometric relationships between them.
[0246] The physical world is not only an interactive element in spatial computing - it is also one of the primary inputs to computing architectures. In order for this data to be streamlined, we need to understand the environment in which the user or experience (location-based experience LBE) is located.
[0247] Figure 46The content delivery system further includes: a rendering function executable by a mobile device processor to connect the mobile device to a plurality of resource devices, send one or more rendering requests, wherein each resource device receives a corresponding rendering request, receives a rendering from each of the remote devices based on the corresponding rendering request; compares the renderings to determine a preferred rendering; and, under control of the mobile device processor, selects the preferred rendering of a first content transmitted by a first resource device transmitter using a mobile device communication interface. The rendering forms a system with a polynomial prediction for rendering frames into the future that the mobile device is predicted to be placed or viewed.
[0248] Figure 47 The steps required to simplify the physical world are shown, where we use the framework described above for different processes in spatial computing at different levels. The physical world is complex and random, so a dynamic framework is required to create accurate and precise reconstructions of the world that are simplified and infer meaning from the grid abstraction layer (498). To map the physical world for this process, people generally prefer hardware that recreates the geometry of the world (500). Previous work on this has been shown to be effective, but fully convolutional neural networks have been shown to increase the accuracy of this and reduce the computational load. The hardware (502) that people can use to geometrically recreate the physical world can include RGB cameras, RGB-D cameras, thermal cameras, short wave infrared (SWIR) cameras, medium wave infrared (MWIR) cameras, laser scanning devices, structured light modules, arrays and combinations of these, as well as process reproduction and completion. These images, point clouds and geometric representations of the world can be captured with many different devices for many different areas and environments. They can be stored in a variety of modern image formats, at full image resolution, sub-sampled or super-sampled.
[0249] Once one has the raw information from one or more sensors, at runtime, or saved and then processed, or some combination thereof, one can find points of interest. As described above, one can use superpoint techniques (504) that employ self-supervised interest point detectors and descriptors, allowing features of interest to be identified.
[0250] In order to optimize the ingestion or processing of data from multiple sources into a single grid (506) that is a digital representation of the world, one may need to perform several backend processes (512, 514, 516, 518). One can be used to add any new regions of the world to what can be called an "Authoritative Intelligent Grid Server" (or "AIMS"), which can also be called a "spatial atlas." If the data does exist, one can then perform sensor fusion to merge the information into a single grid. This type of sensor fusion can be done with traditional methods, such as various variations of the Kalman filter, or we can use deep learning techniques to use the superpoint features of interest found and perform feature-level fusion for each sensor type and format. By using a superpoint fully convolutional neural network structure, we can create a synthetic dataset with each type of sensor and then use this dataset to train each of the separate CNNs. Once the superpoint algorithm has been tuned for each sensor type, feature-level fusion can be performed following the general pattern of the image below, where one has implemented the neural network for a specific pattern.
[0251] Once the unified mesh is computed, we are then faced with the challenge of what to do with the mesh. In one framework, one can seek further contextualization, in another framework, the system can be configured to create a smart mesh (507). A superpoint smart mesh can be employed using a superpoint algorithm and utilizing the implemented homomorphic adaptation to create smart interpolations of the main mesh and add data points by enabling additional features with greater probability. To achieve this, one can follow the same process as for the superpoint feature detector, thereby increasing the training set to include the 3D mesh of the 2D shapes in the original superpoints. The reason this may be needed is that after the mesh is unified, the same superpoints may not be completely consistent and one may want all features in the frame to be extracted as the system will use them to overlay (508) a texture on the wireframe. Homographic adaptation of the feature plane may produce a set of superpoints and since the entire map is in 3D space, one can also rotate the user's perspective about the superpoint along the arc created by the depth from the user to that point to create more perspectives of the features identified by the superpoint. We will then use all the points we created to attach a texture (510) to the surface of the 3D map and pseudo-depth image.
[0252] refer to Figures 48 to 66 , showing various exemplary embodiments of various XR devices used in various scenarios. For example, referring to Figure 48 , shows an office environment where spatial computing is used to aggregate company information for collaboration, scheduling, and production by various users with various XR devices.
[0253] refer to Figure 49, the XR user is shown interrupted by a remotely located physician, who is alerted by the XR user's integrated health sensing capabilities, such that the remote physician can inform the local XR user that he appears to be experiencing relatively high cardiovascular stress levels.
[0254] refer to Figure 50 , showing XR users working with smart appliances and wearable systems integrated into a spatial computing paradigm, making it possible to perform analysis on the user's health and fitness routine so that improvement suggestions can be made based on such data.
[0255] refer to Figure 51 , shows an illustrated outdoor living scene, pulling data from multiple sources together so that many users can experience the same scene physically or virtually.
[0256] refer to Figure 52 , so-called “lifestream” data about a particular user identifies the user’s hypoglycemia based on current and previous activity, including caloric intake, sleep, physical activity, and other factors.
[0257] refer to Figure 53 , showing various XR users through their virtual presence in a common physical space, where spatial computing integration enables them to manipulate physical objects, such as chess pieces, in the common physical space around them.
[0258] refer to Figure 54 , shows a shopping scenario that enables users of XR devices to compare clothes they have at home or in some other location with clothes they can visualize in a physical store, and simultaneously virtually “try on” the clothing items in the store as their virtual self in the physical store.
[0259] refer to Figure 55 In an example, an XR user is sitting down to play an actual physical piano and asks for instructions on how to play a particular sonata. Information about the performance of the sonata is mapped to the appropriate keys of the physical piano to assist the user, and the user is even advised on how to improve his or her performance.
[0260] refer to Figure 56 , shows a collaborative scenario in which five XR users are able to visualize a virtual home model retrieved from a remote company database, they can change aspects of the home model at runtime, and save these changes on a remote computing resource for future use.
[0261] refer to Figure 57In an example, some XR users are shown trying to find the nearest physical public restroom. Regression analysis of their current location information helps the system determine the location of the nearest available restroom, and the system can then be configured to provide virtual visual guidance to the user as they walk to that restroom.
[0262] refer to Figure 58 , shows two XR device users in a prenatal yoga instructor scenario, where the instructor, once granted access by the client, can monitor the mother and child’s vital signs for safety.
[0263] refer to Figure 59 , connected spatial computing systems can, with appropriate permissions, leverage the data of many connected XR users to assist these users in their daily interactions by leveraging not only their own data but the aggregation of all their data, for example, for traffic and congestion analysis.
[0264] refer to Figure 60 , shows an XR user with mobility challenges speaking with her XR device, where her XR device is configured to respond to voice commands and queries so that it can return information about the query and also preferably assist in controlling and even navigating her mobility assistance device. In other words, she can ask the spatial computing system to take her to the restroom in an autonomous or semi-autonomous manner while also preferably avoiding not only structures but also traffic, such as other people’s movements.
[0265] refer to Figure 61 In an example, an XR user visits their primary care physician and discusses and leverages aggregated data from the last visit to simulate and provide visualizations of different courses of action for the patient. For example, the physician can assist the XR user in understanding what he or she will look like after knee replacement surgery.
[0266] refer to Figure 62 , the user is shown speaking through a verbal interface (i.e., not necessarily a visual XR system, but simply an audio-based spatial computing system), he asks the system details about the specific wine he is drinking, and receives response information based not only on cloud-based internet format information but also on his own specific metadata (such as his own wine interests).
[0267] refer to Figure 63 By utilizing GPS positioning devices (i.e., small GPS transmitters that are typically battery powered and portable and can be coupled to clothing, backpacks, shoes, etc.), spatial computing systems can be used to track other people, such as children or the elderly, in real time or near real time.
[0268] refer to Figure 64, multiple XR users are shown participating in a virtual volleyball game such that each user, while viewing through their XR device, must intersect with a virtual ball and exert virtual forces on the ball based on the physical posture changes and rates of change of their XR device (i.e., they can use their smartphone or tablet with their racket).
[0269] refer to Figure 65 , two XR users are shown sitting outside a sculpture garden and wishing to know more about a particular sculpture, they can use the visual search tool to gather and examine information on the internet using the visualization tools of their XR devices, and can share and / or save their findings.
[0270] refer to Figure 66 , the system can be configured to have an optical system that uses various optical relays (i.e. waveguides, "bird bath" type optics, volume phase holograms, etc.) to take time-multiplexed images that are stitched together (i.e. aggregated) and displayed to the user in one configuration. This same system can be used to relay eye information back to the camera, where in some instances one can view the gaze vector, fovea, in a manner that is not exclusively dependent on the diameter of the user's pupil. One can create one or more red, green, and blue light sources and reflect, refract, diffract off / through a phase modulation device such as liquid crystal on silicon ("LCOS", forming a spatial light modulator) or a 2D or 3D light valve. Each of these laser + modulator pairs can be coupled to a scanning or beamforming array of mirrors where the range of the field of view can be managed by the angle of entry (i.e. angle of incidence) of the system. One embodiment can use this information to time-multiplex the full field of view at a range of frame rates to include a field of view that creates a comfortable viewing experience. The time-multiplexed images have the ability to be presented to the user at different depths (ie, perceived focal planes or focal depths).
[0271] Various example embodiments of the present invention are described herein. Reference is made to these examples in a non-restrictive sense. They are provided to illustrate the wider applicable aspects of the present invention. Various changes can be made to the described invention and equivalents can be substituted without departing from the true spirit and scope of the present invention. In addition, many modifications can be made to adapt specific circumstances, materials, compositions of matter, processes, (one or more) process actions or (one or more) steps to the (one or more) purposes, spirit or scope of the present invention. Moreover, as will be understood by those skilled in the art, each individual variation described and shown herein has discrete components and features that can be easily separated or combined with the features of any of the other several embodiments without departing from the scope or spirit of the present disclosure. All such modifications are intended to be within the scope of the claims associated with the present disclosure.
[0272] The present invention includes methods that can be performed using the subject devices. The methods can include the act of providing such suitable devices. Such provisioning can be performed by the end user. In other words, the act of "providing" simply requires the end user to obtain, access, access, locate, configure, activate, power on, or otherwise provide the necessary devices for the method. The methods described herein can be performed in any logically possible order of the described events, and in the order in which the events are described.
[0273] The exemplary aspects of the present invention have been described above, along with details regarding material selection and manufacturing. As for other details of the present invention, these can be understood in conjunction with the patents and publications mentioned above and what is generally known or understood by those skilled in the art. The same applies to the method-based aspects of the present invention, as long as additional actions are used as is commonly or logically possible.
[0274] In addition, although the present invention has been described with reference to several examples that optionally include various features, the present invention will not be limited to the invention described or indicated as expected relative to each variant of the present invention. Various changes can be made to the described invention and equivalents (whether recorded in this article or not included for the sake of some brevity) can be substituted without departing from the true spirit and scope of the present invention. In addition, in the case of providing a range of values, it should be understood that each intermediate value between the upper and lower limits of the range and any other claimed value or intermediate value in the claimed range are encompassed in the present invention.
[0275] Moreover, it should be expected that any optional feature of the invention variant described can be set forth and claimed independently or in combination with any one or more of the features described herein. Reference to a singular item includes the possibility of having multiple identical items. More particularly, as used herein and in the claims associated therewith, unless otherwise specifically stated, the singular forms "a / an", "said" and "the" include plural indicators. In other words, the use of articles allows "at least one" in the subject item in the above description and the claims associated with this disclosure. It should also be noted that such a claim can be written to exclude any optional element. In this way, the statement is intended to be used as a precedent basis for using special terms such as "only", "only" etc. or using "negative" restrictions in conjunction with the recording of the claim elements.
[0276] In the absence of use of such specific terms, the term "comprising" in the claims associated with the present disclosure should allow for the inclusion of any additional elements—regardless of whether a given number of elements are recited in such claims, or whether the addition of features can be considered to transform the nature of the elements recited in such claims. Except as specifically defined herein, all technical and scientific terms used herein are to be given the broadest commonly understood meaning possible while maintaining claim validity.
[0277] The breadth of the present invention should not be limited by the examples provided and / or this description, but instead should be defined only by the scope of the claims language associated with this disclosure.
[0278] While certain exemplary embodiments have been described and shown in the drawings, it should be understood that these embodiments are merely illustrative and not restrictive of the invention, and that the invention is not limited to the specific construction and arrangements shown and described, as modifications may be made by those skilled in the art.
Claims
1. A content providing system, comprising: A first resource device at a first location having: a first resource device processor; a first resource device storage medium; a first resource device data set comprising first content on a storage medium of the first resource device; as well as a first resource device communications interface forming part of the first resource device and connected to the first resource device processor and under the control of the first resource device processor; A mobile device having: Mobile device processors; a mobile device receiver coupled to the mobile device processor; and a first resource device communication interface of the first resource device being at a first location and receiving, under control of the mobile device processor, first content sent by the first resource device communication interface, such that the mobile device communication interface creates a first connection with the first resource device, wherein the first content is specific to a first geographic parameter of the first connection; a mobile device output device connected to the mobile device processor and capable of providing an output capable of being sensed by a user under control of the mobile device processor; The second resource device has: a second resource device processor; a second resource device storage medium; a second resource device data set comprising second content on a storage medium of the second resource device; and a second resource device communication interface forming part of the second resource device and connected to the second resource device processor and under the control of the second resource device processor, wherein the second resource device is at a second location, wherein the mobile device communication interface establishes a second connection with the second resource device, and The second content is specific to a second geographic parameter of the second connection, and the second geographic parameter is different from the first geographic parameter to distinguish the second connection from the first connection within the system of distributed computing resources, and the second content is distinguished from the first content due to the distinction between the second connection and the first connection.
2. The content providing system according to claim 1, wherein: The mobile device includes a head-mounted viewing assembly coupleable to a head of the user, and the first and second content provide the user with at least one of additional content, augmented content, and information regarding a particular view of the world as seen by the user.
3. The content providing system according to claim 1, further comprising: A location island for the user to enter, wherein specific features have been preconfigured to be located and interpreted by the mobile device to determine the geographic parameter relative to the world around the user.
4. The content providing system according to claim 3, wherein: The specific feature is a visually detectable feature.
5. The content providing system according to claim 3, wherein: The specific feature is a wireless connection related feature.
6. The content providing system according to claim 2, further comprising: A plurality of sensors are connected to the head mounted viewing assembly, the plurality of sensors being used by the mobile device to determine the geographic parameters relative to the world around the user.
7. The content providing system according to claim 3, further comprising: A user interface is configured to allow the user to at least one of: ingest, utilize, view, and bypass certain information of the first or second content.
8. The content providing system according to claim 1, wherein: The connection is a wireless connection.
9. The content providing system according to claim 1, wherein: in, The mobile device has a sensor that detects a first feature at the first location, and the first feature is used to determine a first geographic parameter associated with the first feature, and wherein the first content is specific to the first geographic parameter.
10. The content providing system according to claim 9, wherein: The second resource device is at a second location, wherein the mobile device has a sensor that detects a second feature at the second location, and the second feature is used to determine a second geographic parameter associated with the second feature, and wherein the first content is updated with second content specific to the second geographic parameter.
11. The content providing system according to claim 10, wherein: The mobile device includes a head-mounted viewing assembly coupleable to a head of the user, and the first and second content provide the user with at least one of additional content, augmented content, and information regarding a particular view of the world as seen by the user.
12. The content providing system according to claim 1, further comprising: A spatial computing layer, which is between the mobile device and the resource layer having multiple data sources, and is programmed to: Receive data resources; integrating the data resources to determine an integration profile; as well as The first content is determined based on the integration profile.
13. The content providing system according to claim 12, wherein: The spatial computing layer includes: A spatial computing resource device having: Spatial computing resource device processor; Spatial computing resource device storage medium; and A spatial computing resource device data set, which is on the spatial computing resource device storage medium and is executable by the processor to: receiving the data resource; integrating the data resources to determine an integration profile; and The first content is determined based on the integration profile.
14. The content providing system according to claim 12, further comprising: An abstraction and arbitration layer, which is inserted between the mobile device and the resource layer and is programmed to: Make workload decisions; as well as Tasks are distributed based on the workload decision.
15. The content providing system according to claim 14, further comprising: A camera device that captures images of the physical world surrounding the mobile device, wherein the images are used to make the workload decision.
16. The content providing system according to claim 12, further comprising: A camera device that captures images of the physical world surrounding the mobile device, wherein the images form one of the data resources.
17. The content providing system according to claim 1, wherein: The first resource device is an edge resource device having a first delay, and The mobile device communication interface includes one or more mobile device receivers connected to the mobile device processor and connected to a second resource device communication interface in parallel with the connection to the first resource device to receive the second content.
18. The content providing system according to claim 17, wherein: The second resource device is a fog resource device having a second delay that is slower than the first delay.
19. The content providing system according to claim 18, wherein: The mobile device communication interface includes one or more mobile device receivers connected to the mobile device processor and connected to a third resource device communication interface in parallel with the connection of the second resource device to receive third content sent by the third resource device transmitter, wherein the third resource device is a cloud resource device having a third latency that is slower than the second latency.
20. The content providing system according to claim 18, wherein: The connection to the edge resource device is through a cellular tower, and the connection to the fog resource device is through a Wi-Fi connected device.
21. The content providing system according to claim 20, wherein: The cellular tower is connected to the fog resource device.
22. The content providing system according to claim 20, wherein: The Wi-Fi connection device is connected to the fog resource device.
23. The content providing system according to claim 18, further comprising: At least one camera that captures at least a first image and a second image, wherein the mobile device processor sends the first image to the edge resource device for faster processing and sends the second image to the fog resource device for slower processing.
24. The content providing system according to claim 23, wherein: The at least one camera is a room camera that captures the first image of the user.
25. The content providing system according to claim 17, further comprising: a sensor that provides sensor input to a processor; a pose estimator executable by a processor to calculate a pose of the mobile device based on the sensor input, including at least one of a position and an orientation of the mobile device, a steerable wireless connector that creates a steerable wireless connection between the mobile device and the edge resource device; as well as A steering system is coupled to the pose estimator and has an output for providing input to the steerable wireless connection to steer the steerable wireless connection to at least improve the connection.
26. The content providing system according to claim 25, wherein: The steerable wireless connector is a phased array antenna.
27. The content providing system according to claim 25, wherein: The steerable wireless connector is a radar hologram type transmission connector.
28. The content providing system according to claim 17, further comprising: An arbiter function, which can be performed by the processor to: determining how much edge and fog resources are available via the edge and fog resource devices, respectively; Sending processing tasks to the edge and fog resources based on the determination of available resources; as well as Receive results returned from the edge and fog resources.
29. The content providing system according to claim 28, wherein: The arbiter function is executable by the processor to: Combine the results from the edge and fog resources.
30. The content providing system according to claim 28, further comprising: A runtime controller function executable by the processor to: Determine if the process is a runtime process; If a determination is made that the task is a runtime process, immediately executing the task without utilizing the arbiter function to make the determination; and If a determination is made that the task is not a runtime process, then the determination is made using the arbiter function.
31. The content providing system according to claim 18, further comprising: a plurality of edge resource devices, data being exchanged between the plurality of edge resource devices and the fog resource device, the data comprising spatial points captured by different sensors and sent to the edge resource devices; as well as A super point calculation function is executable by the processor to determine a super point, where the super point is a point selected from two or more points where data from the edge resource devices overlap.
32. The content providing system according to claim 31, further comprising: A plurality of mobile devices, wherein each super point is used in each mobile device for location, orientation, or pose estimation of the corresponding mobile device.
33. The content providing system according to claim 32, further comprising: A context trigger function is executable at a processor to generate a context trigger for a set of said super points and store said context trigger on a computer readable medium.
34. The content providing system according to claim 33, further comprising: A rendering engine executable by the mobile device processor, wherein the contextual trigger serves as a handle for rendering an object based on the first content.
35. The content providing system according to claim 1, further comprising: Rendering functionality, executable by the mobile device processor to: connecting the mobile device to a plurality of resource devices; Sending one or more rendering requests, wherein each resource device receives a corresponding rendering request; receiving a rendering from each of the remote devices based on the corresponding rendering request; comparing the renderings to determine a preferred rendering; and The preferred rendering is selected, using the mobile device processor, as first content sent by the first resource device transmitter.
36. The content providing system according to claim 35, wherein: The rendering forms a system with polynomial predictions for rendering frames into the future predicted to be where the mobile device will be placed or viewed.
Citation Information
Patent Citations
Systems and methods for augmented and virtual reality
US10262462B2
Augmented reality systems and methods with variable focus lens elements
US10459231B2
Virtual and augmented reality systems and methods
US20150205126A1
System and method for augmented and virtual reality
US9215293B2
Planar waveguide apparatus with diffraction element(s) and system employing same
US9671566B2