Large geometric model: physics-supervised spatiotemporal foundational mode for gps-denied autonomous systems
The physics-supervised LGM addresses the lack of universal robotic architectures by inferring and maintaining environmental geometry using physics-based data, enhancing cross-platform autonomy and interoperability in sensor-degraded conditions.
Patent Information
- Authority / Receiving Office
- GB · GB
- Patent Type
- Applications
- Current Assignee / Owner
- AARON JIN KIAT TAN
- Filing Date
- 2025-11-24
- Publication Date
- 2026-07-22
AI Technical Summary
Existing robotic systems lack a universal operating architecture that enables autonomous operation across diverse environments without relying on external positioning systems or continuous high-quality sensing, leading to fragmented software frameworks and high integration costs, and fail to maintain spatial coherence in sensor-degraded conditions.
A physics-supervised spatial foundational model (LGM) that infers and maintains an internal representation of environmental geometry using physics-based temporal data, independent of platform-specific coordinate systems or external signals, enabling coherent operation in GPS-denied and sensor-degraded environments.
The LGM provides a common spatial representation that supports cross-platform autonomy, reduces fragmentation, and enhances interoperability by maintaining spatial coherence and enabling reliable navigation and decision-making in diverse environments.
Abstract
Description
[0001] The present invention relates to spatial inference, localisation, and environmental modelling for autonomous or semi-autonomous systems. More particularly, the invention concerns machine-learning architectures configured to infer geometric structure and maintain spatial coherence using physics-supervised temporal data. The invention further relates to platform-agnostic spatial representations for use across heterogeneous robotic or autonomous systems. It is applicable to mobile robotic platforms, underwater or subterranean devices, wearable navigation equipment, and other systems operating in environments where conventional sensing or positional signals are degraded, unavailable, or unreliable. Background of the Invention
[0002] Robotic systems have historically been developed as task-specific and domainspecific platforms, each using bespoke hardware interfaces, sensing pipelines, localisation methods, and control architectures. Ground robots, aerial vehicles, marine systems, and industrial robots typically operate on incompatible software frameworks that do not generalise across embodiments or environments. This fragmentation has been tolerated because, for most of the history of robotics, such systems functioned under continuous or intermittent human supervision. In effect, the human operator supplied the unifying cognitive layer—perception, geometric reasoning, causal inference, and contextual judgement— allowing each specialised robotic platform to operate despite underlying architectural divergence.
[0003] As autonomy increases, the absence of a universal operating architecture becomes a fundamental limitation. Unlike digital computing, which consolidated around shared instruction sets and general-purpose operating systems, robotics has remained a collection of isolated vertical stacks. Manufacturers maintain separate software, mapping, and sensorintegration layers for each robot class, and sensor suppliers must support numerous incompatible interfaces. This results in duplicated engineering effort, poor portability of autonomy components, and high integration costs across the sector.
[0004] These limitations become critical as robots are deployed in environments where human supervision, GPS / GNSS, or clean sensor signals cannot be relied upon. Conventional localisation and mapping methods depend on external positioning systems or on continuous availability of cameras, LIDAR, radar, sonar, magnetometers, and inertial measurement units (IMUs). In degraded or cluttered settings—such as underground spaces, underwater environments, enclosed structures, smoke-filled areas, or regions with dust, silt, fog, or electromagnetic interference—these sensors may provide incomplete, distorted, or contradictory measurements. GPS / GNSS becomes unavailable or unreliable, optical sensing degrades, sonar and radar suffer multipath interference, and IMUs accumulate drift.
[0005] When external sensing degrades, contemporary autonomy stacks typically lose localisation or fail to maintain a coherent environmental map. Existing SLAM pipelines, sensor-fusion frameworks, and neural perception models treat geometry as a by-product of sensor readings rather than as an intrinsic, learned capability. As a result, autonomous systems lack a persistent internal representation of spatial structure that remains stable under partial sensor loss, intermittent data, or contradictory observations.
[0006] There is therefore a technological gap in providing autonomous robots with a unifying spatial foundation that is not dependent on absolute coordinates, external signals, or continuous high-quality sensing. No existing architecture enables robots across land, air, marine, subterranean, underwater, or space environments to infer and maintain geometric coherence using a shared, physics-grounded representation that generalises across embodiments.
[0007] Accordingly, there is a need for a foundational architecture that provides a universal internal representation of geometry, motion, and environmental structure, enabling autonomous systems to operate coherently in GPS-denied, sensor-degraded, and dynamically changing environments. Such an architecture would reduce fragmentation, support crossplatform autonomy, and allow sensors, autonomy modules, and multi-robot systems to interoperate through a common geometric and causal substrate. Summary of the Invention
[0008] The present invention provides a physics-supervised spatial foundational model, hereinafter referred to as the Large Geometric Model (LGM), configured to infer, maintain, and update an internal representation of environmental geometry even under conditions of degraded or unavailable sensor inputs. In addition to addressing the limitations of conventional localisation and mapping pipelines, the invention further provides a unifying geometric architecture capable of supporting autonomous operation across heterogeneous robotic platforms. The LGM supplies a common internal substrate for spatial reasoning that is independent of platform-specific coordinate systems, proprietary mapping stacks, or reliance on external positioning signals.
[0009] During training, the model receives time-varying physical measurements and motionstate signals from one or more agent platforms. In certain embodiments, such inputs may include inertial data, environmental interactions, energy-related variables, or other physics-associated telemetry. Absolute positional labels, such as GPS or GNSS coordinates, are intentionally omitted. As a result, the model is compelled to infer relative spatial relationships, environmental structure, and its own positional context using only the underlying physical regularities present in the input signals.
[0010] Through this physics-supervised training process, the LGM learns a persistent internal representation of geometry that captures environmental topology, spatial continuity, and physically plausible movement. The resultant model remains stable even when individual sensor modalities become unreliable, contradictory, occluded, or absent, thereby enabling autonomous agents to preserve spatial coherence in GPS-denied, cluttered, subterranean, underwater, or otherwise degraded environments. When deployed, the LGM produces a coherent spatial state that may serve as a substrate for localisation, mapping, hazard estimation, navigation, or higher-level decision-making.
[0011] In certain embodiments, multiple devices or platforms may contribute local geometric updates to a shared model through a federated learning process. Raw telemetry need not be exchanged; instead, learned parameters or gradient information may be aggregated in a bandwidth-efficient manner. As more devices operate across diverse environments—land, air, maritime, subterranean, underwater, or space—the shared foundation model may improve in robustness and generality without requiring centralised storage of raw sensor data.
[0012] By providing a common geometric latent space that reflects underlying physical structure, the invention enables disparate robotic and autonomous systems to operate using a unified spatial representation rather than bespoke, platform-specific architectural stacks. This unifying foundation may be integrated with downstream perception modules, navigation controllers, action-selection modules, sensor-fusion pipelines, or multi-robot coordination systems, thereby reducing fragmentation and enabling cross-platform interoperability. In certain embodiments, the LGM may be embedded within mobile robotic platforms, unmanned aerial vehicles, underwater or subterranean vehicles, wearable rescue equipment, or other autonomous or semi-autonomous systems requiring reliable spatial inference under degraded sensing conditions. Description of the Figures
[0013] Figure 1 illustrates an example training protocol for the Large Geometric Model.
[0014] Figure 2 illustrates an example a federated learning training protocol for the Large Geometric Model. Detailed Description of the Invention
[0015] The invention concerns a physics-supervised spatial foundational model (the “Large Geometric Model” or “LGM”) configured to infer, maintain, and update an internal representation of environmental geometry under varying sensing conditions. The LGM is implemented using one or more machine-learning architectures capable of processing multi modal temporal inputs and generating a persistent latent spatial state that reflects physically plausible environmental structure and agent movement.
[0016] In certain embodiments, the LGM further serves as a platform-independent geometric substrate capable of unifying spatial reasoning across heterogeneous robotic systems. Because the model infers geometry using physics-based temporal regularities rather than platform-specific coordinate frames or handcrafted mapping pipelines, the same foundational representation may be deployed across ground vehicles, aerial systems, marine robots, subterranean agents, humanoids, or space-based platforms. The LGM therefore provides a common spatial interface that abstracts away hardware differences and allows downstream autonomy modules to operate using a shared latent state, regardless of sensing configuration or embodiment.
[0017] This unifying geometric layer enables the LGM to function analogously to an operating-system level spatial foundation for autonomous systems. Existing robotics architectures require bespoke localisation stacks, mapping modules, and sensor-fusion pipelines for each platform and environment. By contrast, the LGM supplies a persistent, physics-grounded geometric representation that can be consumed by multiple subsystems— including perception, navigation, hazard detection, and motion control—without requiring duplicated engineering effort. This abstraction reduces fragmentation within robotic software ecosystems and permits cross-platform reuse of autonomy components.
[0018] In general, the LGM receives a sequence of time-varying physical measurements obtained from an agent or platform traversing an environment. Such measurements may include, but are not limited to, inertial data (e.g., accelerometer or gyroscope traces), environmental interaction signals, mechanical or energy-related telemetry, or other physical state variables associated with the agent’s motion. In certain embodiments, the LGM may additionally receive degraded or intermittent signals from optical, acoustic, electromagnetic, or other exteroceptive sensors. However, the presence of such sensors is not required for training or operation.
[0019] During training, the LGM is not provided with absolute positional labels such as GPS / GNSS coordinates, global reference frames, or handcrafted topological maps. Instead, the model is trained to infer relative spatial relationships using only the temporal evolution of physical measurements and the constraints imposed by the underlying physics of motion. This training methodology compels the model to learn geometric consistency, spatial continuity, and physically plausible environmental structure as emergent properties of its internal representation.
[0020] The LGM may be implemented using a recurrent neural architecture, a transformerbased temporal encoder, a latent-state spatial graph, a physics-informed neural network, or any suitable combination thereof. The architecture produces, updates, or maintains a latent spatial state vector or tensor that encodes inferred geometry, relative pose, topological relationships, or structural features of the surrounding environment. This latent representation evolves as new measurements are received, enabling the model to maintain a persistent geometric memory even in the absence of reliable external sensing.
[0021] In certain implementations, the training data supplied to the LGM naturally encodes cause-and-effect relationships arising from the agent’s physical interaction with its environment. As an agent moves, accelerates, decelerates, or changes orientation, these actions produce corresponding changes in inertial measurements, environmental constraints, and relative geometric structure. The sequential evolution of sensor inputs therefore reflects the causal effects of motion through space. By learning from these physics-governed state transitions—rather than from static positional labels—the LGM internalises the underlying causal regularities linking movement to geometric change. This enables the model to maintain spatial coherence, recognise physically impossible transitions, and reject sensor anomalies that conflict with learned causal structure.
[0022] In certain embodiments, the LGM may incorporate physics priors, inductive biases, or constraint-enforcing modules that ensure geometric plausibility. Such priors may include conservation of momentum, smoothness of motion, structural continuity, or energy-based constraints. These elements enable the LGM to reject inconsistent measurements, detect anomalies, and stabilise the inferred geometry when sensor quality degrades.
[0023] During deployment, the LGM generates one or more spatial inference outputs. Examples of such outputs include: (a) estimated relative position of the agent; (b) inferred environmental geometry or topology; (c) identification of geometric hazards or inconsistencies; (d) predictions of future spatial states; or (e) latent variables suitable for downstream navigation or control modules. The specific form of the output may vary depending on the application, but all outputs derive from the model’s internally maintained spatial representation.
[0024] In some implementations, the LGM may be executed on a device that includes a processor, memory, local sensors, and optional communication hardware. In these cases, a local inference controller may orchestrate the delivery of sensor inputs to the LGM, manage the update of the model’s latent state, or fuse the LGM outputs with other perception or control systems. The controller is not essential to the invention and may be omitted in embodiments where only the model’s internal representation is of interest.
[0025] In further embodiments, multiple devices, platforms, or agents may each maintain a local instance of the LGM. These instances may participate in a federated learning process in which parameter updates or gradient information are shared with a central or distributed aggregation node. Raw sensor data need not be transmitted; model updates may be exchanged in a privacy-preserving or bandwidth-efficient manner. Aggregation of such updates enables the global spatial foundational model to improve in resolution, generality, or robustness as additional devices operate in diverse environments.
[0026] The LGM may also support hybrid operational modes in which local inference is performed on-device while long-term model refinement occurs via periodic federated updates. In such scenarios, the learned spatial representation benefits from both individual device adaptation and cross-device generalisation without requiring continuous connectivity or centralised sensing infrastructure.
[0027] Downstream modules may consume the LGM outputs to enhance perception, localisation, mapping, hazard detection, path planning, motion control, or decision-making. For example, an action-sei ection module may combine LGM-derived spatial state vectors with task-specific objectives to generate movement commands; a navigation system may rely on the geometric latent state to maintain orientation under sensor degradation; or a sensorfusion pipeline may use LGM consistency checks to detect anomalous or corrupted sensor readings. Embodiments
[0028] The invention will now be described through a series of non-limiting embodiments. These examples illustrate how the Large Geometric Model (LGM) may be deployed across heterogeneous platforms, including mobile ground robots, underwater vehicles, subterranean agents, indoor GPS-denied systems, aerial drones, humanoid robots, multi-agent networks, wearable devices, and space-based platforms. The embodiments demonstrate that the same physics-supervised geometric foundation may be applied across diverse environments and embodiments.
[0029] In certain embodiments, the geometric-learning process may further incorporate an optional energy-coupled supervisory channel in which agents utilise voltage, current, impedance, or electromagnetic field perturbations propagated through a shared medium as an additional source of information for inferring spatial relationships or hidden topology. This enhancement is compatible with all embodiments described herein and does not limit the scope of the invention.
[0030] It will be understood that the following examples are illustrative and non-exhaustive. Energy-Coupled Embodiment
[0031] In certain embodiments, the LGM is augmented by an energy-coupled supervisory channel. Each agent is electrically coupled—via wired, inductive, sliding-contact, or other conductive interfaces—to a shared DC microgrid, overhead rail, conductive guideway, power bus, or other electrical backbone. The agent measures voltage, current, impedance, harmonic distortion, or electromagnetic field variations at the point of coupling.
[0032] Agents may modulate their electrical load, excitation pattern, or local impedance during operation. These modulations generate spatiotemporally structured electrical disturbances, such as AV / AI fluctuations, harmonic signatures, or propagation delays, which travel across the shared medium. Other agents detect these disturbances and obtain a secondary supervisory signal encoding information about electrical distance, branching structure, impedance distribution, or concealed environmental features.
[0033] The LGM incorporates these energy-derived measurements as additional inputs within the geometric-learning process. By analysing correlations, attenuations, delays, and frequency-domain features, the system may infer aspects of hidden structure not directly observable through inertial, optical, acoustic, or other exteroceptive sensors. This embodiment is particularly relevant in environments with electrical infrastructure, including industrial facilities, subterranean installations, spacecraft, pressurised habitats, and manufacturing plants.
[0034] The energy-coupled embodiment remains fully compatible with the baseline sensor-only geometric learning approach and does not require global coordinates, GPS, or absolute positional information. Mobile Robotic Platform for Subterranean or Tunnel-Based Search and Rescue
[0035] In certain embodiments, the LGM is deployed on a mobile robot operating in subterranean tunnels or collapsed structures. The platform may include inertial sensors, optional acoustic or limited-visibility optical sensors, and a processor executing the LGM.
[0036] As the robot traverses narrow tunnels, bifurcations, debris-filled voids, or unstable structures, the LGM infers geometry from inertial and environmental interaction signals. Even under intermittent or unreliable sensing, the LGM maintains a persistent latent representation of tunnel layout, structural discontinuities, and regions of reduced stability. This enables navigation through confined, occluded, or highly degraded environments.
[0037] Multiple subterranean robots may periodically exchange LGM parameter updates via a federated learning process, allowing subsequent missions to benefit from improved geometric priors without sharing raw sensor data. Underwater or Submersible Platform for Low-Visibility Rescue and Inspection
[0038] In certain embodiments, the LGM is integrated into an underwater or submersible device. Aquatic environments frequently exhibit low visibility due to turbidity, silt, or suspended particulate matter, and acoustic sensing may be degraded by multipath reflections or scattering.
[0039] The submersible collects inertial, pressure, and hydrodynamic measurements during movement. These measurements encode changes in orientation, buoyancy, velocity, and local constraints. The LGM infers geometry—such as submerged obstacles, enclosed chambers, or irregular voids—even when optical and acoustic sensors are unreliable.
[0040] The persistent spatial state maintained by the LGM supports navigation, anomaly detection, and structural inspection under extreme visibility degradation. Federated parameter updates from multiple submersibles may enhance global model robustness across diverse aquatic environments. Wearable Spatial-Inference Device for Confined-Space Rescue
[0041] In certain embodiments, the LGM is incorporated into a wearable device—such as a helmet or head-mounted system—used in confined-space rescue operations. Such environments frequently involve reduced visibility, smoke, darkness, obstructed passages, irregular geometry, and lack of GPS reception.
[0042] The wearable device includes inertial and interaction sensors coupled with a processor executing the LGM. As rescuers crawl, climb, or move through collapsed buildings, ducts, crawlspaces, drainage systems, or industrial vessels, the LGM constructs and maintains an internal representation of local geometry. This representation may include inferred boundaries, voids behind debris, corridor orientation, and relative positioning of other team members using similar devices. Periodic federated learning updates may refine the shared model, enabling improved robustness in future rescue missions. General Applicability
[0043] The invention is not limited to any particular physical platform, sensor suite, communication architecture, or model implementation. The LGM may be deployed in singleagent systems, distributed robotic networks, mobile platforms, fixed installations, wearable devices, aerial vehicles, or space-based systems. Likewise, it may operate in structured indoor spaces, unstructured natural environments, subterranean domains, aquatic regions, or any setting in which geometry must be inferred under degraded or unreliable sensing.
Claims
1. A computer-implemented model for spatial inference, comprising a machine-learning architecture configured to:(a) receive time-varying physical measurements associated with motion of an agent;(b) update a latent spatial state based on the temporal evolution of the physical measurements; and(c) infer one or more geometric properties of an environment from the latent spatial state;wherein the model is trained without absolute positional labels and learns spatial relationships using physics-governed state transitions.
2. A method for training a spatial foundational model, the method comprising:(a) obtaining sequences of physical measurements associated with motion of an agent;(b) omitting absolute positional labels from the training data;(c) processing the sequences using a machine-learning architecture;(d) learning an internal latent spatial state that represents geometric relationships, spatial continuity, or environmental structure;(e) learning cause-and-effect relationships between the agent’s motion and resulting changes in the physical measurements;wherein the model thereby acquires a physics-supervised representation of geometry.
3. A method for spatial inference in an environment, comprising:(a) receiving, at a trained model, a sequence of time-varying physical measurements;(b) updating a latent spatial state based on the received measurements; and(c) generating one or more spatial inference outputs including at least one of:(i) estimated relative position;(ii) inferred environmental geometry;(iii) detection of geometric inconsistencies; or(iv) predicted future spatial states.
4. A device comprising:(a) one or more sensors configured to obtain time-varying physical measurements associated with motion;(b) a processor; and(c) a trained spatial foundational model executed by the processor;wherein the trained model maintains a latent spatial state and infers one or more geometric properties of an environment without reliance on absolute positional signals.
5. The computer-implemented model of claim 1, wherein the physical measurements include inertial measurements, environmental interaction signals, or energy-related telemetry.
6. The computer-implemented model of claim 1, wherein the model learns causal relationships between agent motion and resulting changes in inferred geometry.
7. The computer-implemented model of claim 1, wherein the machine-learning architecture comprises a recurrent neural network, a transformer-based temporal encoder, a physics-informed neural network, a latent-state spatial graph, or any combination thereof.
8. The method of claim 2, wherein learning the latent spatial state comprises enforcing one or more physics constraints including conservation of momentum, smoothness of motion, or structural continuity.
9. The method of claim 3, wherein the spatial inference remains coherent when one or more sensor modalities provide degraded, intermittent, or contradictory measurements.
10. The device of claim 4, wherein the model detects and rejects physically impossible spatial transitions.
11. The computer-implemented model of claim 1, wherein multiple instances of the model participate in a federated learning process by exchanging parameter updates or gradient information without sharing raw sensor data.
12. The method of claim 2, further comprising aggregating model updates from a plurality of devices to refine the spatial foundational model.
13. The method of claim 3, wherein the spatial inference output is provided to a navigation system, perception module, hazard-detection module, or action-sei ection module.
14. The device of claim 4, wherein the device comprises one of:(a) a mobile robotic platform;(b) an underwater or submersible platform;(c) a wearable device; or(d) a fixed monitoring installation.
15. The computer-implemented model of claim 1, wherein the model is configured to operate in an environment having degraded visibility, absence of GPS or GNSS signals, or constrained geometry.
16. The system of Claim 1, wherein each agent is further coupled to a shared electrical or electromagnetic medium and comprises circuitry configured to measure one or more electrical parameters selected from the group consisting of: voltage, current, impedance, harmonic distortion, electromagnetic field variation, and propagation delay; wherein the geometric-learning model incorporates said electrical measurements as a supplementary supervisory signal for inferring relative position, hidden topology, or spatial structure based on perturbations generated by load modulation, impedance variation, or naturally occurring disturbances propagating through the shared medium.
17. A spatiotemporal foundational system comprising a trained spatial model according to any of claims 1-3, wherein the latent spatial state provides a platform-agnostic geometric representation configured to be consumed by heterogeneous downstream modules across a plurality of distinct robotic or autonomous platforms, such that localisation, navigation, perception, or control modules operate using a shared geometric interface independent of platform-specific coordinate frames, sensing configurations, or mapping pipelines.A