Gaze interaction system based on gravity vector constraint and spatial topology association
The gaze interaction system, which uses gravity vector constraints and spatial topology association, solves the problems of unified coordinate expression and object addressing in gaze interaction under different spatial environments. It enables seamless switching across scenarios and secure interaction in high-risk industries, meeting the audit requirements of heavily regulated scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN HUACHUANG HOLOGRAPHIC IMAGING TECHNOLOGY CO LTD
- Filing Date
- 2026-01-28
- Publication Date
- 2026-05-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies cannot achieve a unified coordinate representation and object addressing for gaze interaction in different spatial environments, cannot work across scenarios, and lack a security protocol layer for high-risk scenarios, which prevents the system from being used in industries with strong regulatory oversight.
By associating gravity vector constraints with spatial topology, a gaze interaction system is constructed, including gaze acquisition, attitude correction, spatial object indexing, collision determination, and protocol communication modules. This enables the broadcasting of spatial objects, transmission of gaze vector streams, and command distribution. Combined with the SGIP protocol and XR spatial interaction black-box evidence storage, the closed-loop and security of the interaction are ensured.
It achieves seamless switching and continuous object locking in different indoor and outdoor spatial environments, has cross-platform deployment capabilities, meets the security audit and incident tracing requirements of high-risk industries, significantly reduces the probability of false locking and false triggering, and provides traceable interactive process auditing.
Smart Images

Figure CN122018687A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of spatial computing, embodied intelligence, human-computer interaction and Internet of Things security control technology, specifically a gaze interaction system based on gravity vector constraints and spatial topology association. Background Technology
[0002] With the development of spatial computing and embodied intelligence, spatial computing terminals (including but not limited to AI glasses, AR glasses, head-mounted devices, mobile terminals, and embodied intelligent robot terminals) are gradually shifting from "information display / content consumption" to "physical world operation entry points." Among these, gaze-based interaction is considered one of the most intuitive input methods: users only need to "look at an object" to express their selection, attention, and operational intentions, thereby achieving a natural "what you see is what you get" interactive experience.
[0003] Starting from first principles, for gaze interaction to become an "entry point for physical world operations," it must simultaneously meet three underlying conditions: geometric alignment: the gaze ray must be computable, landable, and reproducible in three-dimensional space; object addressability: the gazed entity must have a unique identity and spatial index in the system; and command execution: the interaction result must be able to form a protocol loop.
[0004] Current technologies cannot simultaneously meet the above conditions: If a stable reference frame that doesn't change with the environment is lacking, gaze-based interaction cannot become an executable physical control method and can only remain at a weak interaction level of "prompting / displaying"; existing solutions typically rely on GNSS / RTK outdoors and local positioning such as UWB / WiFi indoors, while visual SLAM forms another set of local coordinate systems. If heterogeneous coordinates cannot be unified into a single "object addressing mechanism," gaze-based interaction will not be able to work across scenarios or become a universal industry protocol; existing IoT protocols emphasize network connectivity rather than spatial semantics, and without a spatial indexing layer, there is no unified entry point for physical world interaction.
[0005] Furthermore, most existing interaction solutions lack a "security protocol layer" for high-risk scenarios, and even more so, they lack a "XR spatial interaction black box" mechanism similar to aviation / automotive event recorders, preventing the system from being applied in heavily regulated, high-risk industries. Spatial interaction protocols without auditable and documented records are unlikely to become industry infrastructure. Summary of the Invention
[0006] To address the shortcomings of existing technologies, the present invention aims to provide a gaze interaction system based on gravity vector constraints and spatial topology association, thereby solving the problems mentioned in the background. The present invention can construct a unified spatial coordinate expression and object addressing mechanism through multi-source positioning / sensing information in different indoor and outdoor spatial environments, and complete spatial object broadcasting, gaze vector stream transmission, interaction locking acknowledgment and command distribution through standardized protocols, ultimately realizing a closed-loop interaction of "what you see is what you control" for physical entities.
[0007] To achieve the above objectives, the present invention provides a gaze interaction system based on gravity vector constraints and spatial topology association, comprising the following modules:
[0008] The gaze acquisition module is used to acquire user gaze information and generate an original gaze direction vector; the attitude correction module is used to acquire a gravity vector and construct an attitude correction operator based on the gravity vector to correct the original gaze direction vector and obtain a gravity-aligned gaze direction vector.
[0009] The spatial object indexing module is used to maintain a spatial object library, wherein the spatial object includes at least object identifier, spatial location and geometric bounding body information; the collision determination and locking module is used to construct a gaze ray based on the gravity-aligned gaze direction vector and perform collision determination between the gaze ray and the geometric bounding body information to generate a target locking result;
[0010] The protocol communication module is used to distribute control commands to the target space object and receive execution receipts according to the standardized interaction protocol.
[0011] Furthermore, the gaze interaction system is used to implement the following steps:
[0012] S1. The spatial computing terminal collects user gaze information and generates the original gaze direction vector;
[0013] S2. The gravity vector is obtained by the inertial measurement unit of the space computing terminal, and an attitude correction operator is constructed based on the gravity vector to correct the original gaze direction vector, so as to obtain the gravity-aligned gaze direction vector.
[0014] S3. Construct a gaze ray based on the gravity-aligned gaze direction vector, and obtain the geometric bounding volume information of at least one spatial object in the spatial object library;
[0015] S4. Perform spatial collision determination between the gaze ray and the geometric bounding volume information. When the hit condition and gaze dwell time threshold are met, generate target locking result.
[0016] S5. Distribute control instructions to the spatial object corresponding to the target locking result according to the standardized interaction protocol, and receive the execution receipt fed back by the spatial object to complete the gaze interaction closed loop.
[0017] Furthermore, the gaze interaction system also includes a gaze interaction protocol (which may be referred to as SGIP (Spatial Gaze Interaction Protocol)). This protocol specifies the data interaction process and message format for gaze interaction between the spatial computing terminal and spatial objects. The protocol includes at least the following:
[0018] A spatial object broadcast message is used by a spatial object to announce its spatial existence and control capabilities. The spatial object broadcast message includes at least one or more of the following: object identifier, spatial anchor point information, semantic tag, geometric bounding volume parameters, and set of executable instructions.
[0019] A gaze vector stream message is used by the space computing terminal to continuously output gaze ray information. The gaze vector stream message includes at least one or more of the following: gaze ray origin information, gaze direction vector information, gravity alignment identifier, and timestamp information. A target lock receipt message is used to feed back the lock result to the space computing terminal when the target object is gazed upon and locked.
[0020] Command distribution message, used to send control commands to the target object after the target object is locked;
[0021] An execution receipt message is used to provide feedback on the execution result after the target object executes the control command; wherein, the protocol also stipulates that: the spatial computing terminal performs attitude correction based on the gravity vector on the gaze direction vector, and performs spatial collision determination based on the corrected gaze ray and the geometric bounding volume parameters to complete the gaze interaction closed loop.
[0022] Furthermore, the gaze interaction system is also used to implement a spatial object broadcasting and registration method, including:
[0023] Obtain the object identifier, semantic label, geometric bounding volume parameters, and set of executable instructions for a spatial object;
[0024] Obtain spatial anchor point information of spatial objects, wherein the spatial anchor point information includes any one or more of absolute coordinate anchor points, local coordinate anchor points, or visual feature anchor points;
[0025] Generate a spatial object broadcast message and broadcast the spatial object broadcast message to the spatial index center or spatial computing terminal;
[0026] The spatial indexing center writes the spatial objects into the spatial object library based on the spatial object broadcast message, thereby enabling the spatial computing terminal to perform collision determination and lock the spatial objects based on the gaze ray and the spatial object library.
[0027] Furthermore, the original gaze direction vector includes a compensation term for the difference between the user's individual visual axis and optical axis, which is obtained through user calibration parameters, online adaptive estimation, or preset model parameters;
[0028] S2 further includes constructing a yaw alignment operator based on magnetometer output or visual orientation reference to lock the horizontal orientation of the gravity-aligned gaze direction vector.
[0029] The direction vector of the gaze ray is obtained by the combined action of the gravity alignment operator and the yaw alignment operator, so that the gaze ray satisfies the three-degree-of-freedom attitude stability.
[0030] The geometric bounding box information of the spatial object includes, but is not limited to, any one or more of the following: axis-aligned bounding box (AABB), oriented bounding box (OBB), sphere, cylinder, or polyhedron.
[0031] Furthermore, when the gaze ray hits multiple spatial objects, the objects are sorted by depth according to the distance between the ray and the nearest intersection point, and the spatial object with the smallest distance is locked first.
[0032] The multi-target hit further combines the results of depth sensor, binocular vision, structured light, ToF, UWB ranging or visual depth estimation to perform occlusion culling, so as to reduce the locking probability of occluded objects or cullate occluded objects.
[0033] The multi-target hit further combines spatial object semantic correlation determination for locking and disambiguation. The semantic correlation determination includes any one or more of the following: the matching degree between the target category and the current task, the correlation between the target and the user's action, the topological adjacency relationship between the target and historically locked objects, or the consistency of visual semantic confidence.
[0034] Furthermore, the gaze dwell time threshold in S4 is used to switch between the candidate hit state and the stable lock state.
[0035] The step prior to S5 further includes an intent confirmation step, wherein the intent confirmation is achieved through any one of voice confirmation, blink confirmation, gesture confirmation, secondary gaze confirmation, or multimodal combination confirmation.
[0036] The spatial location of the spatial object is obtained by fusing multi-source positioning / sensing information, which includes any one or more of satellite positioning, RTK, UWB, WiFi-FTM, Bluetooth AoA, or visual SLAM.
[0037] The fusion of multi-source localization / sensing information is achieved through Kalman filtering or factor graph optimization to suppress pose jumps or jitter caused by switching of heterogeneous localization sources;
[0038] The object identifiers of the spatial object library are generated by a spatial domain name resolution mechanism, which is used to uniformly map heterogeneous coordinate representations into spatial semantic object identifiers.
[0039] The standardized interaction protocol includes spatial object broadcast messages, which contain at least one or more of the following: object identifier, spatial anchor point information, semantic tags, geometric parameters, and a set of executable instructions.
[0040] When the spatial object is a dynamic target, timing alignment is performed based on the target pose timestamp and the gaze ray timestamp, and motion prediction compensation is performed based on the target velocity or acceleration information before performing the spatial collision determination in step S4.
[0041] Before distributing control instructions in S5, spatial permission verification is performed, which includes at least one or more of the following: gaze hit validity verification, operator identity authentication verification, and operator spatial legitimacy verification.
[0042] Furthermore, when an abnormal command is detected, a high-risk accidental touch, an impact event, or an access denial event is detected, XR space interaction black box evidence storage is triggered. The interaction black box evidence storage records at least one or more of the following: physical sensor data, interaction trajectory data, and semantic evidence data.
[0043] The XR space interaction black box evidence storage generates a data digest and performs signature encapsulation in the trusted execution environment (TEE) on the edge side, and asynchronously synchronizes the data digest to the edge node or trusted evidence storage service to form multiple copies of the evidence storage;
[0044] The semantic evidence data supports data anonymization processing, which includes saving only one or more of the target object feature descriptors, semantic segmentation ROIs, bounding box information, or semantic category confidence, without saving the complete original image or video stream.
[0045] The XR spatial interaction black box evidence storage obtains a reliable timestamp through satellite time synchronization, network time synchronization, or a clock synchronization server, and binds the timestamp with the data digest to achieve the integrity of a verifiable chain of evidence.
[0046] Furthermore, when a dynamic target is occluded, semantic recognition fails, or the positioning signal is momentarily interrupted, the system performs trajectory extrapolation based on the pose sequence and motion vector before the target disappears within a preset lock-out time threshold ΔT_hold to maintain the target lock state; and performs recapture matching to restore the lock when the target reappears, and falls back to the candidate hit or exit state when ΔT_hold is exceeded.
[0047] The spatial domain name resolution mechanism supports multi-level resolution methods, including full resolution in the cloud, regional caching at edge nodes, and local pre-fetching caching on the terminal. This enables the spatial computing terminal to complete local collision detection and target locking based on a pre-downloaded subset of spatial object topology under weak network or offline conditions, and supports the real-time discovery of temporarily appearing local spatial objects through broadcast protocols.
[0048] Furthermore, the spatial domain name resolution mechanism (sDNS) supports pose subscription updates for dynamic spatial objects; when a spatial object is a dynamic object, its geometric bounding volume parameters are pushed to the terminal cache at a preset update frequency or threshold triggering method as the object pose changes, and carry the object pose timestamp for time-series alignment with the gaze ray timestamp to ensure consistency in spatial collision determination.
[0049] When multiple spatial computing terminals lock the same spatial object simultaneously, the spatial domain name resolution mechanism (sDNS) or edge nodes issue the object locking status based on preset permission priorities, first-occupancy logic, or queuing strategies, and return an occupation prompt or rejection receipt to the non-occupying terminals to avoid concurrent command conflicts and misoperations.
[0050] The beneficial effects of this invention are:
[0051] 1. This gaze interaction system based on gravity vector constraints and spatial topology establishes an absolute vertical reference through gravity vector correction, enabling the gaze ray to stably align with physical entities even under dynamic posture, thus solving the problem of pure visual gaze drift; through a heterogeneous anchor point fusion mechanism, it achieves seamless switching between indoor and outdoor environments, ensuring the continuity of object locking and coordinate output.
[0052] 2. This gaze interaction system based on gravity vector constraints and spatial topology uses sDNS spatial domain name resolution to "upgrade" coordinates to semantic object addressing, establishing a unified rule for "gaze intent → object identity → instruction execution"; and defines spatial object broadcasting, gaze vector stream, lock receipt and instruction distribution through the SGIP protocol, enabling the system to have cross-platform deployment and ecosystem expansion capabilities.
[0053] 3. This gaze interaction system based on gravity vector constraints and spatial topology association meets the rigid requirements of high-risk industries such as industry, medical care, and public safety for security auditing and accident tracing through Spatial-ACL permission verification and XR spatial interaction black box evidence storage, forming an irreplaceable industry barrier.
[0054] 4. Based on the absolute vertical reference of gravity vector, this invention enables the wearable terminal to maintain a stable gaze point under dynamic conditions such as walking, vibration, and tilting, significantly reducing the probability of false locking and false triggering; by establishing the unique identity and occupancy status of spatial objects through sDNS addressing and object lock (ObjectLock) mechanism, the interaction is upgraded from "seeing" to "addressable, lockable, and executable" spatial sovereignty control.
[0055] 5. This invention achieves auditing, evidence preservation, and liability determination of the interaction process through an XR spatial interaction black box, meeting the rigid requirements of accident tracing and evidence chain in highly regulated scenarios such as industry and healthcare. Attached Figure Description
[0056] Figure 1 This is a general logical block diagram of a gaze interaction system based on gravity vector constraints and spatial topology association according to the present invention.
[0057] Figure 2 This is a schematic diagram of the gravity alignment reference frame transformation of the present invention;
[0058] Figure 3 This is a curve diagram showing the smooth switching of multi-source coordinate fusion in this invention;
[0059] Figure 4 This is a geometric model diagram of the collision between the gaze ray and the topological bounding volume in this invention;
[0060] Figure 5 This is a flowchart of the interactive state machine transition of the present invention;
[0061] Figure 6 This is a diagram of the XR spatial interaction black box evidence storage data structure of the present invention;
[0062] Figure 7 This is a timing diagram of the SGIP protocol interaction in this invention. Detailed Implementation
[0063] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.
[0064] Please see Figures 1 to 7This invention provides the following technical solution: a gaze interaction system based on gravity vector constraints and spatial topology association. This system abstracts "gaze" as a ray in three-dimensional space, uses a "gravity vector" to establish an absolute physical benchmark that does not drift with the wearing posture, and then uses a "spatial topology object library" to determine the collision between the ray and real physical entities, ultimately forming a complete closed loop of "locking—command distribution—acknowledgment feedback—security audit." Simultaneously, the system can access multi-source positioning / sensing information in different indoor and outdoor environments, achieving seamless and consistent spatial addressing across all scenarios. Furthermore, it ensures traceability and compliant implementation capabilities in high-risk scenarios through permission verification and black-box evidence storage for XR spatial interaction.
[0065] Terminology definition: XR spatial interaction black box
[0066] The "XR spatial interaction black box" described in this invention refers to an interaction auditing and evidence storage module deployed on a spatial computing terminal and / or edge node, used to record, encapsulate, sign, and trace the spatial interaction process based on gaze vectors. For ease of description, it is also referred to as the "interaction black box" below.
[0067] The XR spatial interaction black box is used to record at least one or more of the following data: gravity vector and inertial measurement data, gaze ray parameters and hit results, spatial object identifiers and their pose / bounding volume parameters, semantic snapshots or their feature descriptors, interaction commands and acknowledgment information, and timestamps and integrity verification information associated with the above data; wherein, the timestamps may be provided by the terminal's local clock or an external time source, and the integrity verification information includes any one or more of hash digests, digital signatures, or message authentication codes.
[0068] In one implementation, the XR space interaction black box generates a storage digest within a Trusted Execution Environment (TEE) and synchronizes the digest or encrypted storage data to edge nodes / cloud storage to form an auditable and traceable chain of evidence.
[0069] This embodiment provides the following specific technical solution:
[0070] 1. Overall Solution and System Architecture (Unified Closed Loop of Protocol and System)
[0071] The SGIP protocol and system in this embodiment preferably include the following functional modules: (1) A gaze acquisition module is used to collect key interactive parameters such as the wearer's gaze direction, pupil information, and dwell time, and to generate a gaze direction vector.
[0072] (2) Gravity vector correction module
[0073] Using an inertial measurement unit (IMU) to extract gravity vectors in real time, a gravity alignment reference system is constructed to dynamically correct the gaze ray, so that the gaze direction remains geometrically consistent with the physical world even when the device is tilted, moving, or vibrating. (3) The visual core object locking module (core interaction main line) extracts visual features and performs semantic recognition through camera images, and prioritizes the identification and semantic confirmation of the "object being gazed upon", ensuring that the interaction main line is centered on "what you see is the object".
[0074] (4) Heterogeneous coordinate fusion module (full-scene base)
[0075] Outdoors, it integrates absolute positioning such as GNSS / RTK, while indoors it integrates local positioning such as UWB, WiFi-FTM, and Bluetooth AoA. It then aligns and fuses these local coordinates with those of visual SLAM to output unified fused coordinates.
[0076] (5) Spatial Topology Object Index Module
[0077] Maintain a spatial object library, which abstracts real physical entities into spatial objects, including at least object ID, semantic label, spatial location and geometric bounding volume, for subsequent collision determination and addressing resolution.
[0078] (6) Collision determination and interaction state machine module
[0079] Real-time collision / distance determination is performed between the gaze ray after gravity correction and the bounding volume of spatial objects, and interactive process control such as candidate selection, locking, confirmation, execution, and cancellation is completed based on the state machine.
[0080] (7) SGIP protocol communication module
[0081] It provides standardized messaging interfaces, including at least spatial object broadcasting, gaze vector stream transmission, lock receipts, and command distribution, enabling cross-device and cross-platform interconnection.
[0082] (8) Security permissions and XR space interaction black box evidence storage module (high barrier capability) In high-risk control scenarios, it performs permission verification on gaze-triggered behavior and forms XR space interaction black box evidence storage data when abnormal or shock events occur, ensuring auditability and traceability.
[0083] 2. The underlying logic of mathematics and physics (first principles: gravity is the only absolute benchmark)
[0084] 2.1 Coordinate System Definition
[0085] To achieve stable and reproducible spatial interaction, the present invention preferably defines the following coordinate system:
[0086] Local coordinate system of the spatial computing terminal body: ;
[0087] Gravity-aligned coordinate system: ;
[0088] World Unified Coordinate System: (Can be an absolute geographic coordinate system or a uniformly mapped coordinate system)
[0089] in, The introduction of this concept is key to the invention: the gravity vector exists constantly and has a unique direction in the physical world, and can be used to establish an "absolute vertical reference," thereby eliminating gaze drift caused by changes in the wearer's posture. The gravity vector... The physical parameters characterizing the direction of gravity can be obtained in ways including but not limited to direct acquisition by the hardware sensors of the inertial measurement unit, or by calculation in an equivalent sense using a vision / pose calculation algorithm.
[0090] 2.2 Gravity Alignment Correction Operator (R_align)
[0091] The inertial measurement unit (IMU) outputs accelerometer measurements, from which the system extracts the gravity direction vector. The methods for obtaining the gravity vector include, but are not limited to, direct measurement using accelerometers, or virtual gravity references calculated by reverse engineering using visual inertial odometry (VIO) and visual semantic features (such as horizontal / vertical wall features).
[0092] This invention constructs a gravity alignment operator This is used to map the gaze vector in the device's local coordinate system to the gravity-aligned coordinate system.
[0093]
[0094] in:
[0095] Let be the unit vector representing the gaze direction in the device's local coordinate system. This is the gaze direction vector after gravity alignment.
[0096] With this correction operator, even if the user tilts their head down, leans to the side, or experiences vibrations while walking, the gaze ray can still maintain its geometric alignment with the real physical entity.
[0097] 2.2.1 Yaw Alignment Operator: To further improve the spatial attitude determinism in dynamic scenes, this invention introduces a yaw alignment operator based on gravity vector alignment. This invention addresses the limitation of gravity vectors in constraining yaw angles. Since gravity vectors only provide an "absolute vertical" reference, the direction of gravity remains unchanged when the wearer rotates in place. Without a horizontal reference, the orientation of the gaze ray on the horizontal plane may drift, affecting target locking accuracy. This invention preferably uses one or more of the following methods to generate a horizontal orientation reference and constructs a yaw alignment operator. .
[0098] (1) Magnetometer North Reference Method: The magnetometer outputs the geomagnetic direction vector through the IMU. Combined with the direction of gravity Orthogonalization is performed to construct a stable horizontal reference axis, thereby obtaining an absolute azimuth reference.
[0099] (2) Visual orientation reference method: using the initial pose output by visual SLAM during the initialization phase. Alternatively, the orientation of a stable set of feature points can be used as a reference for azimuth locking;
[0100] (3) Hybrid fusion method: Weighted fusion or gating selection of magnetometer and visual orientation is performed to improve robustness in weak magnetic interference or weak texture environments. Finally, this invention adopts a combined attitude alignment operator. Perform absolute pose mapping on the gaze direction:
[0101]
[0102] in, Used to establish an absolute vertical reference for gravity alignment. Used to lock the horizontal orientation, enabling the gaze ray to have complete three-degree-of-freedom attitude stability, thereby constructing a true "absolute attitude gaze ray".
[0103] 2.3 The gaze ray equation (Ray Model) abstracts gaze behavior into a three-dimensional spatial ray:
[0104] in:
[0105] The ray origin can be the center point of the eyeball, the center point of the pupil, or its external parameter mapping point in the coordinate system of the spatial computing terminal. This is the unit vector for the gaze direction.
[0106] To eliminate the deviation between the visual axis and optical axis caused by the different physical structures of users' eyes, the present invention preferably makes the direction vector The view axis offset compensation term includes individual user characteristics, namely:
[0107]
[0108] in, This is the original direction vector output by the eye-tracking module. The compensation term for individual differences among users can be obtained through one-time user calibration, online adaptive estimation, or default model parameters (including but not limited to pupil-corneal reflex models, fixation point regression fitting models, etc.). This ray serves as the geometric basis input for subsequent "hit determination, locking, and command triggering".
[0109] 3. Spatial topology object indexing and collision determination (look-through hit = computable spatial event)
[0110] 3.1 Space Object Model This invention abstracts devices, facilities, and targets in the physical world into "space objects". It includes at least:
[0111] Unique identifier for the object: ;
[0112] Semantic tags: (e.g., switches, valves, fire extinguishers, etc.);
[0113] Spatial location: (Can be absolute or merged coordinates);
[0114] Triggering bounding volume: (Bounding Volume) It can be implemented using any or a combination of these, including but not limited to: AABB bounding box, OBB bounding box, sphere, cylinder, polyhedron, etc.
[0115] 3.2 Fixation Hit Test
[0116] This invention targets radiation. With surrounding body Perform real-time intersection or distance determination. One of the following rules is preferred:
[0117] If the ray intersects with the surrounding volume, it is considered a hit;
[0118] If the minimum distance between the ray and the surrounding volume satisfies the threshold:
[0119]
[0120] Then it is determined to be a candidate target.
[0121] To reduce accidental touches, this invention further introduces a "dwell time" condition: when the duration of gaze hit reaches a threshold. Only then does it enter a "stable lock" state.
[0122] 3.3 Multi-target hit disambiguation and occlusion elimination mechanisms In real-world space environments, gaze rays It may hit the bounding volume of multiple spatial objects simultaneously. This includes situations such as looking at a target outside a window through glass, looking at devices in front or behind, or multiple objects overlapping in space. To avoid misjudgment by the system leading to erroneous command issuance, this invention further introduces a multi-target disambiguation and occlusion elimination mechanism on the basis of collision determination, used to determine the final locked object from the candidate hit set.
[0123] The present invention preferably constructs a candidate set when the gaze ray hits multiple objects:
[0124]
[0125] And perform one or more of the following decision strategies on the candidate set: (1) Depth Priority operator: calculate the nearest intersection parameter between the ray and the object for all candidate objects. Preferred The smallest one is used as the locking target, that is, the object closest to the wearer is locked first;
[0126] (2) Occlusion Culling: Combine the results of depth sensor, binocular vision, structured light, ToF, UWB ranging or visual depth estimation to determine whether the candidate object is occluded by other objects; if the object is occluded at the depth level, reduce its locking probability or cull it directly.
[0127] (3) Probabilistic Locking: Constructing a comprehensive scoring function for candidate objects:
[0128] in Distance priority is used. For visual semantic confidence,
[0129] This is a safety level penalty item (to reduce the probability of accidental touch on high-risk equipment). This is an enhancement for the residency requirement.
[0130] The system selects the object with the highest score as the final target and outputs a lock confirmation. If the score difference between candidate objects is insufficient or a high-risk object exists, a secondary confirmation process is triggered or the lock is rejected, improving the security and certainty of the interaction.
[0131] (4) Semantic Affinity Determination: In extreme scenarios such as transparent / semi-transparent media occlusion, nested targets, or complex background penetration, the system further combines the semantic affinity of spatial objects with contextual consistency to optimize target selection. The semantic affinity includes, but is not limited to: the matching degree between the target category and the current task, the affinity between the target and user actions (such as hand approach / operation posture), the topological adjacency relationship between the target and the most recently locked object, and the consistency of visual semantic recognition confidence, etc.
[0132] For example, when a ray simultaneously penetrates both the "wall behind the glass" and the "switch object in front of the glass," if the system detects that the user's hand gesture is semantically consistent with the switch object, it will prioritize locking the switch object to reduce the risk of accidental touch under extreme obstruction conditions.
[0133] 4. Full-Scene Heterogeneous Coordinate Fusion Engine (Seamless Indoor and Outdoor + Visual Sovereignty) To meet the "full-scene adaptation" requirement, this invention proposes a heterogeneous anchor fusion mechanism to unify different positioning sources under the same spatial addressing logic.
[0134] 4.1 Outdoor Absolute Layer: In outdoor environments, the system can access GNSS / RTK (including but not limited to BeiDou, GPS, and multi-system fusion) to provide a globally unique absolute position index. The core value of this layer is to provide spatial objects and interactive behaviors with globally reproducible coordinate anchor points that can be reproduced across regions.
[0135] 4.2 Indoor Local Patch Layer: In indoor or GNSS unavailable scenarios, the system can access local positioning signals such as UWB, WiFi-FTM, and Bluetooth AoA to provide local coordinates or distance constraints, achieving centimeter-level or sub-meter-level positioning enhancement.
[0136] 4.3 Visual Core Layer: In this invention, vision is not a patch, but the core thread of gaze interaction. When a user looks at an object, the system first identifies the object through visual feature points and semantic recognition (e.g., identifying it as "Power_Switch_01"), and associates the visual anchor point with the spatial object library, thereby ensuring that the interaction understanding conforms to the user's intuition.
[0137] 4.4 Smooth Transition Fusion Output: The system dynamically allocates weights based on the quality of each signal source and outputs unified fused coordinates.
[0138]
[0139] The weights satisfy the following: It automatically adjusts according to signal quality, ensuring continuous coordinates and no target loss when switching between indoor and outdoor environments.
[0140] 4.5 Heterogeneous accuracy jump suppression and continuous pose optimization mechanism Since there are significant differences in error models and accuracy levels between outdoor absolute positioning sources (e.g., GNSS / RTK) and indoor local positioning sources (e.g., UWB), if only simple weighted fusion is used, visual instantaneous jumps, jitter, or highlight box drift may occur at the switching boundary, affecting the continuity of interaction and user trust.
[0141] To this end, the present invention introduces a continuous pose optimization mechanism in the heterogeneous coordinate fusion module, preferably using Kalman filtering or factor graph optimization to perform smooth constraints on the fused pose.
[0142] (1) Kalman filtering method: establish state vector (Position and velocity) and construct observation models for different signal sources, adaptively adjusting and updating the gain based on the variance of each observation to achieve a continuous transition from prediction to correction. Larger measurement noise is assigned to low-precision observations (such as GNSS), while smaller measurement noise is assigned to high-precision observations (such as UWB), thereby suppressing jumps.
[0143] (2) Factor graph optimization method: GNSS, UWB, visual odometry, and other observations are added to the same optimization framework as different factor constraints to form a sliding window optimization, ensuring that pose updates meet the continuity and consistency constraints, and robustly reducing the weight of abnormal observations. Through the above mechanism, this invention can guarantee the output fused coordinates under conditions of indoor / outdoor signal switching, signal transient failure, or multi-source conflict. It maintains continuity, preventing drastic visual jitter in interactive objects, thereby improving gaze-locking stability and user experience reliability.
[0144] 5. sDNS Spatial Domain Name System (Upgrading "Coordinates" to "Object Addressing") Most existing IoT control systems lack a spatial indexing layer. This invention proposes the sDNS (Spatial Domain Name System) spatial domain name resolution mechanism: uniformly mapping entities under heterogeneous coordinates to "spatial semantic object IDs," achieving universal addressing for gaze-based interactions.
[0145] 5.1 Core Link (Must be closed loop)
[0146] The present invention preferably solidifies the interaction path into the following closed-loop link: Ray → Resolve(sDNS) → HitTest → Lock → Dispatch → Ack
[0147] Right now:
[0148] Determine spatial orientation using the gaze ray.
[0149] It is resolved to the candidate object list using sDNS Resolve.
[0150] Perform HitTest collision detection.
[0151] Achieve stable locking.
[0152] Distribute the Dispatch command.
[0153] Return Ack receipt feedback.
[0154] 5.2 Cloud-Edge-Device Multi-Level Caching and Real-Time Object Discovery Mechanism The spatial domain name resolution mechanism (sDNS) preferably supports a cloud-edge-device multi-level object resolution and caching mechanism:
[0155] (1) Cloud-based global resolution: The cloud maintains the full spatial object index and topology relationship, providing cross-regional object identity resolution and global consistency verification; (2) Edge-side regional caching: Edge nodes are deployed in local areas such as parks / factories / buildings to cache the object topology subset and frequently updated objects in the area, providing low-latency resolution and nearby services;
[0156] (3) End-side prefetching and local hit: The spatial computing terminal predicts based on the current location and field of view, and prefetches the topological subset of spatial objects and bounding volume parameters of the surrounding area. Even under weak network or offline conditions, it can still achieve millisecond-level local collision determination and locking.
[0157] Meanwhile, the sDNS supports real-time discovery and access to temporarily appearing local spatial objects via broadcast protocols. These temporary objects include, but are not limited to, mobile robots, temporary turnstiles, mobile devices, or temporary construction facilities, thereby ensuring the real-time performance and availability of spatial object resolution in dynamic scenarios.
[0158] 6. SGIP Protocol Standard (Data Structure and Interface Definition)
[0159] The message interaction mechanism of this invention can be called SGIP (Spatial Gaze Interaction Protocol), which is used to define standardized message types, fields and interaction processes between spatial computing terminals and spatial pairs.
[0160] Furthermore, the present invention also provides a message interaction mechanism / interface specification (SGIP) for the gaze interaction, including object broadcasting, gaze vector stream uploading, interaction lock receipt, instruction distribution, and execution receipt, so as to realize standardized linkage and closed-loop control between the spatial computing terminal and the spatial object.
[0161] This invention proposes an SGIP protocol standard, which includes at least the following message types:
[0162] 6.1 Full-Scene Spatial Object Broadcast
[0163] The device or spatial object broadcasts its spatial presence and capability set to the Spatial Index Center: { "protocol_version": "SGIP-3.0",
[0164] "anchor_mode": "Hybrid(Visual+UWB)",
[0165] "object_id": "Power_Switch_01",
[0166] "global_position": { "lat": 31.235, "lon": 121.474, "alt": 42.10},
[0167] "local_offset": { "dx": 1.2, "dy": -0.5, "dz": 0.0},
[0168] "local_visual_hash": "v_7ed29b",
[0169] "semantic_tag": "Power_Switch_01",
[0170] "shape": "box",
[0171] "dimensions": { "x": 0.3, "y": 0.2, "z": 0.1},
[0172] "command_set": ["on", "off", "query"],
[0173] "safety_level": "Critical",
[0174] "cache_ttl_ms": 3000,
[0175] "region_topology_hash": "r_a91f2c"}
[0176] in:
[0177] global_position is used for absolute anchor points outdoors.
[0178] local_offset is used for local anchor points indoors.
[0179] local_visual_hash is used for visual object fingerprint anchors.
[0180] command_set defines the set of controllable capabilities.
[0181] safety_level is used for access control of high-risk devices.
[0182] cache_ttl_ms is used for endpoint cache invalidation control.
[0183] region_topology_hash is used for region topology version consistency verification and incremental updates.
[0184] 6.2 The gaze vector stream spatial computing terminal continuously outputs gaze ray information after gravity correction:
[0185] {
[0186] "header": { "seq": 1024, "stamp": 1705270341},
[0187] "origin_fusion": [31.2351, 121.4742, 43.5],
[0188] "direction_vector": [0.707, 0.0, -0.707],
[0189] "gravity_aligned": true,
[0190] "gaze_metadata": { "dwell_time": 450,
[0191] "confidence": 0.92,
[0192] "pupil_diameter": 3.2
[0193] }
[0194] }
[0195] 6.3 Interaction ACK: When the system determines that a gaze has been hit and the object has been locked, an ACK is sent.
[0196] {
[0197] "status": "TARGET_LOCKED",
[0198] "target_id": "Power_Switch_01",
[0199] "feedback": {
[0200] "visual": "highlight_edge",
[0201] "audio": "click",
[0202] "haptic": "pulse"
[0203] }
[0204] }
[0205] 6.4 Command Dispatch: After confirming the intent, the device control command is issued:
[0206] {
[0207] "target_id": "Power_Switch_01",
[0208] "command": "off",
[0209] "parameters": {}, "timeout_ms": 1200,
[0210] "require_confirm": true
[0211] }
[0212] Furthermore, to prevent the broadcast and control commands for space objects from being forged or replayed, SGIP messages may optionally include a unified security header field, which includes at least a nonce, a signature, a key identifier (key_id), and a timestamp; the signature is generated based on the message body digest and is used by the receiving end for integrity verification and anti-replay verification.
[0213] 6.5 Dynamic Object Timing Alignment and Motion Prediction Compensation Mechanism To support dynamic target interaction across all scenarios (including but not limited to industrial robotic arms, mobile inspection vehicles, unmanned equipment, dynamic gates, etc.), this invention introduces a timing alignment and motion prediction compensation mechanism (Motion Prediction) into the SGIP protocol to address the latency mismatch problem of "low object coordinate update frequency and high gaze vector flow frequency." Preferably, this invention introduces a target motion state field into the spatial object broadcast protocol, including at least velocity, acceleration, or prediction model parameters:
[0214] {
[0215] "target_id": "RobotArm_01",
[0216] "pose": [ ... ],
[0217] "velocity": [vx, vy, vz],
[0218] "acc": [ax, ay, az],
[0219] "stamp": 1705270341
[0220] }
[0221] At the same time, a high-precision timestamp is forcibly included in the gaze vector stream to form a unified time alignment basis:
[0222] {
[0223] "header": { "seq": 2048, "stamp": 1705270341.123},
[0224] "direction_vector": [ ... ],
[0225] "origin_fusion": [...]
[0226] }
[0227] When the system calculates gaze hits at the edge or in the cloud, the following compensation logic is preferably executed: for any target object Report the time Extrapolate the position state to the current gaze time t to obtain the predicted position. :
[0228]
[0229] The predicted position is used instead of the original position in collision determination, thereby avoiding the "missed shot" phenomenon and ensuring the real-time performance and stability of dynamic interaction. In addition, this invention can also introduce time window backtracking alignment in the gaze locking stage, that is, to check the consistency between the gaze ray and the predicted target position in multiple frames within a certain time window, so as to improve the anti-shake capability of dynamic target locking.
[0230] Furthermore, the motion prediction compensation mechanism also includes trajectory estimation and lock-maintaining logic after a short-term loss of target lock. When a dynamic target is occluded, visual semantic recognition fails briefly, or the positioning / sensor signal is momentarily interrupted, causing the target object to fail to be parsed in the current frame, the system performs inertial extrapolation based on the pose sequence and motion vectors (including but not limited to velocity, acceleration, or historical trajectory curves) before the target disappears, maintaining the locked state of the gaze ray and consistency with the candidate target identity within a preset loss-of-lock time threshold ΔT_hold; when the target reappears, the system performs recapture matching and restores stable lock; if the ΔT_hold is exceeded, it automatically degrades to a candidate hit or exit state to avoid erroneous lock-up and false triggering.
[0231] 7. Interactive State Machine (Engineering Closed Loop from "Looking" to "Execution") To ensure stable, controllable, and reversible interaction, this invention introduces a state machine mechanism:
[0232] S0 Idle scan: Target not hit.
[0233] S1 Candidate Hit: The ray enters the target's surrounding body, but the dwell time does not reach the threshold.
[0234] S2 Stable Lock-in: Dwell Time Reached And meet the confidence threshold.
[0235] S3 Intent Confirmation: Confirm the action via voice / blink / gesture / secondary gaze.
[0236] S4 instruction execution: Sends control instructions to the target object.
[0237] S5 Feedback: Receives execution results and provides multimodal feedback.
[0238] S6 Exit / Undo / Abnormal Interruption.
[0239] This state machine can significantly reduce accidental touches and ensure consistent engineered behavior in the interaction process.
[0240] The execution instruction (S4) includes sending control messages to external IoT devices, as well as triggering virtual and real interaction events on the terminal's local side (such as AR interface pop-up, local logic jump, or bioelectric signal trigger).
[0241] 7.1 End of sDNS related paragraphs
[0242] Furthermore, the spatial domain name resolution mechanism (sDNS) supports a real-time coordinate subscription and push mechanism: when the spatial object is a dynamic object, the pose information and geometric bounding volume parameters of the spatial object are updated as it moves, and pushed to edge nodes and terminal cache at a preset update frequency f_update or event-triggered mode (pose change exceeds the threshold Δp / Δθ); the push message carries the object pose timestamp t_obj, and when the terminal performs gaze ray collision determination, it performs timing alignment based on t_obj and gaze ray timestamp t_gaze, thereby ensuring the spatial consistency and interaction continuity of dynamic objects.
[0243] 7.2 Interactive Handshake Feedback
[0244] Multi-user conflict suppression mechanism: This protocol supports spatial object state locking (ObjectLock). When the first operator enters the stable locking (S2) or execution (S4) state, the sDNS or edge side broadcasts the 'occupancy' flag to other terminals watching the object; if high-frequency concurrent requests occur, the system executes preemption or queuing logic according to the user's permission level (Spatial-ACL Priority).
[0245] 8. Security Audit and Black-Box Evidence Preservation of XR Spatial Interactions (Protocol Layer Moat)
[0246] 8.1 Spatial-ACL View Permission Validation (Physical Level Admission Control) A Spatial-ACL permission table is maintained for each spatial object, preferably containing at least:
[0247] List of permitted operator identities
[0248] Allowed action set
[0249] Allowed trigger distance / area
[0250] Security Level and Secondary Confirmation Rules
[0251] The system can perform "triple verification" before executing instructions:
[0252] (1) Whether the gaze vector actually hits the object (2) Whether the operator has passed identity / biometric authentication (3) Whether the operator is in a legal spatial position and within a safe distance If any of these conditions are not met, the control command will be refused and a rejection receipt will be returned.
[0253] 8.2 Black-box evidence preservation of XR spatial interaction (traceable incidents)
[0254] The XR spatial interaction black box is used to record key physical data, interaction states, and semantic evidence data during gaze interaction to form an auditable and traceable chain of evidence. XR spatial interaction black box evidence storage is triggered when one of the following events occurs:
[0255] Abnormal command triggered (high-risk accidental touch);
[0256] Impact / severe acceleration (potential collision);
[0257] Permission denied (suspected illegal operation);
[0258] Target instantaneous drift (suspected anchor point attack or environmental change) XR space interaction black box optimal recording before and after and The data for the time window, such as 180 seconds before and after, must include at least:
[0259] L1 physical data: gravity vector, acceleration, angular velocity, attitude.
[0260] L2 interaction data: target ID, hit / lock time, gaze trajectory.
[0261] L3 Semantic Snapshot: Visual sensor frame capture and semantic recognition results.
[0262] XR spatial interaction black box data can be encrypted and encapsulated with trusted timestamps, and written to an immutable storage area for accident determination and audit tracing. Furthermore, to reduce the data volume of XR spatial interaction black box evidence storage and meet privacy compliance requirements, this invention preferably supports data anonymization and minimal evidence storage mechanisms. Specifically, the L3 semantic snapshot does not necessarily need to save the complete original image / video stream; instead, it only extracts key feature descriptors related to the target object (including but not limited to visual feature vectors, semantic segmentation ROIs, target bounding box coordinates, target category, and confidence level), and crops, blurs, or anonymizes irrelevant regions.
[0263] The anonymization process is preferably performed in an edge-side TEE environment and asynchronously distributed to edge nodes or trusted evidence storage services in the form of encrypted digests. This ensures the verifiability of the evidence chain while reducing storage pressure and mitigating compliance risks associated with collecting information from irrelevant personnel. The 180s evidence storage duration is a preferred embodiment, and the trigger threshold and backtracking duration of the XR space interaction black box can be dynamically adjusted according to the safety level of the target object.
[0264] 8.3 Distributed Trusted Evidence Storage Architecture and Evidence Chain Enhancement Mechanism To enhance the legal evidentiary effect of XR space interactive black box evidence storage in accident identification, liability determination and other scenarios, this invention further proposes a distributed trusted evidence storage architecture to solve the defects of "local storage alone is easily damaged and cloud storage alone is affected by network outages", and realize the integrity, non-repudiation and verifiability of interactive evidence.
[0265] This invention preferably employs a collaborative mechanism of "Trusted Execution Environment (TEE) on the edge + edge node synchronization + trusted time synchronization":
[0266] (1) End-side trusted encapsulation: In the TEE (Trusted Execution Environment) isolated environment, the spatial computing terminal performs layered encapsulation of the black box data of XR spatial interaction and generates data digest hash and digital signature;
[0267] (2) Asynchronous distribution of evidence storage: The digest and signature of the black box data of XR space interaction can be sent asynchronously to edge nodes or trusted evidence storage services to form multiple copies of storage;
[0268] (3) Trusted time synchronization and timestamps: Trusted timestamps are added to each XR spatial interaction black box record through satellite time synchronization, network time synchronization server or dedicated time synchronization module to ensure that the time sequence of the evidence chain is verifiable;
[0269] Edge nodes can write the digest to an immutable storage medium or to a distributed ledger with tamper-proof properties to ensure that subsequent audits are verifiable and accountable. Through the above mechanism, this invention enables XR space interactive black-box data to possess the following characteristics:
[0270] Even if the equipment is damaged, it can still be traced (multiple copies);
[0271] Local trusted signing (TEE encapsulation) is still possible even when the network is offline.
[0272] The authenticity and integrity can be verified afterward (Hash + signature + trusted time synchronization), thus making the black-box evidence storage of XR space interaction not only a functional description, but also a closed-loop evidence chain capability that can be implemented in engineering.
[0273] This embodiment also provides the following specific scenario examples:
[0274] I. Standardized Process for Full-Scene Attention Interaction
[0275] (1) System initialization and benchmark establishment: The space computing terminal IMU extracts the gravity vector to establish the vertical benchmark, and at the same time combines geomagnetic or visual features to lock the horizontal yaw angle. Synthesize absolute attitude alignment operator .
[0276] (2) Full-scene coordinate perception: Automatic switching according to signal environment: GNSS / RTK is used outdoors, and UWB or visual SLAM is used indoors. Kalman filtering is used to suppress jumps between different precision sources and output continuous and stable fused coordinates. .
[0277] (3) Spatial Object Addressing (sDNS): The spatial computing terminal uses... The device requests the spatial object library within its current field of view from the Spatial Domain Name System (sDNS). The device informs its bounding volume via Object Broadcast. With control instruction set.
[0278] (4) Collision determination and disambiguation: The system constructs the gaze ray after gravity correction. Real-time computing and The result of the intersection is obtained. If multiple targets are hit, the depth-priority and occlusion culling mechanisms are activated to lock the final object. .
[0279] (5) Interactive state machine transition:
[0280] Lock-on: The duration of gaze reaches Once locked, the spatial computing terminal displays highlighted feedback at the edges of the spatial interaction.
[0281] Confirmation and execution: The operator issues instructions via a second gaze or voice.
[0282] Audit: The system verifies Spatial-ACL permissions and triggers XR space interaction black box evidence storage when high-risk commands are executed.
[0283] II. Application Scenarios Examples
[0284] Scenario 1: Industrial Anti-interference Maintenance (Strong Magnetic Field / Weak Indoor GNSS) When operators enter complex environments such as substations where satellite signals are unavailable or unstable, the system automatically switches to "Visual Semantics + UWB" mode. If the operator focuses on a high-voltage switch, the system will lock onto it and indicate the safety level, requiring secondary confirmation before executing the disconnect command. If the operator's gaze drifts to an adjacent energized area, the system will trigger a warning and freeze control permissions. Simultaneously, XR spatial interaction black-box data can be recorded for post-event analysis.
[0285] Scenario 2: Smart City Governance and Law Enforcement Evidence Collection (Outdoor Absolute Index) Urban management personnel wearing spatial computing terminals observe illegal buildings or municipal facilities. The system matches the target entity based on absolute geographic coordinates and a spatial object database, automatically retrieving planning base maps and historical work order information to achieve "retrieval upon observation".
[0286] If it is necessary to generate law enforcement records, the system can automatically associate the target ID, location and timestamp after confirmation to form a standardized event record.
[0287] Scenario 3: Age-Friendly Assistance (Remote Appliance Control) When elderly people are focused on appliances such as TVs / air conditioners, the system uses gravity correction to reduce accidental touches caused by posture changes and confirms the user's intention by observing their dwell time before completing the control. For high-risk devices (such as gas valves), identity verification and multiple confirmations are required to reduce risk.
[0288] The foregoing has shown and described the basic principles and main features of the present invention and its advantages. It will be apparent to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention.
[0289] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. A gaze-based interactive system based on gravity vector constraints and spatial topology association, characterized in that: Includes the following modules: The gaze acquisition module is used to acquire user gaze information and generate the original gaze direction vector; The attitude correction module is used to acquire the gravity vector and construct an attitude correction operator based on the gravity vector to correct the original gaze direction vector and obtain the gravity-aligned gaze direction vector. The spatial object indexing module is used to maintain a spatial object library, wherein the spatial object includes at least object identifier, spatial location, and geometric bounding volume information; The collision determination and locking module is used to construct a gaze ray based on the gravity-aligned gaze direction vector and perform collision determination between the gaze ray and the geometric bounding volume information to generate a target locking result; The protocol communication module is used to distribute control commands to the target space object and receive execution receipts according to the standardized interaction protocol.
2. The gaze interaction system based on gravity vector constraints and spatial topology association according to claim 1, characterized in that, The gaze interaction system is used to implement the following steps: S1. The spatial computing terminal collects user gaze information and generates the original gaze direction vector; S2. The gravity vector is obtained by the inertial measurement unit of the space computing terminal, and an attitude correction operator is constructed based on the gravity vector to correct the original gaze direction vector, so as to obtain the gravity-aligned gaze direction vector. S3. Construct a gaze ray based on the gravity-aligned gaze direction vector, and obtain the geometric bounding volume information of at least one spatial object in the spatial object library; S4. Perform spatial collision determination between the gaze ray and the geometric bounding volume information. When the hit condition and gaze dwell time threshold are met, generate target locking result. S5. Distribute control instructions to the spatial object corresponding to the target locking result according to the standardized interaction protocol, and receive the execution receipt fed back by the spatial object to complete the gaze interaction closed loop.
3. The gaze interaction system based on gravity vector constraints and spatial topology association according to claim 1, characterized in that: The gaze interaction system also includes a gaze interaction protocol, which specifies the data interaction process and message format for gaze interaction between the spatial computing terminal and spatial objects. The protocol includes at least: A spatial object broadcast message is used by a spatial object to announce its spatial existence and control capabilities. The spatial object broadcast message includes at least one or more of the following: object identifier, spatial anchor point information, semantic tag, geometric bounding volume parameters, and set of executable instructions. A gaze vector stream message is used by the space computing terminal to continuously output gaze ray information. The gaze vector stream message includes at least one or more of the following: gaze ray origin information, gaze direction vector information, gravity alignment identifier, and timestamp information. A target lock receipt message is used to feed back the lock result to the space computing terminal when the target object is gazed upon and locked. Command distribution message, used to send control commands to the target object after the target object is locked; An execution receipt message is used to provide feedback on the execution result after the target object executes the control command; wherein, the protocol also stipulates that: the spatial computing terminal performs attitude correction based on the gravity vector on the gaze direction vector, and performs spatial collision determination based on the corrected gaze ray and the geometric bounding volume parameters to complete the gaze interaction closed loop.
4. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 3, characterized in that: The gaze interaction system is also used to implement spatial object broadcasting and registration methods, including: Obtain the object identifier, semantic label, geometric bounding volume parameters, and set of executable instructions for a spatial object; Obtain spatial anchor point information of spatial objects, wherein the spatial anchor point information includes any one or more of absolute coordinate anchor points, local coordinate anchor points, or visual feature anchor points; Generate a spatial object broadcast message and broadcast the spatial object broadcast message to the spatial index center or spatial computing terminal; The spatial indexing center writes the spatial objects into the spatial object library based on the spatial object broadcast message, thereby enabling the spatial computing terminal to perform collision determination and lock the spatial objects based on the gaze ray and the spatial object library.
5. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 2, characterized in that: The original gaze direction vector includes a compensation term for the difference between the user's individual visual axis and optical axis, which is obtained through user calibration parameters, online adaptive estimation, or preset model parameters; S2 further includes constructing a yaw alignment operator based on magnetometer output or visual orientation reference to lock the horizontal orientation of the gravity-aligned gaze direction vector. The direction vector of the gaze ray is obtained by the combined action of the gravity alignment operator and the yaw alignment operator, so that the gaze ray satisfies the three-degree-of-freedom attitude stability. The geometric bounding box information of the spatial object includes, but is not limited to, any one or more of the following: axis-aligned bounding box (AABB), oriented bounding box (OBB), sphere, cylinder, or polyhedron.
6. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 5, characterized in that: When the gaze ray hits multiple spatial objects, the objects are sorted by depth according to the distance between the ray and the nearest intersection point, and the spatial object with the smallest distance is locked first. The multi-target hit further combines the results of depth sensor, binocular vision, structured light, ToF, UWB ranging or visual depth estimation to perform occlusion culling, so as to reduce the locking probability of occluded objects or cullate occluded objects. The multi-target hit further combines spatial object semantic correlation determination for locking and disambiguation. The semantic correlation determination includes any one or more of the following: the matching degree between the target category and the current task, the correlation between the target and the user's action, the topological adjacency relationship between the target and historically locked objects, or the consistency of visual semantic confidence.
7. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 2, characterized in that: The gaze dwell time threshold in S4 is used to switch between the candidate hit state and the stable lock state. The step prior to S5 further includes an intent confirmation step, wherein the intent confirmation is achieved through any one of voice confirmation, blink confirmation, gesture confirmation, secondary gaze confirmation, or multimodal combination confirmation. The spatial location of the spatial object is obtained by fusing multi-source positioning / sensing information, which includes any one or more of satellite positioning, RTK, UWB, WiFi-FTM, Bluetooth AoA, or visual SLAM. The fusion of multi-source localization / sensing information is achieved through Kalman filtering or factor graph optimization to suppress pose jumps or jitter caused by switching of heterogeneous localization sources; The object identifiers of the spatial object library are generated by a spatial domain name resolution mechanism, which is used to uniformly map heterogeneous coordinate representations into spatial semantic object identifiers. The standardized interaction protocol includes spatial object broadcast messages, which contain at least one or more of the following: object identifier, spatial anchor point information, semantic tags, geometric parameters, and a set of executable instructions. When the spatial object is a dynamic target, timing alignment is performed based on the target pose timestamp and the gaze ray timestamp, and motion prediction compensation is performed based on the target velocity or acceleration information before performing the spatial collision determination in step S4. Before distributing control instructions in S5, spatial permission verification is performed, which includes at least one or more of the following: gaze hit validity verification, operator identity authentication verification, and operator spatial legitimacy verification.
8. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 2, characterized in that: When an abnormal command is detected, a high-risk accidental touch, an impact event, or an access denial event is detected, the XR space interaction black box evidence is triggered. The interaction black box evidence records at least one or more of the physical sensor data, interaction trajectory data, and semantic evidence data. The XR space interaction black box evidence storage generates a data digest and performs signature encapsulation in the trusted execution environment (TEE) on the edge side, and asynchronously synchronizes the data digest to the edge node or trusted evidence storage service to form multiple copies of the evidence storage; The semantic evidence data supports data anonymization processing, which includes saving only one or more of the target object feature descriptors, semantic segmentation ROIs, bounding box information, or semantic category confidence, without saving the complete original image or video stream. The XR spatial interaction black box evidence storage obtains a reliable timestamp through satellite time synchronization, network time synchronization, or a clock synchronization server, and binds the timestamp with the data digest to achieve the integrity of a verifiable chain of evidence.
9. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 8, characterized in that: When a dynamic target is occluded, semantic recognition fails, or the positioning signal is momentarily interrupted, the system performs trajectory extrapolation based on the pose sequence and motion vector before the target disappears within a preset lock-out time threshold ΔT_hold to maintain the target lock state; and performs recapture matching to restore the lock when the target reappears, and falls back to the candidate hit or exit state when ΔT_hold is exceeded. The spatial domain name resolution mechanism supports multi-level resolution methods, including full resolution in the cloud, regional caching at edge nodes, and local pre-fetching caching on the terminal. This enables the spatial computing terminal to complete local collision detection and target locking based on a pre-downloaded subset of spatial object topology under weak network or offline conditions, and supports the real-time discovery of temporarily appearing local spatial objects through broadcast protocols.
10. A gaze interaction system based on gravity vector constraints and spatial topology association according to claim 9, characterized in that: The spatial domain name resolution mechanism (sDNS) supports the pose subscription update of dynamic spatial objects. When the spatial object is a dynamic object, its geometric bounding volume parameters are pushed to the terminal cache with the object pose change at a preset update frequency or threshold triggering method, and carry the object pose timestamp for time sequence alignment with the gaze ray timestamp to ensure the consistency of spatial collision determination. When multiple spatial computing terminals lock the same spatial object simultaneously, the spatial domain name resolution mechanism (sDNS) or edge nodes issue the object locking status based on preset permission priorities, first-occupancy logic, or queuing strategies, and return an occupation prompt or rejection receipt to the non-occupying terminals to avoid concurrent command conflicts and misoperations.