Multi-source physical true value and space-time alignment cross-space embodiment collaborative control system and method
By using a cross-space embodied cooperative control system that aligns multi-source physical truth with spatiotemporal data, the problems of insufficient spatial consistency, temporal consistency, semantic task mapping, security auditing, and traceability in existing cross-space cooperative control technologies are solved, enabling stable, reliable, and transferable task execution under complex environments and strong interference conditions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUHAN HUACHUANG HIGHLIGHT DIGITAL TECHNOLOGY CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-07-21
Smart Images

Figure CN122431218A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of spatial computing, embodied intelligence and remote collaborative control technology, specifically to a cross-space embodied collaborative control system and method based on multi-source physical truth fusion, spatiotemporal alignment, spatial anchor point sharing, task-level instruction mapping, anti-interference auditing and tamper-proof evidence storage.
[0002] Furthermore, the present invention also relates to technologies such as multi-source positioning / attitude fusion, time anchor synchronization, semantic task control, standardized task interface, low-latency feedback transmission, disconnection autonomy, edge-cloud collaborative computing, and multi-entity collaborative scheduling between wearable sensing end and embodied execution end, which can be applied to remote rescue, hazardous operations, industrial inspection, smart agriculture, unmanned interactive experience and other cross-space embodied collaborative control scenarios. Background Technology
[0003] With the development of wearable terminals, extended reality devices, AI glasses, unmanned vehicles, drones, robots, robotic arms, and special-operation equipment, remote collaborative control is gradually expanding from traditional joystick-style remote control to human-machine collaborative control based on first-person perspective perception, spatial positioning, semantic recognition, and task-level commands. Operators can obtain images of the remote environment, target semantics, and execution status through wearable sensing devices, and issue task-level control commands such as search, approach, grab, deliver, inspect, obstacle avoidance, and formation to the embodied execution device.
[0004] However, existing cross-space collaborative control systems mainly revolve around single teleoperation links, single actuator control, or motion control within a local coordinate system. They have yet to form a unified foundation capable of simultaneously addressing spatial consistency, temporal consistency, semantic task mapping, security auditing, autonomous disconnection, and traceable evidence storage. Especially in scenarios such as remote rescue, hazardous operations, industrial inspection, smart agriculture, and unmanned interactive experiences, the virtual content, human intentions, and environmental semantic targets on the operator's side often operate under different coordinate systems, time bases, and control protocols than the physical motion of the actuator, making it difficult for the system to operate stably, reliably, and transferably.
[0005] The existing technology has at least the following problems:
[0006] 1. Coordinate system fragmentation leads to spatial inaccuracies across ends.
[0007] In existing teleoperation or remote collaboration systems, the operator typically relies on wearable terminals, visual SLAM, local maps, or video footage to form a local spatial understanding; while the execution end relies on its own GNSS / RTK, UWB, odometry, VIO, SLAM, or local sensors to establish another set of positioning and motion control coordinates. Because the two lack a unified spatial reference that can be shared, verified, and version-managed, problems such as local coordinate drift, scale inconsistency, anchor point mismatch, or discontinuous spatial reference switching can easily occur.
[0008] In this scenario, the target location, obstacle boundaries, or task area seen by the operator in the wearable terminal may not match the actual reachable path, robotic arm operating range, or unmanned vehicle trajectory, causing problems such as "visible but inaccurate control," "pointing to the target but execution deviation," and "misalignment between virtual annotations and physical objects." Especially in environments involving indoor / outdoor transitions, GNSS obstruction, weak texture environments, numerous dynamic obstacles, or strong interference, a single positioning source or local coordinate system struggles to consistently provide stable cross-terminal spatial consistency.
[0009] 2. The lack of a unified time base between the video and control links can easily lead to operational drift.
[0010] In immersive remote control, operators typically make decisions based on video frames, sensor observations, or semantic summaries transmitted from the execution end. However, in existing systems, the video transmission link and the control command link are often transmitted independently, resulting in varying degrees of latency and jitter between video acquisition, encoding, transmission, decoding, rendering, and control command generation, transmission, and execution.
[0011] Due to the lack of a unified time anchor and command-feedback phase alignment mechanism, the image upon which the operator issues commands may lag behind the current physical state of the execution terminal, while the execution terminal actually executes the commands at a different point in time. This problem can lead to phenomena such as excessive direction correction, wobbling in narrow spaces, failure to approach the target, missed capture opportunities, and unstable path following. Under low-bandwidth, multi-hop networks, long-distance transmission, or highly interference-prone links, these timing inconsistencies will be further amplified and directly affect security.
[0012] 3. Lack of semantic priority in feedback transmission under low bandwidth or complex environments.
[0013] Current video feedback primarily relies on whole-frame compression or fixed-bitrate transmission, with limited integration of feedback information layering based on operator gaze area, environmental semantic targets, risk areas, or task objectives. When bandwidth is limited or the link is unstable, critical target areas may degrade along with the background, preventing operators from promptly identifying injured personnel, hazards, obstacle boundaries, robotic arm contact points, or task objects.
[0014] Meanwhile, existing feedback feedback typically lacks a time stamp directly linked to control commands, making it difficult to determine which frame, semantic digest, or execution state a particular control command was generated from. This results in a lack of reliable evidence for subsequent auditing, review, and accountability, and also makes it difficult for the system to perform effective phase compensation or forward state prediction under conditions of latency and jitter.
[0015] 4. There is a gap between low-level control and task-level collaboration.
[0016] Existing remote control systems mostly focus on low-level control parameters such as speed, steering, joystick input, and robotic arm joint input, requiring operators to perform continuous, high-frequency micro-manipulation. In complex environments or high-risk scenarios, this approach demands high levels of operator attention, experience, and reaction speed, making it prone to errors or excessive cognitive load.
[0017] With the development of AI glasses, spatial computing, and embodied intelligence, operators increasingly need to issue task-level commands such as "search and approach," "follow the target," "maintain a safe distance," "grab and deliver," and "avoid danger sources" through gaze, gestures, posture, voice, virtual content events, and environmental semantic goals. However, existing systems lack a multimodal command mapping mechanism that unifies virtual content event streams, human intentions, and environmental semantic goals into task-level instructions, dynamic instructions, or local planning constraints. They also lack a unified control layer that drills down from high-order semantic tasks to specific motion constraints at the execution end.
[0018] 5. Fragmented protocols across different execution endpoints, lacking standardized, portable task interfaces.
[0019] In practical applications, actuators such as unmanned vehicles, drones, quadruped robots, robotic arms, tracked platforms, and special-operation equipment typically have different kinematic structures, control interfaces, task capabilities, and safety constraints. Existing systems are often custom-developed for a single actuator or a single scenario, making it difficult to reuse task descriptions, interface protocols, and control commands.
[0020] Therefore, it is difficult for the same wearable sensing terminal or remote command system to adapt to different types of embodied execution terminals simply by switching task templates. It is also difficult for different devices to share unified target identifiers, risk levels, spatial anchors, temporal anchors, geofences, safety corridors, action constraints or path constraints, resulting in high deployment costs and poor system scalability across devices, industries, and scenarios.
[0021] 6. Lack of a provable security loop when there is strong interference, disconnection, or descent of the stationary signal.
[0022] In rescue operations, underground spaces, disaster sites, hazardous industrial operations, environments with strong electromagnetic interference, or large-scale agricultural operations, communication link degradation, packet loss, bit errors, latency jitter, insufficient positioning information, and sensor anomalies are common problems. Existing remote control systems typically employ simple strategies such as shutdown, return to home, or continuing to execute the last command when link quality deteriorates. These systems lack a mechanism for comprehensive auditing of link quality, positioning residuals, anchor point consistency, sensor confidence, command risk level, and the status of the executing end.
[0023] Furthermore, existing systems typically only enter a disconnection strategy after a complete link failure, lacking the ability to predict disconnection based on link deterioration trends, latency jitter, packet loss rate, bit error rate, or location signal degradation. They also lack mechanisms to preemptively issue autonomous scripts, fallback paths, security boundaries, task continuation status, or last effective strategies before a complete disconnection. This results in the execution end being in an uncertain state during a sudden link failure, making it difficult to maintain task continuity, avoid obstacles, fall back, or wait safely.
[0024] 7. Lack of unified conflict arbitration in multi-entity collaboration scenarios.
[0025] In collaborative scenarios involving multiple unmanned vehicles, drones, robots, robotic arms, or operators, multiple execution ends may simultaneously receive control commands from the same task template, the same operator, or different operators. Without a unified shared anchor point coordinate system, spatial topological constraints, and priority arbitration mechanisms, problems such as path overlap, formation convergence, insufficient minimum safe distance, task area conflicts, or interference between high-risk actions can easily arise.
[0026] Existing systems' local obstacle avoidance or single-entity path planning are insufficient to fully address global consistency issues across entities, operators, and task phases. They also struggle to record arbitration reasons, demotion results, and alternative trajectories when conflicts occur, resulting in inadequate security, interpretability, and traceability in multi-entity collaboration.
[0027] 8. The accident liability and high-risk actions lack a complete chain of evidence.
[0028] In scenarios such as remote rescue, hazardous materials handling, industrial inspection, emergency response, special operations, agricultural automation, and unmanned interactive experiences, once a collision, boundary crossing, mis-grabbing, mis-triggering, disconnection, autonomous failure, or high-risk abnormal action occurs, it is necessary to be able to trace back what the operator saw at the time, what the system identified, what task description or control instructions were issued, why the audit module passed / downgraded / rejected / shut down, what the actual feedback from the execution end was, and whether the anchor point version and time base are consistent.
[0029] Existing systems typically only contain scattered logs, video recordings, or device operational status records. They lack a spatiotemporal black-box mechanism that correlates and records virtual content status, environmental semantic targets, task description data, control commands, audit receipts, execution feedback, link / confidence summaries, spatial anchor version numbers, and temporal anchors. They also lack tamper-proof protection measures such as chained hashing, digital signatures, or trusted execution environments. Therefore, they are significantly inadequate in incident review, liability determination, third-party auditing, and compliance monitoring.
[0030] 9. Wearable terminals suffer from limitations in computing power, power consumption, and heat dissipation, resulting in insufficient edge-cloud collaboration.
[0031] Wearable sensing devices, especially AI glasses or lightweight head-mounted devices, are limited by size, power consumption, heat dissipation, and battery life, making it difficult to run high-computational tasks such as high-precision SLAM, real-time semantic recognition, global map optimization, loop closure detection, high-density point cloud processing, prediction compensation, and security auditing stably for extended periods. Existing systems either migrate a large amount of computing to the cloud, leading to excessive reliance on network quality, or they over-rely on simplified edge-side algorithms, resulting in insufficient stability in localization, recognition, and control.
[0032] There is a lack of a collaborative mechanism that can dynamically allocate tasks among wearable sensing devices, cloud / edge computing devices, and embodied execution devices by considering factors such as remaining battery power, core temperature, link latency, link jitter, computing load, and positioning / perception confidence. Therefore, existing systems struggle to maintain a minimum available closed loop under different network conditions, different computing power levels, and different security levels.
[0033] In summary, current technologies lack a cross-space embodied collaborative control system and method that, based on multi-source physical truth fusion and spatiotemporal alignment, can simultaneously achieve spatial anchor sharing, temporal anchor synchronization, multimodal / semantic task mapping, standardized task interfaces, anti-interference auditing, disconnection prediction, disconnection autonomy, low-latency ROI feedback, multi-entity conflict arbitration, and tamper-proof evidence storage. Therefore, it is necessary to propose a new technical solution to improve the spatial consistency, temporal consistency, security, reliability, portability, and traceability of cross-space embodied collaborative control in complex environments. Summary of the Invention
[0034] To address the shortcomings of existing technologies, this invention provides a multi-source physical truth and spatiotemporal aligned cross-space embodied collaborative control system and method. By establishing a unified multi-source physical truth coordinate system between the wearable sensing end and the embodied execution end, and combining spatial anchor point sharing, temporal anchor points, task-level instruction mapping, anti-interference auditing, disconnection autonomy, standardized task interfaces, and a spatiotemporal black-box evidence storage mechanism, the virtual content event flow, human intention flow, environmental semantic targets on the operator side, and the physical motion flow on the execution end are mapped, audited, executed, and stored under the same spatiotemporal reference. This achieves reliable, transferable, auditable, and traceable cross-space embodied collaborative control capabilities.
[0035] To achieve the above objectives, the present invention adopts the following technical solution:
[0036] This invention provides a multi-source physical truth and spatiotemporal alignment cross-space embodied collaborative control system. The control system includes a wearable sensing terminal, an embodied execution terminal, a unified spatiotemporal alignment engine, a multimodal instruction mapping engine, a security protection layer, and an evidence storage module.
[0037] The wearable sensing terminal is used to acquire the operator's spatial pose and / or environmental perception data, and to acquire human features such as gaze, gestures, and postures, and is used to output virtual content event streams or control intentions.
[0038] The embodied execution terminal is used to execute task-level instructions or dynamic instructions and to return execution status and environmental feedback.
[0039] The unified spatiotemporal alignment engine is used to fuse at least two types of heterogeneous positioning / attitude information to establish a multi-source physical truth coordinate system, and to synchronize spatial anchor points between the wearable sensing end and the embodied execution end to maintain shared coordinate mapping.
[0040] The multimodal instruction mapping engine is used to map the virtual content event stream and / or human features to the task-level instructions or dynamic instructions of the embodied execution end, and to attach time anchors to the instructions.
[0041] The security protection layer includes an anti-interference audit module, which is used to gate the execution of the instruction based on link quality, location / perception confidence and / or instruction risk level.
[0042] The evidence storage module is used to associate and record virtual content status, instructions, audit decisions and execution feedback, and generate an immutable evidence package when a trigger event occurs.
[0043] Furthermore, the multi-source physical ground truth coordinate system is established by fusing at least two types of observations from RTK / GNSS, UWB, VPS, SLAM / VIO, and IMU. Moreover, the multi-source physical ground truth coordinate system can further fuse at least one of the following: sensing measurements from the 6G integrated sensing network side, channel state information derived observations from Wi-Fi links, and Bluetooth ranging or angle measurement derived observations, as a redundancy constraint on the RTK / GNSS, UWB, VPS, SLAM / VIO, or IMU observations. The unified spatiotemporal alignment engine performs confidence assessment, dynamic weighting, or anomaly removal on various types of observations based on at least one of observation residuals, anchor point consistency, reprojection error, link delay, or jitter, to enhance the spatial alignment continuity and robustness under conditions of occlusion, visual blind spots, or strong interference.
[0044] Furthermore, the multimodal instruction mapping engine includes at least one or more of the following: narrative / rhythm-driven mapping, human capture / gaze-driven mapping, and a semantic control layer. Specifically, the narrative / rhythm-driven mapping is used to extract temporal features such as beats, audio envelopes, key events, or virtual trigger points from virtual content and map them as execution-end dynamic parameters, path segment constraints, or formation constraints; the human capture / gaze-driven mapping is used to map the operator's gaze, gestures, posture, gait, bioelectrical signals, or combinations thereof, as task-level actions, fine-grained operation instructions, or action trajectories at the execution end; and the semantic control layer is used to convert environmental semantic targets, task roles, risk levels, or interactive object recognition results into task-level commands and further map them as motion constraints, operation constraints, or path constraints at the execution end.
[0045] Furthermore, the control system also includes a standardized task interface and / or an execution end adaptation layer; the standardized task interface is used to generate or receive task description data, the task description data including at least one or more of the following: target identifier, task type, risk level, spatial anchor point, time anchor point, geofence, safety corridor, permission level, action constraint, or path constraint; the execution end adaptation layer is used to convert the task description data into task-level instructions, dynamic instructions, or local planning constraints executable by the corresponding execution end based on the capability profile of different embodied execution ends, so that the same task description can be adapted to one or more execution ends of unmanned vehicles, drones, robots, robotic arms, or special operation equipment.
[0046] Furthermore, the anti-interference audit module is used to perform real-time evaluation of at least one of the following: link quality, sensor confidence, positioning residual, anchor point consistency, command risk level, or execution end status, and to gate the execution of control commands; the gate result includes at least one of passing, downgrading, rejecting, or safe shutdown; wherein, the downgrading includes at least one of rate limiting, geofence tightening, safe corridor constraint, control domain switching, task suspension, request confirmation, or switching to a local autonomous policy.
[0047] Furthermore, the anti-interference audit module is also used to predict connection failures. When the link quality shows a continuous downward trend, the latency or jitter exceeds a preset threshold, the packet loss rate or bit error rate exceeds a preset threshold, or the positioning / perception confidence is lower than a preset threshold, the control system sends an autonomous script, fallback path, security boundary, task continuation status, or last effective strategy to the embodied execution terminal before the link is completely interrupted. When the communication link is damaged or the positioning confidence is insufficient, the embodied execution terminal enters the disconnection autonomous mode based on the autonomous script, fallback path, security boundary, task continuation status, or last effective strategy to maintain task continuity, obstacle avoidance, fallback, or safe waiting under the security boundary.
[0048] Furthermore, the control system also includes a spatial anchor point sharing mechanism, which includes anchor point generation, anchor point synchronization and version control, cross-terminal coordinate mapping, and multi-entity consistency maintenance. Specifically, the anchor point generation is used to generate a set of spatial anchor points at the wearable sensing end and / or the embodied execution end based on environmental features, manually calibrated points, UWB base station geometry, RTK reference points, 6G integrated sensing derived observations, or Wi-Fi channel state derived observations. Each spatial anchor point includes at least an anchor point identifier (AnchorID), anchor point pose, anchor point confidence, anchor point version number, and validity period. When either end detects anchor point drift, anchor point matching residual exceeding a preset threshold, anchor point confidence below a preset threshold, or anchor point version inconsistency, the control system triggers an anchor point update, realignment request, or control degradation.
[0049] Furthermore, when there are multiple embodied executors and / or multiple operators, the control system maintains the geometric consistency shared by multiple entities based on a unified set of anchor points. When multiple embodied executors receive potentially conflicting control commands, the control system constructs spatial topological constraints based on the shared anchor point coordinate system and arbitrates according to preset priorities, risk levels, task stages, or minimum safe distances. Among these, executors with higher priorities or higher risk levels retain the original commands, while the remaining executors automatically downgrade their priority and enter obstacle avoidance waiting mode, follow-up mode, alternative trajectory execution mode, or delayed execution mode, and the arbitration results are written into the audit receipt and evidence index.
[0050] The present invention also provides a cooperative control method using the above-described control system, comprising:
[0051] A multi-source physical true coordinate system is established by fusing at least two types of positioning / attitude observations, and spatial anchor points are synchronized to form a shared coordinate mapping;
[0052] Analyze virtual content event streams, environmental semantic targets, and / or operator human characteristics to generate control intent or task description data;
[0053] Based on the control intent or task description data, generate task-level instructions, dynamic instructions, or local planning constraints, and attach time anchors to the instructions;
[0054] When the sensing end receives a delayed video frame carrying a temporal anchor, it calculates the difference between the current 6-DoF pose of the sensing end and the pose corresponding to the temporal anchor, performs asynchronous spatial warp (ASW) reprojection, and compensates for the viewpoint of the video frame.
[0055] The command is gated based on link quality, location / perception confidence, command risk level, or execution end status, and the command is executed by at least one of the following: pass, downgrade, reject, or stop.
[0056] The virtual content state, environmental semantic targets, task description data, instructions, gating decisions and execution feedback are recorded to form a traceable chain of evidence, and the evidence package is sealed when a triggering event occurs.
[0057] Furthermore, the execution feedback includes at least one of video frames, sensor observations, state telemetry, or semantic summaries; wherein the video frames and / or semantic summaries are ROI-layered encoded or semantically prioritized backhaul based on the gaze region, semantic target, or risk region, so that key regions, key targets, or key semantic information are backhauled with higher quality or priority than background regions, and time anchors are embedded in the corresponding video frames and / or semantic summaries; the control instructions are bound to the time anchors of the video frames seen by the operator when the instructions are generated, and the execution end performs phase verification, state forward prediction, or instruction compensation based on the time anchors, link delay estimates, and local motion states; when the phase deviation, prediction uncertainty, link jitter, or reprojection uncertainty exceeds a preset threshold, degradation, confirmation request, execution rejection, or safe shutdown is triggered; the evidence package includes virtual rendering summaries or narrative events, environmental semantic targets, task description data, instructions, audit receipts, execution feedback, link / confidence summaries, anchor version numbers, and time anchors, and uses chain hashing, digital signatures, and / or a trusted execution environment to ensure that the evidence is tamper-proof.
[0058] The beneficial effects of this invention are:
[0059] 1. This invention aligns the wearable sensing end and the embodied execution end in the same physical coordinate system through multi-source physical truth fusion, confidence assessment, dynamic weighting, anomaly elimination and spatial anchor point sharing. This reduces positioning deviations caused by scale drift, anchor point mismatch, inconsistency of reference or sensor anomalies, and improves the accessibility and repeatability of actions such as approaching, grasping, grouping, path following, remote rescue, and industrial inspection.
[0060] 2. This invention, through time anchoring, ROI layered backhaul, command-video frame binding, ASW reprojection, phase verification, state forward prediction, and command compensation, ensures that the perceived information used by the operator to make decisions and the actions of the execution end are in an aligned temporal relationship. In the presence of latency, jitter, narrowband backhaul, or feedback lag, it can reduce the risks of operation drift, overcorrection, and misjudgment, and improve the stability and safety of cross-space immersive control.
[0061] 3. This invention transforms virtual content event streams, operator human intentions, and environmental semantic goals into task-level instructions, dynamic instructions, motion constraints, or operational constraints through narrative / rhythm-driven mapping, human body capture / gaze-driven mapping, and a semantic control layer. This enables the system to upgrade from traditional low-level teleoperation to embodied collaborative control driven by content, intention, and semantic tasks.
[0062] 4. This invention enables the same task description data to be adapted to different specific execution terminals such as unmanned vehicles, drones, robots, robotic arms, or special operation equipment through a standardized task interface and execution terminal adaptation layer. The upper-layer task switching does not require changes to the underlying spatiotemporal alignment and security audit core, thereby enhancing the migration capability and large-scale deployment capability across scenarios, industries, and execution entities.
[0063] 5. This invention, through anti-interference auditing, gating, disconnection prediction, advance distribution of autonomous scripts, disconnection autonomy, and multi-level degradation strategies, can still maintain task continuity, obstacle avoidance, rollback, or safe waiting under conditions of link damage, insufficient positioning information, sensor anomalies, or strong interference, thus avoiding unpredictable actions at the execution end.
[0064] 6. This invention uses audit receipts, black-box indexes, chained hashes, digital signatures, and / or trusted execution environments to associate and record virtual content states, environmental semantic targets, task description data, control instructions, gating decisions, execution feedback, link / confidence summaries, anchor version numbers, and time anchors in an unalterable manner, enabling the traceability, verification, and auditing of accident responsibility, abnormal actions, disconnection autonomy, and high-risk operation processes. Attached Figure Description
[0065] Figure 1 This is a block diagram of the overall architecture of the multi-source physical truth and spatiotemporal alignment cross-space embodied collaborative control system of the present invention;
[0066] Figure 2 This is a logical diagram of the four-layer architecture of the present invention;
[0067] Figure 3 This is a flowchart of the multi-source physical truth fusion and spatial anchor point sharing of the present invention;
[0068] Figure 4 This is a schematic diagram of the multimodal instruction mapping engine and standardized task interface of the present invention;
[0069] Figure 5 This is a flowchart illustrating the anti-interference audit gating and degradation decision-making process of this invention.
[0070] Figure 6 This is a state machine diagram of the disconnection prediction and disconnection self-governance of the present invention;
[0071] Figure 7 This is a flowchart of the low-latency backhaul and command-video phase alignment process of the present invention;
[0072] Figure 8 This is a diagram of the spatiotemporal black box and tamper-proof evidence storage structure of the present invention;
[0073] Figure 9 This is a schematic diagram of the standardized task interface and execution end adaptation layer of the present invention;
[0074] Figure 10 This is a schematic diagram of the edge-cloud collaborative computing and task offloading of the present invention. Detailed Implementation
[0075] To make the technical solution, technical means, technical effects and implementation path of the present invention clearer, the following description is provided in conjunction with the appendix. Figure 1 To be continued Figure 10 The present invention will be further described below. It should be understood that the following embodiments are only used to explain the present invention and are not intended to limit the scope of protection of the present invention. Without departing from the core concept of the present invention, those skilled in the art can make equivalent substitutions or combinations of modules, steps, data fields, communication links, sensor types, fusion algorithms, audit rules and evidence storage methods in various embodiments.
[0076] This invention provides a multi-source physical truth and spatiotemporal aligned cross-space embodied collaborative control system and method. The system comprises a wearable sensing terminal, a cloud / edge computing terminal, and an embodied execution terminal, forming a cross-space collaborative control closed loop. Through multi-source physical truth fusion, spatial anchor point sharing, temporal anchor point synchronization, multimodal command mapping, standardized task interfaces, anti-interference auditing, disconnection autonomy, low-latency feedback transmission, and spatiotemporal black-box evidence storage, the system enables the virtual content event flow, human intention flow, environmental semantic targets, and physical motion flow of the execution terminal to be mapped, audited, executed, and stored under the same spatiotemporal reference.
[0077] The multi-source physical truth value refers to a shareable and verifiable spatial reference formed by at least two types of heterogeneous positioning / attitude observations. These heterogeneous positioning / attitude observations include, but are not limited to, RTK / GNSS, UWB, VPS, SLAM / VIO, IMU, odometry, visual / laser sensors, 6G integrated sensing derived observations, Wi-Fi channel state information derived observations, and Bluetooth ranging or angle measurement derived observations. The system performs time alignment, confidence assessment, dynamic weighting, anomaly removal, and spatial anchor point constraints on these observations, enabling the wearable sensing end and the embodied execution end to maintain spatial consistency within a unified truth coordinate system.
[0078] The time anchor point refers to a data identifier used to identify the time reference of virtual content events, human intention input, task description data, control commands, video frames, sensor observations, semantic summaries, and execution feedback. Through the time anchor point, the system can correlate the feedback screen seen when the operator generates a command, the corresponding environmental semantic target, the time the control command is generated, and the actual execution state of the execution end, thereby providing a timing basis for command-video phase alignment, ASW reprojection, state forward prediction, anti-interference auditing, and black-box evidence preservation.
[0079] The spatial anchor point refers to a data object used to maintain a shared spatial reference among wearable sensing devices, cloud / edge computing devices, and embodied execution devices. Each spatial anchor point includes at least an anchor point identifier, anchor point pose, anchor point confidence level, anchor point version number, and validity period. The system can maintain consistent spatial understanding across different terminals in dynamic, occluded, or strong interference environments through anchor point generation, anchor point synchronization, version control, anchor point drift detection, and realignment triggering mechanisms.
[0080] The multimodal instruction mapping refers to the process of converting virtual content event streams, human features, environmental semantic targets, or task inputs into task-level instructions, dynamic instructions, or local planning constraints that can be executed by the embodied execution end. The virtual content event stream may include narrative beats, audio envelopes, key events, virtual trigger points, or plot stages; the human features may include gaze, gestures, posture, gait, bioelectrical signals, or combinations thereof; the environmental semantic targets may include wounded personnel, hazards, obstacles, passageways, valves, meters, tools, work areas, or other identifiable objects.
[0081] The aforementioned anti-interference audit refers to the process by which the system, based on link quality, sensor confidence, positioning residuals, anchor point consistency, instruction risk level, execution end status, permission level, or task stage, determines whether a control instruction should be approved, downgraded, rejected, or safely shut down. The anti-interference audit can also output audit receipts, which record the decision type, triggering reason, key input indicators, downgrade action, associated instruction identifier, time anchor point, anchor point version number, and black-box index.
[0082] The aforementioned disconnection autonomy refers to the system's operational mechanism that, when a communication link is damaged, location information is insufficient, or the link is about to be interrupted, enables the executor to maintain task continuity, avoid obstacles, roll back, or wait safely within the security boundary, based on pre-downloaded autonomous scripts, fallback paths, security boundaries, task continuation status, or last effective strategy.
[0083] The spatiotemporal black-box evidence storage refers to a mechanism that associates and records virtual content state, environmental semantic targets, task description data, control instructions, audit receipts, execution feedback, link / confidence summaries, anchor version numbers, and time anchors, and generates an immutable evidence package when a trigger event occurs. The immutable evidence package can be protected using chained hashing, digital signatures, trusted execution environments, trusted timestamps, or a combination thereof.
[0084] Implementation Method 1: System Overall Architecture
[0085] like Figure 1 As shown, the control system of the present invention includes a wearable sensing terminal, a cloud / edge computing terminal, an embodied execution terminal, a unified spatiotemporal alignment engine, a multimodal instruction mapping engine, a security protection layer, and an evidence storage module.
[0086] The wearable sensing terminal is used to acquire the operator's spatial pose and / or environmental perception data, and to acquire the operator's gaze, gestures, posture, gait, bioelectrical signals, or combinations thereof, and is used to output virtual content event streams, environmental semantic targets, task selection inputs, or control intentions. The wearable sensing terminal can be AI glasses, AR / MR / VR headsets, helmet-type devices, handheld spatial perception terminals, or other terminals with spatial perception capabilities. The wearable sensing terminal can also receive video frames, sensor observations, state telemetry, or semantic summaries transmitted back from the embodied execution terminal to form an immersive decision-making loop for the operator.
[0087] The cloud / edge computing terminal is used to undertake one or more functions including global scheduling, computation offloading, model inference, spatial anchor synchronization, link quality assessment, task policy distribution, semantic recognition, ROI encoding, audit management, and evidence aggregation. The cloud / edge computing terminal can be an edge server, a field computing node, a cloud platform, a private network computing node, or a local computing unit deployed near the execution terminal. In some implementations, the cloud / edge computing terminal is an optional module; when the link is restricted or the security level is high, the critical control closed loop can be devolved to the embodied execution terminal or wearable sensing terminal for local execution.
[0088] The embodied execution end is used to execute task-level instructions, dynamic instructions, or local planning constraints, and to transmit execution status and environmental feedback. The embodied execution end includes, but is not limited to, unmanned vehicles, drones, quadruped robots, wheeled robots, tracked platforms, robotic arms, gimbals, special-purpose equipment, agricultural equipment, or industrial inspection equipment. The embodied execution end may include a self-localization and perception unit, a motion control unit, an execution mechanism, a local obstacle avoidance unit, a local autonomous unit, and a safe stopping unit; wherein, the self-localization and perception unit may include a camera, depth camera, LiDAR, IMU, UWB, RTK / GNSS, odometry, force sensor, gas sensor, temperature sensor, or a combination thereof.
[0089] The unified spatiotemporal alignment engine is used to fuse at least two types of heterogeneous positioning / attitude information to establish a multi-source physical truth coordinate system, and to synchronize spatial anchor points between the wearable sensing end and the embodied execution end to maintain shared coordinate mapping. The unified spatiotemporal alignment engine is also used to attach or associate temporal anchor points to virtual content events, human intention input, task description data, control commands, video frames, sensor observations, semantic summaries, and execution feedback, enabling the above data to be compared, compensated, audited, and stored under the same temporal reference.
[0090] The multimodal instruction mapping engine is used to map virtual content event streams, human features, environmental semantic targets, or task inputs into task-level instructions, dynamic instructions, or local planning constraints of the embodied execution end, and to attach time anchors to the instructions. The multimodal instruction mapping engine may include one or more of the following: a narrative / rhythm-driven mapping unit, a human capture / gaze-driven mapping unit, a semantic control layer, a task description generation unit, and an execution end adaptation unit. The semantic control layer is used to convert environmental semantic targets, task roles, risk levels, or interactive object recognition results into task-level commands, and further map them into motion constraints, operational constraints, or path constraints of the execution end.
[0091] The standardized task interface is used to generate or receive task description data. The task description data includes at least one or more of the following: target identifier, task type, risk level, spatial anchor point, temporal anchor point, geofence, safety corridor, permission level, action constraints, or path constraints. The execution end adaptation layer is used to convert the task description data into task-level instructions, dynamic instructions, or local planning constraints executable by the corresponding execution end, based on the capability profiles of different execution ends. Through this standardized task interface, the same task description data can be adapted to one or more execution ends, including unmanned vehicles, drones, robots, robotic arms, or special-operation equipment.
[0092] The security protection layer includes an anti-interference audit module. This module performs real-time evaluation of at least one of the following: link quality, sensor confidence, positioning residual, anchor point consistency, command risk level, execution end status, permission level, or task stage, and gates the execution of control commands. The gating result includes at least one of: pass, downgrade, deny, or secure shutdown. The downgrade includes at least one of: rate limiting, geofence tightening, security corridor constraints, control domain switching, task suspension, request confirmation, permission escalation, or switching to a local autonomous policy.
[0093] The anti-interference audit module is also used for connection loss prediction. When the link quality shows a continuous downward trend, latency or jitter exceeds a preset threshold, packet loss rate or bit error rate exceeds a preset threshold, or the positioning / perception confidence is lower than a preset threshold, the control system sends an autonomous script, fallback path, security boundary, task continuation status, or last effective strategy to the embodied execution terminal before the link is completely interrupted. When the communication link is damaged or the positioning confidence is insufficient, the embodied execution terminal enters the disconnection autonomous mode based on the autonomous script, fallback path, security boundary, task continuation status, or last effective strategy to maintain task continuity, obstacle avoidance, fallback, or safe waiting under the security boundary.
[0094] The evidence storage module is used to associate and record virtual content status, environmental semantic targets, task description data, control instructions, audit receipts, execution feedback, link / confidence summaries, anchor version numbers, and time anchors, and generate an immutable evidence package when a triggering event occurs. The triggering events include one or more of the following: collision, boundary crossing, emergency stop, refusal to execute, high-risk action, disconnection prediction, disconnection autonomy, link recovery, anchor realignment, multi-entity conflict arbitration, permission escalation, robotic arm gripping, proximity of a hazard source, or manual takeover.
[0095] like Figure 2 As shown, in one embodiment, the present invention can be divided into a four-layer logical architecture:
[0096] The first layer is a unified spatiotemporal base, used to establish a multi-source physical truth coordinate system, a set of spatial anchor points, and a unified temporal anchor point;
[0097] The second layer is the narrative / intent / semantic driven control layer, which is used to map virtual content event streams, human intentions and environmental semantic goals into task-level instructions, dynamic instructions or local planning constraints.
[0098] The third layer is the extreme environment security layer, which is used to perform anti-interference auditing, gating, degradation, disconnection prediction, disconnection autonomy and safe shutdown.
[0099] The fourth layer is the full-scenario application mapping layer, which is used to achieve cross-execution end, cross-scenario and cross-industry reuse through standardized task interfaces and execution end adaptation layers.
[0100] In this implementation, the wearable sensing end, cloud / edge computing end, and embodied execution end exchange spatial anchor synchronization data, task description data, control command streams, command receipt streams, video / sensor feedback streams, status telemetry streams, semantic summary streams, audit receipts, and evidence summary via a cross-end communication link. The cross-end communication link can be a cellular network, private network, Wi-Fi, Bluetooth, UWB, 6G sensor network, or a combination thereof, and can employ multi-path redundancy to enhance reliability and low-latency determinism.
[0101] Through the aforementioned overall system architecture, this invention integrates operator-side spatial perception, virtual events, human intentions, environmental semantic targets, and the actual pose, motion state, and execution feedback of the execution end into a unified spatiotemporal reference. A complete control loop is formed through task interfaces, audit gating, disconnection autonomy, and black-box evidence storage. This overall architecture provides the foundation for subsequent multi-source physical truth fusion, spatial anchor sharing, semantic task mapping, low-latency ROI backhaul, phase compensation, conflict arbitration, and tamper-proof evidence storage.
[0102] Implementation Method 2: Multi-Source Physical Truth Fusion and Spatial Anchor Point Sharing
[0103] like Figure 3 As shown, this embodiment illustrates how the unified spatiotemporal alignment engine establishes a multi-source physical truth coordinate system based on multi-source physical observations, and through a spatial anchor point sharing mechanism, enables the wearable sensing end and the embodied execution end to maintain a consistent coordinate mapping relationship under the same physical space reference.
[0104] The multi-source physical true coordinate system is not a coordinate system directly output by a single positioning source, but a shared spatial reference formed by the fusion of at least two types of heterogeneous positioning / attitude observations. These heterogeneous positioning / attitude observations include, but are not limited to, one or more of the following: RTK / GNSS, UWB, VPS, SLAM / VIO, IMU, odometer, visual sensor, LiDAR, 6G integrated sensing derived observations, Wi-Fi channel state information derived observations, and Bluetooth ranging or angle measurement derived observations.
[0105] In one implementation, RTK / GNSS is used to provide regional or global absolute coordinate references; UWB is used to provide local relative distance, ranging, or angular constraints; VPS and SLAM / VIO are used to provide local environmental pose, visual feature points, and continuous motion trajectories; IMU is used to provide attitude, acceleration, and angular velocity continuity constraints at short time scales; odometry is used to provide the distance traveled by the actuator or motion increment; and visual sensors or lidar are used to provide environmental geometry, obstacle boundaries, and local point cloud constraints.
[0106] In one alternative implementation, when the embody execution terminal is located in a GNSS obstruction area, a weakly textured environment, or a visual blind spot, the system can also utilize 6G integrated sensing network-side sensing measurements to form redundant spatial constraints. The 6G integrated sensing derived observations include, but are not limited to, arrival delay, angle of arrival, Doppler, path loss, multipath reflection characteristics, or base station-side sensing results. The system can generate sparse point clouds of the environment, equivalent environmental feature sets, or spatial constraint terms based on these derived observations, and use them as supplementary constraints for RTK / GNSS, UWB, VPS, SLAM / VIO, or IMU.
[0107] In another alternative implementation, the system can also utilize Wi-Fi link channel state information, Bluetooth ranging or angle measurement results to form local spatial constraints. The Wi-Fi channel state information can be used to extract local electromagnetic fingerprints, distance variation trends, or spatial stability characteristics; the Bluetooth ranging or angle measurement results can serve as low-power redundant observations to maintain basic spatial alignment capabilities under conditions of weak network, weak vision, or partial occlusion.
[0108] 2.1 Time Alignment and Coordinate Unification of Multi-Source Observations
[0109] After receiving localization / attitude observations from different sources, the unified spatiotemporal alignment engine first performs time alignment on all types of observations. Observation data generated by different sensors or links can be appended with observation timestamps and mapped to a unified time reference.
[0110] In one implementation, the unified time reference can be achieved through GNSS / RTK timing, PTP, NTP, UWB two-way ranging calibration, an execution-side local monotonic clock, a cloud / edge synchronization clock, or a combination thereof. When external timing is unavailable, the system can use the execution-side local monotonic clock as a backup time reference and record the offset between this backup time reference and other time sources.
[0111] The system describes the relationship between the local coordinate system of the wearable sensing end, the local coordinate system of the embodied execution end, the map coordinate system of the cloud / edge end, and the unified truth coordinate system through spatial anchor points and extrinsic parameter transformations. For each type of observation, the system converts it into pose, distance, orientation, constraint terms, or confidence inputs in the unified truth coordinate system.
[0112] 2.2 Optional State Vectors and Fusion Computation Model
[0113] In one alternative implementation, the system can maintain a unified state vector to represent the state of the wearable sensing end and / or the embodied execution end in a unified truth coordinate system. The state vector can be expressed as:
[0114]
[0115] in, Indicates location, Indicates speed, Indicates a gesture, This indicates the bias of the inertial measurement unit or other sensors. This represents the scale factor, clock offset, sensor extrinsic parameters, or other optional state parameters.
[0116] In practical implementation, the aforementioned state vector is not limited to all the variables mentioned above; some variables can be selected based on the type of execution device, sensor configuration, and application scenario. For example, for unmanned ground vehicles, position, heading angle, speed, and odometer deviation can be used as states; for drones, three-dimensional attitude, acceleration, and altitude can be further added; for robotic arms, end effector pose, joint angle, torque state, or gripper state can be used as part of the execution state.
[0117] Multi-source observation can be abstracted as:
[0118]
[0119] in, Indicates RTK / GNSS observations. Indicates UWB distance or angle measurement observation. Indicates VPS, SLAM, or VIO observation. Indicates IMU observations, This indicates 6G integrated sensing derived observation. This indicates a derived observation of Wi-Fi channel state information. This indicates Bluetooth ranging or angle measurement derived observation.
[0120] The system can fuse the above observations using extended Kalman filtering, factor graph optimization, sliding window optimization, particle filtering, learned fusion, or equivalent methods. When expressed as weighted residuals, the system can perform state estimation using the following objective function:
[0121]
[0122] in, Indicates the first The residual of the observation relative to the current state estimate, Indicates the first The dynamic weights of observations are adjusted based on observation confidence, sensor status, link quality, historical stability, and scene conditions.
[0123] The above formula is only one optional fusion expression. In other embodiments, the system can also use a filter prediction-update structure, a graph optimization structure, a deep learning fusion model, or a rule fusion model to achieve multi-source physical truth estimation. As long as a shared spatial reference can be formed based on at least two types of heterogeneous observations, it is an equivalent implementation of the multi-source physical truth fusion described in this invention.
[0124] 2.3 Confidence assessment, dynamic weighting, and outlier removal
[0125] To avoid the failure of overall spatial alignment due to a single sensor malfunction, the unified spatiotemporal alignment engine performs confidence assessments on various observations. These confidence assessments may be based on at least one of the following metrics:
[0126] • RTK / GNSS fixed solution status, number of satellites, or positioning accuracy factor;
[0127] • UWB ranging residuals, non-line-of-sight indicators, or base station geometric distribution;
[0128] • SLAM / VIO reprojection error, number of tracked feature points, or relocation status;
[0129] • IMU integral drift, zero-bias stability, or short-term motion continuity;
[0130] • Stability of angle of arrival, stability of arrival delay, stability of multipath characteristics, or Doppler consistency of 6G integrated sensing and observation;
[0131] • Spatial fingerprint stability of Wi-Fi channel state information, ranging confidence, or link jitter;
[0132] • Signal strength stability and angle estimation confidence level for Bluetooth ranging or angle measurement;
[0133] • Spatial anchor point matching residuals, anchor point confidence, or anchor point version consistency;
[0134] • Continuity of historical trajectory, consistency of kinematic constraints, or consistency of geofence constraints.
[0135] In one optional decision logic, the system can calculate the observation confidence level for each type of observation:
[0136]
[0137] in, This represents the residual or error of the observation. This indicates the sensor's own quality indicators. Indicates link quality metrics, Indicators representing historical stability This represents the confidence calculation function. The confidence calculation function can be a rule function, a weighting function, a normalization function, a machine learning model, or a combination thereof.
[0138] When the confidence level of an observation is lower than a preset threshold, the system reduces the weight of that observation; when the residual of an observation continues to exceed a preset threshold, the system can mark the observation as abnormal and remove it; when an inexplicable conflict occurs between multiple source observations, the system triggers relocation, anchor point reconstruction, alignment degradation, or security audit.
[0139] In one alternative decision expression, it can be represented as:
[0140]
[0141]
[0142] in, The confidence threshold. This refers to the residual threshold. The above threshold can be preset based on the scenario risk level, sensor type, execution end type, or task status, or it can be dynamically adjusted by the system based on historical stability.
[0143] 2.4 Spatial Anchor Point Generation
[0144] Spatial anchors are used to establish a shared spatial reference between wearable sensing devices, cloud / edge computing devices, and embodied execution devices. Each spatial anchor includes at least an anchor identifier, anchor pose, anchor confidence level, anchor version number, and validity period.
[0145] In one implementation, spatial anchor points can be generated in one or more of the following ways:
[0146] (1) Generated based on environmental features, such as stable corners, door frames, road signs, trees, building boundaries, equipment edges, ground markings or other repeatable physical features;
[0147] (2) Generated based on artificial calibration points, such as preset marker points, QR codes, visual markers, reflective markers, positioning reference points, or artificially deployed calibration devices;
[0148] (3) Based on the geometry of UWB base stations, for example, local spatial anchor points are generated from the relative positional relationship of multiple UWB base stations;
[0149] (4) Generate based on RTK reference points, such as generating absolute spatial anchor points from regional reference stations or geographic control points;
[0150] (5) Based on 6G integrated sensing derived observation generation, such as generating equivalent spatial anchor points from environmental reflection points, angle of arrival stability features or multipath stability structures obtained by base station sensing;
[0151] (6) Generate local anchor points based on Wi-Fi channel state information, such as by generating stable electromagnetic fingerprint regions or ranging / angle measurement derived constraints;
[0152] (7) Generate auxiliary anchor points based on Bluetooth ranging or angle measurement, such as those generated from BLE beacons, AoA / AoD angle measurement results, or short-range positioning nodes.
[0153] 2.5 Anchor point synchronization, version control, and realignment triggering
[0154] The anchor set is synchronized across wearable sensing devices, cloud / edge computing devices, and embodied execution devices. To avoid spatial inaccuracies caused by different versions of the anchor set used by different terminals, each spatial anchor is configured with an anchor version number and an expiration date. The system determines whether the anchor sets across different terminals are consistent by using the anchor version number, timestamp, confidence level, and expiration date.
[0155] The system triggers an anchor update, realignment request, or control degradation when any of the following conditions are detected at either end:
[0156] • Anchor point matching residuals exceed a preset threshold;
[0157] • Anchor confidence level is lower than preset threshold;
[0158] • Anchor version numbers are inconsistent;
[0159] • Anchor point validity period expired;
[0160] • Environmental changes have rendered the original anchor points unusable;
[0161] • The relative pose deviation between the wearable sensing device and the embodied execution device exceeds a preset threshold;
[0162] Unexplained conflicts have emerged between multi-source observations;
[0163] • The feedback trajectory from the execution end is inconsistent with the shared coordinate mapping.
[0164] In one alternative decision expression, it can be represented as:
[0165]
[0166] in, This indicates the anchor point matching error. Indicates the anchor point error threshold. Indicates the anchor confidence level. This represents the anchor confidence threshold. and These represent the anchor version numbers used by different terminals.
[0167] 2.6 Cross-end coordinate mapping
[0168] The unified spatiotemporal alignment engine uses spatial anchor points as a common reference to calculate the mapping relationship between the local coordinate system of the wearable sensing end, the local coordinate system of the embodied execution end, and the unified truth coordinate system.
[0169] In one implementation, the local coordinate system of the wearable sensing end is denoted as FAF_AFA, the local coordinate system of the embodied execution end is denoted as FCF_CFC, and the unified truth coordinate system is denoted as FGF_GFG. The system can estimate the transformation relationship from FAF_AFA to FGF_GFG, and the transformation relationship from FCF_CFC to FGF_GFG, respectively, and thereby establish a shared coordinate mapping between the wearable sensing end and the embodied execution end.
[0170] When the local coordinates of the wearable sensing device drift, the system realigns it to a unified true coordinate system through spatial anchor constraints. When the local positioning of the embodied execution device drifts, the system pulls the execution device's state back to a consistent spatial reference through multi-source observation and anchor version verification.
[0171] In one alternative representation, the cross-end coordinate mapping can be expressed as:
[0172]
[0173]
[0174] in, This represents a point or pose in the local coordinate system of the wearable sensing device. This represents a point or pose in the local coordinate system of the embodied execution end. Represents the position or pose of a point in a unified truth coordinate system. and These represent the transformation relationships from the local coordinate system of the wearable sensing end, the local coordinate system of the embodied execution end, to the unified truth coordinate system, respectively.
[0175] The above transformation relationship can be determined by spatial anchor point matching, multi-source positioning fusion, manual calibration, external parameter calibration or a combination thereof, and can be dynamically corrected as anchor points are updated, the environment changes or confidence levels change.
[0176] 2.7 Alignment Degradation and Security Audit Output
[0177] When the multi-source physical truth fusion result, spatial anchor point status, or cross-end coordinate mapping status does not meet the current task's safety requirements, the unified spatiotemporal alignment engine can output an alignment degradation flag to the anti-interference audit module. The alignment degradation flag may include: normal spatial alignment, slightly abnormal spatial alignment, unstable spatial alignment, inconsistent anchor point versions, anchor point relocation in progress, insufficient reliability of the execution end position, insufficient reliability of the wearable sensing end position, task degradation required, rejection of high-risk actions required, and safe shutdown required.
[0178] The anti-interference audit module can process subsequent control commands by granting, downgrading, rejecting, or safely halting them based on the alignment degradation flag, task risk level, and execution end status. For example, when the system detects only a slight alignment error, it can limit the speed or tighten the safety corridor; when the anchor point version is inconsistent or repositioning is incomplete, it can refuse robotic arm grasping, close approach, or high-risk operations; when spatial alignment completely fails, it can trigger a safe halt or disconnection autonomous strategy.
[0179] Through the above-mentioned multi-source physical truth fusion and spatial anchor point sharing mechanism, this implementation method can continuously maintain the spatial consistency between the wearable sensing end and the embodied execution end in GNSS occlusion, visual blind spots, weak texture environment, strong interference environment or multi-execution end collaborative scenarios, providing a verifiable physical spatial foundation for subsequent multimodal command mapping, standardized task interface, anti-interference audit, ROI low-latency backhaul, disconnection autonomy and black-box evidence storage.
[0180] Implementation Method 3: Multimodal Instruction Mapping Engine, Semantic Control Layer, and Standardized Task Interface
[0181] like Figure 4 and Figure 9 As shown, this embodiment illustrates how a multimodal instruction mapping engine converts virtual content event streams, operator human characteristics, environmental semantic targets, or task inputs into task-level instructions, dynamic instructions, or local planning constraints that can be executed by the embodied execution end, and achieves cross-execution end, cross-scenario, and cross-industry reuse through standardized task interfaces and execution end adaptation layers.
[0182] The multimodal instruction mapping engine includes at least one or more of the following: a narrative / rhythm-driven mapping unit, a human body capture / gaze-driven mapping unit, a semantic control layer, a task description generation unit, and an execution-end adaptation layer. These units can be deployed in wearable sensing devices, cloud / edge computing devices, embodied execution devices, or combinations thereof.
[0183] The inputs to the multimodal instruction mapping engine include, but are not limited to:
[0184] Virtual content event stream, operator human characteristics, environmental semantic targets, task templates, spatial anchors, temporal anchors, execution end capability profiles, link quality status, anti-interference audit strategies, and scene security boundaries.
[0185] The output of the multimodal instruction mapping engine includes, but is not limited to: task-level instructions, dynamic instructions, local planning constraints, task description data, permission levels, auditing strategies, black-box indexes, and execution-end adaptation instructions.
[0186] 3.1 Parsing of Virtual Content Event Flow and Narrative / Rhythm-Driven Mapping
[0187] In one implementation, the virtual content event stream is a set of time-series events generated by a wearable sensing device, a scene engine, an interactive script, streaming media content, a performance control system, or a cloud / edge computing device. The virtual content event stream may include one or more of the following: narrative beats, plot stages, audio envelopes, music rhythms, BPM, keyframe events, virtual object appearance or disappearance events, collision events, climax events, virtual trigger points, script nodes, event priorities, or interactive feedback states.
[0188] The multimodal instruction mapping engine parses the virtual content event stream and extracts temporal features that can drive the motion or action changes of the embodied execution end. These temporal features may include:
[0189] BeatID (beat identifier); BPM (tempo); EventID (event identifier); EventType (event type); EventPhase (event phase); AudioEnvelope (audio energy envelope); TriggerPoint (virtual trigger point); Priority (event priority); Duration (duration); TemporalAnchor (temporal anchor point).
[0190] In one optional mapping logic, the system can generate dynamic parameters or path constraints for the execution end based on the virtual content event stream. For example, when the plot enters an acceleration phase, the system generates acceleration or speed targets; when the audio envelope reaches its peak, the system triggers formation changes or lighting / panel actions; when a virtual object appears near a specific spatial anchor point, the system generates actions such as approaching, surrounding, avoiding, or stopping.
[0191] The narrative / rhythm-driven mapping can be represented as:
[0192]
[0193] in, This indicates the output dynamic commands or path segment constraints. Represents the virtual content event stream, Indicates spatial anchor points or scene constraints. Represents the set of security constraints. This represents the narrative / rhythm-driven mapping function.
[0194] The aforementioned mapping function can be implemented using rule tables, state machines, parameterized curves, event-triggered logic, machine learning models, or combinations thereof. The output of the mapping function does not directly bypass the security protection layer, but rather serves as a candidate instruction in subsequent task description generation, execution-end adaptation, and anti-interference auditing processes.
[0195] In one implementation, the dynamic parameters output by the narrative / rhythm-driven map include one or more of the following: velocity, acceleration, angular velocity, braking, steering, formation spacing, following distance, action start and end time, or action intensity. The path segment constraints include one or more of the following: start point, end point, intermediate waypoint, safety corridor, geofence, maximum speed, maximum acceleration, maximum turning angle, minimum safe distance, or action duration.
[0196] 3.2 Human body capture / gaze-driven mapping
[0197] In another implementation, the wearable sensing device acquires the operator's human body features and inputs these features into a multimodal instruction mapping engine. These human body features include, but are not limited to, gaze direction, gaze target, gestures, posture, gait, head movements, hand trajectories, voice input, bioelectrical signals, electromyographic signals, or combinations thereof.
[0198] Human body capture / gaze-driven mapping is used to translate operator intentions into task-level actions, fine-grained commands, or motion trajectories that can be executed by the embodied execution unit. For example:
[0199] When the operator looks at the target and performs a confirmation gesture, the system generates task-level instructions such as "approach the target", "mark the target", "follow the target" or "maintain a safe distance".
[0200] When the operator drags the path with gestures, the system generates path segment constraints for the execution end;
[0201] When the operator makes a rotation gesture, the system generates a twisting motion of the robotic arm, a rotation motion of the gimbal, or a turning motion of the actuator.
[0202] When the operator opens and closes their hand, the system generates a gripper opening, gripper closing, or gripping confirmation command.
[0203] When the operator inputs gestures or gait, the system generates actions such as following, stopping, circling, approaching, or retreating.
[0204] In one optional mapping logic, cross-scale mapping can be used between human input and execution-end actions. This cross-scale mapping is used to convert small movements or intentional inputs from the human end into action constraints at different scales, degrees of freedom, or kinematic structures at the execution end. For example, minute displacements of the operator's hand can be mapped to fine displacements at the end of a robotic arm; the operator's gaze direction can be mapped to the path target of an autonomous vehicle or the heading of a drone; and operator gesture confirmation can be mapped as a second trigger condition for high-risk actions.
[0205] Cross-scale mapping can be represented as:
[0206]
[0207] in, This indicates task-level instructions, fine-grained operation instructions, or motion trajectories generated from human input. Represents a set of human body features. Indicates the object being gazed upon or the semantics of the object. Indicates spatial anchor points or scene constraints. Indicates the risk level or access level. This represents the human body capture / gaze-driven mapping function.
[0208] The human body capture / gaze-driven mapping function can be implemented using rule mapping, gesture recognition models, gaze target selection models, action templates, dynamic scaling models, or combinations thereof. The mapping results are also used as candidate instructions in the task description generation and anti-interference auditing processes.
[0209] 3.3 Semantic Control Layer
[0210] In another implementation, the multimodal instruction mapping engine includes a semantic control layer. This semantic control layer is used to convert environmental semantic targets, task roles, risk levels, or interaction object identification results into task-level commands, and further map them into motion constraints, operational constraints, or path constraints at the execution end.
[0211] The environmental semantic targets can be generated by wearable sensing devices, embodied execution devices, or cloud / edge computing devices through visual recognition, semantic segmentation, target detection, sensor fusion, manual annotation, or task templates. The environmental semantic targets include, but are not limited to, casualties, hazards, obstacles, passages, valves, meters, tools, work areas, plot boundaries, waypoints, hazardous chemical containers, leak points, vehicles, pedestrians, animals, mechanical equipment, or virtual interactive objects.
[0212] The semantic control layer performs at least the following steps:
[0213] (1) Identify or receive environmental semantic targets;
[0214] (2) Associate environmental semantic targets with spatial anchor points or a unified truth coordinate system;
[0215] (3) Determine the task roles, risk levels, and interaction types of the semantic objectives;
[0216] (4) Generate task-level commands based on the task template;
[0217] (5) Convert task-level commands into motion constraints, operation constraints, or path constraints;
[0218] (6) The conversion results are handed over to the standardized task interface to generate task description data.
[0219] For example, in a remote rescue scenario, after the system identifies a "wounded" target, it can generate task commands such as "search and approach", "maintain a safe distance", "mark the location", "deliver supplies" or "call for manual confirmation"; after identifying a "hazard", it can generate task commands such as "bypass the hazard", "establish a safe boundary", "limit the approach distance" or "request permission confirmation".
[0220] In industrial inspection scenarios, after the system identifies a "meter", it can generate a task command to "approach and read the meter"; after identifying a "valve", it can generate a task command to "approach the valve and wait for confirmation by turning"; after identifying a "hazardous chemical drum" or "leak point", it can generate a task command to "maintain a safe distance, mark the location and send back a semantic summary".
[0221] In smart agriculture scenarios, after the system identifies "plot boundaries", "crop rows", "obstacles" or "spraying areas", it can generate task commands such as "cruise along the boundary", "generate operation route", "avoid obstacles" and "spray in zones".
[0222] In one alternative representation, the mapping from semantic target to task command can be expressed as:
[0223]
[0224] in, Indicates a task-level command. Representing the semantic target of the environment, Indicates the spatial anchor point or target spatial location. Indicates the risk level. This indicates a task template or strategy template. This represents a semantic control mapping function.
[0225] The semantic control mapping function can be implemented by task rules, semantic templates, knowledge bases, state machines, policy models, or combinations thereof.
[0226] 3.4 Standardized Task Interface
[0227] like Figure 9 As shown, the standardized task interface is used to generate, receive, or parse task description data. Through the standardized task interface, the system can uniformly encapsulate control intentions from virtual content event streams, human input, and environmental semantic targets into task description data that is understandable, auditable, and adaptable to the executor.
[0228] In one implementation, the task description data includes, but is not limited to, the following fields:
[0229] TaskID, Task Identifier; TargetID, Target Identifier; TaskType, Task Type; RiskLevel, Risk Level; Priority, Task Priority; AnchorID, Spatial Anchor Identifier; AnchorVersion, Anchor Version Number; TemporalAnchor, Temporal Anchor; GeoFence, Geofence; SafeCorridor, Safe Corridor; PermissionLevel, Permission Level; MotionConstraint, Motion Constraint; OperationConstraint, Operation Constraint; PathConstraint, Path Constraint; ROIRegion, Region of Interest or Semantic Target Region; AuditPolicy, Audit Policy; FailOperationalPolicy, Failure Policy; BlackBoxIndex, Black-box Index; CallbackPolicy, Callback Policy.
[0230] The above fields can be expanded, trimmed, or replaced with equivalents depending on the application scenario. Not every task description data must contain all fields, but it should at least include one or more of the following: task objective, time reference, spatial reference, and execution constraints, to ensure that the task description can be parsed by the execution-side adaptation layer and enter the security audit process.
[0231] In one implementation, the task description data can be encapsulated in a structured format, such as key-value pairs, JSON, binary protocol fields, task template objects, control message bodies, or other machine-readable formats. The structured format does not limit the scope of this invention; any format capable of expressing objectives, tasks, spatial references, temporal references, risk levels, and execution constraints can serve as an equivalent implementation of the standardized task interface.
[0232] 3.5 Execution-side capability profile and adaptation layer
[0233] The execution-end adaptation layer is used to convert the task description data generated by the standardized task interface into task-level instructions, dynamic instructions or local planning constraints that can be executed by the corresponding execution end, based on the capability profile of different execution ends.
[0234] The capability profile of the actuator includes, but is not limited to: actuator type; motion dimension; maximum speed; maximum acceleration; minimum turning radius; maximum climbing ability; hovering ability; load capacity; robotic arm degrees of freedom; end effector type; sensor configuration; obstacle avoidance ability; local planning ability; positioning ability; communication ability; disconnection autonomy ability; safe shutdown ability; permitted action types; prohibited action types.
[0235] In one implementation, the execution-side adaptation layer generates corresponding execution instructions based on fields such as TaskType, TargetID, RiskLevel, GeoFence, SafeCorridor, and MotionConstraint in the task description data, combined with the execution-side capability profile. For example:
[0236] For autonomous vehicles, the adaptation layer can transform the task description of "approaching the target and maintaining distance" into path following, speed limits, steering control, braking strategies, and obstacle avoidance constraints;
[0237] For drones, the adaptation layer can convert the same task description into waypoints, altitude, heading, obstacle avoidance radius, hovering point, and return-to-home strategy;
[0238] For quadruped robots, the adaptation layer can convert the same task description into gait planning, foothold constraints, obstacle crossing strategies, and low-speed approach strategies.
[0239] For robotic arms, the adaptation layer can convert the same task description into end-effector pose, joint trajectory, force control threshold, gripper action, and collision detection constraints;
[0240] For special operation equipment, the adaptation layer can convert the same task description into a dedicated action sequence, operating radius, safety boundary, and access confirmation process.
[0241] In one alternative expression, execution-side adaptation can be represented as:
[0242]
[0243] in, This represents the adapted execution-end instructions or local planning constraints. This represents task description data. This represents a profile of the execution capabilities. Represents the set of security constraints. This indicates the execution-side adaptation function.
[0244] The execution-side adaptation function can be implemented using a rule table, device driver, kinematic model, local planner, controller interface, task template, or a combination thereof.
[0245] 3.6 Access Control Levels and Confirmation of High-Risk Actions
[0246] In one implementation, the multimodal instruction mapping engine and standardized task interface also support hierarchical permissions. For high-risk actions, such as robotic arm grasping and closing, valve turning, approaching a hazard source, crossing a safety boundary, releasing a load, performing an emergency stop release, or entering a densely populated area, the system may require multiple triggering conditions to be met before generating an executable task description or allowing it to pass auditing.
[0247] The multiple triggering conditions include, but are not limited to: the gaze target lasts for a preset time; the operator performs a confirmation gesture; the permission level meets the task requirements; the spatial anchor point version is consistent; the positioning / perception confidence meets the threshold; the link quality meets the threshold; the execution end status meets the security conditions; and the security protection layer audit is passed.
[0248] In one alternative decision formula, the conditions for generating high-risk actions can be expressed as:
[0249]
[0250] in, This indicates that the fixation condition is met. This indicates that the gesture confirmation condition has been met. This indicates that the permission level is met. This indicates that the positioning / perception confidence level is met. This indicates that high-risk action instructions are allowed to be generated or enter the audit process.
[0251] If any of the above conditions are not met, the system may refuse to generate high-risk action instructions, or only generate low-risk alternative tasks, such as approaching but not touching, marking but not operating, waiting for manual confirmation, or entering a safe observation mode.
[0252] 3.7 Command Generation and Auditing Pre-processing
[0253] The task description data, task-level instructions, dynamic instructions, or local planning constraints generated by the multimodal instruction mapping engine do not directly drive the execution of the embodied execution end. Instead, they first enter the security protection layer for anti-interference auditing.
[0254] Before entering the audit process, the system can attach the following metadata to each candidate instruction:
[0255] CommandID, command identifier; TemporalAnchor, temporal anchor; AnchorID and AnchorVersion, spatial anchor identifier and version number; SourceType, command source type; RiskLevel, risk level; PermissionLevel, permission level; InputEvidence, input evidence summary; TargetID, target identifier; AuditPolicy, audit policy; BlackBoxIndex, black-box index.
[0256] The InputEvidence can include the operator's gaze target, gesture recognition results, virtual event ID, semantic target recognition results, task template version, or corresponding video frame time anchor point. In this way, the system can retrospectively trace the input that generated the instruction, what the operator saw at the time, the spatial location of the target, the risk level corresponding to the instruction, and why execution was approved or rejected during subsequent auditing and evidence preservation.
[0257] 3.8 Relationship with spatial anchors, temporal anchors, and black-box evidence storage
[0258] The task description data generated by the standardized task interface is interconnected with spatial anchors, temporal anchors, and black-box evidence storage.
[0259] Spatial anchors are used to determine the location of task objectives, geofences, safety corridors, path constraints, and execution areas in a unified truth coordinate system. Temporal anchors are used to determine the temporal relationship between task generation, operator input, environmental semantic recognition, video frame feedback, and execution actions. Black-box indexes are used to link task description data, candidate instructions, audit receipts, and execution feedback to the same chain of evidence.
[0260] For example, in a rescue scenario, the task description data for "approaching the wounded and maintaining distance" might include the wounded target's TargetID, the target's AnchorID, the TemporalAnchor of the video frame seen by the operator upon confirmation, the task risk level, the permissible approach distance, the safe corridor, and the black-box index. Before execution, the anti-interference audit module determines whether the link quality, positioning reliability, and anchor version meet the requirements. After execution, the execution feedback and audit receipt are written into the evidence package under the same black-box index.
[0261] Through the aforementioned multimodal instruction mapping engine, semantic control layer, standardized task interface, and execution end adaptation layer, this implementation can unify virtual content events, human intentions, environmental semantic targets, and different execution end capabilities into an auditable, adaptable, and traceable task description system. This enables the same set of wearable sensing terminals and control bases to be reused among unmanned vehicles, drones, robots, robotic arms, agricultural equipment, industrial inspection equipment, and special operation equipment, thereby improving the portability and scalable deployment capabilities of the cross-space embodied collaborative control system.
[0262] Implementation Method 4: Anti-interference auditing, gating, degradation, disconnection prediction and autonomous disconnection
[0263] like Figure 5 and Figure 6 As shown, this embodiment illustrates how the security protection layer performs anti-interference auditing on candidate instructions based on link quality, sensor confidence, positioning residual, anchor point consistency, instruction risk level, execution end status, permission level, and task stage, and executes pass, downgrade, rejection, security shutdown, disconnection prediction, or disconnection autonomy under conditions of link damage, insufficient positioning confidence, or high risk.
[0264] The anti-interference audit module is located within the security protection layer and can be deployed in wearable sensing devices, cloud / edge computing devices, embodied execution devices, or combinations thereof. In scenarios with high security levels, unstable links, or where a local security loop needs to be maintained, the key decision logic of the anti-interference audit module is preferably deployed locally on the embodied execution device to ensure that basic security judgments can still be performed when the communication link is interrupted.
[0265] 4.1 Anti-interference audit input
[0266] The inputs to the anti-interference audit module include, but are not limited to: link quality indicators; sensor confidence; positioning residual; anchor point consistency; instruction risk level; execution end status; permission level; task stage; geofence status; security corridor status; environmental risk level; historical audit results; disconnection prediction status; and multi-entity conflict status.
[0267] The link quality indicators include, but are not limited to, link latency, link jitter, packet loss rate, bit error rate, number of acknowledgment timeouts, signal strength, bandwidth margin, path switching status, or multi-link redundancy status.
[0268] The sensor confidence levels include, but are not limited to, RTK / GNSS fixed solution status, UWB ranging confidence level, SLAM / VIO reprojection error, IMU drift status, visual target recognition confidence level, semantic target recognition confidence level, 6G integrated sensing derived observation confidence level, Wi-Fi channel status derived observation confidence level, and Bluetooth ranging or angle measurement confidence level.
[0269] The positioning residuals include, but are not limited to, multi-source positioning fusion residuals, anchor point matching residuals, execution end trajectory prediction residuals, wearable sensing end pose drift, and deviations between execution end local coordinates and unified true coordinates.
[0270] The anchor consistency includes whether the anchor version number is consistent, whether the anchor confidence level meets the threshold, whether the anchor validity period has expired, whether the anchor matching error exceeds the limit, and whether it is in a realignment state.
[0271] The risk level of the instruction can be determined based on factors such as task type, action type, target object, environmental semantic target, execution capability, distance from personnel or hazard source, whether it involves contact with the robotic arm, whether it involves crossing boundaries, whether it involves high-speed movement, or whether it involves load release.
[0272] 4.2 Gating Results
[0273] The anti-interference audit module gates candidate instructions, and the gate result includes at least one of passing, downgrading, rejecting, or safe shutdown.
[0274] When the link quality, sensor confidence, positioning residual, anchor point consistency, execution end status, and permission level all meet the requirements of the current instruction risk level, the anti-interference audit module outputs a pass result, allowing the embodied execution end to execute the corresponding task-level instruction, dynamic instruction, or local planning constraint.
[0275] When some indicators fail to meet the original instruction requirements but still meet the minimum conditions for safe operation, the anti-interference audit module outputs a degradation result. This degradation includes, but is not limited to: reducing speed; limiting acceleration; limiting angular velocity; limiting robotic arm torque; tightening geofencing; tightening safety corridors; reducing task priority; reducing video background bitrate; switching to semantic digest feedback; switching to a low-risk action template; pausing high-risk sub-actions; requesting secondary confirmation from the operator; increasing permission level requirements; switching to a local autonomy strategy; delaying execution; executing an alternative trajectory; and entering a safe waiting state.
[0276] When key indicators fail to meet execution conditions, the anti-interference audit module outputs a rejection result. The rejection result is used to prevent the embodied execution terminal from executing the corresponding candidate instruction and generates a rejection reason. These reasons include, but are not limited to, insufficient link quality, insufficient location confidence, inconsistent anchor version, insufficient target semantic confidence, insufficient permission level, abnormal execution terminal status, geofence conflict, or excessively high instruction risk level.
[0277] When the system detects severe spatial inaccuracy, link unavailability, abnormal execution end status, excessive collision risk, loss of control of high-risk actions, breach of safety boundaries, or inability of audit logic to continue reliable judgment, the anti-interference audit module outputs a safe shutdown result. The safe shutdown may include one or more of the following: braking, hovering, emergency stop, power failure protection, retraction to a safe point, maintaining the current position of the robotic arm, releasing the safety lock, or entering a safe waiting state.
[0278] 4.3 Comprehensive safety score and threshold determination
[0279] In one optional implementation, the anti-interference audit module can construct a comprehensive security score for quantitatively judging candidate instructions. The comprehensive security score can be determined jointly by indicators such as link quality, sensor confidence, positioning residual, anchor point consistency, instruction risk level, execution end status, and permission level.
[0280] In one alternative expression, the comprehensive safety score can be represented as:
[0281]
[0282] in, This indicates the overall safety score. Indicates the link quality score. This indicates the sensor confidence score. Indicates the anchor point consistency score. Indicates the execution end status score. Indicates the risk level of the instruction. , , , , These are the weighting coefficients.
[0283] In an optional decision logic:
[0284]
[0285]
[0286]
[0287]
[0288] in, , , These are the pass threshold, downgrade threshold, and safety shutdown threshold. These thresholds can be preset based on the task scenario, execution terminal type, task risk level, environmental risk level, or regulatory requirements, or they can be dynamically adjusted by the system based on historical operating status.
[0289] The scoring formula described above is only one possible judgment method. In other embodiments, the anti-interference audit module can also use rule trees, state machines, expert systems, machine learning models, risk matrices, or combinations thereof to implement gating judgment. As long as it can judge the execution of instructions based on factors such as link quality, location / perception confidence, anchor point consistency, instruction risk level, and execution end status, it is an equivalent implementation of the present invention.
[0290] 4.4 Audit Receipt
[0291] The anti-interference audit module generates an audit receipt when outputting gating results. The audit receipt is used to record why the system passed, downgraded, rejected, or shut down safely, and is used for subsequent black-box evidence preservation, accountability tracing, and post-mortem analysis.
[0292] In one implementation, the audit receipt includes at least one or more of the following fields:
[0293] AuditID, audit receipt identifier; CommandID, associated instruction identifier; DecisionType, decision type; TriggerReason, trigger reason; KeyMetrics, key input metrics; DegradeAction, degrade action; RiskLevel, instruction risk level; PermissionLevel, permission level; TemporalAnchor, temporal anchor; AnchorID, spatial anchor identifier; AnchorVersion, anchor version number; ExecutorState, execution end state summary; LinkState, link state summary; ConfidenceSummary, confidence summary; BlackBoxIndex, black box index; Signature, signature or summary field.
[0294] The DecisionType includes Pass, Degrade, Deny, or Safe Shutdown. The TriggerReason can record reasons such as excessive link jitter, excessive positioning residual, inconsistent anchor version, insufficient target confidence, insufficient permissions, execution end abnormality, geofence conflict, multi-entity conflict, or disconnection prediction trigger.
[0295] Audit receipts can be transmitted back to the wearable sensing device in real time, allowing operators to see why the system rejects or downgrades a certain operation; audit receipts can also be written into the evidence storage module, forming a chain of evidence together with control instructions, task description data, execution feedback, and time anchors.
[0296] 4.5 Prediction of Disconnection
[0297] like Figure 6As shown, the anti-interference audit module is also used to perform disconnection prediction. Disconnection prediction is not triggered only after the communication link is completely interrupted, but rather it is used to predict the risk of disconnection in advance when the link quality shows a continuous deterioration trend, and to issue autonomous scripts, fallback paths, security boundaries, task continuation status or last effective strategies to the embodied execution end before the link is completely interrupted.
[0298] Disconnection prediction can be based on one or more of the following indicators: continuously increasing link latency; continuously increasing link jitter; packet loss rate exceeding a preset threshold; bit error rate exceeding a preset threshold; number of receipt timeouts exceeding a preset threshold; bandwidth margin below a preset threshold; instability in both the main link and backup link; decreased confidence in positioning / perception; increased residual from multi-source observations; decreased confidence in anchor point matching; the execution end entering a weak signal area; and increased environmental risk level.
[0299] In one possible decision formula, the risk of disconnection can be expressed as:
[0300]
[0301] in, Indicates the risk score of disconnection. Indicates link latency. Indicates link jitter. This indicates the packet loss rate or bit error rate. This indicates that the receipt has timed out. Indicates the confidence level of location / perception. This represents the function for assessing the risk of disconnection.
[0302] When the disconnection risk score exceeds a preset threshold, it can be represented as:
[0303]
[0304] in, This indicates the threshold for predicting disconnection. The autonomous strategy includes one or more of the following: autonomous script, fallback path, security boundary, task continuation status, or last effective strategy.
[0305] The above formula is only one possible expression. In other implementations, the system may use rule conditions, trend judgment, state machine, link prediction model, or a combination thereof to predict connection failures.
[0306] 4.6 Issue the autonomous script in advance
[0307] When a disconnection prediction is triggered, the system sends an autonomous script to the embodied execution terminal before the link is completely interrupted. The autonomous script is used to guide the embodied execution terminal on how to maintain safe operation when the communication link is damaged or completely interrupted.
[0308] In one implementation, the autonomous script includes one or more of the following: current task stage; completed task status; target location; safe point location; backtracking path; geofence; safe corridor; maximum speed; maximum acceleration; maximum turning angle; minimum safe distance; obstacle avoidance rules; waiting strategy; low-risk task continuation strategy; continuation strategy after link recovery; safe shutdown strategy under abnormal conditions; black-box recording strategy.
[0309] Autonomous scripts can be bound to task description data, spatial anchor version numbers, and temporal anchors to ensure that the executor continues to run based on the most recent trusted spatiotemporal reference during the disconnection period.
[0310] In one implementation, the system can also save the last valid policy as LastValidPolicy. The last valid policy includes the most recently audited task status, path constraints, security boundaries, speed limits, and an execution end status summary. When the link is suddenly interrupted and the complete autonomous script cannot be delivered in time, the execution end can enter a conservative autonomous mode based on the last valid policy.
[0311] 4.7 Disconnected Autonomous Mode
[0312] When the communication link is damaged, interrupted, the positioning signal is insufficient, or remote real-time control is no longer safe, the embodied executor enters the disconnected autonomous mode. The disconnected autonomous mode includes, but is not limited to: low-speed cruising; local obstacle avoidance; continuing to execute low-risk sub-tasks along a safe corridor; pausing high-risk actions; returning to a safe point; hovering; waiting in place; maintaining the robot arm's safe posture; returning to home; stopping work; and safe shutdown.
[0313] In one implementation, the executor selects different autonomous strategies based on the task risk level. For low-risk tasks, the executor can continue to execute low-speed tasks within geofences and safe corridors; for medium-risk tasks, the executor can pause high-risk sub-actions and enter a safe waiting state; for high-risk tasks, the executor preferably stops the current action, retreats to a safe point, or triggers a safe shutdown.
[0314] During the disconnection process, the execution end continuously records the local sensor status, execution trajectory, obstacle avoidance decisions, speed, attitude, task stage, abnormal events, and security policy execution results. When the link is restored, the execution end sends back the status summary, execution feedback, and audit receipt during the disconnection period to the wearable sensing end or cloud / edge computing end, and writes them into the spatiotemporal black box evidence chain.
[0315] 4.8 The Interaction Between Gating, Disconnection Autonomy, and the Spatiotemporal Black Box
[0316] Anti-interference auditing, disconnection prediction, and disconnection autonomy are all linked to the evidence storage module. When the system experiences events such as approval, downgrade, rejection, security shutdown, disconnection prediction, early distribution of autonomous scripts, entry into disconnection autonomy, link recovery, or manual takeover, the evidence storage module can generate corresponding black-box indexes and record relevant data.
[0317] For example, in a remote rescue scenario, if the system detects a continuous increase in link jitter and a decrease in location reliability, the anti-interference audit module outputs a degradation result and triggers a disconnection prediction. The system pre-issues a fallback path and safety boundary, and records the audit receipt. When the link is completely interrupted, the execution end slowly falls back to a safe point according to the autonomous script. After the link is restored, the execution end sends back the path, obstacle avoidance events, and status summary during the disconnection period. All of the above processes are written under the same black-box index, forming a traceable chain of evidence.
[0318] In hazardous industrial operation scenarios, if the robotic arm is about to perform a high-risk grasping action, but the system detects that the anchor point version is inconsistent or the target confidence is insufficient, the anti-interference audit module outputs a rejection result, prevents the grasping action, and writes the rejection reason, instruction source, operator input, target semantic state, and audit receipt into the evidence package.
[0319] Through the aforementioned anti-interference auditing, gating, degradation, disconnection prediction and disconnection self-governance mechanisms, this implementation method can provide a provable, traceable and degradeable security closed loop for cross-space embodied collaborative control under conditions of link damage, strong interference, decreased positioning information, anchor mismatch or high-risk actions, and prevent the execution end from continuing to execute high-risk actions in an untrusted state.
[0320] Implementation Method 5: Low-latency feedback return, ROI encoding, time anchoring, ASW reprojection, and instruction-video phase alignment
[0321] like Figure 7 As shown, this embodiment illustrates how, in the process of cross-space embodied collaborative control, the system performs low-latency feedback processing on video frames, sensor observations, state telemetry, or semantic summaries transmitted back from the execution end, and reduces operation drift caused by link latency, video lag, and control execution asynchrony through ROI hierarchical coding, time anchoring, ASW reprojection, phase verification, state forward prediction, and instruction compensation.
[0322] In remote rescue, hazardous operations, industrial inspection, unmanned interactive experiences, smart agriculture, or other cross-space control scenarios, operators typically make control decisions based on video feeds transmitted from the actuator, environmental semantic targets, sensor status, or the actuator's location. If there is a significant time delay between the feedback feed and the actual status of the actuator, the operator may issue control commands based on the lagging feed, leading to failure to approach the target, excessive directional correction, misalignment of the robotic arm's grasp, untimely obstacle avoidance, or accidental triggering of high-risk actions.
[0323] To address the aforementioned issues, this implementation method unifies the association of feedback data, operator input, control commands, and execution actions to a time anchor point, and establishes an auditable phase alignment relationship between the feedback link and the control link.
[0324] 5.1 Execution Feedback Data Types
[0325] The execution feedback in this embodiment includes, but is not limited to, one or more of the following: video frames, sensor observations, status telemetry, semantic summaries, and audit status.
[0326] The video frames may include RGB images, depth images, infrared images, thermal imaging images, panoramic images, first-person view videos, images from the end-effector camera of a robotic arm, aerial images from a drone, or front-view images from an unmanned vehicle.
[0327] The sensor observations may include lidar point clouds, millimeter-wave radar data, IMU data, odometer data, UWB ranging data, RTK / GNSS positioning data, torque sensor data, gripper status, gas sensor data, temperature sensor data, pressure sensor data, or battery status data.
[0328] The status telemetry may include the position, speed, attitude, acceleration, angular velocity, path stage, task stage, control mode, disconnected autonomous state, obstacle avoidance state, robotic arm joint state, end effector state, remaining battery power, core temperature, and communication status.
[0329] The semantic summary may include the location of the injured person, the location of the hazard source, the boundary of the obstacle, the passable passage, the location of the valve, the meter reading, the location of the tool, the work area, the geofence, the risk area, the target recognition confidence or the semantic segmentation result.
[0330] 5.2 ROI Layered Encoding and Semantic-First Backhaul
[0331] In one implementation, the system performs ROI hierarchical encoding on video frames or semantic summaries transmitted from the execution end based on the operator's gaze region, environmental semantic targets, risk regions, or task objectives. The ROI can be determined by wearable sensing devices, embodied execution devices, or cloud / edge computing devices.
[0332] ROI areas include, but are not limited to: the area currently being viewed by the operator; the target object that the operator is about to operate on; semantically identified casualties, hazards, valves, meters, tools, obstacles, or passageways; the contact area of the robotic arm end effector; the passable area in the direction of the actuator's movement; the boundary of a geofence or safety corridor; areas determined by the system to be high-risk; and the area corresponding to the TargetID in the current task description data.
[0333] In low-bandwidth, link jitter, or high-risk scenarios, the system enables ROI regions to be transmitted back at a higher quality, bitrate, resolution, frame rate, or priority than background regions; background regions can have their resolution reduced, frame rate lowered, compression rate increased, or be converted to semantic summaries to reduce transmission pressure and maintain the visibility of critical regions.
[0334] In one alternative representation, video frames can be divided into ROI regions and non-ROI regions:
[0335]
[0336] in, Indicates the first Video frames at any given moment Indicates the ROI region. This represents the background region. The system can assign different encoding qualities to the ROI region and the background region respectively:
[0337]
[0338] in, Indicates the coding quality of the ROI region. This indicates the encoding quality of the background area. The encoding quality can be reflected in one or more of the following: resolution, bit rate, frame rate, compression quality, retransmission priority, or transmission priority.
[0339] In one implementation, when the link quality deteriorates, the system prioritizes the transmission of ROI regions, semantic summaries, task target regions, and audit receipts; when the link deteriorates further, the system can reduce the quality of background video transmission, retaining only key target regions, status telemetry, and semantic summaries; when the link is close to disconnection, the system triggers a disconnection prediction and enters the autonomous script advance distribution process described in Implementation 4.
[0340] 5.3 Generation and Binding of Time Anchors
[0341] To ensure a comparable temporal relationship between feedback data, operator input, and control commands, the system adds or associates time anchors to video frames, sensor observations, semantic summaries, operator inputs, task description data, control commands, and execution actions.
[0342] A time anchor may include one or more of the following information: acquisition time; encoding time; transmission time; reception time; rendering time; operator input time; control command generation time; control command transmission time; executor reception time; executor execution time; local monotonic clock value; timestamp under a unified time base; and offset from other time bases.
[0343] In one implementation, each video frame or semantic summary carries a corresponding TemporalAnchor. When the wearable sensing device displays the feedback content, it associates the TemporalAnchor with the screen currently seen by the operator. When the operator issues gaze confirmation, gesture confirmation, voice input, or task selection input based on the screen, the system binds the operator's input to the TemporalAnchor of the video frame they see.
[0344] Subsequently, when the multimodal instruction mapping engine generates task description data or control instructions based on the operator's input, it writes the TemporalAnchor into the task description data or control instructions, so that the execution end can identify which moment's feedback screen or semantic state the instruction was generated from.
[0345] In an optional field structure, control instructions may include: CommandID; TaskID; TargetID; TemporalAnchor_view; TemporalAnchor_cmd; AnchorID; AnchorVersion; RiskLevel; MotionConstraint; AuditPolicy; BlackBoxIndex.
[0346] Among them, TemporalAnchor_view represents the time anchor point of the feedback screen seen when the operator generates the command, and TemporalAnchor_cmd represents the time anchor point when the control command is generated. Through the above fields, the system can reconstruct the time sequence chain of "what the operator sees, when they see it, when the command is issued, and when the execution end executes it" during subsequent auditing and evidence preservation.
[0347] 5.4 ASW Reprojection and Display Viewpoint Compensation
[0348] In immersive control scenarios, there is typically a transmission and rendering delay between video frames captured by the execution end and displayed on the wearable sensing end. If the operator's head changes position during this delay, directly displaying the delayed video frame will cause image lag, spatial misalignment, or dizziness, thus affecting the operator's judgment.
[0349] Therefore, in one implementation, after receiving a delayed video frame carrying a Temporal Anchor, the wearable sensing device reads the 6-DoF pose of the sensing device at the current moment and calculates the difference between the 6-DoF pose and the pose of the sensing device at the corresponding time anchor point of the video frame. Based on this pose difference, the system performs asynchronous spatial warp (ASW) reprojection on the video frame to compensate the displayed image as closely as possible to the viewpoint corresponding to the current head pose.
[0350] In one alternative representation, the perceptual end pose at the corresponding temporal anchor point of the video frame is denoted as:
[0351]
[0352] The current sensor pose at the moment of display is recorded as follows:
[0353]
[0354] The pose difference between the two can be expressed as:
[0355]
[0356] The system is based on Perform reprojection on the original video frame or depth-assisted image to generate a compensated display frame:
[0357]
[0358] in, Represents the original delayed video frame. This represents the display frame after reprojection. This represents the ASW reprojection or equivalent viewpoint compensation function.
[0359] The aforementioned reprojection function can be implemented based on depth maps, sparse features, optical flow estimation, head pose difference, local geometric models, or combinations thereof. ASW reprojection is only used to reduce visual lag at the display end and does not replace security auditing at the execution end. For high-risk actions, gating is still required by the anti-interference audit module, taking into account link latency, positioning / perception confidence, and command risk level.
[0360] When the uncertainty of reprojection exceeds a preset threshold, the system may reduce interaction permissions, request operator confirmation, prompt screen lag, restrict high-risk actions, or trigger security audit downgrade.
[0361] In one of the optional decision expressions:
[0362]
[0363] in, This indicates the uncertainty of reprojection. This represents the threshold of reprojection uncertainty.
[0364] 5.5 Instructions—Video Phase Alignment
[0365] Command-video phase alignment is used to determine whether there is an excessive phase difference between the feedback state on which the operator generates control commands and the current state of the execution end.
[0366] In one implementation, after receiving a control command, the execution terminal parses the TemporalAnchor_view and TemporalAnchor_cmd carried in the control command and compares them with the execution terminal's current local time, state telemetry, and motion state. The system calculates the time delay difference or phase difference between the screen seen by the operator and the current execution time of the execution terminal.
[0367] In one alternative expression, the phase deviation can be represented as:
[0368]
[0369] in, This indicates the time anchor point of the feedback screen seen by the operator. This indicates the time when the executing end is ready to execute the instruction. If If the delay exceeds the safety threshold for the corresponding task, it means that the instruction may be generated based on an expired image and needs to be compensated, downgraded, or refused to be executed.
[0370] In one of the optional decision expressions:
[0371]
[0372] in, This represents the latency threshold related to the risk level of the instruction. For low-risk actions, this threshold can be larger; for high-risk actions such as robotic arm grasping, proximity to hazards, high-speed movement, or operation near personnel, this threshold should be more stringent.
[0373] In addition to time difference, the system can also determine phase based on the spatial deviation between the current state of the execution end and the state seen by the operator. For example:
[0374]
[0375] in, Indicates the current pose of the execution end. This indicates the position and orientation of the execution end as seen by the operator at that moment. When When the spatial threshold is exceeded, the system can consider that the state seen by the operator is no longer suitable for directly driving the current execution end action.
[0376] 5.6 Forward State Prediction and Command Compensation
[0377] To reduce the impact of feedback lag on control execution, the execution end can perform state forward prediction or command compensation based on the time anchor point carried by the control command, the estimated link delay, and the local motion state.
[0378] In one implementation, the execution end depends on its... Given the current state, velocity, attitude, acceleration, and local planning state, the system predicts the state within the current moment or a short future time window. The prediction results can be used to determine whether the original control commands are still valid and whether compensation is needed for the path, velocity, motion amplitude, or target point.
[0379] In one alternative representation, the forward prediction of the execution state can be expressed as:
[0380]
[0381] in, This indicates the current execution state predicted based on the state as seen by the operator at that moment. This indicates the status of the execution terminal at that moment, as seen on the feedback screen by the operator. Indicates the time delay difference. This represents the state prediction function.
[0382] The system can further compare the predicted state with the actual current state at the execution end:
[0383]
[0384] in, This indicates the prediction error. When the prediction error is below a preset threshold, the system can compensate for the original instruction based on the prediction status; when the prediction error exceeds the threshold, the system triggers a degradation, requests confirmation, refuses execution, or performs a safe shutdown.
[0385] In one of the optional decision expressions:
[0386]
[0387] in, This represents the prediction error threshold.
[0388] The state prediction function can be a uniform velocity model, a uniform acceleration model, a vehicle kinematics model, a robotic arm kinematics model, a UAV dynamics model, a local trajectory prediction model, a learning prediction model, or a combination thereof. This prediction function does not limit the scope of protection of this invention; as long as it is used for phase compensation or risk assessment based on time anchor points and the local state of the execution end, it is considered an equivalent implementation of this embodiment.
[0389] 5.7 Phase restrictions for high-risk actions
[0390] For high-risk actions, the system can be configured with stricter phase alignment requirements. These high-risk actions include, but are not limited to: robotic arm grasping and closing; robotic arm turning valves; approaching injured persons or personnel; crossing the boundary of a safety corridor; entering the vicinity of a hazard source; high-speed turning or braking; releasing loads; crossing geofences; performing emergency stop releases; and approaching flammable, explosive, or toxic areas.
[0391] When high-risk actions correspond to , , or If any of the above exceeds the preset threshold, the system will preferably not directly execute the high-risk action, but will instead perform one or more of the following actions: request the operator for secondary confirmation; switch to a low-risk alternative action; limit the speed or range of motion; approach but do not make contact; maintain a safe distance; suspend the task; enter a safe waiting state; trigger a refusal to execute; or trigger a safe shutdown.
[0392] By employing the above methods, the system prevents operators from performing high-risk actions based on delayed visuals, thereby improving the security of remote embodied collaborative control.
[0393] 5.8 Linkage between ROI feedback, phase alignment, and audit receipts
[0394] ROI backhaul, time anchoring, ASW reprojection, phase verification, and instruction compensation are all linked to the anti-interference audit module.
[0395] When the link quality is good, the phase deviation is small, the prediction error is low, and the target confidence meets the requirements, the anti-interference audit module allows the execution end to execute the corresponding instructions. When the link quality deteriorates but the ROI area can still be stably transmitted back, the system can degrade the background video quality, but maintain priority transmission of key target areas and semantic summaries. When the phase deviation or prediction error exceeds the threshold, the anti-interference audit module outputs a degradation, rejection, or safe shutdown result and generates an audit receipt.
[0396] The audit receipt can record the following: the video frame time anchor point to which the instruction is bound; the ROI region seen by the operator when the instruction is generated; phase deviation; prediction error; reprojection uncertainty; link latency and jitter; semantic target confidence; gating decision; degradation action; black box index.
[0397] The above information is used to explain why the system was approved, downgraded, rejected, or shut down, and to provide evidence for subsequent incident review and liability determination.
[0398] 5.9 Relationship with Spatiotemporal Black Box Evidence Preservation
[0399] In this embodiment, the system writes the feedback video frames, ROI regions, semantic summaries, operator inputs, control commands, time anchors, phase deviations, prediction errors, audit receipts, and execution feedback into the spatiotemporal black-box evidence storage chain.
[0400] In one implementation, when an operator performs gaze confirmation or gesture confirmation based on a video frame, the system records the Temporal Anchor, ROI region, target recognition result, operator input, generated task description data, and control commands for that video frame. Upon receiving the command, the execution end records the phase verification result, state forward prediction result, command compensation result, and final execution feedback.
[0401] When collisions, boundary violations, misoperations, refusal to execute, disconnections from autonomous systems, or abnormal high-risk actions occur, the evidence storage module can generate an evidence package based on the above data, enabling subsequent retrospective analysis: which frame the operator saw at the time; what the ROI region was in that frame; which environmental semantic targets the system identified; what input the operator used to generate the instruction; which time anchor point the instruction carried; what the phase deviation was during execution; whether the system performed state prediction or instruction compensation; why the anti-interference audit was passed, downgraded, rejected, or shut down; and what the final feedback from the execution end was.
[0402] Through the aforementioned low-latency feedback transmission, ROI encoding, time anchoring, ASW reprojection, and command-video phase alignment mechanisms, this implementation method can improve the consistency between the information seen by the operator and the actual state of the execution end under conditions of low bandwidth, multi-hop network, long-distance transmission, link jitter, or feedback lag, reduce the risk of operation drift and misjudgment, and provide verifiable timing evidence for subsequent anti-interference auditing, disconnection autonomy, and tamper-proof evidence storage.
[0403] Implementation Method Six: Multi-Entity Consistency Maintenance, Spatial Topology Constraints, and Conflict Arbitration
[0404] like Figure 3 , Figure 5 and Figure 8 As shown, this embodiment illustrates how, when there are multiple embodied execution terminals and / or multiple operators in the system, geometric consistency maintenance, task conflict identification, priority arbitration, demotion processing, and traceable evidence storage can be achieved based on a unified anchor point set, shared coordinate mapping, spatial topology constraints, and anti-interference auditing mechanisms.
[0405] In scenarios such as remote rescue, smart agriculture, industrial inspection, unmanned interactive experiences, or multi-robot collaborative operations, multiple embodied execution terminals may simultaneously reside in the same physical space and receive task instructions from the same operator, multiple operators, cloud / edge scheduling terminals, or virtual content event streams. If different execution terminals lack a shared spatial reference and a unified conflict arbitration mechanism, problems such as path overlap, formation convergence, insufficient minimum safety distance, task area conflicts, overlapping robotic arm operating ranges, or interference between high-risk actions can easily occur.
[0406] To address the aforementioned issues, this implementation method unifies the set of anchor points and spatial topology constraints, thereby enabling collaborative control of multiple embodied execution ends and multiple operators under the same spatial reference and the same audit logic.
[0407] 6.1 Multi-entity shared space reference
[0408] In one implementation, the system establishes a shared spatial reference for multiple embodied execution ends based on the multi-source physical truth coordinate system and spatial anchor point sharing mechanism described in Implementation 2. Each embodied execution end has a corresponding pose, motion state, task state, and safety boundary under a unified truth coordinate system.
[0409] The multiple entities include, but are not limited to: multiple unmanned vehicles; multiple drones; multiple robots; multiple robotic arms; multiple agricultural operation equipment; multiple industrial inspection equipment; multiple special operation equipment; multiple wearable sensing terminals; multiple operators; and cloud / edge scheduling terminals.
[0410] Each executor can be associated with one or more of the following state information: ExecutorID (executor identifier); ExecutorType (executor type); Pose (pose in a unified truth coordinate system); Velocity (velocity); TaskID (current task identifier); RiskLevel (current task risk level); AnchorID (associated spatial anchor); AnchorVersion (anchor version number); SafetyEnvelope (safety envelope); GeoFence (geofence); SafeCorridor (safety corridor); CapabilityProfile (capability profile); ControlMode (control mode); AuditState (audit state); BlackBoxIndex (black-box index).
[0411] The poses and task states of multiple execution ends are all mapped to a unified true coordinate system, enabling the system to calculate the distance, path intersection, area occupancy, action range overlap and potential collision risk between execution ends based on the same spatial reference.
[0412] 6.2 Spatial Topology Constraint Construction
[0413] In one implementation, the system constructs spatial topology constraints based on a shared anchor point coordinate system and task description data. These spatial topology constraints describe the spatial relationships between multiple execution endpoints, multiple task objectives, obstacles, geofences, safety corridors, hazard sources, and work areas.
[0414] Spatial topology constraints include, but are not limited to: minimum safe distance between actuators; minimum safe distance between actuators and personnel; minimum safe distance between actuators and hazard sources; intersection relationships between actuator paths; overlap relationships between actuator action ranges; drone flight path altitude layer constraints; unmanned vehicle path priority; robot formation spacing; robotic arm operating space envelope; geofence boundaries; safety corridor boundaries; task area occupancy status; hazardous area no-entry constraints; and high-risk action mutual exclusion constraints.
[0415] In one alternative expression, the minimum safe distance constraint between multiple execution ends can be represented as:
[0416]
[0417] in, and This indicates two different embodied execution ends. This represents the distance between the two in a unified truth coordinate system. This indicates the minimum safe distance determined based on the task risk level, execution end type, speed status, or safety envelope of both parties.
[0418] When there is an intersection or the distance between the execution end path segments is less than the minimum safe distance, the system identifies it as a potential spatial conflict and submits the conflict to the anti-interference audit module for gating or arbitration.
[0419] 6.3 Potential Conflict Identification
[0420] The system can identify potential conflicts based on the current state of the execution end, task description data, predicted trajectory, and spatial topological constraints.
[0421] The potential conflicts include, but are not limited to: path overlap; flight path intersection; conflict in task area occupancy; insufficient minimum safety distance; overlapping workspace of robotic arms; conflict in drone altitude layers; multiple execution ends entering the same narrow passage at the same time; multiple execution ends approaching the same hazard source at the same time; multiple operators issuing conflicting instructions to the same execution end; the same execution end receiving conflicting instructions from virtual content event streams and manual operation inputs; conflict between high-risk actions and personnel proximity status; and conflict between disconnected autonomous paths and other execution end paths.
[0422] In one optional decision logic, the system can predict the trajectories of multiple execution ends within a future time window and calculate whether there are collisions or spatial conflicts:
[0423]
[0424] or:
[0425]
[0426] in, and These represent the predicted trajectories of the two execution ends within a future time window. This represents the minimum distance between the predicted trajectories of the two trajectories. This represents the minimum safe distance threshold.
[0427] The aforementioned conflict determination can be achieved using rule-based judgment, path geometry calculation, time window prediction, local planner output, space occupancy grid, collision detection model, or a combination thereof.
[0428] 6.4 Conflict Arbitration Strategy
[0429] When the system identifies a potential conflict, the anti-interference audit module arbitrates based on preset priority, risk level, task stage, execution capability, task urgency, minimum safe distance, permission level, and disconnection autonomous status.
[0430] Arbitration results include, but are not limited to: maintaining the original instruction; degrading execution; obstacle avoidance waiting; following mode; alternative trajectory execution; delayed execution; speed-limited passage; replanning the route; requesting manual confirmation; refusing high-risk actions; entering safe waiting; triggering safe shutdown.
[0431] In one implementation, the system prioritizes execution devices with higher risk levels, greater task urgency, or higher safety priority. For example, in rescue scenarios, rescue robots closer to the injured may have higher priority than material transport robots; in agricultural scenarios, drones returning to base with low battery levels may have priority over ordinary spraying drones; and in hazardous industrial operations, robotic arms or robots near the hazard source may have priority in obtaining safe evacuation routes.
[0432] In another implementation, when multiple execution terminals receive potentially conflicting instructions, the execution terminal with higher priority or higher risk level retains the original instruction, while the remaining execution terminals automatically downgrade their priority and enter obstacle avoidance waiting mode, follow-up mode, alternative trajectory execution mode, or delayed execution mode.
[0433] In one alternative expression, the arbitration strategy can be represented as:
[0434]
[0435] in, Indicates the first Arbitration priority for each execution end or task Indicates the risk level. Indicates the urgency of the task. Indicates the permission level. Indicates the execution end status. Indicates the task phase. This indicates the priority calculation function.
[0436] When a conflict occurs between two execution ends:
[0437]
[0438] The priority calculation function described above can be implemented using rule tables, task templates, risk matrices, manual strategies, scheduling algorithms, or combinations thereof.
[0439] 6.5 Handling of Multi-Operator Instruction Conflicts
[0440] In scenarios with multiple operators, different operators may issue different control intentions towards the same execution end, the same target, or the same spatial area. The system can arbitrate multi-operator instructions based on operator permission level, task role, time anchor point, instruction risk level, and current task stage.
[0441] For example, when one operator issues a "approach target" command and another operator issues a "stop" command, the system can determine which command to execute first based on the operator's permission level, task risk level, and the current state of the execution end; when one operator issues a high-risk robotic arm grab command and another operator issues a evacuation command, the system can prioritize executing the safer or higher-level command.
[0442] The results of multi-operator instruction conflict handling are also written into the audit receipt, recording the instruction source, permission level, cause of conflict, arbitration result, and black-box index.
[0443] 6.6 Multi-entity conflict handling under disconnected autonomous status
[0444] When a particular execution terminal is in a disconnected autonomous state, the system still needs to prevent its autonomous path from conflicting with the paths of other execution terminals.
[0445] In one implementation, the executor receives the autonomous script, fallback path, security boundary, and task continuation status before entering the disconnected autonomous state. The autonomous script may contain a security envelope and obstacle avoidance rules during the disconnection period. When the cloud / edge or other executors are still online, the system can incorporate the expected autonomous path of the disconnected executor into spatial topology constraints and implement avoidance or demotion strategies for other executors.
[0446] Once the disconnected execution end link is restored, the system transmits back the actual execution trajectory, obstacle avoidance events, and task status during the disconnection period and compares them with the original autonomous script. If path deviation, boundary crossing, collision risk, or multi-entity conflict events occur during the disconnection period, the system writes the relevant data into the spatiotemporal black-box evidence storage chain.
[0447] 6.7 Arbitration Result and Audit Receipt
[0448] Each multi-entity conflict identification and arbitration generates an audit receipt. The audit receipt includes, but is not limited to: ConflictID, conflict identifier; involved ExecutorID; involved TaskID; conflict type; conflict location; predicted conflict time; relevant spatial anchors; anchor version number; risk level; priority calculation result; arbitration decision; demotion target; alternative trajectory or waiting strategy; operator permission level; time anchor; black-box index.
[0449] Through audit receipts, the system can explain why one execution end maintains the original instruction, and why another execution end is demoted, waits, rerouted, or refuses to execute. This makes the multi-entity collaborative process interpretable and traceable.
[0450] 6.8 Collaboration with Standardized Task Interfaces
[0451] Multi-entity consistency maintenance and conflict arbitration can work in conjunction with the standardized task interface described in Implementation Method 3. Fields such as GeoFence, SafeCorridor, RiskLevel, Priority, MotionConstraint, PathConstraint, PermissionLevel, AuditPolicy, and FailOperationalPolicy in the task description data can be used for conflict detection and arbitration.
[0452] For example, in smart agriculture scenarios, multiple drones or unmanned vehicles receive task description data for the same plot of land. The system can generate spatial topology constraints based on plot boundaries, flight paths, work areas, spraying sequence, and minimum safe distance. When two flight paths intersect, the system determines which device goes first based on task priority and remaining battery power, while the other device either waits or reroutes.
[0453] In remote rescue scenarios, when multiple robots enter the same narrow passage, the system can arbitrate based on the urgency of the mission, the distance to the target, the risk level, and the width of the safety corridor. The robot closest to the injured person maintains its original path, while the material delivery robot enters a waiting or alternative path.
[0454] In industrial inspection scenarios, when multiple robotic arms or robots are working near the same workstation, the system can perform mutual exclusion control based on the robotic arm's workspace envelope, task stage, and action risk level to prevent two robotic arms from entering overlapping workspaces at the same time.
[0455] 6.9 Relationship with Spacetime Black Box Evidence Preservation
[0456] Multi-entity conflict identification, arbitration, and demotion processing are all linked to the spatiotemporal black-box evidence storage module. When path conflicts, task area conflicts, robotic arm workspace conflicts, multi-operator command conflicts, disconnected autonomous path conflicts, or insufficient minimum safety distances occur, the system can write the following data into the evidence package: execution end identifiers; operator identifiers or permission levels; task description data; spatial anchor points and version numbers; temporal anchor points; predicted trajectory; conflict type; spatial topology constraints; priority calculation results; arbitration results; demotion actions; alternative trajectories; execution feedback; audit receipts; and black-box indexes.
[0457] During incident debriefing or regulatory review, the system can trace back the spatial location, task status, conflict judgment, arbitration basis, and final actions of each execution end in the multi-entity collaboration process, avoiding the need to rely solely on a single device log or manual description for liability determination.
[0458] Through the aforementioned multi-entity consistency maintenance, spatial topology constraints, and conflict arbitration mechanism, this implementation method can maintain collaborative security under a unified spatial reference in complex scenarios where multiple embodied execution terminals, multiple operators, or multiple tasks are running simultaneously. It reduces path conflicts, task conflicts, and high-risk action conflicts, and makes the arbitration process explainable, auditable, and traceable.
[0459] Implementation Method Seven: Spatiotemporal Black Box, Audit Receipts, and Tamper-proof Evidence Preservation
[0460] like Figure 8 As shown, this embodiment illustrates how the system associates and records virtual content status, environmental semantic targets, task description data, control instructions, audit receipts, execution feedback, link / confidence summaries, anchor version numbers, time anchors, and abnormal events, and generates an immutable evidence package when a trigger event occurs, thereby forming a spatiotemporal black box evidence chain that is traceable, verifiable, and auditable.
[0461] In cross-space embodied collaborative control scenarios, there are complex relationships between the virtual content seen by the operator, the physical space where the execution terminal is located, the environmental semantic targets identified by the system, the link status at the time of instruction generation, the spatial anchor version, the time anchor, the anti-interference audit results, and the actual actions of the execution terminal. Simply saving ordinary operation logs or video clips is insufficient to fully prove under what spatiotemporal reference, input basis, audit decision, and execution feedback a particular action was generated. Therefore, this implementation method uses a spatiotemporal black box mechanism to structurally associate and record key data objects in the control closed loop.
[0462] 7.1 Spatiotemporal Black Box Recording Objects
[0463] The spatiotemporal black box is used to record key data objects related to cross-space embodied collaborative control. These key data objects include, but are not limited to: virtual content state; virtual narrative events; environmental semantic targets; task description data; operator input; control commands; audit receipts; execution feedback; link quality summaries; location / perception confidence summaries; spatial anchor point identifiers; spatial anchor point version numbers; temporal anchor points; multi-entity conflict arbitration results; disconnection prediction state; disconnection autonomous scripts; execution end state; manual takeover records; abnormal events; and evidence package indexes.
[0464] The virtual content state includes one or more of the following: virtual scene, narrative beat, event identifier, virtual trigger point, plot stage, audio envelope, or interactive state.
[0465] The environmental semantic targets include one or more of the following: casualties, hazards, obstacles, passageways, valves, meters, tools, work areas, geofences, robotic arm targets, agricultural plot boundaries, waypoints, or other identifiable objects.
[0466] The task description data includes one or more of the following: target identifier, task type, risk level, spatial anchor, temporal anchor, geofence, security corridor, permission level, action constraints, path constraints, audit strategy, and black-box index.
[0467] The operator input includes one or more of the following: gaze, gesture, posture, gait, voice input, bioelectrical signals, confirmation actions, task selection input, or manual takeover input.
[0468] The execution feedback includes one or more of the following: video frame, ROI region, semantic summary, sensor observation, state telemetry, execution trajectory, robotic arm joint state, end effector state, torque feedback, gripper state, obstacle avoidance event, disconnected autonomous trajectory, or safe shutdown state.
[0469] 7.2 Black-box indexes and data association
[0470] To ensure traceability between different data objects, the system can generate a black-box index for each control loop event. This black-box index is used to associate virtual content states, operator inputs, task description data, control instructions, audit receipts, and execution feedback within the same event chain.
[0471] In one implementation, the black-box index may be generated from one or more of the following information: task identifier; instruction identifier; time anchor; spatial anchor identifier; anchor version number; executor identifier; operator identifier or permission level; event type; trigger event number; system local serial number.
[0472] In one alternative representation, the black-box index can be represented as:
[0473]
[0474] in, Indicates a black-box index. This represents a hash function or index generation function. This expression is only an optional implementation; black-box indexes can also be generated from sequence numbers, timestamps, task numbers, database primary keys, cryptographic digests, or combinations thereof.
[0475] Through black-box indexing, the system can bind a control command to the context data before and after its generation. For example, a robotic arm grasping command can be linked to the same chain of evidence as the video frame seen by the operator, the gaze target, the gesture confirmation result, the target semantic confidence, the spatial anchor version, the anti-interference audit result, the robotic arm execution feedback, and the gripper torque change.
[0476] 7.3 Evidence Fields for Audit Receipts
[0477] Audit receipts are an important component of the spacetime black box, used to record why the system approves, downgrades, rejects, or safely shuts down a candidate instruction.
[0478] In one implementation, the audit receipt includes one or more of the following fields: AuditID (audit receipt identifier); CommandID (associated instruction identifier); TaskID (task identifier); DecisionType (decision type); TriggerReason (trigger reason); KeyMetrics (key input metrics); RiskLevel (risk level); PermissionLevel (permission level); DegradeAction (degrade action); TemporalAnchor (temporal anchor); AnchorID (spatial anchor identifier); AnchorVersion (anchor version number); ExecutorID (executor identifier); ExecutorState (executor state summary); LinkState (link state summary); ConfidenceSummary (confidence summary); ConflictID (multi-entity conflict identifier); FailOperationalState (disconnection autonomous state); BlackBoxIndex (black box index); and Signature (signature or summary field).
[0479] DecisionType includes Pass, Degrade, Deny, or Safe Shutdown. TriggerReason can include link jitter exceeding limits, positioning residual exceeding limits, anchor point version inconsistency, insufficient target confidence, insufficient permissions, abnormal execution end status, geofence conflict, multi-entity conflict, disconnection prediction trigger, or phase deviation exceeding limits.
[0480] KeyMetrics may include link latency, link jitter, packet loss rate, positioning residual, anchor confidence, reprojection uncertainty, prediction error, semantic target confidence, execution speed, execution attitude, torque state or gripper state, etc.
[0481] By recording audit receipts, the system can explain afterward why a command was allowed to be executed, why it was downgraded, why it was rejected, or why a safety shutdown was triggered.
[0482] 7.4 Evidence Package Trigger Event
[0483] The system can generate an immutable evidence package when a triggered event occurs. These triggered events include, but are not limited to: collision; boundary crossing; emergency stop; refusal to execute; safe shutdown; high-risk actions; robotic arm gripping; valve turning; proximity of a hazard source; operation near personnel; connection loss prediction; advance issuance of autonomous scripts; entering a disconnected autonomous system; link recovery; anchor point realignment; anchor point version inconsistency; multi-entity conflict arbitration; alternative trajectory execution; permission escalation; manual takeover; abnormal task termination; abnormal decrease in target semantic confidence; phase deviation or prediction error exceeding limits.
[0484] In one implementation, the system can set different evidence package triggering strategies based on the task risk level. For low-risk tasks, the system can seal the evidence package only when an abnormal event occurs; for high-risk tasks, the system can generate a continuous chain of evidence for each key instruction, each audit decision, or each execution feedback; for scenarios with high regulatory requirements, the system can continuously record summaries and seal data within a time window before and after the triggering event occurs.
[0485] 7.5 Contents of the Evidence Package
[0486] The evidence package may include one or more of the following: evidence package identifier; black-box index; event type; event occurrence time; event occurrence location; spatial anchor identifier; anchor version number; temporal anchor; virtual content status; environmental semantic target; task description data; operator input; operator-seen video frame summary; ROI region; control instructions; audit receipt; execution feedback; execution end trajectory; sensor observation summary; link quality summary; positioning / perception confidence summary; disconnected autonomous script version; disconnected autonomous execution record; multi-entity conflict arbitration result; manual takeover record; digital signature; chained hash digest; trusted timestamp; trusted execution environment proof information.
[0487] In one implementation, the evidence package does not necessarily store the complete video or raw sensor data, but may store keyframe summaries, hash summaries, semantic summaries, time anchors, anchor version numbers, audit receipts, and execution status summaries. For high-risk scenarios requiring complete traceability, the system can simultaneously save video clips, raw sensor data, or high-frequency status logs within the corresponding time window.
[0488] 7.6 Tamper-proof protection mechanism
[0489] To ensure the credibility of the evidence package in incident review, liability determination, third-party audits, or regulatory inspections, the system can use chained hashing, digital signatures, trusted execution environments, trusted timestamps, encrypted storage, or a combination thereof to protect the evidence package from being tampered with.
[0490] In one alternative implementation, the system performs chain hashing on the continuously generated evidence records:
[0491]
[0492] in, Indicates the first The hash value of each piece of evidence record. This represents the hash value of the previous evidence record, and DatanData_nDatan represents the data content of the nth evidence record. This indicates the time anchor or credible timestamp corresponding to the evidence record.
[0493] When any piece of evidence is tampered with, its corresponding hash value and subsequent chain hashes will change, thus making it detectable.
[0494] In another implementation, the system can use digital signatures to sign the evidence package. The signing entity can be an embodied execution device, a security protection layer, a cloud / edge computing device, a trusted execution environment, or a trusted third-party node. The signature of the evidence package can be used to prove that the evidence package was generated by a legitimate system module and has not been tampered with after being signed.
[0495] In another implementation, the system can deploy critical audit logic or evidence generation logic within a trusted execution environment. This trusted execution environment can be used to protect audit receipts, black-box indexes, chained hashes, and signing keys, making it difficult for external application layers to forge or modify critical evidence data.
[0496] The aforementioned immutability protection mechanism is merely an optional implementation and does not limit the scope of protection of this invention. Any mechanism that can achieve integrity verification, tamper-proof protection, or reliable traceability of the evidence package constitutes an equivalent implementation of the immutable evidence storage mechanism described in this invention.
[0497] 7.7 Time Window and Backtracking Range
[0498] In one implementation, the system can configure different backtracking time windows for different triggering events. The backtracking time window is used to determine which virtual content states, feedback screens, instructions, audit receipts, and execution end states should be sealed before and after the triggering event occurs.
[0499] For example, for ordinary degradation events, the system can save summary data within a short time window before and after the triggering event; for collision, boundary crossing, high-risk robotic arm actions, or disconnection autonomous events, the system can save video frame summaries, sensor observations, task description data, instruction chains, and execution feedback within a longer time window before and after the triggering event.
[0500] In one alternative representation, the time window corresponding to the evidence package can be expressed as:
[0501]
[0502] in, This indicates the evidence package backtracking window. Indicates the time when the triggering event occurred. This indicates the length of time the item needs to be sealed before the event occurs. This indicates the length of time that needs to be sealed after the event occurs.
[0503] The time window can be dynamically adjusted based on event type, risk level, storage capacity, regulatory requirements, or system configuration.
[0504] 7.8 Evidence-based association with instruction-video phase alignment
[0505] Regarding the low-latency feedback return and command-video phase alignment described in Implementation 5, the spatiotemporal black box in this implementation can record the time anchor point, ROI region, semantic target, ASW reprojection state, phase deviation, state prediction error, and reprojection uncertainty of the video frame seen when the operator generates a command.
[0506] For example, when an operator issues a "move closer to target" command based on a video frame, the system records the Temporal Anchor, ROI region, target semantic confidence, and operator input for that video frame. Upon receiving the command, the execution terminal records the phase deviation between the current execution terminal and the view seen by the operator, the state prediction result, and the audit decision. Whether the command is executed, downgraded, or rejected, the relevant results are all written to the same black-box index.
[0507] Therefore, during subsequent review, it can be determined whether the operator issued instructions based on the delayed screen, whether the system performed phase compensation, why the anti-interference audit module allowed or denied the instruction, and whether the final action of the execution end complied with the security policy.
[0508] 7.9 Evidence related to the disconnection and autonomy
[0509] For the disconnection prediction and disconnection autonomy described in Implementation Method 4, the spatiotemporal black box can record the link deterioration trend, disconnection risk score, autonomous scripts issued in advance, fallback path, security boundary, task continuation status, last effective strategy, and the actual running status of the execution end during the disconnection period.
[0510] Once the link is restored, the execution end will transmit back the trajectory, obstacle avoidance events, speed status, task stage, and abnormal events during the disconnection period, and compare them with the original autonomous script. The system can write the comparison results into an evidence package to prove whether the execution end operated according to the preset security policy during the disconnection period.
[0511] For example, in a rescue scenario, if the execution end retreats to a safe point during a disconnection, the system records the link status before the disconnection, the reason for triggering the disconnection prediction, the issued retreat path, the actual trajectory of the execution end, and the status after the link is restored. If a deviation or collision risk occurs during the disconnection, the system records the corresponding abnormal event and local obstacle avoidance decision.
[0512] 7.10 Evidence Preservation Related to Multi-Entity Conflict Arbitration
[0513] For the multi-entity consistency maintenance and conflict arbitration described in Implementation Method 6, the spatiotemporal black box can record information involving the execution end identifier, task description data, predicted trajectory, spatial topology constraints, conflict type, minimum safe distance, priority calculation result, arbitration result, demotion action, alternative trajectory, and execution feedback.
[0514] For example, when the paths of two execution terminals conflict, the system records the predicted trajectories, spatial anchor version numbers, minimum safe distance thresholds, and arbitration results of both execution terminals in a unified truth coordinate system. If one execution terminal is demoted and enters waiting mode, the system records the reason for the demotion and the waiting strategy. If a collision or task delay subsequently occurs, the arbitration basis and execution process can be traced back through the evidence package.
[0515] 7.11 Evidence Package Inquiry and Audit Output
[0516] In one implementation, the system can query the evidence package based on task identifier, instruction identifier, execution terminal identifier, operator permission level, spatial anchor point, time anchor point, event type, or black-box index.
[0517] The system can output audit reports. These audit reports include, but are not limited to: event overview; triggering cause; relevant video frames or semantic summaries; task description data; control instruction chain; audit receipt chain; execution-end feedback chain; spatial anchor version changes; temporal anchor sequence; phase deviation or prediction error; disconnection autonomous process; multi-entity arbitration process; and evidence integrity verification results.
[0518] This audit report can be used for incident review, liability determination, regulatory inspection, training improvement, model optimization, or operation and maintenance assessment.
[0519] Through the aforementioned spatiotemporal black box, audit receipt, and tamper-proof evidence storage mechanism, this implementation method can link key data objects such as virtual content, spatial anchors, temporal anchors, semantic targets, task descriptions, control instructions, audit decisions, execution feedback, disconnected autonomy, and multi-entity arbitration in the cross-space embodied collaborative control process into a complete chain of evidence, enabling the system to have traceability, verifiability, auditability, and tamper-proof capabilities.
[0520] Implementation Method 8: Edge-Cloud Collaborative Computing, Task Offloading, and Computing Power Degradation
[0521] like Figure 10 As shown, this embodiment illustrates how wearable sensing terminals, cloud / edge computing terminals, and embodied execution terminals dynamically allocate tasks such as positioning, identification, map optimization, low-latency backhaul, auditing decisions, disconnection autonomy, and evidence aggregation based on computing power status, link quality, remaining power, core temperature, task risk level, and positioning / perception confidence, in order to maintain the minimum available closed loop for cross-space embodied collaborative control under different network conditions, different computing power conditions, and different security levels.
[0522] In cross-space embodied collaborative control systems, wearable sensing devices are typically limited by size, power consumption, heat dissipation, and battery life, making it difficult to handle high-computational tasks such as high-precision SLAM, global map optimization, real-time semantic recognition, high-density point cloud processing, prediction compensation, ROI encoding, and black-box evidence aggregation for extended periods. Cloud / edge computing devices possess strong computing power but rely on communication links; embodied execution devices are closest to the physical execution mechanism and are suitable for retaining critical closed loops such as low-latency control, security auditing, obstacle avoidance, disconnection autonomy, and secure shutdown. Therefore, this implementation method achieves dynamic deployment and degraded operation of different computing tasks through an end-edge-cloud collaborative computing and task offloading mechanism.
[0523] 8.1 Role Division of End-Edge-Cloud Computing
[0524] In one implementation, the wearable sensing device primarily undertakes low-latency tasks related to real-time interaction with the operator, including but not limited to: operator 6-DoF pose tracking; gaze detection; gesture recognition; basic posture recognition; lightweight VIO; virtual content display; ASW reprojection; operator input acquisition; task selection interface; ROI region marking; and audit result prompts.
[0525] Cloud / edge computing terminals primarily undertake global or high-computing-power tasks, including but not limited to: global map optimization; loop closure detection; high-density point cloud processing; semantic recognition; semantic segmentation; object detection; ROI encoding; global scheduling of multiple execution terminals; spatial anchor point synchronization; task template management; standardized task interface parsing; model inference; evidence aggregation; evidence package management; and historical data analysis.
[0526] The embodied execution end mainly undertakes low-latency tasks directly related to physical motion safety, including but not limited to: local positioning fusion; local obstacle avoidance; local path planning; motion control; robotic arm control; end effector control; anti-interference audit key judgment; disconnection autonomy; safe shutdown; execution feedback collection; status telemetry; local black box summary recording.
[0527] The above task division is not fixed. The system can dynamically migrate tasks between the three types of endpoints based on link status, computing load, task risk level, and device status.
[0528] 8.2 Task Unloading Triggering Conditions
[0529] The system can trigger task offloading or computing power degradation based on one or more of the following conditions: the remaining battery power of the wearable sensing terminal is lower than a preset threshold; the core temperature of the wearable sensing terminal is higher than a preset threshold; the computing power load of the wearable sensing terminal exceeds a preset threshold; the cloud / edge link latency is lower than a preset threshold and stable; the link jitter exceeds a preset threshold; there is insufficient bandwidth; the confidence level of positioning / perception decreases; the complexity of semantic recognition tasks increases; the execution terminal enters a high-risk task stage; the execution terminal enters a weak network area; the system enters a disconnection prediction state; the priority of user interaction tasks is higher than that of background mapping tasks; and the evidence storage event requires additional storage or computing resources.
[0530] For example, when the remaining battery power of the wearable sensing device is lower than a preset threshold or the core temperature is higher than a preset threshold, the system can offload high-performance SLAM, global map optimization, loop closure detection, high-density feature extraction or semantic recognition tasks to the cloud / edge computing end, while retaining lightweight VIO, gaze detection, ASW reprojection and task input acquisition on the wearable sensing device.
[0531] When link quality degrades or link jitter increases, the system can reduce its real-time reliance on the cloud / edge, decentralize key auditing, obstacle avoidance, local planning, and disconnection autonomy to the local execution end, reduce the bitrate of video background areas, and prioritize the transmission of ROI areas and semantic summaries.
[0532] 8.3 Optional Task Unloading Determination Model
[0533] In one alternative implementation, the system can establish an offload score for various computing tasks. The offload score is determined based on task computational load, link quality, energy consumption, thermal status, task risk level, and real-time requirements.
[0534] Task unloading rating can be expressed as:
[0535]
[0536] in, Indicates the first Unloading rating for similar tasks This indicates the computational complexity of the task. Indicates the current link quality. Indicates the remaining power or energy consumption status on the terminal side. Indicates the end-side core temperature. Indicates the task risk level. This indicates the real-time requirements of the task. This indicates the unloading decision function.
[0537] When the unloading score of a task exceeds a preset threshold, the system migrates the task from the wearable sensing end to the cloud / edge computing end or the embodied execution end:
[0538]
[0539] in, This indicates the uninstallation threshold.
[0540] The aforementioned unloading decision function can be implemented using a rule table, a weighted function, a state machine, an optimization model, a reinforcement learning model, or a combination thereof. This formula is merely one possible expression and does not limit the scope of protection of this invention.
[0541] 8.4 Computing Power Degradation Strategy
[0542] When the system detects insufficient computing power, insufficient power, excessive temperature, or insufficient link quality on the edge side, it can trigger a computing power degradation strategy. This degradation strategy includes, but is not limited to: switching from high-precision SLAM to lightweight VIO; switching from global map optimization to local map tracking; switching from high-density point cloud processing to sparse feature tracking; switching from full-frame semantic segmentation to ROI region semantic recognition; switching from high-frame-rate video backhaul to low-frame-rate video backhaul; switching from background video backhaul to semantic summary backhaul; switching from complex motion prediction to conservative motion models; switching from multi-target tracking to key target tracking; switching from cloud inference to edge inference; and switching to an execution-side local security strategy when edge inference is unavailable.
[0543] High-risk actions will be switched to requesting confirmation or prohibiting execution.
[0544] In one implementation, when the remaining battery power of the wearable sensing device is below a preset threshold, the system shuts down or reduces non-critical rendering, background semantic recognition, and high-density mapping tasks, retaining only pose tracking, gaze detection, task input, ROI marking, and audit prompts. When the core temperature exceeds a preset threshold, the system reduces the continuous inference load on the edge, migrating global optimization and semantic recognition to the cloud / edge.
[0545] In another implementation, when link quality degrades, the system reduces its reliance on real-time inference at the cloud / edge, prioritizes local obstacle avoidance, local secure shutdown, and disconnection autonomy at the execution end, and switches video backhaul to ROI-priority and semantic summary-priority modes.
[0546] 8.5 Minimum Available Closed Loop
[0547] To ensure the system can still operate safely under conditions of limited computing power or damaged links, this invention establishes a minimum availability closed loop. The minimum availability closed loop means that even when the cloud / edge computing terminal is unavailable, the computing power of the wearable sensing terminal is degraded, or the communication link is damaged, the system still retains at least the basic functions directly related to safe operation.
[0548] In one implementation, the minimum available closed loop includes: local positioning or short-term pose estimation at the embodied executor; local obstacle avoidance; local motion control; geofencing or security boundary constraints; the most recent valid spatial anchor version; the most recent valid task description data; the last valid policy; basic determination of interference resistance audit; disconnection autonomy; secure shutdown; and local black-box summary recording.
[0549] When the system enters the minimum available closed loop, high-risk actions can be paused, non-critical semantic recognition can be stopped, video transmission quality can be reduced, and some virtual content interactions can be frozen, retaining only necessary security prompts, status transmissions, and low-risk controls. After the link or computing power status is restored, the system will gradually restore global map optimization, semantic recognition, complete video transmission, multi-entity scheduling, and evidence aggregation.
[0550] 8.6 Task Priority Scheduling
[0551] In the edge-cloud collaborative computing process, the system can schedule tasks based on their priority. Task priority can be determined by the task risk level, real-time requirements, operator interaction needs, security relevance, and evidence storage requirements.
[0552] In one implementation, the tasks, from highest to lowest priority, include: safe shutdown; local obstacle avoidance; anti-interference auditing; disconnection autonomy; motion control; positioning / attitude tracking; high-risk action confirmation; ROI critical area backhaul; state telemetry; time anchor point synchronization; black box summary recording; task-level instruction parsing; semantic target recognition; global map optimization; background video backhaul; and non-critical virtual content rendering.
[0553] When computing power is insufficient, the system prioritizes ensuring safe shutdown, local obstacle avoidance, anti-interference auditing, disconnection autonomy, and motion control; when bandwidth is insufficient, the system prioritizes ensuring control commands, audit receipts, status telemetry, ROI key areas, and semantic digests; when storage resources are insufficient, the system prioritizes saving evidence related to triggering events, audit receipts, time anchors, and hash digests.
[0554] 8.7 Collaboration with Standardized Task Interfaces
[0555] Edge-cloud collaborative computing can work in conjunction with the standardized task interface described in Implementation Method 3. The task description data may include fields related to computing power scheduling, such as task risk level, real-time requirements, semantic recognition requirements, ROI region, disconnection autonomy strategy, audit strategy, evidence storage strategy, and execution end capability requirements.
[0556] For example, in remote rescue missions, mission description data can mark "wounded targets" as high-priority ROI areas and require edge devices to prioritize wounded identification and ROI encoding; when the link deteriorates, the system retains high-quality backhaul of wounded areas while reducing the video quality of background areas.
[0557] In industrial inspection tasks, task description data can mark "valve turning" as a high-risk action, requiring the execution end to retain torque monitoring, permission confirmation, anti-interference audit, and black-box summary records locally. Even if the cloud / edge end is temporarily unavailable, the local security judgment must not be bypassed.
[0558] In multi-entity agricultural operations, task description data can mark the operating areas, routes, and priorities of different drones, with the cloud / edge end responsible for global scheduling. When the link drops, each execution end continues to execute low-risk sub-tasks according to the local route and geofence or enters waiting mode.
[0559] 8.8 Synergy with Low-Latency Backhaul and ROI Encoding
[0560] Edge-cloud collaborative computing can also work in conjunction with the ROI hierarchical coding and low-latency feedback backhaul described in Implementation Method 5.
[0561] When the link quality is good and the edge computing power is sufficient, the system can perform semantic segmentation, object detection and ROI encoding at the edge, and simultaneously transmit high-quality ROI video, background compressed video and semantic summary back to the wearable sensing end.
[0562] When link quality degrades, the system can reduce the bitrate in the background area, increase the transmission priority of the ROI area, and offload some semantic recognition tasks to the embodied execution end locally. If the link deteriorates further, the system can stop the backhaul of background video and only backhaul key target screenshots, semantic summaries, status telemetry, and audit receipts. If the link is close to interruption, the system triggers a disconnection prediction and prioritizes the distribution of autonomous scripts, fallback paths, and the last effective strategy to the embodied execution end.
[0563] 8.9 Synergy with Interference-resistant Auditing and Disconnection Autonomy
[0564] End-edge-cloud collaborative computing also works in conjunction with anti-interference auditing and disconnection-based autonomous operation as described in Implementation Method 4.
[0565] In one implementation, the cloud / edge computing endpoint can execute more complex global audit strategies, such as multi-entity conflict detection, global path risk assessment, and historical state analysis; the embodied execution endpoint retains basic audit strategies locally, such as link failure determination, location information insufficiency determination, geofencing constraints, local obstacle avoidance, and secure shutdown.
[0566] When the cloud / edge endpoint is available, the system can combine the audit results from the cloud / edge endpoint and the local audit results from the execution end; when the cloud / edge endpoint is unavailable, the local audit results from the execution end take precedence. For high-risk actions, the system can require local audits to pass, and even if the cloud / edge endpoint returns a permitted result, the local security protection of the execution end must not be bypassed.
[0567] When the disconnection prediction is triggered, the edge-cloud collaborative computing mechanism prioritizes the transmission and local storage of autonomous scripts, fallback paths, security boundaries, task continuation status, and black-box summaries.
[0568] 8.10 Collaboration with Spatiotemporal Black Box Evidence Preservation
[0569] The edge-cloud collaborative computing also works in conjunction with the spatiotemporal black-box evidence storage described in Implementation Method 7.
[0570] In one implementation, the wearable sensing device records the operator's input, the time anchor point of the seen video frame, the gaze area, and the interaction state; the embodied execution device records control commands, execution feedback, local audit results, disconnection autonomous status, and security shutdown events; and the cloud / edge computing device aggregates task description data, semantic recognition results, global scheduling results, multi-entity arbitration results, and evidence package summaries.
[0571] When the link is normal, the evidence data from the three ends can be aggregated to the cloud / edge to form a complete evidence package; when the link is damaged, each end can save the evidence summary locally, and merge it according to the time anchor, anchor version number, task identifier and black box index after the link is restored.
[0572] In one implementation, to avoid evidence loss due to link interruption, the executing end can retain a minimal black-box record locally, including control commands, audit receipts, execution status, disconnection autonomy strategies, key sensor summaries, and chained hash digests from the most recent period. After the link is restored, this local black-box record is reconciled with the cloud / edge record to verify the integrity of the evidence chain.
[0573] 8.11 Equivalent Implementation of Collaborative Computing
[0574] The aforementioned edge-cloud collaborative computing, task offloading, and computing power degradation can be achieved through rule tables, state machines, schedulers, containerized task migration, edge inference services, device driver adaptation layers, task queues, priority queues, or combinations thereof.
[0575] In other implementations, the functions of the wearable sensing end, cloud / edge computing end, and embodied execution end can also be partially combined. For example, the edge computing end can be deployed near the embodied execution end or integrated into the execution end itself; some tasks of the cloud / edge computing end can be undertaken by on-site private network nodes; and the wearable sensing end can also undertake more semantic recognition or map optimization tasks in a high-performance form. As long as it can dynamically allocate, downgrade, or migrate cross-space embodied collaborative control tasks based on computing power, link, energy consumption, temperature, and task risk status, it constitutes an equivalent implementation of the end-edge-cloud collaborative computing mechanism described in this invention.
[0576] Through the aforementioned end-edge-cloud collaborative computing, task offloading, and computing power degradation mechanisms, this implementation method can maintain the operation of key functions such as positioning, auditing, control, disconnection autonomy, and evidence storage in a prioritized manner even when the computing power of the wearable sensing end is limited, the link quality fluctuates, the cloud / edge end is unavailable, or the execution end enters a high-risk task phase. This improves the robustness, maintainability, and engineering implementation capability of the cross-space embodied collaborative control system.
[0577] Implementation Method Nine: Specific Application Examples
[0578] This embodiment is used to further illustrate the deployment method and workflow of the present invention in different application scenarios. The following embodiments are only used to illustrate the feasibility and scope of application of the present invention, and do not limit the present invention to be applied only to the following scenarios. The technical features in each embodiment, such as wearable sensing terminal, cloud / edge computing terminal, embodied execution terminal, multi-source physical truth fusion, spatial anchor point sharing, temporal anchor point, standardized task interface, anti-interference audit, disconnection autonomy, ROI backhaul, and black-box evidence storage, can be combined or equivalently replaced according to specific scenarios.
[0579] Example 1: Narrative / Rhythm-Driven Unmanned Entity Collaborative Performance Control
[0580] In cultural and tourism interactive scenarios, immersive performances, or unmanned physical experience scenarios, operators wear wearable sensing devices to view or control virtual content that includes plot beats, audio envelopes, virtual trigger points, and key events. The system extracts temporal features such as BeatID, BPM, EventID, AudioEnvelope, TriggerPoint, and TemporalAnchor from the virtual content event stream and inputs them into a multimodal instruction mapping engine.
[0581] The multimodal instruction mapping engine uses narrative / rhythm-driven mapping logic to map beats, plot phases, or audio peaks in virtual content to the speed, acceleration, steering, braking, formation changes, or path constraints of autonomous vehicles, robots, or formation devices. For example, when the plot enters an acceleration phase, the system generates a corresponding speed target; when a virtual object appears near a spatial anchor point, the system generates actions such as approaching, circling, avoiding, or stopping.
[0582] The unified spatiotemporal alignment engine establishes a multi-source physical ground truth coordinate system based on at least two types of observations, such as RTK / GNSS, UWB, SLAM / VIO, or IMU, and ensures that the location of virtual events is consistent with the physical movement location of unmanned entities through a spatial anchor point sharing mechanism. The safety protection layer audits candidate commands based on link quality, positioning reliability, anchor point consistency, and path congestion. When link jitter increases, positioning residuals rise, or visitors are too close, the system can implement speed limiting, tighten geofencing, delay actions, or perform a safe shutdown.
[0583] The evidence storage module records the complete chain of "virtual beat - task description - control command - audit receipt - execution feedback". When an emergency stop, boundary violation, collision, or manual takeover occurs, the system generates an immutable evidence package for subsequent review and operational safety audits.
[0584] Example 2: Immersive Semantic Command for Remote Rescue
[0585] In fire scenes, collapses, underground spaces, tunnels, or disaster sites, operators use wearable sensing devices to view first-person perspective videos, sensor observations, and semantic summaries transmitted from embodied actuators. Embodied actuators can be quadruped robots, tracked robots, unmanned vehicles, or drones, and possess local obstacle avoidance, low-speed autonomy, and safe shutdown capabilities.
[0586] The execution unit acquires video frames, thermal imaging, depth information, gas sensor data, IMU data, and location status. The cloud / edge computing end or the execution unit locally identifies casualties, obstacles, hazards, passable pathways, and safe areas. The system performs hierarchical ROI encoding based on the operator's gaze area, the casualty target, or the hazard source area, ensuring that critical areas are transmitted with higher quality or priority than background areas. Each video frame or semantic summary is embedded with a time anchor.
[0587] When the operator observes the wounded and performs a confirmation gesture, the multimodal command mapping engine and semantic control layer generate task-level commands such as "approach the target and maintain a safe distance," "mark the location," or "deliver supplies." A standardized task interface encapsulates these commands as task description data, with fields including TargetID, RiskLevel, AnchorID, TemporalAnchor, SafeCorridor, GeoFence, and PermissionLevel.
[0588] The anti-interference audit module performs gating based on link latency, link jitter, positioning residual, anchor point version, semantic target confidence, and executor status. If the audit passes, the executor performs an approach or delivery task; if the link deteriorates but remains controllable, the system downgrades to low-speed approach or requests confirmation; if the link is about to be interrupted, the system preemptively issues autonomous scripts, fallback paths, and security boundaries. After disconnection, the executor enters low-speed cruise, obstacle avoidance, fallback, or safe waiting modes based on the most recent valid task status.
[0589] The evidence storage module records the time anchor points of video frames seen by the operator, ROI areas, casualty target identification results, task description data, audit receipts, disconnection and autonomous process, and execution feedback. When an emergency stop, boundary violation, refusal to execute, disconnection and autonomous process, or collision risk occurs, the system generates an evidence package to support rescue review, liability determination, and safety supervision.
[0590] Example 3: Remote Hazardous Operations and Access Control in High-Interference Environments
[0591] In scenarios involving strong electromagnetic interference, hazardous materials handling, disaster response, or high-risk industrial operations, operators remotely control robots, robotic arms, tracked platforms, or special operation equipment through wearable sensing devices to detect, mark, grab, twist, block, or transport hazardous sources.
[0592] The system determines the spatial relationship between the operator's view, the target object, and the execution endpoint position through a multi-source physical true coordinate system and a spatial anchor point sharing mechanism. The semantic control layer identifies valves, hazardous chemical containers, leak points, tools, switches, or hazardous areas and converts them into task-level commands. The standardized task interface generates task description data, including target identifiers, risk levels, access levels, action constraints, operational constraints, spatial anchor points, and time anchor points.
[0593] For high-risk actions, such as closing the robotic arm gripper, turning a valve, approaching a hazard source, releasing a load, or canceling an emergency stop, the system employs a hierarchical access control and dual-condition triggering mechanism. Candidate commands must simultaneously meet conditions including gaze target confirmation, gesture confirmation, access control level, link quality, positioning / perception confidence, and anchor point consistency before being allowed to enter the execution process.
[0594] The anti-interference audit module determines whether candidate instructions should be executed, downgraded, rejected, or shut down safely. When link jitter exceeds limits, target semantic confidence is insufficient, anchor point versions are inconsistent, or the torque state at the execution end is abnormal, the system can reject high-risk actions or downgrade them to low-risk actions such as approaching but not touching, marking but not operating, or waiting for manual confirmation.
[0595] The evidence storage module records operator input, task description data, permission level, audit receipt, target semantic state, robotic arm torque, gripper status, execution feedback, and time anchors. When a refusal to execute, a safety shutdown, abnormal contact, or manual takeover occurs, the system archives the corresponding evidence package for subsequent auditing.
[0596] Example 4: Multi-entity collaborative operation in smart agriculture
[0597] In smart agriculture scenarios, operators use wearable sensing devices or cloud / edge computing devices to mark plot boundaries, obstacles, work areas, flight paths, crop rows, and collection points in a unified true coordinate system. The system can generate a set of spatial anchor points using RTK / GNSS, UWB base stations, visual anchor points, drone aerial images, Wi-Fi or Bluetooth derived observations, and synchronize them to multiple drones, unmanned vehicles, or agricultural robots.
[0598] Standardized task interfaces generate task description data for spraying, field inspection, transportation, spreading, sampling, or recovery. Task description data may include plot number, work area, waypoints, risk level, geofence, safety corridor, task priority, work sequence, AnchorID, and TemporalAnchor. The execution-side adaptation layer, based on different device capability profiles, converts the same TaskDescriptor into UAV waypoints, unmanned vehicle paths, robot work trajectories, or transportation tasks.
[0599] When multiple execution units operate simultaneously, the system constructs spatial topology constraints based on a shared anchor point coordinate system, calculating potential conflicts such as route intersections, path overlaps, insufficient minimum safe distances, conflicting work areas, or congestion at recovery points. The anti-interference audit module arbitrates based on task priority, remaining battery power, work stage, risk level, and minimum safe distance. Execution units with higher priority or higher risk levels maintain their original instructions, while the remaining execution units enter obstacle avoidance waiting, follow-up, alternative trajectory, or delayed execution modes.
[0600] When link quality degrades, the system can retain geofencing, local flight paths, last effective policies, and local obstacle avoidance capabilities, allowing the executor to continue executing low-risk subtasks or enter a safe waiting state. The evidence storage module records the parcel anchor point version, task description data, multi-entity conflicts, arbitration results, demotion actions, alternative trajectories, boundary crossing events, and execution feedback to facilitate post-mortem analysis of operational efficiency, safety incidents, and liability attribution.
[0601] Example 5: Industrial Inspection and Hazardous Materials Handling
[0602] In scenarios such as industrial inspection, energy facility inspection, hazardous materials storage, pipeline inspection, or equipment maintenance, operators use wearable sensing devices to view on-site equipment, instruments, valves, pipelines, hazardous chemical containers, leak points, or other target objects. The embodied execution end can be an inspection robot, robotic arm, mobile platform, or hybrid robot.
[0603] The system aligns the semantic targets of equipment in the operator's view with the actual spatial location of the execution end through multi-source physical truth fusion and spatial anchor point sharing. The semantic control layer identifies meters, valves, tools, hazardous material containers, warning signs, or leakage areas, and generates task-level commands such as "read meter", "approach valve", "turn valve", "grab tool", "mark anomaly", and "maintain safe distance".
[0604] The standardized task interface encapsulates task-level commands into task description data, including TargetID, TaskType, RiskLevel, AnchorID, TemporalAnchor, OperationConstraint, PermissionLevel, and AuditPolicy. The execution-side adaptation layer, based on the robot or robotic arm's capability profile, converts the task description data into path planning, end-effector pose, joint trajectories, force control parameters, gripper movements, or local obstacle avoidance constraints.
[0605] For high-risk actions such as grasping and closing, turning, pressing, opening valves, and approaching hazardous chemical containers, the system enables hierarchical access control and anti-interference auditing. If the positioning residual, anchor point version, torque feedback, gripper status, target confidence, or link quality does not meet safety requirements, the system can downgrade, request confirmation, refuse execution, or safely shut down.
[0606] During execution, the evidence storage module records the target semantics, task description, operator input, control commands, audit receipts, sensor feedback from the execution end, torque changes, gripper status, anchor point version, and time anchor point. When a false trigger, abnormal torque, refusal to execute, manual takeover, or safe shutdown occurs, the system generates an immutable evidence package for use in equipment maintenance, incident review, and compliance auditing.
[0607] Implementation Method 10: Explanation of Equivalent Substitution and Combination
[0608] This embodiment is used to illustrate the equivalent substitution relationships of various modules, steps, algorithms, data fields, communication links, terminal forms, and application scenarios in this invention. It should be understood that the foregoing embodiments one to nine are merely specific implementation examples given to facilitate understanding of the technical solution of this invention, and are not intended to limit the scope of protection of this invention. Those skilled in the art can replace, combine, tailor, or extend the specific implementation methods without departing from the core concept of this invention: "multi-source physical truth fusion, spatiotemporal alignment, task-level instruction mapping, anti-interference auditing, disconnection autonomy, and tamper-proof evidence storage".
[0609] 10.1 Equivalent Replacement of Multi-Source Fusion Algorithm
[0610] In the aforementioned embodiments, multi-source physical truth fusion can be achieved using extended Kalman filtering, factor graph optimization, sliding window optimization, particle filtering, rule fusion, learning-based fusion, or a combination thereof.
[0611] In other implementations, other state estimation, sensor fusion, map matching, graph optimization, visual-inertial fusion, wireless sensing fusion, or multimodal sensing fusion methods may also be used, as long as they can establish a shared spatial reference between the wearable sensing end and the embodied execution end based on at least two types of heterogeneous positioning / attitude observations, which is an equivalent implementation of the multi-source physical truth fusion described in this invention.
[0612] The heterogeneous positioning / attitude observations are not limited to RTK / GNSS, UWB, VPS, SLAM / VIO, IMU, odometry, visual sensors, LiDAR, 6G integrated sensing derived observations, Wi-Fi channel state information derived observations, and Bluetooth ranging or angle measurement derived observations. Any observation data that can provide absolute position, relative position, attitude, distance, angle, environmental features, spatial constraints, or time synchronization reference can be used as input for multi-source physical truth fusion.
[0613] 10.2 Equivalent replacement of spatial anchor points and temporal anchor points
[0614] In the aforementioned embodiments, spatial anchor points can be generated from environmental characteristics, artificial calibration points, UWB base station geometry, RTK reference points, 6G integrated sensing derived observations, Wi-Fi channel state derived observations, or Bluetooth ranging / angle measurement results.
[0615] In other embodiments, spatial anchors can also be generated from visual markers, natural feature points, laser point cloud features, building boundaries, ground markings, map control points, robot-generated map features, radio fingerprints, acoustic positioning points, or other repeatably identifiable physical spatial features. As long as they can maintain a shared spatial reference between the wearable sensing end, the cloud / edge computing end, and the embodied execution end, they constitute an equivalent implementation of the spatial anchors described in this invention.
[0616] The time anchor points in the aforementioned embodiments can be generated by GNSS / RTK timing, PTP, NTP, UWB bidirectional ranging calibration, local monotonic clocks, cloud / edge synchronization clocks, or combinations thereof. In other embodiments, hardware timestamps, frame numbers, event numbers, logical clocks, synchronization pulses, network synchronization messages, or other data identifiers that can be used to identify timing relationships can also be used. As long as it can place virtual content events, operator inputs, task description data, control commands, video frames, sensor observations, semantic summaries, and execution feedback under a comparable timing reference, it belongs to an equivalent implementation of the time anchor points described in this invention.
[0617] 10.3 Equivalent Replacement of Multimodal Instruction Mapping and Semantic Control Layer
[0618] In the aforementioned embodiments, the multimodal instruction mapping engine may include narrative / rhythm-driven mapping, human body capture / attention-driven mapping, and a semantic control layer.
[0619] In other implementations, virtual content event streams, human features, environmental semantic targets, and task inputs can be parsed and mapped by rule tables, state machines, semantic templates, behavior trees, task planners, machine learning models, multimodal recognition models, semantic understanding models, expert systems, or combinations thereof.
[0620] The human characteristics mentioned are not limited to gaze, gestures, posture, gait, bioelectrical signals or electromyographic signals, but may also include voice, touch, head movements, body orientation, controller input or other data that can express the operator's intention.
[0621] The environmental semantic targets are not limited to casualties, hazards, obstacles, passages, valves, meters, tools, work areas or plot boundaries, but may also include any physical object, virtual object, spatial area or status information that can be identified, labeled or tasked.
[0622] As long as it can convert virtual content event streams, human intentions, or environmental semantic targets into task-level instructions, dynamic instructions, motion constraints, operational constraints, or path constraints, it belongs to the equivalent implementation of the multimodal instruction mapping and semantic control layer described in this invention.
[0623] 10.4 Equivalent replacement of standardized task interfaces and execution-side adaptation layers
[0624] In the aforementioned embodiments, the standardized task interface may include fields such as TaskID, TargetID, TaskType, RiskLevel, AnchorID, TemporalAnchor, GeoFence, SafeCorridor, PermissionLevel, MotionConstraint, OperationConstraint, PathConstraint, AuditPolicy, and BlackBoxIndex.
[0625] In other implementations, the task description data can be encapsulated using different field names, data structures, encoding formats, or protocols, such as key-value pairs, JSON, XML, binary messages, protocol buffers, task template objects, control message bodies, or device driver commands. As long as it can express one or more of the following: task objective, task type, spatial reference, time reference, risk level, permission level, execution constraints, and auditing strategy, it constitutes an equivalent implementation of the standardized task interface described in this invention.
[0626] The execution end adaptation layer is not limited to adapting to unmanned vehicles, drones, robots, robotic arms, agricultural equipment, industrial inspection equipment, or special operation equipment. Other physical entities that possess mobility, manipulation, perception, execution, or feedback capabilities, and are able to receive task description data and convert it into task-level instructions, dynamic instructions, or local planning constraints that they can execute themselves, can also serve as equivalent implementations of the embodied execution end described in this invention.
[0627] 10.5 Equivalent replacements for interference-resistant auditing, gating, and disconnection-based autonomy
[0628] In the aforementioned implementation, the anti-interference audit module performs gating judgments based on link quality, sensor confidence, positioning residual, anchor point consistency, instruction risk level, execution end status, permission level, and task stage.
[0629] In other implementations, interference resistance auditing can be achieved using rule trees, state machines, risk matrices, expert systems, machine learning models, formal security rules, task security envelopes, or combinations thereof. As long as it can determine whether a control instruction should be granted, downgraded, rejected, or safely halted before or during execution, it constitutes an equivalent implementation of the interference resistance auditing and gating mechanism described in this invention.
[0630] Disconnection prediction can be based on link latency, jitter, packet loss rate, bit error rate, receipt timeout, bandwidth margin, location / perception confidence or a combination thereof, or it can be achieved by trend detection, state machine, link prediction model, network quality assessment model or manual policy configuration.
[0631] Disconnected autonomy can manifest as low-speed cruising, local obstacle avoidance, retreating to a safe point, hovering, waiting in place, ceasing high-risk actions, executing low-risk sub-tasks, safe shutdown, or other conservative strategies. As long as it can enable the executor to maintain task continuity, avoid obstacles, retreat, or wait safely within a safe boundary when the communication link is damaged or the positioning signal is insufficient, it belongs to the equivalent implementation of disconnected autonomy described in this invention.
[0632] 10.6 Equivalent replacement of ROI backhaul, ASW reprojection, and phase alignment
[0633] In the aforementioned embodiments, the system performs ROI hierarchical encoding based on the gaze region, semantic target, or risk region, and reduces operation drift through temporal anchors, ASW reprojection, phase verification, state forward prediction, and instruction compensation.
[0634] In other implementations, the ROI region can be determined by operator gaze, target detection results, semantic segmentation results, task target, risk region, execution end path, robotic arm end position, or manually specified region. ROI encoding can be represented by different resolutions, bitrates, frame rates, compression qualities, retransmission priorities, or transmission priorities.
[0635] ASW reprojection can be achieved using depth maps, optical flow, sparse features, head pose difference, local geometric models, viewpoint interpolation, predictive rendering, or other equivalent display compensation methods.
[0636] Phase alignment can be achieved based on time difference, spatial deviation, state prediction error, link delay estimation, current state of the executor, feedback frame seen by the operator, or a combination thereof. As long as it can determine the consistency between the feedback state on which the operator generates the command and the current state of the executor, and trigger compensation, degradation, request confirmation, refuse execution, or safe shutdown when there is a discrepancy, it is an equivalent implementation of the command-video phase alignment mechanism described in this invention.
[0637] 10.7 Equivalent Replacement of Spacetime Black Box and Tamper-proof Evidence
[0638] In the aforementioned embodiments, the spatiotemporal black box achieves tamper-proof evidence storage through black box indexing, audit receipts, chained hashing, digital signatures, trusted execution environments, and trusted timestamps.
[0639] In other implementations, immutability protection can also be achieved through encrypted storage, access control, trusted hardware, security chips, trusted timestamps, log signatures, integrity verification, or a combination thereof.
[0640] The data fields of the evidence package can be expanded or tailored according to different scenarios. As long as it can associate and record one or more of the following: virtual content state, environmental semantic target, task description data, control instructions, audit decisions, execution feedback, link / confidence summary, spatial anchor version number, and time anchor, and form a verifiable chain of evidence when the triggering event occurs, it is an equivalent implementation of the spatiotemporal black box and tamper-proof evidence storage mechanism described in this invention.
[0641] 10.8 Equivalent Replacement of End-Edge-Cloud Collaborative Computing
[0642] In the aforementioned embodiments, edge-cloud collaborative computing is used to dynamically allocate tasks such as positioning, identification, map optimization, ROI encoding, audit decision-making, disconnection autonomy, and evidence aggregation among wearable sensing terminals, cloud / edge computing terminals, and embodied execution terminals.
[0643] In other implementations, the cloud / edge computing endpoint can be deployed as a field edge server, an in-vehicle computing unit, a robot's local computing unit, a private network computing node, a cloud platform, or a distributed computing cluster. Task allocation between the wearable sensing endpoint, the cloud / edge computing endpoint, and the embodied execution endpoint can be achieved using rule tables, schedulers, containerized task migration, task queues, priority queues, resource managers, model services, or combinations thereof.
[0644] As long as it can dynamically allocate, downgrade, or migrate cross-space embodied collaborative control tasks based on computing power, link, energy consumption, temperature, task risk level, or location / perception confidence, it belongs to the equivalent implementation of the end-edge-cloud collaborative computing mechanism described in this invention.
[0645] 10.9 Equivalent replacement of terminal form and application scenarios
[0646] The wearable sensing terminal in the aforementioned embodiments can be AI glasses, AR / MR / VR headsets, helmet-mounted devices, handheld spatial sensing terminals, or other terminals with spatial sensing capabilities. In other embodiments, it can also be an in-vehicle display terminal, a console, a mobile terminal, a remote control, or other devices capable of acquiring operator input and displaying or processing feedback information.
[0647] The physical execution end in the aforementioned embodiments can be an unmanned vehicle, drone, quadruped robot, wheeled robot, tracked platform, robotic arm, gimbal, agricultural operation equipment, industrial inspection equipment, or special operation equipment. In other embodiments, it can also be a shipborne platform, underwater robot, warehousing robot, patrol robot, construction robot, medical assistance robot, or other physical equipment with physical execution capabilities.
[0648] This invention can be applied to cultural tourism interaction, remote rescue, hazardous operations, industrial inspection, smart agriculture, unmanned performances, multi-robot collaboration, warehousing and logistics, energy facility inspection, emergency response, and other cross-space embodied collaborative control scenarios. Implementation examples in different scenarios can be combined with each other, as long as they are still based on multi-source physical truth and spatiotemporal alignment, and achieve cross-space embodied collaborative control through task-level instruction mapping, security auditing, and traceable evidence storage, they fall within the scope of this invention.
[0649] 10.10 Combined Implementation Instructions
[0650] The embodiments described in this specification are not isolated from each other. The technical features in embodiments one through nine can be combined according to actual needs.
[0651] For example, a multi-source physical truth fusion mechanism can be combined with a low-latency ROI backhaul mechanism to ensure that key target areas and spatial anchor points are consistent in remote rescue; a standardized task interface can be combined with a multi-entity conflict arbitration mechanism for collaborative operations of multiple drones in smart agriculture; a disconnection autonomy mechanism can be combined with a spatiotemporal black box evidence storage mechanism to record the autonomous rollback process of the execution end in a strong interference environment; and a semantic control layer can be combined with a robotic arm permission hierarchical mechanism for auditing high-risk actions in the handling of industrial hazardous materials.
[0652] Therefore, those skilled in the art can select, combine, replace or adjust the aforementioned modules according to specific equipment, scenarios and task requirements. As long as the technical solution still revolves around establishing a shared control base based on multi-source physical truth and spatiotemporal alignment, and realizes task mapping, security auditing, disconnection autonomy and traceable evidence storage in cross-space embodied collaborative control, it should be considered to fall within the protection scope of this invention.
Claims
1. A multi-source physical truth and spatiotemporal aligned cross-space embodied collaborative control system, characterized in that: The control system includes a wearable sensing terminal, an embodied execution terminal, a unified spatiotemporal alignment engine, a multimodal instruction mapping engine, a security protection layer, and an evidence storage module. The wearable sensing terminal is used to acquire the operator's spatial pose and / or environmental perception data, and to acquire human features such as gaze, gestures, and posture, and is used to output virtual content event streams or control intentions. The embodied execution end is used to execute task-level instructions or dynamic instructions, and to return execution status and environmental feedback; The unified spatiotemporal alignment engine is used to fuse at least two types of heterogeneous positioning / attitude information to establish a multi-source physical truth coordinate system, and to synchronize spatial anchor points between the wearable sensing end and the embodied execution end to maintain shared coordinate mapping. The multimodal instruction mapping engine is used to map the virtual content event stream and / or human features to the task-level instructions or dynamic instructions of the embodied execution end, and to attach time anchors to the instructions; The security protection layer includes an anti-interference audit module, which is used to gate the execution of the instruction based on link quality, location / perception confidence and / or instruction risk level; The evidence storage module is used to associate and record virtual content status, instructions, audit decisions and execution feedback, and generate an immutable evidence package when a trigger event occurs.
2. The multi-source physical truth and spatiotemporal alignment cross-space embodied collaborative control system according to claim 1, characterized in that: The multi-source physical true coordinate system is established by fusing at least two types of observations from RTK / GNSS, UWB, VPS, SLAM / VIO, and IMU. The multi-source physical true coordinate system can further integrate at least one of the following: sensing measurements from the 6G integrated sensing network side, channel state information derived observations from the Wi-Fi link, Bluetooth ranging or angle measurement derived observations, as a redundant constraint on the RTK / GNSS, UWB, VPS, SLAM / VIO or IMU observations. The unified spatiotemporal alignment engine assesses the confidence level, dynamically weights, or removes anomalies for various observations based on at least one of observation residuals, anchor point consistency, reprojection error, link delay, or jitter, in order to enhance the continuity and robustness of spatial alignment under conditions of occlusion, visual blind spots, or strong interference.
3. The multi-source physical truth and spatiotemporal alignment cross-space embodied cooperative control system according to claim 2, characterized in that: The multimodal instruction mapping engine includes at least one or more of the following: narrative / rhythm-driven mapping, human body capture / attention-driven mapping, and semantic control layer; The narrative / rhythm-driven mapping is used to extract temporal features such as beats, audio envelopes, key events or virtual trigger points from virtual content and map them as execution-end dynamic parameters, path segment constraints or formation constraints. The human body capture / gaze-driven mapping is used to map the operator's gaze, gestures, posture, gait, bioelectrical signals or combinations thereof into task-level actions, fine operation instructions or action trajectories of the execution end. The semantic control layer is used to convert the identification results of environmental semantic targets, task roles, risk levels, or interactive objects into task-level commands, and further map them into motion constraints, operation constraints, or path constraints at the execution end.
4. The multi-source physical truth and spatiotemporal alignment cross-space embodied cooperative control system according to claim 1, characterized in that: The control system also includes a standardized task interface and / or an execution-end adaptation layer; The standardized task interface is used to generate or receive task description data, which includes at least one or more of the following: target identifier, task type, risk level, spatial anchor point, temporal anchor point, geofence, security corridor, permission level, action constraint, or path constraint. The execution end adaptation layer is used to convert the task description data into task-level instructions, dynamic instructions or local planning constraints that can be executed by the corresponding execution end based on the capability profile of different execution ends, so that the same task description can be adapted to one or more execution ends among unmanned vehicles, drones, robots, robotic arms or special operation equipment.
5. The multi-source physical truth and spatiotemporal alignment cross-space embodied collaborative control system according to claim 1, characterized in that: The anti-interference audit module is used to perform real-time evaluation of at least one of the following: link quality, sensor confidence, positioning residual, anchor point consistency, instruction risk level, or execution end status, and to gate the execution of control instructions. The gating result includes at least one of pass, downgrade, rejection, or safe shutdown; The degradation includes at least one of the following: speed limit, geofence tightening, security corridor constraint, control domain switching, task pause, request for confirmation, or switching to a local autonomous policy.
6. The multi-source physical truth and spatiotemporal alignment cross-space embodied collaborative control system according to claim 5, characterized in that: The anti-interference audit module is also used to predict connection failures. When the link quality shows a continuous downward trend, the latency or jitter exceeds a preset threshold, the packet loss rate or bit error rate exceeds a preset threshold, or the positioning / perception confidence is lower than a preset threshold, the control system sends an autonomous script, fallback path, security boundary, task continuation status or last effective strategy to the embodied execution terminal before the link is completely interrupted. When the communication link is damaged or the location information is insufficient, the embodied execution terminal enters the disconnected autonomous mode based on the autonomous script, fallback path, security boundary, task continuation status or last effective strategy, in order to maintain task continuity, obstacle avoidance, fallback or safe waiting under the security boundary.
7. The multi-source physical truth and spatiotemporal alignment cross-space embodied cooperative control system according to claim 1, characterized in that: The control system also includes a spatial anchor point sharing mechanism, which includes anchor point generation, anchor point synchronization and version control, cross-end coordinate mapping, and multi-entity consistency maintenance. The anchor point generation is used to generate a set of spatial anchor points at the wearable sensing end and / or the embodied execution end based on environmental features, artificial calibration points, UWB base station geometry, RTK reference points, 6G integrated sensing derived observations or Wi-Fi channel state derived observations. Each spatial anchor point must include at least the anchor point identifier AnchorID, anchor point pose, anchor point confidence, anchor point version number, and validity period; When any end detects anchor point drift, anchor point matching residual exceeding a preset threshold, anchor point confidence level below a preset threshold, or anchor point version inconsistency, the control system triggers anchor point update, realignment request, or control degradation.
8. The multi-source physical truth and spatiotemporal alignment cross-space embodied cooperative control system according to claim 7, characterized in that: When there are multiple embodied executors and / or multiple operators, the control system maintains the geometric consistency shared by multiple entities based on a unified set of anchor points; When multiple embodied execution ends receive potentially conflicting control commands, the control system constructs spatial topological constraints based on a shared anchor point coordinate system and arbitrates according to preset priority, risk level, task stage, or minimum safe distance. Among them, the execution terminals with higher priority or higher risk level retain the original instructions, while the remaining execution terminals are automatically downgraded to obstacle avoidance waiting mode, follow-up mode, alternative trajectory execution mode or delayed execution mode, and the arbitration result is written into the audit receipt and evidence index.
9. A cooperative control method using the control system as described in any one of claims 1 to 8, characterized in that, include: A multi-source physical true coordinate system is established by fusing at least two types of positioning / attitude observations, and spatial anchor points are synchronized to form a shared coordinate mapping; Analyze virtual content event streams, environmental semantic targets, and / or operator human characteristics to generate control intent or task description data; Based on the control intent or task description data, generate task-level instructions, dynamic instructions, or local planning constraints, and attach time anchors to the instructions; When the sensing end receives a delayed video frame carrying a temporal anchor, it calculates the difference between the current 6-DoF pose of the sensing end and the pose corresponding to the temporal anchor, performs asynchronous spatial warp (ASW) reprojection, and compensates for the viewpoint of the video frame. The command is gated based on link quality, location / perception confidence, command risk level, or execution end status, and the command is executed by at least one of the following: pass, downgrade, reject, or stop. The virtual content state, environmental semantic targets, task description data, instructions, gating decisions and execution feedback are recorded to form a traceable chain of evidence, and the evidence package is sealed when a triggering event occurs.
10. The cooperative control method according to claim 9, characterized in that: The execution feedback includes at least one of video frames, sensor observations, state telemetry, or semantic summarization. The video frames and / or semantic summaries are ROI hierarchical encoding or semantically prioritized backhaul based on gaze regions, semantic targets or risk regions, so that key regions, key targets or key semantic information are backhauled with higher quality or priority than background regions, and time anchors are embedded in the corresponding video frames and / or semantic summaries. The control command is bound to the time anchor point of the video frame seen when the operator generates the command. The execution end performs phase verification, forward state prediction or command compensation based on the time anchor point, the estimated link delay value and the local motion state. When phase deviation, prediction uncertainty, link jitter, or reprojection uncertainty exceeds a preset threshold, it triggers downgrade, request confirmation, refuse to execute, or safe shutdown. The evidence package includes virtual rendering summaries or narrative events, environmental semantic targets, task description data, instructions, audit receipts, execution feedback, link / confidence summaries, anchor version numbers and time anchors, and uses chain hashing, digital signatures and / or trusted execution environments to ensure that the evidence is tamper-proof.