Physical law constraint-based semantic mask enhancement system and spatial calculation reasoning method

By establishing a semantic masking enhancement system based on physical law constraints, the problem of edge instability in visual semantic segmentation under complex environments is solved, the stability and realism of the interaction between virtual and real objects are realized, and the system's anti-drift and anti-forgery capabilities are enhanced.

CN122023739APending Publication Date: 2026-05-12WUHAN HUACHUANG HIGHLIGHT DIGITAL TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610134893.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-30
Publication Date
2026-05-12

Smart Images

  • Figure CN122023739A_ABST
    Figure CN122023739A_ABST
Patent Text Reader

Abstract

The invention provides a physical law constraint-based semantic mask enhancement system and a spatial calculation reasoning method. The system part comprises the following modules: an initial semantic mask generation module, a physical base construction module, a semantic mask enhancement module, a time sequence consistency module, a dynamic confidence evaluation and mode switching module, a space calculation reasoning module, a physical consistency auditing module, a rendering and safety access module and a space reference alignment module. According to the invention, a physical-visual double closed-loop space computing system is constructed, the system can be deployed in a head-mounted terminal, a mobile terminal, a vehicle-mounted terminal or an end-cloud cooperative system, and an auditing and executable space truth layer is established between a physical entity and AI vision, so that boundary stability and interaction authenticity are kept in a complex environment, and the system can be widely applied to the field of computing. The method has consistent applicability to different physical field sensing modes, the mask position is still kept stable during high-speed movement and vision blurring, and cross-frame drift and flicker are remarkably reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of spatial computing, artificial intelligence, and augmented / mixed reality (AR / MR), specifically to a semantic mask enhancement system based on physical laws and a spatial computing reasoning method. Background Technology

[0002] With the development of spatial computing and augmented reality / mixed reality (AR / MR) applications, the system usually needs to complete the following in real scenes: (1) semantic understanding and region division of real objects for occlusion relationships, interaction area limitation and content fitting; (2) reasoning and prediction of the spatial pose, motion trajectory and interactive behavior of virtual objects for realizing realistic interaction such as edge touching, collision, wall touching, grabbing.

[0003] To address the aforementioned needs, existing technologies primarily employ the following approaches: pure visual semantic segmentation / instance segmentation. However, under conditions of low light, shadows, reflections, low texture, occlusion, or image blurring caused by rapid motion, the semantic mask edges output by pure visual semantic segmentation are prone to jitter, adhesion, missing, or drift, making it difficult to accurately align with the physical boundaries of real objects. This leads to problems such as incorrect occlusion relationships, drifting of edge-fitting content, and misjudgment of interactive areas.

[0004] Existing technologies often exhibit issues such as clipping, boundary crossing, hovering, and unrealistic contact when virtual objects interact with real objects. Particularly in scenarios requiring strict boundary constraints, such as "leaning against walls, touching edges, grabbing, and collision bounce," AI inference outputs are prone to producing pose changes that defy physical laws, impacting immersion and usability. Furthermore, even when physical boundaries are introduced in existing technologies, mask corrections are only performed at the single-frame level. Summary of the Invention

[0005] To address the shortcomings of existing technologies, the present invention aims to provide a semantic mask enhancement system and spatial computation reasoning method based on physical laws to solve the problems mentioned in the background. The present invention establishes an auditable and executable spatial truth layer between physical entities and AI vision, thereby maintaining boundary stability and interaction authenticity in complex environments.

[0006] To achieve the above objectives, the present invention is implemented through the following technical solution: a semantic mask enhancement system based on physical law constraints, comprising the following modules:

[0007] Initial semantic mask generation module: Used to perform semantic segmentation / instance segmentation on the current frame image, outputting pixel-level or instance-level semantic masks. ;

[0008] Physical Base Building Module: Used to build physical base data from at least one non-visual physical field data source. and output a set of physical credibility indicators. ;

[0009] Semantic masking enhancement module: used to... Map to physical coordinate system and utilize Perform topology cropping / restoration and edge reinforcement to obtain a single-frame enhanced mask. ;

[0010] Timing consistency module: used for physical motion vector-based... Perform cross-frame prediction and update of the mask to achieve temporal coherence and output a temporally consistent enhanced mask. ;

[0011] Dynamic confidence assessment and mode switching module: used to calculate the mixed confidence factor. It switches between physical-dominated mode and visual + physical pruning mode based on a threshold τ;

[0012] Spatial computation inference module: used to generate object pose / trajectory / contact state predictions based on enhanced masks, physical bases, and interactive commands. ;

[0013] Physical consistency audit module: used for... Perform physical conflict detection, and if necessary, trigger forced correction / projection back to the feasible domain / rollback strategy, and output the audited rendering and interaction instructions;

[0014] Rendering and security access control module: used for... Implement occlusion, collision, and rendering instruction verification, and perform physical consistency admission control on third-party rendering / interaction instructions.

[0015] Spatial reference alignment module: used to extract the geocentric gravity vector output by the inertial measurement unit (IMU) in real time and define it as the absolute vertical reference axis for spatial calculation; when the virtual vertical axis obtained by visual positioning / mapping or pose estimation has an angular residual with the geocentric gravity vector that exceeds a threshold, the system performs dynamic alignment on the visual coordinate system or rendering coordinate system through a rotation compensation operator to forcibly eliminate the angular drift of the virtual horizontal plane relative to the real ground plane and keep the enhanced semantic mask consistent with the physical topology.

[0016] Furthermore, the system performs unified time base management for visual data and physical field data: it assigns a timestamp to each frame of image and the corresponding physical observation, and estimates and corrects the clock drift of multiple sensors to ensure that they meet the alignment requirements of the same moment or equivalent moment.

[0017] Furthermore, the system performs extrinsic parameter calibration on the visual sensor coordinate system and the physical sensor coordinate system to obtain the coordinate transformation relationship, and uses this transformation relationship to achieve spatial alignment under a unified physical coordinate system when performing mask projection / mapping and boundary constraints.

[0018] When synchronization quality or external parameter reliability falls below a threshold, the system triggers a degradation strategy: reducing... Or enter the physical-dominated / static alignment mode.

[0019] Furthermore, it also includes a physical base. Methods for obtaining and representing:

[0020] The physical base data The acquisition methods include: real-time sensor acquisition, spatiotemporal snapshot backtracking based on historical sensing data, and pre-stored digital twin priors retrieved from the cloud or local storage.

[0021] Furthermore, the aforementioned The sources include:

[0022] Electromagnetic / radio physical field sensing: UWB impulse response (CIR), Wi-Fi / 6G multipath fingerprinting, millimeter wave / microwave radar echo, RSSI / phase / angle of arrival / Doppler;

[0023] Acoustic echo sensing: distance / contour information of ultrasonic / sonar echo formation;

[0024] Inertial measurement envelope: the velocity, displacement, and motion envelope formed by the IMU / odometer;

[0025] Pre-stored digital twin priors: hard boundaries provided by CAD / BIM / point cloud / 3D map / structured model;

[0026] Depth / point cloud sensing: LiDAR / ToF / structured light / binocular depth sensing, etc.

[0027] Furthermore, the aforementioned Represented in any of the following forms or combinations thereof:

[0028] Distance field / Signed distance field ;

[0029] Occupied raster / voxel representation ;

[0030] Hard boundary set ;

[0031] The set of boundary surfaces or feasible regions obtained by inversion of radio / acoustic reflection characteristics .

[0032] Furthermore, the system is also used to output a set of physical credibility indicators. This includes: signal-to-noise ratio (SNR), multipath exponent / echo stability, observation consistency, and matching residuals with the prior twin model;

[0033] The Used to calculate the physical weight operator γt and the mixed confidence factor ηt.

[0034] A spatial computational reasoning method used in the aforementioned semantic masking enhancement system includes a reasoning generation process:

[0035] The spatial computing inference module uses user interaction commands and enhanced masks. Physical base Generate interactive prediction output .

[0036] Furthermore, the aforementioned This includes: virtual object pose, motion trajectory or velocity, contact point / contact normal / collision depth, and interaction action parameters.

[0037] Furthermore, it also includes physical conflict detection: audit filters for... Perform physical conflict detection, including at least: boundary conflict detection, gravity consistency detection, infeasible contact detection, and velocity mutation detection, and the audit filter is independent of the AI ​​model structure.

[0038] The beneficial effects of this invention are:

[0039] 1. This invention uses whether it has spatial scale dimensions and whether it is not derived from pure visual pixels as the judgment criteria. It has consistent applicability to different physical field perception methods. It maintains the stability of the mask position even when there is high-speed motion and visual blur, significantly reducing cross-frame drift and flicker. The stability of the mask comes from physical displacement constraints rather than post-processing beautification.

[0040] 2. This invention addresses the problem of boundary illusion caused by unstable semantic mask boundaries. It significantly reduces mask jitter, sticking and drift, ensuring that occluded and edge-fitted content remains consistent with the real boundary, and maintaining usable boundary stability even under weak texture / shadow / reflection conditions.

[0041] 3. This invention can still output usable semantics and boundaries when visual degradation occurs, ensuring engineering usability and stability, reducing "failure when the scene changes", and possessing the industrial practicality emphasized by PCT review; the enhanced mask is upgraded from a pixel area to a digital entity with physical attributes, which can support highly realistic interaction, while using physical fingerprint anchoring to improve anti-drift and anti-forgery capabilities.

[0042] 4. This invention uses the geocentric gravity vector obtained by the inertial measurement unit as the absolute vertical reference for spatial calculation, and performs dynamic rotation compensation and correction on the visual coordinate system / rendering coordinate system to ensure the physical consistency and interactive realism of virtual-real fusion. Attached Figure Description

[0043] Figure 1 This is a block diagram of the overall architecture of the semantic mask enhancement system based on physical law constraints of the present invention;

[0044] Figure 2 The physical base of this invention A diagram illustrating the sources of generalization;

[0045] Figure 3 This is a schematic diagram of single-frame semantic mask enhancement and boundary topology clipping / repair in this invention;

[0046] Figure 4 This is a schematic diagram illustrating the temporal coherence update of the present invention.

[0047] Figure 5 For the dynamic confidence switching and degradation protection of this invention ( (Mode switching) diagram;

[0048] Figure 6 This is a flowchart illustrating the physical consistency audit filter and conflict detection / forced correction process of the present invention.

[0049] Figure 7 This is a rendering and security access diagram of the present invention;

[0050] Figure 8 This is a schematic diagram of a scenario according to an embodiment of the present invention. Detailed Implementation

[0051] To make the technical means, creative features, objectives and effects of this invention easier to understand, the invention will be further described below in conjunction with specific embodiments.

[0052] Please see Figures 1 to 8 This invention provides the following technical solution: a semantic mask enhancement system based on physical law constraints. This system can be deployed on head-mounted terminals, mobile terminals, vehicle terminals, or end-to-cloud collaborative systems, and the system consists of the following modules:

[0053] Initial semantic mask generation module: Used to perform semantic segmentation / instance segmentation on the current frame image, outputting pixel-level or instance-level semantic masks. ;

[0054] Physical Base Building Module: Used to build physical base data from at least one non-visual physical field data source. and output a set of physical credibility indicators. ;

[0055] Semantic masking enhancement module: used to... Map to physical coordinate system and utilize Perform topology cropping / restoration and edge reinforcement to obtain a single-frame enhanced mask. ;

[0056] Timing consistency module: used for physical motion vector-based... Perform cross-frame prediction and update of the mask to achieve temporal coherence and output a temporally consistent enhanced mask. (or written as) );

[0057] Dynamic confidence assessment and mode switching module: used to calculate the mixed confidence factor. It switches between physical-dominated mode and visual + physical pruning mode based on a threshold τ;

[0058] Spatial computation inference module: used to generate object pose / trajectory / contact state predictions based on enhanced masks, physical bases, and interactive commands. ;

[0059] Physical consistency audit module: used for... Perform physical conflict detection, and if necessary, trigger forced correction / projection back to the feasible domain / rollback strategy, and output the audited rendering and interaction instructions;

[0060] Rendering and security access control module: used for... Implement occlusion, collision, and rendering instruction verification, and perform physical consistency admission control on third-party rendering / interaction instructions.

[0061] Spatial reference alignment module: used to extract the geocentric gravity vector output by the inertial measurement unit (IMU) in real time and define it as the absolute vertical reference axis for spatial calculation; when the virtual vertical axis obtained by visual positioning / mapping or pose estimation has an angular residual with the geocentric gravity vector that exceeds a threshold, the system performs dynamic alignment on the visual coordinate system or rendering coordinate system through a rotation compensation operator to forcibly eliminate the angular drift of the virtual horizontal plane relative to the real ground plane and keep the enhanced semantic mask consistent with the physical topology.

[0062] In this embodiment, the relationship between the data flow and the dual closed loop is as follows:

[0063] Boundary closure (perception closure): →(Semantic Masking Enhancement Module / Temporal Consistency Module)→ ,Depend on and Physical truth calibration and timing stabilization of the mask boundaries;

[0064] Logical closed loop (reasoning closed loop): → (Physical Consistency Audit Module) → Audit / Correction → Rendering instructions to ensure that the final output meets physical consistency requirements such as gravity, collision, and inaccessible boundaries.

[0065] This embodiment also provides timing synchronization and coordinate system calibration processes:

[0066] The system implements unified time base management for visual and physical field data: it assigns a timestamp to each frame of image and its corresponding physical observation, and estimates and corrects multi-sensor clock drift to ensure alignment at the same or equivalent time. The system performs extrinsic parameter calibration on the visual and physical sensor coordinate systems to obtain coordinate transformation relationships (e.g., Tcam→phys), and uses these transformation relationships for mask projection / mapping and boundary constraints to achieve spatial alignment under a unified physical coordinate system. When synchronization quality or extrinsic parameter reliability falls below a threshold, the system triggers a degradation strategy: reducing... Alternatively, it can enter a physically dominant / static alignment mode to avoid boundary jitter or mis-clipping caused by synchronization errors.

[0067] This embodiment also provides a physical base. Acquisition and representation

[0068] In this embodiment, physical base data It refers to spatial constraint information derived from non-pure visual pixels and having spatial scale dimensions, used to provide physical boundaries, geometric topology, or feasible domain constraints of the real environment.

[0069] The physical base data The acquisition methods include: real-time sensor acquisition, spatiotemporal snapshot backtracking based on historical sensing data, and pre-stored digital twin priors retrieved from the cloud or local storage.

[0070] The The sources include:

[0071] Electromagnetic / radio physical field sensing: UWB impulse response (CIR), Wi-Fi / 6G multipath fingerprinting, millimeter wave / microwave radar echo, RSSI / phase / angle of arrival / Doppler;

[0072] Acoustic echo sensing: distance / contour information of ultrasonic / sonar echo formation;

[0073] Inertial measurement envelope: the velocity, displacement, and motion envelope formed by the IMU / odometer;

[0074] Pre-stored digital twin priors: hard boundaries provided by CAD / BIM / point cloud / 3D map / structured model;

[0075] Depth / point cloud sensing: LiDAR / ToF / structured light / binocular depth sensing, etc.

[0076] The Represented in any of the following forms or combinations thereof:

[0077] Distance field / Signed distance field ;

[0078] Occupied raster / voxel representation ;

[0079] Hard boundary set (polygon / surface mesh) ;

[0080] The set of boundary surfaces or feasible regions obtained by inversion of radio / acoustic reflection characteristics .

[0081] This embodiment also provides a physical reliability index. Explanation: Physical reliability index Used to output a set of physical credibility indicators This includes: signal-to-noise ratio (SNR), multipath exponent / echo stability, observation consistency, and matching residuals with the prior twin model;

[0082] The Used to calculate the physical weight operator γt and the mixed confidence factor ηt.

[0083] This embodiment also provides a method flow for mask enhancement and boundary reinforcement, the specific steps of which are as follows:

[0084] S101: Acquire the current frame image data (RGB or RGB-D) and necessary sensor synchronization data.

[0085] S102: The initial semantic mask generation module outputs a pixel mask for the current frame. .

[0086] S103: Output from the physical base construction module and its credibility index .

[0087] S104: By semantic masking enhancement module Mapped to the physical coordinate system, and based on Perform topology trimming / repair and edge reinforcement to obtain an enhanced mask. .

[0088] S104 includes at least one or more of the following operations:

[0089] Projection / mapping: Mapping a pixel mask to a physical coordinate system (such as world coordinates / device coordinates);

[0090] Topological clipping: Clipping the mask in The overflow area outside the boundary;

[0091] Topology repair / completion: based on The feasible domain / occupancy constraint is used to perform connectivity repair, hole filling, or equivalent completion processing on the missing regions of the mask within the physical boundary.

[0092] Edge reinforcement: Employ boundary preservation constraints (e.g., TV(M) or equivalent regularization terms) to enhance edge sharpness and suppress edge jitter;

[0093] Spatial ID Binding: Assigns an object-level spatial ID (Object-ID) to objects in the enhancement mask for subsequent collision and safety access.

[0094] S104-1: Based on Calculate the physical weight operator γt to dynamically adjust the strength of physical boundary constraints.

[0095] When physical reliability is high (e.g., high SNR, high stability, low multipath), increase γt to enhance physical clipping / completion strength;

[0096] When the physical reliability is low, reduce γt to avoid excessive constraints on the mask caused by physical noise.

[0097] Final output: Single-frame enhanced mask .

[0098] S105: Physical motion vectors are acquired from IMU / odometer / radio Doppler, etc. Used to estimate the device's displacement / velocity between adjacent frames.

[0099] S106: Enhance the mask of the previous frame in the physical coordinate system Perform displacement prediction to obtain the prediction mask. .

[0100] S107: Predict the mask With the current frame pixel mask as well as The merged update yields an enhanced mask with temporal coherence. And maintain the continuity of object-level spatial ID binding.

[0101] Among them, the To enhance the mask of the previous frame according to The prediction results obtained by performing motion compensation / deformation propagation in the physical coordinate system; This is the pixel mask output by the initial semantic mask generation module for the current frame image.

[0102] This embodiment also provides a timing update formula:

[0103]

[0104] in and It can be implemented in any way (filtering, optimization, fusion update, etc.), the key is to explicitly introduce... It serves as a time coherence constraint and updates the mask within the physical coordinate system.

[0105] S108: Calculate the mixed confidence factor ηt, whose inputs include at least: visual quality indicators (brightness, contrast, texture richness, motion blur, occlusion rate, etc.).

[0106] Physical credibility index (SNR, multipath index, stability, etc.)

[0107] S109: Compare ηt with the threshold τ to select the mode:

[0108] If ηt > τ, the body enters the dominant physical mode: Directly generate / reconstruct semantic regions or boundaries, with visual masks serving as weak constraints or used only for semantic label mapping;

[0109] If ηt < τ, enter the "visual recognition + physical boundary trimming" mode: execute according to mask enhancement, boundary reinforcement and temporal consistency enhancement.

[0110] This degradation mechanism ensures that usable semantic boundaries and interaction constraints can still be output in visually ineffective scenarios such as darkness, smoke, white walls, and strong glare.

[0111] This embodiment also provides a process for spatial computational inference and physical consistency auditing:

[0112] S110: Spatial computing inference module based on user interaction commands and enhanced mask. Physical base Generate interactive prediction output The aforementioned This includes: virtual object pose, motion trajectory or velocity, contact point / contact normal / collision depth, and interaction action parameters (wall-hugging, grabbing, placing, etc.).

[0113] S111: Audit Filters Performing physical collision detection includes at least:

[0114] Boundary conflict detection: Determines whether an object crosses a boundary. or Defined hard boundaries;

[0115] Gravity consistency detection: Determines whether the direction of motion / acceleration violates the gravity vector or stability threshold;

[0116] Infeasible contact detection: Determines whether the collision depth exceeds a threshold or whether the contact normal is consistent with the feasible region;

[0117] Velocity jump detection: Determines whether there are abnormal jumps in velocity / displacement (used to suppress drift).

[0118] Furthermore, the audit filter is independent of the AI ​​model structure. Regardless of whether it is end-to-end, rule-based, or any type of network, it must pass through this audit before the output rendering instructions are executed.

[0119] This embodiment also provides forced correction and rollback:

[0120] S112: If a conflict is detected, at least one forced correction strategy is triggered:

[0121] Newtonian dynamics / rigid body collision correction: calculate reaction forces and reconstruct pose;

[0122] Projecting back to the feasible region: Projected onto by Within the defined feasible domain;

[0123] Rollback strategy: Switch to physical-dominated mode, static alignment mode, or restrict interaction freedom to ensure safety and stability.

[0124] S113: Output audited and corrected rendering and interaction commands, ensuring at least the following:

[0125] It does not cross the boundary (does not exceed the limit); it does not clip through the real boundary (does not penetrate the real boundary); it does not drift (does not exhibit non-physical jumps).

[0126] This embodiment also provides rendering and security access control:

[0127] Occlusion and Collision Rendering: The rendering module is based on enhanced masks This implements the occlusion relationship between virtual and real objects, as well as feedback effects such as collision particles, friction, and bounce.

[0128] Security access control and command verification: Rendering / interaction commands from third-party applications are executed before entering the rendering pipeline.

[0129] Physical consistency check: If the audit filter fails or the physical fingerprint check is missing, execution of instructions in the specified masked area is prohibited;

[0130] Physical ID binding verification: Enhanced mask objects carry a spatial ID and at least one set of physical attribute parameters (mass, friction, elasticity, etc.), and are anchored to the environmental physical fingerprint for anti-spoofing and access control. The set of physical attribute parameters is generated after the object-level spatial ID is assigned and is written to the rendering pipeline whitelist after passing the audit filter; instructions that do not carry this set or fail verification are rejected.

[0131] The present invention also provides the following specific embodiments to illustrate the above technical solutions:

[0132] This embodiment uses the Zhao Wangcheng Ruins as an example: In the low-light environment of the Zhao Wangcheng Ruins, the rammed earth walls and the ground are similar in color, which can easily lead to "semantic adhesion" in visual segmentation, resulting in unclear wall boundaries. When the virtual "Zhao warrior" performs a wall-leaning action, without physical constraints and audit corrections, clipping can easily occur. In addition, the Zhao Wangcheng Ruins have complex terrain features such as irregular rammed earth boundaries, undulating slopes, and gravel obstruction, causing the semantic boundaries that rely solely on vision to drift at different viewpoints / heights. At the same time, as a high-value digital twin asset used for guidance and interaction, the virtual warrior's actions such as leaning against walls, touching, and patrolling have higher requirements for the realism of collisions / occlusions. Physical bases and audit filters must be used to ensure the safety and stability of the interaction.

[0133] The system constructs a hard boundary for the wall using radio multipath / echo (e.g., UWB CIR), forming And output credibility index .

[0134] In one alternative implementation, radio multipath / echo can be used to calculate the distance range and height profile of the wall hard boundary (e.g., interface features with a distance of approximately 1.25m and a height of approximately 0.8m), and this scale dimension can be used for boundary clipping and collision feasible region construction; under good synchronization and reliability conditions, the fitting error between the enhanced boundary and the real profile can reach the centimeter level (e.g., approximately 1cm).

[0135] Masking enhancement and timing stability:

[0136] The visual model outputs the raw mask Mraw (or The semantic mask enhancement module calls RefineMaskWithPhysics to perform clipping and completion within the physical coordinate system. and through Perform cross-frame prediction and updates to maintain stable wall boundary alignment.

[0137] Reasoning, auditing, and correction:

[0138] The inference module outputs the predicted pose of the samurai leaning against the wall. The audit filter detected the samurai's shoulder crossing. The defined hard boundary triggers rigid body collision correction, pushes the pose back and outputs the collision feedback effect, ultimately achieving stable wall-mounting without clipping.

[0139] The foregoing has shown and described the basic principles and main features of the present invention and its advantages. It will be apparent to those skilled in the art that the present invention is not limited to the details of the above exemplary embodiments, and that the present invention can be implemented in other specific forms without departing from the spirit or basic features of the present invention.

[0140] Furthermore, it should be understood that although this specification describes embodiments, not every embodiment contains only one independent technical solution. This narrative style is merely for clarity. Those skilled in the art should consider the specification as a whole, and the technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.

Claims

1. A semantic masking enhancement system based on physical law constraints, characterized in that, Includes the following modules: Initial semantic mask generation module: Used to perform semantic segmentation / instance segmentation on the current frame image, outputting pixel-level or instance-level semantic masks. ; Physical Base Building Module: Used to build physical base data from at least one non-visual physical field data source. It outputs a set of physical credibility indicators. ; Semantic mask enhancement module: used to enhance the initial semantic mask. Mapping to the physical coordinate system and utilizing physical base data Perform topology cropping / restoration and edge reinforcement to obtain a single-frame enhanced mask. ; Timing consistency module: used for physical motion vector-based... Perform cross-frame prediction and update of the mask to achieve temporal coherence and output a temporally consistent enhanced mask. ; Dynamic confidence assessment and mode switching module: used to calculate the mixed confidence factor. It switches between physical-dominated mode and visual + physical pruning mode based on a threshold τ; Spatial computation inference module: used to generate object pose / trajectory / contact state predictions based on enhanced masks, physical bases, and interactive commands. ; Physical consistency audit module: Used to perform physical conflict detection on the predicted pose / trajectory / contact state or rendering / interaction instructions output by the spatial calculation inference module, and trigger forced correction / projection back to feasible domain / rollback strategy when necessary, and output the rendering and interaction instructions that pass the audit; Rendering and security admission module: used for enhanced mask-based rendering Implement occlusion, collision, and rendering instruction verification, and perform physical consistency admission control on third-party rendering / interaction instructions; Spatial reference alignment module: used to extract the geocentric gravity vector output by the inertial measurement unit (IMU) in real time and define it as the absolute vertical reference axis for spatial calculation; when the virtual vertical axis obtained by visual positioning / mapping or pose estimation has an angular residual with the geocentric gravity vector that exceeds a threshold, the system performs dynamic alignment on the visual coordinate system or rendering coordinate system through a rotation compensation operator to forcibly eliminate the angular drift of the virtual horizontal plane relative to the real ground plane and keep the enhanced semantic mask consistent with the physical topology.

2. The semantic masking enhancement system based on physical law constraints according to claim 1, characterized in that: The system performs unified time base management for visual data and physical field data: it assigns a timestamp to each frame of image and the corresponding physical observation, and estimates and corrects the clock drift of multiple sensors to meet the alignment requirements of the same moment or equivalent moment.

3. The semantic mask enhancement system based on physical law constraints according to claim 2, characterized in that: The system performs extrinsic parameter calibration on the visual sensor coordinate system and the physical sensor coordinate system to obtain the coordinate transformation relationship, and uses the transformation relationship to achieve spatial alignment under a unified physical coordinate system when performing mask projection / mapping and boundary constraints. When synchronization quality or external parameter reliability falls below a threshold, the system triggers a degradation strategy: reducing... Or enter the physical-dominated / static alignment mode.

4. The semantic masking enhancement system based on physical law constraints according to claim 1, characterized in that, Also includes a physical base Methods for obtaining and representing: The physical base data The acquisition methods include: real-time sensor acquisition, spatiotemporal snapshot backtracking based on historical sensing data, and pre-stored digital twin priors retrieved from the cloud or local storage.

5. The semantic masking enhancement system based on physical law constraints according to claim 4, characterized in that: The The sources include, but are not limited to: Electromagnetic / radio physical field sensing: UWB impulse response (CIR), Wi-Fi / 6G multipath fingerprinting, millimeter wave / microwave radar echo, RSSI / phase / angle of arrival / Doppler; Acoustic echo sensing: distance / contour information of ultrasonic / sonar echo formation; Inertial measurement envelope: the velocity, displacement, and motion envelope formed by the IMU / odometer; Pre-stored digital twin priors: hard boundaries provided by CAD / BIM / point cloud / 3D map / structured model; Depth / point cloud sensing: LiDAR / ToF / structured light / binocular depth; Monocular vision motion recovery structure: Based on the depth or structural features calculated by monocular vision SfM / monocular VIO combined with gravity constraints, it is used to form physical boundaries / feasible regions with spatial scale dimensions.

6. The semantic mask enhancement system based on physical law constraints according to claim 5, characterized in that: The physical base data Represented in any of the following forms or combinations thereof: Distance field / Signed distance field ; Occupied raster / voxel representation ; Hard boundary set ; The set of boundary surfaces or feasible regions obtained by inversion of radio / acoustic reflection characteristics .

7. The semantic masking enhancement system based on physical law constraints according to claim 1, characterized in that: The physical consistency audit module / audit filter performs virtual-physical consistency verification based on gravity benchmarks: enhancing semantic masks. The spatial identifier (spatial ID) bound to the virtual object is determined as physical boundary constraint information, and the gravitational acceleration direction reference is determined based on the geocentric gravity vector. The predicted pose / trajectory / contact state output by the spatial calculation inference module is subject to mandatory physical consistency verification. When the prediction result violates the physical boundary constraint, gravity direction consistency constraint or contact feasibility constraint, the audit filter has a veto right and forcibly intercepts and corrects the rendering / interaction instructions through at least one of the following methods: projection back to feasible domain, rigid body collision correction, dynamic constraint correction or rollback strategy, so as to suppress clipping, drift or unreasonable displacement of virtual objects relative to physical entities.

8. A spatial computational reasoning method used in the semantic masking enhancement system as described in claim 1, characterized in that, Including the reasoning generation process: The spatial computing inference module uses user interaction commands and enhanced masks. Physical base Generate interactive prediction output .

9. The spatial computing inference module generates interactive prediction output based on user interaction commands, enhancement mask and physical base, and inputs the interactive prediction output to the physical consistency audit module / audit filter to perform consistency verification based on enhancement mask boundary constraints and gravity benchmark; When the verification passes, output the approved rendering and interaction commands; When the verification fails, at least one forced correction strategy is triggered: rigid body collision correction, Newtonian dynamics constraint correction, projection back to feasible region or rollback strategy, and after correction, the rendering and interaction instructions that have passed the audit are output.

10. The spatial computational reasoning method according to claim 8, characterized in that: The interaction prediction output includes at least: virtual object pose, motion trajectory or velocity, contact point / contact normal / collision depth, and interaction action parameters.

11. The spatial computational reasoning method according to claim 9, characterized in that: It also includes physical conflict detection: the audit filter performs physical conflict detection on the predicted pose / trajectory / contact state or rendering / interaction instructions output by the spatial computing inference module, including at least: boundary conflict detection, gravity consistency detection, infeasible contact detection and velocity change detection, and the audit filter is independent of the AI ​​model structure.