An augmented reality space painting system and method based on AI assistance and edge computing

The augmented reality spatial painting system based on AI and edge computing solves the problems of insufficient environmental perception precision, real-time rendering performance and multi-user collaboration efficiency in existing technologies. It realizes high-precision spatial interaction and multi-user collaborative creation, improves the robustness and energy efficiency of the system, and has a multi-level fault tolerance mechanism to ensure the stability of user experience.

CN122176161APending Publication Date: 2026-06-09HEFEI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
HEFEI UNIV
Filing Date
2026-02-11
Publication Date
2026-06-09

AI Technical Summary

Technical Problem

Existing augmented reality spatial painting technology has shortcomings in terms of environmental perception precision, real-time rendering performance, multi-user collaboration efficiency, and system robustness. It is difficult to achieve high-precision gesture painting and multi-user collaborative creation in complex physical spaces, and the system is unstable under dynamic network and computing resource conditions.

Method used

Employing an AI-assisted and edge computing-based approach, sparse view images are acquired using an RGB-D camera. High-resolution RGB images, surface normal maps, and semantic segmentation maps are generated by combining the TriMap video diffusion model. Multimodal alignment and feature fusion are performed using a generalized spatial propagation network. High-fidelity rendering is achieved by combining a 3D Gaussian sputtering neural rendering engine. An adaptive rendering decision-maker and a multi-agent resource allocator are also introduced to enable millimeter-level precision spatial interaction and multi-user collaborative creation.

Benefits of technology

It achieves second-level reconstruction from sparse input to dense, semantic environment models, supports millimeter-level precision spatial interactive positioning, ensures high reliability and low latency of the system under dynamic networks, improves the smoothness of multi-user collaborative creation and system energy efficiency, and has a multi-level fault tolerance mechanism to ensure smooth recovery of user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122176161A_ABST
    Figure CN122176161A_ABST
Patent Text Reader

Abstract

This invention discloses an augmented reality spatial painting system and method based on AI assistance and edge computing. The method includes: acquiring sparse views using an RGB-D camera and inputting them into a TriMap video diffusion model to generate a multimodal environment perception field containing RGB, normal, and semantic segmentation maps; simultaneously acquiring user gesture poses and keypoint data; inputting the perception field and gesture data into a feature fusion engine based on a generalized spatial propagation network to achieve spatiotemporal alignment and fusion, outputting a high-precision spatial painting coordinate sequence; inputting this sequence and the perception field into a neural rendering engine based on 3D Gaussian sputtering, combining gaze-driven differential rendering for real-time image synthesis; and dynamically adjusting system parameters through an adaptive rendering decision-maker based on near-end strategy optimization. The system comprises five subsystems: intelligent perception, edge computing, real-time rendering, intelligent decision-making, and collaborative fault tolerance. This invention achieves rapid environment reconstruction, accurate interaction, and multi-user collaboration, exhibiting high system resilience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of augmented reality spatial painting technology, and in particular to an augmented reality spatial painting system and method based on AI assistance and edge computing. Background Technology

[0002] Augmented reality (AR) spatial painting technology aims to allow users to create and immerse themselves in virtual 3D content using natural gestures in real physical spaces. However, existing solutions have significant limitations in environmental perception, interactive experience, system performance, and collaborative capabilities. First, at the level of environmental understanding, mainstream methods mostly rely on feature-point-based visual SLAM or sparse point clouds for scene reconstruction. The resulting maps lack fine geometric details and high-level semantic information, making it difficult for virtual handwriting to be accurately and reasonably integrated and interacted with in complex environments (such as irregular curved surfaces and transparent objects), and the initialization process is cumbersome.

[0003] Secondly, in terms of real-time rendering and interactive performance, existing systems face a severe "quality-latency-power consumption" triangle. To achieve high-fidelity visual effects, cloud rendering or high-performance local computing is often employed, but this introduces unpredictable network latency or high terminal power consumption, compromising the immersive interactive experience. Simultaneously, the lack of utilization of human visual characteristics leads to wasted computing power. In multi-user collaborative scenarios, simple client-server architectures struggle to dynamically balance the load of each user, resulting in rigid resource allocation and easily causing problems such as rendering asynchrony and increased interactive latency, severely impacting the smoothness and fairness of collaborative creation.

[0004] Finally, existing solutions are generally fragile in terms of system architecture robustness and intelligent management. They lack the ability to perceive and adapt to dynamic network conditions and heterogeneous computing resources, and cannot gracefully degrade and quickly recover under performance pressure, network fluctuations, or node failures. The conflict detection and resolution mechanisms for multi-user concurrent operations are also relatively primitive, usually employing simple locking mechanisms or last-in-first-out strategies. This not only undermines creative freedom but also lacks respect for physical semantic rules, limiting the reliable application of this technology in serious industrial collaborations or large-scale public interactions. Summary of the Invention

[0005] In view of this, the present invention addresses the shortcomings of existing technologies, and its main objective is to provide an augmented reality spatial painting system and method based on AI assistance and edge computing. It aims to overcome the deficiencies of existing augmented reality spatial painting technologies in terms of environmental perception precision, real-time rendering performance, multi-user collaborative efficiency, and system robustness, providing a systematic solution that deeply integrates artificial intelligence, edge computing, and advanced rendering technologies. It enables rapid construction from sparse visual input to high-fidelity, strong semantic environmental understanding, ensuring users can perform natural gesture painting with millimeter-level precision in complex physical spaces, supporting real-time collaborative creation by multiple users, while guaranteeing high reliability, low latency, and high energy efficiency under dynamic network and computing resource conditions.

[0006] To achieve the above objectives, the present invention adopts the following technical solution: An augmented reality spatial painting method based on AI assistance and edge computing includes the following steps: S1. Acquire a sequence of sparse environmental view images from at least two different perspectives using an RGB-D camera mounted on the user terminal; S2. Input the sparse view image sequence into the pre-trained TriMap video diffusion model to simultaneously generate a high-resolution RGB image, surface normal map and pixel-level semantic segmentation map of the current environment, forming a multimodal environment perception field. S3. Through a high-precision inertial measurement unit and a visual hand key point detection model, collect the six-degree-of-freedom pose data stream and skeletal key point data stream of the user's hand gestures; S4. Input the multimodal environment perception field and the gesture action data stream into a multimodal alignment and feature fusion engine constructed based on a generalized spatial propagation network; the generalized spatial propagation network processes high-resolution visual data with significantly lower computational complexity than linear growth through its linear propagation mechanism, realizes spatiotemporal alignment and confidence-weighted fusion of gesture action features and environmental geometric and semantic features, and outputs a spatial painting coordinate sequence with millimeter-level accuracy; S5. Input the spatial painting coordinate sequence and the multimodal environment perception field into a neural rendering engine based on 3D Gaussian sputtering; the neural rendering engine includes a dynamically pruned Gaussian meta-scene representation and integrates an efficiency-aware gaze-point rendering pipeline; the gaze-point rendering pipeline uses a complete set of Gaussian meta-units for high-fidelity rendering of the visual focal area based on real-time tracked user eye movement focus, and uses a subset of Gaussian points pruned for efficiency-aware rendering of the visual periphery. S6. An adaptive rendering decision maker based on a near-end strategy optimization algorithm monitors the system's end-to-end latency, rendering frame rate, and user interaction intent in real time, and dynamically and collaboratively adjusts the computation path of the feature fusion engine, the pruning rate of the Gaussian set, and the focal zone parameters of the foveated rendering pipeline to achieve preset real-time performance, fidelity, and energy efficiency goals.

[0007] As a preferred approach: the TriMap video diffusion model described in step S2 is obtained through a four-stage progressive training strategy, which specifically includes the following ordered steps: Basic interpolation training phase: On the network image dataset, using the image frame order as a condition, the model is trained to learn the ability to interpolate high-fidelity keyframes and generate intermediate transition frames. 3D Consistency Injection Stage: On a video dataset with accurate camera pose annotations, the model is trained to generate continuous images with three-dimensional geometric consistency by using camera pose changes as a diffusion condition, ensuring the geometric structure of the generated content from different perspectives is coherent. Multimodal supervised training phase: Simultaneously inject supervision signals for the generation of surface normal maps and semantic segmentation maps; wherein, the supervision of surface normal maps is achieved by calculating the cosine loss between generated normals and real normals, and the supervision of semantic segmentation maps is achieved by introducing a semantic decoder branch with cross-entropy loss; Collaborative generation fine-tuning stage: The model obtained from the first three training stages is jointly fine-tuned end-to-end to optimize the weighted sum of RGB reconstruction loss, normal cosine loss and semantic cross-entropy loss, and finally obtain a unified model that can receive sparse view input and collaboratively output high-quality RGB images, normal maps and semantic segmentation maps.

[0008] As a preferred option: the multimodal alignment and feature fusion engine based on the generalized spatial propagation network in step S4 performs the following steps in sequence: Unified feature encoding: The input gesture action skeletal key point sequence is encoded into an action feature vector through a temporal convolutional network. At the same time, the RGB image, normal map, and semantic segmentation map are encoded into a multi-channel environment feature map through a convolutional backbone network with shared weights. The action feature vector is then concatenated with the environment feature map in the channel dimension through spatial broadcasting to form an initial fusion feature map. Multi-directional stable propagation: On the initial fused feature map, linear scanning propagation is performed sequentially in four directions: horizontal, vertical, main diagonal, and secondary diagonal. The propagation process in each direction is constrained by a learnable row random matrix to ensure that the information remains numerically stable during long-distance transmission, avoid gradient explosion or vanishing, and thus establish dense contextual connections between global pixels of the image with low computational complexity. Hidden state iterative update: For each spatial location in the feature map, its current hidden state value is calculated by weighted aggregation of the hidden state values ​​of all neighboring locations in the four propagation directions of the previous location, through a lightweight multilayer perceptron, thereby achieving feature fusion with efficient parameters and global receptive field coverage. High-precision coordinate regression: The densely fused feature map obtained after multiple rounds of iterative propagation and updates is input into a coordinate regression head composed of fully connected layers. The regression head outputs a sequence of three-dimensional spatial coordinates relative to the global coordinate system of the multimodal environment perception field for each timestamp.

[0009] As a preferred approach: the construction and real-time operation of the neural rendering engine based on 3D Gaussian sputtering in step S5 specifically includes the following sub-steps: Rapid initialization of Gaussian scenes: The RGB image and surface normal map in the multimodal environment perception field are input into a pre-trained feedforward encoder-decoder network; the encoder extracts multi-scale features of the image, and the decoder directly regresses the initial 3D Gaussian unit set representing the scene. The attributes of each Gaussian unit include 3D position, covariance matrix, opacity, and spherical harmonic function coefficient color; more than one million Gaussian units can be generated within 1 second, completing the scene construction in seconds; Adding dynamic handwriting Gaussian primitives: Based on the spatial painting coordinate sequence output in step S4, new dynamic Gaussian primitives are created in real time; each new primitive is centered at the current coordinates, and its initial covariance, color, and opacity are determined by the preset style parameters of the handwriting, and it participates in subsequent optimization and rendering together with the static scene primitives. Foveat-driven differential rendering: During the rendering of each frame, the efficiency-aware foveat rendering pipeline first determines the visual focal zone in screen space based on eye-tracking data; for pixels falling within the visual focal zone, a complete, unpruned list of Gaussian elements is retrieved from GPU memory for forward rasterization, and its cumulative color and depth are calculated; for pixels in the peripheral region, a simplified list that has been ordered by importance and pruned, retaining only the top K most relevant Gaussian elements, is retrieved for rendering, where the pruning rate K is dynamically determined by the adaptive rendering decision-maker based on the current frame rate; Online adaptive optimization: During system idle periods or background threads, continuous online optimization is performed on the Gaussian primitive set, including gradient descent updates based on view reconstruction loss, as well as primitive densification operations on newly emerging regions and pruning operations on primitives with low contribution, in order to continuously improve the scene representation quality.

[0010] As a preferred embodiment, the method supports multi-user collaborative creation and specifically includes the following collaborative control steps: User instantiation branch creation: When a new user joins, the system creates an independent instantiation Gaussian scene representation branch for the user; this branch copies the initial static Gaussian primitive set from the main scene representation through a deep copy and assigns an independent dynamic handwriting Gaussian primitive list; Main scene update synchronization mechanism: The online adaptive optimization results of the main scene are incrementally updated and broadcast in real time to all online user instantiation branches through the state synchronization engine to ensure that the static scene representation of each branch is consistent with the main scene; Semantic scene map construction: Based on the pixel-level semantic segmentation map generated in step S2, the two-dimensional semantic labels are back-projected into the three-dimensional space using the occupancy grid mapping algorithm to construct a three-dimensional voxelized scene occupancy grid map with different semantic labels; the semantic labels include at least passable areas, walls, and furniture; Intelligent dynamic resource allocation: The multi-agent resource allocator based on near-end strategy optimization models each user and its corresponding computing and rendering needs as an agent; the allocator's state space includes the number of primitives in each user branch, real-time interaction latency, spatial location, and the semantic scene map; through offline pre-completed centralized training, it learns a resource allocation strategy that can minimize the global average latency and maximum latency under total resource constraints, and dynamically allocates CPU / GPU time slots and memory bandwidth of edge computing nodes to each instantiated branch; Semantic-aware anti-collision response: Real-time calculation of the 3D position or bounding box projection of dynamic handwriting Gaussian primitives in all user instantiation branches onto the 3D semantic occupancy grid map; when overlapping grids occupied by handwriting primitives from different users are detected, and the grid is marked with the semantic label "not passable", an anti-collision mechanism is immediately triggered: First, a semi-transparent wavy warning mask based on a physical field is overlaid and rendered in the collision area in the corresponding user's display screen; second, directional warning sound effects are provided to the user's ears through spatial sound field technology; finally, a slight avoidance guidance vector is provided for the subsequent generation of handwriting coordinates through a path planning algorithm.

[0011] As a preferred solution, the specific execution process of adaptive rendering tasks includes the following steps: Layered task modeling: The complete rendering task is divided into high-priority tasks and low-priority tasks; the high-priority tasks are defined as rasterization and shading of Gaussian primitives in the current user's visual focal area and near-focal area; the low-priority tasks are defined as rendering of the peripheral area, background area and non-critical dynamic objects, as well as incremental optimization calculation of the Gaussian primitive set. Unloading Decision Generation: The rendering task decision-maker continuously monitors the real-time frame rendering duration, GPU utilization, and network round-trip latency and available bandwidth to the cloud server of the edge nodes; the decision-maker uses a lightweight machine learning model to output a binary decision in real time based on these state characteristics: whether to unload low-priority tasks, and if unloading, determine the data block size and compression level of the unloaded task. Collaborative rendering execution: When the decision is to offload, the edge node sends the pruned and compressed Gaussian metadata blocks and rendering parameters corresponding to the low-priority task to the cloud through a dedicated streaming channel; the cloud server uses its powerful parallel computing capabilities to quickly complete the rendering or optimization calculation of the specified area, and sends back the generated sub-image or the updated model parameter difference after efficient video encoding; the lightweight rendering client of the edge node receives the data, decodes it, and performs alpha mixing between the sub-image rendered by the cloud and the focal area image rendered locally to synthesize the final image.

[0012] As a preferred option: it features a multi-layered, resilient fault-tolerance mechanism, executing a differentiated three-level response process based on the severity of the anomaly. Level 1: When the adaptive rendering decision-maker detects that the generation latency of N consecutive frames exceeds the first threshold but is lower than the second threshold, it determines that there is mild performance pressure and automatically triggers a performance enhancement command: instructing the foveated rendering pipeline to adopt a more aggressive pruning strategy, reducing the primitive sampling number K value of the peripheral region by one level; at the same time, instructing the feature fusion engine to temporarily skip non-critical feature refinement layers; where N is an integer greater than or equal to 3; Level 2: When the communication quality diagnostic unit detects that the continuous packet loss rate of the link with a certain edge computing node or user terminal exceeds the set limit, it determines that the communication is interrupted. The system immediately freezes the latest state snapshot of the affected module and stores it in the persistent cache. On the user terminal side, it switches to using the most recent complete scene data that has been cached locally for rendering, and temporarily stores the new painting operation data in the local queue with a high-precision timestamp. Level 3: When a complete interruption of connections to the cloud is diagnosed, or the load on the edge master computing node exceeds the safety threshold, a global service degradation is initiated; the system control terminal device activates the built-in lightweight visual inertial odometry module and loads an extremely simplified, parameter-fixed micro Gaussian sputtering model; the user performs restricted spatial drawing based on the pose and micro model provided by the local VIO, on the basis of the last known scene anchor point, and all operations are recorded locally as a time-stamped instruction stream; after communication is restored, the instruction stream is uploaded to the edge node for offline replay and state synchronization, achieving lossless recovery.

[0013] An augmented reality spatial painting system based on AI assistance and edge computing for implementing the method includes the following interconnected subsystems: The intelligent sensing subsystem includes the RGB-D camera, the high-precision inertial measurement unit, and a terminal processor that runs the hand key point detection model, which is responsible for the synchronous acquisition and preliminary processing of raw multimodal data; The edge fusion and computing subsystem consists of multiple microservices deployed on a physically distributed edge server cluster; the subsystem includes at least: TriMap model inference service, responsible for executing step S2; GSPN feature fusion service, responsible for executing step S4; and scene management service, responsible for maintaining and managing the 3D Gaussian scene representation; The real-time rendering and output subsystem consists of a foveated rendering service deployed on edge nodes with high-performance GPUs and a lightweight rendering client running on the user terminal. The two are connected through a low-latency streaming protocol and work together to complete the differential rendering and image compositing in step S5. The intelligent decision-making and coordination subsystem, as the control center of the system, is physically deployed on the management node of the edge computing cluster. This subsystem integrates an adaptive rendering decision-maker, a multi-agent resource allocator, a rendering task decision-maker, and a communication quality diagnostic unit, and sends control commands to other subsystems through a publish-subscribe mechanism. The collaboration and fault tolerance management subsystem, as a module to ensure system resilience, includes a state synchronization engine, a multi-level state manager, and an offline fault tolerance client running on the user terminal, which together implement multi-user management and exception handling processes.

[0014] As a preferred embodiment, the internal communication between the edge fusion and computing subsystem, the real-time rendering and output subsystem, and the intelligent decision-making and coordination subsystem is interconnected using a low-latency data center network based on RDMA technology to ensure that the latency of data transmission between microservices is less than 100 microseconds; the end-to-edge communication between the intelligent sensing subsystem and the terminal rendering client and the edge cluster is achieved using a highly reliable, low-latency wireless network based on 5G URLLC slicing or Wi-Fi 6E to ensure that the end-to-end round-trip latency of uplink control data and downlink video stream is stable within 20 milliseconds.

[0015] As a preferred solution, the state synchronization engine in the collaboration and fault tolerance management subsystem adopts an optimistic synchronization and conflict resolution strategy, which works as follows: During normal multi-user collaboration, the operation instructions of each user terminal are marked with a logical timestamp in real time and sent to the edge server for global sorting and execution; When a user reconnects from offline fault-tolerant mode, their local timestamped operation instruction stream is uploaded to the state synchronization engine. The state synchronization engine first adopts a "replay-compare" mechanism: the offline instruction stream is re-executed in the current global scene state, and compared with the final scene state caused by the operations of other users during this period as recorded by the server, and possible state conflict areas are identified. For non-conflicting handwriting operations, their Gaussian elements are directly merged into the global scene; For conflict areas, a semantically guided conflict resolution algorithm is activated: based on the three-dimensional semantic occupancy grid map, handwriting in the "creative" semantic area is retained first, while for conflict handwriting generated in the "non-occupiable" semantic area, a negotiation request is initiated or a soft transparency blending process is carried out according to preset rules, ultimately achieving a smooth and consistent fusion of all user states.

[0016] Compared with existing technologies, this invention has significant advantages and beneficial effects. Specifically, as can be seen from the above technical solution, in terms of technical performance, it achieves second-level reconstruction of dense, semantic environment models from sparse input, as well as millimeter-level precision spatial interaction positioning. Simultaneously, end-to-end latency is stabilized at an extremely low level through foveated rendering and adaptive decision-making. In terms of application experience, the system supports smooth collaborative creation by a large number of users in the same virtual scene, and ensures the naturalness and rationality of interaction through semantic-aware collision avoidance and intelligent conflict resolution. Regarding system efficiency and reliability, edge-cloud collaborative computing and dynamic resource scheduling significantly improve overall energy efficiency and resource utilization. Its built-in multi-level fault tolerance mechanism ensures that user experience can smoothly degrade and quickly recover without loss when facing network fluctuations or node failures, greatly enhancing the system's practicality and applicability.

[0017] To more clearly illustrate the structural features and effects of the present invention, a detailed description is provided below in conjunction with the accompanying drawings and specific embodiments. Attached Figure Description

[0018] Figure 1 This is an overall flowchart of the augmented reality spatial painting method of the present invention; Figure 2 This is a schematic diagram illustrating the core principle of multimodal environment perception and neural rendering of the present invention; Figure 3 This is a schematic diagram of the multi-user collaboration and resource management architecture of the present invention; Figure 4 This is a block diagram of the overall architecture of the augmented reality spatial painting system of the present invention. Detailed Implementation

[0020] The present invention is as follows Figures 1 to 4 As shown, an augmented reality spatial painting system and method based on AI assistance and edge computing are described, wherein: The method includes the following steps: As attached Figure 1 As shown, attached Figure 1 This is a flowchart illustrating the overall process of the augmented reality spatial painting method described in this invention, showing the complete processing flow from data acquisition to final image output. The process begins in step S1, where the user terminal's RGB-D camera acquires a sequence of sparse environmental view images from at least two different perspectives. Then, in step S2, a pre-trained TriMap video diffusion model is used to generate a multimodal environmental perception field containing high-resolution RGB images, surface normal maps, and semantic segmentation maps, providing the system with rich prior environmental information. Simultaneously, step S3 acquires the user's six-DOF pose and skeletal keypoint data stream using an inertial measurement unit and a hand keypoint detection model. Step S4 inputs the environmental perception field and gesture data into a multimodal alignment and feature fusion engine built on a generalized spatial propagation network, achieving spatiotemporal alignment and confidence-weighted fusion of gesture actions and environmental features, outputting a spatial painting coordinate sequence with millimeter-level precision. Step S5 inputs this coordinate sequence and the perception field into a 3D Gaussian sputtering neural rendering engine, combined with a gaze-driven differential rendering strategy, to achieve high-fidelity and efficient image synthesis. The entire process operates under the dynamic control of the adaptive collaborative decision optimization module in step S6. This module adjusts the parameters of each stage in real time based on the near-end policy optimization algorithm to ensure that the system achieves the best balance between latency, frame rate, and energy efficiency. This flowchart intuitively illustrates the closed-loop architecture of the invention, namely "perception-alignment-rendering-decision," highlighting the synergistic effect of the AI ​​model and edge computing, as well as the system's technical advantages in real-time performance, accuracy, and adaptability.

[0021] Step S1: Acquisition of a sparse environmental view image sequence; The user operates their user terminal (e.g., AR glasses, smartphone, or tablet equipped with a depth sensor) in the real physical environment to be rendered. The terminal's RGB-D camera module is activated, enabling simultaneous acquisition of color (RGB) images and depth information. The user simply holds the terminal and briefly takes or scans the environment from at least two perspectives with significant parallax (i.e., different shooting angles) to obtain a sparse environmental view image sequence. Here, "sparse" refers to the availability of only a small number (e.g., 2-5) of images from different perspectives, compared to the hundreds or thousands required for dense 3D reconstruction. This provides sufficient reconstruction data for the subsequent AI model, greatly simplifying user operation and initialization time.

[0022] Step S2: Multimodal Environment Awareness Field Generation; The sparse view image sequence (including RGB images and corresponding depth maps) acquired in Step S1 is uploaded via network to a dedicated AI inference service deployed on edge computing nodes (such as 5G MEC servers close to the user). This service loads and runs a pre-trained TriMap video diffusion model. This model is a generative AI model based on a denoising diffusion probability model, specializing in inferring and generating high-density, high-quality, and multimodal consistent outputs from extremely sparse video inputs. After receiving the sparse sequence, the model performs forward inference and simultaneously outputs three high-resolution images aligned with the input view: 1) a high-resolution RGB image, which fills in the visual information between sparse viewpoints, generating a coherent and realistic color view of the scene; 2) a surface normal map, where each pixel value (usually RGB three channels) encodes the three-dimensional information of the surface orientation (normal vector) of the corresponding scene point, which is key to understanding the geometric structure of the environment; and 3) a pixel-level semantic segmentation map, which assigns a predefined semantic category label (such as "floor," "wall," "table," "person," etc.) to each pixel, thereby adding high-level semantic understanding to the scene. These three elements—color (RGB), geometry (normal), and semantics (segmentation)—together constitute an information-rich multimodal environment perception field, serving as an accurate environmental prior for all subsequent processing and interaction.

[0023] Appendix Figure 2 The data flow and functional relationships between the core processing modules of this invention are further refined, and can be regarded as supplementary. Figure 1 The diagram illustrates the unfolding of steps S2 to S5. As shown, the sparse view image sequence is first input into the TriMap video diffusion model, which is obtained through four-stage progressive training. This model can collaboratively output high-quality RGB images, surface normal maps, and semantic segmentation maps, forming a multimodal environment perception field. This perception field, along with motion data from user gestures, is input into a feature fusion engine based on a generalized spatial propagation network (GSPN). This engine achieves efficient and accurate feature alignment and fusion through multi-directional stable propagation and iterative update mechanisms of hidden states, ultimately outputting a spatial painting coordinate sequence. This coordinate sequence and the perception field are then input into a 3D Gaussian sputtering neural rendering engine. This engine sequentially executes four sub-steps: rapid Gaussian scene initialization, dynamic handwriting Gaussian primitive addition, gaze-driven differential rendering, and online adaptive optimization, ultimately generating and outputting an AR overlay image. This diagram clearly illustrates the end-to-end processing chain from sparse input to dense rendering, emphasizing the core technical roles of the TriMap model in environment understanding, GSPN in motion alignment, and 3D Gaussian sputtering in real-time rendering, highlighting the system's innovation in multimodal fusion and neural rendering.

[0024] Step S3: Gesture motion data stream acquisition; while being aware of the environment, the system needs to accurately capture the user's drawing intent. This is achieved through multi-sensor fusion on the user terminal: Six-DOF pose data stream: continuously provided by the high-precision inertial measurement unit built into the terminal. IMU data (gyroscope, accelerometer) is filtered and fused using algorithms (such as Kalman filtering) to output the real-time position (X, Y, Z) and rotational attitude (roll, pitch, yaw) of the terminal device in three-dimensional space, forming a continuous data stream.

[0025] Skeletal keypoint data stream: Provided by a lightweight visual hand keypoint detection model (such as MediaPipe Hands or a proprietary optimized model) running on the terminal. This model processes hand images captured by the terminal's RGB camera in real time, identifies and outputs a 2D / 3D coordinate sequence of 21 or more keypoints of the hand, accurately describing the postures of the fingers, such as bending and extension.

[0026] After these two data streams (pose stream and keypoint stream) are synchronized in time, they jointly represent the motion trajectory and posture details of the user's hand in global space, which is the basis for analyzing drawing actions.

[0027] Step S4: Multimodal Alignment and Spatial Painting Coordinate Generation; This step is crucial for connecting user intent (gestures) with the environment (perceptual field). The multimodal environmental perception field generated in Step S2 and the gesture action data stream acquired in Step S3 are input into a multimodal alignment and feature fusion engine located on an edge server. The core architecture of this engine is based on a generalized spatial propagation network (GSPN). GSPN's core innovation lies in its linear propagation mechanism, which establishes long-range dependencies between global pixels in an image through a series of learnable linear scanning and information transfer operations performed in row / column / diagonal directions. Compared to traditional methods based on Transformer or large kernel convolution, this mechanism only exhibits approximately linear growth in computational complexity when processing high-resolution images, rather than quadratic or exponential growth, thus significantly improving efficiency while maintaining accuracy.

[0028] The engine's goal is to achieve spatiotemporal alignment and confidence-weighted fusion: Spatiotemporal alignment: Accurately match the dynamically changing sequence of hand key points (temporal dimension) with static / quasi-static environmental multimodal features (spatial dimension) to determine the precise position and orientation of the hand in the environmental coordinate system.

[0029] Confidence-weighted fusion: Taking into account the reliability of gesture data (such as whether it is occluded), the clarity of environmental features (such as the geometry of textured areas is more reliable) and semantic constraints (such as handwriting should not pass through the semantic region of "wall"), features from different sources are weighted and fused.

[0030] Ultimately, the engine outputs a spatial drawing coordinate sequence with millimeter-level precision. Every key point of the user's gesture in the air is precisely mapped to the three-dimensional coordinates (X, Y, Z) in the global environment coordinate system defined in step S2, forming the spatial path of the virtual handwriting.

[0031] Step S5: Foveatory-driven neural rendering; To overlay the virtual handwriting onto the user's field of vision in real time with high fidelity, this invention employs advanced neural rendering technology based on 3D Gaussian sputtering. The coordinate sequence generated in step S4 and the environmental perception field in step S2 are fed into the neural rendering engine.

[0032] The core data structure of the engine is a dynamically pruned Gaussian primitive scene representation. The scene is represented by hundreds of thousands to millions of Gaussian primitives (i.e., three-dimensional ellipsoids with position, covariance, opacity, and spherical harmonic color coefficients), which can represent both static environments and dynamic handwriting.

[0033] The engine integrates a key efficiency-aware foveated rendering pipeline, whose workflow is as follows: 1. Eye-tracking focus: By integrating an eye tracker or software prediction algorithm, the position of the user's current gaze point on the screen is obtained in real time, and a visual focal zone with high visual acuity is determined (e.g., a circular area with a viewing angle of 10°-20° centered on the gaze point).

[0034] 2. Differential rendering strategy: For pixels within the visual focal area: the rendering pipeline calls the complete, unpruned set of Gaussian pixels from GPU memory for rasterization computation. This ensures that the image in the area of ​​the user's attention has the highest detail, resolution, and fidelity.

[0035] For pixels in the visual periphery: a subset of Gaussian points that has undergone efficiency-aware pruning is invoked. The pruning process quickly sorts and filters based on the color contribution of each Gaussian primitive to the target pixel, retaining only the top K primitives with the largest contributions (K value is adjustable). This significantly reduces computational load, given that peripheral vision is not sensitive to detail.

[0036] Through this foveated rendering, the system intelligently concentrates rendering resources on the areas that users care about most with limited computing power, achieving the optimal balance between image quality and performance.

[0037] Step S6: Adaptive Collaborative Decision Optimization; To ensure the stable and efficient operation of the entire complex system under various conditions (such as rapid user movement, sudden changes in scene complexity, and the addition of multiple users), this invention introduces an adaptive rendering decision maker based on a near-end policy optimization algorithm. This decision maker acts as an intelligent controller, continuously monitoring the system status in three dimensions: 1) end-to-end latency (total time from gesture input to screen display); 2) rendering frame rate; 3) user interaction intent inferred through analysis of user interaction patterns (e.g., whether it's detailed rendering or rapid smudging).

[0038] Based on this real-time feedback, the decision-maker dynamically and collaboratively adjusts the key parameters of the aforementioned modules: Adjust the computation path of the feature fusion engine in step S4, for example, skipping some unnecessary feature refinement layers when it is necessary to reduce latency.

[0039] Adjust the pruning rate (i.e., K value) of the Gaussian set used for peripheral rendering in step S5, and adopt more aggressive pruning (smaller K) when the frame rate is insufficient.

[0040] Adjust the focal area parameters of the gaze point rendering pipeline in step S5, such as dynamically expanding or shrinking the focal area range based on whether the user is in "observation" mode.

[0041] The decision-maker's strategy is trained on massive simulation and real data using the PPO algorithm. Its optimization goal is to maximize the overall user experience score while satisfying multiple constraints such as real-time performance (latency <20ms, frame rate >90fps), fidelity (focus area image quality PSNR >30dB), and energy efficiency (edge ​​node power consumption below the threshold).

[0042] Training Strategy for the TriMap Video Diffusion Model: To obtain a TriMap model capable of generating high-quality, geometrically consistent, and semantically rich multimodal outputs from sparse views, this invention designs a four-stage progressive training strategy, as shown in the appendix. Figure 2 As shown: Phase 1: Basic Interpolation Training. On a massive static image dataset, using two random frames as conditions, the model is trained to predict and generate the transition frames between them, thereby enabling the model to master basic image content generation and temporal coherence.

[0043] Phase 2: 3D Consistency Injection. The model is trained on a dynamic video dataset with precise camera intrinsic and extrinsic parameter annotations. In this phase, camera pose changes between adjacent frames are used as additional conditional input to the model. This forces the model to consider not only content coherence but also adherence to 3D geometric projection rules when generating continuous images, thus outputting images with 3D geometric consistency—that is, the generated content has a reasonable and coherent 3D structure when viewed from different perspectives.

[0044] Phase 3: Multimodal Supervised Training. In this phase, the training dataset not only contains RGB video frames but also provides the corresponding ground surface normal map and pixel-level semantic segmentation map for each frame. The model generates the normal map and semantic segmentation map in parallel while denoising and generating RGB images. Supervision is achieved by introducing two additional loss functions: cosine loss for the surface normal map (measuring the difference in direction between the generated and ground normals) and cross-entropy loss for the semantic segmentation map (measuring the difference between the generated segmentation class and the ground class).

[0045] Phase 4: Collaborative Fine-tuning. The pre-trained model from the first three phases is then jointly fine-tuned end-to-end. The optimization objective is a weighted sum of the RGB reconstruction loss (e.g., L1 or L2 loss), normal cosine loss, and semantic cross-entropy loss. Through this phase, the model learns how to optimally utilize the information from the sparse input, simultaneously outputting results that are best in visual, geometric, and semantic terms.

[0046] TriMap model co-training loss function: ; The total loss during model training; RGB image reconstruction loss, such as L1 or L2 loss; Cosine similarity loss of normal maps ; Cross-entropy loss for semantic segmentation; I, N, S: represent the predicted value (pred) and the ground truth value (gt) of the RGB image, surface normal map, and semantic segmentation map, respectively; λ rgb , λ normal ,λ semantic: The weighting coefficients for each loss term are used to balance the strength of different monitoring signals.

[0047] The internal workflow of the GSPN fusion engine: Based on a generalized spatial propagation network, this multimodal alignment and feature fusion engine follows a sophisticated data processing flow. Step a: Unified feature encoding. The gesture skeleton keypoint sequence (time T × keypoint N × coordinate C) is first encoded through a temporal convolutional network (TCN) to extract temporal dynamic features and output a compact action feature vector. Simultaneously, the high-resolution RGB image, normal map, and semantic segmentation map (all H × W × C) are encoded through a weighted convolutional backbone network (such as a lightweight version of ResNet) to obtain a set of multi-channel environment feature maps with reduced spatial resolution but increased channel count. Then, the action feature vector is broadcast spatially, copied to the same spatial size (H' × W') as the environment feature map, and concatenated along the channel dimension to form the initial fused feature map.

[0048] Step b: Multi-directional stable propagation. This is the core of GSPN. The system performs linear scans sequentially along four directions on the initial feature map: horizontal (left to right / right to left), vertical (top to bottom / bottom to top), main diagonal, and secondary diagonal. In each scan propagation in each direction, information propagation is constrained and modulated by a learnable row-random matrix. This matrix ensures the numerical stability of information during propagation, effectively avoiding the gradient explosion or vanishing problem common in long-range dependency modeling, thus enabling the establishment of dense contextual connections between global pixels in the image with low computational cost.

[0049] Step c: Hidden State Iterative Update. For each location (i, j) in the feature map, a hidden state is maintained. In each propagation iteration, the new hidden state value at that location is calculated by weighting and aggregating the hidden states of all its neighboring locations (defined by the propagation directions) in the previous time step into a lightweight multilayer perceptron. Through multiple iterations, the features at each location are fused with information from various regions of the entire image, achieving parameter-efficient deep feature fusion with a global receptive field.

[0050] The GSPN hidden state update formula is as follows: ; : The hidden state vector at position (i,j) and time t; D: Set of propagation directions, namely {horizontal, vertical, main diagonal, secondary diagonal}; The set of all neighboring locations of position (i,j) along direction d; W d : The learnable row stochastic propagation weight matrix corresponding to direction d; ⨁: Vector concatenation operation; MLP θ: A lightweight multilayer perceptron with parameter θ.

[0051] Step d: High-precision coordinate regression. After several rounds of iterative updates, a dense fused feature map rich in global context information is obtained. This feature map is input into a coordinate regression head consisting of several fully connected layers. At each time stamp t, this regression head outputs a three-dimensional coordinate (x, y, t). t ,y t ,z t This coordinate is relative to the global coordinate system of the multimodal environment perception field established in step S2. Connecting the coordinates of all timestamps forms the final spatial painting coordinate sequence.

[0052] Building and running a 3D Gaussian sputtering neural rendering engine: Sub-step 4.1: Rapid Initialization of the Gaussian Scene. To quickly transform the 2D environment-aware field into a 3D renderable representation, the system uses a pre-trained feedforward encoder-decoder network. The encoder (e.g., a CNN) extracts multi-scale features from the input environment RGB and normal maps. The decoder (e.g., a series of deconvolutional layers or an MLP) directly regresses an initial set of 3D Gaussian primitives from these features. The attributes of each Gaussian unit include: 3D position (mean), covariance matrix (controlling the shape and orientation of the ellipsoid), opacity (alpha value), and spherical harmonic function coefficients to represent view-dependent colors. This network is optimized to generate over one million meaningful initial primitives from an image from a single viewpoint within one second, achieving "second-level scene construction."

[0053] Sub-step 4.2: Adding Dynamic Handwriting Hexagonal Primitives. When step S4 outputs new spatial painting coordinates, the rendering engine creates new dynamic Hexagonal primitives in real time. Each new primitive is centered at the current coordinates, and its initial attributes (such as covariance determining stroke thickness, spherical harmonics determining color, and opacity determining sharpness) are determined by the user-selected preset style parameters (brush type). These dynamic primitives are added to the scene representation and participate in subsequent frame-by-frame rendering and optimization along with the static environment primitives.

[0054] Sub-step 4.3: Foveate-Driven Differential Rendering. As described in step S5 above, in each frame rendering loop, the efficiency-aware foveate rendering pipeline first determines the visual focal zone (e.g., an elliptical region) in screen space based on eye-tracking data. Rendering is divided into two passes: for pixels within the focal zone, standard Gaussian sputtering forward rasterization is performed, accumulating the color and depth of all relevant primitives; The 3D Gaussian sputtering color rendering formula is as follows: ; C(p): The final rendered color of pixel p; N is the number of Gaussian pixels that contribute to this pixel (all in the focal region, and K pruned pixels in the peripheral region). c i The color calculated by the i-th Gaussian element based on its spherical harmonic coefficients and the viewing angle; α i The opacity of the i-th Gaussian pixel is evaluated by its two-dimensional projected Gaussian distribution at pixel p. Transmittance in volumetric rendering represents the degree to which the first i-1 primitives obstruct light.

[0055] For pixels in the outer region, rasterization is performed using a simplified list that has been sorted by importance and pruned, retaining only the top K most relevant primitives. The value of K (pruning rate) is not fixed, but is dynamically determined by the adaptive rendering decision-maker in step S6 based on the frame rate that needs to be maintained.

[0056] Sub-step 4.4: Online Adaptive Optimization. To continuously improve rendering quality and adapt to new perspectives, the system utilizes rendering gaps or background threads to perform continuous online adaptive optimization of the Gaussian primitive set. This includes: 1) Gradient descent update based on view reconstruction loss: Compare the rendering result of the current frame with possible reference inputs (such as a real RGB image transmitted from the user terminal), calculate the loss, and backpropagate to update the primitive attributes (position, color, opacity, etc.); 2) Densification operation: For blank areas or areas with large reconstruction errors appearing under new perspectives, large primitives are split into multiple smaller primitives to increase detail; 3) Pruning operation: Periodically remove primitives with extremely low opacity or minimal contribution to any perspective to maintain the efficiency of the representation.

[0057] Online adaptive optimization reconstruction loss formula: ; View reconstruction loss; Crender: The image rendered from the current Gaussian set; C target The target image can be a real RGB frame transmitted from the user terminal (for static scene optimization) or the rendering result of the previous frame (for dynamic handwriting stabilization). The square of the L2 norm is the mean squared error loss. Alternatively, the L1 loss can be used depending on the specific circumstances.

[0058] The specific implementation of multi-user collaborative creation: As attached Figure 3 As shown, attached Figure 3This diagram illustrates the system architecture and interaction mechanism of this invention in a multi-user collaborative creation scenario. In the diagram, the main scene management service is deployed on edge nodes, responsible for maintaining the global Gaussian scene representation and coordinating the states of each user terminal using an optimistic synchronization strategy through a state synchronization engine. Users A, B, and other terminal devices access the system via AR devices. Each user is assigned an independent instantiated Gaussian scene branch, which is initialized from the main scene through deep copying and independently manages the user's dynamic handwriting. The PPO multi-agent resource allocator dynamically allocates computing and rendering resources based on each user's state, location, and the global semantic scene map, ensuring fairness and smoothness in the collaborative experience. The semantic scene map stores environmental semantic information in voxel grid form, supporting the semantic perception anti-collision response mechanism. This mechanism detects the overlap of different users' handwriting in non-passable semantic areas in real time and triggers multi-layered anti-collision responses guided by vision, hearing, and behavior. This diagram illustrates the system's architectural design in supporting real-time multi-user collaboration, including key mechanisms such as state synchronization, resource scheduling, semantic perception, and conflict handling, highlighting the system's technical characteristics in terms of collaboration, semantic understanding, and natural interaction.

[0059] User-instantiated branch creation: When a new user requests to join a collaborative session through their terminal, the system creates an independent instantiated Gaussian scene representation branch for them on the edge server. This branch is not created entirely independently, but rather by copying the initial static Gaussian primitive set from the current main scene representation using a deep copy method. Simultaneously, a completely new and independent dynamic handwriting Gaussian primitive list is assigned to the user to store their personal drawing creations.

[0060] Main Scene Update Synchronization Mechanism: After online adaptive optimization, the state of the main scene (managed by the first user or server) will change (such as fine-tuning static primitive attributes). These changes are encapsulated into incremental update packages and broadcast in real time to all online user instantiation branches through a dedicated state synchronization engine. This ensures that the static environment background seen by all users remains consistent.

[0061] Semantic scene map construction: Using the pixel-level semantic segmentation map generated in step S2, the system executes an occupancy raster mapping algorithm. This algorithm combines the semantic labels in the 2D image with depth information, back-projects them into 3D space, and fills a 3D voxel grid. Each voxel (3D pixel) is marked as occupied or vacant; if occupied, it is assigned a semantic label (at least including passable areas, walls, furniture, etc.). This constructs a globally usable 3D voxelized scene occupancy raster map.

[0062] Intelligent Dynamic Resource Allocation: In collaborative scenarios, computing resources become a bottleneck. This invention employs a multi-agent resource allocator based on near-end policy optimization. It models each user and its corresponding computing and rendering needs as an agent. The allocator's state space is very rich, including: the current number of Gaussian elements in each user branch, its real-time interaction latency, its position in physical / virtual space, and a global semantic scene map. This allocator learns an optimal policy through offline pre-completion centralized training. At runtime, it can dynamically allocate resources to each instantiated branch based on the real-time state, under the constraints of total CPU / GPU time slots and memory bandwidth, aiming to minimize the global average latency and maximum latency, ensuring fairness and overall smoothness.

[0063] Semantic-aware collision avoidance response: This is key to achieving natural interaction. The system calculates the 3D position or bounding box of all users' dynamic handwriting primitives in real time and projects them onto the aforementioned 3D semantic occupancy grid map. When overlapping grids occupied by handwriting primitives from different users are detected, and the overlapping grid is marked with a "non-passable" semantic label (such as walls or the interior of physical furniture), the system immediately triggers the collision avoidance mechanism, alerting the user from multiple sensory levels: Visual warning: In the corresponding user's display screen, a semi-transparent wavy warning mask based on the physical field is overlaid and rendered in the collision area to simulate a virtual "collision" effect.

[0064] Auditory warning: Using spatial sound field technology, a directional warning sound effect is played in the user's headphones, with the direction of the sound source consistent with the direction of the collision.

[0065] Behavior guidance: Through a simple path planning algorithm, a slight avoidance guidance vector is provided for the subsequent generation of the coordinates of the currently ongoing handwriting, so that the handwriting automatically deviates slightly from the conflict area without interrupting the user's drawing process.

[0066] Adaptive rendering task offloading in edge-cloud collaboration: To cope with extremely complex large-scale scenarios or a large number of users, this invention designs an adaptive task unloading mechanism.

[0067] Step a: Layered task modeling. The system intelligently divides the complete real-time rendering and computation tasks: High-priority task: Defined as the rasterization and shading calculation of Gaussian elements in the current user's visual focal area and near-focal area. This part is directly related to the core quality of the user experience and must be completed with low latency at the edge.

[0068] Low-priority tasks include: 1) rendering of peripheral and background areas; 2) rendering of non-critical dynamic objects; and 3) incremental optimization calculations of Gaussian meta-sets (such as densification, pruning, and attribute fine-tuning). These tasks have relatively low real-time requirements or minimal impact on the current view quality.

[0069] Step b: Unloading Decision Generation. A lightweight rendering task decision maker (e.g., a small neural network or gradient boosting decision tree model) runs continuously. Key state features it monitors include: real-time frame rendering duration at edge nodes, GPU utilization, network round-trip latency to the cloud server, and available bandwidth. Based on a real-time combination of these features, the decision maker outputs a binary decision: 1) whether to unload the low-priority task of the current frame to the cloud; 2) if unloading, determine the task block size and compression level for this unloading (e.g., using lossy compression to reduce data transfer).

[0070] Step c: Co-rendering execution. When the decision is "Unload": 1. Edge nodes package and compress the data corresponding to low-priority tasks (e.g., the pruned Gaussian list of the outer region, gradient information required for optimization calculation).

[0071] 2. The data is sent to the cloud server through a dedicated high-priority data stream channel.

[0072] 3. Cloud servers utilize their powerful parallel computing capabilities (such as multiple high-end GPUs) to quickly complete the rendering of a specified area (generating a sub-image) or optimization calculations (obtaining the model parameter update amount).

[0073] 4. The cloud performs efficient video encoding (such as H.265 / HEVC) on the results (sub-images or parameter differences) and then sends them back.

[0074] 5. The lightweight rendering client at the edge node receives and decodes the data, then performs alpha blending between the sub-image rendered in the cloud and the focal area image rendered locally, seamlessly synthesizing the final image for output to the user.

[0075] 6. Multi-layered flexible fault-tolerance mechanism; To ensure the availability of the system under abnormal conditions, this invention designs a three-level progressive fault-tolerant response process.

[0076] Level 1 Response (Performance Adaptive): Trigger Condition – The adaptive rendering decision-maker detects that the generation latency of N consecutive frames (N is an integer greater than or equal to 3, such as 5) exceeds the first threshold T1 (e.g., 30ms) but is below the more severe second threshold T2 (e.g., 50ms). This is determined to be a mild performance stress. Response Action: The system automatically triggers performance enhancement commands, specifically including: 1) instructing the foveated rendering pipeline to adopt a more aggressive pruning strategy, immediately reducing the primitive sample count K value of the outer region by one level (e.g., from 500 to 300); 2) instructing the feature fusion engine to temporarily skip internal non-critical feature refinement layers to reduce computational load. This level of response aims to quickly reduce its own load and restore smoothness when the system overheats or experiences a sudden high load.

[0077] Level 2 Response (Local Communication Interruption Handling): Triggering Condition—The communication quality diagnostic unit detects that the continuous packet loss rate of the communication link with a certain edge computing node or user terminal exceeds a set limit (e.g., 10%). This is determined to be a local communication interruption. Response Actions: 1) The system immediately freezes the latest state of the affected module (e.g., the instantiated branch of the user) and generates a state snapshot, storing it in a persistent cache (e.g., local SSD). 2) On the user terminal side where the interruption occurred, the client automatically switches to offline fault-tolerant mode: rendering is performed using the most recent complete scene data cached locally, maintaining basic display; simultaneously, new user drawing operation data is temporarily stored in a local queue, and a high-precision timestamp is attached to each operation.

[0078] Level 3 Response (Global Service Degradation): Triggering Condition – Diagnosis of a complete interruption of connections to the cloud, or the load (CPU / memory) of the edge master computing node exceeding a safety threshold. Response Action: Initiate global service degradation. 1) The system control terminal device activates its built-in lightweight visual inertial odometry module, relying solely on the local camera and IMU to provide basic pose tracking. 2) The terminal loads a pre-stored, fixed-parameter miniature Gaussian sputtering model, which is extremely simplified and can only represent simple geometry and handwriting. 3) In this mode, the user performs restricted spatial drawing (e.g., supporting only basic lines of a single color) based on the pose and miniature model provided by the local VIO, at the last known scene anchor point (i.e., the last valid scene position synchronized before the interruption). 4) All operations are locally recorded as a timestamped instruction stream (a compact, serialized command describing "adding a primitive with a certain attribute at position (x,y,z) at time t"). 5) After the network connection or edge node service is restored, the locally stored instruction stream is uploaded to the edge node's state synchronization engine. The engine replays these instructions offline and compares the results with the server's current global state to synchronize the state, thereby achieving lossless recovery of all operations performed by the user while offline and ensuring data integrity.

[0079] Based on the above method, the present invention provides a hardware and software collaborative system.

[0080] The system is a distributed, modular collection of hardware and software, consisting of the following five logically separate but physically interconnected subsystems: Intelligent Sensing Subsystem: Deployed on the user terminal hardware. Physically, it includes: the sensor hardware of the RGB-D camera and the high-precision inertial measurement unit, and a terminal processor (such as a mobile SoC) running the hand keypoint detection model, data preprocessing, and sensor synchronization firmware. Its responsibility is to strictly follow the timing requirements to complete the synchronous acquisition and preliminary processing of raw multimodal data (such as image distortion correction and IMU data filtering), and package and send the processed data stream.

[0081] Edge Fusion and Computing Subsystem: This is the system's "AI brain," consisting of a set of microservices deployed on a physically distributed edge server cluster. These microservices are containerized for easy management and expansion. It includes at least three core services: 1) TriMap Model Inference Service: Dedicated to loading and executing step S2, i.e., generating the multimodal environment perception field. 2) GSPN Feature Fusion Service: Dedicated to loading and executing step S4, i.e., generating the coordinate sequence. 3) Scene Management Service: This is a continuously running service responsible for maintaining and managing the 3D Gaussian scene representation for the entire session, including performing initialization, addition, and optimization operations, and responding to queries from other services.

[0082] Real-time Rendering and Output Subsystem: This is the system's "graphics engine," consisting of two parts: 1) Foveated Rendering Service: Deployed on edge nodes equipped with high-performance GPUs. It obtains Gaussian scene data from the scene management service and executes step S5 and the foveated-driven differential rendering algorithm. 2) Lightweight Rendering Client: The application software running on the user terminal. It is primarily responsible for receiving rendering instructions or compressed image / video streams from the edge rendering service, decoding and compositing them (such as overlaying with local camera video streams), and finally displaying them on the screen. A dedicated connection is established between the two via a low-latency streaming protocol (such as a proprietary protocol based on WebRTC or optimized UDP).

[0083] Intelligent Decision-Making and Coordination Subsystem: This is the system's "control center" and "scheduling center," physically deployed on the management node of the edge computing cluster (typically with more powerful CPUs and memory). This subsystem is a software system integrating multiple intelligent decision-making modules, including: an adaptive rendering decision-maker, a multi-agent resource allocator, a rendering task decision-maker, and a communication quality diagnostic unit. These modules exchange information through shared memory or message queues and uniformly send control commands and configuration parameters to all other subsystems through a publish-subscribe mechanism (such as using Redis Pub / Sub or MQTT).

[0084] Collaboration and Fault Tolerance Management Subsystem: This is the "safety net" ensuring system resilience, existing as an independent protection module. It includes: 1) State Synchronization Engine: Its core function is to achieve multi-user synchronization and conflict resolution. 2) Multi-level State Manager: Responsible for maintaining snapshots and states of the system at different levels, supporting fault-tolerant recovery. 3) Offline Fault-Tolerant Client: As an extension function of the intelligent perception subsystem and lightweight rendering client, it runs on the user terminal and is specifically responsible for executing second- and third-level response processes in abnormal situations.

[0085] As attached Figure 4 As shown, attached Figure 4The system of this invention is presented in the form of a logical architecture diagram, illustrating the complete closed-loop process from data perception to intelligent decision-making, clearly revealing the data interaction and control relationships between the various subsystems. As shown in the diagram, the process begins with "User Information of the Intelligent Perception Solution Group," corresponding to the intelligent perception subsystem in the patent. This subsystem is responsible for collecting raw video signals, user gestures, poses, and other multimodal data through an RGB-D camera, IMU, and hand keypoint model, and performing preliminary processing and analysis. This data is uploaded to the edge via a high-reliability, low-latency communication link (such as 5G URLLC or Wi-Fi 6E) represented by "RSA 4G(V)." At the edge, the "Functional Link Analysis and Development Model" module represents the core services of the edge fusion and computing subsystem, such as TriMap model inference service, GSPN feature fusion service, and scene management service. These services perform in-depth processing on the uploaded data, generating environmental perception fields and spatial drawing coordinates. The processed data and system status information converge to the "Automatic Monitoring and Decision Maker" and "Data Analysis" modules, which together constitute the intelligent decision-making and coordination subsystem in the patent. This subsystem continuously monitors comprehensive information, including external environmental data, system performance metrics (such as frame rate and latency), and multi-user behavior. Through integrated adaptive rendering decision-makers and multi-agent resource allocators, it performs efficient task decision-making and resource scheduling. Ultimately, the generated "control strategy" is distributed in real time, dynamically adjusting parameters in key areas such as feature fusion, rendering pruning, and task unloading. This forms an intelligent closed loop of "perception-computation-decision-control," ensuring the entire system maintains optimal performance and robustness in complex and ever-changing application scenarios.

[0086] System Communication Architecture: To ensure efficient collaboration between subsystems, the system adopts a layered, heterogeneous network communication scheme. Internal communication (between edge subsystems): Communication between the edge fusion and computing subsystem, the real-time rendering and output subsystem (server-side), and the intelligent decision-making and coordination subsystem is extremely sensitive to latency. Therefore, a low-latency data center network based on RDMA technology (such as RoCEv2) is used for interconnection. RDMA allows data to be transferred directly from the memory of one server to the memory of another server, bypassing the operating system kernel and CPU, thereby reducing the latency of critical control commands and data transfer between microservices to less than 100 microseconds.

[0087] Edge-to-edge communication (between terminal and edge cluster): Communication between the intelligent sensing subsystem and the terminal rendering client and the edge cluster faces uncertainties in the wireless environment. Therefore, a highly reliable, low-latency wireless network based on 5G URLLC slicing (ultra-reliable low-latency communication) or Wi-Fi 6E (supporting the 6GHz band, with less interference and low latency) is adopted. Through dedicated network slicing or QoS guarantees, the end-to-end round-trip latency between uplink control data (gestures, poses) and downlink video stream / rendering commands is ensured to be stable within 20 milliseconds, meeting the real-time requirements of AR interaction.

[0088] The conflict resolution strategy of the state synchronization engine: The state synchronization engine in the coordination and fault tolerance management subsystem is the core for handling multi-user concurrency and offline recovery. It adopts an optimistic synchronization and conflict resolution strategy, which works as follows: During normal collaboration: In a network environment free of anomalies, every operation command from each user terminal (such as "add a handwriting primitive at time t") is marked with a local logical timestamp in real time and immediately sent to the state synchronization engine on the edge server. The engine does not execute immediately but collects all user commands, globally sorts them according to timestamps and preset rules (such as Lamport clocks or hybrid logical clocks), obtains a command execution sequence that is universally agreed upon by all users, and then applies it sequentially to the global scenario. This ensures that all online users see a consistent order of events.

[0089] Offline recovery period: When a user restores network connection from offline fault-tolerant mode, their local timestamped operation command stream is fully uploaded to the state synchronization engine.

[0090] The "Replay-Compare" mechanism: The engine first initiates this mechanism: 1) Replay: The uploaded offline command stream is re-executed in the current global scene state of the server, simulating a "hypothetical" scene branch that would occur if the user were always online. 2) Compare: The state of this "hypothetical" branch is compared item by item with the final scene state caused by the actual operations of other users during this period, as recorded by the server (mainly comparing the position and attributes of Gaussian elements), thereby accurately identifying possible state conflict areas (i.e., the same spatial area that both operations are trying to modify).

[0091] Semantic-guided conflict resolution algorithm: For conflicts identified during matching, instead of simply rejecting or overwriting them, it initiates intelligent resolution. 1. Based on the constructed 3D semantic occupancy raster map, query the semantic labels assigned to the voxels where the conflict areas are located.

[0092] 2. Priority Retention Principle: For handwriting conflicts occurring in "creative" semantic areas (such as airspace or specially marked canvas areas), the system tends to retain all creations. Preset rules, such as a soft transparency blending process, can be used to overlay the handwriting of different users in a semi-transparent manner, creating a collaborative visual effect.

[0093] 3. Negotiation or Rejection Principle: For conflicting handwriting generated in "unoccupiable" semantic areas (such as walls or the interior of physical furniture), this indicates possible user misoperation or malicious damage. The system can initiate a negotiation request (such as displaying a prompt to the relevant user) or directly process it according to preset rules (such as "later arrivals are invalid"), and fade or remove invalid handwriting.

[0094] This process ensures a smooth and consistent fusion of all user states (including those of recently recovered users), guaranteeing the continuity and fairness of the collaborative experience.

[0095] Application examples: Example 1: Cross-regional industrial equipment maintenance guidance: A turbine at a large hydroelectric power station requires maintenance. Headquarters experts and on-site engineers wear AR devices. The on-site engineer scans the equipment room (S1), and the system quickly constructs a detailed 3D model containing pipes, valves, and instruments, recognizing their semantics (S2). The expert, in the headquarters office, uses gestures to circle the sequence of bolts that need to be disassembled on the virtual model and annotates with precautions (S3-S4). Red virtual arrows and annotations are overlaid in real-time on the on-site engineer's view (S5). The system allocates resources to both parties to ensure low latency (S6). When the on-site engineer's view obstructs part of the area annotated by the expert, the annotation automatically shifts slightly to remain visible. Suddenly, the on-site network fluctuates, triggering plaintext level 2 fault tolerance. The model within the engineer's view freezes but can still be viewed, and new inspection records are saved locally. After the network recovers, the records are automatically synchronized to the model at headquarters.

[0096] Example 2: AR Interactive Marketing in Large Shopping Malls An AR treasure hunt was held in the mall atrium, with hundreds of customers participating simultaneously. Customers scanned the atrium with their phones (S1-S2), and the system generated a 3D map of the entire atrium. Customers searched for and tapped virtual treasure chests hidden at the entrances of different shops (S3-S4). All treasure chests and customer tapping effects (lighting effects) were rendered in real-time on each person's phone (S5). Faced with a massive number of users and a complex scene, the intelligent decision-maker (S6) dynamically adjusted the rendering quality and offloaded most of the rendering tasks for peripheral users and backgrounds to the cloud. The resource allocator prioritized ensuring a smooth experience for users tapping treasure chests. When the density of users in a certain area caused edge nodes to overload, a third-level fault tolerance plan was triggered, and users in that area automatically switched to a simplified local LBS AR mode, seamlessly switching back after the load decreased.

[0097] The key design focus of this invention is as follows: First, by constructing an end-to-end intelligent processing pipeline, the core of which utilizes a TriMap video diffusion model trained through a four-stage strategy to generate a multimodal environment perception field rich in geometric and semantic information from sparse views. Furthermore, it achieves precise and efficient spatiotemporal alignment and fusion of gesture actions and multimodal environment features based on a generalized spatial propagation network. Second, it employs a dynamically optimizeable neural rendering representation based on 3D Gaussian sputtering and deeply integrates efficiency-aware foveated rendering technology to achieve an intelligent balance between rendering quality and computational load. Finally, at the system architecture level, it innovatively designs edge-cloud collaborative adaptive task offloading, dynamic resource allocation based on multi-agent reinforcement learning, and a multi-layered elastic fault-tolerance mechanism, thereby constructing an efficient, collaborative, and robust computing and rendering platform.

[0098] The above description is merely a preferred embodiment of the present invention and does not constitute any limitation on the technical scope of the present invention. Therefore, any minor modifications, equivalent changes, and alterations made to the above embodiments based on the technical essence of the present invention shall still fall within the scope of the technical solution of the present invention.

Claims

1. An augmented reality spatial painting method based on AI assistance and edge computing, characterized in that, Includes the following steps: S1. Acquire a sequence of sparse environmental view images from at least two different perspectives using an RGB-D camera mounted on the user terminal; S2. Input the sparse view image sequence into the pre-trained TriMap video diffusion model to simultaneously generate a high-resolution RGB image, surface normal map and pixel-level semantic segmentation map of the current environment, forming a multimodal environment perception field. S3. Through a high-precision inertial measurement unit and a visual hand key point detection model, collect the six-degree-of-freedom pose data stream and skeletal key point data stream of the user's hand gestures; S4. Input the multimodal environment perception field and the gesture action data stream into a multimodal alignment and feature fusion engine constructed based on a generalized spatial propagation network; The generalized spatial propagation network processes high-resolution visual data with significantly lower computational complexity than linear growth through its linear propagation mechanism, achieving spatiotemporal alignment and confidence-weighted fusion of gesture action features with environmental geometric and semantic features, and outputting a spatial painting coordinate sequence with millimeter-level accuracy. S5. Input the spatial painting coordinate sequence and the multimodal environment perception field into a neural rendering engine based on 3D Gaussian sputtering; the neural rendering engine includes a dynamically pruned Gaussian meta-scene representation and integrates an efficiency-aware gaze-point rendering pipeline; the gaze-point rendering pipeline uses a complete set of Gaussian meta-units for high-fidelity rendering of the visual focal area based on real-time tracked user eye movement focus, and uses a subset of Gaussian points pruned for efficiency-aware rendering of the visual periphery. S6. An adaptive rendering decision maker based on a near-end strategy optimization algorithm monitors the system's end-to-end latency, rendering frame rate, and user interaction intent in real time, and dynamically and collaboratively adjusts the computation path of the feature fusion engine, the pruning rate of the Gaussian set, and the focal zone parameters of the foveated rendering pipeline to achieve preset real-time performance, fidelity, and energy efficiency goals.

2. The augmented reality spatial painting method based on AI assistance and edge computing according to claim 1, characterized in that, The TriMap video diffusion model described in step S2 is obtained through a four-stage progressive training strategy, which specifically includes the following ordered steps: Basic interpolation training phase: On the network image dataset, using the image frame order as a condition, the model is trained to learn the ability to interpolate high-fidelity keyframes and generate intermediate transition frames. 3D Consistency Injection Stage: On a video dataset with accurate camera pose annotations, the model is trained to generate continuous images with three-dimensional geometric consistency by using camera pose changes as a diffusion condition, ensuring the geometric structure of the generated content from different perspectives is coherent. Multimodal supervised training phase: Simultaneously inject supervision signals for the generation of surface normal maps and semantic segmentation maps; wherein, the supervision of surface normal maps is achieved by calculating the cosine loss between generated normals and real normals, and the supervision of semantic segmentation maps is achieved by introducing a semantic decoder branch with cross-entropy loss; Collaborative generation fine-tuning stage: The model obtained from the first three training stages is jointly fine-tuned end-to-end to optimize the weighted sum of RGB reconstruction loss, normal cosine loss and semantic cross-entropy loss, and finally obtain a unified model that can receive sparse view input and collaboratively output high-quality RGB images, normal maps and semantic segmentation maps.

3. The augmented reality spatial painting method based on AI assistance and edge computing according to claim 1, characterized in that, In step S4, the multimodal alignment and feature fusion engine based on the generalized spatial propagation network performs the following steps in sequence: Unified feature encoding: The input gesture action skeletal key point sequence is encoded into an action feature vector through a temporal convolutional network. At the same time, the RGB image, normal map, and semantic segmentation map are encoded into a multi-channel environment feature map through a convolutional backbone network with shared weights. The action feature vector is then concatenated with the environment feature map in the channel dimension through spatial broadcasting to form an initial fusion feature map. Multi-directional stable propagation: On the initial fused feature map, linear scanning propagation is performed sequentially in four directions: horizontal, vertical, main diagonal, and secondary diagonal. The propagation process in each direction is constrained by a learnable row random matrix to ensure that the information remains numerically stable during long-distance transmission, avoid gradient explosion or vanishing, and thus establish dense contextual connections between global pixels of the image with low computational complexity. Hidden state iterative update: For each spatial location in the feature map, its current hidden state value is calculated by weighted aggregation of the hidden state values ​​of all neighboring locations in the four propagation directions of the previous location, through a lightweight multilayer perceptron, thereby achieving feature fusion with efficient parameters and global receptive field coverage. High-precision coordinate regression: The densely fused feature map obtained after multiple rounds of iterative propagation and updates is input into a coordinate regression head composed of fully connected layers. The regression head outputs a sequence of three-dimensional spatial coordinates relative to the global coordinate system of the multimodal environment perception field for each timestamp.

4. The augmented reality spatial painting method based on AI assistance and edge computing according to claim 1, characterized in that, Step S5, which describes the construction and real-time operation of the neural rendering engine based on 3D Gaussian sputtering, specifically includes the following sub-steps: Rapid initialization of Gaussian scenes: The RGB image and surface normal map in the multimodal environment perception field are input into a pre-trained feedforward encoder-decoder network; the encoder extracts multi-scale features of the image, and the decoder directly regresses the initial 3D Gaussian unit set representing the scene. The attributes of each Gaussian unit include 3D position, covariance matrix, opacity, and spherical harmonic function coefficient color; more than one million Gaussian units can be generated within 1 second, completing the scene construction in seconds; Adding dynamic handwriting Gaussian primitives: Based on the spatial painting coordinate sequence output in step S4, new dynamic Gaussian primitives are created in real time; each new primitive is centered at the current coordinates, and its initial covariance, color, and opacity are determined by the preset style parameters of the handwriting, and it participates in subsequent optimization and rendering together with the static scene primitives. Foveat-driven differential rendering: During the rendering of each frame, the efficiency-aware foveat rendering pipeline first determines the visual focal zone in screen space based on eye-tracking data; for pixels falling within the visual focal zone, a complete, unpruned list of Gaussian elements is retrieved from GPU memory for forward rasterization, and its cumulative color and depth are calculated; for pixels in the peripheral region, a simplified list that has been ordered by importance and pruned, retaining only the top K most relevant Gaussian elements, is retrieved for rendering, where the pruning rate K is dynamically determined by the adaptive rendering decision-maker based on the current frame rate; Online adaptive optimization: During system idle periods or background threads, continuous online optimization is performed on the Gaussian primitive set, including gradient descent updates based on view reconstruction loss, as well as primitive densification operations on newly emerging regions and pruning operations on primitives with low contribution, in order to continuously improve the scene representation quality.

5. The augmented reality spatial painting method based on AI assistance and edge computing according to claim 1, characterized in that, The method supports multi-user collaborative creation and specifically includes the following collaborative control steps: User instantiation branch creation: When a new user joins, the system creates an independent instantiation Gaussian scene representation branch for the user; this branch copies the initial static Gaussian primitive set from the main scene representation through a deep copy and assigns an independent dynamic handwriting Gaussian primitive list; Main scene update synchronization mechanism: The online adaptive optimization results of the main scene are incrementally updated and broadcast in real time to all online user instantiation branches through the state synchronization engine to ensure that the static scene representation of each branch is consistent with the main scene; Semantic scene map construction: Based on the pixel-level semantic segmentation map generated in step S2, the two-dimensional semantic labels are back-projected into the three-dimensional space using the occupancy grid mapping algorithm to construct a three-dimensional voxelized scene occupancy grid map with different semantic labels; the semantic labels include at least passable areas, walls, and furniture; Intelligent dynamic resource allocation: The multi-agent resource allocator based on near-end strategy optimization models each user and its corresponding computing and rendering needs as an agent; the allocator's state space includes the number of primitives in each user branch, real-time interaction latency, spatial location, and the semantic scene map; through offline pre-completed centralized training, it learns a resource allocation strategy that can minimize the global average latency and maximum latency under total resource constraints, and dynamically allocates CPU / GPU time slots and memory bandwidth of edge computing nodes to each instantiated branch; Semantic-aware anti-collision response: Real-time calculation of the 3D position or bounding box projection of dynamic handwriting Gaussian primitives in all user instantiation branches onto the 3D semantic occupancy grid map; when overlapping grids occupied by handwriting primitives from different users are detected, and the grid is marked with the semantic label "not passable", an anti-collision mechanism is immediately triggered: First, a semi-transparent wavy warning mask based on a physical field is overlaid and rendered in the collision area in the corresponding user's display screen; second, directional warning sound effects are provided to the user's ears through spatial sound field technology; finally, a slight avoidance guidance vector is provided for the subsequent generation of handwriting coordinates through a path planning algorithm.

6. The augmented reality spatial painting method based on AI assistance and edge computing according to claim 1, characterized in that, The specific execution process of an adaptive rendering task includes the following steps: Layered task modeling: The complete rendering task is divided into high-priority tasks and low-priority tasks; the high-priority tasks are defined as rasterization and shading of Gaussian primitives in the current user's visual focal area and near-focal area; the low-priority tasks are defined as rendering of the peripheral area, background area and non-critical dynamic objects, as well as incremental optimization calculation of the Gaussian primitive set. Unloading Decision Generation: The rendering task decision-maker continuously monitors the real-time frame rendering duration, GPU utilization, and network round-trip latency and available bandwidth to the cloud server of the edge nodes; the decision-maker uses a lightweight machine learning model to output a binary decision in real time based on these state characteristics: whether to unload low-priority tasks, and if unloading, determine the data block size and compression level of the unloaded task. Collaborative rendering execution: When the decision is to offload, the edge node sends the pruned and compressed Gaussian metadata blocks and rendering parameters corresponding to the low-priority task to the cloud through a dedicated streaming channel; the cloud server uses its powerful parallel computing capabilities to quickly complete the rendering or optimization calculation of the specified area, and sends back the generated sub-image or the updated model parameter difference after efficient video encoding; the lightweight rendering client of the edge node receives the data, decodes it, and performs alpha mixing between the sub-image rendered by the cloud and the focal area image rendered locally to synthesize the final image.

7. The augmented reality spatial painting method based on AI assistance and edge computing according to claim 1, characterized in that, It features a multi-layered, resilient fault-tolerance mechanism, executing a differentiated three-level response process based on the severity of the anomaly: Level 1: When the adaptive rendering decision-maker detects that the generation latency of N consecutive frames exceeds the first threshold but is lower than the second threshold, it determines that there is mild performance pressure and automatically triggers a performance enhancement command: instructing the foveated rendering pipeline to adopt a more aggressive pruning strategy, reducing the primitive sampling number K value of the peripheral region by one level; at the same time, instructing the feature fusion engine to temporarily skip non-critical feature refinement layers; where N is an integer greater than or equal to 3; Level 2: When the communication quality diagnostic unit detects that the continuous packet loss rate of the link with a certain edge computing node or user terminal exceeds the set limit, it determines that the communication is interrupted. The system immediately freezes the latest state snapshot of the affected module and stores it in the persistent cache. On the user terminal side, it switches to using the most recent complete scene data that has been cached locally for rendering, and temporarily stores the new painting operation data in the local queue with a high-precision timestamp. Level 3: When a complete interruption of connections to the cloud is diagnosed, or the load on the edge master computing node exceeds the safety threshold, a global service degradation is initiated; the system control terminal device activates the built-in lightweight visual inertial odometry module and loads an extremely simplified, parameter-fixed micro Gaussian sputtering model; the user performs restricted spatial drawing based on the pose and micro model provided by the local VIO, on the basis of the last known scene anchor point, and all operations are recorded locally as a time-stamped instruction stream; after communication is restored, the instruction stream is uploaded to the edge node for offline replay and state synchronization, achieving lossless recovery.

8. An augmented reality spatial painting system based on AI assistance and edge computing for implementing the method of any one of claims 1-7, characterized in that, It includes the following interconnected subsystems: The intelligent sensing subsystem includes the RGB-D camera, the high-precision inertial measurement unit, and a terminal processor that runs the hand key point detection model, which is responsible for the synchronous acquisition and preliminary processing of raw multimodal data; The edge fusion and computing subsystem consists of multiple microservices deployed on a physically distributed edge server cluster; the subsystem includes at least: TriMap model inference service, responsible for executing step S2; GSPN feature fusion service, responsible for executing step S4; and scene management service, responsible for maintaining and managing the 3D Gaussian scene representation; The real-time rendering and output subsystem consists of a foveated rendering service deployed on edge nodes with high-performance GPUs and a lightweight rendering client running on the user terminal. The two are connected through a low-latency streaming protocol and work together to complete the differential rendering and image compositing in step S5. The intelligent decision-making and coordination subsystem, as the control center of the system, is physically deployed on the management node of the edge computing cluster. This subsystem integrates an adaptive rendering decision-maker, a multi-agent resource allocator, a rendering task decision-maker, and a communication quality diagnostic unit, and sends control commands to other subsystems through a publish-subscribe mechanism. The collaboration and fault tolerance management subsystem, as a module to ensure system resilience, includes a state synchronization engine, a multi-level state manager, and an offline fault tolerance client running on the user terminal, which together implement multi-user management and exception handling processes.

9. The system according to claim 8, characterized in that, The internal communication between the edge fusion and computing subsystem, the real-time rendering and output subsystem, and the intelligent decision-making and coordination subsystem is interconnected using a low-latency data center network based on RDMA technology, ensuring that the latency of data transmission between microservices is less than 100 microseconds. The end-to-edge communication between the intelligent sensing subsystem and the terminal rendering client and the edge cluster adopts a highly reliable, low-latency wireless network based on 5G URLLC slicing or Wi-Fi 6E, ensuring that the end-to-end round-trip latency of uplink control data and downlink video stream is stable within 20 milliseconds.

10. The system according to claim 8, characterized in that, The state synchronization engine in the collaboration and fault tolerance management subsystem adopts an optimistic synchronization and conflict resolution strategy, which works as follows: During normal multi-user collaboration, the operation instructions of each user terminal are marked with a logical timestamp in real time and sent to the edge server for global sorting and execution; When a user reconnects from offline fault-tolerant mode, their local timestamped operation instruction stream is uploaded to the state synchronization engine. The state synchronization engine first adopts a "replay-compare" mechanism: the offline instruction stream is re-executed in the current global scene state, and compared with the final scene state caused by the operations of other users during this period as recorded by the server, and possible state conflict areas are identified. For non-conflicting handwriting operations, their Gaussian elements are directly merged into the global scene; For conflict areas, a semantically guided conflict resolution algorithm is activated: based on the three-dimensional semantic occupancy grid map, handwriting in the "creative" semantic area is retained first, while for conflict handwriting generated in the "non-occupiable" semantic area, a negotiation request is initiated or a soft transparency blending process is carried out according to preset rules, ultimately achieving a smooth and consistent fusion of all user states.