Webgpu-based digital twin three-dimensional scene modeling method

Through the WebGPU-based dual-channel pipeline of graph computing and graphics rendering and the dynamic perception module, the shortcomings of existing 3D modeling methods in adapting to dynamic scenes and multi-source data fusion are solved, and efficient and real-time 3D digital twin scene reconstruction and interaction are achieved, which is suitable for highly dynamic scenes such as smart cities.

CN120580367BActive Publication Date: 2025-10-17ZHEJIANG ZHEFENG YUNZHI TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202511081654.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-04
Publication Date
2025-10-17
Estimated Expiration
2045-08-04

AI Technical Summary

Technical Problem

Existing three-dimensional modeling methods are difficult to quickly adapt to dynamically changing real scenes, lack the ability to directly and parallelly analyze multi-source semantic features and spatial hierarchical structures, and cannot meet the modeling requirements of multi-source state coupling, dynamic interaction and structural adaptive updating in digital twin environments.

Method used

Based on WebGPU, a dual-channel pipeline of graph computing and graphics rendering is built. Modeling is driven by multi-dimensional structural unit graphs, and combined with dynamic perception modules to capture physical scene changes in real time, achieving incremental reconstruction and real-time reconstruction of three-dimensional models, and supporting micro-semantic queries, entity-level interactions, and multi-layer data linkage.

Benefits of technology

It achieves efficient fusion of heterogeneous perception data, continuous topological modeling of three-dimensional scenes, and real-time response reconstruction in dynamic environments, improves modeling concurrency capabilities and boundary fitting accuracy, and is suitable for three-dimensional digital twin applications in highly dynamic scenarios such as smart cities, industrial operation and maintenance, building information modeling, and intelligent transportation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120580367B_ABST
    Figure CN120580367B_ABST
Patent Text Reader

Abstract

The application discloses a digital twin three-dimensional scene modeling method based on webGPU, relates to the technical field of three-dimensional scene modeling and graph computing, assembles multi-source sensing data from sparse point clouds, video texture sequences, scene semantic label graphs and structural boundary element groups asynchronously, and generates six-dimensional structural element groups through normalization operators; a structural element graph is constructed, in which node represents component entity and edge represents constraint relationship, and a tension balance mechanism and a multi-scale constraint are introduced to generate a modeling path prior model; a graph computing and graph rendering double-channel pipeline is constructed in WebGPU to execute parallel texture mapping and boundary fitting operations; a dynamic sensing module is used to capture scene disturbance and drive graph response to realize incremental reconstruction of the model; finally, the model is mapped to a Web terminal to support micro semantic query and multi-layer data linkage; the method improves the response speed, semantic consistency and structural adaptability of three-dimensional modeling in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of three-dimensional scene modeling and graphics computing, in particular to a digital twin three-dimensional scene modeling method based on webGPU. BACKGROUND

[0002] Traditional three-dimensional modeling methods usually rely on pre-modeling and offline rendering processes, and their modeling data mostly comes from structured perception devices such as laser radars or photogrammetry results, which are difficult to quickly adapt to dynamically changing real scenes. In addition, existing modeling methods based on WebGL, OpenGL and other graphics interfaces only provide image pixel stream-oriented graphics processing capabilities, lack direct parallel analysis capabilities for multi-source semantic features and spatial hierarchical structures, and are difficult to support multi-dimensional expression of entity cascading relationships in high-complexity digital twin scenes.

[0003] WebGPU, as a new generation of graphics and computing interface, has more efficient low-level instruction compilation capabilities and more parallel computing characteristics closer to hardware. However, the current technology has not established a real-time three-dimensional modeling mechanism based on WebGPU, which is driven by multi-level spatial semantic perception and scene behavior constraint features. Existing methods only call WebGPU in the graphics rendering stage, lack of adaptation and reconstruction of modeling data structures and construction logic, and cannot meet the modeling needs of multi-source state coupling, dynamic interaction and structure adaptive update in digital twin environments.

[0004] Therefore, there is an urgent need for a new three-dimensional modeling method with the following characteristics: first, it can fuse multi-dimensional semantic fields and spatial topological structures before data enters the model; second, it can build a heterogeneous data-driven parallel modeling logic flow graph based on WebGPU, rather than just graphics rendering; third, it can realize the dynamic binding of structure continuity and state update. This method will significantly improve the mapping accuracy and response initiative of digital twin systems to real physical environments, and has high novelty and breakthrough. SUMMARY

[0005] The purpose of the present application is to provide a digital twin three-dimensional scene modeling method based on webGPU to solve the problems in the background art.

[0006] In order to achieve the above purpose, the present application provides the following technical scheme: a digital twin three-dimensional scene modeling method based on webGPU, comprising:

[0007] obtaining data streams from multiple heterogeneous sources, including sparse point clouds, video texture sequences, scene semantic label maps and structure boundary element tuples, and parsing the data streams into six-dimensional structure element tuples aggregated by spatial semantics;

[0008] A multi-dimensional structure unit graph is constructed, a node represents a scene component entity, and an edge represents a constraint relationship between entities, and the graph topology is taken as a prior model input for driving modeling path generation;

[0009] Based on the graph model, a unified graph computation and graphics rendering double-channel pipeline is constructed through WebGPU to generate an initial three-dimensional model framework, and the pipeline can perform texture mapping, geometry fusion, and boundary fitting processes in parallel;

[0010] A dynamic perception module is used to capture the state of the physical scene in real time, generate a state disturbance vector set, and dynamically correct the three-dimensional model node structure through a WebGPU asynchronous scheduling mechanism to realize incremental reconstruction of the model;

[0011] Finally, the dynamic three-dimensional model is mapped to a Web terminal platform to realize real-time reconstruction of a digital twin scene supporting micro semantic queries, entity-level interactions, and multi-layer data linkage.

[0012] Preferably, the data stream obtained from the plurality of heterogeneous sources includes:

[0013] An asynchronous parallel sampling module is used to extract characteristic elements from sparse point clouds, video texture sequences, scene semantic label graphs, and structure boundary tuples, and a cross-modal collaborative normalization operator is used to scale and align different data dimensions in time;

[0014] A dynamic semantic mapping matrix is constructed, taking entity classification, spatial orientation, texture gradient, boundary continuity, time sequence response, and data source confidence as six elements to generate a structure candidate unit;

[0015] Based on a tensor configuration path scheduling algorithm, the candidate unit is spatially and semantically clustered and mapped to a six-dimensional structure unit group, which is represented as a reconfigurable topology node set in a GPU graph computation graph;

[0016] A WebGPU underlying data stream control module is used to realize continuous updating and boundary correction of the structure unit group, so that the heterogeneous source information is deeply coupled at the initial stage of structure construction.

[0017] Preferably, the graph is constructed and the prior model input is generated by:

[0018] Based on the six-dimensional structure unit group, a multi-dimensional structure unit graph is constructed, each structure unit is taken as a graph node, and a graph edge set is generated according to orientation consistency, semantic label compatibility, and boundary fitting gradient to form an initial topology structure;

[0019] An intra-graph tension balancing mechanism is introduced to the initial topology structure, the edge weight distribution is adjusted through the internal tension coefficient of the node group, and the implicit dependency relationship between complex spatial components is expressed in a high-weight path;

[0020] Embedding multi-scale constraint layers in the graph, including spatial constraint tensors, semantic preservation tensors, and hierarchical propagation operators, and constructing a graph prior structure for path generation oriented modeling;

[0021] Embedding the graph prior structure as input into the path determination module of the WebGPU pipeline to drive the structural scheduling logic and entity disassembly sequence in the modeling process, and realizing topology-driven path adaptive modeling.

[0022] Preferably, the WebGPU graph computation and graphics rendering dual-channel pipeline constructed based on the graph includes:

[0023] Initializing the dual-channel pipeline structure in WebGPU, where the graph computation channel receives the graph model as the topology control input, and the rendering channel loads the bound texture vector and geometry control pointer;

[0024] Based on the node distribution characteristics in the structural unit graph, a graph execution topology table is constructed, the node granularity is mapped to the GPU thread cluster, and the dynamic load routing graph is constructed using the cross-node edge weight to allocate parallel computing blocks;

[0025] By binding the cascading texture buffer and the geometry constraint tensor, the synchronization access between the graph computation result and the graphics rendering unit is realized, and the texture mapping, geometry fusion and boundary fitting operations are realized in different rendering frames. Asynchronous iteration linkage;

[0026] Injecting a boundary continuity dynamic adjustment operator in the rendering channel, based on the edge fitting residual of the node graph to control the boundary convergence, thereby constructing an initial three-dimensional model framework driven by topology structure and optimized by consistent constraints.

[0027] Preferably, the dual-channel pipeline further includes a graph enhancement mechanism based on disturbance feedback, specifically:

[0028] In each rendering frame period, a rendering disturbance vector composed of geometry fusion error, texture stretching rate and boundary fitting residual is extracted;

[0029] The disturbance vector is input into the graph enhancement module, and an improved graph attention mechanism is used to dynamically adjust the edge weight distribution and node activity in the structural unit graph, wherein the mechanism is based on sparse edge screening and node confidence reconstruction algorithm;

[0030] The graph enhancement result is redeployed to the execution topology table through the WebGPU graph computation channel, forming a structure and rendering bidirectional coupling feedback loop;

[0031] By adaptively adjusting the graph execution frequency and texture loading priority based on the disturbance evolution trend in a plurality of consecutive frames, the dynamic optimization of the intra-frame graphics output quality is realized.

[0032] Preferably, the dynamic correction of the three-dimensional model node structure comprises:

[0033] The state tracking unit in the dynamic perception module is used to capture the spatial displacement, material variation and occlusion relationship change in the physical scene in real time, and generate a set of state disturbance vectors;

[0034] The disturbance vectors are input into the graph structure mapping engine, the affected structure unit group is located through the disturbance feature pointer, and the disturbance propagation path is constructed in combination with the time decay weight;

[0035] The disturbance response model is deployed in the WebGPU graph calculation channel, the graph edge set and node state are reconstructed based on the nonlinear structure disturbance function, and the updated state graph is formed;

[0036] Through the asynchronous scheduling mechanism, the interframe node replacement and local topology rearrangement between the graph calculation and the graph rendering channel are realized, and the non-destructive incremental reconstruction of the three-dimensional model is realized.

[0037] In the above technical solution, the technical effects and advantages provided by the present application are:

[0038] 1、The present application realizes efficient fusion of heterogeneous perception data, topological continuous modeling of three-dimensional scene and real-time response reconstruction in dynamic environment by constructing a WebGPU-based graph calculation and graph rendering double-channel pipeline structure, combining structure unit graph atlas driving, dynamic disturbance feedback mechanism and incremental reconstruction method. Compared with the traditional modeling scheme relying on static graphics interface, the present application can significantly improve the modeling concurrency capability, improve the boundary fitting accuracy, and has low-delay local model adjustment capability, and is especially suitable for complex entity organization and continuous state evolution expression in large-scale digital twin scene.

[0039] 2、The present application realizes the deep fusion of structure expression and data expression by deploying the modeling result with structure graph as the core on the Web side, supporting micro semantic query, entity level interaction and multi-layer data linkage. This scheme breaks through the limitations of model static, rough interaction and response lag in existing Web modeling systems, has strong scalability and engineering landing ability, and is especially suitable for three-dimensional digital twin application requirements in high dynamic scenes such as smart city, industrial operation and maintenance, building information modeling (BIM) and intelligent transportation. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments described in the present application, and other drawings can also be obtained by those skilled in the art based on these drawings.

[0041] Figure 1The mind map of the method of the present application. DETAILED DESCRIPTION

[0042] In order to make the purposes, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative labor fall within the protection scope of the present application.

[0043] Embodiment 1, please refer to Figure 1 The webGPU-based digital twin three-dimensional scene modeling method described in this embodiment includes:

[0044] Obtaining data streams from multiple heterogeneous sources, including sparse point clouds, video texture sequences, scene semantic label maps and structure boundary tuples, and parsing the data streams into six-dimensional structure unit groups aggregated according to spatial semantics;

[0045] Constructing a multi-dimensional structure unit graph, taking nodes to represent scene component entities and edges to represent constraint relationships between entities, and taking the graph topology as a prior model input for driving modeling path generation;

[0046] Based on the graph model, a unified graph computation and graphics rendering double-channel pipeline is constructed through WebGPU to generate an initial three-dimensional model framework, and the pipeline can perform texture mapping, geometry fusion and boundary fitting processes in parallel;

[0047] A dynamic perception module is used to capture the physical scene change state in real time, generate a state disturbance vector set, and dynamically correct the three-dimensional model node structure through a WebGPU asynchronous scheduling mechanism to realize incremental reconstruction of the model;

[0048] Finally, the dynamic three-dimensional model is mapped to a Web terminal platform to realize real-time reconstruction of a digital twin scene supporting microscopic semantic queries, entity-level interactions and multi-layer data linkage.

[0049] In order to realize unified modeling of multi-source heterogeneous scene perception data and provide a semantic structure basis for subsequent graph-driven modeling path construction, the present application proposes a high-concurrency heterogeneous data sampling and semantic aggregation mechanism based on WebGPU, specifically including the following steps:

[0050] The present application first sets up an asynchronous parallel sampling module for real-time acquisition and extraction of key perception elements from multiple data channels. The module supports the following data sources:

[0051] Sparse point cloud data (such as obtained based on laser radar or structured light scanning);

[0052] Video texture sequence (including multi-view RGB video or depth map sequence);

[0053] Scene semantic label map (can be derived from deep learning segmentation model or image annotation system);

[0054] Structural boundary element tuple (generated by architectural plan, GIS boundary data or CAD model).

[0055] The asynchronous parallel sampling module performs feature element extraction operations on the above data sources simultaneously based on the concurrent processing capability of WebGPU scheduling. Each channel sampling thread uses a differentiated sampling strategy to set different sampling frequencies and window sizes according to the spatial density and temporal variation characteristics of the input data. For example, density-guided sampling is used for sparse point cloud channels, and time sequence key frame extraction strategy is used for texture channels.

[0056] The multi-modal data after sampling has problems such as inconsistent dimension structure and inaccurate time alignment, so it is uniformly processed by a cross-modal collaborative normalization operator. The operator includes two core processes:

[0057] Scale calibration: transform matrix is used to unify each modal data to a common coordinate system, and structural strength regularization term is used to adjust the scale difference;

[0058] Temporal alignment: through an event-driven timestamp remapping mechanism, an isochronous frame index table is constructed under the framework of GPU unified clock to ensure the synchronization of different data channels within the modeling window period.

[0059] This normalization processing method can make all kinds of data have a unified geometric semantic reference in the subsequent processing process, improving the subsequent aggregation accuracy.

[0060] After the normalization processing of multi-modal data is completed, the structure candidate unit generation stage is entered. The present application introduces a dynamic semantic mapping matrix for expressing unit semantic features and their interaction weights in a six-dimensional space.

[0061] The mapping matrix takes each data segment as an input unit, and takes the following six elements as its core dimensions:

[0062] Entity classification (C): according to the label map and point cloud clustering result, the component type (such as wall, road, tree, building edge, etc.) is identified;

[0063] Spatial orientation (P): based on three-dimensional coordinate position and direction vector, the global orientation and local orientation are extracted;

[0064] Texture gradient (T): calculate the gray scale gradient change rate in the local texture map to represent the material characteristics;

[0065] Boundary continuity (B): measure whether the current cell boundary is closed in space and whether there is a topological gap with the adjacent cell;

[0066] Temporal response (R): indicates whether the cell remains stable in multiple time frames (used to distinguish temporary occlusion from real components);

[0067] Data source confidence (W): combines the perception quality, data density and reconstruction consistency during sampling to assign a confidence weight.

[0068] The dynamic semantic mapping matrix is organized in the form of a tensor mapping, stored in sparse tensor blocks in a GPU, and shares a memory channel with a graph computing engine, providing efficient input for subsequent graph construction.

[0069] To extract structured expression units from the dynamic semantic mapping matrix, the present application proposes a tensor configuration path scheduling algorithm for converting candidate units into a set of modelable structure units.

[0070] The algorithm includes the following three steps:

[0071] Path extraction: perform a path search operation on the semantic mapping tensor, and construct a semantic connected path based on entity classification and spatial continuity;

[0072] Configuration judgment: analyze the configuration similarity of the path nodes in the six-dimensional feature space to determine whether the structure consistency condition is met;

[0073] Scheduling mapping: mark the path segments that meet the conditions as a set of structure units, and map them to a set of graph computing nodes according to the GPU processing cluster.

[0074] This process is essentially a structural semantic clustering of input heterogeneous data, ultimately generating a set of six-dimensional structure units, each representing a constructable and reconstructable entity component in space.

[0075] It is worth noting that this clustering method does not rely on traditional Euclidean distance or K-Means algorithms, but is based on graph-semantic fusion path expression, which is a non-conventional clustering mechanism in the field, with stronger semantic perception ability and topological adaptability.

[0076] The generated six-dimensional structure unit set needs to be injected as input data stream into the WebGPU pipeline for modeling operation. To ensure the consistency and responsiveness of the structure units during modeling, the present application proposes a structure unit set binding mechanism and boundary correction process.

[0077] Binding mechanism: through the WebGPU underlying data flow control module, map the structure unit set to the graph nodes of the graph computing pipeline, and simultaneously register its texture and boundary information to the texture binding table of the rendering channel;

[0078] Boundary correction: The boundary alignment error between adjacent units is monitored in real time by the built-in geometry deviation monitoring unit of the GPU during initial construction, the boundary regression operator is used for vector correction, and local topological clipping is performed to ensure model continuity.

[0079] The structure unit group binding mechanism ensures that heterogeneous perception information can be deeply coupled at the beginning of modeling, avoiding the error accumulation problem of traditional methods of "modeling first and then splicing", and has the advantages of high modeling accuracy, strong boundary processing continuity, high atlas expression consistency and the like.

[0080] The present application firstly constructs a multi-dimensional structure unit atlas based on six-dimensional structure unit groups. In the atlas:

[0081] Each structure unit group is abstracted as a node in the atlas;

[0082] The connection relationship between the nodes is represented by a set of constructive atlas edges, and the generation of the edges is based on the following three heterogeneous structure criteria:

[0083] Orientation consistency: Calculate the included angle of the direction vectors of two structure units in three-dimensional space, and if it is less than a set threshold, it is considered that there is a geometric extension relationship;

[0084] Semantic label compatibility: Determine whether the entity classification corresponding to the structure unit can be combined to form a higher-level semantic unit (such as "wall + window frame" combined as "building facade");

[0085] Boundary fitting gradient: Extract the fitting error rate at the boundary of adjacent units, and judge the boundary continuity and topological transition smoothness through the boundary gradient tensor.

[0086] Through the above three criteria, the system generates a sparse atlas edge set in parallel in the GPU, and together with the node set, it forms an initial topological structure atlas. This structure not only expresses the direct adjacency relationship between the structure units, but also can capture the potential semantic connection relationship, providing a structural basis for subsequent path optimization.

[0087] The initial topological graph only expresses the static connection relationship, and it is difficult to accurately depict the structure complexity and entity dependence strength. To solve this problem, the present application introduces an intra-graph tension balancing mechanism in the structure atlas, which is used to dynamically adjust the edge weight distribution to enhance the expression ability of the atlas.

[0088] The basic principle is:

[0089] For each node group (i.e. a cluster of structure units closely related in geometry or semantics), set the tension coefficient matrix T(i,j) representing the structural dependence tension between node i and node j;

[0090] The tension coefficient is generated by fitting the error of the aggregation boundary, the texture consistency score and the dynamic disturbance response to generate a tension strength index.

[0091] The atlas edge weight w(i,j) is updated in real time in the WebGPU graph computation engine according to the following formula: ; wherein, is the initial edge weight, and alpha and beta are dynamic adjustment coefficients for controlling the weight balance between static connection and dynamic tension.

[0092] Through this mechanism, components with strong coupling relationships will be preferentially connected by high-weight paths in the graph, thereby obtaining higher structural scheduling priority during modeling path generation, effectively improving the accuracy of structural organization in complex scenarios.

[0093] After completing the tension balance, the application further embeds a multi-scale constraint layer in the graph structure to constrain the structural consistency and semantic coherence during the modeling path generation process.

[0094] The constraint layer is composed of the following three types of tensors:

[0095] The spatial constraint tensor (SCT) represents the relative position and expansion direction of the structural unit in the spatial organization, preventing the modeling path from appearing in space;

[0096] The semantic preservation tensor (SPT) records the entity semantic category of each node and its propagation path, ensuring that the semantic label is not damaged during the model reconstruction process;

[0097] The hierarchical propagation operator (HPO) simulates the hierarchical structure propagation relationship of the nodes in the graph, dynamically adjusting the depth-first or breadth-first strategy of the modeling path.

[0098] The above constraint layer participates in the path judgment process during graph computation, controls the path growth direction and node access order, and makes the modeling process take into account the geometric integrity, semantic coherence and execution efficiency.

[0099] Finally, the constraint layer and the graph structure are fused to form a graph prior structure, which is essentially a high-dimensional structural semantic graph that supports scheduling reasoning, used to drive the path planning of the GPU modeling process.

[0100] The graph prior structure is input as input and injected into the path judgment module in the WebGPU pipeline, which is composed of two cooperative sub-modules:

[0101] The structural scheduling unit performs modeling path growth according to the edge weight distribution in the GPS and the tension feedback, uses a scheduling algorithm based on belief propagation, and preferentially selects path segments with high tension and consistent constraints;

[0102] Entity deconstruction manager: receives the modeling path segment output by the scheduling unit, parses the geometric control points and texture parameters of the corresponding node, and issues them to the graphics rendering pipeline.

[0103] Throughout the entire execution process, the system forms a structure execution chain from the atlas → path determination → rendering modeling, realizing a topology-driven path adaptive modeling mechanism, which has the following advantages:

[0104] The modeling path can be dynamically adjusted according to the complexity of the actual scene;

[0105] Maintain spatial consistency and semantic stability between structural components;

[0106] Support interruptive incremental update to improve the sustainability and response speed of the modeling task.

[0107] To achieve efficient and dynamically adaptable three-dimensional scene modeling driven by structural atlas, the present application proposes a double-channel pipeline structure of graph computing and graphics rendering based on WebGPU, and further combines a graph atlas enhancement mechanism based on rendering disturbance feedback to build a unified modeling execution framework with structure scheduling capability, parallel rendering capability and adaptive optimization capability.

[0108] The pipeline structure breaks down the data barrier between the structural expression of the graph model and the geometric rendering, enabling them to work together and form a complete optimization closed loop through the feedback mechanism.

[0109] In the present application, WebGPU is configured to run two cooperative channels simultaneously:

[0110] Graph computing channel: mainly responsible for processing the graph structure and modeling scheduling path analysis. This channel takes the structural unit graph as input, performs structural level deconstruction, node activity judgment, edge weight reasoning, etc.

[0111] Graphics rendering channel: used to perform texture mapping, geometric fusion and boundary fitting operations, etc., to build a visual three-dimensional structure.

[0112] In the initialization stage, the system first inputs the graph model structure into GCC through the WebGPU graph binding interface as a topology control reference. At the same time, the system loads the texture vector data and geometric control pointer to be bound into the GRC channel through the buffer mapping mechanism, establishing the initial state of modeling execution.

[0113] To realize the effective parallelization of the graph model on the GPU, the present application introduces a graph execution topology table to map the node distribution characteristics of the structural unit graph to the GPU thread resources.

[0114] The key steps are as follows:

[0115] Node mapping: According to the semantic independence and tension strength of nodes in the spatial structure, each structure node is mapped to a thread cluster in WebGPU, ensuring that the thread granularity matches the complexity of the component;

[0116] Edge weight routing construction: A dynamic load routing graph is constructed with the edge weight between nodes as the weight value, which is used to dynamically allocate computing blocks in the GPU kernel to achieve optimal scheduling of computing resources;

[0117] Path reasoning scheduling: An optimal path priority strategy is adopted to perform modeling path growth on GETT to dynamically schedule node activation order and modeling progress.

[0118] Through this mechanism, semantic-driven asynchronous scheduling execution can be realized at the GPU layer, improving parallel efficiency and supporting controllable reconstruction of complex scene structures.

[0119] To achieve the consistency between structure data and graphical output, the present application realizes the synchronous access of the graph channel and the rendering channel through the following mechanisms:

[0120] Texture buffer binding: The texture vector in the structure graph is bound to the cascading texture buffer in WebGPU to ensure consistency between texture block loading and GPU thread access;

[0121] Geometric constraint tensor transmission: The geometric control tensor generated by the graph computing channel during node deconstruction is transmitted to the rendering channel in real time to participate in the geometric fusion process;

[0122] Asynchronous frame cascading mechanism: A "main graph and secondary rendering" double-frame iteration strategy is adopted, in which the main frame performs path scheduling while the secondary frame performs rendering tasks, reducing the probability of frame blocking and forming an asynchronous linked modeling mode.

[0123] Under this cooperative mechanism, texture mapping, geometric fusion, and boundary fitting operations of three-dimensional models can be performed in parallel in different GPU frame periods to maximize rendering efficiency.

[0124] To ensure the topological coherence and visual integrity of three-dimensional structures, the present application introduces a boundary continuity dynamic adjustment operator in the rendering channel. The technical points are as follows:

[0125] Based on the fitting residual of the node boundary in the graph, a local continuity index is calculated;

[0126] A dynamic convergence function is used to control the boundary convergence rate to avoid geometric distortion caused by overfitting;

[0127] In each rendering round, the boundary tensor is fine-tuned to achieve local topological smoothing and boundary convergence control.

[0128] The mechanism effectively enhances the boundary consistency between components, avoiding the common problems of cracks, sections and skip in traditional rendering.

[0129] To further improve the robustness and responsiveness of the model, the present application introduces a graph enhancement mechanism based on rendering disturbance feedback on the basis of the double-channel pipeline, forming a closed-loop optimization path of structure and rendering.

[0130] The core steps are as follows:

[0131] In each rendering frame period, the system extracts three types of disturbance indicators from the rendering output: geometric fusion error; texture stretching rate; boundary fitting residual.

[0132] The above three indicators constitute the rendering disturbance vector, which is used to represent the stability and error distribution of the model in the current frame.

[0133] The RDV is input into the graph enhancement module, and the following operations are performed:

[0134] Sparse edge screening is used to remove edges with extremely low edge weight and redundant edge connection.

[0135] The node confidence reconstruction algorithm is used to adjust the node activity distribution.

[0136] The graph attention update operator is executed to strengthen the structure edge weight of high disturbance area to improve the structure expression ability.

[0137] The graph enhancement result is redeployed to GETT through the WebGPU graph calculation channel, forming a new scheduling path. The update shares the disturbance feedback state with the graphics rendering channel, forming a closed-loop structure of structure-rendering-feedback in the modeling process.

[0138] According to the evolution trend of RDV in a plurality of consecutive frames, the system performs graph execution frequency adaptive control and texture resource loading priority adjustment, so that the high error area obtains higher update frequency and GPU resource allocation, thereby continuously optimizing the modeling accuracy and in-frame image quality.

[0139] To adapt to the real-time changes of the spatial component state in the dynamic physical environment, the present application designs a structure disturbance driven three-dimensional model incremental reconstruction mechanism, which realizes the local correction and non-destructive update of the node structure of the three-dimensional model through the dynamic sensing module, disturbance propagation path analysis, graph response model and asynchronous topology rearrangement mechanism. The mechanism fully utilizes the asynchronous scheduling and graph calculation parallelism of WebGPU to realize the high responsiveness, stability and continuity of the modeling system.

[0140] Firstly, the continuous state changes of the physical scene are collected by the dynamic sensing module in the present application. The DSM includes a state tracking unit for capturing the following three types of changes from real-time sensing streams:

[0141] Spatial displacement variation: identify the coordinate drift, rotation or deformation of structural units in three-dimensional space, such as component movement caused by environmental disturbance;

[0142] Material variation information: detect significant changes in texture maps or surface materials, such as pollution, damage or abnormal lighting;

[0143] Occlusion relationship reconstruction: Real-time capture of occlusion relationship changes between front and back components, identify new appearing or disappearing structural areas.

[0144] The above change data is mapped into a unified state disturbance vector, which is encoded in the form of a sparse tensor block in GPU for subsequent graph structure processing module calls.

[0145] In the present application, a graph structure mapping engine is provided for receiving SDV and performing disturbance influence positioning. The main steps are as follows:

[0146] Disturbance feature pointer generation: map each component in SDV to the node feature space in the structural unit graph, forming a disturbance feature pointer (DFP);

[0147] Structural unit group positioning: quickly determine the affected node set and its surrounding high dependence substructure through pointer lookup mechanism, and generate candidate structural unit groups;

[0148] Disturbance path construction: According to the edge weight and dynamic confidence between nodes, combined with time decay weight function (TDWF), construct the disturbance propagation path graph. TDWF takes the disturbance occurrence time and frequency as input, dynamically decays the influence of long-time disturbance, and ensures that the disturbance response space is focused and time convergent.

[0149] This path graph logically constitutes a local dynamic influence graph, which is the input basis for subsequent graph reconstruction.

[0150] In order to realize the intelligent response adjustment of graph structure, the present application deploys a nonlinear structure disturbance response model in the WebGPU graph calculation channel, and its technical details are as follows:

[0151] Disturbance function modeling: Construct a structure disturbance mapping function D(i) = f(SDV(i), T(i), L(i)) in the form of a polynomial kernel function, where SDV(i) is the disturbance vector of node i, T(i) is its topological position tensor, and L(i) is the hierarchical semantic attribute;

[0152] Edge set reconstruction: Recalculate the edge weight between nodes using the above disturbance function, and perform edge screening mechanism to update the graph edge set, remove broken edges, and strengthen high response edges;

[0153] Node state reset: Recalculate the geometric control parameters and semantic activation values for the disturbed nodes, forming an updated state graph (USG) that will serve as the path input for the next frame of modeling execution.

[0154] NSDRM provides a graph reconstruction mechanism based on disturbance feature reasoning, which is significantly different from traditional static model switching or full graph redrawing methods, and has strong structural continuity and computational efficiency advantages.

[0155] After the generation of the updated state graph, the invention dynamically introduces the new graph into the dual-channel modeling process through an asynchronous scheduling mechanism, achieving non-destructive incremental reconstruction. Specifically, the following operations are included:

[0156] Inter-frame node replacement: Using inter-frame version control mechanism, hot replacement operation is performed on the disturbed nodes in the GPU graph calculation graph in the current rendering frame, replacing the old node state and maintaining the memory structure continuity;

[0157] Local topology rearrangement: Local topology rearrangement logic is executed in the graph structure, only reconstructing the disturbed area and its first-order adjacent area to avoid full graph calculation;

[0158] Inter-channel synchronous broadcast: The structure adjustment results are pushed to the graphics rendering channel through the GPU shared buffer, updating the geometric texture and control pointer to ensure structure-rendering data consistency;

[0159] Execution state feedback marker: Set a feedback marker frame for each graph update, which is used to determine whether further rearrangement is needed in subsequent disturbance trend tracking.

[0160] Through this mechanism, the three-dimensional structure model can respond to environmental changes in real time without interrupting the rendering task, achieving micro-granularity reconstruction at the component and semantic levels, and building a digital twin modeling system with dynamic adaptability and topological stability.

[0161] To achieve efficient visualization and real-time interactive application of digital twin modeling results, the invention finally maps and deploys the dynamically constructed three-dimensional model to the Web terminal platform, building an interactive digital twin visualization system for user operation, achieving micro-semantic query, entity-level interaction, and multi-layer data linkage.

[0162] The three-dimensional model completed in the modeling phase is dynamically generated by the structure unit graph, and the model structure includes entity node topology, texture mapping parameters, boundary fitting attributes, and state disturbance markers. To ensure efficient loading and smooth rendering of the model in the Web terminal environment, the invention uses the following encapsulation and mapping mechanisms:

[0163] Model slice packaging mechanism: the complete model is divided into micro granularity component units according to the structure level, and is packaged into a Web compatible format (such as glTF+ binary texture block), while the structure index is retained;

[0164] State synchronization channel establishment: through the state flow buffer shared by WebGPU and the backend graph computing engine, the model state (position, semantic label, change identification) is synchronized to the front-end rendering module in real time;

[0165] Mapping distribution mechanism: the system adopts a server-side sharding scheduling + client-side incremental loading mechanism, dynamically distributes the model region according to the user view range, and reduces the first loading pressure.

[0166] This step ensures the smooth analysis and structural integrity of the modeling results in the Web environment.

[0167] The application provides a micro semantic query mechanism based on a structure map, and a user can perform semantic attribute query on any structure unit in the model through point selection, frame selection or keyword input in a Web interface.

[0168] The technical process is as follows:

[0169] All structure unit nodes are bound with semantic label vectors (including entity type, function attribute, historical state, etc.) when the graph is constructed;

[0170] After the query request is triggered, the system matches the target node through the graph index module and returns the semantic vector;

[0171] The semantic relationship in the graph structure is supported, and when the building facade is queried, the wall, window and decoration structure nodes can be traced.

[0172] The query mechanism enables the user to obtain semantic information of any component in the scene at a fine granularity, and enhances the explainability and application expansion capability of the model.

[0173] In the Web terminal platform, a user can perform entity-level interactive operation on a three-dimensional model, including rotation, scaling, attribute viewing, local reconstruction triggering, etc. In order to realize accurate entity-level control, the application constructs the following control mechanisms at the deployment end:

[0174] Interactive event mapping layer: a quick mapping table of user interactive events (such as mouse click, touch drag) and structure unit ID is established, and event-component bidirectional binding is realized;

[0175] Attribute window linkage mechanism: after clicking the component, the bound attribute window is automatically popped up, and real-time state, texture information and historical disturbance record are displayed;

[0176] Local update trigger interface: allows users to mark model errors or change areas, and after triggering, the feedback is written to the backend disturbance scheduling module, restarting the local atlas construction and node incremental reconstruction process, realizing closed-loop control from the front end to the modeling link.

[0177] This interaction mode not only enhances the visual experience, but also provides an artificial intervention entry for system adaptive optimization.

[0178] Considering that digital twin systems often involve multiple types of data (such as real-time sensor data, operation logs, historical simulation results, etc.), the present application sets up a multi-layer data linkage architecture in Web deployment, and the core design includes:

[0179] Model nodes establish an index relationship with external data sources and bind states by node ID through a unified data bus;

[0180] Different data layers can be interactively displayed with the three-dimensional model through view superposition, time axis sliding, or atlas association;

[0181] Supporting users to perform full-dimensional information linkage operations on the same component, including structure visualization, semantic analysis, state monitoring, and historical trend.

[0182] In order to verify the effectiveness of the WebGPU-based digital twin three-dimensional scene modeling method proposed in the present application, the applicant has built a prototype system and conducted comparative test experiments based on actual scenarios. The test results show that the present application is superior to the prior art in terms of model structure consistency, rendering response speed, and multi-source data fusion accuracy.

[0183] Test platform: Chrome Canary 116 supporting WebGPU is used on the browser side and runs on Windows 11 system;

[0184] Graphics processing hardware: NVIDIA RTX 4070 GPU with 8 GB video memory;

[0185] Point cloud data: derived from Velodyne HDL-32E scanning of building environment, containing about 1.5 million sparse points;

[0186] Video texture sequence: multi-view building facade real video with 1080p resolution, frame rate 30 fps;

[0187] Semantic label map: based on DeepLabV3+ semantic segmentation of image sequence, outputting 13 types of semantic labels;

[0188] Boundary element set: boundary lines and node pairs from a real building CAD file.

[0189] The scheme A is compared with a modeling method (labeled as scheme B) under a traditional WebGL architecture, the latter adopts an OpenGL style structure splicing and a static texture rendering strategy, and does not have a graph structure, a disturbance feedback or an incremental reconstruction capability.

[0190] Table 1, test index and result

[0191] Indicator category Test indicator Scheme A (the present application) Scheme B (comparative scheme) Modeling accuracy Structural unit boundary alignment error (px) 1.73 4.86 Rendering response capability Average frame rate (fps) 58.9 41.2 Model update efficiency Node disturbance response delay (ms) 37.5 119.7 (requires overall reconstruction) Data fusion performance Semantic mapping consistency rate (%) 94.2 79.6 GPU resource utilization Average GPU utilization (%) 78.4 53.9

[0192] The asynchronous parallel sampling module is used to collect point clouds, image sequences, semantic graphs and boundary element groups respectively, and the time sequence alignment and scale calibration are completed through a cooperative normalization operator, and 1431 six-dimensional structure element groups are generated.

[0193] After the structure graph is constructed, a graph computing and graph rendering double-channel pipeline is established under a WebGPU environment, and the structure nodes are preferentially scheduled to the modeling path according to the edge weight.

[0194] After the model is deployed, a dynamic occlusion test (artificially occluding the camera and then removing it) is performed, the scheme A completes local reconstruction within 2 frames, and the scheme B cannot respond locally and needs to be refreshed as a whole, which shows obvious delay.

[0195] The model is deployed to the Web end, and touch interaction and semantic query are used for actual measurement, the scheme A supports query response <100ms, and the scheme B does not have entity-level response capability.

[0196] The following technical effects can be explicitly verified through the embodiment:

[0197] The structure graph modeling mechanism of the application can greatly improve the modeling parallel efficiency without sacrificing the model accuracy, after adopting the graph computing and rendering double-channel linkage mechanism, the Web end modeling response speed is significantly better than the traditional WebGL method, the local incremental reconstruction driven by the disturbance feedback greatly reduces the modeling delay in the dynamic change scene, the micro semantic query and multi-layer linkage function improve the digital twin interaction ability of the user, and under the same GPU resource configuration, the scheme has higher hardware utilization rate and improves the energy efficiency ratio.

[0198] The above is only a specific embodiment of the application, but the protection scope of the application is not limited thereto, any person skilled in the art can easily think of changes or replacements within the technical range disclosed in the application, which should be covered within the protection scope of the application.

Claims

1. A digital twin 3D scene modeling method based on webGPU, characterized by: include: Obtaining data streams from multiple heterogeneous sources, including sparse point clouds, video texture sequences, scene semantic label maps, and structure boundary tuples, and parsing the data streams into a set of six-dimensional structure units aggregated by spatial semantics; Construct a multi-dimensional structural unit graph, using nodes to represent scene component entities and edges to represent the constraint relationships between entities, and use the graph topology as a priori model input to drive modeling path generation; Based on the graph, a unified dual-channel pipeline of graph computation and graphics rendering is built through webGPU to generate an initial 3D model framework. The pipeline can perform texture mapping, geometry fusion, and boundary fitting processes in parallel. Specifically, it initializes a dual-channel pipeline structure in webGPU, where the graph computation channel receives the graph as topology control input, and the rendering channel loads the texture vectors and geometry control pointers to be bound; Based on the node distribution characteristics in the structural unit graph, a graph execution topology table is constructed, node granularity is mapped to GPU thread clusters, and cross-node edge weights are used to build a dynamic load routing graph to allocate parallel computing blocks; By binding cascaded texture buffers and geometric constraint tensors, we can achieve synchronous access between graph computation results and graphics rendering units, enabling asynchronous iterative linkage of texture mapping, geometric fusion, and boundary fitting operations within different rendering frames. The method of binding the cascaded texture buffer and the geometric constraint tensor includes: Texture buffer binding: binds the texture vector in the structure unit atlas to the cascaded texture buffer in webGPU; Geometric constraint tensor transfer: The geometric control tensors generated by the graph computation pipeline during node deconstruction are transferred to the rendering pipeline in real time to participate in the geometric fusion process. While the primary frame performs path scheduling, the secondary frame performs rendering tasks; A boundary continuity dynamic adjustment operator is injected into the rendering pipeline, and boundary convergence is controlled based on the edge fitting residual of the node graph, thereby constructing an initial 3D model framework driven by topology structure and optimized for constraint consistency. The dynamic perception module is used to capture the changing state of the physical scene in real time, generate a state disturbance vector set, and dynamically correct the 3D model node structure through the webGPU asynchronous scheduling mechanism to achieve incremental reconstruction of the model. Finally, the dynamic three-dimensional model is mapped to the Web terminal platform to achieve real-time reconstruction of the digital twin scene that supports micro-semantic queries, entity-level interactions, and multi-layer data linkage.

2. The webGPU-based digital twin 3D scene modeling method according to claim 1, characterized in that: Acquiring data streams from multiple heterogeneous sources includes: Using asynchronous parallel sampling modules, we extract characteristic elements from sparse point clouds, video texture sequences, scene semantic label maps, and structural boundary tuples, and perform scale calibration and temporal alignment of different data dimensions through a cross-modal collaborative normalization operator. A dynamic semantic mapping matrix is ​​constructed, using entity classification, spatial orientation, texture gradient, boundary continuity, temporal response, and data source confidence as six elements to generate structural candidate units; Based on the tensor configuration path scheduling algorithm, the candidate units are spatially semantically clustered and mapped into a six-dimensional structural unit group, and the unit group is represented as a reconfigurable topological node set in the GPU graph calculation graph; The webGPU underlying data flow control module is used to achieve continuous updating and boundary correction of the structural unit group, enabling deep coupling of heterogeneous source information at the early stage of structure construction.

3. The webGPU-based digital twin 3D scene modeling method according to claim 1, characterized in that: Constructing the graph and generating prior model input includes: Based on the six-dimensional structural unit group, a multi-dimensional structural unit graph is constructed, each structural unit is used as a graph node, and a graph edge set is generated according to orientation consistency, semantic label compatibility and boundary fitting gradient to form an initial topological structure; An intra-graph tension balance mechanism is introduced into the initial topological structure, and the edge weight distribution is adjusted by the internal tension coefficient of the node group, so that the implicit dependency relationship between complex spatial components is strengthened and expressed with high-weight paths; Embed a multi-scale constraint layer in the graph, including a spatial constraint tensor, a semantic preservation tensor, and a hierarchical propagation operator, and use this to construct a graph prior structure for modeling path generation; The graph prior structure is embedded as input into the path determination module of the webGPU pipeline to drive the structure scheduling logic and entity deconstruction order in the modeling process, realizing topology-driven path adaptive modeling.

4. The webGPU-based digital twin 3D scene modeling method according to claim 1, characterized in that: The dual-channel pipeline further includes a graph enhancement mechanism based on disturbance feedback, specifically: In each rendering frame cycle, a rendering perturbation vector consisting of geometric fusion error, texture stretching rate and boundary fitting residual is extracted; Inputting the perturbation vector into a graph enhancement module, and using an improved graph attention mechanism to dynamically adjust the edge weight distribution and node activity in the structural unit graph, wherein the mechanism is based on a sparse edge screening and node confidence reconstruction algorithm; The graph enhancement results are redeployed to the execution topology table through the webGPU graph computing channel, forming a bidirectional coupled feedback loop between structure and rendering; By adaptively adjusting the graph execution frequency and texture loading priority based on the perturbation evolution trend within multiple frames, dynamic optimization of the intra-frame graphics output quality is achieved.

5. The webGPU-based digital twin 3D scene modeling method according to claim 1, characterized in that: Among the dynamic Modifying the 3D model node structure includes: The state tracking unit in the dynamic perception module is used to capture the spatial displacement, material variation, and occlusion relationship changes in the physical scene in real time, and generate a state disturbance vector set; Input the disturbance vector into the graph structure mapping engine, locate the affected structural unit group through the disturbance feature pointer, and construct the disturbance propagation path in combination with the time decay weight; Deploy a disturbance response model in the webGPU graph computing channel, reconstruct the graph edge set and node state based on the nonlinear structural perturbation function, and form an updated state graph; Through an asynchronous scheduling mechanism, inter-frame node replacement and local topology rearrangement are achieved between graph computing and graphics rendering channels, realizing non-destructive incremental reconstruction of 3D models.

Citation Information

Patent Citations

  • Geographic knowledge graph-guided twin modeling method for complex three-dimensional scene

    CN119445015A

  • Multi-dimensional digital twinning modeling method

    CN119848600A

  • High-performance volume rendering method and system based on autonomous controllable GPU pre-rendering

    CN120259518A