A dynamic semantic scene reconstruction method for industrial digital twinning

CN122244332BActive Publication Date: 2026-09-18天津龙创恒盛实业有限公司
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610678334.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-05-18
Publication Date
2026-09-18
Estimated Expiration
2046-05-18

AI Technical Summary

Technical Problem

[0006]本发明的目的在于提出一种面向工业数字孪生的动态语义场景重建方法,以解决现有面向数字孪生的3DGS重建方法无法编辑交互、动态适应差以及重建数据量大的问题,通过对实际动态工厂进行重建并去除时间冗余,然后提出双不透明度动态语义高斯表示来联合训练高斯的颜色属性与语义属性,实现动态数字孪生系统的可编辑、可交互与动态更新,网络友好

Benefits of technology

[0053] This invention first distills the semantic features of multi-view images into 3D Gaussian semantic attributes to support editing interaction; secondly, it proposes a dual-opacity dynamic semantic Gaussian representation to jointly train the color and semantic attributes of the 3D Gaussian; finally, it proposes a deformation field based on time interpolation to model realistic dynamic scenes while minimizing the amount of data. This method improves the accuracy of open semantic queries by 9%, while reducing the data size to 43KB and achieving almost no loss in reconstruction quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122244332B_ABST
    Figure CN122244332B_ABST
Patent Text Reader

Abstract

The application discloses a dynamic semantic scene reconstruction method for industrial digital twinning, and relates to the field of industrial digital twinning reconstruction, and comprises the following steps: collecting multi-camera continuous video frames for dynamic and static Gaussian decomposition, and obtaining a dynamic and static Gaussian set; performing multi-level semantic feature extraction on the multi-camera continuous video frames and the dynamic and static Gaussian set by using a SAM and a CLIP model, and obtaining a multi-dimensional semantic supervision graph; obtaining a double-opacity dynamic semantic Gaussian model according to the dynamic and static Gaussian set and the multi-dimensional semantic supervision graph; and constructing a time difference value-based deformation field by using the double-opacity dynamic semantic Gaussian model, and obtaining a dynamic semantic scene reconstruction result. According to the application, the precision of open semantic query is improved by 9%, the data volume is reduced to 43 KB, and the reconstruction quality is almost not lost.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of industrial digital twin reconstruction, and in particular to a dynamic semantic scene reconstruction method for industrial digital twins. Background Technology

[0002] Digital twins, as a core technology for achieving bidirectional mapping and real-time interaction between physical systems and digital space, are not merely static 3D models of physical assets, but dynamic systems integrating geometric, physical, logical, and semantic information. They can support monitoring, simulation, prediction, and optimization throughout the entire lifecycle of a digital factory. Currently, building high-fidelity, interactive industrial digital twins with deep semantic understanding capabilities still faces the challenge of bridging the "physical-digital" gap. This is mainly reflected in how to efficiently and accurately recover the three-dimensional geometric structure from complex, dynamic, and noisy industrial field data, and transform it into semantically dynamic objects that computers can understand.

[0003] Traditional 3D reconstruction techniques, such as LiDAR scanning and multi-view geometric photogrammetry, while achieving significant success in surveying and geographic information systems, have obvious limitations when facing highly reflective metal surfaces, complex pipeline topologies, extreme occlusion, and real-time update requirements prevalent in industrial environments. With the development of deep learning and neural rendering technologies, 3D reconstruction is gradually shifting from explicit point cloud and mesh representations to implicit neural representations. Among these, 3D Gaussian Splatting (3DGS) has become the mainstream solution for industrial digital twin reconstruction due to its high reconstruction accuracy and fast rendering speed. However, existing 3D Gaussian-based reconstruction methods still suffer from three major drawbacks:

[0004] First, 3D Gaussian is essentially an unordered Gaussian function, making editing and interaction impossible, which makes digital twins difficult to manage and maintain. Second, real industrial environments are highly dynamic, and static 3D reconstruction models are prone to becoming outdated. Existing methods have poor dynamic adaptability and cannot achieve incremental updates and real-time synchronization of scenes. Third, the amount of factory-level 3D Gaussian data is huge, and coupled with the need for dynamic updates, it places extremely high demands on network transmission latency and bandwidth.

[0005] Therefore, there is an urgent need for a dynamic semantic scene reconstruction method for industrial digital twins to address the shortcomings of existing technologies. Summary of the Invention

[0006] The purpose of this invention is to propose a dynamic semantic scene reconstruction method for industrial digital twins, in order to solve the problems of existing 3DGS reconstruction methods for digital twins, such as lack of editability and interactivity, poor dynamic adaptation, and large amount of reconstruction data. By reconstructing an actual dynamic factory and removing temporal redundancy, a dual-opacity dynamic semantic Gaussian representation is proposed to jointly train Gaussian color attributes and semantic attributes, thereby realizing the editability, interactivity, and dynamic updating of the dynamic digital twin system, which is also network-friendly.

[0007] To achieve the above objectives, this invention provides a dynamic semantic scene reconstruction method for industrial digital twins, comprising:

[0008] S1. Acquire continuous video frames from multiple cameras and perform dynamic and static Gaussian decomposition to obtain dynamic and static Gaussian sets;

[0009] S2. Based on the continuous video frames from the multi-camera system and the dynamic and static Gaussian set, multi-level semantic feature extraction is performed using SAM and CLIP models to obtain a multi-dimensional semantic supervision graph.

[0010] S3. Based on the dynamic and static Gaussian set and the multidimensional semantic supervision graph, obtain a dynamic semantic Gaussian model with double opacity;

[0011] S4. Using the aforementioned dual-opacity dynamic semantic Gaussian model, construct a deformation field based on time difference to obtain the dynamic semantic scene reconstruction result.

[0012] Optionally, S1, acquire continuous video frames from multiple cameras, perform dynamic and static Gaussian decomposition, and obtain dynamic and static Gaussian sets, including:

[0013] Acquire continuous video frames from multiple cameras and calculate the pixel temporal variance of the continuous video frames from multiple cameras;

[0014] Based on the pixel temporal variance of the continuous video frames from the multi-camera system and a preset threshold, dynamic pixel labels and static pixel labels are obtained, thereby obtaining the multi-camera dynamic labels.

[0015] Learnable dynamic indication attributes are introduced based on the multi-camera dynamic labels to obtain a predicted dynamic map;

[0016] The predicted dynamic map is supervised by the multi-camera dynamic labels, and the dynamic indicator attributes are optimized by binary cross-entropy loss to obtain the optimized dynamic indicator attributes.

[0017] Based on the optimized dynamic indicator attributes and the preset threshold, obtain the dynamic and static Gaussian sets.

[0018] Optionally, the formula for calculating the predicted dynamic graph is:

[0019]

[0020] in, Let x be the predicted dynamic graph. pixels, This is the predicted dynamic graph for pixel x. Let d be a sigmoid, single-valued, continuously differentiable nonlinear activation function, m be the m-th 3D Gaussian currently being computed, N be the set of 3D Gaussians covering the pixels, and d be the sigmoid activation function. m For the dynamic indicator property of the m-th 3D Gaussian, o m Let o be the opacity property of the m-th 3D Gaussian. m (x) represents the opacity attribute of the m-th 3D Gaussian at pixel x. j (x) represents the opacity attribute of the j-th 3D Gaussian at pixel x, where j is the index of all Gaussians that have been rendered before the m-th 3D Gaussian.

[0021] Optionally, the formula for calculating the binary cross-entropy loss is:

[0022]

[0023] Where L is the binary cross-entropy loss, Let M(x) be the mathematical expectation function, and M(x) be the dynamic label of pixel x. This is a predicted dynamic graph for pixel x.

[0024] Optionally, S2, based on the continuous video frames from the multi-camera setup and the dynamic / static Gaussian set, multi-level semantic feature extraction is performed using SAM and CLIP models to obtain a multi-dimensional semantic supervision map, including:

[0025] The scene spatial range is determined using the aforementioned dynamic and static Gaussian sets;

[0026] The continuous video frames from the multi-camera system are divided based on the scene space range to obtain video frames of the image group.

[0027] The video frames of the image group are input into the SAM model for segmentation processing, and the segmentation results output by the SAM model are obtained.

[0028] The segmentation result output by the SAM model is input into the CLIP model to obtain the semantic graph output by the CLIP model;

[0029] Tensor concatenation is performed on the semantic graph output by the CLIP model to obtain a multidimensional semantic supervision graph.

[0030] Optionally, S3, based on the dynamic and static Gaussian set and the multidimensional semantic supervision graph, obtain a dual-opacity dynamic semantic Gaussian model, including:

[0031] Gaussian assignment is performed on the dynamic and static Gaussian sets to obtain language feature vectors;

[0032] Based on the language feature vector and the opacity parameter, a double-opacity Gaussian representation is constructed, and path rendering is performed to obtain the language feature rendering result;

[0033] Based on the rendering results of the language features, an L1 loss function is used to construct a language feature supervised loss.

[0034] The rendering results of the language features are supervised by the multidimensional semantic supervision graph, and a dynamic semantic Gaussian model with double opacity is obtained by combining the language feature supervision loss.

[0035] Optionally, the calculation formula for the path rendering is:

[0036]

[0037]

[0038] Among them, R c For the rendered color feature map, c m For the m-th Gaussian color feature, Let the opacity be the color feature of the m-th Gaussian. Let R be the opacity of the j-th Gaussian color feature. l For rendering the language feature map, f m For the m-th Gaussian linguistic feature, Let the opacity of the m-th Gaussian language feature be . Let M be the linguistic feature opacity of the j-th Gaussian, and M be the set of all 3D Gaussians for the current pixel.

[0039] Optionally, S4, using the aforementioned dual-opacity dynamic semantic Gaussian model, a deformation field based on time difference is constructed to obtain the dynamic semantic scene reconstruction result, including:

[0040] Based on the aforementioned dual-opacity dynamic semantic Gaussian model, a deformation field based on time difference is constructed, and key time features are maintained.

[0041] Based on the key time features, an interpolator is used to obtain non-key time features;

[0042] The non-critical temporal features are input into the Gaussian attribute residual decoder and combined with the normalized space 3D Gaussian attributes to obtain the temporal dynamic Gaussian attribute residual as the dynamic semantic scene reconstruction result.

[0043] The process of maintaining key temporal features also includes: reordering the key temporal features into a 2D feature image, compressing it into a feature video using a video codec, and introducing a temporal regularization term to smooth feature changes between adjacent keyframes.

[0044] Optionally, the method further includes:

[0045] The parameters of the converged double-opacity dynamic semantic Gaussian model of the first image group are used as the initialization of the second image group to obtain the normal space of the second image group, wherein the first image group is the previous frame image group of the second image group.

[0046] Based on the canonical space of the second image group, a voxel grid is constructed;

[0047] Based on the canonical space of the second image group, trilinear interpolation is performed using the voxel grid to obtain spatial features;

[0048] The spatial features are input into a multilayer perceptron for decoding to obtain the dynamic and static Gaussian attribute residuals of the image group.

[0049] Optionally, the formula for calculating the dynamic and static Gaussian attribute residuals of the image group is:

[0050]

[0051] in, The property residuals of a static Gaussian. For the property residuals of a dynamic Gaussian, For a multilayer perceptron, spatial features f g =Sigmoid(interp(X,V)), For activation function, is the trilinear interpolation operation function, X is the position coordinate of the Gaussian point in three-dimensional space, and V is the voxel mesh.

[0052] Compared with the closest existing technology, the present invention has the following advantages:

[0053] This invention first distills the semantic features of multi-view images into 3D Gaussian semantic attributes to support editing interaction; secondly, it proposes a dual-opacity dynamic semantic Gaussian representation to jointly train the color and semantic attributes of the 3D Gaussian; finally, it proposes a deformation field based on time interpolation to model realistic dynamic scenes while minimizing the amount of data. This method improves the accuracy of open semantic queries by 9%, while reducing the data size to 43KB and achieving almost no loss in reconstruction quality.

[0054] This invention can reconstruct dynamic real-world scenes into dynamic digital twin systems, enabling the digital twins to perform downstream tasks such as editing, interaction, and spatial intelligence. This invention proposes a dual-opacity dynamic semantic Gaussian representation to jointly integrate Gaussian color and semantic attributes, avoiding gradient conflicts between color and semantics during joint training, thereby reducing training and storage costs. This invention proposes a deformation field based on time interpolation, which maintains only key temporal features, while non-key temporal features are obtained through interpolation, effectively reducing temporal redundancy. This invention proposes a per-image-group static-dynamic Gaussian compensation strategy to solve the scene flickering phenomenon across image groups. Attached Figure Description

[0055] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0056] Figure 1 This is a flowchart of a dynamic semantic scene reconstruction method for industrial digital twins according to an embodiment of the present invention;

[0057] Figure 2 This is a diagram of the dynamic semantic Gaussian reconstruction architecture proposed in an embodiment of the present invention. Detailed Implementation

[0058] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions in the embodiments of this invention will be clearly and completely described below with reference to specific embodiments and corresponding drawings. Obviously, the described embodiments are only a part of the embodiments of this invention, and not all of them. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.

[0059] The terminology used in the embodiments section of this invention is for the purpose of explaining specific embodiments of the invention only, and is not intended to limit the invention.

[0060] This invention proposes a dynamic semantic scene reconstruction method for industrial digital twins, aiming to enable dynamic digital twin systems to be editable, interactive, dynamically updated, and capable of knowledge reasoning. This allows dynamic digital twin systems to not only perform simulation functions but also support downstream tasks such as spatial intelligence and autonomous robot navigation. First, the dynamic scene is decomposed into static and dynamic elements to reduce unnecessary Gaussian updates and achieve incremental updates of the digital twin system. To support editing, interaction, and knowledge reasoning, a Segment Anything Model (SAM) is used to segment the captured multi-view images. Then, a Contrastive Language-Image Pre-training (CLIP) model is used to extract the semantic features of the target, resulting in a final semantic map used to supervise the Gaussian semantic rendering result. Furthermore, a dual-opacity dynamic semantic Gaussian representation is proposed to jointly represent Gaussian color and semantic attributes to avoid gradient conflicts between color and semantics during joint training.

[0061] Each 3D Gaussian has 59-dimensional features and the amount of data ranges from hundreds of thousands to millions. Dynamic updates require the transmission of a large number of Gaussian parameters, typically requiring bandwidth of up to 3Gbps. To address this, a deformation field based on time interpolation is proposed. This field maintains only key time features, while non-key time features are obtained through interpolation, thereby eliminating temporal redundancy and significantly reducing the amount of data that needs to be transmitted.

[0062] like Figure 1 As shown, this embodiment of the invention provides a dynamic semantic scene reconstruction method for industrial digital twins, including:

[0063] S1. Acquire continuous video frames from multiple cameras and perform dynamic and static Gaussian decomposition to obtain dynamic and static Gaussian sets;

[0064] A 3D Gaussian set is a collection of 3D Gaussian primitives that possess geometric and visual attributes such as position, size, rotation, opacity, and color, and can also be supplemented with dynamic indicator attributes and linguistic semantic feature vectors. It is the core data carrier for representing the three-dimensional structure of industrial scenes in 3DGS reconstruction technology. It can also be decomposed into static Gaussian sets and dynamic Gaussian sets according to the characteristics of scene elements, corresponding to static and dynamic elements in industrial scenes, respectively.

[0065] This step involves deploying a multi-camera array to acquire continuous video frames covering the target scene. Using pixel temporal variance analysis as the core, and combining it with binary cross-entropy loss optimization based on learnable dynamic indicator attributes, the 3D Gaussian set is precisely decomposed into a static Gaussian set and a dynamic Gaussian set. The static Gaussian set corresponds to fixed facilities in the scene with stable positions and shapes (such as factory buildings and equipment bases). Its geometric and attribute parameters are fixed during training to reduce unnecessary computation. The dynamic Gaussian set corresponds to objects in the scene that are moving or undergoing state changes (such as production line workpieces and mobile robots). Subsequent update operations are performed only on the dynamic Gaussian set, achieving static and dynamic hierarchical modeling of the industrial scene and laying the data foundation for incremental dynamic updates and semantic fusion.

[0066] S2. Based on the continuous video frames from the multi-camera system and the dynamic and static Gaussian set, multi-level semantic feature extraction is performed using SAM and CLIP models to obtain a multi-dimensional semantic supervision graph.

[0067] To train the Gaussian language features F, this step uses SAM and CLIP models to extract semantic maps from multi-view images to supervise the rendering results of language features F. Specifically, using continuous video frames from multiple cameras and dynamic / static Gaussian sets as input, SAM performs coarse, medium, and fine-grained scene segmentation on the images to accurately distinguish different levels of industrial objects, such as the entire production line, a single piece of equipment, and equipment components. The three-level segmentation results are then input into the CLIP model to extract corresponding semantic features, and the images are stitched together according to pixel coordinates to generate a multi-dimensional semantic supervision map. This supervision map corresponds one-to-one with the pixels of the original image, providing accurate and multi-granular supervision for subsequent training of the semantic attributes of 3D Gaussians. This achieves deep integration of geometric structure and semantic information, endowing the digital twin with the core capability of semantic interaction.

[0068] S3. Based on the dynamic and static Gaussian set and the multidimensional semantic supervision graph, obtain a dynamic semantic Gaussian model with double opacity;

[0069] This step is based on a dynamic and static Gaussian set. Each Gaussian is assigned a 9-dimensional language feature vector consisting of cascaded coarse, medium, and fine-grained 3-dimensional sub-feature channels. While sharing the geometric attributes of Gaussian position, rotation, and size, independent and learnable opacity parameters are configured for the color and language features of the Gaussian. A dual-opacity representation is constructed to form two independent Alpha-Blending rendering paths, achieving complete decoupling of the color and semantic rendering processes and avoiding gradient conflicts during joint training. At the same time, a multi-dimensional semantic supervision map is used as a supervision signal, and an L1 loss function is used to construct a language feature supervision loss. The color and semantic attributes of the Gaussian are jointly trained, ultimately resulting in a dual-opacity dynamic semantic Gaussian model that combines high-fidelity visual reconstruction with accurate semantic expression capabilities. This provides core model support for subsequent dynamic scene rendering, editing interaction, and semantic querying.

[0070] S4. Using the aforementioned dual-opacity dynamic semantic Gaussian model, construct a deformation field based on time difference to obtain the dynamic semantic scene reconstruction results;

[0071] This step involves constructing a deformation field based on time interpolation for the obtained dual-opacity dynamic semantic Gaussian model, maintaining only the set of key time features, and significantly compressing the amount of original time data. For any non-key timestamp, corresponding time features are generated through linear interpolation, and dynamic Gaussian attribute residuals are obtained by decoding the deformation field. The Gaussian attributes in the normal space are summed with the residuals to generate dynamic semantic Gaussian attributes for any timestamp as the result of dynamic semantic scene reconstruction, thereby achieving real-time and high-fidelity reconstruction of industrial dynamic scenes.

[0072] In summary, steps S1 to S4 are geared towards industrial digital twin scenarios. First, 3D Gaussian static and dynamic hierarchical decomposition is performed using continuous video frames from multiple cameras to obtain static and dynamic Gaussian sets, providing a foundation for incremental modeling. Then, SAM and CLIP models are combined to extract multi-level semantic features from the video frames, generating a multi-dimensional semantic supervision graph to fuse geometric and semantic information. Subsequently, a dual-opacity dynamic semantic Gaussian model is constructed based on the static and dynamic Gaussian sets and the semantic supervision graph. By decoupling color and semantic rendering paths through independent opacity, the joint training gradient conflict is eliminated. Finally, a time-interpolated deformation field is built based on this model, maintaining only key time features and generating dynamic Gaussian attributes at any time through interpolation. While eliminating time redundancy and compressing data volume, it outputs editable, interactive, and low-latency dynamic semantic scene reconstruction results, providing a complete technical solution for the real-time construction, updating, and application of industrial digital twins.

[0073] Real-world videos contain numerous static elements, and representing these static elements with a dynamic Gaussian would incur unnecessary computational and storage costs. Therefore, a static-dynamic decomposition method is employed, using a static Gaussian to reconstruct static elements and a dynamic Gaussian to reconstruct dynamic elements. As a possible implementation, in the above embodiment, step S1 may specifically include the following steps:

[0074] S1-1. Acquire continuous video frames from multiple cameras and calculate the pixel temporal variance of the continuous video frames from multiple cameras.

[0075] A multi-camera array is deployed in the target industrial scene to ensure that each camera's viewpoint covers the scene without blind spots, capturing continuous video frames simultaneously from multiple cameras in the industrial area. First, the temporal variance of pixels in all video frames of each camera is calculated, i.e., the pixel temporal variance of continuous video frames from multiple cameras, which is used to quantify the magnitude of pixel change over time and distinguish between static and dynamic areas.

[0076] S1-2. Based on the pixel time variance of the continuous video frames of the multi-camera system and a preset threshold, obtain dynamic pixel labels and static pixel labels, and then obtain multi-camera dynamic labels.

[0077] Set preset threshold (Default is 0.03), the temporal variance of each pixel in each camera is compared with a preset threshold. Comparison:

[0078] If the pixel time variance is greater than Marked as dynamic pixels;

[0079] If the pixel time variance is less than or equal to , marked as static pixels.

[0080] Based on the single-camera dynamic-static pixel labels obtained above, the label results of all cameras are summarized to obtain the dynamic labels M of all cameras, i.e., the multi-camera dynamic labels, which serve as the ground truth for subsequent prediction of the dynamic map.

[0081] S1-3. Based on the multi-camera dynamic labels, learnable dynamic indication attributes are introduced to obtain the predicted dynamic map;

[0082] To enable the 3D Gaussian to learn the prior information of the dynamic regions corresponding to the dynamic labels M of all cameras, a learnable dynamic indicator attribute d is introduced into each 3D Gaussian, and this attribute replaces the original color attribute c of the Gaussian, thereby calculating the predicted dynamic map. The calculation formula is as follows:

[0083]

[0084] Where x is the predicted dynamic graph pixels, This is the predicted dynamic graph for pixel x. Let d be a sigmoid, single-valued, continuously differentiable nonlinear activation function, m be the m-th 3D Gaussian currently being computed, N be the set of 3D Gaussians covering the pixels, and d be the sigmoid activation function. m For the dynamic indicator property of the m-th 3D Gaussian, o m Let o be the opacity property of the m-th 3D Gaussian. m (x) represents the opacity attribute of the m-th 3D Gaussian at pixel x. j (x) represents the opacity attribute of the j-th 3D Gaussian at pixel x, where j is the index of all Gaussians that have been rendered before the m-th 3D Gaussian, and the value range is j=1,2,…,m-1.

[0085] S1-4. Use the multi-camera dynamic labels to supervise the predicted dynamic map, and use binary cross-entropy loss to optimize the dynamic indicator attributes to obtain the optimized dynamic indicator attributes.

[0086] Dynamic labels M can be used for supervised prediction of dynamic graphs. During training, all attributes of the 3D Gaussian model except for the dynamic indicator attribute d are fixed, and the dynamic labels M of all cameras are used as the supervised ground truth to constrain the predicted dynamic graph. The binary cross-entropy loss is used to optimize the dynamic indicator attribute d, making the prediction of the dynamic graph more accurate. The pixel values ​​are close to 0 or 1, thus optimizing the distribution of the dynamic indicator attribute d in the range of (-∞, +∞). The binary cross-entropy loss L is defined as follows:

[0087]

[0088] in, Let M(x) be the mathematical expectation function, and M(x) be the dynamic label of pixel x.

[0089] S1-5. Obtain a dynamic and static Gaussian set based on the optimized dynamic indicator attribute and the preset threshold.

[0090] After optimizing and obtaining the dynamic indicator properties, the method proposed by Swift4D is used to decompose the complete 3D Gaussian set into a static Gaussian set and a dynamic Gaussian set, using a preset threshold. As the basis for judgment, traverse all 3D Gaussians:

[0091] If the optimized dynamic indicator attribute exceeds the preset threshold It is classified as a dynamic Gaussian set G. d ;

[0092] If the optimized dynamic indicator attribute does not exceed the preset threshold It is classified as a static Gaussian set G. s .

[0093] Finally, we obtain the dynamic Gaussian set G. d With static Gaussian set G s The dynamic Gaussian set is composed of static elements and dynamic elements, which can be reconstructed using static Gaussian and dynamic Gaussian, thus avoiding the additional computation and storage costs caused by static elements occupying dynamic Gaussian.

[0094] In summary, the dynamic and static Gaussian decomposition in steps S1-1 to S1-5 aims to solve the problem of unnecessary computation and storage costs caused by reconstructing static elements with dynamic Gaussian in industrial scenarios. By first calculating the temporal variance of each pixel in continuous video frames from multiple cameras, dividing static and dynamic pixels according to a preset threshold and obtaining dynamic labels for the entire camera, then introducing learnable dynamic indicator attributes for each 3D Gaussian and calculating the predicted dynamic map in combination with opacity, and optimizing the dynamic indicator attributes with binary cross-entropy loss while keeping other Gaussian attributes fixed, the 3D Gaussian is finally divided into a set of static Gaussians and a set of dynamic Gaussians according to the threshold, thus achieving the goal of reconstructing static elements with static Gaussians and reconstructing dynamic elements with dynamic Gaussians.

[0095] As one possible implementation, in the above embodiments, step S2 may specifically include the following steps:

[0096] S2-1. Use the dynamic and static Gaussian sets to determine the scene space range;

[0097] The dynamic Gaussian set G obtained from dynamic and static Gaussian decomposition d With static Gaussian set G s As a spatial reference, the dynamic Gaussian set G d Represented as G d =(X d ,S d ,R d O d C d ), static Gaussian set G s Represented as G s =(X s ,S s ,R s O s C s X, S, R, O, and C represent the five attributes of the Gaussian set: position, size, rotation, opacity, and color, respectively. This is achieved by statistically analyzing the dynamic Gaussian set G. d With static Gaussian set G s The effective spatial range of the industrial scene is determined by the 3D coordinate range of all Gaussians. This range defines the effective area for semantic extraction, model rendering, and feature supervision, excludes invalid background areas, ensures strict alignment of semantic features with the 3D Gaussian space, and avoids interference from irrelevant areas during training.

[0098] S2-2. Divide the continuous video frames of the multi-camera system based on the scene space range to obtain the video frames of the image group;

[0099] Within a defined scene space, consecutive video frames from multiple cameras are grouped by camera, forming image groups based on each camera. Each image group T contains G consecutive temporal frames, denoted as T={t0,t1,…,t…}. g}.like Figure 2 As shown, the dynamic scene in the image group is composed of a static 3D Gaussian set G. s =(X s ,S s ,R s O s C s Dynamic 3D Gaussian set G d =(X d ,S d ,R d O d C d It consists of a deformation field based on time interpolation. When grouping, only the effective pixel area falling within the scene space is retained, and redundant frames and invalid pixels that exceed the range are removed to ensure that the content of the image group is consistent with the 3D Gaussian reconstruction space, thus achieving temporal and spatial alignment.

[0100] S2-3. Input the video frames of the image group into the SAM model for segmentation processing, and obtain the segmentation result output by the SAM model;

[0101] Each consecutive video frame within the image group is input into SAM, and three-level granular pixel-level segmentation is performed on the corresponding scene spatial range within each frame to obtain coarse-grained segmented images. Medium-grained image segmentation fine-grained image segmentation Specifically:

[0102] Coarse-grained segmentation: Set a low segmentation threshold (e.g., 0.2) to extract only large, highly salient global objects in the scene (e.g., the main body of the factory, the entire production line, and the base of large equipment) to generate a coarse-grained segmented image, corresponding to the macroscopic semantic level of the industrial scene;

[0103] Medium-grained segmentation: Set a medium segmentation threshold (e.g., 0.5) to further break down medium-sized objects such as single industrial equipment, production line workstations, and mobile robots based on coarse-grained segmentation, generating medium-grained segmented images that correspond to the medium semantic level of the industrial scene.

[0104] Fine-grained segmentation: Set a high segmentation threshold (e.g., 0.8) to further break down a single device that has been segmented in a medium-grained manner into small detail objects such as device components (e.g., robotic arm joints, sensors, valves, workpieces) to generate a fine-grained segmented image that corresponds to the micro-semantic level of the industrial scene.

[0105] The SAM model performs classless general segmentation of industrial targets such as equipment, pipes, goods, robots, and workbenches within a frame, and outputs pixel-level masks to locate interactive and editable semantic entities within the scene, providing accurate regions for subsequent CLIP semantic extraction.

[0106] S2-4. Input the segmentation result output by the SAM model into the CLIP model to obtain the semantic map output by the CLIP model;

[0107] The coarse, medium, and fine-grained segmented images output by SAM are used as region constraints input to the CLIP model to extract linguistic semantic features from each segmented region:

[0108] For coarse-grained segmented images, global scene-level semantic features (i.e., coarse-grained semantic features) are extracted to represent the semantic attributes of macroscopic objects such as factories and production lines, and a coarse-grained semantic map is output. ;

[0109] For medium-granularity segmented images, extract device-level semantic features (i.e., medium-granularity semantic features) to represent the semantic attributes of medium-level objects such as individual devices and workstations, and output a medium-granularity semantic map. ;

[0110] For fine-grained segmented images, component-level semantic features (i.e., fine-grained semantic features) are extracted to characterize the semantic attributes of micro-objects such as equipment components and workpieces, and a fine-grained semantic map is output. .

[0111] Each semantic graph has a dimension of 1. Where M and N are the image height and width, and the 3D corresponds to the semantic feature vector output by CLIP. This step realizes the transformation from pixel-level mask to understandable semantic features, enabling the 3D Gaussian to obtain editable and queryable language attributes.

[0112] S2-5. Perform tensor concatenation on the semantic graph output by the CLIP model to obtain a multidimensional semantic supervision graph;

[0113] For each frame, the coarse, medium, and fine semantic maps are concatenated using tensors aligned to the pixel dimensions along the feature channel dimension, resulting in a single semantic map with dimension [missing information]. The multidimensional semantic supervision graph is used to supervise the rendering results of language features F. This supervision graph is directly used to supervise the rendering and training of subsequent 3D Gaussian semantic features, so that each Gaussian feature has three levels of semantic attributes: coarse, medium, and fine, and supports open vocabulary query, editing, and interaction.

[0114] In summary, steps S2-1 to S2-5 determine the scene spatial range by using a set of static and dynamic Gaussian elements. This clarifies that the dynamic scene within an image group is composed of static Gaussian elements, dynamic Gaussian elements, and a temporal interpolation deformation field. Based on this spatial range, continuous video frames from multiple cameras are divided into temporal image groups according to a single viewpoint. These image groups are then fed into a Semantic Image Processing (SAM) module to obtain coarse, medium, and fine segmentation results. The CLIP model extracts the corresponding three-level semantic maps, and finally, the multi-dimensional semantic supervision map is obtained by stitching the aligned pixel tensors along the channel dimension. This process precisely limits the semantic extraction space, ensures Gaussian attributes are aligned with the temporal sequence of the images, and achieves efficient fusion of pixel-level segmentation and multi-level semantic features. It provides strong supervision signals for subsequent training with double-opacity Gaussian elements, significantly improving semantic rendering accuracy and scene alignment.

[0115] As one possible implementation, in the above embodiments, step S3 may specifically include the following steps:

[0116] S3-1. Perform Gaussian allocation on the dynamic and static Gaussian sets to obtain language feature vectors;

[0117] Instead of a multi-model architecture, a high-dimensional language feature vector is assigned to each dynamic Gaussian and static Gaussian in the scene. and It is used to uniformly represent language features at three levels: coarse-grained, medium-grained, and fine-grained. The language feature vector is a dynamic Gaussian vector. R is a static Gaussian language feature vector. 9 Let F be a 9-dimensional real space. The vector is structurally partitioned along the feature channel dimension, assuming F ∈ R. 9 Let be an arbitrary Gaussian language feature vector, which is defined as a channel concatenation of three sub-feature vectors:

[0118]

[0119] Where F is the language feature vector, f small For coarse-grained linguistic semantic features, f middle For medium-granularity linguistic semantic features, f large For fine-grained linguistic semantic features. The first three dimensions f of the vector. small ∈R 3 The three middle dimensions f middle ∈R 3 And the last three dimensions f large ∈R 3 These correspond to coarse-grained, medium-grained, and fine-grained linguistic semantic features, respectively.

[0120] This compact language feature design leads to performance improvements. During the rendering stage, it allows the three levels of language features to be treated as additional attributes of the same Gaussian unit, requiring only one Gaussian splashing operation to project F onto the two-dimensional image plane. Subsequently, multi-level semantic information can be recovered by slicing the feature map in screen space. This mechanism successfully avoids repetitive rasterization calculations, achieving a significant increase in frames per second (FPS) and optimization of data volume while maintaining the feature representation.

[0121] S3-2. Based on the language feature vector and the opacity parameter, construct a double-opacity Gaussian representation and perform path rendering to obtain the language feature rendering result;

[0122] In constructing dynamic scene reconstructions using language embeddings, a strategy of jointly optimizing high-dimensional language features and color attributes was first explored to reduce training time. However, experimental analysis shows that directly jointly training language features and color faces significant challenges, stemming from their fundamental differences in physical properties: color is typically modeled using spherical harmonics to capture anisotropic viewpoint-dependent effects (such as specular highlights and reflections), while language features are essentially isotropic viewpoint-independent representations. When these two signals with different geometric dependencies are trained together, gradient conflicts due to shared Gaussian properties lead to severe performance degradation, significantly reducing the accuracy of open-vocabulary queries and compromising the fidelity of color rendering.

[0123] Further experiments revealed that opacity is a key factor in rendering color and speech features. Therefore, to address this issue, a dual-opacity dynamic Gaussian representation of speech is proposed. This representation introduces independent, learnable opacity parameters o for color and speech features while sharing geometric properties such as position, rotation, and size. c ∈O c and o l ∈O l , where o c For color feature opacity, O c Let o be the set of color feature opacities. l For the opacity of language features, O l This represents the set of language feature opacities. This design decouples the rendering path during the rasterization alpha-blending stage, completing path rendering according to a predetermined rendering formula, and ultimately obtaining the language feature rendering result, i.e., the rendered language feature map R. l The specific path rendering definition is as follows:

[0124]

[0125]

[0126] Among them, R c For the rendered color feature map, c m For the m-th Gaussian color feature, Let the opacity be the color feature of the m-th Gaussian. Let R be the opacity of the j-th Gaussian color feature. l For rendering the language feature map, f m For the m-th Gaussian linguistic feature, Let the opacity of the m-th Gaussian language feature be . Let M be the linguistic feature opacity of the j-th Gaussian, and M be the set of all 3D Gaussians for the current pixel.

[0127] S3-3. Based on the rendering results of the language features, the L1 loss function is used to construct a language feature supervised loss;

[0128] To achieve accurate training of Gaussian-progressive language features F, a rendering-based language feature map R is used. l The language feature supervision loss L is constructed using the L1 loss function. l This is used to measure the error between rendered semantic features and real semantic labels, constraining language features to converge towards real labels, ensuring training stability, and suppressing noise. Language feature supervision loss L l It can be represented as:

[0129]

[0130] Where L1 is the loss function, This is a tensor concatenation operation, which concatenates the language feature labels from three levels together.

[0131] S3-4. Supervise the language feature rendering results through the multidimensional semantic supervision graph, and obtain a dual-opacity dynamic semantic Gaussian model by combining the language feature supervision loss.

[0132] The SAM and CLIP models are used to process multi-view images, extracting coarse, medium, and fine-grained semantic maps and stitching them together to form a multi-dimensional semantic supervision map. This map is then used as the ground truth label to process the rendered language feature map R. l End-to-end supervision is performed. A multi-dimensional semantic supervision graph is used as the supervision signal, combined with the constructed language feature supervision loss L. l Joint training was conducted to iteratively optimize the color attributes, language feature vectors, and double opacity parameters of the Gaussian model, ultimately resulting in a dynamic semantic Gaussian model with double opacity that supports editing interaction, query reasoning, and gradient conflict-free operation.

[0133] In summary, steps S3-1 to S3-4 abandon the multi-model architecture and assign a 9-dimensional high-dimensional language feature vector to each dynamic and static Gaussian in the scene. This vector is then structured in the channel dimension into a cascaded form of sub-features corresponding to coarse, medium, and fine semantic levels. This compact design can complete the two-dimensional projection of multi-level language features through a single Gaussian sputtering, and then recover semantic information through feature map slicing, effectively avoiding repeated rasterization calculations. While maintaining feature expressive power, it significantly improves frame rate and optimizes data volume. Furthermore, it addresses the issue of gradient conflicts and performance degradation that easily occur when color and language features are jointly trained. Based on shared geometric attributes, independent and learnable opacity parameters are set for color and language features respectively, and a dual-opacity dynamic Gaussian representation of language is constructed. In the Alpha-Blending stage, the two rendering paths are decoupled, which not only alleviates the performance degradation problem, but also ensures the accuracy of open vocabulary query and the fidelity of color rendering. At the same time, with the help of multi-dimensional semantic supervision graph, combined with the L1 loss function, the rendering results of Gaussian language features are supervised and trained. Finally, a dual-opacity dynamic semantic Gaussian model that can be stably optimized and has both efficient rendering and accurate semantic representation capabilities is obtained.

[0134] As one possible implementation, in the above embodiments, step S4 may specifically include the following steps:

[0135] S4-1. Based on the aforementioned dual-opacity dynamic semantic Gaussian model, construct a deformation field based on time difference and maintain key time features;

[0136] Based on the existing dual-opacity dynamic semantic Gaussian model, to address the issues of large 3D Gaussian data volume, high temporal redundancy, and high network transmission pressure in industrial dynamic scenes, a deformation field based on temporal interpolation is constructed. This deformation field uses the total number of frames T within a group of images (GOP) as a basis, and performs key temporal feature sampling according to a set keyframe sampling interval n, maintaining the key temporal features K={k0,k1,…,k T / n (T is the total number of frames in the GOP), reducing the amount of temporal feature data T to T / n, thus eliminating temporal redundancy at its source.

[0137] S4-2. Based on the key time features, use the interpolator to obtain non-key time features;

[0138] For any given timestamp t in the deformation field, first perform a floor operation. Locate its two adjacent keyframes k on the timeline. i and k i+1 Then, based on the key time features, an interpolator is invoked to obtain non-key time features. This embodiment uses a linear interpolator to obtain the non-key time feature k. t The calculation formula is:

[0139]

[0140] in, For linear interpolation function, k i For the i-th keyframe, k i+1 Let w be the (i+1)th keyframe, w be the time difference weight, and i be the keyframe index. The time difference weight w is the interpolation weight of the current timestamp t relative to two adjacent keyframes. Using this linear interpolation method, full-time feature information can be obtained without separately storing non-key time features. The calculation formula is:

[0141]

[0142] Where t is the time index and n is the keyframe sampling interval.

[0143] S4-3. Input the non-critical temporal features into the Gaussian attribute residual decoder, and combine them with the normalized space 3D Gaussian attributes to obtain the temporal dynamic Gaussian attribute residual as the dynamic semantic scene reconstruction result.

[0144] Obtain non-critical time features k t Afterwards, the Gaussian attribute residual decoder D a With k t As input, the temporal features are decoded into 3D Gaussian attribute residuals. The 3D Gaussian attribute residual obtained from this decoding is the temporal dynamic Gaussian attribute residual, which directly serves as the final result of dynamic semantic scene reconstruction. This temporal dynamic Gaussian attribute residual is compared with the 3D Gaussian attribute a in the norm space. d Superposition yields the dynamic Gaussian properties at time t. This provides temporal constraints on the scene's spatial properties for the deformation field, enabling the mapping from temporal features to spatial deformation. The dynamic Gaussian properties at time t... The formula for calculation is:

[0145]

[0146] in, , Let t be the dynamic 3D Gaussian attribute residual space. The proposed deformation field based on time feature interpolation is universal and suitable for representation using various Gaussian elements. For example, within the standard 3DGS framework, the dynamic 3D Gaussian attribute space at time t... This includes position, rotation, size, opacity, and spherical harmonics; while in variants such as Scaffold-GS, these correspond to anchor point position, features, size, and offset attributes.

[0147] Furthermore, after maintaining the key temporal features, the process includes: reordering the key temporal features into a 2D feature image, compressing it into a feature video using a video codec, and introducing a temporal regularization term to smooth feature changes between adjacent keyframes. Specifically:

[0148] To compress Gaussian properties and temporal features as much as possible, non-critical temporal features are reordered into a series of 2D feature images. These feature images are then compressed into a feature video using conventional video codecs (such as H.264 and HEVC), significantly reducing redundancy between temporal features. Simultaneously, to make the feature video smaller, a temporal regularization term L is introduced. tsr The calculation formula is:

[0149]

[0150] This loss function forces adjacent keyframes k i and k i+1 The characteristic changes between them remain smooth, ensuring the continuity of temporal features and the compression effect.

[0151] In summary, steps S4-1 to S4-3, based on a dual-opacity dynamic semantic Gaussian model, first construct a deformation field based on time differences and maintain key temporal features. The key temporal features are then rearranged into 2D feature images and compressed into feature videos by a video codec. Simultaneously, a temporal regularization term is introduced to smooth the feature changes of adjacent key frames. Then, non-key temporal features are obtained through a linear interpolator based on the key temporal features. These non-key temporal features are input into the decoder and combined with the normalized 3D Gaussian properties to obtain temporal dynamic Gaussian properties as the result of dynamic semantic scene reconstruction. This approach can achieve efficient compression of temporal features while ensuring the temporal continuity of the deformation field, effectively suppressing scene flicker and jitter, and improving the stability and smoothness of dynamic semantic scene reconstruction.

[0152] As one possible implementation, in the above embodiments, the present invention further includes per-image-group dynamic and static Gaussian compensation, the steps of which are as follows:

[0153] S5-1. Use the converged double-opacity dynamic semantic Gaussian model parameters of the first image group as the initialization of the second image group, and obtain the normal space of the second image group, wherein the first image group is the previous frame image group of the second image group.

[0154] To avoid noticeable flickering between different image groups, the parameters of the final converged double-opacity dynamic semantic Gaussian model of the (g-1)th image group are directly used as the initialization for the g-th image group. The current image group no longer learns the scene distribution from scratch, but focuses on learning incremental updates relative to the previous image group, while using the static Gaussian model converged in the (g-1)th image group. Dynamic Gaussian Based on this, the normalization space of the (g-1)th image group is generated, and finally the static Gaussian space of the (g-1)th image group is... and dynamic Gaussian The learned incremental update is added to the result, and the calculation formula is as follows:

[0155] ,

[0156] Where R is the learned Gaussian attribute incremental update parameter. This step utilizes the temporal continuity of dynamic scenes to improve the model convergence speed, suppress flickering across image groups, and provide an accurate canonical spatial benchmark for subsequent multi-resolution voxel mesh construction.

[0157] S5-2. Construct a voxel grid based on the canonical space of the second image group;

[0158] The normal space generated by the convergence parameters of the (g-1)th image group (i.e. and Using the corresponding spatial range as the spatial benchmark, an encoding scheme based on binary feature voxel grids is introduced to construct a set of multi-resolution voxel grids V. Unlike the traditional storage of high-precision floating-point features, the parameter space of the grid vertices is strictly constrained within the binary set {+1,-1}, which fundamentally compresses the storage overhead of features. At the same time, relying on the multi-resolution structure, it takes into account both the global scene structure and the ability to express local details, providing an efficient encoding carrier for subsequent spatial feature extraction.

[0159] S5-3. Perform trilinear interpolation using the voxel grid according to the normalized space of the second image group to obtain spatial features;

[0160] For the standard space and Taking any position X as a Gaussian point, a trilinear interpolation operation is performed on a multi-level binary feature voxel grid to capture the spatial correlation of Gaussians and obtain the spatial feature f. g The calculation formula is:

[0161] f g =Sigmoid(interp(X,V))

[0162] in, For activation function, Here, X is the trilinear interpolation function, and X represents the coordinates of a Gaussian point in three-dimensional space. This step uses a normalized Gaussian distribution as the query reference and a voxel grid as the feature carrier to convert discrete binary grid features into continuous spatial features, providing multi-scale, robust spatial context information for subsequent attribute decoding.

[0163] S5-4. Input the spatial features into a multilayer perceptron for decoding to obtain the dynamic and static Gaussian attribute residuals of the image group;

[0164] Spatial features f g The input is fed into a small multilayer perceptron (MLP), which maps the highly compressed binary space context back to high-precision continuous attribute residuals, thereby decoding the attribute residuals of dynamic and static Gaussians in the normalized space. and That is, the dynamic and static Gaussian attribute residuals of the image group The calculation formula is:

[0165]

[0166] in, The property residuals of a static Gaussian. For the property residuals of a dynamic Gaussian, It is a multilayer perceptron. The residual can be used to correct the parameters of the initialized Gaussian model, complete the static and dynamic Gaussian completion of the current image group, and ensure the reconstruction accuracy and temporal stability of dynamic semantic scenes while significantly reducing storage and computing costs.

[0167] In summary, the parameters of the dual-opacity dynamic semantic Gaussian model converged in the previous frame image group from steps S5-1 to S5-4 are used as the initialization of the current image group, and a normalized space is generated based on this. On this basis, a multi-resolution voxel grid is constructed. For Gaussian points in the normalized space, trilinear interpolation is performed on the voxel grid to obtain spatial features. Then, the features are input into a multilayer perceptron for decoding to obtain the attribute residuals of static and dynamic Gaussian, thus realizing static and dynamic Gaussian completion for each image group. This process not only utilizes temporal continuity to suppress flickering across image groups and improve convergence speed, but also significantly compresses feature storage overhead through multi-resolution binarized voxel grids. Finally, high-precision attribute residual learning is achieved through interpolation and MLP decoding. While ensuring the quality of dynamic semantic scene reconstruction, it achieves efficient, stable and low-overhead Gaussian model update for each image group.

[0168] like Figure 2 As shown, this invention uses multiple groups of images (GOPs) as the time-phased full-process reconstruction logic: First, in the initial GOP0 stage, dynamic Gaussian attribute residuals are generated based on key temporal features through interpolation deformation fields, and then the static Gaussian set G... s With dynamic Gaussian set G d Merged into a dual-opacity dynamic language Gaussian model, using color feature opacity o c With language feature opacity lThe process involves rendering and outputting color maps and semantic feature maps separately. In subsequent GOP1 and update phases, temporal regularization constraints and binary voxel encoding are introduced. After MLP decoding, static / dynamic Gaussian incremental residuals across GOPs are obtained. The incremental residuals are then fused into the preceding model through a transformation operation T. Simultaneously, the temporal interpolation deformation field is reused to update the dynamic Gaussian attributes, achieving incremental dynamic updates. Finally, after multiple iterations in the GOPg phase, the updated static and dynamic Gaussian sets are merged, outputting an industrial digital twin that supports semantic interaction and dynamic updates. The overall process embodies the core technical logic of static and dynamic hierarchical modeling, temporal redundancy elimination, dual opacity decoupling, and cross-group incremental updates, with each step progressing progressively and forming a logical closed loop.

[0169] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0170] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0171] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0172] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process.Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that modifications or equivalent substitutions can still be made to the specific implementation of the present invention. Any modifications or equivalent substitutions that do not depart from the spirit and scope of the present invention should be covered within the scope of protection of the claims of the present invention.

Claims

1. A dynamic semantic scene reconstruction method for industrial digital twins, characterized in that, include: S1. Acquire continuous video frames from multiple cameras and perform dynamic and static Gaussian decomposition to obtain dynamic and static Gaussian sets, including: Acquire continuous video frames from multiple cameras and calculate the pixel temporal variance of the continuous video frames from multiple cameras; Based on the pixel temporal variance of the continuous video frames from the multi-camera system and a preset threshold, dynamic pixel labels and static pixel labels are obtained, thereby obtaining the multi-camera dynamic labels. Learnable dynamic indication attributes are introduced based on the multi-camera dynamic labels to obtain a predicted dynamic map; The predicted dynamic map is supervised by the multi-camera dynamic labels, and the dynamic indicator attributes are optimized by binary cross-entropy loss to obtain the optimized dynamic indicator attributes. Based on the optimized dynamic indicator attributes and the preset threshold, obtain a dynamic and static Gaussian set; S2. Based on the continuous video frames from the multi-camera system and the dynamic and static Gaussian set, multi-level semantic feature extraction is performed using SAM and CLIP models to obtain a multi-dimensional semantic supervision graph. S3. Based on the dynamic and static Gaussian sets and the multidimensional semantic supervision graph, obtain a dual-opacity dynamic semantic Gaussian model, including: Gaussian assignment is performed on the dynamic and static Gaussian sets to obtain language feature vectors; Based on the language feature vector and the opacity parameter, a double-opacity Gaussian representation is constructed, and path rendering is performed to obtain the language feature rendering result; Based on the rendering results of the language features, an L1 loss function is used to construct a language feature supervised loss. The rendering results of the language features are supervised by the multidimensional semantic supervision graph, and a dynamic semantic Gaussian model with double opacity is obtained by combining the language feature supervision loss. S4. Using the aforementioned dual-opacity dynamic semantic Gaussian model, construct a deformation field based on time difference to obtain the dynamic semantic scene reconstruction results; The method further includes: The parameters of the converged double-opacity dynamic semantic Gaussian model of the first image group are used as the initialization of the second image group to obtain the normal space of the second image group, wherein the first image group is the previous frame image group of the second image group. Based on the canonical space of the second image group, a voxel grid is constructed; Based on the canonical space of the second image group, trilinear interpolation is performed using the voxel grid to obtain spatial features; The spatial features are input into a multilayer perceptron for decoding to obtain the dynamic and static Gaussian attribute residuals of the image group. The formula for calculating the dynamic and static Gaussian attribute residuals of the image group is: in, The property residuals of a static Gaussian. For the property residuals of a dynamic Gaussian, For a multilayer perceptron, spatial features f g =Sigmoid(interp(X,V)), For activation function, is the trilinear interpolation operation function, X is the position coordinate of the Gaussian point in three-dimensional space, and V is the voxel mesh.

2. The dynamic semantic scene reconstruction method for industrial digital twins according to claim 1, characterized in that, The formula for calculating the predicted dynamic graph is: in, Let x be the predicted dynamic graph. pixels, This is the predicted dynamic graph for pixel x. Let d be a sigmoid, single-valued, continuously differentiable nonlinear activation function, m be the m-th 3D Gaussian currently being computed, N be the set of 3D Gaussians covering the pixels, and d be the sigmoid activation function. m For the dynamic indicator property of the m-th 3D Gaussian, o m Let o be the opacity property of the m-th 3D Gaussian. m (x) represents the opacity attribute of the m-th 3D Gaussian at pixel x. j (x) represents the opacity attribute of the j-th 3D Gaussian at pixel x, where j is the index of all Gaussians that have been rendered before the m-th 3D Gaussian.

3. The dynamic semantic scene reconstruction method for industrial digital twins according to claim 1, characterized in that, The formula for calculating the binary cross-entropy loss is: Where L is the binary cross-entropy loss, Let M(x) be the mathematical expectation function, and M(x) be the dynamic label of pixel x. This is a predicted dynamic graph for pixel x.

4. The dynamic semantic scene reconstruction method for industrial digital twins according to claim 1, characterized in that, S2. Based on the continuous video frames from the multi-camera setup and the dynamic / static Gaussian set, multi-level semantic feature extraction is performed using SAM and CLIP models to obtain a multi-dimensional semantic supervision map, including: The scene spatial range is determined using the aforementioned dynamic and static Gaussian sets; The continuous video frames from the multi-camera system are divided based on the scene space range to obtain video frames of the image group. The video frames of the image group are input into the SAM model for segmentation processing, and the segmentation results output by the SAM model are obtained. The segmentation result output by the SAM model is input into the CLIP model to obtain the semantic graph output by the CLIP model; Tensor concatenation is performed on the semantic graph output by the CLIP model to obtain a multidimensional semantic supervision graph.

5. The dynamic semantic scene reconstruction method for industrial digital twins according to claim 1, characterized in that, The calculation formula for path rendering is: Among them, R c For the rendered color feature map, c m For the m-th Gaussian color feature, Let the opacity be the color feature of the m-th Gaussian. Let R be the opacity of the j-th Gaussian color feature. l For rendering the language feature map, f m For the m-th Gaussian linguistic feature, Let the opacity of the m-th Gaussian language feature be . Let M be the linguistic feature opacity of the j-th Gaussian, and M be the set of all 3D Gaussians for the current pixel.

6. The dynamic semantic scene reconstruction method for industrial digital twins according to claim 1, characterized in that, S4. Using the aforementioned dual-opacity dynamic semantic Gaussian model, construct a deformation field based on time difference to obtain the dynamic semantic scene reconstruction results, including: Based on the aforementioned dual-opacity dynamic semantic Gaussian model, a deformation field based on time difference is constructed, and key time features are maintained. Based on the key time features, an interpolator is used to obtain non-key time features; The non-critical temporal features are input into the Gaussian attribute residual decoder and combined with the normalized space 3D Gaussian attributes to obtain the temporal dynamic Gaussian attribute residual as the dynamic semantic scene reconstruction result. The process of maintaining key temporal features also includes: reordering the key temporal features into a 2D feature image, compressing it into a feature video using a video codec, and introducing a temporal regularization term to smooth feature changes between adjacent keyframes.