This invention relates to a multimodal spatiotemporal alignment and interactive
reconstruction method for digital twins of intangible cultural heritage skills, belonging to the interdisciplinary field of
digital protection of intangible cultural heritage and
computer graphics. It simultaneously acquires five modalities: optical and
inertial motion capture,
hyperspectral imaging, panoramic
sound field, multi-view images, and physiological data of inheritors. Temporal alignment is achieved through cubic spline interpolation and
Gaussian mixture model expectation-maximization
algorithm, and spatial alignment is completed using
point cloud registration. A dynamic graph spatiotemporal
interaction network is used to generate cross-
modal fusion features. Based on this, a three-level digital twin is constructed: geometric, technological, and knowledge-based. The technological twin extracts temporal evolution features based on
Transformer, while the knowledge twin expresses causal transmission relationships using process hyperedges. Finally, interactive experiences are provided through dynamic geometric optimization, multimodal immersive rendering, and gesture / controller interaction, and user behavior feedback is used for
incremental learning and
knowledge graph updates, forming a self-evolving
closed loop.