Non-genetic skill motion capture and digital reproduction method based on deep learning

By dynamically adjusting the weights of multimodal data fusion using deep learning technology, a three-level motion feature map is constructed and semantically labeled, solving the feature conflict and hand-tool collaborative modeling problems in the digital reproduction of intangible cultural heritage skills, and achieving high-precision, natural and coherent motion reproduction.

CN121808327APending Publication Date: 2026-04-07SHANDONG POLYTECHNIC COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-04-07

AI Technical Summary

Technical Problem

Existing digital reproduction technologies for intangible cultural heritage skills suffer from feature conflicts and weight mismatches when fusing multimodal data. They cannot take into account cross-scale information such as macroscopic posture and microscopic jitter, and the hand-tool collaborative modeling is insufficient, resulting in inaccurate reproduction and physical inconsistencies.

Method used

By employing a deep learning-based approach, raw motion data is acquired through multimodal sensors, the fusion weights are dynamically adjusted, a three-level motion feature map is constructed, key operation points of the skill are identified, and semantic annotation is performed in conjunction with a knowledge graph, thereby achieving detailed motion reconstruction and tool-hand collaborative modeling.

Benefits of technology

It improves the fidelity of digital reproduction, ensures comprehensive information contribution, eliminates feature distortion and physical inconsistencies, and supports cultural identity and skill dissemination in the transmission of intangible cultural heritage.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121808327A_ABST
    Figure CN121808327A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of digital cultural heritage protection, and discloses a deep learning-based non-abandoned skill motion capture and digital reproduction method, which comprises the following steps of: obtaining original motion data and carrying out multi-modal fusion on the original motion data to obtain fused motion data; constructing a three-level action characteristic spectrum based on the fused action data, and constructing a fine action key frame sequence for a characteristic sequence in a fine operation layer in the three-level action characteristic spectrum; the key frame sequence is optimized in combination with time continuity and space coordination constraint, and enhanced action feature representation is obtained; semantic annotation and digital action reconstruction are carried out on the enhanced action feature representation; according to the method, the fidelity of skill details and the interpretability of cultural semantics are remarkably improved, and high-quality technical support is provided for non-abandoned inheritance and digital protection.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of digital cultural heritage protection, and more particularly, to a non-heritage skill motion capture and digital reproduction method based on deep learning. BACKGROUND

[0002] With the in-depth development of digital technology, the protection of intangible cultural heritage is undergoing a transformation from traditional text and image recording to immersive digital reproduction. Traditional intangible cultural heritage relies on oral transmission and physical instruction between teachers and apprentices, and is at risk of being lost as inheritors gradually die out. In recent years, motion capture and three-dimensional reconstruction technology has made significant breakthroughs in digital humanities, and the multi-modal perception ability combined with deep learning provides a technical possibility for the digital inheritance of intangible cultural heritage.

[0003] However, existing digital technology for intangible cultural heritage faces multi-dimensional technical bottlenecks in fine motion reproduction, especially in the fusion of multi-modal data in processing heterogeneous sensor information, which encounters feature conflict and weight mismatch problems. With the increase in the types of acquisition devices, the physical dimensions, sampling frequencies and noise characteristics of different modal data differ significantly, and traditional fixed weight fusion strategies cannot adapt to fluctuations in data quality in actual collection. In complex skill scenarios, when a sensor is blocked or interfered, the system still forcibly fuses low-quality data according to the preset weight, resulting in serious distortion of feature representation. To maintain basic reproduction accuracy, the system is forced to rely on a single dominant modality (usually visual data), artificially ignoring the complementary information in other dimensions. This simplification strategy causes significant loss of details, especially the intensity changes of fine operations of the hands, subtle differences in muscle control and other key skill features, which cannot be captured at all, and expert reviews often point out that the motion looks similar but lacks charm.

[0004] At the same time, existing methods generally use a single scale feature extraction architecture, which cannot take into account cross-scale information of macro gestures and micro tremors. Traditional convolutional networks are limited by fixed receptive fields, either focusing on overall motion and losing fine details, or focusing on local operations and ignoring global coordination. In particular, skill micro-motions at the micro-tremor level (such as the lifting and pressing tremors of calligraphy brush movements, and the control micro-tremors of embroidery puncture) are mixed with physiological tremors in high-frequency signals, and existing filtering schemes use global noise reduction, which suppresses physiological interference while also eliminating key features that reflect skill levels.

[0005] In addition, in the aspect of hand-tool co-modeling, the prior art often processes hand actions and tool operations separately, lacks overall modeling of interactive features such as contact relationship and mechanical transmission, and causes unrealistic phenomena such as tool suspension and holding through the model in the reproduced virtual character. Existing solutions such as adding constraint rules or increasing sampling density not only cannot solve the fundamental problem of semantic understanding, but also introduce additional computational overhead and parameter tuning complexity, which seriously restricts the practical application value of the technology in the field of intangible cultural heritage protection.

[0006] In view of this, the present application proposes a deep learning-based intangible cultural heritage skill action capture and digital reproduction method to solve the above problems. SUMMARY

[0007] In order to overcome the above-mentioned defects of the prior art, in order to achieve the above-mentioned purpose, the present application provides the following technical solutions:

[0008] The deep learning-based intangible cultural heritage skill action capture and digital reproduction method comprises:

[0009] Obtaining original action data in the execution process of intangible cultural heritage skills, wherein the original action data comprises multi-view high-speed video stream, six-degree-of-freedom motion data, pressure distribution data and electromyographic signal data;

[0010] Determining the fusion weight of each modal data in the original action data according to the skill type characteristics of the intangible cultural heritage action, and performing weighted fusion on each modal data based on the fusion weight to obtain fused action data;

[0011] Performing action feature analysis of different granularities based on the fused action data to obtain a three-level action feature map; the three-level action feature map comprises a macro posture layer, a fine operation layer and a micro jitter layer;

[0012] Identifying and extracting skill key operation points in the fine action region of the hand in the fine operation layer, and constructing a fine action key frame sequence according to the time sequence distribution of the skill key operation points;

[0013] Performing action refinement reconstruction according to the fine action key frame sequence and combining the spatio-temporal correlation between each level in the three-level action feature map to obtain an enhanced action feature representation;

[0014] Performing semantic annotation on the fine action feature representation based on a preset intangible cultural heritage knowledge graph to obtain skill semantic annotation, and driving digital action reproduction based on the skill semantic annotation.

[0015] Further, the process of obtaining the fused action data comprises:

[0016] According to the skill type characteristics of the non-heritage skill action, modal importance prior knowledge corresponding to the skill type characteristics is retrieved from a preset non-heritage skill knowledge graph;

[0017] Each modal data in the original action data is subjected to data preprocessing, and a basic fusion weight of each modal data is initialized based on the modal importance prior knowledge;

[0018] A signal-to-noise ratio and a data integrity index corresponding to each modal data in a data acquisition process are obtained, and the basic fusion weight is dynamically corrected based thereon to obtain a fusion weight;

[0019] Each modal data after preprocessing is subjected to data fusion based on the fusion weight to obtain fused action data.

[0020] Further, the process of constructing a three-level action feature graph includes:

[0021] A hierarchical action analysis network is constructed, and the hierarchical action analysis network includes a macro posture branch, a fine operation branch, and a micro jitter branch;

[0022] The fused action data is input into the constructed hierarchical action analysis network for action analysis to obtain a macro posture level feature graph, a fine operation level feature graph, and a micro jitter level feature graph;

[0023] Based on the macro posture level feature graph, the fine operation level feature graph, and the micro jitter level feature graph, different level features are analyzed and organized in a hierarchical manner to obtain a three-level action feature graph.

[0024] Further, the process of constructing a fine action key frame sequence includes:

[0025] A feature sequence of the fine operation layer in the three-level action feature graph is extracted, and a feature change rate of each frame feature vector in the feature sequence relative to a previous frame is calculated;

[0026] A time domain derivative calculation is performed on the feature change rate to obtain a feature acceleration curve, and extreme points in the feature acceleration curve are obtained;

[0027] According to the spatial gradient characteristics of the pressure distribution data, a force gradient vector at each time is calculated, and an amplitude analysis is performed thereon to obtain a force turning point;

[0028] A velocity vector sequence is extracted from the preprocessed six-degree-of-freedom motion data, and a module length change curve corresponding thereto is calculated; a start and end time of a high-variance interval in the module length change curve is analyzed and identified as a velocity change point;

[0029] Extracting motion trajectories of the hand and the tool based on the multi-view high-speed video stream, and obtaining a curvature mutation position by performing curvature difference calculation on the motion trajectories;

[0030] Performing time alignment on the extreme value points, the intensity turning points, the speed change points and the curvature mutation positions, and performing key point aggregation by using a time clustering algorithm;

[0031] Performing skill importance scoring on the aggregated key points, and retaining key points with skill importance scores greater than a preset importance threshold as the skill key operation points;

[0032] According to the time sequence of the skill key operation points, extracting action data frames corresponding to the time, and constructing a fine action key frame sequence.

[0033] Further, the process of obtaining the enhanced action feature representation comprises:

[0034] Inputting the fine action key frame sequence into a preset constraint optimization network, and constructing a comprehensive constraint objective function; the constraint optimization network comprises a spatial coordination constraint module and a time continuity constraint module;

[0035] By minimizing the comprehensive constraint objective function, iteratively updating the action features in the corresponding fine action key frame sequence; and in the iterative updating process, dynamically monitoring the convergence speed of the comprehensive constraint objective function; when the convergence speed is lower than a preset speed threshold, dynamically scaling the learning rate according to the historical statistical information of the current gradient;

[0036] After the comprehensive constraint objective function converges, extracting the optimized fine action key frame sequence features, and refining and reconstructing the features of non-key frames between the key frames by a space-time interpolation network to obtain the enhanced action feature representation.

[0037] Further, the process of constructing the comprehensive constraint objective function comprises:

[0038] In the time continuity constraint module, interpolating and predicting the action features between adjacent key frames in the fine action key frame sequence; calculating the residual between the intermediate frame features obtained by interpolation and prediction and the corresponding actual frame features in the fused action data, and constructing a time continuity loss function based on the residual;

[0039] In the spatial coordination constraint module, a spatial coordination rule library is constructed according to human kinematics constraints and tool operation constraints;

[0040] For each frame in the fine action key frame sequence, the corresponding skeleton joint position and tool position are extracted; and whether the skeleton joint position and tool position satisfy the rules in the spatial coordination rule library is verified; if not, the violation degree of the corresponding key frame violating the rules is quantified, and it is taken as a spatial coordination penalty term;

[0041] The time continuity loss function and the spatial coordination penalty term are combined by weighting to construct a comprehensive constraint objective function.

[0042] Further, the process of obtaining the skill semantic annotation includes:

[0043] An action category label set related to the intangible skill action is extracted from the intangible skill knowledge graph, and the action category label set is encoded into a semantic embedding vector;

[0044] The enhanced action feature representation is input into a preset skill semantic understanding model, and the enhanced action feature representation is mapped from the action feature space to the semantic feature space through the cross-modal alignment layer in the skill semantic understanding model to obtain an action semantic feature;

[0045] The similarity between the action semantic feature and the semantic embedding vector of each action category label is calculated, and the top K action category labels with the highest similarity are selected as candidate semantic annotations, where K is an integer representing the number of candidates;

[0046] The action specification compliance degree corresponding to each candidate semantic annotation is calculated based on the action specification standard in the intangible skill knowledge graph;

[0047] According to the weighted score of the similarity and the action specification compliance degree, and based on the weighted score, the candidate semantic annotations are sorted, and the annotation with the highest weighted score is selected as the skill semantic annotation;

[0048] The cultural semantic association in the intangible skill knowledge graph is read, the cultural semantic association content associated with the skill semantic annotation is retrieved, and it is attached to the skill semantic annotation as an auxiliary annotation.

[0049] Further, the process of driving digital action reproduction based on the skill semantic annotation includes:

[0050] According to the skill semantic annotation, a corresponding driving parameter template is retrieved from a preset action driving parameter library; and the enhanced action feature representation is converted into a skeleton joint angle sequence at the same time;

[0051] Based on the driving parameter template and the skeleton joint angle sequence, the intangible skill action is executed in a pre-constructed digital role model; and during the action execution process, the position and posture of the virtual tool are synchronously driven according to the pre-acquired tool-hand coordination feature;

[0052] Meanwhile, the micro-shaking feature is taken as a high-frequency disturbance signal, superimposed on the key point position of the digital role model, and a micro-shaking effect is generated through a programmed animation technology to output a motion reproduction result based on the digital role model.

[0053] Further, the tool-hand coordination feature acquisition process includes:

[0054] extracting a hand region and a tool region from the multi-view high-speed video stream;

[0055] extracting three-dimensional coordinates of part key points in the hand region, and constructing a hand skeleton topology graph based thereon;

[0056] Meanwhile, a tool three-dimensional model is constructed according to the tool type required in the non-heritage skill action execution process, and the position and posture of the tool three-dimensional model in the world coordinate system are calculated through six-degree-of-freedom motion data;

[0057] labeling functional key points on the tool three-dimensional model, and extracting three-dimensional coordinates of the functional key points;

[0058] Further, a hand-tool contact relationship graph is constructed according to the hand skeleton topology graph and the tool three-dimensional model; the hand-tool contact relationship graph is composed of nodes composed of hand key points and functional key points and contact edges between the nodes;

[0059] According to the pressure distribution data, the contact force intensity of each contact edge is calculated, and the contact force intensity is taken as the weight of the contact edge;

[0060] For the hand-tool contact relationship graph, feature propagation is performed through a graph neural network, and a node embedding vector of the graph neural network is extracted as a tool-hand coordination feature.

[0061] Further, the method further includes separating and enhancing the micro-shaking feature, including:

[0062] extracting shaking feature data in the micro-shaking layer from the three-level action feature graph, the shaking feature data including physiological tremor and skill micro-motion of the hand;

[0063] The shaking feature data is decomposed into a plurality of frequency components through frequency domain analysis, and the corresponding frequency components are divided into physiological tremor frequency bands and skill micro-motion frequency bands according to the frequency range;

[0064] The frequency components of the physiological tremor frequency band are suppressed through a notch filter; at the same time, the micro-motion type label and the micro-motion intensity value corresponding to the frequency components of the skill micro-motion frequency band are identified and extracted;

[0065] According to the micro-motion type label, a standard micro-motion template corresponding to the micro-motion type is retrieved from a preset non-heritage skill knowledge graph, the standard micro-motion template including typical time-frequency characteristics of the micro-motion type;

[0066] A frequency component of the technical micro-motion frequency band is calculated to be similar to a feature of the standard micro-motion template, and when the feature similarity is lower than a preset similarity threshold, a micro-motion feature enhancement operation is performed, the micro-motion feature enhancement operation including: migrating high-confidence features of the standard micro-motion template to current micro-motion features through a feature migration network;

[0067] The enhanced technical micro-motion features and the inhibited physiological tremor features are recombined to obtain separated enhanced micro-dithering features, and the separated enhanced micro-dithering features are superimposed into the enhanced action feature representation.

[0068] The technical effects and advantages of the non-heritage skill action capture and digital reproduction method based on deep learning of the present application are as follows:

[0069] The present application scientifically fuses multi-modal heterogeneous data to ensure that information from the visual, motion, mechanical to physiological dimensions can truly contribute to skill representation, significantly improving the degree of restoration of digital reproduction, and enhancing the cultural identity and skill communication of non-heritage inheritance; and in practical application, even in the face of complex and delicate high-difficulty skills (such as cloisonné dot blue, nuclear carving micro-carving, etc.), the cross-scale feature integrity from macro-posture to micro-dithering can still be maintained, and artificial compensation or simplification is no longer needed, making the action sequence more natural and coherent. In particular, the tool-hand collaborative modeling breaks through the limitations of traditional fragmented processing, eliminating the physical unreasonable phenomena of tool suspension, holding and wearing the mold.

[0070] In terms of data quality management, the adaptive fusion weight mechanism improves the robustness of the system to fluctuations in the collection environment, and when the sensor is blocked or disturbed, there is no longer obvious feature distortion, and it performs stably in a long-time collection scene, providing reliable protection for the one-time complete recording of the skills of precious inheritors.

[0071] In terms of cultural heritage value, the micro-dithering separation and enhancement technology first realizes the intelligent differentiation of technical micro-motion and physiological tremor, truly preserves the core detail features reflecting the skill level, and enables the digital archive to permanently retain the exquisite skill essence of the inheritors. Combined with knowledge graph driven semantic annotation, the system can automatically associate the cultural connotation of skill actions, support quantitative evaluation of the skill level of apprentices, and provide an objective and scientific tool for non-heritage education. BRIEF DESCRIPTION OF DRAWINGS

[0072] Figure 1 The present application is a schematic diagram of a non-heritage skill action capture and digital reproduction method based on deep learning. DETAILED DESCRIPTION

[0073] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work are within the protection scope of the present application.

[0074] Embodiment 1

[0075] Please refer to Figure 1 The method for capturing and digitizing the non-heritage skill action based on deep learning in the embodiment includes:

[0076] The original action data in the non-heritage skill action execution process is acquired through a multi-modal sensor array. The multi-modal sensor array includes a multi-view high-speed camera array, a wearable inertial measurement unit, a hand contact force sensor, a surface electromyography signal collector and other key devices. Synchronous data acquisition is realized through a high-speed API interface. The original action data includes multi-view high-speed video stream, six-degree-of-freedom motion data in kinematics dimension, pressure distribution data in mechanics dimension, and electromyography signal data in physiology dimension, and other key information. The multi-view high-speed video stream captures the visual details of the skill execution at a fixed frame rate. The six-degree-of-freedom motion data of the inertial measurement unit records the acceleration and angular velocity changes of the limbs in real time. The pressure distribution data of the contact force sensor reflects the interactive mechanical characteristics of the hands and tools. The electromyography signal data records the physiological state of muscle activation.

[0077] The fusion weight of each modal data in the original action data is determined according to the skill type characteristics of the non-heritage skill action, and each modal data is weighted and fused based on the fusion weight to obtain the fused action data. Through the preprocessing operations such as time alignment, coordinate system unification and noise suppression on each modal data in the original action data, the basic fusion weight of each modal data is initialized based on the modal importance prior knowledge. At the same time, the fluctuation of data quality in the actual acquisition process, further obtains the signal-to-noise ratio and data integrity index of each modal data, through the adaptive weight adjustment strategy to dynamically correct the basic fusion weight, generates the fused action data which is time and space aligned and multi-dimensional complementary

[0078] Based on the fused action data, action feature analysis of different granularities is performed to obtain a three-level action feature graph; the three-level action feature graph includes a macro posture layer, a fine operation layer, and a micro jitter layer. By adopting a parallel multi-branch network architecture, including a macro posture branch, a fine operation branch, and a micro jitter branch, three independent channels are formed; the macro posture branch uses a three-dimensional convolution kernel with a large receptive field to extract the overall motion trajectory of the main joints of the body, and generates low-frequency posture features through global average pooling; the fine operation branch focuses on the hand and tool regions through an attention mechanism, and uses a combination of multi-scale convolution kernels to capture fine action patterns of different time spans; the micro jitter branch performs high-pass filtering on the data, and uses a gated recurrent unit to extract high-frequency jitter time-dependent features. The output features of the three branches are associated across levels through a graph convolution network to form a tree-structured three-level action feature graph, which completely represents multi-scale action information from a macro to a micro level.

[0079] Key operation points in the hand fine action region in the fine operation layer are identified and extracted, and a fine action key frame sequence is constructed according to the time sequence distribution of the key operation points. First, the fine operation layer features are extracted from the three-level action feature graph, the inter-frame feature change rate and the feature acceleration curve are calculated, and the action turning points are identified through extreme value detection; the strength turning points are identified by combining the spatial gradient mutation of the pressure distribution data, the speed variance analysis of the inertia data is used to identify the speed change points, and the trajectory mutation points are identified by combining the curvature second-order difference of the video trajectory. The multiple types of key points are aggregated on the time axis, time clustering algorithm is used to merge similar points, then the key points with high matching degree to the standard action specification in the intangible cultural heritage knowledge graph are reserved through importance scoring screening, and a refined set of key operation points is formed; the action data frames corresponding to the key points are extracted in time sequence, and a fine action key frame sequence is constructed.

[0080] According to the fine action key frame sequence, and combining the spatio-temporal correlation between the levels in the three-level action feature graph, action refinement reconstruction is performed to obtain enhanced action feature representation. The refinement reconstruction process includes two core components: a time continuity constraint module and a spatial coordination constraint module; the time continuity constraint module performs physical motion model interpolation prediction on the key frame gaps, calculates the residual between the predicted frame and the actual frame to construct a time loss function; the spatial coordination constraint module verifies the rationality of the joint angles and tool positions according to the human kinematics and tool operation constraint rule library, and calculates a penalty term for the frames that violate the constraints. The time loss and spatial penalty are combined into a comprehensive constraint objective function, the key frame features are optimized through gradient descent iteration, and the spatio-temporal interpolation reconstruction is performed on the non-key frames using a bidirectional long short-term memory network, the forward and backward propagation are fused in the intermediate frames to generate enhanced fine action feature representation, ensuring the spatio-temporal coherence and physical rationality of the action sequence.

[0081] Based on a pre-defined intangible cultural heritage skills knowledge graph, semantic annotations are performed on the fine motion feature representations to obtain skills semantic annotations, which then drive digital motion reproduction. First, a pre-trained Transformer encoder semantically encodes the motion features. A cross-modal alignment layer maps the motion features to a semantic space, performing similarity matching with the skills motion category labels in the knowledge graph. Combined with motion standard compliance scores, the optimal semantic annotation is selected and appended with cultural connotation descriptions. A digital character model is then driven by these semantic annotations. This model includes a hierarchical skeleton system and a muscle system. An inverse kinematics solver converts motion features into joint angle sequences, and simultaneously infers muscle activation timing based on electromyography signals, bidirectionally driving virtual character posture changes. A dual quaternion skinning algorithm is used to update the skin surface in real time, overlaying multi-layered deformation and micro-jitter effects in fine areas to output a high-fidelity 3D motion reproduction animation sequence.

[0082] In an embodiment of the present invention, the process of acquiring fused action data includes:

[0083] Based on the skill type characteristics of intangible cultural heritage (ICH) techniques, prior knowledge of modal importance for corresponding skill type characteristics is retrieved from a pre-defined ICH knowledge graph. Data preprocessing is performed on each modal data within the original action data, and basic fusion weights for each modality are initialized based on the prior knowledge of modal importance. Prior knowledge of modal importance is a structured expression of domain expert experience, reflecting the dependence characteristics of different skill types on each modality data. The retrieval process identifies the skill type label of the current skill action (e.g., "Jingdezhen hand-throwing," "Suzhou embroidery," "Yixing purple clay teapot shaping") based on the preprocessed original action data, searches for the corresponding node in the ICH knowledge graph, and extracts the modal importance parameters for that skill type as basic fusion weights. Data preprocessing includes time alignment of multi-view high-speed video streams, coordinate transformation of six-free motion data, bandpass filtering and rectification of electromyographic signal data, and acquisition of spatial gradient features corresponding to pressure distribution data.

[0084] The signal-to-noise ratio (SNR) and data integrity index of each modality are acquired during the data acquisition process. Based on these, the basic fusion weights are dynamically adjusted to obtain the fusion weights. During the acquisition process, the SNR and data integrity index of each modality are calculated in real time. When the quality of a certain modality deteriorates, its weight is dynamically reduced. Weight adjustment is employed, using the fusion effect and subsequent recognition accuracy as reward signals to optimize the weight allocation online. The mathematical formula for weight adjustment is as follows: In the formula, Indicates the final fusion weight; Indicates the basic fusion weights; SNR i C represents the signal-to-noise ratio of the i-th modal data; iThe i-th modal data represents the integrity index; j represents the modal data involved in the fusion (traversing all modalities);

[0085] α0∈(0.3,0.5) is the confidence coefficient of the prior weight;

[0086] The preprocessed modal data are fused based on fusion weights to obtain fused motion data. The fusion process first aligns the modal data along the time dimension to ensure that data at the same time step corresponds to the same action state; then, the feature vectors of each modality are normalized to eliminate dimensional differences; finally, they are weighted and summed according to the calculated fusion weights to obtain a fused feature vector. The fused feature vectors form a continuous sequence along the time dimension, constituting the fused motion data. This fused data retains the advantages of each modality (such as the visual richness of the video, the trajectory accuracy of the motion data, the force details of the pressure data, and the physiological authenticity of the electromyographic signals), while eliminating the limitations of a single modality (such as viewpoint occlusion in the video, drift errors in the motion data, insufficient spatial resolution of the pressure data, and individual differences in the electromyographic signals).

[0087] In embodiments of the present invention, the data preprocessing implementation process includes:

[0088] Time synchronization processing is performed on high-speed video streams from multiple perspectives, extracting hardware timestamps from each perspective and calculating the time offset between perspectives. Time synchronization ensures consistency across different data sources in the time dimension. The process first reads the timestamps from each camera, calculates the delay difference between adjacent perspectives through cross-correlation analysis, and then uses a time interpolation algorithm to align the video frames of all perspectives to the master clock reference. For video streams with inconsistent frame rates, a dynamic time warping algorithm is used for variable-speed alignment to ensure that critical actions remain synchronized across all perspectives. Synchronization accuracy is verified through sub-pixel-level feature point matching, with errors controlled within 1 millisecond, providing a reliable time reference for subsequent multi-view feature fusion.

[0089] Coordinate transformation and error correction are performed on the six-DOF motion data of the inertial measurement unit (IMU) to map the sensor's local coordinate system data to the world coordinate system. Coordinate transformation eliminates the influence of differences in sensor installation position and orientation. The transformation process first uses the rotation matrix and translation vector obtained during the calibration phase to transform acceleration and angular velocity from the sensor coordinate system to the body coordinate system. Then, based on visually captured body posture estimation, it is further mapped to the world coordinate system. Due to the inherent drift error of the inertial sensor, an extended Kalman filter is used for real-time correction, fusing visual position observations as a reference. Through state prediction and observation update iterative optimization, the accumulated error is significantly reduced, ensuring the accuracy of data acquired over long periods.

[0090] Spatial mapping is performed on pressure distribution data from contact force sensors to construct pressure distribution heatmaps and extract spatial gradient features. The pressure heatmaps visually represent the force distribution during hand-tool contact. The mapping process generates a continuous two-dimensional heatmap from the discrete pressure values ​​of the sensor array using bilinear interpolation. Pixel brightness corresponds to pressure intensity, and color coding reflects pressure level. Spatial gradient features are extracted from the heatmaps, and the Sobel operator is used to calculate the horizontal and vertical gradient components. Gradient magnitude characterizes the steepness of pressure change, and gradient direction indicates the force transmission path. These spatial distribution vectors provide a mechanical dimension representation for identifying details of the technique, such as grip patterns and force application methods.

[0091] Signal processing was performed on electromyographic (EMG) data to extract time-series curves of muscle activation. The processing flow first removed motion artifacts and power frequency interference using a bandpass filter, preserving the effective frequency bands for muscle activation. Then, the filtered signal was fully rectified to convert the bipolar signal into a unipolar amplitude representation. Finally, the envelope was extracted using a low-pass filter to generate a smooth time-series curve of muscle activation. The peak value of the curve corresponds to the moment of maximum muscle contraction, and the area under the curve reflects the sustained contraction intensity, providing physiological evidence for analyzing muscle control patterns in technical movements.

[0092] In an embodiment of the present invention, the process of constructing a three-level action feature map includes:

[0093] A hierarchical action parsing network is constructed, which includes three parallel feature extraction channels: macroscopic pose branch, fine operation branch, and microscopic jitter branch.

[0094] Then, the fused action data is input into the hierarchical action parsing network. The network input layer receives the fusion matrix and performs shallow feature extraction through the initial convolutional layer in the input layer. Then, it is split into three parallel feature extraction channels for deep processing. The convolution kernel size and pooling strategy of each feature extraction channel are designed differently according to the target feature size to ensure that each branch focuses on action patterns of a specific granularity and avoids feature confusion and information loss.

[0095] The overall motion trajectory features are extracted in the macro-pose branch, and spatiotemporal feature extraction is performed using a 3D convolutional kernel. Macro-pose reflects the large-amplitude movements of the body's major joints. In the feature extraction channel design, the temporal dimension of the 3D convolutional kernel is set to cover the number of frames of a complete skill movement cycle, and the spatial dimension covers the joint regions of the whole body. High-level semantic features are gradually extracted through multi-layer convolution and pooling. Finally, the spatial dimension is compressed through a global average pooling layer to extract low-frequency features of overall pose changes, generating a macro-pose feature map. The macro-pose feature map represents macro-movement patterns such as body center of gravity shift and major limb swings.

[0096] In the fine manipulation branch, fine hand movement features are extracted. First, key manipulation areas are located using an attention mechanism, and then local feature extraction is performed using a combination of multi-scale convolutional kernels. Fine manipulation is the core of intangible cultural heritage skills, requiring a focus on subtle changes in the hands, fingertips, and tool contact areas. During the execution of the attention mechanism, high-pressure areas are automatically focused based on pressure distribution data, and a spatial attention weight map is calculated to highlight important areas of hand manipulation. For the located fine areas, a combination of multi-scale convolutional kernels is deployed, including three spatiotemporal scales: 1×3×3, 3×3×3, and 5×5×5, to capture instantaneous movements, short sequence patterns, and medium-span manipulation sequences, respectively. Feature maps at each scale are fused through a feature pyramid network to construct a multi-resolution feature representation from the bottom up, generating a fine manipulation feature map that records in detail the key skill details such as finger movements and tool manipulation.

[0097] High-frequency jitter features are extracted from the micro-jitter branch. After high-pass filtering of the fused motion data, the data is input into a recurrent neural network (GRU) to capture temporal dependencies. Micro-jitter includes physiological tremors and skillful micro-movements. The feature extraction channel first extracts high-frequency jitter components through a high-pass filter, removing low-frequency overall motion trends. Then, the high-frequency signal is input into a gated recurrent unit (GRU) network, utilizing the GRU's memory capability to capture the temporal patterns and rhythmic patterns of the jitter. The hidden state sequence of the GRU encodes the dynamic evolution information of the jitter, serving as the core representation of the micro-jitter feature map. This micro-jitter feature map reveals the stability of the artist's hand control, the frequency and amplitude of skillful micro-movements, and other micro-features, providing an important basis for distinguishing different skill levels.

[0098] A three-level feature dependency relationship is established, and a graph convolutional network is used to propagate and fuse feature information. The association process uses macroscopic, fine, and microscopic feature maps as graph nodes, and establishes directed graph edges based on temporal causal relationships (different level features at the same time) and spatial collaborative relationships (inter-level influence at adjacent times). Feature information is propagated through multi-layer graph convolution, with each node aggregating features from neighboring nodes for updating, achieving information interaction and complementary enhancement between levels. The message passing mechanism of graph convolution ensures the rationality of macroscopic posture constraints on fine operations, fine operations refine the local details of macroscopic posture, and microscopic jitter adds realism by superimposing the two. The graph convolution output is hierarchically organized and stored in a tree structure to construct a complete three-level action feature map. The root node is the macroscopic posture feature layer, used to store content related to macroscopic posture features; the middle nodes are the fine operation feature layers, used to store content related to fine operation features; and the leaf nodes are the microscopic jitter feature layers, used to store content related to microscopic jitter features, to comprehensively represent the multi-scale information of intangible cultural heritage techniques.

[0099] In embodiments of the present invention, the process of constructing a sequence of fine-grained motion keyframes includes:

[0100] Fine-grained operation feature sequences within the fine-grained operation layer are extracted from the three-level motion feature map, and the inter-frame feature change rate and feature acceleration curve are calculated. Feature change analysis identifies key moments by quantifying the feature evolution speed. The analysis process first calculates the Euclidean distance between the feature vectors of adjacent frames as the feature change rate to reflect the instantaneous speed of motion change; then, the first derivative of the obtained feature change rate in the time domain is calculated to obtain the feature acceleration curve, which is used to reveal the acceleration or deceleration trend of motion change; extreme points, including local maxima and minima, are detected in the acceleration curve. These extreme points correspond to turning points in motion change, such as the initial acceleration, reaching the peak, and the beginning of deceleration. Extreme points are identified using a sliding window comparison method, with the window size adaptively adjusted according to the average duration of the skill movement to ensure the accuracy and robustness of detection.

[0101] This method combines multimodal data to identify different types of key points: force inflection points from pressure distribution data, velocity change points from six-degree-of-freedom motion data, and curvature abrupt change points from multi-view high-speed video streams. Specifically, in the force inflection point identification process, spatial gradient vectors are extracted based on the spatial distribution characteristics corresponding to the pressure distribution data. The temporal changes in gradient amplitude are analyzed, and when the temporal changes in gradient amplitude exceed a dynamic threshold, it is marked as a force inflection point, reflecting a change in the force application method. In the velocity change point identification, velocity vector magnitude sequences are extracted from six-degree-of-freedom motion data, and local variance is calculated using a sliding window. High variance intervals correspond to periods of drastic velocity changes, and the start and end times of these intervals are extracted as velocity change points. In the curvature abrupt change point identification, the three-dimensional motion trajectories of the hand and tool are extracted from the multi-view high-speed video stream, the curvature function of the trajectory is calculated, and second-order difference analysis is performed on the curvature to detect abrupt change locations, corresponding to sharp changes in trajectory direction. These three types of key points comprehensively capture the key features of the skill movement from mechanical, kinematic, and geometric perspectives.

[0102] Extreme points, intensity inflection points, velocity change points, and curvature abrupt change locations are time-aligned, and keypoints are aggregated using a temporal clustering algorithm. The keypoint aggregation process involves aligning these points along the time axis and then using a temporal clustering algorithm to aggregate keypoints with similar times. Aggregation is a necessary step to remove redundant keypoints and prevent the same action event from being repeatedly marked. The aggregation process arranges all types of keypoints in chronological order and uses the density-based DBSCAN clustering algorithm, with the cluster radius set to the minimum recognizable action duration of the intangible cultural heritage skill. Multiple keypoints within a cluster radius are merged into a single aggregation point. The temporal position of the aggregation point is calculated as a weighted average of its members, with weights allocated based on the keypoint's confidence level. The aggregated keypoints retain multimodal comprehensive information while eliminating temporal duplication, forming a concise set of candidate keypoints.

[0103] The aggregated key points are scored for their technical importance. Based on the action specification definitions in the intangible cultural heritage skills knowledge graph, important key points are selected through semantic matching. Importance scoring is the core step in extracting truly key operational points from a large number of candidate points, ensuring that the recognition results conform to the technical specifications. The scoring process first extracts the action fragment features before and after each key point, including the posture, trajectory, and force patterns near that point; then, it performs semantic matching with the specification definitions of standard technical actions in the knowledge graph, calculating a feature similarity score. Similarity calculation uses cosine distance or dynamic time-warped distance to quantify the degree of consistency between the current action fragment and the standard action. The technical importance score comprehensively considers the similarity score and the multimodal consistency of the key points, using the formula: score k =β0×Sim(F k F standard )+(1-β0)×Con k In the formula, scote k Assess the skill importance of key point k, Sim(F) k F standard The similarity between the current action segment and the standard action is represented by Con. k β0 is a multimodal consistency index (identifying the degree of consistency of a point based on different modalities), and β0 is a balance coefficient (usually set to 0.6-0.7). Key points with scores greater than a preset importance threshold are retained as the final key operation points of the skill.

[0104] Data frames are extracted according to the time sequence of key operation points in the technique to construct a refined keyframe sequence. The keyframe sequence can be used to retain the most critical action information with the fewest frames. The construction process extracts the complete data frame corresponding to the time of the key operation point from the original fused action data, including the multimodal feature vector, three-level feature map representation and contextual information at that time. The keyframes are organized into a sequence structure in chronological order, which significantly compresses the amount of data while retaining the core elements of the technique.

[0105] In embodiments of the present invention, the process of obtaining enhanced action feature representations includes:

[0106] The obtained fine-grained motion keyframe sequences are input into a pre-constructed constraint optimization network, which includes a temporal continuity constraint module and a spatial coordination constraint module. The network architecture of the constraint optimization network adopts an encoder-optimizer-decoder structure. The encoder extracts the latent representation of the keyframes and iteratively solves the problem by applying spatiotemporal constraints through the optimizer. The decoder generates the optimized feature representation. The corresponding modules work in parallel, applying physical laws and kinematic constraints from the time and space dimensions respectively to ensure that the reconstructed motion features conform to the motion laws of the real world.

[0107] Physical motion interpolation is performed in the temporal continuity constraint module to predict trajectories based on a physical model between keyframes. Temporal continuity is a fundamental characteristic of natural motion, and physical model interpolation better reflects real motion patterns than simple linear interpolation. The interpolation process establishes a physical motion model, considering the influence of gravity, inertia, and friction on the movement of the hand and tool. Based on the position and velocity boundary conditions of the keyframes, the motion trajectory of intermediate frames is solved. Numerical integration is used to simulate the physical process, generating smooth intermediate frame features that conform to physical laws. The residual between the interpolated predicted frame and the actual frame in the fused data is calculated to construct a temporal continuity loss function. A weighted mean square error form is used, with weights dynamically adjusted according to the time interval between frames; the larger the interval, the smaller the weight, reflecting increased uncertainty. The mathematical expression of the temporal continuity loss function is as follows: In the formula, L temp For the loss of time continuity, Let be the predicted features of the a-th intermediate frame. For the corresponding actual characteristics, Δt a The time interval between the corresponding frame and the most recent keyframe is given by A, and the total number of frames is given by A. This loss function drives the optimization process to generate a temporally coherent sequence of actions.

[0108] Kinematic verification is performed in the spatial coordination constraint module. Based on the human kinematics constraint and tool operation constraint rule base, the rationality of the spatial configuration is checked. Spatial coordination is an inherent constraint of human movement; joint angles, limb lengths, and tool contact must meet physical limitations. The constraint rule base includes professional rules such as joint angle ranges, constant bone length, and consistency of tool-hand contact, established based on ergonomics and operational standards. During verification, the skeletal joint positions and tool positions of each frame are extracted to check whether they meet the constraints in the rule base. For configurations that violate constraints, the degree of deviation is calculated as a spatial coordination penalty. The penalty is designed using a barrier function form, with small penalties for minor violations and large penalties for serious violations, guiding the optimization process towards the feasible region.

[0109] The time loss and space penalty are weighted and combined into a comprehensive constraint objective function, and the action features are iteratively optimized through gradient descent. Comprehensive optimization is a solution process that balances temporal continuity and spatial consistency, seeking the optimal feature configuration that simultaneously satisfies both types of constraints. The formula for the comprehensive constraint objective function is: L total =λ1×L temp +λ2×L spatial In the formula, L total To synthesize the objective function with constraints, L spatial λ1 and λ2 are weighting coefficients, representing the spatial coordination penalty term.

[0110] The optimization algorithm employs Adam gradient descent to iteratively update keyframe feature representations and dynamically monitor the convergence speed of the objective function. When the convergence speed falls below a threshold, an adaptive learning rate adjustment mechanism is introduced to dynamically scale the step size based on historical gradient statistics, accelerating the convergence process. After convergence, a bidirectional long short-term memory network is used to perform spatiotemporal interpolation reconstruction on non-keyframes. The forward LSTM propagates from the starting keyframe, and the backward LSTM propagates from the ending keyframe. Bidirectional features are fused at intermediate frame positions to generate refined features that incorporate both preceding and following contextual information, resulting in an enhanced action feature representation.

[0111] In embodiments of the present invention, the process of obtaining craft semantic annotations includes:

[0112] Extract action category label sets related to intangible cultural heritage skills from the knowledge graph of intangible cultural heritage skills, and encode the action category label sets into semantic embedding vectors;

[0113] The enhanced, refined action features are input into a pre-defined skill semantic understanding model, which maps them to the semantic space via a Transformer encoder and a cross-modal alignment layer. The skill semantic understanding model employs a pre-trained Transformer encoder architecture, capturing long-range temporal dependencies and internal correlations of action features through a multi-head self-attention mechanism. The cross-modal alignment layer, trained through contrastive learning, maps the action feature space to the semantic feature space, ensuring that similar actions and semantics are close in distance after mapping. The cosine similarity between the aligned action semantic features and the embedded labels of each category is calculated, and the top K labels with the highest similarity are selected as candidate semantic annotations.

[0114] Based on the action standard criteria within the intangible cultural heritage skills knowledge graph, the action standard compliance degree corresponding to each candidate semantic label is calculated. The candidate semantic labels are then ranked according to a weighted score of similarity and action standard compliance, with the label having the highest weighted score selected as the skill semantic label. The compliance evaluation process involves retrieving the corresponding action standard criteria from the knowledge graph for each candidate label, including the standard ranges for key parameters such as action duration, peak force, and trajectory deviation. The actual parameters of the captured action are compared with the standard ranges to calculate the degree of deviation and obtain a compliance score. The weighted score, combining similarity and compliance, is used to select the label with the highest score as the final skill semantic label. Simultaneously, based on the cultural semantic associations within the knowledge graph, the associated cultural semantic content is retrieved and added as an auxiliary label to the skill semantic label.

[0115] In embodiments of the present invention, the process of digital action reproduction driven by skill semantic annotation includes:

[0116] Based on the semantic annotation of the skill, the corresponding driving template is retrieved from the action driving parameter library, and the skeletal driving mode and muscle activation mode are defined. The action features are converted into a sequence of skeletal joint angles through an inverse kinematics solver, and the cyclic coordinate descent method is used for iterative solution to satisfy the position constraints of the end effector while optimizing the joint configuration.

[0117] This approach leverages skill-based semantic annotation to drive the execution of actions in digital character models, constructing virtual characters that include a layered skeleton, muscular system, and skin. Digital character-driven animation is a core technology that transforms abstract features into visual animation, enabling the virtual reproduction of skill-based movements. The skeletal system of the digital character model employs a layered structure, with the main skeleton controlling major parts such as the torso and limbs, and the finer skeleton controlling finer parts such as fingers and face. Movement is transmitted through parent-child hierarchical relationships. Furthermore, fine constraints such as inter-finger coupling and range of motion are applied to the finer skeleton, and angle configurations satisfying all constraints are obtained through constraint optimization.

[0118] By combining electromyography (EMG) signals to drive a virtual muscle system, high-fidelity motion reproduction with bidirectional musculoskeletal actuation is achieved. Muscle actuation is a crucial mechanism for enhancing the realism of movement, simulating the biomechanics of human muscles. The actuation process infers the activation sequence of muscle groups based on EMG signal data, mapping it to the virtual muscle system of the digital character to drive muscle contraction and relaxation. The biomechanical behavior of the virtual muscles is simulated using the Hill muscle model, calculating the driving force generated by the muscles. Combined with joint angle sequences, muscle force and angle bidirectionally drive changes in the character's posture. When the two are inconsistent, a force-angle coordination algorithm is used to reconcile them, prioritizing the angle accuracy of fine-grained areas. Virtual tools are synchronously driven, updating their position and posture based on pre-acquired tool-hand coordination features to ensure correct contact between the tool and the hand.

[0119] The skin of the digital character model is updated in real time, employing a dual quaternion skinning algorithm and a multi-layered deformation strategy to reveal details. Skinning updates are the surface processing that generates the final visual effect, transferring the movement of the internal skeleton and muscles to the skin surface. The skinning algorithm calculates the skin vertex positions based on the deformation of the bones and muscles, and the dual quaternion method effectively avoids the candy-wrapping effect of linear blending skinning. In areas of fine motion, multi-layered skinning is used: the first layer is based on bone deformation, the second layer on muscle deformation, and the third layer superimposes detail deformation based on micro-jitter features, refining the skin layer by layer to enhance realism.

[0120] Meanwhile, the micro-shaking features are used as high-frequency perturbation signals to generate micro-shaking effects through programmed animation technology. The shaking amplitude and frequency are modulated based on the noise function and superimposed on the key point positions. A continuous three-dimensional animation sequence and corresponding semantic annotation text of the skills are output, which fully presents the digital reproduction results of the intangible cultural heritage skills.

[0121] In an embodiment of the present invention, the process of obtaining tool-hand collaborative features includes:

[0122] The process of obtaining tool-hand collaborative features includes:

[0123] Hand and tool regions are extracted separately from multi-view high-speed video streams. Accurate separation of hands and tools is achieved through an instance segmentation network. The segmentation process employs a pre-trained MaskR-CNN instance segmentation network, specifically trained on a large-scale hand-tool interaction dataset, capable of simultaneously detecting and segmenting multiple instance objects. The network input consists of multi-view synchronized high-speed video frames. A backbone feature extraction network generates multi-scale feature maps and candidate bounding boxes, finally outputting pixel-level classification masks through a segmentation head. For the hand region, the network recognizes the complete palm and five-finger outline, even in cases of partial occlusion, by using contextual reasoning to complete the image. For the tool region, the network identifies tools based on their shape, texture, and color features, supporting various intangible cultural heritage tool types such as brushes, carving knives, potter's wheels, and scissors. The segmentation results are output as binary masks, with pixels within the mask representing the target object.

[0124] The 3D coordinates of key points within the hand region were extracted, and a hand skeleton topology map was constructed based on these coordinates. The extraction process first involved locating 21 key points on the 2D images from each viewpoint using a heatmap regression-based keypoint detection algorithm. These key points included 5 fingertips, 12 interphalangeal joints, and 4 metacarpophalangeal joints. Each key point's position probability distribution was represented by a Gaussian heatmap, with the heatmap peak corresponding to the most probable key point location. Then, 3D reconstruction was performed using multi-view geometric constraints. Triangulation was employed, and the projection rays of the 2D key points were intersected in 3D space based on the camera intrinsic and extrinsic parameter matrices from each viewpoint to obtain the 3D coordinates of the key points in the world coordinate system. To improve reconstruction accuracy, bundle adjustment optimization was introduced to minimize reprojection errors and iteratively optimize the 3D positions of the key points. Based on the 3D coordinates of the 21 key points, a hand skeleton topology map was constructed. This map uses key points as nodes and hand bone connections as edges, forming a tree-like hierarchical structure. The root node is the wrist, and the five branches correspond to the bone chains of the five fingers. The length of each edge in the topology map corresponds to the physical length of the bone, and the direction vector of the edge represents the bone orientation.

[0125] Simultaneously, based on the types of tools required in the execution of intangible cultural heritage techniques, corresponding 3D models of the tools are constructed, and the position and orientation of the 3D tool models in the world coordinate system are calculated using six-degree-of-freedom motion data. First, sparse point clouds are reconstructed from the tool region in multi-view high-speed video streams, and 3D point sets on the tool surface are generated using stereo matching or multi-view geometry methods. The reconstructed point cloud is then registered with the tool CAD model, and the optimal rotation matrix and translation vector are solved using an iterative nearest-point algorithm to minimize the distance between the model point cloud and the observed point cloud. Based on the six-degree-of-freedom motion data, the corresponding 3D position coordinates and 3D rotation angles are obtained, fully describing the spatial transformation of the tool relative to the world coordinate system.

[0126] Functional key points are marked on the 3D model of the tool, and the 3D coordinates of the functional key points are extracted. The marking process is based on the functional semantics and operating specifications of the tool. The functional key points include three types of core points: the point of action, that is, the position where the tool performs its main function, such as the tip of a brush, the blade of a carving knife, and the spout of a teapot; the grip point, that is, the conventional position where the hand grasps the tool, such as the grip area of ​​a brush handle, the grip surface of a knife handle, and the handle of a teapot; and the force transmission point, that is, the key node where the force applied by the hand is transmitted to the tool, such as the base of a brush handle and the end of a carving knife handle. The marking adopts a semi-automatic method. First, experts in intangible cultural heritage techniques manually mark typical functional points on the tool model to establish a tool functional point template library. In actual use, the functional point configuration of the current tool is retrieved from the template library, and the template functional points are mapped to the current world coordinate system through rigid body transformation according to the estimated tool posture.

[0127] Furthermore, a hand-tool contact relationship diagram is constructed based on the hand skeleton topology and the tool's 3D model. This diagram consists of nodes composed of key hand points and functional key points, along with contact edges between these nodes. The construction process first initializes the node set of the relationship diagram, containing 21 key hand points and N functional tool points (N varies depending on the tool type, typically 3-8). Each node has additional attributes including 3D coordinates, a node type label (hand point / tool ​​point), and a timestamp. Then, a distance-based algorithm is used to identify the contact relationships between nodes. For any hand key point and tool functional point, the Euclidean distance between them in 3D space is calculated. When the Euclidean distance is less than a preset contact distance threshold, a contact relationship is determined, and a five-way contact edge is established between the corresponding nodes, resulting in the corresponding hand-tool contact relationship diagram.

[0128] Simultaneously, the contact force intensity of each contact edge is calculated based on the pressure distribution data, and this contact force intensity is used as the weight of the contact edge. The calculation process first obtains pressure distribution data of the hand surface from the contact force sensor; then, a spatial mapping relationship between the pressure sensor coordinates and key points of the hand is established, and the anatomical position of the hand corresponding to each sensing unit is determined through a calibration program, such as the thumb pad, index finger tip, and palm. For each identified contact edge, the sensor area corresponding to its hand endpoint is located, and the pressure reading of that area is extracted. If the area contains multiple sensing units, the average pressure or maximum pressure is calculated as a representative value. For pressure transmission points on the tool side, if the tool itself is also equipped with a pressure sensor (such as the pressure sensing of the penholder of a smart brush), a weighted average of the pressure from both sides is used; if the tool has no sensor, only the pressure from the hand side is used. The dynamic update frequency of the edge weights is synchronized with the video frame rate, reflecting changes in mechanical interaction in real time.

[0129] For the hand-tool contact relationship graph, feature propagation is performed using a graph neural network (Graph Neural Network), and the node embedding vectors of the Graph Neural Network are extracted as tool-hand collaborative features. Graph Neural Network processing is the core step in transforming the relationship graph into high-level semantic features, capturing the hand-tool collaborative patterns through a message passing mechanism. The processing adopts a graph attention network architecture, which can adaptively learn the importance weights of different neighboring nodes. The network input is the hand-tool contact relationship graph. Each node initializes a feature vector. The initial features of hand nodes include geometric and kinematic information such as the 3D coordinates of key points, movement velocity, and skeletal angles. The initial features of tool nodes include attribute information such as function point coordinates, tool type embedding, and pose parameters. The graph neural network comprises multiple message-passing layers. The computation process of each layer is as follows: First, for each node, the features of all its neighboring nodes are aggregated. Attention weights are calculated using an attention mechanism, considering both node feature similarity and edge contact force weights. Then, the neighboring features are weighted and summed according to the attention weights to update the node features. After multiple propagation layers (typically 2-4 layers), the network learns node representations that integrate multi-hop neighborhood information from both the hand and tool. The features of hand nodes encode their contact patterns and mechanical relationships with the tool, while the features of tool nodes encode the collaborative configuration of multiple hand contact points. The node embedding vectors of the final layer are extracted, and all hand and tool node embeddings are concatenated or pooled to form a global tool-hand collaborative feature. This feature vector not only contains geometric spatial information but also integrates the strength of mechanical interactions and temporal evolution patterns.

[0130] At the same time, the obtained tool-hand coordination features and fine motor features can be spliced ​​and fused together to make the features expressed more complete.

[0131] One embodiment of the present invention further includes separating and enhancing micro-jitter features, the specific implementation process of which includes:

[0132] Microscopic tremor features are separated and enhanced by extracting tremor layer data from a three-level feature map and using frequency domain analysis to distinguish between physiological tremors and skillful micro-movements. Accurate processing of microscopic tremors is crucial for showcasing the subtlety of a craft, requiring the differentiation between beneficial skillful micro-movements and irrelevant physiological tremors. The processing first involves performing a Fast Fourier Transform on the tremor features, decomposing them into multiple frequency components; then, based on the frequency range, physiological tremor frequency bands and skillful micro-movement frequency bands are defined. Notch filters are applied to suppress physiological tremor frequencies. The notch filter center frequency is adjusted in real time according to the individual tremor dominant frequency, effectively removing physiological interference. The notch filter design adopts a second-order IIR notch filter structure. The adjustment process first identifies the position of the maximum energy peak in the corresponding frequency band through peak detection and uses the corresponding frequency as the individualized tremor dominant frequency. Then, during the execution of the movement, a sliding window is used to monitor the dominant frequency drift in real time. When the dominant frequency shift is detected to exceed 0.5Hz, the notch filter center frequency is dynamically updated. The filtering process is performed in the frequency domain, and the frequency component amplitude of the physiological tremor frequency band is attenuated according to the frequency response curve of the notch filter, effectively suppressing physiological tremor. Meanwhile, the technique micro-motion frequency band and other frequency bands are basically unaffected, preserving technique-related jitter information.

[0133] This study identifies and extracts the micro-motion type labels and micro-motion intensity values ​​corresponding to the frequency components of the micro-motion frequency bands. The identification process first converts the frequency components of the micro-motion frequency bands into a time-frequency representation, using a short-time Fourier transform to generate a spectrogram. The horizontal axis of the spectrogram represents time, the vertical axis represents frequency, and the color represents energy intensity, visually displaying the time-frequency evolution pattern of the micro-motion. The spectrogram is then used as a two-dimensional image input to a pre-trained micro-motion pattern recognition convolutional neural network. This network, trained on a large-scale intangible cultural heritage micro-motion dataset, can identify various typical micro-motion patterns, such as the "lifting and pressing tremor" in calligraphy, the "piercing tremor" in embroidery, the "shaping and light tapping" in pottery, and the "fine-tuning with scissors" in paper cutting. The network uses a ResNet or EfficientNet backbone architecture, extracting time-frequency pattern features through multiple convolutional layers. Finally, a fully connected classification layer outputs the probability distribution of the micro-motion type, selecting the category with the highest probability as the micro-motion type label. Simultaneously, the network includes a regression branch that outputs micro-motion intensity values; these values ​​reflect the proportion and significance of the micro-motion in the overall shaking.

[0134] Based on the micro-movement type label, standard micro-movement templates corresponding to the corresponding micro-movement type are retrieved from the pre-set intangible cultural heritage skills knowledge graph. These templates include typical time-frequency characteristics of the micro-movement type. The retrieval process first accesses the micro-movement feature library subgraph of the intangible cultural heritage skills knowledge graph. This subgraph uses micro-movement types as nodes and stores standardized feature representations of various micro-movements. The micro-movement templates in the knowledge graph are extracted and statistically established from a large number of demonstration movements by intangible cultural heritage inheritors. Each template contains typical time-frequency characteristics of that micro-movement type, specifically including: the dominant frequency range, i.e., the energy concentration range of the micro-movement in the frequency spectrum; the frequency modulation mode, i.e., the change law of the dominant frequency over time, such as linear increase, periodic oscillation, or step change; amplitude envelope characteristics, i.e., the time evolution curve of the micro-movement intensity, such as gradual increase, gradual decrease, or pulse; and phase consistency index, i.e., the phase relationship between multiple repeated micro-movements, reflecting the rhythmic stability of the movement. The templates are stored in the form of multi-dimensional feature vectors, learned from a large number of samples through a variational autoencoder, preserving key features while achieving dimensionality reduction and compression. During retrieval, precise matching is performed based on the identified micro-motion type label, and the feature vector and time-frequency representation of the corresponding template are extracted from the knowledge graph. If there are multiple subcategories of the micro-motion type in the knowledge graph (e.g., "lift-press tremor" is divided into "heavy lift-light press" and "light lift-heavy press"), secondary matching is performed based on the intensity value and frequency distribution of the current micro-motion, and the most similar subcategory template is selected. The retrieved standard template provides an authoritative reference benchmark for subsequent similarity evaluation and feature transfer.

[0135] Feature enhancement is triggered when the similarity is less than a preset similarity threshold. Similarity assessment is the decision-making step to determine whether feature optimization is needed. It determines the necessity of optimization by comparing the difference between the current micro-motion and the standard template. The calculation process first converts the frequency components of the current micro-motion into feature vectors of the same dimension as the standard template, using the same feature extraction method as when the template was constructed to ensure consistency in the feature space. Then, the cosine similarity between the current feature vector and the feature vector of the standard template is calculated. When the cosine similarity is less than the preset similarity threshold, it is determined that there is a significant difference between the current micro-motion and the standard, which may be due to motion capture errors, skill execution deviations, or data noise interference, requiring feature enhancement to correct it. The setting of the similarity threshold balances feature fidelity and standardization optimization. If the threshold is too high, it will lead to excessive intervention and damage to the original features; if the threshold is too low, it will not effectively improve the standardization of the skill.

[0136] The process involves enhancing the micro-motion features by transferring high-confidence features from the standard micro-motion template to the current micro-motion features via a feature transfer network. Feature transfer is a deep learning-based feature optimization technique that integrates typical features from the standard template while preserving the personalized features of the current micro-motion, achieving standardized enhancement. The transfer network employs an encoder-transferrer-decoder architecture. The encoder encodes the current micro-motion features and the standard template features into latent representations, respectively. The transferr identifies high-confidence feature components in the template through an attention mechanism. High-confidence features refer to feature dimensions that consistently appear in multiple expert demonstration samples and have low variance, representing the core features of this micro-motion type. The transferr calculates the differences between each dimension of the current feature and the high-confidence features of the template. For dimensions with significant differences, a transfer operation is applied, and the template features are blended into the current feature through weighted fusion. The fusion weight is inversely proportional to the similarity; the lower the similarity, the higher the fusion ratio. The decoder decodes the transferred latent representation back to the frequency domain representation, generating enhanced frequency components. The entire transfer process is optimized through end-to-end training. The loss function includes three terms: reconstruction loss, ensuring the main consistency between the enhanced features and the original features; normalization loss, ensuring the similarity between the enhanced features and the standard template is improved; and smoothing loss, ensuring that the enhancement process does not introduce unnatural abrupt changes. The strength of the enhancement operation is adaptively adjusted through similarity: slight enhancement when the similarity is close to a threshold, and strong enhancement when the similarity is far below the threshold, achieving a dynamic balance between personalization and normalization.

[0137] The enhanced skillful micro-motion characteristics and the suppressed physiological tremor characteristics are reconstructed and converted back to the time domain using an inverse Fourier transform to obtain the separated enhanced micro-jitter characteristics. Reconstruction is the restoration step from frequency domain processing back to the time domain, integrating the differentiated frequency band components to form a complete signal. The reconstruction process first merges the processed frequency components in the frequency domain, including the suppressed physiological tremor frequency band (significantly reduced energy), the enhanced skillful micro-motion frequency band (characteristics closer to the standard), and other unprocessed frequency bands (such as the very low-frequency trend term and the mid-frequency transition component). Merging is achieved by frequency domain superposition, adding the complex spectral values ​​of each frequency band at their corresponding frequency positions to form a complete processed spectrum. Then, an inverse fast Fourier transform is performed on the complete spectrum to convert it back to a time-domain jitter signal. The time-domain signal obtained by the inverse transform contains the micro-jitter characteristics after suppressing physiological tremor and enhancing skillful micro-motion. This feature removes non-skillful physiological interference and enhances the fine control actions conforming to skill specifications, more purely representing the micro-characteristics of the skill. Furthermore, the separated and enhanced micro jitter features are superimposed onto the enhanced action feature representation, making the micro-detail representation in the enhanced action feature representation more complete.

[0138] It should be further noted that the neural networks and models involved in this invention are all pre-built, and the construction process is existing technology, which will not be elaborated on in this application.

[0139] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

[0140] All formulas in this manual are dimensionless and calculated numerically. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.

[0141] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.

Claims

1. A method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning, characterized in that, include: The original motion data during the execution of intangible cultural heritage skills are acquired, including multi-view high-speed video streams, six-degree-of-freedom motion data, pressure distribution data, and electromyographic signal data. Based on the skill type characteristics of intangible cultural heritage skills, the fusion weights of each modal data in the original action data are determined, and the modal data are weighted and fused based on the fusion weights to obtain fused action data; Based on the fused motion data, motion features at different granularities are analyzed to obtain a three-level motion feature map; the three-level motion feature map includes a macroscopic posture layer, a fine operation layer, and a microscopic jitter layer. Identify and extract key operation points in the fine hand movement region within the fine operation layer, and construct a fine movement key frame sequence based on the temporal distribution of the key operation points; Based on the refined action keyframe sequence and combined with the spatiotemporal correlation between each level in the three-level action feature map, the action is refined and reconstructed to obtain an enhanced action feature representation; Based on a pre-defined intangible cultural heritage skills knowledge graph, the fine motor feature representation is semantically annotated to obtain skills semantic annotations, and digital motor reproduction is driven based on the skills semantic annotations.

2. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The process of acquiring fused motion data includes: Based on the skill type characteristics of the intangible cultural heritage skills, the modal importance prior knowledge of the corresponding skill type characteristics is retrieved from the preset intangible cultural heritage skills knowledge graph; Data preprocessing is performed on each modal data within the original action data, and basic fusion weights for each modal data are initialized based on the prior knowledge of modal importance. The signal-to-noise ratio and data integrity index of each modality data are obtained during the data acquisition process; and the basic fusion weights are dynamically adjusted based on these in order to obtain the fusion weights. The preprocessed modal data are fused based on the fusion weights to obtain fused action data.

3. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The process of constructing a three-level action feature map includes: A hierarchical motion parsing network is constructed, which includes a macroscopic pose branch, a fine operation branch, and a microscopic jitter branch. The fused motion data is input into the constructed hierarchical motion parsing network for motion parsing to obtain macroscopic posture hierarchical feature maps, fine operation hierarchical feature maps, and microscopic jitter hierarchical feature maps. Based on the macroscopic posture hierarchical feature map, the fine operation hierarchical feature map, and the microscopic jitter hierarchical feature map, correlation analysis and hierarchical organization of features at different levels are performed to obtain a three-level motion feature map.

4. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The process of constructing a sequence of fine-grained motion keyframes includes: Extract the feature sequence of the fine operation layer in the three-level action feature map, and calculate the feature change rate of each frame feature vector in the feature sequence relative to the previous frame; The time-domain derivative of the characteristic rate of change is calculated to obtain the characteristic acceleration curve, and the extreme points in the characteristic acceleration curve are obtained. Based on the spatial gradient characteristics of the pressure distribution data, the force gradient vector at each moment is calculated; and its amplitude is analyzed to obtain the force inflection point. Extract the velocity vector sequence from the preprocessed six-degree-of-freedom motion data and calculate the corresponding modulus variation curve; analyze and identify the start and end times of the high variance interval in the modulus variation curve and use them as velocity change points; The motion trajectories of the hand and tool are extracted based on the multi-view high-speed video stream, and the curvature abrupt change positions are obtained by calculating the curvature difference of the motion trajectory. The extreme points, force inflection points, velocity change points, and curvature change points are time-aligned, and a time clustering algorithm is used to aggregate key points. The aggregated key points are scored for their technical importance; and key points with technical importance scores greater than a preset importance threshold are retained as the technical key operation points. Based on the time sequence of the key operation points of the technique, extract the motion data frames at the corresponding times to construct a fine motion key frame sequence.

5. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The process of obtaining enhanced action feature representations includes: The sequence of keyframes representing the fine motion is input into a preset constraint optimization network, and a comprehensive constraint objective function is constructed. The constraint optimization network includes a spatial coordination constraint module and a temporal continuity constraint module. By minimizing the comprehensive constraint objective function, the action features in the corresponding fine action keyframe sequence are iteratively updated; and during the iterative update process, the convergence speed of the comprehensive constraint objective function is dynamically monitored; when the convergence speed is lower than the preset speed threshold, the learning rate is dynamically scaled according to the historical statistical information of the current gradient. Once the comprehensive constraint objective function converges, the optimized fine-grained action keyframe sequence features are extracted, and the non-keyframe features between keyframes are refined and reconstructed through a spatiotemporal interpolation network to obtain an enhanced action feature representation.

6. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 5, characterized in that, The process of constructing the comprehensive constraint objective function includes: In the temporal continuity constraint module, interpolation prediction is performed on the motion features between adjacent key frames in the fine motion key frame sequence; the residual between the interpolated intermediate frame features and the corresponding actual frame features in the fused motion data is calculated, and a temporal continuity loss function is constructed based on it. In the spatial coordination constraint module, a spatial coordination rule base is constructed based on human kinematic constraints and tool operation constraints; For each frame in the fine motion keyframe sequence, the corresponding skeletal joint position and tool position are extracted; and it is verified whether the skeletal joint position and tool position satisfy the rules in the spatial coordination rule base; if not, the degree of violation of the rules in the corresponding keyframe is quantified and used as a spatial coordination penalty item. The temporal continuity loss function and the spatial consistency penalty term are weighted and combined to construct a comprehensive constraint objective function.

7. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The process of obtaining semantic annotations for skills includes: Extract a set of action category tags related to the actions of intangible cultural heritage skills from the knowledge graph of intangible cultural heritage skills, and encode the set of action category tags into a semantic embedding vector; The enhanced action feature representation is input into a preset skill semantic understanding model, and the enhanced action feature representation is mapped from the action feature space to the semantic feature space through the cross-modal alignment layer in the skill semantic understanding model to obtain action semantic features. Calculate the similarity between the semantic features of the action and the semantic embedding vectors of each action category label, and select the top K action category labels with the highest similarity as candidate semantic labels, where K is an integer representing the number of candidates; Calculate the action specification compliance degree of each candidate semantic label based on the action specification standards within the intangible cultural heritage skills knowledge graph; Based on the weighted score of similarity and action specification compliance, the candidate semantic labels are sorted according to the weighted score, and the label with the highest weighted score is selected as the skill semantic label. Read the cultural semantic associations in the knowledge graph of intangible cultural heritage skills, retrieve the cultural semantic association content associated with the semantic annotation of skills, and attach it as an auxiliary annotation to the semantic annotation of skills.

8. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The process of digital motion reproduction driven by semantic annotation of the aforementioned techniques includes: Based on the semantic annotation of the technique, the corresponding driving parameter template is retrieved from the preset action driving parameter library; and the enhanced action feature representation is simultaneously converted into a skeletal joint angle sequence. Based on the aforementioned driving parameter template and skeletal joint angle sequence, intangible cultural heritage skills movements are performed in a pre-constructed digital character model; and during the movement execution, the position and posture changes of the virtual tool are synchronously driven according to the pre-acquired tool-hand coordination features. Simultaneously, the micro-shaking features are used as high-frequency disturbance signals and superimposed on the key points of the digital character model. Micro-shaking effects are generated through procedural animation technology, and the motion reproduction results based on the digital character model are output.

9. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 8, characterized in that, The process of obtaining tool-hand collaborative features includes: The hand region and tool region are extracted from the multi-view high-speed video stream, respectively. The three-dimensional coordinates of key points in the hand area are extracted, and a topological map of the hand skeleton is constructed based on them. At the same time, based on the types of tools required in the execution of intangible cultural heritage skills, corresponding 3D models of tools are constructed, and the position and orientation of the 3D models of tools in the world coordinate system are calculated using six-degree-of-freedom motion data. Mark functional key points on the tool's 3D model and extract the 3D coordinates of the functional key points; Furthermore, a hand-tool contact relationship diagram is constructed based on the hand skeleton topology diagram and the tool 3D model; the hand-tool contact relationship diagram consists of nodes composed of key points of the hand and functional key points, and contact edges between nodes; The contact force intensity of each contact edge is calculated based on the pressure distribution data, and the contact force intensity is used as the weight of the contact edge. For the hand-tool contact relationship graph, feature propagation is performed through a graph neural network, and the node embedding vectors of the graph neural network are extracted as tool-hand collaborative features.

10. The method for motion capture and digital reproduction of intangible cultural heritage skills based on deep learning according to claim 1, characterized in that, The method further includes separating and enhancing micro-jitter features, including: The shaking feature data within the micro-shaking layer is extracted from the three-level motion feature map, and the shaking feature data includes physiological hand tremors and skillful micro-movements; The jitter feature data is decomposed into multiple frequency components through frequency domain analysis, and the corresponding frequency components are divided into physiological tremor frequency band and skillful micro-motion frequency band according to the frequency range. The frequency components of the physiological tremor band are suppressed by using a notch filter; at the same time, the micro-motion type label and micro-motion intensity value corresponding to the frequency components of the technical micro-motion band are identified and extracted. Based on the micro-motion type label, a standard micro-motion template corresponding to the micro-motion type is retrieved from a preset intangible cultural heritage skills knowledge graph. The standard micro-motion template includes typical time-frequency features of the micro-motion type. Calculate the feature similarity between the frequency components of the technical micro-motion band and the standard micro-motion template; when the feature similarity is lower than a preset similarity threshold, perform a micro-motion feature enhancement operation, which includes: transferring the high-confidence features of the standard micro-motion template to the current micro-motion feature through a feature transfer network; The enhanced skillful micro-movement features and the suppressed physiological tremor features are recombined to obtain the separated enhanced micro-shaking features, which are then superimposed on the enhanced motion feature representation.