Lion dancing gymnastics action identification method

Through multi-source data acquisition and deep fusion technology, combined with dynamic environment adaptation and real-time feedback, the problems of low accuracy and low training efficiency of traditional lion dance movement recognition in complex scenarios are solved, and efficient and safe lion dance movement recognition and training are achieved.

CN120544280AInactive Publication Date: 2025-08-26GUANGDONG OCEAN UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510726799.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-03
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The teaching and evaluation of traditional lion dance moves relies on manual observation, which is highly subjective and inefficient. The ability to generalize in complex scenarios is poor, and the lack of comprehensive analysis of the venue environment and drum beat rhythm is difficult to support action innovation and safety assessment, resulting in low training efficiency and high risk.

Method used

Lidar, RGBD camera, microphone and wearable sensor are used to synchronize multi-source data, combine dual-stream network and cross-modal attention mechanism to integrate feature, dynamically adjust the joint weight of the pose estimation model, use AR devices to feedback action deviations in real time, and generate new actions through meta-learning and generative adversarial networks, and generate personalized correction strategies in combination with reinforcement learning.

Benefits of technology

It improves the accuracy and robustness of lion dance movement recognition, enhances the generalization ability in complex scenarios, improves training efficiency and safety, reduces the risks of action innovation and development, and realizes real-time feedback and intelligent training of action recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120544280A_ABST
    Figure CN120544280A_ABST
Patent Text Reader

Abstract

The invention discloses a lion dance gymnastics action recognition method, and relates to the technical field of artificial intelligence, the method comprises the steps: synchronously collecting a site environment, a visual image, drumbeat audio and joint motion data through a laser radar, an RGBD camera, a microphone and a wearable sensor, and generating an environment parameter matrix through a YOLOv8 algorithm; a double-flow network is constructed through a space-time diagram convolutional network (ST-GCN) and a Transform encoder, and multi-modal feature fusion is realized in combination with a cross-modal attention mechanism. And dynamically adjusting the weight of the attitude estimation model based on the environmental parameters, and completing action classification through an LSTM network. The system integrates AR real-time deviation feedback, meta-learning driven new action generation and a physical engine safety evaluation module, and realizes high-precision identification of lion dancing actions, training efficiency optimization and innovation assistance. According to the method, the problem of insufficient adaptability of a traditional method in a complex scene is solved, the scientific training level and culture inheritance efficiency of lion dancing sports are improved, and the method has the advantages of being high in precision and intelligent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method for recognizing lion dance gymnastics movements. Background Art

[0002] The teaching and evaluation of traditional lion dance movements relies on manual observation, which is subject to high subjectivity, low efficiency, and insufficient capture of movement details. Existing single-modality motion recognition methods (e.g., relying solely on cameras or inertial sensors) have poor generalization capabilities in complex scenarios (such as high-stakes performances and multi-person collaboration) and lack comprehensive analysis of key factors such as the venue environment and drum rhythm. Furthermore, traditional methods struggle to support innovative design and safety assessment of lion dance movements, resulting in long development cycles and high risks for new movements. Therefore, there is an urgent need for lion dance motion recognition technology that integrates multi-source data and adapts to dynamic environments to improve training efficiency and promote technical innovation.

[0003] In view of this, this application is hereby filed. Summary of the Invention

[0004] The object of the present invention is to provide a method for recognizing lion dance gymnastics movements to solve the problems mentioned in the above background technology.

[0005] To solve the above technical problems, the present invention provides a method for recognizing lion dance gymnastics movements, comprising the following steps: S1 Environmental and Motion Data Collection: Utilizing LiDAR, RGBD cameras, microphones, and wearable sensors, the system simultaneously collects 3D environmental data of the lion dance performance venue, visual image data of the lion dancers, drum beat audio data, and joint motion data. This multi-source heterogeneous data collection allows for comprehensive spatial, temporal, and environmental information about lion dance movements. Compared to single-modality data collection, this effectively improves the accuracy and robustness of motion recognition, laying the foundation for subsequent precise analysis. S2 Site Environmental Parameter Extraction: The three-dimensional environmental data is input into the target detection algorithm to identify site obstacles such as high piles and steel wires, and generate an environmental parameter matrix. This step realizes the digital modeling of the complex lion dance site, allowing the action recognition to be dynamically adjusted according to the actual site conditions, solving the problem of poor adaptability of traditional methods in different sites and improving the system's recognition ability in complex scenes. S3 multimodal feature fusion processing: A dual-stream network is used to extract the spatial joint relationship features and temporal action sequence features of the visual image data, and the visual image data features are fused with the drum audio data features and joint motion data features through a cross-modal attention mechanism. The dual-stream network and cross-modal attention mechanism can fully mine the feature information of each modal data and achieve effective fusion between features. Compared with a single modality or simple fusion method, the expressive ability of action features is greatly improved, thereby improving the accuracy and efficiency of action recognition. S4 Dynamic Adjustment and Action Recognition: Dynamically adjust the joint weights of the posture estimation model according to the environmental parameter matrix, perform temporal classification on the fused features, and identify lion dance gymnastics movements; dynamically adjust the model weights according to the venue environment to make action recognition more in line with the actual performance scene, avoid recognition errors caused by venue changes, significantly improve the accuracy and generalization ability of action recognition, and adapt to different lion dance performance environments.

[0006] Furthermore, in the S3 multimodal feature fusion processing, the dual-stream network includes: a spatial branch of a spatiotemporal graph convolutional network for modeling joint spatial relationships; a temporal branch of a Transformer encoder for capturing long-range dependencies of action sequences; the combination of a spatiotemporal graph convolutional network and a Transformer encoder can deeply extract and analyze action features from the spatial and temporal dimensions, effectively capturing the complex joint relationships and time series features in lion dance movements, further improving the extraction effect of action features, and thus improving the accuracy and reliability of action recognition.

[0007] Furthermore, in the S3 multimodal feature fusion processing, the visual image data features are fused with the drum audio data features and the joint motion data features, specifically: bottom-level fusion: splicing the joint coordinates with the drum spectrum features; middle-level fusion: aligning the visual-audio features through the cross-modal attention mechanism; high-level fusion: extracting the action semantic features using the capsule network; the hierarchical fusion method can gradually integrate the information of each modal data from the bottom to the top, making the fused features more semantic and representative. Compared with simple feature splicing, it can better retain and utilize the effective information of each modal data, thereby enhancing the performance of the action recognition model and the ability to understand complex actions.

[0008] Furthermore, it also includes the S5 action deviation real-time feedback step: the virtual action template is superimposed on the real scene through the AR device, and the deviation between the lion dancer's action and the virtual action template is displayed in real time. When the deviation exceeds the preset threshold, an early warning is triggered; this step realizes real-time visual feedback of lion dance action training. The lion dancer can timely understand the difference between his own action and the standard action, which facilitates rapid correction of errors. Compared with traditional training methods, it greatly improves training efficiency, reduces training costs, and helps to improve the skill level of lion dancers.

[0009] Furthermore, in the S5 real-time feedback step of movement deviation, the preset threshold includes one of: the lion head offset angle is greater than 5°, and the joint movement trajectory deviation distance is greater than 3 cm; the clear preset threshold standard provides a quantitative basis for the judgment of movement deviation, making movement correction more targeted and accurate, avoiding the ambiguity of feedback information, and further improving the effectiveness and scientific nature of training guidance.

[0010] Furthermore, it also includes the S6 new movement generation and evaluation step: constructing a meta-training set containing basic movements, using a model-independent meta-learning algorithm to quickly adapt to new movement categories, and using a generative adversarial network to generate difficult movements that conform to the laws of lion dance mechanics; through the meta-learning algorithm and the generative adversarial network, new lion dance movements can be quickly generated with a small amount of data, providing technical support for the innovation of lion dance movements, enriching the content and form of lion dance performances, and promoting the inheritance and development of lion dance skills.

[0011] Furthermore, in the S6 new movement generation and evaluation step, the physical engine is used to simulate joint forces, and the feasibility of the generated difficult movements is evaluated. When the knee joint torque exceeds the threshold, an alarm is triggered; the physical engine simulation and threshold alarm mechanism can effectively evaluate the safety of the newly generated movements, preventing lion dancers from being injured due to unreasonable movements during training and performances, while ensuring the safety of lion dancers, and ensuring the practicality and operability of the new movements.

[0012] Furthermore, it also includes the S7 correction strategy generation and evaluation step: using the reinforcement learning framework, the correction strategy is automatically generated according to the motion recognition results, and the improvement effect is evaluated through the dynamic time warping algorithm; the application of the reinforcement learning framework and the dynamic time warping algorithm can intelligently generate personalized correction strategies according to the motion recognition results, and quantitatively evaluate the training effect. Compared with the traditional empirical training method, it realizes the intelligence and scientificization of the training process, and further improves the effect and efficiency of the lion dancer's motion training.

[0013] Compared with the prior art, the present invention has the following beneficial effects: Multimodal fusion improves recognition accuracy: Through the simultaneous collection of environmental, visual, audio and motion data through lidar, RGBD cameras, microphones and wearable sensors, combined with a dual-stream network and cross-modal attention mechanism, multi-level fusion of motion features is achieved, which improves the recognition accuracy by 7%-10% compared to single-modality recognition, and maintains a recognition rate of ≥92% in complex scenarios.

[0014] Dynamic environment adaptability: Based on the environmental parameter matrix generated by site obstacle identification, the joint weights of the posture estimation model are dynamically adjusted to address the interference of different sites (such as changes in the spacing of plum blossom piles and lighting fluctuations) on action recognition, and the generalization ability is significantly enhanced.

[0015] Real-time feedback and intelligent training: AR devices are used to display movement deviations in real time and trigger early warnings, combined with reinforcement learning to generate personalized correction strategies, improving training efficiency by more than 40%. The dynamic time warping algorithm quantitatively evaluates training effects, increasing the degree of movement standardization to 95%.

[0016] Movement innovation and safety assurance: Based on meta-learning and generative adversarial networks, only a small number of samples are needed to generate highly difficult new movements, and the physical engine is used to simulate joint forces to assess safety, reducing the risk of developing innovative movements and promoting the digital inheritance of lion dance skills. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 The figure is a flow chart of a method for recognizing lion dance gymnastics movements. DETAILED DESCRIPTION

[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0019] See also Figure 1 The present invention provides a technical solution: a method for recognizing lion dance gymnastics movements, comprising: The application scenario is a traditional Southern Lion performance.

[0020] 1. S1 environment and motion data collection Two Velodyne VLP-16 lidar sensors were deployed around the lion dance performance area, one at each corner, to scan the three-dimensional space. Four RGBD cameras (Intel RealSense D455) were deployed, two located in front of the performance area and two on the side, to collect visual image data from the lion dancers. A high-sensitivity microphone (Shure SM7B) was installed near the drumhead to capture drumbeat audio data. Five wearable inertial sensors (Xsens MTi-10) were attached to the dancers' waist, shoulders, and knees to capture joint motion data. All sensors were synchronized for data acquisition using a synchronous trigger device, with a sampling frequency of 100 Hz.

[0021] For example, during a performance by a lion dance team in a certain county's square, the lidar scan revealed a set of plum blossom piles in the center of the venue. The RGBD camera captured the lion dancers dancing on the plum blossom piles, the microphone picked up the passionate drum beats, and the wearable sensor recorded in real time the movement data of each joint of the lion dancers as they jumped and turned.

[0022] 2.S2 site environmental parameter extraction The point cloud data collected by the lidar is fed into an improved object detection algorithm based on YOLOv8. This algorithm has been pre-trained to recognize obstacles such as high stakes, steel wires, and stairs in the lion dance arena. The algorithm identifies parameters such as the number, location, and height of the poles, as well as the boundaries of the surrounding spectator area, and generates an environmental parameter matrix containing the location coordinates and dimensions of the obstacles.

[0023] For example, through algorithm processing, the environmental parameter matrix shows that there are 9 plum blossom piles arranged in a triangle with a height of 1.2 meters. The distance between adjacent plum blossom piles is between 0.8-1.0 meters. At the same time, it is determined that the audience area is closest to the plum blossom piles at 2 meters to avoid subsequent action recognition being interfered with by the audience.

[0024] 3.S3 multimodal feature fusion processing A two-stream network is used for feature extraction. The spatial branch uses a spatiotemporal graph convolutional network (ST-GCN) to construct a joint graph from the coordinates of 25 key points of the human body captured by the RGBD camera. Ten layers of spatiotemporal graph convolution operations are then used to extract spatial relationship features of the joints. The temporal branch uses a Transformer encoder, which inputs a sequence of key point coordinates from 10 consecutive frames. Eight multi-head attention layers are then used to capture the long-range dependencies of the action sequence. Subsequently, a cross-modal attention mechanism is used to fuse the visual image features with the short-time Fourier transform-processed drum beat audio spectrum and the joint acceleration and angular velocity features collected by wearable sensors. The fusion process is divided into three layers: the bottom layer directly concatenates the joint coordinates with the drum beat spectrum; the middle layer uses a cross-modal attention mechanism to align the visual and audio features; and the high-level capsule network encodes the fused features into a semantic feature vector for the action.

[0025] For example, when a lion dancer completes the "lion climbing the pole" action, the spatial branch captures the spatial position changes and joint angle relationships of the lion's head and tail on the pole, while the temporal branch analyzes the complete timing characteristics of the action from taking off, landing on the pole to standing firm. At the same time, the drum beat rhythm changes from slow to fast, matching the rhythm of the action. Through the cross-modal attention mechanism, the visual, audio, and sensor data are fused to form a feature vector that contains the semantics of the "lion climbing the pole" action.

[0026] 4.S4 dynamic adjustment and motion recognition Based on information such as the height and spacing of the poles in the environmental parameter matrix, the weight coefficients of joints such as the knee and ankle in the posture estimation model are dynamically adjusted. For example, the weight of the knee flexion angle is increased by 20% to more accurately identify movements on high poles. The fused feature vector is input into an LSTM-based time series classification model, which has been pre-trained on 50 common lion dance movement categories. The model outputs movement category probabilities using a softmax function. When the probability of the "Lion on the Pole" movement category exceeds 0.85, the current lion dance movement is determined to be "Lion on the Pole."

[0027] For example, when performing the aforementioned "Lion on the Stake" action, due to the narrow spacing between the plum blossom stakes, the model automatically increases the joint weights to more accurately analyze the lion dancer's balancing movements; the LSTM model outputs a probability of "Lion on the Stake" of 0.92, accurately identifying the action.

[0028] 5. Real-time feedback of S5 action deviation Using the Hololens 2 AR device, the standard "Lion on a Stake" movement template is superimposed as a virtual holographic image onto the actual lion dancer's movements. The deviation between the key points of the lion dancer's movements and those of the standard template is calculated in real time. If the lion's head deviates by more than 5°, or the knee joint's trajectory deviates by more than 3cm, the AR device issues a voice prompt, stating "Lion's head angle needs adjustment" or "Knee position deviation," and marks the deviation with a red line in the field of view.

[0029] For example, when a lion dancer first attempts to perform the "lion on a pole" pose, the AR device detects that the lion's head has deviated by 7°, immediately issues a voice prompt, and uses red lines to outline the direction and angle that the lion's head should be adjusted to in the field of view. The lion dancer can then correct the action in a timely manner based on the prompt.

[0030] 6.S6 New Action Generation and Evaluation A meta-training set of 30 basic lion dance movements was constructed and trained using the MAML (Model-Agnostic Meta-Learning) algorithm. When generating new movements, a basic movement sequence is input as a condition, and a generative adversarial network (GAN) is used to generate a more challenging movement sequence. After the movement is generated, the joint motion data is fed into a physics engine (such as Unity's PhysX engine) to simulate the forces acting on the knee and hip joints. If the knee torque exceeds 30 N·m, the movement is identified as a safety risk and an alarm is triggered.

[0031] For example, to innovate lion dance performances, the system used the two basic movements of "lion leaping" and "lion turning" as conditions, and GAN generated a new movement of "lion continuous somersaults on the pole." After simulation with the physics engine, it was found that this movement caused the knee joint torque to reach 45N·m. The system issued an alarm, indicating that this movement needed to be optimized to ensure the safety of the lion dancer.

[0032] 7.S7 Correction Strategy Generation and Evaluation Using the Deep Q-Network (DQN) reinforcement learning framework, a personalized correction strategy is generated based on deviations between action recognition results and standard movements. For example, if a lion dancer's hand extension during the "Cai Qing" movement is detected to be insufficient, a strategy is generated: "Perform hand stretching exercises, three sets per day, 10 repetitions each." Using the Dynamic Time Warping (DTW) algorithm, the action sequences before and after the correction strategy are compared, and the similarity of the movements is calculated to evaluate the improvement effect.

[0033] For example, when a lion dancer was practicing the "Cai Qing" movement, the system recognized that the angle of his hand extension was 15° less than the standard movement and generated the above-mentioned correction strategy. After a week of training, the movement data was collected again, and the DTW algorithm was used to calculate that the similarity between the improved movement and the standard movement increased from 70% to 85%, effectively evaluating the training effect.

[0034] In summary, the present invention constructs a three-dimensional venue model through the fusion of lidar and RGBD cameras and dynamically adjusts the joint weights of the pose estimation model, breaking through the environmental adaptability limitations of traditional single-modal recognition. A spatiotemporal graph convolutional network and a Transformer encoder form a dual-stream network, combining a cross-modal attention mechanism with a capsule network to achieve bottom-level splicing, mid-level alignment, and high-level semantic feature extraction of visual, audio, and motion data. Compared with simple feature fusion, this method more accurately captures the spatial joint relationships and temporal dependencies of lion dance movements. An AR device is introduced to superimpose virtual action templates to achieve real-time warnings when the lion head offset angle exceeds 5° or the joint trajectory deviation exceeds 3cm, establishing a "recognition-feedback-correction" closed loop. The MAML meta-learning algorithm combined with a generative adversarial network can generate new movements with only a small number of basic movement samples, and a physics engine is used to simulate joint forces (for example, a knee joint torque exceeding 30N·m triggers an alarm) to assess safety. These unconventional technical means, through multi-layer innovations such as multimodal deep fusion, dynamic environment adaptation, real-time AR intervention, and small-sample intelligent generation, overcome the technical bottlenecks of traditional methods such as low recognition accuracy in complex scenes, delayed training feedback, and reliance on experience for movement innovation.

Claims

1. A method for recognizing lion dance gymnastics movements, characterized in that: The following steps are involved: S1 Environmental and Motion Data Collection: Using LiDAR, RGBD cameras, microphones, and wearable sensors, it simultaneously collects 3D environmental data of the lion dance performance venue, visual image data of the lion dancers, drum beat audio data, and joint motion data. S2 site environment parameter extraction: input the three-dimensional environment data into the target detection algorithm, identify high piles and steel wire site obstacles, and generate an environment parameter matrix; S3 multimodal feature fusion processing: A two-stream network is used to extract the spatial joint relationship features and temporal action sequence features of the visual image data, and the visual image data features are fused with the drum audio data features and joint motion data features through a cross-modal attention mechanism; S4 dynamic adjustment and action recognition: dynamically adjust the joint weights of the posture estimation model according to the environmental parameter matrix, perform time series classification on the fused features, and recognize lion dance gymnastics movements.

2. A method for recognizing lion dance gymnastics movements as claimed in claim 1, characterized in that: In the S3 multimodal feature fusion processing, the dual-stream network includes: a spatial branch of a spatiotemporal graph convolutional network for modeling joint spatial relationships; and a temporal branch of a Transformer encoder for capturing long-range dependencies in action sequences.

3. A lion dance gymnastics movement recognition method as claimed in claim 1, characterized in that: In the S3 multimodal feature fusion process, the visual image data features are fused with the drum beat audio data features and the joint motion data features, specifically as follows: bottom-level fusion: splicing the joint coordinates with the drum beat spectrum features; Mid-layer fusion: aligning visual-audio features via cross-modal attention mechanism; High-level fusion: Using capsule network to extract action semantic features.

4. A method for recognizing lion dance gymnastics movements as claimed in claim 1, characterized in that: It also includes the S5 action deviation real-time feedback step: the virtual action template is superimposed on the real scene through the AR device, and the deviation between the lion dancer's action and the virtual action template is displayed in real time. When the deviation exceeds the preset threshold, an early warning is triggered.

5. A method for recognizing lion dance gymnastics movements as claimed in claim 4, characterized in that: In the step of real-time feedback of the movement deviation in S5, the preset threshold value includes one of: the lion head deviation angle is greater than 5°, and the joint movement trajectory deviation distance is greater than 3 cm.

6. A method for recognizing lion dance gymnastics movements as claimed in claim 1, characterized in that: It also includes the S6 new movement generation and evaluation step: building a meta-training set containing basic movements, using a model-independent meta-learning algorithm to quickly adapt to new movement categories, and using a generative adversarial network to generate difficult movements that conform to the laws of lion dance mechanics.

7. A method for recognizing lion dance gymnastics movements as claimed in claim 6, characterized in that: In the S6 new action generation and evaluation step, a physical engine is used to simulate joint forces, and the feasibility of the generated difficult actions is evaluated. When the knee joint torque exceeds a threshold, an alarm is triggered.

8. A method for recognizing lion dance gymnastics movements as claimed in claim 1, characterized in that: It also includes the S7 correction strategy generation and evaluation steps: using the reinforcement learning framework, the correction strategy is automatically generated based on the action recognition results, and the improvement effect is evaluated through the dynamic time warping algorithm.