Teaching implementation method based on combination of motion perception and AI driving
Through multimodal data collection and AI-driven teaching methods, the problems of high hardware cost, limited scenarios and lack of real-time performance in action-related subject teaching have been solved, cross-disciplinary adaptation and personalized feedback have been achieved, and teaching efficiency and applicability have been improved.
Patent Information
- Application Number
- CN202510903342.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-01
- Publication Date
- 2025-10-03
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing action-based subject teaching has problems such as high hardware costs, limited scenarios, insufficient real-time performance, poor interdisciplinary adaptability, and lack of personalized feedback mechanisms. It is difficult to reuse in different fields and lacks real-time interaction and personalized evaluation.
Using multimodal data acquisition technology and combining it with AI-driven teaching methods, we collect motion, physiological, and environmental data through ordinary mobile phone cameras and built-in sensors, use the MediaPipe posture estimation algorithm and DBSCAN algorithm for motion evaluation, and build a cross-scenario learning graph combined with graph neural networks to provide personalized feedback and cross-disciplinary adaptation, thus achieving cross-scenario data interoperability and system optimization.
It reduces hardware costs, achieves full-scene coverage, improves real-time performance to millisecond level, provides personalized action teaching feedback and cross-disciplinary adaptation capabilities, and shortens the time to get started with new subjects.
Smart Images

Figure CN120745931A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of education and teaching technology, and specifically to a teaching implementation method based on the combination of motion perception and AI drive. Background Art
[0002] There are many technical bottlenecks in the current teaching of movement subjects: reliance on dedicated hardware leads to high costs and limited scenarios; lack of real-time performance; most systems only support post-playback analysis; poor interdisciplinary adaptability and difficulty in reuse in different fields such as dance, sports, and experimental operations; lack of personalized feedback mechanisms and inability to provide dynamic evaluation based on individual student characteristics; After searching, the applicant discovered that a Chinese patent disclosed "A dance teaching auxiliary method and device", with the publication number "CN112309181A". This patent mainly collects the teacher's dance movements, obtains reference dance videos, obtains standard standing pictures of the dance teacher, constructs an initial picture set, and reconstructs the dance teacher in three dimensions based on the elements in the initial picture set and the MATLAB algorithm. The reference dance video is passed into a preset spatiotemporal feature extraction model to identify the movement changes of different parts of the teacher in the dance video. Based on the recording time in the reference dance video, the movement changes of different parts of the human body at the same time are combined to obtain a reference movement unit. The reference movement units corresponding to different times in the reference dance video are obtained in turn, and the reference movement units corresponding to different times are displayed in turn on the preset teaching display interface as teaching auxiliary movements. However, this teaching method relies on pre-recorded videos by the teacher and lacks real-time interaction. To this end, we propose a teaching implementation method based on the combination of motion perception and AI drive. Summary of the Invention
[0003] The purpose of this invention is to provide a teaching implementation method based on the combination of motion perception and AI drive.
[0004] To achieve the above objectives, the present invention provides the following technical solutions: a teaching implementation method based on the combination of motion perception and AI drive, including the implementation of a teaching method, wherein the teaching method implementation includes multimodal data expansion acquisition, cross-dimensional dynamic assessment and cognitive diagnosis, emotional adaptive feedback and intervention, cross-scenario learning map construction and ability transfer, and full-process data closure and system evolution. The multimodal data expansion acquisition includes dynamic data acquisition, physiological data acquisition, environmental behavior data acquisition, and data preprocessing. The specific steps of the teaching implementation method based on the combination of motion perception and AI drive are as follows: Step 1: Collect motion data, physiological state data, and environmental and behavioral data, unify the timestamps of the three types of data, filter out outliers, and process the collected data; Step 2: Evaluate the action execution layer and the physiological state layer, diagnose the cognitive ability layer, and generate a comprehensive ability score for the three layers for real-time analysis; Step 3: Identify and categorize students' emotional states, provide intelligent feedback based on vision, hearing, and touch, and generate personalized intervention strategies, including progressive training paths and cognitive reinforcement training, and customize feedback preferences according to the five-level feedback mechanism; Step 4: Unify the data standard system to conduct cross-scenario correlation analysis, personalized learning path planning, and teaching strategy optimization; Step 5: Analyze the data, optimize teaching strategies based on the collected data, communicate data across scenarios, and optimize system performance.
[0005] As a further solution of the present invention: the motion data collection captures skeletal key points, uses a camera with a resolution of ≥720P, and combines the MediaPipe posture estimation algorithm to obtain the coordinates of 21 skeletal key points of the human body in real time, with a sampling frame rate of ≥30fps, covering the motion characteristics of dance joint angles and sports force trajectories. The physiological state data collection includes eye movement trajectory detection, heart rate variability collection and real-time analysis of micro-expressions. The environmental and behavioral data collection includes audio and text data and environmental parameter detection. After the data collection is completed, the motion, physiological and audio data are unified with a timestamp, and the error is controlled to be ≤50ms. A synchronous trigger is used to achieve multimodal data alignment, and the 3σ principle is used to eliminate mutation points in the motion trajectory. The heart rate data is smoothed with a sliding window algorithm to eliminate motion artifact interference.
[0006] As a further solution of the present invention: the action execution layer evaluation includes standard quantification and error pattern clustering, calculation of joint angle compliance rate, use of DBSCAN algorithm to classify action errors, generation of dance scene error action heat map, and centralized review and comparison of common problems; the physiological state layer evaluation includes fatigue index calculation and attention concentration calculation; the cognitive ability layer diagnosis includes cognitive dimension mapping and three-dimensional ability matrix, dance action memory error → associated short-term memory capacity, experimental step omission → associated executive dysfunction, three-dimensional ability matrix to construct "spatial perception-working memory-logical reasoning" radar chart, based on 5000+ sample training (using a three-layer fully connected neural network (input layer: 21 bone point coordinates + 6-dimensional physiological indicators, hidden layer: 128 ReLU units, output layer: 3-dimensional cognitive score), loss function is mean square error) association model, input action error data to automatically generate cognitive ability evaluation.
[0007] As a further solution of the present invention: the comprehensive evaluation model includes a weighted score of 50% action execution layer + 30% physiological state layer + 20% cognitive ability layer.
[0008] As a further solution of the present invention: the visual feedback is AR superimposed with a translucent blue standard action skeleton and red real-time action, and joint deviation is highlighted when it is greater than 5°; the auditory feedback is to convert the key points of the action into language instructions, and the instruction speed is dynamically adjusted according to the student's heart rate; the tactile feedback is the vibration feedback force deviation of the mobile phone, and the vibration frequency is positively correlated with the degree of error; the personalized intervention strategy includes progressive training paths and cognitive reinforcement training; the five-level feedback mechanism includes silent observation, slight prompts, active suggestions, forced intervention and emergency pause, students customize feedback preferences, and the system dynamically adjusts the intervention strategy according to the configuration.
[0009] As a further solution of the present invention: the unified data standard system is divided into a capability label system and error coding rules. The capability label system defines 128 basic capability labels, and the error coding rules establish unified error coding. The cross-scenario association analysis includes graph neural network modeling, and uses a graph convolutional network (GCN) to construct a capability association graph. Node features are extracted through a spatiotemporal graph convolution module. The nodes are actions in each scenario, and the edge weights represent the strength of capability migration. When the scenario capability is improved, the capability improvement probability of the associated scenario is automatically calculated.
[0010] As a further solution of the present invention: the personalized learning path planning includes efficient migration path recommendation and full-cycle capability tracking; the teaching strategy optimization includes teacher-side decision support and system self-optimization mechanism; the teacher-side outputs a class capability shortcoming report and recommends an interdisciplinary integrated teaching plan; the system self-optimization mechanism continuously iterates the capability association model based on 100,000+ student data and updates the migration coefficient threshold every quarter.
[0011] As a further solution of the present invention: the data precipitation and analysis automatically generates a holographic learning archive containing action records, cognitive assessments and emotional feedback, which is classified and archived according to subject-action-cognitive roots, supporting teachers to retrieve typical errors. The teaching strategy iteration is divided into AI-driven optimization and teacher-system collaboration. AI drives the analysis of the characteristics of outstanding students, generates new teaching standards according to their characteristics, teachers mark AI assessment deviation cases, and the system records and strengthens learning to correct the emotion recognition model.
[0012] As a further solution of the present invention: the cross-scene data intercommunication includes capability matrix synchronization and multi-terminal data synchronization, which displays the capability migration trajectory of dance, experiment and sports venue data under a unified coordinate. The multi-terminal data supports real-time synchronization of cross-terminal data of mobile phones, teacher large screens and VR devices, with a delay of <200ms. The system performance optimization includes lightweight deployment and adaptive learning. The lightweight deployment uses model distillation technology to compress the AI engine, and achieves action recognition delay of <100ms on mobile phones. The system automatically adjusts the algorithm complexity according to the device performance.
[0013] By adopting the above technical solution, compared with the prior art, the beneficial effects of the present invention are: 1. This invention uses a multimodal data acquisition solution using an ordinary mobile phone camera and built-in sensors to replace the traditional dance teaching method that relies on dedicated 3D modeling equipment and sensor gloves for boxing teaching. This reduces hardware costs and can achieve full coverage in homes, classrooms, and outdoors without the need for professional site deployment. This solves the core problems of existing technologies, namely "high hardware costs and limited scenarios", and enables movement teaching to move from professional venues to daily learning and life. 2. This invention leverages millisecond-level multimodal data fusion technology and AR real-time feedback to upgrade the traditional solution's "post-playback analysis" to "synchronous diagnostic intervention." In dance instruction, when a student's joint angle deviation >5° is detected, the AR system immediately highlights the error area and pushes a decomposition training plan. Compared to existing technologies, feedback latency is shortened from minutes to less than 100ms, resolving the technical pain point of "lack of real-time performance" and improving the efficiency of teaching intervention. 3. The present invention uses a system of 128 basic ability labels and graph neural network modeling to achieve plug-in reuse of evaluation rules for more than 10 subjects such as dance, experiments, and physical education, thereby improving the reuse rate of single-field models. At the same time, based on a dynamic cognitive diagnosis model, movement errors are accurately mapped to cognitive dimensions such as working memory and spatial perception, generating a personalized plan of "ability shortcomings-training path-migration prediction". The system can automatically recommend an efficient migration path of "chemical experiment operation precision training → dance equipment holding stability improvement", shortening the introduction time to new subjects and solving the problem of "poor interdisciplinary adaptation and lack of personalized feedback". BRIEF DESCRIPTION OF THE DRAWINGS
[0014] Figure 1 This is a general flow chart of the implementation method in the embodiment of the present invention. DETAILED DESCRIPTION
[0015] The specific embodiments of the present invention will be further described below in conjunction with the accompanying drawings. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0016] In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0017] Please see the attached Figure 1The present invention provides a teaching implementation method based on the combination of motion perception and AI drive, including the implementation of the teaching method, which includes multimodal data expansion acquisition, cross-dimensional dynamic assessment and cognitive diagnosis, emotional adaptive feedback and intervention, cross-scenario learning map construction and ability transfer, and full-process data closed loop and system evolution. The multimodal data expansion acquisition includes dynamic data acquisition, physiological data acquisition, environmental behavior data acquisition and data preprocessing. The specific steps of the teaching implementation method based on the combination of motion perception and AI drive are as follows: Step 1: Collect motion data, physiological state data, and environmental and behavioral data, unify the timestamps of the three types of data, filter out outliers, and process the collected data; Step 2: Evaluate the action execution layer and the physiological state layer, diagnose the cognitive ability layer, and generate a comprehensive ability score for the three layers for real-time analysis; The basis for comprehensive ability scoring is: Comprehensive score S=0.5×A+0.3×B+0.2×C, where A is the score of action execution layer, B is the score of physiological state layer, and C is the score of cognitive ability layer. Step 3: Identify and categorize students' emotional states, provide intelligent feedback based on vision, hearing, and touch, and generate personalized intervention strategies, including progressive training paths and cognitive reinforcement training, and customize feedback preferences according to the five-level feedback mechanism; Step 4: Unify the data standard system to conduct cross-scenario correlation analysis, personalized learning path planning, and teaching strategy optimization; Step 5: Analyze the data, optimize teaching strategies based on the collected data, communicate data across scenarios, and optimize system performance.
[0018] In one embodiment of the present invention: motion data collection includes skeletal key point capture, using a camera with a resolution of ≥720P, combined with the MediaPipe posture estimation algorithm to obtain the coordinates of 21 skeletal key points of the human body in real time, with a sampling frame rate of ≥30fps, covering the motion characteristics of dance joint angles and sports force trajectories, physiological state data collection includes eye movement trajectory detection, heart rate variability collection and real-time analysis of micro-expressions, environmental and behavioral data collection includes audio and text data and environmental parameter detection, after data collection is completed, the motion, physiological and audio data are unified with a time stamp, the error is controlled to be ≤50ms, a synchronous trigger is used to achieve multimodal data alignment, and the 3σ principle is used to eliminate mutation points in the motion trajectory, and the sliding window algorithm is used to smooth the heart rate data to eliminate motion artifact interference.
[0019] In one embodiment of the present invention: in the standardized quantification of action execution layer evaluation, fatigue index = (resting HRV - real-time HRV) / resting HRV × 100%, effective gaze rate = (key area gaze time / total operation time) × 100%.
[0020] In one embodiment of the present invention: comprehensive dance score = joint compliance rate × 0.5 + attention index × 0.3 + spatial perception score × 0.2.
[0021] In one embodiment of the present invention: in the emotional adaptive feedback and intervention operation, taking dance turning imbalance as an example, the progressive training path is single-leg standing balance training (10s / group) in the first week - adding turning decomposition movements in the second week - complete turning movements in the fourth week (the goal is to maintain balance for 20s), and cognitive reinforcement training is divided into insufficient working memory and weak spatial perception (working memory training is triggered when the short-term memory test score is <4, and VR space training is triggered when the spatial perception score is <60). For the problem of insufficient working memory, digital span training is pushed, gradually increasing from 3 digits to 7 digits, 3 sets × 5 minutes per day, and for the problem of weak spatial perception, three-dimensional space rotation training is performed using VR, 2 times a week × 15 minutes.
[0022] In one embodiment of the present invention: In the unified data labeling system, the capability labeling system: 128 basic capability tags are defined, using a three-tier classification structure: Level 1 tags (6): movement control, cognitive ability, emotional management, etc.; Secondary tags (32): fine motor skills, spatial perception, working memory, etc.; Level 3 tags (90): wrist flexibility, joint angle control, attention maintenance, etc. Error coding rules: Establish a five-digit encoding rule: E-XX-YY-Z E: error category (E0: action error, E1: cognitive error); XX: subject code (01: mathematics, 02: experiment, 03: physical education); YY: Action type (01: Attention, 02: Measurement, 03: Dribbling); Z: severity level (1-5); E0-01-01-3 indicates a level 3 control error in a dance turn.
[0023] Example 1 Specific situation: The system detected a 15% decrease in class concentration in the 20th minute compared to the previous 15 minutes (real-time concentration 48%, standard ≥60%). This was manifested as: 30% of the time spent looking down at the phone (standard ≤10%), 55% of the time spent eye movement away from the blackboard (standard ≤30%), frowning frequency 2 times / minute (confused state), body leaning back angle >25° (relaxed posture), and heart rate variability (HRV) = 72ms (normal range). Action and behavior data collection: Hardware configuration: A 4K resolution smart camera (60fps) with built-in AI expression recognition and posture analysis algorithms is installed at the front of the classroom; student desktops are equipped with smart terminals (with integrated inertial sensors and a sampling rate of 200Hz); Technical parameters: Concentration calculation: AI algorithms are used to identify head posture (a face angle of less than 15° toward the blackboard is considered "looking down") and phone usage behavior (the camera detects screen on and hand gripping movements), and calculate the percentage of effective concentration time in real time. Eye tracking: The camera tracks pupil movement at 120Hz and combines it with corneal reflection to determine whether the gaze is focused on the key area of the blackboard; Body posture: Extract the coordinates of eight skeletal points, including the shoulder and spine, and calculate the body's backward leaning angle (the standard is ≤15° for a focused posture); Physiological and environmental data collection: Micro-expression analysis: AI algorithm identifies the frequency of corrugator muscle contraction (>1 times / minute is considered confusion); Environmental noise: When the classroom microphone detects background noise > 55dB, it will automatically filter out irrelevant voices; Data preprocessing: Spatiotemporal alignment: The error between motion data (09:35:18.234) and eye movement data (09:35:18.260) is controlled within 26ms; Outlier filtering: Use the 3σ principle to eliminate sudden action data such as sudden head turns; Action execution layer evaluation: Effective concentration rate = 48% (standard ≥ 60%), mobile phone use exceeding the standard rate = 20% (standard ≤ 10%); Error pattern clustering: The DBSCAN algorithm identified the problem as "distraction in class" (similar behaviors accounted for 78%); Cognitive ability level diagnosis: Mobile phone use → associated with insufficient self-control, as verified by a behavioral scale test, with a self-control score of <40 points; Frowning frequency → associated with "difficulty understanding the principles of engine fault diagnosis", pre-test accuracy rate <45%; Comprehensive evaluation score: Action execution layer: 48 points (50% weight) → 24 points; Physiological state layer: 55 points (confused micro-expression but normal HRV) → 16.5 points; Cognitive ability level 45 points → 9 points; Comprehensive score of 49.5 points (must be ≥60 points to meet the standard); Attention state classification: Low arousal + cognitive confusion → triggers third-level feedback (active suggestions); Multimodal feedback implementation: Visual feedback: The teacher's large screen displays a real-time warning message: "Concentration drops 15% in the 20th minute." The blackboard AR displays a yellow border marking key points: "Fault code reading steps." If eye movement deviates for more than 10 seconds, the border flashes and displays "Attention to key fault diagnosis steps." Auditory feedback: The smart speaker plays a prompt tone, "Please pay attention to the fault code example on the blackboard." The speaking speed is dynamically adjusted based on the student's heart rate (current heart rate is 78 beats / minute, using a medium speaking speed); Tactile feedback: The student's smart terminal vibrates (80Hz) for 1.5 seconds, prompting the student to look up. Personalized intervention strategies: Progressive training: Phase 1 (3 minutes): Play a 20-second dynamic animation of "Typical Engine Fault Codes" and simultaneously conduct "Look Up - Look at the Blackboard" training; Phase 2 (8 minutes): Complete two fault code matching exercises, with the intelligent terminal marking incorrect options in real time; Cognitive enhancement: 1.5-minute audio mini-lesson on "Engine Fault Diagnosis Formulas" is delivered, simulating fault code reading operations using VR scenarios. Unified data standards: Ability tag: Associate "Classroom Concentration" (ID-032) with "Technical Principles Understanding" (ID-017); Error code: E1-07-04-2 (E1 cognitive error, 07 auto repair subject, 04 classroom behavior, level 2 severe); Cross-scenario correlation analysis: The GNN graph shows that the migration coefficient between "concentration in auto repair theory classes" and "efficiency in auto repair practical classes" is 0.72. The similarity of the behavioral characteristics of the two is extracted through the LSTM network. Prediction: If current concentration increases by 25%, practical class fault diagnosis efficiency can be increased by 20%; Learning path planning: It is recommended to first conduct "Mechanical Drawing Reading" training (transfer coefficient 0.68) to strengthen visual focus and understanding of technical principles, and then return to the study of auto repair theory; Full-cycle tracking: Generate a 45-minute concentration change curve (from 65% → 48% → 58%); Data precipitation: Holographic files: records 300 sets of eye movement data, 15 concentration curves, and 12 feedback records; Error library archive: classified by "Auto Repair - Theory Course - E1-07-04-2", associated with VR training plan; Teaching strategy iteration: AI-driven: Analyzing the characteristics of highly focused students (with a concentration level of ≥70%), we found that their average eye movement dwell time in key areas was ≥2 seconds, which was then updated as the standard for classroom focus training. Teacher collaboration: The system flagged a misjudgment case where a student lowered their head to take notes. The system optimized the posture recognition model and incorporated "hand writing movements" into the concentration judgment criteria. Subsequently, the teacher inserted a "auto repair fault answering session" according to the system prompts, and classroom concentration returned to 62% within 5 minutes.
[0024] Example 2: Intelligent error correction scenario for chemical titration experiments Specific situation: Student Wang Qiang titrated too quickly three times in the "0.1 mol / L NaOH titration of HCl" experiment. The standard was ≤5 drops / second, but the actual measured speed was 8-10 drops / second, resulting in an endpoint error of >0.5 mL. The system detected that his gaze stayed on the burette scale for only 35% of the time, while the standard was ≥70%. His HRV was 35ms, indicating fatigue. Key operating steps, Multimodal acquisition: The camera identifies the burette holding angle, the standard is 45°, and the actual angle is 30°; PPG monitored the heart rate at 108 beats / min, and eye tracking showed that the gaze frequently strayed from the scale line; Cognitive diagnosis: Titration speed is too fast → associated with insufficient "fine motor control", the nine-hole peg test score is 28 points, the standard is ≥40 points; Step omission → executive dysfunction, task switching test error rate 25%; Feedback intervention: The AR virtual scale line turns red and flashes when it approaches the end point, and the phone's vibration frequency increases with the dripping speed; Push the "Drop-by-Drop Practice" task (automatically triggered when a drip rate > 6 drops / second is detected): Complete 10 precise titrations in virtual simulation with an error of ≤ 0.1 mL); Cross-scenario migration: It is recommended to first practice "calligraphy brushwork" to improve wrist control, and then return to the experiment. The efficiency is expected to increase by 35%; The ability matrix shows that the migration coefficient between “fine movements” and “dancing fan swing” is 0.68.
[0025] Example 3: Basketball three-step layup action optimization scenario, Specific situation: Student Zhao Gang over-strided in his second step during a three-step layup four times in a row. The standard was 70cm ± 5cm, but the actual measurement was 85-90cm. His layups were also off target. The system detected that his knees buckled inward when he exerted force, with an angle deviation of >10°. His gaze was not focused on the basket, and he remained there for 40% of the time. Key operating steps, Dynamic data collection: MediaPipe captures six key points, including the knee and ankle joints, and calculates the average stride distance to be 87 cm. The average knee inward angle detected by the inertial sensor is 12°; Three-dimensional assessment: The achievement rate of the action execution layer is 55%, the attention index of the physiological state layer is 60%, and the "spatial distance judgment" score of the cognitive layer is 45 points; Comprehensive score 52 points (must be ≥60 points); Feedback Strategy: Visual feedback: AR marks the starting and ending points of the stride, showing the ideal trajectory in blue and the actual trajectory in red; Tactile feedback: When you land in the second step, the phone vibrates strongly at 300Hz to indicate that the distance is too long; Progressive training: stick stickers on the ground to control stride distance, then add shooting movements; Cross-scenario association: Error code E0-03-02-3 (Level 3 error in sports layup); We recommend the “standing long jump distance control” training, which has a transfer coefficient of 0.71 and improves spatial perception ability.
[0026] Table 1 is a common operation flow chart of the two groups of embodiments;
[0027] Table 1 Through the full-process operation of the above embodiments, through multimodal data collection, cross-dimensional evaluation, emotional adaptive feedback, cross-scenario migration and data closure, it is specifically demonstrated how to solve the problems of hardware dependence, poor real-time performance, insufficient interdisciplinary adaptation and so on in the existing technology. Each operation link contains clear technical parameters, trigger conditions and intervention strategies, which verifies the effectiveness and operability of the invented method.
[0028] Although the present invention is disclosed above with reference to preferred embodiments, this is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications without departing from the spirit and scope of the present invention. Therefore, any modifications, equivalent variations, and modifications made to the above embodiments in accordance with the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection defined by the claims of the present invention.
Claims
1. A teaching implementation method based on the combination of motion perception and AI drive, including a teaching method implementation, characterized by: The implementation of the teaching method includes multimodal data expansion collection, cross-dimensional dynamic assessment and cognitive diagnosis, emotional adaptive feedback and intervention, cross-scenario learning map construction and ability transfer, and full-process data closure and system evolution. The multimodal data expansion collection includes dynamic data collection, physiological data collection, environmental behavior data collection, and data preprocessing. The specific steps of the teaching implementation method based on the combination of motion perception and AI drive are as follows: Step 1: Collect motion data, physiological state data, and environmental and behavioral data, unify the timestamps of the three types of data, filter out outliers, and process the collected data; Step 2: Evaluate the action execution layer and the physiological state layer, diagnose the cognitive ability layer, and generate a comprehensive ability score for the three layers for real-time analysis; Step 3: Identify and categorize students' emotional states, provide intelligent feedback based on vision, hearing, and touch, and generate personalized intervention strategies, including progressive training paths and cognitive reinforcement training, and customize feedback preferences according to the five-level feedback mechanism; Step 4: Unify the data standard system to conduct cross-scenario correlation analysis, personalized learning path planning, and teaching strategy optimization; Step 5: Analyze the data, optimize teaching strategies based on the collected data, communicate data across scenarios, and optimize system performance.
2. The teaching implementation method realizes ability transfer through graph neural network, and the system updates the migration coefficient threshold every quarter.
3. The teaching implementation method based on the combination of motion perception and AI drive according to claim 1 is characterized by: The motion data collection captures skeletal key points, uses a camera with a resolution of ≥720P, and combines the MediaPipe posture estimation algorithm to obtain the coordinates of 21 skeletal key points of the human body in real time, with a sampling frame rate of ≥30fps, covering the motion characteristics of dance joint angles and sports force trajectories. The physiological state data collection includes eye movement trajectory detection, heart rate variability collection, and real-time analysis of micro-expressions. The environmental and behavioral data collection includes audio and text data and environmental parameter detection. After the data collection is completed, the motion, physiological, and audio data are unified with a timestamp, and the control error is ≤50ms. A synchronous trigger is used to achieve multimodal data alignment, and the 3σ principle is used to eliminate mutation points in the motion trajectory. The sliding window algorithm is used to smooth the heart rate data to eliminate motion artifact interference. The cognitive ability layer diagnosis is based on a pre-trained cognitive association model. The model establishes a mapping relationship between motion errors and cognitive dimensions through an LSTM network. A three-layer Bi-LSTM is used, and the input layer dimension = 21 skeletal points × 3 coordinates.
4. The teaching implementation method based on the combination of motion perception and AI drive according to claim 2 is characterized by: The action execution layer assessment includes standard quantification and error pattern clustering, calculation of joint angle compliance rate, classification of action errors using the DBSCAN algorithm, generation of dance scene error action heat map, and centralized review and comparison of common problems. The physiological state layer assessment includes fatigue index calculation and attention concentration calculation. The cognitive ability layer diagnosis includes cognitive dimension mapping and a three-dimensional ability matrix. Dance action memory errors → associated short-term memory capacity, experimental step omissions → associated executive dysfunction. The three-dimensional ability matrix constructs a "spatial perception-working memory-logical reasoning" radar chart. Based on the association model trained with 5,000+ samples, the input action error data automatically generates a cognitive ability assessment.
5. The teaching implementation method based on the combination of motion perception and AI drive according to claim 3 is characterized by: The comprehensive evaluation model includes a weighted score of 50% action execution layer + 30% physiological state layer + 20% cognitive ability layer.
6. The teaching implementation method based on the combination of motion perception and AI drive according to claim 4 is characterized by: The visual feedback is AR superimposed with a translucent blue standard action skeleton and red real-time action, and joint deviations greater than 5° are highlighted. The auditory feedback converts the key points of the action into language instructions, and the instruction speed is dynamically adjusted according to the student's heart rate. The tactile feedback is the vibration feedback force deviation of the mobile phone, and the vibration frequency is positively correlated with the degree of error. The personalized intervention strategy includes progressive training paths and cognitive reinforcement training. The five-level feedback mechanism includes silent observation, slight prompts, active suggestions, forced intervention, and emergency pause. Students customize their feedback preferences, and the system dynamically adjusts the intervention strategy according to the configuration.
7. The teaching implementation method based on the combination of motion perception and AI drive according to claim 5 is characterized by: The unified data standard system is divided into a capability label system and error coding rules. The capability label system defines 128 basic capability labels, and the error coding rules establish unified error coding. The cross-scenario association analysis includes graph neural network modeling, and uses a graph convolutional network to construct a capability association graph. Node features are extracted through a spatiotemporal graph convolution module. The nodes are actions in each scenario, and the edge weights represent the strength of capability migration. When the scenario capability is improved, the capability improvement probability of the associated scenario is automatically calculated.
8. The teaching implementation method based on the combination of motion perception and AI drive according to claim 6 is characterized by: The personalized learning path planning includes efficient migration path recommendations and full-cycle capability tracking. The teaching strategy optimization includes teacher-side decision support and system self-optimization mechanism. The teacher-side outputs a class capability shortcoming report and recommends interdisciplinary integrated teaching plans. The system self-optimization mechanism continuously iterates the capability association model based on 100,000+ student data and updates the migration coefficient threshold every quarter.
9. The teaching implementation method based on the combination of motion perception and AI drive according to claim 7 is characterized by: The data precipitation and analysis automatically generates a holographic learning archive containing action records, cognitive assessments, and emotional feedback, which is classified and archived according to subject-action-cognitive root causes to support teachers in retrieving typical errors. The teaching strategy iteration is divided into AI-driven optimization and teacher-system collaboration. AI drives the analysis of the characteristics of outstanding students and generates new teaching standards based on their characteristics. Teachers mark AI assessment deviation cases, and the system records and reinforces learning to correct the emotion recognition model.
10. The teaching implementation method based on the combination of motion perception and AI drive according to claim 8, characterized in that: The cross-scenario data intercommunication includes capability matrix synchronization and multi-terminal data synchronization, which displays the capability migration trajectory of dance, experiment and sports venue data under a unified coordinate. Multi-terminal data supports real-time synchronization of cross-terminal data of mobile phones, teacher large screens and VR devices, with a delay of less than 200ms. The system performance optimization includes lightweight deployment and adaptive learning. The lightweight deployment uses model distillation technology to compress the AI engine, and achieves action recognition delay of less than 100ms on mobile phones. The system automatically adjusts the algorithm complexity according to the device performance.
Citation Information
Patent Citations
Dance teaching assisting method and device
CN112309181A
Cited By
Layered learning condition diagnosis and intervention system based on AI big data driving
CN121391556A
Peking Opera interactive experience method and device based on motion capture
CN121438396A