Auxiliary ethnic and folk dance movement teaching method based on intelligent somatosensory recognition

By adjusting the acquisition time resolution and machine learning model analysis in real time, the problem of identifying breakpoints in the somatosensory recognition system during high-speed movements is solved, and efficient and stable assistance is achieved for ethnic folk dance teaching.

CN120296280AInactive Publication Date: 2025-07-11CHONGQING TOURISM VOCATIONAL COLLEGE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510363382.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-11
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the teaching of folk dances, existing somatosensory recognition systems are difficult to accurately capture key information during high-speed movements, resulting in identification of breakpoints and misjudgment, affecting the stability and efficiency of the teaching auxiliary system.

Method used

By adjusting the acquisition time resolution in real time, analyzing learner's action characteristics using machine learning models, dynamically switch high-resolution modes to capture high-speed movements, and recovering the initial frequency when the movement is stable, optimizing resource utilization.

Benefits of technology

It improves the integrity and system stability of action recognition, reduces data redundancy, ensures high accuracy and efficient performance of the teaching process, and improves the intelligence level of ethnic folk dance teaching.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120296280A_ABST
    Figure CN120296280A_ABST
Patent Text Reader

Abstract

The invention discloses a national and folk dance movement teaching assistance method based on intelligent somatosensory recognition, and relates to the technical field of intelligent somatosensory recognition, and the method comprises the following steps: when dance movement teaching assistance is carried out, a somatosensory recognition system carries out the real-time collection of the body movement of a learner at an initial time resolution, and forms a continuous movement data stream; preprocessing the obtained original action data, and organizing the preprocessed data according to a time sequence mode to form a data set; according to the method, the acquisition time resolution is adjusted in real time according to the action state of a learner, accurate capture of high-speed actions and resource optimization acquisition of stable actions are realized, the problems of recognition breakpoints, key frame omission and the like are effectively avoided, and the action recognition integrity and the system stability are improved; and when the action is stable, the initial sampling frequency is recovered, the data redundancy and the calculation burden are reduced, high precision and high efficiency are both considered, and the intellectualization and reliability of the folk dance teaching are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent body sensory recognition, and in particular to a method for teaching ethnic folk dance movements based on intelligent body sensory recognition. Background Art

[0002] "Teaching assistance for ethnic and folk dance movements based on intelligent body recognition" refers to the use of modern intelligent sensing technology (such as body recognition, motion capture, artificial intelligence analysis, etc.) to assist the teaching process of ethnic and folk dances. By collecting learners' body movement data in real time through sensors (such as cameras, depth sensors, wearable devices, etc.), the system can automatically identify and analyze the differences between their movements and standard dance movements, thereby providing learners with instant movement correction feedback, evaluation results or personalized guidance. This method not only improves the interactivity and fun of teaching, but also helps learners master complex ethnic and folk dance movements more efficiently. It is especially suitable for scenarios or online teaching environments that lack professional teacher guidance.

[0003] The body sensing recognition system is an intelligent perception system based on computer vision, sensor technology and artificial intelligence algorithms. It can collect the human body's movement posture, bone structure and movement trajectory in real time without contact, and identify and analyze them. In the teaching assistance of ethnic folk dance movements, the body sensing recognition system continuously collects and dynamically tracks the learner's body movements, compares them with the preset standard dance movement model, and thus judges the correctness, coherence and completion of the movements, and provides visual feedback and correction suggestions. The system can significantly improve teaching efficiency and interactivity, helping learners to more accurately master complex dance movements with ethnic characteristics. It is especially suitable for self-study scenarios without teacher guidance, online teaching environments, and the digital inheritance of ethnic dances.

[0004] The existing technology has the following deficiencies: In the existing technology, the body recognition system usually continuously collects learners' motion data in the process of assisting the teaching of ethnic folk dances, so as to fully capture their dynamic changes in the entire dance movement process, thereby realizing a comprehensive analysis of elements such as movement rhythm, coherence and timing structure. The so-called continuous motion data collection means that the system records the learner's body posture, bone joint position and movement angle at each moment in real time with a certain time resolution, thereby forming a continuous data stream with time sequence.

[0005] However, when learners perform high-speed actions such as rapid rotation, jumping, or complex twists, the time resolution of the system often fails to adapt to the high-speed changes of the actions, easily resulting in the problem of insufficient sampling density. This leads to the omission of important motion information during key action nodes, thus forming "recognition breakpoints" or "empty time periods", causing truncation or distortion of the action time series. In this case, it will be difficult for the system to accurately compare the complete actions of the learners with the standard action model, and it is prone to recognition errors such as "action recognition failure", "unable to score", or "action missing". More seriously, the recognition model falls into an abnormal state due to the lack of key time-series data, even causing the subsequent processing flow to interrupt, thus seriously affecting the stability and effectiveness of the dance teaching assistance system.

[0006] The above information disclosed in the background art section is only used to enhance the understanding of the background of the present disclosure. Therefore, it may include information that does not constitute the prior art known to those of ordinary skill in the art. Summary of the Invention

[0007] The object of the present invention is to provide a teaching assistance method for ethnic and folk dance movements based on intelligent somatosensory recognition, which adjusts the acquisition time resolution in real time according to the action state of the learner, realizes accurate capture of high-speed actions and resource-optimized acquisition of stable actions, effectively avoids problems such as recognition breakpoints and omission of key frames, improves the integrity of action recognition and the stability of the system; restores the initial sampling frequency when the action is stable, reduces data redundancy and computational burden, achieves both high precision and high efficiency, and enhances the intelligence and reliability of ethnic and folk dance teaching to solve the problems in the above background art.

[0008] To achieve the above object, the present invention provides the following technical solution: A teaching assistance method for ethnic and folk dance movements based on intelligent somatosensory recognition, including the following steps:

[0009] When carrying out dance movement teaching assistance, the somatosensory recognition system collects the body movements of the learner in real time at the initial time resolution to form a continuous action data stream;

[0010] Preprocess the obtained original action data and organize the preprocessed data in chronological order to form a data set;

[0011] Extract key indicators reflecting the rapid changes of the learner's actions from the data set through feature engineering, and comprehensively analyze the extracted key indicators to identify whether the learner is in a state of rapid action change;

[0012] Use the key indicators after analysis and processing as feature vectors and input them into a pre-trained machine learning model. Through the machine learning model, intelligently evaluate the current action characteristics of the learner to determine whether the learner's action is in a state of rapid change;

[0013] When the evaluation result of the machine learning model indicates that the learner's actions are in a state of rapid change, an adaptive regulation mechanism is triggered based on the evaluation result to switch the time resolution to the high-resolution mode and increase the sampling density of the action frames; when it is detected that the learner's action state has returned to the stable stage, the time resolution is automatically switched back to the initial value to reduce the data volume, relieve the computational pressure, and maintain the overall operation efficiency.

[0014] Preferably, when assisting in dance movement teaching, the somatosensory recognition system collects the learner's body movements in real time at the initial time resolution. The specific steps of the process are as follows:

[0015] First, continuously obtain the original action data of the learner at each moment through the somatosensory device, including but not limited to body postures, bone joint positions, and movement trajectories;

[0016] Next, mark the collected data with timestamps to ensure that the data is arranged in chronological order;

[0017] Subsequently, store the action data at each time point in sequence according to the sampling interval to form a structured and continuously tracked action data stream, providing a basic support for subsequent action recognition, analysis, and feedback.

[0018] Preferably, key indicators reflecting the rapid changes in the learner's actions are extracted from the data set through feature engineering. The extracted indicators include the mutation frequency of the movement direction per unit time and the change rate of the orientation angles of the upper and lower body joint groups of the human body per unit time. After comprehensively analyzing the mutation frequency of the movement direction per unit time and the change rate of the orientation angles of the upper and lower body joint groups of the human body under the detection window, a direction switching reference value and a skeleton torsion reference value are respectively generated. Whether the learner is in a state of rapid action change is identified through the direction switching reference value and the skeleton torsion reference value.

[0019] Preferably, the specific steps for comprehensively analyzing the mutation frequency of the movement direction per unit time under the detection window to generate the direction switching reference value are as follows:

[0020] In the detection window, first obtain the velocity vector of each frame, denoted as representing the spatial movement direction of the key body parts of the current frame. For two consecutive frames of velocity vectors, calculate the direction angle between the two consecutive frames. The calculation expression is as follows:

[0021]

[0022] , where is the velocity vector of the i-th frame, is the velocity vector of the (i - 1)-th frame, is the modulus of the velocity vector of the i-th frame, is the modulus of the velocity vector of the (i - 1)-th frame, and ∈ is a very small positive number set to prevent division-by-zero errors, and θ i is the direction change angle;

[0023] To capture the frequency of the "direction switching event" in the detection window, set the angle threshold to δ, which is used to determine whether a "mutation switch" has occurred. If θ i > δ, it is recorded as a valid direction mutation event; count the number of such mutation events in the statistical window, and calculate the weight in combination with the density of consecutive mutations within the window to generate a direction switching reference value. The generation formula is as follows:

[0024]

[0025] , where DC is the direction switching reference value, and N δ is the number of direction mutations, indicating the number of mutations that satisfy θ i > δ within the current detection window; κ j is the frame interval between the j-th mutation and the (j + 1)-th mutation.

[0026] Preferably, the specific steps for comprehensively analyzing the change rate of the orientation angle between the upper body and lower body joint groups of the human body within the detection window to generate a skeleton torsion reference value are as follows:

[0027] First, respectively define the overall orientation space vectors for the upper body and lower body of the human body, and reflect the severity of the action torsion by calculating the spatial torsion energy generated by the change between two orientation vectors at adjacent times. The calculation expression is as follows:

[0028]

[0029] , where is the overall orientation vector of the upper body joint group of the human body at the q-th moment, is the overall orientation vector of the lower body joint group of the human body at the q-th moment, is the overall orientation vector of the upper body joint group of the human body at the (q + 1)-th moment, is the overall orientation vector of the lower body joint group of the human body at the (q + 1)-th moment, is the amplitude of the upper body orientation change, is the amplitude of the lower body orientation change, and θ q is the angle between the upper and lower body orientation vectors, M is the total number of action frames sampled by the detection window, and E twist is the instantaneous change energy of the joint group orientation angle, representing the cumulative torsion energy of the upper and lower body orientation changes per unit time;

[0030] After calculating the instantaneous change energy E of the joint group orientation angle twistAfter that, a non-linear enhancement function is further adopted to highlight the sensitivity of the high-value section of the instantaneous change energy, and the reference value of the skeleton torsion is obtained. The calculation expression is as follows:

[0031]

[0032] , where ST is the reference value of the skeleton torsion, α is the non-linear enhancement weight coefficient, and β is the non-linear exponential enhancement coefficient.

[0033] Preferably, the processed direction switching reference value and the skeleton torsion reference value are used as feature vectors and input into a pre-trained machine learning model. The machine learning model generates an action amplitude change coefficient, and the action amplitude change coefficient is used to intelligently evaluate the current action characteristics of the learner to determine whether the learner's action is in a rapid change state.

[0034] Preferably, the action amplitude change coefficient generated when the pre-trained machine learning model intelligently evaluates the current action characteristics of the learner is compared and analyzed with a pre-set reference threshold of the action amplitude change coefficient to determine whether the learner's action is in a rapid change state. The judgment logic is as follows:

[0035] If the action amplitude change coefficient is greater than the pre-set reference threshold of the action amplitude change coefficient, it is determined that the current learner's action is in a rapid change state; if the action amplitude change coefficient is less than or equal to the pre-set reference threshold of the action amplitude change coefficient, it is determined that the current learner's action is not in a rapid change state.

[0036] Preferably, when the evaluation result of the machine learning model shows that the learner's action is in a rapid change state, an adaptive regulation mechanism is triggered based on the evaluation result, and the time resolution is switched to the high-resolution mode; when it is monitored that the learner's action state returns to the stable stage, the time resolution is automatically switched back to the initial value. The specific steps are as follows:

[0037] During the teaching assistance process of dance movements, the somatosensory recognition system collects the learner's actions in real time at the initial time resolution, and outputs the action amplitude change coefficient MAV at the current moment through the pre-trained model. The action amplitude change coefficient is used to characterize the overall amplitude change intensity of the learner's actions. To dynamically respond to the change trend of the action state, a regulation driving function is introduced to measure the growth rate and trend of MAV. The calculation expression is as follows:

[0038]

[0039] , where Φ(t) is the regulation driving function, which is used to measure the change trend intensity of the action amplitude change coefficient, t is the time variable, and ln(1 + MAV 2 ) is the non-linear enhancement mapping;

[0040] After obtaining the regulation driving function Φ(t), the time resolution is dynamically adjusted according to the value of the regulation driving function Φ(t). The time resolution is adjusted through an exponential enhancement adjustment function, and the adjustment formula is as follows:

[0041] f(t) = f default ·(1 + γ·e Φ(t) )

[0042] , where f(t) is the adjusted sampling time resolution, f default is the initial time resolution, γ is the adjustment gain coefficient used to control the amplitude of resolution change, and e is the natural base;

[0043] To avoid wasting resources due to running in a high time resolution mode for a long time, when the action change intensity weakens, that is, the value of Φ(t) approaches 0, a self-stabilizing recovery mechanism is started. The current frame rate is smoothly called back through the following exponential decay function, and the calculation expression is as follows:

[0044] f(t + Δt) = f(t) - η·(f(t) - f default )·e -λ|Φ(t)|

[0045] , where f(t + Δt) is the time resolution at the next moment, η is the decay rate coefficient to control the backward speed, λ is the decay sensitivity coefficient to adjust the influence degree on small changes, and e -λ|Φ(t)| is the exponential decay function.

[0046] In the above technical solution, the technical effects and advantages provided by the present invention are as follows:

[0047] The present invention dynamically adjusts the time resolution of data collection according to the real-time change of the learner's action state, realizing accurate capture of high-speed complex actions and resource-optimized collection of low-speed stable actions. This method effectively avoids problems such as action recognition breakpoints, key frame omissions, or model misjudgments caused by insufficient sampling density when the learner performs high-dynamic actions such as rapid rotation, jumping, or twisting, improving the integrity of action recognition and the stability of system operation; at the same time, it automatically restores the initial sampling frequency when the action state tends to be stable, significantly reducing data redundancy and system load, ensuring both high-precision recognition and efficient performance operation in the teaching process, thereby improving the intelligent level of ethnic and folk dance teaching and the reliability of teaching assistance. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments described in the present invention. For those of ordinary skill in the art, other drawings can also be obtained based on these drawings.

[0049] Figure 1 This is the method flow chart of the teaching assistance method for ethnic and folk dance movements based on intelligent somatosensory recognition of the present invention. Specific embodiments

[0050] Now, example embodiments will be described more fully with reference to the accompanying drawings. However, the example embodiments can be implemented in various forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that this disclosure will be more thorough and complete, and will fully convey the concept of the example embodiments to those skilled in the art.

[0051] The present invention provides a Figure 1 teaching assistance method for ethnic and folk dance movements based on intelligent somatosensory recognition as shown below, including the following steps:

[0052] When conducting teaching assistance for dance movements, the somatosensory recognition system collects the learner's body movements in real time at the initial time resolution to form a continuous action data stream.

[0053] Provide a basic and continuous action data stream for the system, so as to ensure that no matter how the learner's action rhythm changes, relatively complete initial timing information can be obtained first. This timing information lays a foundation for data preprocessing, feature extraction, and action state evaluation in the subsequent steps.

[0054] When conducting teaching assistance for dance movements, the somatosensory recognition system collects the learner's body movements in real time at the initial time resolution. The specific steps of the process are as follows: First, continuously obtain the original action data of the learner at each moment through somatosensory devices (such as depth cameras, infrared sensors, or inertial sensors), including but not limited to body postures, bone joint positions, and movement trajectories; then, mark the collected data with time stamps to ensure that the data is arranged in time sequence; subsequently, store the action data at each time point in sequence according to the sampling interval to form a structured and continuously tracked action data stream, providing basic support for subsequent action recognition, analysis, and feedback. In this process, the initial time resolution determines the number of data frames collected per unit time, thereby affecting the fineness of action capture and the system operation efficiency.

[0055] Preprocess the obtained original action data and organize the preprocessed data in chronological order to form a data set.

[0056] Preprocessing mainly includes operations such as removing noise, correcting missing frames, and unifying data formats. Removing noise aims to eliminate the influence of external environmental light, sensor interference, etc. on the recognition data; correcting missing frames can fill in the missing frames caused by device jitter or short-term occlusion to ensure the continuity of subsequent analysis; unifying data formats is to convert information that may come from different devices or different dimensions (such as RGB images, depth information, skeletal coordinates) into a standard data structure that can be directly processed by the system. Through these preprocessing steps, the data becomes purer and more consistent, facilitating subsequent feature extraction and analysis.

[0057] In the data set, each record corresponds to key indicators such as the body posture, skeletal joint positions, and angles of the learner at a certain time point or during a certain action segment. The purpose of establishing the data set is to facilitate the system to uniformly manage and count the same action cycle or the same dance passage, and also to enable quick retrieval and analysis of a series of consecutive action frame data as needed when inputting into subsequent feature engineering and machine learning models.

[0058] Through feature engineering, key indicators reflecting the rapid change of the learner's actions are extracted from the data set, and the extracted key indicators are comprehensively analyzed to identify whether the learner is in a state of rapid action change;

[0059] Through feature engineering, key indicators reflecting the rapid change of the learner's actions are extracted from the data set. The extracted indicators include the mutation frequency of the movement direction per unit time and the change rate of the orientation angle between the upper body and lower body joint groups of the human body per unit time. After comprehensively analyzing the mutation frequency of the movement direction per unit time and the change rate of the orientation angle between the upper body and lower body joint groups of the human body per unit time under the detection window, a direction switching reference value and a skeleton torsion reference value are generated respectively. Whether the learner is in a state of rapid action change is identified through the direction switching reference value and the skeleton torsion reference value.

[0060] An increase in the mutation frequency of the movement direction within a unit time usually indicates that the current learner's movements are in a state of rapid change. This is because during natural human movement, the change in the direction of the velocity vector reflects the degree of adjustment of body posture and movement path. When a learner performs high-dynamic dance movements such as quick turns, jump landings with turns, rapid arm swings, or full-body twists, the movement directions of multiple body parts will switch violently within a short time, resulting in a significant increase in the mutation frequency of the velocity vector. Especially in ethnic and folk dances, many movements are characterized by strong rhythms, large amplitudes, and frequent direction switches. At this time, the spatial trajectory of the movement is often no longer smooth and linear, but full of complex spatial changes and multi-directional energy releases. Therefore, an increase in the mutation frequency of the movement direction within a unit time can be used as a key indicator reflecting the explosiveness and complexity of movements, quantifying whether the learner is in the stage of rapid movements, thereby assisting the somatosensory recognition system to dynamically adjust the sampling strategy, capture movement details, and improve recognition accuracy.

[0061] The specific steps for comprehensively analyzing the mutation frequency of the movement direction within a unit time under the detection window to generate a direction switch reference value are as follows:

[0062] In the detection window, first obtain the velocity vector of each frame, denoted as represents the spatial movement direction of the key body parts (such as the center of the chest and the mid-axis of the pelvis) in the current frame. For two consecutive frames of velocity vectors, calculate the direction angle between the two consecutive frames. The calculation formula is as follows:

[0063]

[0064] , where is the velocity vector of the i-th frame, is the velocity vector of the (i - 1)-th frame, is the magnitude of the velocity vector of the i-th frame, representing the magnitude of the movement speed in the current frame, is the magnitude of the velocity vector of the (i - 1)-th frame, ∈ is a very small positive number (such as 1e-6) set to prevent division-by-zero errors, and θ i is the direction change angle, representing the amplitude of the velocity direction change of the current frame (the i-th frame) relative to the previous frame (the (i - 1)-th frame);

[0065] The angle θ i reflects whether there is a significant change in direction for the learner in this frame. When θ i is close to 0, it indicates a weak direction change; when θ i is large (such as close to 90° or 180°), it indicates the existence of violent changes such as sharp turns and reversals, providing a basis for subsequent mutation recognition.

[0066] To capture the frequency of the "direction switching event" in the detection window, an angular threshold δ is set to determine whether a "mutation switch" has occurred. If θ i > δ, it is recorded as a valid direction mutation event; count the number of such mutation events in the window, and calculate the weight in combination with the density of consecutive mutations within the window to generate a direction switching reference value. The generation formula is as follows:

[0067]

[0068] , where DC is the direction switching reference value, and N δ is the number of direction mutations, indicating the number of mutations that satisfy θ i > δ within the current detection window; κ j is the frame interval between the j-th mutation and the (j + 1)-th mutation.

[0069] By counting the frequency of direction mutation events and combining their density in time, a direction switching reference value is constructed to quantify the intensity of the current action change of the learner. The direction switching reference value not only reflects the number of action direction changes but also emphasizes the concentration of mutations, thus more sensitively identifying the rapid change state.

[0070] The larger the direction switching reference value generated by comprehensively analyzing the mutation frequency of the movement direction within a unit time under the detection window, the more significant the movement direction changes the learner has undergone during this period, usually corresponding to high-intensity, explosive, or coherent transition dance movements (such as continuous turns, sudden stops and turns, swinging jumps, etc.). In ethnic and folk dances, such movements frequently appear in action segments with sudden rhythm changes or emotional expressions. Therefore, an increase in the value of the direction switching reference value can be used as a reliable representation that the current action of the learner is in a rapid change state. On the contrary, when this reference value is small, it indicates that the movement direction change of the learner within the monitoring window is small or relatively stable, and the action is closer to the normal rhythm or slow transition state, belonging to the normal or low-intensity change stage.

[0071] An increase in the rate of change of the orientation angle between the upper and lower body joint groups of the human body per unit time usually indicates that the current learner's movement is in a state of rapid change. This phenomenon often occurs in high-dynamic movement phases such as twisting, spinning, jumping and turning, and quickly shifting the center of gravity in dance movements. When the human body's upper body (such as the shoulders and thoracic vertebra) and lower body (such as the pelvis and knees) move smoothly, their orientations are relatively consistent or change slowly. However, when the movement changes drastically, such as when the upper body twists quickly and the lower body pushes off the ground and turns, the movement directions of the two parts will show significant differences in a short period of time, resulting in a rapid change in the orientation angle per unit time. This increase in the angle change rate is a signal that the coordinated movement structure between different parts of the body is broken or reorganized, representing a transition of the movement state from stable to high-speed transformation. Therefore, monitoring the rate of change of the upper-lower body orientation angle can not only identify the explosiveness of the movement but also capture subtle body movement trends, which is one of the important features for determining whether the learner is in a "rapidly changing movement state".

[0072] The specific steps for comprehensively analyzing the rate of change of the orientation angle between the upper and lower body joint groups of the human body per unit time under the detection window to generate a skeleton torsion reference value are as follows:

[0073] First, define the overall orientation space vectors for the upper body (such as the thoracic vertebra and scapula) and lower body (such as the pelvis and hip joint) of the human body respectively. The spatial torsion energy generated by the change of the two orientation vectors between adjacent moments is calculated to reflect the severity of the movement torsion. The calculation formula is as follows:

[0074]

[0075] , where is the overall orientation vector of the upper body joint group of the human body at the qth moment, determined by key upper body joints such as the shoulder-thoracic vertebra-spinal vertex, and is used to represent the overall orientation of the upper part of the body, is the overall orientation vector of the lower body joint group of the human body at the qth moment, determined by key lower body joints such as the pelvis-hip joint-bottom of the coccyx, and is used to represent the overall orientation of the lower part of the body, is the overall orientation vector of the upper body joint group of the human body at the (q + 1)th moment, is the overall orientation vector of the lower body joint group of the human body at the (q + 1)th moment, is the change amplitude of the upper body orientation, indicating the change amount of the upper body orientation vector (i.e., the modulus of the spatial difference vector) from moment q to q + 1, is the change amplitude of the lower body orientation, indicating the change amount of the lower body orientation vector (i.e., the modulus of the spatial difference vector) from moment q to q + 1, θ q is the angle between the upper and lower body orientation vectors, indicating the upper body vector at the qth moment, and the lower body vector The spatial angle between them, M is the total number of action frames sampled by the detection window, and E twist It is the energy of instantaneous change in the angle of the joint group, which represents the cumulative torsional energy of the change in the direction of the upper and lower body per unit time;

[0076] The degree of coordinated change of the upper and lower body at adjacent moments is quantified by multiplying the spatial vector modulus and the sine angle, and the cumulative result of the instantaneous change energy of all adjacent action moments in the detection window is obtained, which intuitively reflects the intensity of the skeletal muscle torsional action.

[0077] The energy E of the instantaneous change in the angle of the joint group is calculated. twist After that, a nonlinear enhancement function (such as exponential enhancement or logarithmic enhancement) is further used to highlight the sensitivity of the high-value segment of instantaneous change energy, and the skeleton torsion reference value is obtained. The calculation expression is as follows:

[0078]

[0079] , where ST is the skeleton torsion reference value, and α is the nonlinear enhancement weight coefficient, which is used to amplify or reduce the input energy value. The weight in the overall index, β is the nonlinear exponential enhancement coefficient, which is used to adjust the amplification or compression amplitude of the instantaneous change energy and highlight the high energy change area.

[0080] Through nonlinear mapping (such as logarithmic-power function mapping), the originally nonlinear energy change value is converted nonlinearly, highlighting the sensitivity of the indicator in the violent torsion stage, so that the indicator can better reflect the obvious characteristics of the learner's movement in the rapid change stage, and effectively enhance the system's ability to identify the movement change state.

[0081] The skeleton torsion reference value is generated by comprehensively analyzing the rate of change of the orientation angle of the joint groups of the upper and lower body of the human body in a unit time under the detection window. The larger the reference value, the greater the direction deviation or angle change of the learner's upper and lower body in a short period of time, which usually corresponds to high-dynamic dance movements such as twisting the waist, turning, rotating in the air, and rapid center of gravity switching, which is a typical manifestation of the body in a rapidly changing state. Therefore, when the skeleton torsion reference value increases, it can be judged that the learner is currently performing a complex action that changes rapidly; on the contrary, if the reference value remains at a low level within a certain range, it means that the learner's movement structure is relatively stable, the direction is consistent, and the overall change is in a normal or stable state.

[0082] The analyzed key indicators are input into the pre-trained machine learning model as feature vectors. The machine learning model is used to intelligently evaluate the learner's current action characteristics to determine whether the learner's action is in a state of rapid change.

[0083] The direction switching reference value and skeleton torsion reference value after analysis and processing are input into the pre-trained machine learning model as feature vectors. The motion amplitude variation coefficient is generated by the machine learning model. The current motion characteristics of the learner are intelligently evaluated through the motion amplitude variation coefficient to determine whether the learner's motion is in a rapidly changing state.

[0084] "Pre-trained machine learning model" refers to a discriminative model that is trained by machine learning algorithms based on a large amount of labeled historical action data before the system is officially put into use. The model has the ability to identify and classify action states. During the training phase, developers will collect a variety of sample data including "fast-changing actions" and "smooth actions", and label each set of data (such as "fast" or "non-fast" state), and then input these data into the selected machine learning algorithm (such as support vector machine, random forest, decision tree, or neural network, etc.), and optimize the parameters of the model through continuous iteration so that it can learn the mapping relationship between feature vectors and action states. The trained model will be saved and deployed to the somatosensory recognition system. When the system extracts new feature vectors (such as "direction switching reference value" and "skeleton torsion reference value") during actual operation, these features can be directly input into the model, and the model can quickly determine the state category of the current action.

[0085] Unlike traditional rule-based judgment systems, pre-trained machine learning models have stronger generalization and adaptability. It can dig out hidden patterns and laws in complex, changeable, and nonlinear motion data. For example, in the teaching assistance of ethnic folk dances, the movement changes of different dances vary greatly. Some fast movements do not completely rely on the increase in joint speed, but may be manifested as a drastic switch of the overall direction or an increase in local torsion. Therefore, relying solely on threshold judgment is prone to misjudgment. The use of pre-trained models can more accurately output a "motion amplitude change coefficient" to quantify the intensity of the change in the current action by modeling the nonlinear combination of comprehensive indicators such as "direction switching reference value" and "skeleton torsion reference value". This coefficient is essentially a prediction result of the model output, which can be used as the core basis for the system to determine whether to switch to a high-time resolution sampling mode, thereby realizing intelligent perception and response to the fast action stage.

[0086] The machine learning model is not limited here, and any machine learning model that can perform a comprehensive analysis of the direction switching reference value DC and the skeleton torsion reference value ST to generate the motion amplitude variation coefficient MAV is acceptable. To implement the technical solution of the present invention, the present invention provides a specific implementation method:

[0087] The formula for generating the Motion Amplitude Variation coefficient MAV is as follows: MAV = α1·DC + α2·ST, where α1 and α2 are the preset proportionality coefficients of the Direction Change reference value DC and the Skeleton Twist reference value ST respectively, and both α1 and α2 are greater than 0.

[0088] The preset proportionality coefficients (α1 and α2) refer to the weighting factors or weight parameters used to measure the influence degrees of the Direction Change reference value DC and the Skeleton Twist reference value ST on the final result respectively during the calculation of generating the Motion Amplitude Variation coefficient MAV. The settings of these coefficients can be constants determined based on empirical rules, data statistics, or after model training. Their core function is: to assign reasonable influence to different features, making the comprehensive index MAV more accurately reflect the intensity of the current motion. For example, if the "skeleton twist" contributes more to the intensity of the motion in a specific dance scenario, then α2 should be set larger than α1. Since both α1 and α2 in the formula are greater than 0, they not only ensure that the feature influence is positive, but also can control the sensitivity of the model to different reference values through fine-tuning, thereby improving the accuracy and adaptability of judging the motion state.

[0089] From the Motion Amplitude Variation coefficient, it can be seen that the larger the Direction Change reference value generated by comprehensively analyzing the mutation frequency of the motion direction within a unit time under the detection window, and the larger the Skeleton Twist reference value generated by comprehensively analyzing the change rate of the orientation angle between the upper body and lower body joint groups of the human body within a unit time under the detection window, the larger the Motion Amplitude Variation coefficient generated when the intelligent evaluation of the current motion characteristics of the learner is performed by the pre-trained machine learning model, indicating that the current motion of the learner is in a rapid change state. On the contrary, it indicates that the motion structure of the learner is relatively stable, the orientations are consistent, and the overall change is in a normal or stable state.

[0090] Compare and analyze the Motion Amplitude Variation coefficient generated when the intelligent evaluation of the current motion characteristics of the learner is performed by the pre-trained machine learning model with the preset reference threshold of the Motion Amplitude Variation coefficient to judge whether the learner's motion is in a rapid change state. The judgment logic is as follows:

[0091] If the Motion Amplitude Variation coefficient is greater than the preset reference threshold of the Motion Amplitude Variation coefficient, it is judged that the current motion of the learner is in a rapid change state; if the Motion Amplitude Variation coefficient is less than or equal to the preset reference threshold of the Motion Amplitude Variation coefficient, it is judged that the current motion of the learner is not in a rapid change state.

[0092] When the evaluation result of the machine learning model indicates that the learner's actions are in a rapidly changing state, an adaptive regulation mechanism is triggered based on the evaluation result to switch the time resolution to the high-resolution mode and increase the sampling density of action frames; when it is detected that the learner's action state has returned to the stable stage, the time resolution is automatically switched back to the initial value to reduce the data volume, relieve the computing pressure, and maintain the overall operation efficiency;

[0093] The main function of the above steps is to achieve dynamic sampling adjustment for the complexity of dance movements, so as to ensure the complete acquisition of key action information and optimize the utilization of system resources at the same time. During the process of dance movement teaching assistance, the changes in the learner's movements show significant phased characteristics, that is, there are high-intensity dynamic movements such as rapid rotation, jumping, and twisting in some periods, while in other periods, they are relatively slow or stable transitional movements. If the somatosensory recognition system always uses a fixed time resolution for data collection, it is easy to miss key frames, break action recognition, or make evaluation mistakes due to insufficient frame rate during the stage of rapidly changing movements, affecting the accuracy of teaching assistance. By introducing a machine learning model to evaluate the learner's current action state, once it is judged as "rapidly changing", the system triggers an adaptive regulation mechanism to temporarily switch the time resolution to the high-resolution mode, so as to collect more action frames per unit time, enhance the continuous capture ability of high-speed action processes, and effectively improve the integrity and accuracy of time series analysis. When it is subsequently detected that the action has entered the stable stage, the system automatically switches the time resolution back to the initial value, reduces the unnecessary data sampling density, reduces the data transmission and processing burden, and maintains the overall operation efficiency and stability of the system. This dynamic adjustment mechanism not only improves the recognition ability of complex dance movements, but also improves the utilization rate of the system's computing resources, providing key support for realizing high-performance and intelligent dance teaching assistance.

[0094] When the evaluation result of the machine learning model indicates that the learner's actions are in a rapidly changing state, an adaptive regulation mechanism is triggered based on the evaluation result to switch the time resolution to the high-resolution mode; when it is detected that the learner's action state has returned to the stable stage, the time resolution is automatically switched back to the initial value. The specific steps are as follows:

[0095] During the process of dance movement teaching assistance, the somatosensory recognition system collects the learner's actions in real time at the initial time resolution, and outputs the movement amplitude variation coefficient MAV at the current moment through a pre-trained model. The movement amplitude variation coefficient is used to characterize the overall amplitude change intensity of the learner's actions. To dynamically respond to the change trend of the action state, a regulation driving function is introduced to measure the growth rate and trend of MAV. The calculation expression is as follows:

[0096]

[0097] , where Φ(t) is the control driving function, which is used to measure the intensity of the change trend of the action amplitude change coefficient, t is the time variable, which represents the continuous time points in the action acquisition process, ln(1+MAV 2 ) is a nonlinear enhancement mapping that enhances the dynamic responsiveness of the amplitude, making the growth trend of high MAV values ​​more sensitive while avoiding over-response to low values;

[0098] This step is used to monitor the dynamic trend of motion changes in real time. It does not use classification labels, but instead uses the continuous trend-driven adjustment mechanism of the MAV to provide a continuous response reference for subsequent time resolution adjustments, ensuring that the adjustment process is sensitive and stable.

[0099] After obtaining the control driving function Φ(t), the time resolution is dynamically adjusted according to the value of the control driving function Φ(t). The time resolution is adjusted by the exponential enhancement adjustment function. The adjustment formula is as follows:

[0100] f(t)=f default ·(1+γ·e Φ(t) )

[0101] , where f(t) is the adjusted sampling time resolution, f default is the initial time resolution, γ is the adjustment gain coefficient used to control the resolution change amplitude, and e is the natural base;

[0102] The above steps allow the temporal resolution to be dynamically improved according to the changes in the MAV without the need for preset classification labels or hard thresholds. The exponential enhancement form allows the system to respond faster to sudden and violent movements, and increases the sampling density during unstable movement stages, thereby capturing more frame data and improving recognition and analysis accuracy.

[0103] In order to avoid wasting resources by operating in high time resolution mode for a long time, when the intensity of the motion change weakens, that is, the Φ(t) value approaches 0, the self-stabilization recovery mechanism is started, and the current frame rate is smoothly callback through the following exponential fallback function. The calculation expression is as follows:

[0104] f(t+Δt)=f(t-η·(f(t)-f default )·e -λ|Φ(t)|

[0105] , where f(t+Δt) is the time resolution of the next moment, η is the fallback rate coefficient, which controls the fallback speed, λ is the fallback sensitivity coefficient, which adjusts the degree of influence on small changes, and e -λ|Φ(t)| is an exponential damping function that rapidly approaches the default value as Φ(t)→0.

[0106] The above mechanism is used for the system to automatically identify the "recovery" behavior when the action change tends to be stable, realizing the smooth fallback of the time resolution from high frequency to the initial frequency, without relying on external judgment tags, ensuring the balance between resource utilization and recognition accuracy, and avoiding redundancy caused by long-term high-density sampling.

[0107] The present invention dynamically adjusts the time resolution of data acquisition according to the real-time changes of the learner's action state, realizing the accurate capture of high-speed complex actions and the resource-optimized acquisition of low-speed stable actions. This method effectively avoids problems such as action recognition breakpoints, omission of key frames, or model misjudgment caused by insufficient sampling density when the learner performs high-dynamic actions such as rapid rotation, jumping, or twisting, improving the integrity of action recognition and the stability of system operation; at the same time, it automatically restores the initial sampling frequency when the action state tends to be stable, significantly reducing data redundancy and system load, ensuring the balance between high-precision recognition and high-efficiency performance operation in the teaching process, thereby improving the intelligent level of ethnic and folk dance teaching and the reliability of teaching assistance.

[0108] The above formulas are all dimensionless and take their numerical calculations. The formula is a formula obtained by software simulation of a large amount of collected data to approximate the real situation as closely as possible. The preset parameters in the formula are set by those skilled in the art according to the actual situation.

[0109] Only some exemplary embodiments of the present invention have been described by way of illustration above. Undoubtedly, for those of ordinary skill in the art, the described embodiments can be modified in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and descriptions are illustrative in nature and should not be construed as limiting the scope of protection of the claims of the present invention.

[0110] It should be noted that in this article, if there are relational terms such as first and second, they are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of another identical element in the process, method, article or device including the element.

[0111] It should be understood that in various embodiments of the present application, the sequence numbers of the above processes do not imply the order of execution, and the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.

[0112] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0113] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0114] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0115] In addition, the functional units in various embodiments of the present application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0116] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

[0117] Only some exemplary embodiments of the present invention have been described above by way of illustration. Undoubtedly, those of ordinary skill in the art can modify the described embodiments in various different ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the protection scope of the claims of the present invention.

Claims

1. An auxiliary method for teaching ethnic and folk dance movements based on intelligent somatosensory recognition, characterized in that, Including the following steps: When conducting dance movement teaching assistance, the somatosensory recognition system collects the learner's body movements in real time at the initial time resolution to form a continuous action data stream; Preprocess the obtained original action data and organize the preprocessed data in chronological order to form a data set; Extract key indicators reflecting the rapid changes in the learner's movements from the data set through feature engineering, and comprehensively analyze the extracted key indicators to identify whether the learner is in a state of rapid movement change; Use the key indicators after analysis and processing as feature vectors and input them into a pre-trained machine learning model. Through the machine learning model, intelligently evaluate the current action characteristics of the learner to determine whether the learner's actions are in a state of rapid change; When the evaluation result of the machine learning model indicates that the learner's actions are in a state of rapid change, trigger the adaptive regulation mechanism based on the evaluation result, switch the time resolution to the high-resolution mode, and increase the sampling density of the action frames; When it is detected that the learner's action state returns to the stable stage, automatically switch the time resolution back to the initial value to reduce the data volume, relieve the computational pressure, and maintain the overall operation efficiency.

2. The ethnic and folk dance movement teaching assistance method based on intelligent somatosensory recognition according to claim 1, wherein When conducting dance movement teaching assistance, the somatosensory recognition system collects the learner's body movements in real time at the initial time resolution. The specific steps are as follows: First, continuously obtain the learner's original action data at each moment through the somatosensory device, including but not limited to body postures, bone joint positions, and movement trajectories; Next, mark the collected data with timestamps to ensure that the data is arranged in chronological order; Subsequently, store the action data at each time point in sequence according to the sampling interval to form a structured and continuously tracked action data stream, providing a basic support for subsequent action recognition, analysis, and feedback.

3. The teaching assistance method for ethnic and folk dance movements based on intelligent somatosensory recognition according to claim 1, characterized in that, Extract key indicators reflecting the rapid changes in the learner's movements from the data set through feature engineering. The extracted indicators include the mutation frequency of the movement direction per unit time and the change rate of the orientation angles of the upper and lower body joint groups of the human body per unit time. After comprehensively analyzing the mutation frequency of the movement direction per unit time and the change rate of the orientation angles of the upper and lower body joint groups of the human body per unit time under the detection window, generate a direction switching reference value and a skeleton torsion reference value respectively. Identify whether the learner is in a state of rapid movement change through the direction switching reference value and the skeleton torsion reference value.

4. The method for assisting the teaching of ethnic and folk dance movements based on intelligent somatosensory recognition according to claim 3, characterized in that The specific steps for comprehensively analyzing the mutation frequency of the movement direction per unit time under the detection window to generate a direction switching reference value are as follows: In the detection window, first obtain the velocity vector of each frame, denoted as which represents the spatial motion direction of the key body parts in the current frame. For two consecutive frames of velocity vectors, calculate the direction angle between the two consecutive frames. The calculation expression is as follows: , In the formula, is the velocity vector of the i-th frame, is the velocity vector of the i-1th frame, is the modulus of the velocity vector of the i-th frame, is the modulus of the velocity vector of the i-1th frame, ∈ is a very small positive number set to prevent division by zero errors, and θ i is the angle of direction change; To capture the frequency of the "direction switching event" in the detection window, an angular threshold δ is set to determine whether a "mutation switch" has occurred. If θ i > δ, it is recorded as a valid direction mutation event; count the number of such mutation events in the window, and calculate the weight in combination with the density of continuous mutations within the window to generate a direction switching reference value. The generation formula is as follows: , Wherein, DC is the direction switching reference value, N δ is the number of direction mutations, indicating the number of mutations that satisfy θ i >δ within the current detection window; κ j is the frame interval between the j-th mutation and the (j + 1)-th mutation.

5. The method for teaching and assisting ethnic and folk dance movements based on intelligent somatosensory recognition according to claim 3, characterized in that, The specific steps for comprehensively analyzing the change rate of the orientation angles of the upper and lower body joint groups of the human body per unit time under the detection window to generate a skeleton torsion reference value are as follows: First, define the overall orientation space vectors for the upper and lower body of the human body respectively, and reflect the severity of the action torsion by calculating the spatial torsion energy generated by the change between two orientation vectors at adjacent moments. The calculation expression is as follows: , In the formula, is the overall orientation vector of the upper body joint group at the q-th moment, is the overall orientation vector of the lower body joint group at the q-th moment, is the overall orientation vector of the upper body joint group at the (q + 1)-th moment, is the overall orientation vector of the lower body joint group at the (q + 1)-th moment, is the upper body orientation change amplitude, is the lower body orientation change amplitude, θ q is the angle between the upper and lower body orientation vectors, M is the total number of action frames sampled in the detection window, E twist is the instantaneous change energy of the joint group orientation angle, representing the cumulative torsional energy of the upper and lower body orientation changes per unit time; After calculating the instantaneous change energy E of the joint group orientation angle twist further, a non-linear enhancement function is adopted to highlight the sensitivity of the high-value section of the instantaneous change energy, and the skeleton torsion reference value is obtained. The calculation expression is as follows: , In the formula, ST is the skeleton torsion reference value, α is the non-linear enhancement weight coefficient, and β is the non-linear exponential enhancement coefficient.

6. The teaching assistance method for ethnic and folk dance movements based on intelligent somatosensory recognition according to claim 3, characterized in that, The direction switching reference value and the skeleton torsion reference value after analysis and processing are input into a pre-trained machine learning model as feature vectors. The machine learning model generates an action amplitude change coefficient, and the current action characteristics of the learner are intelligently evaluated through the action amplitude change coefficient to determine whether the learner's action is in a rapidly changing state.

7. The method for teaching and assisting ethnic folk dance movements based on intelligent somatosensory recognition according to claim 6, wherein, The action amplitude change coefficient generated when the current action characteristics of the learner are intelligently evaluated by the pre-trained machine learning model is compared and analyzed with the pre-set reference threshold of the action amplitude change coefficient to determine whether the learner's action is in a rapidly changing state. The judgment logic is as follows: If the action amplitude change coefficient is greater than the pre-set reference threshold of the action amplitude change coefficient, it is determined that the current learner's action is in a rapidly changing state; if the action amplitude change coefficient is less than or equal to the pre-set reference threshold of the action amplitude change coefficient, it is determined that the current learner's action is not in a rapidly changing state.

8. The method for assisting in the teaching of ethnic and folk dance movements based on intelligent somatosensory recognition according to claim 7, characterized in that, When the evaluation result of the machine learning model indicates that the learner's action is in a rapidly changing state, an adaptive regulation mechanism is triggered based on the evaluation result, and the time resolution is switched to the high-resolution mode; when it is monitored that the learner's action state returns to the stable stage, the time resolution is automatically switched back to the initial value. The specific steps are as follows: During the teaching assistance process of dance movements, the somatosensory recognition system collects the learner's actions in real time at the initial time resolution, and outputs the action amplitude change coefficient MAV at the current moment through the pre-trained model. The action amplitude change coefficient is used to characterize the overall amplitude change intensity of the learner's actions. To dynamically respond to the change trend of the action state, a regulation driving function is introduced to measure the growth rate and trend of MAV. The calculation expression is as follows: , where φ(t) is a regulation driving function used to measure the change trend intensity of the action amplitude change coefficient, t is a time variable, and ln(1 + MAV 2 ) is a non-linear enhancement mapping; After obtaining the regulation driving function Φ(t), the time resolution is dynamically adjusted according to the value of the regulation driving function Φ(t), and the time resolution is adjusted through an exponential enhancement type adjustment function. The adjustment formula is as follows: f(t) = f default ·(1 + γ·e Φ(t) ) where \(f(t)\) is the adjusted sampling time resolution, \(f\) default is the initial time resolution, \(\gamma\) is the adjustment gain coefficient used to control the amplitude of the resolution change, and \(e\) is the natural base. To avoid resource waste caused by running in the high time resolution mode for a long time, when the action change intensity weakens, that is, the value of Φ(t) approaches 0, a self-stable recovery mechanism is started, and the current frame rate is smoothly called back through the following exponential decay function. The calculation expression is as follows: f(t + Δt) = f(t - η·(f(t) - f default )·e -λ|Φ(t)| , where f(t + Δt) is the time resolution at the next moment, η is the fallback rate coefficient that controls the fallback speed, λ is the fallback sensitivity coefficient that adjusts the influence degree on small changes, and e -λ|Φ(t)| is the exponential decay function.