An intelligent analysis system for martial arts learning process based on multimodal deep learning
Through the intelligent analysis system of martial arts learning process based on multimodal deep learning, combined with computer vision and deep learning technology, the motion strength, speed and expression analysis are carried out, and the comprehensive and objective evaluation and in-depth feedback of martial arts learners are solved.
Patent Information
- Application Number
- CN202410712673.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-06-04
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2044-06-04
AI Technical Summary
The comprehensive application of the prior art in motor skills assessment and sentiment analysis has not been fully developed, and traditional methods rely on high-cost equipment and professionals, resulting in biased assessment results and difficult to apply in individual training and small-scale teaching environments.
A martial arts learning process intelligent analysis system based on multimodal deep learning is adopted, combined with computer vision and deep learning technology, and comprehensive concentration assessment is carried out through motion strength, speed and expression analysis. The system includes a motion velocity analysis module, a motion velocity analysis module, an expression analysis module and a concentration comprehensive evaluation module, and uses technical components such as YOLOPose, OpenSim, OpenFace 2.0 and MER-Net.
A comprehensive and objective assessment of the strength, speed and concentration of martial arts learners' movements is achieved, and the accuracy and consistency of the assessment is improved. It provides in-depth training feedback for martial arts teaching and personal training, and is suitable for personal training and small-scale teaching environments.
Smart Images

Figure CN118486085B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computer vision and deep learning, and in particular to a martial arts learning process intelligent analysis system based on multimodal deep learning. Background Art
[0002] The application of computer vision in motion analysis is mainly focused on motion capture, skill evaluation and training optimization of martial arts learners. Traditional methods rely on high-cost sensor equipment and subjective judgment of professionals, which is not only costly but may also lead to deviations in evaluation results. To solve these problems, researchers have begun to use deep learning algorithms such as convolutional neural networks (CNN) and recurrent neural networks (RNN) to process motion data obtained from videos. For example, the YOLOPose algorithm has shown significant results in athlete motion detection, but it still has limitations in complex motion and fine skill evaluation. In addition, the previous application areas of OpenSim software were biomechanical research and rehabilitation science, and its application in this field provides a new dimension for motion analysis.
[0003] On the other hand, emotion recognition technology is also gradually becoming an important part of motion analysis. Emotion recognition mainly judges the emotional state of an individual by analyzing facial expressions, voice, body language, etc. In recent years, with the advancement of deep learning technology, facial expression recognition technology based on images and videos has made significant progress. For example, OpenFace uses facial expression analysis technology and has been applied in the fields of advertising and market research. However, the application of these technologies in action training fields such as martial arts is still in its infancy.
[0004] Aerobics has absorbed a lot of upper and lower limb, torso, head and neck, and foot movements from disco, jazz and breakdancing, especially hip movements, which adds vitality to aerobics, and is also beneficial to reduce the accumulation of fat in the buttocks and abdomen, and improve the coordination and flexibility of movements. Gymnastics is a general term for all gymnastic events, mainly including competitive gymnastics, artistic gymnastics, trampoline, as well as competitive aerobics, skill sports and other events. Martial arts is a sport that has gradually formed and developed by the Chinese nation in the long-term production labor, struggle with nature and war in the era of cold weapons. It has the functions of fitness, body protection, defense against enemies and victory.
[0005] In summary, although computer vision and emotion recognition technologies have made significant progress in their respective fields, their comprehensive applications in sports skill assessment and emotion analysis have not yet been fully developed. Existing technologies mostly focus on single-aspect analysis, such as focusing only on movement accuracy or only analyzing facial expressions, and lack a comprehensive assessment of the athlete's overall performance. In addition, existing systems often require complex equipment support and professional operation, which limits their application in personal training and small-scale teaching environments.
[0006] Based on this background, this paper aims to provide a comprehensive, objective, and easy-to-use analysis system for strength, speed, and concentration of martial arts learners by integrating multimodal deep learning techniques, including action-based activities such as aerobics, gymnastics, and martial arts. Summary of the invention
[0007] The purpose of the invention provided by the present invention is to provide a martial arts learning process intelligent analysis system based on multimodal deep learning to solve the problems in the above-mentioned background technology.
[0008] To achieve the above objectives, the present invention is implemented through the following technical solutions: a martial arts learning process intelligent analysis system based on multimodal deep learning, comprising:
[0009] Processor: responsible for program execution sequence control, data arithmetic, logical operations and time control to ensure the orderly operation of the system;
[0010] Movement force analysis module: to clarify the force used by martial arts learners in each movement during martial arts practice and to quantify it;
[0011] Action speed analysis module: Use the ratio of displacement between human joints between adjacent key frames in martial arts learners' training videos to time as speed, and use this to perform speed analysis;
[0012] Expression analysis module: Analyze the facial expressions of martial arts learners during martial arts practice and classify the expressions;
[0013] Comprehensive concentration assessment module: integrates muscle strength and expression classification to obtain the concentration of martial arts learners;
[0014] Control and display module: used for operation control during system use and display of various data and results;
[0015] The action strength analysis module includes a video preprocessing module, a joint point sequence extraction module, a parameter sequence determination module, a parameter vector splicing module and a strength evaluation module, and the video preprocessing module, the joint point sequence extraction module, the parameter sequence determination module, the parameter vector splicing module and the strength evaluation module are all bidirectionally connected by signals. The action speed analysis module includes a data calculation module, a data summary module and a fully connected layer module, and the data calculation module, the data summary module and the fully connected layer module are all bidirectionally connected by signals. The expression analysis module includes an analysis model module, a feature vector splicing module and an expression classification processing module, and the analysis model module, the feature vector splicing module and the expression classification processing module are all bidirectionally connected by signals.
[0016] The output signal of the strength assessment module in the action strength analysis module is connected to the input of the fully connected layer module in the action speed analysis module, and the output signal of the strength assessment module in the action strength analysis module is connected to the input of the expression classification processing module in the expression analysis module.
[0017] Furthermore, the processor includes an algorithm processing module, an information encryption module, an information interaction module and a data backup storage module, and the algorithm processing module, the information encryption module, the data backup storage module and the information interaction module are all bidirectionally signal-connected.
[0018] Furthermore, the comprehensive concentration evaluation module includes a data merging module, an attention mechanism module and a result output module, and the data merging module, the attention mechanism module and the result output module are all bidirectionally signal connected.
[0019] Furthermore, the control display module includes an operation control module and an information display module, and there is a bidirectional signal connection between the operation control module and the information display module.
[0020] Furthermore, the output signal of the force evaluation module in the action force analysis module is connected to the input signal of the information interaction module in the processor.
[0021] Furthermore, the output signal of the fully connected layer module in the motion speed analysis module is connected to the input signal of the information interaction module in the processor.
[0022] Furthermore, the output terminal signal of the expression classification processing module in the expression analysis module is connected to the input terminal of the information interaction module in the processor.
[0023] Furthermore, the output end of the result output module in the comprehensive concentration evaluation module is bidirectionally connected to the input end of the information interaction module in the processor.
[0024] Furthermore, the bidirectional signal of the output end of the operation control module in the control display module is connected to the input end of the information interaction module in the processor.
[0025] Furthermore, the output terminal signal of the operation control module in the control and display module is connected to the input terminal of the result output module in the concentration comprehensive evaluation module.
[0026] The present invention provides a martial arts learning process intelligent analysis system based on multimodal deep learning. It has the following beneficial effects:
[0027] (1) This intelligent analysis system for martial arts learning process based on multimodal deep learning innovatively combines computer vision and deep learning technology to build a multimodal system. By extracting joint point data from martial arts action videos, a personalized action comparison model is constructed to analyze the learner's action strength. Then, the displacement data of human joint points between adjacent key frames is used to analyze the speed of the action. At the same time, emotion recognition technology is used to evaluate the learner's facial expressions and micro-expressions, so as to accurately judge his or her emotional state and conduct a comprehensive concentration analysis based on strength and speed.
[0028] (2) This intelligent analysis system of martial arts learning process based on multimodal deep learning, through the application of multimodality, on the one hand analyzes the technical details of the action to improve the accuracy and consistency of the intelligent evaluation of action learning and practice; on the other hand, it evaluates the learner's emotional state, thereby providing more in-depth training feedback for martial arts teaching and personal training, so as to provide more comprehensive and scientific guidance.
[0029] (3) The intelligent analysis system of martial arts learning process based on multimodal deep learning has the core intelligent comprehensive analysis capability, which can consider the learner's physical movements and emotional expressions at the same time, and provides a new and scientific intelligent solution for the teaching and evaluation of martial arts and other action-based sports. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 This is a general system diagram of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0031] Figure 2 A schematic diagram of a processor of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0032] Figure 3 This is a schematic diagram of an action strength analysis module of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0033] Figure 4 A schematic diagram of a motion speed analysis module of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0034] Figure 5 This is a schematic diagram of an expression analysis module of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0035] Figure 6 A schematic diagram of a comprehensive evaluation module of concentration of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0036] Figure 7 A schematic diagram of a control and display module of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0037] Figure 8 This is a flow chart of a martial arts learning process intelligent analysis system based on multimodal deep learning of the present invention;
[0038] Fig. 9 A transformer flow diagram of a martial arts learning process intelligent analysis system based on multimodal deep learning according to the present invention;
[0039] Fig.10 This is a structural diagram of the attention mechanism of a martial arts learning process intelligent analysis system based on multimodal deep learning in the present invention.
[0040] In the figure: 1. Processor; 2. Movement intensity analysis module; 3. Movement speed analysis module; 4. Expression analysis module; 5. Concentration comprehensive evaluation module; 6. Control display module. DETAILED DESCRIPTION
[0041] The technical solutions in the embodiments of the present invention will be described clearly and completely below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.
[0042] Examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements having the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are intended to be used to explain the present invention, and should not be construed as limiting the present invention.
[0043] The present invention will be further described below in conjunction with the accompanying drawings and embodiments:
[0044] See also Figure 1-10 , the present invention provides a technical solution:
[0045] The invention discloses a martial arts learning process intelligent analysis system based on multimodal deep learning, which aims to provide an objective and quantitative martial arts practice status evaluation system. The system innovatively combines computer vision and deep learning technology to construct a multimodal system for action strength, speed judgment and facial expression analysis. By extracting joint point data in action videos, the system constructs a personalized action model to analyze the learner's action strength, and then uses the displacement of human joint points between adjacent key frames to analyze the movement speed of the learner. At the same time, using emotion recognition technology, the system can evaluate the learner's facial expression and micro-expression, so as to accurately judge and classify his emotional state. Finally, combined with strength, speed and expression, the overall concentration of the learner is comprehensively analyzed and evaluated. The application of this multimodal system can not only improve the accuracy and consistency of intelligent evaluation of martial arts practice, but also provide coaches and learners with more in-depth training feedback and guidance. The core innovation of the invention lies in its comprehensive analysis ability, which can simultaneously consider the learner's body movements and emotional expressions, and provide a new and scientific method for martial arts teaching and evaluation.
[0046] Analytical methods to determine the power, speed and concentration of martial arts movements
[0047] The present invention focuses on constructing an intelligent analysis system for the martial arts learning process based on multimodal deep learning. In order to make the analysis more comprehensive, systematic and professional, the present invention constructs three modules to evaluate and measure the practice of martial arts learners from three aspects: movement strength analysis, movement speed analysis and expression analysis. Finally, the feature vectors of the three modules are spliced, and the attention mechanism method is used to make a comprehensive judgment on the concentration, which not only retains the details but also has comprehensive judgment.
[0048] 1. Movement force analysis: In the movement force analysis module, the present invention intends to conduct a scientific and professional analysis of the learner's movement force through the key frames of the martial arts learner's training video, and classify the learner's movement force into: "no force", "neutral", "force" and "very force". This requires the human posture detection model to cooperate with the scientific biomechanical and kinematic simulation model to complete the extraction of feature vectors, and input the feature vectors into a large deep learning network, so that the network can perform feature learning and finally perform movement force classification. Therefore, the present invention selects a posture estimation model based on deep learning-YOLOPose, which has the advantage of being able to identify and locate the key points of the human body from the video frame in real time and accurately. The reason for selecting YOLOPose is its efficient processing speed and good recognition ability for complex human postures, which is crucial for real-time analysis of movements and extraction of key information from movements. The present invention selects OpenSim as a biomechanical and kinematic simulation model, which has the knowledge of multi-body dynamics and the ability of musculoskeletal modeling, which allows researchers to build detailed biomechanical models and perform motion simulations. The use of OpenSim can calculate key biomechanical parameters such as joint angles, muscle strength and acceleration based on the human joint point data extracted from the key frames by YOLOPose. These parameters provide the necessary quantitative basis for understanding and analyzing the biomechanical characteristics of the action. The two work together, which is equivalent to pre-feature extraction for large deep learning networks, ensuring that feature extraction is more scientific and rich. In the analysis of motion intensity, the present invention selects the Transformer network, which is a deep learning model based on the self-attention mechanism, is particularly suitable for complex features, and has excellent learning ability. In the present invention, the introduction of the Transformer network is mainly to utilize its ability to effectively process time series data, so as to capture and learn complex patterns and relationships in the motion process. By training the Transformer network to classify the features obtained from YOLOPose and OpenSim, the motion intensity can be effectively quantified and classified.
[0049] 2. Speed analysis: In the speed analysis module, the method adopted by the present invention is: calculate the displacement of human joints between adjacent key frames, obtain the time difference through the frame number of the extracted key frame and the unified frame rate of the preprocessed video, and then perform speed analysis. This method is selected in speed analysis based on its multiple advantages in accuracy, efficiency and application applicability. First, this method is based on accurate human joint positioning. Compared with the traditional speed estimation method based on pixel changes, the joint-based method can more directly reflect the movement of various parts of the body, thereby improving the accuracy of the analysis. Secondly, the human joint sequence obtained by the YOLOPose model in the action intensity module is used as the input of the speed analysis module, which greatly reduces the computational overhead in the speed analysis, and by analyzing the displacement of the joints between adjacent key frames, the method can directly calculate the instantaneous speed of the joints. This speed calculation method is not only concise and clear, but also can effectively reflect the dynamic changes of motion, especially when analyzing fast and complex motion. Compared with the optical flow method or other model-based motion estimation techniques that require a lot of calculations and processing, this displacement-based speed analysis method is more efficient. In summary, this method is efficient and applicable, and is particularly suitable for the analysis of fast and complex movements, such as martial arts movement speed assessment.
[0050] 3. Expression analysis: In the field of expression analysis, the present invention selects OpenFace 2.0 and MER-Net as key technical components to jointly complete expression analysis based on their excellent performance and accuracy in facial feature detection and micro-expression recognition. The former analyzes the overall facial expression of martial arts learners, which helps to make a comprehensive judgment on their expressions and determine the main emotional tendencies; the latter, as a micro-expression recognition network, can supplement the missing details of the former, enrich the details of expression analysis, and make expression analysis more scientific and detailed.
[0051] OpenFace 2.0 is an advanced facial behavior analysis tool designed specifically for real-time facial feature point detection, head pose estimation, eye gaze estimation, and facial action unit (Action Units) analysis. Based on deep learning technology, the tool can accurately capture and analyze facial expressions from video or real-time image sources. The key advantage of OpenFace 2.0 is that it can provide detailed facial feature data, including micro-movements of various facial regions, which is essential for understanding complex human emotions and expressions.
[0052] MER-Net (Micro-Expression Recognition Network) is a network specifically designed to recognize and analyze micro-expressions. Micro-expressions are brief, subtle changes in facial expressions that often occur when individuals try to hide or suppress their true emotions. MER-Net can detect and analyze these rapid and subtle changes in expression through highly refined algorithms. Its introduction is particularly suitable for capturing those subtle but informative changes in facial expressions, providing a powerful tool for in-depth understanding of human emotions and non-verbal communication.
[0053] The application of OpenFace 2.0 and MER-Net enables the present invention to capture and analyze facial expressions comprehensively and accurately. OpenFace 2.0 provides basic facial feature points and action unit information, while MER-Net supplements the detailed recognition of micro-expressions. This combined method not only improves the accuracy of expression analysis, but also increases the ability to recognize subtle emotional changes, making the present invention more effective and comprehensive in processing complex human emotional expressions.
[0054] Movement force analysis
[0055] The action strength analysis module of the present invention is to clarify the force used in each action of a martial arts learner during martial arts practice and to quantify it.
[0056] In the present invention, the martial arts learner training videos input into each module are all subjected to uniform preliminary preprocessing, that is, the frame rates of all input videos are unified (the present invention converts the videos uniformly into 25 frames per second); key frames are selected using a frame difference algorithm, and K frames are selected as key frames per second in the present invention; the extracted key frames are combined into a key frame sequence for use by subsequent modules.
[0057] The quantitative evaluation of action strength is mainly divided into three steps:
[0058] 1. Use the YOLOPose model to locate human joints;
[0059] 2. Import human joint data into OpenSim software to build a personalized musculoskeletal model and quantify related parameters.
[0060] 3. Build a transformer network, input the musculoskeletal model parameters calculated in the OpenSim software as extracted feature values into the transformer network, and finally output the action force evaluation results.
[0061] With the development of human posture detection technology, models for detecting human posture by locating human joints emerge in an endless stream, and the accuracy of the models is also constantly improving. Among the many models, the accuracy and efficiency of the YOLOPose model are very good, and it has unique advantages in motion recognition and motion analysis. Therefore, the motion analysis module of the present invention is further implemented using the YOLOPose model as the basic framework.
[0062] YOLOv8 is the leading detector in terms of accuracy and complexity. Therefore, YOLOPose chose it as the basis and built on top of it. YOLO Pose uses CSP-darknet53 as the backbone, and uses PANet to fuse features of various scales from the backbone; then there are four detection heads of different scales; finally, there are two decoupled heads for predicting boxes and key points. For human pose estimation, it boils down to a single category of person detection problem, each person has 17 associated key points, each key point is again identified by position and confidence: {x, y, conf}, and finally determines the sequence of human joint points.
[0063] OpenSim is a musculoskeletal modeling and simulation application and a library specifically designed for these purposes. This technology completes the model building by providing musculoskeletal modeling elements such as biomechanical joints, muscle actuators, ligament forces, compliant contacts, and controllers; in addition, this technology is also used to fit general models to specific object data and perform inverse kinematics and forward dynamics simulations.
[0064] These tools include inverse kinematics to resolve internal coordinates from available spatial marker positions corresponding to known landmarks on the rigid segments; inverse dynamics to determine the set of generalized forces necessary to match the estimated accelerations; static optimization to factorize the net generalized forces between redundant actuators (muscles); and forward dynamics to generate trajectories of states by integrating the system dynamics equations in response to input controls and external forces.
[0065] The present invention builds a personalized musculoskeletal model through OpenSim, and inputs the human joint point sequence obtained from the YOLOPose model into OpenSim. OpenSim uses a series of inverse kinematics and forward dynamics simulation tools to finally output three dynamic parameters: joint angle, muscle strength and acceleration.
[0066] In addition, the present invention constructs a martial arts video database. The martial arts video database constructed by the present invention is obtained by dividing a large number of martial arts training videos into short video samples of fixed length, and labeling these samples for use in the neural network training in subsequent modules. Each short video sample is double-labeled with "expression label" and "action intensity label". Among them, there are four expression labels: "not serious", "neutral", "serious" and "very serious". These labels are used to train the classifier of the expression analysis module to identify and classify the facial expressions of martial arts learners; there are also four action intensity labels: "no force", "neutral", "force" and "very forceful". These labels are used to train the Transformer network of the action intensity analysis module to quantify the intensity of the action.
[0067] The Transformer network is a widely used model in deep learning. It has achieved remarkable results in fields such as natural language processing (NLP) and has gradually been applied to other fields such as computer vision and audio processing. The core feature of the Transformer is that it uses the "Self-Attention" mechanism. This mechanism enables the model to effectively weight the importance of different parts of the sequence when processing sequence data, thereby capturing the complex relationships within the sequence. Compared with traditional recurrent neural networks (RNNs) and long short-term memory networks (LSTMs), Transformers can handle long-distance dependency problems more effectively and are easy to parallelize, so they are more efficient when processing large-scale data.
[0068] Specifically, the Transformer network structure consists of two main parts: the encoder and the decoder. The encoder consists of multiple identical layers, each with two sublayers. The first sublayer is a multi-head self-attention mechanism that allows the model to simultaneously focus on different positions of the input sequence. The second sublayer is a simple, position-by-position fully connected feedforward network. The decoder also consists of multiple identical layers, each with three sublayers. In addition to the two sublayers of the encoder, the decoder inserts a third sublayer between the two sublayers for multi-head self-attention, focusing on the output of the encoder. The output of each sublayer is connected through a residual connection, and each sublayer output is followed by layer normalization. Finally, the output of the decoder is passed through a linear layer and a softmax layer to generate the final output probability. This structure uses the self-attention mechanism to allow the model to model direct dependencies between the input and output of any two positions, no matter how far apart they are in the sequence.
[0069] In this paper, a Transformer network-based action intensity analysis method is proposed to classify the action intensity of martial arts learners. This method first preprocesses and extracts features from martial arts video samples, then uses the Transformer network for learning and classification, and finally achieves effective evaluation of action intensity.
[0070] Video samples are selected from the martial arts video database constructed by the present invention for training the Transformer network. The selected sample videos in the database are preprocessed, and the N key frames contained therein are input into the YOLOpose model in chronological order to extract the human joint point sequence of each frame of the video. The human joint point sequence of each frame is input into the OpenSim software, and the three parameter sequences of the joint angle sequence, muscle strength sequence and acceleration sequence are calculated. The joint angle feature vector, muscle strength feature vector and acceleration feature vector are spliced together to form a new feature vector, which is used as the input of the Transformer network, and the output of the Transformer network is the classification of the action strength: "no force", "neutral", "force" or "very force". During the training process, for each sample video, its pre-annotated strength classification ("no force", "neutral", "force" or "very force") is used as a training label. And the cross entropy loss function is selected as the optimization target. For this multi-classification problem, the cross entropy loss function is expressed as:
[0071]
[0072] Among them, y is the one-hot encoding of the true label, is the model prediction output.
[0073] In the present invention, the Transformer network is used to analyze and classify the action intensity, which effectively utilizes the encoder and decoder architecture of the Transformer model. The core feature of the Transformer network is its multi-head self-attention mechanism, which can dynamically focus on the relationship between different parts of the input sequence. This mechanism is particularly important when performing action intensity analysis because it enables the model to recognize and explain the complex interactions of various time points in martial arts movements.
[0074] Specifically, when the concatenated feature vector (containing joint angles, muscle strength, and acceleration data) is input into the Transformer network, the model's multi-head self-attention layer first evaluates the key features of the input data at different time points and identifies the relationship between these features. By calculating the attention scores between different features, the Transformer network can determine which features are most important for the final classification of movement intensity. This weighted attention method allows the model to not only process a single feature, but also comprehensively consider multiple biomechanical parameters to more accurately evaluate the overall intensity of the movement. Next, the data processed by the self-attention layer flows to the position-by-position feedforward network. This part of the network further processes the data, enhancing the model's understanding of the characteristics of the movement, enabling it to more accurately classify the intensity of the movement. In this process, the Transformer network uses its deep learning capabilities to effectively extract meaningful patterns and features from complex biomechanical time series. Finally, the output of the Transformer network is used to classify the intensity of the movement. According to the categories defined during training (such as "no force", "neutral", "force", or "very forceful"), the model can accurately classify each movement sample. This deep learning-based classification method not only improves the accuracy of motion intensity assessment, but is also suitable for processing large amounts of video data due to its efficient computing power.
[0075] Analysis of martial arts movement speed
[0076] The speed analysis in the present invention uses the ratio of the displacement between the human joints between adjacent key frames in the pre-processed training video of the martial arts learner to the time as the speed and performs the analysis accordingly. The speed analysis process for the training video of the martial arts learner is as follows:
[0077] 1. Calculation of joint point vector difference: First, pre-process the martial arts video of martial arts learners to obtain the key frame sequence and frame number. Input the key frame sequence into the YOLOpose model to extract the human joint point sequence of these key frames. In this way, the human joint point vector sequence Pt = {P1,t, P2,t,…, Pi,t,…Pn,t} corresponding to the frame number t can be obtained, where Pi,t is the vector of the i-th joint point in the frame with frame number t, and n is the number of human joint points. The displacement of the i-th human joint point vector of adjacent key frames can be expressed as: ΔPi = Pi,t + α-Pi,t, where α is the frame number difference between adjacent key frames. Then there is a sequence of human joint point vector displacements of adjacent key frames
[0078] {ΔP1,t,ΔP2,t,…,ΔPi,t,…,ΔPn,t}
[0079] 2. Calculation of time interval: For the time interval between adjacent key frames, the present invention uses the frame number and frame rate of the key frame for calculation, and the formula is as follows:
[0080]
[0081] Among them, α is the frame number difference between adjacent key frames, and FPS is the frame rate set during the preprocessing of the training video of the martial arts learner, which is 25 frames per second.
[0082] 3. Speed calculation: After having the vector displacement sequence of human joints of adjacent key frames and the time interval between adjacent key frames, the moving speed sequence of human joints can be obtained.
[0083] {V1,t,V2,t,…,Vi,t,…Vn,t}, the calculation formula is as follows:
[0084]
[0085] The maximum value V in the moving speed sequence of the human joint point is selected in the present invention. max,t The corresponding displacement of the key human joint point as the key frame with frame number t is the key speed of the key frame with frame number t.
[0086] 4. Application of the fully connected layer: The key speed data of all key frames are aggregated to form a key speed feature vector V = {Vmax,1, Vmax,2,…, Vmax,t,…}. This vector is then input into a fully connected neural network layer, which is used to extract the speed evaluation of the entire martial arts learner's training video from the key speed data of all key frames. The formula is as follows:
[0087] Vtotal=FC(V)
[0088] Among them, V is the feature vector containing the key speeds of all key frames, V total is the overall video speed evaluation of the fully connected layer output, and FC represents the fully connected layer.
[0089] Through this process, the present invention can comprehensively consider the motion information of all key frames in the video and provide a comprehensive speed evaluation for the entire video. This method combines the advantages of time series analysis and neural networks and provides an effective means to evaluate and understand the motion characteristics in the video.
[0090] Analysis of facial expressions during martial arts practice
[0091] Different from the quantitative analysis of the action strength analysis module, the expression analysis module of the present invention is to analyze and classify the facial expressions of martial arts learners during martial arts practice, and divide facial expressions into four categories: "not serious", "neutral", "serious" and "very serious". In the model construction of expression analysis, the present invention uses a combination of facial behavior analysis tools and micro-expression facial recognition systems to identify and analyze the facial expressions of martial arts learners. The facial recognition tool selected by the present invention is OpenFace 2.0, which is a facial behavior analysis toolkit for computer vision and machine learning research, emotional computing, and building interactive applications based on facial behavior analysis. It is based on advanced machine learning technology and is suitable for real-time emotional analysis. The selected micro-expression facial recognition system is MER-Net (Micro-Expression Recognition Network), which is specially designed to capture and analyze short and subtle facial expression changes. These micro-expressions usually reveal the unconscious real emotions of martial arts learners in their sports state, which is crucial for in-depth analysis of the learners' sports seriousness.
[0092] First, it is necessary to build an expression analysis model. In the present invention, the expression analysis model is mainly divided into three modules: 1. The preprocessing module in the OpenFace 2.0 model and its convolutional neural network with weights and parameters retained. OpenFace2.0 uses computer vision algorithms to analyze facial behaviors, including facial key point detection, head posture tracking, eye gaze estimation, and facial action unit recognition. These behaviors play an important role in understanding the facial expressions of martial arts learners and classifying facial expressions; 2. The preprocessing module in the MER-Net model and its two-dimensional and three-dimensional convolutional neural networks with weights and parameters retained. MER-Net is a network model for micro-expression recognition. The use of MER-net in the present invention is to identify micro-expressions by analyzing tiny facial movements in martial arts learners' training videos, making up for the gap in micro-expression extraction and recognition when only the OpenFace2.0 model is used. In addition, MER-Net itself has advantages in extracting facial features and classifying expressions. It can identify the presence of micro-expressions and the categories of expressions, such as anger, happiness, sadness, etc.; 3. Construct a classifier for the expression analysis model: concatenate the feature vectors obtained from the convolutional neural network of OpenFace 2.0 and the two-dimensional and three-dimensional convolutional neural networks of MER-Net, and input the new complete feature vector into the classifier for classification. The feature vector concatenation formula is as follows:
[0093]
[0094] F cis the newly concatenated feature vector, which contains the facial behavior features obtained by the OpenFace 2.0 model and the micro-expression features obtained by the MER-Net model. openface Represents facial behavior characteristics, F mer-net Represents micro-expression characteristics. Represents the concatenation of features.
[0095] This classifier consists of a fully connected layer and a Sigmoid activation function.
[0096] The formula of the fully connected layer is as follows:
[0097] z=WFc+b
[0098] Among them, W and b are the weight and bias of the fully connected layer respectively.
[0099] The Sigmoid activation function formula is as follows:
[0100]
[0101] Here, z is the output of the fully connected layer.
[0102] The present invention regards this as a multi-classification task, and regards the classification task as "not serious", "neutral", "serious" and "very serious". The final classification output is θ, where θ 0 stands for "not serious", θ 1 stands for "neutral", θ 2 stands for "serious", θ 3 Stands for "very serious".
[0103] We stipulate that
[0104] In the present invention, supervised learning is required to train the classifier in the constructed expression analysis model. The small sample data set constructed by the present invention is still used, and the expression classification labels of the martial arts videos in the data set are utilized, including "not serious", "serious" and "very serious".
[0105] The loss function used is the cross entropy loss function
[0106]
[0107] Among them, y is the one-hot encoding of the true label, is the model prediction output.
[0108] The process of expression analysis is as follows: 1. Input the video frames into the preprocessing module in the OpenFace 2.0 model and the preprocessing module in the MER-Net model respectively. 2. The OpenFace 2.0 model and the MER-Net model are processed in parallel. The former is processed by a 2D convolutional neural network, and the latter is processed by a special neural network that combines a 2D convolutional neural network with a 3D convolutional neural network. The feature vectors corresponding to the OpenFace 2.0 model and the feature vector corresponding to the MER-Net model are obtained respectively. The two feature vectors are concatenated as a new feature vector. 3. The new feature vector is input into the trained expression analysis model classifier to classify the facial expressions of martial arts learners, and the classification results are output as "easy-going", "neutral", "serious" or "all-out".
[0109] Comprehensive Assessment of Concentration in Martial Arts Practice
[0110] At the end of the entire model, the present invention introduces an attention mechanism layer to fuse the muscle strength obtained by the action strength analysis module and the expression classification obtained by the expression analysis module, thereby obtaining the concentration of martial arts learners. The attention mechanism is a technology widely used in deep learning, which can increase the model's attention to important parts of the input data. By assigning different weights to different parts, the attention mechanism enables the model to learn and process complex data structures more effectively.
[0111] The concentration evaluation process in the present invention is as follows:
[0112] 1. First, perform feature fusion to combine the strength data obtained from the action strength analysis, the speed data obtained from the speed analysis, and the emotion classification data obtained from the facial expression analysis into a comprehensive feature vector. Among them, the action strength feature vector, speed feature vector, and expression feature vector are all taken from the feature vectors in the corresponding network input classifier. Let the action strength feature be F motion , facial expression feature is F expression , the velocity characteristic of the human joint point is F V , the fused feature vector F combined It can be expressed as:
[0113]
[0114] in, Represents a feature concatenation operation.
[0115] This comprehensive feature vector contains the comprehensive information of each key frame image in three dimensions: intensity, speed and emotion.
[0116] 2. Then the attention mechanism layer is introduced, and the attention mechanism is used to weight important features. The attention mechanism layer is a widely used structure in deep learning, which is mainly used to enhance the attention of the neural network to the important parts of the input data. This mechanism is implemented by assigning different weights to different parts of the input data, enabling the network to prioritize more important or relevant information. In traditional neural networks, all input data is usually processed in the same way, regardless of its relative importance in a specific task. In contrast, the attention mechanism allows the model to dynamically focus on the information that is most critical to the current task, thereby improving processing efficiency and performance. The introduction of the attention mechanism has significantly improved the ability of neural networks to process complex data, especially in fields such as computer vision. It enables the model to process various types of data more flexibly and efficiently, thereby achieving excellent performance in various complex tasks.
[0117] In the attention mechanism, the attention function is particularly important. The attention function can be described as mapping a query and a set of key-value pairs to an output, where the query, key, value, and output are all vectors. The output is calculated as a weighted sum of the values, where the weight assigned to each value is calculated by the compatibility function of the query and the corresponding keyword.
[0118] The present invention uses the attention mechanism to calculate and classify the concentration of martial arts learners. The present invention converts the fused feature vector F combined Input to the attention mechanism layer. In this layer, the calculation formula of the self-attention mechanism is used:
[0119]
[0120] Among them, Q (query), K (key), and V (value) are all derived from feature vectors. Through this mechanism, the model calculates the weights between different features to determine which features are more important in concentration judgment.
[0121] The attention score is converted into a probability distribution through the softmax function, so that each feature is assigned a weight to indicate its importance in the focus assessment. Features with higher weights in the feature vector will have a greater proportion in the focus assessment.
[0122] Through the self-attention mechanism, the model can not only focus on a single feature, but also gain insight into the interrelationships and influences between different features. For example, when performing an action, the combination of certain muscle strength and joint angles, as well as the speed of the action, may be key indicators of the level of concentration. Similarly, in expression analysis, subtle changes in micro-expressions may also reflect the learner's level of concentration.
[0123] Finally, the attention-weighted feature vectors are used to evaluate the learner's concentration and are classified into four categories: "very inattentive", "relatively focused", "focused" and "very focused". This classification result can provide a basis for martial arts learners to improve themselves.
[0124] This application provides "an intelligent analysis system for martial arts learning process based on multimodal deep learning", but it does not only protect martial arts action learning, but also includes all action learning other than martial arts, such as dance, aerobics and gymnastics, etc., by analyzing the individual's action learning process to provide intelligent guidance and feedback, usually using multimodal deep learning technology, which can simultaneously process and analyze data from different modalities, such as video, audio, sensor data, etc., to more comprehensively understand the individual's action learning process.
[0125] This system has broad application prospects in the fields of dance, aerobics, gymnastics and martial arts, and can enable students in the fields of dance, aerobics, gymnastics and martial arts to better understand and improve their movement skills.
[0126] The application of the attention mechanism in this invention significantly improves the accuracy and meticulousness of concentration assessment, and pushes the depth and breadth of multimodal data analysis to a new level. Through refined weight allocation, it ensures that in the overall evaluation process, each feature can be appropriately considered according to its actual importance in affecting concentration. This method not only shows innovation in technology, but also has important guiding significance in practical applications.
[0127] The above are only preferred embodiments of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A martial arts learning process intelligent analysis system based on multimodal deep learning, characterized in that: include: Processor (1): responsible for program execution sequence control, data arithmetic, logical operations and time control to ensure the orderly operation of the system; Movement force analysis module (2): Identify the force used by martial arts learners in each movement during martial arts practice and quantify it; Action speed analysis module (3): using the ratio of the displacement between the joints of the human body between adjacent key frames in the martial arts learner's training video to the time as the speed, and using this to perform speed analysis; Expression analysis module (4): Analyze the facial expressions of martial arts learners during martial arts practice and classify the expressions; Comprehensive concentration evaluation module (5): integrates action strength data, action speed data and expression classification to obtain the concentration of martial arts learners; Control and display module (6): used for operation control during system use and display of various data and results; The action strength analysis module (2) comprises a video preprocessing module, a joint point sequence extraction module, a parameter sequence determination module, a parameter vector splicing module and a strength evaluation module, wherein the video preprocessing module, the joint point sequence extraction module, the parameter sequence determination module, the parameter vector splicing module and the strength evaluation module are all bidirectionally connected by signals. The action speed analysis module (3) comprises a data calculation module, a data aggregation module and a fully connected layer module, wherein the data calculation module, the data aggregation module and the fully connected layer module are all bidirectionally connected by signals. The expression analysis module (4) comprises an analysis model module, a feature vector splicing module and an expression classification processing module, wherein the analysis model module, the feature vector splicing module and the expression classification processing module are all bidirectionally connected by signals. The output signal of the strength evaluation module in the action strength analysis module (2) is connected to the input of the fully connected layer module in the action speed analysis module (3), and the output signal of the strength evaluation module in the action strength analysis module (2) is connected to the input of the expression classification processing module in the expression analysis module (4).
2. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The processor (1) comprises an algorithm processing module, an information encryption module, an information interaction module and a data backup storage module, wherein the algorithm processing module, the information encryption module, the data backup storage module and the information interaction module are all bidirectionally signal-connected.
3. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The comprehensive concentration evaluation module (5) comprises a data merging module, an attention mechanism module and a result output module, and the data merging module, the attention mechanism module and the result output module are all bidirectionally signal-connected.
4. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The control display module (6) comprises an operation control module and an information display module, and the operation control module and the information display module are both bidirectionally signal-connected.
5. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The output end signal of the strength evaluation module in the action strength analysis module (2) is connected to the input end of the information interaction module in the processor (1).
6. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The output terminal signal of the fully connected layer module in the motion speed analysis module (3) is connected to the input terminal of the information interaction module in the processor (1).
7. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The output terminal signal of the expression classification processing module in the expression analysis module (4) is connected to the input terminal of the information interaction module in the processor (1).
8. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1 is characterized by: The output end of the result output module in the comprehensive concentration evaluation module (5) is bidirectionally signal-connected to the input end of the information interaction module in the processor (1).
9. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1, characterized in that: The output end bidirectional signal of the operation control module in the control display module (6) is connected to the input end of the information interaction module in the processor (1).
10. The martial arts learning process intelligent analysis system based on multimodal deep learning according to claim 1, characterized in that: The output terminal signal of the operation control module in the control display module (6) is connected to the input terminal of the result output module in the concentration comprehensive evaluation module (5).
Citation Information
Patent Citations
Attention detection method fusing multiple discrimination technologies
CN111814718A
Method for evaluating learning concentration through video based on expression and behavior feature extraction
CN112287891A
Cited By
Student classroom concentration intelligent evaluation system and method based on multi-modal data
CN122336858A