Physical education evaluation system based on artificial intelligence
Through the artificial intelligence physical education teaching evaluation system, using technologies such as YOLOv8 and OpenPose algorithms, accurate assessment and personalized feedback on students' movements are achieved, and the problems of inaccurate assessment and insufficient safety monitoring in traditional physical education teaching are solved, and teaching quality and safety are improved.
Patent Information
- Application Number
- CN202510424058.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2045-04-07
AI Technical Summary
In traditional physical education teaching, teaching evaluation lacks accuracy, cannot monitor sports safety in real time, and it is difficult to provide personalized guidance, resulting in insufficient teaching quality and safety.
The physical education teaching and evaluation system based on artificial intelligence is adopted, including data acquisition module, multi-objective detection module, key point modeling module, action trajectory analysis module, multi-modal fusion module and dynamic feedback module, and real-time action evaluation and personalized feedback are achieved using YOLOv8 algorithm, OpenPose algorithm, LSTM network and other technologies.
It realizes accurate assessment of students' movements, improves teaching quality and exercise safety, provides personalized training suggestions, and improves teaching efficiency and safety guarantees.
Smart Images

Figure CN120339004A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of physical education teaching, and specifically to a physical education teaching evaluation system based on artificial intelligence. Background Art
[0002] In traditional physical education teaching, there are many limitations in the teaching evaluation method, which are difficult to meet the needs of modern physical education. From the perspective of teaching quality evaluation, in the past, it mainly relied on teachers' subjective observation and experience judgment to evaluate students' movement performance. When teaching gymnastics, martial arts and other projects, teachers need to pay attention to multiple students at the same time, and it is difficult to accurately capture the subtle deviations of each student's movements. Taking broadcast gymnastics as an example, for details such as the angle of students' arm extension and the amplitude of body torsion, teachers may miss them due to the observation angle or distraction of attention, which leads to the inability to accurately evaluate students' mastery of movement norms, and the improvement of teaching quality is restricted.
[0003] In terms of sports safety, there is a lack of an effective real-time monitoring mechanism. In some competitive sports projects, such as basketball, football games or training, students are prone to injury risks due to excessive exercise intensity, improper movements, etc. Since teachers cannot obtain students' sports physiological data and movement states in real time, they cannot discover and stop dangerous movements in time, and it is difficult to ensure the safety of students during sports. For example, if a student's heart rate is too high due to excessive fatigue during a basketball game and cannot be monitored in time, it may lead to more serious health problems.
[0004] From the perspective of teaching personalization, in the traditional teaching mode, it is very difficult for teachers to provide personalized teaching guidance for each student's characteristics. Each student has different physical fitness, sports talent and learning ability, but teachers usually adopt unified teaching standards and methods, which cannot meet the different needs of students. For example, in track and field teaching, it is difficult to give targeted training plans for students with strong explosive power but insufficient endurance and students with better endurance but lacking explosive power.
[0005] With the rapid development of artificial intelligence technology, its application in the field of education has gradually deepened. The progress of computer vision technology, such as the continuous optimization of object detection and pose recognition algorithms, makes it possible to accurately capture human movements; the innovation of sensor technology makes it a reality to obtain high-precision sports physiological data. However, in the field of physical education teaching evaluation, how to organically integrate these advanced technologies to form a complete, efficient and intelligent evaluation system still faces many challenges. Based on such a background, the present invention is committed to filling this technical gap and promoting the development of physical education teaching evaluation towards intelligence and precision. Summary of the Invention
[0006] The purpose of the present invention is to provide a physical education teaching evaluation system based on artificial intelligence to solve the problems raised in the above background art.
[0007] To achieve the above object, the present invention provides the following technical solutions: A sports teaching evaluation system based on artificial intelligence, the system comprising:
[0008] A data acquisition module for receiving real-time video streams from camera devices and motion physiological data from wearable sensors;
[0009] A multi-object detection module for performing object localization on teachers and students, sports equipment, and balls in the real-time video stream based on the YOLOv8 algorithm, and outputting detection box coordinates and object categories;
[0010] A key point modeling module for extracting skeletal key points of the detected target human body using the OpenPose algorithm to generate a three-dimensional motion pose model;
[0011] An action trajectory analysis module for calculating the Euclidean distance deviation between the motion trajectory and a preset standard action according to the three-dimensional motion pose model, and generating an action standardness evaluation result;
[0012] A multi-modal fusion module for performing temporal alignment on the motion physiological data and the action standardness evaluation result to construct a multi-dimensional feature matrix;
[0013] A dynamic feedback module for generating real-time scores and personalized action correction instructions based on the multi-dimensional feature matrix through an LSTM network optimized by transfer learning.
[0014] Preferably, the execution steps of the multi-object detection module include:
[0015] Inputting the real-time video stream into a pre-trained YOLOv8 model, and extracting multi-scale feature maps through a convolutional neural network;
[0016] Performing anchor box matching and non-maximum suppression processing on the feature maps, and outputting object detection results.
[0017] Preferably, the execution steps of the key point modeling module include:
[0018] Performing local image cropping on the detected human target, and inputting it into the OpenPose network for heat map prediction;
[0019] Extracting the two-dimensional coordinates of the skeletal key points through a Gaussian filtering and peak detection algorithm;
[0020] Combining the depth information of binocular cameras, mapping the two-dimensional coordinates to a three-dimensional space, and generating an action pose sequence with timestamps.
[0021] Preferably, the execution steps of the action trajectory analysis module include:
[0022] Dynamically time-warp and match the three-dimensional action pose model with the standard action template to calculate the joint angle deviation;
[0023] Filter the motion trajectory noise based on the Laplace smoothing algorithm;
[0024] Integrate the trajectory deviation of consecutive action frames and output an action coherence score;
[0025] The Euclidean distance deviation calculation formula is:
[0026]
[0027] where (x i , y i , z i ) represents the three-dimensional coordinates of the i-th joint point, (x std , y std , z std ) represents the corresponding coordinates of the standard action, and n is the total number of joint points.
[0028] Preferably, the calculation formula for dynamic time warping and matching is:
[0029]
[0030] where A = (a1,..., a m ) and B = (b1,..., b n ) are the joint angle sequences of the actual action and the standard action respectively, P is the set of all warping paths, (p, q) is a pair of matching points in the path π, satisfying p ∈ A, q ∈ B, and d(a p , b q ) is the difference metric function for two angle values.
[0031] Preferably, the execution steps of the multi-modal fusion module include:
[0032] Perform wavelet transform denoising on the heart rate, acceleration, and angular velocity in the motion physiological data;
[0033] Perform feature-level fusion on the denoised physiological data and the action standardness evaluation result to generate a fusion feature vector containing spatio-temporal correlation;
[0034] Dynamically allocate the weight ratio of visual data and sensor data through the attention mechanism.
[0035] Preferably, the execution steps of the dynamic feedback module include:
[0036] Adopt the curriculum learning strategy to train the LSTM network in stages, freeze the parameters of the convolutional layer in the first stage, and jointly optimize the fully connected layer in the second stage;
[0037] Perform sliding window segmentation on the multi-dimensional feature matrix and input it into the LSTM network to predict the action scores for the next 3 frames;
[0038] Generate an instruction set containing joint angle correction amounts and suggestions for adjusting the movement rhythm based on the prediction results;
[0039] The loss function used for LSTM training is:
[0040]
[0041] where, is the total loss value for model training, is the predicted output at time t, y t is the true value, T is the time step, θ is the model parameter, and λ is the regularization coefficient.
[0042] Preferably, the system further includes:
[0043] A model adaptation module for dynamically adjusting the confidence threshold of the YOLOv8 model according to real-time detection results;
[0044] The execution steps of the model adaptation module include:
[0045] Statistical false positive rate and missed detection rate of target detection in the current scene;
[0046] When the false positive rate exceeds the first threshold, use the exponential decay algorithm to reduce the confidence threshold;
[0047] When the missed detection rate exceeds the second threshold, locally update the model weights through gradient backpropagation;
[0048] The formula of the exponential decay algorithm is:
[0049] c t = c0·e -kt
[0050] where, c t is the confidence threshold at time t, c0 is the initial threshold, k is the decay rate coefficient, and t is the number of iterations.
[0051] Preferably, the multi-object detection module further includes:
[0052] An abnormal action recognition sub-module for detecting abnormal trajectories of sports equipment;
[0053] The execution steps of the sub-module include:
[0054] Perform polynomial fitting on the trajectory of the ball game and calculate the fitting residuals;
[0055] When the residual exceeds the preset threshold, trigger the high-speed camera device to perform slow-motion playback analysis at 1000fps;
[0056] The calculation formula for the polynomial fitting residual is:
[0057]
[0058] where is the trajectory fitting residual, p i is the coordinate of the i-th trajectory point, t i is the timestamp, a j are the polynomial coefficients, d is the degree of the polynomial, and N is the total number of trajectory points.
[0059] Preferably, the system further includes:
[0060] A distributed computing module for implementing parallel data processing of multiple cameras; the execution steps of the distributed computing module include:
[0061] Build a message queue using Apache Kafka to synchronize timestamps for multiple video streams;
[0062] Distribute the detection tasks to the GPU cluster through the Dask framework to achieve real-time feedback with sub-second latency.
[0063] Compared with the prior art, the beneficial effects of the present invention are:
[0064] In terms of precise teaching evaluation, the system obtains real-time video streams and sports physiological data through the data acquisition module. The multi-object detection module locates teachers, students, sports equipment, and balls based on the YOLOv8 algorithm. The key point modeling module uses the OpenPose algorithm to build a three-dimensional action pose model. The action trajectory analysis module calculates the Euclidean distance deviation from the standard action to generate an action standard evaluation result. The application of this series of technical means makes the evaluation of students' actions more accurate and objective. Taking swimming teaching as an example, the system can accurately analyze details such as the amplitude and frequency of students' arm strokes, and the strength and angle of leg kicks. After comparing with the standard action, a detailed and accurate evaluation report can be obtained, helping teachers comprehensively understand the learning situation of students and providing a strong basis for teaching improvement.
[0065] From the perspective of improving sports safety guarantee, the multi-modal fusion module combines sports physiological data with the action standard evaluation result. If the student's exercise intensity is too high, resulting in an abnormal increase in heart rate, and at the same time the action is deformed, the dynamic feedback module will issue a warning in a timely manner. Teachers can adjust the teaching arrangement accordingly to avoid students being injured due to excessive fatigue or incorrect actions. For example, in long-distance running training, the system monitors the students' heart rate and running posture in real time. Once an abnormality is detected, the teacher can quickly let the students stop exercising or adjust the running method to ensure the students' physical health.
[0066] In terms of implementing personalized teaching, the system generates personalized action correction instructions through the dynamic feedback module based on the action evaluation results and physiological data of each student. For students with different physical fitness levels and learning progress, targeted training suggestions are provided. For example, in basketball shooting teaching, for students with insufficient arm strength, the system recommends strengthening arm strength training and adjusting the force application method during shooting; for students with poor coordination, a special coordination training plan is given to meet the different learning needs of students and improve learning effects.
[0067] In terms of improving teaching efficiency, the distributed computing module uses the Apache Kafka and Dask frameworks to achieve parallel data processing of multiple cameras, achieving real-time feedback with sub-second latency. Teachers can obtain students' assessment information in real time and adjust teaching strategies in a timely manner without spending a lot of time on manual observation and recording. At the same time, the model adaptation module can dynamically adjust the confidence threshold of the YOLOv8 model according to the detection results, and the abnormal action recognition sub-module can timely detect abnormal trajectories of sports equipment, ensuring the smooth progress of teaching and comprehensively improving the overall efficiency and quality of physical education teaching. Brief Description of the Drawings
[0068] Figure 1 It is the working principle diagram of the physical education teaching assessment system described in the present invention;
[0069] Figure 2 It is the flow chart of the key point modeling module;
[0070] Figure 3 It is the flow chart of multi-modal data fusion processing;
[0071] Figure 4 It is the flow chart of scoring and instruction generation of the dynamic feedback module. Detailed Embodiments
[0072] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0073] Please refer to Figures 1-4 , the present invention provides a technical solution: a physical education teaching assessment system based on artificial intelligence, aiming to comprehensively and accurately assess the physical education teaching process by using advanced technical means and provide strong support for teaching. The system includes:
[0074] Data Acquisition Module: This module is responsible for receiving the real-time video stream from the camera device and the motion physiological data collected by the wearable sensors. The camera device captures the activities of teachers and students, sports equipment, and ball games in the sports teaching scene in all directions; the wearable sensors monitor the physiological information such as heart rate, acceleration, and angular velocity of students during exercise in real time, providing a multi-dimensional data basis for subsequent comprehensive analysis.
[0075] Multi-object Detection Module: Based on the YOLOv8 algorithm, this module accurately locates the teachers, students, sports equipment, and balls in the real-time video stream. Through its powerful object location function, it can quickly determine the positions of each object in the video frame and output the detection box coordinates and object categories, such as distinguishing different objects like students, teachers, basketballs, footballs, etc., providing a prerequisite for subsequent in-depth analysis of actions and movement trajectories.
[0076] Key Point Modeling Module: Using the OpenPose algorithm, it extracts the skeletal key points of the detected target human body. Starting from each joint point of the human body, it constructs a three-dimensional action posture model, accurately recording the posture changes of the human body during movement, providing key data support for action analysis.
[0077] Action Trajectory Analysis Module: Based on the generated three-dimensional action posture model, it calculates the Euclidean distance deviation between the movement trajectory and the preset standard action. By comparing the differences between the actual action and the standard action, it generates an action standardization evaluation result, accurately judging the standardization of students' actions, and providing a quantitative basis for teaching improvement.
[0078] Multi-modal Fusion Module: It aligns the motion physiological data and the action standardization evaluation results in time series and constructs a multi-dimensional feature matrix. By organically combining different types of data, it fully explores the potential correlations between the data, providing richer and more comprehensive information for subsequent dynamic feedback.
[0079] Dynamic Feedback Module: Based on the multi-dimensional feature matrix, it generates real-time scores and personalized action correction instructions with the help of the LSTM network optimized by transfer learning. The real-time scores enable students and teachers to timely understand the learning and teaching effects; the personalized action correction instructions provide targeted improvement suggestions according to the specific situations of each student, helping students improve their action levels faster.
[0080] The present invention will be further described below in conjunction with Embodiments 1 to 6:
[0081] Embodiment 1:
[0082] In this embodiment, the specific execution process of the multi-object detection module is elaborated in detail. The multi-object detection module based on the YOLOv8 algorithm is the key link for the entire system to achieve accurate object location.
[0083] First, input the real-time video stream obtained from the camera device into the pre-trained YOLOv8 model. The YOLOv8 model is an advanced convolutional neural network with powerful feature extraction capabilities. In this process, the model processes the image data in the video stream through multiple convolutional layers to extract multi-scale feature maps. These feature maps contain image information at different levels. Small-scale feature maps can capture the detailed information of the target, while large-scale feature maps help identify the overall contour and position of the target. For example, when identifying a basketball, the small-scale feature map can clearly show the texture details on the surface of the basketball, and the large-scale feature map can quickly locate the approximate position of the basketball in the entire scene.
[0084] Next, perform anchor box matching and non-maximum suppression on the extracted feature maps. Anchor box matching refers to matching the bounding boxes predicted by the model with the pre-set anchor boxes. By calculating metrics such as the Intersection over Union (IoU) between the two, the bounding boxes that best match the target object are selected. Non-maximum suppression is used to remove redundant detection boxes. In the actual detection process, there may be multiple detection boxes pointing to the same target object. Non-maximum suppression will sort the detection boxes according to their confidence levels, retain the detection box with the highest confidence level, and remove those detection boxes with a high overlap with the highest-confidence detection box and a lower confidence level, thus ensuring that each target object is detected only once.
[0085] After this series of processes, the target detection results are output, including the coordinates of the detection box and the target category. The coordinates of the detection box accurately describe the position of the target in the video frame. For example, (x1, y1, x2, y2) respectively represent the coordinates of the upper left and lower right corners of the detection box; the target category clarifies what the detected object is, such as "student", "basketball", etc. In this way, the multi-target detection module can accurately and efficiently locate various targets in the sports teaching scene, providing a reliable data basis for subsequent analysis work.
[0086] Example 2:
[0087] When the multi-target detection module detects a human target, the key point modeling module first performs local image cropping on the detected human target. The entire human body is cropped out from the video frame, and only the image area related to the human body is retained, which can reduce the amount of data for subsequent processing and improve processing efficiency. Then, the cropped local image is input into the OpenPose network for heatmap prediction. The OpenPose network analyzes the input image and predicts the heatmaps of the skeletal key points of the human body. The value of each pixel point in the heatmap represents the probability of the skeletal key point appearing at that position. The higher the value, the more likely that position is where the skeletal key point is located.
[0088] Next, the two-dimensional coordinates of the skeletal key points are extracted through Gaussian filtering and peak detection algorithms. Gaussian filtering is a commonly used image smoothing method, which can smooth the heat map, remove noise interference, and make the heat map clearer and more stable. After Gaussian filtering, a peak detection algorithm is used to find the peak points in the heat map, and the coordinates corresponding to these peak points are the two-dimensional coordinates of the skeletal key points. For example, when detecting the key points of the human wrist, the two-dimensional coordinate position of the wrist in the image can be accurately found through the peak detection algorithm.
[0089] Finally, combining the depth information of the binocular camera, the two-dimensional coordinates are mapped to the three-dimensional space to generate an action pose sequence with timestamps. The binocular camera can obtain the depth information of the scene. Using this depth information, the previously obtained two-dimensional coordinates can be converted into three-dimensional coordinates, thereby constructing a three-dimensional action pose model of the human body. At the same time, a timestamp is added to each action pose to form an action pose sequence, recording the pose changes of the human body at different time points. In this way, through a series of operations of the key point modeling module, a three-dimensional action pose model of the human body can be accurately generated, providing accurate data support for subsequent action analysis.
[0090] Embodiment 3:
[0091] Based on the three-dimensional action pose model, the action trajectory analysis module comprehensively evaluates the actions of students. First, the three-dimensional action pose model is matched with the standard action template by dynamic time warping. In physical education teaching, there are corresponding standard actions for different sports events, and these standard action templates are pre-stored in the system. Dynamic Time Warping (DTW) can solve the problem of inconsistent time series data on the time axis. It finds an optimal warping path to align the joint angle sequence of the actual action with the joint angle sequence of the standard action in time, so as to make a more accurate comparison. Its calculation formula is:
[0092]
[0093] where A = (a1, …, a m ) and B = (b1, …, b n ) are the joint angle sequences of the actual action and the standard action respectively, P is the set of all warping paths, (p, q) is a pair of matching points in the path π, satisfying p ∈ A, q ∈ B, and d(a p , b q ) is the difference metric function of two angle values. Through this formula, the difference in joint angles between the actual action and the standard action is calculated.
[0094] After calculating the joint angle deviation, the Laplace smoothing algorithm is used to filter the noise in the motion trajectory. During the movement, due to the interference of various factors, the collected motion data may contain noise. The Laplace smoothing algorithm can effectively remove this noise, making the motion trajectory smoother and more accurate.
[0095] Then, the integral operation is performed on the trajectory deviation of consecutive action frames to output the action coherence score. Action coherence is an important indicator to measure the quality of students' actions. By integrating the trajectory deviation, the degree of action coherence can be quantified. For example, when a student is performing a shooting action, if the transition between each action link is smooth, the integral value of the trajectory deviation will be small, and the corresponding action coherence score will be high; on the contrary, if the action is jerky and not smooth, the integral value of the trajectory deviation will be large, and the action coherence score will be low.
[0096] At the same time, when calculating the action standardization, the Euclidean distance deviation calculation formula is also used:
[0097]
[0098] Among them, (x i , y i , z i ) represents the three-dimensional coordinates of the i-th joint point, (x std , y std , z std ) represents the corresponding coordinates of the standard action, and n is the total number of joint points. This formula is used to calculate the distance deviation between each joint point in the actual action and the corresponding joint point in the standard action in three-dimensional space. By synthesizing the distance deviations of all joint points, the overall action standardization evaluation result is obtained. Through these calculations and analyses, the action trajectory analysis module can comprehensively and accurately evaluate the standardization and coherence of students' actions, providing valuable references for teaching.
[0099] Example 4:
[0100] The multimodal fusion module aims to deeply fuse the motion physiological data with the action standardization evaluation results to obtain more comprehensive and accurate information. Wavelet transform denoising is performed on the heart rate, acceleration, and angular velocity in the motion physiological data. During the actual acquisition process, the motion physiological data is easily interfered by the external environment and the device itself, generating noise. Wavelet transform is an effective signal processing method. It can decompose the signal into different frequency bands, and by processing different frequency bands, the noise components are removed, and the useful signal features are retained. For example, when processing heart rate data, wavelet transform can remove the noise generated by factors such as motion artifacts, making the heart rate data more accurately reflect the real physiological state of students.
[0101] Perform feature-level fusion on the denoised physiological data and the action standard assessment results to generate a fused feature vector containing spatio-temporal correlation. Feature-level fusion directly fuses the features of different data sources to fully explore the potential connections between data. In the sports teaching scenario, the exercise physiological data reflects the physical function state of students, and the action standard assessment results reflect the sports skill level of students. There is a close connection between the two. Through feature-level fusion, this information is integrated together to form a fused feature vector containing spatio-temporal correlation. For example, in basketball teaching, when a student is dribbling quickly, if the heart rate is too high and the action standard is poor, the fused feature vector can reflect both aspects of information, providing a more comprehensive basis for subsequent analysis.
[0102] Dynamically allocate the weight ratio of visual data and sensor data through the attention mechanism. In multi-modal data fusion, different types of data may have different importance for the final analysis result. The attention mechanism can dynamically adjust the weights of visual data (such as image information in a video stream) and sensor data (such as exercise physiological data) according to the characteristics of the data and the current analysis requirements. In some cases, such as when analyzing a student's shooting action, the information about hand movements and the basketball trajectory in the visual data may be more important, and the attention mechanism will increase the weight of this part of the visual data; while when evaluating a student's physical energy consumption, information such as heart rate in the exercise physiological data is more crucial, and at this time the attention mechanism will increase the weight of the sensor data. In this way, the multi-modal fusion module can more flexibly and effectively integrate multi-source data, providing better data support for subsequent dynamic feedback.
[0103] Example 5:
[0104] This example elaborates in detail the specific implementation process of the dynamic feedback module, which provides real-time scoring and personalized action correction instructions for students.
[0105] The dynamic feedback module is one of the core modules for the system to implement the teaching assistance function. It generates real-time scoring and personalized action correction instructions based on the multi-dimensional feature matrix. The curriculum learning strategy is used to train the LSTM network in stages. The curriculum learning strategy simulates the process of humans learning new knowledge, learning from simple to complex gradually. In the initial stage of training, the parameters of the convolutional layer are frozen in the first stage. Since the convolutional layer is mainly responsible for extracting the features of images, these features are relatively stable in the initial stage of training. Freezing the parameters of the convolutional layer can speed up the training and avoid model instability caused by over-adjusting the convolutional layer in the initial stage of training. In the second stage, the fully connected layer is jointly optimized. The fully connected layer is responsible for classifying and predicting the extracted features. By jointly optimizing the fully connected layer, the parameters of the model can be better adjusted to improve the prediction accuracy of the model.
[0106] Perform sliding window segmentation on the multi-dimensional feature matrix and input it into the LSTM network to predict the action scores for the next 3 frames. Sliding window segmentation divides the multi-dimensional feature matrix according to a certain time step, and each window contains feature information for a period of time. Input these window data into the trained LSTM network in sequence. The LSTM network can capture the long-term dependencies in time series data and predict the action scores for the next 3 frames by learning from historical data.
[0107] Generate an instruction set containing joint angle correction amounts and motion rhythm adjustment suggestions based on the prediction results. For example, if the predicted shooting action score of a student in the next 3 frames is low and it is found through analysis that it is due to incorrect wrist joint angles, the system will generate corresponding joint angle correction amounts to guide the student to adjust the wrist angles; if it is also found that the student's shooting rhythm is too fast, the system will give suggestions for adjusting the motion rhythm to let the student slow down appropriately.
[0108] During the LSTM training process, the loss function used is:
[0109]
[0110] Among them, is the total loss value of model training, is the predicted output at time t, y t is the true value, T is the time step, θ is the model parameter, and λ is the regularization coefficient. This loss function optimizes the model parameters by calculating the difference between the predicted value and the true value and combining the regularization term, enabling the model to better fit the data, improve the prediction accuracy, and thus provide more reliable real-time scores and personalized action correction instructions for students.
[0111] Example 6:
[0112] This example covers the specific implementation manners of the model adaptation module, the abnormal action recognition sub-module in the multi-object detection module, and the distributed computing module.
[0113] The model adaptation module dynamically adjusts the confidence threshold of the YOLOv8 model according to the real-time detection results. First, count the false positive rate and missed detection rate of object detection in the current scene. The false positive rate refers to the proportion of misjudging non-target objects as target objects, and the missed detection rate refers to the proportion of not detecting actual existing target objects. Obtain these two key indicators through statistical analysis of the detection results over a period of time.
[0114] When the false positive rate exceeds the first threshold, use the exponential decay algorithm to reduce the confidence threshold. The exponential decay algorithm formula is:
[0115] c t = c0·e -kt
[0116] Among them, c t is the confidence threshold at time t, c0 is the initial threshold, k is the decay rate coefficient, and t is the number of iterations. By reducing the confidence threshold, the occurrence of false alarms can be reduced. For example, in a football teaching scenario, if the system frequently misjudges the sundries on the sidelines as footballs, when the false alarm rate exceeds the set first threshold, the exponential decay algorithm is used to reduce the confidence threshold, making the system more strict in judging the target in subsequent detections, thereby reducing false alarms.
[0117] When the missed detection rate exceeds the second threshold, the model weights are locally updated through gradient backpropagation. Gradient backpropagation is a commonly used neural network training optimization method. It calculates the gradient of the loss function with respect to the model parameters and updates the model parameters along the opposite direction of the gradient, enabling the model to optimize in the direction of reducing the loss function during training. In this case, locally updating the model weights can improve the model's detection ability for target objects and reduce the missed detection rate.
[0118] In the multi-object detection module, the abnormal action recognition sub-module is used to detect the abnormal trajectories of sports equipment. This sub-module first performs polynomial fitting on the ball movement trajectory and calculates the fitting residual. Polynomial fitting is to approximate the actual movement trajectory with a polynomial function, and its fitting residual calculation formula is:
[0119]
[0120] Among them, is the trajectory fitting residual, p i is the coordinate of the i-th trajectory point, t i is the timestamp, a j is the polynomial coefficient, d is the polynomial degree, and N is the total number of trajectory points. When the residual exceeds the preset threshold, it indicates that the ball movement trajectory is abnormal. At this time, the high-speed camera device is triggered to perform slow-motion playback analysis at 1000fps to more clearly observe and analyze the abnormal situation.
[0121] The distributed computing module is used to implement parallel data processing of multiple cameras. First, an Apache Kafka message queue is built to synchronize the timestamps of multiple video streams. Apache Kafka is a high-throughput distributed messaging system that can effectively manage and transmit large amounts of data streams. By adding timestamps to multiple video streams and synchronizing them, it can ensure that the data collected by different cameras is consistent in time, providing accurate time series data for subsequent analysis.
[0122] Then, the detection task is distributedly deployed to the GPU cluster through the Dask framework to achieve real-time feedback with sub-second latency. The Dask framework is an open-source framework for distributed computing, which can decompose and distribute computing tasks to multiple computing nodes for execution. The GPU cluster has powerful parallel computing capabilities. By deploying the detection task to the GPU cluster through the Dask framework, the computing resources of the GPU can be fully utilized, greatly improving the data processing speed, achieving real-time feedback with sub-second latency, and enabling the system to monitor and evaluate the sports teaching process in a timely manner. Through the collaborative work of these modules, the system can operate more intelligently and efficiently, providing more comprehensive and accurate support for sports teaching evaluation.
[0123] It should be noted that in this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device.
[0124] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. An artificial intelligence-based sports teaching evaluation system, characterized in that, Including: A data acquisition module for receiving real-time video streams from camera devices and motion physiological data from wearable sensors; A multi-object detection module that locates teachers and students, sports equipment, and balls in the real-time video stream based on the YOLOv8 algorithm, and outputs detection box coordinates and target categories; A key point modeling module that uses the OpenPose algorithm to extract skeletal key points from the detected target human body and generates a three-dimensional action pose model; An action trajectory analysis module that calculates the Euclidean distance deviation between the motion trajectory and a preset standard action according to the three-dimensional action pose model, and generates an action standardness evaluation result; A multi-modal fusion module that temporally aligns the motion physiological data and the action standardness evaluation result to construct a multi-dimensional feature matrix; A dynamic feedback module that generates real-time scores and personalized action correction instructions through an LSTM network optimized by transfer learning based on the multi-dimensional feature matrix.
2. The system according to claim 1, wherein The execution steps of the multi-object detection module include: Input the real-time video stream into a pre-trained YOLOv8 model, and extract multi-scale feature maps through a convolutional neural network; Perform anchor box matching and non-maximum suppression processing on the feature maps, and output the target detection results.
3. The system according to claim 1, wherein The execution steps of the key point modeling module include: Perform local image cropping on the detected human target and input it into the OpenPose network for heat map prediction; Extract the two-dimensional coordinates of the skeletal key points through Gaussian filtering and peak detection algorithms; Combine the depth information of the binocular camera to map the two-dimensional coordinates to the three-dimensional space and generate an action pose sequence with timestamps.
4. The system according to claim 1, wherein The execution steps of the action trajectory analysis module include: Perform dynamic time warping matching between the three-dimensional action pose model and a standard action template, and calculate the joint angle deviation; Perform filtering processing on the motion trajectory noise based on the Laplacian smoothing algorithm; Perform integral operation on the trajectory deviation of consecutive action frames and output the action coherence score; The calculation formula for the Euclidean distance deviation is: Among them, (x i , y i , z i ) represents the three-dimensional coordinates of the i-th joint point, (x std , y std , z std ) represents the corresponding coordinates of the standard action, and n is the total number of joint points.
5. The system according to claim 4, wherein The calculation formula for the dynamic time warping matching is: where A = (a1, …, a m ) and B = (b1, …, b n ) are the joint angle sequences of the actual action and the standard action respectively, P is the set of all regularized paths, (p, q) is a pair of matching points in the path π, satisfying p ∈ A, q ∈ B, and d(a p , b q ) is the difference metric function of two angle values.
6. The system according to claim 1, wherein The execution steps of the multi-modal fusion module include: Perform wavelet transform denoising on the heart rate, acceleration, and angular velocity in the motion physiological data; Perform feature-level fusion on the denoised physiological data and the action standardness evaluation result to generate a fusion feature vector containing spatio-temporal correlation; Dynamically allocate the weight ratio of visual data and sensor data through an attention mechanism.
7. The system according to claim 1, wherein The execution steps of the dynamic feedback module include: Adopt a curriculum learning strategy to train the LSTM network in stages. Freeze the convolutional layer parameters in the first stage and jointly optimize the fully connected layer in the second stage; Perform sliding window segmentation on the multi-dimensional feature matrix and input it into the LSTM network to predict the action scores of the next 3 frames; Generate an instruction set containing joint angle correction amounts and motion rhythm adjustment suggestions according to the prediction results; The loss function used for LSTM training is: Among them, is the total loss value of model training, is the predicted output at time t, y t is the true value, T is the time step, θ is the model parameter, and λ is the regularization coefficient.
8. The system according to claim 1, wherein It also includes: A model adaptation module for dynamically adjusting the confidence threshold of the YOLOv8 model according to real-time detection results; The execution steps of the model adaptation module include: Statistically calculate the false alarm rate and missed detection rate of target detection in the current scene; When the false alarm rate exceeds the first threshold, the exponential decay algorithm is used to reduce the confidence threshold; When the missed detection rate exceeds the second threshold, the model weights are locally updated through gradient backpropagation; The formula of the exponential decay algorithm is: c t = c0·e -kt ; where c t is the confidence threshold at time t, c0 is the initial threshold, k is the decay rate coefficient, and t is the number of iterations.
9. The system according to claim 1, wherein The multi-object detection module further includes: An abnormal action recognition sub-module for detecting abnormal trajectories of sports equipment; The execution steps of the sub-module include: Performing polynomial fitting on the ball movement trajectory and calculating the fitting residual; When the residual exceeds the preset threshold, triggering a high-speed camera device to perform slow-motion playback analysis at 1000fps; The formula for calculating the polynomial fitting residual is: Among them, is the trajectory fitting residual, p i is the coordinate of the i-th trajectory point, t i is the timestamp, a j is the polynomial coefficient, d is the polynomial degree, and N is the total number of trajectory points.
10. The system according to claim 1, characterized in that, It further includes: A distributed computing module for implementing parallel data processing of multiple cameras; The execution steps of the distributed computing module include: Using Apache Kafka to build a message queue for timestamp synchronization of multiple video streams; Distributing the detection tasks to the GPU cluster through the Dask framework to achieve real-time feedback with sub-second latency.
Citation Information
Patent Citations
Student classroom state identification method based on multi-source information fusion
CN118013389A
Human body action recognition system and method based on deep learning
CN118397692A
Intelligent classroom student posture key point detection method based on improved YOLOv8-Pose
CN118918341A
Cited By
Classroom behavior analysis method and analysis system for autistic children
CN120748050A
Multi-modal dynamic weight action evaluation method and system and storage medium
CN120771525A
Multimodal dynamic weighted action evaluation method, system and storage medium
CN120771525B
Competitive sports training monitoring system and method based on multi-modal data
CN120893001A
AI identification physical exercise trajectory data analysis teaching system
CN121388500A