Method for decomposing actions and extracting key frames during chin-ups and computer device

By extracting and analyzing the movement trajectory and velocity curve of the human body key points in the video data, and comparing it with the standard action model, automatically divide the pull-up action stage and selecting keyframes, the problems of capturing the critical moments of movement and evaluating the coordination of movement in the existing technology are solved, and the precise decomposition and personalized improvement suggestions are achieved.

CN119068555BActive Publication Date: 2025-05-30ZHUHAI XINWEI TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411484752.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-23
Publication Date
2025-05-30
Estimated Expiration
2044-10-23

AI Technical Summary

Technical Problem

The prior art is difficult to accurately capture critical moments in pull-up movements, and there is a lack of a comprehensive assessment of the overall coordination and efficiency of the movement in terms of action comparison analysis.

Method used

By collecting and preprocessing video data, the coordinate sequence of human key points is extracted, the motion trajectory and velocity curve are calculated, and the preset standard action model library is compared, and the action phase is automatically divided and the keyframes are selected.

Benefits of technology

Automatic capture and precise decomposition of pull-up actions is realized, the most representative keyframes can be extracted from each stage, detailed static breakdown diagrams of the key points of action, and personalized improvement suggestions are provided.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119068555B_ABST
    Figure CN119068555B_ABST
Patent Text Reader

Abstract

This application relates to a method for decomposing the movements during a chin-up and extracting key frames, as well as a computer device. The method includes: obtaining a sequence of key point coordinates; calculating the movement trajectories and velocity curves of the human body key points; selecting the most matching standard model from a preset standard action model library; determining the key stage division criteria for the student's chin-up action; dividing the student's chin-up action into N key stages; selecting candidate key frames corresponding to the key stages; performing optimized screening on the candidate key frames to obtain the final key frames; comparing the static decomposition diagrams with the most matching standard model, and giving improvement suggestions based on the comparison results. This application can automatically capture the complete movement process, accurately divide the movement stages, extract the most representative key frames from each stage, generate detailed static decomposition diagrams of the movement essentials, and consider the individual characteristics of students to provide personalized improvement suggestions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of machine vision technology, and in particular to a method and computer device for decomposing and extracting key frames during a pull-up process. Background Art

[0002] In the process of teaching pull-ups, teachers need to demonstrate and explain the students' movements, but due to the continuity and rapidity of the movement process, students often find it difficult to accurately grasp the key movements. At the same time, when correcting students' movements, teachers also need to frequently pause and replay the demonstration video to find the key mistakes, which greatly reduces the teaching efficiency.

[0003] Existing motion analysis methods mainly rely on manual observation and subjective judgment, which makes it difficult to provide accurate and objective evaluation. Although some systems use a single camera to capture motion, they cannot fully reflect the motion characteristics in three-dimensional space. In addition, existing technologies often use fixed time intervals or simple motion thresholds in key frame extraction, which makes it difficult to accurately capture the key moments in the action.

[0004] In terms of action comparison and analysis, current methods are mostly limited to simple trajectory comparison or key point position comparison, lacking a comprehensive evaluation of the overall coordination and efficiency of the action. At the same time, when generating improvement suggestions, existing systems can only provide general guidance and it is difficult to give precise improvement plans based on individual differences of students. Summary of the invention

[0005] Based on this, it is necessary to provide a pull-up process action decomposition and key frame extraction method and computer equipment that can automatically capture the complete action process to address the above technical problems.

[0006] A method for decomposing and extracting key frames of a pull-up process comprises the following steps:

[0007] Collect and preprocess the video data of students doing pull-ups, extract the key points of the human body from the preprocessed video data, and obtain the key point coordinate sequence;

[0008] According to the key point coordinate sequence, calculate the motion trajectory and speed curve of the key points of the human body;

[0009] Compare the motion trajectory and speed curve with the preset standard motion model library, and select the most matching standard model from the preset standard motion model library;

[0010] According to the motion trajectory, speed curve and the most matching standard model, determine the key stage division criteria of students' pull-up movements;

[0011] According to the key stage division criteria, the key point coordinate sequence, and the speed curve, the student's chin-up movement is divided into N key stages;

[0012] According to the key point coordinate sequence, the movement trajectory, and the speed curve, candidate key frames corresponding to the key stages are selected;

[0013] The candidate key frames are optimized and screened to obtain the final key frames; each key stage corresponds to at least one final key frame, and the number of final key frames is within a preset range.

[0014] In one embodiment, in the step of calculating the movement trajectory and speed curve of the human body key points according to the key point coordinate sequence, the steps include:

[0015] Using a low-pass filter to slide in the time dimension, locally polynomial fitting is performed on the coordinate values in the key point coordinate sequence to obtain the movement trajectory;

[0016] The differential method is used to process the movement trajectory to obtain a preliminary speed curve;

[0017] Using a low-pass filter to slide in the time dimension, locally polynomial fitting is performed on the speed values in the preliminary speed curve to obtain the speed curve.

[0018] In one embodiment, the preset standard action model library includes standard models of N chin-up movements; the standard models are divided into high-level standard models, medium-level standard models, and low-level standard models;

[0019] In the step of comparing the movement trajectory and speed curve with the preset standard action model library and selecting the most matching standard model from the preset standard action model library, the steps include:

[0020] Preprocessing is performed on the movement trajectory, speed curve, and standard model; the preprocessing includes time normalization and space normalization;

[0021] Using the dynamic time warping algorithm, calculate the first dynamic time warping distance between the preprocessed movement trajectory and the standard movement trajectory in the preprocessed standard model, and calculate the second dynamic time warping distance between the preprocessed speed curve and the standard speed curve in the preprocessed standard model;

[0022] According to the first dynamic time warping distance, the second dynamic time warping distance, and the dynamic time warping weight corresponding to the human body key points, a comprehensive distance index is obtained;

[0023] The standard model corresponding to the smallest comprehensive distance index is used as the most matching standard model.

[0024] In one embodiment, in the step of determining the key stage division criteria for the student's chin-up motion according to the motion trajectory, speed curve, and the most-matched standard model, the steps include:

[0025] Extract key parameters from the most-matched standard model, where the key parameters include speed thresholds, position thresholds, and time ratios corresponding to each stage of the chin-up motion.

[0026] Compare the motion trajectory, speed curve with the speed thresholds, position thresholds, and time ratios to obtain the key stage division criteria; the key stage division criteria include the start and end determination conditions for the key stages, the time ratio range of the key stages in the entire chin-up motion, and the dynamic adjustment rules.

[0027] In one embodiment, the key stages include a preparation stage, a pulling-up stage, a transition stage, and a lowering stage.

[0028] In the step of dividing the student's chin-up motion into N key stages according to the key stage division criteria, key point coordinate sequence, and speed curve, the steps include:

[0029] Starting from the first frame of the key point coordinate sequence, continuously monitor the first vertical speed of the student's wrist and hip according to the speed curve, and mark the frames where the first vertical speed meets the corresponding start and end determination conditions as the end point of the preparation stage and the start point of the pulling-up stage.

[0030] After identifying the end point of the preparation stage and the start point of the pulling-up stage, continuously monitor the student's head position and the second vertical speed of the wrist according to the key point coordinate sequence, and mark the frames where the head position and the second vertical speed meet the corresponding start and end determination conditions as the end point of the pulling-up stage and the start point of the transition stage.

[0031] After identifying the end point of the pulling-up stage and the start point of the transition stage, continuously monitor the third vertical speed of the student's hip according to the speed curve, and mark the frames where the third vertical speed meets the corresponding start and end determination conditions as the end point of the transition stage and the start point of the lowering stage.

[0032] After identifying the end point of the transition stage and the start point of the lowering stage, continuously monitor the fourth vertical speed of the student's wrist and hip according to the speed curve, and mark the frames where the fourth vertical speed meets the corresponding start and end determination conditions as the end point of the lowering stage.

[0033] Check and adjust the end point of the preparation stage, the start point of the pulling-up stage, the end point of the pulling-up stage, the start point of the transition stage, the end point of the transition stage, and the start point of the lowering stage according to the time ratio range to divide the student's chin-up motion into a preparation stage, a pulling-up stage, a transition stage, and a lowering stage.

[0034] During the critical stage division process, according to the dynamic adjustment rules, the key point coordinate sequence, and the velocity curve, dynamically adjust the start and end determination conditions and the time ratio range.

[0035] In one of the embodiments, in the step of selecting candidate key frames corresponding to the critical stage according to the key point coordinate sequence, the motion trajectory, and the velocity curve, it includes the steps:

[0036] In the preparation stage, according to the motion trajectory, calculate the average positions of the student's hip and shoulder in the preparation stage; according to the motion trajectory, calculate the offset distances between the positions of the student's hip and shoulder in each frame of the preparation stage and the average position, and take the frame corresponding to the smallest offset distance as the candidate key frame corresponding to the preparation stage;

[0037] In the upward pull stage, calculate the average vertical velocity in the upward pull stage according to the velocity curve, and mark the frame on the velocity curve that is closest to the average vertical velocity as the first feature frame; analyze the change speed of the student's elbow angle according to the motion trajectory, and mark the frame corresponding to the maximum change speed as the second feature frame; analyze the vertical relative position between the student's wrist and shoulder according to the motion trajectory, and take the frame with a vertical relative position of zero as the third feature frame; select the candidate key frame corresponding to the upward pull stage from the first feature frame, the second feature frame, and the third feature frame;

[0038] In the transition stage, mark the frame with the highest position of the student's head as the fourth feature frame according to the motion trajectory; calculate the student's elbow angle according to the motion trajectory, and mark the frame corresponding to the smallest elbow angle as the fifth feature frame; mark the frame when the student's hip starts to move forward as the sixth feature frame according to the motion trajectory; select the candidate key frame corresponding to the transition stage from the fourth feature frame, the fifth feature frame, and the sixth feature frame;

[0039] In the downward release stage, calculate the average vertical velocity in the downward release stage according to the velocity curve, and mark the frame on the velocity curve that is closest to the average vertical velocity as the seventh feature frame; analyze the change speed of the student's elbow angle according to the motion trajectory data, and mark the frame corresponding to the maximum change speed as the eighth feature frame; calculate the change in the included angle between the student's hip and shoulder according to the motion trajectory, and mark the frame when the student starts to lean forward according to the included angle change as the ninth feature; select the candidate key frame corresponding to the downward release stage from the seventh feature frame, the eighth feature frame, and the ninth feature frame.

[0040] In one of the embodiments, in the step of optimizing and screening the candidate key frames to obtain the final key frames, it includes the steps:

[0041] Sort the candidate key frames in chronological order to obtain a preliminary sorting;

[0042] Based on the preliminary sorting, identify the first candidate key frame in the preparation stage as the final key frame;

[0043] Traverse the remaining candidate key frames in the initial sorting, and calculate the time interval between the current candidate key frame and the last selected final key frame during the traversal;

[0044] Based on the preliminary judgment condition that the time interval is greater than the preset time interval threshold, and combined with judging the key features included in the candidate key frame, filter the remaining final key frames, and form the first final key frame and the remaining final key frames into a final key frame sequence.

[0045] In one of the embodiments, the method further includes the following steps:

[0046] Detect the time interval between adjacent final key frames in the final key frame sequence;

[0047] When it is detected that the time interval between two adjacent final key frames exceeds the preset maximum time interval threshold: obtain the speed values of all skipped candidate key frames between the two adjacent final key frames whose time interval exceeds the preset maximum time interval threshold; for each candidate key frame, calculate the sum of the absolute value of the speed difference between the current candidate key frame and the previous candidate key frame and the absolute value of the speed difference between the current candidate key frame and the next candidate key frame, as the speed change significance index of the current candidate key frame; select the candidate key frame with the largest speed change significance index as the additional key frame; add the additional key frame to the final key frame sequence.

[0048] A computer device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the steps of the above method are implemented.

[0049] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the above method are implemented.

[0050] One of the above technical solutions has the following advantages and beneficial effects:

[0051] This application includes the following steps: collecting video data of students doing chin-ups and preprocessing it, extracting human key points from the preprocessed video data to obtain a key point coordinate sequence; calculating the movement trajectory and speed curve of the human key points according to the key point coordinate sequence; comparing the movement trajectory and speed curve with a preset standard action model library, and selecting the most matching standard model from the preset standard action model library; determining the key stage division criteria for the students' chin-up actions according to the movement trajectory, speed curve, and the most matching standard model; dividing the students' chin-up actions into N key stages according to the key stage division criteria, key point coordinate sequence, and speed curve; selecting candidate key frames corresponding to the key stages according to the key point coordinate sequence, movement trajectory, and speed curve; optimizing and screening the candidate key frames to obtain the final key frames; at least one final key frame corresponds to each key stage, and the number of final key frames is within a preset range. This application can automatically capture the complete action process, accurately divide the action stages, extract the most representative key frames from each stage, generate detailed static decomposition diagrams of action essentials, and consider the individual characteristics of students to provide personalized improvement suggestions. Brief Description of the Drawings

[0052] Figure 1 It is a schematic flowchart of the method for action decomposition and key frame extraction during the chin-up process in an embodiment of this application.

[0053] Figure 2 It is an internal structure diagram of a computer device in an embodiment of this application. Detailed Embodiments

[0054] In order to make the purpose, technical solutions, and advantages of this application clearer, the following further details this application in combination with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0055] To achieve the above object, in one embodiment, as Figure 1 shown, a method for action decomposition and key frame extraction during the chin-up process is provided, including the following steps:

[0056] Step S110: Collect video data of students doing chin-ups and preprocess it, extract human key points from the preprocessed video data to obtain a key point coordinate sequence.

[0057] In the campus environment, fixed high-definition network cameras are installed near the chin-up equipment. Each piece of equipment is equipped with two cameras, one for frontal shooting and one for side shooting to obtain a comprehensive action perspective. The system uses network cameras with a resolution of 1080p and a frame rate of 30fps, which are connected to the school's local area network through a wireless network to ensure stable data transmission.

[0058] When students use the system, they need to enter their student IDs on the touch - screen terminal, and the system retrieves the students' basic information for subsequent analysis. After the students are ready, they click the "Start" button, and the system automatically starts recording for a preset duration of 30 seconds. After the recording is completed, the system quickly checks the video to ensure that the whole body of the student is within the frame.

[0059] Immediately after the video is captured, it enters the pre - processing stage. First, noise removal is performed. The Gaussian filtering algorithm is used to remove the noise in the frame. Then, background separation is carried out. The Mixture of Gaussian (MOG2) model is used for background modeling, and then background subtraction is applied to obtain the foreground mask.

[0060] In the motion detection step, the frame difference method is used to detect the motion area, and morphological operations are applied to the detected area to remove small noise points. The system determines the smallest rectangular area (ROI) that contains the student and crops the original video to this area to reduce the data volume for subsequent processing.

[0061] To improve the image quality, the system applies Contrast Limited Adaptive Histogram Equalization (CLAHE) to the ROI area. Subsequently, the frame rate is adjusted to ensure that the video frame rate is 30fps. If the frame rate is higher than 30fps, downsampling is performed; if it is lower than 30fps, linear interpolation is used to fill in the frames.

[0062] The system uniformly adjusts the resolution of the ROI area to 640x480 and uses the bicubic interpolation algorithm for scaling. For the convenience of subsequent analysis, a millisecond - level timestamp is added to the upper - left corner of each frame.

[0063] Finally, the system converts the processed video into the MP4 format encoded with H.264 and sets the bit rate to 2Mbps. After the pre - processing is completed, the system automatically checks the video quality, including brightness uniformity, clarity of the student's contour, and correctness of the timestamp. Qualified video files are stored on the school's central server, and relevant metadata is recorded in the database.

[0064] When extracting human key points from the preprocessed video data to obtain the key point coordinate sequence, the system first loads a pre-trained deep learning model. This model is based on an improved OpenPose framework and is trained with a large-scale human motion dataset, featuring high accuracy and robustness. This model has been specifically optimized to effectively handle various variations in the pull-up actions of students in the school sports scenario. The model takes the preprocessed video frames with a resolution of 640x480 as input and outputs the two-dimensional coordinates of 8 human key points. These 8 key points include the left and right wrists, left and right shoulders, left and right elbows, the center point of the hip, and the vertex of the head. Their selection is based on the principles of human kinematics and the characteristics of the pull-up action, ensuring that the key motion features in the action can be accurately captured. For example, the key points of the wrists and elbows are crucial for analyzing the movements of the upper and lower arms, while the center point of the hip helps to evaluate the overall body posture and balance.

[0065] The key point extraction process uses the sliding window technique, with 5 frames as a processing unit, and processes the entire video sequence step by step with a step size of 1 frame. This method utilizes temporal information to improve the accuracy and stability of key point detection. By considering the relationships between adjacent frames, the system can better handle detection difficulties caused by changes in perspective, lighting, or partial occlusion. The selection of the window size of 5 frames is based on the analysis results of the characteristics of the pull-up action, achieving a balance between temporal resolution and computational efficiency, being able to capture the continuity of the action while not introducing excessive delay.

[0066] The entire processing flow includes three core steps: feature extraction and heatmap generation, temporal integration, and key point localization and skeleton connection. In the feature extraction and heatmap generation stage, the system uses the ResNet-50 network to extract features from each frame of the 640x480 resolution image. ResNet-50 contains 50 convolutional layers and pooling layers, and outputs a feature map with 2048 channels. These feature maps are then input into a transposed convolutional network, which consists of 3 transposed convolutional layers, with a convolutional kernel size of 4x4 and a stride of 2 for each layer. The output channel number of the last layer is 8, corresponding to 8 key points. A heatmap is generated for each channel, with the same size as the input image, representing the presence probability of the corresponding key point at each position in the image.

[0067] In the temporal integration stage, the system uses a 3D convolutional neural network to process the heatmaps of 5 consecutive frames. This network contains 3 layers of 3D convolutional layers, with a convolutional kernel size of 3x3x3 and a stride of (1,1,1) for each layer. The input channel number of the first layer is 8 (corresponding to 8 key points), and the output channel number is 16; the output channel number of the second layer is 32; the output channel number of the third layer is still 8. Each 3D convolutional layer is followed by a ReLU activation function and a batch normalization layer. This process integrates the information in the time dimension into the heatmap of each frame, further improving the accuracy of key point detection.

[0068] The final key point localization and skeleton connection stage is divided into two sub-steps. First, the system applies the non-maximum suppression (NMS) algorithm to the temporally integrated heatmaps for key point localization. NMS uses a 3x3 sliding window on each heatmap, retains the local maximum, and suppresses the non-maxima around it. For each key point, the position with the highest probability value on the heatmap is selected as the coordinate of that key point. Then, the system uses a predefined skeleton model to connect the detected key points. The skeleton model defines 7 connections: left shoulder to left elbow, left elbow to left wrist, right shoulder to right elbow, right elbow to right wrist, left shoulder to right shoulder, midpoint between the two shoulders to the center of the hip, and center of the hip to the vertex of the head. The system traverses these predefined connections and connects the corresponding key point coordinates to form a complete skeleton structure.

[0069] Through this processing flow, the system finally outputs the two-dimensional coordinates of 8 key points in each frame (i.e., the key point coordinate sequence). The structure of the output data is a three-dimensional tensor with dimensions of number of frames × 8 × 2, where 8 represents the 8 key points and 2 represents the x and y coordinates of each key point.

[0070] Step S120, calculate the motion trajectory and velocity curve of the human body key points according to the key point coordinate sequence.

[0071] For each key point, its motion trajectory during the entire action can be represented as a series of two-dimensional coordinate points. For example, the motion trajectory of the left wrist can be represented as {(x_t, y_t)|t = 1, 2,..., T}, where T is the total number of frames of the video, and (x_t, y_t) is the coordinate of the left wrist at the t-th frame. The system performs the same processing on all 8 key points to obtain 8 sets of key point coordinate sequences.

[0072] In one example, in the step of calculating the motion trajectory and velocity curve of the human body key points according to the key point coordinate sequence, the steps include: sliding a low-pass filter in the time dimension, performing local polynomial fitting on the coordinate values in the key point coordinate sequence to obtain the motion trajectory; using the difference method to process the motion trajectory to obtain a preliminary velocity curve; sliding a low-pass filter in the time dimension, performing local polynomial fitting on the velocity values in the preliminary velocity curve to obtain the velocity curve.

[0073] To obtain a more accurate and smooth motion trajectory, the system applies the Savitzky-Golay filter to the human body key point coordinate sequence. This filtering method can effectively smooth the data while retaining the high-order moments of the signal, and is particularly suitable for processing motion data that may have noise but needs to retain fast-changing features. The window size of the filter is set to 15 frames, and the polynomial order is 3.

[0074] The filtering process is as follows: For each human body key point, the system applies the Savitzky-Golay filter to its x-coordinate sequence and y-coordinate sequence respectively. The filter slides in the time dimension and performs local polynomial fitting on the coordinate values at each time point. Specifically, for time point t, the filter considers a total of 15 points before and after it (if t is close to the start or end of the sequence, the number of points considered is reduced accordingly), fits these points with a third-order polynomial, and then uses the value of the fitting polynomial at point t as the filtered coordinate value. This process can be expressed as:

[0075] x'_t = Σ(c_n * x_{t + n}) (n ranges from -7 to 7)

[0076] y'_t = Σ(c_n * y_{t + n}) (n ranges from -7 to 7)

[0077] where c_n are the coefficients of the Savitzky-Golay filter, determined by the window size and polynomial order. (x'_t, y'_t) are the filtered coordinates.

[0078] Through this filtering process, the system obtains smooth motion trajectories. These motion trajectories not only eliminate the noise caused by detection errors or minor jitters but also retain the key features in the chin-up movement, such as the changes in rapid upward pulling and slow downward lowering.

[0079] The motion trajectories are stored in the form of two-dimensional coordinate sequences. For each human body key point k (k = 1, 2,..., 8, corresponding to the left and right wrists, left and right shoulders, left and right elbows, the center point of the hip, and the vertex of the head respectively), there is a sequence Tk = {(x'_k,t, y'_k,t) | t = 1, 2,..., T}. Where T is the total number of frames in the video. These motion trajectories intuitively show the spatial position changes of each human body key point during the chin-up process.

[0080] Next, the system calculates the velocity curve based on the smoothed trajectory data. The velocity is calculated using the central difference method, which is more accurate than simple forward difference or backward difference and can better handle non-linear changes. For any time point t, the velocity v_t is calculated as follows:

[0081] v_x_t = (x'_{t + 1} - x'_{t - 1}) / (2 * Δt)

[0082] v_y_t = (y'_{t + 1} - y'_{t - 1}) / (2 * Δt)

[0083] where Δt is the time interval between adjacent frames, usually 1 / 30 second (for a 30fps video). v_x_t and v_y_t are the velocity components in the x-direction and y-direction at time t respectively.

[0084] For points near the start and end of the sequence, the system uses forward differences or backward differences to calculate the velocity to ensure a complete velocity curve. Specifically:

[0085] For t = 1:

[0086] v_x_1=(x'_2 - x'_1) / Δt

[0087] v_y_1=(y'_2 - y'_1) / Δt

[0088] For t = T:

[0089] v_x_T=(x'_T - x'_{T - 1}) / Δt

[0090] v_y_T=(y'_T - y'_{T - 1}) / Δt

[0091] To further improve the quality of the velocity curve, the system applies the Savitzky - Golay filter to the calculated velocity data again. This filtering uses a smaller window size (7 frames) and a lower polynomial order (2nd order) to smooth out the minor fluctuations that may exist in the velocity curve while retaining the main features of the velocity changes. The filtering process is similar to that of the coordinate data but uses different parameters:

[0092] v'_x_t = Σ(c'_n * v_x_{t + n}) (n from - 3 to 3)

[0093] v'_y_t = Σ(c'_n * v_y_{t + n}) (n from - 3 to 3)

[0094] where c'_n are the new Savitzky - Golay filter coefficients, and (v'_x_t, v'_y_t) are the filtered velocity values.

[0095] Through these steps, the system obtains the smooth motion trajectories and corresponding velocity curves of each human body key point during the entire pull - up process. The velocity curve data is stored in the form of two one - dimensional arrays, corresponding to the velocity in the x - direction and the velocity in the y - direction respectively. For each human body key point k, there is a corresponding velocity curve Vk={(v'_x_k,t, v'_y_k,t)|t = 1,2,...,T}, where T is the total number of frames of the video.

[0096] These motion trajectories and velocity curve data comprehensively describe the motion state of the human body key points in space, including the changes in position and velocity. The motion trajectories reflect the movement paths of the human body key points in the two - dimensional plane, while the velocity curves represent the changes in the motion speed of these human body key points in the x and y directions.

[0097] For example, by observing the movement trajectory of the wrist, the rising height and movement path of a chin-up can be visually seen; while its speed curve reflects the speed changes during the upward and downward pulling processes. The movement trajectory of the hip can show the movement of the body's center of gravity, and its speed curve reflects the overall movement rhythm. The movement trajectory and speed curve of the shoulder can be used to analyze the movement characteristics of the upper body.

[0098] Step S130: Compare the movement trajectory and speed curve with a preset standard action model library, and select the most matching standard model from the preset standard action model library.

[0099] Among them, the preset standard action model library includes standard models of N chin-up actions; the standard models are divided into high-level standard models, medium-level standard models, and low-level standard models. The preset standard action model contains standard models of different-level chin-up actions. These models are divided into three levels: high, medium, and low. The low-level standard model reflects the basic action form, the medium-level standard model embodies the standard action requirements, and the high-level standard model demonstrates the action performance with high efficiency and high stability. This division enables the system to select the most suitable reference standard for students at different levels, thereby improving the accuracy of analysis and guidance.

[0100] Each standard model m contains data in the same format:

[0101] Movement trajectory Tm,k = {(x_m,k,t, y_m,k,t)|t = 1, 2,..., T_m}

[0102] And speed curve Vm,k = {(v_x_m,k,t, v_y_m,k,t)|t = 1, 2,..., T_m}, where T_m is the number of frames of this standard model.

[0103] In an example, in the step of comparing the movement trajectory and speed curve with the preset standard action model library and selecting the most matching standard model from the preset standard action model library, it includes the steps of: preprocessing the movement trajectory, speed curve, and standard model; the preprocessing includes time normalization and space normalization; using the dynamic time warping algorithm, calculating the first dynamic time warping distance between the preprocessed movement trajectory and the standard movement trajectory in the preprocessed standard model, and calculating the second dynamic time warping distance between the preprocessed speed curve and the standard speed curve in the preprocessed standard model; obtaining a comprehensive distance index according to the first dynamic time warping distance, the second dynamic time warping distance, and the dynamic time warping weight corresponding to the human key points; taking the standard model corresponding to the smallest comprehensive distance index as the most matching standard model.

[0104] 1. For effective comparison, the system first preprocesses the student data (including motion trajectories and velocity curves) and the standard model, including time normalization and space normalization:

[0105] (1) Time normalization:

[0106] Since the time for students to complete actions may be different from that of the standard model, all data needs to be resampled to the same number of time points N (e.g., N = 100). This process uses linear interpolation.

[0107] For student data Tk and Vk, for each human key point k, the data at T time points is resampled to N time points:

[0108] Tk_n = {(x_k,τ,y_k,τ)|τ = 1,2,...,N}

[0109] Vk_n = {(v_x_k,τ,v_y_k,τ)|τ = 1,2,...,N}

[0110] Among them, (x_k,τ,y_k,τ) and (v_x_k,τ,v_y_k,τ) are the positions and velocities of the τ-th sampling point obtained by linear interpolation. Here, τ (the Greek letter tau) is used to represent the index of the normalized time point to distinguish it from 'n' in the sequence name.

[0111] For each human key point k in the standard model m:

[0112] Tm,k_n = {(x_m,k,τ,y_m,k,τ)|τ = 1,2,...,N}

[0113] Vm,k_n = {(v_x_m,k,τ,v_y_m,k,τ)|τ = 1,2,...,N}

[0114] The 'n' in Vk_n and Vm,k_n both represents "normalized time" (time normalization)

[0115] (2) Space normalization:

[0116] To eliminate the influence of height differences, all spatial coordinates are divided by the corresponding height.

[0117] For student data, assuming the student's height is H:

[0118] Tk_s = {(x_k,τ / H,y_k,τ / H)|τ = 1,2,...,N}

[0119] For the standard model m, assuming the corresponding standard height is H_m:

[0120] Tm,k_s = {(x_m,k,τ / H_m,y_m,k,τ / H_m) | τ = 1,2,...,N}

[0121] The "s" in Tk_s and Tm,k_s represents "spatially normalized".

[0122] Note: The velocity data (Vk_n and Vm,k_n) do not need to be spatially normalized because they are already relative quantities.

[0123] 2. After the preprocessing is completed, the system starts data matching. The matching process uses the Dynamic Time Warping (DTW) algorithm. For each human key point k, the system calculates the DTW distance between its motion trajectory and velocity curve and the corresponding data of each standard model m.

[0124] The DTW distance of the trajectory is calculated as follows:

[0125] DTW_traj(k,m) = DTW(Tk_s,Tm,k_s)

[0126] The DTW distance of the velocity curve is calculated as follows:

[0127] DTW_vel(k,m) = DTW(Vk_n,Vm,k_n)

[0128] Where the DTW function calculates the dynamic time warping distance between two time series. Specifically, for two sequences A = {a_1,a_2,...,a_N} and B = {b_1,b_2,...,b_N}, the DTW distance is defined as:

[0129] DTW(A,B) = min(Σd(w_i) for all warping paths W = w_1,w_2,...,w_L)

[0130] Where d(w_i) = ||a_i - b_j|| is the Euclidean distance and W is all possible warping paths.

[0131] 3. Next, the system calculates the comprehensive distance metric for each standard model m:

[0132] D(m) = Σ(w_traj,k * DTW_traj(k,m) + w_vel,k * DTW_vel(k,m)) for k = 1 to 8

[0133] Where w_traj,k and w_vel,k are the weights of the trajectory and velocity of each human key point. These weights can be adjusted according to the importance of different human key points in the chin-up action.

[0134] The system calculates the comprehensive distance metric between the student data and each standard model in the model library, and then selects the model with the minimum distance as the most matching standard model:

[0135] best_model=argmin(D(m))forallminmodel_library

[0136] Step S140: Determine the key phase division criteria for the student's chin-up action based on the motion trajectory, speed curve, and the most matching standard model.

[0137] In an example, in the step of determining the key phase division criteria for the student's chin-up action based on the motion trajectory, speed curve, and the most matching standard model, it includes the steps of: extracting key parameters from the most matching standard model, where the key parameters include the speed thresholds, position thresholds, and time ratios corresponding to each phase of the chin-up action; comparing the motion trajectory, speed curve with the speed thresholds, position thresholds, and time ratios to obtain the key phase division criteria; the key phase division criteria include the start and end determination conditions of the key phases, the time ratio range of the key phases in the entire chin-up action, and the dynamic adjustment rules.

[0138] The system first extracts key parameters from the selected most matching standard model. The key parameters include the speed thresholds, position thresholds, and time ratios for each phase. For example, in the standard model, it may define that the speed threshold at the start of the upward pull phase is 1.2 m / s, the distance threshold between the wrist and the horizontal bar at the end of the upward pull phase is 20 cm, and the typical time ratio of the upward pull phase is 30% - 45% of the total time. The key parameters will serve as the initial reference values, and the system will adjust them according to the student's actual data.

[0139] Next, the system deeply analyzes the student's actual motion trajectory Tk and speed curve Vk, and compares them with the parameters of the most matching standard model. This process will conduct a detailed analysis for each phase of the chin-up action. In the analysis of the preparation phase, the system carefully examines the speed changes of the wrist and hip in Vk to find the time point when the speed starts to increase significantly. The system pays special attention to the inflection point of the speed curve, which usually indicates that the student starts to generate the intention of upward pull. Suppose the standard model defines the speed threshold at the end of the preparation phase as 1.2 m / s, but after analyzing the student's Vk data, the system finds that the actual starting speed is 0.9 m / s. This difference is recorded by the system as an important basis for adjusting the speed threshold at the end of the preparation phase. The system not only considers this speed value but also analyzes the slope of the speed increase to ensure capturing the real action start point. At the same time, the system will also deeply analyze the position data in Tk to check the posture changes of the student in the preparation phase, including observing the relative positions of the body joint points, especially the angle changes between the wrist, shoulder, and hip.

[0140] In the analysis of the upward pull phase, the system will more comprehensively examine the speed curve in Vk and the position changes in Tk. First, the maximum upward pull speed of the student is extracted from Vk, assumed to be 1.8 m / s, while the maximum upward pull speed of the standard model is 2.2 m / s. The system not only records this peak difference but also analyzes the speed change pattern during the entire upward pull process, including the change in acceleration, the shape of the speed curve (such as whether there are multiple peaks), and the duration of speed maintenance. Then, the system deeply analyzes the time point when the speed starts to significantly decline. If it is found that the student's speed starts to rapidly decline when it reaches 70% of the maximum value (while it may be 50% for the standard model), the system will carefully consider the reasons behind this difference, which may reflect the student's ability to maintain at the high point or the pattern of strength decay.

[0141] In terms of position, the system details the position changes in Tk when the head or wrist approaches the crossbar, including the movement in the vertical and horizontal directions. Suppose the system finds that the student starts to decelerate when they are 22 cm away from the crossbar, while in the standard model, this distance is 20 cm. The system will carefully consider the impact of this 2-cm difference on the overall movement quality and at the same time analyze the body posture of the student at this position, including the tilt angle of the torso and the degree of arm bending.

[0142] In the analysis of the transition phase, the system focuses on the movement characteristics of the hips in Tk and Vk. The system mainly analyzes the change in hip speed in Vk, especially paying attention to the time point when the speed changes from positive to negative. Suppose the system finds that this change point for the student occurs at 58% of the entire movement duration, while it may be 55% for the standard model. The system will carefully evaluate the impact of this 3% difference, considering whether the student has more staying time at the high point or whether there are unnecessary movement delays.

[0143] In the analysis of the downward phase, the system comprehensively examines the downward speed curve in Vk and the position recovery in Tk. The system will detail the change pattern of the downward speed, including the initial falling speed, the fluctuation of the speed, and the deceleration control near the end. Suppose the system finds that the average downward speed of the student is 1.3 m / s, while it may be 1.5 m / s for the standard model. The system will carefully evaluate the impact of this speed difference on the movement quality and safety, and at the same time analyze the position recovery in Tk, including the order and time for each part of the body to return to the initial position.

[0144] The system also determines the typical time proportion ranges for each phase in the entire movement. The preparation phase usually accounts for 10% - 20% of the total time, the upward pull phase accounts for 30% - 45%, the transition phase accounts for 10% - 25%, and the downward phase accounts for 25% - 40%. The setting of these time proportion ranges refers to both the standard model and takes into account the actual situation of ordinary students, providing a flexible standard for subsequent movement evaluation.

[0145] To adapt to the individual differences of different students, the system also formulates a set of dynamic adjustment rules and ranges. In terms of the speed threshold, if it is detected that the maximum upward pulling speed of a student is lower than 60% of the standard model, the system will reduce all speed-related thresholds by 25%. On the contrary, if the maximum upward pulling speed of a student exceeds 130% of the standard model, the system will increase the speed threshold by 15%. In terms of the height threshold, the system will adjust according to the student's height. If the student's height exceeds or is lower than 7% of the standard model, the system will correspondingly adjust all height-related thresholds by 5%. In terms of the time ratio, if the overall completion time of a student exceeds 30% of the standard model, the system allows the time ratio of each stage to float up and down by 7 percentage points.

[0146] Finally, the key stage division criteria for the chin-up action determined by the system include: the determination conditions for the start and end of each stage (preparation, upward pull, transition, downward release), the typical time ratio range of each stage in the whole action, and the rules and ranges of dynamic adjustment. These division criteria and dynamic adjustment rules comprehensively consider the ideal characteristics of the standard model and the actual performance of students, and can adapt to students of different levels and body types, providing personalized action analysis and evaluation.

[0147] Step S150, divide the student's chin-up action into N key stages according to the key stage division criteria, the key point coordinate sequence, and the speed curve. Among them, the key stages include the preparation stage, the upward pull stage, the transition stage, and the downward release stage;

[0148] In an example, in the step of dividing the student's chin-up action into N key stages according to the key stage division criteria, the key point coordinate sequence, and the speed curve, it includes the steps:

[0149] Starting from the first frame of the key point coordinate sequence, continuously monitor the first vertical speed of the student's wrist and hip according to the speed curve, and mark the frames where the first vertical speed meets the corresponding start and end determination conditions as the end point of the preparation stage and the start point of the upward pull stage;

[0150] After identifying the end point of the preparation stage and the start point of the upward pull stage, continuously monitor the second vertical speed of the student's head position and wrist according to the key point coordinate sequence, and mark the frames where the head position and the second vertical speed meet the corresponding start and end determination conditions as the end point of the upward pull stage and the start point of the transition stage;

[0151] After identifying the end point of the upward pull stage and the start point of the transition stage, continuously monitor the third vertical speed of the student's hip according to the speed curve, and mark the frames where the third vertical speed meets the corresponding start and end determination conditions as the end point of the transition stage and the start point of the downward release stage;

[0152] After identifying the end point of the transition phase and the start point of the lowering phase, continuously monitor the fourth vertical velocity of the student's wrist and hip according to the velocity curve, and mark the frames where the fourth vertical velocity meets the corresponding start and end determination conditions as the end point of the lowering phase;

[0153] According to the time ratio range, check and adjust the end point of the preparation phase, the start point of the pulling-up phase, the end point of the pulling-up phase, the start point of the transition phase, the end point of the transition phase, and the start point of the lowering phase, so as to divide the student's chin-up movement into a preparation phase, a pulling-up phase, a transition phase, and a lowering phase;

[0154] During the key phase division process, dynamically adjust the start and end determination conditions and the time ratio range according to the dynamic adjustment rules, the key point coordinate sequence, and the velocity curve.

[0155] The system uses a sliding window of 15 frames to analyze the entire action sequence frame by frame. Within each window, the system comprehensively uses the key point coordinate sequence and velocity data to calculate key metrics. Specifically:

[0156] First, the system identifies the preparation phase. According to the start and end determination conditions, the system starts detecting from the first frame of the action. Using the velocity curve, the system continuously monitors the vertical velocity of the wrist and hip. When it detects that the vertical velocities of both parts exceed 1% / second of the height simultaneously, the system marks this frame as the end of the preparation phase and also the start of the pulling-up phase.

[0157] Next, the system starts to identify the pulling-up phase. During this process, the system continuously monitors the vertical velocity of the wrist and the head position. The system first records the maximum vertical velocity during the pulling-up process, and then uses the velocity curve to detect the decrease in velocity. At the same time, the system uses the key point coordinate sequence to calculate the distance between the head and the horizontal bar. When the vertical velocity of the wrist drops below 50% of the peak velocity and the distance between the head and the horizontal bar is less than 20 cm, the system marks this frame as the end of the pulling-up phase and also the start of the transition phase.

[0158] The identification of the transition phase mainly depends on the movement state of the hip. The system uses the velocity curve to continuously monitor the vertical velocity of the hip. When it detects that the vertical velocity of the hip changes from positive to negative, the system marks this frame as the end of the transition phase and also the start of the lowering phase.

[0159] Finally, the system identifies the lowering phase. During this process, the system monitors the vertical velocities of the wrist and hip simultaneously, as well as their positions. Using the velocity curve, the system detects the change in velocity; using the key point coordinate sequence, the system calculates the difference between the current position and the initial position. When the vertical velocities of the wrist and hip are both less than 1% / second of the height and their positions are close to the initial position, the system marks this frame as the end of the lowering phase and also the end of the entire action.

[0160] Throughout the partitioning process, the system applies the defined dynamic adjustment rules. For example, if it is detected during the upward pull phase that the student's maximum upward pull speed is lower than 60% of the standard model, the system will reduce all speed-related thresholds by 25%. This means that the determination speed threshold for the end of the upward pull phase may be reduced from 50% of the peak speed to 37.5%. Similarly, if the student's height significantly deviates from the standard model, the system will adjust all height-related thresholds accordingly. For example, for a student whose height exceeds the standard model by 7%, the determination threshold for the distance from the horizontal bar may be increased from 20 cm to 21 cm.

[0161] After the preliminary partitioning is completed, the system checks and adjusts according to the specified typical time proportion ranges for each phase. The system calculates the time proportion occupied by each phase and compares it with the expected range. If the time proportion of a certain phase significantly deviates from the expectation, the system will recheck the boundary of that phase. For example, if the proportion of the upward pull phase in the total time exceeds 45% (higher than the expected 30% - 45%), the system will re-analyze the end point of this phase and may end the upward pull phase earlier, correspondingly extending the start time of the transition phase.

[0162] When making these adjustments, the system repeatedly utilizes the key point coordinate sequence and the speed curve. For example, when adjusting the end point of the upward pull phase, the system will re-analyze the vertical speed change (speed curve) and body posture change (motion trajectory) during this time period to find a more appropriate demarcation point. If the student's overall completion time exceeds 30% of the standard model, the system will allow a greater floating space for the time proportion of each phase, which can float up and down by 7 percentage points.

[0163] Through this comprehensive and refined process, the system successfully divides the continuous action sequence into four key phases: preparation, upward pull, transition, and lowering. Finally, the system generates a concise data structure as output. This data structure contains the start frame and end frame of each phase, as well as the total number of frames of the entire action sequence. For example:

[0164] {"Total number of frames": 180, "Preparation phase": [1, 30], "Upward pull phase": [31, 90], "Transition phase": [91, 105], "Lowering phase": [106, 180]}

[0165] Step S160, select candidate key frames corresponding to the key phases according to the key point coordinate sequence, motion trajectory, and speed curve.

[0166] In an example, in the step of selecting candidate key frames corresponding to the key phases according to the key point coordinate sequence, motion trajectory, and speed curve, it includes the steps:

[0167] In the preparation stage, according to the movement trajectory, calculate the average positions of the student's hip and shoulder in the preparation stage; according to the movement trajectory, calculate the offset distances between the hip position and shoulder position of each frame in the preparation stage and the average position, and take the frame corresponding to the minimum offset distance as the candidate key frame corresponding to the preparation stage;

[0168] In the upward pull stage, calculate the average vertical speed in the upward pull stage according to the speed curve, and mark the frame closest to the average vertical speed on the speed curve as the first feature frame; analyze the change speed of the student's elbow angle according to the movement trajectory, and mark the frame corresponding to the maximum change speed as the second feature frame; analyze the vertical relative position of the student's wrist and shoulder according to the movement trajectory, and take the frame corresponding to the vertical relative position of zero as the third feature frame; select the candidate key frame corresponding to the upward pull stage from the first feature frame, the second feature frame and the third feature frame;

[0169] In the transition stage, mark the frame with the highest head position of the student as the fourth feature frame according to the movement trajectory; calculate the elbow angle of the student according to the movement trajectory, and mark the frame corresponding to the minimum elbow angle as the fifth feature frame; mark the frame when the student's hip starts to move forward as the sixth feature frame according to the movement trajectory; select the candidate key frame corresponding to the transition stage from the fourth feature frame, the fifth feature frame and the sixth feature frame;

[0170] In the lowering stage, calculate the average vertical speed of the lowering stage according to the speed curve, and mark the frame closest to the average vertical speed on the speed curve as the seventh feature frame; analyze the change speed of the student's elbow angle according to the movement trajectory data, and mark the frame corresponding to the maximum change speed as the eighth feature frame; calculate the change in the included angle between the student's hip and shoulder according to the movement trajectory, and mark the frame when the student starts to lean forward according to the change in the included angle as the ninth feature; select the candidate key frame corresponding to the lowering stage from the seventh feature frame, the eighth feature frame and the ninth feature frame.

[0171] Specifically, in the preparation stage, the main goal of the system is to find the frame that best represents the stable preparation posture. The system first uses the movement trajectory to calculate the average positions of the hip and shoulder throughout the preparation stage. Assuming there are 30 frames in the preparation stage, the system will extract the hip and shoulder coordinates of these 30 frames from the movement trajectory data, accumulate their x coordinates and y coordinates, and then divide them by 30 respectively to obtain the average position. This average position represents an ideal static state of the body throughout the preparation stage.

[0172] Next, the system calculates the distances between the hip and shoulder positions in each frame of the preparation stage and this average position. This calculation process directly uses the motion trajectory data. Specifically, the system calculates the differences in the x-coordinate of the hip, the y-coordinate of the hip, the x-coordinate of the shoulder, and the y-coordinate of the shoulder for each frame. Then, the squares of these four differences are added together and the square root is taken to obtain a comprehensive distance value. This distance value reflects the degree of deviation of the current frame from the ideal stationary state.

[0173] For example, if the hip position in a certain frame is shifted 2 centimeters to the right and 1 centimeter up from the average position, and the shoulder position is shifted 1 centimeter to the left and 2 centimeters down from the average position, then the total deviation distance for this frame is the square root of the sum of the squares of these four deviation values, which is approximately 3.16 centimeters. The system calculates this deviation distance for each frame of the preparation stage.

[0174] The system selects the frame with the smallest deviation distance as the candidate key frame. If there are multiple frames with very close deviation distances, such as a difference less than 0.5 centimeters, the system selects the frame that is closest in time to the midpoint of the preparation stage. For example, if the preparation stage ranges from frame 1 to frame 30, then the midpoint is frame 15. If the deviation distances of frame 14 and frame 16 are both small and not very different, the system selects frame 15 as the candidate key frame.

[0175] In the pulling-up stage, the system needs to capture multiple key features because this stage contains the most complex and crucial parts of the pull-up action. First, the system uses the velocity curve data to calculate the average vertical velocity for the entire pulling-up stage. Assuming there are 60 frames in the pulling-up stage, the system extracts the vertical velocity values for these 60 frames from the velocity curve, accumulates them, and divides by 60 to obtain the average velocity. This average velocity represents the typical velocity of the entire pulling-up process.

[0176] Then, the system finds the point on the velocity curve where the velocity is closest to this average value and marks the corresponding frame as the feature frame representing the average velocity. For example, if the average velocity is 2 meters per second, the velocity of a certain frame is 1.98 meters per second, and the velocities of other frames vary more, then this frame may be selected as the feature frame representing the average velocity. This frame usually appears in the middle section of the pulling-up stage and reflects the average intensity of the action.

[0177] Next, the system will analyze the motion trajectory data to determine the rate of change of the elbow angle. Based on the coordinates of the shoulder, elbow, and wrist in the motion trajectory, the system calculates the elbow angle for each frame, then computes the change in angle between each adjacent pair of frames to find the frame with the largest change. This frame typically represents the moment of most concentrated effort and the most intense movement. For example, if in a certain frame, the elbow angle decreases by 15 degrees compared to the previous frame, while the changes in other frames are less than 10 degrees, then this frame may be identified as the frame with the fastest angle change. This frame usually appears in the early or middle stage of the pull-up, reflecting the moment of main force application.

[0178] The system also analyzes the relative position of the wrist and shoulder using the motion trajectory data to find the frame where the wrist is exactly level with the shoulder. This is achieved by comparing the y-coordinates of the wrist and shoulder in each frame. If in a certain frame, the difference between the y-coordinate of the wrist and the y-coordinate of the shoulder is no more than 2 centimeters, then this frame may be considered the frame where the wrist is level with the shoulder. This frame represents a key transition point in the pull-up action and usually indicates that the upper body is about to cross the horizontal bar.

[0179] If these three characteristic frames (average speed frame, frame with the fastest angle change, wrist-shoulder level frame) are close in time, for example, the difference in their frame numbers does not exceed 20% of the total number of frames in the pull-up stage, then the system will select the middle one of them as the candidate key frame. If these three frames are more dispersed, the system will select all three frames as candidate key frames to ensure that different characteristics of the pull-up stage can be captured.

[0180] In the transition stage, although this stage is usually short, it contains the key turning points of the pull-up action. The system first uses the motion trajectory data to find the frame with the highest head position. This is achieved by comparing the y-coordinates of the head in each frame. This frame represents the moment when the pull-up action reaches the highest point and is the peak of the movement.

[0181] Then, the system calculates the elbow angle for each frame using the motion trajectory data and finds the frame with the smallest angle, which represents the moment of maximum arm flexion. This frame reflects the limit state of the action and usually corresponds to the maximum force output.

[0182] The system also analyzes the hip motion trajectory to find the frame when the hip starts to move forward. This is achieved by detecting the horizontal speed of the hip. The system calculates the horizontal speed of the hip for each frame, and when the horizontal speed for several consecutive frames (such as 3 frames) exceeds a small positive value (such as 10 centimeters per second), it is considered that the hip starts to move forward. This frame marks the beginning of the body's preparation for the downward movement and is the turning point from the upward to the downward movement.

[0183] If these three feature frames are close in time, for example, the difference in their frame numbers does not exceed 30% of the total number of frames in the transition phase, then the system will select the middle frame among them as the candidate key frame. Otherwise, all three frames will be selected as candidate key frames to ensure that all key features in the transition phase are captured.

[0184] In the lowering phase, the system also relies on the velocity curve and motion trajectory data for analysis. First, the system calculates the average vertical velocity of the entire lowering phase using the velocity curve data. This velocity is usually negative because the body is descending. Then, the system locates the point on the velocity curve where the velocity is closest to this average value and marks the corresponding frame as the feature frame representing the typical lowering velocity. This frame reflects the overall performance of the athlete in controlling the descent.

[0185] Next, the system analyzes the change in the elbow angle using the motion trajectory data and finds the frame with the fastest angle change. This usually occurs at the beginning of the lowering phase and reflects the rapid transition from the bent - arm state to the extended - arm state. The system calculates the change in the elbow angle between each adjacent pair of frames and finds the frame with the largest rate of change. This frame is very important for understanding the force control during the lowering process.

[0186] The system also calculates the change in the angle between the hip and the shoulder using the motion trajectory data and finds the frame when the body starts to lean forward. This can be achieved by comparing the angle of each frame with the average angle in the preparation phase. The system first calculates the hip - shoulder angle for all frames in the preparation phase and takes the average value as a reference. Then, in the lowering phase, when the difference between the angle of a certain frame and this reference value exceeds a certain degree (such as 10 degrees), it is considered that the body starts to lean forward. This forward - leaning movement is an important feature in the lowering process, reflecting the skills of maintaining balance and controlling the descent speed.

[0187] If these three feature frames (the average - velocity frame, the frame with the fastest angle change, and the forward - leaning frame) are close in time, for example, the difference in their frame numbers does not exceed 25% of the total number of frames in the lowering phase, then the system will select the middle frame among them as the candidate key frame. Otherwise, all three frames will be selected as candidate key frames to ensure that all key features in the lowering phase are captured.

[0188] Step S170: Optimally screen the candidate key frames to obtain the final key frames; at least one final key frame corresponds to the key phase, and the number of final key frames is within a preset range.

[0189] In one example, in the step of optimizing and screening candidate key frames to obtain the final key frames, the steps include: sorting the candidate key frames in chronological order to obtain a preliminary sorting; based on the preliminary sorting, identifying the first candidate key frame in the preparation stage as the final key frame; traversing the remaining candidate key frames in the preliminary sorting, and calculating the time interval between the current candidate key frame and the most recently selected final key frame during the traversal process; based on the preliminary judgment condition that the time interval is greater than the preset time interval threshold, and combining with judging the key features included in the candidate key frame, screening the remaining final key frames.

[0190] The system first performs a preliminary sorting on all candidate key frames. This sorting is based on the frame numbers to ensure that the system can process these frames in chronological order. After the sorting is completed, the system creates a new list to store the finally screened key frames.

[0191] Next, the system sets a preset time interval threshold. The setting of this threshold needs to comprehensively consider the total duration of the pull-up action and the expected number of key frames. For example, if the entire action lasts for 3 seconds (assuming a video frame rate of 30fps, a total of 90 frames), and it is expected to finally have 8 - 10 key frames, then the minimum time interval can be set to approximately 0.3 seconds (9 frames). This setting is not fixed, and the system will make dynamic adjustments according to specific situations to adapt to the movement characteristics and speeds of different students.

[0192] The optimization and screening process starts with selecting the first candidate key frame. Usually, the system selects the first candidate frame in the preparation stage as the initial key frame and adds it to the final key frame list. This frame represents the standard preparation posture before the action starts and is crucial for understanding the entire action. Then, the system starts to traverse the remaining candidate key frames and conducts a detailed analysis and judgment on each frame.

[0193] During the traversal process, the system first calculates the time interval between the current candidate key frame and the most recently selected final key frame during the traversal. If this interval is less than the set preset time interval threshold, the system usually temporarily skips this frame and continues to check the next candidate frame. The purpose of doing this is to avoid selecting overly dense key frames, which may lead to redundant information in the final action decomposition diagram. However, the system does not apply this rule completely mechanically. If the current frame represents a particularly important action feature, even if its interval from the previous key frame is small, the system may choose to retain it.

[0194] If the time interval meets the requirements, the system will further evaluate the key features of the current frame. This includes checking whether the frame represents important features of a specific stage, such as the maximum speed point in the upward pull stage or the highest point in the transition stage. The system will refer to the selection reasons and feature data of each candidate key frame and comprehensively consider its importance. At the same time, the system will also compare the current frame with the selected key frames to check for information redundancy. For example, if the current frame and the previous key frame both represent similar action features (such as both being force-applying frames in the upward pull stage), the system will carefully compare the feature data of the two frames and select the one that better represents the feature.

[0195] During this process, the system always maintains attention to the overall action structure. It will ensure that there is at least one key frame for each key stage (preparation, upward pull, transition, downward release). If it is found that a certain stage does not have a key frame and the current frame belongs to this stage, the system will give priority to selecting this frame, even if it may not be the most representative frame in this stage. This approach ensures that the final set of key frames can completely cover the entire action process.

[0196] Based on the above comprehensive judgment, if the current frame passes all the checks, the system will add it to the final key frame list. If it does not pass, the system will continue to check the next candidate frame. This process will continue until all candidate frames have been evaluated.

[0197] After traversing all candidate frames, the system will perform an overall check and adjustment. First, the system will confirm again whether there is at least one final key frame for each key stage. If it is found that a certain stage does not have a final key frame, the system will go back to the candidate key frames of this stage and select the most representative one to add to the final list. This ensures that the integrity of the action will not be affected by the screening process.

[0198] Next, the system will check whether the total number of final key frames is within the expected range. If the total number exceeds the preset upper limit (for example, 10 frames), the system will perform further screening. This may involve deleting some relatively less important key frames while ensuring that the overall performance of the action is not affected. During this process, the system will give priority to retaining those frames that represent key action turning points, such as the frame at the start of the upward pull, the frame at the highest point, the frame at the start of the downward release, etc. On the contrary, if the number of key frames is insufficient, the system will re-examine the previously skipped candidate frames and may relax some screening criteria to increase the number of key frames. This dynamic adjustment ensures that the final set of key frames can comprehensively reflect the action features without affecting the efficiency of subsequent processing due to excessive quantity.

[0199] In an example, the method for decomposing the action process and extracting key frames during a chin-up in this application further includes the following steps:

[0200] Detect the time interval between adjacent final key frames in the final key frame sequence;

[0201] When it is detected that the time interval between two adjacent final key frames exceeds the preset maximum time interval threshold: Obtain the speed values of all skipped candidate key frames between the two adjacent final key frames whose time interval exceeds the preset maximum time interval threshold; For each candidate key frame, calculate the sum of the absolute value of the speed difference between the current candidate key frame and the previous candidate key frame and the absolute value of the speed difference between the current candidate key frame and the next candidate key frame, as the speed change significance index of the current candidate key frame; Select the candidate key frame with the largest speed change significance index as the additional key frame; Add the additional key frame to the final key frame sequence.

[0202] The system also pays special attention to the time interval between adjacent key frames. If it is found that the interval between two adjacent key frames is too large (for example, exceeding 25% of the entire action duration), the system will try to add an additional key frame in this interval. This additional frame is usually selected from the previously skipped candidate frames, aiming to better describe the continuity of the action and avoid the "jumping" feeling in the action description. When selecting the additional key frame, the system mainly focuses on the speed change of the action. Specifically, the system first obtains the speed values of all skipped candidate key frames within the large time interval. For each candidate frame, the system calculates the speed difference v1 between it and the previous frame and the speed difference v2 between it and the next frame. Then, the system calculates the sum of the absolute values of these two differences, that is, |v1| + |v2|, as the speed change significance index of this frame. This calculation method takes into account the speed changes of the candidate frame with its previous and next frames, and can effectively capture the mutation points of the speed. The system traverses all candidate frames and records the significance index values of each frame. Finally, the system compares the significance indices of all candidate frames and selects the frame with the largest index value as the frame with the most significant speed change. If there are multiple maximum index values that are the same, the system will select the candidate frame that is closest to the midpoint of the two original key frames in time. After selection, the system inserts this frame as an additional key frame into the final key frame sequence.

[0203] After completing all these optimizations and adjustments, the system will sort the finally selected key frame sequence once to ensure that they are strictly arranged in chronological order. At the same time, the system will also add additional information to each key frame, including the action phase it belongs to, its importance in this phase, and the specific action characteristics it represents. These additional information not only helps with subsequent action analysis and teaching guidance, but also provides important context information for generating coherent action decomposition diagrams.

[0204] For example, the final key frame list may include: the stable posture frame in the preparation stage, the initial force application frame at the start of the pulling-up stage, the maximum speed frame in the middle of the pulling-up stage, the wrist approaching the horizontal bar frame at the end of the pulling-up stage, the highest point of the head frame in the transition stage, the hip starting to move forward frame in the transition stage, the controlled falling frame at the start of the lowering stage, the typical falling speed frame in the middle of the lowering stage, and the frame approaching the starting position at the end of the lowering stage. Such a set of key frames not only covers all the key stages of the pull-up action but also captures important action details within each stage.

[0205] Through this meticulous optimization and screening process, the set of key frames finally obtained by the system can comprehensively and accurately reflect the entire pull-up action while maintaining a reasonable time distribution between the key frames.

[0206] The method for decomposing the pull-up process action and extracting key frames in this application further includes the steps of: visualizing the final key frames to generate a static decomposition diagram of the student's pull-up action, comparing the static decomposition diagram with the most matching standard model, and giving improvement suggestions based on the comparison results.

[0207] In an example, in the steps of visualizing the final key frames to generate a static decomposition diagram of the student's pull-up action, comparing the static decomposition diagram with the most matching standard model, and giving improvement suggestions based on the comparison results, it includes the steps of: preprocessing the final key frames, adding corresponding text information to the preprocessed final key frames to generate a static decomposition diagram; generating a standard static decomposition diagram based on the most matching standard model; presenting the static decomposition diagram and the standard static decomposition diagram side by side and highlighting the different parts; comparing the static decomposition diagram and the standard static decomposition diagram, giving improvement suggestions, and displaying the improvement suggestions in the static decomposition diagram.

[0208] First, the system reads the set of final key frames. This set usually contains 8 - 10 final key frames, and each final key frame represents an important moment or posture in the pull-up action. The system processes these frames in chronological order to ensure that the finally generated decomposition diagram can clearly show the progress of the action.

[0209] For each key frame, the system performs a series of image processing and enhancement operations. This includes improving the clarity, adjusting the contrast, and removing noise from the original image, etc. These processes are aimed at ensuring that each key frame in the final decomposition diagram is clearly distinguishable, facilitating students and coaches to observe details. At the same time, the system marks key body parts on each key frame, such as the wrist, elbow, shoulder, hip, and head. These markings usually use eye-catching colors and appropriate graphic symbols, such as dots or small arrows, to highlight the positions and movement trajectories of these key points.

[0210] Next, the system adds concise text descriptions to each keyframe. These descriptions include the action phase to which the frame belongs (such as "preparation phase", "early pull-up phase", etc.), as well as the specific action characteristics represented by the frame (such as "maximum speed point", "transition start", etc.). The text descriptions use clear, easy-to-read fonts and are placed in a position that does not obstruct the key action area. In addition, the system will add arrows or lines between adjacent keyframes to indicate the flow direction of the action and enhance the coherence of the decomposition diagram.

[0211] In order to better demonstrate the continuity of the action, the system may insert some semi-transparent transition frames between the main key frames. These transition frames will not contain detailed annotations, but can help viewers better understand the continuous change process of the action. At the same time, the system will add a timeline at the top or bottom of the entire decomposition diagram, indicating the relative time position of each key frame in the entire action, to help viewers understand the rhythm and time allocation of the action.

[0212] After the static decomposition diagram of the student's action is completed, the system will read the selected standard model that best matches it. The system will use the same visualization method to generate a corresponding standard static decomposition diagram for the best matching standard model. This standard static decomposition diagram will be displayed side by side with the student's static decomposition diagram for direct comparison. In order to highlight the difference between the two, the system will use different color schemes, such as blue for the student's static decomposition diagram and green for the standard static decomposition diagram.

[0213] Next, the system will conduct a detailed comparative analysis. This analysis process will take into account multiple aspects: first, posture comparison. The system will compare the body postures of the students and the standard model at each critical moment, including the relative positions and angles of key points (such as wrists, elbows, shoulders, and hips). Secondly, time allocation comparison. The system will analyze whether the time allocation of the students at each stage is reasonable and whether it is consistent with the standard model. Thirdly, motion trajectory comparison. The system will compare the key point motion trajectories of the students and the standard model to identify the parts with large deviations.

[0214] Based on these comparative analyses, the system will generate a series of improvement suggestions. These suggestions will be directly marked on the student's decomposition diagram, and presented in a way that is eye-catching but does not affect the overall appearance of the picture. For example, if it is found that the angle of the student's elbow in the pull-up phase is significantly different from that of the standard model, the system will add a suggestion mark to the corresponding keyframe, accompanied by a concise text description, such as "Note: When pulling up, you should tighten your elbow more to reduce the angle between the upper arm and the torso." If the student's time allocation at a certain stage is obviously unreasonable, the system will also mark it on the timeline to prompt the student to adjust the rhythm.

[0215] An example to illustrate the whole process:

[0216] High school student Wang is learning pull-ups but has been unable to complete the standard movements. The PE teacher decides to use the system of this invention to help Wang improve his technique.

[0217] Step 1: When Wang stands in front of the pull-up equipment, the system automatically activates two high-definition cameras to start recording. One camera shoots from the front and the other from the side to ensure a full-angle view of Wang's movements is captured. After the recording is completed, the system automatically preprocesses the video, removing background noise and highlighting Wang's body contour.

[0218] Step 2: The system uses an improved human pose estimation algorithm to accurately extract 8 key points of Wang's body from the preprocessed video, including the left and right wrists, shoulders, elbows, hip center, and the top of the head. Even during the rapid upward pull phase of Wang's movement, the system can accurately track the positions of these key points.

[0219] Step 3: Based on the extracted key point data, the system calculates the movement trajectories and speed curves of various key parts of Wang during the entire pull-up process. Through filtering, the system eliminates the data fluctuations caused by the slight shaking of Wang's body and obtains smooth movement curves.

[0220] Step 4: The system compares Wang's movement data with a pre-established standard movement model library. Considering Wang's height (170 cm) and weight (60 kg), the system selects an intermediate-level standard model of a similar body type as a reference.

[0221] Step 5: Based on the selected standard model and Wang's actual data, the system formulates a standard for dividing Wang's movement phases. Since Wang is a beginner, the system appropriately relaxes the speed requirement for the upward pull phase, reducing the standard from 1.5 m / s to 1.2 m / s.

[0222] Step 6: Using the standard formulated in Step 5, the system automatically divides Wang's entire movement sequence into four phases: preparation, upward pull, transition, and downward release. The system accurately identifies the moment when Wang starts to pull upward forcefully and marks it as the starting point of the upward pull phase.

[0223] Step 7: In each divided phase, the system selects the most representative frames as candidate key frames. For example, during the upward pull phase, the system selects the frames when Wang's speed reaches the average value, the frames with the fastest elbow angle change, and the frames when the wrists are level with the shoulders.

[0224] Step 8: The system optimally screens the candidate key frames to ensure that the time intervals between them are appropriate and that there is at least one key frame in each phase. Finally, the system selects 9 key frames for Wang's movement, comprehensively reflecting his pull-up process.

[0225] Step 9: The system processes the filtered key frames into static decomposition diagrams and compares them with the standard model. Through analysis, the system finds that the elbow angle of Xiao Wang is too large during the upward pull phase, resulting in a decrease in the power transfer efficiency. Based on this finding, the system generates targeted improvement suggestions, such as "should tighten the elbows more when doing the upward pull to reduce the angle between the upper arm and the torso". According to these suggestions, the physical education teacher formulates a targeted training plan for Xiao Wang to help him improve his chin-up technique.

[0226] This application includes the following steps: collecting video data of students doing chin-ups and preprocessing it, extracting human key points from the preprocessed video data, and obtaining the key point coordinate sequence; calculating the movement trajectory and speed curve of the human key points according to the key point coordinate sequence; comparing the movement trajectory and speed curve with a preset standard action model library, and selecting the most matching standard model from the preset standard action model library; determining the key stage division criteria for the students' chin-up actions according to the movement trajectory, speed curve, and the most matching standard model; dividing the students' chin-up actions into N key stages according to the key stage division criteria, the key point coordinate sequence, and the speed curve; selecting candidate key frames corresponding to the key stages according to the key point coordinate sequence, movement trajectory, and speed curve; optimizing and screening the candidate key frames to obtain the final key frames; at least one final key frame corresponds to each key stage, and the number of final key frames is within a preset range; performing visualization processing on the final key frames to generate static decomposition diagrams of the students' chin-up actions, comparing the static decomposition diagrams with the most matching standard model, and giving improvement suggestions based on the comparison results. This application can automatically capture the complete action process, accurately divide the action stages, extract the most representative key frames from each stage, generate detailed static decomposition diagrams of the action essentials, and consider the individual characteristics of students to provide personalized improvement suggestions.

[0227] 1. This application realizes personalized action evaluation through an intelligent standard model matching and dynamic adjustment mechanism. The system establishes a standard model library containing various body types and skill levels. The intelligent matching algorithm comprehensively considers the static parameters and dynamic factors of students and selects the most matching standard model from the model library. This method overcomes the limitations of the traditional "one-size-fits-all" evaluation method. The dynamic adjustment mechanism can adjust the evaluation criteria in real time according to the actual performance of students. For example, it appropriately reduces the requirement for the upward pull speed for beginners and increases the action stability standard for high-level students. This ensures that the evaluation criteria always match the actual abilities of students and improves the accuracy and adaptability of action evaluation.

[0228] 2. This application adopts the technologies of automatic action segmentation and intelligent key frame extraction, providing a scientific and systematic method for action decomposition. The automatic segmentation algorithm uses signal processing and pattern recognition technologies to accurately divide the entire chin-up action into key stages such as preparation, pulling up, transition, and lowering. The key frame extraction technology makes intelligent selections based on the importance of action features, rather than simply selecting at fixed time intervals. For example, during the pulling-up stage, the system will preferentially select frames at critical moments when the speed reaches its peak, the elbow angle changes most significantly, and the wrist is level with the shoulder. This method ensures that the selected key frames can reflect the key features of the action to the greatest extent, while improving the efficiency and consistency of analysis.

[0229] 3. This application introduces the key frame optimization technology, further enhancing the accuracy and practicality of action analysis. The optimization process balances the time intervals between key frames through intelligent algorithms, ensuring that each action stage has sufficient key frame descriptions while avoiding information redundancy. The system also considers the importance weights of different action stages and adjusts the key frame distribution of each stage according to the level and characteristics of the learner. For example, more key frames may be selected during the pulling-up stage for beginners, while the number of key frames is increased during the transition stage for high-level athletes. This intelligent optimization ensures that the final set of selected key frames can comprehensively and accurately reflect the characteristics of the entire chin-up action, providing reliable data support for subsequent action evaluation and guidance.

[0230] It should be understood that although Figure 1 the steps in the flowchart of Figure 1 are shown sequentially in the direction of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless there is a clear indication in this article, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Moreover,

[0231] In one embodiment, a computer device is provided. This computer device can be a server, and its internal structure diagram can be as Figure 2As shown in the figure. The computer device includes a processor, a memory, a network interface, and a database connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data such as video data, motion trajectories, and speed curves. The network interface of the computer device is used to communicate with an external terminal via a network connection. When the computer program is executed by the processor, it implements a method for decomposing the actions of a chin-up process and extracting key frames.

[0232] Those skilled in the art can understand that Figure 2 the structure shown in the figure is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0233] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:

[0234] Collect video data of students doing chin-ups and perform preprocessing, extract human key points from the preprocessed video data, and obtain a sequence of key point coordinates;

[0235] According to the sequence of key point coordinates, calculate the motion trajectory and speed curve of the human key points;

[0236] Compare the motion trajectory and speed curve with a preset standard action model library, and select the most matching standard model from the preset standard action model library;

[0237] According to the motion trajectory, speed curve, and the most matching standard model, determine the key stage division criteria for the students' chin-up actions;

[0238] According to the key stage division criteria, the sequence of key point coordinates, and the speed curve, divide the students' chin-up actions into N key stages;

[0239] According to the sequence of key point coordinates, motion trajectory, and speed curve, select candidate key frames corresponding to the key stages;

[0240] Optimize and screen the candidate key frames to obtain the final key frames; at least one final key frame corresponds to a key stage, and the number of final key frames is within a preset range;

[0241] Visualize the final key frames to generate static decomposition diagrams of the student's chin-up action, compare the static decomposition diagrams with the most matching standard model, and give improvement suggestions based on the comparison results.

[0242] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0243] Collect video data of the student doing chin-ups and preprocess it, extract human key points from the preprocessed video data, and obtain the key point coordinate sequence;

[0244] According to the key point coordinate sequence, calculate the movement trajectories and speed curves of the human key points;

[0245] Compare the movement trajectories and speed curves with the preset standard action model library, and select the most matching standard model from the preset standard action model library;

[0246] According to the movement trajectories, speed curves and the most matching standard model, determine the key stage division criteria for the student's chin-up action;

[0247] According to the key stage division criteria, the key point coordinate sequence and the speed curve, divide the student's chin-up action into N key stages;

[0248] According to the key point coordinate sequence, movement trajectories and speed curves, select candidate key frames corresponding to the key stages;

[0249] Optimize and screen the candidate key frames to obtain the final key frames; at least one final key frame corresponds to each key stage, and the number of final key frames is within the preset range;

[0250] Visualize the final key frames to generate static decomposition diagrams of the student's chin-up action, compare the static decomposition diagrams with the most matching standard model, and give improvement suggestions based on the comparison results.

[0251] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the various embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in many forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0252] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0253] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application shall be subject to the appended claims.

Claims

1. A method for decomposing and extracting key frames during a pull-up process, characterized in that: The following steps are involved: Collecting and preprocessing video data of students doing pull-ups, extracting key points of the human body from the preprocessed video data, and obtaining a key point coordinate sequence; According to the key point coordinate sequence, calculating the motion trajectory and speed curve of the key points of the human body; Comparing the motion trajectory and the speed curve with a preset standard motion model library, and selecting the most matching standard model from the preset standard motion model library; Determining the key stage division criteria of the student's pull-up action according to the motion trajectory, the speed curve and the best matching standard model; According to the key stage division standard, the key point coordinate sequence and the speed curve, the student's pull-up action is divided into N key stages; the key stages include a preparation stage, a pull-up stage, a transition stage and a lowering stage; Selecting candidate key frames corresponding to the key stage according to the key point coordinate sequence, the motion trajectory and the speed curve; Optimizing and screening the candidate key frames to obtain the final key frames; The key stage corresponds to at least one final key frame, and the number of the final key frames is within a preset range; Processing the final key frame into a static decomposition diagram, and comparing the static decomposition diagram with the body posture, time allocation and movement trajectory of the standard model at each key moment, so as to analyze the student's pull-up and obtain improvement suggestions; The step of determining the key stage division standard of the student's pull-up action according to the motion trajectory, the speed curve and the best matching standard model includes the following steps: Extracting key parameters from the best matching standard model, wherein the key parameters include a speed threshold, a position threshold, and a time ratio corresponding to each stage of the pull-up action; The motion trajectory, the speed curve, the speed threshold, the position threshold and the time ratio are compared to obtain the key stage division standard; the key stage division standard includes the start and end judgment conditions of the key stage, the time ratio range of the key stage in the entire pull-up action and the dynamic adjustment rule; The step of dividing the student's pull-up action into N key stages according to the key stage division standard, the key point coordinate sequence and the speed curve includes the steps of: Starting from the first frame of the key point coordinate sequence, the first vertical speed of the student's wrist and hip is continuously monitored according to the speed curve, and the frame where the first vertical speed meets the corresponding start and end judgment conditions is marked as the end point of the preparation phase and the start point of the pull-up phase; After identifying the end point of the preparation phase and the start point of the pull-up phase, continuously monitoring the student's head position and the second vertical speed of the wrist according to the key point coordinate sequence, marking the frame where the head position and the second vertical speed meet the corresponding start and end judgment conditions as the end point of the pull-up phase and the start point of the transition phase; After identifying the end point of the pull-up phase and the start point of the transition phase, continuously monitoring the third vertical speed of the student's hip according to the speed curve, marking the frame where the third vertical speed meets the corresponding start and end judgment conditions as the end point of the transition phase and the start point of the lowering phase; After identifying the end point of the transition phase and the start point of the lowering phase, continuously monitoring the fourth vertical speed of the student's wrist and hip according to the speed curve, and marking the frame where the fourth vertical speed meets the corresponding start and end judgment conditions as the end point of the lowering phase; According to the time ratio range, the end point of the preparation phase, the start point of the pull-up phase, the end point of the pull-up phase, the start point of the transition phase, the end point of the transition phase and the start point of the lowering phase are checked and adjusted to divide the student's pull-up action into the preparation phase, the pull-up phase, the transition phase and the lowering phase; In the key stage division process, dynamically adjusting the start and end judgment conditions and the time ratio range according to the dynamic adjustment rules, the key point coordinate sequence and the speed curve; The dynamic adjustment rule is that if the student's maximum pull-up speed is detected to be lower than 60% of the standard model during the pull-up stage, the system will reduce all speed-related thresholds by 25%; if the student's overall completion time exceeds 30% of the standard model, the system will allow a larger floating space for the time proportion of each stage, which can fluctuate up or down by 7 percentage points; if the student's height exceeds or falls below the standard model by 7%, the system will adjust all height-related thresholds by 5% accordingly; The step of optimizing and screening the candidate key frames to obtain the final key frames includes the following steps: Sorting the candidate key frames in chronological order to obtain a preliminary sorting; Based on the preliminary sorting, identifying the first candidate key frame in the preparation stage as the final key frame; Traversing the remaining candidate key frames in the preliminary sorting, and calculating the time interval between the current candidate key frame and the last selected final key frame in the traversal process; Based on the preliminary judgment condition that the time interval is greater than a preset time interval threshold, and in combination with judging the key features contained in the candidate key frames, the remaining final key frames are screened, and the first final key frame and the remaining final key frames are combined into a final key frame sequence; The following steps are also included: Detecting the time intervals between adjacent final key frames in the final key frame sequence; When it is detected that the time interval between two adjacent final key frames exceeds the preset maximum time interval threshold: obtain the speed values ​​of all skipped candidate key frames between the two adjacent final key frames whose time interval exceeds the preset maximum time interval threshold; for each candidate key frame, calculate the sum of the absolute value of the speed difference between the current candidate key frame and the previous candidate key frame and the absolute value of the speed difference between the current candidate key frame and the next candidate key frame as the speed change significance index of the current candidate key frame; select the candidate key frame with the largest speed change significance index as an additional key frame; and add the additional key frame to the final key frame sequence.

2. The pull-up process action decomposition and key frame extraction method according to claim 1 is characterized in that: The step of calculating the motion trajectory and speed curve of the key points of the human body according to the key point coordinate sequence comprises the steps of: Using a low-pass filter to slide in the time dimension, local polynomial fitting is performed on the coordinate values ​​in the key point coordinate sequence to obtain the motion trajectory; Processing the motion trajectory using a differential method to obtain a preliminary velocity curve; The speed curve is obtained by sliding a low-pass filter in the time dimension and performing local polynomial fitting on the speed values ​​in the preliminary speed curve.

3. The pull-up process action decomposition and key frame extraction method according to claim 1 is characterized in that: The preset standard action model library includes N standard models of pull-up actions; the standard models are divided into high-level standard models, medium-level standard models and low-level standard models; The step of comparing the motion trajectory and the speed curve with a preset standard motion model library and selecting the most matching standard model from the preset standard motion model library comprises the steps of: Preprocessing the motion trajectory, the velocity curve and the standard model; the preprocessing includes time normalization and space normalization; Using a dynamic time warping algorithm, calculate a first dynamic time warping distance between the preprocessed motion trajectory and the preprocessed standard motion trajectory in the standard model, and calculate a second dynamic time warping distance between the preprocessed velocity curve and the preprocessed standard velocity curve in the standard model; Obtaining a comprehensive distance index according to the first dynamic time warping distance, the second dynamic time warping distance, and the dynamic time warping weights corresponding to the key points of the human body; The standard model corresponding to the smallest comprehensive distance index is used as the best matching standard model.

4. The pull-up process action decomposition and key frame extraction method according to claim 1, characterized in that: The step of selecting candidate key frames corresponding to the key stage according to the key point coordinate sequence, the motion trajectory and the speed curve comprises the steps of: In the preparation stage, according to the motion trajectory, the average position of the student's hip and shoulder in the preparation stage is calculated; according to the motion trajectory, the offset distance between the student's hip position and shoulder position of each frame in the preparation stage and the average position is calculated, and the frame corresponding to the smallest offset distance is used as the candidate key frame corresponding to the preparation stage; In the pull-up stage, an average vertical speed in the pull-up stage is calculated according to the speed curve, and a frame closest to the average vertical speed on the speed curve is marked as a first feature frame; Analyze the change speed of the student's elbow angle according to the motion trajectory, and mark the frame corresponding to the maximum change speed as the second feature frame; Analyze the vertical relative position of the student's wrist and shoulder according to the motion trajectory, and use the frame corresponding to the vertical relative position of zero as the third feature frame; select the candidate key frame corresponding to the pull-up stage from the first feature frame, the second feature frame and the third feature frame; In the transition phase, marking the frame with the highest position of the student's head according to the motion trajectory as the fourth feature frame; Calculating the student's elbow angle according to the motion trajectory, and marking the frame corresponding to the smallest elbow angle as the fifth feature frame; Marking the frame where the student's hip starts to move forward as the sixth feature frame according to the motion trajectory; Selecting a candidate key frame corresponding to the transition stage from the fourth feature frame, the fifth feature frame, and the sixth feature frame; In the lowering stage, the average vertical speed of the lowering stage is calculated according to the speed curve, and the frame closest to the average vertical speed on the speed curve is marked as the seventh feature frame; Analyze the change speed of the student's elbow angle according to the motion trajectory data, and mark the frame corresponding to the maximum change speed as the eighth feature frame; calculate the change of the angle between the student's hip and shoulder according to the motion trajectory, and mark the frame where the student starts to lean forward according to the change of the angle as the ninth feature frame; A candidate key frame corresponding to the dropping stage is selected from the seventh feature frame, the eighth feature frame and the ninth feature frame.

5. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.

6. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 4 are implemented.

Citation Information

Patent Citations

  • Sports test method and system, electronic equipment, chip and storage medium

    CN117292288A