Motion evaluation method, device, electronic device, and storage medium
By combining human figure detection and human posture estimation models with multi-level cascade and multi-modal keyframe localization models, the subjectivity and accuracy issues in standing long jump evaluation are solved, realizing an automated, accurate and convenient evaluation solution.
Patent Information
- Application Number
- CN202411209616.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-30
- Publication Date
- 2025-12-26
- Estimated Expiration
- 2044-08-30
AI Technical Summary
Existing standing long jump evaluation methods suffer from the subjectivity of manual measurement, the complexity of infrared equipment, and the insufficient accuracy of keyframe positioning in computer vision solutions, which affect the accuracy and efficiency of the evaluation.
By employing a human figure detection model, a human posture estimation model, and a keyframe localization model, key actions such as take-off and landing in the standing long jump are automatically identified and located through image processing. The keyframe localization model is combined with multi-level cascade and multi-modal models to improve the localization accuracy of keyframes.
It has achieved automation, objectivity and accuracy in standing long jump evaluation, simplified equipment deployment, improved the accuracy and efficiency of evaluation, and reduced human intervention.
Smart Images

Figure CN119097892B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of motion evaluation, in particular to a motion evaluation method and device, electronic equipment and storage medium. BACKGROUND
[0002] The existing evaluation of standing long jump project usually relies on manual measurement by physical education teachers or the use of infrared measuring equipment. Manual measurement is determined by the teacher according to visual observation or the use of measuring tools to determine the long jump distance, while the infrared equipment captures and calculates the long jump result through a complex system. These two methods are widely used in physical education, aiming to ensure the accuracy and fairness of student physical fitness tests.
[0003] However, the existing standing long jump project evaluation has some problems. First, manual measurement is easily affected by the subjective judgment of the teacher, resulting in differences in the results of the same testers. Second, although the infrared measuring equipment is relatively accurate, it is complex and cumbersome to deploy, which is not conducive to the simplicity and flexibility of physical education. In addition, the computer vision-based evaluation scheme that has emerged in recent years, although it uses advanced technology, often relies on single-frame pose comparison, and the key frame positioning accuracy is insufficient, resulting in a large measurement error of the result, which is difficult to meet the requirements of high-precision national standards. These problems affect the accuracy and efficiency of standing long jump evaluation. SUMMARY
[0004] The present application provides a motion evaluation method, device, electronic equipment and storage medium, aiming to solve the problems of subjectivity of manual measurement, complexity of infrared equipment and insufficient key frame positioning accuracy of computer vision scheme in standing long jump evaluation.
[0005] In a first aspect, the present application provides a motion evaluation method applied to standing long jump motion, covering a preparation phase, a take-off phase, a flight phase and a landing phase, the motion evaluation method comprising:
[0006] In the preparation phase, a preset human detection model is called to detect the input image frame to locate the human body frame, and the tester located in the take-off area is determined according to the human body frame;
[0007] In the take-off phase, a preset human pose estimation model and a preset key frame positioning model are called to determine whether the tester has started to take off to determine the take-off frame;
[0008] In the flight phase and the landing phase, based on the take-off frame, the human pose estimation model and the key frame positioning model are called to determine whether the tester has started to land to determine the landing frame, and the landing point of the tester is determined according to the landing frame, and the standing long jump result is calculated.
[0009] In a second aspect, the present application also provides a motion evaluation device applied to standing long jump, covering a preparation stage, a take-off stage, a flight stage and a landing stage, the motion evaluation device comprising:
[0010] A tester positioning module is configured to, in the preparation stage, call a preset human body detection model to detect input image frames to locate a human body frame, and determine a tester located in a take-off area according to the human body frame;
[0011] A take-off detection module is configured to, in the take-off stage, call a preset human body posture estimation model and a preset key frame positioning model to determine whether the tester has started to take off to determine a take-off frame;
[0012] A landing detection module is configured to, in the flight stage and the landing stage, call the human body posture estimation model and the key frame positioning model based on the take-off frame to determine whether the tester has started to land to determine a landing frame;
[0013] A result calculation module is configured to determine a landing position of the tester according to the landing frame to calculate a standing long jump result.
[0014] In a third aspect, the present application provides an electronic device comprising a memory, a processor and a computer program stored in the memory and running on the processor, wherein the processor implements the steps of the motion evaluation method according to any one of the first aspect.
[0015] In a fourth aspect, the present application provides a non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program is executed by a processor to implement the steps of the motion evaluation method according to any one of the first aspect.
[0016] The motion evaluation method, device, electronic device and storage medium provided by the present application aim to solve the problems of subjectivity of manual measurement, complexity of infrared equipment and insufficient accuracy of key frame positioning in standing long jump evaluation. The method can objectively identify and locate the key actions of the tester in the standing long jump process, such as take-off and landing, through the automatic human body detection model and human body posture estimation model. Moreover, the method only relies on image input, such as video frames, without the need for additional high-cost hardware devices, simplifying the deployment and maintenance of the evaluation system. The key frame positioning model is used to accurately identify and locate the key action frames in the standing long jump process, such as the take-off frame and the landing frame. By combining human body posture estimation and key frame positioning technology, the method can more accurately capture the subtle changes in the tester's actions, thereby improving the accuracy of the evaluation.
[0017] Therefore, the application effectively solves the problems in the traditional standing long jump evaluation by means of automation and intelligent technology, and provides a more accurate, objective and convenient evaluation scheme. BRIEF DESCRIPTION OF DRAWINGS
[0018] In order to more clearly illustrate the technical solutions in the application or the prior art, the drawings needed to be used in the embodiments or the prior art description will be briefly introduced. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative labor.
[0019] Figure 1 is one of the flowcharts of the motion evaluation method provided by the application;
[0020] Figure 2 is the second flowchart of the motion evaluation method provided by the application;
[0021] Figure 3 is a schematic diagram of calibrating a motion field provided by the embodiment of the application;
[0022] Figure 4 is a schematic diagram of a first key frame positioning model provided by the embodiment of the application;
[0023] Figure 5 is a schematic diagram of a second key frame positioning model provided by the embodiment of the application;
[0024] Figure 6 is a schematic diagram of state machine determination logic provided by the embodiment of the application;
[0025] Figure 7a is a schematic diagram of a target tracking model provided by the embodiment of the application;
[0026] Figure 7b is a schematic diagram of a human body frame and human body key points provided by the embodiment of the application;
[0027] Figure 8 is a schematic diagram of calculating a standing long jump result provided by the embodiment of the application;
[0028] Figure 9 is a structural schematic diagram of a motion evaluation device provided by the application;
[0029] Figure 10 is a structural schematic diagram of an electronic device provided by the embodiment of the application. DETAILED DESCRIPTION
[0030] In order to make the purposes, technical solutions and advantages of the present application clearer, the technical solutions of the present application will be described clearly and completely below in conjunction with the drawings in the present application. Obviously, the described embodiments are some of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative work belong to the scope of protection of the present application.
[0031] The terms "first", "second", etc. in the specification and claims of the present application and the above drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data thus used can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.
[0032] In order to solve the problems of subjectivity of manual measurement, complexity of infrared equipment and insufficient positioning accuracy of key frames in computer vision scheme in standing long jump evaluation, the present application provides a sports evaluation method, device, electronic equipment and storage medium, which realizes automatic evaluation of the whole process of standing long jump by integrating human body detection, posture estimation and key frame positioning technology. Not only the subjectivity of manual measurement is eliminated, but also the limitation of relying on complex infrared equipment is avoided, and the positioning accuracy of key action frames is improved. Through the comprehensive application of these technologies, the present application provides an objective, accurate and easy-to-implement evaluation solution for standing long jump, thereby optimizing the evaluation process and improving the reliability of the evaluation results.
[0033] The following will be described in conjunction with Figures 1-10 The present application provides a sports evaluation method, device, electronic equipment and storage medium.
[0034] Please refer to Figure 1 , Figure 2 , Figure 1 is a flowchart of the sports evaluation method provided by the present application, Figure 2 is a flowchart of the sports evaluation method provided by the present application. A sports evaluation method is applied to standing long jump, covering preparation stage, take-off stage, flight stage and landing stage, and the sports evaluation method comprises:
[0035] S110, in the preparation stage, a preset human body detection model is called to detect the input image frame to locate the human body frame, and the test personnel located in the take-off area are determined according to the human body frame.
[0036] Specifically, a current frame can be obtained from a video stream or a series of images as input. A preset human pose estimation model is called to analyze the input image frame. The human pose estimation model can be a deep learning-based object detection algorithm, such as YOLO (You Only Look Once), SSD (Single Shot MultiBox Detector), or Faster R-CNN (a deep learning model for object detection), etc., which can identify multiple human bodies in the image and give the position of each human body in the image, i.e., a rectangular frame surrounding the human body, also known as a human body frame or human body bounding box. According to the position information of the human body frame, it is determined whether the human bodies are located within a predefined take-off area. The take-off area is a specific ground area where the test personnel prepare and take off. The purpose of this step is to automatically identify and locate the human bodies in the video or image and further determine whether these human bodies are located within the specified take-off area, providing basic data for the subsequent evaluation of the take-off stage, the flight stage, and the landing stage, which can reduce the need for manual monitoring and marking and improve the efficiency and accuracy of the evaluation.
[0037] S120, in the take-off stage, a preset human pose estimation model and a preset key frame positioning model are called to determine whether the test personnel have started to take off to determine the take-off frame.
[0038] Specifically, a preset human pose estimation model is called to analyze the posture of the test personnel. This model can identify and track the key points of the human body, thereby estimating the posture and action of the test personnel. At the same time, a preset key frame positioning model is called to identify the key frames in the video. Key frames refer to those frames that can represent important action changes, such as the moment when the test personnel take off. Combining the outputs of the human pose estimation model and the key frame positioning model, the system determines whether the test personnel have started to take off, for example, by identifying specific action patterns such as leg and body acceleration, off the ground, etc. Once the system determines that the test personnel have started to take off, the key frame positioning model will further determine the specific frame at which the take-off occurs, i.e., the take-off frame. The take-off frame is used for subsequent analysis, such as calculating the distance of the jump and evaluating the compliance of the action. The purpose of this step is to automatically capture the key moment of the test personnel taking off, providing accurate time points for the subsequent evaluation of the flight stage and the landing stage, improving the degree of automation of the evaluation, and reducing the subjectivity and errors of manual judgment.
[0039] S130, in the flight stage and the landing stage, based on the take-off frame, the human pose estimation model and the key frame positioning model are called to determine whether the test personnel have started to land to determine the landing frame, and to determine the landing point of the test personnel according to the landing frame, and to calculate the standing long jump result.
[0040] Specifically, the system will continue analyzing the tester's movements from the previously determined takeoff frame. During the flight and landing phases, a human pose estimation model is used to track the tester's body pose and the position of key points. This helps the system understand the tester's movements and pose changes in the air. A keyframe localization model is used to identify the important moment that marks the tester's landing, i.e., the landing frame. By analyzing the outputs of the human pose estimation model and the keyframe localization model, the system will determine whether the tester has started to contact the ground. Once the system confirms that the tester has started to land, the keyframe localization model will accurately determine the specific frame in which the landing occurs, i.e., the landing frame. Based on the tester's position information in the landing frame, the system will determine the tester's landing location. For example, identify the position where the tester's feet contact the ground. Finally, the system will calculate the standing long jump score based on the tester's landing location. For example, measure the straight-line distance between the takeoff point and the landing location, and make necessary adjustments according to the competition rules. The purpose of this step is to track the tester's landing movements through automation technology and calculate the tester's score accordingly, improving the accuracy and efficiency of the evaluation and reducing the need for manual intervention.
[0041] In some embodiments, the method further comprises:
[0042] S140, after determining the tester located in the takeoff area, calling a preset target tracking model, and starting the person tracking logic after the tester reaches the preparation state.
[0043] Specifically, in the aforementioned step S110, a human detection model has been used to identify and locate the tester located in the takeoff area. When the tester enters the preparation state, i.e., starts to make the preparation movements before taking off (such as squatting, swinging arms, etc.), the system will start the person tracking logic. This means that the system will start to closely follow the tester's movements in order to accurately evaluate the tester when he / she takes off, flies and lands. The purpose of step S140 is to ensure that the system is ready and able to continuously track the tester's position and movements before the tester starts the key movements of the standing long jump, so as to carry out subsequent evaluation and analysis, improving the continuity and accuracy of the evaluation, and reducing the risk of tracking loss due to the tester's movement.
[0044] The above steps S110 to S140 are described in detail as follows.
[0045] In some embodiments, the sports evaluation method further comprises:
[0046] S10, training a target detection model to calibrate the sports field, the target detection model comprising a first detection model and a second detection model, and the training steps comprising S11 to S14:
[0047] S11, obtain standing long jump images and label each scale point in each image using a preset labeling tool to form dense labeling.
[0048] Specifically, a data set for training the standing long jump automatic calibration needs to be prepared, and each scale point needs to be marked out using a rectangular box in each image using a labeling tool (such as labelme),
[0049] S12, train a preset first detection model to detect the four vertices of the sports field, which are the key points for determining the location of the field.
[0050] Specifically, a preset first detection model (such as a YOLOX model) is used to detect the field in the image, and the purpose is to detect the four vertices corresponding to the sports field.
[0051] S13, according to the four detected vertices, crop the image to extract the sports field.
[0052] Specifically, according to the four detected vertices, the image is cropped to extract the sports field from the original image, which helps to reduce background interference and make subsequent dense labeling detection more accurate.
[0053] S14, train a preset second detection model to perform dense labeling detection on the cropped image to identify all scale points to complete the training of the target detection model.
[0054] Specifically, a preset second detection model (such as another YOLOX model) is used to perform dense labeling detection on the cropped image. The purpose of this step is to identify all scale points on the sports field, thereby obtaining a trained target detection model.
[0055] Therefore, through the above steps S11 to S14, the system can train a target detection model that can accurately detect and calibrate the standing long jump sports field. This model can be used in practical applications to automatically identify and measure the distance of standing long jump.
[0056] In some embodiments, the method further includes an inference stage of the target detection model, which inference steps include S20 to S22:
[0057] S20, according to the detection result of the second detection model, sort the detected scale point boxes according to their positions in the image, and divide the boxes into upper and lower rows of boxes by identifying the distance mutation.
[0058] Specifically, the four vertices of the sports field are detected using the trained first detection model. According to the detected four vertices, the sports field is cropped from the original image. And the cropped sports field is densely annotated and detected using the trained second detection model to identify all the scale points and output a series of rectangular boxes, each box corresponding to a detected scale point.
[0059] The detected rectangular boxes are sorted according to their positions in the image (such as x-axis and y-axis coordinates), which helps to identify and process the layout of the boxes. And in the sorted boxes, find those points whose positions on the x-axis or y-axis change significantly, these mutations mark the boundaries between rows of boxes. According to the distance from the mutation, the boxes are divided into upper row boxes and lower row boxes, each row of boxes corresponds to a row of scale points in the standing long jump area.
[0060] Further, for the upper row boxes and the lower row boxes, the intersection over union (IoU) between the boxes is calculated, if the IoU of two boxes exceeds a preset threshold, it is considered that the two boxes overlap too much, which may be a false detection (false alarm), so one of the boxes is deleted.
[0061] S21, for the upper row boxes and the lower row boxes, the width of each box and the spacing between the boxes are counted, and from left to right, according to the width of the box, the spacing size and the distance between the two boxes, it is judged whether there is a missing box; if there is a missing box, linear interpolation is performed according to the known box information to complete the missing box.
[0062] S22, judge whether the number of completed boxes meets the preset requirement.
[0063] Specifically, after completing all the missing boxes, a final check is performed to ensure that the number of completed boxes meets the set requirement of scale points, for example, please refer to Figure 3 , Figure 3 is a schematic diagram of calibrating a sports field provided by an embodiment of the present application. The set requirement is 61 boxes in each of the upper and lower rows, a total of 122 boxes, to ensure that all scale points are correctly detected and labeled.
[0064] Therefore, through the above steps S20 to S22, the output of the target detection model is optimized to ensure that all scale points are accurately identified and labeled in dense point detection, while reducing false detection and missing detection.
[0065] In some embodiments, the method further comprises training a human detection model, a target tracking model, a human pose estimation model, and a landing point detection model, wherein:
[0066] S30, using a preset target detection framework, using a dataset containing a first number of key points and labeled long jump scene human box data for training, obtaining a trained human detection model.
[0067] For example, the YOLOX target detection framework can be used for training based on the MS COCO (Microsoft Common Objects in Context) public dataset (with 17 key points) and the labeled long jump scene human frame data (with 30 key points), and the trained model can detect the test personnel in the image and give the rectangular frame of the position of the test personnel.
[0068] For example, the 30 key points are: top of head, nose, left ear, right ear, chin, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left palm, right palm, left middle finger, right middle finger, left hip, right hip, left knee, right knee, left ankle, right ankle, left heel, right heel, left toe, right toe, left eye, right eye, left thumb, right thumb.
[0069] S31, using a preset single target tracking model, using a preset human body tracking dataset and a labeled long jump scene motion dataset for training, obtaining a trained target tracking model.
[0070] For example, based on the OS-Track single target tracking model, the public human body tracking dataset LaSOT (Large-scale Single Object Tracking) and GOT-10k (Generic Object Tracking Benchmark) are used for training, and the collected and labeled long jump scene motion dataset is used for training, and the trained model can continuously track the position of the test personnel in the video sequence.
[0071] S32, based on a preset single person pose estimation framework, using a dataset containing a first number of key points and a labeled dataset containing a second number of key points for mixed training, obtaining a trained human body pose estimation model.
[0072] For example, based on the SimDR (Similarity-based Domain Randomization) single person pose estimation framework of one-dimensional heat map paradigm, the MS COCO public dataset (containing 17 key points) and the labeled long jump scene human body skeleton point data (containing 30 key points) are used for mixed training. The trained model can output 30 key points of human body pose estimation results, which helps to analyze the action of the test personnel in more detail.
[0073] S33, using a preset key point detection framework, training using the labeled heel key point data of the long jump scene to obtain a trained key point detection model.
[0074] For example, based on the RLEPose (a method of representing or estimating the pose of an object, especially a human body, using RLE encoding technology) key point detection framework, the experimentally labeled heel key point data of the long jump scene can be used for training. The trained key point detection model can detect the position of the heel when the test person lands.
[0075] Therefore, in the above steps S30 to S33, a series of models are trained by using different data sets and algorithm frameworks, which can detect, track and analyze the actions of the standing long jump test person, thereby realizing automated motion evaluation. Each model is optimized for a specific task to improve the accuracy and efficiency of the evaluation.
[0076] In the standing long jump, the landing frame refers to the video frame at the moment when the test person's feet touch the ground. Accurate capture of this video frame is crucial for calculating the test person's long jump score, as the score is calculated based on the distance between the take-off point and the landing point. If the image resolution of the camera is 1080 pixels vertically and 1920 pixels horizontally, i.e. 1080x1920 resolution, and the frame rate is 25 FPS (Frames Per Second), i.e. 25 frames of images are taken per second. According to actual statistics, if there is a deviation of one frame in the positioning of the landing frame, it may cause an error of about 3 centimeters in the calculation of the score. This is because in the case of 25 frames taken in one second, each frame represents a time interval of about 0.04 seconds, and the test person's time in the air is usually very short, so the deviation of each frame can lead to a large distance error. Therefore, in order to find a stable and reliable key frame, the present application proposes an improved key frame positioning model, which is different from the existing scheme that only relies on the motion speed of the human body skeletal points to determine the take-off and landing. It uses a multi-level cascade and multi-modal method, divided into two stages: coarse positioning stage (using the first key frame positioning model) and fine positioning stage (using the second key frame positioning model), as follows:
[0077] The goal of the coarse positioning stage is high recall, i.e. to capture as many possible key frames as possible. In order to speed up the inference, the first key frame positioning model uses a human key point sequence (i.e. a skeletal point sequence) as input. The key point sequence is obtained by analyzing the data of the position of the human body joints over time in the video. During the evaluation process, the coarse positioning stage will be used multiple times to ensure that no key frame is missed.
[0078] The goal of the fine localization stage is high accuracy, i.e., accurately localizing to the position of the key frame. The pure skeleton point input can be affected by the quality of skeleton point extraction and is not robust enough in some scenarios (such as shadow, uneven lighting, etc.). Therefore, the second key frame localization model adopts multi-modal input of key point sequence and picture sequence, combining the motion information of skeleton points and the visual information of images to improve the robustness and accuracy of the model.
[0079] The first key frame localization model and the second key frame localization model are described below.
[0080] In some embodiments, the method further includes an inference stage of the key frame localization model, the key frame localization model including the first key frame localization model and the second key frame localization model, and the inference steps thereof including S40 and S41:
[0081] S40, the first key frame localization model is used to identify whether to contain a key frame; it uses the normalized key point sequence as input, and predicts the action category through pose encoding, time sequence fusion and action classification.
[0082] Specifically, key poses such as jumping and landing are considered as micro-actions because they have very short duration in the video, usually only 2-3 frames. For example, the landing action from the tip of the foot landing to the sole flattening may only occupy about 3 frames of time in a 25 FPS (25 frames per second) video. The first key frame localization model abstracts the key frame localization problem into a micro-action recognition problem, i.e., determining what the micro-action category corresponding to the input pose estimation sequence is. The input of the first key frame localization model is the normalized key point sequence, which represents the poses of the human body in different frames of the video.
[0083] Please refer to Figure 4 , Figure 4is a schematic diagram of a first keyframe positioning model provided by an embodiment of the present application. The first keyframe positioning model comprises a pose encoding module, a temporal fusion network, and an action classification head. First, the pose encoding module encodes the key points of each frame into a high-dimensional vector (referred to as token), i.e., a key point vector, which helps to capture the spatial features of the key points. The pose encoding module is a multi-layer perceptron. Second, the input high-dimensional vector sequence (tokens) of each frame, together with a learnable class vector token, is input into a temporal fusion network. The temporal fusion network is a multi-layer Transformer (based on the self-attention mechanism) encoder. The temporal fusion network is responsible for fusing the temporal information, i.e., the dynamic changes between different frames, and outputting the class vector token. Finally, the class vector token is input into an action classification head. The action classification head is also a multi-layer perceptron, and the action classes can include jumping, landing, and others. The action classification head is responsible for predicting the final action class, i.e., the action class of the key pose. For example, the action classification head can determine whether the input sequence contains a jump, a landing, or other micro-motions.
[0084] Exemplarily, the key points of each frame are normalized as follows:
[0085] x min =min(x 1,1 , x 1,2 , …, x 1,n , …, x 1,30 );
[0086] y min =min(y 1,1 , y 1,2 , …, y 1,n , …, y 1,30 );
[0087]
[0088]
[0089] wherein x 1,n and y 1,n represent the horizontal and vertical coordinates of the nth key point of the first frame, respectively; and represent the horizontal and vertical coordinate sets of all key points of the tth frame, respectively; W and H represent the horizontal and vertical resolutions of the original image, respectively; and represent the horizontal and vertical coordinate sets of all key points of the tth frame after normalization.
[0090] S41, the second keyframe positioning model is used for positioning the position of the keyframe; it uses multimodal input of the key point sequence and the picture sequence, and predicts the position of the keyframe through pose encoding, picture encoding, modal fusion, time sequence fusion and keyframe multi-classification.
[0091] Specifically, the first keyframe positioning model can roughly judge whether the current video segment contains a keyframe (such as a take-off or landing frame), but cannot accurately locate the specific position of the keyframe. Since only the key point segment is relied on as input, when the key point quality is poor (for example, in the case of insufficient light or occlusion), the model is difficult to accurately locate the position of the landing frame.
[0092] In order to improve the accuracy of positioning, more information is needed in the fine positioning stage, including the field information and the detailed description of the human body action. Therefore, in addition to the key point information, the input of the picture modality is also needed. However, relying only on the picture sequence as input has challenges, because there may be a lack of effective feature alignment between pictures, which makes it difficult to quickly locate the motion information. Therefore, the key point information is used as an aid to help the model better understand the motion pattern.
[0093] The keyframe positioning of the fine positioning stage is abstracted into a multi-classification problem. Specifically, the second keyframe positioning model receives a time sequence information containing T frames (including the key point sequence and the picture sequence), and outputs the position t of the frame where the keyframe is located, where t is an integer between 0 and T-1. The goal of this multi-classification problem is to determine the specific position of the keyframe in the input sequence, so as to improve the accuracy of the evaluation.
[0094] Please refer to Figure 5 , Figure 5 is a schematic diagram of the second keyframe positioning model provided by the embodiment of the application. The second keyframe positioning model includes a pose encoding module, a picture encoding module, a modal fusion module, a time sequence fusion network and a keyframe classification head. First, the pose encoding module is used to encode the key point sequence into a high-dimensional vector, i.e. a key point vector, denoted as skeleton token, representing the skeleton point modality. At the same time, the model uses the picture encoding module to encode the picture sequence into a high-dimensional vector, i.e. a picture vector, denoted as image token, representing the picture modality.
[0095] Next, the model employs a modal fusion module to cross-modally fuse the skeleton token and the image token. In order to better align the different modal information of each frame input during the fusion process, the modal fusion module performs a concatenation operation, i.e., multi-modal token = concat(skeleton token, image token). This means that the skeleton point information and the image information are concatenated into a single high-dimensional vector, i.e., the multi-modal vector, denoted as multi-modal token. The multi-modal vector contains information of the two modalities.
[0096] Then, the multi-modal vector multi-modal token and a learnable class vector token are input into a temporal fusion network for temporal feature fusion. In this process, the class vector token fully fuses the multi-modal information and the temporal information through an attention mechanism to capture the features of the key frame, and outputs the class vector token.
[0097] Finally, the class vector token is input into a key frame multi-classification head to output the final position of the key frame. In this application, the number of multi-classification classes is equal to the input temporal length, representing the specific position of the key frame in the input sequence. For example, if the input sequence has T frames, then the second key frame positioning model will predict the specific position of the key frame in the 0 to T-1 frames.
[0098] As can be seen, the first key frame positioning model abstracts the key frame positioning problem into a classification problem, i.e., determining whether the current video clip contains a key frame (such as a take-off or landing frame). By simplifying the problem into a classification problem, the model reduces the difficulty of learning, thereby being able to obtain a classification model with high recall rate (i.e., detecting as many key frames as possible) and high accuracy. The first key frame positioning model is used for preliminary screening to avoid missing key frames.
[0099] The second key frame positioning model, on the other hand, pays more attention to the subtle differences between frames, especially those features that can indicate a key frame (such as a landing frame). By utilizing the multi-modal information of the skeleton points and the images, the model can more accurately identify and locate the landing frame. The second key frame positioning model aims to provide more accurate key frame positions, thereby providing a more reliable and accurate basis for subsequent performance calculation. The two models work together to ensure the accuracy and reliability of the entire evaluation process.
[0100] In some embodiments, the motion evaluation method uses a state machine to logically determine the preparation phase, the take-off phase, the flight phase, and the landing phase, and the logical determination steps thereof include:
[0101] S50, in the preparation stage, call the human detection model every certain time interval, identify the human body in the video frame sequence, and screen out the human body rectangular frame located in the take-off area; if only one test person is detected in the take-off area for several times in succession, it is determined that the test person enters the preparation state.
[0102] Please refer to Figure 6 , Figure 6 is a schematic diagram of the state machine determination logic provided by the embodiment of the application. The system is in the waiting-to-stand state when starting, that is, waiting for the test person to be ready for the long jump. For example, the system calls the human detection model every detection interval of 25 frames at a certain time interval, which is used to identify the human body in the video. According to the output result of the human detection model, the system screens out the rectangular frame of the person in the take-off area, that is, determines which people are located in the take-off area. When the detection result shows that there is only one person in the take-off area for three times in succession, the system determines to enter the preparation state, that is, the test person has prepared to take off.
[0103] S51, in the take-off stage, when the test person is ready to take off, a first key frame positioning model is used to identify whether the take-off action occurs; if the take-off action is identified, a second key frame positioning model is switched to for positioning the take-off frame.
[0104] Specifically, the determination of the take-off state first uses the first key frame positioning model in the coarse positioning stage. When the classification result of the first key frame positioning model is the take-off category, the system switches to the second key frame positioning model in the fine positioning stage for accurate positioning of the take-off frame. The take-off frame is used to judge whether the test person violates the rules when taking off, for example, the positioning of the take-off frame is used to determine whether the line stepping phenomenon occurs in the evaluation process.
[0105] S53, in the take-off stage and the landing stage, a first key frame positioning model is used to identify whether the landing action occurs; if the landing action is identified, a second key frame positioning model is switched to for positioning the landing frame.
[0106] Specifically, the determination rule of the landing state is similar to that of the take-off state, and both are determined by using the cascade model. The first key frame positioning model in the coarse positioning stage is used for preliminary determination of the landing state. The second key frame positioning model in the fine positioning stage is used for accurate positioning of the landing frame. The landing frame is used to judge whether the test person violates the rules when landing.
[0107] In some embodiments, in step S140, the steps of the personnel tracking logic include:
[0108] S141, capturing an initial image of the test person when the test person is located in the take-off area, and setting the initial image as a tracking template; continuously obtaining a current frame from the video stream as a test area, and applying a template matching algorithm in the current frame to search for the location of the test person based on the tracking template; and locating the position of the test person in the current frame according to the output result of the template matching algorithm, and outputting a human body frame.
[0109] Specifically, personnel tracking is crucial to ensure the accuracy of the evaluation results. If tracking is wrong, it may lead to wrong landing frame and landing site detection, and thus calculate the wrong evaluation results. After the system determines that the test person reaches the preparation state, the personnel tracking logic will be started to ensure accurate tracking of the test person during the entire long jump process.
[0110] Traditional tracking solutions are usually based on multi-person pose estimation, which requires pose estimation for all persons in the picture every frame, and then tracks the test person through post-processing such as field matching and frame matching. The traditional solution is prone to mis-tracking when there are a large number of spectators, because a large amount of data needs to be processed and complex matching needs to be performed.
[0111] Please refer to Figure 7a , Figure 7b , Figure 7a is a schematic diagram of a target tracking model provided by an embodiment of the present application, Figure 7b is a schematic diagram of a human body frame and human body key points. The target tracking model provided by the present application adopts a single target tracking model to improve the reliability of tracking. The single target tracking model can simplify the tracking process, because it only needs to focus on a single target, rather than all persons in the picture, such as the rectangular frame shown in Figure 7b Specifically, as described in Figure 7a , when the test person is located in the take-off area, the system captures an initial image of the test person, and a target tracking model uses a template as a tracking module, which will be used as a reference for searching for the position of the test person in subsequent frames. The system continuously obtains a current frame from the video stream, which will be used as a test area for searching for the position of the test person in each frame. A template matching algorithm is applied in the current frame to search for the position of the test person based on the tracking template. The template matching algorithm compares the image in the current frame with the tracking template to determine the position of the test person in the current frame. According to the output result of the template matching algorithm, the system locates the position of the test person in the current frame and outputs a human body frame (bounding box), i.e., a rectangular frame surrounding the test person.
[0112] As can be seen, the single target tracking model ensures the accuracy of tracking the test person, and simplifies the complexity of the entire system, improving the running efficiency of the system.
[0113] In some embodiments, in step S130, the landing position of the test person is determined according to the landing frame, and the step of calculating the standing long jump result comprises:
[0114] In step S131, all the calibration points on the sports field are obtained.
[0115] For example, based on the field calibration algorithm, the system can obtain a calibration distribution with an interval of 5 cm in the range of 0-300 cm. These calibrations are used to determine the specific position of the landing position of the test person on the field.
[0116] In step S132, in the landing frame, whether the test person's foot contour point is located in a certain calibration interval is determined to determine the initial result interval.
[0117] Specifically, in the landing frame, the system uses the test person's foot contour point to determine whether it is located in a certain large result interval. This is the initial result interval determination, which is used to determine the approximate result range.
[0118] In step S133, according to the distance of the foot contour point from the left and right calibration lines, the result in the small interval is calculated according to the equal proportion division strategy, and the final result is determined by adding the large interval and the small interval results.
[0119] That is, the result is calculated according to the relative position of the landing position between the calibration lines.
[0120] The following illustrates how to calculate the standing long jump result through an example.
[0121] Suppose we have a standing long jump field, and the calibration line is from 0 cm to 300 cm, with a calibration every 5 cm. After the test person jumps, the system needs to calculate his final result.
[0122] Suppose the landing position coordinate of the test person is 123 cm.
[0123] In the landing frame, the system determines the initial result interval according to whether the test person's foot contour point is located in a certain calibration interval. For example, the landing position 123 cm is located between the calibration line 120 cm (x1) and 125 cm (x2).
[0124] According to the distance of the foot contour point from the left and right calibration lines, the result in the small interval is calculated according to the equal proportion division strategy. For example, the result calculation formula is: 120+(123-120) / (125-120)*5=120+3 / 5*5=120+3=123 cm.
[0125] The final score is determined by accumulating the large-interval score and the small-interval score. The large interval refers to the range between two scale lines where the landing point coordinates are located.
[0126] In this example, the large interval is the range between 120 cm (x1) and 125 cm (x2). The system first determines the initial score interval according to whether the landing point coordinates of the tester are located in this large interval. Specifically, when the landing point coordinates of the tester are 123 cm, the coordinates are located between the two scale lines of 120 cm and 125 cm, so the large interval is 120 cm to 125 cm. In this large interval, the system calculates the score in the small interval according to the distance of the tester's foot contour point from the left and right scale lines, according to the equal proportion division strategy. Finally, the large interval score is 120 cm, the small interval score is 3 cm, and the final score is determined by accumulating the large interval score and the small interval score, which is 123 cm.
[0127] Therefore, the above steps S131 to S133 provide the overall method of score calculation, that is, the final score is determined by accumulating the large interval score and the small interval score. However, the specific calculation method, that is, how to calculate the score of each landing point according to the coordinates of the landing point and the position of the scale line, includes:
[0128] S134, determine which two scale lines on the field calibration scale the landing point coordinates of the left and right feet of the tester fall between, and check whether the landing point exceeds the test area.
[0129] Specifically, please refer to Figure 8 , Figure 8 is a schematic diagram for calculating the standing long jump score provided by the embodiment of the present application. The system uses human posture estimation and landing point detection model to locate the landing point, i.e., the tester's heel key point, according to the landing state determination strategy. For example, the system determines which two scale lines [x1, x2] the left and right foot landing point coordinates fall between, and checks whether it exceeds the test area.
[0130] S135, for each landing point, apply a preset score calculation formula to determine the score of the landing point, and calculate the scores corresponding to all landing points, from which the smallest score is selected as the final standing long jump score of the current evaluation.
[0131] Exemplarily, the score calculation formula for each landing point is:
[0132] Scale line x1+(x-x1) / (x2-x1)*s.
[0133] where x1 represents the starting scale line of the interval where the landing point is located, x represents the coordinate of the actual landing point of the tester, x2-x1 represents the length of the interval where the landing point is located, and s represents the length of each scale interval (for example, 5 cm). The meaning of the calculation formula is to start from the starting scale x1 of the interval, and add the offset of the landing point x relative to the starting scale x1, which is calculated according to the relative position of the landing point in the interval, and the length of each scale interval s is multiplied by the proportion.
[0134] Finally, the system calculates the scores corresponding to all possible landing points, and selects the smallest score from these scores as the final standing long jump score of the current evaluation.
[0135] The following illustrates how to calculate the standing long jump score through an example.
[0136] For example, in a standing long jump field, the scale line is from 0 cm to 300 cm, and there is a scale every 5 cm. After the tester jumps, the system needs to calculate his final score.
[0137] Suppose the coordinate of the landing point of the tester's left foot is 123 cm, and the coordinate of the landing point of the tester's right foot is 127 cm. The system uses the human pose estimation model and the key frame positioning model to locate the two landing points.
[0138] For the left foot landing point (123 cm):
[0139] It falls between the scale lines 120 cm (x1) and 125 cm (x2).
[0140] The score calculation formula is: 120+(123-120) / (125-120)*5=120+3 / 5*5=120+3=123 cm.
[0141] For the right foot landing point (127 cm):
[0142] It falls between the scale lines 125 cm (x1) and 130 cm (x2).
[0143] The score calculation formula is: 125+(127-125) / (130-125)*5=125+2 / 5*5=125+2=127 cm.
[0144] The system takes the minimum value of the two landing point scores as the final score, which is 123 cm.
[0145] In summary, the key frame positioning model provided in the application improves the key frame recognition accuracy of key actions (such as take-off and landing) through a multi-stage cascading manner. In the coarse positioning stage, the key frame is quickly recognized using the human key point sequence, and the purpose is to have a high recall rate, that is, as many possible key frames are recognized as possible. In the fine positioning stage, the multi-modal input of the key point sequence and the picture sequence is combined to further improve the accuracy and robustness of key frame positioning, that is, it can work stably under various conditions. In addition, in the fine positioning stage, the modal fusion module is used to fuse the skeleton points and picture information to overcome the limitations of single modal input in certain conditions, such as shadows, uneven lighting, etc. In addition, a single target tracking model is applied to improve the reliability of the system and avoid system abnormal termination or score errors caused by irrelevant personnel interference.
[0146] The motion evaluation device provided in the application is described below, and the motion evaluation device and the motion evaluation device described below can be mutually corresponding with reference to the motion evaluation method described above.
[0147] Please refer to Figure 9 , Figure 9 is a structural schematic diagram of the motion evaluation device provided in the application. A motion evaluation device 900 includes a test personnel positioning module 910, a take-off detection module 920, a landing detection module 930, and a score calculation module 940.
[0148] Illustratively, the test personnel positioning module 910 is used to call a preset human shape detection model to detect the input image frame in the preparation stage to locate the human body frame, and determine the test personnel located in the take-off area according to the human body frame.
[0149] Illustratively, the take-off detection module 920 is used to call a preset human pose estimation model and a preset key frame positioning model in the take-off stage to determine whether the test personnel has started to take off to determine the take-off frame.
[0150] Illustratively, the landing detection module 930 is used to call the human pose estimation model and the key frame positioning model based on the take-off frame in the take-off stage and the landing stage to determine whether the test personnel has started to land to determine the landing frame.
[0151] Illustratively, the score calculation module 940 is used to determine the landing point of the test personnel according to the landing frame to calculate the standing long jump score.
[0152] Illustratively, the motion evaluation device 900 further includes a target tracking module, which is used to call a preset target tracking model after determining the test personnel located in the take-off area to start the personnel tracking logic after the test personnel reaches the preparation state.
[0153] Exemplarily, the motion evaluation device 900 is further configured to:
[0154] acquire standing long jump images, and label each scale point in each image using a preset labeling tool to form dense labeling;
[0155] train a preset first detection model to detect four vertices of a sports field, and the vertices are key points for determining the position of the field;
[0156] crop images according to the detected four vertices to extract the sports field;
[0157] train a preset second detection model to perform dense labeling detection on the cropped images to identify all scale points, so as to complete the training of the target detection model.
[0158] Exemplarily, the motion evaluation device 900 is further configured to:
[0159] According to the detection result of the second detection model, the detected scale point boxes are sorted according to their positions in the image, and the boxes are divided into two rows of boxes by identifying the distance mutation in the sorted boxes;
[0160] For the upper row of boxes and the lower row of boxes, the width of each box and the interval size between the boxes are counted, and the boxes are traversed from left to right, and according to the width of the boxes, the interval size and the interval between the two boxes, it is judged whether there is a missed box; if there is a missed box, linear interpolation is performed according to the known box information to complete the missing box;
[0161] determine whether the number of completed boxes meets a preset requirement.
[0162] Exemplarily, the motion evaluation device 900 is further configured to:
[0163] using a preset target detection framework, using a data set containing a first number of key points and a labeled long jump scene human body box data to train a trained human body detection model;
[0164] using a preset single target tracking model, using a preset human body tracking data set and a labeled long jump scene motion data set to train a trained target tracking model;
[0165] based on a preset single-person pose estimation framework, using a data set containing a first number of key points and a labeled data set containing a second number of key points to perform mixed training, to obtain a trained human body pose estimation model;
[0166] using a preset key point detection framework, using labeled long jump scene heel key point data to train a trained key point detection model.
[0167] Exemplarily, the motion evaluation device 900 is further configured to:
[0168] The first key frame positioning model is configured to identify whether a key frame is contained; it uses the normalized key point sequence as input, and predicts the action category through pose encoding, temporal fusion and action classification;
[0169] The second key frame positioning model is configured to locate the position of the key frame; it uses the multi-modal input of the key point sequence and the picture sequence, and predicts the position of the key frame through pose encoding, picture encoding, modal fusion, temporal fusion and key frame multi-classification.
[0170] Exemplarily, the first key frame positioning model comprises a pose encoding module, a temporal fusion network and an action classification head; the pose encoding module encodes the key point sequence into a key point vector, and inputs the high-dimensional vector and a learnable category vector into the temporal fusion network to fuse the temporal information and input the output category vector into the action classification head to predict the final action category.
[0171] Exemplarily, the second key frame positioning model comprises a pose encoding module, a picture encoding module, a modal fusion module, a temporal fusion network and a key frame classification head; the pose encoding module encodes the key point sequence into a key point vector, and the picture encoding module encodes the picture sequence into a picture vector; the modal fusion module splices the key point vector and the picture vector into a multi-modal vector that fuses multi-modal information, and inputs it and a learnable category vector into the temporal fusion network to fuse the multi-modal information and the temporal information to output the category vector and input it into the key frame multi-classification head to predict the position of the key frame.
[0172] Exemplarily, the motion evaluation device 900 is further configured to:
[0173] In the preparation stage, the human body detection model is called every certain time interval to identify the human body in the video frame sequence, and the human body rectangular frame located in the take-off area is screened out; if the take-off area is detected to have only one test person for multiple times in succession, it is determined that the test person enters the preparation state;
[0174] In the take-off stage, when the test person is ready to take off, the first key frame positioning model is used to identify whether the take-off action occurs; if the take-off action is identified, the second key frame positioning model is switched to locate the take-off frame;
[0175] In the take-off stage, when the test person is ready to take off, the first key frame positioning model is used to identify whether the take-off action occurs; if the take-off action is identified, the second key frame positioning model is switched to locate the take-off frame;
[0176] The take-off frame is used to determine whether the tester violates the rules when taking off, and the landing frame is used to determine whether the tester violates the rules when landing.
[0177] The target tracking module is further configured to:
[0178] When the tester is in the take-off area, an initial image of the tester is captured and set as a tracking template.
[0179] A current frame is continuously obtained from the video stream as a test area, and a template matching algorithm is applied in the current frame to search for the position based on the tracking template.
[0180] According to the output result of the template matching algorithm, the position of the tester in the current frame is located, and a human body frame is output.
[0181] The score calculation module 940 is further configured to:
[0182] All scale points on the sports field are obtained.
[0183] In the landing frame picture, whether the tester's foot contour point is located in a certain scale interval is determined to determine the initial score interval.
[0184] According to the distance between the foot contour point and the left and right scale lines, the scores in the small intervals are calculated according to the equal proportion division strategy.
[0185] The final score is determined by adding the scores of the large intervals and the small intervals.
[0186] The score calculation module 940 is further configured to:
[0187] The landing point coordinates of the tester's left and right feet are determined to be between which two scale lines on the field scale, and whether the landing point exceeds the test area is checked.
[0188] For each landing point, a preset score calculation formula is applied to determine the score of the landing point, and the scores corresponding to all landing points are calculated, and the smallest score is selected as the final standing long jump score of the current evaluation.
[0189] It should be noted that the above sports evaluation device provided by the embodiments of the present application can realize all the method steps realized by the above method embodiments, and can achieve the same technical effects. Therefore, the same parts and beneficial effects of the method embodiments in the embodiments will not be described in detail.
[0190] Figure 10 is a structural schematic diagram of an electronic device provided by the embodiments of the present application, such as Figure 10As shown, the electronic device can include a processor 1010, a communications interface 1020, a memory 1030, and a communications bus 1040, wherein the processor 1010, the communications interface 1020, and the memory 1030 complete mutual communication through the communications bus 1040. The processor 1010 can invoke a logical instruction in the memory 1030 to execute the motion evaluation method.
[0191] In addition, the logical instruction in the memory 1030 described above can be implemented in the form of a software function unit and sold or used as an independent product, and can be stored in a computer-readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or part of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a plurality of instructions to ensure that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of the methods described in the various embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
[0192] On the other hand, the present application also provides a computer program product, which includes a computer program stored on a non-transitory computer-readable storage medium, and the computer program includes program instructions, and when the program instructions are executed by a computer, the computer can execute the motion evaluation method provided by the above-mentioned methods.
[0193] In yet another aspect, the present application also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the motion evaluation method provided by the above-mentioned methods.
[0194] The electronic device, the computer program product, and the processor-readable storage medium provided by the embodiments of the present application have computer programs stored thereon, which enable the processor to implement all the method steps of the above-mentioned method embodiments and achieve the same technical effects. Therefore, the same parts and beneficial effects of the embodiments of the present application as the method embodiments will not be described in detail here.
[0195] The device embodiments described above are merely illustrative, wherein the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed to multiple network units. Part or all of the modules can be selected to achieve the purposes of the embodiments according to actual needs. Those skilled in the art can understand and implement without creative labor.
[0196] Through the description of the above embodiments, those skilled in the art can clearly understand that the embodiments can be realized by means of software and necessary universal hardware platforms, and of course can also be realized by hardware. Based on such understanding, the above technical solutions can be embodied in the form of software products, and the computer software products can be stored in a computer readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and include a plurality of instructions to ensure that a computer device (which can be a personal computer, a server, or a network device, etc.) executes the methods described in each embodiment or some parts of the embodiments.
[0197] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method of motion assessment, characterized by, The motion evaluation method is applied to standing long jump, covering a preparation stage, a take-off stage, a flight stage and a landing stage, and comprises the following steps: In the preparation stage, a preset human body detection model is called to detect input image frames to locate a human body frame, and a testee located in a take-off area is determined according to the human body frame; In the take-off stage, a preset human body posture estimation model and a preset key frame positioning model are called to determine whether the testee has started to take off, so as to determine a take-off frame; In the flight stage and the landing stage, based on the take-off frame, the human body posture estimation model and the key frame positioning model are called to determine whether the testee has started to land, so as to determine a landing frame, and a landing position of the testee is determined according to the landing frame, and a standing long jump result is calculated; The method further comprises training a target detection model to calibrate a motion field, the target detection model comprising a first detection model and a second detection model, and the training steps comprising: Obtaining standing long jump images, and using a preset annotation tool to annotate each scale point in each image to form dense annotation; Training a preset first detection model to detect four vertices of the motion field, the vertices being key points for determining the position of the field; According to the four detected vertices, the image is cropped to extract the motion field; Training a preset second detection model to detect the cropped image to identify all scale points, so as to complete the training of the target detection model.
2. The motion assessment method of claim 1, wherein, The motion evaluation method further comprises: After determining the testee located in the take-off area, a preset target tracking model is called to start the personnel tracking logic after the testee reaches the preparation state.
3. The motion assessment method of claim 1, wherein, The method further comprises an inference stage of the target detection model, and the inference steps comprising: According to the detection result of the second detection model, the detected scale point frame is sorted according to its position in the image, and the frame is divided into two rows of frames by identifying the distance mutation; For the upper row of frames and the lower row of frames, the width of each frame and the interval size between the frames are counted, and the frames are traversed from left to right, and whether there is a missed frame is determined according to the width of the frame, the interval size and the distance between the two frames; if there is a missed frame, linear interpolation is performed according to the known frame information to complete the missing frame; Determine whether the number of completed frames meets the preset requirement.
4. The motion assessment method of claim 1, wherein, The method further comprises training a human body detection model, a target tracking model, a human body posture estimation model and a landing position detection model, wherein: A preset target detection framework is used to train a human body detection model by using a data set containing a first number of key points and a human body frame data labeled in a long jump scene; A preset single target tracking model is used to train a target tracking model by using a preset human body tracking data set and a labeled long jump scene motion data set; Based on a preset single human posture estimation framework, a human body posture estimation model is obtained by mixed training using a data set containing a first number of key points and a data set labeled containing a second number of key points. The preset key point detection framework is used to train a labeled heel key point data of a long jump scene to obtain a trained key point detection model.
5. The motion assessment method of claim 1, wherein, The method further comprises an inference stage of a key frame positioning model, the key frame positioning model comprising a first key frame positioning model and a second key frame positioning model, and the inference steps of the key frame positioning model comprising: The first key frame positioning model is used to identify whether a key frame is contained; the first key frame positioning model uses the normalized key point sequence as an input, and predicts an action category through pose encoding, temporal fusion and action classification; The second key frame positioning model is used to locate a position of the key frame; the second key frame positioning model uses a multi-modal input of the key point sequence and the picture sequence, and predicts the position of the key frame through pose encoding, picture encoding, modal fusion, temporal fusion and key frame multi-classification.
6. The motion assessment method of claim 5, wherein, The first key frame positioning model comprises a pose encoding module, a temporal fusion network and an action classification head; the pose encoding module encodes the key point sequence into a key point vector, and inputs the key point vector and a learnable category vector into the temporal fusion network to fuse the temporal information and input a category vector output by the temporal fusion network into the action classification head to predict a final action category.
7. The motion assessment method of claim 5, wherein, The second key frame positioning model comprises a pose encoding module, a picture encoding module, a modal fusion module, a temporal fusion network and a key frame classification head; The pose encoding module encodes the key point sequence into a key point vector, and the picture encoding module encodes the picture sequence into a picture vector; the modal fusion module splices the key point vector and the picture vector into a multi-modal vector that fuses multi-modal information, and inputs the multi-modal vector and a learnable category vector into the temporal fusion network to fuse the multi-modal information and the temporal information, and then outputs a category vector and inputs the category vector into the key frame multi-classification head to predict the position of the key frame.
8. The motion assessment method of claim 1, wherein, The motion evaluation method uses a state machine to logically determine the preparation stage, the take-off stage, the flight stage and the landing stage, and the logical determination steps of the state machine comprising: In the preparation stage, a human body detection model is called every certain time interval to identify a human body in a video frame sequence, and a human body rectangular frame located in a take-off area is screened out; if the take-off area is detected to have only one test person for multiple times in succession, it is determined that the test person enters a preparation state; In the take-off stage, when the test person is ready to take off, a first key frame positioning model is used to identify whether a take-off action occurs; if the take-off action is identified, a second key frame positioning model is switched to to locate a take-off frame; In the flight stage and the landing stage, the first key frame positioning model is used to identify whether a landing action occurs; if the landing action is identified, the second key frame positioning model is switched to to locate a landing frame; The take-off frame is used to determine whether the test person violates a rule when taking off, and the landing frame is used to determine whether the test person violates the rule when landing.
9. The motion assessment method of claim 2, wherein, The personnel tracking logic steps comprising: When the test person is located in the take-off area, an initial image thereof is captured and set as a tracking template; A current frame is continuously obtained from a video stream as a test area, and a template matching algorithm is applied in the current frame to perform a position search based on the tracking template; According to the output result of the template matching algorithm, the position of the test person in the current frame is located, and a human body frame is output.
10. The motion assessment method of claim 1, wherein, The step of determining the landing position of the test person according to the landing frame and calculating the standing long jump score comprises: Obtaining all scale points on the sports field; In the landing frame picture, according to whether the test person's foot contour point is located in a certain scale interval, the initial score interval is determined; According to the distance of the foot contour point from the left and right scale lines, the score in the small interval is calculated according to the equal proportion division strategy; The final score is determined by adding the scores of the large and small intervals.
11. The motion assessment method of claim 1, wherein, The step of determining the landing position of the test person according to the landing frame and calculating the standing long jump score further comprises: Determine which two scale lines on the field calibration scale the left and right foot landing position coordinates of the test person fall between, and check whether the landing position exceeds the test area; For each landing position, apply a preset score calculation formula to determine the score of the landing position, and calculate the scores corresponding to all landing positions, and select the smallest score as the final standing long jump score of the current evaluation.
12. A motion evaluation apparatus characterized by comprising: Applied to standing long jump movement, covering preparation stage, take-off stage, flight stage and landing stage, the movement evaluation device comprises: A test person positioning module is used to locate the human body frame in the preparation stage by calling a preset human detection model to detect the input image frame, and to determine the test person located in the take-off area according to the human body frame; A take-off detection module is used to determine the take-off frame by calling a preset human pose estimation model and a preset key frame positioning model to determine whether the test person has started to take off in the take-off stage; A landing detection module is used to determine whether the test person has started to land by calling the human pose estimation model and the key frame positioning model based on the take-off frame in the flight stage and the landing stage, to determine the landing frame; A score calculation module is used to determine the landing position of the test person according to the landing frame, and to calculate the standing long jump score. The movement evaluation device is also used to: Obtain standing long jump images and use a preset annotation tool to annotate each scale point in each image to form dense annotation; Train a preset first detection model to detect the four vertices of the sports field, which are the key points for determining the position of the field; According to the four detected vertices, the image is cropped to extract the sports field; Train a preset second detection model to perform dense annotation detection on the cropped image to identify all scale points to complete the training of the target detection model.
13. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The processor executes the computer program to realize the steps of the movement evaluation method according to any one of claims 1 to 11.
14. A non-transitory computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program is executed by the processor to realize the steps of the movement evaluation method according to any one of claims 1 to 11.
Citation Information
Patent Citations
Long jump evaluation method and device, electronic equipment and storage medium
CN115546688A
Long jump measuring method based on key points of human skeleton
CN118038549A