A standard method and scoring device for judging basic human motor skills

By refining motion standards and using a lightweight OpenPose network, combined with target object recognition, the problems of scale inconsistency, spatiotemporal dimension differences, and occlusion in the assessment of basic human motor skills are solved, achieving efficient and accurate assessment on embedded devices.

CN116246347BActive Publication Date: 2026-05-26HUAZHONG UNIV OF SCI & TECH

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HUAZHONG UNIV OF SCI & TECH
Filing Date
2023-03-09
Publication Date
2026-05-26

AI Technical Summary

Technical Problem

Existing technologies for assessing basic human motor skills suffer from inconsistent scales, neglect of spatiotemporal differences in movements, occlusion issues, and high computational complexity, resulting in low accuracy and efficiency and making them difficult to apply in real time on resource-constrained devices.

Method used

By refining motion standards into detailed requirements, combining a lightweight OpenPose network with target object recognition, it can evaluate human key points and object positions in real time, and use pruning techniques to optimize the network structure, making it suitable for embedded devices.

Benefits of technology

It improves the accuracy and efficiency of scoring, enabling real-time assessment of basic human motor skills on embedded devices, identifying substandard details, reducing computational complexity, and making it suitable for devices with limited resources.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246347B_ABST
    Figure CN116246347B_ABST
Patent Text Reader

Abstract

This invention relates to the field of computer vision technology, and provides a standard judgment method and scoring device for basic human motor skills. The method includes: recording relevant information of key points of the human body; identifying the person with the smallest change in human body indicator features in two adjacent frames as the same person; recording the position information of the target object; processing the relevant information of the key points and the position information of the target object to obtain action detail information; comparing and analyzing the action detail information with the detail requirements, and synchronously displaying the preliminary evaluation results on the image; after completing the processing and analysis of all frames of the video, obtaining the final evaluation result corresponding to each detail requirement; the evaluation result of the action standard is derived from the final evaluation result of the corresponding detail requirement through a bitwise AND operation; this invention improves the accuracy of scoring by breaking down each action standard into different detail requirements; and makes sports training more targeted and scientific by judging human movements in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a standard method and scoring device for judging basic human motor skills. Background Technology

[0002] Promoting a healthy China is a priority development task. School sports, leisure sports, competitive sports, and mass fitness will all play important roles. Standard and correct sports movements can effectively improve physical fitness, and vice versa. Meanwhile, professionals mainly evaluate sports movements through visual observation and experience, and provide corresponding suggestions. Although professional evaluation through direct observation has a certain degree of accuracy, it is highly subjective. On the one hand, human observation has various limitations; on the other hand, prolonged evaluation can lead to fatigue, both of which seriously affect the accuracy of recording and evaluating the complete movement process. Therefore, researching a technology that can accurately and intelligently evaluate the standardization of sports movements has significant application value.

[0003] Some existing motion assessment or analysis techniques are listed below:

[0004] Patent application number 201710922613.4, entitled "A Video-Based Motion Detection Method and Apparatus," provides an evaluation model based on global and local motion features. The global motion evaluation model is used to detect the degree of completion of human movements in a video relative to standard movements, while the local motion evaluation model is used to detect the degree of completion of human joint movements in video frames relative to standard joint movements.

[0005] Patent application number 201910008331.2, entitled "A Motion Evaluation Method, Device, Mobile Terminal, and Readable Storage Medium," provides a motion evaluation method. This method first identifies the joints of the object to be tested, then determines the rotation angle based on the different colors marked on the joints, thereby creating a motion joint diagram. Finally, it compares this diagram with a pre-stored standard motion joint diagram to determine whether the motion is standard.

[0006] Patent application number 202110104619.7, entitled "A Method and System for Evaluating Sports Movements Based on Skeletal Point Depth Images," provides a method for evaluating sports movements based on human skeletal point depth images. This method involves obtaining the three-dimensional coordinates of human skeletal points corresponding to the sports movement to be evaluated based on the depth image, determining the actual feature values ​​of the sports movement to be evaluated (actual distances between key skeletal points and actual angles of vectors between skeletal points), and then comparing these values ​​with corresponding standard feature values.

[0007] Patent application number 202110250096.7, entitled "A Temporal Action Evaluation Method Based on Keyframe Preferences," provides an action evaluation method based on keyframe preferences. The method involves: collecting skeletal temporal data of standard actions and user actions; extracting keyframe region lists and non-keyframe region lists from the standard action sequence and user action sequence, respectively; performing dynamic time warping on the keyframe regions and non-keyframe regions of both the user action sequence and the standard action sequence; calculating the similarity between keyframe regions and non-keyframe regions using a cosine similarity algorithm; and finally, weighting the keyframe region similarity and non-keyframe region similarity to obtain a comprehensive action score.

[0008] Patent application number 202110578414.2, entitled "A Personalized Motion Posture Estimation and Analysis Method and System Based on Temporal Consistency," provides a personalized motion posture estimation and analysis method based on temporal consistency. This method involves: identifying actions in instructional videos and student self-study videos, extracting key skeletal coordinates; aligning asynchronous actions in the instructional and student self-study videos to ensure consistent action speeds; measuring the actions based on the ratio of key skeletal coordinate angle deviations and joint distances in each frame of the aligned instructional and student self-study videos, and calculating the action difference degree; and generating multi-scale scores from local to overall action based on the action difference degree and a motion pyramid model.

[0009] In summary, the problems with existing technologies are:

[0010] 1) Most of the methods mentioned above are matching-based evaluation methods, which obtain the action of the target to be evaluated through object detection algorithms and then compare it with a standard action or template to obtain an action score. However, these methods have strict scaling requirements for the action. The scaling formula is expressed as follows:

[0011]

[0012] Where σ represents scale, S represents the size of the target person (which can be volume, area, skeleton size, etc.), L represents the vertical distance from the camera lens, and F(·) represents linear transformation.

[0013] Therefore, such methods require matching and evaluation of movements under the premise of the same scale. Typically, L is controllable, but S is difficult to keep consistent. For example, if the standard movement is demonstrated by an adult, and the test subject is a child, the movement will be difficult to match. Especially when the measurement standard includes factors such as joint distance, it will lead to a decrease in evaluation accuracy, as seen in patent documents with application numbers 202110104619.7 and 202110578414.2.

[0014] 2) A continuous action has spatiotemporal dimensions; that is, an action at different moments is composed of sub-actions made up of different parts of the body (spatiality), while a continuous action is composed of actions at multiple moments (temporality). Most existing methods focus on the temporal dimension of the action. For example, patent application number 202110250096.7 uses keyframe preferences to obtain actions at important moments; patent application number 202110578414.2 compares the action to be evaluated with a standard action by aligning the time. These methods ignore the differences in movement frequency. The same action might take an adult only three or four seconds to complete, while a child might need six or seven seconds or even longer. Therefore, neither keyframe preference selection nor time alignment can effectively solve this problem. A very small number of methods focus on the spatial dimension of the action. For example, patent application number 201710922613.4 proposes using one or more of the following as behavioral characteristics: the position of the human joint relative to the body, the angle of the human joint, the body's orientation, and the body's tilt angle. However, it ignores the periodicity of actions, such as running or horse stance. In particular, this method will no longer be applicable to the basic motor skills that this application targets.

[0015] 3) During movement, occlusion problems inevitably occur, such as background occlusion or self-occlusion. Background occlusion can be solved by setting a clean physical environment. Self-occlusion is usually solved by the following techniques: 1. Using a multi-view camera to acquire images or videos of the target from different angles; 2. Using a depth camera to acquire current depth information to construct a three-dimensional coordinate system, such as the RGB-D camera used in patent document application number 202110104619.7. However, these methods require specialized equipment for data acquisition (which is costly), and data acquisition technology based on multi-view cameras also requires design and construction schemes, which hinders its widespread application.

[0016] 4) There are many different types of daily movements, so there is currently a lack of a scientific and complete scheme for assessing basic human motor skills and corresponding evaluation techniques.

[0017] Furthermore, for pose estimation networks, while increasing the number of layers enhances representation and modeling capabilities and leads to superior performance, it also increases network complexity and redundancy, resulting in significant efficiency issues. Many pose estimation methods proposed in recent years tend to use large-scale, deep, and complex network structures. The enormous number of parameters and computational demands place increasingly high demands on the computing power and memory of hardware processing platforms, making deployment difficult on edge devices with limited memory and computing resources, such as embedded platforms or mobile devices. Even when deployed on highly integrated, high-performance cloud environments and transmitting data over a network, the massive bandwidth consumption significantly increases industrial costs, and transmission latency prevents real-time feedback, hindering widespread practical application. Therefore, designing a fast, lightweight pose estimation network that can efficiently run on localized, real-time devices with limited resources has become an urgent problem to solve.

[0018] Therefore, overcoming the shortcomings of the existing technology is an urgent problem to be solved in this technical field. Summary of the Invention

[0019] The technical problem to be solved by this invention is:

[0020] Currently, in most cases, the human body key point information extracted by key point extraction devices is manually compared with the motion standards specified in TGMD-3. However, given the insufficient precision of the TGMD-3 motion standards, the accuracy of the scoring is inadequate, making it difficult to identify subtle motion defects. Furthermore, manual scoring reduces objectivity and is time-consuming, resulting in low cost-effectiveness for motion capture. Therefore, it is necessary to explore how to improve the OpenPose network to quickly acquire human body key information.

[0021] The present invention achieves the above objectives through the following technical solutions:

[0022] In a first aspect, the present invention provides a standard method for judging basic human motor skills, comprising: if a human body is identified in a video, generating a JSON file corresponding to each frame, wherein the JSON file is used to record relevant information of key points of the human body; determining the person whose human body indicator features change the least in two adjacent frames as the same person and assigning them the same number; if a target object corresponding to the motor skill specified by the target human body is identified in the video, recording the position information of the target object;

[0023] By processing the relevant information of the key points of the target human body and the position information of the target object in each frame, motion detail information is obtained; each motion standard specified by TGMD-3 is broken down into different detail requirements; the motion detail information is compared and analyzed with the corresponding detail requirements, and the preliminary evaluation results of the corresponding detail requirements are displayed on each frame image in sync with the motion; after completing the processing and analysis of all frames of the video, the final evaluation result of the target human body's motion skill corresponding to each of the detailed requirements is obtained, and the evaluation result of the motion standard is obtained by AND operation with the final evaluation result of the corresponding detail requirements.

[0024] Secondly, the present invention provides a scoring device for basic human motor skills, using the standard judgment method for basic human motor skills described in the first aspect. The device includes a human posture prediction unit, a movement standardization judgment unit, and a motor skill scoring unit, wherein:

[0025] The human pose prediction unit includes a key point extraction unit and a tracking processing unit. The key point extraction unit is used to perform human body recognition in the video and record the relevant information of the key points of the human body in each frame in the corresponding JSON file. The tracking processing unit is used to read the JSON file in pairs according to the video frame number and determine the person with the smallest change in human body indicator features in two adjacent frames as the same person and assign the same number.

[0026] The motion standardization judgment unit includes a target object recognition unit, a motion detail analysis unit, and a process synchronization unit. The target object recognition unit is used to identify the target object corresponding to the specified motion skill of the target human body and record the position information of the target object. The motion detail analysis unit is used to process the relevant information of the key points of the target human body and the position information of the target object in each frame to obtain motion detail information. The process synchronization unit is used to connect the identified key points with line segments in a specified order and synchronously display the body skeleton of the target human body, the hand skeleton of the target human body, the number of the target human body, the position coordinates of the target human body, the name of the target object, and the position coordinates of the target object in each frame.

[0027] The motor skill scoring unit is used to compare and analyze the action details in each frame with the corresponding detail requirements to obtain a preliminary evaluation result of the corresponding detail requirements; after processing and analyzing all frames of the video, the final evaluation result of the target human's motor skills corresponding to each of the detail requirements is obtained; the evaluation result of the action standard is obtained by performing an AND operation on the final evaluation result of the detail requirements.

[0028] In a second aspect, the present invention provides a scoring device for basic human motor skills, comprising at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the processor to perform the standard judgment method for basic human motor skills as described in the first aspect.

[0029] Compared with the prior art, the beneficial effects of the present invention are:

[0030] This invention refines the original action standards by breaking down each action standard into different detailed requirements, thereby improving the accuracy of scoring. The action detail information in this invention includes relevant information about the key points of the target human body and the position information of the target object. The position information of the target object is used to assist in calculating the standardization of the human body's movements, further improving the accuracy of scoring. In this invention, the completion degree of each detailed requirement is displayed synchronously with the movement in each frame of the image, allowing for real-time judgment of whether the human body's movements are standard. This helps to discover previously difficult-to-detect substandard details, making sports training more targeted and scientific.

[0031] Furthermore, removing redundant structures from pose estimation networks using pruning techniques is an effective lightweight method. On the one hand, networks with an appropriate number of layers and parameters can generally meet accuracy requirements in practical applications; on the other hand, deep neural networks often suffer from over-parameterization, with the network size often far exceeding task requirements, and a considerable portion of parameters not playing a role in the network model. Using the method designed in this invention to prune the OpenPose network can improve detection speed and reduce the number of parameters and computational load with almost no loss of accuracy. It can also be deployed in embedded systems or mobile devices for real-world applications, achieving fast and lightweight pose estimation. Attached Figure Description

[0032] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments of the present invention will be briefly described below. Obviously, the drawings described below are merely some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without any creative effort.

[0033] Figure 1 This is a flowchart of a standard method for judging basic human motor skills provided in an embodiment of the present invention;

[0034] Figure 2 This is an example of the JSON file content in a standard method for judging basic human motor skills provided in an embodiment of the present invention;

[0035] Figure 3 This is a diagram showing the positions of 25 key points on the human body in a standard method for judging basic human motor skills provided in an embodiment of the present invention.

[0036] Figure 4 This is a schematic diagram showing the synchronous display of the evaluation results of the detailed requirements for running skills in a standard judgment method for basic human motor skills provided in an embodiment of the present invention;

[0037] Figure 5 This is a schematic diagram showing the final evaluation result of running skills in a standard judgment method for basic human motor skills provided in an embodiment of the present invention.

[0038] Figure 6 This is a diagram showing the positions of 21 key points on the human hand in a standard method for judging basic human motor skills provided in an embodiment of the present invention.

[0039] Figure 7 This is a flowchart of a multi-granularity adaptive pruning process provided in an embodiment of the present invention;

[0040] Figure 8 This is a schematic diagram of the lightweight OpenPose backbone network structure provided in this embodiment of the invention;

[0041] Figure 9 This is a schematic diagram of the lightweight OpenPose backbone network structure provided in this embodiment of the invention;

[0042] Figure 10 This is a schematic diagram of the refined network structure of the lightweight OpenPose provided in this embodiment of the invention;

[0043] Figure 11 This is a schematic diagram of the refined network structure of the lightweight OpenPose provided in this embodiment of the invention;

[0044] Figure 12 This is a structural block diagram of a scoring device for basic human motor skills provided in an embodiment of the present invention;

[0045] Figure 13 This is a structural block diagram of another scoring device for basic human motor skills provided in an embodiment of the present invention;

[0046] Figure 14 This is a structural block diagram of the motion detail analysis unit in a scoring device for basic human motor skills provided in an embodiment of the present invention;

[0047] Figure 15 This is a partial structural schematic diagram of a scoring device for basic human motor skills provided in an embodiment of the present invention. Detailed Implementation

[0048] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0049] In various embodiments of the present invention, the symbol “ / ” indicates that it has two functions at the same time, while the symbol “A and / or B” indicates that the combination between the preceding and following objects connected by the symbol includes three cases: “A”, “B”, and “A and B”.

[0050] Furthermore, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other. The present invention will now be described in detail with reference to the accompanying drawings and embodiments.

[0051] Example 1:

[0052] Embodiment 1 of the present invention provides a standard method for judging basic human motor skills, such as... Figure 1 As shown, it includes:

[0053] In steps 202 and 203, subsequent steps are performed based on whether a human body is detected in the video. If a human body is detected in the video, a JSON file is generated for each frame, and the JSON file is used to record relevant information about the key points of the human body. If no human body is detected in the video, the corresponding error is displayed on the console and the process is stopped.

[0054] The key points include 25 key points on the human body and 42 key points on both hands. The relevant information of the key points includes the predicted position coordinates and the predicted confidence level of the key points. The predicted position coordinates of the key points are the coordinates of the suspected key point positions in each frame of the image. The probability that the coordinates of the suspected key point positions are exactly the coordinates of the actual key point positions is the predicted confidence level of the key point. The predicted confidence level of the key point ranges from 0 to 1. The higher the predicted confidence level of the key point, the higher the predicted accuracy of the predicted position coordinates of the corresponding key point.

[0055] The above mainly describes the extraction of key points from the human body in a video. Taking OpenPose as an example, the key point extraction process involves installing the underlying Caffe deep learning framework, the CUDA platform for running deep neural networks, and its dependency library cuDNN, etc. Then, referring to the official documentation installation.md on GitHub, compilation is performed, creating the OpenPose_ROOT / build path, and generating Makefiles. Since Caffe and OpenCV are already installed, CMake needs to be provided with library and include paths. For OpenCV, OpenCV_INCLUDE_DIRS and OpenCV_LIBS_DIR are used to specify the paths to the installed OpenCV libraries and include paths, or OpenCV_CONFIG_FILE is used to specify the path to OpenCVConfig.cmake. For Caffe, Caffe_INCLUDE_DIRS and Caffe_LIBS are used to specify the paths to the installed Caffe libraries and include paths. Finally, the `make -j`nproc` command is used to complete the compilation.

[0056] After compilation, human recognition can be performed on the target video. Execute `OpenPose.bin` in the ` / build / examples / OpenPose` directory using the command line, and use the command `video "${PATH_TO_VIDEO}"` to process the specified video in the `PATH_TO_VIDEO` directory. Then, perform another check. If a human is detected in the video, proceed with further processing; otherwise, if no human is detected, display "No people detected" in the console output window and terminate the process.

[0057] In the command line for processing videos that identify the presence of human bodies, add a command `write_json"${PATH_TO_SAVE}"`. This command causes OpenPose to generate a JSON file while recognizing human bodies. The rule for generating this file is that each frame of the video will generate a corresponding JSON file, named "(video name)_(frame number minus one)_keypoints". For example, when processing a video whose full name is "xxx.avi", the full name of the JSON file generated for the first frame of the video will be "xxx_000000000000_keypoints.json".

[0058] The content of the output JSON file is as follows Figure 2As shown, the entire structure is of dictionary type, and the stored data all have mapping relationships; in Figure 2 In the diagram, "version" is mapped to 1.3, indicating that the human pose estimation method is version 1.3; the mapping for "people" is represented by a list type, where each list element corresponds to a dictionary-type result of the skeleton keypoint recognition for a single person in a single image or video frame; the mapping for "person_id" is [-1], meaning no actual ID; the mapping for "pose_keypoints_2d" contains the required body keypoint information, with 25 groups of 3 data points each, and the first two data points of each group are the keypoints. Figure 3 The BODY_25 model shown has keypoint coordinates (x, y) corresponding to the index numbers. The second data is the prediction confidence score, ranging from [0, 1], with a higher score indicating a higher prediction accuracy. Similarly, the mappings "hand_left_keypoints_2d" and "hand_right_keypoints_2d" contain information about the left and right hand keypoints. The 21 points on a single hand correspond to keypoints as follows: Figure 4 As shown; in addition, using commands such as --face and --3d, you can create mappings of facial key point coordinates and 3D coordinates of various parts of each person's face.

[0059] In step 201, to reduce the impact of noise on the extraction of human key point information, each frame of the video is preprocessed before human recognition is performed. The specific video preprocessing process includes: firstly, analyzing the target video path (one of the inputs), using Python's `os` module to determine if a corresponding video exists at that path; if so, using OpenCV to read and obtain basic video information, including total number of frames, frame rate, frame width, and frame height; then, denoising the video. If no corresponding video exists, an error message "Error opening file ${PATH_TO_VIDEO}" is displayed in the console output window, and the process is terminated. The denoising method chosen for the video is non-local mean filtering (NMR). Means; it uses global information of the image for denoising, taking image patches as units. The smaller the Gaussian weighted Euclidean distance between pixels, the more similar the structure. Then, the average value of these similar image patches is taken to achieve the purpose of removing noise in the image relatively well. Compared with image smoothing methods that use local information, such as mean filtering, this algorithm has higher time complexity, but can better preserve image details. Since the objects to be identified are all color images and videos, they need to be converted to the LaB color space first, and then the luminance L and color channels ab are denoised separately.

[0060] In step 204, human body tracking processing is performed. In two adjacent frames, the person with the smallest change in human body indicator features is identified as the same person and assigned the same number.

[0061] The human body indicator features include one or more of the following: relative displacement of the human body, human body posture, and human body prediction confidence.

[0062] The relative displacement of the human body refers to the relative displacement between the position coordinates of the same target human body in two adjacent frames; the human body's posture refers to the angles between different parts of the target human body and the relative positions between these parts; the prediction confidence of the human body refers to the probability that the position of the suspected target human body is exactly the position of the actual target human body in the next frame.

[0063] To address the issue that some human pose estimation algorithms cannot perform continuous recognition in multi-person scenarios, in this embodiment of the invention, the person with the smallest change in human body indicator features in two adjacent frames is determined to be the same person and assigned the same number. The specific determination method is as follows: the determination is performed according to the priority order from the relative displacement of the human body, the human body's action posture to the prediction confidence of the human body.

[0064] Taking OpenPose as an example, human body tracking is processed as follows: Although it comes with a tracking module, according to the OpenPose usage guide quick_start.md on GitHub, the value of number_people_max can only be set to 1, meaning that a maximum of one person can be tracked. Furthermore, by default, without using the tracking module, the numbering of each person in each frame is random. To re-determine the numbering of each person, we first need to start with the existing list of key points on the human skeleton, and reset their position in the list according to the relative displacement of the person in adjacent frames. Because in the extremely short time interval between two adjacent frames, the person with the smallest coordinate change is usually the same person. Considering that in scenarios with occlusion or high-speed movement, a person's skeleton cannot be completely identified, especially the limbs, the coordinates of the neck are used as the judgment standard. Figure 3The first location point is identified. Assume that *a* people were identified in the current frame and *b* people in the next frame. Based on the extracted indexes from the output JSON file, obtain all human keypoints in the current and next frames. Then, iterate through these keypoints and calculate the Euclidean distance between pairs of abdominal coordinates between the two frames, resulting in a distance list *dist* of length *a*×*b*. For the coordinates of the *m*th person in the *people* list in the current frame and the *n*th person in the next frame, the generated *dist* will be formatted as [ρ, m, n]. At this point, you cannot directly sort the points by distance from smallest to largest and then take the first *min* (*a*, *b*) *m* pairs to adjust the order, because this might include cases where two people in the next frame are very close to someone in the current frame. Directly taking the first few pairs would lead to duplicate numbering. For pairing, to ensure that each mn pair is unique, the values ​​of the mn pairs can be bubble sorted, starting from the minimum value and recording the corresponding m or n values ​​sequentially. If a duplicate is encountered, the distance values ​​are compared, and the larger one is deleted. If no duplicate is encountered, the pointer is moved one position to the left to continue recording. If a special case is encountered where the distance ρ is the same, the corresponding human body posture of the mn pair is analyzed, specifically the angles and relative positions between different parts of the human body. A pair of human body postures that are the same between two frames is considered to be the same person. Similarly, if a special case is encountered where multiple pairs of postures are also the same, the change in confidence of the corresponding human body parts of the mn pair is judged. A pair with a smaller change in average confidence between two frames is considered to be the same person.

[0065] By prioritizing the relative displacement of the human body, the posture of the human body, and the prediction confidence of the human body, the correspondence of the target human body in two adjacent frames can be guaranteed. Finally, the positions of the min(a,b) mn pairs with the smallest distance are adjusted in the list. When the number of people in the next frame is equal to or less than the number of people in the current frame, the nth position coordinate set of the people list in the next frame is adjusted to the mth position in turn according to the mn pairs. When the number of people in the next frame is more than the number of people in the current frame, in addition to adjusting the position, the unadjusted coordinate set is randomly assigned empty positions to keep the total list length unchanged.

[0066] In steps 205 and 206, if a target object corresponding to the specified motor skill of the target human body is identified in the video, the position information of the target object is recorded to assist in judging the standard of the human body's movements; if no target object corresponding to the specified motor skill of the target human body is identified in the video, a corresponding error is displayed on the console, and the evaluation items involving the target object in the final evaluation result are not evaluated, and a corresponding error is displayed.

[0067] The target human body refers to the human body in the video whose movements need to be judged for their standardization. The motor skills include six types of movement skills and seven types of ball skills. The six types of movement skills include: running, horse stance running, single-leg hopping, running jump, side slide, and standing long jump. The seven types of ball skills include: hitting a stationary ball with both hands, hitting a tossed ball with a forehand, dribbling a ball with one hand in place, catching a ball with both hands, kicking a stationary ball, overhead throw, and low throw. The motor skills specified by the target human body are those for which the target human body needs to be judged for their standardization. The target objects corresponding to the motor skills specified by the target human body include various sports balls, bats, and rackets involved in the seven types of ball skills.

[0068] The above content mainly focuses on the recognition of target objects corresponding to the specified motor skills of the target human body. This requires calling relevant models and weight files to recognize the target objects in each frame of the video, obtain their positions, and output them in list format. The methods for creating the dataset and training the model are as follows:

[0069] Regarding the creation of the dataset, firstly, a large number of existing videos demonstrating various sports skills were used, and the video frames containing the target object were saved as images. Then, these images were labeled using annotation tools such as LabelImg to generate label files containing the target object's location information. These files were then uniformly converted to txt format, with each line of data in the file representing ground truth information: the category number, and the normalized values ​​of the coordinates of the four points of the bounding box (top left, bottom left, top right, and bottom right). After the annotation was completed, the dataset was divided into training, validation, and test sets, placed in a specified path, and corresponding path sets were created as txt files.

[0070] For model training, YOLOv3 was used to train the labeled dataset. YOLO is a target recognition method that balances speed and accuracy. In particular, version v3 improves the detection rate of small targets and can better complete the task of recognizing various small target objects. The target objects are divided into three types: sports balls, baseball bats, and tennis rackets. A class information file, classes.names, was created based on this, with each class name on a separate line. Then, based on these categories, the script file create_custom_model.sh generated a custom network model configuration information file, yolov3-custom.cfg, which includes convolutional layer information such as the number of convolutional kernels, kernel size, and stride, as well as configuration information for YOLO layers and other layers. Finally, YOLOv3.train.py was run to start training using the dataset in the specified path.

[0071] After training, the pth file with the smallest loss is selected as the model's weight file. Code is written to load the network configuration, weight information, and label categories. Each frame of image is converted into blob format and sent to the network input layer. Information such as the detection boxes in the network output layer is obtained. Forward propagation is set to filter out detection boxes with low confidence. Finally, non-maximum suppression is applied to further filter them out, thus obtaining the rectangular range in which the target object is most likely to appear in the image.

[0072] In step 207, motion detail information is obtained by processing the relevant information of the key points of the target human body and the position information of the target object in each frame. Specifically, processing the relevant information of the key points of the target human body and the position information of the target object in each frame involves: connecting the key points identified in each frame with line segments in a specified order to obtain the body skeleton and hand skeleton of the target human body, thereby calculating the angles formed by the limbs of the target human body; obtaining the distances between the joints of the target human body and their relative positions based on the relevant information of the identified key points of the target human body; and calculating the distances between the joints of the target human body and their relative positions based on the position information of the target object. The motion detail information includes one or more of the following: the angles formed by the limbs of the target human body, the distances between the joints of the target human body, the relative positions between the joints of the target human body, the distances between the joints of the target human body and their relative positions.

[0073] In step 208, each action standard specified in TGMD-3 is broken down into different detailed requirements; the action details are compared and analyzed with the corresponding detailed requirements, and the preliminary evaluation results of the corresponding detailed requirements are displayed on each frame of the image in sync with the action.

[0074] In step 209, after processing and analyzing all frames of the video, the final evaluation results of the target human's motor skills corresponding to each of the detailed requirements are obtained.

[0075] In step 210, the evaluation result of the action standard is obtained by AND operation of the final evaluation result of the corresponding detailed requirements, and the evaluation result of the action standard can be output in the form of a list.

[0076] TGMD-3 stands for Test of Gross Motor Development-3, 3rd Edition. It is an internationally advanced assessment standard developed by Professor Dale A. Ulrich of the University of Michigan over 30 years, based on natural human development and aimed at developing students' basic motor skills. The test includes 13 items: running, horse stance running, single-leg hopping, running jump, side slide, standing long jump, hitting a stationary ball with both hands, hitting a dropped ball with a forehand, throwing a ball with one hand, catching a ball with both hands, kicking a stationary ball, overhead throwing, and low throwing. TGMD-3 specifies different scoring rules for different types of motor skills. Specifically, each type of motor skill corresponds to a set of scoring rules, and each set of scoring rules includes multiple motor standards.

[0077] It should be noted that the sports standards included in the scoring rules specified by TGMD-3 are known technical content in the field. The innovation of this invention lies in further refining each sports standard into different detailed requirements, the evaluation method of the detailed requirements, the processing of the evaluation results of the detailed requirements, and how to evaluate the movement standards through the detailed requirements to determine the standard of the sports skills specified by the target human body.

[0078] Specifically, to avoid the scope of each exercise standard being too narrow and affecting the accuracy of the scoring, each exercise standard is further refined by breaking it down into different detailed requirements. Each detailed requirement is evaluated, and the evaluation results are processed to obtain the evaluation result for each exercise standard. This evaluation is then conducted using multiple exercise standards of the same type of exercise skill to determine the standardization of the specified exercise skill for the target individual. The evaluation results for the detailed requirements include preliminary evaluation results and final evaluation results. The evaluation result for the movement standard is derived from the final evaluation results of the corresponding detailed requirements through a bitwise AND operation.

[0079] Taking running as an example, in each frame of the target human running, according to Figure 4 The score for each detailed requirement is displayed as shown; after the target human finishes running, the score is... Figure 5 The score for each action is displayed as shown, and the final score is output as a list in the console; Figure 4 and Figure 5 Taking running as an example, when the final evaluation result of all the detailed requirements corresponding to the same movement standard is 1, the evaluation result of the corresponding movement standard is 1; otherwise, the evaluation result of the corresponding movement standard is 0. Figure 4 and Figure 5In the process, an evaluation result of 1 is displayed as True, and an evaluation result of 0 is displayed as False. After the entire synchronous evaluation process is completed, the final evaluation result is displayed on the console. Figure 4 and Figure 5 Taking the running example shown, the final evaluation result is [0,1,1,0], which is [False,True,True,False] shown in the figure.

[0080] Specifically, the process first evaluates whether the action details in each frame meet the various detail requirements, obtains preliminary evaluation results for each detail requirement, and displays them synchronously on each frame. After processing and comparing all frames of the video, the final evaluation results for each detail requirement are obtained. Then, by classifying, organizing, and calculating the final evaluation results for each detail requirement, the evaluation results for the action standard are obtained and displayed in the final evaluation results.

[0081] Specifically: if the action detail information meets the detail requirements, the final evaluation result of the corresponding detail requirement is 1; otherwise, the final evaluation result of the corresponding detail requirement is 0; when the final evaluation results of all the detail requirements corresponding to the same action standard are 1, the evaluation result of the corresponding action standard is 1; otherwise, the evaluation result of the corresponding action standard is 0.

[0082] The detailed requirements are further divided into global requirements and local requirements, wherein:

[0083] For the global requirement, the global requirement is determined in each frame, specifically as follows:

[0084] In each frame, if the action details meet the global requirements, the preliminary evaluation result of the corresponding global requirements is displayed as 1 on each frame image; otherwise, the preliminary evaluation result of the corresponding global requirements is displayed as 0.

[0085] After processing and analyzing all frames of the video, when the proportion of frames with a preliminary evaluation result of 1 for the corresponding global requirement exceeds a first preset value, the final evaluation result of the corresponding global requirement is displayed as 1; otherwise, the final evaluation result of the corresponding global requirement is displayed as 0. The first preset value is the pre-set proportion of frames with a preliminary evaluation result of 1 for the global requirement to all frames, generally set between 50% and 100%, and adjusted according to actual requirements.

[0086] For local requirements, the determination is made when the corresponding condition is met, specifically as follows:

[0087] In each frame, if the action detail information meets the local requirements, the preliminary evaluation result of the corresponding local requirements is displayed as 1 on the frame image after the corresponding conditions are completed; otherwise, the preliminary evaluation result of the corresponding local requirements is displayed as 0.

[0088] After processing and analyzing all frames of the video, when the proportion of the number of times the preliminary evaluation result of the corresponding local requirement is 1 exceeds the total number of times the corresponding condition is completed, the final evaluation result of the corresponding local requirement is displayed as 1; otherwise, the final evaluation result of the corresponding local requirement is displayed as 0. The second preset value is the proportion of the number of times the preliminary evaluation result of the local requirement is 1 to the total number of times the corresponding condition is completed, which is generally set between 50% and 100% and adjusted according to actual requirements.

[0089] The method for calculating the evaluation result of the corresponding action standard from the evaluation results of the global requirements and the local requirements is as follows: when the final evaluation results of all global requirements and local requirements corresponding to the same action standard are 1, the evaluation result of the corresponding action standard is 1; otherwise, the evaluation result of the corresponding action standard is 0.

[0090] In this embodiment of the invention, the key points include 25 key points on the human body and 42 key points on both hands. The relevant information of the key points includes the predicted position coordinates of the key points and the prediction confidence of the key points. The predicted position coordinates of the key points are the coordinates of the suspected key point positions in each frame of the image. The probability that the coordinates of the suspected key point positions are exactly the coordinates of the actual key point positions is the prediction confidence of the key point. The prediction confidence of the key point ranges from 0 to 1. The higher the prediction confidence of the key point, the higher the prediction accuracy of the predicted position coordinates of the corresponding key point.

[0091] In this embodiment of the invention, the processing of the relevant information of the key points of the target human body and the position information of the target object in each frame specifically includes:

[0092] The key points identified in each frame are connected by line segments in a specified order to obtain the body skeleton and hand skeleton of the target human body, thereby calculating the angles formed by the limbs of the target human body.

[0093] Based on the relevant information of the key points of the target human body, the distance between the joints of the target human body and the relative position between the joints are obtained, and the number of the target human body and the position of the target human body in each frame are marked.

[0094] Based on the location information of the target object, the name of the target object is marked, and the distance between the joints of the target human body and the target object, as well as the relative position between the joints of the target human body and the target object, are calculated.

[0095] Taking OpenPose as an example, regarding the synchronized display of the target human body, after identifying the key points of the human body, the write_video command is used to connect them, resulting in the following: Figure 3 The body skeleton diagram shown and as Figure 6 The image shows a hand skeleton. Then, using the OpenCV function cv2.rectangle (formatted as "image, (top left corner coordinates), (bottom right corner coordinates), (rectangle color), line thickness"), a square marker box is used to mark the area where the human body appears on the video frame image based on the coordinates of the human body key points. Finally, the OpenCV function cv2.putText (formatted as "image, content to be written, (text coordinates), font, font size, (font color), font thickness") is used to display the human body number set after tracking processing at the top left corner of the marker box.

[0096] Regarding the synchronous display of target objects, the cv2.rectangle function is used to mark the range of the target object's appearance on the video frame image with a square marker based on the identified target object's position information, and the cv2.putText function is used to display the name of the target object in the upper left corner of the marker box.

[0097] Regarding the synchronous display of motion details, the cv2.putText function is used to display the completion status of various detailed requirements corresponding to the specified motion skill standards of the target human body in the upper left position of each video frame as True or False. True indicates that the evaluation result is 1, and False indicates that the evaluation result is 0.

[0098] In this embodiment of the invention, the action details information includes one or more of the following: the angle formed by the limbs of the target human body, the distance between the joints of the target human body, the relative position between the joints of the target human body, the distance between the joints of the target human body and the target object, and the relative position between the joints of the target human body and the target object.

[0099] In this embodiment of the invention, the detailed requirements are divided into global requirements and local requirements. The global requirements are determined in each frame, and the local requirements are determined when the corresponding conditions are met. When the final evaluation results of all global requirements and local requirements corresponding to the same action standard are 1, the evaluation result of the corresponding action standard is 1; otherwise, the evaluation result of the corresponding action standard is 0.

[0100] In this embodiment of the invention, the global requirement is determined in each frame, specifically as follows:

[0101] In each frame, if the action detail information meets the global requirement, the preliminary evaluation result of the corresponding global requirement is displayed as 1 on each frame image; otherwise, the preliminary evaluation result of the corresponding global requirement is displayed as 0.

[0102] After processing and analyzing all frames of the video, if the proportion of frames with a preliminary evaluation result of 1 for the corresponding global requirement exceeds a first preset value, the final evaluation result for the corresponding global requirement is displayed as 1; otherwise, the final evaluation result for the corresponding global requirement is displayed as 0.

[0103] In this embodiment of the invention, the local requirement is determined when the corresponding condition is met, specifically as follows:

[0104] In each frame, if the action detail information meets the local requirements, the preliminary evaluation result of the corresponding local requirements is displayed as 1 on the frame image after the corresponding conditions are completed; otherwise, the preliminary evaluation result of the corresponding local requirements is displayed as 0.

[0105] After processing and analyzing all frames of the video, if the number of times the preliminary evaluation result of the corresponding local requirement is displayed as 1 exceeds the proportion of the total number of times the corresponding condition is completed, the final evaluation result of the corresponding local requirement is displayed as 1; otherwise, the final evaluation result of the corresponding local requirement is displayed as 0.

[0106] Example 2:

[0107] The detailed requirements for the six types of movement skills and the seven types of ball skills in Example 1 above, and the analysis method for these detailed requirements, are as follows:

[0108] The standard running technique and corresponding detailed requirements are shown in Table 1. Standard technique 1 includes three detailed requirements: the first is that "the angle between the upper arm and forearm should not exceed 160 degrees," which means... Figure 3The first requirement is the angle between line segments 2-3 and 3-4, and 5-6 and 6-7. The second requirement is that "the angle between the upper arm and the torso should not be less than 20 degrees." This detail involves judging the extreme value of the arm swinging to the front or back. Therefore, an adaptive sliding window method is used, with a window size of 4 frames. A flag is set for each extreme value detected. If a more extreme value is found within the next 10 frames, it is replaced; otherwise, the value is determined. Determining the extreme value indicates that one arm swing or one step has been completed. Then, the angles between line segments 2-3 and 1-8, and 5-6 and 1-8 are calculated for the video frame image corresponding to this extreme value. The third requirement is "do not use the same hand and foot at the same time." When the angles between line segments 2-3 and 1-8, and 5-6 and 1-8 are greater than a certain value, it indicates that both hands have been swung out. The analysis is then performed to determine whether the opposite hand and foot are simultaneously in front of or behind the body.

[0109] The detailed requirement of action standard 2 is "both feet leave the ground simultaneously for a short period of time". The ground is determined by the position of the sole of the foot, i.e., line segment 22-24 or line segment 19-21, when each step becomes horizontal. The video frame image corresponding to the maximum angle formed by line segments 9-10 and 12-13 when the legs are at their foremost or last position is determined by the sliding window method. The relative positions of points 22 and 19 (toes) and points 21 and 24 (heels) with the ground are analyzed.

[0110] The detailed requirement of action standard 3 is that "the angle between the line connecting the toes and heels and the ground is not less than 20 degrees". Determine the angle between the lower leg and the torso when the lower leg is in front of the torso, i.e., the angle between line segment 9-10 or line segment 12-13 and line segment 1-8, which decreases to a certain value and this angle decreases monotonically over several consecutive frames. Calculate the angle between line segment 22-24 or line segment 19-21 and the horizontal plane.

[0111] The detailed requirement of action standard 4 is that "the swing leg bends at nearly 90 degrees when it swings back". Similarly, the sliding window method is used to determine the video frame image corresponding to when both legs swing to the front or back, and the angle between line segments 9-10 and 10-11, or line segments 12-13 and 13-14 is calculated.

[0112] Table 1:

[0113]

[0114] Table 2 shows the standard and corresponding detailed requirements for the horse stance running movement. Among them, the movement standard 1 includes three detailed requirements: the first detailed requirement is analyzed in the same way as the first detailed requirement in the running movement standard 1; the second detailed requirement is analyzed in the same way as the second detailed requirement in the running movement standard 1; the third is "do not use the same hand and foot at the same time". When the angle formed by line segments 2-3 and 1-8, 5-6 and 1-8 is greater than a certain value, it indicates that the hands have been swung out. Analyze whether the two arms are in front of or behind the body at the same time.

[0115] The detailed requirement of action standard 2 is that "the toe of the rear foot should not extend beyond the arch of the front foot". The sliding window method is used to determine the video frame image corresponding to the maximum value of the angle formed by line segments 9-10 and 12-13 when both legs are at their most forward or rearward. For each complete step between two maximum values, the analysis is performed to see if the toe of the rear foot, i.e., point 19, extends beyond the arch of the front foot, i.e., the midpoint of line segment 22-24, or if point 22 extends beyond the midpoint of line segment 19-21.

[0116] The analysis method for the detailed requirements in Movement Standard 3 and Running Movement Standard 2 is the same.

[0117] The detailed requirement of action standard 4 is "uninterrupted, four consecutive horse stance runs". Similarly, the sliding window method is used to determine the video frame image corresponding to when the legs are at their front or back. The distance between the ankles, i.e., points 11 and 14, and the completion status are analyzed to determine whether it can be regarded as running a step that meets the detailed requirements. If so, it is accumulated and recorded; otherwise, the count is reset to zero.

[0118] Table 2:

[0119]

[0120] The action standard and corresponding detailed requirements for single-leg hopping are shown in Table 3. Among them, the detailed requirement of action standard 1 is that "the angle between the swinging leg and the supporting leg is not less than 30 degrees". The sliding window method is used to determine the video frame image corresponding to the maximum value when both legs swing to the front or back, that is, the angle formed by line segments 9-10 and 12-13. The maximum value is analyzed to see if the maximum value is ≥30 degrees.

[0121] The detailed requirement of action standard 2 is that "the toe of the swinging leg should not exceed the take-off leg". Similarly, the sliding window method is used to determine the video frame image corresponding to when both legs swing to the front or back. For each complete step between two times, the toe, i.e., point 19, is analyzed to see if it exceeds the arch of the take-off leg, i.e., the midpoint of line segment 22-24, or whether point 22 exceeds the midpoint of line segment 19-21.

[0122] Movement Standard 3 includes three detailed requirements: the first is analyzed using the same method as the first detailed requirement in Running Movement Standard 1; the second is analyzed using the same method as the second detailed requirement in Running Movement Standard 1; and the third is analyzed using the same method as the third detailed requirement in Horse Stance Running Movement Standard 1.

[0123] The detailed requirement of action standard 4 is "uninterrupted, four consecutive single-leg jumps". Similarly, the sliding window method is used to determine the video frame image corresponding to when both legs are at their front or back. The distance between the ankles, i.e., points 11 and 14, and the completion status are analyzed to determine whether it can be regarded as a step that meets the detailed requirements. If so, the count is accumulated; otherwise, the count is reset to zero.

[0124] Table 3:

[0125]

[0126] The action standards and corresponding detailed requirements for the running and jumping steps are shown in Table 4. Among them, action standard 1 includes three detailed requirements: The first is "one foot takes a step and jumps, then the other foot takes a step and jumps." The sliding window method is used to determine the video frame image corresponding to the maximum angle formed by line segments 9-10 and 12-13 when both legs are at their front or back. The analysis is conducted to see if there is a stepping action in the position of the leg in front of the body (line segment 9-10 or 12-13) relative to the torso (line segment 1-8). At the same time, the analysis is conducted to see if there is a continuous upward displacement of the ankle of the corresponding leg behind the body (point 11 or 14). The second is "alternating legs." Similarly, the sliding window method is used to determine the video frame image corresponding to the maximum or back position of both legs. For each running and jumping step, the analysis is conducted to see if the leg that is in front of the body is a different one.

[0127] Movement Standard 2 includes three detailed requirements: the first is analyzed using the same method as the first detailed requirement in Running Movement Standard 1; the second is analyzed using the same method as the second detailed requirement in Running Movement Standard 1; and the third is analyzed using the same method as the third detailed requirement in Horse Stance Running Movement Standard 1.

[0128] The detailed requirement of action standard 3 is that it must be uninterrupted and consist of four consecutive running and jumping steps. Similarly, the sliding window method is used to determine the video frame image corresponding to when both legs are at their front or back. The distance between the ankles, i.e., points 11 and 14, and the completion status are analyzed to determine whether it can be regarded as a step that meets the detailed requirements. If so, it is accumulated and recorded; otherwise, the number of steps is reset to zero.

[0129] Table 4:

[0130]

[0131] The standard and corresponding details of the side slide step are shown in Table 5. Among them, the standard 1 includes two details: the first is "sideways body", which means that the perspective in the video actually requires facing the camera direction, that is, analyzing whether the head organs are arranged in the order of "right ear (point 17) - right eye (point 15) - left eye (point 16) - left ear (point 18)" from left to right; the second is "the angle between the body and the marker line does not exceed 20 degrees", that is, calculating the angle between the right shoulder (line segment 1-2), the left shoulder (line segment 1-5) and the horizontal plane.

[0132] Action Standard 2 includes two detailed requirements: The first is "slide with the dominant foot, and quickly follow with the non-dominant foot, without jumping." The sliding window method is used to determine the video frame image corresponding to the maximum angle formed by line segments 9-10 and 12-13 when both legs are at their front or back. For each side slide, the analysis is performed to see if there are multiple frames of changes in the distance between the ankles of both legs, i.e., points 11 and 14, decreasing and then increasing, and to analyze whether there is continuous upward displacement in multiple frames. The second is analyzed using the same method as the detailed requirements in Running Action Standard 2.

[0133] The detailed requirement of action standard 3 is "four consecutive leftward side slides". Similarly, the sliding window method is used to determine the video frame image corresponding to when both legs are at their front or back. The distance between the ankles, i.e., points 11 and 14, and the completion status are analyzed to determine whether it can be regarded as running a leftward step that meets the detailed requirements. If so, it is accumulated and recorded; otherwise, the count is reset to zero.

[0134] The detailed requirement of action standard 4 is "four consecutive rightward side slides". Similar to action standard 3, it is determined whether it can be regarded as running a rightward step that meets the detailed requirement; if so, it is accumulated and recorded; otherwise, the count is reset to zero.

[0135] Table 5:

[0136]

[0137] The standard movements and corresponding detailed requirements for the standing long jump are shown in Table 6. Since each standard movement is aimed at a specific stage of the entire movement, in addition to analyzing the detailed requirements, it is also necessary to distinguish the various stages of the movement. The angles formed by the thigh and calf at the knee, i.e., the angles between line segments 9-10 and 10-11, or 12-13 and 13-14, are selected to divide the stages of the standing long jump.

[0138] First, when the angle between the thigh and calf decreases to less than a certain angle, it indicates that the body has entered the power preparation phase. In this phase, the first detail requirement in action standard 1, "knee flexion," is analyzed to determine the video frame image corresponding to the maximum degree of leg flexion, i.e., the angle between line segments 9-10 and 10-11, or 12-13 and 13-14, and to determine whether this angle is less than the specified value. Then, the second detail requirement in action standard 1, "the angle between the forearm and torso is not less than 20 degrees during the backswing," is analyzed to determine the video frame image corresponding to the maximum angle between the arms and torso, i.e., the angle between line segments 2-3 or 5-6 and 1-8, and to determine whether this angle is less than the specified value.

[0139] Secondly, when the angle between the thigh and calf increases to a certain value, the take-off phase begins. This phase analyzes the first detail requirement in action standard 3, "both feet leave the ground simultaneously," and determines whether the feet leave the ground simultaneously based on the changes in the distance between the soles of the feet (line segments 19-21 and 22-24) and the ground (the position before take-off).

[0140] Then, when the angle between the thigh and calf decreases again to less than a certain angle, the airborne phase begins. In this phase, the detailed requirements of "swinging forcefully forward and upward with the elbow not lower than the chin" in action standard 2 are analyzed. It is determined whether the relative position of the elbow (point 3 or 6) relative to the chin (point 1) is higher than the chin when it is at its highest, and whether the two forearms (line segments 3-4 or 6-7) are swinging upward at this time.

[0141] Next, when the angle between the thigh and calf increases again to a certain angle, the landing phase begins. This phase analyzes the second detail requirement in action standard 3, "both feet land simultaneously," and determines whether the feet leave the ground simultaneously based on the changes in the distance between the soles of the feet (line segments 19-21 and 22-24) and the ground (the landing point).

[0142] Finally, when the angle between the thigh and calf is detected to decrease to less than a certain angle for the third time, the buffering end phase begins. In this phase, the detailed requirement of "both arms swing naturally to the sides when landing" in action standard 4 is analyzed to determine whether the two hands, i.e., points 4 or 7, follow the swing to the position of the hip bone, i.e., points 9 or 12.

[0143] Table 6:

[0144]

[0145] The standard and corresponding detailed requirements for hitting a stationary ball with both hands are shown in Table 7. Since each standard refers to a specific stage of the entire action, in addition to analyzing the detailed requirements, it is also necessary to distinguish each stage of the action. The distance changes from the midpoint of both hands to the torso (i.e., the midpoint of points 4 and 7) to line segments 1-8 are selected to divide the stages of hitting a stationary ball with both hands.

[0146] Before conducting the analysis, we traversed the entire video and identified the two frames on the left and right sides of the body where the hands were furthest from the torso. The frame number from the beginning of the video to the smaller frame number represents the pre-hit preparation phase, the frame number from the smaller frame number to the larger frame number represents the swing and hit phase, and the frame number from the larger frame number to the end represents the recovery phase.

[0147] First, if entering the power preparation stage, analyze the detail requirement of "the hand on the same side as the direction of the shot is below" in action standard 1. Determine the direction of the shot based on the position of both hands relative to the torso in the frame with the larger frame number, and judge whether the hand on the same side of the direction meets the detail requirement of being below the other side. Analyze the detail requirement of "both shoulders facing the direction of the shot, with an angle of inclination not exceeding 20 degrees" in action standard 2. Analyze whether the organs of the head are arranged in the order of "right ear (point 17) - right eye (point 15) - left eye (point 16) - left ear (point 18)" from left to right, and calculate the angle between the right shoulder (line segment 1-2), the left shoulder (line segment 1-5) and the horizontal plane.

[0148] Secondly, if the swing phase begins, the first detail requirement in action standard 3, "leg push-off and hip rotation," is analyzed. The continuous change of the angle between the hip (line segment 9-12) and the horizontal plane in the direction of the hit is calculated. The detail requirement in action standard 4, "the non-dominant foot moves forward and the ball of the foot leaves its original position," is analyzed. The position of the sole of the foot on the side in the direction of the hit (line segment 19-21 or 22-24) is analyzed to see if there is a continuous change, and it is determined whether the corresponding position in the previous phase meets the detail requirement of being in front. The detail requirement in action standard 5, "the bat contacts the ball and the ball is hit," is analyzed to determine whether the object detection frame corresponding to the bat contacts the object detection frame corresponding to the ball, and whether the position of the object detection frame corresponding to the ball changes continuously.

[0149] Finally, if entering the finishing stage, analyze the second detail requirement in action standard 3, "the body faces the direction of the shot after hitting the ball," and determine whether the shoulder on the same side of that direction, i.e., point 2 or 5, meets the detail requirement of being in the upper position relative to the other side; for the third detail requirement in action standard 3, "the heel of the foot on the opposite side of the shot direction leaves the ground," calculate the angle between the sole of the foot on the opposite side of the shot direction and the horizontal and compare it with the corresponding angles in the previous two stages, and at the same time determine whether the toe on that side meets the detail requirement of being in the lower position relative to the heel.

[0150] Table 7:

[0151]

[0152] The standard and corresponding detailed requirements for the forehand toss and drop shot are shown in Table 8. Since each standard refers to a specific stage of the entire action, in addition to analyzing the detailed requirements, it is also necessary to distinguish each stage of the action. In this skill, the two hands, namely points 4 and 7, can be divided into the hitting hand and the tossing hand. The distance change from the hitting hand to the torso, i.e. to line segments 1-8, is selected to divide the stages of the forehand toss and drop shot.

[0153] Before conducting the analysis, we traversed the entire video to determine that the hand holding the ball and raising it at the beginning of the action was the ball-tossing hand, and the hand holding the racket and spreading it out was the hitting hand. We then identified the two frames on the left and right sides of the body where the hitting hand was furthest from the torso. From the beginning of the video to the frame with the smaller frame number, we observed the pre-hit power-building preparation phase. From the frame with the smaller frame number to the frame with the larger frame number, we observed the swing and hitting phase. From the frame with the larger frame number to the end of the video, we observed the follow-through and finishing phase.

[0154] First, if the player enters the power preparation stage, the detailed requirement in action standard 1, "there is a backswing motion, from the front of the body to the side of the body, and the ball can be hit directly without rebound," is analyzed. Since the specific trajectory of the ball after leaving the hand is not restricted, the position change of the hitting hand in this stage is directly detected to determine whether there is a movement of extending to the side of the body and gradually moving away from the torso.

[0155] Secondly, if the swing and hitting phase begins, the detailed requirement of "the non-dominant foot moves forward and the ball of the foot leaves its original position" in action standard 2 is analyzed, using the same method as the detailed requirement in action standard 4 for hitting a stationary ball with both hands. The second detailed requirement in action standard 3, "swinging and hitting the ball," is analyzed to determine whether the object detection frame corresponding to the racket is in contact with the object detection frame corresponding to the ball, and whether the position of the object detection frame corresponding to the ball changes continuously, while recording the position of the hitting point. The detailed requirement of "the vertical line between the farthest end of the racket and the ground exceeds the shoulder" in action standard 4 is analyzed to detect the positional change of the object detection frame corresponding to the racket and determine the relative positional relationship between its highest point and the shoulder area, i.e., points 2 and 5.

[0156] Finally, if the game enters the finishing stage, the first detail requirement in action standard 3, "the ball landing point does not exceed 30 degrees to the left or right of the front of the ball", is analyzed. If the ball lands within the video field of view, the landing point is determined according to the trajectory, and the angle between the line connecting the landing point and the hitting point and the horizontal plane is calculated. If the ball flies out of the field of view, the speed of the ball is calculated according to the position change within the field of view, and the possible landing point is predicted. Based on this, the relevant angle is calculated.

[0157] Table 8:

[0158]

[0159] The standard and corresponding details of the stationary one-handed ball-bouncing action are shown in Table 9. Among them, the details of action standard 1 are "not higher than the chest, only one hand is required, and the left and right hands can be used alternately to bounce the ball." Since the ball-bouncing hand is not fixed, the height of the two hands, i.e., points 4 and 7, is analyzed. If it is detected that at a certain moment, the height of one hand is higher than the other hand and is also higher than the positions of several frames in front and several frames behind, and is also higher than the other hand, then it is determined that the hand is the ball-bouncing hand for that moment. Since the hand and the ball will not separate immediately, the contact range is taken as the midpoint between the highest and lowest points in the current ball-bouncing process, and the positional relationship between it and the chest, i.e., the trisection of points 1 and 8, is determined.

[0160] The detailed requirement of action standard 2 is that the palm should not touch the ball; the angle between the back of the hand and the ball within the contact area needs to be calculated. Figure 6 The angle between line segments 0-9 and 9-12 is less than a specified value if the ball is touched by five fingers, and greater than a specified value if the ball is touched by the palm.

[0161] Action Standard 3 includes two detailed requirements: the first is "the pivot foot remains stationary," which detects the positional changes of the soles of the feet, i.e., line segments 19-21 and 22-24, throughout the entire process; the second is "dribble the ball 4 times consecutively," which detects whether both hands meet the condition of being the dribbling hands and complete multiple frames of continuous descent to a certain position throughout the entire process; if so, the count is accumulated; otherwise, the count is reset to zero.

[0162] Table 9:

[0163]

[0164] Table 10 shows the standard and corresponding details of the two-handed ball receiving action. Since each action standard is aimed at a specific stage of the whole action, in addition to the analysis of the details, it is also necessary to distinguish each stage of the action. The angle change between the upper arm (line segment 2-3 or 5-6) and the torso (line segment 1-8) is selected to divide the stages of the forehand ball toss.

[0165] First, if the angle between the arm and torso remains relatively constant for a certain number of consecutive frames, the pre-receive preparation phase begins. This phase analyzes the first detail requirement in Action Standard 1: "both hands placed in the area between the chest and abdomen in front of the chest." The angle between the upper arm and torso is recorded, and the height relationship between the hands (points 4 or 7) and the torso (points 1 and 8) at the trisection near point 1 is checked. The second detail requirement in Action Standard 1: "fingers naturally spread," is analyzed, calculating the extension of each finger. Figure 6 The angles formed by line segments 2-3 and 3-4, 5-6 and 6-8, 9-10 and 10-12, 13-14 and 14-16, 17-18 and 18-20 shown.

[0166] Secondly, when the angle between the upper and lower legs gradually increases, the ball-receiving phase begins. In this phase, the first detail requirement in action standard 2, "the angle between the upper arm and the torso is not less than 20 degrees," is analyzed, and the angle between the upper arm and the torso is recorded. The first detail requirement in action standard 2, "there is a ball-receiving action," is also analyzed, and the change in the angle between the upper arm and the torso is analyzed. The detail requirement in action standard 3, "receiving the ball and not using one hand or the body to receive the ball," is also analyzed, determining whether the object detection box corresponding to the ball remains within the human body area, and analyzing the positional relationship between the ball and the human body's hands and torso when it contacts the human body.

[0167] Table 10:

[0168]

[0169] The standard and corresponding details of the action of kicking a fixed ball are shown in Table 11. Since each action standard is aimed at a specific stage of the whole action, in addition to the analysis of the details, it is also necessary to distinguish each stage of the action. The position change of the midpoint of the middle part of the body, i.e., line segment 1-8, is selected to divide the stages of the forehand ball toss.

[0170] First, if the human body position remains basically unchanged for a certain number of consecutive frames, then the pre-kick preparation stage begins; no judgment is made in this stage, and the initial positions of the human body and the ball are recorded.

[0171] Secondly, if the displacement of the human body from the initial position is greater than a certain value and the ball is not within the range of the human body, then the pre-kick run-up phase begins. In this phase, the detailed requirement of "no small, shuffling steps and no interruptions in the run-up" in action standard 1 is analyzed. The sliding window method is used to determine the video frame image corresponding to the maximum value when the legs are at their foremost or last position, i.e., the angle formed by line segments 9-10 and 12-13. The stride length, i.e., the distance between the ankles, i.e., the distance between points 11 and 14, is determined each time, and the interval between the occurrence of the maximum value is calculated. The first detailed requirement of action standard 2, "the stride of the last step is greater than the previous step," is analyzed. The last step corresponds to the video frame image of the last time the legs are at their foremost or last position in this phase, and the step before the last step corresponds to the second to last step. The stride lengths of these two steps are compared. The second detailed requirement of action standard 2, "the angle between the swinging leg and the supporting leg is not less than 20 degrees" is analyzed. The angle formed by the two thighs, i.e., line segments 9-10 and 12-13, is calculated at the last step.

[0172] Finally, if the ball enters the body's range, the kicking phase begins. This phase requires identifying the kicking leg and the supporting leg. The foot that initially moves from one side of the ball to the other corresponds to the kicking leg, while the foot that remains on one side throughout the movement corresponds to the supporting leg. This phase analyzes the first detail requirement in action standard 3, "the supporting leg stands behind the ball." The example given is that if the left foot is supporting, the foot is in the third quadrant. Therefore, due to the perspective, if the supporting leg is the leg facing the camera, its sole must be lower than the bottom of the ball in the image; if the supporting leg is the leg facing the camera, its sole must be lower than the bottom of the ball in the image. Therefore, it is necessary to determine the positional relationship between the supporting leg's sole and the bottom of the ball. This relates to the action standard. Analyzing the second detail requirement in Standard 3, "the distance between the toe and the ball should not exceed one foot length," we calculate the distance between the leftmost or rightmost point of the ball on the side closest to the player's initial position and the toe of the supporting leg, as well as the length of the sole of the supporting leg. Analyzing the detail requirement in Standard 4, "do not touch the ball with the toes," if the contact point is within the 1 / 4 arc from the top of the ball to the leftmost or rightmost point on the side closest to the player's initial position, it can be considered a toe contact. If the contact point is within the 1 / 4 arc from the bottom of the ball to the leftmost or rightmost point on the side closest to the player's initial position, it can be considered an instep contact. Therefore, it is necessary to determine the specific point of contact, i.e., the point where the foot of the striking leg enters the initial range of the ball.

[0173] Table 11:

[0174]

[0175] The standard and corresponding details of the overhand throw are shown in Table 12. Since each standard refers to a specific stage of the entire action, in addition to analyzing the details, it is also necessary to distinguish each stage of the action. In this skill, the two hands, namely points 4 and 7, can be divided into the throwing hand and the free hand. The changes in the throwing hand's arm movements and the changes in its distance relative to the torso, namely the distance to line segments 1-8, are selected to divide the stages of the overhand throw.

[0176] Before analysis, the entire video is traversed to identify the hand with the most significant movement throughout the process as the pitching hand, and the other hand, which remains relatively still, as the free hand. At the beginning of the video, a preparation phase is initiated. If, for a certain number of consecutive frames, the upper arm (segment 2-3 or 5-6) gradually moves away from the torso until the angle between the upper arm and the torso exceeds a specified value, it indicates that the arm has swung up, and the pitching phase begins. Next, if it is detected that the pitching hand's elbow (point 3 or 6) is in front of the torso and the hand is in front of the elbow, and the angle between the upper arm and the torso exceeds another specified value, the end phase begins.

[0177] First, proceed to the pre-pitch preparation stage; no analysis is performed in this stage, but the initial orientation of the body and the initial position of the pitcher are recorded.

[0178] Secondly, we move to the pitching phase. This phase analyzes the first detail requirement in Action Standard 1, "naturally swinging back to the pitching position." This detail requires the pitcher to swing back to build power, meaning the upper arm moves to a position relative to the torso and free hand behind the body, rather than throwing directly in front. The position of the pitcher's arm relative to the torso needs to be determined during this process. Next, we analyze the second detail requirement in Action Standard 1, "drawing a circle from bottom to top, rather than swinging from top to behind." This detail indicates that the pitcher should swing back with their hand below elbows, then swing the ball forward from bottom to top, rather than directly swinging back with their hand above elbows and then throwing the ball forward. To determine the changes in the pitcher's arm movement during the process; analyze the first detail requirement in action standard 2, "body rotation," and check the distance between the shoulders, i.e., the length of line segment 2-5. If there is body rotation, the distance will increase due to the angle of view; analyze the second detail requirement in action standard 2, "the non-pitching side faces the wall," i.e., check which shoulder is closer to the body's initial orientation. If the non-pitching side faces the wall, then the shoulder on the free hand side is closer to the body's initial orientation than the shoulder on the pitcher side; analyze the detail requirement in action standard 3, "the non-dominant foot has forward displacement and the ball of the foot leaves its original position," using the same analysis method as the detail requirement in action standard 4 for hitting a stationary ball with both hands.

[0179] Finally, we move on to the finishing stage; this stage analyzes the detailed requirement in action standard 4 that "the hand follows through to the opposite hip bone, at least the fingertips touch," and determines that after the pitch, the fingertips of both hands should be... Figure 6 Does point 12 follow the swing to the vicinity of point 9 or 12 on the hip bone?

[0180] Table 12:

[0181]

[0182]

[0183] The standard and corresponding details of the low-hand throw are shown in Table 13. Since each standard refers to a specific stage of the entire action, in addition to analyzing the details, it is also necessary to distinguish the various stages of the action. In this skill, the two hands, i.e., points 4 and 7, can be divided into the throwing hand and the free hand. The movement changes of the throwing hand's arm and the distance changes of its relative to the torso, i.e., to line segments 1-8, are selected to divide the stages of the overhead throw. The division method is the same as that of the overhead throw, which is divided into the preparation stage, the throwing stage, and the finishing stage.

[0184] First, proceed to the pre-pitch preparation stage; no analysis is performed in this stage, but the initial orientation of the body and the initial position of the pitcher are recorded.

[0185] Next, we enter the pitching phase; this phase analyzes the first detail requirement in action standard 1, "swing naturally back to the pitching position," using the same method as the analysis of the detail requirement in action standard 4 for overhand pitching; it also analyzes the detail requirement in action standard 2, "the non-dominant foot moves forward and the ball of the foot leaves its original position," using the same method as the analysis of the detail requirement in action standard 4 for two-handed fixed ball hitting.

[0186] Finally, the final stage begins. This stage analyzes the detailed requirement in action standard 3 that "the ball must touch a wall at a distance of 4.5 meters," meaning the ball must travel at least 4.5 meters before landing. The stage detects changes in the position of the object detection frame corresponding to the ball. If the ball lands within the video field of view, the landing point is determined based on its trajectory, and the horizontal distance between the landing point and the initial position of the human body is calculated. If the ball flies out of the field of view, the ball's speed is calculated based on changes in its position within the field of view, and possible landing points are predicted, with the relevant distance calculated accordingly. The stage also analyzes the detailed requirement in action standard 4 that "the elbow should be used as the observation point, and the movement should reach chest height." This determines whether the thrower's elbow (points 3 or 6) follows through to the chest height (points 1 and 8, roughly corresponding to point 1).

[0187] Table 13:

[0188]

[0189] Example 3:

[0190] The aforementioned embodiment mainly describes the proposed adoption of the Test of Gross Motor Development-3 (TGMD-3). This scheme has been successfully applied internationally and has received unanimous praise from domestic and international peers. It includes 6 movement skills (running, horse stance running, single-leg hop, running jump, side slide, and standing long jump) and 7 operational skills (two-handed hitting of a stationary ball, forehand hitting of a tossed ball, two-handed catching of a ball, overhead throw, low throw, and standing one-handed dribbling), totaling 13 movements. For the movement standard corresponding to each skill, the human body key point information extracted by the key point extraction device is compared item by item with the movement standard specified in TGMD-3 to obtain the final evaluation result. Under certain circumstances, it can replace referees in evaluating sports events, not only freeing up human resources but also achieving the goal of fair and impartial scoring.

[0191] In this embodiment, each frame of video is processed using a lightweight OpenPose network to obtain key points of the human body, and the relevant information of the key points of the human body is recorded in a JSON file.

[0192] For attitude estimation networks, while increasing the number of layers enhances representation and modeling capabilities and leads to better performance, it also increases network complexity and redundancy, causing significant efficiency issues. Many attitude estimation methods proposed in recent years tend to use large-scale, deep, and complex network structures. The enormous number of parameters and computational demands place increasingly high demands on the computing power and memory of hardware platforms, making deployment difficult on edge devices with limited memory and computing resources, such as embedded platforms or mobile devices. Even when deployed on highly integrated, high-performance cloud environments and transmitted via network, the massive bandwidth consumption significantly increases industrial costs, and transmission latency prevents real-time feedback, hindering widespread practical application. Therefore, designing a fast, lightweight attitude estimation network that can efficiently run on localized, real-time devices with limited resources has become a pressing problem.

[0193] The following section primarily explains how to obtain a lightweight OpenPose network model. To lighten the pose estimation network through pruning to meet practical applications, this embodiment designs a fast and lightweight pose estimation network and proposes a multi-granularity adaptive network pruning method to obtain a lightweight OpenPose network model. Multi-granularity refers to the fusion of fine-grained unstructured pruning and coarse-grained structured pruning. Adaptive means that some hyperparameters, such as pruning threshold and pruning rate, do not need to be pre-set but can be obtained during the pruning process based on the actual situation. The innovation lies in the proposed unstructured pruning, which addresses the drawback of current methods requiring fixed thresholds and indirectly enhances the strength of structured pruning by increasing sparsity. The proposed structured pruning avoids the reliance on sparse matrices when using unstructured pruning alone and addresses the drawback of current methods requiring large amounts of data for auxiliary judgment.

[0194] The process of obtaining the lightweight OpenPose network is as follows: The structure of the convolutional kernels of the original OpenPose network is improved to reduce the number of parameters and adapt to subsequent pruning; the zeroing threshold of the convolutional kernels is determined based on the mean and variance of the weight parameters, and weights below the zeroing threshold are zeroed to sparsify the weight matrix; the information richness and information independence of the convolutional kernels of each filter are calculated, and weighted fusion is performed based on the information richness and information independence; the importance score of each filter is determined based on the weighted fusion result, and the pruning rate of each layer is determined based on the importance score; the corresponding filters are removed based on the pruning rate scores to obtain the lightweight OpenPose network.

[0195] The overall process diagram of pruning is as follows: Figure 7As shown, the target object is the already trained OpenPose network model. After pruning and retraining, a lightweight OpenPose network model is obtained. The whole process can be divided into two stages: unstructured pruning and structured pruning.

[0196] The first stage performs fine-grained pruning at the level of convolutional kernel weights, which is the most basic computational unit in the forward and backward propagation of the network. The weight matrix of the convolutional kernels in each layer of the network is statistically analyzed. An adaptive pruning threshold is calculated by using the mean and variance of the weight matrix, which measure the distribution of the overall data. Convolutional kernel weights below the threshold are set to zero, thus sparsifying the convolutional kernels.

[0197] The second stage performs coarse-grained pruning on the feature dimensions corresponding to the amount of information carried (i.e., the input channel level). Information richness and information independence are calculated for each filter combination in each layer of the network, thereby calculating the importance score of each filter and the pruning rate of each layer. Filters with low pruning rates are removed, and pruning is performed layer by layer. Finally, the pose estimation ability of the network is restored through retraining.

[0198] The first stage, the unstructured pruning process, is as follows:

[0199] Generally, convolutional layers use filters Input tensor Convert to output tensor Where C is the number of input feature channels, N is the number of filters, and also the number of output channels, K h and K w These are the height and width of the filter, respectively. The convolution operation can be represented as shown in equation (1):

[0200]

[0201] Where * represents convolution operation, X c and Y b This represents the input features having C and N channels respectively. Figure X and output feature map Y. Commonly used structured pruning methods apply to all weight parameters in the target layer l. Set an artificial threshold based on experience or actual situation. For example, when using this method to iteratively prune the first layer of the AlexNet network, the threshold was finally set to 0.006 based on experiments, that is, to set all weights with an absolute value lower than 0.006 to zero during pruning.

[0202] The unstructured pruning method proposed in this embodiment follows the previous approach, resetting weights in convolutional kernels whose absolute values ​​are significantly smaller than normal to zero. However, to adaptively prune according to the characteristics of each layer, the filter structure is first adjusted. Specifically, the OpenPose pose estimation network extensively uses 7×7 convolutional kernels during the optimization phase. This results in many parameters, few layers, and weak expressive power. Therefore, one approach is to replace each large 7×7 convolutional kernel with three smaller 3×3 convolutional kernels. This maintains the receptive field and increases the number of network layers. The activation functions between layers increase the nonlinear expressive power of the network. Residual connections are added here to avoid gradient vanishing, allowing gradient backpropagation to shallower layers. Finally, and most importantly, the number of parameters is reduced. The original number of parameters was 7×7 = 49, while the current number is 3×3×3 = 27, a reduction of about half.

[0203] On the other hand, introducing depthwise separable convolution in the feature extraction stage achieves the same effect in two steps compared to standard convolution which changes both the feature map size and the number of channels simultaneously: first, using depthwise convolution to split the convolution kernel into single-channel forms, and then performing convolution on each channel without changing the input depth to obtain an output with the same number of input channels; then, using pointwise convolution, which is the normal 1×1 convolution, to modify the number of channels and obtain the desired number of channels for the output. This can greatly reduce the number of parameters in the network model.

[0204] Next, unstructured pruning is performed. Instead of assigning fixed weight thresholds to each layer, the mean and variance of the statistical indicators representing the overall discrete sample distribution are used to determine the threshold. Specifically, the threshold T... i The set of all weight parameters of the i-th layer of the network The weighted sum of the mean and variance is used to determine the value, as shown in formula (2).

[0205]

[0206] Here, L represents the total number of layers in the network, and α represents a negative sensitivity coefficient, which is set to -0.2 in this paper. When α = 0, only the mean is considered, but there may be data distributions with the same mean but different variances, so variance needs to be considered. According to experiments with CLR-RNF, setting the threshold to the mean or higher will lead to irrecoverable accuracy loss, so a negative sensitivity coefficient is needed to appropriately reduce it. If α is too small, the threshold will be too small, resulting in a low pruning rate for low-variance distributions with relatively similar data distributions; if α is too large, the threshold will be too large, resulting in excessive pruning for high-variance distributions that may have many small weights. Finally, based on experiments, α was determined to be -0.2, which can uniformly sparsify networks with different weight distributions.

[0207] Based on threshold T i For convolution kernel Each weight parameter in The update strategy is shown in formula (3).

[0208]

[0209] By adaptively determining the pruning threshold based on the relevant statistics of each convolutional kernel parameter, the need for repeatedly iterating and setting a fixed threshold, and the inability to provide different references for different convolutional kernels, can be avoided. Furthermore, since unstructured pruning of the convolutional kernel increases the proportion of zero parameters, improving the kernel's sparsity, and the subsequent structured pruning method needs to utilize this sparsity, unstructured pruning can be performed before structured pruning. This is incorporated into practical applications as a new way to enhance the strength of subsequent structured pruning, achieving the goal of promoting network lightweighting without designing special sparse matrix storage or implementing special sparse convolution calculations.

[0210] The second stage, the structured pruning process, is as follows:

[0211] In structured pruning, the importance of corresponding filters is typically determined based on the importance of the input feature map to decide whether to prune. To determine the importance of filters in the network, related work in the quantization experiments of deep visual representations maps the corresponding feature maps to the resolution of the input image using bilinear interpolation. Then, a fixed activation threshold is set, and the input image is binarized according to the threshold, which can be viewed as a binary mask of the input image. The portion above the threshold represents the range of information contained in the feature map, i.e., the receptive field range of each pixel on the feature map within the input image. Similarly, in experiments on localization and classification capabilities in convolutional neural networks, the corresponding feature maps are superimposed along the channel dimension and upsampled to map to the resolution of the input image, resulting in a heatmap of the input image. This heatmap reflects the distribution of information, and each pixel value on the heatmap represents the importance of that pixel in the input image.

[0212] The structured pruning method proposed in this embodiment follows the previous method, which also removes filters with relatively low importance in each layer. However, it avoids the disadvantage of relying on a large number of feature maps for auxiliary judgment. It calculates the information richness and information independence of the filter using single-channel information and cross-channel information respectively, and calculates the joint importance score of the filter by weighted fusion. Thus, it can perform structured pruning directly based on the characteristics of the filter itself without a large amount of data.

[0213] (1) The information richness of the filter is obtained as follows:

[0214] In structured pruning, this invention uses entropy to measure the information richness of filters. In information theory, entropy is a general and effective measure of information, representing the uncertainty of a system, which reflects the amount of information contained in the system. In deep learning, Rényi's alpha entropy is an effective measure of entropy. Given a mini-batch of size s, the feature map generated by the i-th filter in the j-th layer is... Simplified representation: M = {m 1 ,m 2 ,…,m s Here, m can be considered a random variable. According to the definition of Renyi entropy, the entropy of m can be expressed by a normalized Gram matrix. The eigenspectrum is represented as shown in formula (4), where K m (i,j)=σ(m i ,m j ), and σ is a Gaussian kernel.

[0215]

[0216] Where, N m It is the normalized K m N m =K m / tr(K m ). λ i (N m ) represents N m The m-th feature value, α has a default value of 1.01. Since the calculation of entropy in deep neural networks requires the marginal probability of high-dimensional variables, it is difficult to directly calculate the entropy of the filter according to formula (4) and then obtain the corresponding information richness. However, based on the definition of entropy, the uncertainty of the filter can be used to measure its own information richness. This invention uses the distribution of the convolution kernels contained in the filter to represent the uncertainty of the filter. The more complex the distribution of the convolution kernels, the greater the uncertainty of the filter, and the more diverse the features that the filter can extract, the more it should be retained in pruning. Since the essence of entropy is uncertainty, the entropy of the i-th filter in the j-th layer can be similarly defined as:

[0217]

[0218]

[0219]

[0220] Where, sim qp represents the sum of Euclidean distances between the q-th convolutional kernel and all other convolutional kernels in the filter. Euclidean distance indicates the similarity between two convolutional kernels. q This represents the probability distribution of the q-th convolutional kernel calculated using the softmax function. The higher the probability, the less similar the convolutional kernel is to other convolutional kernels.

[0221] To simplify the calculation, we can simply calculate the sum of similarities between each convolutional kernel and its M nearest neighbors to measure the distribution around that kernel.

[0222]

[0223] The smaller the entropy of the filter defined by formula (5), the greater the amount of information contained in the corresponding feature map. Therefore, we can define the information richness index of the i-th filter in the j-th layer as formula (9):

[0224] IM r (F i, )=1-H f (F i, (9)

[0225] (2) The method for obtaining the information independence of the filter is as follows:

[0226] To avoid losing certain special information that is difficult to relearn, this embodiment uses information independence to measure the substitutability of the information contained in the filter. The higher the information independence of a filter, the lower its substitutability, and the more difficult it is for other filters to relearn it during retraining; therefore, it should be retained. Thus, similar to information richness, information independence can also be calculated using the characteristics of the filter itself. It can be considered that when a filter is very similar to its surrounding filters, the resulting feature maps are also very similar, and the corresponding information substitutability is stronger. Similarly, we can use Euclidean distance as a similarity metric. Therefore, we define the information independence of the i-th filter in the j-th layer as:

[0227]

[0228] Finally, we use both information richness and information independence to measure the importance of filters. Therefore, the importance score of the i-th filter in the j-th layer can be weighted as follows:

[0229]

[0230] Here, σ is a weighting parameter used to control the influence of information richness and information independence on the importance of the filter. According to experiments, a value of 0.8 can achieve the best pruning effect.

[0231] After calculating the importance score of the filters in each layer, a pruning rate is given for each layer. n represents the total number of filters in this layer. The filters can be sorted by score, and finally, filters with lower pruning rates are removed, retaining the remaining important filters. This invention prunes all layers at once; the pruned filters are removed from the network structure, and then... -2 The learning rate was used to fine-tune and retrain the pruned network for 400 rounds. It's important to note that this invention treats fully connected layers as convolutional layers, with each individual neuron acting as a filter with a kernel size of 1 and a count of 1, thus allowing for pruning of fully connected layers just like convolutional layers.

[0232] Applying the multi-granularity adaptive pruning method proposed in this embodiment to OpenPose yields a fast and lightweight pose estimation network. The overall network structure of OpenPose consists of a backbone network and a refinement network. The backbone networks before and after lightweighting are shown below. Figure 8 and Figure 9 As shown, the refined networks before and after lightweighting are respectively as follows: Figure 10 and Figure 11 As shown.

[0233] Table 14 compares the number of parameters, model size, computational cost, and inference speed before and after network lightweighting. The data in the table includes a comparison of the number of parameters, model size, computational cost, and inference speed before and after network pruning and network structure improvement, which can prove the effectiveness of the fast and lightweight attitude estimation network designed in this invention.

[0234] Table 14:

[0235]

[0236] Example 4:

[0237] This invention provides a scoring device for basic human motor skills, using the human basic motor skills described in Example 1.

[0238] The standard method for judging this motor skill is as follows: Figure 12 As shown, the device includes a human posture prediction unit, a movement standardization judgment unit, and a motor skill scoring unit, wherein:

[0239] The human pose prediction unit includes a key point extraction unit and a tracking processing unit. The key point extraction unit is used to perform human body recognition in the video and record the relevant information of the key points of the human body in each frame in the corresponding JSON file. The tracking processing unit is used to read the JSON file in pairs according to the video frame number and determine the person with the smallest change in human body indicator features in two adjacent frames as the same person and assign the same number.

[0240] The motion standardization judgment unit includes a target object recognition unit, a motion detail analysis unit, and a process synchronization unit. The target object recognition unit is used to identify the target object corresponding to the specified motion skill of the target human body and record the position information of the target object. The motion detail analysis unit is used to process the relevant information of the key points of the target human body and the position information of the target object in each frame to obtain motion detail information. The process synchronization unit is used to connect the identified key points with line segments in a specified order and synchronously display the body skeleton of the target human body, the hand skeleton of the target human body, the number of the target human body, the position coordinates of the target human body, the name of the target object, and the position coordinates of the target object in each frame.

[0241] The motor skill scoring unit is used to calculate scores according to the completion of detailed requirements and corresponding rules in order to display the evaluation results of the corresponding action standards. After processing and analyzing all frames of the video, the evaluation results of the target human's motor skills corresponding to each action standard are obtained and added to the end of the set of all frame images in the video. The results are then saved as a complete analysis and scoring video in the results folder under the original video file path.

[0242] In embodiments of the present invention, such as Figure 13 As shown, the human pose prediction unit also includes a video preprocessing unit, which uses a non-local averaging denoising method to smooth each frame of the image.

[0243] like Figure 14 As shown, the motion detail analysis unit includes parallel mobile skill assessment units and ball skill assessment units. The mobile skill assessment units include assessment units for running, horse stance running, single-leg hopping, running jump, side slide, and standing long jump. The ball skill assessment units include assessment units for hitting a stationary ball with both hands, hitting a tossed ball with a forehand, dribbling a ball with one hand in place, catching a ball with both hands, kicking a stationary ball, overhead throw, and low toss.

[0244] In embodiments of the present invention, such as Figure 15 As shown, another scoring device for the aforementioned basic human motor skills is also provided, including one or more processors 41 and a memory 42; wherein, Figure 15 Take a processor 41 as an example.

[0245] Processor 41 and memory 42 can be connected via a bus or other means. Figure 15 Taking the example of a connection between China and Israel via a bus.

[0246] The memory 42 is a non-volatile computer-readable storage medium that stores instructions executable by the at least one processor. These instructions can be understood as being used to store non-volatile software programs and non-volatile computer-executable programs. After being executed by the processor, the instructions are used to complete the standard judgment method for basic human motor skills as described in Embodiment 1. The processor 41 executes the standard judgment method for basic human motor skills by running the non-volatile software programs and instructions stored in the memory 42.

[0247] The memory 42 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other non-volatile solid-state storage device; in some embodiments, the memory 42 may optionally include memory remotely located relative to the processor 41, and these remote memories may be connected to the processor 41 via a network; examples of such networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.

[0248] The program instructions / modules are stored in the memory 42. When executed by one or more processors 41, they perform the standard judgment method for basic human motor skills in Embodiment 1.

[0249] Those skilled in the art will understand that all or part of the steps in the various methods of the embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk or optical disk, etc.

[0250] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A standard method for judging basic human motor skills, characterized in that, include: If a human body is detected in the video, a JSON file will be generated for each frame. The JSON file is used to record relevant information about the key points of the human body. The person with the smallest change in human body key features between two adjacent frames is identified as the same person and assigned the same number. This includes: not using OpenPose's tracking module, using the coordinates of the neck area as the judgment criterion, assuming that a people are identified in the current frame and b people are identified in the next frame, after obtaining all human body key points in the current and next frames from the output JSON file, iterating through and calculating the Euclidean distance between the pairs of abdominal coordinate points between these two frames, resulting in a distance list dist of length a×b; for the coordinates of the m-th person in the people list in the current frame and the coordinates of the n-th person in the next frame, the format in the generated distance list dist is [ For the pair [m, n], perform bubble sort on the values ​​of m and n, starting from the minimum value and recording the corresponding m or n values ​​sequentially. If a duplicate is encountered, compare the distance values ​​and delete the larger one. If no duplicate is encountered, move the pointer one position to the left and continue recording. If a distance is encountered... In the same special case, the corresponding human body postures of the mn pairs are analyzed. Pairs of human body postures that are the same between two frames are considered to be the same person. If multiple pairs of postures are also the same, the confidence changes of each part of the human body corresponding to the mn pairs are judged. Pairs of people with smaller average confidence changes between two frames are considered to be the same person. When the number of people in the next frame is equal to or less than the number of people in the current frame, the nth position coordinate set of the people list in the next frame is adjusted to the mth position according to the mn pairs. When the number of people in the next frame is more than the number of people in the current frame, in addition to adjusting the position, the unadjusted coordinate sets are randomly assigned empty positions. If a target object corresponding to the specified movement skill of the target human body is identified in the video, the position information of the target object is recorded. By processing the relevant information of the key points of the target human body and the position information of the target object in each frame, motion detail information is obtained; each motion standard specified by TGMD-3 is broken down into different detail requirements; the motion detail information is compared and analyzed with the corresponding detail requirements, and the preliminary evaluation results of the corresponding detail requirements are synchronized to the motion display on each frame image; after completing the processing and analysis of all frames of the video, the final evaluation result of the target human body's motion skill corresponding to each of the detailed requirements is obtained, and the evaluation result of the motion standard is obtained by AND operation of the final evaluation result of the corresponding detail requirement.

2. The standard method for judging basic human motor skills according to claim 1, characterized in that, Each frame of video is processed using a lightweight OpenPose network model to obtain key points of the human body, and the relevant information of the key points is recorded in a JSON file.

3. The standard method for judging basic human motor skills according to claim 2, characterized in that, The process of acquiring the lightweight OpenPose network is as follows: The structure of the convolution kernels in the original OpenPose network is improved to reduce the number of parameters and adapt to subsequent pruning. The zeroing threshold of the convolution kernel is determined based on the mean and variance of the weight parameters. Weights below the zeroing threshold are set to zero to make the weight matrix sparse. Calculate the information richness and information independence of the convolution kernel for each filter, perform weighted fusion based on information richness and information independence, determine the importance score of each filter based on the weighted fusion result, and determine the pruning rate of each layer based on the importance score; The corresponding filters are removed based on the pruning rate score to obtain a lightweight OpenPose network.

4. The standard method for judging basic human motor skills according to claim 1, characterized in that, The processing of the relevant information of the key points of the target human body and the position information of the target object in each frame specifically involves: The key points identified in each frame are connected by line segments in a specified order to obtain the body skeleton and hand skeleton of the target human body, thereby calculating the angles formed by the limbs of the target human body. Based on the relevant information of the key points of the target human body identified, the distance between the joints of the target human body and the relative position between the joints are obtained; Based on the position information of the target object, the distance between the joints of the target human body and the target object, as well as the relative position between the joints of the target human body and the target object, are calculated.

5. The standard method for judging basic human motor skills according to claim 1, characterized in that, The action details include one or more of the following: the angles formed by the limbs of the target human body, the distances between the joints of the target human body, the relative positions between the joints of the target human body, the distances between the joints of the target human body and the target object, and the relative positions between the joints of the target human body and the target object.

6. The standard method for judging basic human motor skills according to claim 1, characterized in that, The detailed requirements are divided into global requirements and local requirements. The global requirements are determined in each frame, and the local requirements are determined when the corresponding conditions are met. When the final evaluation results of all global and local requirements corresponding to the same action standard are 1, the evaluation result of the corresponding action standard is 1; otherwise, the evaluation result of the corresponding action standard is 0.

7. The standard method for judging basic human motor skills according to claim 6, characterized in that, The global requirement is determined in each frame, specifically as follows: In each frame, if the action detail information meets the global requirement, the preliminary evaluation result of the corresponding global requirement is displayed as 1 on each frame image; otherwise, the preliminary evaluation result of the corresponding global requirement is displayed as 0. After processing and analyzing all frames of the video, if the proportion of frames with a preliminary evaluation result of 1 for the corresponding global requirement exceeds a first preset value, the final evaluation result for the corresponding global requirement is displayed as 1; otherwise, the final evaluation result for the corresponding global requirement is displayed as 0.

8. The standard method for judging basic human motor skills according to claim 6, characterized in that, The local requirement is determined when the corresponding condition is met, specifically as follows: In each frame, if the action detail information meets the local requirements, the preliminary evaluation result of the corresponding local requirements is displayed as 1 on the frame image after the corresponding conditions are completed; otherwise, the preliminary evaluation result of the corresponding local requirements is displayed as 0. After processing and analyzing all frames of the video, if the number of times the preliminary evaluation result of the corresponding local requirement is displayed as 1 exceeds the proportion of the total number of times the corresponding condition is completed, the final evaluation result of the corresponding local requirement is displayed as 1; otherwise, the final evaluation result of the corresponding local requirement is displayed as 0.

9. A scoring device for basic human motor skills, characterized in that, Using the standard judgment method for basic human motor skills as described in any one of claims 1-8, the device includes a human posture prediction unit, a movement standardization judgment unit, and a motor skill scoring unit, wherein: The human pose prediction unit includes a key point extraction unit and a tracking processing unit; the tracking processing unit is used to identify the person whose human body indicator features change the least in two adjacent frames as the same person and assign them the same number. The motion standard judgment unit includes a target object recognition unit, a motion detail analysis unit, and a process synchronization unit; the motion detail analysis unit is used to process the relevant information of the key points of the target human body and the position information of the target object in each frame to obtain motion detail information; The motor skill scoring unit is used to compare and analyze the action details in each frame with the corresponding detail requirements to obtain a preliminary evaluation result of the corresponding detail requirements; after processing and analyzing all frames of the video, the final evaluation result of the target human's motor skills corresponding to each of the detail requirements is obtained; the evaluation result of the action standard is obtained by performing an AND operation on the final evaluation result of the detail requirements.

10. A scoring device for basic human motor skills, characterized in that, It includes at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which, when executed by the processor, are used to perform the standard judgment method for basic human motor skills as described in any one of claims 1-8.