Motion evaluation method, device, electronic device, and storage medium

By analyzing video images of subjects for identity recognition and motion monitoring, and using a state machine to evaluate motion compliance, this system solves the problems of low efficiency and high resource consumption in existing motion assessment systems for multi-person testing. It achieves efficient and real-time motion monitoring and feedback, and improves the accuracy of user motion.

CN119097891BActive Publication Date: 2025-12-26IFLYTEK CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411177227.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-26
Publication Date
2025-12-26
Estimated Expiration
2044-08-26

AI Technical Summary

Technical Problem

Existing motion assessment systems are inefficient, resource-intensive, and lack accurate monitoring of movement standards when testing multiple people. They also fail to provide timely feedback on violations and rely on expensive and complex sensor devices.

Method used

By analyzing video images of subjects for identity recognition and readiness assessment, continuously monitoring key points and motion attributes of the human body, using a state machine for motion compliance assessment, and outputting evaluation results in real time, the system reduces traditional hardware resources and uses computer vision technology for multi-person pose estimation and motion recognition.

Benefits of technology

It improves evaluation efficiency and user action accuracy, reduces resource consumption, and enables efficient, real-time action monitoring and interactive feedback for multi-person testing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119097891B_ABST
    Figure CN119097891B_ABST
Patent Text Reader

Abstract

The application provides a motion evaluation method and device, electronic equipment and storage medium, belonging to the motion evaluation technical field, the method comprises the following steps: analyzing the video image of the subject entering the test area to realize identity recognition, and judging the preparation state according to the preset evaluation mode to start the evaluation process; in the evaluation process, the video image of the subject is continuously analyzed, the human body key points and motion attributes are extracted, and they are input into the preset state machine to evaluate the motion compliance; according to the result of motion compliance evaluation, the evaluation result of the subject is output. The application optimizes the use of resources, strengthens the motion monitoring and provides real-time feedback, effectively solves the existing problems of the existing system, and improves the overall efficiency of the evaluation and the accuracy of the user motion.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of motion evaluation, in particular to a motion evaluation method and device, electronic equipment and storage medium. BACKGROUND

[0002] With the enhancement of public health awareness, intelligent motion evaluation systems have been widely used in education and daily activities, aiming to provide more efficient, fair and timely motion data collection and performance evaluation. Current motion evaluation systems usually use manual counting, sensor counting and vision-based counting technology to record and evaluate individual motion performance.

[0003] However, current motion evaluation systems are mostly single-person mode, which is inefficient when dealing with multiple-person tests, increasing the demand for manpower and equipment. Manual counting is affected by subjective judgment, and sensor counting is costly and complex to operate due to the use of specialized equipment. Although there are multi-person counting solutions based on vision, there are still improvements to be made in key point detection, human recognition, pose estimation and counting logic. For example, the CN115482580A solution lacks counting function, the CN116957868A solution relies on image similarity rather than traditional counting logic, the CN118230244A solution counts by analyzing the change of the included angle, and the CN117643718A solution does not explicitly describe the implementation of its counting logic. SUMMARY

[0004] The present application provides a motion evaluation method, device, electronic equipment and storage medium, aiming to solve the problems of existing motion evaluation systems in at least one of resource consumption, motion standard monitoring and real-time feedback of violations, to improve evaluation efficiency and accuracy of user motion.

[0005] In a first aspect, the present application provides a motion evaluation method, comprising:

[0006] analyzing video images of a subject entering a test area to realize identity recognition, and judging the readiness state according to a preset evaluation mode to start the evaluation process;

[0007] In the evaluation process, the video images of the subject are continuously analyzed, the human key points and motion attributes are extracted and input into a preset state machine for motion compliance evaluation;

[0008] According to the result of motion compliance evaluation, the evaluation result of the subject is output.

[0009] In a second aspect, the present application also provides a motion evaluation device, comprising:

[0010] An identity detection module is configured to analyze video images of a subject entering a test area to achieve identity recognition, and determine a preparation state of the subject according to a preset evaluation mode to start an evaluation process.

[0011] A compliance evaluation module is configured to continuously analyze video images of the subject in the evaluation process, extract human key points and action attributes, and input the human key points and the action attributes into a preset state machine to perform action compliance evaluation.

[0012] An achievement output module is configured to output an evaluation achievement of the subject according to a result of the action compliance evaluation.

[0013] In a third aspect, the present application provides an electronic device, which includes a memory, a processor, and a computer program stored in the memory and running on the processor, and the processor implements the steps of the motion evaluation method according to any one of the first aspect when executing the computer program.

[0014] In a fourth aspect, the present application provides a non-transitory computer-readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the steps of the motion evaluation method according to any one of the first aspect.

[0015] The motion evaluation method, device, electronic device, and storage medium provided by the present application aim to solve the problems of the existing motion evaluation system in resource consumption, action standard monitoring, and real-time feedback of violations. The method achieves identity recognition by analyzing video images of a subject entering a test area, reduces the large amount of hardware resources required in the traditional motion evaluation system, such as sensors and special equipment, thereby reducing resource consumption. In the evaluation process, the video images of the subject are continuously analyzed, human key points and action attributes are extracted, and the human key points and the action attributes are input into a preset state machine, which can monitor the actions of the subject in real time, ensure that the subject meets the preset standard, and improve the accuracy and real-time performance of action monitoring. Action compliance evaluation is performed according to the changes of the human key points, and the evaluation achievement of the subject is output according to the evaluation result. This real-time feedback mechanism enables the subject to immediately understand whether his / her action is compliant, thereby timely adjusting and improving the interactivity of the evaluation and the accuracy of the user action.

[0016] Therefore, the present application optimizes resource use, strengthens action monitoring, and provides real-time feedback, effectively solving the problems of the existing system and improving the overall efficiency of the evaluation and the accuracy of the user action. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the application or the prior art, the following will briefly introduce the drawings needed to be used in the embodiments or prior art description. Obviously, the drawings in the following description are some embodiments of the application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0018] Figure 1 is a flowchart of the motion evaluation method provided by the application;

[0019] Figure 2 is a schematic diagram of the site arrangement provided by the embodiment of the application;

[0020] Figure 3 is a schematic diagram of the human key points provided by the embodiment of the application;

[0021] Figure 4 is a schematic diagram of the human posture estimation model provided by the embodiment of the application;

[0022] Figure 5 is a schematic diagram of the evaluation process provided by the embodiment of the application;

[0023] Figure 6 is a schematic diagram of the counting logic provided by the embodiment of the application;

[0024] Figure 7 is a structural schematic diagram of the motion evaluation device provided by the application;

[0025] Figure 8 is a structural schematic diagram of the electronic device provided by the embodiment of the application. DETAILED DESCRIPTION

[0026] In order to make the purpose, technical solutions and advantages of the application clearer, the technical solutions in the application will be described clearly and completely in the following with reference to the drawings in the application. Obviously, the described embodiments are some embodiments of the application, not all the embodiments. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative effort belong to the protection scope of the application.

[0027] The terms "first", "second", etc. in the specification and claims of the application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0028] To solve the problems of existing systems in resource consumption, action standard monitoring, and real-time feedback of violations, the present application provides a sports evaluation method, device, electronic equipment and storage medium. The method realizes identity recognition by analyzing video images of the subject entering the test area, reduces the large amount of hardware resources required by traditional sports evaluation systems such as sensors and special equipment, thereby effectively reducing resource consumption. In the evaluation process, the system continuously analyzes the video images of the subject, accurately extracts the key points of the human body and action attributes, and inputs them into the preset state machine. This process can monitor the subject's actions in real time, ensure that they meet the preset standards, and improve the accuracy and real-time performance of action monitoring. According to the changes in the key points of the human body, the action compliance is evaluated, and the evaluation results of the subject are output in real time according to the evaluation results. This real-time feedback mechanism enables the subject to immediately understand whether their actions are compliant, thereby adjusting in a timely manner and enhancing the interactivity of the evaluation and the accuracy of the user's actions.

[0029] Therefore, the present application not only optimizes resource usage, strengthens action monitoring, and provides real-time feedback, but is also applicable to the evaluation scene of multiple sit-ups, and can effectively improve the overall efficiency of the evaluation and the accuracy of the user's actions.

[0030] The following will be described in conjunction with Figures 1-8 A sports evaluation method, device, electronic equipment and storage medium provided by the present application are described.

[0031] Please refer to Figure 1 , Figure 1 is a flowchart of the sports evaluation method provided by the present application. A sports evaluation method includes the following steps:

[0032] S110, analyze the video images of the subject entering the test area to realize identity recognition, and judge the readiness state of the subject according to the preset evaluation mode to start the evaluation process.

[0033] Specifically, by analyzing the video images of the subject entering the test area, the system can identify the identity of the subject. This involves using computer vision technology, such as face recognition or human pose estimation, to determine the individual entering the test area. After identifying the subject, the system will judge whether the subject is ready to start the evaluation according to the preset evaluation mode. For example, check whether the subject has stood in the correct position, whether the subject has taken the correct starting posture, or whether the subject has completed other necessary preparations. Once the system confirms that the subject is ready, it will start the evaluation process. This means that the system will start collecting data, monitoring the subject's actions, and evaluating according to the predetermined standards.

[0034] S120, in the evaluation process, continuously analyze the video images of the subject, extract human key points and action attributes, and input them into a preset state machine for action compliance evaluation.

[0035] Specifically, in the evaluation process, the system continuously analyzes the video images and action attributes of the subject, monitors their actions in real time, and ensures that every detail in the evaluation process is captured. The system accurately extracts human key point data, including joint positions, limb angles, etc., which are key criteria for evaluating the accuracy of the subject's actions. The extracted human key points are then input into a preset state machine.

[0036] The state machine analyzes the input human key points and action attributes, and determines whether the subject's actions meet the requirements according to the preset action standards and rules. It checks whether the human key points and action attributes meet specific conditions such as angles, positions, etc. at each state transition, thereby evaluating the compliance of the action. The state machine can determine whether the subject has met the evaluation standard by comprehensively evaluating the accuracy of the subject's actions based on changes in human key points. In this way, the state machine can ensure the accuracy and fairness of the evaluation results.

[0037] S130, according to the results of the action compliance evaluation, output the evaluation results of the subject.

[0038] Specifically, in the previous step (i.e. S120), the system has used the state machine to evaluate the compliance of the subject's actions by analyzing the video images of the subject and extracting human key points. This includes determining whether the subject's actions meet the predetermined standards or rules. After completing the action compliance evaluation, the system will generate the evaluation results of the subject according to the evaluation results. The evaluation results will reflect the performance of the subject in the evaluation process, including the accuracy of the action, etc. For example, the evaluation results can be presented in the form of numerical scores, grades, charts or other forms, so that the subject or the evaluator can intuitively understand the performance of the subject.

[0039] The above steps S110 to S130 are described in detail in the following embodiments.

[0040] Please refer to Figure 2 , Figure 2 is a schematic diagram of the arrangement of the site provided by the embodiments of the present application. The present application uses video acquisition, and specifically can use a single RGB camera to directly acquire the video of multiple subjects in the picture. There should be no occlusion between the subjects to ensure that the actions of each subject can be clearly captured. Since the sit-up evaluation needs to focus on the sit-up action and whether the legs are bent, each subject is required to face the side of the camera. This setting is to ensure that the camera can clearly capture the abdominal and leg movements of the subject.

[0041] The number of testees can be randomly set within the camera shooting area without blocking the subjects. The system supports setting each test area on the visual panel. For example, if there are three subjects by default, the arrangement of the camera and the test area is as shown in Figure 2 If it is necessary to increase the number of testees, for example, from three to five, the test area can be reset on the visual panel. Five test areas are determined by setting five quadrilaterals to accommodate the increased number of subjects.

[0042] In some embodiments, in S110, the step of analyzing the video image of the subject entering the test area to achieve identity recognition comprises:

[0043] S111, input each frame image of the video into a preset human pose estimation model to extract the position of each human body box and the human body key point, determine whether the subject has arrived at the test area through the intersection over union of the human body box position and each test area, and further obtain the face position of the subject who has arrived at the test area through the corresponding human body key point.

[0044] Specifically, the system inputs each frame of the video image of each subject entering the test area into a preset human pose estimation model. The human pose estimation model is a multi-human pose estimation model that can identify human bodies in an image and extract the bounding box (i.e., human body box) and human body key points (such as joint positions) of each human body. The intersection over union (IoU) between each human body box and the predefined test area is calculated. The IoU is an indicator that measures the degree of overlap between two regions, and the calculation formula is the intersection area divided by the union area. If the IoU of a certain human body box and the test area exceeds a certain threshold, the system determines that the subject has arrived at the test area. Through the human pose estimation model, not only the key point data of the human body can be extracted, but also the face position of the subject who has arrived at the test area can be further determined.

[0045] Please refer to Figure 3 , Figure 3 which is a schematic diagram of human body key points provided by an embodiment of the present application. Figure 3 The 30 key points of the human body are shown, including any one or a combination of the following: top of head, nose, left ear, right ear, chin, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left palm, right palm, left middle finger, right middle finger, left hip, right hip, left knee, right knee, left ankle, right ankle, left heel, right heel, left toe, right toe, left eye, right eye, left thumb, right thumb, etc. These human body key points cover the main joints and facial features of the human body, which helps to improve the accuracy and comprehensiveness of human body posture analysis in sports events.

[0046] When performing multi-person keypoint annotation in a scene, due to the huge workload, the present application adopts two methods of machine annotation and manual annotation.

[0047] Method one:

[0048] Pre-training: First, manually annotate the keypoint information of the single person motion main target, including the human keypoint and the human box (i.e. the rectangular bounding box surrounding the human body).

[0049] Human detection: Existing human detection models can be used to detect human boxes in images, and the IoU between the detected human boxes and the manually annotated human boxes is calculated.

[0050] Keypoint detection and screening: When the IoU is less than the set threshold, it means that the detected human box does not match the manually annotated human box. In this case, use the existing single keypoint detection large model to detect the human keypoint on the detected human box, and use the existing human keypoint confidence evaluation model to measure the quality of these key points. In this way, the human box and human keypoint of non-motion main target with high quality are screened out.

[0051] Method two:

[0052] Manual annotation: On completely unannotated image data, manually annotate the human box. Keypoint detection and quality evaluation: Use the single keypoint detection large model to detect the human keypoint on the manually annotated human box, and use the human keypoint confidence evaluation model to measure the quality of these human keypoints. In this way, the multi-human box and human keypoint with high quality are obtained.

[0053] The combination of these two methods aims to improve the annotation efficiency and accuracy. Method one uses existing benchmark data and machine detection to assist annotation, while method two manually annotates on completely new data and combines machine detection to ensure the quality of key points. In this way, the workload of manual annotation can be reduced while ensuring the quality of annotation.

[0054] Please refer to Figure 4 , Figure 4FIG. 1 is a schematic diagram of a human pose estimation model provided by an embodiment of the present application. The architecture of the human pose estimation model includes Backbone, PANet, and Detection head. The video frame is input to the Backbone, which is a basic part of the deep learning model and can be a pre-trained convolutional neural network (such as ResNet, VGG, etc.), used to extract features from the input image. The PANet is a network structure used to improve the feature propagation and aggregation in the Feature Pyramid Network (FPN). It can help the model better utilize features at different levels, improving the performance of detection and pose estimation. The Detection head is responsible for detecting the position and category of the human body, the position and confidence of the human key points from the features extracted by the Backbone. The Detection head can include a series of convolutional layers to generate human boxes and category scores, point positions and scores. Therefore, the processing flow of the human pose estimation model includes the processing of the input image, feature extraction (through Backbone), further processing and aggregation of features (through PANet), regression detection of human bodies and key points (through Detection head), and finally obtaining the human box and human key points of each human body in the image through post-processing after conversion. These components work together to accurately estimate the poses of multiple people from the image.

[0055] The human pose estimation model is based on computer vision technology and uses an end-to-end architecture design, which can simultaneously identify the positions (human boxes) and key points (joints) of all human bodies in an image at one time. Compared with the traditional two-stage strategy, this method has significant advantages. The traditional top-down strategy first performs human detection and framing, and then performs single-person pose estimation on the human body in each frame; while the bottom-up strategy first detects all key points, and then assembles these key points into the poses of multiple people. In contrast, the one-time end-to-end human pose estimation model achieves higher processing efficiency and real-time performance by directly jointly detecting human boxes and human key points, and is suitable for scenarios that require fast and accurate processing of large-scale human data.

[0056] Specifically, each anchor (i.e., a predefined reference box) in the Detection head predicts a human bounding box (for locating the human body in the image), a target confidence (indicating the probability that the human body is contained within the bounding box predicted by the anchor), a class confidence (related to the class of the human, indicating the probability that the predicted bounding box is a human body), and 30 keypoint information (i.e., the predicted human keypoint coordinates and confidence scores). Subsequently, NMS (Non-Maximum Suppression) is performed on the human bounding boxes, i.e., overlapping bounding boxes are removed. When multiple bounding boxes that may belong to the same person are detected, NMS selects the one with the highest confidence and suppresses the other overlapping boxes. Finally, the information of each human body, i.e., the human box and the human keypoint, is output, as follows:

[0057]

[0058] The first four columns (x1, y1, x2, y2) represent the coordinates of the human box, which are the coordinates of the top-left corner and the bottom-right corner of the bounding box. Following the human box coordinates are the coordinates of the 30 keypoints, which define the pose of the human body.

[0059] Therefore, the Detection head predicts the human bounding box and keypoint through each anchor, then removes the overlapping boxes using NMS, and finally outputs the formatted results, which include the human box coordinates and keypoint coordinates of each human body. This output format facilitates subsequent processing and analysis, such as pose estimation and action recognition.

[0060] S112, comparing the face image with the records in the subject database to achieve identity recognition.

[0061] Specifically, after obtaining the face position, the system compares the extracted face image with the records in the subject database. The database contains the facial feature information of registered subjects. Through comparison, the system can identify the identity of the subject, ensuring accurate identity verification during the evaluation process.

[0062] After the multi-person pose estimation task is completed, the system has identified all individuals in the scene. However, these individuals may contain many non-test objects, and ensuring that multiple testers do not interfere with each other during the evaluation process is also an important consideration. Therefore, in the multi-person evaluation system of the present application, in order to ensure that each test point can be smoothly evaluated, the present application also provides a human body matching scheme. This scheme mainly relies on the accurate identification and positioning of the human body box in the matching process to ensure the independence of each tester and the accuracy of the evaluation. Through this scheme, the tester and the non-tester can be effectively distinguished, and the smooth progress of the test process can be ensured. The following are two schemes, which can be selected according to the hardware and actual test requirements:

[0063] Scheme 1:

[0064] This scheme is a rule-based matching scheme, which is suitable for fixed-point motion, and the subject's activity is basically unchanged, such as sit-ups, sit-and-reach and other movements. This scheme identifies the subjects in each test area, and the identification steps include:

[0065] S113, input the video image of each subject entering the test area into a pre-set human pose estimation model to output the human body box of each subject.

[0066] S114, for each test area, calculate the proportion of the overlapping part of the human body box of each subject and the test area to the entire test area.

[0067] Specifically, each test area is assigned according to the following formula:

[0068]

[0069] Where bbox 人体 represents the detected human body box, bbox 测试区域 represents the test area box, and for example, the threshold value in sit-ups can be set to 0.2. When the proportion of the detected human body box in the test area is greater than or equal to 20%, the human body box is considered to belong to this test area. This threshold is set to ensure that when the subject performs sit-ups, even if part of the body is outside the test area, it will not be mistakenly excluded.

[0070] S115, if the proportion is greater than or equal to the pre-set threshold, the subject is assigned to the test area to identify the subject in each test area.

[0071] For example, when the proportion of the detected human body box in the test area is greater than or equal to 0.2, the object is assigned to the test area.

[0072] If a test area is assigned multiple bounding boxes (e.g., possibly because a foot presser also entered the test area), the system proceeds with further processing:

[0073] S116, when a test area is assigned multiple bounding boxes, calculate the Euclidean distance between the center point of each object and the center point of the test area in the horizontal direction.

[0074] S117, select the subject closest to the center point of the test area and confirm it as the subject of the test area.

[0075] That is, the system determines which individual is closest to the center of the test area by calculating the distance, thereby determining that the individual is the correct object participating in the test.

[0076] Therefore, the purpose of the above S113 to S117 is to assign the detected bounding box to the correct test area, ensuring that only one subject is identified in each test area. And in the case of multiple people entering the same test area, the correct subject can be selected. For example, when a non-tester (such as a foot presser) accidentally enters the test area, the system can accurately identify and confirm the true test subject by calculating the distance. In this way, the system ensures the accuracy and reliability of the test.

[0077] Scheme II:

[0078] This scheme uses image-based matching technology and is suitable for moving object recognition in dynamic scenes. The specific implementation steps are as follows:

[0079] Initial object assignment step: According to the test area matching strategy, assign an initial test object to each test area. Through bounding box positioning, obtain the corresponding human body cropped image, and use the similarity perception model to extract the features of these images. Then, store these features in a special feature bank.

[0080] Real-time feature matching step: During video stream processing, after each frame of image is processed by the multi-person pose estimation model, the human body bounding box and its cropped image of each object are output. These cropped images are input into the similarity perception model one by one to extract real-time image features. By calculating the distance between the features extracted in the current frame and the features stored in the feature bank, it is determined whether each object should be assigned to the corresponding test area.

[0081] Continuous matching and isolation step: Through continuous human body matching process, ensure that the bounding box and key point information of each test area are accurately matched, and maintain the independence of each test area during the entire evaluation and counting process, avoiding mutual interference.

[0082] Through the initial object allocation step, the real-time feature matching step, and the continuous matching and isolation step described above, the system can efficiently and accurately manage object matching in dynamic scenes, ensuring smooth progress of the evaluation process.

[0083] Please refer to Figure 5 , Figure 5 is a schematic diagram of the evaluation process provided by the embodiments of the present application. The evaluation mode includes a fixed-time evaluation mode and a non-fixed-time evaluation mode.

[0084] In the fixed-time evaluation mode, the system confirms that all subjects have entered their respective test areas and that each person has completed the preset preparation action (for example, lying on the side and preparing to start the sit-up, the system announces "please prepare"). Once all subjects are ready, the system starts a unified timer (for example, the system announces "start", indicating the start of the timer), marking the official start of the evaluation process. The subjects then start the sit-up action, and the system continuously monitors and records their performance until the timer reaches the preset time endpoint (for example, the system announces "time's up"), at which point the system stops the evaluation and outputs the score of each subject.

[0085] In the non-fixed-time evaluation mode, the system confirms that a subject has entered any test area and completed the preset preparation action, and then starts independent counting for that subject to start the evaluation process of that subject. This mode allows subjects to enter the test area and start the evaluation according to their own pace, without the need to wait for all subjects to start simultaneously. The system will time each subject entering the test area individually and output the score of the subject immediately when the subject leaves the test area or the timing ends. This flexible evaluation process design is suitable for groups of subjects who cannot start the evaluation at the same time.

[0086] In some embodiments, in S120, during the evaluation process, the step of continuously analyzing the video images of the subjects to extract the human key points includes:

[0087] S121, the video images of the subjects are processed frame by frame, and each frame of image is input into a human pose estimation model to extract the human key points and human box of each subject.

[0088] Specifically, the human pose estimation model can extract the human key points and human box of each subject in one pass and end-to-end.

[0089] S122, input the human key points of each subject into a preset action attribute classification model to obtain the action attribute of each subject.

[0090] Specifically, the action attribute classification model aims to identify and output the current action features of the subject, which can include the following aspects: whether to touch the knee, whether to lie flat, whether the knee angle is within the normal range (e.g., between about 65° and 135°), whether to hold the head with both hands, and whether the direction of sitting during the test is facing left or right.

[0091] The action attribute classification model is a multi-layer perceptron network model. The input data of this model is processed as follows: first, normalize the 30 key points of the human body using the thigh length and hip center point, and then straighten these key point information into a 60-dimensional vector. The middle layer of the model is set to 64 dimensions to extract and convert features. The output layer is a 7-dimensional vector corresponding to the classification of the above action attributes.

[0092] In the selection of the loss function, for the four binary classification attributes of touching the knee, lying flat, holding the head, and orientation, binary cross-entropy loss (BCE Loss) is used for supervised learning. For the three classification attribute of knee angle (including too large, normal, and too small), cross-entropy loss (CE Loss) is used for supervision.

[0093] Through the design of the above action attribute classification model, the action attributes of the subject can be effectively classified, ensuring the accuracy and reliability of the evaluation results. Moreover, if the subject's action does not meet the pre-set standards or rules (i.e., violation), the system will make an announcement, i.e., through a certain way (which can be sound, text prompt, etc.) to inform the relevant personnel that the subject's action has a problem and needs to be adjusted or corrected. Such an announcement function helps to monitor and guide the subject in real time, ensuring that they perform the test according to the correct action.

[0094] S123, input the human key points and action attributes of each subject into a state machine, which defines a preparation state, a lying flat state, and a sitting up state, wherein the switching between the lying flat state and the sitting up state depends on the key inflection point determined by the human key points.

[0095] S124, determine the key inflection point according to the input key point data and action attributes, and make a compliance judgment accordingly, output the compliance judgment result and the counting result.

[0096] Specifically, the key inflection points include a sit-up inflection point and a lie-down inflection point. When the subject sits up from a lying-down state to a moving speed less than a preset threshold, the state machine counts a minimum body bending angle and determines the sit-up inflection point according to the angle. When the subject lies down from a sit-up state to a moving speed less than a preset threshold, the state machine counts a maximum body bending angle and determines the lie-down inflection point according to the angle; according to the determined sit-up inflection point and lie-down inflection point, the state machine combines the action attribute to make a compliance judgment, and outputs a compliance judgment result and a counting result; the compliance judgment result is used to confirm whether the action of the subject conforms to the evaluation standard, and the counting result is counted according to the action switching in compliance, so as to obtain the number of sit-ups of the subject.

[0097] Reference is made to Figure 6 , Figure 6 FIG. 1 is a schematic diagram of counting logic provided by an embodiment of the present application. Figure 6 FIG. 2 shows a counting logic flow of a multi-person sit-up evaluation system, in which each group of human key points and action attributes are input into corresponding state machines, as follows:

[0098] First, images are extracted from the video stream, and human key point information of each subject is identified through a human posture estimation model. Then, the detected human key point information and its corresponding human box are put into a queue and matched with the corresponding test area queue. The human key point information matched to the area queue is input into an action attribute classification model, which outputs the action attributes of the objects in each area. These action attributes are also put into a queue and matched with the area queue. The updated human key points and updated action attributes are input into the state machine. The state machine is responsible for state judgment and counting for each area (such as area 1, area 2, area 3, etc.). Specifically, it includes:

[0099] In the preparation state, the system checks whether the subject has completed the preparation action and the action is compliant. If compliant, the object enters the lying-down state. In the lying-down state, the system detects whether the subject touches the knee and whether there is a violation action. If the object touches the knee and does not violate, the count is incremented by one. The transition between the lying-down state and the sit-up state is determined according to the lie-down inflection point and the sit-up inflection point. In the sit-up state, the system detects whether the subject lies down.

[0100] Further, for the state machine counting logic flow of the multi-person sit-up evaluation system, the specific steps are as follows:

[0101] Process 1: Preparation action detection.

[0102] The system first enters an initial loop to detect whether the preparation action of the test object meets the standard. Once it is confirmed that the preparation action is compliant, the system will switch to the lying-down state and start a new sit-up-sit-down cycle.

[0103] Flow 2: Sit-up state detection.

[0104] During the supine-sit-up cycle, the system continuously detects the sit-up key frame and determines whether the object touches the knee through the action attribute classification model. When the sit-up key frame is detected and there is no violation behavior from the start of the cycle to the current time (compliance judgment through action attributes and key points), the system will add one to the count. In addition, the system also continuously detects the sit-up inflection point. If the sit-up inflection point is detected but the sit-up key frame is not detected or there is a violation behavior, it will not be counted, and the system switches to the sit-up state.

[0105] Flow 3: Supine state detection.

[0106] The system continues to detect the supine key frame and determines whether the object is lying down through the action attribute classification model. At the same time, the system also continuously detects the lying down inflection point. If the lying down inflection point is detected but the supine key frame is not detected, the system will record the not lying down state and start the next supine-sit-up cycle, while archiving the not lying down information, and then switch back to the lying down state.

[0107] Flow 2 and Flow 3 are executed in a loop.

[0108] Flow 4: Output.

[0109] After completing the count, the system outputs the relevant information of each cycle, including the time from the start (lying down) to the sit-up and then to the end (lying down), the flag item whether the knee is lying down, the number of each violation, whether the cycle is counted, etc. According to this information, the system can obtain video clips of a certain cycle for count verification, violation review, or statistical frequency of each violation item, thereby providing comprehensive and comprehensive test evaluation.

[0110] Through the above counting logic flow, the system can accurately track the action state of the object in each test area and count when certain conditions are met, thereby realizing the automatic evaluation of the sit-up action of multiple people.

[0111] In summary, the motion evaluation method provided in the application is different from the traditional bottom-up scheme and top-down scheme. The method jointly detects the human body frame and the human body key points in the entire test area, and then processes through a matching scheme based on the test area, without the need for an additional human body detection step, thereby saving computing resources and improving evaluation efficiency. In addition, the state machine in the application archives key information during the counting process. For example, in the case where the user does not completely lie down or does not touch the knee, the system can capture and record the key frame closest to completely lying down or touching the knee, so that the user can view and analyze after the test. Finally, by combining the action attribute classification judgment and the state transition of the state machine, the system can detect and report, in real time, for example, the violation behaviors including the leg not being straightened, not holding the head, the foot being off the ground, etc., which helps the subject to correct the action in time during the test, and ensures the accuracy and effectiveness of the test.

[0112] The motion evaluation device provided in the application is described below. The motion evaluation device and the motion evaluation device described below can be mutually corresponding with reference to the motion evaluation method described above.

[0113] Please refer to Figure 7 , Figure 7 is a structural schematic diagram of the motion evaluation device provided in the application. A motion evaluation device 700 includes an identity detection module 710, a compliance evaluation module 720, and a score output module 730.

[0114] Illustratively, the identity detection module 710 is used to analyze the video image of the subject entering the test area to realize identity recognition, and to judge the readiness state according to the preset evaluation mode to start the evaluation process.

[0115] Illustratively, the compliance evaluation module 720 is used to continuously analyze the video image of the subject in the evaluation process, extract the human body key points and the action attributes, and input them into the preset state machine to perform action compliance evaluation.

[0116] Illustratively, the score output module 730 is used to output the evaluation score of the subject according to the result of the action compliance evaluation.

[0117] Illustratively, the identity detection module 710 is also used to:

[0118] input each frame of the video image into a preset human posture estimation model to extract each human body frame position and human body key point, determine whether the subject has reached the test area through the intersection over union of the human body frame position and each test area, and further obtain the face position of the subject reaching the test area through the corresponding human body key point;

[0119] compare the face image with the record in the subject database to realize identity recognition.

[0120] Exemplarily, the identity detection module 710 is further configured to:

[0121] In the fixed-time evaluation mode, after confirming that all the subjects have entered the test area and completed the preset preparation action, a timer is started to start the evaluation process.

[0122] In the non-fixed-time evaluation mode, after confirming that a subject has entered the test area and completed the preset preparation action, the subject is counted to start the evaluation process.

[0123] Exemplarily, the compliance evaluation module 720 is further configured to:

[0124] The video images of the subjects are processed frame by frame, and each frame of image is input into a human pose estimation model to extract the human key points and the human box of each subject.

[0125] The human key points of each subject are input into a preset action attribute classification model to obtain the action attribute of each subject.

[0126] Exemplarily, the compliance evaluation module 720 is further configured to:

[0127] The human key points and the action attribute of each subject are input into a state machine, and the state machine defines a preparation state, a lying state and a sitting state, wherein the switching between the lying state and the sitting state depends on a key inflection point determined by the human key points.

[0128] According to the input key point data and the action attribute, the key inflection point is determined, and the compliance judgment is made according to the key inflection point, and the compliance judgment result and the counting result are output.

[0129] Exemplarily, the compliance evaluation module 720 is further configured to:

[0130] When the subject sits up from the lying state to the moving speed less than the preset threshold, the state machine counts the minimum body bending angle, and determines the sitting-up inflection point according to the angle.

[0131] When the subject lies from the sitting state to the moving speed less than the preset threshold, the state machine counts the maximum body bending angle, and determines the lying inflection point according to the angle.

[0132] According to the determined sitting-up inflection point and lying inflection point, the state machine makes compliance judgment combined with the action attribute, and outputs the compliance judgment result and the counting result; the compliance judgment result is used to confirm whether the action of the subject conforms to the evaluation standard, and the counting result is counted according to the compliant action switching to obtain the sit-up number of the subject.

[0133] Exemplarily, the identity detection module 710 is further configured to:

[0134] The video image of each subject entering the test area is input into a preset human pose estimation model to output the human bounding box of each subject;

[0135] For each test area, calculate the proportion of the overlap between the human body frame of each subject and its test area to the entire test area;

[0136] If the proportion is greater than or equal to a preset threshold, the subject is assigned to the test area to identify the subject in each test area.

[0137] For example, the identity detection module 710 is also used for:

[0138] When a test area is assigned multiple human bounding boxes, calculate the Euclidean distance between the center point of each object and the center point of the test area in the horizontal direction.

[0139] Select the subject closest to the center of the test area and identify them as the subject for that test area.

[0140] For example, the key points of the human body include any one or a combination of the following:

[0141] Top of head, nose, left ear, right ear, chin, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left palm, right palm, left middle finger, right middle finger, left hip, right hip, left knee, right knee, left ankle, right ankle, left heel, right heel, left toe, right toe, left eye, right eye, left thumb, right thumb.

[0142] It should be noted that the motion evaluation device provided in this application embodiment can implement all the method steps implemented in the above method embodiment and can achieve the same technical effect. Here, the parts that are the same as those in the method embodiment and the beneficial effects will not be described in detail.

[0143] Figure 8 This is a schematic diagram of the structure of the electronic device provided in the embodiments of this application, such as... Figure 8 As shown, the electronic device may include a processor 810, a communications interface 820, a memory 830, and a communication bus 840, wherein the processor 810, the communications interface 820, and the memory 830 communicate with each other via the communication bus 840. The processor 810 can call logical instructions in the memory 830 to execute the motion evaluation method.

[0144] Further, the logic instructions in the memory 830 described above can be implemented in the form of software functional units and sold or used as independent products, and can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the parts that make contributions to the prior art or parts of the technical solutions can be embodied in the form of a software product, and the computer software product is stored in a storage medium, including a number of instructions for ensuring that a computer device (which can be a personal computer, a server, or a network device, etc.) executes all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.

[0145] In another aspect, the present application also provides a computer program product, which comprises a computer program stored on a non-transitory computer readable storage medium, and the computer program comprises program instructions, and when the program instructions are executed by a computer, the computer can execute the motion evaluation method provided by the above-mentioned methods.

[0146] In yet another aspect, the present application also provides a non-transitory computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to implement the motion evaluation method provided by the above-mentioned methods.

[0147] The electronic device, the computer program product, and the processor readable storage medium provided by the embodiments of the present application store the computer program which enables the processor to implement all the method steps realized by the above-mentioned method embodiments and achieve the same technical effects. Therefore, the same parts and beneficial effects of the embodiments of the present application as the method embodiments are not described in detail here.

[0148] The device embodiments described above are only schematic, and the units described as separate components can or can not be physically separate, and the components displayed as units can or can not be physical units, that is, they can be located in one place, or distributed on multiple network units. Part or all of the modules can be selected according to actual needs to achieve the purpose of the embodiments of the present application. Those skilled in the art can understand and implement it without creative labor.

[0149] Those skilled in the art can clearly understand the implementation of the various embodiments by means of software and necessary general hardware platforms through the description of the above embodiments, and of course, the embodiments can also be implemented by hardware. Based on such understanding, the above technical solutions, essentially or in other words, the part of the prior art that contributes to the technical solutions can be embodied in the form of a software product. The computer software product can be stored in a computer readable storage medium, such as a ROM / RAM, a magnetic disk, an optical disk, and the like, and includes a plurality of instructions for ensuring that a computer device (which can be a personal computer, a server, or a network device, etc.) executes the methods described in the various embodiments or some parts of the embodiments.

[0150] Finally, it should be noted that: the above examples are only used to illustrate the technical solutions of the present application, and not to limit them; although the present application has been described in detail with reference to the foregoing examples, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing examples, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method of motion assessment, characterized by, The method comprises: analyzing video images of subjects entering a test area to achieve identity recognition, and judging a preparation state according to a preset evaluation mode to start an evaluation process; wherein the evaluation mode comprises a fixed-time evaluation mode and a non-fixed-time evaluation mode; in the fixed-time evaluation mode, after confirming that all subjects have entered the test area and completed a preset preparation action, a timer is started to start the evaluation process; in the non-fixed-time evaluation mode, after confirming that a subject has entered the test area and completed a preset preparation action, counting is started on the subject to start the evaluation process; in the evaluation process, video images of the subjects are continuously analyzed, human body key points and action attributes are extracted, and the human body key points and the action attributes are input into a preset state machine to perform action compliance evaluation; according to a result of the action compliance evaluation, an evaluation result of the subjects is output.

2. The motion assessment method of claim 1, wherein, The step of analyzing video images of subjects entering a test area to achieve identity recognition comprises: inputting each frame of the video into a preset human body posture estimation model to extract human body frame positions and human body key points, determining whether the subjects have arrived at the test area through an intersection-over-union of the human body frame positions and each test area, and further obtaining a face position of the subject who has arrived at the test area through a corresponding human body key point; comparing the face image with records in a subject database to achieve identity recognition.

3. The motion assessment method of claim 1, wherein, The step of continuously analyzing video images of the subjects and extracting human body key points comprises: frame-by-frame processing of the video images of the subjects, inputting each frame of the video images into the human body posture estimation model to extract human body key points and human body frames of each subject; inputting the human body key points of each subject into a preset action attribute classification model to obtain action attributes of each subject.

4. The motion assessment method of claim 3, wherein, The step of inputting the human body key points and the action attributes of each subject into a preset state machine to perform action compliance evaluation comprises: inputting the human body key points and the action attributes of each subject into the state machine, the state machine defining a preparation state, a lying state and a sitting state, wherein the lying state and the sitting state are switched depending on a key inflection point determined by the human body key points; determining the key inflection point according to the input key point data and the action attributes, and performing compliance judgment according to the key inflection point to output a compliance judgment result and a counting result.

5. The motion assessment method of claim 4, wherein, The key inflection point comprises a sitting inflection point and a lying inflection point, and the step of inputting the key inflection point into the preset state machine to perform action compliance evaluation comprises: when the subject sits up from the lying state to a moving speed less than a preset threshold, the state machine counts a minimum body bending angle, and determines the sitting inflection point according to the angle; when the subject lies down from the sitting state to a moving speed less than a preset threshold, the state machine counts a maximum body bending angle, and determines the lying inflection point according to the angle; according to the determined sitting inflection point and the lying inflection point, the state machine performs compliance judgment in combination with the action attributes, and outputs a compliance judgment result and a counting result; the compliance judgment result is used to confirm whether the action of the subject conforms to the evaluation standard, and the counting result is counted according to the compliant action switching to obtain the number of sit-ups of the subject.

6. The motion assessment method of claim 1, wherein, The method further comprises identifying the subject in each test area, the identifying step comprising: inputting the video image of each subject entering the test area into a preset human pose estimation model to output a human box of each subject; for each test area, calculating the proportion of the overlapping part of the human box of each subject and its test area in the whole test area; if the proportion is greater than or equal to a preset threshold, assigning the subject to the test area to identify the subject in each test area.

7. The motion assessment method of claim 6, wherein, The method further comprises processing the case of multiple people being assigned to a test area, the processing step comprising: when a test area is assigned to multiple human boxes, calculating the Euclidean distance between the center point of each object and the center point of the test area in the horizontal direction; selecting the subject closest to the center point of the test area and confirming it as the subject of the test area.

8. The motion assessment method of claim 1, wherein, The human key points include any one or a combination thereof: top of head, nose, left ear, right ear, chin, neck, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left palm, right palm, left middle finger, right middle finger, left hip, right hip, left knee, right knee, left ankle, right ankle, left heel, right heel, left toe, right toe, left eye, right eye, left thumb, right thumb.

9. A motion evaluation apparatus, characterized by comprising: The device comprises: an identity detection module for analyzing the video image of the subject entering the test area to achieve identity recognition, judging the readiness state according to a preset evaluation mode to start the evaluation process; wherein the evaluation mode includes a fixed-time evaluation mode and a non-fixed-time evaluation mode; in the fixed-time evaluation mode, after confirming that all subjects have entered the test area and completed the preset preparation action, a timer is started to start the evaluation process; in the non-fixed-time evaluation mode, after confirming that a subject has entered the test area and completed the preset preparation action, the subject is counted to start the evaluation process; a compliance evaluation module for continuously analyzing the video image of the subject in the evaluation process, extracting human key points and action attributes, and inputting them into a preset state machine to evaluate the action compliance; a score output module for outputting the evaluation score of the subject according to the result of the action compliance evaluation.

10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that, The processor executes the computer program to realize the steps of the motion evaluation method according to any one of claims 1 to 8. 11.A non-transitory computer-readable storage medium having stored thereon a computer program. The computer program is executed by the processor to realize the steps of the motion evaluation method according to any one of claims 1 to 8.

Citation Information

Patent Citations

  • Target tracking method, device, equipment and medium

    CN115115670A

  • Motion evaluation method and device, electronic equipment and storage medium

    CN115590504A

  • Physical education teaching test method and system based on human body posture estimation

    CN116503954A