Sit-up counting method, system, equipment and product based on YOLO and ResNet
Through the combination of YOLO and ResNet model, efficient, accurate and standardized judgment and counting of sit-up movements is achieved, and the problem of low accuracy of sit-up movement counting in the prior art is solved, and it is suitable for large-scale applications and other actions detection.
Patent Information
- Application Number
- CN202510538263.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-27
- Publication Date
- 2025-08-15
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art has low accuracy in the standardized judgment and counting of sit-up movements. Sensor methods are susceptible to position changes and are prone to cheating, manual monitoring is highly subjective, and the detection method of human key point detection may lead to inaccurate positioning of key points.
Using a combination of YOLO and ResNet model, the video data is converted into image frames, the target part is detected and the area image is intercepted, and the deep features are extracted using the ResNet model to judge the action normativeness, and then count it.
It improves the accuracy and standardized judgment of sit-up movement counting, is suitable for large-scale applications, is efficient and flexible, and is suitable for detection and analysis of other movements.
Smart Images

Figure CN120496170A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computer vision and deep learning technology, and specifically relates to a sit-up counting method, system, device and product based on YOLO and ResNet. Background Art
[0002] Sit-ups are a common physical training exercise, widely used in physical fitness tests, fitness, and rehabilitation training. Sit-ups can effectively enhance the body's core strength and muscle endurance by mobilizing muscle groups in the abdomen, hips, and other parts of the body. However, in actual training and testing, the standardization and standardized execution of sit-ups are key factors affecting the training effect and evaluation accuracy. If the movement is not standardized, it will not only affect the training effect, but may also cause damage to the lumbar spine, neck, and other parts of the body. Therefore, how to effectively monitor the standardization of sit-ups and how to accurately count them have become important issues in the fitness field and physical fitness tests.
[0003] At present, the standardization judgment and counting of traditional sit-up movements usually rely on manual monitoring and sensor technology. Among them, manual monitoring requires manual judgment of whether the movements are standard and counting. However, this method has problems such as strong subjectivity, low efficiency, and easy error, and it is difficult to guarantee accuracy. The sensor-based sit-up counting system uses sensor equipment such as accelerometers, gyroscopes, pressure sensors, infrared sensors, etc., which judge whether a sit-up is completed by monitoring the movement of the body. However, in actual application, sensor monitoring may affect the detection results due to changes in the position of the sensor, and it is also prone to cheating, thereby reducing the accuracy of monitoring.
[0004] At the same time, with the rapid development of machine learning, in practical applications, there have also been methods of analyzing sit-up movements by identifying key points of the human body, such as the prior art with application number 202010191868.X, which discloses a sit-up test counting method and system based on the Quick-OpenPose model. The technology identifies key points of the human body and calculates the lines and angles between the key points to judge the standardization of sit-up movements and count them. However, this method of identifying key points of the human body may have the problem that the key points cannot be accurately located, which may lead to errors in the movement analysis and affect the judgment of the standardization of the movement. Therefore, based on the aforementioned deficiencies, how to provide a method that can accurately judge and count the standardization of sit-up movements has become a problem that needs to be solved urgently. Summary of the Invention
[0005] The purpose of the present invention is to provide a sit-up counting method, system, device and product based on YOLO and ResNet to solve the problem of low counting accuracy in the prior art.
[0006] In order to achieve the above object, the present invention adopts the following technical solutions:
[0007] First, a sit-up counting method based on YOLO and ResNet is provided, including:
[0008] Obtain video data of the target person performing sit-ups;
[0009] Converting the sit-up video data into continuous image frames to obtain a plurality of action images;
[0010] Each action image is input into the YOLO target detection model to obtain the region image corresponding to the target part in each action image, wherein the target parts include the lying mat, hands, knees, head, elbows, shoulder blades, buttocks and upper body;
[0011] Input the regional image corresponding to the target part in each action image into the action detection model to obtain the sit-up action category corresponding to each action image, wherein the action detection model is a trained ResNet model, and the sit-up action category corresponding to any action image is used to indicate whether the sit-up action of the target person in any action image is a standard action;
[0012] According to the sit-up action category corresponding to each action image, the sit-up counting result of the target person is obtained.
[0013] Based on the above-disclosed content, the present invention first obtains the sit-up video data of the target person, and then converts it into continuous image frames to obtain a number of action images; then, the YOLO target detection model is used to detect the target parts such as shoulder blades, knees, and hands for judging the standardization of sit-up movements, and the regional images corresponding to each target part are cut out; then, the cropped images are further processed using the ResNet model, that is, by extracting the features of the regional images of the target parts, the sit-up action category corresponding to each action image is obtained to determine whether the movement is standard; finally, the sit-up count can be performed according to the sit-up action category corresponding to each action image, thereby obtaining the sit-up count result.
[0014] Through the above design, the present invention is different from the traditional sensor method and human key point detection method. It uses the method of target detection and image feature extraction to obtain the sit-up action category corresponding to each action image, and based on this, it judges whether the sit-up action corresponding to each action image is standard, thereby completing the counting of sit-ups of people; thus, the present invention uses the YOLO model to perform target detection and the ResNet model to extract deep features of various parts of the body. In this way, the difference in actions can be accurately judged, so that the standardization of sit-up actions can be more accurately identified, thereby improving the accuracy of sit-up counting, and therefore, it is very suitable for large-scale application and promotion.
[0015] In one possible design, the region image corresponding to the target part in each action image is input into the action detection model to obtain the sit-up action category corresponding to each action image, including:
[0016] Merge the area images corresponding to the hands and head, the area images corresponding to the elbows and knees, the area images corresponding to the shoulder blades and the mat, and the area images corresponding to the buttocks and the mat in each action image to obtain a detection image of the hands holding the head, a detection image of the elbows touching the knees, a detection image of the shoulder blades touching the mat, and a detection image of the buttocks touching the mat corresponding to each action image;
[0017] Inputting the regional images corresponding to the knees and the regional images corresponding to the upper body in each action image, as well as the detection images of hands holding the head, elbows touching the knees, shoulder blades touching the pads, and buttocks touching the pads corresponding to each action image into the action detection model, thereby respectively obtaining the target person's knee flexion detection category, upper body state category, hands holding the head detection category, elbows touching the knees detection category, shoulder blades touching the pads detection category, and buttocks leaving the pad detection category in each action image;
[0018] The sit-up action category of each action image is generated by using the target person's knee bending detection category, upper body state category, hands holding head detection category, elbows touching knees detection category, shoulder blades touching pad detection category and buttocks leaving pad detection category in each action image.
[0019] In one possible design, the sit-up action categories corresponding to any action image include a knee flexion detection category, a hands-on-head-holding detection category, a double-elbow-to-knee-touching detection category, a shoulder blade-to-pad-touching detection category, a hip-off-pad detection category, and a human upper body state category during the target person's sit-up exercise. The sit-up action categories include 42 categories, each of which includes a first standard action category, a second standard action category, and a third standard action category. The first standard action category, the second standard action category, the third standard action category, the second standard action category, and the first standard action category constitute the human action categories corresponding to a complete sit-up action.
[0020] The target person's sit-up count result is obtained based on the sit-up action category corresponding to each action image, including:
[0021] determining whether there is a target image group among the plurality of action images, wherein the target image group includes at least two consecutive action images, and each action image in the at least two consecutive action images has the same sit-up action category;
[0022] If so, deleting the designated image in the target image group to obtain an action image set after deleting the designated image, wherein the designated image is all action images after the first action image in the target image group;
[0023] Initialize the number of sit-ups i to 0;
[0024] Extract the first n action images from the action image set, and determine whether the sit-up action categories corresponding to the first n action images are the first standard action category, the second standard action category, the third standard action category, the second standard action category, and the first standard action category, respectively, where the value of n is 5
[0025] If so, add 1 to i and delete the first n-1 action images from the action image set to obtain a new action image set;
[0026] The action image set is updated to the new action image set, and the first n action images are extracted from the action image set again until the action image set is polled. After the polling is completed, the value of i is used as the sit-up counting result.
[0027] In one possible design, the method further includes:
[0028] If not, the output prompt message is that the action is not standardized.
[0029] In one possible design, after obtaining the plurality of action images, the method further includes:
[0030] Performing data preprocessing on the plurality of action images to obtain the plurality of preprocessed action images, so as to input each of the preprocessed action images into a YOLO target detection model to obtain a region image corresponding to a target part in each action image, wherein the data preprocessing includes resizing and data augmentation processing;
[0031] Before inputting the region image corresponding to the target part in each action image into the action detection model, the method further includes:
[0032] The area image corresponding to the target part in each action image is cropped to obtain the cropped area image corresponding to the target part in each action image, so as to input the cropped area image corresponding to the target part in each action image into the action detection model to obtain the sit-up action category corresponding to each action image.
[0033] In one possible design, the YOLO target detection model is trained using the following method;
[0034] Obtain sample video data of sit-up performances of several sample persons;
[0035] Convert the sit-up sample video data of each sample person into continuous sample image frames to obtain a number of sample action images corresponding to each sample person, and use the sample action images corresponding to each sample person to form a first initial training data set;
[0036] Performing data preprocessing on the first initial training data set to obtain a preprocessed initial first training data set;
[0037] performing labeling processing on the target part in each sample action image in the preprocessed initial first training data set to obtain a first training data set after the labeling processing, wherein the label data of any sample action image in the first training data set includes the category and position of the target part in the any sample action image;
[0038] The YOLO model is trained using each sample action image of each sample person in the first training data set as input and the predicted category, confidence, and predicted position of the regional image corresponding to the target part in each sample action image of each sample person as output, so as to obtain the YOLO target detection model after the training is completed.
[0039] In a possible design, the action detection model is trained using the following method:
[0040] Based on the predicted category, confidence, and predicted position of the regional image corresponding to the target part in each sample action image of each sample person output by the YLOL model, a sample regional image of the target part in each sample action image of each sample person is cut out from each sample action image of each sample person;
[0041] performing a resizing process on the sample region image of the target part in each sample action image of each sample person, so as to obtain an adjusted region image of the target part in each sample action image after the resizing process;
[0042] For any sample person, the adjustment area images corresponding to the hands and head, the adjustment area images corresponding to the elbows and knees, the adjustment area images corresponding to the shoulder blades and the lying pad, and the adjustment area images corresponding to the buttocks and the lying pad in each sample action image of any sample person are merged respectively, so as to obtain, after the merging process, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image corresponding to each sample action image of the any sample person, and after the sample action images of all the sample persons are polled, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image corresponding to each sample action image of all the sample persons are obtained;
[0043] A second initial training dataset is formed using the adjustment area images corresponding to the knees, the adjustment area images corresponding to the upper body, the hands-on-head detection sample images, the elbows-on-knees detection sample images, the shoulder blades-on-pads detection sample images, and the buttocks-on-pads detection sample images in each sample action image of each sample person;
[0044] performing action category labeling processing on each image in the second initial training data set to obtain a second training data set after the action category labeling processing;
[0045] The ResNet model is trained using the adjustment area image corresponding to the knee, the adjustment area image corresponding to the upper body, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image in each sample action image corresponding to each sample person in the second training data set as input, and the sit-up action category of each sample person as output, so as to obtain the action detection model after the training is completed.
[0046] Secondly, a sit-up counting system based on YOLO and ResNet is provided, including:
[0047] an acquisition unit, configured to acquire video data of a target person performing sit-ups;
[0048] A frame segmentation unit, configured to convert the sit-up video data into continuous image frames to obtain a plurality of action images;
[0049] A target detection unit is used to input each action image into a YOLO target detection model to obtain a region image corresponding to a target part in each action image, wherein the target part includes a lying mat, hands, knees, head, elbows, shoulder blades, buttocks, and upper body;
[0050] An action detection unit is configured to input a region image corresponding to a target part in each action image into an action detection model to obtain a sit-up action category corresponding to each action image, wherein the action detection model is a trained ResNet model, and the sit-up action category corresponding to any action image is used to indicate whether the sit-up action performed by the target person in the action image is a standard action;
[0051] The counting unit is used to obtain the sit-up counting result of the target person according to the sit-up action category corresponding to each action image.
[0052] In the third aspect, a sit-up counting device based on YOLO and ResNet is provided. Taking the device as an electronic device as an example, it includes a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the sit-up counting method based on YOLO and ResNet as described in the first aspect or any possible design of the first aspect.
[0053] In a fourth aspect, a storage medium is provided, on which instructions are stored. When the instructions are run on a computer, the sit-up counting method based on YOLO and ResNet as described in the first aspect or any possible design of the first aspect is executed.
[0054] In a fifth aspect, a computer program product comprising instructions is provided, which, when executed on a computer, causes the computer to execute the sit-up counting method based on YOLO and ResNet as described in the first aspect or any possible design of the first aspect.
[0055] Beneficial effects:
[0056] (1) The present invention is different from the traditional sensor method and the human key point detection method. It uses the target detection and image feature extraction method to obtain the sit-up action category corresponding to each action image, and based on this, it determines whether the sit-up action corresponding to each action image is standard, thereby completing the sit-up counting of the person; thus, the present invention uses the YOLO model to perform target detection and the ResNet model to extract the deep features of various parts of the body. In this way, the difference in actions can be accurately judged, so that the standardization of the sit-up action can be more accurately identified, thereby improving the accuracy of the sit-up counting, and therefore, it is very suitable for large-scale application and promotion.
[0057] (2) YOLO, as a lightweight but efficient target detection model, can quickly detect the positions of multiple body parts in real-time videos. At the same time, ResNet can extract deep features of various parts of the body and accurately judge the differences in movements. Therefore, the present invention has higher processing speed and accuracy than traditional human posture estimation algorithms.
[0058] (3) The present invention can be further extended to the detection and analysis of other actions, such as push-ups, squats, etc. Therefore, by simply modifying the data and labels of network training, it can quickly adapt to new action categories, thereby improving the flexibility of use. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 A schematic flow chart of the steps of a sit-up counting method based on YOLO and ResNet provided in an embodiment of the present invention;
[0060] Figure 2 A schematic diagram of the structure of a sit-up counting system based on YOLO and ResNet provided in an embodiment of the present invention;
[0061] Figure 3 A schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0062] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the present invention will be briefly introduced below in conjunction with the drawings and the description of the embodiments or the prior art. Obviously, the following description of the structure of the drawings is only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative work. It should be noted that the description of these embodiments is used to help understand the present invention, but does not constitute a limitation of the present invention.
[0063] It should be understood that although the terms "first," "second," etc. may be used herein to describe various elements, these elements should not be limited by these terms. These terms are only used to distinguish one element from another. For example, a first element can be referred to as a second element, and similarly, a second element can be referred to as a first element without departing from the scope of the exemplary embodiments of the present invention.
[0064] It should be understood that the term "and / or" that may appear in this document is merely a description of the association relationship between associated objects, indicating that three relationships may exist. For example, A and / or B may indicate three situations: A exists alone, B exists alone, and A and B exist at the same time. The term " / and" that may appear in this document describes another type of association object relationship, indicating that two relationships may exist. For example, A / and B may indicate two situations: A exists alone, and A and B exist alone. In addition, the character " / " that may appear in this document generally indicates that the previous and subsequent associated objects are in an "or" relationship.
[0065] Example:
[0066] See also Figure 1 As shown, the sit-up counting method based on YOLO and ResNet provided in this embodiment is different from the traditional sensor method and the human key point detection method. It uses the target detection and image feature extraction method to obtain the sit-up action category corresponding to each action image, and based on this, it determines whether the sit-up action corresponding to each action image is standardized, thereby completing the sit-up counting of the person; thus, the present invention uses the YOLO model for target detection and the ResNet model to extract deep features of various parts of the body, thereby accurately judging the difference in actions, thereby being able to more accurately identify the standardization of sit-up actions, thereby improving the accuracy of sit-up counting, and therefore, is very suitable for large-scale application and promotion; wherein, for example, this method can be, but is not limited to, running on the sit-up counting terminal side. Optionally, the sit-up counting terminal can be, but is not limited to, a personal computer, a tablet computer, or a smart phone. It can be understood that the aforementioned execution subject does not constitute a limitation of the embodiment of the present application. Accordingly, the running steps of this method can be, but are not limited to, as shown in the following steps S1 to S5.
[0067] S1. Obtain video data of the target person doing sit-ups. In a specific implementation, for example, but not limited to, a camera may be used to collect video data of the target person doing sit-ups, and then the collected video data may be transmitted to a sit-up counting terminal for image processing, thereby achieving the target person's sit-up counting.
[0068] Optionally, for example, when collecting video data, the camera is placed on one side of the target person and the target person's body is ensured to be in the center of the picture. This ensures the accuracy of subsequent image processing.
[0069] After acquiring the sit-up video data of the target person, the video data may be framed to obtain a sit-up action image, and the process is shown in the following step S2.
[0070] S2. Convert the sit-up video data into continuous image frames to obtain a number of action images; in specific implementation, for example, but not limited to, using video processing tools such as OpenCV to decode the video data into a continuous image frame sequence for subsequent image processing; after obtaining a number of action images, images of human body parts that need to be judged for normativeness (i.e., regional images corresponding to the target parts described below) can be extracted, so that the normativeness of the sit-up action can be judged based on the extracted images of the human body parts; wherein, the extraction process of the target part image can be, but not limited to, as shown in the following step S3.
[0071] S3. Input each action image into the YOLO target detection model to obtain a regional image corresponding to the target part in each action image, wherein the target parts include the mat, hands, knees, head, elbows, shoulder blades, buttocks and upper body of the human body; in a specific implementation, the YOLO target detection model is trained by taking the sample action images corresponding to the sit-up sample video data of each sample person as input, and the predicted category, confidence and predicted position of the regional image corresponding to the target part in each sample action image of each sample person as output. Therefore, when any action image is input into the YOLO target detection model, the predicted category, confidence and predicted position of the target part in any action image will be output; specifically, the predicted category refers to the category of the target part, such as hands, buttocks or shoulder blades, etc. (which can be represented by different category numbers), and the predicted position includes the center coordinates of the bounding box of the target part, the height and width of the bounding box, etc.; therefore, based on the aforementioned predicted category and predicted position, the regional images corresponding to different target parts can be cut out from the action image.
[0072] Optionally, for example, before inputting each action image into the YOLO target detection model, each action image may be subjected to data preprocessing to obtain several preprocessed action images, so that each preprocessed action image can be input into the YOLO target detection model to obtain a regional image corresponding to the target part in each action image; wherein, for example, the data preprocessing may include but is not limited to size adjustment and data enhancement processing.
[0073] Specifically, the size of each action image can be adjusted to (416, 146, 3) to adapt to the input of the YOLO target detection model. At the same time, data augmentation processing is to perform various transformations on the original image, such as image rotation, flipping, brightness adjustment, etc., to increase data diversity and avoid overfitting of the model.
[0074] Optionally, the following discloses one of the training methods for the aforementioned YOLO object detection model, as shown in the following steps:
[0075] A. Obtain sit-up sample video data of several sample persons; in this embodiment, the sit-up sample video data is the same as the acquisition process of the sit-up video data, which will not be repeated here. At the same time, the sit-up sample video data contains at least one complete sit-up action of the target person; after obtaining the sit-up sample video data of the sample person, it also needs to be processed frame by frame, and the process is shown in the following step B.
[0076] B. Convert the sit-up sample video data of each sample person into continuous sample image frames, obtain a number of sample action images corresponding to each sample person, and use the sample action images corresponding to each sample person to form a first initial training data set; in this embodiment, video processing tools are also used to convert the sit-up sample video data into a continuous image frame sequence, and the process can be seen in the aforementioned step S2; then, after obtaining a number of sample action images corresponding to each sample person, they can be used to form an initial first training data set, and the process is shown in the following step C.
[0077] C. Perform data preprocessing on the first initial training data set to obtain a preprocessed initial first training data set. In this embodiment, the data preprocessing also includes size adjustment and data enhancement processing, and the process can be found in the aforementioned step S3, which will not be repeated here.
[0078] After completing the preprocessing of the first initial training data set, labeling can be performed to facilitate subsequent model training. The labeling process can be, but is not limited to, as shown in step D below.
[0079] D. Label the target part in each sample action image in the preprocessed initial first training data set to obtain a first training data set after the label labeling process, wherein the label data of any sample action image in the first training data set includes the category and position of the target part in the any sample action image; in the specific implementation, take any sample action image as an example to illustrate the label labeling process, that is, label the target part in any sample action image, that is, label the category and position of the target part.
[0080] Specifically, for the category number class_id, the mat is 0, the sample person's knees are 1, the hands are 2, the head is 3, the elbows are 4, the shoulder blades are 5, the buttocks are 6, and the upper body is 7. For the position, it is necessary to mark the x-axis coordinate x_center, the y-axis coordinate y_center, and the width (normalized to the ratio of the image width) width and the height (normalized to the ratio of the image height) height of the bounding box center point of each target part; in this way, after completing the data annotation, these labeled data can be used to train the YOLO model; based on this, the model is trained with a sufficient number of training samples so that the model can learn the characteristics of each target part, and then identify the target part in real-time detection.
[0081] The training process is shown in the following step E.
[0082] E. Using each sample action image of each sample person in the first training data set as input and the predicted category, confidence, and predicted position of the regional image corresponding to the target part in each sample action image of each sample person as output, the YOLO model is trained to obtain the YOLO target detection model after the training is completed; in this embodiment, the output of the YOLO model is: the bounding box of the target part (i.e., the predicted position), the confidence, and the category label (i.e., the predicted category), which can be expressed as (x, y, w, h, con, C); where x and y are the center coordinates of the bounding box of the prediction result, w and h are the height and width of the bounding box, con is the confidence, and C is the category.
[0083] In this way, through the above steps A to E, the training of the YOLO model can be completed and the YOLO target detection model can be obtained; based on this, in actual use, it is only necessary to input each action image of the target person into the above YOLO target detection model, and the center coordinates of the bounding box of the lying mat, hands, knees, head, elbows, shoulder blades, buttocks and upper body in each action image, as well as the height and width of the bounding box can be obtained; based on this, the area image corresponding to the target part in each action image can be cut out according to the center coordinates, height and width of the above bounding box.
[0084] After completing the capture of the regional image of the part that needs to be judged for the standardization of the sit-up action, the trained ResNet model can be used to perform action recognition on the regional image corresponding to the target part in each action image, thereby obtaining the action detection category of each target part in each action image, and based on this, determining the sit-up action category corresponding to each action image; wherein, the action recognition process is shown in the following step S4.
[0085] S4. Input the area image corresponding to the target part in each action image into the action detection model to obtain the sit-up action category corresponding to each action image, wherein the action detection model is a trained ResNet model, and the sit-up action category corresponding to any action image is used to characterize whether the sit-up action of the target person in any action image is a standard action; in specific implementation, for example, but not limited to, first merging the area images corresponding to the hands and head, the area images corresponding to the elbows and knees, the area images corresponding to the shoulder blades and the mat, and the area images corresponding to the buttocks and the mat in each action image to obtain the hands holding the head detection image, the elbows touching the knees detection image, the shoulder blades touching the pad detection image, and the buttocks touching the pad detection image corresponding to each action image. then, the area images corresponding to the knees and the area images corresponding to the upper body in each action image, as well as the hands-holding-head detection image, the elbows-touching-knees detection image, the shoulder blades-touching-pad detection image, and the buttocks-touching-pad detection image corresponding to each action image are input into the action detection model to respectively obtain the knee flexion detection category, the upper body state category, the hands-holding-head detection category, the elbows-touching-knees detection category, the shoulder blades-touching-pad detection category, and the buttocks-off-pad detection category of the target person in each action image; finally, the sit-up action category of each action image can be generated by using the knee flexion detection category, the upper body state category, the hands-holding-head detection category, the elbows-touching-knees detection category, the shoulder blades-touching-pad detection category, and the buttocks-off-pad detection category of the target person in each action image.
[0086] In this way, the sit-up action category corresponding to any one of the aforementioned action images includes the knee flexion detection category, the hands holding the head detection category, the elbows touching the knees detection category, the shoulder blades touching the pad detection category, the buttocks leaving the pad detection category and the upper body state category of the target person during the sit-up exercise; based on this, it is equivalent to using the state of each target part to form the sit-up action category of each action image, thereby completing the normative recognition of the sit-up action.
[0087] Optionally, in this embodiment, before performing action recognition, for example, but not limited to, the regional image corresponding to the target part in each action image may be cropped to obtain a cropped regional image corresponding to the target part in each action image. Finally, the cropped regional image corresponding to the target part in each action image may be input into the action detection model to obtain the sit-up action category corresponding to each action image; wherein, each regional image may be adjusted to (224, 224, 3) to adapt to the input of ResNet.
[0088] In this embodiment, for example, the aforementioned knee bending detection category, hands holding head detection category, elbows touching knees detection category, shoulder blades touching pad detection category, buttocks leaving pad detection category, and upper body status category can be represented as follows:
[0089]
[0090] Among them, the above-mentioned state 1 indicates that the lifting angle range of the upper body of the human body is (0°, 10°) (that is, the upper body is lying flat, that is, the head, shoulders, back, etc. are all close to the mat), state 2 indicates that the lifting angle range of the upper body of the human body is [10°, 80°], and state 3 indicates that the lifting angle range of the upper body of the human body is (80°, 90°).
[0091] Based on this, for any action image, the sit-up action category corresponding to the action image can be formed according to the detection categories of the six actions in the action image.
[0092] Optionally, for example, the sit-up action categories include 42 categories, the 42 categories include the first standard action category, the second standard action category and the third standard action category, and the first standard action category, the second standard action category, the third standard action category, the second standard action category and the first standard action category constitute the human action category corresponding to a complete sit-up action.
[0093] Among them, the first standard action category, the second standard action category and the third standard action category can be expressed as:
[0094] p 动作1 =(y 屈膝 =1,y 抱头 =1,y 触膝 =0,y 肩胛骨 =1,y 臀部 =1,y 上身 =0);
[0095] p 动作2 =(y 屈膝 =1,y 抱头 =1,y 触膝 =0,y 肩胛骨 =0,y 臀部 =1,y 上身 =1);
[0096] p 动作3 =(y 屈膝 =1,y 抱头 =1,y 触膝 =1,y 肩胛骨 =0,y 臀部 =1,y 上身 =2).
[0097] Among them, the aforementioned p 动作1 、p 动作2 、p 动作3 They represent the first standard action category, the second standard action category, and the third standard action category in sequence.
[0098] At the same time, the remaining sit-up action categories can be expressed as:
[0099] p 动作4 =(y 屈膝 =0,y 抱头 =1,y 触膝 =0,y 肩胛骨 =1,y 臀部 =0,y 上身 =0);
[0100] p 动作5 =(y 屈膝 =0,y 抱头 =1,y 触膝 =0,y 肩胛骨 =1,y 臀部 =1,y 上身 =0);
[0101] p 动作6 =(y 屈膝 =0,y 抱头 =1,y 触膝 =0,y 肩胛骨 =0,y 臀部 =0,y 上身 =0);
[0102] p 动作7 =(y 屈膝 =0,y 抱头 =1,y 触膝 =0,y 肩胛骨 =0,y 臀部 =1,y 上身 =0);
[0103] p 动作8 =(y 屈膝 =0,y 抱头 =0,y 触膝 =0,y 肩胛骨 =1,y 臀部 =0,y 上身 =0);
[0104] p 动作9 =(y 屈膝 =0,y 抱头 =0,y 触膝 =0,y 肩胛骨 =1,y 臀部 =1,y上身 =0);
[0105] p 动作10 =(and 屈膝 =0,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =0,and 上身 =0);
[0106] p 动作11 =(and 屈膝 =0,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =0);
[0107] p 动作12 =(and 屈膝 =1,and 抱头 =1,and 触膝 =0,and 肩胛骨 =1,and 臀部 =0,and 上身 =0);
[0108] p 动作13 =(and 屈膝 =1,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =0,and 上身 =0);
[0109] p 动作14 =(and 屈膝 =1,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =0);
[0110] p 动作15 =(and 屈膝 =1,and 抱头 =0,and 触膝 =0,and 肩胛骨 =1,and 臀部 =0,and 上身 =0);
[0111] p 动作16 =(and 屈膝 =1,and 抱头 =0,and 触膝 =0,and 肩胛骨 =1,and臀部 =1,and 上身 =0);
[0112] p 动作17 =(and 屈膝 =1,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =0,and 上身 =0);
[0113] p 动作18 =(and 屈膝 =1,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =0);
[0114] p 动作19 =(and 屈膝 =2,and 抱头 =1,and 触膝 =0,and 肩胛骨 =1,and 臀部 =0,and 上身 =0);
[0115] p 动作20 =(and 屈膝 =2,and 抱头 =1,and 触膝 =0,and 肩胛骨 =1,and 臀部 =1,and 上身 =0);
[0116] p 动作21 =(and 屈膝 =2,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =0,and 上身 =0);
[0117] p 动作22 =(and 屈膝 =2,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =0);
[0118] p 动作23 =(and 屈膝 =2,and 抱头 =0,and 触膝 =0,and肩胛骨 =1,and 臀部 =0,and 上身 =0);
[0119] p 动作24 =(and 屈膝 =2,and 抱头 =0,and 触膝 =0,and 肩胛骨 =1,and 臀部 =1,and 上身 =0);
[0120] p 动作25 =(and 屈膝 =2,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =0,and 上身 =0);
[0121] p 动作26 =(and 屈膝 =2,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =0);
[0122] p 动作27 =(and 屈膝 =0,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =1);
[0123] p 动作28 =(and 屈膝 =0,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =1);
[0124] p 动作29 =(and 屈膝 =1,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =1);
[0125] p 动作30 =(and 屈膝 =2,and 抱头 =0,and触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =1);
[0126] p 动作31 =(and 屈膝 =2,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =1);
[0127] p 动作32 =(and 屈膝 =0,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =2);
[0128] p 动作33 =(and 屈膝 =0,and 抱头 =0,and 触膝 =1,and 肩胛骨 =0,and 臀部 =1,and 上身 =2);
[0129] p 动作34 =(and 屈膝 =0,and 抱头 =1,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =2);
[0130] p 动作35 =(and 屈膝 =0,and 抱头 =1,and 触膝 =1,and 肩胛骨 =0,and 臀部 =1,and 上身 =2);
[0131] p 动作36 =(and 屈膝 =1,and 抱头 =0,and 触膝 =0,and 肩胛骨 =0,and 臀部 =1,and 上身 =2);
[0132] p 动作37 =(and 屈膝 =1,and抱头 =0,y 触膝 =1,y 肩胛骨 =0,y 臀部 =1,y 上身 =2);
[0133] p 动作38 =(y 屈膝 =1,y 抱头 =1,y 触膝 =0,y 肩胛骨 =0,y 臀部 =1,y 上身 =2);
[0134] p 动作39 =(y 屈膝 =2,y 抱头 =0,y 触膝 =0,y 肩胛骨 =0,y 臀部 =1,y 上身 =2);
[0135] p 动作40 =(y 屈膝 =2,y 抱头 =0,y 触膝 =1,y 肩胛骨 =0,y 臀部 =1,y 上身 =2);
[0136] p 动作41 =(y 屈膝 =2,y 抱头 =1,y 触膝 =0,y 肩胛骨 =0,y 臀部 =1,y 上身 =2);
[0137] p 动作42 =(y 屈膝 =2,y 抱头 =1,y 触膝 =1,y 肩胛骨 =0,y 臀部 =1,y 上身 =2).
[0138] Of course, the aforementioned 动作4 to p 动作42 There are 39 remaining sit-up action categories.
[0139] Therefore, the sit-up action normativeness of the target person in each action image can be identified based on the sit-up action category of each action image.
[0140] Furthermore, one of the training processes of the following public action detection model can be, but is not limited to, as shown in the following steps.
[0141] The first step: based on the predicted category, confidence and predicted position of the regional image corresponding to the target part in each sample action image of each sample person output by the YLOL model, the sample regional image of the target part in each sample action image of each sample person is cut out from each sample action image of each sample person.
[0142] Step 2: Resize the sample region image of the target part in each sample action image of each sample person to obtain an adjusted region image of the target part in each sample action image after the resizing process. In this embodiment, the size of each sample region image is simply resized to (224, 224, 3) to adapt to the input of ResNet.
[0143] After the image size adjustment is completed, the images can be merged to obtain motion detection images of different human body parts. The merging process is shown in the third step below.
[0144] Step 3: For any sample person, the adjustment area images corresponding to the hands and head, the adjustment area images corresponding to the elbows and knees, the adjustment area images corresponding to the shoulder blades and the pad, and the adjustment area images corresponding to the buttocks and the pad in each sample action image of any sample person are merged separately, so as to obtain, after the merging process, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image and the buttocks touching the pad detection sample image corresponding to each sample action image of any sample person, and after polling the sample action images of all sample persons, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image and the buttocks touching the pad detection sample image corresponding to each sample action image of all sample persons are obtained.
[0145] After merging the adjustment area images corresponding to the hands and head, the adjustment area images corresponding to the elbows and knees, the adjustment area images corresponding to the shoulder blades and the mat, and the adjustment area images corresponding to the buttocks and the mat, a second initial training data set can be constructed based on this. The process is shown in the fourth step below.
[0146] Step 4: Use the adjustment area images corresponding to the knees, the adjustment area images corresponding to the upper body, the sample images of detection of hands holding the head, the sample images of detection of elbows touching knees, the sample images of detection of shoulder blades touching pads, and the sample images of detection of buttocks touching pads in each sample action image of each sample person to form the second initial training data set.
[0147] After obtaining the second initial training data set, the action category of the target part can be labeled, and the process is as follows.
[0148] Step 5: Perform action category labeling processing on each image in the second initial training data set to obtain the second training data set after the action category labeling processing; in the specific implementation, for any image in the second initial training data set, the action category labeling process is as follows: ① The degree of knee flexion, when the knee flexion angle is less than 80°, it is marked as insufficient knee flexion (indicated by the number 0), when the knee flexion angle is greater than 100°, it is marked as excessive knee flexion (indicated by the number 2), and the rest of the cases are marked as qualified knee flexion (indicated by the number 1); ② Whether the hands are holding the head, when the palms are close to the head, it is marked as holding the head (indicated by the number 1), otherwise it is marked as not holding the head (indicated by the number 0); ③ Whether the shoulder blade touches the pad, when the lower area of the shoulder blade is in full contact with the surface of the mat, there is no obvious lifting or hanging of the torso, it is marked as shoulder blade touching the pad (indicated by the number 1), otherwise it is marked as shoulder blade not touching the pad (indicated by the number 0); ④ Whether the buttocks leave the mat, when the buttocks When the elbows clearly touch the knees or are very close to the knees during the process of getting up, it is marked as touching the knees (indicated by the number 1), otherwise it is marked as not touching the knees (indicated by the number 0); ⑥ The state of the upper body, when the upper body is lying flat, that is, the head, shoulders, back, etc. are all close to the mat, it is marked as state 1, the upper body has just been raised between 10°-80°, the shoulder blades and back have left the mat, and the head and shoulders are raised during the movement, it is marked as state 2, the upper body is close to or reaches a vertical state between 80°-90°, and the back is almost in a vertical state, it is marked as state 3; of course, each image in the second initial training data set can be displayed, but is not limited to, so as to obtain the action category labeling result of the target part in each image in response to the action category labeling human-computer operation.
[0149] After completing the action category labeling, a second training data set can be obtained. Then, the ResNet model can be trained based on the second training data set to obtain an action recognition model after the training is completed. The training process is shown in the sixth step below.
[0150] Step 6: Use the adjustment area image corresponding to the knee, the adjustment area image corresponding to the upper body, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image in each sample action image corresponding to each sample person in the second training data set as input, and the sit-up action category of each sample person as output to train the ResNet model, so as to obtain the action detection model after the training is completed.
[0151] In this embodiment, after an image in the second training data set is input into the ResNet model, its output is one of the aforementioned 42 sit-up action categories. Therefore, after continuous training, the ResNet model can accurately identify the action status of each part; in this way, after the aforementioned training steps, the training of the ResNet model can be completed. Then, in actual use, it is only necessary to input the regional image of the target part in step S4 into the model to obtain the sit-up action category corresponding to each action image.
[0152] After obtaining the sit-up action category corresponding to each action image, the number of sit-ups of the target person can be counted based on the category, and the process is shown in the following step S5.
[0153] S5. According to the sit-up action category corresponding to each action image, the sit-up count result of the target person is obtained; in a specific implementation, for example, but not limited to, the following steps S51 to S56 can be used to obtain the sit-up count result of the target person.
[0154] S51. Determine whether there is a target image group among the plurality of action images, wherein the target image group includes at least two consecutive action images, and the sit-up action category of each action image in at least two consecutive action images is the same; in this embodiment, it is necessary to filter out action images with the same sit-up action category and which are continuous, so as to subsequently remove the same action category that appears continuously therein; optionally, the removal process is shown in the following step S52.
[0155] S52. If yes, delete the designated image in the target image group to obtain an action image set after deleting the designated image, wherein the designated image is all action images after the first action image in the target image group; in this embodiment, for consecutive action images with the same sit-up action category, only the first one is retained. In this way, after removing the consecutive action images corresponding to the same action category, the number of sit-ups can be counted, and the process is shown in the following steps S53 to S56.
[0156] S53. Initialize the number of sit-ups i to 0.
[0157] S54. Extract the first n action images from the action image set, and determine whether the sit-up action categories corresponding to the first n action images are the first standard action category, the second standard action category, the third standard action category, the second standard action category, and the first standard action category, respectively, wherein the value of n is 5; in this embodiment, when p 动作1 、p 动作2 、p 动作3 、p动作2 、p 动作1 When it appears continuously, it means that a standard sit-up action has been completed. Therefore, it is necessary to determine whether the sit-up action category corresponding to the first n action images extracted meets the above conditions. If so, it means that the target person has completed a standard and complete sit-up action, and a count is performed; otherwise, it means that the sit-up action is not standard and is not counted; then, the first n-1 action images extracted can be removed, and the first n action images can be re-extracted from the action image set for judgment, until all action image sets are polled, and the sit-up count result can be obtained; wherein, the cyclic counting process is shown in the following steps S55 and S56.
[0158] S55. If yes, increment i by 1 and delete the first n-1 action images from the action image set to obtain a new action image set.
[0159] S56. Update the action image set to the new action image set, and re-extract the first n action images from the action image set until the action image set is polled, so that after the polling is completed, the value of i is used as the sit-up counting result; in this embodiment, when the number of action images in the action image set is less than n or all the images in the action image set are polled, the counting process can be ended; at this time, the value of i can be used as the sit-up counting result.
[0160] In addition, in this embodiment, when it is determined in the aforementioned step S54 that the sit-up action category corresponding to the first n action images extracted does not meet the aforementioned conditions, an action non-standard prompt message can be output to prompt the target person; of course, an action non-standard prompt can also be given when it is identified that the sit-up action category corresponding to any action image does not belong to the aforementioned three standard action categories.
[0161] In this way, through the aforementioned steps S51 to S56 , the sit-up counting of the target person can be achieved based on the correspondence between each action image and the sit-up action category.
[0162] Therefore, the sit-up counting method based on YOLO and ResNet described in detail in the above steps S1 to S5 has the following beneficial effects:
[0163] (1) Use the YOLO algorithm to detect key areas, and then use ResNet to learn various postures of key areas. In this way, the action status of each area can be obtained, and the action can be analyzed in detail and the details of the action can be accurately analyzed.
[0164] (2) The present invention obtains a variety of sit-up action categories and subdivides the sit-up actions, thereby making the results of normative judgment more accurate.
[0165] (3) The local feature extraction method of non-key points is used to improve the accuracy and generalization ability of the model.
[0166] like Figure 2 As shown, the second aspect of this embodiment provides a hardware system for implementing the sit-up counting method based on YOLO and ResNet described in the first aspect of the embodiment, including:
[0167] The acquisition unit is used to acquire sit-up video data of the target person.
[0168] The frame segmentation unit is used to convert the sit-up video data into continuous image frames to obtain a plurality of action images.
[0169] The target detection unit is used to input each action image into the YOLO target detection model to obtain a regional image corresponding to the target part in each action image, wherein the target part includes the lying mat, hands, knees, head, elbows, shoulder blades, buttocks and upper body.
[0170] The action detection unit is used to input the area image corresponding to the target part in each action image into the action detection model to obtain the sit-up action category corresponding to each action image, wherein the action detection model is a trained ResNet model, and the sit-up action category corresponding to any action image is used to characterize whether the sit-up action of the target person in any action image is a standard action.
[0171] The counting unit is used to obtain the sit-up counting result of the target person according to the sit-up action category corresponding to each action image.
[0172] The working process, working details and technical effects of the system provided in this embodiment can be found in the first aspect of the embodiment and will not be described in detail here.
[0173] like Figure 3 As shown, the third aspect of this embodiment provides a sit-up counting device based on YOLO and ResNet. Taking the device as an electronic device as an example, it includes: a memory, a processor and a transceiver that are communicatively connected in sequence, wherein the memory is used to store computer programs, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the sit-up counting method based on YOLO and ResNet as described in the first aspect of the embodiment.
[0174] For example, the memory may include, but is not limited to, random access memory (RAM), read-only memory (ROM), flash memory, first-in first-out memory (FIFO), and / or first-in last-out memory (FILO); specifically, the processor may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor may be implemented in at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field-Programmable Gate Array), and PLA (Programmable Logic Array). Furthermore, the processor may include a main processor and a coprocessor. The main processor is a processor for processing data in an awake state, also known as a CPU (Central Processing Unit); and the coprocessor is a low-power processor for processing data in a standby state.
[0175] In some embodiments, the processor may be integrated with a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. For example, the processor may be, but is not limited to, a microprocessor of the STM32F105 series, a reduced instruction set computer (RISC) microprocessor, an X86 architecture processor, or a processor with an integrated embedded neural network processing unit (NPU); the transceiver may be, but is not limited to, a wireless fidelity (WIFI) wireless transceiver, a Bluetooth wireless transceiver, a general packet radio service technology (GPRS) wireless transceiver, a ZigBee protocol (a low-power local area network protocol based on the IEEE802.15.4 standard, ZigBee) wireless transceiver, a 3G transceiver, a 4G transceiver, and / or a 5G transceiver. In addition, the device may also include, but is not limited to, a power module, a display screen, and other necessary components.
[0176] The working process, working details and technical effects of the electronic device provided in this embodiment can be found in the first aspect of the embodiment and will not be described in detail here.
[0177] A fourth aspect of this embodiment provides a storage medium storing instructions for the sit-up counting method based on YOLO and ResNet described in the first aspect of the embodiment, that is, the storage medium stores instructions, and when the instructions are run on a computer, the sit-up counting method based on YOLO and ResNet described in the first aspect of the embodiment is executed.
[0178] The storage medium refers to a carrier for storing data, which may include but is not limited to a floppy disk, an optical disk, a hard disk, a flash memory, a USB flash drive and / or a memory stick, and the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0179] The working process, working details and technical effects of the storage medium provided in this embodiment can be found in the first aspect of the embodiment and will not be described in detail here.
[0180] A fifth aspect of this embodiment provides a computer program product comprising instructions, which, when executed on a computer, causes the computer to execute the sit-up counting method based on YOLO and ResNet as described in the first aspect of the embodiment, wherein the computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device.
[0181] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention shall be included in the scope of protection of the present invention.
Claims
1. A sit-up counting method based on YOLO and ResNet, characterized in that: include: Obtain video data of the target person performing sit-ups; Converting the sit-up video data into continuous image frames to obtain a plurality of action images; Each action image is input into the YOLO target detection model to obtain the region image corresponding to the target part in each action image, wherein the target parts include the lying mat, hands, knees, head, elbows, shoulder blades, buttocks and upper body; Input the regional image corresponding to the target part in each action image into the action detection model to obtain the sit-up action category corresponding to each action image, wherein the action detection model is a trained ResNet model, and the sit-up action category corresponding to any action image is used to indicate whether the sit-up action of the target person in any action image is a standard action; According to the sit-up action category corresponding to each action image, the sit-up counting result of the target person is obtained.
2. The method according to claim 1, characterized in that The region image corresponding to the target part in each action image is input into the action detection model to obtain the sit-up action category corresponding to each action image, including: Merge the area images corresponding to the hands and head, the area images corresponding to the elbows and knees, the area images corresponding to the shoulder blades and the mat, and the area images corresponding to the buttocks and the mat in each action image to obtain a detection image of the hands holding the head, a detection image of the elbows touching the knees, a detection image of the shoulder blades touching the mat, and a detection image of the buttocks touching the mat corresponding to each action image; Inputting the regional images corresponding to the knees and the regional images corresponding to the upper body in each action image, as well as the detection images of hands holding the head, elbows touching the knees, shoulder blades touching the pads, and buttocks touching the pads corresponding to each action image into the action detection model, thereby respectively obtaining the target person's knee flexion detection category, upper body state category, hands holding the head detection category, elbows touching the knees detection category, shoulder blades touching the pads detection category, and buttocks leaving the pad detection category in each action image; The sit-up action category of each action image is generated by using the target person's knee bending detection category, upper body state category, hands holding head detection category, elbows touching knees detection category, shoulder blades touching pad detection category and buttocks leaving pad detection category in each action image.
3. The method according to claim 1, characterized in that The sit-up action categories corresponding to any action image include a knee flexion detection category, a hands-on-head detection category, a double-elbow-to-knee-touching detection category, a shoulder blade-to-pad detection category, a buttocks-off-pad detection category, and a human upper body state category during the target person's sit-up exercise. The sit-up action categories include 42 categories, each of which includes a first standard action category, a second standard action category, and a third standard action category. The first standard action category, the second standard action category, the third standard action category, the second standard action category, and the first standard action category constitute the human action categories corresponding to a complete sit-up action. The target person's sit-up count result is obtained based on the sit-up action category corresponding to each action image, including: determining whether there is a target image group among the plurality of action images, wherein the target image group includes at least two consecutive action images, and each action image in the at least two consecutive action images has the same sit-up action category; If so, deleting the designated image in the target image group to obtain an action image set after deleting the designated image, wherein the designated image is all action images after the first action image in the target image group; Initialize the number of sit-ups i to 0; Extracting first n action images from the action image set, and determining whether the sit-up action categories corresponding to the first n action images are, in order, the first standard action category, the second standard action category, the third standard action category, the second standard action category, and the first standard action category, where the value of n is 5; If so, add 1 to i and delete the first n-1 action images from the action image set to obtain a new action image set; The action image set is updated to the new action image set, and the first n action images are extracted from the action image set again until the action image set is polled. After the polling is completed, the value of i is used as the sit-up counting result.
4. The method according to claim 3, characterized in that The method further comprises: If not, the output prompt message is that the action is not standardized.
5. The method according to claim 1, wherein After obtaining the plurality of action images, the method further includes: Performing data preprocessing on the plurality of action images to obtain the plurality of preprocessed action images, so as to input each of the preprocessed action images into a YOLO target detection model to obtain a region image corresponding to a target part in each action image, wherein the data preprocessing includes resizing and data augmentation processing; Before inputting the region image corresponding to the target part in each action image into the action detection model, the method further includes: The area image corresponding to the target part in each action image is cropped to obtain the cropped area image corresponding to the target part in each action image, so as to input the cropped area image corresponding to the target part in each action image into the action detection model to obtain the sit-up action category corresponding to each action image.
6. The method according to claim 1, characterized in that The YOLO target detection model is trained using the following method: Obtain sample video data of sit-up performances of several sample persons; Convert the sit-up sample video data of each sample person into continuous sample image frames to obtain a number of sample action images corresponding to each sample person, and use the sample action images corresponding to each sample person to form a first initial training data set; Performing data preprocessing on the first initial training data set to obtain a preprocessed initial first training data set; performing labeling processing on the target part in each sample action image in the preprocessed initial first training data set to obtain a first training data set after the labeling processing, wherein the label data of any sample action image in the first training data set includes the category and position of the target part in the any sample action image; The YOLO model is trained using each sample action image of each sample person in the first training data set as input and the predicted category, confidence, and predicted position of the regional image corresponding to the target part in each sample action image of each sample person as output, so as to obtain the YOLO target detection model after the training is completed.
7. The method according to claim 6, characterized in that The action detection model is trained using the following method: Based on the predicted category, confidence, and predicted position of the regional image corresponding to the target part in each sample action image of each sample person output by the YLOL model, a sample regional image of the target part in each sample action image of each sample person is cut out from each sample action image of each sample person; performing a resizing process on the sample region image of the target part in each sample action image of each sample person, so as to obtain an adjusted region image of the target part in each sample action image after the resizing process; For any sample person, the adjustment area images corresponding to the hands and head, the adjustment area images corresponding to the elbows and knees, the adjustment area images corresponding to the shoulder blades and the lying pad, and the adjustment area images corresponding to the buttocks and the lying pad in each sample action image of any sample person are merged respectively, so as to obtain, after the merging process, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image corresponding to each sample action image of the any sample person, and after the sample action images of all the sample persons are polled, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image corresponding to each sample action image of all the sample persons are obtained; A second initial training dataset is formed using the adjustment area images corresponding to the knees, the adjustment area images corresponding to the upper body, the hands-on-head detection sample images, the elbows-on-knees detection sample images, the shoulder blades-on-pads detection sample images, and the buttocks-on-pads detection sample images in each sample action image of each sample person; performing action category labeling processing on each image in the second initial training data set to obtain a second training data set after the action category labeling processing; The ResNet model is trained using the adjustment area image corresponding to the knee, the adjustment area image corresponding to the upper body, the hands holding the head detection sample image, the elbows touching the knees detection sample image, the shoulder blades touching the pad detection sample image, and the buttocks touching the pad detection sample image in each sample action image corresponding to each sample person in the second training data set as input, and the sit-up action category of each sample person as output, so as to obtain the action detection model after the training is completed.
8. A sit-up counting system based on YOLO and ResNet, characterized in that: include: an acquisition unit, configured to acquire video data of a target person performing sit-ups; A frame segmentation unit, configured to convert the sit-up video data into continuous image frames to obtain a plurality of action images; A target detection unit is used to input each action image into a YOLO target detection model to obtain a region image corresponding to a target part in each action image, wherein the target part includes a lying mat, hands, knees, head, elbows, shoulder blades, buttocks, and upper body; An action detection unit is configured to input a region image corresponding to a target part in each action image into an action detection model to obtain a sit-up action category corresponding to each action image, wherein the action detection model is a trained ResNet model, and the sit-up action category corresponding to any action image is used to indicate whether the sit-up action performed by the target person in the action image is a standard action; The counting unit is used to obtain the sit-up counting result of the target person according to the sit-up action category corresponding to each action image.
9. An electronic device, characterized in that: include: A memory, a processor, and a transceiver that are sequentially communicatively connected, wherein the memory is used to store a computer program, the transceiver is used to send and receive messages, and the processor is used to read the computer program and execute the sit-up counting method based on YOLO and ResNet according to any one of claims 1 to 7.
10. A computer program product comprising instructions, characterized in that When the instruction is executed on a computer, the computer is caused to execute the sit-up counting method based on YOLO and ResNet as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Sit-up test counting method and sit-up test counting system based on Quick-OpenPose model
CN111401260A