An automated pull-up image detection method based on computer vision
Patent Information
- Application Number
- CN202310830144.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-07
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-07-07
Smart Images

Figure CN116844233B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image detection method and the field of artificial intelligence, in particular to an automatic pull-up image detection method based on computer vision. Background Art
[0002] Computer vision technology is currently widely used in sports-related fields such as standardized sports guidance and unsupervised testing. The current unsupervised pull-up assessment methods or systems are cumbersome and complex. Generally, the physical test begins after the user logs in with their information. To meet the requirements for individual sub-movement video image segments in movement quality assessment, the current subject's movement images and physical test results are usually recorded, and ultimately, continuous video frames of a single pull-up movement are obtained through manual screening and cropping. This automated assessment and image acquisition method is not only inefficient and fails to meet the requirements of campus physical testing for speed, convenience, and practicality, but also fails to automatically determine and record qualified movements and individual pull-up movement images, presenting a technical challenge that urgently needs to be addressed. Summary of the Invention
[0003] In order to solve the problems existing in the background technology, the present invention provides a method for automatic pull-up image detection based on computer vision. In view of the problems of high time and labor costs, complicated physical test procedures and lack of standardization in the current physical test scenarios with or without supervision, and the need for automatic collection of motion image clips of a large number of single actions, there is an urgent need for an automated pull-up physical test assessment and image acquisition method to standardize the pull-up physical test process in unsupervised scenarios and realize the positioning and image acquisition of a single qualified pull-up action. The method of the present invention can meet the requirements of the pull-up physical test assessment for standardization and efficiency of the automatic testing process, and realize the qualified judgment statistics of the pull-up action and the automated detection of a single qualified action image.
[0004] The technical solution adopted in the present invention is:
[0005] The computer vision-based pull-up automated image detection method of the present invention comprises the following steps:
[0006] Step 1: Collect the pull-up action video to be detected, calibrate the regional position of each video frame in the pull-up action video to be detected to obtain a calibrated video frame, and crop each calibrated video frame according to the calibrated area to obtain a cropped video frame.
[0007] Step 2: For each calibrated video frame, perform posture detection, contact detection, and straight arm detection on each cropped video frame of the calibrated video frame to obtain the detection results. According to the detection results and the scene state of the previous cropped video frame, the scene state of the current cropped video frame is judged, and the number of pull-ups in the pull-up action video to be detected is detected and image processing is performed according to the scene state of the current cropped video frame until the pull-up action is completed, completing the pull-up automated image detection.
[0008] In the step 1, the regional position of each video frame in the pull-up action video to be detected is calibrated to obtain a calibrated video frame. Specifically, the area where the pull-up horizontal bar equipment is located in each video frame is divided into a movement area and a contact detection area. The movement area is specifically a rectangular area surrounded by the support rods on both sides of the horizontal bar equipment and the horizontal bar on the upper side of the horizontal bar equipment, as well as a rectangular area within a preset height above the horizontal bar. The contact detection area is specifically a rectangular area where the horizontal bar on the upper side of the horizontal bar equipment is located, and the contact detection area is located within the movement area; each calibrated video frame is cropped according to the calibrated movement area to obtain a cropped video frame.
[0009] Continuous video frames of a motion scene are input at a preset time interval of 30ms. When calibrating the area position within the video frame, the upper left corner coordinates and the lower right corner coordinates of the motion area and the contact detection area are calibrated.
[0010] In step 2, the scene states include no-one-waiting-entry state, pull-up waiting state, straight-arm hanging state, pull-up action state, and separation completion state. Step 2 is specifically as follows:
[0011] 2.1) The cropped video frames of the motion area of each calibration video frame are input into the human posture estimation model in sequence. For each calibration video frame, the human posture estimation model outputs a set of posture points of the cropped video frames of the motion area of the calibration video frame as the detection result of posture detection. According to the detection result of posture detection, the scene state of the current calibration video frame is judged as no one waiting to enter or the body waiting to start, until the scene state of the current calibration video frame is the body waiting to start state.
[0012] The posture detection algorithm is used to determine whether the person being tested in the current video frame is within the preset motion area. If this condition is met, further contact detection is performed.
[0013] 2.2) The cropped video frames of the motion area of the current frame of the pull-up waiting state and each subsequent frame of the calibration video frame are input into the human posture estimation model in sequence. For each frame of the calibration video frame, the human posture estimation model outputs a set of posture points of the cropped video frames of the motion area of the calibration video frame as the input of contact detection and straight arm detection, until the scene state of the current calibration video frame is judged to be the straight arm hanging state according to the detection results of the contact detection and straight arm detection.
[0014] 2.3) Continue to perform contact detection and straight arm detection based on the posture detection results of the cropped video frames of the motion area of the current frame and the subsequent calibration video frames. For each calibration video frame, determine the scene state of the current calibration video frame as the straight arm hanging state or the pull-up action state based on the detection result of the straight arm detection of the current calibration video frame and the scene state of the previous calibration video frame, and then perform the detection of the number of pull-ups in the pull-up action video to be detected, until the detection result of the contact detection of the current calibration video frame is that the contact condition is not met, that is, the left-hand node coordinates and the right-hand node coordinates in the posture point set of the posture detection result are not located in the contact detection area at the same time, then the scene state of the current calibration video frame is the separation end state, the pull-up action is ended, and the pull-up automatic image detection is completed.
[0015] In the step 2.1), the scene state of the current calibration video frame is judged as the no-entry state or the body-waiting-to-start state according to the detection result of the posture detection. Specifically, when the detection result of the posture detection is an empty set, the scene state of the current frame calibration video frame is the no-entry state, and the cropped video frame of the motion area of the next frame calibration video frame is continued to be input until the detection result of the posture detection is not an empty set, then the scene state of the current frame calibration video frame is the body-waiting-to-start state.
[0016] In the step 2.2), the scene state of the current calibration video frame is judged to be the straight-arm hanging state based on the detection results of contact detection and straight-arm detection. Specifically, when the left-hand node coordinates and the right-hand node coordinates of the posture detection result are within the contact detection area, it is mainly judged whether the person being tested in the current video frame is in a state of movement with both feet off the ground and hands holding the horizontal bar. Then, the detection results of the straight-arm detection of the cropped video frames of the current frame and several subsequent movement areas are continued to be obtained until the detection results of the straight-arm detection of the cropped video frames of the movement area meet the straight-arm condition, that is, the absolute value of the difference between the angle formed by the line connecting the left shoulder node, left elbow node and left wrist node in the posture point set of the posture detection result and the angle formed by the line connecting the right shoulder node, right elbow node and right wrist node and the posture angle of the preset straight-arm action are all less than the first preset threshold value, then the scene state of the current frame calibration video frame is the straight-arm hanging state.
[0017] If the detection result obtained during contact detection is that the contact condition is not met, and the scene state of the previous frame is the no-entry state or the pull-up state, then the scene state of the current frame is the pull-up state, and the contact detection area and the cropped video frame of the motion area of the next calibration video frame are input.
[0018] In the step 2.3), for each frame of the calibration video frame, the scene state of the current calibration video frame is determined to be the straight-arm hanging state or the pull-up action state according to the detection result of the straight-arm detection of the current calibration video frame and the scene state of the previous calibration video frame. Specifically, when the scene state of the previous calibration video frame is the pull-up waiting state, the straight-arm hanging state or the pull-up action state, and the detection result of the straight-arm detection is that the straight-arm condition is met, then the scene state of the current calibration video frame is the straight-arm hanging state; if the straight-arm condition is not met, then the scene state of the current calibration video frame is the pull-up action state.
[0019] The scene state attribute judgment of the current video frame realizes the automation of the physical test process. The scene state attribute of the current video frame is judged in combination with the algorithm judgment result of the current cropped video frame and the scene state attribute of the previous frame, and the pull-up automation process is standardized. The scene state attribute of the preset video frame is the no-one-waiting-entry state. The scene state attribute of the current frame is judged by the result of the algorithm output of the video frame and the scene state attribute constraint of the previous frame. The main scene state attributes are no-one-waiting-entry, pull-up waiting, straight-arm hanging, pull-up movement and end of separation. By setting the scene attribute state of the pull-up physical test video frame of a single test person, the pull-up automation process is standardized, and the pull-up physical test exercise is performed on a large number of people.
[0020] The specific conditions for judging the scene state attributes of the current frame are as follows:
[0021] The conditions for no one waiting to enter are met: posture detection fails; the conditions for pull-ups are met: posture detection is successful, horizontal bar contact detection fails, and the scene state attribute of the previous video frame is no one waiting to enter or pull-ups; the conditions for straight-arm hanging are met: posture detection is successful, contact detection is successful, and straight-arm detection is successful, and the scene state attribute of the previous video frame is pull-ups, straight-arm hanging, or pull-up movement; the conditions for pull-up movement are met: posture detection is successful and contact detection is successful, straight-arm detection fails, and the scene state attribute of the previous video frame is pull-ups, straight-arm hanging, or pull-up movement; the conditions for separation end are met: posture detection is successful. If contact detection is successful, the scene state attribute of the previous video frame must be separation end. If contact detection fails, the scene state attribute of the previous video frame must be pull-up movement or separation end.
[0022] Combining the above detection results of the video frames and the scene state attribute description of the previous video frame, the scene state attribute judgment is performed on each video frame to achieve standardization and automation of the pull-up physical test process.
[0023] In the above step 2.3), the number of pull-ups in the pull-up action video to be detected is detected. Specifically, based on the scene state of the current calibration video frame and the scene state of the previous calibration video frame, it is determined whether the current frame meets the conditions of the start frame, judgment frame, and end frame of the pull-up number detection, and then the number of pull-ups is obtained, as follows:
[0024] When the scene state of the previous calibration video frame is the straight-arm hanging state, and the scene state of the current calibration video frame is the pull-up movement state, the current calibration video frame is marked as the start frame. If the current frame is the start frame, start recording each frame of the image until the end frame is detected; if the vertical coordinate of the facial posture point in the posture point set of the cropped video frame of the motion area of the current calibration video frame is greater than the preset horizontal bar vertical coordinate, the current calibration video frame is used as the judgment frame; when the scene state of the previous calibration video frame is the pull-up movement state, and the scene state of the current calibration video frame is the straight-arm hanging state or the separation end state, the current calibration video frame is used as the end frame.
[0025] Determine whether there is a judgment frame between each adjacent start frame and end frame. If the current calibration video frame is the start frame, start recording the current calibration video frame until the current calibration video frame is used as the end frame. If the current calibration video frame is the end frame, determine whether the current end frame matches the adjacent start frame and whether there is a judgment frame between the start frame and the end frame. If the adjacent start frame is matched and the judgment frame exists, stop recording the marked video frame and save the video frame between the start frame and the end frame, and increase the number of pull-ups by one. If the adjacent start frame is matched but the judgment frame does not exist, stop recording the marked video frame and increase the number of pull-ups by zero. It is also necessary to determine whether the pull-up movement of the person being tested is qualified. The conditions include: the arms remain in a natural drooping state at the beginning of the exercise, and the specific arm posture angle is greater than 165 degrees; the arms remain in a natural drooping state at the end of the exercise, and the specific arm posture angle is greater than 165 degrees; the chin is above the horizontal bar during the exercise.
[0026] After the start frame, if the scene state attribute judgment result of the current frame is a pull-up movement and matches the start frame, the current frame is written to the video file; if the above-mentioned start frame is not matched, the current action is judged to be an invalid pull-up action, and the judgment frame and the end frame are reset; if the above-mentioned start frame is matched and there is a judgment frame between the start frame and the end frame, the qualified action count is increased by one, the image recording is ended, and the start frame and the end frame are reset after the video is saved; otherwise, the current action is judged to be an unqualified action, the image recording is ended, the video file is not saved, and the start frame and the end frame are reset; by judging whether the input video frame is a start frame, an end frame or a judgment frame, the qualified action is counted, and the qualified pull-up action image of the test person is recorded and saved as a video file.
[0027] The beneficial effects of the present invention are:
[0028] The method of the present invention can realize the image acquisition and recording of physical examinations of a large number of people in an unsupervised situation, solving the problems of cumbersome processes, low efficiency, and inability to accurately judge, locate, and capture images of qualified movements in current unsupervised physical examination scenarios. It mainly includes standardizing and automating the automatic physical examination process of pull-ups through the scene state attributes of video frames, significantly improving physical examination efficiency, and accurately judging pull-up movements that meet physical examination standards and capturing motion images. The Mediapipe algorithm is used as the human posture estimation model in the entire test process, combined with the proposed automated detection method to realize automated and standardized physical examination assessment of multiple people's pull-ups and record and save motion images of qualified movements of a single person that meet physical examination standards.
[0029] The method of the present invention solves the problems of low efficiency and high time cost caused by the large number of personnel involved and complex situations in the current pull-up physical test. At the same time, it solves the problems of high labor cost and inability to achieve automation caused by the current manual performance of a single pull-up motion image segment. It realizes automated time domain positioning and motion image collection and storage for a single pull-up sub-action, thereby improving the efficiency of the automated pull-up physical test process in an unsupervised scenario. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1 is a flow chart of scene state attribute determination of a video frame in an embodiment of the present invention;
[0031] Figure 2 Schematic diagram of human joints predicted by the Mediapipe Blazepose model for posture estimation in an embodiment of the present invention;
[0032] Figure 3 This is a flowchart of automated detection in an embodiment of the present invention;
[0033] Figure 4 is a schematic diagram of the position of the pre-calibrated area in the video frame in an embodiment of the present invention;
[0034] Figure 5 This is a flowchart of pull-up action determination, counting, and image acquisition and storage in an embodiment of the present invention;
[0035] Figure 6 1. It is a schematic diagram of the temporal positioning of the start frame and the end frame of the pull-up sub-movement in an embodiment of the present invention;
[0036] Figure 7 It is a schematic diagram of the results of the accuracy of the video frames of the automatic recording of the pull-up sub-movements in the present invention. DETAILED DESCRIPTION
[0037] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0038] The computer vision-based pull-up automated image detection method of the present invention comprises the following steps:
[0039] Step 1: Collect the pull-up action video to be detected, calibrate the regional position of each video frame in the pull-up action video to be detected to obtain a calibrated video frame, and crop each calibrated video frame according to the calibrated area to obtain a cropped video frame.
[0040] In step 1, the regional position of each video frame in the pull-up action video to be detected is calibrated to obtain a calibrated video frame. Specifically, the area where the pull-up horizontal bar equipment is located in each video frame is divided into a motion area and a contact detection area. The motion area is specifically a rectangular area surrounded by the support rods on both sides of the horizontal bar equipment and the horizontal bar on the upper side of the horizontal bar equipment, as well as a rectangular area within a preset height above the horizontal bar. The contact detection area is specifically a rectangular area where the horizontal bar on the upper side of the horizontal bar equipment is located, and the contact detection area is located within the motion area; each calibrated video frame is cropped according to the calibrated motion area to obtain a cropped video frame.
[0041] Continuous video frames of a motion scene are input at a preset time interval of 30ms. When calibrating the area position within the video frame, the upper left corner coordinates and the lower right corner coordinates of the motion area and the contact detection area are calibrated.
[0042] Step 2: For each calibrated video frame, perform posture detection, contact detection, and straight arm detection on each cropped video frame of the calibrated video frame to obtain the detection results. According to the detection results and the scene state of the previous cropped video frame, the scene state of the current cropped video frame is judged, and the number of pull-ups in the pull-up action video to be detected is detected and image processing is performed according to the scene state of the current cropped video frame until the pull-up action is completed, completing the pull-up automated image detection.
[0043] In step 2, the scene states include no one waiting to enter state, pull-up waiting state, straight arm hanging state, pull-up action state and separation end state. Step 2 is as follows:
[0044] 2.1) The cropped video frames of the motion area of each calibration video frame are input into the human posture estimation model in sequence. For each calibration video frame, the human posture estimation model outputs a set of posture points of the cropped video frames of the motion area of the calibration video frame as the detection result of posture detection. According to the detection result of posture detection, the scene state of the current calibration video frame is judged as no one waiting to enter or the body waiting to start, until the scene state of the current calibration video frame is the body waiting to start state.
[0045] The posture detection algorithm is used to determine whether the person being tested in the current video frame is within the preset motion area. If this condition is met, further contact detection is performed.
[0046] In step 2.1), the scene state of the current calibration video frame is judged as no-one-waiting-entry state or body-waiting-to-start state according to the detection result of the posture detection. Specifically, when the detection result of the posture detection is an empty set, the scene state of the current frame calibration video frame is no-one-waiting-entry state, and the cropped video frame of the motion area of the next frame calibration video frame is input until the detection result of the posture detection is not an empty set, then the scene state of the current frame calibration video frame is body-waiting-to-start state.
[0047] 2.2) The cropped video frames of the motion area of the current frame of the pull-up waiting state and each subsequent frame of the calibration video frame are input into the human posture estimation model in sequence. For each frame of the calibration video frame, the human posture estimation model outputs a set of posture points of the cropped video frames of the motion area of the calibration video frame as the input of contact detection and straight arm detection, until the scene state of the current calibration video frame is judged to be the straight arm hanging state according to the detection results of the contact detection and straight arm detection.
[0048] In step 2.2), the scene state of the current calibration video frame is judged to be the straight-arm hanging state according to the detection results of contact detection and straight-arm detection. Specifically, when the left-hand node coordinates and the right-hand node coordinates of the posture detection result are within the contact detection area, it is mainly judged whether the person under test in the current video frame is in a state of movement with both feet off the ground and hands holding the horizontal bar, then the detection results of the straight-arm detection of the cropped video frames of the current frame and several subsequent movement areas are continued to be obtained until the detection results of the straight-arm detection of the cropped video frames of the movement area meet the straight-arm condition, that is, the absolute value of the difference between the angle formed by the connecting line of the left shoulder node, left elbow node and left wrist node in the posture point set of the posture detection result and the angle formed by the connecting line of the right shoulder node, right elbow node and right wrist node and the posture angle of the preset straight-arm action are all less than the first preset threshold value, then the scene state of the current frame calibration video frame is the straight-arm hanging state; in the specific implementation, the size of the preset straight-arm action posture angle is 180°, and the first preset threshold value is 15°.
[0049] If the detection result obtained during contact detection is that the contact condition is not met, and the scene state of the previous frame is the no-entry state or the pull-up state, then the scene state of the current frame is the pull-up state, and the contact detection area and the cropped video frame of the motion area of the next calibration video frame are input.
[0050] 2.3) Continue to perform contact detection and straight arm detection based on the posture detection results of the cropped video frames of the motion area of the current frame and the subsequent calibration video frames. For each calibration video frame, determine the scene state of the current calibration video frame as the straight arm hanging state or the pull-up action state based on the detection result of the straight arm detection of the current calibration video frame and the scene state of the previous calibration video frame, and then perform the detection of the number of pull-ups in the pull-up action video to be detected, until the detection result of the contact detection of the current calibration video frame is that the contact condition is not met, that is, the left-hand node coordinates and the right-hand node coordinates in the posture point set of the posture detection result are not located in the contact detection area at the same time, then the scene state of the current calibration video frame is the separation end state, the pull-up action is ended, and the pull-up automatic image detection is completed.
[0051] In step 2.3), for each frame of the calibration video frame, the scene state of the current calibration video frame is determined to be the straight-arm hanging state or the pull-up action state according to the detection result of the straight-arm detection of the current calibration video frame and the scene state of the previous calibration video frame. Specifically, when the scene state of the previous calibration video frame is the pull-up waiting state, the straight-arm hanging state or the pull-up action state, and the detection result of the straight-arm detection is that the straight-arm condition is met, then the scene state of the current calibration video frame is the straight-arm hanging state; if the straight-arm condition is not met, then the scene state of the current calibration video frame is the pull-up action state.
[0052] The scene state attribute judgment of the current video frame realizes the automation of the physical test process. The scene state attribute of the current video frame is judged in combination with the algorithm judgment result of the current cropped video frame and the scene state attribute of the previous frame, and the pull-up automation process is standardized. The scene state attribute of the preset video frame is the no-one-waiting-entry state. The scene state attribute of the current frame is judged by the result of the algorithm output of the video frame and the scene state attribute constraint of the previous frame. The main scene state attributes are no-one-waiting-entry, pull-up waiting, straight-arm hanging, pull-up movement and end of separation. By setting the scene attribute state of the pull-up physical test video frame of a single test person, the pull-up automation process is standardized, and the pull-up physical test exercise is performed on a large number of people.
[0053] The specific conditions for judging the scene state attributes of the current frame are as follows:
[0054] The conditions for no one waiting to enter are met: posture detection fails; the conditions for pull-ups are met: posture detection is successful, horizontal bar contact detection fails, and the scene state attribute of the previous video frame is no one waiting to enter or pull-ups; the conditions for straight-arm hanging are met: posture detection is successful, contact detection is successful, and straight-arm detection is successful, and the scene state attribute of the previous video frame is pull-ups, straight-arm hanging, or pull-up movement; the conditions for pull-up movement are met: posture detection is successful and contact detection is successful, straight-arm detection fails, and the scene state attribute of the previous video frame is pull-ups, straight-arm hanging, or pull-up movement; the conditions for separation end are met: posture detection is successful. If contact detection is successful, the scene state attribute of the previous video frame must be separation end. If contact detection fails, the scene state attribute of the previous video frame must be pull-up movement or separation end.
[0055] Combining the above detection results of the video frames and the scene state attribute description of the previous video frame, the scene state attribute judgment is performed on each video frame to achieve standardization and automation of the pull-up physical test process.
[0056] In step 2.3), the number of pull-ups in the pull-up action video to be detected is detected. Specifically, based on the scene state of the current calibration video frame and the scene state of the previous calibration video frame, it is determined whether the current frame meets the conditions of the start frame, judgment frame, and end frame of the pull-up number detection, and then the number of pull-ups is obtained, as follows:
[0057] When the scene state of the previous calibration video frame is the straight-arm hanging state and the scene state of the current calibration video frame is the pull-up movement state, the current calibration video frame is marked as the start frame. If the current frame is the start frame, start recording each frame of the image until the end frame is detected; if the vertical coordinate of the facial posture point in the posture point set of the cropped video frame of the motion area of the current calibration video frame is greater than the preset horizontal bar vertical coordinate, the current calibration video frame is used as the judgment frame; when the scene state of the previous calibration video frame is the pull-up movement state and the scene state of the current calibration video frame is the straight-arm hanging state or the separation end state, the current calibration video frame is used as the end frame.
[0058] Determine whether there is a judgment frame between each adjacent start frame and end frame. If the current calibration video frame is the start frame, start recording the current calibration video frame until the current calibration video frame is used as the end frame. If the current calibration video frame is the end frame, determine whether the current end frame matches the adjacent start frame and whether there is a judgment frame between the start frame and the end frame. If the adjacent start frame is matched and the judgment frame exists, stop recording the marked video frame and save the video frame between the start frame and the end frame, and increase the number of pull-ups by one. If the adjacent start frame is matched but the judgment frame does not exist, stop recording the marked video frame and increase the number of pull-ups by zero. It is also necessary to determine whether the pull-up movement of the person being tested is qualified. The conditions include: the arms remain in a natural drooping state at the beginning of the exercise, and the specific arm posture angle is greater than 165 degrees; the arms remain in a natural drooping state at the end of the exercise, and the specific arm posture angle is greater than 165 degrees; the chin is above the horizontal bar during the exercise.
[0059] After the start frame, if the scene state attribute judgment result of the current frame is a pull-up movement and matches the start frame, the current frame is written to the video file; if the above-mentioned start frame is not matched, the current action is judged to be an invalid pull-up action, and the judgment frame and the end frame are reset; if the above-mentioned start frame is matched and there is a judgment frame between the start frame and the end frame, the qualified action count is increased by one, the image recording is ended, and the start frame and the end frame are reset after the video is saved; otherwise, the current action is judged to be an unqualified action, the image recording is ended, the video file is not saved, and the start frame and the end frame are reset; by judging whether the input video frame is a start frame, an end frame or a judgment frame, the qualified action is counted, and the qualified pull-up action image of the test person is recorded and saved as a video file.
[0060] In the specific implementation of pull-up detection, the position range of the exercise area and the horizontal bar detection area within the video frame is first calibrated. When performing a single pull-up test, the subject enters the exercise area, lifts their feet off the ground, grasps the horizontal bar, and lets their arms hang naturally. After waiting for approximately 2 seconds with their arms straight, they perform a pull-up. After the pull-up is completed, their hands leave the bar, the current subject leaves the exercise area, and the next subject enters. During the test, person information capture and matching are also required. After the scene state attributes transition from a non-pull-up state to a straight-arm hang or pull-up state, the current subject is matched through facial capture pre-information, and their corresponding database is obtained to store the pull-up test results and related data. For qualified action judgment and image acquisition, the scene attribute states of the current and previous video frames are further determined to determine whether the current frame belongs to the start or end frame, achieving sub-action frame-level positioning. Combined with the formulation of qualified action rules in the pull-up physical test, the jaw posture point coordinates output by the posture detection algorithm and the preset horizontal bar position coordinates are used to determine whether the jaw over-the-bar condition is met, and whether there is a judgment frame between the start frame and the end frame. Finally, OpenCV is used to realize the acquisition and storage of the input video frames.
[0061] The present invention proposes a standardized pull-up physical test and image acquisition method based on computer vision, which relates to the field of artificial intelligence technology, including but not limited to standardized pull-up physical test, qualified movement judgment and image recording and preservation. The purpose is to solve the problems of high cost and low efficiency in the existing pull-up automatic physical test and assessment technology and automatic recording and preservation of qualified movement images.
[0062] The present invention mainly realizes the standardized physical test of pull-ups under unsupervised operation by inputting input video frames into a deep learning model or detection algorithm in sequence, combining the output results of the detection with the scene state attributes of the previous video frame to judge the scene state attributes of the current input video frame.
[0063] The scene state attributes of the video frame are used to describe the specific stage of the pull-up test process in the current frame. The proposed scene state attributes include no one waiting to enter, pull-up waiting to start, straight arm hanging, pull-up movement and separation end; Figure 1 As shown, Figure 1 The following is a schematic diagram of a scene state attribute determination process for a video frame provided in an embodiment, wherein a series of detections and scene state attributes of an input video frame are specifically described as follows:
[0064] First, the original input video frame is cropped through the preset motion area to obtain a cropped frame; the cropped frame is detected, including posture detection, contact detection and straight arm detection, and the above detections have only two results: success or failure; posture detection is mainly achieved by inputting the cropped frame into the human posture estimation model to detect whether there is a person under test in the current motion area. The condition for successful detection is that the posture point set output by the human posture estimation model is not an empty set; the above human posture model is implemented by posture estimation algorithm training, the input is a video frame image, and the output result is a set of posture point coordinates. If a human body is detected in the image, the 17 human posture point coordinate sets of the current posture estimation are output. If no human body is detected, the output set is an empty set; such as Figure 2 The figure shows a schematic diagram of human joint points predicted based on the posture estimation algorithm.
[0065] The prerequisite for contact detection is the success of posture detection. Its main purpose is to detect whether a single person in the current motion area is holding the horizontal bar with both hands. The condition for successful contact detection is that the left hand node coordinate (x 10 、y 10 ), the right hand node coordinates (x9, y9) are located in the preset horizontal bar contact detection area, combined with Figure 2 , the left-hand node coordinate (x 10 、y 10 ), the right hand node coordinates (x9, y9) are the coordinates of node 10 and node 9 respectively.
[0066] The prerequisite for straight arm detection is that the conditions for successful contact detection have been met. The main purpose is to detect whether the current person being tested is in a straight arm hanging state. The success condition must meet the following conditions: the absolute value of the difference between the angle formed by the connection line of the left shoulder node, left elbow node and left wrist node in the posture point output by the human posture estimation model and the angle formed by the connection line of the right shoulder node, right elbow node and right wrist node and the corresponding posture angle in the standard straight arm action is less than the first preset threshold, wherein the corresponding posture angle in the standard straight arm action is 180°, combined with Figure 2 The left shoulder node, left elbow node and left wrist node are specifically node 6, node 8 and node 10 respectively, and the right shoulder node, right elbow node and right wrist node are specifically node 5, node 7 and node 9 respectively.
[0067] After the cropped frame is tested and the detection results are obtained, the scene state attributes of the current frame are determined in combination with the scene state attributes of the previous frame. The specific judgment conditions and scene state descriptions of the above-mentioned scene state attributes are as follows:
[0068] The criteria for the no-person-waiting-entry state are: posture detection failure, specifically, the scenario is that the subject is not detected in the current frame or the subject is not in the preset pull-up movement area.
[0069] The pull-up waiting state meets the following criteria: posture detection succeeds, contact detection fails, and the scene state attribute of the previous frame is no one waiting state or pull-up waiting state; the specific scene is that the subject is in the pull-up movement area in the current frame, and his hands are not in contact with the horizontal bar.
[0070] The straight-arm hang test meets the following criteria: successful posture detection, successful contact detection, and successful straight-arm detection, and the scene state attributes of the previous video frame are pull-up waiting state, straight-arm hang, or pull-up movement; the specific scene is that in the current frame, the subject's feet are off the ground, with both hands tightly gripping the horizontal bar, and is in the state of preparing to do a pull-up or just finishing a pull-up.
[0071] The pull-up movement meets the following criteria: posture detection and contact detection are successful, straight arm detection fails, and the scene state attributes of the previous video frame are pull-up waiting state, straight arm hanging, or pull-up movement; specifically, in the current frame, the subject is exerting force with both arms and performing a pull-up movement;
[0072] The preset standard for the end of disengagement is: posture detection is successful. If the contact detection is successful, the scene state attribute of the previous video frame must be disengagement end; if the contact detection fails, the scene state attribute of the previous video frame must be pull-up movement or disengagement end; the specific scene is that the subject has completed this round of physical examination, his hands are not in contact with the horizontal bar, and are still in the preset pull-up movement area.
[0073] like Figure 3 FIG. 1 is a flowchart of an automated pull-up physical test provided by an embodiment, which specifically includes the following steps:
[0074] S100: Receive a video frame transmitted by a capture device, and pre-calibrate the positions of a current motion area and a horizontal bar area in the transmitted video frame.
[0075] Specifically, turn on the acquisition device, adjust the transmitted video frame image, pre-set the position of the pull-up movement area and contact detection area in the image, obtain the vertical coordinate Y0 of the horizontal bar position in the image, the tested person should be outside the pre-set movement area, and the scene state attribute of the video frame should be no one waiting to enter, such as Figure 4 FIG. 1 is a schematic diagram of the position of a preset area in a video frame according to an embodiment of the present invention.
[0076] S200: After the setting of the exercise area and the horizontal bar contact detection area is completed, a test subject in a waiting state enters the exercise area alone.
[0077] Specifically, after confirming that the system is turned on, the person in the unmanned state enters the motion area and looks directly at the image acquisition device; the scene state attribute of the current frame is the pull-up state, and the video frame is cropped according to the pre-calibrated motion area coordinates (x1, y1), (x2, y2) and input into the posture estimation model. The human posture point coordinate set output by the model is not an empty set. At the same time, it is further judged whether the current frame meets the conditions for successful contact detection, that is, the left hand posture point coordinates (x1, y1) output by the human posture estimation model are 10 ,y 10 ) and the right hand posture point coordinates (x9, y9) are located in the preset horizontal bar contact detection area (x3, y3), (x4, y4), and the specific formula is as follows:
[0078] x3<x 10 <x9<x4
[0079] y4<y 10 <y3
[0080] y4<y9<y3
[0081] S300: The testee lifts his feet off the ground, holds the horizontal bar tightly with both hands, and lets his arms hang naturally, ready to start the pull-up movement.
[0082] Specifically, the current subject holds the bar with both hands, meeting the contact detection success condition, performs face capture and information matching, and at the same time determines whether the current frame meets the condition that the scene state attribute is a pull-up movement. That is, when the scene state attribute of the current frame is not the end of separation, the straight arm detection is successful. Specifically, the absolute value of the difference between the angle formed by the line connecting the left shoulder node, left elbow node, and left wrist node in the posture point output by the human posture estimation model and the angle formed by the line connecting the right shoulder node, right elbow node, and right wrist node and the corresponding posture angle in the standard straight arm movement is less than the first preset threshold of 15°, and the formula is satisfied as follows:
[0083] 180°-∠ 5-7-9 <15°
[0084] 180°-∠ 6-8-10 <15°
[0085] If the straight arm detection success condition is met, the scene state attribute of the current frame is set to straight arm suspension; if not, the scene state attribute of the current frame is set to pull-up movement.
[0086] S400, the subject performs a pull-up action, and at the same time determines the start frame, end frame and judgment frame in the input frame according to the scene state attributes of the input video frame, thereby achieving qualified judgment of the pull-up action and image acquisition and storage.
[0087] In the pull-up test, the conditions for a qualified pull-up action are: the arms remain in a natural drooping state at the beginning of the exercise, the arms remain in a natural drooping state at the end of the exercise, and the chin is above the horizontal bar during the exercise; if the input video frame does not meet the conditions for successful straight arm detection, the current subject is performing a pull-up action. By combining the scene state attributes of each input video frame with the scene state attributes of its previous frame, the start frame, end frame, and judgment frame are detected to achieve the judgment, positioning, and image acquisition and storage of a single qualified pull-up action, such as Figure 5 The figure is a flowchart of the qualified pull-up action judgment counting and image acquisition and storage provided by an embodiment of the present invention, which is specifically as follows:
[0088] The scene state attributes of each input video frame are judged, and combined with the scene state attributes of the previous frame, it is detected whether the current frame is a start frame, end frame or judgment frame.
[0089] Specifically, if the scene state attribute of the previous frame is straight-arm hanging and the scene state attribute judgment result of the current frame is pull-up movement, the current frame is marked as the start frame; if the scene state attribute of the previous frame is pull-up movement and the scene state attribute judgment result of the current frame is not pull-up movement, the current frame is marked as the end frame; if the vertical coordinate y of the facial posture point obtained by the posture estimation model in the current frame is greater than the preset horizontal bar vertical coordinate Y0, the current frame is marked as the judgment frame, where y is calculated based on the vertical coordinates of nodes 3 and 4 output by the posture estimation model, and the specific formula is:
[0090] y=(y3+y4) / 2
[0091] y>Y0
[0092] If the current frame is detected as the start frame, it is determined whether the person under test in the start frame has his arms straight, and each frame of the image is recorded to a video file until the end frame is detected.
[0093] If the current frame is detected as the end frame, check whether there is a start frame that matches it. If a start frame that matches it is found, further determine whether there is a judgment frame and the person being tested is in a straight arm position in the matched start frame and the current end frame. If so, the current action is judged to be qualified, the current image acquisition record is ended, and the start frame, end frame and judgment frame are saved and reset; if they do not exist, the current image acquisition record is ended and the video file is not saved; if the end frame does not detect a start frame that matches it, the judgment frame and end frame are directly reset.
[0094] If it is detected that the current frame is not the start frame or the end frame, the next video frame is input in a loop.
[0095] like Figure 6 and Figure 7 As shown, Figure 6It is a schematic diagram of the time domain positioning of the start frame and end frame of the pull-up sub-action in an embodiment of the present invention, and the positioning of the sub-action and the recording of continuous video frames are achieved through the start frame and end frame. Figure 7 It is a schematic diagram of the results of the accuracy of the video frames of the automatic recording of the pull-up sub-movements in the present invention.
[0096] S500, the subject finishes the pull-up movement and removes both hands from the horizontal bar;
[0097] Specifically, after the pull-up assessment action of the current person being tested is completed, the number of qualified actions of the current person being tested is recorded, and the motion images of the qualified actions are saved as video files, so as to realize automatic recording of the motion images as the original data for subsequent action evaluation.
[0098] S600: The current person being tested leaves the exercise area, and the next person being tested enters the exercise area.
[0099] Specifically, after the current test subject completes the physical test, the test results and the collected images are saved, and the test subject leaves the sports venue. The next test subject enters and the cycle repeats in sequence, thereby realizing the automation of the pull-up physical test.
[0100] The standardized pull-up physical test assessment and automated image acquisition method based on computer vision proposed in the present invention can solve the technical problems of low efficiency and non-standardization of the testing process of existing unsupervised pull-up physical test equipment and the acquisition and preservation of motion images; the implementation of the technical solution of the present invention can solve the problem of low efficiency of pull-up physical tests for large-scale physical test personnel, and realize the automated recording and preservation of motion images of qualified movements, so as to facilitate the subsequent analysis and evaluation of qualified movements.
Claims
1. A computer vision-based automatic image detection method for pull-ups, characterized by: The method comprises the following steps: Step 1: Capture a pull-up action video to be detected, perform regional position calibration on each video frame in the pull-up action video to obtain a calibrated video frame, and crop each calibrated video frame according to the calibrated area to obtain a cropped video frame; Step 2: For each calibration video frame, perform posture detection, contact detection, and straight arm detection on each cropped video frame of the calibration video frame to obtain detection results. The scene state of the current cropped video frame is determined based on the detection results and the scene state of the previous cropped video frame. The number of pull-ups in the pull-up action video to be detected is detected and image processing is performed based on the scene state of the current cropped video frame until the pull-up action is completed, completing the pull-up automated image detection; In the step 1, the regional position of each video frame in the pull-up action video to be detected is calibrated to obtain a calibrated video frame. Specifically, the area where the pull-up horizontal bar equipment is located in each video frame is divided into a movement area and a contact detection area. The movement area is specifically a rectangular area surrounded by the support rods on both sides of the horizontal bar equipment and the horizontal bar on the upper side of the horizontal bar equipment, and a rectangular area within a preset height above the horizontal bar. The contact detection area is specifically a rectangular area where the horizontal bar on the upper side of the horizontal bar equipment is located. The contact detection area is located within the movement area; each calibrated video frame is cropped according to the calibrated movement area to obtain a cropped video frame; In step 2, the scene states include no-one-waiting-entry state, pull-up waiting state, straight-arm hanging state, pull-up action state, and separation completion state. Step 2 is specifically as follows: 2.1) The cropped video frames of the motion region of each calibration video frame are sequentially input into the human pose estimation model. For each calibration video frame, the human pose estimation model outputs a set of pose points of the cropped video frames of the motion region of the calibration video frame as the pose detection result. Based on the pose detection result, the scene state of the current calibration video frame is determined to be either the no-entry state or the vehicle-initiating state, until the scene state of the current calibration video frame is the vehicle-initiating state. 2.2) The cropped video frames of the motion region of the current frame and subsequent calibration video frames in the pull-up standby state are sequentially input into the human pose estimation model. For each calibration video frame, the human pose estimation model outputs a set of pose points of the cropped video frames of the motion region of the calibration video frame as input for contact detection and straight arm detection, until the scene state of the current calibration video frame is determined to be the straight arm hanging state based on the detection results of contact detection and straight arm detection. 2.3) Continue contact detection and straight-arm detection based on the posture detection results of the cropped video frames of the motion area of the current frame and subsequent calibration video frames. For each calibration video frame, determine the scene state of the current calibration video frame as a straight-arm hanging state or a pull-up action state based on the detection result of the straight-arm detection of the current calibration video frame and the scene state of the previous calibration video frame, and then detect the number of pull-ups in the pull-up action video to be detected. Until the detection result of the contact detection of the current calibration video frame is that the contact condition is not met, that is, the left-hand node coordinates and the right-hand node coordinates in the posture point set of the posture detection result are not simultaneously located in the contact detection area, then the scene state of the current calibration video frame is the disengagement end state, the pull-up action is completed, and the pull-up automated image detection is completed; After the settings of the exercise area and the horizontal bar contact detection area are completed, the tested person in the waiting state enters the exercise area alone; Specifically, after confirming that the system is turned on, the person in the unmanned state enters the motion area and looks directly at the image acquisition device; the scene state attribute of the current frame is the pull-up state, and the video frame is cropped according to the pre-calibrated motion area coordinates (x1, y1), (x2, y2) and input into the posture estimation model. The human posture point coordinate set output by the model is not an empty set. At the same time, it is further judged whether the current frame meets the conditions for successful contact detection, that is, the left hand posture point coordinates (x1, y1) output by the human posture estimation model are 10 ,y 10 ) and the right hand posture point coordinates (x9, y9) are located in the preset horizontal bar contact detection area (x3, y3), (x4, y4), and the specific formula is as follows: x3<x 10 <x9<x4 y4<y 10 <y3 y4<y9<y3.
2. The computer vision-based automatic pull-up image detection method according to claim 1, characterized in that: In the step 2.1), the scene state of the current calibration video frame is judged as the no-entry state or the body-waiting-to-start state according to the detection result of the posture detection. Specifically, when the detection result of the posture detection is an empty set, the scene state of the current frame calibration video frame is the no-entry state, and the cropped video frame of the motion area of the next frame calibration video frame is continued to be input until the detection result of the posture detection is not an empty set, then the scene state of the current frame calibration video frame is the body-waiting-to-start state.
3. The computer vision-based automatic pull-up image detection method according to claim 1, characterized in that: In the step 2.2), the scene state of the current calibration video frame is judged to be the straight-arm hanging state based on the detection results of contact detection and straight-arm detection. Specifically, when the left-hand node coordinates and the right-hand node coordinates of the posture detection result are within the contact detection area, the detection results of the straight-arm detection of the cropped video frames of the current frame and several subsequent motion areas are continued to be obtained until the detection results of the straight-arm detection of the cropped video frames of the motion area meet the straight-arm condition, that is, the absolute value of the difference between the angle formed by the connecting line of the left shoulder node, the left elbow node and the left wrist node in the posture point set of the posture detection result and the angle formed by the connecting line of the right shoulder node, the right elbow node and the right wrist node and the posture angle of the preset straight-arm action are both less than the first preset threshold value, then the scene state of the current frame calibration video frame is the straight-arm hanging state.
4. The computer vision-based pull-up automated image detection method according to claim 1, characterized in that: In the step 2.3), for each frame of the calibration video frame, the scene state of the current calibration video frame is determined to be the straight-arm hanging state or the pull-up action state according to the detection result of the straight-arm detection of the current calibration video frame and the scene state of the previous calibration video frame. Specifically, when the scene state of the previous calibration video frame is the pull-up waiting state, the straight-arm hanging state or the pull-up action state, and the detection result of the straight-arm detection is that the straight-arm condition is met, then the scene state of the current calibration video frame is the straight-arm hanging state; if the straight-arm condition is not met, then the scene state of the current calibration video frame is the pull-up action state.
5. The computer vision-based automatic pull-up image detection method according to claim 1, characterized in that: In the above step 2.3), the number of pull-ups in the pull-up action video to be detected is detected. Specifically, based on the scene state of the current calibration video frame and the scene state of the previous calibration video frame, it is determined whether the current frame meets the conditions of the start frame, judgment frame, and end frame of the pull-up number detection, and then the number of pull-ups is obtained, as follows: When the scene state of the previous calibration video frame is the straight-arm hanging state and the scene state of the current calibration video frame is the pull-up movement state, the current calibration video frame is marked as the start frame; if the vertical coordinate of the facial posture point in the posture point set of the cropped video frame of the movement area of the current calibration video frame is greater than the preset horizontal bar vertical coordinate, the current calibration video frame is used as the judgment frame; when the scene state of the previous calibration video frame is the pull-up movement state and the scene state of the current calibration video frame is the straight-arm hanging state or the breakaway end state, the current calibration video frame is used as the end frame; Determine whether there is a judgment frame between each adjacent start frame and end frame. If the current calibration video frame is the start frame, start recording the current calibration video frame until the current calibration video frame is the end frame. If the current calibrated video frame is the end frame, determine whether the current end frame matches the adjacent start frame and whether there is a judgment frame between the start frame and the end frame. If the adjacent start frame is matched and the judgment frame exists, stop recording the marked video frame and save the video frames between the start frame and the end frame, and the number of pull-ups is increased by one; if the adjacent start frame is matched but the judgment frame does not exist, stop recording the marked video frame and the number of pull-ups is increased by zero.