Pull-up counting method, device, medium, product and system

By introducing a pull-up human body detection network of multiple feature parser and feature extraction block, the problems of high cost and complex installation of sensor counting equipment are solved, and accurate and low-cost pull-up action counting in ordinary teaching places are achieved.

CN120260130APending Publication Date: 2025-07-04CHINA TELECOM CORP LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510371367.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

In the existing pull-up counting methods, sensor counting equipment is expensive and cumbersome to install, so it cannot be applied to ordinary teaching places.

Method used

A pull-up human body detection network based on a multi-channel feature parser and a multi-channel feature extraction block is used to generate a heat map through convolution processing, feature extraction and upsampling processing, identify target points and calculate the number of actions, and use mobile phones or ordinary camera equipment to collect images to lower the equipment threshold.

Benefits of technology

It improves the accuracy and robustness of key point detection, reduces equipment costs, is suitable for ordinary teaching places, and realizes real-time or near-real-time pull-up action counting and quality feedback.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260130A_ABST
    Figure CN120260130A_ABST
Patent Text Reader

Abstract

The invention provides a pull-up counting method and device, a medium, a product and a system. According to the method, more details and global information are captured through a multi-path feature analyzer and a multi-path feature extraction block, so that the detection precision of key point locations is improved, and the multi-path feature analyzer enhances extraction of local features through parallel strategies such as maximum pooling and average pooling; the convolutional structure in the multi-path feature extraction block is beneficial to further fusion and extraction of features, the accuracy of key point detection is ensured, the network can more accurately judge the completion standard of the pull-up action by calculating the angle change between target point locations, image acquisition and action recognition can be completed by using a mobile phone or common camera equipment, and the accuracy of key point detection is improved. Compared with the prior art, expensive special equipment is not needed, the application threshold is lowered, and therefore the problems that a sensor is adopted for counting in a pull-up counting mode in an existing scheme, but the sensor is high in price, tedious in installation and not suitable for a common teaching place are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video processing. Specifically, it relates to a method for counting pull-ups, a device for counting pull-ups, a computer-readable storage medium, a computer program product, and a system for counting pull-ups. Background Art

[0002] Pull-ups are a common upper body strength exercise in physical fitness training, which can exercise the upper limbs and shoulder muscles. The pull-up action consists of multiple steps. The person being tested needs to complete actions such as tightly gripping the horizontal bar, suspending the body in the air, swinging the body backward, exerting force with both arms, pulling the body up, pausing, swinging the body downward, and returning to the starting position. Among them, exerting force with both arms is the key action and also the basis for judging the count of pull-ups.

[0003] Currently, the methods for counting pull-ups include manual counting and sensor counting. In manual counting, the teacher needs to keep an eye on the person being tested all the time, which is very laborious and prone to subjective judgment errors. In sensor counting, the sensors are relatively expensive, and the installation of sensors is cumbersome, making it inapplicable to ordinary teaching places. Summary of the Invention

[0004] The main objective of the present application is to provide a method for counting pull-ups, a device for counting pull-ups, a computer-readable storage medium, a computer program product, and a system for counting pull-ups, so as to at least solve the problem that the existing pull-up counting method uses sensor counting, but the sensors are relatively expensive and the installation of sensors is cumbersome, making it inapplicable to ordinary teaching places.

[0005] To achieve the above objective, according to one aspect of the present application, a method for counting pull-ups is provided. The method includes: acquiring multiple frames of images including pull-up actions; performing preset processing on the multiple frames of images to obtain a heat map for predicting each human body point position of the person being tested, and extracting the target point positions of the person being tested from the heat map, where the pull-up human detection network is constructed based on a multi-path feature resolver and a multi-path feature extraction block, and the preset processing includes convolutional processing, feature extraction processing, and upsampling processing; identifying whether the person being tested has completed a pull-up action according to the association information between the target point positions of the person being tested in the multiple frames of images, and calculating and displaying the number of pull-up actions completed by the person being tested, where the association information includes the angular change between the target point positions.

[0006] Optionally, performing a preset process on the multi-frame images to obtain a heat map for predicting each human body point of the subject, including: respectively performing two convolutional processes on the multi-frame images to obtain a first convolutional result and a second convolutional result, wherein the target parameters of the convolutional structure for outputting the second convolutional result are twice the target parameters of the convolutional structure for outputting the first convolutional result, and the target parameters of the convolutional structure include the stride and the number; using a first multi-path feature extraction block to process the first convolutional result and the second convolutional result to obtain two first-layer feature extraction results, namely a first feature extraction result and a second feature extraction result; performing different convolutional processes on the second feature extraction result to obtain a corresponding number of intermediate convolutional results, and using a second multi-path feature extraction block to process the first feature extraction result and all the intermediate convolutional results to obtain a corresponding number of second-layer feature extraction results, the number of the intermediate convolutional results is determined according to the number of channels of the second multi-path feature extraction block, and the number of channels of the second multi-path feature extraction block is a preset multiple of the number of channels of the first multi-path feature extraction block; using a feature extraction module to process the second-layer feature extraction results respectively to obtain a corresponding number of extraction results, and using an upsampling layer to perform upsampling processing on at least part of the extraction results to obtain a corresponding number of upsampling results, and performing a summation process on the extraction results that are not subjected to upsampling processing and the upsampling results to obtain a final prediction result, and generating a heat map of the human body points based on the final prediction results of the human body points.

[0007] Optionally, using a first multi-path feature extraction block to process the first convolutional result and the second convolutional result to obtain two first-layer feature extraction results, including: using two feature extraction modules to process the first convolutional result and the second convolutional result to obtain a first processing result and a second processing result; using two downsampling layers to perform downsampling processing on the first processing result to obtain a first downsampling processing result, and using the upsampling layer to perform upsampling processing on the second processing result to obtain a first upsampling processing result; performing an averaging process on the first upsampling processing result and the first processing result to obtain a first average result, and using the ReLU function to process the first average result to obtain the first feature extraction result; performing an averaging process on the first downsampling processing result and the second processing result to obtain a second average result, and determining the second feature extraction result as the second average result.

[0008] Optionally, in the process of using the first multi-channel feature extraction block to process the first convolution result and the second convolution result to obtain two-way first-layer feature extraction results, the method includes: using a convolution structure to process the first convolution result to obtain a sixth convolution result; using a multi-channel feature parser to process the sixth convolution result to obtain a first processing result.

[0009] Optionally, using a multi-channel feature parser to process the sixth convolution result to obtain the first processing result includes: performing channel maximum pooling on the sixth convolution result to obtain a maximum pooling result, and performing channel average pooling on the sixth convolution result to obtain an average pooling result; respectively performing convolution processing on the sixth convolution result, the maximum pooling result, and the average pooling result to obtain a seventh convolution result, an eighth convolution result, and a ninth convolution result; using the Sigmoid function to process the seventh convolution result, the eighth convolution result, and the ninth convolution result respectively to obtain a first function result, a second function result, and a third function result; performing dot multiplication on the first function result and the sixth convolution result to obtain a first dot multiplication result, performing dot multiplication on the second function result and the sixth convolution result to obtain a second dot multiplication result, and performing dot multiplication on the third function result and the sixth convolution result to obtain a third dot multiplication result; determining the first processing result as the average value of the first dot multiplication result, the second dot multiplication result, and the third dot multiplication result.

[0010] Optionally, according to the association information between the target points of the subject in the multi-frame images, identifying whether the subject has completed a chin-up action includes: connecting the shoulder point, elbow point, and wrist point in the target points into a target broken line; when the included angle of the target broken line is less than a first preset included angle, determining that the subject's arm is in a bent state; when the included angle of the target broken line is greater than a second preset included angle, determining that the subject's arm is in a straight state, and the second preset included angle is less than the first preset included angle; when it is determined that the state of the subject's arm is converted between the bent state and the straight state, and there is no point in the target points that coincides with the ground, determining that the subject has completed the chin-up action.

[0011] According to another aspect of the present application, a pull-up counting device is provided, including: an acquisition unit for acquiring multiple frames of images including pull-up actions; a first processing unit for performing preset processing on the multiple frames of images to obtain a heat map for predicting each human body point position of the person to be measured, and extracting the target point positions of the person to be measured from the heat map, wherein the pull-up human body detection network is constructed based on a multi-path feature parser and a multi-path feature extraction block, and the preset processing includes convolutional processing, feature extraction processing, and upsampling processing; a second processing unit for identifying whether the person to be measured has completed a pull-up action according to the association information between the target point positions of the person to be measured in the multiple frames of images, and calculating and displaying the number of pull-up actions completed by the person to be measured, where the association information includes the angular change between the target point positions.

[0012] According to still another aspect of the present application, a computer-readable storage medium is provided, the computer-readable storage medium includes a stored program, wherein when the program runs, it controls the device where the computer-readable storage medium is located to execute any one of the methods.

[0013] According to yet another aspect of the present application, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the steps of any one of the methods.

[0014] According to yet another aspect of the present application, a pull-up counting system is provided, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods.

[0015] Applying the technical solution of the present application, through the multi-path feature parser and the multi-path feature extraction block, the network can parse images from different angles and scales, capture more details and global information, thereby improving the detection accuracy of key point positions. The multi-path feature parser enhances the extraction of local features through strategies such as parallel max pooling and average pooling; the convolutional structure in the multi-path feature extraction block helps further fusion and extraction of features, ensuring the accuracy of key point detection. By calculating the angular change between target point positions, the network can more accurately judge the completion standard of pull-up actions. Image acquisition and action recognition can be completed using a mobile phone or a common imaging device, without the need for expensive dedicated equipment, reducing the application threshold, thus solving the problem that the existing pull-up counting method uses sensors for counting, but the sensors are expensive and the installation is cumbersome, and it is not applicable to ordinary teaching places. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The accompanying drawings of the specification, which form a part of this application, are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation to this application. In the drawings:

[0017] Figure 1 A flowchart showing a method for counting pull-ups provided according to an embodiment of this application is shown;

[0018] Figure 2 A schematic diagram of a pull-up human body detection network provided according to an embodiment of this application is shown;

[0019] Figure 3 A schematic diagram of key points of a human body provided according to an embodiment of this application is shown;

[0020] Figure 4 A schematic diagram showing the change of key points when a human body performs a pull-up action provided according to an embodiment of this application is shown;

[0021] Figure 5 A schematic diagram of a partial structure in a pull-up human body detection network provided according to an embodiment of this application is shown;

[0022] Figure 6 A structural block diagram of a device for counting pull-ups provided according to an embodiment of this application is shown. Detailed implementation manners

[0023] It should be noted that, without conflict, the embodiments in this application and the features in the embodiments can be combined with each other. The following will refer to the accompanying drawings and combine with the embodiments to detail this application.

[0024] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0025] It should be noted that the terms "first", "second", etc. in the description, claims and above-mentioned drawings of this application are used to distinguish similar objects, and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances for the embodiments of this application described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device comprising a series of steps or units does not necessarily limit to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products or devices.

[0026] As introduced in the background art, the current pull-up counting methods include manual counting and sensor counting. In manual counting, the teacher needs to keep an eye on the person being tested all the time, which is very laborious and prone to subjective judgment errors. In sensor counting, the sensors are relatively expensive and the installation is cumbersome, and it cannot be applied to ordinary teaching places. To solve the problem that the existing pull-up counting method uses sensor counting, but the sensors are relatively expensive and the installation is cumbersome and cannot be applied to ordinary teaching places, the embodiments of this application provide a pull-up counting method, a pull-up counting device, a computer-readable storage medium, a computer program product and a pull-up counting system.

[0027] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention.

[0028] In this embodiment, a pull-up counting method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0029] Figure 1 is a flowchart showing a pull-up counting method provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0030] Step S101, obtaining multiple frames of images containing pull-up actions;

[0031] Step S102, performing preset processing on the above-mentioned multiple frames of images to obtain a heat map for predicting each human body point of the person being tested, and extracting the target points of the person being tested from the above-mentioned heat map. Among them, the above-mentioned pull-up human body detection network is constructed based on a multi-path feature resolver and a multi-path feature extraction block, and the above-mentioned preset processing includes convolution processing, feature extraction processing and upsampling processing;

[0032] Among them, performing preset processing on the above-mentioned multi-frame images to obtain a heat map for predicting each human body point of the subject includes the following steps:

[0033] Step S201: Perform two types of convolution processing on the above-mentioned multi-frame images respectively to obtain a first convolution result and a second convolution result. Among them, the target parameters of the convolution structure for outputting the second convolution result are twice the target parameters of the convolution structure for outputting the first convolution result. The above-mentioned target parameters of the convolution structure include the stride and the number;

[0034] Step S202: Use a first multi-channel feature extraction block to process the first convolution result and the second convolution result to obtain two first-layer feature extraction results, namely a first feature extraction result and a second feature extraction result;

[0035] Among them, using a first multi-channel feature extraction block to process the first convolution result and the second convolution result to obtain two first-layer feature extraction results includes: using two feature extraction modules to process the first convolution result and the second convolution result to obtain a first processing result and a second processing result; using two downsampling layers to perform downsampling processing on the first processing result to obtain a first downsampling processing result, using the above-mentioned upsampling layer to perform upsampling processing on the second processing result to obtain a first upsampling processing result; performing averaging processing on the first upsampling processing result and the first processing result to obtain a first average result, and using the ReLU function to process the first average result to obtain the first feature extraction result; performing averaging processing on the first downsampling processing result and the second processing result to obtain a second average result, and determining the second feature extraction result as the second average result.

[0036] Specifically, the first convolution result and the second convolution result are processed by two different feature extraction modules respectively, where the parameters of the second convolution result are twice those of the first convolution result, which helps the network learn and extract features of different scales. The first convolution result retains more details, while the second convolution result focuses on more global features. The combination of the two can enhance the network's understanding of the details of the chin-up action. The downsampling layer processes the first processing result, which can reduce the size of the feature map, help extract more abstract and high-level features, and improve the computational efficiency of the model. The upsampling layer processes the second processing result, which can restore or enhance the resolution of the feature map, thus maintaining fine detail information. Averaging the results of upsampling and downsampling processing with the original processing results and then activating through the ReLU function can effectively fuse feature information of various scales and enhance the accuracy of the model for key point detection. Averaging the first upsampling processing result and the first processing result, as well as the first downsampling processing result and the second processing result, can dynamically fuse information from different processing stages, ensure the consistency and integrity of information when the network processes multi-scale features, and thus obtain a more comprehensive and balanced feature representation. This multi-level feature extraction and processing strategy can make the model more robust, adapt to key point detection under different shooting conditions, and at the same time, through flexible upsampling and downsampling operations, enable the model to maintain good detection performance on input images of different resolutions. The serial and parallel design of the feature extraction modules, as well as the use of downsampling and upsampling layers, can increase the depth and width of the network, which not only helps capture more complex action patterns but also improves the generalization ability of the model, enabling it to handle more variable action details when detecting chin-ups. Averaging processing combined with ReLU function activation is a simple and effective feature extraction method. While maintaining the complexity of the model, it improves computational efficiency, reduces resource consumption, enables the model to run on mobile devices, reduces application costs, and improves practicality.

[0037] In addition, in the process of using the first multi-path feature extraction block to process the above first convolution result and second convolution result to obtain two first-layer feature extraction results, the above method includes: processing the above first convolution result using a convolution structure to obtain a sixth convolution result; processing the above sixth convolution result using a multi-path feature parser to obtain a first processing result.

[0038] Multiple applications of the convolutional structure can increase the depth of the network, which helps to extract deeper features. The processing from the first convolutional result to the sixth convolutional result enables the network to learn more complex image features, which is particularly important in human pose analysis and action recognition, and can help the model to more accurately understand the human body structure and dynamic changes. By processing the first convolutional result through the convolutional structure, specific features can be enhanced while retaining the detailed information in the image. This retention of detailed information is crucial for key point detection because human key points often need to be accurately identified from the image background, and the retention of detailed information can improve the accuracy of detection. The multi-path feature parser can parse and fuse features from different dimensions by processing different pooling strategies (such as channel max pooling and channel average pooling) in parallel, resulting in the first processing result. This fusion mechanism enables the model to consider both local and global features simultaneously, enhancing the sensitivity to information at different scales and improving the robustness and precision of key point detection. The multi-path feature parser integrates the feature outputs of different channels through an averaging operation, avoiding the over-prominence of certain features and ensuring the balance and comprehensiveness of feature representation. This balance helps the model to evenly consider all key points when processing dynamic human actions, reducing misjudgments and missed detections. Through multi-level feature extraction and fusion, the model can learn and understand more complex and abstract feature representations, which helps to improve the generalization ability of the model on unseen data. Even for pull-up actions captured from different environments and angles, the model can accurately identify and count them. The designs of the convolutional structure and the multi-path feature parser are usually highly computationally efficient and suitable for real-time or near-real-time key point detection tasks, such as sports exercise monitoring and motion capture. This means that even when processing multiple frames of images in a video stream, a relatively fast processing speed can be maintained, providing immediate feedback.

[0039] Specifically, perform channel max pooling on the above-mentioned sixth convolutional result to obtain the max pooling result, and perform channel average pooling on the above-mentioned sixth convolutional result to obtain the average pooling result; perform convolutional processing on the above-mentioned sixth convolutional result, the above-mentioned max pooling result, and the above-mentioned average pooling result respectively to obtain the seventh convolutional result, the eighth convolutional result, and the ninth convolutional result; use the Sigmoid function to process the above-mentioned seventh convolutional result, the eighth convolutional result, and the ninth convolutional result respectively to obtain the first function result, the second function result, and the third function result; perform element-wise multiplication on the above-mentioned first function result and the above-mentioned sixth convolutional result to obtain the first element-wise multiplication result, perform element-wise multiplication on the above-mentioned second function result and the above-mentioned sixth convolutional result to obtain the second element-wise multiplication result, and perform element-wise multiplication on the above-mentioned third function result and the above-mentioned sixth convolutional result to obtain the third element-wise multiplication result; determine the above-mentioned first processing result as the average of the above-mentioned first element-wise multiplication result, the above-mentioned second element-wise multiplication result, and the above-mentioned third element-wise multiplication result.

[0040] Through channel maximum pooling and channel average pooling, the network can enhance certain features and retain key information. Maximum pooling helps to highlight prominent features, while average pooling can smooth the feature map and reduce the impact of noise. The combination of these two pooling methods can ensure that the network extracts representative and relatively smooth features from the input.

[0041] Performing convolution operations on the sixth convolution result, the maximum pooling result, and the average pooling result respectively can learn image features at different scales. Features at different scales are crucial for accurately identifying human key points, especially in action recognition such as pull-ups that involve multiple body parts, as they can capture subtle changes in key points. The application of the Sigmoid function can convert the convolved feature map into probability values to evaluate the importance of each feature. This evaluation is crucial for the subsequent dot product operation, ensuring that the model pays more attention to important features and ignores irrelevant or low-quality feature information in the final feature fusion. By performing a dot product operation to fuse the sixth convolution result with the results processed by the Sigmoid function (the first function result, the second function result, and the third function result), the weights of the features can be dynamically adjusted to ensure that the final feature map contains both the original information and the features processed by pooling and convolution. This fusion method helps to improve the comprehensive representation ability of the features. Determining the average of the first dot product result, the second dot product result, and the third dot product result as the first processing result, the average operation avoids the excessive influence of a certain feature and ensures the balanced fusion of features from different sources, contributing to improving the robustness and detection accuracy of the model. This mechanism of multi-channel feature processing and dynamic fusion can make the model more robust to small changes in the input. Especially when processing dynamic images and videos, it can effectively handle lighting changes, occlusions, and perspective transformations, improving the stability and accuracy of key point detection.

[0042] The present application also provides a specific usage scenario of using a multi-channel feature parser to process the above-mentioned sixth convolution result to obtain a first processing result: In a physical education class, a physical education teacher needs to supervise and count the chin-up movements of students to evaluate their strength and endurance. The traditional manual counting method is time-consuming and laborious, and it is difficult to be accurate in a class with a large number of students. Introducing the chin-up counting technology based on human key point detection can completely change this situation. And the present application includes the following processes: The physical education teacher uses his own mobile phone or any device with a camera function to record the front or side video of the students doing chin-ups. The recorded video is segmented into multiple frames of images, and each frame of image is processed through a network model. After the sixth convolution result of each frame of image is processed by channel maximum pooling and channel average pooling, a maximum pooling result and an average pooling result are obtained. Subsequently, these three results are respectively subjected to convolution processing to obtain seventh, eighth, and ninth convolution results. Through the Sigmoid function conversion, these convolution results are converted into first, second, and third function results. The sixth convolution result is fused with the three function results after Sigmoid conversion using a dot product operation to obtain first, second, and third dot product results. Finally, by calculating the average value of these three dot product results, the first processing result, that is, a more refined human key point feature representation, is determined. Based on the first processing result, the system can accurately identify each key point (shoulder point, elbow point, wrist point, etc.) of the students during the chin-up process, and judge whether the action is standard and count the number of completions through the angle change.

[0043] Benefits of a specific usage scenario of using a multi-channel feature parser to process the above-mentioned sixth convolution result to obtain a first processing result: The system can process the video stream in real time and immediately give the count and quality feedback of the chin-up action, which is faster and more accurate than manual counting, helping the physical education teacher immediately understand the training effect of the students and provide immediate guidance and motivation. By accurately identifying action details, such as the bending and straightening states of the arms, the system can help students understand the correct posture of the chin-up action, promote the standardization of motor skills, and avoid injuries caused by improper techniques. The automated detection and counting reduce the workload of the physical education teacher, allowing them to focus more on teaching quality and individual student guidance rather than the cumbersome counting work, improving teaching efficiency. The system can automatically record and store the training data of students, including the number of completions, action quality, etc., facilitating subsequent data analysis and long-term tracking of students' physical fitness, providing a basis for customized training plans. It can be achieved with only an ordinary mobile phone, without additional hardware devices, reducing the technical threshold and cost of physical education teaching and making it possible to popularize the technology. For remote teaching or home exercise scenarios, students can record the video of their chin-ups at home through their mobile phones and upload it to the system for counting and quality assessment, making physical education teaching no longer limited by geographical location.

[0044] In step S203, different convolution processes are performed on the second feature extraction result to obtain corresponding numbers of intermediate convolution results, and a second multi-channel feature extraction block is used to process the first feature extraction result and all the intermediate convolution results to obtain corresponding numbers of second-layer feature extraction results. The number of the intermediate convolution results is determined according to the number of channels of the second multi-channel feature extraction block, and the number of channels of the second multi-channel feature extraction block is a preset multiple of the number of channels of the first multi-channel feature extraction block;

[0045] Specifically, different convolution processes are performed on the second feature extraction result to obtain a third convolution result, a fourth convolution result, and a fifth convolution result. The third convolution result is the convolution result obtained by performing one convolution process on the second feature extraction result. The fourth convolution result is the convolution result obtained by performing two convolution processes on the second feature extraction result. The fifth convolution result is the convolution result obtained by performing three convolution processes on the second feature extraction result. The second multi-channel feature extraction block is used to process the first feature extraction result, the third convolution result, the fourth convolution result, and the fifth convolution result respectively to obtain four second-layer feature extraction results, namely a third feature extraction result, a fourth feature extraction result, a fifth feature extraction result, and a sixth feature extraction result. The number of channels of the second multi-channel feature extraction block is twice the number of channels of the first multi-channel feature extraction block.

[0046] The principle of the second multi-channel feature extraction block is the same as that of the first multi-channel feature extraction block, only the number of channels is different. As Figure 2 shown, the downward-slanting arrow represents downsampling processing, and the upward-slanting arrow represents upsampling processing.

[0047] In step S204, a feature extraction module is used to process the second-layer feature extraction results respectively to obtain corresponding numbers of extraction results, and an upsampling layer is used to perform upsampling processing on at least part of the extraction results to obtain corresponding numbers of upsampling results, and the extraction results that have not undergone upsampling processing are summed with the upsampling results to obtain a final prediction result, and a heat map of the human body points is generated based on the final prediction results of the human body points.

[0048] Specifically, the feature extraction module processes the above third feature extraction result, the above fourth feature extraction result, the above fifth feature extraction result, and the above sixth feature extraction result respectively to obtain four extraction results. Then, the upsampling layer performs upsampling processing on three of the four extraction results to obtain three upsampling results. Next, the one extraction result that has not undergone upsampling processing among the four extraction results is summed with the three upsampling results to obtain the final prediction result. Finally, the heat map of the human body points is generated based on the final prediction results at each point.

[0049] By performing two types of convolution processing on the image respectively (obtaining the first convolution result and the second convolution result), where the target parameters (including the stride and quantity) of the convolution structure of the second convolution result are twice those of the first convolution result. This design can ensure that the network can capture features from different-sized perspectives, thus adapting to key points and local details of different sizes, and enhancing the sensitivity and recognition ability of the model to human body points of different scales. The first multi-path feature extraction block processes the two convolution results to obtain two first-layer feature extraction results. This not only increases the diversity of features but also further processes the first convolution result and all intermediate convolution results through the second multi-path feature extraction block to obtain the second-layer feature extraction results, which can integrate and optimize feature information from different layers and channels, improving the robustness and prediction accuracy of the model. The quantity of the intermediate convolution results is determined according to the number of channels of the second multi-path feature extraction block, and this number of channels is a preset multiple (such as 2 times) of the number of channels of the first multi-path feature extraction block. This dynamic adjustment mechanism can adaptively adjust the depth and breadth of feature extraction according to the complexity of the input image and the distribution of key points, enabling the model to better adapt to the changing detection environment and diverse pull-up actions. After the feature extraction module, the upsampling layer performs upsampling processing on some of the extraction results to obtain upsampling results. Upsampling can restore or enhance the resolution of detailed features, which is crucial for capturing subtle changes in human body points and helps improve the accuracy of key point localization. The extraction result that has not undergone upsampling processing is summed with the upsampling results to obtain the final prediction result. Generating the heat map of human body points through the final prediction result not only intuitively shows the detection positions of each key point but also facilitates further analysis and evaluation. The generation of the heat map helps the model output more explicit key point localization information and also enables users or systems to understand the confidence level and position of each point. This architecture of multi-level processing and feature fusion can achieve real-time or near-real-time key point detection and pull-up counting, improving the system response speed and processing efficiency, and is especially suitable for dynamic sports analysis and evaluation scenarios.

[0050] Step S103: Based on the correlation information between the target points of the subject in the above-mentioned multiple frames of images, identify whether the subject has completed a chin-up motion, and calculate and display the number of chin-up motions completed by the subject. The correlation information includes the angular changes between the target points.

[0051] In the above device, through the multi-channel feature parser and the multi-channel feature extraction block, the network can parse images from different angles and scales, capture more details and global information, thereby improving the detection accuracy of key points. The multi-channel feature parser enhances the extraction of local features through strategies such as parallel max pooling and average pooling; the convolutional structure in the multi-channel feature extraction block helps further fusion and extraction of features, ensuring the accuracy of key point detection. By calculating the angular changes between target points, the network can more accurately judge the completion criteria of the chin-up motion. Image acquisition and motion recognition can be completed using a mobile phone or a common camera device, without the need for expensive dedicated equipment, reducing the application threshold, thus solving the problem that the existing chin-up counting method uses a sensor for counting, but the sensor is expensive and the installation is cumbersome, and it cannot be applied to ordinary teaching places.

[0052] In addition, in different environments and shooting angles, multi-channel feature processing can make the network more robust to light, occlusion, and perspective changes. Even in complex backgrounds or low-light conditions, it can stably identify and locate human body points, ensuring the reliability of counting chin-up motions.

[0053] Among them, identifying whether the subject has completed a chin-up motion based on the correlation information between the target points of the subject in the above-mentioned multiple frames of images includes: connecting the shoulder point, elbow point, and wrist point in the target points into a target broken line; when the included angle of the target broken line is less than the first preset included angle, determining that the subject's arm is in a bent state; when the included angle of the target broken line is greater than the second preset included angle, determining that the subject's arm is in a straight state, and the second preset included angle is less than the first preset included angle; when it is determined that the state of the subject's arm is converted between the bent state and the straight state, and there is no point in the target points that coincides with the ground, determining that the subject has completed the chin-up motion.

[0054] Specifically, by calculating the included angles of the broken lines formed by the shoulder point, elbow point, and wrist point, it is possible to accurately determine whether the arm is in a bent or straight state. This method is direct and effective, capable of avoiding misjudgments caused by the positioning error of a single key point and improving the accuracy of action recognition. By using preset included angle thresholds (the first preset included angle and the second preset included angle), the completion criteria for the chin-up action can be quantitatively defined, making the evaluation more objective and standardized. The first preset included angle (e.g., less than 90 degrees) is used to judge the bent state of the arm, while the second preset included angle (e.g., greater than 170 degrees) is used to judge the straight state of the arm, ensuring the consistency and fairness of action recognition. By monitoring the conversion of the arm state, that is, the cycle from bent to straight and then from straight to bent, the interference of non-chin-up actions can be effectively excluded, enhancing the robustness of the recognition process. At the same time, by checking whether there are no points coinciding with the ground in the target points, misidentifications caused by non-standard postures or occlusions can be excluded, further improving the accuracy of counting. The change in the included angle reflects the dynamic change of the arm, and the capture of this dynamic change is crucial for recognizing continuous actions such as chin-ups. By monitoring the change in the included angle in consecutive frames, the system can track the progress of the action in real time, ensuring the real-time and coherence of counting. The criteria for judging the bent and straight states of the arm are based on angle changes and are less affected by environmental factors such as external light and background complexity, enabling the system to work stably in various environments and improving its practicality.

[0055] Multi-channel feature parser: Processes the input according to different pooling strategies, and uses channel maximum pooling, average pooling, and the original information to construct a more comprehensive feature representation. Finally, averaging the outputs of the three branches can avoid a certain feature being too prominent and ensure a more balanced feature fusion.

[0056] Multi-channel feature extraction block: Each multi-channel feature extraction block includes a convolutional structure and a multi-channel feature parser, aiming to perform more detailed parsing and processing on the input features. The input of each branch will pass through different pooling layers (maximum pooling and average pooling) to extract different information.

[0057] First feature extraction layer: Consists of multiple multi-channel feature extraction blocks, used to further extract and fuse the input features, strengthen the relationship between different features, and extract multi-scale features.

[0058] Second feature extraction layer: The second feature extraction layer performs feature extraction on the four-branch outputs obtained from the first feature extraction layer again, and further fuses them through multi-channel feature extraction blocks. By increasing the number of multi-channel feature extraction blocks and more convolutional operations, the network's processing ability for different scales and different features is further enhanced, helping the network focus on the details and complexity of the chin-up action.

[0059] Obtain pictures of professional athletes on the horizontal bar and students' daily pull-up training pictures, mark the skeletal point information of athletes and students during exercise in the two types of pictures, and obtain a pull-up point data set; input the pull-up point data set into the pull-up human body detection network for training to obtain a pull-up human body detection network; place a mobile phone or other camera device in front of the student to shoot a video, and capture the video and send it to the pull-up human body detection network at X frames per second; obtain the predicted human skeletal point information of the student for all images through the model; judge whether the conditions are met according to the preset human skeletal point information of the student; such as Figure 3 As shown in FIG. 1 , the human skeleton point information obtained includes 17 points in total, including left and right eye points, left and right ear points, nose, left and right shoulder points, left and right elbow points, left and right wrist points, left and right hip points, left and right knee points, and left and right ankle points. Taking the right arm as an example, the right shoulder point, right elbow point, and right wrist point are trained into a broken line. When the included angle of the preset broken line is less than 90 degrees, the subject's right arm is in a bent state. When the included angle is greater than 170 degrees, the subject's right arm is in a straight state. The same is true for the left arm. Figure 4 As shown, the arms of the student being tested are in a straight state. If the left and right wrist points are at the highest position among all points and the left and right elbow points are higher than the nose point of the student being tested, this state is marked as state one; the arms of the student being tested are in a bent state. When the left and right eye points, left and right ear points, and nose positions are all higher than the left and right wrist points, this is recorded as state two; a ground position is set up, and when no point of the subject overlaps with the ground and the subject completes a transition from state one to state two, it is recorded as one pull-up.

[0060] According to the embodiments of the present disclosure, first, by acquiring images of professional athletes and students during daily pull-up training, and marking the skeletal point information of athletes and students in these images, a pull-up point data set is constructed. These data sets will be input into the pull-up human body detection network for training, so as to obtain an accurate pull-up key point detection model. In actual use, a camera device (such as a mobile phone) is placed in front of the student to shoot a video of the student performing pull-up training. The video is intercepted as X frames of images per second and is sequentially input into the trained pull-up key point detection model. The model processes each frame of the image and predicts the human skeletal point information of the student. By analyzing the predicted skeletal point information, the system can determine whether the student's movements meet the preset standards.

[0061] The training method of the pull-up human body detection network includes: selecting multiple pull-up competition and teaching professional videos, shooting professional and student pull-up videos on campus, intercepting the videos into multiple frames of pictures, marking the human body bone point information in the pictures, and combining professional actions with the actual campus physical fitness test pull-up scenarios to form a special dataset for pull-up in multiple scenarios, so as to enhance the robustness of the model to recognize each key point in pull-ups and adapt to the actual pull-up detection scenarios.

[0062] Among them, the pull-up human body detection network is as Figure 2 and Figure 5 shown. The steps implemented by this network include:

[0063] Step 1, the input of the above pull-up human body detection network first passes through a convolutional structure to increase the dimension. The output of this convolutional structure passes through two convolutional structures respectively to obtain two branch outputs. The stride and number of convolutional kernels in the convolutional layer of one convolutional structure are twice that of the other. The convolutional layer structure consists of a convolutional layer, a batch normalization layer, and a ReLU activation function. The formula is as follows:

[0064] F o = ReLU(BatchNorm(Conv(F i )));

[0065] Conv(F i ) represents the convolutional result of the input value (F i ), and BatchNorm(Conv(F i ) represents the batch normalization processing result of Conv(F i ).

[0066] Step 2: Pass the outputs of the two branches in Step 1 through Feature Extraction Layer 1, which is composed of two cascaded multi-path feature extraction blocks. In the multi-path feature extraction block, the output of each branch in the first step will pass through two cascaded feature extraction blocks respectively to obtain new output branches 1 and 2. Branch 2 passes through an upsampling layer with a two-fold upsampling, and then calculates the average value with the output of Branch 1 and activates it through the ReLU function to obtain the final output of Branch 1 in Feature Extraction Layer 1. Branch 1 passes through a two-fold downsampling layer and then calculates the average value with Branch 2 to obtain the final output of Branch 2. Among them, the upsampling layer consists of a convolutional layer, a batch normalization layer, and an upsampling structure. The two-fold downsampling layer consists of a convolutional layer with a kernel size of 3, a stride of 2, and a padding of 0 and a batch normalization layer. If it is four-fold downsampling, there is one more convolutional layer with the same configuration. Among them, the feature extraction block is composed of a convolutional structure and a multi-path feature parser in series. Among them, the multi-path feature parser contains three branches. The input of the multi-path feature parser first passes through a channel max pooling layer, and then passes through a convolutional layer and a Sigmoid activation function to obtain the output. This output is multiplied by the input of this branch to obtain the output of Branch 1 of the multi-path feature parser. The configuration of Branch 2 changes the channel max pooling layer to a channel average pooling layer, and the remaining configuration and process are the same as those of Branch 1. Branch 3 has no channel max pooling layer and channel average pooling layer, and the average value of the outputs of the three branches is calculated to obtain the final output of the multi-path feature parser.

[0067] Step 3: Pass the output Branch 2 in Step 2 through the branches composed of Convolutional Structure 1, Convolutional Structure 2, and Convolutional Structure 3, and the branches composed of Convolutional Structure 4, 5, and 6 respectively to obtain three new branches, which together with the output Branch 1 in Step 2 form four output branches. Among them, the number of convolutional kernels and the stride of Convolutional Structure 1, 2, and 4 are the same. The number of convolutional kernels and the stride of Convolutional Structure 3 and 5 are twice that of Convolutional Structure 2. The number of convolutional kernels and the stride of Convolutional Structure 6 are 4 times that of Convolutional Structure 2.

[0068] Step 4: Input the four branches in Step 3 into Feature Extraction Layer 2, which contains 4 multi-path feature extraction blocks described in Step 2, and then obtain the four output branches of Feature Extraction Layer 2.

[0069] Step 5: Pass the outputs of the four branches in Step 4 through eight feature extraction modules respectively to obtain four new output branches. Output Branches 2, 3, and 4 pass through the upsampling structure mentioned in Step 2 respectively, and finally the four branches are added to obtain the output.

[0070] Step 6: Pass the output in Step 5 through a convolutional layer with 17 convolutional kernels to obtain the final output result of the pull-up human detection network. Calculate the heatmaps of the 17 output channels of the pull-up human detection network, and the predicted detection results of each key point of the student can be obtained.

[0071] Through the embodiments of the present application, in order to enhance the robustness of the model in different scenarios, the system will select multiple chin-up competition and teaching videos, and combine them with the actual chin-up test scenario on campus to construct a multi-scenario dataset. This dataset covers professional training videos and actual test videos, thereby improving the model's ability to recognize human body bone points in different environments. The input of the network first undergoes dimensionality improvement through a convolutional structure and then outputs two branches. The convolution kernel stride and the number of one branch are twice those of the other branch. Each convolutional structure includes a convolutional layer, a batch normalization layer, and a ReLU activation function. The two output branches will be further processed through a feature extraction layer. The feature extraction layer consists of multiple feature extraction blocks, and each block contains multiple convolutional layers and feature parsers. The feature parser uses operations such as max pooling and average pooling to parse human feature information from different dimensions. After feature extraction, the output branches will be fused through multiple convolutional structures. In this way, the model can more effectively integrate features at different levels and extract more accurate human feature data. Finally, the model generates a heat map for each key point through a convolutional layer with 17 output channels. These heat maps can intuitively display the positions of each key point of the student.

[0072] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0073] The embodiments of the present application also provide a chin-up counting device. It should be noted that the chin-up counting device of the embodiments of the present application can be used to execute the chin-up counting method provided by the embodiments of the present application. This device is used to implement the above embodiments and preferred implementation manners, and those that have been described will not be repeated. As used hereinafter, the term "module" can be a combination of software and / or hardware that can achieve a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.

[0074] The following introduces the chin-up counting device provided by the embodiments of the present application.

[0075] Figure 6 is a structural block diagram of a chin-up counting device provided according to an embodiment of the present application. As Figure 6 shown, the device includes:

[0076] An acquisition unit 61, configured to acquire multiple frames of images including chin-up actions;

[0077] The first processing unit 62 is configured to perform preset processing on the above-mentioned multi-frame images to obtain a heat map for predicting each human body point of the subject, and extract the target points of the subject from the above-mentioned heat map. The pull-up human body detection network is constructed based on a multi-channel feature resolver and a multi-channel feature extraction block. The preset processing includes convolution processing, feature extraction processing, and upsampling processing;

[0078] The second processing unit 63 is configured to identify whether the subject has completed a pull-up action according to the association information between the above-mentioned target points of the subject in the above-mentioned multi-frame images, and calculate and display the number of pull-up actions completed by the subject. The association information includes the angular change between the above-mentioned target points.

[0079] In an embodiment of the present application, the first processing unit includes a first processing module, a second processing module, a third processing module, and a fourth processing module. The first processing module is configured to perform two types of convolution processing on the above-mentioned multi-frame images respectively to obtain a first convolution result and a second convolution result. Among them, the target parameters of the convolution structure for outputting the second convolution result are twice the target parameters of the convolution structure for outputting the first convolution result. The above-mentioned target parameters of the convolution structure include the stride and the number; the second processing module is configured to use a first multi-channel feature extraction block to process the first convolution result and the second convolution result to obtain two first-layer feature extraction results, namely a first feature extraction result and a second feature extraction result; the third processing module is configured to perform different convolution processing on the second feature extraction result to obtain a corresponding number of intermediate convolution results, and use a second multi-channel feature extraction block to process the first feature extraction result and all the above-mentioned intermediate convolution results to obtain a corresponding number of second-layer feature extraction results. The number of the above-mentioned intermediate convolution results is determined according to the channel number of the second multi-channel feature extraction block. The above-mentioned channel number of the second multi-channel feature extraction block is a preset multiple of the channel number of the first multi-channel feature extraction block; the fourth processing module is configured to use a feature extraction module to process the above-mentioned second-layer feature extraction results respectively to obtain a corresponding number of extraction results, and use an upsampling layer to perform upsampling processing on at least part of the above-mentioned extraction results to obtain a corresponding number of upsampling results, and perform a summation process on the above-mentioned extraction results that have not undergone upsampling processing and the above-mentioned upsampling results to obtain a final prediction result, and generate a heat map of the above-mentioned human body points based on the above-mentioned final prediction results of each of the above-mentioned human body points.

[0080] In an embodiment of the present application, the second processing module includes a first processing sub-module, a second processing sub-module, a third processing sub-module, and a fourth processing sub-module. The first processing sub-module is configured to process the first convolution result and the second convolution result by using two feature extraction modules to obtain a first processing result and a second processing result; the second processing sub-module is configured to perform downsampling processing on the first processing result by using two downsampling layers to obtain a first downsampling processing result, and perform upsampling processing on the second processing result by using the upsampling layer to obtain a first upsampling processing result; the third processing sub-module is configured to perform averaging processing on the first upsampling processing result and the first processing result to obtain a first average result, and process the first average result by using the ReLU function to obtain the first feature extraction result; the fourth processing sub-module is configured to perform averaging processing on the first downsampling processing result and the second processing result to obtain a second average result, and determine the second feature extraction result as the second average result.

[0081] In an embodiment of the present application, the second processing module includes a fifth processing sub-module and a sixth processing sub-module. The fifth processing sub-module is configured to process the first convolution result by using a convolutional structure to obtain a sixth convolution result during the process of processing the first convolution result and the second convolution result by using the first multi-channel feature extraction block to obtain two first-layer feature extraction results; the sixth processing sub-module is configured to process the sixth convolution result by using a multi-channel feature parser to obtain a first processing result.

[0082] In an embodiment of the present application, the sixth processing sub-module includes a seventh processing sub-module, an eighth processing sub-module, a ninth processing sub-module, a tenth processing sub-module, and a determination sub-module. The seventh processing sub-module is configured to perform channel maximum pooling processing on the sixth convolution result to obtain a maximum pooling result, and perform channel average pooling processing on the sixth convolution result to obtain an average pooling result; the eighth processing sub-module is configured to perform convolution processing on the sixth convolution result, the maximum pooling result, and the average pooling result respectively to obtain a seventh convolution result, an eighth convolution result, and a ninth convolution result; the ninth processing sub-module is configured to process the seventh convolution result, the eighth convolution result, and the ninth convolution result respectively by using the Sigmoid function to obtain a first function result, a second function result, and a third function result; the tenth processing sub-module is configured to perform dot multiplication on the first function result and the sixth convolution result to obtain a first dot multiplication result, perform dot multiplication on the second function result and the sixth convolution result to obtain a second dot multiplication result, and perform dot multiplication on the third function result and the sixth convolution result to obtain a third dot multiplication result; the determination sub-module is configured to determine the first processing result as the average value of the first dot multiplication result, the second dot multiplication result, and the third dot multiplication result.

[0083] In an embodiment of the present application, the second processing unit includes a fifth processing module, a first determination module, a second determination module, and a third determination module. The fifth processing module is configured to connect the shoulder point, elbow point, and wrist point in the above-mentioned target points into a target broken line; the first determination module is configured to determine that the arm of the above-mentioned subject is in a bent state when the included angle of the above-mentioned target broken line is less than a first preset included angle; the second determination module is configured to determine that the arm of the above-mentioned subject is in a straight state when the included angle of the above-mentioned target broken line is greater than a second preset included angle, and the second preset included angle is less than the first preset included angle; the third determination module is configured to determine that the above-mentioned subject has completed the above-mentioned chin-up movement when it is determined that the state of the arm of the above-mentioned subject is switched between the bent state and the straight state, and there is no point in the above-mentioned target points that coincides with the ground.

[0084] The above-mentioned chin-up counting device includes a processor and a memory. The above-mentioned acquisition unit, the first processing unit, the second processing unit, etc. are all stored in the memory as program units, and the processor executes the above-mentioned program units stored in the memory to implement corresponding functions. The above-mentioned modules are all located in the same processor; or, the above-mentioned modules are respectively located in different processors in any combination form.

[0085] The processor contains a kernel, and the kernel retrieves the corresponding program unit from the memory. One or more kernels can be set, and by adjusting the kernel parameters, the problem that the existing chin-up counting method uses a sensor for counting, however, the sensor is expensive and the installation of the sensor is cumbersome and cannot be applied to ordinary teaching places can be solved.

[0086] The memory may include non-permanent memory in a computer-readable medium, forms such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM), and the memory includes at least one storage chip.

[0087] An embodiment of the present invention provides a computer-readable storage medium. The above-mentioned computer-readable storage medium includes a stored program, wherein when the above-mentioned program runs, it controls the device where the above-mentioned computer-readable storage medium is located to execute the above-mentioned chin-up counting method.

[0088] An embodiment of the present invention provides a processor. The above-mentioned processor is used to run a program, wherein when the above-mentioned program runs, it executes the above-mentioned chin-up counting method.

[0089] An embodiment of the present invention provides a device, which includes a processor, a memory, and a program stored on the memory and executable on the processor. When the processor executes the program, it implements at least the following steps: obtaining multiple frames of images containing chin-up actions; performing preset processing on the multiple frames of images to obtain a heat map for predicting each human body point of the person to be measured, and extracting the target points of the person to be measured from the heat map. Among them, the chin-up human detection network is constructed based on a multi-channel feature parser and a multi-channel feature extraction block, and the preset processing includes convolution processing, feature extraction processing, and upsampling processing; according to the correlation information between the target points of the person to be measured in the multiple frames of images, identifying whether the person to be measured has completed a chin-up action, and calculating and displaying the number of chin-up actions completed by the person to be measured. The correlation information includes the angle change between the target points. The device in this article can be a server, a PC, a PAD, a mobile phone, etc.

[0090] The present application also provides a computer program product, which, when executed on a data processing device, is adapted to execute a program initialized with at least the following method steps: obtaining multiple frames of images containing chin-up actions; performing preset processing on the multiple frames of images to obtain a heat map for predicting each human body point of the person to be measured, and extracting the target points of the person to be measured from the heat map. Among them, the chin-up human detection network is constructed based on a multi-channel feature parser and a multi-channel feature extraction block, and the preset processing includes convolution processing, feature extraction processing, and upsampling processing; according to the correlation information between the target points of the person to be measured in the multiple frames of images, identifying whether the person to be measured has completed a chin-up action, and calculating and displaying the number of chin-up actions completed by the person to be measured. The correlation information includes the angle change between the target points.

[0091] The present application also provides a chin-up counting system, including: one or more processors, a memory, and one or more programs. Among them, the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The one or more programs include those for executing any one of the above methods.

[0092] Obviously, those skilled in the art should understand that the various modules or steps of the present invention described above can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed over a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately fabricated into individual integrated circuit modules, or multiple modules or steps among them can be fabricated into a single integrated circuit module for implementation. In this way, the present invention is not limited to any specific combination of hardware and software.

[0093] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) containing computer-usable program codes.

[0094] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0095] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the flows Figure 1 or multiple flows and / or blocks

[0096] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process. Thus, the instructions executed on the computer or other programmable device provide for implementing the functions in the flowFigure 1 one or more processes and / or blocks Figure 1 steps of the functions specified in one or more blocks

[0097] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.

[0098] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory such as read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0099] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology for information storage. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory media such as modulated data signals and carrier waves.

[0100] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover a non-exclusive inclusion, such that a process, method, commodity or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, commodity or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, commodity or device comprising the element.

[0101] The above description is only the preferred embodiment of the present application and is not intended to limit the present application. For those skilled in the art, the present application may have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A method for counting pull-ups, characterized in that, Including: Obtaining multiple frames of images including pull-up actions; Performing preset processing on the multiple frames of images to obtain a heat map for predicting each human body point of the person to be measured, and extracting the target points of the person to be measured from the heat map. Among them, the pull-up human detection network is constructed based on a multi-path feature resolver and a multi-path feature extraction block, and the preset processing includes convolution processing, feature extraction processing, and upsampling processing; According to the association information between the target points of the person to be measured in the multiple frames of images, identifying whether the person to be measured has completed a pull-up action, and calculating and displaying the number of pull-up actions completed by the person to be measured. The association information includes the angle change between the target points.

2. The method according to claim 1, characterized in that, Performing preset processing on the multiple frames of images to obtain a heat map for predicting each human body point of the person to be measured, including: Performing two types of convolution processing on the multiple frames of images respectively to obtain a first convolution result and a second convolution result. Among them, the target parameters of the convolution structure outputting the second convolution result are twice the target parameters of the convolution structure outputting the first convolution result. The target parameters of the convolution structure include the stride and the number; Using a first multi-path feature extraction block to process the first convolution result and the second convolution result to obtain two first-layer feature extraction results, namely a first feature extraction result and a second feature extraction result; Performing different convolution processing on the second feature extraction result to obtain a corresponding number of intermediate convolution results, and using a second multi-path feature extraction block to process the first feature extraction result and all the intermediate convolution results to obtain a corresponding number of second-layer feature extraction results. The number of intermediate convolution results is determined according to the channel number of the second multi-path feature extraction block, and the channel number of the second multi-path feature extraction block is a preset multiple of the channel number of the first multi-path feature extraction block; Using a feature extraction module to process the second-layer feature extraction results respectively to obtain a corresponding number of extraction results, and using an upsampling layer to perform upsampling processing on at least part of the extraction results to obtain a corresponding number of upsampling results, and performing a summation process on the extraction results that have not undergone upsampling processing and the upsampling results to obtain a final prediction result, and generating a heat map of the human body points based on the final prediction results of each human body point.

3. The method according to claim 2, wherein Using a first multi-path feature extraction block to process the first convolution result and the second convolution result to obtain two first-layer feature extraction results, including: Using two feature extraction modules to process the first convolution result and the second convolution result to obtain a first processing result and a second processing result; Performing downsampling processing on the first processing result using two downsampling layers to obtain a first downsampling processing result, and performing upsampling processing on the second processing result using the upsampling layer to obtain a first upsampling processing result; Performing an averaging process on the first upsampling processing result and the first processing result to obtain a first average result, and using the ReLU function to process the first average result to obtain the first feature extraction result; Perform an averaging process on the first downsampling processing result and the second processing result to obtain a second average result, and determine the second feature extraction result as the second average result.

4. The method according to claim 2, wherein In the process of using the first multi-channel feature extraction block to process the first convolution result and the second convolution result to obtain two first-layer feature extraction results, the method includes: Process the first convolution result using a convolution structure to obtain a sixth convolution result; Process the sixth convolution result using a multi-channel feature parser to obtain a first processing result.

5. The method according to claim 4, wherein Processing the sixth convolution result using a multi-channel feature parser to obtain the first processing result includes: Perform channel maximum pooling on the sixth convolution result to obtain a maximum pooling result, and perform channel average pooling on the sixth convolution result to obtain an average pooling result; Perform convolution processing on the sixth convolution result, the maximum pooling result, and the average pooling result respectively to obtain a seventh convolution result, an eighth convolution result, and a ninth convolution result; Use the Sigmoid function to process the seventh convolution result, the eighth convolution result, and the ninth convolution result respectively to obtain a first function result, a second function result, and a third function result; Perform dot multiplication on the first function result and the sixth convolution result to obtain a first dot multiplication result, perform dot multiplication on the second function result and the sixth convolution result to obtain a second dot multiplication result, and perform dot multiplication on the third function result and the sixth convolution result to obtain a third dot multiplication result; Determine the first processing result as the average value of the first dot multiplication result, the second dot multiplication result, and the third dot multiplication result.

6. The method according to claim 1, wherein According to the association information between the target points of the subject in the multi-frame images, identifying whether the subject has completed a chin-up action includes: Connect the shoulder point, elbow point, and wrist point in the target points to form a target broken line; When the included angle of the target broken line is less than a first preset included angle, determine that the subject's arm is in a bent state; When the included angle of the target broken line is greater than a second preset included angle, determine that the subject's arm is in a straight state, and the second preset included angle is less than the first preset included angle; When it is determined that the state of the subject's arm is converted between the bent state and the straight state, and there is no point in the target points that coincides with the ground, determine that the subject has completed the chin-up action.

7. A pull-up counting device, characterized in that, Includes: An acquisition unit for acquiring multi-frame images including a chin-up action; A first processing unit for performing a preset process on the multi-frame images to obtain a heat map for predicting each human body point of the subject, and extracting the target points of the subject from the heat map, wherein the chin-up human detection network is constructed based on a multi-channel feature parser and a multi-channel feature extraction block, and the preset process includes convolution processing, feature extraction processing, and upsampling processing; A second processing unit, configured to identify whether the subject has completed a chin-up action according to the association information between the target points of the subject in the multi-frame images, and calculate and display the number of chin-up actions completed by the subject, where the association information includes the angular change between the target points.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program runs, it controls the device where the computer-readable storage medium is located to execute the method according to any one of claims 1 to 6.

9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 6.

10. A pull-up counting system, characterized in that, Comprising: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing the method according to any one of claims 1 to 6.