Method and device for detecting sit-up and readable storage medium
By combining the key point recognition model and the space-time graph convolution model, the spatial and temporal characteristics of sit-up movements are identified and extracted, and the problem of inaccurate detection of sit-ups in the prior art is solved, and high accuracy detection in complex scenarios is achieved.
Patent Information
- Application Number
- CN202411963301.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-30
- Publication Date
- 2025-05-06
AI Technical Summary
When detecting sit-ups, the prior art relies on artificial experience and is susceptible to subjective factors. In complex scenes, factors such as light changes and occlusion lead to insufficient clear images, affecting detection accuracy.
The key point recognition model recognizes the human key points in each frame of the video, and preliminarily determines whether sit-ups are standard. The space-time graph convolution model is used to extract features of multi-frame action images, fuse spatial and temporal features, and then accurately detects the number of sit-ups.
Improves the accuracy of sit-up detection, reduces interference from other factors, and accurately counts sit-up movements in complex scenarios.
Smart Images

Figure CN119942639A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of image processing technology. More specifically, the present disclosure relates to a method and apparatus for detecting sit-ups and a readable storage medium. Background Art
[0002] Sit-ups are a common physical education and track and field competition. As a fitness exercise, sit-ups are important for strengthening abdominal muscles, improving posture and promoting physical health.
[0003] In the prior art, the number of sit-ups performed by a person is detected by manually counting the movements that meet the sit-up standards. This detection method relies heavily on manual experience and is easily affected by subjective factors. It may not be possible to accurately determine whether the sit-up movements meet the sit-up standards, thus resulting in inaccurate detection.
[0004] Moreover, in the prior art, the sit-up images are mainly detected by target detection and posture recognition to count or evaluate the sit-up actions. In complex scenes, changes in lighting conditions and occlusion of sports personnel may cause the images captured when taking sit-ups to be unclear or lack key information. Therefore, when the above detection method detects the image, it will be interfered by other factors, resulting in the inability to accurately detect or count sit-ups.
[0005] In view of this, there is an urgent need to provide a method for detecting sit-ups so as to improve the accuracy of detecting sit-ups. Summary of the invention
[0006] In order to at least solve one or more technical problems mentioned above, the present disclosure proposes a method and an apparatus for detecting sit-ups and a readable storage medium in multiple aspects.
[0007] In a first aspect, the present disclosure provides a method for detecting sit-ups, comprising: acquiring a video of a person performing sit-ups; identifying one or more human key points of the person performing the sit-ups in each frame of the video through a key point recognition model; preliminarily determining whether the sit-ups are standard based on the positional relationship between each of the human key points; if the sit-ups are standard, determining multiple frames of action images corresponding to the sit-ups; extracting features of multiple human key points in multiple frames of the action images and the times corresponding to the multiple frames of the action images using a spatiotemporal graph convolution model to obtain action features of the sit-ups; and determining the number of sit-ups in the video based on the action features.
[0008] In a second aspect, the present disclosure provides an apparatus for detecting sit-ups, comprising: a processor configured to execute program instructions; and a memory configured to store the program instructions, wherein when the program instructions are loaded and executed by the processor, the apparatus executes the method provided in the first aspect.
[0009] In a third aspect, the present disclosure provides a computer-readable storage medium having program instructions stored thereon, and when the program instructions are loaded and executed by a processor, the method provided in the first aspect is implemented.
[0010] Through the method for detecting sit-ups provided above, the disclosed embodiment identifies one or more human key points of a moving person in each frame of a video through a key point recognition model; preliminarily determines whether a sit-up is standard based on the positional relationship between each human key point; if the sit-up is standard, determines multiple frames of action images corresponding to the sit-up; uses a spatiotemporal graph convolution model to extract features of multiple human key points in multiple frames of the action images and the time corresponding to the multiple frames of the action images to obtain action features, so that the spatial features and time features of multiple human key points can be extracted and fused, thereby obtaining action features including information in two dimensions of space and time; based on the action features, the number of sit-ups in the video is detected, so that complex sit-ups can be detected from two dimensions of space and time, which can reduce interference from other factors, thereby accurately determining the number of sit-ups. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] By reading the detailed description below with reference to the accompanying drawings, the above and other objects, features and advantages of the exemplary embodiments of the present disclosure will become readily understood. In the accompanying drawings, several embodiments of the present disclosure are shown in an exemplary and non-limiting manner, and the same or corresponding reference numerals represent the same or corresponding parts, wherein:
[0012] Figure 1 An exemplary flow chart showing a method for detecting sit-ups according to some embodiments of the present disclosure;
[0013] Figure 2 An exemplary flow chart showing a method for identifying key points of a human body in each frame of a video according to some embodiments of the present disclosure;
[0014] Figure 3 An exemplary flow chart showing a method for obtaining human body enhancement features according to some embodiments of the present disclosure is shown;
[0015] Figure 4 An exemplary flow chart of a method for obtaining channel width human body characteristics according to some embodiments of the present disclosure is shown;
[0016] Figure 5 An exemplary flow chart showing a method for obtaining a high-passage human body feature according to some embodiments of the present disclosure is shown;
[0017] Figure 6 An exemplary flow chart showing a method for determining whether a sit-up is standard according to some embodiments of the present disclosure;
[0018] Figure 7 An exemplary flow chart showing a method for obtaining the motion characteristics of sit-ups according to some embodiments of the present disclosure;
[0019] Figure 8 An exemplary structural block diagram showing an apparatus for detecting sit-ups according to an embodiment of the present disclosure DETAILED DESCRIPTION
[0020] The following will be combined with the drawings in the embodiments of the present disclosure to clearly and completely describe the technical solutions in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments in the present disclosure, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of the present disclosure.
[0021] It should be understood that the terms "include" and "comprising" used in the specification and claims of the present disclosure indicate the presence of described features, integers, steps, operations, elements and / or components, but do not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or collections thereof.
[0022] It should also be understood that the terms used in this disclosure are only for the purpose of describing specific embodiments and are not intended to limit the disclosure. As used in this disclosure and claims, the singular forms of "a", "an", and "the" are intended to include the plural forms unless the context clearly indicates otherwise. It should also be further understood that the term "and / or" used in this disclosure and claims refers to any combination of one or more of the associated listed items and all possible combinations, including these combinations.
[0023] As used in this specification and claims, the term "if" may be interpreted as "when" or "upon" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" may be interpreted as meaning "upon determination" or "in response to determining" or "upon detection of [described condition or event]" or "in response to detecting [described condition or event]," depending on the context.
[0024] The specific implementation of the present disclosure is described in detail below with reference to the accompanying drawings.
[0025] Exemplary application scenarios.
[0026] An existing technical solution for counting sit-ups specifically includes: manually counting the actions that meet the sit-up standards to detect the number of sit-ups performed by a person. This detection method relies heavily on manual experience and is easily affected by subjective factors. It may not be able to accurately determine whether the sit-up actions meet the sit-up standards, thus resulting in inaccurate detection.
[0027] Another existing technical solution for counting sit-ups specifically includes: detecting sit-up images by means of target detection and posture recognition, so as to count or evaluate the sit-up actions. In complex scenes, changes in lighting conditions and occlusion of sports personnel may cause the image captured when taking sit-ups to be unclear or lack key information. Therefore, when the above detection method detects the image, it will be interfered by other factors, resulting in the inability to accurately detect or count sit-ups.
[0028] Exemplary application scenarios.
[0029] In view of this, the disclosed embodiment provides a method for detecting sit-ups, which identifies one or more human key points of a moving person in each frame of a video through a key point recognition model; preliminarily determines whether a sit-up is standard based on the positional relationship between each human key point; if the sit-up is standard, determines multiple frames of action images corresponding to the sit-up; uses a spatiotemporal graph convolution model to extract features of multiple human key points in multiple frames of the action images and the time corresponding to the multiple frames of the action images to obtain action features, so that the spatial features and time features of multiple human key points can be extracted and fused, thereby obtaining action features including information in two dimensions of space and time; based on the action features, the number of sit-ups in the video is detected, so that complex sit-ups can be detected from two dimensions of space and time, which can reduce interference from other factors, thereby accurately determining the number of sit-ups.
[0030] Figure 1 An exemplary flow chart of a method for detecting sit-ups according to some embodiments of the present disclosure is shown.
[0031] As shown in the figure, in step S100, a video of a sit-up of an athlete is obtained; in step S200, one or more human key points of the athlete in each frame of the video are identified by a key point recognition model; in step S300, whether the sit-up is standard is preliminarily determined based on the positional relationship between each human key point; in step S400, if the sit-up is standard, multiple frames of action images corresponding to the sit-up are determined; in step S500, a spatiotemporal graph convolution model is used to extract features of multiple human key points in the multiple frames of action images and the time corresponding to the multiple frames of action images to obtain the action features of the sit-up; and in step S600, the number of sit-ups in the video is determined based on the action features.
[0032] In the embodiment of the present disclosure, the athlete refers to a person who performs sit-ups. In step S100, the athlete who performs sit-ups is filmed by a camera or other filming device to obtain a video of the athlete performing sit-ups. The video includes the whole body of the athlete.
[0033] In some embodiments, each athlete may be assigned a user number, such as an ID (Identity document), to distinguish different athletes.
[0034] In certain embodiments, the key point recognition model can be a target detection model, specifically, the key point recognition model can be a pose estimation model (You Only Look Once v8 pose, yolov8-pose for short), the key point recognition model is used to identify the human key points of the sports personnel in the image, "human key points" here refer to the key parts or joints of the human body used to describe the long jump personnel, and the human key points include hands, head, ankles, knees and hip joints, etc. It will be appreciated by those skilled in the art that the human key points are not subject to any restrictions. In step S200, the above-mentioned video is identified by the key point recognition model to obtain one or more human key points.
[0035] In some embodiments, the positional relationship between the key points of the human body in step S300 refers to the distance between the key points of the human body, for example, the distance between the hand and the ear. The positional relationship between the key points of the human body can also include the set relationship between the key points of the human body, for example, the straight line formed by the head and the hip joint is called the first straight line, the straight line formed by the knee and the hip joint is called the second straight line, and the angle formed by the first straight line and the second straight line.
[0036] For ease of description, the coordinates of the key points of the human body in each frame of the video are referred to as "key point coordinates of the human body". In step S300, the positional relationship between the key points can be determined by the key point coordinates of the human body. Then, through the positional relationship between the key points, it can be preliminarily determined whether the sit-up is standard. For example, in each frame, if the distance between the head and the hand exceeds 10 mm, it is determined that the sit-up is not standard. In the embodiments of the present disclosure, the specific standard of sit-ups can be set by oneself without any limitation.
[0037] It should be noted that the "standard" here refers to whether the athlete is doing sit-ups, not other sports. Even if a sit-up is not a standard action and is not countable, it is still a standard sit-up. For example, if the athlete does a small amount of sit-ups, it is still a standard sit-up.
[0038] In some embodiments, if it is determined that the sit-ups are not standard, then there is no need to count the sit-ups. If the sit-ups are standard, then the sit-ups are further counted, i.e., the number of sit-ups performed by the athlete is detected. In step S400, when the sit-ups are standard, multiple frames of images corresponding to the standard sit-ups in the above video are determined, where "images corresponding to qualified sit-ups in the above video" are referred to as "action images".
[0039] Then, in step S500, multiple frames of action images and the time corresponding to the multiple frames of action images are used as a spatiotemporal graph convolution model, so as to extract spatial features of key points of the human body in the multiple frames of action images through the spatiotemporal graph convolution model, and extract temporal features of the time corresponding to the multiple frames of action images, and fuse the spatial features and the temporal features to obtain the action features corresponding to the sit-ups. The spatiotemporal graph convolution model is SpatialTempora lGraph ConvolutionalNetworks, referred to as stgcn. Those skilled in the art can understand that the spatiotemporal graph convolution model stgcn can extract temporal features and spatial features, and effectively fuse the temporal features and spatial features.
[0040] The above-mentioned time feature represents the time information of each human body key point, and the above-mentioned space feature represents the spatial relationship between each human body key point. Here, the feature after the fusion of time feature and space feature is called "action feature". The action feature can represent the dynamic change of each human body key point within a certain period of time, that is, the action feature can represent the position change, change speed and movement trend of each human body key point.
[0041] In step S500, the spatiotemporal graph convolution model can be used to extract features from multiple action images in two dimensions, time and space, and obtain the spatial and temporal relationships of key points of the human body in the multiple action images. Therefore, the action features include information in two dimensions, space and time.
[0042] After obtaining the action features in step S500, in step S600, the number of sit-ups in the video is determined according to the action features. Specifically, the action execution of the athlete in the multi-frame action image can be judged according to the action features, that is, the dynamic changes of each human key point within a certain period of time; if the action of the athlete meets the standard sit-ups, the number of sit-ups corresponding to the athlete increases by 1; if the action of the athlete does not meet the standard sit-ups, the number of sit-ups corresponding to the athlete remains unchanged. Finally, the spatiotemporal graph convolution model outputs whether the sit-ups of the athlete are standard and the number of sit-ups.
[0043] Through the above process, the dynamic changes of each human key point within a certain period of time can be fully utilized to accurately count the sit-ups. Even if one or more frames of images are missing in the video corresponding to the sit-ups, the sit-ups can still be counted by the change trend of the human key points, which increases the robustness of counting sit-ups.
[0044] In summary, the method for detecting sit-ups provided by the disclosed embodiment identifies one or more human key points of a moving person in each frame of a video through a key point recognition model; preliminarily determines whether a sit-up is standard based on the positional relationship between each human key point; if the sit-up is standard, determines multiple frames of action images corresponding to the sit-up; uses a spatiotemporal graph convolution model to perform feature extraction on multiple human key points in multiple frames of the action images and the time corresponding to the multiple frames of the action images to obtain action features, thereby extracting and fusing the spatial features and time features of multiple human key points to obtain action features including information in two dimensions of space and time; based on the action features, detects the number of sit-ups in the video, thereby detecting complex sit-ups from two dimensions of space and time, reducing interference from other factors, and accurately determining the number of sit-ups.
[0045] Figure 2 An exemplary flow chart of a method for identifying key points of a human body in each frame of a video according to some embodiments of the present disclosure is shown. It can be understood that the method for identifying key points of a human body in each frame of a video is a specific implementation of the aforementioned step S200, so the aforementioned method is combined with the method described in the preceding paragraphs. Figure 1 and Figure 2 The features described can analogously apply here.
[0046] As shown in the figure, in step S210, the human features of the athlete are extracted by the human feature extraction module to obtain the human features of the athlete; in step S220, the human features are semantically enhanced by the feature enhancement module to obtain human enhanced features; in step S230, one or more human key points are determined according to the human enhanced features using the key point determination module.
[0047] In some embodiments, the key point recognition model includes a human feature extraction module, a feature enhancement module and a key point determination module. For example, the key point recognition model is yolov8-pose, and the human feature extraction module is the backbone module in yolov8-pose, that is, the backbone network. The human feature extraction module is used to extract human features of athletes to obtain the human features of athletes. Here, "human features" refer to the features obtained after the human feature extraction module extracts features of athletes. The feature enhancement module is a high-level semantic enhancement model, which is used to perform high semantic extraction of human features, that is, feature enhancement, to obtain human enhanced features. Here, "human enhanced features" refer to the features after human features are enhanced. The key point determination module is the output network of yolov8-pose, that is, the head network, which is used to determine one or more of the human key points according to the human enhanced features. It should be noted that those skilled in the art can understand the structure and specific functions of yolov8-pose, backbone network and head network, which will not be repeated here.
[0048] Through steps S210-S230, the human features of the athlete can be extracted first, and the extracted human features can be enhanced to obtain human enhanced features, which can capture more high-level semantic features, enrich the feature information, and expand the ability to capture human features. Therefore, when determining one or more human key points based on the human enhanced features, one or more human key points can be more accurately determined based on the rich feature information, that is, the accuracy of identifying human key points is improved.
[0049] Figure 3 An exemplary flow chart of a method for obtaining human body enhancement features in some embodiments of the present disclosure is shown. It can be understood that the method for obtaining human body enhancement features is a specific implementation of the aforementioned step S220, so the aforementioned method is combined with Figure 1 , Figure 2 and Figure 3 The features described can analogously apply here.
[0050] As shown in the figure, in step S221, the dependency between the channel feature and the width feature is extracted by the first dependency extraction unit to obtain the channel width human body feature; in step S222, the dependency between the height feature and the channel feature is extracted by the second dependency extraction unit to obtain the height channel human body feature; in step S223, the channel width human body feature and the height channel human body feature are feature aggregated to obtain the human body enhancement feature.
[0051] In some embodiments, due to the features of different dimensions of the extracted features, the human body features include height features, width features and channel features. The height feature refers to the feature on the height dimension in the feature map, which can represent the height of the local area. The width feature refers to the feature on the width dimension in the feature map, which can represent the width of the local area. The channel feature refers to the information of each channel in the feature map, which represents the different feature types extracted from the input image. Those skilled in the art can understand that the features extracted by the backbone network include height features, width features and channel features.
[0052] In some embodiments, the feature enhancement module includes a first dependency extraction unit and a second dependency extraction unit, and the first dependency extraction unit and the second dependency extraction unit have the same structure. The first dependency extraction unit and the second dependency extraction unit can use an attention mechanism to achieve feature enhancement.
[0053] In some embodiments, the first dependency extraction unit is used to extract the dependency between the channel feature and the width feature to obtain the channel width human body feature. "The dependency between the channel feature and the width feature" refers to the mapping relationship between the channel feature and the width feature. After extracting the dependency between the channel feature and the width feature, the human body feature is enhanced according to the dependency between the channel feature and the width feature, and the feature obtained after the feature enhancement in step S221 is called the "channel width human body feature".
[0054] In some embodiments, the second dependency extraction unit is used to extract the dependency between the height feature and the channel feature to obtain the height channel human body feature. "The dependency between the height feature and the channel feature" refers to the mapping relationship between the height feature and the channel feature. After extracting the dependency between the height feature and the channel feature, the human body feature is enhanced according to the dependency between the height feature and the channel feature. The feature obtained after the feature enhancement in step S222 is called the "height channel human body feature".
[0055] Figure 4An exemplary flow chart of a method for obtaining channel width human body features in some embodiments of the present disclosure is shown. It can be understood that the method for obtaining channel width human body features is a specific implementation of the aforementioned step S221, so the aforementioned method is combined with Figure 1 , Figure 2 and Figure 3 The features described can analogously apply here.
[0056] As shown in the figure, in step S2211, the human body features are transposed by the first dependency extraction unit to obtain the height channel width human body features; and in step S2212, the height channel width human body features are pooled to obtain the channel width pooled features; and in step S2213, the channel width pooled features are convolved to obtain the channel width convolution features; in step S2214, the channel width convolution features are mapped by the activation function to obtain the channel width attention weights between the channel features and the width features; in step S2215, the human body features are weighted according to the channel width attention weights to obtain preliminary channel width human body features; in step S2216, the preliminary channel width human body features are transposed to obtain the channel width human body features.
[0057] In some embodiments, the human body features include channel features, height features, and dimension features, and the feature sequence in the human body features is channel features, height features, and dimension features. In step S2211, the human body features are transposed, thereby converting the human body features into features whose feature sequence is height features, channel features, and dimension features. Here, the features whose feature sequence is height features, channel features, and dimension features are referred to as height channel width human body features.
[0058] Secondly, in step S2212, the height-channel-width human body features can be pooled through a pooling layer to obtain channel-width pooled features, which can reduce redundancy in the height-channel-width human body features. The "channel-width pooled features" refer to the features obtained after the height-channel-width human body features are pooled.
[0059] Next, in step S2213, the channel width pooling feature can be convolved through a convolution layer to obtain a channel width convolution feature, thereby further extracting features from the channel width pooling feature to enrich the information in the channel width convolution feature. The channel width convolution feature refers to a feature obtained after convolution processing is performed on the channel width pooling feature.
[0060] Then, in step S2214, the channel width convolution feature is mapped by the activation function to obtain the channel width attention weight between the channel feature and the width feature. The activation function here can be sigmoid, softmax, etc. The channel width attention weight refers to the feature obtained after the activation function maps the channel width convolution feature, and the channel width attention weight is a number greater than 0 and less than or equal to 1. The channel width attention weight is used to represent the dependency between the channel feature and the width feature, which can be represented by similarity. If the channel width attention weight is closer to 1, it means that the correlation between the channel feature and the width feature is greater, that is, the stronger the dependency; conversely, the farther the channel width attention weight is from 1, the smaller the correlation between the channel feature and the width feature, that is, the weaker the dependency.
[0061] Next, in step S2215, the human body features are weighted according to the channel width attention weights, that is, weighted operations are performed on the human body features according to the channel width attention weights, and the weighted operation results are called preliminary channel width human body features.
[0062] Finally, in step S2216, the preliminary channel width human body feature is transposed, and the feature after the preliminary channel width human body feature is transposed is called the "channel width human body feature".
[0063] Through steps S2221-S2216, the dependency between the channel feature and the width feature can be extracted, and the human body feature can be weighted according to the channel width attention weight, so that the dependency between the channel feature and the width feature is integrated into the human body feature to obtain the channel width human body feature.
[0064] Figure 5 FIG. 1 is an exemplary flow chart of a method for obtaining a human body feature in a height channel according to some embodiments of the present disclosure. It can be understood that the method for obtaining a human body feature in a height channel is a specific implementation of the aforementioned step S222, so the aforementioned method is combined with Figure 1 - Figure 4 The features described can analogously apply here.
[0065] As shown in the figure, in step S2221, the human body features are transposed by the second dependency extraction unit to obtain the width-height channel human body features; and in step S2222, the width-height channel human body features are pooled to obtain the height channel pooled features; and in step S2223, the height channel pooled features are convolved to obtain the height channel convolution features; in step S2224, the height channel convolution features are mapped by an activation function to obtain the height channel attention weight between the height features and the channel features; in step S2225, the human body features are weighted according to the height channel attention weight to obtain the preliminary height channel human body features; in step S2226, the preliminary height channel human body features are transposed to obtain the height channel human body features.
[0066] First, in step S2221, the human body features are transposed by the second dependency extraction unit to convert the human body features into features whose feature order is width feature, height feature and channel feature. The features whose feature order is width feature, height feature and channel feature are called "width-height-channel human body features".
[0067] Secondly, in step S2222, the width and height channel human body features can be pooled through the pooling layer to obtain the height channel pooling features, which can reduce the redundancy in the width and height channel human body features. The "height channel pooling features" refer to the features obtained after the width and height channel human body features are pooled.
[0068] Next, in step S2223, the height channel pooling feature can be convolved through a convolution layer to obtain a height channel convolution feature, thereby further extracting features from the height channel pooling feature to enrich the information in the height channel convolution feature. The height channel convolution feature refers to a feature obtained after convolution processing is performed on the height channel pooling feature.
[0069] Then, in step S2224, the high-level channel convolution feature is mapped by the activation function to obtain the high-level channel attention weight between the high-level feature and the channel feature. The activation function here can be sigmoid, softmax, etc. The high-level channel attention weight refers to the feature obtained after the activation function maps the high-level channel convolution feature, and the high-level channel attention weight is a number greater than 0 and less than or equal to 1. The high-level channel attention weight is used to represent the dependency between the high-level feature and the channel feature, which can be represented by similarity. If the high-level channel attention weight is closer to 1, it means that the correlation between the high-level feature and the channel feature is greater, that is, the stronger the dependency; conversely, the farther the high-level channel attention weight is from 1, the smaller the correlation between the high-level feature and the channel feature, that is, the weaker the dependency.
[0070] Next, in step S2225, the human body features are weighted according to the height channel attention weights, that is, weighted operations are performed on the human body features according to the height channel attention weights, and the weighted operation results are called preliminary height channel human body features.
[0071] Finally, in step S2226, the preliminary height channel human body features are transposed, and the features after the preliminary height channel features are transposed are called "height channel human body features".
[0072] Through steps S2221-S2216, the dependency between the height feature and the channel feature can be extracted, and the human body feature is weighted according to the height channel attention weight, so that the dependency between the height feature and the channel feature is integrated into the human body feature to obtain the height channel human body feature.
[0073] In some embodiments, the key points of the human body include at least one of the hands, head, ankles, knees and hip joints of the athlete. It will be appreciated by those skilled in the art that the key points of the human body disclosed herein are not subject to any limitation. The starting point of the first vector is the ankle, the end point of the first vector is the ankle, the starting point of the second vector is the ankle, and the end point of the second vector is the hip joint.
[0074] Figure 6 An exemplary flow chart of a method for determining whether a sit-up is standard in some embodiments of the present disclosure is shown. It can be understood that the method for determining whether a sit-up is standard is a specific implementation of the aforementioned step S300, so the aforementioned method is combined with Figure 1 and Figure 6 The features described can analogously apply here.
[0075] As shown in the figure, in step S310, the distance between the hand and the head is compared with the first threshold, and the angle between the first vector and the second vector is compared with the preset angle threshold; if the distance between the hand and the head is less than the first threshold, and the angle between the first vector and the second vector is less than the angle threshold, then in step S320, the sit-up standard is determined; otherwise, in step S330, it is determined that the sit-up is not standard.
[0076] In some embodiments, the first threshold is used to determine whether the distance between the hand and the head is too large. The first threshold can be set by oneself. Specifically, the first threshold can be the length of the arm. If the distance between the hand and the head is greater than the first threshold, it indicates that the distance between the hand and the head is too large, and it does not meet the standard of standard sit-ups; conversely, if the distance between the hand and the head is less than or equal to the first threshold, it indicates that the distance between the hand and the head is small, which meets the standard of standard sit-ups. The angle between the first vector and the second vector refers to the angle between the thigh and the calf. The angle threshold is used to determine whether the angle between the thigh and the calf meets the standard of standard sit-ups. The angle threshold can be set by oneself. For example, the angle threshold can be 30 degrees. If the angle between the first vector and the second vector is greater than the angle threshold, it indicates that the angle between the thigh and the calf meets the standard of standard sit-ups.
[0077] It should be noted that when the distance between the hand and the head is less than the first threshold, and the angle between the first vector and the second vector is less than the angle threshold, the sit-up meets the standard sit-up criteria, and the sit-up count is increased by 1.
[0078] Through steps S310-S330, according to the distance between the hand and the head and the angle between the first vector and the second vector, it can be determined whether the sit-up meets the standard sit-up standard. If the sit-up meets the standard sit-up standard, the sit-up count is increased by 1.
[0079] Figure 7 FIG. 1 is an exemplary flow chart of a method for obtaining the motion characteristics of a sit-up in some embodiments of the present disclosure. It can be understood that the method for obtaining the motion characteristics of a sit-up is a specific implementation of the aforementioned step S500, so the aforementioned method is combined with the method described in FIG. Figure 1 and Figure 5 The features described can analogously apply here.
[0080] As shown in the figure, in step S510, spatial feature extraction is performed on multiple human key points in multiple frames of action images to obtain key point spatial features of the human key points; in step S520, temporal feature extraction is performed on the time corresponding to the multiple frames of action images to obtain image temporal features of each frame of image; in step S530, feature fusion of key point spatial features and image temporal features is performed to obtain action features.
[0081] First, in the aforementioned step S510, the spatial features of multiple human key points are extracted through the spatiotemporal graph convolution model, and the extracted spatial features of the human key points are called “key point spatial features”.
[0082] In step S520, the time features of each human key point in the multiple frames of the action images are extracted through the spatiotemporal graph convolution model to obtain the time features of each frame of the image. The extracted time features of each frame of the image are referred to as "image time features". It should be noted that the time of each human key point in each frame of the action image is the same, and the time features of each human key point are also the same.
[0083] It should be noted that the above step S520 and step S530 are performed alternately, and the execution order of step S520 and step S530 can be set by oneself.
[0084] Next, in step S530, the key point spatial features and the image temporal features are fused, and the fused features are used as action features.
[0085] Through steps S520-S530, the spatiotemporal graph convolution model can be used to extract features of action images and key points of the human body from two dimensions of time and space, so that the obtained action features include spatial information and time information, thereby enriching the feature information, facilitating the subsequent detection of complex sit-ups from two dimensions of space and time based on the action features, and reducing interference from other factors, thereby accurately determining the number of sit-ups in the video.
[0086] In some embodiments, the method for detecting sit-ups further comprises: performing motion analysis according to the motion features to obtain motion analysis results of the athlete; the motion analysis results include the amplitude and frequency of sit-ups. It can be seen from the aforementioned embodiments that since the motion features can represent the position change, change speed and motion trend of each key point of the human body, motion analysis can be performed according to the motion features to obtain the motion analysis results of the athlete.
[0087] In summary, the disclosed embodiments provide a method for detecting sit-ups, which uses a key point recognition model to identify one or more human key points of a moving person in each frame of a video; based on the positional relationship between each human key point, preliminarily determine whether the sit-up is standard; if the sit-up is standard, determine multiple frames of action images corresponding to the sit-up; use a spatiotemporal graph convolution model to extract features of multiple human key points in multiple frames of the action images and the time corresponding to the multiple frames of the action images to obtain action features, so that the spatial features and time features of multiple human key points can be extracted and fused, thereby obtaining action features including information in two dimensions of space and time; based on the action features, the number of sit-ups in the video is detected, so that complex sit-ups can be detected from two dimensions of space and time, which can reduce interference from other factors, thereby accurately determining the number of sit-ups.
[0088] The disclosed embodiment also provides a device for detecting sit-ups.
[0089] Figure 8 An exemplary structural block diagram of an apparatus for detecting sit-ups according to an embodiment of the present disclosure is shown.
[0090] As shown in the figure, the device 800 for detecting sit-ups includes: a processor 810 and a memory 820 .
[0091] In the embodiment of the present disclosure, the processor 810 is configured to execute program instructions. The memory 820 is configured to store program instructions, and when the program instructions are loaded and executed by the processor 810, the apparatus 800 for detecting sit-ups executes the aforementioned method for detecting long jump.
[0092] In summary, in the device for detecting sit-ups of the disclosed embodiment, when the program instructions are loaded by the processor 810 and the aforementioned method for detecting long jump is executed, the device 800 for detecting sit-ups can identify one or more human key points of the athlete in each frame of the video through the key point recognition model; preliminarily determine whether the sit-ups are standard based on the positional relationship between each human key point; if the sit-ups are standard, determine the multiple frames of action images corresponding to the sit-ups; use the spatiotemporal graph convolution model to extract features of multiple human key points in the multiple frames of the action images and the time corresponding to the multiple frames of the action images to obtain action features, so that the spatial features and time features of the multiple human key points can be extracted and fused, thereby obtaining action features including information in two dimensions of space and time; according to the action features, the number of sit-ups in the video is detected, so that the sit-ups with complex processing can be detected from two dimensions of space and time, which can reduce the interference of other factors, thereby accurately determining the number of sit-ups.
[0093] The disclosed embodiment further provides a computer-readable medium having a computer program stored thereon, and when the computer program is executed by a processor, the processor executes the method for detecting long jump provided in the first aspect.
[0094] Although multiple embodiments of the present disclosure have been shown and described herein, it will be apparent to those skilled in the art that such embodiments are provided by way of example only. Those skilled in the art may think of many changes, modifications, and alternatives without departing from the thought and spirit of the present disclosure. It should be understood that in the process of practicing the present disclosure, various alternatives to the embodiments of the present disclosure described herein may be adopted. The attached claims are intended to define the scope of protection of the present disclosure, and therefore cover equivalents or alternatives within the scope of these claims.
Claims
1. A method for detecting sit-ups, comprising: Get videos of exercisers doing sit-ups; Identify one or more key points of the human body of the moving person in each frame of the video by a key point recognition model; Preliminarily determining whether the sit-up is standard based on the positional relationship between the key points of the human body; In the case of the sit-up standard, determining a plurality of frames of action images corresponding to the sit-up; Using a spatiotemporal graph convolution model, feature extraction is performed on a plurality of human key points in a plurality of frames of the action images and the time corresponding to the plurality of frames of the action images to obtain the action features of the sit-ups; and The number of sit-ups in the video is determined according to the action feature.
2. The method according to claim 1, characterized in that The key point recognition model includes a human feature extraction module, a feature enhancement module and a key point determination module, and the step of identifying one or more human key points of the moving person in each frame of the video by the key point recognition model includes: Extracting human features of the athlete through the human feature extraction module to obtain human features of the athlete; Performing semantic enhancement on the human body features by means of the feature enhancement module to obtain human body enhanced features; The key point determination module is used to determine one or more key points of the human body according to the human body enhancement feature.
3. The method according to claim 2, characterized in that The feature enhancement module includes a first dependency extraction unit and a second dependency extraction unit, wherein the first dependency extraction unit and the second dependency extraction unit have the same structure, the human body features include height features, width features and channel features, wherein the semantic enhancement of the human body features by the feature enhancement module to obtain human body enhancement features includes: Extracting the dependency between the channel feature and the width feature by the first dependency extraction unit to obtain a channel width human body feature; Extracting the dependency between the height feature and the channel feature by the second dependency extraction unit to obtain a height channel human body feature; The channel width human body feature and the height channel human body feature are subjected to feature aggregation to obtain the human body enhancement feature.
4. The method according to claim 3, characterized in that The extracting the dependency between the channel feature and the width feature by the first dependency extraction unit to obtain the channel width human body feature includes: Transposing the human body feature by the first dependency extraction unit to obtain a height channel width human body feature; and Performing pooling processing on the height channel width human body features to obtain channel width pooling features; and Convolution is performed on the channel width pooling feature to obtain the channel width convolution feature; The channel width convolution feature is mapped by an activation function to obtain a channel width attention weight between the channel feature and the width feature; Performing weighted processing on the human body feature according to the channel width attention weight to obtain a preliminary channel width human body feature; The preliminary channel width human body feature is transposed to obtain the channel width human body feature.
5. The method according to claim 3, characterized in that: The extracting the dependency between the height feature and the channel feature by the second dependency extraction unit to obtain the height channel human body feature comprises: Transposing the human body features by the second dependency extraction unit to obtain human body features of width and height channels; and Performing pooling processing on the width-height channel human body features to obtain height channel pooling features; and Perform convolution processing on the height channel pooling features to obtain the height channel convolution features; Mapping the height channel convolution feature through an activation function to obtain a height channel attention weight between the height feature and the channel feature; Performing weighted processing on the human features according to the height channel attention weights to obtain preliminary height channel human features; The preliminary height channel human body features are transposed to obtain the height channel human body features.
6. The method according to claim 1, characterized in that The human body key points include at least one of the hands, head, ankles, knees and hip joints of the athlete, wherein the starting point of the first vector is the ankle, the end point of the first vector is the ankle, the starting point of the second vector is the ankle, and the end point of the second vector is the hip joint, wherein the preliminary determination of whether the sit-up is standard based on the positional relationship between the key points of the human body includes: Determine the relationship between the distance between the hand and the head and a first threshold value and the relationship between the angle between the first vector and the second vector and a preset angle threshold value; If the distance between the hand and the head is less than the first threshold, and the angle between the first vector and the second vector is less than the angle threshold, determining the sit-up standard; Otherwise, it is determined that the sit-ups are not standard.
7. The method according to claim 1, characterized in that The step of extracting features of a plurality of human body key points in a plurality of frames of the action images and the time corresponding to the plurality of frames of the action images to obtain the action features of the sit-ups comprises: Extracting spatial features of a plurality of human key points in a plurality of frames of the action images to obtain key point spatial features of the human key points; Extracting time features of the times corresponding to the multiple frames of the action images to obtain image time features of each frame of the image; The key point spatial feature and the image temporal feature are fused to obtain the action feature.
8. The method according to claim 1, characterized in that The method further comprises: Performing motion analysis according to the motion characteristics to obtain the motion analysis result of the athlete; the motion analysis result includes the amplitude and frequency of sit-ups.
9. A device for detecting sit-ups, comprising: a processor configured to execute program instructions; as well as A memory configured to store the program instructions, which, when loaded and executed by the processor, causes the apparatus to perform the method according to any one of claims 1 to 8.
10. A computer-readable storage medium having program instructions stored thereon, wherein when the program instructions are loaded and executed by a processor, the method according to any one of claims 1 to 8 is implemented.