A Skipping Rope Counting Method Based on Semi-Supervised Video Object Segmentation
Through the method based on semi-supervised video target segmentation, the STCN segmentation network and feature pixel-by-pixel random inactivation mechanism are used to solve the problem of large and complex data processing volume in the prior art, and the fast and accurate counting of single and multi-person skipping ropes is achieved.
Patent Information
- Application Number
- CN202210836154.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-15
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-07-15
AI Technical Summary
The prior art has problems such as large data processing volume and complex data processing process in the rope skipping counting method. Especially in the multi-person rope skipping scenario, the counting error is large, and rapid detection cannot be achieved.
Using a method based on semi-supervised video target segmentation, the STCN segmentation network and feature pixel-by-pixel random inactivation mechanism is used to segment the rope jump character image, and count it through the relative position relationship between the hand and the rope, and a memory bank is constructed to record the position changes and realize the rope jump count.
It realizes fast and accurate counting in single-player and multi-player rope skipping scenarios, reduces data processing volume, and improves the robustness and speed of counting.
Smart Images

Figure CN115171019B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of skipping exercise monitoring methods, and specifically to a skipping counting method based on semi-supervised video object segmentation. Background Art
[0002] Currently, there are many applications for automatically counting fitness exercises such as skipping rope, bouncing a ball, and kicking a shuttlecock using video, audio, etc. However, there are still defects such as easy mis-counting through body part recognition or only being able to count single-person skipping.
[0003] A skipping counting method based on image information and a skipping counting method based on video image object recognition disclosed in Chinese patent documents with publication numbers CN109876416A and CN110210360A count by identifying the positions of the rope and the face. A skipping counting method based on the cross-correlation coefficient method disclosed in Chinese patent document with publication number CN110102040A judges the number of skipping times by identifying the sound of the rope touching the ground. A skipping counting method based on deep learning disclosed in Chinese patent document with publication number CN112044046A preprocesses the obtained image data, then classifies using a trained model, judges the current motion state based on the classification result, and finally counts the number of changes in the skipping state for counting. These methods can only count single-person skipping and cannot achieve counting in a multi-person skipping scenario.
[0004] A skipping counting method based on multi-object tracking disclosed in Chinese patent document with publication number CN112044047A determines the number of skipping times by identifying changes in the face position. During the skipping process, when a person approaches or moves away from the camera, the position and proportion of the face in the frame often change over time, thus causing counting errors.
[0005] A skipping counting method based on health monitoring disclosed in Chinese patent document with publication number CN 113627396 A discloses a technical means for skipping counting by processing skipping images through a PoseNet model. Although the technical means of this method can be used for single-person and multi-person skipping counting, it needs to select a relatively large number of (17 in the literature) body pose key points for heatmap encoding and decoding, and also perform jitter suppression calculations for jitter, in order to obtain the skipping count. Therefore, there are problems of large data processing volume and complex data processing process, and rapid detection of skipping counting cannot be achieved. Summary of the Invention
[0006] The object of the present invention is to provide a skipping rope counting method based on semi-supervised video object segmentation, so as to solve the problems of large data processing volume and complex data processing process existing in the skipping rope counting method based on image processing in the prior art.
[0007] In order to achieve the above object, the technical solution adopted by the present invention is as follows:
[0008] A skipping rope counting method based on semi-supervised video object segmentation includes the following steps:
[0009] Step 1, obtain the original video data of the skipping rope movement, and extract the image data of each frame from the original video data;
[0010] Step 2, based on the semi-supervised video object segmentation method, adopt the STCN segmentation network, introduce the feature per-pixel random inactivation mechanism, and segment the skipping rope person image from the image data of each frame obtained in Step 1, where:
[0011] When there is a single person in the image data of each frame, based on the semi-supervised video object segmentation method and the feature per-pixel random inactivation mechanism, a single skipping rope person image is segmented from the image data of each frame;
[0012] When there are multiple people in the image data of each frame, based on the semi-supervised video object segmentation method and the feature per-pixel random inactivation mechanism, first segment the pixel distribution of all person Mask images, and then block each skipping rope person from the Mask image data of each frame, thereby obtaining the skipping rope person images of each person in the image data of each frame;
[0013] Step 3, determine the relative position of the person's hand and the rope from each skipping rope person image corresponding to each frame obtained in Step 2, and then count the number of times the relative position of the hand and the rope changes in each skipping rope person image of all frames, thereby obtaining the skipping rope count of each skipping rope person.
[0014] Further, in Step 2, first train the STCN segmentation network using static images and continuous / intermittent frames in various video sequences to enable it to segment a large number of different types of people / objects, and then apply it to the skipping rope video sequence to segment single-person / multi-person skipping Mask images;
[0015] The training is divided into two stages. The first stage is the pre-training stage. In the pre-training stage, several image data are extracted from the image data of each frame obtained in Step 1, and then trained based on the feature per-pixel inactivation mechanism; the second stage is the public dataset training stage. The public dataset training stage uses a public dataset, and the feature per-pixel inactivation mechanism is incorporated in both training stages.
[0016] Further, when there are multiple people in each frame of image data in step 2, according to the pixel distribution of the Mask map, pixel statistics are performed column by column from left to right to complete the segmentation. The process is as follows:
[0017] Step 2-1: Starting from the first column in the left-to-right direction in the Mask map obtained by segmenting the STCN segmentation network, which contains people and ropes, when there are no person pixel points in a certain column, continue to detect the right column. When person pixel points are first detected, mark them as person 0, and continue to detect the presence of person and rope pixels to the right until the person disappears. When the column density value of the rope pixel presence exceeds the preset threshold, determine that the person is a rope jumper; otherwise, determine that the person is a bystander or a counter, etc., who is not a rope jumper. The determination of person 0 ends;
[0018] Step 2-2: Continue to detect to the right until person pixel points are detected again, mark them as person 1, and continue to the right until the person disappears in the pixel column. Count the number of columns containing rope pixels to determine whether the person is a rope jumper or other irrelevant people;
[0019] Step 2-3: Repeat the above steps 2-1 and 2-2 until the last pixel column, and finally complete the segmentation of the rope-jumping people.
[0020] Further, in step 3, record the relative positions of the hands of each rope-jumping person in each corresponding frame of the image and the rope, and construct a memory bank for the video sequence based on this. The memory bank is a class created immediately in the program, which is used to store and access the number of times the relative positions of the hands and the rope in each rope-jumping person image in the video frame change, thereby obtaining the rope-jumping count of each rope-jumping person.
[0021] Further, the construction process of the memory bank is as follows:
[0022] (1) In each rope-jumping person image corresponding to each frame, find the corresponding height values h1 and h2 of the outermost pixels of the left and right hands in the image, and take the average value as the height position h of the hand in the image, that is
[0023] (2) Divide each rope-jumping person image into upper and lower parts through the height position h, and collect the number of rope pixels c1 and c2 corresponding to the upper and lower parts;
[0024] (3) Use the difference (c1 - c2) between c1 and c2 as the final determination value. A positive number indicates that the rope is above the hand of the person, and a negative number indicates that the rope is below the hand of the person, thereby obtaining the relative position state of the hand and the rope in each rope-jumping person image;
[0025] (4) Record the relative position state obtained in step (3) and form a memory bank based on this.
[0026] Compared with the prior art, the advantages of the present invention are as follows:
[0027] The present invention processes video information by adopting a video image segmentation strategy based on semi-supervised video object segmentation, and then counts the number of changes in the skipping rope state according to the judgment result of the relative position relationship between the hand and the rope. The whole process has strong robustness, high accuracy and fast counting speed. Compared with the prior art, it has the advantages of small data processing volume and simple data processing process. It can not only count single-person skipping ropes, but also count multi-person skipping ropes in real time, and has high application value. Brief Description of the Drawings
[0028] Figure 1 It is a flowchart of the skipping rope counting method based on semi-supervised video object segmentation in an embodiment of the present invention.
[0029] Figure 2 It is a schematic diagram of the per-pixel random inactivation architecture (taking the first Residual Block of ResNet50 as an example) in an embodiment of the present invention.
[0030] Figure 3 It is a schematic diagram of the segmented person and rope Mask images in an embodiment of the present invention, where the left is person 0 and the right is person 1.
[0031] Figure 4 It is a schematic diagram of the relative position relationship between the hand and the rope in an embodiment of the present invention.
[0032] Figure 5 It is a schematic diagram of multi-person partitioning in an embodiment of the present invention. Detailed Embodiments
[0033] To make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described below in conjunction with the embodiments and their accompanying drawings. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the described embodiments fall within the scope of protection of the present invention.
[0034] Unless otherwise defined, the technical terms or scientific terms used in this invention shall have the ordinary meanings as understood by those of ordinary skill in the field to which this invention belongs. The words such as "including" or "comprising" used in this invention mean that the elements or objects appearing before this word cover the elements or objects listed after this word and their equivalents, without excluding other elements or objects. The words such as "connected" or "coupled" are not limited to physical or mechanical connections, but may include electrical connections, whether direct or indirect. The terms such as "upper", "lower", "left", and "right" are only used to indicate relative positional relationships. When the absolute position of the object being described changes, the relative positional relationship may also change accordingly.
[0035] As Figure 1 shown, a skipping rope counting method based on semi-supervised video object segmentation in this embodiment includes the following steps:
[0036] S100. Acquire the original video data of the skipping rope movement through a camera, and extract each frame of image data from the original video data.
[0037] S200. Based on the semi-supervised video object segmentation method (Semi-Supervised Video Object Segmentation, SVOS), adopt the STCN (Space-Time Correspondence Networks) segmentation network, introduce the feature pixel-wise random inactivation mechanism, and segment the skipping rope person image from each frame of image data obtained in step 1.
[0038] The semi-supervised video object segmentation task assumes that the complete Mask of the object of interest (the skipping rope person in this embodiment) is input in the first frame of the video sequence, and requires the algorithm to generate a segmentation Mask for this object in all subsequent frames.
[0039] As Figure 2 shown, this embodiment introduces the feature pixel-wise random inactivation mechanism (Pixel-Wise Stochastic Deactivation, PWSD) into the semi-supervised segmentation algorithm. The algorithm adopts the STCN (Space-Time Correspondence Networks) segmentation network, introduces the feature pixel-wise random inactivation mechanism, and applies it to the counting of single-person skipping rope repetitive actions. That is, when there is only one person in each frame of image data, based on the feature pixel-wise random inactivation mechanism in the semi-supervised video object segmentation method, a single skipping rope person image is segmented from each frame of image data.
[0040] Figure 2It is an illustration of a pixel-by-pixel random inactivation architecture. In any Residual Block in ResNet, for the feature map obtained from the fourth convolutional kernel (1×1 convolution), all pixels are subjected to the following operations according to a predetermined probability:
[0041]
[0042] Finally, the feature map after the first convolutional kernel (7×7 convolution) and MaxPool processing is added to the feature map after the fourth convolutional kernel (1×1 convolution) through the above operations to obtain the final output feature map of this Residual Block.
[0043] Specifically, first, the STCN segmentation network is trained using static images and continuous / intermittent frames in various video sequences for segmentation ability, enabling it to segment a large number of different types of people / objects, and then it is applied to the skipping rope video sequence to segment single / multi-person skipping Mask images;
[0044] The training is divided into two stages, s0 and s3. The first stage, s0, is the pre-training stage. In the pre-training stage, several image data are extracted from the static image data of each frame obtained in step 1, and then training is carried out based on the feature pixel-by-pixel inactivation mechanism; the second stage, s3, is the public dataset training stage, and the public dataset is used in the public dataset training stage. The feature pixel-by-pixel inactivation mechanism is incorporated in both training stages.
[0045] When there are multiple people in each frame of image data, each frame of image data is segmented based on the semi-supervised video object segmentation method, and the Mask image as shown in Figure 3 will be generated. Figure 3 On the left is person 0 and on the right is person 1. According to the pixel distribution of the Mask map, each skipping person in each frame of image data is segmented, and thus the image of each skipping person in each frame of image data is obtained.
[0046] Since the skipping rope counting scheme based on semi-supervised segmentation can segment multiple people in a frame more quickly according to the pixel distribution of the Mask map and complete image cutting without introducing an additional person detection model, it can more conveniently perform simultaneous counting of multiple people skipping rope. Since the Mask map person pixel blocks of video frames are easier to segment and usually exist continuously, while the rope target is smaller and the pixel blocks are discontinuous, pixel statistics can be performed column by column from left to right, specifically as shown in Figure 5 and the process is as follows:
[0047] (a) Starting from the first column in the left-to-right direction in the Mask image containing the person and the rope obtained by STCN network segmentation, when there are no person pixels in a certain column, continue to detect the right column. When person pixels are first detected, mark it as person 0, and continue to detect the presence of person and rope pixels to the right until the person disappears. When the column density value of the presence of rope pixels exceeds the preset threshold, determine that this person is a rope skipper; otherwise, determine that it is a bystander or a counter, etc., a non-rope skipper. The determination of person 0 ends;
[0048] (b) Continue to detect to the right until person pixels are detected again, mark it as person 1, and continue to the right until the person disappears in the pixel column, and count the number of columns containing rope pixels to determine whether the person is a rope skipper or other irrelevant persons;
[0049] (c) Repeat the above steps 2-1 and 2-2 until the last column of pixel columns, and finally complete the segmentation of the rope-skipping persons. The segmented rope-skipping person images are as Figure 5 shown, Figure 5 and both of the two persons in it are rope-skipping persons.
[0050] S300. Determine the relative position between the person's hand and the rope in each rope-skipping person image corresponding to each frame obtained from step S200, and then count the number of changes in the relative position between the hand and the rope in each rope-skipping person image of all frames, thereby obtaining the rope-skipping count of each rope-skipping person.
[0051] In this embodiment, the counting is based on the judgment result of the relative position relationship between the person's hand and the rope. During the rope-skipping process, due to the influence of many factors (such as camera shaking, the person approaching or moving away from the camera during rope-skipping), the position and proportion of the person in the picture often change with time. However, in a complete rope-skipping action, the rope will always appear above and below the hand. Since the hand is the intersection point between the person and the rope, no matter how the person's position changes in the picture, the relative position relationship between the hand and the rope will always be in a state where the rope is above the hand or the rope is below the hand. Therefore, this is used to judge the number of rope-skipping times.
[0052] Therefore, in this embodiment, by recording the relative position relationship between the person's hand and the rope in each rope-skipping person image of each frame in the video frame sequence, a memory bank for this video sequence is constructed accordingly. This memory bank is a class created immediately in the program, used to store and retrieve the relative position relationship between the person's hand and the rope corresponding to the video frame, and has a statistical function.
[0053] As Figure 4 shown, the construction process of the memory bank is as follows:
[0054] (S3001) In each skipping person image corresponding to each frame, find the corresponding height values h1 and h2 at which the outermost pixels of the left and right hands are located in the image, and take the average as the height position h of the hand in the image.
[0055] (S3002) Divide each skipping person image into upper and lower parts through the height position h, and collect the corresponding rope pixel numbers c1 and c2 of the upper and lower parts.
[0056] (S3003) Use the difference (c1 - c2) between c1 and c2 as the final determination value. A positive number indicates that the rope is above the person's hand, and a negative number indicates that the rope is below the person's hand, thereby obtaining the relative position state of the person's hand and the rope in each skipping person image.
[0057] (S3004) Record the relative position state obtained in step (S3003) and form a memory bank with it.
[0058] In the memory bank, count the number of times the relative position of the hand and the rope changes in each skipping person image of all frames, and thus the skipping count of each skipping person can be obtained.
[0059] S400. Output the skipping count obtained in step S300.
[0060] The embodiments described in the present invention are merely descriptions of the preferred embodiments of the present invention, and do not limit the concept and scope of the present invention. Without departing from the design idea of the present invention, various modifications and improvements made by those skilled in the art to the technical solutions of the present invention shall fall within the protection scope of the present invention. The technical content claimed by the present invention has been fully recorded in the claims.
Claims
1. A skipping count method based on semi-supervised video object segmentation, characterized in that, Including the following steps: Step 1: Obtain the original video data of rope skipping exercise, and extract each frame of image data from the original video data; Step 2: Based on the semi-supervised video object segmentation method, use the STCN segmentation network, introduce the feature per-pixel random inactivation mechanism, and segment the rope skipping person images from each frame of image data obtained in Step 1, where: When there is a single person in each frame of image data, based on the semi-supervised video object segmentation method and the feature per-pixel random inactivation mechanism, a single rope skipping person image is segmented from each frame of image data; When there are multiple people in each frame of image data, based on the semi-supervised video object segmentation method and the feature per-pixel random inactivation mechanism, first segment the pixel distribution of all people's Mask images, and then block each rope skipping person from each frame of Mask image data, thereby obtaining each rope skipping person image in each frame of image data; Step 3: Determine the relative position of the person's hand and the rope from each rope skipping person image corresponding to each frame obtained in Step 2, and then count the number of times the relative position of the hand and the rope changes in each frame of each rope skipping person image, thereby obtaining the rope skipping count of each rope skipping person.
2. The skipping rope counting method based on semi-supervised video object segmentation according to claim 1, wherein In Step 2, first train the STCN segmentation network using static images and continuous / intermittent frames in various video sequences for segmentation ability, so that it can segment a large number of different types of people / objects, and then apply it to the rope skipping video sequence to segment single-person / multi-person rope skipping Mask images; The training is divided into two stages. The first stage is the pre-training stage. In the pre-training stage, several image data are extracted from each frame of image data obtained in Step 1, and then trained based on the feature per-pixel inactivation mechanism; The second stage is the public dataset training stage. The public dataset training stage uses a public dataset, and the feature per-pixel inactivation mechanism is incorporated in both training stages.
3. A skipping rope counting method based on semi-supervised video object segmentation according to claim 1, characterized in that, In Step 2, when there are multiple people in each frame of image data, according to the pixel distribution of the Mask image, perform pixel statistics column by column from left to right to complete the blocking. The process is as follows: Step 2-1: Starting from the first column in the left-to-right direction in the Mask image containing people and ropes segmented by the STCN segmentation network, when there are no person pixel points in a certain column, continue to detect the right column. When person pixels are first detected, mark it as person 0, and continue to detect the presence of person and rope pixels to the right until the person disappears. When the column density value of the rope pixel exceeds the preset threshold, determine that the person is a rope skipper, otherwise determine that the person is a bystander or a counter, etc., a non-rope skipper, and the determination of person 0 ends; Step 2-2: Continue to detect to the right until person pixels are detected again, mark it as person 1, and continue to the right until the person disappears in the pixel column, and count the number of columns containing rope pixels to determine whether the person is a rope skipper or other irrelevant persons; Step 2-3: Repeat the above Steps 2-1 and 2-2 until the last column of pixel columns, and finally complete the blocking of the rope skipping people.
4. A skipping rope counting method based on semi-supervised video object segmentation according to claim 1, characterized in that In step 3, record the relative positions of the hands of the figures and the ropes in each skipping figure image corresponding to each frame, and construct a memory bank for the video sequence accordingly. The memory bank is a class created instantaneously in the program and is used to store and access the number of times the relative positions of the hands and the ropes in each skipping figure image in the video frames change, thereby obtaining the skipping counts of each skipping figure.
5. A skipping rope counting method based on semi-supervised video object segmentation according to claim 4, characterized in that, The construction process of the memory bank is as follows: (1) In each image of a rope-skipping person corresponding to each frame, find the corresponding height values h1 and h2 at which the outermost pixels of the left and right hands are located in the image, and take the average as the height position h of the hand in the image, that is (2) Divide each skipping figure image into upper and lower parts by the height position h, and collect the number of rope pixels c1 and c2 corresponding to the upper and lower parts; (3) Use the difference between c1 and c2, i.e., (c1 - c2), as the final determination value. A positive number indicates that the rope is above the hand of the figure, and a negative number indicates that the rope is below the hand of the figure, thereby obtaining the relative position state of the hand of the figure and the rope in each skipping figure image; (4) Record the relative position state obtained in step (3) and form a memory bank with this.
Citation Information
Patent Citations
Skipping rope counting method based on image information
CN109876416A
Audio rope skipping counting method based on cross-correlation coefficient method
CN110102040A
Rope skipping counting method based on video image target recognition
CN110210360A
Rope skipping counting method based on deep learning
CN112044046A
Rope skipping counting method based on health monitoring
CN113627396A