A YOLO-based 2D pose detection method
Patent Information
- Application Number
- CN202310094627.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-02-10
- Publication Date
- 2026-10-09
- Estimated Expiration
- 2043-02-10
AI Technical Summary
[0009]本发明要解决的技术问题是:提供一种基于YOLO的2D姿态估计方法,解决目前已有的基于YOLO架构在对超大目标的关键点定位时,由于都是基于目标所在anchor回归关键点的偏移量,导致anchor尺度过大、目标末端关键点距离anchor距离较远而产生了较大的误差的问题
[0016] The beneficial effects of this invention are that it solves the defects existing in the background technology, detects key points as individual targets, combines key points into the same target using the matching embedding method, and predicts the embedding of each key point simultaneously; during training, for the same target, the model converges to the shortest distance between the embeddings of each key point, and for different targets, it converges in the direction that increases the embedding distance; it retains the advantages of YOLO-based pose estimation methods, which have faster inference speed and smaller memory usage compared to heatmap methods, while improving the prediction accuracy of key points and adding almost no additional algorithm running time.
Smart Images

Figure CN115953806B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of computer vision detection technology, and in particular to a 2D pose detection method based on YOLO. Background Technology
[0002] Pose estimation is an important research area in computer vision, and it is currently widely used in human activity analysis, human-computer interaction, and video surveillance. Most pose estimation is human pose estimation, with some focusing on hand pose estimation. Human pose estimation refers to locating key points of the human body (such as shoulders, elbows, wrists, hips, knees, and ankles) in images or videos using computer algorithms. Hand pose estimation is divided into labeled and unlabeled pose estimation, used to understand the meaning of hand movements.
[0003] Pose estimation methods can be divided into traditional pose estimation and deep learning-based pose estimation.
[0004] Traditional pose estimation is primarily based on graph structure model methods. These methods consist of three parts: a graph model, an optimization algorithm, and a component appearance model. They provide a classic statistical model of objects, using a graph structure model to identify objects in an image. However, their drawback is that they rely on heuristic local searches and cannot find a globally optimal solution.
[0005] Deep learning is a self-explanatory learning method that is simple, convenient, and powerful, and is used in many fields. Pose estimation based on deep learning utilizes deep convolutional neural networks to enhance the performance of human body estimation systems. Compared to traditional methods, deep learning can obtain deeper image features and more accurately represent data, thus becoming a mainstream research direction. Deep learning methods are categorized into single-person pose estimation and multi-person pose estimation based on the number of people being detected. Single-person pose estimation is further divided into methods based on coordinate regression and heatmap detection; multi-person pose estimation can be divided into top-down and bottom-up methods.
[0006] Top-down approach involves first detecting targets and then using single-target keypoint detection methods to construct poses from the extracted target regions. The advantage of this method is that it avoids the matching and combination problem between multiple keypoints of the same category for multiple targets. The disadvantage is that it is highly dependent on the target detection performance; when targets are not fully detected, not all keypoints of the targets can be detected. Furthermore, the computational cost increases with the number of targets. Bottom-up approach, currently the mainstream method, first calculates all keypoints for all targets and then combines the keypoints to the corresponding targets. The process of matching and combining keypoints to targets increases the algorithm's complexity.
[0007] Commonly used methods for keypoint detection include heatmap-based methods and methods like YOLOPose that directly regress keypoint coordinates using the target model. Early coordinate regression methods were used for single-target keypoint detection, extracting features from the single-target image and directly outputting the coordinates of all keypoints using a fully connected layer. Heatmap methods output the keypoint coordinates as images, generating a number of heatmaps equal to the number of keypoint categories. The disadvantages of this method are high computational cost and high GPU memory usage; typically, the heatmap size is one-quarter the size of the input image, leading to an error of at least 3 pixels. Using heatmaps significantly increases the feature map size as the number of keypoints increases. Furthermore, it's necessary to consider the differentiation and matching of keypoints of the same category for different targets. For example, in human pose estimation, an image may contain two people, each with three keypoints: left shoulder, left elbow, and left hand. The challenge lies in how to connect these six points in the image to correctly form the left arms of both individuals. The OpenPose algorithm uses a keypoint affinity field generation method, which further increases the feature map size, leading to increased computation and memory usage. While heatmaps can also perform matching by generating keypoint embeddings, this method cannot solve the problem of increased computation caused by generating heatmap feature maps.
[0008] YoloPose, on the other hand, calculates each keypoint of each target while predicting the bounding box. This means that the keypoint coordinates are obtained and associated with the target object simultaneously, generating keypoints using the target object's global feature pose. Therefore, it doesn't need to consider the association and matching problem between found keypoints and the target, allowing the algorithm to achieve detection speeds almost identical to those of the Yolo series of object detection models. However, a drawback of YoloPose is that for large targets, Yolo uses low-resolution feature maps for regression, leading to significant accuracy loss during fine-grained operations such as keypoint detection for large targets. Summary of the Invention
[0009] The technical problem to be solved by this invention is to provide a 2D pose estimation method based on YOLO, which solves the problem that existing YOLO-based architectures, when locating key points of ultra-large targets, cause large errors due to the fact that they are all based on the offset of the key points regressed from the anchor where the target is located, resulting in the anchor scale being too large and the distance between the target's end key points and the anchor being too far.
[0010] The technical solution adopted by this invention to solve its technical problem is: a 2D pose detection method based on YOLO, comprising the following steps: 1) Training set annotation: Annotate the bounding boxes of the detected objects in the training set images, the coordinates of all key points of the detected objects, the key point categories, and the connection order of each key point; 2) Train the detection model and perform detection; 3) The input during detection consists of two parts: the image to be detected and the connection order of key points. First, the detection model detects the Bbox of the object to be detected, the embedding value of the Bbox, the coordinates of the key points, and the embedding value of the key points. Then, the key point matching and combination part combines the key points of the same object together according to the embedding value, and then determines the position and pose of the object to be detected according to the connection order of the key points.
[0011] Furthermore, in step 1) of the present invention, the number and category of key points of the same detection object are completely the same, and the connection method of each key point is unique. The connection order of the key points is only created once.
[0012] Furthermore, in step 3) of this invention, the key point detection part uses the YoloX backbone; the output is a point-type target, and a high-resolution feature map is used to improve the localization accuracy. A CSP1 and CBA structure is removed from the backbone, so that the size of the feature map output by the original backbone changes from (W / 8, H / 8), (W / 16, H / 16), (W / 32, H / 32) to (W / 4, H / 4), (W / 8, H / 8), (W / 16, H / 16), where W and H are the width and height of the input image.
[0013] Furthermore, in step 3) of this invention, the method for calculating the embedding head loss is as follows:
[0014] in, , , ; This represents the coordinates of the k-th key point of the n-th target; It is the embedding value of the k-th keypoint of the predicted n-th target. It is the reference embedding of the nth target, and it is the average of all keypoint embeddings of the current target. The calculation method is as follows: .
[0015] Furthermore, in step 3) of this invention, when matching key points, the MeanShift algorithm is used to cluster the embeddings of key points, including the following steps: 1. Randomly select n points from the unlabeled data points as the starting center points (center) for n clusters; 2. Find all data points that appear in the region with center as the center and radius as the radius, and consider these points to belong to the same cluster C; at the same time, increment the access frequency of the data points in this cluster by 1; 3. Using center as the center point, calculate the sum of the vectors from center to each data point in set M, obtaining the vector shift. For a given set of n sample points in d-dimensional space... x i, i =1,...,n, For a point x, the basic form of the MeanShift vector is: (3) 4. The center point moves along the direction of the vector shift by a distance of ||shift||. 5. Iteration: Repeat steps 2, 3, and 4 until ||shift|| is very small, i.e., the iteration has converged. Remember the center at this point. All points encountered during this iteration should be classified into cluster C. 6. If the distance between the center of the current cluster C and the center of another existing cluster C2 is less than the threshold when convergence occurs, then C2 and C are merged, and the occurrence counts of data points are also merged accordingly; otherwise, C is taken as the new cluster. 7. Repeat steps 1, 2, 3, 4, 5, and 6 until all points are marked as visited; 8. Based on the access frequency of each point for each class, take the class with the highest access frequency as the class to which the current point set belongs.
[0016] The beneficial effects of this invention are that it solves the defects existing in the background technology, detects key points as individual targets, combines key points into the same target using the matching embedding method, and predicts the embedding of each key point simultaneously; during training, for the same target, the model converges to the shortest distance between the embeddings of each key point, and for different targets, it converges in the direction that increases the embedding distance; it retains the advantages of YOLO-based pose estimation methods, which have faster inference speed and smaller memory usage compared to heatmap methods, while improving the prediction accuracy of key points and adding almost no additional algorithm running time. Attached Figure Description
[0017] Figure 1This is a schematic diagram of the basic architecture of the present invention; Figure 2 This is a schematic diagram of the model structure. Detailed Implementation
[0018] The present invention will now be described in further detail with reference to the accompanying drawings and preferred embodiments. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0019] like Figures 1-2 The YOLO-based 2D pose detection method shown consists of two parts: keypoint detection and keypoint matching and combination. The basic architecture is as follows: Figure 1 As shown, the process includes the following: 1) Training set annotation. Annotate the bounding box of the detected object in the training set image, the coordinates and categories of all key points of the detected object, and the connection order of each key point (Note: embedding does not need to be annotated; since the number of key points and key point categories of the same detected object are exactly the same and the connection method of each key point is unique, the connection order of the key points only needs to be created once when annotating).
[0020] 2) Model training.
[0021] 3) The testing process. For example... Figure 1 As shown, the input of this detection scheme consists of two parts: the image to be detected and the connection order of key points. First, the detection model detects the bounding box (BBox) and its embedding value of the detected object, as well as the coordinates and embedding values of the key points. The key point matching and combination part combines key points of the same detected object together based on the embedding value (note: there may be multiple detected objects in one image), and then determines the position and pose of the detected object based on the connection order of the key points.
[0022] The keypoint detection part uses the YoloX backbone. Since the output is a point-type target, a high-resolution feature map is needed to improve localization accuracy. Therefore, a CSP1 and CBA structure was removed from the backbone, changing the original backbone output feature map size from (W / 8, H / 8), (W / 16, H / 16), (W / 32, H / 32) to (W / 4, H / 4), (W / 8, H / 8), (W / 16, H / 16), where W and H are the width and height of the input image. Compared to the official version of YoloX, the regression detection head branch adds the output of keypoint embeddings, such as... Figure 2 As shown.
[0023] Since the detected targets are all point types, the model will mainly output from the detection head branch with a resolution of (W / 4, H / 4). In order to further reduce the amount of computation, the PAN structure in the backbone is removed, and the detection head branches with resolutions of (W / 8, H / 8) and (W / 16, H / 16) are removed.
[0024] The loss functions for the class head, region head, and object head branches remain unchanged. The loss calculation method for the embedding head is as follows: (1) in , , . This represents the coordinates of the k-th key point of the n-th target. It is the embedding value of the k-th keypoint of the predicted n-th target. It is the reference embedding of the nth target, and it is the average of all keypoint embeddings of the current target. The calculation method is as follows: (2) The purpose of this loss function is to minimize the distance between the embeddings of each keypoint of the same target during model training, and maximize the distance between the reference embeddings of different targets, so as to achieve matching between keypoints of the same target and distinguish between keypoints of the same type of different targets.
[0025] When performing keypoint matching, the MeanShift algorithm is used to cluster the keypoint embeddings. The basic steps of MeanShift are as follows: 1. Randomly select n points from the unlabeled data points as the starting center points for n clusters.
[0026] 2. Identify all data points within a region centered at center and with a radius of radius, assuming these points belong to the same cluster C. Increment the access frequency of each data point within that cluster by 1.
[0027] 3. Using center as the center point, calculate the sum of the vectors from center to each data point in set M, obtaining the vector shift. For a given set of n sample points in d-dimensional space... x i, i =1,...,n, For a point x, the basic form of the MeanShift vector is: (3) 4. The center point moves along the direction of the vector shift by a distance of ||shift||.
[0028] 5. Iteration: Repeat steps 2, 3, and 4 until ||shift|| is very small (i.e., the iteration has converged), and remember the center at this point. All points encountered during this iteration should be classified into cluster C.
[0029] 6. If, upon convergence, the distance between the center of the current cluster C and the center of another existing cluster C2 is less than a threshold, then C2 and C are merged, and the occurrence counts of data points are also merged accordingly. Otherwise, C is treated as a new cluster.
[0030] 7. Repeat steps 1, 2, 3, 4, 5, and 6 until all points are marked as visited.
[0031] 8. Based on the access frequency of each point for each class, take the class with the highest access frequency as the class to which the current point set belongs.
[0032] This invention employs a method similar to heatmap methods, first detecting all keypoints and then matching them using keypoint embeddings. However, unlike heatmap methods, each keypoint is detected as a target using the target model. This avoids the problem of generating heatmap feature maps requiring a large amount of computational resources. Furthermore, since it detects keypoints, higher feature map resolution is more conducive to regressing accurate coordinates. Also, because the output of the YOLO algorithm for detecting small targets mainly comes from high-resolution feature maps, the network structure is simplified by removing the PAN structure, retaining only the FPN part, and cutting off the other two low-resolution detection heads.
[0033] To address the computational overhead and unsuitability for deployment on edge devices of the Heatmap method used for keypoint detection, this object detection method treats keypoints as targets for detection. The overall architecture of the object detection model is based on YOLOX. The YOLO series of models is widely recognized in the industry for its balance of accuracy and speed. YOLOX, for the first time in YOLO models, uses a decoupled head that separates classification and regression detection heads, increasing the model's convergence speed and improving its scalability. The simOTA label allocation strategy enables the model to achieve excellent detection results even when using an anchor-free mechanism, simplifying the output decoding complexity of previous YOLO models using an anchor-based mechanism. By simplifying the model structure and using int8 quantization during deployment, the model can meet the real-time detection requirements on edge devices.
[0034] To address the high complexity of keypoint matching and combination algorithms, this invention employs a keypoint embedding method. Each keypoint is regressed into an embedding value by the model. During training, the embedding values of keypoints within the same detection object converge towards equality; while the embedding values of keypoints from different detection objects converge towards increasing distance. Finally, during model inference, keypoints with similar embedding values can be directly grouped into the same detection object through clustering, simplifying the keypoint matching and combination process.
[0035] Meanwhile, since it is the detection of key points, the higher the resolution of the feature map, the better it is to regress accurate coordinates. Also, since the output of the YOLO algorithm for small target detection mainly comes from high-resolution feature maps, the network structure is simplified by removing the PAN structure, keeping only the FPN part, and cutting off the other two low-resolution detection heads.
[0036] The above description is only a specific embodiment of the present invention. Various examples and illustrations do not constitute a limitation on the substantive content of the present invention. Those skilled in the art can make modifications or variations to the above-described specific embodiments after reading the specification without departing from the substance and scope of the invention.
Claims
1. A 2D pose detection method based on YOLO, characterized in that: Includes the following steps, S1, Training Set Annotation: Annotate the bounding box of the detected object in the training set image, the coordinates and categories of all keypoints of the detected object, and the connection order of each keypoint; the number of keypoints and keypoint categories of the same detected object are completely the same, and the connection method of each keypoint is unique. The connection order of the keypoints is created only once when annotating. S2. Train the detection model and perform detection; S3. The input during detection consists of two parts: the image to be detected and the connection order of key points. First, the detection model detects the Bbox of the object to be detected, the embedding value of the Bbox, the coordinates of the key points, and the embedding value of the key points. Then, the key point matching and combination part combines the key points of the same object together according to the embedding value, and then determines the position and pose of the object to be detected according to the connection order of the key points. The key point detection part uses the YOLOX backbone; the output is a point-type target, and a high-resolution feature map is used to improve the localization accuracy. When matching keypoints, the MeanShift algorithm is used to cluster the keypoint embeddings, including the following steps: 1) Randomly select n points from the unlabeled data points as the starting center points (center) for n clusters; 2) Find all data points that appear in the region with center as the center and radius as the radius, and consider these points to belong to the same cluster C; at the same time, increment the access frequency of the data points in this cluster by 1; 3) Using center as the center point, calculate the sum of the vectors from center to each data point in set M, obtaining the vector shift. For a given set of n sample points in d-dimensional space... x i, i =1,...,n, For a point x, the basic form of the MeanShift vector is: 4) The center point moves along the direction of the vector shift by a distance of ||shift||. 5) Iteration: Repeat steps 2), 3), and 4) until ||shift|| approaches zero, i.e., the iteration converges. Remember the center at this point. All points encountered during this iteration process should be classified into cluster C. 6) If the distance between the center of the current cluster C and the center of another existing cluster C2 is less than the threshold when convergence, then C2 and C are merged, and the occurrence counts of data points are also merged accordingly; otherwise, C is taken as the new cluster. 7) Repeat steps 1), 2), 3), 4), 5), and 6) until all points are marked as visited; 8) Based on the access frequency of each point according to each class, take the class with the highest access frequency as the class to which the current point set belongs.
2. The YOLO-based 2D pose detection method as described in claim 1, characterized in that: In step S3, a CSP1 and CBA structure is removed from the backbone, so that the size of the feature map output by the original backbone changes from (W / 8, H / 8), (W / 16, H / 16), (W / 32, H / 32) to (W / 4, H / 4), (W / 8, H / 8), (W / 16, H / 16), where W and H are the width and height of the input image.
3. The YOLO-based 2D pose detection method as described in claim 2, characterized in that: In step S3, the embedding head loss is calculated as follows: in, , , ; This represents the coordinates of the k-th key point of the n-th target; It is the embedding value of the k-th keypoint of the predicted n-th target. It is the reference embedding of the nth target, and it is the average of all keypoint embeddings of the current target. The calculation method is as follows: .
Citation Information
Patent Citations
Hand key point detection method and device, and computer equipment
CN111523387A
Hand key point detection method and device and terminal equipment
CN112336342A