A Human Neck Keypoint Recognition Method and System Based on KPNet
By improving the CenterNet network, it is used for the detection of key points on the human neck, and the problem of key points detection in the medical field is solved, achieving efficient and accurate key points recognition.
Patent Information
- Application Number
- CN202210970994.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-14
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-08-14
AI Technical Summary
The prior art is difficult to effectively apply to key point detection tasks in the medical field, especially in the identification of human necks.
Improved CenterNet network, including backbone network, upsampling module and prediction head, is used to identify key points in the human neck and realize key point detection through thermal map prediction.
It realizes accurate identification of key points on the human neck, simplifies the network structure, reduces post-processing steps, and improves detection efficiency and accuracy.
Smart Images

Figure CN115346241B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology. More specifically, it particularly relates to a method and system for identifying human neck key points based on KPNet. Background Art
[0002] Object detection refers to finding the target position and classifying it in an image or video. In recent years, it has attracted attention due to its wide application. However, making a machine learn object detection is a difficult task. Object detection requires identifying and locating all instances of a certain target in the field of view, which is a fundamental problem in the field of computer vision. At the same time, the research on medical image processing is also an important branch in the field of object detection in recent years.
[0003] Currently, the existing detectors are mainly divided into two categories: two-stage detectors and one-stage detectors. If a network has a separate module for generating region proposals, then the network is called a two-stage detector. This model takes a long time in the first stage to find a certain number of target proposals, and has a complex structure and lacks global information. While one-stage detectors directly classify and locate the target through dense sampling. They use predefined boxes / points with different scales and aspect ratios to locate the target, and they exceed two-stage detectors in terms of real-time performance and simpler design.
[0004] CenterNet, as an emerging one-stage detector, adopts a different method, modeling the object as a point instead of the traditional bounding box representation. CenterNet predicts the object as a single point at the center of the bounding box. The input image generates a heatmap through FCN, and the peak of the heatmap corresponds to the center of the detected object. It uses the ImageNet pre-trained Hourglass-101 as the feature extraction network, which has 3 heads: the heatmap head for the center point of the point target, the width and height head of the target size, and the offset head of the target center point. During training, the multi-task loss of the three heads is backpropagated into the feature extractor. During the inference process, the output of the offset head is used to determine the object point, and finally a bounding box is generated. Since the prediction is a point rather than a result, there is no need to use non-maximum suppression (NMS) for post-processing here. Traditional object detection usually identifies the target as a bounding box parallel to the coordinate axes. Most object detectors first list all possible target position bounding boxes and then classify them one by one. This method is extremely wasteful and inefficient and requires additional post-processing.
[0005] Compared with other one-stage detectors, CenterNet is similar to ConearNet and FCOS, both of which are keypoint-based Anchor-free object detection frameworks. The difference is that it sets the keypoint as the center point of the object, and other attributes are obtained through regression. And since there is only one point, there is no need to group the detected keypoints. Due to this characteristic, its framework is extremely simple and there is no NMS post-processing (only removing duplicates through the topK of the heatmap peak, which is similar to the NMS function but faster). The performance comparison when training with the COCO dataset is as Figure 1 shown.
[0006] It can be seen that using the CenterNet network from a novel perspective is more accurate and has a shorter inference time compared to previous first-order detectors. It has high precision and is often used in various tasks such as 3D object detection, keypoint estimation, pose, instance segmentation, and orientation detection. It has been applied to many application scenarios in the past, such as gesture detection, face recognition, and human pose detection, etc., and has been quite mature when integrated into various enterprises.
[0007] However, the traditional CenterNet network is mainly applied to object detection, identifying each object in the image. In order to apply it to the keypoint detection task in the medical field, it is necessary to develop a method and system for identifying human neck keypoints based on KPNet. Summary of the Invention
[0008] The purpose of the present invention is to provide a method and system for identifying human neck keypoints based on KPNet to overcome the defects existing in the prior art.
[0009] To achieve the above purpose, the technical solution adopted by the present invention is as follows:
[0010] A method for identifying human neck keypoints based on KPNet includes the following steps:
[0011] S1. Collect a number of images including human neck keypoints;
[0012] S2. Construct a training dataset and a validation dataset including human neck keypoints;
[0013] S3. Improve the CenterNet network. The improved CenterNet network includes a backbone network, an upsampling module, and a prediction head. The backbone network is a classification model with the fully connected layer removed. The backbone network is used to reduce the resolution of the input human neck image by 32 times and obtain a first feature map. The upsampling module is a multi-layer transposed convolution. The upsampling module is used to restore the resolution of the first feature map to 1 / 4 times that of the human neck image and output a second feature map. The output part of the prediction head only retains the head for predicting the heat map. The prediction head includes a 3*3 convolution, a 1*1 convolution, and a sigmoid activation function;
[0014] S4. Train the improved CenterNet network using the training dataset in step S2 to obtain corresponding weights and a patrol function, and verify based on the validation dataset. If the loss function of the validation set decreases steadily and there is no overfitting or underfitting phenomenon during the network training process, it can be used for testing on the test set. Otherwise, modify and expand the dataset and continue to retrain until the verification is qualified;
[0015] S5. Input the actually acquired human neck image obtained by camera shooting into the network for key point detection;
[0016] S6. Take the pixel values of each predicted feature map as a two-dimensional normal distribution, take the peak value of the two-dimensional normal distribution as the predicted key point, and combine the predicted key point with the human neck point cloud information captured by the camera to obtain the coordinate information of the key point.
[0017] Further, the human neck key points include a bottom center point, a bottom left point, a bottom right point, a top left point, and a top right point.
[0018] Further, the output channels of the prediction result are 5.
[0019] Further, in step S4, Focal Loss and L1 Loss are used to calculate the heat map loss and the feature point offset loss respectively.
[0020] The present invention also provides a system according to the above-mentioned method for identifying human neck key points based on KPNet, including:
[0021] An acquisition module, used to acquire a plurality of images including human neck key points;
[0022] A construction module, used to construct a training dataset and a validation dataset including human neck key points;
[0023] Improvement module, used to improve the CenterNet network. The improved CenterNet network includes a backbone network, an upsampling module, and a prediction head. The backbone network is a classification model with the fully connected layer removed. The backbone network is used to reduce the resolution of the input human neck image by 32 times and obtain a first feature map. The upsampling module is a multi-layer transposed convolution. The upsampling module is used to restore the resolution of the first feature map to 1 / 4 times that of the human neck image and output a second feature map. The output part of the prediction head only retains the head for predicting the heat map. The prediction head includes a 3*3 convolution, a 1*1 convolution, and a sigmoid activation function;
[0024] Training module, used to train the improved CenterNet network with the training dataset in step S2 to obtain corresponding weights and a patrol function, and verify based on the validation dataset. If the loss function of the validation set decreases steadily and there is no overfitting or underfitting phenomenon during the network training process, it can be used for testing on the test set. Otherwise, modify and expand the dataset and continue to retrain until the verification is qualified;
[0025] Input module, used to input several images in the acquisition module into the KPNet network to output a prediction feature map;
[0026] Calculation module, used to make the pixel values of each prediction feature map follow a two-dimensional normal distribution, take the peak value of the two-dimensional normal distribution as the predicted key point, and restore the predicted key point to the pixel coordinates in the human neck image to obtain the coordinate information of the key point.
[0027] Compared with the prior art, the advantages of the present invention are as follows: A method and system for identifying human neck key points based on KPNet provided by the present invention follow the characteristics of the CenterNet network model and improve it. Finally, the KPNet network is obtained. Specifically, compared with the traditional method of recognizing each target in an image and displaying its category and confidence level in text, the present invention outputs the required key points by improving the network heat map processing. Correspondingly, its dataset label also changes from an anchor box to a key point. The present invention uses the improved CenterNet network for key point recognition for the first time, which is convenient for key point recognition in the medical field and even other fields. Description of the Drawings
[0028] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0029] Figure 1It is a comparison chart of the real-time speed-performance balance of various detectors using the COCO dataset in the prior art.
[0030] Figure 2 It is a flowchart of the method for identifying human neck key points based on KPNet in the present invention.
[0031] Figure 3 It is the network structure diagram of KPNet in the present invention.
[0032] Figure 4 It is the diagram of the key point detection result in the present invention.
[0033] Figure 5 It is the heat map output in the present invention.
[0034] Figure 6 It is the schematic diagram of the system for identifying human neck key points based on KPNet in the present invention. Specific Embodiments
[0035] The preferred embodiments of the present invention will be described in detail below with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby making the protection scope of the present invention more clearly defined.
[0036] Refer to Figure 2 As shown, this embodiment discloses a method for identifying human neck key points based on KPNet, including the following steps:
[0037] Step S1: Collect a number of images including human neck key points.
[0038] Specifically, the human neck key points include the bottom center point, the bottom left point, the bottom right point, the top left point, and the top right point.
[0039] Step S2: Construct a training dataset and a validation dataset including human neck key points, and the ratio of the training dataset to the validation dataset can be set according to actual needs.
[0040] Step S3: Improve the CenterNet network. The improved CenterNet network includes a backbone network, an upsampling module, and a prediction head. The backbone network is a classification model with the fully connected layer removed. The backbone network is used to reduce the resolution of the input human neck picture by 32 times and obtain a first feature map. The upsampling module is a multi-layer transposed convolution. The upsampling module is used to restore the resolution of the first feature map to 1 / 4 times that of the human neck picture and output a second feature map. Only the head for predicting the heat map is retained in the output part of the prediction head. The prediction head includes a 3*3 convolution, a 1*1 convolution, and a sigmoid activation function, and the output channel of the prediction result is 5.
[0041] Since the CenterNet network is an Anchor free object detection algorithm, its inference results include three attributes: the center point, width and height of the prediction box, and the center point offset. This embodiment draws on the idea of the CenterNet network to regress the center point, removes the width, height, and center point offset parts, and applies the improved network to the key point detection task.
[0042] As Figure 3 shown, for an image as input, assuming its width and height dimensions are N*N (the number of channels is ignored here), after passing through a backbone network (ResNet50 or Hourglass), a feature map with a resolution reduced by 32 times is obtained; then an upsampling module is needed to restore the resolution to 1 / 4 of the original image, that is, 8 times upsampling is performed; finally, a prediction head for regressing the center point heat map is connected.
[0043] Step S4: Train the improved CenterNet network using the training dataset in Step S2 to obtain the corresponding weights and inspection functions, and verify based on the validation dataset. If the loss function of the validation set drops steadily and there is no overfitting or underfitting phenomenon during the network training process, it can be used for testing on the test set; otherwise, modify and expand the dataset and continue to retrain until the verification is qualified.
[0044] Step S5: Input the actually obtained human neck image captured by the camera into the network for key point detection.
[0045] Step S6: Take the pixel values of each predicted feature map as a two-dimensional normal distribution, take the peak value of the two-dimensional normal distribution as the predicted key point, and combine the predicted key point with the point cloud information of the human neck captured by the camera to obtain the coordinate information of the key point, which can be stored and recorded in the form of a file.
[0046] Since traditional object detection based on the CenterNet model uses the preliminary features extracted by the backbone network to obtain a high-resolution feature map of size 128×128 through three transposed convolutions, then inputs it into three prediction heads to obtain three prediction attributes (heat map, center point offset, and width and height), and finally uses these prediction results to obtain the position of the prediction box in the original image, thereby completing the object recognition task. And this embodiment is for the key point detection task, so only the head of the predicted heat map is left in the output part. Since the positions of 5 key points need to be regressed (as Figure 4 shown), the number of channels of the heat map is set to 5. At this time, after inputting an RGB image into the network, it will first pass through the backbone network for feature extraction, then pass through three transposed convolutions to improve the resolution of the feature map, and finally pass through the predicted heat map head to obtain 5 heat maps (as Figure 5as shown, corresponding to 5 different key points respectively.
[0047] Refer to Figure 6 As shown, the present invention also provides a system according to the above-mentioned method for identifying human neck key points based on KPNet, including: an acquisition module 1 for acquiring a plurality of images including human neck key points; a construction module 2 for constructing a training data set and a validation data set including human neck key points; an improvement module 3 for improving the CenterNet network. The improved CenterNet network includes a backbone network, an upsampling module, and a prediction head. The backbone network is a classification model with the fully connected layer removed. The backbone network is used to reduce the resolution of the input human neck picture by 32 times and obtain a first feature map. The upsampling module is a multi-layer transposed convolution. The upsampling module is used to restore the resolution of the first feature map to 1 / 4 times that of the human neck picture and output a second feature map. The output part of the prediction head only retains the head for predicting the heat map. The prediction head includes a 3*3 convolution, a 1*1 convolution, and a sigmoid activation function; a training module 4 for training the improved CenterNet network with the training data set in step S2 to obtain corresponding weights and a patrol function, and validating based on the validation data set. If the loss function of the validation set decreases steadily and there is no overfitting or underfitting phenomenon during the network training process, it can be used for testing the test set. Otherwise, modify and expand the data set and continue to retrain until the validation is qualified; an input module 5 for inputting a plurality of images in the acquisition module into the KPNet network to output a prediction feature map; a calculation module 6 for taking the pixel values of each prediction feature map as a two-dimensional normal distribution, taking the peak value of the two-dimensional normal distribution as the predicted key point, and restoring the predicted key point to the pixel coordinates in the human neck picture to obtain the coordinate information of the key point.
[0048] The present invention uses the feature information extracted by the network to predict the heat map, thereby realizing key point detection. The inference process of the network is an end-to-end process, that is, input an RGB image of a human neck and output the pixel coordinates of the neck key points, which can be used as an important reference position for subsequent medical detection.
[0049] Although the embodiments of the present invention have been described in conjunction with the accompanying drawings, the patent owner can make various deformations or modifications within the scope of the appended claims. As long as it does not exceed the protection scope described in the claims of the present invention, it should be within the protection scope of the present invention.
Claims
1. A method for identifying human neck key points based on KPNet, characterized in that It includes the following steps: S1. Collect a number of images including the key points of the human neck; S2. Construct a training data set and a validation data set including the key points of the human neck; S3. Improve the CenterNet network. The improved CenterNet network includes a backbone network, an upsampling module and a prediction head. The backbone network is a classification model with the fully connected layer removed. The backbone network is used to reduce the resolution of the input human neck picture by 32 times and obtain a first feature map. The upsampling module is a multi-layer transposed convolution. The upsampling module is used to restore the resolution of the first feature map to 1 / 4 times that of the human neck picture and output a second feature map. The output part of the prediction head only retains the head for predicting the heat map. The prediction head includes a 3*3 convolution, a 1*1 convolution and a sigmoid activation function; S4. Train the improved CenterNet network using the training data set in step S2 to obtain the corresponding weights and inspection functions, and verify based on the validation data set. If the loss function of the validation set decreases steadily and there is no overfitting or underfitting phenomenon during the network training process, it can be used for testing on the test set. Otherwise, modify and expand the data set and continue to retrain until the verification is qualified; S5. Input the actually obtained human neck image captured by the camera into the network for key point detection; S6. Take the pixel values of each predicted feature map as a two-dimensional normal distribution, take the peak value of the two-dimensional normal distribution as the predicted key point, and combine the predicted key point with the point cloud information of the human neck captured by the camera to obtain the coordinate information of the key point; The key points of the human neck include the bottom center point, the bottom left point, the bottom right point, the top left point and the top right point; The output channel of the prediction result is 5; In step S4, the Focal Loss and the L1 Loss are used to calculate the heat map loss and the feature point offset loss respectively.
2. A system for the method of identifying human neck key points based on KPNet according to claim 1, characterized in that, It includes: A collection module for collecting a number of images including the key points of the human neck; A construction module for constructing a training data set and a validation data set including the key points of the human neck; An improvement module for improving the CenterNet network. The improved CenterNet network includes a backbone network, an upsampling module and a prediction head. The backbone network is a classification model with the fully connected layer removed. The backbone network is used to reduce the resolution of the input human neck picture by 32 times and obtain a first feature map. The upsampling module is a multi-layer transposed convolution. The upsampling module is used to restore the resolution of the first feature map to 1 / 4 times that of the human neck picture and output a second feature map. The output part of the prediction head only retains the head for predicting the heat map. The prediction head includes a 3*3 convolution, a 1*1 convolution and a sigmoid activation function; A training module, which is used to train the improved CenterNet network with the training dataset in step S2 to obtain the corresponding weights and inspection functions, and verify based on the validation dataset. If the loss function of the validation set decreases steadily and there is no overfitting or underfitting phenomenon during the network training process, it can be used for testing on the test set; otherwise, modify and expand the dataset and continue to retrain until the verification is qualified. An input module, which inputs the actually obtained human neck image captured by the camera into the network for key point detection. A calculation module, which is used to make the pixel values of each predicted feature map follow a two-dimensional normal distribution, take the peak value of the two-dimensional normal distribution as the predicted key point, and combine the predicted key point with the human neck point cloud information captured by the camera to obtain the coordinate information of the key point.
Citation Information
Patent Citations
Safety helmet wearing convolutional network based on feature fusion, training method and detection method
CN112070043A
Object pose obtaining method, and electronic device
US20210304438A1