Multi-class key point detection method based on fusion of high-resolution features
By fusing high-resolution and secondary resolution feature maps in the deep learning key point detection model, and combining longest alignment technology and image gradient correction, the key point detection accuracy and multi-category detection problems in the existing technology are solved, and a more efficient and flexible key point detection method is achieved.
Patent Information
- Application Number
- CN202510685332.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-27
- Publication Date
- 2025-06-24
- Estimated Expiration
- 2045-05-27
AI Technical Summary
The existing key point detection method is difficult to extract useful features during the convolution process, resulting in errors in position information and affecting detection accuracy; at the same time, the existing method is difficult to realize multi-category key point detection, resulting in system bloat and high computing costs.
By building an improved deep learning key point detection model, integrating high-resolution feature maps and secondary resolution feature maps, combining the longest alignment technology, multi-category key point detection is realized, and image gradient correction is added in the prediction stage.
It improves the accuracy and stability of key point detection, and can realize multi-category key point detection without increasing computing overhead, reducing the complexity and computing cost of the system.
Smart Images

Figure CN120198685A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision technology, and particularly to the field of target key point detection technology. Specifically, it is a multi-class key point detection method based on fusing high-resolution features. Background Art
[0002] In recent years, key point detection has become an important research direction in the fields of image processing and computer vision. The goal of this technology is to determine the positions of key points of the target to be measured, including human pose estimation, facial key point estimation, aircraft key point estimation, etc. The accuracy of this technology directly affects subsequent data-driven decision-making content, such as quantifying human movement (rehabilitation assessment), object movement state (robotic arm grasping target), aircraft vulnerable part strike, etc. With the rapid development of deep learning, many methods for target key point detection based on neural networks have achieved very excellent results. These methods have promoted the development of key point technology from their respective perspectives. However, the following problems remain unsolved: First, in the existing methods, the original image is continuously downsampled through a convolutional neural network. In this process, although the network model can learn high-level features of the image, as the number of convolutional layers increases, the network may be difficult to extract useful features. For the key point detection task, since this task is extremely sensitive to position information, if incorrect position information is extracted in the network, it will seriously affect the accuracy of the final prediction result; In addition, the current key point detections are all based on single-class. In actual application requirements, there are various different types of targets. For example, in the aircraft key point detection task, it may include fixed-wing aircraft, rotary-wing aircraft, etc. If data needs to be re-annotated and the network needs to be trained for each type of aircraft, it will not only consume a lot of time, but also make the system bloated.
[0003] It can be seen that in the key point detection method, there are problems that the accuracy of key point detection is affected by incorrect position information, and it takes a long time to achieve multi-class key point detection and the system is bloated. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-class key point detection method based on fusing high-resolution features, which is used to solve the problems in the prior art that the accuracy of key point detection is affected by incorrect position information, and it takes a long time to achieve multi-class key point detection and the system is bloated.
[0005] The present invention solves the above problems through the following technical solutions: A multi-class key point detection method based on fusing high-resolution features, including: Step S100: Collect the image data of the target to be detected to obtain an image data sample; Step S200: Label the bounding box and key point coordinates of the target to be detected in the image data sample to obtain bounding box data and key point data; Step S300: Determine whether the key point data needs to be aligned by the longest length. If so, perform the alignment by the longest length and then convert the key point data after the alignment by the longest length into a Gaussian distribution map. Otherwise, directly convert the key point data into a Gaussian distribution map; Step S400: Train a target detector using the bounding box data. This target detector is used to predict the position information of the target to be detected for key point detection in the image to be detected; Step S500: Use the Gaussian distribution map as a label and input it into the improved deep learning key point detection model for training, and obtain the trained key point detection model; Step S600: Input the image to be detected into the trained target detector to obtain all the target position information in the image to be detected. Cut out the position of the target in the image to be detected to obtain a sub-image. If the size of the sub-image is not equal to the preset size, stretch or compress the sub-image to the preset size, and then input it into the trained key point detection model to obtain the key point coordinate information; Step S700: Based on the key point coordinate information predicted by the model, calculate the image gradient of the key point and its adjacent pixels, and regard the point with the largest gradient change as the final key point.
[0006] By constructing an improved deep learning key point detection model, the present invention can continuously fuse the high-resolution feature map from the original size and the secondary resolution feature map from the upper-layer convolution during the process of convolutional feature extraction, so as to more accurately capture the effective position information and predict the more accurate key point positions. Through the alignment by the longest length, a network model can be extended from single-category key point detection to multi-category key point detection without adding too much computational overhead.
[0007] Further, the method for determining whether the key point data needs to be aligned by the longest length in the step S300 is as follows: when the number of key points of a certain category in the target to be detected is equal to the maximum number of key points in all categories, there is no need to perform the alignment by the longest length, otherwise, the alignment by the longest length is required.
[0008] For example, when detecting fixed-wing aircraft and quadrotor aircraft, there are two categories: fixed-wing aircraft (the key point is that the number is 5) and quadrotor aircraft (the key point number is 7). The maximum number of key points in the two categories is the key point number of the quadrotor. In practical applications, the target to be detected is not necessarily an aircraft, and it may also be other categories, such as the key points of the human skeleton. In addition, the maximum number of key points is determined according to the actual application, based on the target to be detected. For example, when detecting three categories, A, B, and C, the corresponding key point numbers are 6, 9, and 9 respectively, then the maximum number of key points is 9.
[0009] Furthermore, the longest alignment method in step S300 is as follows: Step S310, if labeling key points of the same category, the labeling order of key points for each target needs to be consistent; Step S320, if labeling key points of different categories and the number of key points between categories is inconsistent, then after the target with fewer key points is labeled in order, a padding operation with 0s needs to be performed at the end until the number is filled to the maximum number of key points among all categories, thereby completing the longest alignment.
[0010] Furthermore, the method for converting key point data into a Gaussian distribution map is as follows: Step A1, convert the coordinates of the key points into a scale factor relative to the width and height of the image; Step A2, pre-fix the width and height of the target to be measured. If the obtained image is inconsistent with the preset width and height, then stretch or compress the image to the preset value; Step A3, generate an image with 1 channel, size of the preset value, and all pixel values being 0; Step A4, generate a two-dimensional Gaussian distribution map at the corresponding position of the image generated in step A3 through the scale factor, and the corresponding function is: ; In the formula, represents the value of the generated two-dimensional Gaussian distribution map at the coordinate, represents the amplitude, and represent the center of the two-dimensional Gaussian distribution map, that is, the position coordinates of the key point in the image, and respectively represent the manually preset variances on the x-axis and y-axis in the Gaussian function; If a certain key point is obtained through the longest alignment, then the amplitude of the Gaussian function corresponding to this key point is 0.
[0011] Furthermore, the improved deep learning keypoint detection model includes two parts: an encoder and a decoder. The encoder is responsible for extracting semantic features from images, and the decoder is responsible for translating the extracted features, where: The encoder has four layers, each layer consisting of a first module. The first module is composed of a resolution-preserving component, a downsampling component, and a first feature fusion component. The resolution-preserving component is composed of a convolutional layer, a BN layer, and an activation function layer. After the data passes through the resolution-preserving component, the size of the feature map remains unchanged, thus obtaining a high-resolution feature map. The downsampling component is composed of a convolutional layer, a BN layer, an activation function layer, and a pooling layer. After the data passes through the downsampling component, the size of the feature map decreases, but more advanced semantic features are obtained. The first feature fusion component is composed of a convolutional layer, a BN layer, an activation function layer, a pooling layer, and a channel concatenation layer. The feature fusion component fuses features of different sizes together to ensure the stability of feature extraction during the encoding process. In the first module of the first layer of the encoder, the first feature fusion component only needs to fuse the output of the resolution-preserving component and the output features of the downsampling component in the first module of this layer. In the first modules of the subsequent three layers of the encoder, in addition to fusing the output of the resolution-preserving component and the output features of the downsampling component in the first module of the corresponding layer, the first feature fusion component also fuses the output of the resolution-preserving component in the first module of the first layer. The decoder has four layers, each layer consisting of a second module. Each second module is divided into two parts: an upsampling component and a second feature fusion component. The upsampling component consists of a transposed convolutional layer. After passing through the transposed convolutional layer, the size of the output feature map will be enlarged to the same size as the feature map output by the resolution-preserving component corresponding to the same layer in the encoder. The second feature fusion component includes a feature channel concatenation layer, a convolutional layer, a BN layer, and an activation function. The second feature fusion layer will fuse the feature map output by the resolution-preserving component in the first module of the same layer in the encoder and the feature map after the upsampling component in the second module of the same layer in the decoder to ensure that the decoder decodes in the correct direction.
[0012] Furthermore, the loss function of the improved deep learning keypoint detection model is defined as: ; where m represents the number of maximum keypoints, n represents the number of pixels in each output feature map, represents the th pixel value of the Gaussian distribution map corresponding to the real th keypoint, represents the th pixel value of the Gaussian distribution map corresponding to the predicted th keypoint.
[0013] Further, the specific method for calculating the image gradients of the key points and their adjacent pixels in step S700 is as follows: Gradient value calculation template in the x direction: ; Gradient value calculation template in the y direction: ; Gradient value of the corresponding pixel point: ; Among them, represents the x-direction template, I represents the corresponding image region, represents the y-direction template; and are the absolute values of the gradient values in the x direction and y direction; Finally, the pixel point with the maximum gradient is regarded as the position of the final key point.
[0014] Compared with the prior art, the present invention has the following advantages and beneficial effects: (1) By constructing a network model, the present invention can continuously fuse the high-resolution feature maps from the original size and the secondary-resolution feature maps from the upper-layer convolution during convolution feature extraction, so as to more accurately capture the effective position information, and thus can predict the more accurate key point positions in practical applications, and to a certain extent, suppress the influence of background interference and object pose transformation.
[0015] (2) The present invention performs the longest key point alignment during the model training stage, enabling the network model to have the ability to detect key points for multiple categories, so that it is not necessary to train a separate model for each category to adapt to the multi-object key point detection in practical engineering, increasing the flexibility in engineering applications.
[0016] (3) The longest key point alignment designed by the present invention, since the pixel values of the feature maps finally generated by the aligned key points are all 0, does not consume too many resources during both the model training stage and the prediction stage, reducing the computational cost.
[0017] (4) The present invention adds image gradient correction based on the prediction results of the model. By calculating the gradients of the predicted pixel points and the adjacent position pixel points, the final prediction results are updated to make the prediction results more accurate. Description of the Drawings
[0018] Figure 1 is the network structure diagram of the improved deep learning key point detection model in the present invention; Figure 2The detection effect diagrams of five key points of a fixed-wing aircraft, where (a) is the detection result under a normal and good viewing angle; (b) is the detection result under a partially occluded viewing angle; (c) is the detection result under a deformed viewing angle caused by the aircraft's roll. Figure 3 The detection effect diagrams of seven key points of a quadrotor aircraft, where (a) is the detection result under a blurred viewing angle; (b) is the detection result under an occluded viewing angle; (c) is the detection result under a good viewing angle. Specific implementation manners
[0019] The present invention will be further described in detail below in conjunction with embodiments, but the implementation manners of the present invention are not limited thereto.
[0020] Embodiment 1:
[0021] Combined with the attached Figure 1 As shown, a multi-class key point detection method based on fusing high-resolution features includes: Step 1, according to actual requirements, a certain number of target images to be detected are collected by a detector. For example, in this embodiment, one thousand images with a width and height of 1024 pixels and a channel number of 1 are collected for each of the fixed-wing aircraft and the quadrotor aircraft for subsequent model training and testing.
[0022] Step 2, the collected images are labeled using the labelme annotation tool. Since the longest alignment is required later, the number of the longest key points needs to be determined before annotation.
[0023] Specifically, in this embodiment, the number of the longest key points is 7. The specific annotation format of a target is as follows: the horizontal and vertical coordinates of the upper left corner of the target bounding box, the horizontal and vertical coordinates of the lower right corner of the target bounding box, the coordinates of key point 1, the coordinates of key point 2, the coordinates of key point 3, the coordinates of key point 4, the coordinates of key point 5, the coordinates of key point 6, the coordinates of key point 7. It should be noted that: First, the annotation order of the key points of the same type of target must be consistent, and the specific order can be customized; Second, if the number of target key points is insufficient, it needs to be padded with 0 for alignment. In this embodiment, the annotation result of the quadrotor aircraft with seven key points is as follows: the horizontal and vertical coordinates of the upper left corner of the aircraft bounding box, the horizontal and vertical coordinates of the lower right corner of the aircraft bounding box, the coordinates of the nose, the coordinates of the fuselage, the coordinates of the tail wing, the coordinates of rotor 1, the coordinates of rotor 2, the coordinates of rotor 3, the coordinates of rotor 4; the annotation result of the fixed-wing aircraft with five key points after the longest alignment is as follows: the horizontal and vertical coordinates of the upper left corner of the aircraft bounding box, the horizontal and vertical coordinates of the lower right corner of the aircraft bounding box, the coordinates of the nose, the coordinates of the fuselage, the coordinates of the tail wing, the coordinates of the left wing, the coordinates of the right wing, (0, 0), (0, 0).
[0024] Step 3: Divide the labeled data into a training set and a test set according to a ratio. Specifically, in this embodiment, the training set and the test set are divided according to a ratio of 7:3. That is, there are 700 images of fixed-wing aircraft and quadrotor aircraft each for training, and 300 images are set aside for testing respectively.
[0025] Step 4: First, use the labeled data to train an object detection model. Since keypoint detection only requires the position of the object in the image and does not need to know the specific category, in the object detection model, the number of categories is set to 1, and the category name can be customized. Specifically, in this embodiment, yolov5 is used as the object detection model to detect the position of the aircraft in the image, and the model outputs the position information as the input information for the next keypoint detection.
[0026] Step 5: In the training stage, preprocess the data, and then use the improved deep learning keypoint detection model to train the preprocessed data. Specifically, in this embodiment, the data preprocessing method is as follows: First, calculate the size of the object according to the coordinates of the bounding box. The calculation method is as follows: , ; where w represents the width of the object, represents the abscissa of the lower right corner of the bounding box, represents the abscissa of the upper left corner of the bounding box; h represents the height of the object, represents the ordinate of the lower right corner of the bounding box, represents the ordinate of the upper left corner of the bounding box.
[0027] Secondly, obtain the width, height and keypoint position information of the object to be detected according to the original annotation information, and based on this, obtain the scale factor of each keypoint relative to the width and height of the object. The calculation method is as follows: , ; where, represents the scale factor of the width, represents the distance from the abscissa of the keypoint to the left side of the bounding box, and w represents the width of the bounding box; represents the scale factor of the height, represents the distance from the ordinate of the keypoint to the upper side of the bounding box, and h represents the height of the bounding box; According to the above calculation, if the width or height of the object is not equal to 255, then the object will be transformed to the original image with a width and height of 255 through stretching or shrinking methods.
[0028] Then, according to the size of the transformed image and the width and height scale factors, calculate the position of the transformed coordinate points. The calculation method is as follows: , ; wherein represents the abscissa of the key point after transformation, and represents the ordinate of the key point after transformation; is the width scale factor obtained above,
[0029] After obtaining the coordinates of the key points after transformation, n heatmaps with Gaussian distribution can be generated, where n is the maximum number of key points defined above. In this embodiment, n is 7. Specifically, first generate an image with a width and height of 255, all values being 0, and a channel number of 1, and then, according to the coordinates after transformation obtained above, set the corresponding positions of the image to 255, and then perform Gaussian filtering with a Gaussian convolution kernel to finally generate a heatmap with Gaussian distribution corresponding to this key point. In this embodiment, the size of the Gaussian kernel is 7. It should be noted that the values of the label map generated from the key points obtained by the longest alignment should all be 0. Aiming at the problem that the key point detection task is sensitive to position information, the present invention makes corresponding improvements on the basis of a general convolutional neural network, integrates high-resolution feature maps, and enables it to more effectively encode and decode position information. The specific improvements are as follows: The model is divided into an encoding part and a decoding part. In the encoding part, there are a total of four layers, and each layer contains a resolution-preserving component, a downsampling component, and a feature fusion component. Among them, the resolution-preserving component consists of a convolutional layer with a kernel size of 1, a BN layer, and an activation function layer. When the data passes through this component, the size of the feature map will not change, so that high-resolution feature maps can be retained. The downsampling component consists of a convolutional layer with a kernel size of 3 and a padding of 1, a BN layer, an activation function layer, and a max pooling layer with a size of 2. When the data passes through this component, the size of the feature map will be halved, so that more advanced semantic features can be obtained. The feature fusion component consists of a convolutional layer with a kernel size of 3 and a padding of 1, a BN layer, an activation function layer, a max pooling layer with an adaptive size, and a data concatenation layer. The data passing through this component can change the size of the feature map to the same size as the feature map to be fused, and then the two data are concatenated by the data concatenation layer according to the number of channels. In the encoder, each module first receives the feature map output from the previous module as the input of this module, and the input of the first module is the original image. Then each module first obtains a feature map with the same resolution as the input of this module through the resolution-preserving component, and then based on the input of the module, obtains a map with halved resolution but containing advanced features through the downsampling component. In the first module, the feature map obtained after downsampling only needs to fuse the feature map output by the resolution-preserving component in this module through the feature fusion component, while in the second, third, and fourth modules, the feature map obtained after downsampling needs to fuse two sets of data, namely the feature map output by the resolution-preserving component in this module and the feature map output by the resolution-preserving component of the first module. This continuous fusion of high-resolution features from the top layer and sub-high-resolution features from the adjacent upper layer enables the network to capture position information features more accurately during the feature extraction process, thereby improving the accuracy and stability of network prediction.
[0030] In the decoding part, it is divided into four layers, and each layer contains an upsampling component and a feature concatenation component. Among them, the upsampling component consists of a transposed convolution with a kernel size of 4, a stride of 2, and a padding of 1. After the data passes through the upsampling component, the size of the feature map will be expanded to twice the original size. The role of the feature concatenation component is to concatenate the feature map of the corresponding same size in the encoder with the feature map after passing through the upsampling component, which ensures that the network can be more stable and accurate during the decoding process.
[0031] After encoding and decoding, a feature map with the same size as the original input image is obtained. Finally, through a convolution with a kernel size of 1, the final feature map is converted into n channels, where n represents the maximum number of key points to be detected, that is, one feature map is responsible for predicting one key point. In this embodiment, the target has a maximum of 7 key points, so n is also 7. Finally, the loss between the predicted value output by the network and the labeled ground truth value is calculated using the following loss function: ; where m represents the number of key points, represents the -th pixel prediction value of the Gaussian distribution map corresponding to the real -th key point, represents the -th pixel value of the Gaussian distribution map corresponding to the predicted -th key point. The meaning of this loss function is to match an output feature map with a key point, calculate the difference between each pixel value of the feature map and the Gaussian distribution heat map corresponding to this key point, and finally take the average as the final loss for subsequent parameter update of the optimizer.
[0032] Step 6, in the prediction stage, first, the position information of the target to be detected in the image is detected by the target detector described above. Then, according to the image preprocessing method described above, the target is input into the key point detection network, and a Gaussian distribution heat map for each key point is predicted. The non-zero values in the heat map indicate that the key points predicted by the network are near this area, and the point with the largest value is the predicted point by the network. Subsequently, the scale factors of the predicted coordinate points are calculated: , ; where represents the scale factor in the x direction, represents the scale factor in the y direction, represents the predicted x coordinate, represents the predicted y coordinate, and 255 is the size of the image output to the key point detection network. According to the above scale factors, the positions of the key points in the original image are calculated: , ; where represents the abscissa of the key point in the original image, represents the ordinate of the key point in the original image, corresponds to the width of the original image, corresponds to the height of the original image.
[0033] Step 7, after obtaining the key point coordinates predicted by the model, calculate the image gradients at this coordinate and adjacent pixel points. In this embodiment, a total of 9 gradient values at the predicted coordinate and adjacent 8 pixel points are calculated. The gradient calculation template is as follows: , ; where represents the x-direction template, I represents the corresponding image region, represents the y-direction template, and the calculated gradients are as follows: ; Compare the gradient values of the nine points, and output the point with the largest gradient value as the final key point prediction point. As shown in Figure 2 、 Figure 3 , the key point positions of the fixed-wing aircraft and the quadrotor aircraft are predicted respectively according to the above steps and visualized in the image. Among them Figure 2 in (a), (b), and (c) are the detection results of the fixed-wing aircraft under normal good view, partial occlusion view, and flipped view respectively, and the number of key points is 5; Figure 3 in (a), (b), and (c) are the detection results of the quadrotor aircraft under blurred view, partial occlusion view, and normal good view respectively, and the number of key points is 7. In the present invention, a resolution preservation component and a feature fusion component are added to a general convolutional network, which ensures that high-resolution features are not lost during feature extraction, thereby extracting more effective position encodings. Further, after the model completes key point prediction and calculates the image gradients of the prediction point and its surrounding adjacent pixel points, the point where the maximum gradient value is located is used as the final key point output. Finally, through the longest alignment method, key point detection of multiple categories can be achieved in one network model, greatly improving the flexibility of the network.
[0034] Although the present invention has been described herein with reference to its explanatory embodiments, the above embodiments are only the preferred embodiments of the present invention. The embodiments of the present invention are not limited by the above embodiments. It should be understood that those skilled in the art can design many other modifications and embodiments, which will fall within the scope and spirit of the principles disclosed in this application.
Claims
1. A multi-class key point detection method based on fusing high-resolution features, characterized in that Including: Step S100: Collect the image data of the target to be detected to obtain an image data sample; Step S200: Label the bounding box and key point coordinates of the target to be detected in the image data sample to obtain bounding box data and key point data; Step S300: Determine whether the key point data needs to be aligned to the longest length. If so, perform the longest alignment and then convert the key point data after the longest alignment into a Gaussian distribution map. Otherwise, directly convert the key point data into a Gaussian distribution map; Step S400: Train an object detector using the bounding box data. This object detector is used to predict the position information of the target to be detected for key point detection in the image to be detected; Step S500: Use the Gaussian distribution map as a label and input it into the improved deep learning key point detection model for training, and obtain the trained key point detection model; Step S600: Input the image to be detected into the trained object detector to obtain all the target position information in the image to be detected. Crop out the position of the target in the image to be detected to obtain a sub-image. If the size of the sub-image is not equal to the preset size, stretch or compress the sub-image to the preset size, and then input it into the trained key point detection model to obtain the key point coordinate information; Step S700: Based on the key point coordinate information predicted by the model, calculate the image gradient of the key point and its adjacent pixels, and regard the point with the largest gradient change as the final key point.
2. The multi-class key point detection method based on fused high-resolution features according to claim 1, wherein The method for determining whether the key point data needs to be aligned to the longest length in step S300 is as follows: When the number of key points of a certain category in the target to be detected is equal to the maximum number of key points in all categories, no longest alignment is required. Otherwise, longest alignment is required.
3. The multi-class key point detection method based on fused high-resolution features according to claim 2, wherein, The longest alignment method in step S300 is as follows: Step S310: If labeling key points of the same category, the labeling order of key points for each target needs to be consistent; Step S320: If labeling key points of different categories and the number of key points between categories is inconsistent, after the target with fewer key points is labeled in order, a padding operation with 0 needs to be performed at the end until the number is filled up to the maximum number of key points in all categories, so as to complete the longest alignment.
4. The multi-class key point detection method based on fused high-resolution features according to claim 3, wherein The method for converting the key point data into a Gaussian distribution map is as follows: Step A1: Convert the coordinates of the key point into a scale factor relative to the width and height of the image; Step A2: Fix the width and height of the target to be measured in advance. If the obtained image is inconsistent with the preset width and height, then stretch or compress the image to the preset value; Step A3: Generate an image with 1 channel, the size of the preset value, and all pixel values being 0; Step A4: Generate a two-dimensional Gaussian distribution map at the corresponding position of the image generated in step A3 through the scale factor. The corresponding function is: ; In the formula, represents the value of the generated two-dimensional Gaussian distribution graph at coordinates, represents the amplitude, and represent the center of the two-dimensional Gaussian distribution graph, that is, the key point coordinates, and respectively represent the variances manually preset on the x-axis and y-axis in the Gaussian function.
5. The multi-class key point detection method based on fused high-resolution features according to claim 1, wherein The improved deep learning keypoint detection model includes an encoder and a decoder. The encoder has four layers, and each layer is composed of a first module. The first module is composed of a resolution-preserving component, a downsampling component, and a first feature fusion component. The resolution-preserving component is composed of a convolutional layer, a BN layer, and an activation function layer. The downsampling component is composed of a convolutional layer, a BN layer, an activation function layer, and a pooling layer. The first feature fusion component is composed of a convolutional layer, a BN layer, an activation function layer, a pooling layer, and a channel concatenation layer. In the first module of the first layer of the encoder, the first feature fusion component only needs to fuse the output of the resolution-preserving component in the first module of this layer and the output features of the downsampling component. In the first modules of the subsequent three layers of the encoder, in addition to fusing the output of the resolution-preserving component in the first module of the corresponding layer and the output features of the downsampling component, the first feature fusion component also fuses the output of the resolution-preserving component in the first module of the first layer. The decoder has four layers, and each layer is composed of a second module. Each second module is divided into two parts: an upsampling component and a second feature fusion component. The upsampling component is composed of a transposed convolutional layer. The second feature fusion component includes a feature channel concatenation layer, a convolutional layer, a BN layer, and an activation function. The second feature fusion layer fuses the feature map output by the resolution-preserving component in the first module of the same layer of the encoder and the feature map after the upsampling component in the second module of the same layer of the decoder to ensure that the decoder decodes in the correct direction.
6. The multi-class key point detection method based on fused high-resolution features according to claim 5, wherein, The loss function of the improved deep learning keypoint detection model is defined as: ; Among them, m represents the number of the maximum key points, and n represents the number of pixels of each output feature map. represents the th pixel value of the Gaussian distribution map corresponding to the real th key point, represents the th pixel value of the Gaussian distribution map corresponding to the predicted th key point.
7. The multi-class key point detection method based on fused high-resolution features according to claim 6, characterized in that The specific method for calculating the image gradients of the keypoints and their adjacent pixels in step S700 is as follows: The gradient value calculation template in the x direction: ; The gradient value calculation template in the y direction: ; The gradient value of the corresponding pixel point: ; Among them, represents the template in the x direction, I represents the corresponding image region, and represents the template in the y direction; , are the absolute values of the gradient values in the x direction and y direction; Finally, the pixel point with the maximum gradient is regarded as the position of the final keypoint.
Citation Information
Patent Citations
Method for detecting tree-shaped structure bifurcation key point in three-dimensional tomography image
CN112541893A
Target key point detection method and device, equipment and storage medium
CN114463534A
Poultry behavior analysis method based on computer vision
CN115147918A
Key point detection model training method, key point detection method and device
CN119273978A