A Pose Estimation Method for Improving the Resolution of Small Targets
By enhancing resolution of the human body frame of small targets and multi-channel attention network with full attention network, the error problem caused by low resolution of small targets is solved, and the accuracy of posture estimation is improved.
Patent Information
- Application Number
- CN202310222976.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-03-09
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2043-03-09
AI Technical Summary
The prior art has failed to effectively solve the error problem caused by low resolution of small targets in human posture estimation, and has not fully considered the global context concern and the connection between channels.
By judging the size of the human body frame, token mark the human body frame of the small target, and input the marked result to the module that improves the resolution for resolution enhancement. The improved full attention network is adopted to maintain high-resolution feature maps for multi-channel attention, reducing the accuracy error caused by small targets and low resolution.
The resolution of small targets is improved, the accuracy error in pose estimation is reduced, the detection effect is enhanced, and the accuracy of pose estimation is effectively improved.
Smart Images

Figure CN116386133B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of pose estimation, and more particularly, to a pose estimation method for improving the resolution of small targets. Background Art
[0002] In recent years, using deep learning for human pose estimation has been one of the issues that computer vision has been concerned about. Human pose estimation refers to estimating the key point parts of the human body, such as the positions of hands, head, shoulder joints, hip joints, ankles, etc., and is usually associated with behavior analysis, gesture analysis, motion capture, etc. Continuously tracking and estimating a person's pose can be used to detect whether a person has a tendency to fall, whether they have a certain disease (such as Parkinson's disease, etc.). In addition, it can also be used to detect whether the technical movements of athletes are standardized.
[0003] After retrieval, the patent number is CN114999002A, which discloses a behavior recognition method that fuses human pose information. Although this invention has strong stability and overcomes the influence of the recognition ability of the graph convolutional neural network being greatly affected by the translation of the bone point coordinates, and in addition, it fuses the information of the front and back frames of the image and the information of human key points, and the fusion of information helps to improve the performance of action recognition. However, this method does not consider the error caused by the low resolution of small targets in action recognition; in addition, this method does not pay attention to the global context, but operates on the local part and does not consider the connection between channels. Summary of the Invention
[0004] According to the above-mentioned technical problems, a pose estimation method for improving the resolution of small targets is provided. The present invention judges the size of the human body frame where the human body part is located in the image, uses the size of the human body frame as the judgment basis, and inputs the picture into the module for improving the resolution, so as to enhance the resolution of small targets. An improved full attention network is adopted to keep the features input into the network Figure 1 in a high-resolution state all the time and perform multi-channel attention, reducing the accuracy error caused by the low resolution of small targets.
[0005] The technical means adopted by the present invention are as follows:
[0006] A pose estimation method for improving the resolution of small targets, comprising:
[0007] Obtain an image, perform human body frame detection on the image, and mark the human body frames of small targets;
[0008] Enhance the resolution of the marked human body frames of small targets;
[0009] Based on the image after resolution enhancement, design an attention network that maintains high resolution.
[0010] Further, the obtaining of the image, performing human body bounding box detection on the image, and marking the human body bounding boxes of small targets includes:
[0011] Input the obtained image into a convolutional neural network with three convolutional layers and one fully connected layer for human body bounding box detection, assign an ID to all the people in the picture information. After the assignment, judge the size of the detected human body bounding boxes, and perform Token marking on the human body bounding boxes of small targets.
[0012] Further, the inputting of the obtained image into a convolutional neural network with three convolutional layers and one fully connected layer for human body bounding box detection, assigning an ID to all the people in the picture information, judging the size of the detected human body bounding boxes after the assignment, and performing Token marking on the human body bounding boxes of small targets specifically includes:
[0013] Input the picture into the convolutional neural network, which includes three convolutional layers, one fully connected layer, and small target marking operations. The first convolutional layer module contains a convolutional layer with 256 3×3 convolutions, a BN layer, and a RELU layer; the second convolutional layer module contains a convolutional layer with 512 3×3 convolutions, a BN layer, and a RELU layer; the third convolutional layer module contains a convolutional layer with 512 3×3 convolutions, a BN layer, and a RELU layer. The input image passes through the three convolutional layers and the fully connected layer in sequence for feature extraction and human body bounding box marking. After obtaining the human body bounding box marking, judge the size of each human body bounding box, perform Token marking on the sitting position of the human body bounding boxes of small targets, and output the marked result.
[0014] Further, the resolution enhancement of the marked human body bounding boxes of small targets includes:
[0015] Perform two deconvolution operations and operations of the fully connected layer on the small target marking area, and at the same time perform bilinear interpolation to finally obtain the result of improved resolution.
[0016] Further, the performing of two deconvolution operations and operations of the fully connected layer on the small target marking area, and at the same time performing bilinear interpolation to finally obtain the result of improved resolution specifically includes:
[0017] Input the marked result into the small target resolution enhancement network, which consists of two deconvolution modules and one fully connected layer. The first deconvolution module has a convolutional layer with 256 3×3 convolutions, a BN layer, and a RELU layer; the second deconvolution module has a deconvolution layer with 512 3×3 convolutions, a BN layer, and a RELU layer; enhance the resolution of the small targets through the deconvolution operations of the two deconvolution modules.
[0018] Furthermore, design an attention network that maintains high resolution based on the image after resolution enhancement, including:
[0019] Perform a multi-head self-attention mechanism operation on the image after resolution enhancement, then use a transposed convolutional layer to increase the resolution to meet the requirement of maintaining high resolution, and finally perform channel self-attention and a multi-layer perceptron to finally generate joint heatmaps.
[0020] Furthermore, the above-mentioned operation of performing a multi-head self-attention mechanism on the image after resolution enhancement, then using a transposed convolutional layer to increase the resolution to meet the requirement of maintaining high resolution, and finally performing channel self-attention and a multi-layer perceptron to finally generate joint heatmaps specifically includes:
[0021] Construct a full attention network that maintains high resolution. The full attention network that maintains high resolution consists of an MHSA layer, a transposed convolutional module, a Channel Self-Attention layer, and an MLP layer; the transposed convolutional module contains a transposed convolutional layer with 3×3 convolution, a BN layer, and a RELU layer; the MHSA multi-head self-attention module is used to extract the structural information in the result of increasing the resolution, establish associations between each element in the abstract feature map, and parallelly calculate and select multiple pieces of information from the input information. Each attention focuses on different parts of the input information, and then they are concatenated. The corresponding receptive field is the entire image; perform a transposed convolutional operation on the obtained information to ensure the resolution of the image, and then perform the Channel Self-Attention layer. The Channel Self-Attention layer will calculate a channel weight, mainly focusing on different channel information of the input. Input the information after increasing the resolution into the MLP layer, and fuse the matrices passed through the Channel Self-Attention layer and the MLP layer to achieve full-channel attention and obtain key point heatmaps.
[0022] Compared with the prior art, the present invention has the following advantages:
[0023] 1. The pose estimation method for improving the resolution of small targets provided by the present invention marks all the people in the picture information through a convolutional neural network, assigns an ID to all the people in the picture information. After the assignment is completed, judge the size of the detected human body box, and perform Token marking on the human body box of small targets whose pixel value is less than 1 / 3 of the pixel of the overall picture. Then input the marked result into the module for improving the resolution of small targets. This module only performs the operation of increasing the resolution on the pixels marked with tokens, reducing the time for increasing the resolution of all people.
[0024] 2. The pose estimation method for improving the resolution of small targets provided by the present invention performs a deconvolution operation after passing through the MHSA multi-head self-attention module to ensure the high-resolution requirements during the prediction process, conducts channel self-attention, focuses on different channel information of the input, and predicts the key point heat map after channel self-attention and through the MLP layer. During the process of generating the key point heat map, high resolution is always maintained, the detection effect is improved, and the accuracy of pose estimation is effectively improved.
[0025] For the above reasons, the present invention can be widely promoted in the fields of pose estimation and the like. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0027] Figure 1 It is a flowchart of the method of the present invention.
[0028] Figure 2 It is a network diagram for marking small targets based on a neural network provided by an embodiment of the present invention.
[0029] Figure 3 It is a network diagram for improving the resolution of small targets provided by an embodiment of the present invention.
[0030] Figure 4 It is a network diagram of full attention provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other. The following will refer to the drawings and combine the embodiments to detail the present invention.
[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. The following description of at least one exemplary embodiment is actually only illustrative and in no way constitutes a limitation to the present invention and its application or use. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.
[0033] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present invention. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, they specify the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0034] Unless otherwise specifically stated, the relative arrangements of components and steps, numerical expressions, and values set forth in these embodiments do not limit the scope of the present invention. At the same time, it should be clear that, for the sake of convenience of description, the dimensions of the various parts shown in the drawings are not drawn in actual proportional relationships. Technologies, methods, and devices known to those of ordinary skill in the relevant art may not be discussed in detail, but where appropriate, such technologies, methods, and devices should be regarded as part of the authorized specification. In all the examples shown and discussed here, any specific values should be construed as merely exemplary and not as a limitation. Therefore, other examples of the exemplary embodiments may have different values. It should be noted that like reference numerals and letters denote like items in the following drawings, and thus, once an item is defined in one drawing, further discussion thereof is not required in subsequent drawings.
[0035] In the description of the present invention, it should be understood that the orientation or positional relationships indicated by orientation words such as "front, rear, upper, lower, left, right", "lateral, vertical, perpendicular, horizontal", and "top, bottom" are generally based on the orientation or positional relationships shown in the drawings, and are only for the convenience of describing the present invention and simplifying the description. Without contrary description, these orientation words do not indicate and imply that the device or element referred to must have a specific orientation or be constructed and operated in a specific orientation, and thus cannot be construed as limiting the protection scope of the present invention: the orientation words "inside, outside" refer to the inside and outside relative to the contour of each component itself.
[0036] For ease of description, spatial relative terms such as "above", "over", "on the upper surface", "upper", etc. can be used here to describe the spatial positional relationship between a device or feature shown in a figure and other devices or features. It should be understood that spatial relative terms are intended to encompass different orientations in use or operation in addition to the orientation depicted in the figure for the device. For example, if the device in the figure is inverted, a device described as "above" or "over" other devices or structures will then be positioned "below" or "under" other devices or structures. Thus, the exemplary term "above" can include both the orientations of "above" and "below". The device can also be positioned in other different ways (rotated 90 degrees or in other orientations), and corresponding interpretations are made for the spatial relative descriptions used here.
[0037] In addition, it should be noted that the use of terms such as "first", "second", etc. to limit components is merely for the convenience of differentiating the corresponding components. Without additional statements, the above terms have no special meanings, and thus should not be construed as limiting the protection scope of the present invention.
[0038] As Figure 1 shown, the present invention provides a pose estimation method for improving the resolution of small targets, including:
[0039] S1. Obtain an image, perform human bounding box detection on the image, and mark the human bounding boxes of small targets; that is, input the obtained image into a convolutional neural network with three convolutional layers and one fully connected layer for human bounding box detection, assign an ID to all persons in the picture information, after the assignment, judge the size of the detected human bounding boxes, and perform Token marking on the human bounding boxes of small targets.
[0040] S2. Enhance the resolution of the marked human bounding boxes of small targets; that is, perform two deconvolution operations and operations of the fully connected layer on the marked area of small targets, and at the same time perform bilinear interpolation to finally obtain the result of improved resolution.
[0041] S3. Design an attention network that maintains high resolution based on the image after resolution enhancement. That is, perform multi-head self-attention mechanism operations on the image after resolution enhancement, then use a deconvolution layer to improve the resolution to meet the requirement of maintaining high resolution, and finally perform channel self-attention and a multi-layer perceptron to finally generate joint point heatmaps.
[0042] Specifically in implementation, as a preferred implementation manner of the present invention, as Figure 2As shown, it is a network diagram for marking small targets based on a neural network provided by an embodiment of the present invention. In the step S1, the specific implementation steps for obtaining an image, performing human box detection on the image, and marking the human boxes of small targets are as follows:
[0043] Input the picture into a convolutional neural network, which includes three convolutional layers, a fully connected layer, and a small target marking operation. The first convolutional layer module contains a convolutional layer with 256 3×3 convolutions, a BN layer, and a RELU layer; the second convolutional layer module contains a convolutional layer with 512 3×3 convolutions, a BN layer, and a RELU layer; the third convolutional layer module contains a convolutional layer with 512 3×3 convolutions, a BN layer, and a RELU layer. The input image passes through the three convolutional layers and the fully connected layer in sequence for feature extraction and human box marking. After obtaining the human box marking, judge the size of each human box, perform Token marking on the position of the human box of the small target, and output the marked result.
[0044] In this embodiment, the small targets required by the present invention refer to human box sizes with pixel points less than 32*32, medium targets refer to human box sizes with pixel points greater than 32*32 and less than 64*64, and large targets refer to human box sizes exceeding 64*64. The input image is processed using Figure 2 a convolutional neural network to detect humans and perform small target marking on the detected humans. The convolutional network includes an input layer, a convolutional layer, a fully connected layer, and an output layer. The kernel size of the convolutional layer is 3*3. In each convolutional layer, ReLU is used as the activation function. Judge the result output after the fully connected layer. If the detected human border is smaller than the small target size in the setting, mark the human box, and subsequent operations to improve the resolution will be performed on this part.
[0045] The loss function used for this human detection is as follows:
[0046] L(p, u, t u , v) = L cls (p, u) + λ|u≥1L loc (t a , v)
[0047]
[0048] in which
[0049]
[0050] where P represents the predicted class score, u represents the true class, t u represents the true bounding box coordinates, v represents the predicted bounding box coordinates, and L cls(p, u) represents the classification loss softmax, λ[u≥1]L loc (t u , v) is the bounding box regression loss SmoothL1 Loss. When u = 0, it is the background class, which does not correspond to any actual class and has no bounding box regression loss.
[0051] The Softmax loss function is as follows:
[0052]
[0053] Among them, z i is the output value of the i-th node, and C is the number of output nodes, that is, the number of classification categories. Through the Softmax function, the output values of multi-classification can be converted into a probability distribution ranging from [0, 1] and summing to 1.
[0054] The Smooth L1 Loss loss function is as follows:
[0055]
[0056] Among them, x = f(x i ) - y i is the difference between the true value and the predicted value. Smooth L1 can limit the gradient in two aspects: when the predicted box is very different from the ground truth, the gradient value will not be too large; when the predicted box is very similar to the ground truth, the gradient value is small enough.
[0057] In specific implementation, as a preferred implementation manner of the present invention, as Figure 3 shown, it is a network diagram for enhancing the resolution of small targets provided by an embodiment of the present invention. In step S2, the specific implementation steps for enhancing the resolution of the human body box of the marked small target are as follows:
[0058] Input the marked result into the small target resolution enhancement network, which consists of two deconvolution modules and a fully connected layer. The first deconvolution module has a convolutional layer with 256 3×3 convolutions, a BN layer, and a RELU layer; the second deconvolution module has a deconvolution layer with 512 3×3 convolutions, a BN layer, and a RELU layer; the resolution of the small target is enhanced through the deconvolution operations of the two deconvolution modules.
[0059] In this embodiment, the small target marking result is used to perform an operation to improve the resolution. For this purpose, a convolutional neural network is established. The convolutional neural network includes an input layer, a transposed convolutional layer, a fully connected layer, and an output layer. The kernel size of the transposed convolutional layer is 3*3, there is no padding, the stride is 1, and for each transposed convolutional layer, ReLU is used as the activation function. The transposed convolution operation is used to increase the size of the upsampled feature map to a preset multiple, and bilinear interpolation is used to generate a high-resolution image from a low-resolution image to restore the lost information in the image. This algorithm reduces certain visual distortions caused by resizing the image to a non-integer scaling factor.
[0060] Formula for calculating the output dimension of the transposed convolution:
[0061] output=(input - 1)*s + k - 2*p
[0062] Where input is the size of the input matrix, p is the padding, k is the kernel size, and s is the convolution stride.
[0063] The goal of bilinear interpolation is to obtain the pixel value of the unknown function f at the point P(x, y), Q 11 (x 1 , y 1 ), Q 12 (x 1 , y 2 ), Q 21 (x 2 , y 1 ), Q 22 (x 2 , y 2 ) are the coordinates and corresponding pixel values of the four adjacent points of point P. First, interpolation is performed in the x direction to obtain R 1 and R 2 , and then interpolation is performed in the y direction to obtain P, where the calculation formulas for R 1 and R 2 are as follows:
[0064]
[0065]
[0066] By interpolating in the y direction, we get:
[0067]
[0068] In specific implementation, as a preferred implementation manner of the present invention, such as Figure 4As shown, it is the network diagram of full attention provided by the embodiment of the present invention. In the step S3, based on the image after resolution enhancement, the specific implementation steps for designing an attention network that maintains high resolution are as follows:
[0069] Construct a full attention network that maintains high resolution. The full attention network that maintains high resolution consists of an MHSA layer, a transposed convolution module, a Channel Self-Attention layer, and an MLP layer; the transposed convolution module includes a transposed convolution layer with a 3×3 convolution, a BN layer, and a RELU layer; the MHSA multi-head self-attention module is used to extract the structural information in the upsampled result, establish associations between each element in the abstract feature map, and parallelly calculate and select multiple pieces of information from the input information. Each attention focuses on different parts of the input information, and then they are concatenated. The corresponding receptive field is the entire image; perform a transposed convolution operation on the obtained information to ensure the resolution of the image. After upsampling, enter the Channel Self-Attention layer. The Channel Self-Attention layer will calculate a channel weight, mainly focusing on different channel information of the input. Input the upsampled information into the MLP layer, and fuse the matrices passed through the Channel Self-Attention layer and the MLP layer to achieve full-channel attention and obtain the key point heat map.
[0070] In this embodiment, the upsampled result is used for full attention operation. For this purpose, a full attention recognition network is designed. The full attention recognition network includes a multi-head self-attention mechanism (MHSA), a transposed convolution layer, channel self-attention (Channel Self-Attn), and a multi-layer perceptron (MLP). The transposed convolution layer uses ReLU as the activation function. The MHSA multi-head self-attention module is used to extract the structural information in the upsampled result and establish associations between each element in the abstract feature map. The corresponding receptive field is the entire image. The transposed convolution layer is used to increase the resolution in the full attention recognition network to avoid losing information during the convolution operation. After the transposed convolution operation, the information is simultaneously transmitted to the channel self-attention and the multi-layer perceptron, and the results obtained from both are fused, enabling the entire network to perform self-attention and channel attention simultaneously and maintaining globality.
[0071] Convert the three feature maps of high, medium, and small targets into one-dimensional feature sequences respectively. Based on the multi-head attention mechanism method, extract the features of the three branches respectively. The calculation formula is as follows:
[0072]
[0073] where i = 1, 2, 3, representing the three branches of small, medium, and large targets respectively, and mha iis the output of the multi-head attention mechanism, Q i represents the query vector, κ i represents the key vector, V i represents the value vector, d i represents the number of dimensions of the mapping, T represents the matrix transpose operation, and softmax is the multi-classification function.
[0074] The calculation formula of the channel attention mechanism is as follows:
[0075] att = sigmoid(fc(gap(X))), X ∈ R C×H×W
[0076] where X ∈ R C×H×W is the image feature tensor, C is the number of channels, H is the height of the feature map, and W is the width of the feature map. where att ∈ R C is the attention vector, sigmoid is the sigmoid function, fc represents the fully connected layer, and gap is the global average pooling.
[0077] Input the feature map of the upsampling result into the full attention recognition network. The full attention recognition network uses the L2 loss function to obtain the predicted value of the key point heat map. If only one predicted value is given for all sample points, then this value is the average of all target values. The formula of the L2 loss function is as follows:
[0078]
[0079] It minimizes the sum of the squares of the differences between the target value Y i and the estimated value f(x i ).
[0080] The human key points specifically include the nose, left eye, right eye, left ear, right ear, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. When training the pose estimation network for improving the resolution of small targets, the dataset used is the human pose estimation dataset in the CoCo dataset. The dataset is divided into a training set, a validation set, and a test set according to the ratio of 7:2:1. The iterative method is used to train this behavior prediction network. After 300 iterations of training, the parameter model is saved. Then, the test set is input into the trained model, and the predicted percentage of the pose data is output. The highest pose data is used as the predicted result for output.
[0081] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A pose estimation method for improving the resolution of small targets, characterized in that, it includes: Obtain an image, perform human body box detection on the image, and mark the human body boxes of small targets; Enhance the resolution of the human body boxes of the marked small targets; Based on the image after resolution enhancement, design an attention network that maintains high resolution, including: Perform multi-head self-attention mechanism operations on the image after resolution enhancement, then use a transposed convolution layer to improve the resolution and meet the requirement of maintaining high resolution, and finally perform channel self-attention and a multi-layer perceptron to finally generate joint heatmaps.
2. The pose estimation method for improving the resolution of small targets according to claim 1, characterized in that, the obtaining of the image, performing human body box detection on the image, and marking the human body boxes of small targets includes: Input the obtained image into a convolutional neural network with three convolutional layers and one fully connected layer for human body box detection, assign an ID to all the people in the picture information, after the assignment, judge the size of the detected human body boxes, and perform Token marking on the human body boxes of small targets.
3. The pose estimation method for improving the resolution of small targets according to claim 2, characterized in that, the inputting of the obtained image into a convolutional neural network with three convolutional layers and one fully connected layer for human body box detection, assigning an ID to all the people in the picture information, after the assignment, judging the size of the detected human body boxes, and performing Token marking on the human body boxes of small targets specifically includes: Input the image into a convolutional neural network, which includes three convolutional layers, a fully connected layer, and a small target marking operation. The first convolutional layer module contains 256 convolutional convolutional layers, a BN layer, and a RELU layer; the second convolutional layer module contains 512 convolutional convolutional layers, a BN layer, and a RELU layer; the third convolutional layer module contains 512 convolutional convolutional layers, a BN layer, and a RELU layer. The input image passes through the three convolutional layers and the fully connected layer in sequence for feature extraction and human body box marking. After obtaining the human body box marking, judge the size of each human body box, perform Token marking on the sitting position of the small target human body box, and output the marked result.
4. The pose estimation method for improving the resolution of small targets according to claim 1, characterized in that, the enhancing of the resolution of the human body boxes of the marked small targets includes: Perform two transposed convolution operations and operations of a fully connected layer on the marked area of the small target, and at the same time perform bilinear interpolation to finally obtain the result of improved resolution.
5. The pose estimation method for improving the resolution of small targets according to claim 4, characterized in that, the performing of two transposed convolution operations and operations of a fully connected layer on the marked area of the small target, and at the same time performing bilinear interpolation to finally obtain the result of improved resolution specifically includes: The marked result is input into the small target resolution enhancement network, which consists of two deconvolution modules and a fully connected layer. The first deconvolution module has 256 convolutional convolutional layers, a BN layer, and a RELU layer; the second deconvolution module has 512 deconvolutional layers with convolution, a BN layer, and a RELU layer; the resolution of the small target is enhanced through the deconvolution operations of the two deconvolution modules.
6. The pose estimation method for improving the resolution of small targets according to claim 4, characterized in that, the performing of multi-head self-attention mechanism operations on the image after resolution enhancement, then using a transposed convolution layer to improve the resolution and meet the requirement of maintaining high resolution, and finally performing channel self-attention and a multi-layer perceptron to finally generate joint heatmaps specifically includes: Construct a full attention network that maintains high resolution. The full attention network that maintains high resolution consists of an MHSA layer, a transposed convolution module, a Channel Self-Attention layer, and an MLP layer; the transposed convolution module contains a transposed convolution layer for convolution, a BN layer, and a RELU layer; the MHSA multi-head self-attention module is used to extract the structural information in the result of improving the resolution, establish associations between each element in the abstract feature map, and calculate in parallel the selection of multiple pieces of information from the input information. Each attention focuses on different parts of the input information, and then they are concatenated. The corresponding receptive field is the entire image; perform a transposed convolution operation on the obtained information to ensure the resolution of the image. The Channel Self-Attention layer will calculate a channel weight, focus on different channel information of the input, input the information after improving the resolution into the MLP layer, and fuse the matrices passed through the Channel Self-Attention layer and the MLP layer to achieve full-channel attention and obtain the joint heat map.
Citation Information
Patent Citations
Behavior recognition method fusing human body posture information
CN114999002A
quick low-illumination target detection method based on convolutional neural network
CN113052210A
Method and device for intelligent estimation of human body movement posture based on convolutional neural network
WO2022036777A1