An Online Cross-Channel Interactive Parallel Distillation Architecture Pose Estimation Method and Device

Through the online cross-channel interactive parallel distillation architecture, combined with the channel and spatial attention mechanism, the problem of failure to effectively combine channel characteristics and spatial characteristics in the prior art is solved, and higher posture prediction accuracy and stability are achieved.

CN115359571BActive Publication Date: 2025-07-18XIAMEN INFORMATION TECH APPL INNOVATION RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211061531.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-01
Publication Date
2025-07-18
Estimated Expiration
2042-09-01

AI Technical Summary

Technical Problem

The existing attention model fails to effectively combine channel feature information and spatial feature information, resulting in low accuracy of human posture estimates.

Method used

A online cross-channel interactive parallel distillation architecture is designed. By introducing a cross-channel interactive attention mechanism and a spatial attention mechanism, combining channel attention and spatial attention, the channel similarity is calculated using the covariance matrix, and an online parallel knowledge distillation method is introduced in the Faster-Pose pose detection model to improve the ability to express feature information.

Benefits of technology

It improves the accuracy and stability of human posture prediction, enhances the ability to express characteristic information, and improves the accuracy of target detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115359571B_ABST
    Figure CN115359571B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of artificial intelligence computer vision technology, and specifically provides an online cross-channel interactive parallel distillation architecture pose estimation method. First, a video acquisition device obtains an external video stream, and the video stream is sliced into frames and input into a feature extraction network for feature extraction; the extracted features are sent to a YOLOV5 object detection model to detect the position of the target human body in each frame of image and mark the detection frame, obtaining the feature data of the target human body; the target human body feature data is passed to the pose detection model Faster-Pose to obtain the human key point feature information; the obtained human key point feature information is mapped to the feature map through a linear transformation to obtain a feature map with human key point annotations. Compared with the prior art, the present invention takes into account the correlation between channel feature information and spatial feature information, and improves the expression ability of the required feature information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence computer vision, and specifically provides an online cross-channel interactive parallel distillation architecture pose estimation method and device. Background Art

[0002] Human pose estimation is an important technical field in artificial intelligence computer vision. By estimating the behavior of people in a scene, better human-computer interaction can be achieved. Currently, human pose estimation is commonly used in worker violation detection, the security field, and VR wearable devices. In human pose estimation algorithms, the accuracy of the target detection model in extracting the human detection box is crucial for the accuracy and stability of human key point positioning.

[0003] Existing attention models do not consider the correlation between channel feature information and spatial feature information, resulting in low accuracy. Summary of the Invention

[0004] The present invention aims at the above-mentioned deficiencies of the prior art and provides a practical online cross-channel interactive parallel distillation architecture pose estimation method.

[0005] A further technical task of the present invention is to provide an online cross-channel interactive parallel distillation architecture pose estimation device with reasonable design, safety and applicability.

[0006] The technical solution adopted by the present invention to solve its technical problems is as follows:

[0007] An online cross-channel interactive parallel distillation architecture pose estimation method, first, a video acquisition device obtains an external video stream, and the video stream is split into frames and input into a feature extraction network for feature extraction;

[0008] The extracted features are fed into the YOLOV5 target detection model to detect the position of the target human body in each frame of the image and mark the detection box, obtaining the feature data of the target human body;

[0009] The target human body feature data is passed to the pose detection model Faster-Pose to obtain human key point feature information; the obtained human key point feature information is mapped into the feature map through linear transformation to obtain a feature map with human key point annotations.

[0010] Further, the feature extraction network is designed as a CSP structure, and a cross-channel interactive attention mechanism is introduced. The cross-channel interactive attention mechanism combines channel attention and spatial attention, and uses the covariance matrix to calculate the similarity between every two channels of the feature map in the channel attention model, and the channels with high similarity are fused;

[0011] In spatial attention, the second-order finite difference method is used to calculate the difference of feature image pixel values and the pixel gradient direction.

[0012] Furthermore, if the calculated covariance value of the two channels is negative, it indicates negative correlation; if the value is 0, it means that the two channels are independent and uncorrelated with each other; if the value is positive, it indicates positive correlation between the two channels for feature fusion.

[0013] First, calculate the mean of each channel as shown in formula (1):

[0014]

[0015] The mean feature of all channels is denoted as

[0016] Calculate the variance of each channel as shown in formula (2):

[0017]

[0018] The variances of all channels are denoted as:

[0019] Calculate the covariance between channels C1 and C2 as shown in formula (3):

[0020]

[0021] All the obtained channel similarity covariance values are denoted as Cov i,k , and covariance is only calculated between different channels. Channels with positive correlation are fused pixel by pixel according to the covariance value.

[0022] Furthermore, the extracted features are fed into the YOLOV5 object detection model to detect the position of the target human body in each frame of the image and mark the detection box, obtaining the feature data of the target human body.

[0023] The YOLOV5 object detection model is trained using the publicly available dataset MSCOCO 2017. The MSCOCO 2017 dataset is randomly sampled according to a preset ratio, and data augmentation preprocessing operations are performed on the sample data.

[0024] Preferably, the data augmentation methods include rotating the image at multiple angles with a rotation angle interval of 30 degrees, randomly masking the image according to the probability P, setting the pixel values under the mask to 0, flipping the image up, down, left, and right, performing different degrees of distortion processing on the image, and performing color perturbation on the image.

[0025] Furthermore, the channel attention model uses the SoftMax function to obtain the channel feature probability matrix, and the spatial attention model uses the SoftMax function to obtain the spatial feature probability matrix;

[0026] The probability matrix and the original feature map are fused by multiplication respectively to add weight information to the feature map.

[0027] Furthermore, the Depth-Wise method is used to extract features from each channel of the feature map to obtain the eigenvalue matrix of each channel, and cross-channel feature fusion is performed.

[0028] Furthermore, an online parallel knowledge distillation method is carried out in the pose detection model Faster-Pose. The online parallel knowledge distillation method continues to use the Teacher-Student knowledge distillation framework in the network structure. The Teacher network consists of 8 Hourglass feature extraction modules, and the Student network consists of 4 Hourglass feature extraction modules;

[0029] The Teacher network is trained using the MSCOCO 2017 dataset, and the Student network is trained using a part of the labeled dataset. During the training process, the KL divergence is used to calculate the loss between the feature maps of the Teacher network and the Student network, and the Teacher feature map information and the Student feature map information are fused according to the channel similarity. The Teacher and Student networks are trained in parallel during the training process;

[0030] During the inference process, the Teacher network is removed and the Student network is directly inferred. A cross-channel interactive attention mechanism is introduced into the Faster-Pose pose detection model. The cross-channel interactive attention mechanism assigns different weight information to the feature maps in the Teacher network. The calculation processes of the Teacher network feature map and the Student network feature map are shown in formula (4):

[0031]

[0032] where, respectively represent the feature map extracted by the second Hourglass module of the Teacher network and the feature map extracted by the first Hourglass module of the Student network;

[0033] The total feature map loss is shown in formula (5):

[0034]

[0035] The final loss function of the Faster-Pose pose estimation model is shown in Equation (6):

[0036]

[0037] Where is the loss of the Student network model, and α and λ are hyperparameters to be learned.

[0038] Furthermore, the human key point Heat Map data information output by the Faster-Pose pose detection model is mapped to the original feature map using linear interpolation, and the pixel point offset that appears during the mapping process is corrected using trilinear interpolation.

[0039] An online cross-channel interactive parallel distillation architecture pose estimation device includes: at least one memory and at least one processor;

[0040] The at least one memory is used to store machine-readable programs;

[0041] The at least one processor is used to call the machine-readable program to execute an online cross-channel interactive parallel distillation architecture pose estimation method.

[0042] Compared with the prior art, the online cross-channel interactive parallel distillation architecture pose estimation method and device of the present invention have the following outstanding beneficial effects:

[0043] The present invention proposes a cross-channel interactive attention mechanism and a new pose detection model, Faster-Pose. In the feature extraction stage, channel attention is used to detect the feature expressions on which channels of the feature map contain the required information, and spatial attention detects which positions on the feature map have the required feature information. In the present invention, the feature information extracted by spatial attention and the feature information extracted by channel attention are fused, taking into account the correlation between channel feature information and spatial feature information, and improving the expression ability of the required feature information. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0045] Att Figure 1 is a schematic flow diagram of an online cross-channel interactive parallel distillation architecture pose estimation method;

[0046] Att Figure 2It is a schematic diagram of the pose estimation algorithm model framework in an online cross-channel interactive parallel distillation architecture pose estimation method;

[0047] Appendix Figure 3 It is a schematic diagram of the preprocessing process of the MSCOCO 2017 dataset in an online cross-channel interactive parallel distillation architecture pose estimation method;

[0048] Appendix Figure 4 It is a schematic diagram of the YOLOV5 object detection algorithm framework in an online cross-channel interactive parallel distillation architecture pose estimation method;

[0049] Appendix Figure 5 It is a schematic diagram of the C-CIAM attention mechanism framework in an online cross-channel interactive parallel distillation architecture pose estimation method;

[0050] Appendix Figure 6 It is a schematic diagram of the Faster-Pose pose detection model architecture in an online cross-channel interactive parallel distillation architecture pose estimation method. Detailed implementation method

[0051] To enable those skilled in the art to better understand the solution of the present invention, the present invention will be further described in detail below in conjunction with specific implementation manners. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0052] The following gives a best embodiment:

[0053] As Figures 1-6 shown, in an online cross-channel interactive parallel distillation architecture pose estimation method in this embodiment, first, a video acquisition device obtains an external video stream, and the video stream is split into frames and input into a feature extraction network for feature extraction;

[0054] The extracted features are sent to the YOLOV5 object detection model to detect the position of the target human body in each frame of the image and mark the detection frame, obtaining the feature data of the target human body;

[0055] The target human body feature data is passed to the pose detection model Faster-Pose to obtain the human key point feature information; the obtained human key point feature information is mapped to the feature map through a linear transformation to obtain a feature map with human key point annotations.

[0056] Among them, the video acquisition device in this embodiment is a 4-channel 2D camera, and the video stream is split into frames at a rate of 60 frames per second and input into the feature extraction network for feature extraction.

[0057] The feature extraction network draws on the CSPNet network to design a CSP structure and introduces a cross-channel interactive attention mechanism. The cross-channel interactive attention mechanism combines channel attention and spatial attention, calculates the similarity between every two channels of the feature map in the channel attention model using the covariance matrix, and fuses the channels with high similarity. In spatial attention, the second-order finite difference method is used to calculate the pixel value difference and pixel gradient direction of the feature image.

[0058] If the calculated covariance value of two channels is negative, it indicates negative correlation; if the value is 0, it means the two channels are independent and uncorrelated; if the value is positive, it means the two channels are positively correlated for feature fusion.

[0059] First, calculate the mean of each channel as shown in formula (1):

[0060]

[0061] The mean feature of all channels is denoted as

[0062] Calculate the variance of each channel as shown in formula (2):

[0063]

[0064] The variance of all channels is denoted as: Calculate the covariance between channels C1 and C2 as shown in formula (3):

[0065]

[0066] All the obtained covariance values of channel similarities are denoted as Cov i,k , only calculate the covariance between different channels, and fuse the positively correlated channels pixel by pixel according to the covariance value.

[0067] The extracted features are fed into the YOLOV5 object detection model to detect the location of the target human body in each frame of the image and mark the detection box, obtaining the feature data of the target human body. The YOLOV5 object detection model is trained using the publicly available dataset MSCOCO 2017. The MSCOCO 2017 dataset is randomly sampled according to a preset ratio, and data augmentation preprocessing operations are performed on the sample data.

[0068] The channel attention model uses the SoftMax function to obtain the channel feature probability matrix, and the spatial attention model uses the SoftMax function to obtain the spatial feature probability matrix. The probability matrix and the original feature map are fused by multiplication respectively to add weight information to the feature map.

[0069] The Depth-Wise method is used to extract features for each channel of the feature map, obtaining the eigenvalue matrix for each channel. The probability matrix and the original feature map are fused by multiplication respectively to add weight information to the feature map.

[0070] The data augmentation methods include rotating the image at multiple angles with a rotation angle interval of 30 degrees; randomly masking the image according to the probability P, and setting the pixel values under the mask to 0; flipping the image up, down, left, and right; performing distortion processing on the image to different degrees; and performing color perturbation on the image.

[0071] The Faster-Pose pose detection model improves the existing FastPose pose detection model and proposes a new distillation method - the Online Parallel Distillation method. The Online Parallel Distillation method continues to use the Teacher-Student knowledge distillation framework in the network structure. The Teacher network consists of 8 Hourglass feature extraction modules, and the Student network consists of 4 Hourglass feature extraction modules. The Teacher network is trained using the MSCOCO 2017 dataset, and the Student network is trained using a part of the labeled dataset. During the training process, the KL (Kullback-Leibler Divergence) divergence is used to calculate the loss between the feature maps of the Teacher network and the Student network, and the Teacher feature map information and the Student feature map information are fused according to the channel similarity. The Teacher and Student networks are trained in parallel during the training process. During the inference process, the Teacher network is removed and the Student network is directly inferred. A cross-channel interactive attention mechanism is introduced in the Faster-Pose pose detection model, and the cross-channel interactive attention mechanism assigns different weight information to the feature maps in the Teacher network. The calculation processes of the Teacher network feature map and the Student network feature map are shown in formula (4):

[0072]

[0073] where represent the feature maps extracted by the second Hourglass module of the Teacher network and the first Hourglass module of the Student network respectively.

[0074] The total feature map loss is shown in formula (5):

[0075]

[0076] The final loss function of the Faster-Pose pose estimation model is shown in Equation (6):

[0077]

[0078] Where is the loss of the Student network model, and α and λ are hyperparameters to be learned.

[0079] The human key point Heat Map data information output by the Faster-Pose pose detection model is mapped to the original feature map using linear interpolation, and the pixel point offset that occurs during the mapping process is corrected using trilinear interpolation.

[0080] Based on the above method, the pose estimation device of the online cross-channel interactive parallel distillation architecture in this embodiment includes: at least one memory and at least one processor;

[0081] The at least one memory is used to store machine-readable programs;

[0082] The at least one processor is used to call the machine-readable program and execute an online cross-channel interactive parallel distillation architecture pose estimation method.

[0083] Among them, the memory in this embodiment is 512GB, the processor selects an 8-core CPU processor, and the device also requires an NVIDIA graphics card with a model number of RTX2080TI or above.

[0084] The present invention fully considers the connection between channel attention and spatial attention, fuses the two features according to the designed channel fusion rule, and the advantage is that it can not only determine the distribution of features through the channel attention model, but also combine the position information of the target features in the spatial dimension, which can further enhance the representation ability of the target features in the spatial dimension.

[0085] Fully considering that when the number of channels of the extracted feature map is small but the spatial features are large, it is easy to cause insufficient generalization of channel features and difficulty in learning sensitive spatial features. Using the second-order finite difference method to calculate the pixel value difference and pixel gradient direction of the feature map on the spatial attention model can improve the positioning performance of the target position in the spatial dimension.

[0086] The original FastPose pose estimation model is improved, and a new pose estimation model Faster-Pose is proposed. A new information interaction and fusion method between the Teacher network and the Student network is designed, and a new loss function is proposed.

[0087] The above specific embodiments are only specific cases of the present invention. The patent protection scope of the present invention includes but is not limited to the above specific embodiments. Any method and device for attitude estimation of an online cross-channel interactive parallel distillation architecture that conforms to the claims of the present invention and any appropriate changes or substitutions made by those of ordinary skill in the relevant technical field shall fall within the patent protection scope of the present invention.

[0088] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An online cross-channel interactive parallel distillation architecture pose estimation method, characterized in that First, the video acquisition device obtains the external video stream, slices the video stream into frames and inputs them into the feature extraction network for feature extraction; The extracted features are fed into the YOLOV5 object detection model to detect the position of the target human body in each frame of the image and mark the detection box, obtaining the feature data of the target human body; The target human body feature data is passed to the pose detection model Faster-Pose to obtain the human body key point feature information; the obtained human body key point feature information is mapped into the feature map through linear transformation to obtain the feature map with human body key point annotations; The feature extraction network is designed as a CSP structure, and a cross-channel interactive attention mechanism is introduced. The cross-channel interactive attention mechanism combines channel attention and spatial attention, and uses the covariance matrix to calculate the similarity between every two channels of the feature map in the channel attention model, and the channels with high similarity are fused; In spatial attention, the second-order finite difference method is used to calculate the pixel value difference and pixel gradient direction of the feature image; If the covariance calculation value of the two channels is negative, it means negative correlation. If the value is 0, it means that the two channels are independent and irrelevant to each other. If the value is positive, it means that the two channels are positively correlated for feature fusion; First, calculate the mean of each channel as shown in formula (1): The mean feature of all channels is denoted as Calculate the variance of each channel as shown in formula (2): The variance of all channels is denoted as: Calculate the covariance between channels C1 and C2 as shown in formula (3): All the obtained channel similarity covariance values by analogy are denoted as Cov i,k , and only the covariance is calculated between different channels. The channels with positive correlation are fused pixel by pixel according to the covariance values; The channel attention model uses the SoftMax function to obtain the channel feature probability matrix, and the spatial attention model uses the SoftMax function to obtain the spatial feature probability matrix; The probability matrix and the original feature map are fused by multiplication respectively to add weight information to the feature map.

2. An online cross-channel interactive parallel distillation architecture pose estimation method according to claim 1, characterized in that, The extracted features are fed into the YOLOV5 object detection model to detect the position of the target human body in each frame of the image and mark the detection box, obtaining the feature data of the target human body; The YOLOV5 object detection model is trained using the publicly available dataset MSCOCO 2017. The MSCOCO 2017 dataset is randomly sampled according to a pre-set ratio, and data augmentation preprocessing operations are performed on the sample data.

3. An online cross-channel interactive parallel distillation architecture pose estimation method according to claim 2, characterized in that, The data augmentation methods include rotating the image at multiple angles, with the rotation angle division interval being 30 degrees, randomly masking the image according to the probability P, setting the pixel values under the mask to 0, flipping the image up, down, left, and right, performing different degrees of distortion processing on the image, and performing color perturbation on the image.

4. An online cross-channel interactive parallel distillation architecture attitude estimation method according to claim 3, characterized in that, The Depth-Wise method is used to extract features from each channel of the feature map to obtain the feature value matrix of each channel, and cross-channel feature fusion is performed.

5. An online cross-channel interactive parallel distillation architecture pose estimation method according to claim 4, characterized in that, In the pose detection model Faster-Pose, the online parallel knowledge distillation method is carried out. The online parallel knowledge distillation method continues to use the teacher-student knowledge distillation framework in the network structure. The Teacher network consists of 8 Hourglass feature extraction modules, and the Student network consists of 4 Hourglass feature extraction modules; The Teacher network is trained using the MSCOCO 2017 dataset, and the Student network is trained using a part of the labeled dataset. During the training process, the KL divergence is used to calculate the loss between the feature maps of the Teacher network and the Student network, and the information of the Teacher feature maps and the Student feature maps is fused according to the channel similarity. The Teacher and Student networks are trained in parallel during the training process; During the inference process, the Teacher network is removed and only the Student network is used for inference. A cross-channel interactive attention mechanism is introduced into the Faster-Pose pose detection model. The cross-channel interactive attention mechanism assigns different weight information to the feature maps in the Teacher network. The calculation processes of the Teacher network feature maps and the Student network feature maps are shown in Equation (4): Among them, respectively represent the feature maps extracted by the second Hourglass module of the Teacher network and the first Hourglass module of the Student network; The total feature map loss is shown in Equation (5): The final loss function of the Faster-Pose pose estimation model is shown in Equation (6): where is the loss of the Student network model, and α and λ are hyperparameters to be learned.

6. An online cross-channel interactive parallel distillation architecture attitude estimation method according to claim 5, characterized in that The human key point Heat Map data information output by the Faster-Pose pose detection model is mapped to the original feature map using linear interpolation, and the pixel point offset that occurs during the mapping process is corrected using trilinear interpolation.

7. An online cross-channel interactive parallel distillation architecture pose estimation device, characterized in that, Including: At least one memory and at least one processor; The at least one memory is used to store machine-readable programs; The at least one processor is used to call the machine-readable program and execute the method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-mode small target detection method based on knowledge distillation

    CN113449680A

  • Real-time robust two-stage attitude estimation method

    CN114842389A