Visual motion capture method based on inertial motion capture information knowledge distillation

By transferring IMU data information to the camera's human posture prediction model, the problem of identifying human joints in complex and occlusion environments is solved, and the accurate identification and generalization ability in occlusion environments is improved.

CN120108039APending Publication Date: 2025-06-06SOUTHEAST UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510256774.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-05
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The existing two-dimensional human motion capture technology has limitations when dealing with complex and occlusion environments, especially because single-camera input is difficult to accurately identify human joint nodes in multi-person scenes and occlusion situations.

Method used

The visual motion capture method of inertial motion capture information knowledge distillation is adopted. By co-calibrating the camera and IMU and designing a multimodal fusion knowledge distillation network, the IMU data information is transferred to the camera's human posture prediction model, and the network's recognition ability in the occlusion environment is improved.

Benefits of technology

It realizes accurate output of human joint node coordinates in an occlusion environment, improves the generalization ability and recognition accuracy of the motion capture network, reduces dependence on IMU, and reduces costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108039A_ABST
    Figure CN120108039A_ABST
Patent Text Reader

Abstract

The invention discloses a visual motion capture method combined with inertia motion capture information distillation. The method provides an inertia motion capture information knowledge distillation network. In the training stage, the camera and the IMU collect data at the same time as input, and after the camera obtains a video image, the probability value that each pixel has an articulation point is detected through a feature extraction network, and a human body articulation point thermodynamic diagram is generated; the IMU is initialized through attitude analysis, records an initial position, performs numerical integration by using the measured acceleration, calculates joint point coordinates of a human body in an image, and maps the joint point coordinates to a thermodynamic diagram. And inputting into a knowledge distillation network together. The knowledge distillation network performs weighted fusion on the thermodynamic diagrams of the camera and the IMU, adds the thermodynamic diagrams of the camera and the IMU element by element, and obtains joint point fusion prediction distribution through a classifier; meanwhile, distribution of the thermodynamic diagram of the camera joint points is predicted through a classifier, and training is supervised through a KL divergence loss function. During verification, only camera images need to be input, and the network can accurately output human body joint point coordinates.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for capturing two-dimensional human motion, and in particular to a visual motion capture method based on inertial motion capture information knowledge distillation, and belongs to the field of computer vision. Background Art

[0002] Two-dimensional human motion capture is an important topic in the field of computer vision. It is of great significance for understanding human behavior and promoting the application of artificial intelligence. It can accurately identify the positions of multiple human bodies and the positions of sparse key points on their skeletons from a single RGB image. It is widely used in motion recognition, animation generation, augmented reality and other fields.

[0003] In the early days of this field, single camera 2D human motion capture mainly focused on regression methods, which distinguish human motion by directly regressing the coordinates of joint points. With the development of machine learning, researchers have designed machine learning algorithms and feature extraction models such as support vector machines, random forests, etc., as well as methods based on feature point matching, but single camera input has limitations in dealing with some complex and occluded situations. Subsequently, researchers used inertial measurement units (IMUs) to complete motion capture, fixed IMUs on the main parts of the human body, and transmitted sensor posture data to the data processing system through wireless transmission methods such as Bluetooth, WIFI, and ZIGBEE to perform full-body posture solution. Its main advantages are easy to carry, simple operation, almost no site restrictions, no occlusion, high accuracy and sampling speed. Its disadvantage is that IMUs are prone to accumulate errors, and when facing multi-person scenes, the cost of using a large number of IMUs is extremely high.

[0004] The camera can provide rich appearance information and global positioning information of the human body, and the IMU can still accurately provide coordinate information when the human body blocks each other. Therefore, it provides a new idea for 2D human motion capture. This method proposes an inertial motion capture information knowledge distillation network. During the training process, the camera and IMU collect data at the same time as input to train the network. During the verification process, the network only needs to input the camera image network to accurately output the coordinates of the human body's joints. Summary of the invention

[0005] Since the camera and IMU can provide complementary information, the present invention proposes a method for IMU-assisted camera motion capture, constructs a motion capture network, and integrates vision and IMU to accurately estimate two-dimensional human posture.

[0006] The technical solution adopted by the present invention is as follows:

[0007] A visual motion capture method combined with inertial motion capture information distillation. This method first calibrates the camera and IMU together to ensure the accuracy and synchronization of the data of the two. A multimodal fusion knowledge distillation network is designed to transfer the IMU data information in the occluded environment to the camera human body posture prediction model. During the training process, the camera and IMU collect data at the same time as the input training network. During the verification process, the network only needs to input the camera image network to accurately output the coordinates of the human body's joints. The specific steps include:

[0008] Step 1: Get the camera image heat map

[0009] The RGB image is input into the feature extraction network for feature extraction. The network input size set in this method is 256×256. Then the output feature map is input into the human joint point detection network. The network contains 18 input channels. The feature map output by the feature extraction network is input as the input channel. Finally, the two-dimensional joint point heat map of the human body is output. The network is composed of two camera image joint point recognition modules connected in series.

[0010] Define the camera image joint point recognition module:

[0011] The camera image is extracted through the feature extraction network to obtain a feature F. The network is divided into two channels: one channel is used to predict the confidence map S: to detect all the joints of the human body; the other channel is used to predict the affinity L between the joints: to determine whether multiple joints are the same target. The first loop channel takes the feature F as input and obtains a set of S 1 , L 1 ; After entering the second loop channel, the output S of the first loop channel is used 1 , L 1 and feature F as input, and finally output the target two-dimensional joint point heat map. The specific formula is as follows:

[0012] S 2 =ρ 2 (F,S 1 ,L 1 )

[0013] L 2 =φ 2 (F,S 1 ,L 1 )

[0014] Where F is the feature map extracted by the feature extraction network and is the input included in both loop channels.

[0015] ρ is the module function for obtaining the joint confidence map, which is used to detect all the joints in the camera image, including five convolution layers and upsampling operations. First, it passes through three convolution layers with a kernel size of 3×3 and a moving step of 1, and then the feature map is restored to the size of the original image through an upsampling operation, and then passes through two convolution layers with a kernel size of 1x1 and a moving step of 1. The convolution layer can generate multi-channel outputs, and each output channel corresponds to a heat map of a joint point.

[0016] After outputting the heat map of each joint point, in order to find the specific location of the joint point, the following formula can be used to find the maximum value point of the heat map:

[0017]

[0018] in is the predicted joint point coordinate, and G(x,y) is the heat map function value.

[0019] It is the module function for judging the affinity between joints. After ρ detects all joints, To determine which joint points belong to the same person, including the convolution layer and the matching module, first pass through three layers of convolution layers with a kernel size of 3×3 and a moving step size of 1, then restore the feature map to the size of the original image through an upsampling operation, and then pass through two layers of convolution layers with a kernel size of 1x1 and a moving step size of 1 to obtain the joint point features.

[0020] After obtaining the positions of all relevant nodes, we use joint point matching to determine whether the joint points belong to the same person. We use the bipartite graph matching method, where V is the joint point set and E is the edge set. The joint point set V can be divided into two non-intersecting subsets V 1 and V 2 , that is, V = V 1 ∪V 2 and Each edge e∈E connects a V 1 The joint point and a V 2 A matching is a subset of edges such that for any two different edges e 1 ,e 2 ∈M, they have no common endpoints. That is, for any e 1 =(u 1 ,v 1 ) and e 2 =(u 2 ,v 2 ), both have u 1 ≠u 2 And v 1 ≠v2 .

[0021] The adjacency matrix A is used to represent the complete set of matches, including all joint points; the matching situation can be represented by the matrix X, where if there is a joint point match, then X ij =1, otherwise X ij =0; then the maximum matching can be expressed as:

[0022]

[0023] Where i and j are the i-th joint point and the j-th joint point respectively; max() means finding the maximum value; by solving this integer linear programming problem, the output of each person's joint point is completed.

[0024] S t , L t By S t and f L t Loss function to supervise:

[0025]

[0026] in It is a real body part position map. is the true part affinity; t represents the tth recurrent channel; W is a binary mask used to avoid penalizing the correct predictions during training; p represents a position on the image, and W(p) = 0 when the annotation at image position p is missing.

[0027] The total loss function is the sum of all module loss functions:

[0028]

[0029] Step 2: IMU data processing

[0030] Sub-step 1: Posture parsing initialization

[0031] The wearer stands still at the initial moment, with his hands raised parallel to the ground in a "T" shape. In this case, the initial quaternion of the sensor state in the geographic coordinate system is calculated:

[0032]

[0033] Where θ, φ and ψ are roll angle (roll), pitch angle (pitch), and yaw angle (yaw) respectively. There is no displacement between the inertial sensor and the human body, so the coordinate relationship between the inertial sensor and the human body will not change. Therefore:

[0034]

[0035] Initialize the human body through IMU and record the conversion relationship between all IMU and human body coordinates Sensor A on the human body i Relative to human body A 0 The expression of the relationship is:

[0036]

[0037] Sub-step 2: IMU obtains joint point coordinates

[0038] The acceleration measured by the calibrated accelerometer is a = [a x ,a y ,a z ], numerically integrate a to get the velocity v:

[0039]

[0040] Then integrate the velocity to get the displacement L:

[0041]

[0042] Through the above calculations, the 3D coordinates (X, Y, Z) in the world coordinate system can be gradually obtained from the initial position and the position and posture of the object can be updated at each moment. The following converts the coordinates into 2D coordinates in the camera image:

[0043] The rotation matrix R is calculated from yaw, pitch, and roll, and is calculated using the following formula:

[0044] R=R z (yaw)·R y (pitch) R x (roll)

[0045] The conversion of world coordinate system coordinates to camera coordinate system coordinates still requires translation vector T. The conversion formula is:

[0046]

[0047] The camera internal parameters obtained by camera calibration include: the focal length f of the camera x and f y , and the center point of the camera imaging plane cx, cy, are all in pixels. c, Y c, Z c ] is converted to 2D coordinates [x 2D, y 2D ]:

[0048]

[0049] Sub-step 3: IMU data mapping

[0050] The IMU joint points are solved and processed to obtain the coordinates of the human joint points in the two-dimensional image. The camera image is passed through the feature extraction network to obtain the joint point heat map. In order to better distill the IMU data knowledge into the camera network parameter training, the IMU data needs to be visualized and converted into a heat map display, where the position of each joint point is represented in the form of a Gaussian distribution in the heat map. First, create a heat map of the same size as the input image, initialize it to an all-zero matrix, and import the camera image. Then draw the heat map. For each point (x, y), determine where it belongs in the grid unit (m x , m y ), can be calculated by the following formula:

[0051]

[0052] Where Δx and Δy are the grid widths along the X-axis and Y-axis respectively. Indicates rounding down.

[0053] For each joint point (x, y), its position on the heat map (m x , m y ) generates a Gaussian distribution on the heat map, calculated by the following formula:

[0054]

[0055] Where H(m x ,m y ) represents the value at the joint point (m x ,m y ), σ is the standard deviation of the Gaussian distribution, which controls the diffusion degree of the Gaussian distribution in the heat map. Finally, the same color mapping as the camera image heat map is used to obtain a heat map with consistent color mapping.

[0056] Step 3: Distill IMU data knowledge into visual network

[0057] The key to this method is to use the auxiliary information of IMU measurement data to improve the performance of camera motion capture in occluded environments and improve the generalization ability of the network. The image captured by the camera is output as a heat map of human joints through the feature extraction network. After IMU correction, the two-dimensional joint coordinates of the human body are obtained through joint coordinate solution. The two are mapped to the same feature space through the multimodal fusion knowledge distillation network proposed by this method. The features of the camera image heat map and IMU data are fused and aligned in one direction, retaining the complete information from the multimodal data. The output joints will reversely correct the parameters of the visual capture network, so that the camera motion capture network can infer the coordinates of the occluded joints in an occluded environment, and has good generalization ability.

[0058] The heat map fusion operation of this method uses the weighted average method, and the heat map obtained from the IMU data is recorded as The heat map obtained from the camera image is denoted as Heat map after fusion It can be expressed by the following formula:

[0059]

[0060] Among them, α is a weight parameter used to balance the influence of the two heat maps. In this method, α = 0.7. After the heat map obtained by converting the IMU data is fused with the heat map extracted by the camera, it is input into the MLP layer again, which contains two hidden layers. Each hidden layer contains multiple neurons, and the neurons are fully connected. The output can be specifically expressed by the formula:

[0061] y=σ(W 3 l 2 +b 3 )

[0062] Where W 3 , b 3 is the weight matrix and bias parameter of the output layer, l 2 is the output of the second hidden layer, which can be expressed as:

[0063] l 2 =σ(W 2 l 1 +b 2 )

[0064] Where W 2 , b 2 is the weight matrix and bias parameter of the second hidden layer, l 1 is the output of the first hidden layer, which can be expressed as:

[0065] l 1 =σ(W 1 x+b 1 )

[0066] Where W 1 , b 1 is the weight matrix and bias parameter of the first hidden layer, and x is the feature output of the fusion heat map.

[0067] The activation function used in this method is the sigmoid function, and its formula is as follows:

[0068]

[0069] As a nonlinear activation function, it allows the network to learn nonlinear relationships between data, thereby better integrating multimodal information.

[0070] If occlusion occurs when predicting human joints, the camera image feature extraction network may identify the wrong joints, so it is necessary to use IMU data to correct the camera's prediction, multiply the heat map after one feature extraction fusion with the heat map after two feature extractions element by element, and weight the image to highlight or suppress specific features to achieve the function of image filtering.

[0071] Element-by-element addition adds two heat maps pixel by pixel, adds different feature maps to retain more information, and enhances the feature expression ability of the network to create a new image or enhance feature map. Add the fused heat map and the heat map converted by IMU element by element to improve the feature expression of the heat map as the output result.

[0072] The knowledge distillation process is as follows: The multimodal fusion knowledge distillation module first converts the IMU data F imu After the joint point coordinates are solved, the coordinates are visualized to obtain the joint point heat map The camera is obtained through the feature extraction network and The same size, and After the fusion operation, we get After the MLP layer, a new channel is introduced and passed through the MLP layer twice. After the activation function, the two are multiplied element by element and added After adding each element together, we get The predicted probability distribution of joint points is obtained through the classifier The predicted probability distribution of the camera's joint points is also obtained through the classifier Finally, KL divergence is used as the loss function of the distillation module, so that the camera network has the ability to infer joint points under occlusion.

[0073] It can be a feature after the fusion of camera and IMU, which can be expressed by the following formula:

[0074]

[0075] Step 4: Train the knowledge distillation network parameters

[0076] The SGD optimizer was used during training, the learning rate was set to 0.24, and 64 rounds were performed.

[0077] KL divergence is used as the loss function of knowledge distillation, as follows:

[0078]

[0079] Where D KL is the mapping relationship symbol of the KL divergence function, The value represents and The smaller the function value, the closer the two probability distributions are.

[0080] The KL divergence loss function represents the difference between the model's predicted distribution and the true distribution, thereby improving the model's accuracy and generalization ability. The parameters of the monocular camera motion capture network are trained based on the motion capture after the camera and IMU are fused, so that the network has good generalization ability.

[0081] After the network training is completed, you only need to input the camera image and the network can accurately output the coordinates of the joint points of the human body.

[0082] Advantages and significant effects of this method:

[0083] This method designs a multimodal fusion knowledge distillation network to transfer the IMU data information in the occluded environment to the camera human posture prediction model. During the training process, the camera and IMU simultaneously collect data as input to the training network. During the verification process, the network only needs to input the camera image network to accurately output the coordinates of the human body's joints. BRIEF DESCRIPTION OF THE DRAWINGS

[0084] Figure 1 It is a flow chart of the visual motion capture method combined with inertial motion capture information distillation;

[0085] Figure 2 This is the network architecture diagram of inertial measurement unit-assisted visual motion capture;

[0086] Figure 3 This is the visual feature extraction network architecture diagram;

[0087] Figure 4 It is a schematic diagram of the multimodal fusion knowledge distillation module. DETAILED DESCRIPTION

[0088] Two-dimensional human motion capture is an important topic in the field of computer vision. It is of great significance for understanding human behavior and promoting the application of artificial intelligence. It can accurately identify the positions of multiple human bodies and the positions of sparse key points on their skeletons from a single RGB image. It is widely used in motion recognition, animation generation, augmented reality and other fields.

[0089] In the early days of this field, single camera 2D human motion capture mainly focused on regression methods, which distinguish human motion by directly regressing the coordinates of joint points. With the development of machine learning, researchers have designed machine learning algorithms and feature extraction models such as support vector machines, random forests, etc., as well as methods based on feature point matching, but single camera input has limitations in dealing with some complex and occluded situations. Subsequently, researchers used inertial measurement units (IMUs) to complete motion capture, fixed IMUs on the main parts of the human body, and transmitted sensor posture data to the data processing system through wireless transmission methods such as Bluetooth, WIFI, and ZIGBEE to perform full-body posture solution. Its main advantages are easy to carry, simple operation, almost no site restrictions, no occlusion, high accuracy and sampling speed. Its disadvantage is that IMUs are prone to accumulate errors, and when facing multi-person scenes, the cost of using a large number of IMUs is extremely high.

[0090] The camera can provide rich appearance information and global positioning information of the human body, and the IMU can still accurately provide coordinate information when the human body occludes each other. Therefore, it provides a new idea for 2D human motion capture. This method uses a small number of IMUs to assist monocular cameras for motion capture. The IMU mutually enhances the features of adjacent key points and improves the detection performance of occluded key points.

[0091] Since the camera and IMU can provide complementary information, the present invention proposes a method for IMU-assisted camera motion capture, constructs a motion capture network, and integrates vision and IMU to accurately estimate two-dimensional human posture.

[0092] The present invention provides a visual motion capture method combined with inertial motion capture information distillation. The camera and IMU are first calibrated together to ensure the accuracy and synchronization of the data of the two. A multimodal fusion knowledge distillation network is designed to transfer the IMU data information in an occluded environment to the camera human posture prediction model. When training the network, a small amount of information provided by the IMU is used as auxiliary information to assist the camera in predicting human movements, so that the network has better generalization ability. In the verification process, only using the camera to capture human movements can achieve accurate recognition effects. The flow chart of this method is shown in the figure. Figure 1 , Figure 2 As shown, the specific steps include:

[0093] Step 1: Get the camera image heat map

[0094] The RGB image is input into the feature extraction network for feature extraction. The network input size set in this method is 256×256. Then the output feature map is input into the human joint point detection network. The network contains 18 input channels. The feature map output by the feature extraction network is used as the input channel. Finally, the two-dimensional joint point heat map of the human body is output. The network is composed of two camera image joint point recognition modules connected in series, such as Figure 3 shown.

[0095] Define the camera image joint point recognition module:

[0096] The camera image is extracted through the feature extraction network to obtain a feature F. The network is divided into two channels: one channel is used to predict the confidence map S: to detect all the joints of the human body; the other channel is used to predict the affinity L between the joints: to determine whether multiple joints are the same target. The first loop channel takes the feature F as input and obtains a set of S 1 , L 1 ; After entering the second loop channel, the output S of the first loop channel is 1 , L 1 and feature F as input, and finally output the target two-dimensional joint point heat map. The specific formula is as follows:

[0097] S 2 =ρ 2 (F,S 1 ,L 1 )

[0098] L 2 =φ 2 (F,S 1 ,L 1 )

[0099] Where F is the feature map extracted by the feature extraction network and is the input included in both loop channels.

[0100] ρ is the module function for obtaining the joint confidence map, which is used to detect all the joints in the camera image, including five convolution layers and upsampling operations. First, it passes through three convolution layers with a kernel size of 3×3 and a moving step of 1, and then the feature map is restored to the size of the original image through an upsampling operation, and then passes through two convolution layers with a kernel size of 1x1 and a moving step of 1. The convolution layer can generate multi-channel outputs, and each output channel corresponds to a heat map of a joint point.

[0101] After outputting the heat map of each joint point, in order to find the specific location of the joint point, the following formula can be used to find the maximum value point of the heat map:

[0102]

[0103] in is the predicted joint point coordinate, and G(x,y) is the heat map function value.

[0104] It is the module function for judging the affinity between joints. After ρ detects all joints, To determine which joint points belong to the same person, including the convolution layer and the matching module, first pass through three layers of convolution layers with a kernel size of 3×3 and a moving step size of 1, then restore the feature map to the size of the original image through an upsampling operation, and then pass through two layers of convolution layers with a kernel size of 1x1 and a moving step size of 1 to obtain the joint point features.

[0105] After obtaining the positions of all relevant nodes, we use joint point matching to determine whether the joint points belong to the same person. We use the bipartite graph matching method, where V is the joint point set and E is the edge set. The joint point set V can be divided into two non-intersecting subsets V 1 and V 2 , that is, V = V 1 ∪V 2 and Each edge e∈E connects a V 1 The joint point and a V 2 A matching is a subset of edges such that for any two different edges e 1 ,e 2 ∈M, they have no common endpoints. That is, for any e 1 =(u 1 ,v 1 ) and e 2 =(u 2 ,v 2 ), both have u 1 ≠u 2 And v 1 ≠v 2 .

[0106] The adjacency matrix A is used to represent the complete set of matches, including all joint points; the matching situation can be represented by the matrix X, where if there is a joint point match, then X ij =1, otherwise X ij =0; then the maximum matching can be expressed as:

[0107]

[0108] Where i and j are the i-th joint point and the j-th joint point respectively; max() means finding the maximum value. By solving this integer linear programming problem, the output of each person's joint point is completed.

[0109] S t , L t By S t and f L t Loss function to supervise:

[0110]

[0111] in It is a real body part position map. is the true part affinity; t represents the tth recurrent channel; W is a binary mask used to avoid penalizing the correct predictions during training; p represents a position on the image, and W(p) = 0 when the annotation at image position p is missing.

[0112] The total loss function is the sum of all module loss functions:

[0113]

[0114] Step 2: IMU data processing

[0115] Sub-step 1: Posture parsing initialization

[0116] The wearer stands still at the initial moment, with his hands raised parallel to the ground in a "T" shape. In this case, the initial quaternion of the sensor state in the geographic coordinate system is calculated:

[0117]

[0118] Where θ, φ and ψ are roll angle (roll), pitch angle (pitch), and yaw angle (yaw) respectively. There is no displacement between the inertial sensor and the human body, so the coordinate relationship between the inertial sensor and the human body will not change. Therefore:

[0119]

[0120] Initialize the human body through IMU and record the conversion relationship between all IMU and human body coordinates Sensor A on the human body i Relative to human body A 0 The expression of the relationship is:

[0121]

[0122] Sub-step 2: IMU obtains joint point coordinates

[0123] The acceleration measured by the calibrated accelerometer is a = [a x ,a y ,az ], numerically integrate a to get the velocity v:

[0124]

[0125] Then integrate the velocity to get the displacement L:

[0126]

[0127] Through the above calculations, the 3D coordinates (X, Y, Z) in the world coordinate system can be gradually obtained from the initial position and the position and posture of the object can be updated at each moment. The following converts the coordinates into 2D coordinates in the camera image:

[0128] The rotation matrix R is calculated from yaw, pitch, and roll, and is calculated using the following formula:

[0129] R=R z (yaw)·R y (pitch) R x (roll)

[0130] The conversion of world coordinate system coordinates to camera coordinate system coordinates still requires translation vector T. The conversion formula is:

[0131]

[0132] The camera internal parameters obtained by camera calibration include: the focal length f of the camera x and f y , and the center point of the camera imaging plane cx, cy, are all in pixels. c, Y c, Z c ] is converted to 2D coordinates [x 2D, y 2D ]:

[0133]

[0134] Sub-step 3: IMU data mapping

[0135] The IMU joint points are solved and processed to obtain the coordinates of the human joint points in the two-dimensional image. The camera image is passed through the feature extraction network to obtain the joint point heat map. In order to better distill the IMU data knowledge into the camera network parameter training, the IMU data needs to be visualized and converted into a heat map display, where the position of each joint point is represented in the form of a Gaussian distribution in the heat map. First, create a heat map of the same size as the input image, initialize it to an all-zero matrix, and import the camera image. Then draw the heat map. For each point (x, y), determine where it belongs in the grid unit (m x , m y), can be calculated by the following formula:

[0136]

[0137] Where Δx and Δy are the grid widths along the X-axis and Y-axis respectively. Indicates rounding down.

[0138] For each joint point (x, y), its position on the heat map (m x , m y ) generates a Gaussian distribution on the heat map, calculated by the following formula:

[0139]

[0140] Where H(m x ,m y ) represents the value at the joint point (m x ,m y ), σ is the standard deviation of the Gaussian distribution, which controls the diffusion degree of the Gaussian distribution in the heat map. Finally, the same color mapping as the camera image heat map is used to obtain a heat map with consistent color mapping.

[0141] Step 3: Distill IMU data knowledge into visual network

[0142] The key to this method is to use the auxiliary information of IMU measurement data to improve the performance of camera motion capture in occluded environments and improve the generalization ability of the network. The image captured by the camera is output as a heat map of human joints through the feature extraction network. After IMU correction, the two-dimensional joint coordinates of the human body are obtained through joint coordinate solution. The two are mapped to the same feature space through the multimodal fusion knowledge distillation network proposed by this method, such as Figure 4 As shown in the figure, the features of the camera image heat map and IMU data are fused and aligned unidirectionally, which retains the complete information from the multimodal data. The output joint points will reversely correct the parameters of the visual capture network, so that the camera motion capture network can infer the coordinates of the occluded joint points in an occluded environment, and has good generalization ability.

[0143] The heat map fusion operation of this method uses the weighted average method, and the heat map obtained from the IMU data is recorded as The heat map obtained from the camera image is denoted as Heat map after fusion It can be expressed by the following formula:

[0144]

[0145] Among them, α is a weight parameter used to balance the influence of the two heat maps. In this method, α = 0.7. After the heat map obtained by converting the IMU data is fused with the heat map extracted by the camera, it is input into the MLP layer again, which contains two hidden layers. Each hidden layer contains multiple neurons, and the neurons are fully connected. The output can be specifically expressed by the formula:

[0146] y=σ(W 3 l 2 +b 3 )

[0147] Where W 3 , b 3 is the weight matrix and bias parameter of the output layer, l 2 is the output of the second hidden layer, which can be expressed as:

[0148] l 2 =σ(W 2 l 1 +b 2 )

[0149] Where W 2 , b 2 is the weight matrix and bias parameter of the second hidden layer, l 1 is the output of the first hidden layer, which can be expressed as:

[0150] l 1 =σ(W 1 x+b 1 )

[0151] Where W 1 , b 1 is the weight matrix and bias parameter of the first hidden layer, and x is the feature output of the fusion heat map.

[0152] The activation function used in this method is the sigmoid function, and its formula is as follows:

[0153]

[0154] As a nonlinear activation function, it allows the network to learn nonlinear relationships between data, thereby better integrating multimodal information.

[0155] If occlusion occurs when predicting human joints, the camera image feature extraction network may identify the wrong joints, so it is necessary to use IMU data to correct the camera's prediction, multiply the heat map after one feature extraction fusion with the heat map after two feature extractions element by element, and weight the image to highlight or suppress specific features to achieve the function of image filtering.

[0156] Element-by-element addition adds two heat maps pixel by pixel, adds different feature maps to retain more information, and enhances the feature expression ability of the network to create a new image or enhance feature map. Add the fused heat map and the heat map converted by IMU element by element to improve the feature expression of the heat map as the output result.

[0157] The knowledge distillation process is as follows: The multimodal fusion knowledge distillation module first converts the IMU data F imu After the joint point coordinates are solved, the coordinates are visualized to obtain the joint point heat map The camera is obtained through the feature extraction network and The same size, and After the fusion operation, we get After the MLP layer, a new channel is introduced and passed through the MLP layer twice. After the activation function, the two are multiplied element by element and added After adding each element together, we get The predicted probability distribution of joint points is obtained through the classifier The predicted probability distribution of the camera's joint points is also obtained through the classifier Finally, KL divergence is used as the loss function of the distillation module, so that the camera network has the ability to infer joint points under occlusion.

[0158] It can be a feature after the fusion of camera and IMU, which can be expressed by the following formula:

[0159]

[0160] Step 4: Train the camera network parameters

[0161] The SGD optimizer was used during training, the learning rate was set to 0.24, and 64 rounds were performed.

[0162] KL divergence is used as the loss function of knowledge distillation, as follows:

[0163]

[0164] Where D KL is the mapping relationship symbol of the KL divergence function, The value represents and The smaller the function value, the closer the two probability distributions are.

[0165] The KL divergence loss function represents the difference between the model's predicted distribution and the true distribution, thereby improving the model's accuracy and generalization ability. The parameters of the monocular camera motion capture network are trained based on the motion capture after the camera and IMU are fused, so that the network has good generalization ability.

[0166] After the network training is completed, you only need to input the camera image and the network can accurately output the coordinates of the joint points of the human body.

[0167] This method designs a multimodal fusion knowledge distillation network to transfer the IMU data information in the occluded environment to the camera human posture prediction model. During the training process, the camera and IMU simultaneously collect data as input to the training network. During the verification process, the network only needs to input the camera image network to accurately output the coordinates of the human body's joints.

Claims

1. A visual motion capture method based on knowledge distillation of inertial motion capture information, characterized by: The camera and IMU are jointly calibrated to ensure the accuracy and synchronization of their data. A multimodal fusion knowledge distillation network is designed to transfer IMU data information in occluded environments to the camera human posture prediction model. When training the network, a small amount of information provided by the IMU is used as auxiliary information to assist the camera in predicting human movements, so that the network has better generalization ability. During the verification process, only using the camera to capture human motion can achieve accurate recognition results; the specific steps include: Step 1: Get the camera image heat map The RGB image is input into the feature extraction network for feature extraction. The input size of the feature extraction network is set to 256×256. Then the output feature map is input into the human joint point detection network. The human joint point detection network contains 18 input channels. The feature map output by the feature extraction network is used as the input channel. Finally, the two-dimensional joint point heat map of the human body is output. The network is composed of two camera image joint point recognition modules connected in series. Define the camera image joint point recognition module: The camera image is extracted through the feature extraction network to obtain a feature F. The network is divided into two channels: one channel is used to predict the confidence map S: to detect all the joints of the human body; the other channel is used to predict the affinity L between the joints: to determine whether multiple joints are the same target; the first loop channel takes the feature F as input and obtains a set of S 1 , L 1 ; After entering the second loop channel, the output S of the first loop channel is used 1 , L 1 and feature F as input, and finally output the target two-dimensional joint point heat map; the specific formula is as follows: S 2 =ρ 2 (F,S 1 ,L 1 ) L 2 =φ 2 (F,S 1 ,L 1 ) Where F is the feature map extracted by the feature extraction network, which is the input included in both loop channels; ρ is the module function for obtaining the joint confidence map, which is used to detect all the joints in the camera image, including five convolution layers and upsampling operations. First, it passes through three convolution layers with a convolution kernel size of 3×3 and a moving step size of 1, and then the feature map is restored to the size of the original image through an upsampling operation. Then, it passes through two convolution layers with a convolution kernel size of 1x1 and a moving step size of 1. This convolution layer can generate multi-channel outputs, and each output channel corresponds to a heat map of a joint point. After outputting the heat map of each joint point, in order to find the specific location of the joint point, use the following formula to find the maximum value point of the heat map: in is the predicted joint point coordinate, G(x,y) is the value of the heat map function; It is the module function for judging the affinity between joints. After ρ detects all joints, To determine which joints belong to the same person, including convolutional layers and matching modules, first pass through three layers of convolutional layers with a kernel size of 3×3 and a moving step of 1, then restore the feature map to the size of the original image through upsampling operations, and then pass through two layers of convolutional layers with a kernel size of 1x1 and a moving step of 1 to obtain joint point features; After obtaining the positions of all relevant nodes, joint point matching is used to determine whether the joint points belong to the same person. The bipartite graph matching method is used. Let V be the joint point set and E be the edge set. The joint point set V can be divided into two non-intersecting subsets V1 and V2, that is, V = V1 ∪ V2 and Each edge e∈E connects a node in V1 and a node in V2; a matching is a subset of edges such that for any two different edges e1, e2∈M, they have no common endpoints; that is, for any e1=(u1, v1) and e2=(u2, v2), u1≠u2 and v1≠v2; The adjacency matrix A is used to represent the complete set of matches, including all joint points; the matching situation is represented by the matrix X, where if there is a joint point match, then X ij =1, otherwise X ij =0; then the maximum matching is expressed as: Where i and j are the i-th joint point and the j-th joint point respectively; max() means finding the maximum value; by solving this integer linear programming problem, the output of each person's joint point is completed; S t , L t By S t and f L t Loss function to supervise: in It is a real body part position map. is the true partial affinity; t represents the tth loop channel; W is a binary mask used to avoid penalizing the correct prediction during training; p represents a position on the image, when the annotation at the image position p is missing, W(p) = 0; The total loss function is the sum of all module loss functions: Step 2: IMU data processing Sub-step 1: Posture parsing initialization The wearer stands still at the initial moment, with his hands raised parallel to the ground, forming a "T" shape; in this case, the initial quaternion of the sensor state in the geographic coordinate system is calculated Where θ, φ and ψ are roll angle roll, pitch angle pitch and yaw angle yaw respectively. There is no displacement between the inertial sensor and the human body, so the coordinate relationship between the inertial sensor and the human body will not change. Therefore: Initialize the human body through IMU and record the conversion relationship between all IMU and human body coordinates Sensor A on the human body i The expression of the relationship with respect to the human body A0 is: Sub-step 2: IMU obtains joint point coordinates The acceleration measured by the calibrated accelerometer is a = [a x ,a y ,a z ], numerically integrate a to get the velocity v: Then integrate the velocity to get the displacement L: After the above calculations, the 3D coordinates (X, Y, Z) in the world coordinate system are gradually obtained from the initial position and the position and posture of the object are updated at each moment; the coordinates are converted into 2D coordinates in the camera image: The rotation matrix R is calculated from yaw, pitch, and roll, and is calculated using the following formula: R=R z (yaw)·R y (pitch)·R x (roll) The conversion of world coordinate system coordinates to camera coordinate system coordinates still requires translation vector T. The conversion formula is: The camera internal parameters obtained by camera calibration include: the focal length f of the camera x and f y , and the center point of the camera imaging plane cx, cy, are all in pixels; the coordinates of the camera coordinate system [X c, Y c, Z c ] is converted to 2D coordinates [x 2D, y 2D ]: Sub-step 3: IMU data mapping The IMU joint points are solved and processed to obtain the coordinates of the human joint points in the two-dimensional image. The camera image is passed through the feature extraction network to obtain the joint point heat map. In order to distill the IMU data knowledge into the camera network parameter training, the IMU data needs to be visualized and converted into a heat map display, where the position of each joint point is represented in the form of a Gaussian distribution in the heat map. First, create a heat map of the same size as the input image, initialize it to an all-zero matrix, and import the camera image; then draw the heat map. For each point (x, y), determine its grid unit (m x , m y ), calculated by the following formula: Where Δx and Δy are the grid widths along the X-axis and Y-axis respectively. Indicates rounding down; For each joint point (x, y), its position on the heat map (m x , m y ) generates a Gaussian distribution on the heat map, calculated by the following formula: Where H(m x ,m y ) represents the value at the joint point (m x ,m y ) point, σ is the standard deviation of the Gaussian distribution, which controls the diffusion degree of the Gaussian distribution in the heat map; finally, the same color mapping as the camera image heat map is used to obtain a heat map with consistent color mapping; Step 3: Distill IMU data knowledge into visual network The auxiliary information of IMU measurement data is used to improve the performance of camera motion capture in occluded environments and enhance the generalization ability of the network. The image captured by the camera is output as a heat map of human joints through the feature extraction network. After IMU correction, the two-dimensional coordinates of human joints are obtained through joint coordinate solution. The two are mapped to the same feature space and passed through a multimodal fusion knowledge distillation network to fuse the features of the camera image heat map and IMU data and align them unidirectionally, retaining the complete information from the multimodal data. The output joints will reversely correct the parameters of the visual capture network, so that the camera motion capture network can infer the coordinates of occluded joints in an occluded environment, and has good generalization ability. Among them, the weighted average method is used in the heat map fusion operation, and the heat map obtained from the IMU data is recorded as The heat map obtained from the camera image is denoted as Heat map after fusion It is expressed by the following formula: Where α is a weight parameter used to balance the influence of the two heat maps, α = 0.7; After the IMU data is converted into a thermal map and fused with the thermal map extracted by the camera, it is input into the MLP layer again, which contains two hidden layers. Each hidden layer contains multiple neurons, and the neurons are fully connected. The output is specifically expressed by the formula: y=σ(W3l2+b3) Where W3 and b3 are the weight matrix and bias parameter of the output layer, and l2 is the output of the second hidden layer, which can be expressed as: l2=σ(W2l1+b2) Where W2 and b2 are the weight matrix and bias parameter of the second hidden layer, and l1 is the output of the first hidden layer, which can be expressed as: l1=σ(W1x+b1) Where W1 and b1 are the weight matrix and bias parameters of the first hidden layer, and x is the feature output of the fusion heat map; The activation function used in this method is the sigmoid function, and its formula is as follows: As a nonlinear activation function, it allows the network to learn nonlinear relationships between data, making it better able to integrate multimodal information. If occlusion occurs when predicting human joints, the camera image feature extraction network may identify the wrong joints, so the camera prediction needs to be corrected through IMU data. The heat map after one feature extraction fusion is multiplied element by element with the heat map after two feature extractions, and the image is weighted to highlight or suppress specific features to achieve the image filtering function. Element-by-element addition adds two heat maps pixel by pixel, adds different feature maps to retain more information, and enhances the feature expression ability of the network to create a new image or enhance the feature map; adds the fused heat map and the heat map converted by IMU element by element to improve the feature expression of the heat map as the output result; The knowledge distillation process is as follows: The multimodal fusion knowledge distillation module first converts the IMU data F imu After the joint point coordinates are solved, the coordinates are visualized to obtain the joint point heat map The camera is obtained through the feature extraction network and The same size, and After the fusion operation, we get After the MLP layer, a new channel is introduced and passed through the MLP layer twice. After the activation function, the two are multiplied element by element and added After adding each element together, we get The predicted probability distribution of joint points is obtained through the classifier The predicted probability distribution of the camera's joint points is also obtained through the classifier Finally, KL divergence is used as the loss function of the distillation module, so that the camera network has the ability to infer joint points under occlusion. Contains the features after camera and IMU fusion, expressed by the following formula: Step 4: Train the camera network parameters The SGD optimizer was used during training, with the learning rate set to 0.24 for 64 rounds; KL divergence is used as the loss function of knowledge distillation, as follows: Where D KL is the mapping relationship symbol of the KL divergence function, The value represents and The smaller the function value, the closer the two probability distributions are. The KL divergence loss function represents the difference between the model's predicted distribution and the true distribution, thereby improving the model's accuracy and generalization ability. The parameters of the monocular camera motion capture network are trained based on the motion capture after the camera and IMU are fused, so that the network has good generalization ability. After the network training is completed, you only need to input the camera image and the network can accurately output the coordinates of the joints of the human body.

Citation Information

Cited By

  • Cleaning robot deviation correction method based on component contour visual extraction

    CN121797704A