Three-dimensional human body posture prediction method based on data fusion and improved long short-term memory network

Through data fusion and improvement of LSTM network, combined with Kinect camera and OpenPose algorithm, the accuracy problem of human posture prediction in complex human-computer collaboration environments is solved, and higher prediction accuracy and stability are achieved.

CN120412093AInactive Publication Date: 2025-08-01CHINA THREE GORGES UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510524697.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-24
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In a complex human-computer collaboration environment, a single Kinect depth camera is difficult to accurately predict human posture, and there are problems of visual occlusion and data loss.

Method used

Combining the Kinect camera and OpenPose algorithm, through data fusion and improving the long and short-term memory network (LSTM), we use limb length and direction consistency constraints to optimize human posture prediction.

Benefits of technology

It significantly improves the accuracy of human posture prediction, reduces the error caused by visual occlusion and data loss, and enhances the feasibility and practicality of the prediction model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120412093A_ABST
    Figure CN120412093A_ABST
Patent Text Reader

Abstract

The invention provides a three-dimensional human body posture prediction method based on data fusion and an improved long short-term memory network, and the method comprises the steps: obtaining three-dimensional skeleton data and joint point pixel coordinate information of a human body at the same time through a Kinect camera and an OpenPose algorithm based on a data fusion technology, limb length and direction consistency constraints and an LSTM network improved model; coordinate alignment and fusion are carried out on the two kinds of data, a unified skeleton data set is constructed, motion prediction is carried out by combining limb length invariance and direction consistency as constraint conditions through a long-short-term memory network, an accurate mapping relation is generated between motion sequences, and the prediction precision of the human body posture is optimized. According to the method, the man-machine cooperation collision prediction accuracy is improved, the man-machine collision accident rate can be reduced, man-machine co-fusion development is promoted, the method is more closely combined with the actual cooperation manufacturing environment, and the benign development of safety production of a man-machine co-fusion manufacturing unit is promoted.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human - robot collaboration, and particularly to a three - dimensional human pose prediction method based on data fusion and improved long - short - term memory network. Background Art

[0002] Human - robot collaboration means that humans and robots jointly complete tasks in the same working environment, sharing spatio - temporal resources. Among them, in the human - robot collaboration scenario, human pose estimation is the basis for many research works in the field of computer vision. Through human pose estimation, the spatial position relationships and mutual relationships of various parts of the human body can be detected and understood from images or videos. Human pose estimation has broad application prospects in fields such as action recognition, behavior analysis, health monitoring, and human - computer interaction. Facing complex human - robot collaboration scenarios, relying solely on a depth camera at a fixed position and angle cannot well recognize the three - dimensional human pose, which has a certain impact on the robot's understanding of human intentions and the avoidance strategies to be adopted. Therefore, the ability to predict human poses has become a research focus in the existing human - robot collaboration environment. However, problems such as self - occlusion or occlusion by the machine of human joints and loss of jumping data in skeleton recognition still exist during human - robot collaboration. Summary of the Invention

[0003] Aiming at the problems of visual occlusion and data loss in complex human - robot collaboration environments, and the problem that it is difficult to accurately predict human poses simply through a Kinect depth camera, the present invention provides a three - dimensional human pose prediction method based on data fusion and improved long - short - term memory (LSTM) network. This method combines the characteristics of human bone coupling, takes the human action sequence as input based on constraint conditions and sends it into the LSTM network for training, reduces the generation of unreasonable data, enhances the feasibility and practicality of the action prediction model, and thus improves the prediction accuracy.

[0004] To achieve the above - mentioned technical features, the object of the present invention is realized as follows: A three - dimensional human pose prediction method based on data fusion and improved long - short - term memory network. The method is based on data fusion technology, limb length and direction consistency constraints, and an improved LSTM network model. After using a Kinect camera and OpenPose algorithm to simultaneously obtain the three - dimensional skeleton data and joint point pixel coordinate information of the human body, the two types of data are aligned and fused in coordinates to construct a unified skeleton data set. Then, the long - short - term memory network is used to perform action prediction with limb length invariance and direction consistency as constraint conditions, and an accurate mapping relationship is generated between action sequences to optimize the prediction accuracy of human poses.

[0005] Preferably, the method specifically includes the following steps:

[0006] Step 1: Perform the assembly work from the perspective of the Kinect camera. While using the camera SDK program to obtain the required human joint point data, save the RGB image of this frame.

[0007] Step 2: Use the OpenPose open-source program to extract the pixel coordinates of the required human joint points from the saved RGB image.

[0008] Step 3: Perform pixel coordinate transformation. The data is unified with the camera as the coordinate system. Use the depth data z value of the joints in the same frame obtained through the SDK and combine it with the visual principle to calculate the remaining two-dimensional x and y coordinates of the joint points.

[0009] Step 4: Use the information between adjacent frames to estimate the coordinates (x, y) of the missing joint points through the relative positions of the joint points.

[0010] Step 5: Perform data fusion to obtain a data set: Adopt the weighted average method to fuse the three-dimensional skeleton data obtained by the SDK and the picture skeleton data processed by OpenPose. During normal tracking, the joint point information of the two data sources is superimposed and fused with a certain weight; for invalid data, the data source with a higher support degree is preferentially used for replacement.

[0011] Step 6: Improve the long short-term memory network by adding constraints on limb length and direction consistency, and train the data set to obtain weights.

[0012] Step 7: Use the weight modules of the unimproved LSTM network and the improved LSTM network to conduct test experiments, and calculate the error between the predicted value and the true value through the final displacement error FDE.

[0013] Step 8: Conduct assembly experiments on the proposed multiple human-machine collaboration scenarios, repeat Steps 1 to 7, and verify the applicability of the method under different human-machine collaboration scenarios.

[0014] Preferably, in Step 1, use the camera SDK program to obtain the three-dimensional coordinate data of the required human joint points in real time, and at the same time save the RGB image of the corresponding frame. By setting the synchronous processing mechanism of the real-time data stream, ensure the synchronization of each frame of data and the image to avoid data loss or misalignment.

[0015] Preferably, in Step 2, the required human joint points include the accurate detection of 8 key joint points on the upper body of the human body, and optimize the model parameters of OpenPose to adapt to the changes of different human postures.

[0016] Preferably, in step 3, the obtained pixel coordinates are converted into a three-dimensional coordinate system with the Kinect camera as a reference through the visual principle. Using the z value of the depth data provided by the camera SDK and combining the corresponding depth camera parameters, the two-dimensional x and y coordinates of each joint point are accurately calculated by applying the mathematical formula based on perspective projection to ensure the unity of coordinates.

[0017] Preferably, in step 4, in view of the possible missing or error in the key point detection in some frames, the data of the previous frame is processed by the sliding window method, the relative position of the joint point with respect to the center of the shoulder is calculated, and based on this relative position relationship, the missing joint point coordinates in the current frame are filled.

[0018] Preferably, in step 5, the weighted average method is used to fuse the three-dimensional skeleton data obtained by the SDK and the joint point data of the RGB image extracted by the OpenPose algorithm. When the data is valid, the weights of the two data sources are both 0.5; when the data of a certain data source is invalid or missing, the data source with higher support is preferentially used for substitution to ensure the stability of data fusion.

[0019] Preferably, in step 6, in the limb length constraint, not only the accuracy of the bone length is considered, but also the relationship between each joint can be dynamically adjusted through the attention mechanism, forming an attention mechanism that not only pays attention to the spatial relationship between joints but also considers the bone length constraint. The direction vector first calculates the vector between each joint and its parent joint, and calculates their Euclidean distance, and the direction information between the parent node and the child node.

[0020] Preferably, in step 7, FDE is used to measure the final error of time series prediction, that is, to measure the error between the last time step of the prediction sequence and the last time step of the true trajectory, and is applicable to evaluating the accuracy of the model at the prediction end position.

[0021] Preferably, in step 8, experimental design is adopted to conduct assembly experiments in five typical human-robot collaboration scenarios, classify and evaluate various working conditions in each scenario, and verify the applicability and robustness of the method in various human-robot collaboration tasks through classification tests of different dynamic obstacles, occlusion situations and complex interaction actions.

[0022] The present invention has the following beneficial effects:

[0023] 1. In view of the limitations of a single Kinect camera in human pose prediction in a complex human-machine collaboration environment, the present invention proposes a three-dimensional human pose prediction method based on data fusion and an improved LSTM network. This method fuses the three-dimensional skeleton data collected by the Kinect camera with the joint point information in the RGB image processed by the OpenPose algorithm, and combines the limb length and direction consistency constraints to optimize the prediction ability of the LSTM network. Under five typical human-machine collaboration scenarios, the experimental results show that data fusion not only significantly improves the prediction accuracy, but also effectively reduces the errors caused by visual occlusion and data loss. This research provides strong technical support for improving the accuracy of human pose prediction and is expected to be widely applied in fields such as virtual assembly design, human-computer interaction, and collaborative robots in the future. Therefore, the method of the present invention has broad scene adaptability.

[0024] 2. In the field of human-machine collaboration, on the basis of improving the prediction accuracy, the present invention gives more operable space for the human body and the robotic arm in terms of advanced obstacle avoidance, making the behavior of the human body and the movement of the robotic arm more flexible. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] The present invention will be further described below with reference to the drawings and embodiments.

[0026] Figure 1 It is a schematic diagram of Kinect human joint points of the present invention.

[0027] Figure 2 It is a 25-point skeleton structure diagram corresponding to OpenPose of the present invention.

[0028] Figure 3 It is a simplified bone model of the present invention.

[0029] Figure 4 It is the conversion relationship between coordinate systems of the present invention. [[ID=Z5]]

[0030] Figure 5 It is the conversion from the camera coordinate system to the image physical coordinate system of the present invention.

[0031] Figure 6 It is the direction information between the parent node and the child node of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0032] The embodiments of the present invention will be further described below with reference to the drawings.

[0033] Embodiment 1:

[0034] Referring to Figure 1-6 , the technical problem of the present invention is that there are problems of visual occlusion and data loss in a complex human-machine collaboration environment, and it is very difficult to accurately predict the human pose simply through the Kincet depth camera.

[0035] The object of the present invention is to address the above problems and provide a three-dimensional human pose prediction method based on data fusion and an improved long short-term memory (LSTM) network. First, the three-dimensional human skeleton data obtained by a Kinect camera is fused with the skeleton data obtained by coordinate transformation of the positions of each joint pixel point in the human RGB image processed by the OpenPose algorithm in the same frame. Using the long short-term memory network as the prediction model framework and taking the length invariance and direction consistency of the human skeleton as constraint conditions, the current action and future action can be coordinated to form an accurate mapping relationship.

[0036] The technical solution of the present invention is a three-dimensional human pose prediction method based on data fusion and an improved long short-term memory (LSTM) network. The method includes the following steps:

[0037] Step 1: Perform assembly work from the perspective of the Kinect camera, and use the camera SDK program to obtain the required human joint point data while saving the RGB image of this frame.

[0038] Step 2: Extract the pixel coordinates of the required human joint points from the saved RGB image through the OpenPose open-source program.

[0039] Step 3: Perform pixel coordinate transformation, and unify the data with the camera as the coordinate system. Use the depth data z value of the joints in the same frame obtained through the SDK and combine the visual principle to calculate the remaining two-dimensional x and y coordinates of the joint points.

[0040] Step 4: In view of the problem of inaccurate key point detection, use the information between adjacent frames to estimate the coordinates (x, y) of the missing joint points through the relative positions of the joint points.

[0041] Step 5: Perform data fusion to obtain a data set. Use the weighted average method to fuse the three-dimensional skeleton data obtained by the SDK and the skeleton data of the image processed by OpenPose. During normal tracking, the joint point information of the two data sources is superimposed and fused with a weight of 0.5. For invalid data, the data source with a higher support degree is preferentially used for replacement.

[0042] Step 6: Improve the long short-term memory (LSTM) network by adding constraints on limb length and direction consistency, and train the data set to obtain weights.

[0043] Step 7: Use the unimproved LSTM network and the weight module of the improved LSTM network to conduct a test experiment. Calculate the error between the predicted value and the true value through the final displacement error FDE, and conduct a comparative analysis to prove that the improved network under the fused data set has a great effect on improving the prediction accuracy.

[0044] Step 8: Conduct assembly experiments on the five proposed human - machine collaboration scenarios, and repeat Steps 1 to 7. Verify the applicability of the method under different human - machine collaboration scenarios.

[0045] Further, in Step 1, by setting up a synchronous processing mechanism for real - time data streams, frame synchronization of joint point data and RGB images is achieved, avoiding data loss of frames.

[0046] Further, in Step 2, in order to improve the accuracy and efficiency of joint point extraction, RGB images with a high resolution of 1920*1080 are input, and the model parameters of OpenPose are optimized to adapt to changes in different human postures.

[0047] Further, in Step 3, a more accurate camera calibration method is adopted. Using the z - value of depth data combined with the internal and external parameter models of the camera, the two - dimensional coordinates of joint points are accurately calculated by applying mathematical formulas based on perspective projection.

[0048] Further, in Step 4, the shoulder center coordinates of the previous frame are used for automatic filling. Specifically, the data of the previous frame is processed by the sliding window method, the relative positions of joint points relative to the shoulder center are calculated, and based on this relative position relationship, the missing joint point coordinates in the current frame are filled.

[0049] Further, in Step 5, the skeletons are fused, and the weights are dynamically adjusted according to the stability of the skeleton data, rather than simply using a fixed weight of 0.5.

[0050] Further, in Step 6, in the limb length constraint, not only the accuracy of bone length is considered, but also the relationship between joints can be dynamically adjusted through the attention mechanism, forming an attention mechanism that not only focuses on the spatial relationship between joints but also considers bone length constraints. The direction vector first calculates the vector between each joint and its parent joint, and calculates their Euclidean distance (bone length) and the direction information between the parent node and the child node.

[0051] Further, in Step 7, FDE is used to measure the final error of time - series prediction, that is, to measure the error between the last time step of the predicted sequence and the last time step of the real trajectory, and is applicable to evaluating the accuracy of the model at the predicted end position.

[0052] Further, in Step 8, a more detailed experimental design is adopted to classify and evaluate various working conditions in each scenario. Through classification tests of different dynamic obstacles, occlusion situations, and complex interaction actions, the applicability and robustness of the method in various human - machine collaboration tasks are verified.

[0053] Example 2:

[0054] To address the occlusion and data loss issues in existing human-machine collaboration scenarios, the present invention provides a 3D human posture prediction method based on data fusion and an improved long short-term memory (LSTM) network. The specific implementation method is as follows:

[0055] 1. Obtain human joint data points through the depth camera Kinect, such as Figure 1 As shown; the OpenPose algorithm extracts joint pixels, such as Figure 2 As shown. By fusing the same skeleton points, the human body joint data in the camera coordinate system is obtained. The fused skeleton points are as follows Figure 3 shown.

[0056] 2. OpenPose extracts joint pixels and performs coordinate transformation to obtain the x and y values in the camera coordinate system. The x and y values in the camera coordinate system are obtained by inverse transformation from the image plane coordinates. c and Y c The formula is:

[0057]

[0058] Among them, f represents the focal length, which represents the distance from the optical center of the camera to the image plane, that is, the difference between the camera coordinate system and the image coordinate system on the Z axis, such as Figure 5 shown.

[0059] 3. Data enhancement and supplementation: the effect is improved by frame-by-frame restoration. Based on the shoulder center position of the previous frame, the present invention calculates the relative position of the joint point to fill in the estimated coordinates (X n ,Y n ), the calculation formula is:

[0060] ΔX=X w -X i ;

[0061] ΔY=Y h -Y j ;

[0062] X n =X m +ΔX;

[0063] Y n =Y k +ΔY;

[0064] Among them, the coordinates of the shoulder center in the previous frame are (X i ,Y j ), the coordinates of the shoulder center in the current frame are (X m ,Y k ), the coordinates of the lost bone joint point in the previous frame are (X w , Yh ).

[0065] 4. To effectively integrate the data from both, it is necessary to set appropriate weights based on the experimental results. For invalid data, data sources with high support are used as replacements first. Assume that the weight of Kinect data is w K , OpenPose data weight is w O , these two weights should satisfy the sum of 1, normally w K is 0.5, w O The weighted average fusion method is used to enhance the stability and accuracy of the skeleton model. For each joint point, the weighted average calculation is performed using the following formula:

[0066] w K +w O =1;

[0067] X n =w K ×X Kinect +w O ×X OPenPose ;

[0068] Y n =w K ×Y Kinect +w O ×Y OpenPose ;

[0069] Z n =Z Kinect ;

[0070] 5. Establish skeleton constraints, combine bone length constraints and pairwise attention mechanism to dynamically adjust the spatial relationship between joints. When offline, calculate the average distance μ in the training set u,v and standard deviation σ u,v As a priori limb length distribution let e u,v =(μ u,v ,σ u,v ) as the predefined parameters of the limb. The formula for defining the joint pair attention weight and the bone length loss function is:

[0071]

[0072] Among them, α is a hyperparameter that adjusts the tolerance of bone length error. The arm length error tolerance is empirically set to 1500. ∈ is used to enhance numerical stability. μ u,v It is expressed as the average distance between two joints calculated in the training set. For each pair of joints (u, v), the vector d between joint u and its parent joint v is calculated. u,v =p u -p v, that is, the Euclidean distance is l u,v = ||d u,v ||2;

[0073] 6. According to Figure 6 As shown, the direction loss calculates the squared error by comparing the direction vector predicted by the model with the true direction vector:

[0074]

[0075] where d pred = (x pred , y pred , z pred ). The direction loss strengthens the model's understanding of the human body structure by considering the direction relationship between joint points, thus generating more reasonable and natural postures.

[0076] 7. By adding bone length and direction constraint conditions to the training of the LSTM model, the present invention introduces two loss errors. The mean squared error inherent in the LSTM mainly focuses on the joint position error. Therefore, the present invention needs to balance these errors during the training process, set a weight for each error, and adjust their contributions to the total loss. By adjusting these weights, a suitable balance point can be found:

[0077] loss = weight_pos * mse_loss + weight_bone * bone_loss + weight_dir * direction_loss;

[0078] where weight_pos, weight_bone, and weight_dir are adjustable hyperparameters. The present invention can be gradually optimized by adopting a phased training strategy. First, optimize the joint positions, and then gradually introduce bone length and direction errors, thereby gradually improving the structural consistency of the model.

[0079] 8. To focus on the accuracy of the final position, most existing studies use the deviation between the predicted distance and the true distance as the evaluation index, and use the final displacement error FDE (Final Displacement Error) to evaluate the model. The calculation formula is:

[0080]

[0081] where x i,T , y i,T , z i,T are the position coordinates of the predicted trajectory at the final time step T; is the position coordinate of the true trajectory at the final time step T, and n is the number of joints.

[0082] 9. Under the fused dataset, the improved and unimproved LSTM network models were compared, and it was found that the prediction accuracy of the improved LSTM network under the fused dataset was higher. Then, the improved LSTM network model was applied to the single Kinect dataset, and the prediction accuracies under the two datasets were compared to verify the feasibility of the present invention.

[0083] 10. The present invention designed 5 specific human-robot collaboration scenarios, which mainly revolved around the cooperation and respective operations of the robotic arm and humans on the assembly table in the factory. The experimental scenarios were that the robotic arm and humans moved in the same direction, the robotic arm and humans moved towards each other, the robotic arm and human hands crossed, the robotic arm and humans worked on both sides separately, and the robotic arm and humans worked on the same side.

Claims

1. A three-dimensional human pose prediction method based on data fusion and improved long short-term memory network, characterized in that: The method is based on data fusion technology, limb length and direction consistency constraints, and an improved LSTM network model. After using a Kinect camera and the OpenPose algorithm to simultaneously obtain the three-dimensional skeleton data and joint pixel coordinate information of the human body, the two types of data are aligned and fused in coordinates to construct a unified skeleton data set. Then, the long short-term memory network combines limb length invariance and direction consistency as constraint conditions for action prediction, and an accurate mapping relationship is generated between action sequences to optimize the prediction accuracy of human postures.

2. The three-dimensional human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that, The method specifically includes the following steps: Step 1: Perform assembly work from the perspective of the Kinect camera. While using the camera SDK program to obtain the required human joint data, save the RGB image of this frame. Step 2: Use the OpenPose open-source program to extract the pixel coordinates of the required human joints from the saved RGB image. Step 3: Perform pixel coordinate transformation. The data is unified with the camera as the coordinate system. Use the z value of the depth data of the same-frame joints obtained through the SDK and combine the visual principle to calculate the remaining two-dimensional x and y coordinates of the joints. Step 4: Use the information between adjacent frames to estimate the coordinates (x, y) of the missing joints by the relative positions of the joints. Step 5: Perform data fusion to obtain a data set: Use the weighted average method to fuse the three-dimensional skeleton data obtained by the SDK and the picture skeleton data processed by OpenPose. During normal tracking, the joint information of the two data sources is superimposed and fused with a certain weight; for invalid data, the data source with a higher support degree is preferentially used for substitution. Step 6: Improve the long short-term memory network by adding limb length and direction consistency constraints, and train the data set to obtain weights. Step 7: Use the unimproved LSTM network and the weight module of the improved LSTM network for test experiments, and calculate the error between the predicted value and the true value through the final displacement error FDE. Step 8: Conduct assembly experiments on the proposed multiple human-machine collaboration scenarios, repeat Steps 1 to 7, and verify the applicability of the method under different human-machine collaboration scenarios.

3. The 3D human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, wherein In Step 1, use the camera SDK program to obtain the three-dimensional coordinate data of the required human joints in real time, and at the same time save the RGB image of the corresponding frame. By setting the synchronous processing mechanism of the real-time data stream, ensure the synchronization of each frame of data and the image to avoid data loss or misalignment.

4. A three-dimensional human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that In Step 2, the required human joints include the accurate detection of 8 key joints on the upper body of the human body, and the model parameters of OpenPose are optimized to adapt to the changes of different human postures.

5. The 3D human pose prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that, In Step 3, convert the obtained pixel coordinates into a three-dimensional coordinate system with the Kinect camera as the reference through the visual principle. Use the z value of the depth data provided by the camera SDK, combine the corresponding depth camera parameters, and apply the mathematical formula based on perspective projection to accurately calculate the two-dimensional x and y coordinates of each joint to ensure the unity of coordinates.

6. The three-dimensional human body posture prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that, In step 4, considering that key point detection in some frames may be missing or inaccurate, the data of the previous frame is processed by the sliding window method, the relative positions of the joints with respect to the shoulder center are calculated, and based on this relative position relationship, the missing joint coordinates in the current frame are filled.

7. A three-dimensional human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that In step 5, the weighted average method is used to fuse the three-dimensional skeleton data obtained by the SDK and the RGB image joint point data extracted by the OpenPose algorithm. When the data is valid, the weights of the two data sources are both 0.5; When the data of a certain data source is invalid or missing, the data source with higher support is preferentially used for replacement to ensure the stability of data fusion.

8. A three-dimensional human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that, In step 6, in the limb length constraint, not only the accuracy of the bone length is considered, but also the relationship between joints can be dynamically adjusted through the attention mechanism, forming an attention mechanism that pays attention to both the spatial relationship between joints and the bone length constraint. The direction vector first calculates the vector between each joint and its parent joint, and calculates their Euclidean distance, and the direction information between the parent node and the child node.

9. A three-dimensional human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, characterized in that, In step 7, FDE is used to measure the final error of time series prediction, that is, to measure the error between the last time step of the prediction sequence and the last time step of the true trajectory, and is applicable to evaluating the accuracy of the model at the prediction end position.

10. The 3D human body pose prediction method based on data fusion and improved long short-term memory network according to claim 1, wherein, In step 8, experimental design is adopted to conduct assembly experiments in five typical human-robot collaboration scenarios, classify and evaluate various working conditions in each scenario, and verify the applicability and robustness of the method in various human-robot collaboration tasks through classification tests of different dynamic obstacles, occlusion situations and complex interaction actions.