Multi-vision sensor human body joint data fusion method based on quality evaluation
Through the data fusion method of multi-vision sensors, the problem of inaccurate human posture estimation of a single sensor in complex environments is solved, and the human posture prediction with higher accuracy and robustness is achieved, supporting the real-time response of the robot in complex environments.
Patent Information
- Application Number
- CN202510532143.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-01
AI Technical Summary
In dynamic and complex human-computer collaboration scenarios, existing single sensor or fixed-angle visual data acquisition methods are difficult to provide accurate human posture estimation, especially in the face of problems such as occlusion, sensor error, joint data beat, etc., the three-dimensional joint position acquisition is inaccurate.
Joint data is obtained through multiple vision sensors, quality evaluation and data fusion are carried out, including sensor calibration, data synchronization, reliability judgment, motion threshold judgment and sliding window technology, and data fusion is used to ensure data reliability and robustness.
It improves the accuracy of three-dimensional human posture prediction and the practicality of the model, enhances the robustness of the system in the case of occlusion and data loss, and supports robots to make real-time responses and decisions in a dynamic human-machine collaboration environment.
Smart Images

Figure CN120412094A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of human motion recognition, prediction and other human - machine collaboration technologies, and particularly relates to a method for fusing human joint data of multiple vision sensors based on quality assessment. Background Art
[0002] Human - machine collaboration refers to humans and robots collaborating to complete tasks in a common working environment, sharing spatio - temporal resources with each other. In this process, human pose estimation technology, as an important research direction in the field of computer vision, plays a crucial role. Through human pose estimation technology, the spatial position relationships and their interactions of each joint of the human body can be extracted and analyzed from images or videos, which has broad prospects for applications such as action recognition, behavior analysis, health monitoring, and human - machine interaction. In complex human - machine collaboration scenarios, a single depth camera at a fixed position and angle is difficult to comprehensively capture the three - dimensional pose information of the human body, which poses challenges for robots to understand human actions, infer human intentions, and adopt appropriate avoidance strategies.
[0003] In summary, in the existing human - machine collaboration scenarios, the technical problem is that in dynamic and complex human - machine collaboration scenarios, the existing single - sensor or fixed - angle visual data acquisition methods are difficult to provide accurate human pose estimation when facing problems such as occlusion, sensor errors, and joint data jitter, such as vibration, bone length change, or self - occlusion, resulting in inaccurate acquisition of three - dimensional joint positions.
[0004] Therefore, how to improve the accuracy of human pose estimation, especially to perform precise data fusion in an environment where multiple vision sensors work together, has become one of the research focuses. Summary of the Invention
[0005] In view of the above problems, the present invention provides a method for fusing human joint data of multiple vision sensors based on quality assessment. This method performs quality assessment on the joint data obtained by multiple vision sensors, and combines the spatial and temporal synchronization information between sensors to perform weighted fusion on the joint data of different sensors. Through this fusion method, the error of a single sensor can be effectively reduced, and the robustness of the system in the case of occlusion and data loss can be enhanced, further improving the accuracy of three - dimensional human pose prediction and the practicality of the model, so as to better support the robot to make real - time responses and decisions in a dynamic human - machine collaboration environment.
[0006] To achieve the above - mentioned technical features, the object of the present invention is achieved as follows: A method for fusing human joint data of multiple vision sensors based on quality assessment, the method comprising the following steps: Step 1: Simultaneously collect human joint data based on multiple Kinect v2 sensors; before data fusion, calibrate each Kinect v2 sensor so that data from different sensors can be processed in a unified coordinate system; Step 2: After unifying the coordinates of the Kinect v2 sensors, conduct real-time communication via TCP / IP. Set the computer PC connected to the main Kinect v2 sensor as the server, and the computer PCs connected to the remaining sensors as clients; Step 3: Judge the credibility of each joint point of the sensor; Step 4: Solve the unified upper body bone model; Step 5: Judge the joint movement amplitude threshold. If the movement amplitude is greater than the set threshold T, it is considered that the position data of this joint point is not credible during the current time period; Step 6: When a single Kinect v2 sensor collects three-dimensional information of human joint points, set a sliding window for it, and adjust the weights of the three-dimensional data of each joint point obtained within the effective range of the Kinect device; Step 7: Fuse the data collected by multiple devices. The data of each joint point is weighted and averaged according to the weights to obtain the fused joint point coordinates; Step 8: Build an experimental scenario with human occlusion in a human-machine collaborative manufacturing environment, conduct an experiment on the fusion of human data information under occlusion, and verify the superiority of the algorithm in a collaborative scenario.
[0007] Preferably, in step 1, by using a method combining geometric calibration and visual calibration, determine that each Kinect v2 sensor processes data in the same global coordinate system.
[0008] Preferably, during the communication process in step 2, by using standard network protocols, ensure that the real-time data of each Kinect v2 sensor can be quickly transmitted to the host and processed synchronously.
[0009] Preferably, during the credibility judgment process in step 3, use the TrackingState enumeration field provided by the Kinect SDK to judge whether each joint is successfully tracked. For each joint point, the SDK will return different tracking states, such as "TRACKED", "INFERRED", and "NOT_TRACKED".
[0010] Preferably, step 4 specifically includes: Before processing human joint data, establish a unified upper body bone model, thereby reducing the offset of joint positions caused by human movement, and providing accurate bone model support for subsequent joint data threshold judgment and data fusion.
[0011] Preferably, in step 5, the movement amplitude of the joint is judged according to the three-dimensional position change of the joint , and compared with the joint movement amplitude threshold T to improve the accurate screening of abnormal data and the determination of valid data.
[0012] Preferably, in step 6, a sliding window with an activity range of 60 is set. Within the sliding window, the weights of each data frame are dynamically adjusted, and the size of the weight depends on the time freshness and credibility of the data. Among them, the latest data frame will be given a larger weight, and as time goes by, the weight of earlier data frames gradually decreases to ensure a more timely response to the rapid changes of the joint.
[0013] Preferably, when fusing the human joint data collected by multiple Kinect v2 sensors in step 7, the data of each joint point is adjusted according to the weight, and a method combining the weighted average method and the quality evaluation mechanism is adopted to ensure the accuracy and robustness of data fusion.
[0014] The present invention has the following beneficial effects: 1. Through steps such as multi-sensor calibration, data synchronization, credibility judgment, motion threshold judgment, sliding window, and data weighting, the present invention ensures that data with higher sensor quality has a greater impact on the final fusion result, thereby enhancing the reliability and robustness of the overall data fusion.
[0015] 2. The method of the present invention can effectively eliminate the errors between devices and ensure more stable and accurate data support for the cooperation between robots and humans in the case of occlusion.
[0016] 3. Through geometric calibration, the present invention can ensure the unity of the coordinate systems of each Kinect v2 sensor, thereby eliminating the spatial errors between different sensors and ensuring the consistency of subsequent data processing.
[0017] 4. Through this preferred experimental design and verification method, the robustness of the algorithm in practical applications can be fully tested, ensuring that in a complex human-robot cooperation environment, even in the face of problems such as occlusion, data loss, or quality degradation, the system can still provide reliable and accurate human pose estimation. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] The present invention will be further described below with reference to the drawings and embodiments.
[0019] Figure 1 Three-sided calibration target point cloud information acquisition.
[0020] Figure 2 Point cloud plane segmentation and extraction.
[0021] Figure 3 Kinect SDK enumeration field representation.
[0022] Figure 4 Upper body bone model.
[0023] Figure 5 Flowchart for judging the validity of joint points. Specific implementation mode
[0024] The following further explains the implementation mode of the present invention in conjunction with the attached drawings.
[0025] Example 1: Refer to Figures 1 - 5 , aiming at the fact that the single sensor or the visual data acquisition method at a fixed angle in the complex human-machine cooperation scenario is difficult to provide accurate human body pose estimation, resulting in inaccurate acquisition of three-dimensional joint positions, a multi-visual sensor human joint data fusion method based on quality assessment is proposed. The specific implementation method is as follows: The method is based on multiple visual sensors, Kinect v2 cameras, to simultaneously collect human joint data, overcoming the measurement errors caused by a single sensor during discontinuous movement, occlusion, or equipment failure. Before data fusion, first, the Kinect v2 sensor needs to be calibrated to unify its coordinates; secondly, the credibility of each joint point is judged; since the bone length between each joint of the same target human body is a fixed value, a movement amplitude threshold T for each joint is established; the movement amplitude of the human body is compared with the threshold T. When the joint movement is less than the threshold T, the data of this joint point can be considered credible, otherwise, data distortion or unreliability may occur; finally, in order to quantitatively describe the phenomenon that the quality of joint point data changes in real time due to the change of joint point tracking state, a data quality control and multi-device data fusion method is designed through a sliding window and weight adjustment. This method is widely used in scenarios that require high-precision human motion capture and analysis, such as human-machine cooperation, virtual reality (VR), augmented reality (AR), and sports health monitoring.
[0026] Specifically, it includes the following steps: Step 1: Before data fusion, first calibrate each Kinect v2 sensor so that the data of different sensors can be processed in a unified coordinate system, thereby eliminating the errors caused by different sensor positions and directions.
[0027] Step 2: After unifying the coordinates of the Kinect v2 sensor, real-time communication needs to be carried out through TCP / IP. Set the computer PC connected to the main sensor Kinect v2 as the server, and the computer PCs connected to the other sensors as the clients.
[0028] Step 3: Determine the credibility of each joint point of the sensor. It is possible to determine whether a human joint point can be tracked through the enumeration fields provided by the built-in algorithm of the Kinect SDK, and the reliability of this position is represented in an enumeration manner.
[0029] Step 4: Since human movement often causes offsets in the same bone length of the same target human body between different frames, it is necessary to solve a unified upper body bone model before using the original data as the joint movement threshold T.
[0030] Step 5: Determine the joint movement amplitude threshold. If the movement amplitude is greater than the set threshold T, it is considered that the position data of this joint point is not credible during the current time period.
[0031] Step 6: When collecting three-dimensional information of human joint points by a single Kinect v2 sensor, set a sliding window with an active range of 60 frames, and adjust the weights of the three-dimensional data of each joint point obtained within the effective range of the Kinect device.
[0032] Step 7: Fuse the data collected by multiple devices, and perform weighted averaging on the data of each joint point according to the weights to obtain the fused joint point coordinates.
[0033] Step 8: Build an experimental scenario with human occlusion in a human-machine collaborative manufacturing environment, conduct an experiment on the fusion of human data information under occlusion, and verify the superiority of the algorithm in a collaborative scenario.
[0034] Preferably, in Step 1, ensure that the coordinate systems of each Kinect v2 sensor are unified. The key to this step is to eliminate the spatial errors between different sensors through geometric calibration to ensure the consistency of subsequent data processing. <l
[0035] Preferably, in Step 2, communicate by setting a master sensor (PC as the server) and a client (PC of other sensors) to ensure that the data of each sensor can be synchronized and transmitted to the main control system for processing. The TCP / IP protocol can ensure low latency and stability and is suitable for real-time data transmission.
[0036] Preferably, in Step 3, the Kinect SDK provides the joint point data of the human skeleton and provides a credibility score for each joint point, which can determine whether the joint point is reliable. If the credibility is low, it means that there may be errors in the coordinate data of this joint point.
[0037] Preferably, in Step 4, reduce the error caused by bone movement through unified processing of the upper body bone model. This step calculates the static bone model of the human body as a reference framework for subsequent data processing. A
[0038] Preferably, in step 5, according to the movement amplitude of the joint , the credibility of the current joint data can be judged. If the movement amplitude exceeds the set threshold T, it may indicate that the data is abnormal and needs to be discarded or recalculated. This step is used to eliminate abnormal movement data and improve data quality.
[0039] Preferably, in step 6, the sliding window technique can help reduce noise and mutation data, and calculate the joint point coordinates of the current frame through the historical data in the window. Weight adjustment can make the recent data have a greater impact on the result, thereby improving the real-time performance and accuracy of the data.
[0040] Preferably, in step 7, the data of each joint point is adjusted according to the weight. The outputs of multiple sensors can be integrated by weighted average of weights, reducing the data error of a single sensor, and generating the accurate position of the human joint points in a global coordinate system.
[0041] Preferably, in step 8, through this preferred experimental design and verification method, the robustness of the algorithm in practical applications can be fully tested, ensuring that in a complex human-machine collaboration environment, even in the face of problems such as occlusion, data loss, or quality degradation, the system can still provide reliable and accurate human pose estimation.
[0042] Embodiment 2: This embodiment provides a method for fusing human joint data of multiple vision sensors based on quality assessment. The specific implementation method is as follows: 1. First, use the three-plane target point cloud information reconstruction technology in the vision and geometric calibration method to calibrate the multi-sensor Kinect v2 to unify its coordinates. Through the collection of the three-plane target point cloud information by each sensor, such as Figure 1 , secondly, segment and extract the point cloud plane of each sensor, such as Figure 2 , and finally use PCL the point cloud library to calculate the rotation matrix R and the translation matrix T between multiple sensors.
[0043] 2. Use the TCP / IP protocol to perform real-time communication and data synchronization between the computers PC connected to the sensors. Specifically, set the computer connected to the main sensor (main PC) as the server, and the computers of the remaining Kinect sensors as the clients. Adopt multi-thread technology and data buffer mechanism to ensure that the data of each sensor can be transmitted to the host in chronological order and synchronized for processing, and will not be lost or disordered due to network delay or transmission blockage.
[0044] 3. When each sensor collects human joint data, an enumeration field based on the built-in algorithm of the Kinect SDK is used to determine whether a human joint point can be tracked, and the enumeration method is used to evaluate the credibility of the joint point. For each joint point, the SDK will return different tracking states, such as Figure 3 . If the joint point status is "TRACKED", the data of this joint point is considered reliable and the data of this joint point can be used; if it is "INFERRED", it indicates that the estimated value of this joint position is not very reliable and may be affected by noise or error; if it is "NOT_TRACKED", it means that this joint point cannot be tracked and the data is not credible. If the enumeration field is "NOT_TRACKED" and "INFERRED", it indicates that the human joint point data obtained by the sensor does not conform to human characteristics, and the coordinates are (0, 0...0).
[0045] 4. First, by modeling the human bone structure, especially the upper body (such as joints of the shoulders, elbows, wrists, necks, etc.), such as Figure 4 , a bone constraint model is established. During the movement of each frame, the offset of the joint is continuously corrected. This process ensures that among the joints of the bone model, consistency is maintained throughout the movement sequence, avoiding the appearance of a skeleton structure that does not conform to biological laws.
[0046] The specific design is that the target human poses in a "T" pose facing the Kinect sensor without occlusion, so that the sensor continuously collects the three-dimensional coordinate data of 8 joint points when the human body is stationary for 500 frames. To prevent the data collected in the middle and at the end from fluctuating greatly, the middle 300 frames of data are used to solve the unified bone model. For the bone length, the following formula is used for calculation: ; where the bone length is the Euclidean distance between two joints, ( , ) represents the three-dimensional information of the joint point k in the i frame, and ( , ) represents the three-dimensional information of the adjacent joint point k in the j frame. Theoretically, the deviation between the same bone lengths in different frames will not be too large. If there is data with a large difference from most of the data, these data are removed, and finally N frames of valid data are left. According to the calculation, the unified upper body bone length in the experiment is obtained, as shown in Table 1.
[0047] Table 1 Upper body bone length
[0048] 5. During the process of judging the joint movement range, it is necessary to consider the movement range of the human body. According to the human movement characteristics, set the movement range thresholds for each joint of the human body. T . Assume that the position of the joint point at the previous moment is . The current position is . The movement range is expressed by the Euclidean distance formula as: ; Among them, represents the movement range. If is greater than the set threshold, it is considered that the position data of this joint point is not credible within the current time period. The flowchart for judging the validity of joint points is shown in Figure 5 .
[0049] 6. When processing the three-dimensional data of human joint points collected by a single Kinect v2 sensor, set a sliding window with an active range of 60 frames. This window can continuously slide in the time series and capture historical data. Whenever a new data frame enters, the window slides forward, discards the oldest frame, and includes the most recent 60 frames of data. In this way, the joint positions can be smoothed and denoised through the data sequence within the window. Assume that there are r frames of valid data in the sliding window, and 60 - r frames of data are not credible due to occlusion and other problems. As the number of frames of credible data in the window increases, the weights of the three-dimensional data of each joint point obtained within the effective range of the Kinect device become larger. The weight w of each joint point is expressed as follows: ; Among them, represents the adjustment factor for adjusting the change rate of the joint point data weight. According to domestic and foreign literature, generally = 3.
[0050] 7. Before fusing the human joint data collected by multiple Kinect v2 sensors, first evaluate the quality of the joint point data from different Kinect sensors and assign different weights. The basic principle of weight assignment is that the higher the data quality of the sensor, the greater its influence on the final fusion result. Among them, the deviation degree of the data collected by a single Kinect sensor device is expressed as: ; Among them, is the coordinate of the joint point collected by the n th device, j and is the weight of the n th device for this joint point. The greater the deviation degree , the more it indicates that the devicen The data is quite different from the collected data of other devices. The degree of deviation of the reciprocal is used as the new weight of the device n to reduce the influence of the device with a large data deviation.
[0051] 8. Finally, the data is fused. The joint point data collected by all devices is weighted and averaged according to the weights to obtain the fused joint point coordinates , and the calculation formula is: ; Among them, for the device with a data deviation the smaller it is, then the larger it is, and it accounts for a greater weight in the fusion result, while the influence of the device with a large data deviation will be reduced.
[0052] 9. In order to verify the effectiveness of this method under occlusion, the present invention conducts a comparative experiment with a single sensor and an unweighted data fusion method to evaluate the performance of the fusion algorithm in dealing with occlusion problems. The comparative experiment can evaluate the performance of different methods under occlusion by calculating indicators such as position error, prediction accuracy, and real-time performance, so as to prove the superiority of the optimized algorithm.
Claims
1. A method for fusing human joint data of multiple vision sensors based on quality assessment, characterized in that The method includes the following steps: Step 1: Simultaneously collect human joint data based on multiple Kinect v2 sensors; before data fusion, calibrate each Kinect v2 sensor so that data from different sensors can be processed in a unified coordinate system; Step 2: After unifying the coordinates of the Kinect v2 sensors, perform real-time communication via TCP / IP. Set the computer PC connected to the main Kinect v2 sensor as the server, and the computer PCs connected to the other sensors as clients; Step 3: Judge the credibility of each joint point of the sensor; Step 4: Solve the unified upper body bone model; Step 5, judge the joint movement amplitude threshold. If the movement amplitude is greater than the set threshold T, it is considered that the joint position data in the current time period is not credible; Step 6: When collecting three-dimensional information of human joint points by a single Kinect v2 sensor, set a sliding window for it, and adjust the weights of the three-dimensional data of each joint point obtained within the effective range of the Kinect device; Step 7: Fuse the data collected by multiple devices. The data of each joint point is weighted and averaged to obtain the fused joint point coordinates; Step 8: Build an experimental scenario with human occlusion in the human-machine collaborative manufacturing environment, conduct an experiment on the fusion of human data information under occlusion, and verify the superiority of the algorithm in the collaborative scenario.
2. The method for fusing human joint data of multiple vision sensors based on quality assessment according to claim 1, wherein, In Step 1, by adopting a method combining geometric calibration and visual calibration, it is determined that each Kinect v2 sensor processes data in the same global coordinate system.
3. The multi-visual sensor human joint data fusion method based on quality assessment according to claim 1, wherein, In Step 2 during the communication process, by using standard network protocols, it is ensured that the real-time data of each Kinect v2 sensor can be quickly transmitted to the host and processed synchronously.
4. A method for fusing human joint data of multi-vision sensors based on quality assessment according to claim 1, characterized in that, In Step 3 during the credibility judgment process, use the TrackingState enumeration field provided by the Kinect SDK to judge whether each joint is successfully tracked. For each joint point, the SDK will return different tracking states, such as "TRACKED", "INFERRED", and "NOT_TRACKED".
5. The multi-vision sensor human joint data fusion method based on quality assessment according to claim 1, characterized in that, Step 4 specifically includes: Before processing human joint data, establish a unified upper body bone model, thereby reducing the joint position offset caused by human movement, and providing accurate bone model support for subsequent joint data threshold judgment and data fusion.
6. The multi-vision sensor human joint data fusion method based on quality assessment according to claim 1, characterized in that, In step 5, the movement amplitude of the joint is judged according to the three-dimensional position change of the joint , and compared with the joint movement amplitude threshold T to improve the accurate screening of abnormal data and the determination of valid data.
7. The method for fusing human joint data of multiple vision sensors based on quality assessment according to claim 1, wherein In Step 6, set a sliding window with an activity range of 60. Within the sliding window, the weight of each data frame is dynamically adjusted. The size of the weight depends on the time freshness and credibility of the data. Among them, the latest data frame will be given a larger weight, and as time goes by, the weight of earlier data frames gradually decreases to ensure a more timely response to the rapid changes of joints.
8. A method for fusing human joint data of multi-vision sensors based on quality assessment according to claim 1, characterized in that, In Step 7 when fusing the human joint data collected by multiple Kinect v2 sensors, adjust the data of each joint point according to the weight, and adopt a method combining the weighted average method and the quality evaluation mechanism to ensure the accuracy and robustness of data fusion.