Mechanical arm operation data acquisition system and method based on multi-source perception

By using multimodal perception and data fusion technology, the problem of incomplete data acquisition by robotic arms in occluded environments has been solved, achieving high-quality data collection and adaptive optimization, and improving the robotic arm's operational capabilities in complex environments.

CN121340281APending Publication Date: 2026-01-16NANCHANG UNIV +1
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202511755526.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-26
Publication Date
2026-01-16

AI Technical Summary

Technical Problem

Existing robotic arm data acquisition methods struggle to obtain complete and accurate operational data in complex operating scenarios, especially under occlusion conditions. They also face difficulties in synchronizing multi-source data and lack the ability to capture real physical interaction characteristics, thus affecting the accuracy of data analysis.

Method used

By combining multimodal perception, active interaction, and online learning, and through isomorphic robotic arm operation module, multi-view visual image acquisition module, optical motion capture module, and occlusion detection and data fusion module, it can autonomously generate highly generalizable embodied datasets and optimize the robotic arm's interaction strategy in real time.

Benefits of technology

Maintaining data integrity in occluded environments improves the adaptability and task execution efficiency of robotic arms in complex operating scenarios, and provides high-quality multimodal datasets to support deep learning and reinforcement learning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121340281A_ABST
    Figure CN121340281A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-source perception-based mechanical arm operation data acquisition system and method, the system comprises an isomorphic mechanical arm operation module, a multi-view visual image acquisition module, an optical motion capture module and a shielding detection and data fusion module, vision, pose and force sense data are cooperatively acquired through multiple sensors, and time-space consistency is synchronously kept by adopting hardware. The system has a multi-dimensional occlusion detection mechanism and can dynamically switch data processing strategies according to the occlusion degree: visual data is preferentially adopted when no occlusion exists, multi-sensor information is fused when partial occlusion exists, and force feedback compensation is predicted and introduced based on a dynamic model when complete occlusion exists. Physical feedback and visual perception are creatively combined, the problem of data missing of a traditional single sensor data acquisition system under the shielding condition is solved, more comprehensive and more reliable data support is provided for mechanical arm algorithm training and performance optimization, and the method is particularly suitable for complex operation scenes with shielding.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of robot perception and data acquisition technology, specifically relating to a multi-source data acquisition system for robotic arm operation. Background Technology

[0002] In today's rapidly developing robotics technology, high-quality operational data is crucial for the algorithm training and performance optimization of robotic arms. However, existing data acquisition methods for robotic arms often struggle to obtain complete and accurate operational data when faced with complex operational scenarios, especially those with occlusion. Traditional single-sensor data acquisition schemes have significant limitations: vision-based acquisition systems experience a sharp decline in data quality when occluded; schemes relying on force / torque sensors cannot provide spatial pose information of the target object; and while motion capture systems offer high accuracy, they are expensive and require strict visibility of marker points.

[0003] Current data acquisition platforms still face challenges in synchronizing multi-source data. Data collected from different sensors encounters technical challenges in areas such as timestamp alignment and coordinate system consistency, directly impacting the accuracy of subsequent data analysis. Furthermore, existing systems often lack the ability to capture real-world physical interaction characteristics, such as force feedback and minute deformations when a robotic arm interacts with its environment. These factors lead to discrepancies between the acquired data and actual application scenarios, making it difficult to meet the high training data requirements of advanced algorithms like deep learning.

[0004] With the continuous expansion of robot applications, especially in fields such as intelligent manufacturing and medical surgery where high precision is required, the demand for high-quality robotic arm operation data is becoming increasingly urgent. Developing a platform capable of overcoming occlusion and achieving high-precision synchronous acquisition of multi-source data is of great significance for promoting the development of robotics technology. Such a platform can not only provide reliable data support for algorithm development but also provide more precise operational guidance in practical applications, thereby improving the robotic arm's ability to operate in complex environments. Summary of the Invention

[0005] This invention addresses the technical shortcomings of existing robotic arm data acquisition systems, which suffer from incomplete and inaccurate data acquisition in occluded environments. It proposes a multi-source sensing robotic arm operation data acquisition system and method. This system achieves adaptive data acquisition and skill optimization for the robotic arm in complex environments through the organic combination of multimodal perception, active interaction, and online learning. The system can autonomously generate highly generalizable embodied datasets and optimize the robotic arm's interaction strategies in real time, significantly improving the robot's adaptability to dynamic environments and task execution efficiency.

[0006] This invention is achieved through the following technical solution.

[0007] The present invention discloses a robotic arm operation data acquisition system based on multi-source perception, comprising a homogeneous robotic arm operation module, a multi-view visual image acquisition module, an optical motion capture module, and an occlusion detection and data fusion module.

[0008] The isomorphic robotic arm operation module consists of a UR5 robotic arm and a scaled-down isomorphic arm with the same structure as the UR5, forming a working unit. The isomorphic arm serves as the master operating arm, and the UR5 robotic arm serves as the slave arm. The UR5 robotic arm is equipped with an adaptive gripper and a high-precision six-dimensional force sensor at its end, enabling real-time sensing of force and torque changes during operation. The isomorphic arm also uses a 3D-printed model gripper at its end, allowing precise control of the opening and closing of the UR5 end gripper. Motion data from the master operating arm flows as control commands to the slave arm, driving it to reproduce the operation. Simultaneously, the six-dimensional force sensor at the slave arm's end collects real-time force / torque data as physical feedback, which flows back to the control system and is uploaded to the occlusion detection and data fusion module.

[0009] The multi-view visual image acquisition module employs three high-resolution industrial cameras, positioned at the end of the robotic arm, above the work area, and to the side of the operating plane, forming a comprehensive observation network. All cameras are synchronized via hardware to ensure consistent data acquisition timing and are equipped with specialized optical components to eliminate ambient light interference. The module uses standard calibration methods to ensure the measurement accuracy of each camera. The synchronized image stream acquired by the three cameras is transmitted in real-time to the processing host via a high-speed network, serving as the raw visual data stream for the occlusion detection and data fusion module, used for target recognition and pose estimation.

[0010] The optical motion capture module consists of multiple high-performance infrared cameras, with dedicated reflective markers placed on key parts of the robotic arm and the manipulated object. The module is precisely calibrated and can track the three-dimensional motion trajectory of the robotic arm and the object in real time. Through multi-sensor data fusion technology, the module can automatically adjust its data acquisition strategy according to environmental changes, ensuring complete and accurate operational data under various working conditions. The infrared camera array captures images of the marker points, and triangulation calculations generate high-precision six-DOF pose data of the robotic arm and the target object. This data serves as a true reference for spatial positioning and is transmitted in real time to the occlusion detection and data fusion module.

[0011] The present invention proposes a complete solution for the occlusion detection and data fusion module. First, a multi-dimensional occlusion detection mechanism is employed, comprehensively analyzing the number of visual feature points, motion capture markers, and force sensor data to determine the occlusion status. The degree of occlusion (no occlusion, partial occlusion, complete occlusion) is determined based on a preset threshold. Then, the module automatically invokes the corresponding data fusion or compensation algorithm based on the determination result: for no occlusion, visual data is prioritized; for partial occlusion, a multi-sensor fusion algorithm based on an extended Kalman filter is activated, and the weights of each sensor are dynamically adjusted; for complete occlusion, a prediction mode based on the robotic arm's dynamics model is switched, and force feedback is introduced for motion compensation and data reconstruction. The module ultimately outputs processed, continuous, high-precision 7-DOF pose data of the robotic arm, and can issue prediction compensation commands to the isomorphic arm operation module as needed, forming an adaptive closed-loop control.

[0012] The present invention discloses a method for acquiring robotic arm operation data based on multi-source sensing, comprising the following steps:

[0013] (1) Synchronous motion control of isomorphic robotic arms: After the system is started, the main control isomorphic arm establishes a communication connection with the execution UR5 robotic arm to perform zero-position calibration and motion parameter configuration. The operator drives the UR5 robotic arm to perform grasping tasks in real time by controlling the main control arm. The system ensures motion synchronization and control accuracy through a high-speed data link.

[0014] (2) Multi-view visual image acquisition: Three industrial cameras are placed at the end of the robotic arm, above and to the side of the working area to acquire 640×480 resolution, 30Hz image data in a hardware synchronous manner. Each frame of the image is accompanied by a high-precision timestamp and is transmitted to the processing host in real time via gigabit network port for target detection, pose estimation and process recording.

[0015] (3) Optical motion capture data acquisition: A capture network consisting of 8 infrared cameras forms a 360° coverage around the working area; dedicated reflective markers are placed on key parts of the robotic arm and the operated object, with at least 3 non-collinear markers placed on each rigid body; the system acquires the marker images synchronously through multiple cameras, calculates the three-dimensional coordinates using triangulation, generates 7-DOF pose data in real time, and transmits the pose data to the processing host in real time through the NatNet protocol.

[0016] (4) Force perception data acquisition: The force / torque change data during the operation process is collected in real time through the six-dimensional force sensor installed on the end effector of the UR5 robotic arm; the sensor data is transmitted to the processing host through the real-time industrial Ethernet protocol, and the sampling frequency is not less than 1kHz to ensure that the transient changes of contact force can be captured.

[0017] (5) Occlusion detection and judgment: The system analyzes the number of feature points, matching success rate and contour integrity in each camera image in real time, and performs multi-dimensional occlusion judgment by combining the visibility and distribution status of marker points in the motion capture system.

[0018] For visual occlusion detection: ORB feature points are extracted in real time from images captured by each camera. If the number of feature points detected by any camera is less than 30, or the feature point matching success rate between the three views is less than 40%, or the outline integrity of the target object is less than 60%, a visual occlusion flag is triggered. For motion capture occlusion verification: The visible proportion of passive reflective markers on the target object is monitored in real time. If the visible proportion of the main markers is less than 80%, or there is an abnormal change in the coplanarity of the spatial distribution of markers, a motion capture occlusion flag is triggered. For force anomaly detection: The readings of the six-dimensional force sensor of the UR5 end effector are monitored. If an unexpected sudden change in contact force is detected and the force value exceeds the dynamic threshold (default setting is 5N), a force anomaly flag is triggered. The final occlusion determination adopts a voting mechanism: if both the visual and motion capture systems report a fault, it is directly determined that occlusion has occurred; if only a single system reports a fault, a third-level verification is performed in conjunction with force data, and occlusion is determined after confirmation.

[0019] (6) Data acquisition strategy switching: If there is no occlusion, the system mainly uses visual data for 3D modeling and pose estimation; if partial occlusion is detected, the multi-sensor fusion strategy is activated, which fuses visual, motion capture and force data through extended Kalman filter and dynamically adjusts the weight of each sensor; if there is complete occlusion, the system switches to the prediction mode based on the kinematic model of the robotic arm and introduces force feedback for motion compensation and data reconstruction.

[0020] (7) Multi-source data fusion and output: Through hardware synchronization and coordinate unification processing, visual, motion capture and force data are spatiotemporally registered, and the time synchronization error is controlled within 2ms. The system outputs fused object pose data, operation command data and multimodal sensing data for subsequent analysis, learning or control.

[0021] (8) Abnormal recovery and mode switching: When the sensor data returns to normal for 5 consecutive frames, the system automatically switches back to the standard vision-dominated mode to ensure the continuity and stability of data acquisition.

[0022] Technical effects of the present invention:

[0023] (1) High robustness and anti-occlusion capability: Through multi-source perception and hierarchical processing strategies, the system can still maintain data integrity in occlusion environment, significantly improving its adaptability in complex operation scenarios.

[0024] (2) High-precision synchronization of multimodal data: Hardware synchronization and unified coordinate transformation are adopted to ensure the consistency of visual, motion capture and force data in time and space, providing high-quality input for subsequent algorithms.

[0025] (3) Adaptive data fusion and compensation mechanism: The system can automatically switch the data processing mode according to the degree of occlusion, and combine the dynamic model and force feedback to perform data prediction and compensation, thereby improving data reliability.

[0026] (4) Applicable to deep learning and reinforcement learning: It provides rich, multimodal, and highly synchronous embodied datasets to support online learning and optimization of robot operation skills and promote the application of intelligent robotic arms in real-world scenarios. Attached Figure Description

[0027] Figure 1 This is a diagram showing the overall architecture of the data acquisition system according to an embodiment of this application.

[0028] Figure 2 This is a schematic diagram of the data acquisition process in an embodiment of this application.

[0029] Figure 3 This is a flowchart illustrating the isomorphic robotic arm operation module according to an embodiment of this application.

[0030] Figure 4 This is a flowchart illustrating the multi-view visual image acquisition module according to an embodiment of this application.

[0031] Figure 5 This is a flowchart illustrating the optical motion capture module according to an embodiment of this application.

[0032] Figure 6 This is a flowchart illustrating the occlusion detection and data fusion module in an embodiment of this application. Detailed Implementation

[0033] The technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings.

[0034] Please refer to Figure 1 This invention provides a robotic arm operation data acquisition platform and method based on multi-source perception, including: a homogeneous robotic arm operation module, a multi-view visual image acquisition module, an optical motion capture module, and an occlusion detection and data fusion module.

[0035] The isomorphic robotic arm operation module is used to control the movement of the UR5 robotic arm, including the isomorphic arm as the main control robotic arm, the UR5 as the execution robotic arm, the Robotiq 2F-140 gripper, and a lightweight gripper printed by 3D.

[0036] The multi-view visual image acquisition module is used to capture images from three perspectives during the grasping process of the robotic arm, including three industrial cameras.

[0037] The optical motion capture module is used to supplement the missing data when there is obstruction. It includes eight PrimeX13 infrared cameras arranged around the work area.

[0038] The occlusion detection and data fusion module is used to determine whether the robotic arm experiences occlusion during the grasping process, resulting in partial data loss, including the host computer and the software program preset on the host computer.

[0039] Further, please refer to Figure 2 This demonstrates the data acquisition process described in this invention. The process begins with a isomorphic robotic arm controlling the synchronous movement of the UR actuator arm and simultaneously activating three cameras to acquire video data. Subsequently, the visual data is analyzed in real time to determine whether occlusion has occurred. If no occlusion is detected, the object pose is calculated using the visual and robotic arm data. If occlusion occurs, the process switches to capturing marker points through an optical motion capture system to supplement the missing data caused by occlusion. Finally, the visual, optical motion capture, and robotic arm data are fused together to output the final object pose information.

[0040] Further, please refer to Figure 3 The isomorphic robotic arm operation module uses the isomorphic arm as the master control end and the UR5 robotic arm as the execution end. Upon system startup, a communication connection is first established between the master control robotic arm and the execution robotic arm. The master control end isomorphic arm connects via USB, while the execution end UR5 robotic arm uses a real-time industrial Ethernet protocol to achieve high-speed data exchange between the two, ensuring timely transmission of control commands. After the operator moves the robotic arm to three specific reference positions as prompted, the system automatically calculates the zero-position compensation value and rotation direction parameters of each joint. This calibration data is encrypted and stored in the system's non-volatile memory to ensure long-term storage and secure use. The system adopts a master-slave control architecture. The motion data of the master control robotic arm, after coordinate transformation and parameter mapping, drives the execution robotic arm to complete the corresponding actions in real time. During control, the system dynamically adjusts motion parameters, including velocity curves and acceleration limits, to ensure the smoothness and accuracy of the movements. The gripper control uses a special normalization processing algorithm to achieve precise force control.

[0041] Further, please refer to Figure 4The multi-view visual image acquisition module uses three cameras arranged in three positions: a main operating camera (located at the end of the UR5 robotic arm), a global scene camera (viewing the entire workspace from above), and a third-view camera (viewing from the side to observe the front of the object being grasped). The main operating camera, located at the end of the UR5 robotic arm, captures the interaction between the robotic arm and the object, providing high-resolution object details for target detection, pose estimation, or motion supervision. The global scene camera, viewing the entire workspace from above, monitors the global state of the working environment, such as the initial position of the object. The third-view camera observes the entire process of the robotic arm grasping the object from the side, recording the complete grasping process for post-analysis or reward calculation in reinforcement learning. All cameras are uniformly set to a resolution of 640×480 and a acquisition frequency of 30Hz. Each frame is accurately timestamped (accuracy 10μs). Image transmission uses the GVSP protocol and is transmitted in real-time to the processing host via a gigabit network port.

[0042] Further, please refer to Figure 5 The optical motion capture module uses eight PrimeX13 infrared cameras arranged around the working area to form a 360° coverage network. The camera installation height and angle must ensure no blind spots. Passive reflective markers with a diameter of 6-12mm are affixed to the surfaces of each joint and end effector of the robotic arm. Each rigid body has at least three non-collinear markers, with redundant markers added to critical areas. An asymmetrical arrangement is used. After camera calibration and coordinate system establishment, motion capture can begin. Once the system is started, images of the markers are simultaneously acquired by multiple cameras, and three-dimensional coordinates are calculated using triangulation. Finally, in the output stage, high-precision 7-DOF pose data of the object is generated in real time and transmitted to the computer in real time via the NatNet protocol.

[0043] Further, please refer to Figure 6The occlusion detection and data fusion module extracts feature points from images captured by each camera. A primary occlusion flag is triggered when any of the following conditions occur: 1) fewer than 30 feature points are detected by any camera; 2) the three-view feature point matching success rate is less than 40%; 3) the outline integrity of the target object (calculated through edge detection) is less than 60%. Simultaneously, cross-validation of the motion capture system is performed: real-time monitoring of the visibility index of marker points on the target object. A secondary occlusion flag is triggered when the visibility ratio of the primary marker points is less than 80%, or when the spatial distribution of marker points is abnormal (e.g., a sudden change in coplanarity). The final occlusion determination uses a voting mechanism: when both the vision system and the motion capture system report a fault, the system immediately switches to occlusion mode; when only a single system reports a fault, a third-level verification is initiated—detecting sudden changes in contact force using the torque sensor of the robotic arm. If an unexpected force change is detected (the threshold is dynamically adjusted according to the load, typically set to 5N), occlusion is confirmed. When the system detects occlusion, it first achieves spatiotemporal registration of the vision system, optical motion capture system, and force sensor through hardware synchronization and coordinate transformation, ensuring that the time synchronization error is less than 2ms and the spatial coordinates are unified. A tiered processing strategy is adopted based on the degree of occlusion: for partial occlusion, visual feature point and optical marker point data are fused using an extended Kalman filter, and an adaptive weighting algorithm is used to dynamically adjust the contribution of each sensor; for complete occlusion, it switches to a prediction mode based on the robot's kinematics model, while introducing force feedback for motion compensation. The system is equipped with an abnormal data filtering mechanism and low-pass filtering smoothing; when sensor data returns to normal for 5 consecutive frames, it automatically switches back to the vision-driven standard tracking mode.

[0044] Example: Block grabbing task

[0045] Taking the example of an operator remotely manipulating the UR5 robotic arm to grasp a cube object on the control panel using a homogeneous arm, the implementation process of this system is explained in detail:

[0046] (1) Task preparation stage: The operator places a 5cm side cube in the center of the operating table. A total of 16 passive reflective markers with a diameter of 12mm are affixed to the base, joint links and end effector surfaces of the UR5 robotic arm to form an asymmetrical layout. The system modules are started to complete the zero-position calibration of the robotic arm, camera calibration and motion capture system coordinate alignment, and the occlusion detection threshold parameters are set (feature point number threshold 30, matching success rate threshold 40%, contour integrity threshold 60%).

[0047] (2) Remote operation execution stage: The operator controls the UR5 robotic arm to move towards the cube by dragging the main control robotic arm. Three industrial cameras synchronously acquire images at a frequency of 30Hz, eight infrared cameras track the robotic arm markers at a frequency of 120Hz, and a six-dimensional force sensor acquires contact force data at a frequency of 1kHz. The isomorphic arm records the operator's control actions in real time.

[0048] (3) Occlusion Handling Example: When the UR5 robotic arm's end effector approaches the cube, the robotic arm itself visually occludes the global scene camera. The system detects that the number of feature points at the robotic arm's end effector in the global camera drops to 25, and the cube's outline integrity drops to 55%, triggering a primary occlusion flag. Simultaneously, the motion capture system shows that the visible percentage of marker points at the robotic arm's end effector drops from 95% to 45%, triggering a secondary occlusion detection flag. The system initiates a third-level verification. The force sensor reading does not exceed the threshold (1.2N), indicating partial occlusion. The system then activates an extended Kalman filter, fusing visual data from the main operating camera, the third-view camera, and the end effector camera with motion capture data, dynamically adjusting the weights of each sensor, and outputting continuous robotic arm pose information.

[0049] (4) Data fusion and output: The system controls the spatiotemporal error of multi-source data to within 2ms through hardware time synchronization and coordinate unification, and outputs the fused UR5 robotic arm 7-DOF pose data, operation isomorphic arm motion data, force data and synchronized image data. When the robotic arm completes the grasping and picking up, the visual occlusion is removed, and the system detects that the data is normal for 5 consecutive frames, and automatically switches back to the vision-dominated mode.

[0050] (5) Data saving and format: After the capture task is completed, the system saves the complete operation dataset in PyTorch tensor format, which includes: 1) Image observation data: RGB image sequence acquired by three-view industrial camera, encoded as MP4 video file, with a resolution of 640x480 and a frame rate of 30Hz; 2) Robotic arm status data: 7-DOF pose data (position, orientation, gripper status) of UR5 robotic arm; 3) Operation command data: motion trajectory and control action data of the master control isomorphic arm; 4) Force sensing data: 6-dimensional force / torque data of UR5 robotic arm end effector, with a sampling frequency of 1kHz; 5) Time series data: frame index, episode index, timestamp, episode end marker; 6) Metadata information: data acquisition frequency, coordinate information, data format version; 7) Occlusion processing log: occlusion detection flag, sensor fusion strategy, data compensation record.

[0051] Although this example uses block grasping as an example, the application scenarios of this system are not limited to this. By adjusting the task configuration, it can be applied to a variety of complex operation scenarios, including but not limited to: 1) Part assembly tasks: collecting interaction data between the robotic arm and small parts during precision assembly; 2) Flexible object manipulation: grasping and manipulating easily deformable objects; 3) Multi-object collaborative operation: synchronous data acquisition when multiple robotic arms work together; 4) Medical operation simulation: simulating the operation of surgical instruments and acquiring high-precision force and visual data.

Claims

1. A robotic arm operation data acquisition system based on multi-source sensing, characterized in that, It includes a homogeneous robotic arm operation module, a multi-view visual image acquisition module, an optical motion capture module, and an occlusion detection and data fusion module; The isomorphic robotic arm operation module consists of a UR5 robotic arm and a scaled-down isomorphic arm with the same structure as the UR5 robotic arm. The isomorphic arm serves as the master operating arm, and the UR5 robotic arm serves as the slave arm. The UR5 robotic arm is equipped with an adaptive gripper and a high-precision six-dimensional force sensor at its end to sense changes in force and torque during operation in real time. The isomorphic arm uses a 3D-printed model gripper at its end to precisely control the opening and closing of the UR5 end gripper. The motion data of the master operating arm is used as control commands to the slave arm to drive it to reproduce the operation. At the same time, the six-dimensional force sensor at the end of the slave arm collects real-time force / torque data as physical feedback, which is fed back to the control system and uploaded to the occlusion detection and data fusion module. The multi-view visual image acquisition module employs three high-resolution industrial cameras, respectively positioned at the end of the robotic arm, above the work area, and to the side of the operating plane, forming an all-around observation network. All cameras are synchronously triggered via hardware to ensure the consistency of data acquisition timing and are equipped with professional optical components to eliminate ambient light interference. The module uses standard calibration methods to ensure the measurement accuracy of each camera. The synchronous image stream acquired by the three cameras is transmitted in real time to the processing host via a high-speed network, serving as the raw visual data stream for the occlusion detection and data fusion module, used for target recognition and pose estimation. The optical motion capture module consists of multiple high-performance infrared cameras, with dedicated reflective markers placed on key parts of the robotic arm and the object being manipulated. The module undergoes precise calibration to track the three-dimensional motion trajectory of the robotic arm and the object in real time. Through multi-sensor data fusion technology, the module automatically adjusts its data acquisition strategy according to environmental changes, ensuring complete and accurate operational data under various working conditions. The infrared camera array captures images of the markers, and triangulation calculations generate high-precision six-degree-of-freedom pose data of the robotic arm and the target object, serving as a true reference for spatial positioning, which is then transmitted in real time to the occlusion detection and data fusion module. The occlusion detection and data fusion module employs a multi-dimensional occlusion detection mechanism, comprehensively analyzing the number of visual feature points, motion capture markers, and force sensor data to determine the occlusion situation. It determines the degree of occlusion based on preset thresholds: no occlusion, partial occlusion, or complete occlusion. Then, based on the determination result, it automatically calls the corresponding data fusion or compensation algorithm: when there is no occlusion, visual data is prioritized; when there is partial occlusion, a multi-sensor fusion algorithm based on an extended Kalman filter is activated, and the weights of each sensor are dynamically adjusted; when there is complete occlusion, it switches to a prediction mode based on the robotic arm's dynamics model and introduces force feedback for motion compensation and data reconstruction. Finally, it outputs processed, continuous, high-precision 7-DOF pose data of the robotic arm, and can issue prediction compensation commands to the isomorphic arm operation module as needed, forming an adaptive closed-loop control.

2. The data acquisition method of the robotic arm operation data acquisition system based on multi-source sensing as described in claim 1, characterized in that, Includes the following steps: (1) Synchronous motion control of isomorphic robotic arm: After the system is started, the main control isomorphic arm establishes a communication connection with the execution UR5 robotic arm to perform zero-position calibration and motion parameter configuration; the operator drives the UR5 robotic arm to perform grasping tasks in real time by controlling the main control arm. The system ensures motion synchronization and control accuracy through a high-speed data link. (2) Multi-view visual image acquisition: Three industrial cameras are arranged at the end of the robotic arm, above and to the side of the work area to acquire image data with a resolution of 640×480 and a frequency of 30Hz in a hardware synchronous manner. Each frame of the image is accompanied by a high-precision timestamp and is transmitted to the processing host in real time via a gigabit network port for target detection, pose estimation and process recording. (3) Optical motion capture data acquisition: A capture network consisting of 8 infrared cameras forms a 360° coverage around the working area; dedicated reflective markers are placed on key parts of the robotic arm and the operated object, with at least 3 non-collinear markers placed on each rigid body; the system acquires the marker images synchronously through multiple cameras, calculates the three-dimensional coordinates using triangulation, generates 7-DOF pose data in real time, and transmits the pose data to the processing host in real time through the NatNet protocol; (4) Force perception data acquisition: The force / torque change data during the operation process is collected in real time through the six-dimensional force sensor installed on the end effector of the UR5 robotic arm; the sensor data is transmitted to the processing host through the real-time industrial Ethernet protocol, and the sampling frequency is not less than 1kHz to ensure that the transient changes of contact force can be captured. (5) Occlusion detection and judgment: The system analyzes the number of feature points, matching success rate and contour integrity in each camera image in real time, and performs multi-dimensional occlusion judgment by combining the visibility and distribution status of marker points in the motion capture system. For visual occlusion detection: ORB feature points are extracted in real time from images captured by each camera. If the number of feature points detected by any camera is less than 30, or the feature point matching success rate between the three views is less than 40%, or the outline integrity of the target object is less than 60%, a visual occlusion flag is triggered. For motion capture occlusion verification: The visible proportion of passive reflective markers on the target object is monitored in real time. If the visible proportion of the main markers is less than 80%, or there is an abnormal change in the coplanarity of the spatial distribution of markers, a motion capture occlusion flag is triggered. For force anomaly detection: The readings of the six-dimensional force sensor of the UR5 end effector are monitored. If an unexpected sudden change in contact force is detected and the force value exceeds the dynamic threshold (default is 5N), a force anomaly flag is triggered. The final occlusion determination adopts a voting mechanism: If both the visual and motion capture systems report a fault, it is directly determined that occlusion has occurred; if only a single system reports a fault, a third-level verification is performed in conjunction with force data, and occlusion is determined after confirmation. (6) Data acquisition strategy switching: If there is no occlusion, the system mainly uses visual data for 3D modeling and pose estimation; if partial occlusion is detected, the multi-sensor fusion strategy is activated, which fuses visual, motion capture and force data through extended Kalman filter and dynamically adjusts the weights of each sensor. If completely occluded, the system switches to a prediction mode based on the robotic arm's kinematics model and introduces force feedback for motion compensation and data reconstruction. (7) Multi-source data fusion and output: Through hardware synchronization and coordinate unification processing, visual, motion capture and force data are spatiotemporally registered, and the time synchronization error is controlled within 2ms; the system outputs fused object pose data, operation command data and multimodal sensing data for subsequent analysis, learning or control. (8) Abnormal recovery and mode switching: When the sensor data returns to normal for 5 consecutive frames, the system automatically switches back to the standard vision-dominated mode to ensure the continuity and stability of data acquisition.

Citation Information

Cited By

  • Load device cooperative control system and method based on visual identification

    CN122115426A

  • A load device cooperative control system and method based on visual recognition

    CN122115426B

  • Head-mounted multi-mode vision method and system for intelligent data acquisition of body

    CN122200269A