This invention discloses a method for
autonomous robot operation in industrial
assembly, relating to the field of
robot control technology. The method is applied to a
system containing multiple RGB-D cameras, a
robotic arm, and an
edge computing platform. It simultaneously acquires images and depth information from multiple perspectives, unifying them to a base coordinate
system through coordinate reconstruction. Based on dual-view deviation and a fusion threshold, it determines whether to merge point clouds or
downgrade single-view images, and performs PCA-OBB solving to obtain target geometric parameters and generate a grasping
pose. It receives
natural language commands, generates
voxel operation code from a large
language model to construct a 3D
voxel cost map, and plans the path using a weighted A*
algorithm. During
robotic arm operation, it calculates
pose error in real time and achieves closed-
loop control through judgment, PID correction, or local replanning. This invention solves problems such as poor robustness of multi-view
perception, poor real-time performance of large models, and large
assembly errors, improving the efficiency and stability of autonomous operation in industrial
assembly, and is suitable for complex assembly scenarios.