The invention discloses a multi-
modal assembly action recognition method for comparative
semantic query, and relates to the technical field of man-
machine cooperation assemblation.The method comprises the steps that a visual sensor is arranged on an
assembly workbench to obtain an operator action video, a sampling
frame sequence is obtained through random frame sampling, a skeleton sequence is obtained through
human body posture
estimation, and a skeleton sequence is obtained through
human body posture
estimation; and inputting an
assembly action recognition model to complete recognition. The model comprises an image coding module, a skeleton coding module, a
feature fusion module, a text coding module and a semantic comparison module which are used for extracting image and skeleton features, fusing features, coding preset category text description, comparing action features with category text features and outputting a result with the highest similarity, and a comparison
loss function is adopted during training. According to the method, multi-
modal information is fused, the problems of single-
modal limitation and multi-modal semantic segmentation are solved, category text
semantics are fully utilized, the fine-grained
action recognition precision is improved, the over-fitting risk is reduced, and the generalization and task migration ability of the model in a dynamic industrial scene is enhanced.