The invention relates to the field of
robot perception and intelligent man-
machine interaction, an equipment platform and a
perception module, a
system is provided with an Intel RealSense D435i depth camera which is used for acquiring an RGB-D
image sequence containing depth information in real time, and meanwhile, an advanced YOLOv8-
Pose deep learning model is embedded, so that the depth information of the RGB-D
image sequence is acquired in real time. High-precision detection,
natural interaction, a gesture guidance mechanism, direction
inference and target candidate aggregation, interactive target screening and judgment, a robustness enhancement strategy, deep deletion and
noise robustness and an incomplete target completion mechanism of human upper body skeleton key points in a front scene are realized; the man-
machine collaborative target object identification and positioning method based on depth vision, attitude
estimation and semantic understanding is mainly realized. The technology can be suitable for service robots, intelligent assistant robots and other practical application scenes with high requirements for autonomous
perception and
natural interaction, and man-
machine co-fusion and
intelligent environment construction are further promoted.