The invention discloses a multi-
modal fusion body-equipped intelligent
system in a complex scene and a use method, and belongs to the technical field of
artificial intelligence and robots. Aiming at the problems of
perception lag, task splitting and incapability of understanding an open instruction of an existing
intelligent agent in a dynamic complex environment, a complete
perception-decision-execution
closed loop is constructed by integrating core modules such as multi-
modal perception, open vocabulary target detection, dynamic
semantic mapping,
natural language instruction analysis and multi-
modal decision. According to the
system, vision-language embedding, cross-modal attention fusion and an intelligent
obstacle avoidance mechanism are adopted, multi-dimensional understanding and self-adaptive path planning of the environment are achieved, and the performance of the model is optimized through a two-stage training strategy. According to the method, the accurate response, multi-task cooperative execution and dynamic
obstacle avoidance capability of the
intelligent agent to a
natural language instruction in an unknown scene are remarkably improved, and the method can be widely applied to the fields of military reconnaissance, disaster rescue,
urban security and protection and the like and has high practicability, adaptability and expandability.