The invention discloses a multi-
modal visual language navigation method based on a dynamic environment
knowledge graph and related equipment, and can be applied to the technical field of
artificial intelligence. After the real-time environment image of the area to which the target
robot belongs is captured through the
single camera,
feature extraction is carried out on the real-time environment image to obtain the image features, the to-be-processed category corresponding to the target object is coded to obtain the
semantic code, and after the spatial features of the target object are extracted, the target
robot is obtained. Constructing a dynamic environment
knowledge graph according to the confidence, semantic coding and spatial features of the target object in combination with an external
knowledge base, extracting high-dimensional graph features of the dynamic environment
knowledge graph, and performing implicit fusion and display modeling in combination with image features and language features corresponding to the target instruction to obtain a navigation
state vector containing long-
term memory; therefore, the moving operation of the target
robot can be controlled based on the navigation
state vector, efficient and explainable navigation reasoning is further realized, and the language navigation accuracy is improved at relatively low cost.