The invention relates to the technical field of
artificial intelligence, can be applied to business scenes such as intelligent, financial science and technology,
medical health and the like, and discloses a multi-
modal fusion feature driven
action control method, device, equipment and medium, and the method comprises the steps: obtaining first
modal input information and second
modal input information, the method comprises the steps of generating a corresponding first modal
feature vector and a corresponding second modal
feature vector, fusing the two modal feature vectors to generate a multi-modal fusion feature, generating an action instruction based on the multi-modal fusion feature, generating an initial
action plan in combination with the current state of equipment, the current environment information and a task target, and performing
motion planning. And generating a
global optimal action sequence based on the initial
action plan, and controlling the equipment to execute the
global optimal action sequence. According to the method, the
global optimal action sequence is generated and equipment execution is controlled by combining the action instruction generated by the multi-modal fusion information with the
equipment state, the environment information and the task target, so that the decision-making effect and the dynamic
adaptive capacity in a complex environment are improved.