The invention discloses a multi-
modal condition-driven closed-loop sensing and generation optimization method and device and a medium, which are applied to an automatic driving long-
tail scene, firstly, multi-
modal conditions such as text description, a
semantic map and a 3D
layout are uniformly represented, a
causal inference network is introduced, a causal structure between scene elements is explicitly inferred from multi-
modal input, and a multi-modal condition-driven closed-loop sensing and generation optimization model is obtained. Generating structured causal embedding to realize causal
perception enhancement of generation conditions; in a
diffusion generation stage, a causal consistency mechanism is deeply fused into a condition control and denoising process, and a
diffusion model is guided to effectively inhibit generation of unreasonable or common sense violating scenes in a
generation process through an anti-fact condition constructed by a
causal inference network, so that semantic reasonability,
spatial consistency and dynamic credibility of long-
tail data are remarkably improved. According to the method, the problem that long-
tail scene data is deficient is solved, causal constraints are introduced to a generation source, and the reliability, robustness and cross-domain generalization ability of an automatic driving
system in an extreme scene are remarkably improved.