The invention relates to an illumination
perception video generation method and device based on renderer proxy reasoning and a storage medium. A user is supported to perform accurate and decoupling control on geometric
layout, illumination conditions and camera tracks in a video
generation process through a
natural language instruction. A renderer agent is introduced to convert a text into a structured three-dimensional scene parameter, and a rendering engine is utilized to generate a two-dimensional scene agent comprising a
diffuse reflection, gloss and roughness multi-channel layer; then, through a lightweight proxy
encoder and an adapter, the physical illumination attribute is used as a strong condition
signal to be injected into the video
diffusion model. According to the method, the end-to-end generation from the text to the physical consistency video is realized, the generated video has accurate shadow, reflection and ambient light shielding effects while keeping realistic visual textures, the control precision and the
automation degree of scene physical attributes are improved, and the method is suitable for the fields of film and television rehearsal, game asset production, virtual film production and the like.