The invention discloses an image adjustment 3D
content creation architecture based on 2D
diffusion, and relates to the technical field of
diffusion models, and the architecture comprises a
diffusion model which is composed of a plurality of cross attention blocks in a U-Net structure and supports effective fusion of various
modes of texts, images and camera parameters. In the present invention, it is devoted to create 3D content using a potential diffusion model. The 3D geometry is not generated directly through a potential diffusion model, but a two-stage formula is employed. First, a potential diffusion model of view
conditioned reflex is organized, and a multi-view image is synthesized with a
monocular image and the outside of a camera as inputs. Next, a neural
radiation field is trained using the synthesized multi-view image, which is easily optimized as
volume rendering is differentiable. After the neural
radiation field training is completed, a 3D geometric model is generated through an advancing cube
algorithm applied to a
density field. The framework eliminates the requirement for
pairing 3D training data and does not require a large amount of computing resources.