This invention discloses a method for generating poses of articulated objects based on physical
perception and graph
diffusion, relating to the field of
computer vision. The method first reconstructs the componentized geometric and physical properties of objects from images using a bi-
branch neural implicit network and initializes component
connectivity relationships. Subsequently, it refines the component relationship graph through kinematic fitting and
temporal consistency checks, and uses a physically enhanced graph
diffusion process to infer the prior
pose distribution of components on an SE(3) manifold. Finally, using this prior and reconstructed information as conditions, a conditional
diffusion model is constructed on the SE(3) manifold, and backsampling is performed through a physically guided two-step backsampling framework to generate a diverse and physically plausible set of
pose assumptions for articulated objects. This invention achieves efficient generation of diverse and highly physically plausible poses of articulated objects by tightly
coupling physical laws with data-driven generation.