This invention relates to the field of
artificial intelligence, specifically disclosing a method,
system, and medium for converting model images into 3D flat lay images based on physical
perception. The method acquires a
model image and text instructions, uses a prediction network to obtain a semantic
mask and a 2.5D normalized deformation field, then extracts pure clothing content features based on the semantic
mask, and uses the deformation field to reconstruct the spatial relationships of the content features, eliminating
pose deformation in the feature space to obtain flattened clothing features. Finally, using the text instructions and flattened features as conditions, a high-fidelity flat lay image is generated through parallel decoding. This invention innovatively introduces a 2.5D physical
perception mechanism and self-
supervised training, effectively solving the problems of feature cross-
contamination, pattern
distortion, and generation speed bottlenecks, achieving high-speed, high-definition automatic generation of clothing flat lay images end-to-end without
manual annotation.