This application belongs to the field of
artificial intelligence technology and relates to an AI-based
image processing method, apparatus,
computer device, and storage medium. The method includes: acquiring input text description and a
reference image; extracting features from the text description and
reference image using a hierarchical
encoder to obtain a text embedding corresponding to the text description and a
pose embedding corresponding to the
reference image; performing
semantic alignment processing on the text embedding and
pose embedding using a hierarchical alignment module to obtain alignment features; performing semantic optimization
processing on the alignment features using a cross-
modal adapter to obtain target features; performing
image generation processing corresponding to the target features using a
pose condition generator to obtain a target image; and outputting the target image. Furthermore, the target image of this application can be stored in a
blockchain. This application can be applied to text-to-image scenarios in the financial and medical fields, achieving precise pose control from text to image and improving the quality of the generated target image.