The invention discloses a self-adaptive
diffusion image editing method and
system based on concept attention, and the method comprises the following steps: constructing a
paired data set; analyzing the editing instruction, and extracting a key concept; a pre-trained T5
language model is utilized to convert the key concept into text embedding, and the text embedding is mapped to an image feature space; modifying a
diffusion model based on a Transform architecture, embedding a concept attention module in an attention layer of a multi-
modal diffusion converter, calculating an attention
score between image features and concept embedding, and generating a concept
saliency map; in the denoising process, the weight of the target area is adjusted by using the concept
saliency map so as to realize accurate editing. According to the method, under the condition that the global
image quality is not affected, the editing precision can be improved, interference to a non-target area is reduced, and meanwhile,
reinforcement learning and real-time feedback are combined, so that the model can be adaptively optimized, and an editing result better meeting the user requirement is generated.