A semantic picture editing method based on diffusion model
By constructing a text encoder and an image encoder to generate image masks, the problem of lack of semantic information utilization in diffusion models during image editing is solved. This enables automatic semantic editing without user mask input, improving user experience and the model's generalization ability.
CN117671084BActive Publication Date: 2026-07-24HEFEI UNIV OF TECH
View PDF 2 Cites 0 Cited by
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- HEFEI UNIV OF TECH
- Filing Date
- 2023-12-13
- Publication Date
- 2026-07-24
AI Technical Summary
Technical Problem
Existing diffusion models lack semantic information utilization in image editing, require users to provide mask input, may discard important information during the editing process, and tend to modify the entire image.
Method used
We construct a text encoder and an image encoder. The text encoder embeds labels into a vector space to generate an image mask. We then use a denoising diffusion probability model for encoding and decoding to generate a new image.
Benefits of technology
It enables automatic identification and semantic editing of image editing regions without user mask input, improving user operation efficiency. The model can generalize to unseen categories without retraining.
✦ Generated by Eureka AI based on patent content.
Smart Images

Figure CN117671084B_ABST
Abstract
The application discloses a semantic picture editing method based on a diffusion model, belongs to the technical field of computer vision and artificial intelligence, and comprises the following steps: constructing a text encoder and an image encoder, embedding labels into a vector space through the text encoder, calculating through the image encoder and the labels, and generating a picture mask; constructing a denoising diffusion probability model, encoding an input image based on the denoising diffusion probability model, and obtaining latent features corresponding to the input image; performing decoding on the denoising diffusion probability model based on the picture mask and the latent features corresponding to the input image, obtaining an inferred mask; and replacing the background of the input image with pixel values in the encoding process based on the inferred mask and mapping the pixel values back to original pixels to obtain a new picture. The application discloses a semantic-guided picture editing method based on a diffusion model. According to the method, which regions of an input image should be edited can be automatically found according to a text query.
Need to check novelty before this filing date? Find Prior Art