场景文字处理方法、装置、设备及产品
By using a multimodal signal feature encoding and a diffusion model superposition reference position encoding mechanism, the problem of inaccurate text spatial layout and style control in existing technologies is solved, achieving high-quality scene text processing, eliminating ghosting artifacts, and ensuring background consistency.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- PEKING UNIV SHENZHEN GRADUATE SCHOOL
- Filing Date
- 2026-05-07
- Publication Date
- 2026-07-17
AI Technical Summary
Existing scene text editing technologies struggle to achieve precise joint control over text spatial layout and visual typography style. Furthermore, scene text replacement tasks often result in ghosting artifacts or damage to background textures, making it difficult to obtain high-quality processing results.
By acquiring multimodal control signals, feature encoding is performed using a preset variational autoencoder and a visual language model. Combined with the superposition reference position encoding mechanism and the region adaptive suppression mechanism in the preset diffusion model, decoupled control and efficient fusion processing of text spatial layout and style are achieved.
It achieves high-quality, natural scene text processing results, eliminates ghosting artifacts, ensures the consistency of background texture, and completes complex text editing tasks in one go.
Smart Images

Figure CN122156390B_ABST