Scene text processing method, apparatus, device, and product

CN122156390AActive Publication Date: 2026-06-05PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
PEKING UNIV SHENZHEN GRADUATE SCHOOL
Filing Date
2026-05-07
Publication Date
2026-06-05

AI Technical Summary

Technical Problem

Existing text editing technologies struggle to achieve precise control over the spatial layout of text and the visual typography style. Furthermore, ghosting artifacts or damage to background textures often occur during text replacement, resulting in unnatural processing results.

Method used

By acquiring multimodal control signals, using a preset variational autoencoder and a visual language model for feature encoding, and combining the superposition reference position encoding mechanism and region adaptive suppression mechanism in the preset diffusion model, decoupling control and feature fusion of text spatial layout and style are achieved to generate high-quality target text images.

Benefits of technology

It achieves precise joint control over text spatial layout and visual typography style, eliminates ghosting artifacts, and generates natural and high-quality scene text processing results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122156390A_ABST
    Figure CN122156390A_ABST
Patent Text Reader

Abstract

The application discloses a scene text processing method and device, equipment and product, and relates to the technical field of data processing. The method comprises the following steps: acquiring a multi-modal control signal of a to-be-edited text image, wherein the multi-modal control signal comprises a source image, an erasing mask image, a structure guide image, a style reference image and a target editing instruction; performing feature coding on the multi-modal control signal based on a preset variational autoencoder and a visual language model to obtain corresponding mixed latent features; and performing fusion processing on the mixed latent features based on a preset diffusion model to generate an edited target text image, wherein the preset diffusion model comprises a plurality of diffusion model modules connected in series, and each diffusion model module comprises a superimposed reference position coding mechanism and a region adaptive suppression mechanism. The application realizes decoupling of space layout control and style control of the to-be-edited text image through the above scheme, and can output a high-quality and natural scene text processing result at one time.
Need to check novelty before this filing date? Find Prior Art

Citation Information

Patent Citations

  • Stylized visual text editing method, system and equipment and storage medium

    CN121095395A

  • Techniques for generating images of object interactions

    US20240161468A1