Stylized visual text editing method, system, device, and storage medium

By combining variational autoencoders and visual language models with diffusion models, style embedding information is extracted for fine-grained control, solving the problem of inconsistent text editing in complex scenarios and achieving text image generation with high readability and style consistency.

CN121095395BActive Publication Date: 2026-02-10UNIV OF SCI & TECH OF CHINA
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511655543.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-11-12
Publication Date
2026-02-10
Estimated Expiration
2045-11-12

AI Technical Summary

Technical Problem

Existing image text editing methods based on diffusion models suffer from insufficient understanding of text styles and glyphs when dealing with complex scenes, resulting in mismatched text with the image background, inconsistent style, and poor readability.

Method used

By mapping the input text image to a structured latent space, style embedding information is extracted and refined by combining a diffusion model to generate highly readable and style-consistent text images. A combination of variational autoencoder, visual language model and diffusion model is used to achieve preservation of the original text style or style transfer of reference image.

Benefits of technology

The generated text images exhibit high readability and style consistency in complex scenes, effectively maintaining or transferring the original text style, significantly improving the accuracy and visual coherence of text editing, and reducing data annotation costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121095395B_ABST
    Figure CN121095395B_ABST
Patent Text Reader

Abstract

The application discloses a style visual text editing method, system, device and storage medium, which are corresponding solutions, and the related solutions aim to solve the style consistency problem existing in the image text editing of the existing diffusion model, extract style embedding information from the visual features of the glyph image and the input text image, and input the style embedding information as an enhanced style condition into the diffusion model, so as to realize fine control of the diffusion process, so that the diffusion model can generate a text image with high readability and style consistency, and can realize the keeping of the original text style or style migration based on a reference image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision and graphics technology, and in particular to a stylized visual text editing method, system, device and storage medium. Background Technology

[0002] Image text editing is an important task in computer vision and graphics. Currently, mainstream image generation techniques include Generative Adversarial Networks (GANs) and Diffusion Models (DFMs).

[0003] For example, References 1 and 2 are schemes based on adversarial networks, while References 3 and 4 are schemes based on diffusion models.

[0004] Document 1: Qu, Y, et al.: Exploring stroke-level modifications for scene textediting. In: Proceedings of the AAAI Conference on Artificial Intelligence.vol. 37, pp. 2119–2127 (2023).

[0005] Reference 2: Su, T et al.: Scene style text editing. arXiv preprint arXiv:2304.10097 (2023).

[0006] Document 3: Chen, H et al.: Diffute: Universal text editing diffusion model. In: Thirty-seventh Conference on Neural Information Processing Systems (NeurIPS) (2023).

[0007] Document 4: Tuo, Y et al.: Anytext: Multilingual visual text generation and editing. In: ICLR (2024).

[0008] While existing GAN-based methods have made some progress, they have limitations when processing text with arbitrary fonts, sizes, and colors, resulting in poor readability and inconsistency with the surrounding background style.

[0009] While DFM-based methods can generate high-quality natural scene images, they still face challenges in generating high-quality text images with a consistent style. These models lack explicit style coordination mechanisms and have limited understanding of text styles and glyphs, leading to mismatches between the generated text and the image background. For example, the scheme provided in Reference 3 (called DiffSTE, a diffusion model specifically improved for text editing tasks, using two encoders to control the content to be replaced and the style to be preserved, respectively) improves the accuracy of generated text to some extent when handling complex scenes, but its style often visually mismatches with the styles of other text originally present in the image.

[0010] Therefore, the main technical problem that exists in existing technologies is: how to solve the technical defects of image text editing methods based on diffusion models when dealing with complex scenes, which result in inconsistent styles and poor readability due to insufficient understanding of text styles and glyphs.

[0011] In view of this, the present invention is hereby proposed. Summary of the Invention

[0012] The purpose of this invention is to provide a stylized visual text editing method, system, device, and storage medium that can generate text images with high readability and stylistic consistency, and can achieve the preservation of the original text style or style transfer based on reference images.

[0013] The objective of this invention is achieved through the following technical solution:

[0014] A stylized visual text editing method includes:

[0015] Step 1: Map the input text image to a structured latent space to obtain the image latent representation vector;

[0016] Step 2: Extract the replacement text from the input text instruction and render it as a glyph image. Extract visual features from both the glyph image and the input text image to obtain corresponding glyph image features and text image features. Combine these glyph image features, text image features, and text instruction prediction style embedding information. Alternatively, extract visual features from both the input text image and a given reference image to obtain corresponding reference image features and text image features. Combine these reference image features, text image features, and input text instruction prediction style embedding information. The text instruction is an instruction to replace the original text in the input text image with the replacement text.

[0017] Step 3: Obtain the text mask and combine it with the text mask to obtain the latent representation vector of the masked image. Input the style embedding information, the latent representation vector of the image, the text mask and the latent representation vector of the masked image into the diffusion model. The diffusion model will gradually predict that the original text will be replaced with the replacement text, while retaining the original text style or the style of the reference image.

[0018] A stylized visual text editing system, comprising:

[0019] Variational autoencoders are used to map input text images to a structured latent space to obtain image latent representation vectors.

[0020] A visual language model is used to extract replacement text from an input text instruction and render it as a glyph image. Visual features are extracted from both the glyph image and the input text image to obtain corresponding glyph image features and text image features. These glyph image features, text image features, and text instruction prediction style embedding information are then combined. Alternatively, visual features are extracted from both the input text image and a given reference image to obtain corresponding reference image features and text image features. These reference image features, text image features, and input text instruction prediction style embedding information are then combined. The text instruction is an instruction to replace the original text in the input text image with the replacement text.

[0021] The predicted text image output unit is used to obtain a text mask and combine it with the text mask to obtain the latent representation vector of the masked image. The style embedding information, the latent representation vector of the image, the text mask and the latent representation vector of the masked image are input into the diffusion model. The diffusion model gradually predicts the text image in which the original text is replaced with replacement text while retaining the style of the original text or the style of the reference image.

[0022] A processing device includes: one or more processors; and a memory for storing one or more programs;

[0023] When the one or more programs are executed by the one or more processors, the one or more processors implement the aforementioned method.

[0024] A readable storage medium storing a computer program that, when executed by a processor, implements the aforementioned method.

[0025] As can be seen from the technical solution provided by the present invention, by combining the visual features of glyph images and input text images, stylistic information (style embedding information) is extracted from them, and thus used as an enhanced style condition input to the diffusion model, thereby achieving fine control over the diffusion process. This enables the diffusion model to generate text images with high readability and style consistency, and to maintain the style of the original text or perform style transfer based on the reference image. Attached Figure Description

[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0027] Figure 1 A flowchart illustrating a stylized visual text editing method provided in an embodiment of the present invention.

[0028] Figure 2 This is a schematic diagram of the overall framework of a stylized visual text editing method provided in an embodiment of the present invention.

[0029] Figure 3 This is a schematic diagram illustrating the implementation effect in an experiment provided in this embodiment of the invention.

[0030] Figure 4 This is a schematic diagram illustrating the visual comparison results provided in an embodiment of the present invention.

[0031] Figure 5 This is a schematic diagram of a stylized visual text editing system provided in an embodiment of the present invention.

[0032] Figure 6 This is a schematic diagram of a processing device provided in an embodiment of the present invention. Detailed Implementation

[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the present invention.

[0034] First, the following explanations are provided for the terms that may be used in this article:

[0035] The terms "comprising," "including," "containing," "having," or other similar semantic descriptions should be interpreted as non-exclusive inclusion. For example, including a technical feature element (such as raw material, component, ingredient, carrier, dosage form, material, size, part, component, mechanism, device, step, process, method, reaction conditions, processing conditions, parameter, algorithm, signal, data, product or article of manufacture, etc.) should be interpreted as including not only the expressly listed technical feature element, but also other technical feature elements that are not expressly listed and are well-known in the art.

[0036] The term "composed of" excludes any technical features not expressly listed. When used in a claim, it closes the claim to exclude all technical features other than those expressly listed, except for associated conventional impurities. If the term appears only in a clause of a claim, it limits the claim to the elements expressly listed in that clause; elements recited in other clauses are not excluded from the overall claim.

[0037] The following provides a detailed description of a stylized visual text editing method, system, device, and storage medium provided by this invention. Contents not described in detail in the embodiments of this invention are prior art known to those skilled in the art. Where specific conditions are not specified in the embodiments of this invention, they are performed according to conventional conditions in the art or conditions recommended by the manufacturer. Where the manufacturers of the instruments used in the embodiments of this invention are not specified, they are all conventional products that can be purchased commercially.

[0038] Example 1

[0039] This invention provides a stylized visual text editing method, such as... Figure 1 As shown, it mainly includes the following steps:

[0040] Step 1: Map the input text image to a structured latent space to obtain the image latent representation vector.

[0041] This step can be implemented using a variational autoencoder (VAE). The structured latent space and image latent representation vector mentioned here are common technical terms in this field. The structured latent space is the low-dimensional feature space inside the variational autoencoder, and the image latent representation vector is the vector information obtained after the image is mapped to the structured latent space.

[0042] Step 2: Combine text images with instruction text to predict style embedding information.

[0043] In this embodiment of the invention, replacement text is extracted from the input text instruction and rendered as a glyph image. Visual features are extracted from both the glyph image and the input text image to obtain corresponding glyph image features and text image features. These glyph image features, text image features, and text instruction prediction style embedding information are then combined. Alternatively, visual features are extracted from both the input text image and a given reference image to obtain corresponding reference image features and text image features. These reference image features, text image features, and input text instruction prediction style embedding information are then combined. The text instruction is an instruction to replace the original text in the input text image with the replacement text.

[0044] In this embodiment of the invention, either the original text style can be preserved or style transfer based on a reference image can be achieved. When preserving the original text style, the replacement text is rendered as a glyph image to provide structural information of the text; a preset standard font is used for rendering. Then, style embedding information can be predicted by combining glyph image features, text image features, and text commands. When performing style transfer based on a reference image, the given reference image is an additional target image to be transferred, which can provide the structure and style of the text. Then, style embedding information can be predicted by combining reference image features, text image features, and text commands.

[0045] Preferably, this step can be implemented using a visual language model (VLM), which includes a visual encoder, a word segmenter, a glyph renderer, and a style abstractor.

[0046] The word segmenter is responsible for segmenting the input text command and outputting it to the style abstractor; the glyph renderer is responsible for rendering the replacement text into glyph images; the visual encoder is responsible for extracting visual features from the glyph images and the input text images respectively; the style abstractor includes a style embedding prediction unit and a text decoder. The style embedding prediction unit is responsible for predicting style embedding information by combining the glyph image features, text image features and text command, or by combining the reference image features, text image features and text command; the text decoder is responsible for using the style embedding information and combining it with the input query information to predict the replacement text and text position, and can be used to assist in training the style embedding prediction unit.

[0047] In this embodiment of the invention, a supervised training method is used to train the style embedding prediction unit in the style abstractor. The training dataset consists of tuples {(x s ,x g ,x i ,q pos ,q text ),x t} constitutes, where xs For text images, x g For glyph images, x i For text commands, q pos Query information for position prediction of the edit area, q text To retrieve information for text recognition in the edit area, x t The target image is shown above. In the tuple structure, the information inside the parentheses is the input information, and the information outside the parentheses is the supervision information. Here, the actual text positions and corresponding replacement texts used as supervision information are omitted. Both of these can be directly determined when constructing the training dataset (i.e., they belong to known information) and are used to calculate the cross-entropy loss.

[0048] In this embodiment of the invention, the text image x s With target image x t For paired images; when the glyph image x g With target image x t When creating a text image with a uniform font style, the text image x s With target image x t They have the same style but different text content; when the glyph image x g When representing a text image of a reference style, the text image x s With target image x t They have the same text content but different styles.

[0049] The core objective of this stage is to optimize the style abstractor's capabilities, enabling it to accurately extract style information. During training, only the parameters of the style abstractor are updated (specifically, only the parameters of the style embedding prediction unit within the style abstractor). During training, the stylized visual text image output by the diffusion model is acquired, and the mean squared error (MSE) loss between it and the target image is calculated. Furthermore, to improve the accuracy of the visual language model in locating edit regions and its text recognition capabilities, thus providing better control information to the diffusion model, the text decoder introduces independent cross-entropy losses for different queries. Specifically, the cross-entropy loss is calculated by the text decoder based on the query q. pos With q text After predicting the text position and the replacement text, the corresponding cross-entropy loss is calculated. Specifically, it includes the cross-entropy loss between the predicted replacement text and the replacement text in the text instruction (the actual replacement text), as well as the cross-entropy loss between the predicted text position and the actual text position.

[0050] Finally, the training loss of the style embedding prediction unit is constructed by combining two cross-entropy losses and mean squared error losses, and the style embedding prediction unit is trained.

[0051] Step 3: Obtain the text mask m, and combine the text mask to obtain the latent representation vector of the masked image. The diffusion model uses style embedding information, the latent representation vector of the image, the text mask m, and the latent representation vector of the masked image to predict the stylized visual text image.

[0052] In this embodiment of the invention, the style embedding information, the image latent representation vector, the text mask m, and the masked image latent representation vector are input into the diffusion model, and the diffusion model gradually predicts the text image in which the original text is replaced with replacement text while retaining the style of the original text or the style of the reference image.

[0053] In this embodiment of the invention, the text mask m can be external input information or obtained by combining the text position predicted by the text decoder.

[0054] In this embodiment of the invention, the diffusion model is trained using a self-supervised method; the training method includes: acquiring unlabeled data (various types of image data containing text, such as document data, natural street view data, etc.), and training the model using tuples {(x... s , x m , m,x g , x i , q pos ), x s} constitutes the text image x s Mapping to a structured latent space yields the image's latent representation vector; this vector is then used by the text decoder in the style abstractor to perform a positional query q. pos The text location is predicted. The style embedding prediction unit in the style abstractor is trained in a supervised manner, using the text location as implicit supervision information, and combined with the text image x. s , character image x g With text command x i Predict style embedding information; combine style embedding information, image latent representation vector, text mask m, and masked image latent representation vector x. m The text image is input into a diffusion model, which predicts the corresponding text image. The predicted text image is then compared with the text image x. s The differences between them are used to construct a loss function (e.g., mean squared error loss function), and the diffusion model is trained using the constructed loss function.

[0055] The solution provided by this invention can solve the style consistency problem in image text editing of existing diffusion models, and has the following main advantages:

[0056] (1) Encoding fine-grained text features: The Visual Language Model (VLM) is introduced, whose core component, the Style Abstractor, can effectively extract and represent cross-language text style information, including glyphs, fonts, and colors.

[0057] (2) Provide enhanced style condition guidance: By inputting the style information extracted by the style abstractor into the diffusion model as enhanced style conditions, fine control of the diffusion process can be achieved.

[0058] (3) Design a combined training strategy: Combine supervised and self-supervised methods to train the model. Among them, the self-supervised training strategy can use unlabeled real data to train the model, effectively overcoming the challenge of collecting and labeling large-scale paired datasets and significantly reducing resource requirements.

[0059] Thanks to the above improvements, the present invention can generate text images with high readability and style consistency, and can achieve style preservation of the original text or style transfer based on reference images.

[0060] To more clearly demonstrate the technical solution and its effects provided by the present invention, the method provided by the embodiments of the present invention will be described in detail below with reference to specific examples.

[0061] I. Detailed description of the present invention.

[0062] 1. Introduction to the overall framework.

[0063] This invention provides a stylized visual text editing method, the overall framework of which is as follows: Figure 2 As shown, this framework, called the Visual Text Editing Framework (DiffCTE), mainly consists of three core parts: Variational Autoencoder (VAE), Diffusion Model (U-Net architecture), and Visual Language Model (VLM). The Diffusion Model in this framework is built on top of the Stable Diffusion Model and additionally introduces style conditional constraints.

[0064] Among them, the U-Net architecture is a U-shaped network architecture. The Stable Diffusion model mentioned above is an existing model. Its principle can be found in reference 5: Rombach, R et al.: High-resolution image synthesis with latentdiffusion models. In: Proceedings of the IEEE / CVF conference on computer vision and pattern recognition. pp. 10684–10695 (2022).

[0065] (1) Variational autoencoder.

[0066] In this embodiment of the invention, the variational autoencoder is responsible for processing the input text image x.s Transform into a low-dimensional latent representation vector z t This part pertains to the basic principles of variational autoencoders, which can be implemented using conventional techniques; therefore, it will not be elaborated upon in this invention.

[0067] (2) Visual language model.

[0068] In visual language models, the input text image x is integrated. s and instruction text x i This enhances style embedding. The Glyph Renderer is responsible for rendering the replacement text in the text instruction into glyph images, while the Vision Encoder is responsible for extracting visual features, with its input including the text image x. s The word segmenter primarily processes text commands by segmenting them into words. The Style Abstractor, whose Style Embedding Prediction Unit combines visual features extracted from two types of images by the visual encoder with the output of the word segmenter, predicts a refined style embedding (i.e., style condition). Its Text Decoder utilizes style embeddings to perform position prediction and text recognition tasks based on different queries. The goal is to enable the Style Abstractor to perceive the rendering position of each character, providing implicit supervision.

[0069] In this embodiment of the invention, style abstraction is the core technical step in style extraction. The function of this style abstraction is to learn and extract style information from the input image that can be used to guide the diffusion process. Specifically, to preserve the original text style, it takes the following three types of information as input:

[0070] (A1) Features f extracted by the visual encoder s It is a visual feature extracted from the input text image, mainly capturing the visual context information of the image itself.

[0071] (A2) Character image encoding f g It is a visual feature extracted from glyph images, containing glyph structure information of the target text.

[0072] (A3) Word segmentation results of text instructions: clearly describe the text content to be edited.

[0073] The style abstractor fuses these three types of information into a comprehensive style embedding, which serves as the conditional input to the diffusion model. This mechanism enables visual language models to abstract complex visual style attributes from images, including font, stroke thickness, color, and texture, and accurately apply them to newly generated text, thus effectively solving the style consistency problem.

[0074] When applied to reference image-based style transfer, glyph image encoding f g The features of the reference image are transformed into the features of the reference image, and the other two pieces of information are the same, thereby generating a style embedding that contains style information from the reference image.

[0075] Figure 2 In the example shown, the text instruction "Change “DINER” to “MODEL” means changing "DINER" in the text image to "MODEL", where "MODEL" is the replacement text. Of course, this is just an example; in practical applications, users can choose the language of the text instruction (e.g., Chinese or other languages) and adjust the content of the replacement text as needed. The text decoder outputs A1 and A2 as q. text q pos The corresponding answer is the predicted replacement text and its location. Furthermore, considering the space constraints of the attached figures, Figure 2 Only an example of achieving original text style preservation is provided.

[0076] (3) Diffusion model.

[0077] Conditional diffusion models are an enhanced form of diffusion models that allow the generation process to be controlled by additional inputs (conditions). In classic diffusion models, images are generated by progressively denoising random noise, resulting in a random outcome. Conditional diffusion models, however, guide the model to generate images that meet specific requirements by injecting conditional information (such as text descriptions, category labels, or, in this invention, style embedding information) into each step of the denoising process. This invention uses style embedding information extracted by a style abstractor as a condition, and combines it with the image latent representation vector z. t The text mask m and the masked image latent representation vector x m The text is fed into the diffusion model to ensure that the generated text image is consistent with the given conditions in terms of font, color, and overall visual style.

[0078] 2. Training plan.

[0079] This invention employs a hybrid training scheme that combines supervised and self-supervised training to achieve efficient training and superior performance.

[0080] (1) Supervised training utilizes a synthetic dataset to train the style extractor (primarily training the style embedding prediction unit) to ensure it can accurately extract style information. This dataset consists of tuples {(x... s ,x g ,x i ,q pos ,q text ),x t} constitutes. Where x s For text images, x g For glyph images, x i For text commands, q pos Predict query information for the location of the edit area, q text To retrieve information from the text in the editing area, x t The target image.

[0081] This training strategy aims to significantly improve the style abstractor's extraction capabilities in VLM. Specifically, when x g When creating text images with a consistent font style, x s and x t These are paired images that share the same style but have different text content. Conversely, when x... g When representing a text image in the style of a reference image, x s and x t These are paired images that share the same text content but differ in style. This allows the model to adapt to x. s or x g Extract style information.

[0082] In this part, in addition to using the difference between the text image predicted by the diffusion model and the target image to constrain the style embedding prediction unit, a text decoder is also set in the style abstractor. The decoder decodes the position information (used to generate the text mask) and the text information to be edited (replacement text) to construct a loss to further constrain the style embedding prediction unit, so that it can learn the text information content and position.

[0083] (2) Self-supervised training utilizes unlabeled real-world data to train the diffusion model. The model reconstructs images within masked regions guided by location prediction queries by leveraging style information extracted from glyph images, text instructions, and the original image. This process is optimized by minimizing the mean squared error (MSE) loss between the reconstructed and original images. This process, which eliminates the need for expensive manual annotation, effectively reduces data preparation costs.

[0084] 3. Example introduction of the solution.

[0085] In this example, Stable Diffusion v1.4 (v1.4 being the model version number) is used as the base diffusion model, and modifications have been made to support joint editing of images and text. The entire framework consists of three parts: a VAE (Variational Autoencoder) is used to compress the input image into a latent space representation and decode the latent representation back to a pixel space image after denoising; the aforementioned base diffusion model is the U-Net architecture, which serves as the core model and is responsible for predicting the denoised latent representation from the noisy latent representation; and a VLM (Visual Language Model) is used to extract and provide additional style embedding information. This VLM is further composed of components such as a Vision Encoder and a Style Abstractor.

[0086] In this example, the Style Abstractor can be implemented using a Transformer-based architecture that fuses features from the visual encoder, glyph image encoding, and text instructions through a cross-attention mechanism. Its key features are its Multilayer Perceptron (MLP) and self-attention layers, which extract fine-grained style information from both visual and text inputs.

[0087] In this example, the Vision Encoder can use a pre-trained ViT (VisionTransformer) model whose parameters are frozen during training to ensure effective encoding of image features.

[0088] The Text Decoder in the Style Abstractor is based on the BERT model (Bidirectional Encoder Representation Model) and is used to process text instructions and generate text-related queries to guide text localization and recognition during the denoising process.

[0089] All models were trained on a server equipped with eight NVIDIA A100 GPUs, and all input images were resized to 512x512 pixels. A two-stage training strategy was employed. The first stage used a paired dataset containing synthetic text images with a uniform font style to supervise the training of a style abstractor, enabling it to effectively extract style information from images or glyphs. The learning rate was 1e-5, and the AdamW optimizer (an adaptive moment estimator with weight decay) was used. The second stage fine-tuned the diffusion model on unlabeled real-world data, adjusting the learning rate to 1e-6, also using the AdamW optimizer.

[0090] III. Effects and Experimental Verification

[0091] (1) Significantly improves text accuracy and style consistency: Existing diffusion models face problems of poor readability and style inconsistency when processing text images. This invention significantly improves the accuracy and style consistency of the model-generated text editing by introducing a style abstractor and providing enhanced style condition guidance, making the generated text more visually coherent, such as... Figure 3 As shown.

[0092] (2) Achieving high-fidelity style transfer: This invention not only maintains the style of the original image, but also performs effective style transfer based on the reference image, such as... Figure 3 As shown, this makes it more practical in various application scenarios such as advertising design.

[0093] Figure 3 In the diagram, the first row contains the input text image, and the second row contains the text image predicted by the diffusion model. The two columns on the left provide reference images for style transfer, while the two columns on the right require maintaining style consistency and generating complex characters.

[0094] (3) Overcoming the data labeling problem: Traditional text and image editing methods require a large number of labeled paired datasets. The self-supervised training strategy proposed in this invention effectively solves this problem, enabling the model to be trained using unlabeled real-world data, which greatly reduces resource investment and manpower costs.

[0095] (4) Superior to existing technologies: Through comprehensive comparative experiments with existing solutions, the present invention demonstrates superior performance in both text accuracy (OCRAcc) and style consistency (Cor) evaluation metrics. For example, experiments on multiple public datasets such as ArT, COCOText, TextOCR, and ICDAR13 have verified its effectiveness. Using text accuracy as the main evaluation metric, the text accuracy of the present invention is significantly higher than existing methods such as AnyText, TextDiffuser, and T2I-Adapter on the ArT, COCOText, TextOCR, and ICDAR13 datasets, as shown in Table 1. The FT following SD1 and SD2 indicates the text fine-tuning scenario. Details of other solutions are provided below. Furthermore, style consistency was evaluated through human evaluation, and the results show that the present invention achieved the highest scores in both style preservation and style transfer tasks.

[0096] Table 1: Quantitative Comparison Results of This Method with Various Existing Solutions

[0097]

[0098] like Figure 4 As shown, a visual comparison result is provided. Figure 4 The first line is the input text image, and the text below it is the replacement text; the second line is the text mask, which covers the area in the text image where the text needs to be replaced; the third to eighth lines are the text images predicted by various schemes, corresponding to the following schemes in order: SRNet (Scene Text Replacement Network, a generative adversarial network-based method used to replace text in an image and attempt to blend the new text with the background), MOSTEL (Stroke-Level Modified Text Editing, a fine-grained text editing method that achieves more realistic style transfer by modifying the "stroke" level features of the text), SD1 (Stable Diffusion Model 1), SD2 (Stable Diffusion Model 2), DiffSTE (Dual Encoder Diffusion Text Editing), and this invention.

[0099] The comparative results demonstrate that the present invention can generate text with a high degree of consistency with the original text style, including font, color, texture, and shadows. Especially in complex scenes, such as images with specific lighting effects or background textures, the present invention exhibits excellent performance, generating edited results with high visual coherence.

[0100] also, Figure 3 and Figure 4 All the illustrations involved are from existing public test sets and do not contain sensitive information.

[0101] In summary, this invention not only surpasses existing technologies in quantitative indicators but also demonstrates significant advantages in qualitative effects, fully proving its innovation and practicality in the field of visual text editing.

[0102] Through the above description of the embodiments, those skilled in the art can clearly understand that the above embodiments can be implemented by software, or by using software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solutions of the above embodiments can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, mobile hard drive, etc.), including several instructions to cause a computer device (such as a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0103] Example 2

[0104] This invention also provides a stylized visual text editing system, which is mainly used to implement the methods provided in the foregoing embodiments, such as... Figure 5 As shown, the system mainly includes:

[0105] Variational autoencoders are used to map input text images to a structured latent space to obtain image latent representation vectors.

[0106] A visual language model is used to extract replacement text from an input text instruction and render it as a glyph image. Visual features are extracted from both the glyph image and the input text image to obtain corresponding glyph image features and text image features. These glyph image features, text image features, and text instruction prediction style embedding information are then combined. Alternatively, visual features are extracted from both the input text image and a given reference image to obtain corresponding reference image features and text image features. These reference image features, text image features, and text instruction prediction style embedding information are then combined. The text instruction is an instruction to replace the original text in the input text image with the replacement text.

[0107] The predicted text image output unit is used to obtain a text mask and combine it with the text mask to obtain the latent representation vector of the masked image. The style embedding information, the latent representation vector of the image, the text mask and the latent representation vector of the masked image are input into the diffusion model. The diffusion model gradually predicts the text image in which the original text is replaced with replacement text while retaining the style of the original text or the style of the reference image.

[0108] Since the main technical details involved in this system have been described in detail in previous embodiments, they will not be repeated here.

[0109] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional modules is used as an example. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the system can be divided into different functional modules to complete all or part of the functions described above.

[0110] Example 3

[0111] The present invention also provides a processing device, such as Figure 6 As shown, it mainly includes: one or more processors; a memory for storing one or more programs; wherein, when the one or more programs are executed by the one or more processors, the one or more processors implement the method provided in the foregoing embodiments.

[0112] Furthermore, the processing device also includes at least one input device and at least one output device; in the processing device, the processor, memory, input device, and output device are connected via a bus.

[0113] In this embodiment of the invention, the specific types of the memory, input device, and output device are not limited; for example:

[0114] Input devices can be touchscreens, image acquisition devices, physical buttons, or mice, etc.

[0115] The output device can be a display terminal;

[0116] The memory can be random access memory (RAM) or non-volatile memory, such as disk storage.

[0117] Example 4

[0118] The present invention also provides a readable storage medium storing a computer program that, when executed by a processor, implements the method provided in the foregoing embodiments.

[0119] In this embodiment of the invention, the readable storage medium is a computer-readable storage medium and can be disposed in the aforementioned processing device, for example, as a memory in the processing device. Furthermore, the readable storage medium can also be any medium capable of storing program code, such as a USB flash drive, portable hard drive, read-only memory (ROM), magnetic disk, or optical disk.

[0120] The above description is merely a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims. The information disclosed in the background section is intended only to enhance the understanding of the overall background technology of the present invention and should not be construed as an admission or implication in any way that such information constitutes prior art known to those skilled in the art.

Claims

1. A stylized visual text editing method, characterized in that, include: Step 1: Map the input text image to a structured latent space to obtain the image latent representation vector; Step 2: Extract the replacement text from the input text instruction and render it as a glyph image. Extract visual features from both the glyph image and the input text image to obtain corresponding glyph image features and text image features. Combine these glyph image features, text image features, and text instruction prediction style embedding information. Alternatively, extract visual features from both the input text image and a given reference image to obtain corresponding reference image features and text image features. Combine these reference image features, text image features, and input text instruction prediction style embedding information. The text instruction is an instruction to replace the original text in the input text image with the replacement text. Step 3: Obtain the text mask and combine it with the text mask to obtain the latent representation vector of the masked image. Input the style embedding information, the latent representation vector of the image, the text mask and the latent representation vector of the masked image into the diffusion model. The diffusion model will gradually predict that the original text will be replaced with the replacement text, while retaining the original text style or the style of the reference image.

2. The stylized visual text editing method according to claim 1, characterized in that, Step 1 is implemented by a variational autoencoder, which maps the input text image to its own internal structured latent space to obtain the image latent representation vector.

3. The stylized visual text editing method according to claim 1, characterized in that, Step 2 is implemented through a visual language model, which includes a visual encoder, a word segmenter, a glyph renderer, and a style abstractor. The word segmenter is responsible for segmenting the input text command and outputting it to the style abstractor; the glyph renderer is responsible for rendering the replacement text into glyph images; the visual encoder is responsible for extracting visual features from the glyph images and the input text image respectively; the style abstractor includes a style embedding prediction unit and a text decoder. The style embedding prediction unit is responsible for predicting style embedding information by combining the glyph image features, text image features and text command, or by combining the reference image features, text image features and text command; the text decoder is responsible for using the style embedding information and combining it with the input query information to predict the replacement text and text position, which is used to assist in training the style embedding prediction unit.

4. The stylized visual text editing method according to claim 3, characterized in that, Also includes: The style embedding prediction unit in the style abstractor is trained using a supervised method. The training dataset consists of tuples {(x s ,x g ,x i ,q pos ,q text ),x t } constitutes, where x s For text images, x g For glyph images, x i For text commands, q pos Predict query information for the location of the edit area, q text To retrieve information from the text in the editing area, x t For the target image; During training, the text image predicted by the diffusion model is acquired, and the result is compared with the target image x. t Mean squared error loss between them; Furthermore, the text decoder utilizes style embedding information and combines it with q pos With q text After predicting the corresponding text position and replacement text, the corresponding cross-entropy loss is calculated. Finally, the training loss of the style embedding prediction unit is constructed by combining the cross-entropy loss and the mean squared error loss, and the style embedding prediction unit is trained.

5. The stylized visual text editing method according to claim 4, characterized in that, The text image x s With target image x t For paired images; when the glyph image x g With target image x t When creating a text image with a uniform font style, the text image x s With target image x t They have the same style but different text content; when the glyph image x g When the text image represents the style of the reference image, the text image x s With target image x t They have the same text content but different styles.

6. The stylized visual text editing method according to claim 4, characterized in that, The diffusion model is trained using a self-supervised method; Training methods include: obtaining unlabeled data, and generating data from tuples {(x... s , x m , m, x g , x i , q pos ), x s } constitutes; where x s For text images, x g For glyph images, x i For text commands, m is the text mask, and x is the text mask. m Let q be the latent representation vector of the masked image. pos Predict query information for the location of the edit area; transform the text image x s Mapping to a structured latent space yields the image's latent representation vector; this vector is then used by the text decoder in the style abstractor to perform a positional query q. pos The text location is predicted. The style embedding prediction unit in the style abstractor is trained in a supervised manner, using the text location as implicit supervision information, and combined with the text image x. s , character image x g With text command x i Predict style embedding information; combine style embedding information, image latent representation vector, text mask m, and masked image latent representation vector x. m The text image is input into a diffusion model, which predicts the corresponding text image. The predicted text image is then compared with the text image x. s The difference between the two is used to construct a loss function, and the constructed loss function is used to train the diffusion model.

7. The stylized visual text editing method according to claim 3, characterized in that, The process of obtaining the text mask and combining it with the text mask to obtain the latent representation vector of the masked image includes: Obtain the text mask from external input, or generate the corresponding text mask after predicting the text position using a text decoder; The text mask is added to the input text image to obtain the masked image; the masked image is then mapped to a structured latent space to obtain the latent representation vector of the masked image.

8. A stylized visual text editing system, characterized in that, include: Variational autoencoders are used to map input text images to a structured latent space to obtain image latent representation vectors. A visual language model is used to extract replacement text from an input text instruction and render it as a glyph image. Visual features are extracted from both the glyph image and the input text image to obtain corresponding glyph image features and text image features. These glyph image features, text image features, and text instruction prediction style embedding information are then combined. Alternatively, visual features are extracted from both the input text image and a given reference image to obtain corresponding reference image features and text image features. These reference image features, text image features, and input text instruction prediction style embedding information are then combined. The text instruction is an instruction to replace the original text in the input text image with the replacement text. The predicted text image output unit is used to obtain a text mask and combine it with the text mask to obtain the latent representation vector of the masked image. The style embedding information, the latent representation vector of the image, the text mask and the latent representation vector of the masked image are input into the diffusion model. The diffusion model gradually predicts the text image in which the original text is replaced with replacement text while retaining the style of the original text or the style of the reference image.

9. A processing device, characterized in that, include: One or more processors; Memory, used to store one or more programs; Wherein, when the one or more programs are executed by the one or more processors, the one or more processors cause the one or more processors to implement the method as described in any one of claims 1 to 7.

10. A readable storage medium storing a computer program, characterized in that, When a computer program is executed by a processor, it implements the method as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Garment style fusion method and system based on diffusion model

    CN117315417A

  • Single stream multi-level alignment for vision-language pretraining

    US20230281963A1