Image generation method and device

By obtaining the art word description text and conditional images, and using the deep learning model to generate art word word and content images that meet user needs, it solves the problems of complex design and limited style of handwritten newspapers, and realizes efficient and personalized handwritten newspaper creation.

CN120495476APending Publication Date: 2025-08-15BEIJING YUANLI WEILAI SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510754907.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-08-15

AI Technical Summary

Technical Problem

During the design process of traditional handwritten newspapers, it is difficult for users to freely choose the artistic style and the design is complex. The operating threshold of existing tools and software is high, resulting in the lack of personality and novelty of handwritten newspapers, and the design process is cumbersome and time-consuming.

Method used

By obtaining art word description text, target content text and conditional images, using art word generation model and image generation model, automatically generate art word image and content images that meet user needs, and combining deep learning technologies such as Transformer and FLUX models to achieve the perfect integration of art word and content.

Benefits of technology

It lowers the threshold for users to design handwritten newspapers, improves the flexibility and personalization of creation, and the generated art-like images and content images perfectly meet user needs in style and details, improving the aesthetics and unity of handwritten newspapers.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120495476A_ABST
    Figure CN120495476A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an image generation method and device, and the method comprises the steps: obtaining a wordart description text, a target content text and a condition image, and enabling the condition image to be used for constraining the font of a target wordart; based on the wordart description text, processing the condition image to generate a target wordart image corresponding to the condition image; and generating a target image corresponding to the target content text based on the target content text and the target wordart image. According to the method, perfect fusion of the wordart and the content is realized, the aesthetic property and expressive power of the image are improved, the creation threshold of the user is reduced, the image generation process is more efficient and intelligent, and diversified creative requirements are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of this specification relate to the field of image processing technology, and more particularly to an image generation method and device. Background Art

[0002] Handwritten newspapers are a common form of school or community activity, often used to present information, promote a theme, or enhance the environment. However, traditional handwritten newspaper design often relies on fixed templates. While these templates offer certain conveniences, they also come with significant limitations, resulting in handwritten newspapers that often become stereotyped, lacking individuality and novelty. While graphics processing software offers powerful features and a wealth of tools, the process is relatively complex and requires considerable time and effort to learn and master. For many beginners, this complexity can become a barrier to their creative process, making designing a handwritten newspaper difficult and frustrating.

[0003] Therefore, it is urgent to provide a solution to solve the above technical problems. Summary of the Invention In view of this, embodiments of this specification provide an image method. One or more embodiments of this specification also relate to a method for training an artistic character generation model, a method for training an image generation model, an image generation device, a method for training an artistic character generation model, a method for training an image generation model, a computing device, a computer-readable storage medium, and a computer program product to address technical deficiencies in the prior art.

[0004] According to a first aspect of the embodiments of this specification, there is provided an image generation method, comprising: Acquire a word art description text, a target content text, and a conditional image, wherein the conditional image is used to constrain the shape of the target word art; Based on the artistic word description text, the conditional image is processed to generate a target artistic word image corresponding to the conditional image; Based on the target content text and the target word art image, a target image corresponding to the target content text is generated.

[0005] According to a second aspect of the embodiments of this specification, a method for training an artistic character generation model is provided, comprising: Acquire a plurality of first training samples, wherein the first training samples include word art description text samples, conditional image samples, and word art image samples; Inputting the word art description text sample and the conditional image sample into an initial word art generation model, generating a first word art image corresponding to the word art description text sample, and adjusting parameters of the initial word art generation model according to the accuracy of the first word art image relative to the word art image sample until a first training stop condition is met, thereby obtaining an intermediate word art generation model; Selecting a first target word art image from a plurality of first word art images and constructing a first target training sample; The intermediate artistic word generation model is trained based on the first target training sample until a second training stop condition is met, thereby obtaining a target artistic word generation model.

[0006] According to a third aspect of the embodiments of this specification, a method for training an image generation model is provided, comprising: Acquire a plurality of second training samples, wherein the second training samples include target content text samples, word art image samples, and target image samples; Inputting the target content text sample and the word art image sample into an initial image generation model to generate a first image corresponding to the target content text sample, and adjusting parameters of the initial image generation model based on the accuracy of the first image relative to the target image sample until a third training stop condition is met, thereby obtaining an intermediate image generation model; Screening a first target image from a plurality of first images and constructing a second target training sample; The intermediate image generation model is trained based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model.

[0007] According to a fourth aspect of the embodiments of this specification, there is provided an image generating apparatus, including: A data acquisition module is configured to acquire a description text of a word art, a target content text, and a conditional image, wherein the conditional image is used to constrain the shape of the target word art; An artistic word generation module is configured to process the conditional image based on the artistic word description text to generate a target artistic word image corresponding to the conditional image; The image generation module is configured to generate a target image corresponding to the target content text based on the target content text and the target artistic word image.

[0008] According to a fifth aspect of the embodiments of this specification, there is provided a device for training an artistic character generation model, comprising: A first word art sample construction module is configured to obtain a plurality of first training samples, wherein the first training samples include word art description text samples, conditional image samples and word art image samples; a first word art model training module configured to input the word art description text sample and the conditional image sample into an initial word art generation model, generate a first word art image corresponding to the word art description text sample, and adjust parameters of the initial word art generation model according to an accuracy rate of the first word art image relative to the word art image sample until a first training stop condition is met, thereby obtaining an intermediate word art generation model; A second artistic word sample construction module is configured to select a first target artistic word image from a plurality of first artistic word images and construct a first target training sample; The second artistic character model training module is configured to train the intermediate artistic character generation model based on the first target training sample until a second training stop condition is met to obtain a target artistic character generation model.

[0009] According to a sixth aspect of the embodiments of this specification, there is provided an image generation model training device, comprising: A first image sample construction module is configured to obtain a plurality of second training samples, wherein the second training samples include target content text samples, word art image samples and target image samples; a first image model training module configured to input the target content text sample and the word art image sample into an initial image generation model, generate a first image corresponding to the target content text sample, and adjust parameters of the initial image generation model based on an accuracy rate of the first image relative to the target image sample until a third training stop condition is met, thereby obtaining an intermediate image generation model; a second image sample construction module, configured to screen a first target image from a plurality of first images and construct a second target training sample; The second image model training module is configured to train the intermediate image generation model based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model.

[0010] According to a seventh aspect of the embodiments of this specification, a computing device is provided, including: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above method are implemented.

[0011] According to an eighth aspect of the embodiments of this specification, a computer-readable storage medium is provided, which stores computer-executable instructions, and the steps of the above method are implemented when the instructions are executed by a processor.

[0012] According to a ninth aspect of the embodiments of this specification, a computer program product is provided, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above method.

[0013] One embodiment of this specification implements the automatic generation of a target word art image that matches the style of the word art description text and the glyphs in the conditional image by receiving user input of word art description text, target content text, and conditional images, thereby enabling the generation of any word art style desired by the user. Furthermore, based on the generated target word art image and target content text, a target image that matches the style of the target content text is generated, resolving the existing issues of users being unable to freely select word art styles and limited hand-copied newspaper design, thereby enhancing creative flexibility and personalization. This lowers the barrier for ordinary users to design high-quality hand-copied newspapers. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flow chart of an image generation method provided by one embodiment of this specification; Figure 2 is a schematic diagram of generating a first target artistic word image in an image generating method provided in one embodiment of this specification; Figure 3 This is a schematic diagram of a target word art generation model structure in an image generation method provided in one embodiment of this specification; Figure 4 is a schematic diagram of generating a first target image in an image generating method provided in one embodiment of this specification; Figure 5 This is a flowchart of the processing process of an image generation method provided by an embodiment of this specification in a handwritten newspaper production scenario; Figure 6 This is a flowchart of a processing process of a method for training an artistic character generation model provided by one embodiment of this specification; Figure 7 This is a flowchart of a processing process of an image generation model training method provided by one embodiment of this specification; Figure 8 This is a flowchart of a processing process of an artistic character generation model training method and an image generation model training method in a handwritten newspaper production scenario provided by an embodiment of this specification; Figure 9 is a schematic diagram of generating a second target word art image in an image generating method provided in one embodiment of this specification; Figure 10 is a schematic diagram of generating a second target image in an image generating method provided in one embodiment of this specification; Figure 11This is a schematic structural diagram of an image generating device provided by one embodiment of this specification; Figure 12 This is a structural diagram of an artistic character generation model training device provided by an embodiment of this specification; Figure 13 This is a schematic diagram of the structure of an image generation model training device provided by one embodiment of this specification; Figure 14 This is a structural block diagram of a computing device 1400 provided in one embodiment of this specification. DETAILED DESCRIPTION

[0015] The following description sets forth many specific details to facilitate a thorough understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar generalizations without violating the scope of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.

[0016] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a," "the," and "the" used in one or more embodiments of this specification and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.

[0017] It should be understood that although the terms first, second, etc. may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining".

[0018] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0019] First, the terms involved in one or more embodiments of this specification are explained.

[0020] Image generation: A major task in computer vision, its goal is to generate a corresponding image based on a given text description.

[0021] SAM (Segment Anything Model): A powerful image segmentation model that can accurately segment objects in various images and separate the target object from the background or other objects.

[0022] Transformer: A deep learning model architecture used for natural language processing and other sequence-to-sequence tasks, typically consisting of an encoder and a decoder.

[0023] The FLUX model is a text-to-image synthesis model developed by the Black Forest Lab, the original team behind Stable Diffusion. Technically, FLUX boasts 12 billion parameters, making it one of the largest open-source text-to-image models. It utilizes several innovative technologies: Flow Matching Training improves upon the traditional diffusion model training process, streamlining it and significantly improving generation quality, resulting in more realistic and detailed images; Rotational Position Embedding enhances the model's ability to recognize features at different locations in the image, helping it better grasp the overall structure and details of the image, resulting in more reasonable and accurate spatial layouts for generated images; and Parallel Attention Layers significantly improve the model's ability to capture long-range dependencies, enabling it to better understand the complex relationships between text and images, and enhancing the accuracy and logic of image generation. FLUX excels in image generation: its powerful text-to-image generation capabilities enable it to generate matching images based on user-provided text descriptions. Stable Diffusion technology ensures image coherence and consistency, while LoRa further enhances image detail. For example, in the field of artistic creation, artists can use it to create paintings with film-level quality and unique style; excellent text rendering effect: it is good at processing complex text in images, including words with repeated letters, which is ideal for designs that require accurate text presentation; advanced composition capabilities: it can follow complex instructions about the position and relationship of objects in the image to create complex scenes.

[0024] In the field of hand-written newspaper creation, traditional methods face many limitations. First, the design process of artistic words is complicated. Ordinary users find it difficult to achieve diverse style designs due to the cumbersome operation and high learning cost of graphic design software. Traditional artistic word generation tools have fixed styles and cannot meet the growing and diverse needs. Secondly, the production process of hand-written newspapers is cumbersome. From overall planning, material collection to typesetting design, all must be completed manually, which not only consumes a lot of time and energy, but also makes it difficult to ensure the uniformity of the style and visual effects of each element. In addition, the requirements for the style of hand-written newspapers vary in different occasions, which even experienced creators find difficult to cope with easily. Therefore, how to automatically generate hand-written newspapers with style and content that meet user requirements has become a core technical problem that needs to be solved urgently.

[0025] When it comes to word art design, current methods often rely on providing a library of preset fonts. Users can only choose from a limited set of fonts and achieve the desired effect through simple rendering. Meanwhile, breakthroughs in deep learning technology have significantly advanced the development of Chinese font generation and its related applications. For example, the HFH-Font method generates word art in a specific style by using a target character image rendered in the style of a source font and a set of images rendered in the target font style to guide a diffusion model. However, these methods require a small number of samples of the target style in advance and are unable to generate word art in any arbitrary style.

[0026] In the field of hand-copied newspaper design, the currently widely used technical approach is to provide users with a relatively simple template framework and a series of insertable assets to assist in completing the design process. These templates typically have a fixed style, a single layout, and generally generic assets, limiting user freedom in the creative process. Deep learning-based methods, on the other hand, require users to provide some background or asset information before designing the layout, and are primarily used in poster design.

[0027] On the other hand, designers can use professional image editing software such as Photoshop to manually draw highly creative artistic wording and hand-written newspapers, but this process involves complex graphic drawing techniques, color matching knowledge, and proficiency in the use of software tools. For ordinary users, the operation threshold is extremely high, and it is difficult to design works that suit their personal preferences.

[0028] In this regard, in this specification, an image method is provided. This specification also involves an artistic character generation model training method, an image generation model training method, an image generation device, an artistic character generation model training device, an image generation model training device, a computing device, a computer-readable storage medium and a computer program product, which are described in detail one by one in the following embodiments.

[0029] See also Figure 1 , Figure 1 This is a flowchart of an image generation method provided by an embodiment of this specification, which specifically includes the following steps.

[0030] Step 102: Obtaining a word art description text, a target content text, and a conditional image, wherein the conditional image is used to constrain the shape of the target word art.

[0031] Specifically, the artistic word description text refers to the text in which the user specifically describes the style, pattern and other characteristics of the artistic word. The target content text is the content description text of the image to be generated by the user. The target content description text may include the specific elements, element positions and style of the image that the user expects to generate. The conditional image provides the basic framework of the glyph to ensure that the generated artistic word maintains the individual style while not losing the standardization and readability of the glyph. In one embodiment of this specification, the conditional image can be understood as the standard font glyph input by the user corresponding to the target artistic word. The target artistic word can be understood as the artistic word that the user needs to generate. The font glyph is the visual presentation form of the text, covering elements such as stroke shape, structural layout and overall style. Different fonts, such as Song font with thin horizontal and thick vertical strokes and serif decoration, bold font with simple and straight strokes, and Kai font with handwritten strokes, all convey unique beauty and recognition through stroke features and component combinations.

[0032] Based on this, in actual applications, users need to first describe the style of the word art they expect according to their needs, provide word art description text, describe the overall content of the image they expect to generate, provide target content text, and enter the conditional image of the word art. By obtaining the basic data such as the word art description text, target content text and conditional image, the word art can be automatically generated.

[0033] For an example, see Figure 2 , Figure 2is a schematic diagram of generating a first target wordart image in an image generation method provided by an embodiment of the present specification, wherein the wordart description is the wordart description text, specifically "This characters are presented in a bold, sans-serif style with thick, uniform strokes. The first character is red with a black outline, the second character is blue with a black outline, the third character is yellow with a black outline, and the fourth character is orange with a black outline. Each character has a consistent stroke width throughout, giving them a balanced and cohesive appearance. The strokes are smooth and rounded at the corners, contributing to a modern and clean aesthetic. The black outlines around each character enhance their visibility and provide a clear separation from thebackeround", that is, "the characters are presented in a bold, sans serif font with thick, even strokes. The first character has a red outline with a black outline, the second character has a blue outline with a black outline, the third character has a yellow outline with a black outline, and the fourth character has an orange outline with a black outline. The stroke width of each character is consistent throughout, giving it a balanced and cohesive appearance. The strokes are smooth and rounded at the corners, which helps create a modern and clean aesthetic. The black outline around each character enhances their visibility and provides a clear separation from the background", the standard glyph image "Thank You Teachers", that is, the conditional image, is used to constrain the glyph or stroke features of the generated artistic text, and the artistic text image is the generated target artistic text image.

[0034] Step 104: Based on the wordart description text, the conditional image is processed to generate a target wordart image corresponding to the conditional image.

[0035] Specifically, target artistic characters refer to characters with specific colors and styles generated according to user needs.

[0036] Based on this, after receiving the artistic word description text, target content text and conditional image provided by the user, the artistic word needs to be generated first. Therefore, the artistic word description text and conditional image are processed first to generate the corresponding target artistic word image. The target artistic word image will also serve as a condition for the generation of the target image and participate in the subsequent generation of the target image.

[0037] Furthermore, in the embodiments of this specification, the target word art can be generated by a target word art generation model. The processing of the conditional image based on the word art description text to generate the target word art image corresponding to the conditional image includes: inputting the word art description text and the conditional image into the target word art generation model; and the target word art generation model processing the word art description text and the conditional image to obtain the target word art image corresponding to the conditional image.

[0038] Specifically, the target WordArt model is trained with a large amount of sample data to accurately understand and generate WordArt images that match user descriptions. The target WordArt generation model can simultaneously accept a conditional image and WordArt description text. The conditional image must be resized to a preset size before input to ensure compatibility with the model.

[0039] Based on this, the adjusted conditional image and the artistic word description text are input into the target artistic word generation model. The target artistic word generation model processes the artistic word description text and conditional image through multi-level feature extraction and fusion to generate a target artistic word image that matches the description.

[0040] Continuing with the previous example, after the user enters a word art description, target content, and a conditional image, the target word art generation model processes the input word art description and conditional image, resizing the conditional image to 512x512 pixels, ensuring that the width and height are integer multiples of 16. The resized conditional image is then processed with the word art description to generate a target word art "Thank you, teacher" that matches the description and the glyphs in the conditional image. The glyphs in the word art image match the standard glyphs in the conditional image, and the color, outline, and stroke characteristics of the text in the word art image match the user's description, ensuring that the style and details of each character are as expected. The target word art generation model ultimately outputs a word art image with a pixel size that matches the input conditional image. This output target word art image will be used as the title for the subsequent hand-written newspaper.

[0041] It should be noted that for the conditional image, the user only needs to input the characters of the standard glyph corresponding to the target word art to be generated. The standard glyph can be understood as a font, for example, Kaiti, Microsoft YaHei, Songti, etc. The user only needs to input characters such as "Thank you teacher" and use it as the conditional image parameter. After the word art generation model receives the characters input by the user, it converts them into a standard glyph image (that is, the conditional image) for processing.

[0042] In summary, the embodiments of this specification, through precise processing of the word art description text and conditional images, ensure that the generated target word art image not only matches the standard word shape, but also perfectly meets user needs in style and details, greatly improving the aesthetics and personalization of the handwritten newspaper title. At the same time, the input operation is simple for the user, providing a better experience.

[0043] For further information, see Figure 3 , Figure 3 This is a schematic diagram of the structure of a target word art generation model in an image generation method provided in an embodiment of this specification. In order to accurately integrate the word art description text with the conditional image, the target word art generation model includes: a first text encoder 302, a first image encoder 304 and a word art generation sub-module 306.

[0044] Correspondingly, the target word art generation model processes the word art description text and the conditional image to obtain the target word art image, including: the first text encoder 302 encodes the word art description text to obtain the corresponding first text feature; the first image encoder 304 encodes the conditional image to obtain the corresponding first image code; obtains a first noise variable, and splices the first noise variable, the first text feature and the first image code to obtain a first splicing feature; inputs the first splicing feature into the word art generation submodule 306 to generate a target word art image.

[0045] Specifically, the first text encoder 302 is a module in the target word art generation model used to encode text-type data and is responsible for extracting key features from the word art description text. The first image encoder 304 specifically processes image data and extracts key feature information such as glyph shape, color, and outline from the conditional image. The first noise variable is used to increase the randomness and diversity of the generated image. The first noise variable is randomly sampled from a normal distribution to ensure the uniqueness of the generated image. The word art generation submodule 306 is used to fuse the first text features, the first image encoding, and the first noise variable. The specific word art generation submodule 306 can be implemented using a transformer architecture.

[0046] Based on this, after obtaining the word art description text and the conditional image, the target word art model inputs the word art description text into the first text encoder 302 for encoding processing to obtain the first text feature, and inputs the conditional image into the first image encoder 304 for encoding processing to obtain the first image code. Subsequently, a first noise variable is obtained by random sampling from a normal distribution, and the first noise variable, the first text feature and the first image code are spliced to form a first splicing feature. Finally, the first splicing feature is input into the word art generation submodule 306 to generate a target word art image with the same pixel size as the input standard font image, ensuring that it perfectly meets user needs in terms of font shape, style and details.

[0047] Continuing with the above example, the target word art generation model obtains the word art description text and the conditional image "Thank you, teacher". The word art description text is input into T5 and CLIP of the first text encoder 302 for processing to obtain the first text feature A1. The conditional image is input into the image encoder VAE for processing to obtain the first image code B1. The first noise variable C1 is randomly sampled from the normal distribution. The three features C1, A1, and B1 are spliced along the second dimension through torch.cat((txt, img), 1), keeping the batch size (the number of samples used in a forward propagation and backpropagation during training or inference) unchanged. The other dimensions are spliced together to obtain the first spliced feature D1. D1 is input into the word art generation submodule 306, and the feature fusion and generation are performed through the transformer architecture. The final output is the target word art image that is consistent with the style of the conditional image and contains the meaning of "Thank you, teacher".

[0048] The word art generation submodule 306 includes: a first feature fusion block 3062, a second feature fusion block 3064 and a decoder 3066; The step of inputting the first splicing feature into the word art generation submodule 306 to generate a target word art image includes: The first splicing feature is input into the first feature fusion block 3062, and the first feature fusion block 3062 performs denoising processing on the first splicing feature to obtain a first fusion feature; the first fusion feature is input into the second feature fusion block 3064, and the second fusion block processes the first fusion feature into a single feature sequence through linear projection, and performs denoising processing on the single feature sequence to obtain a target image feature; the decoder 3066 decodes the target image feature to obtain a target art word image.

[0049] The target word art generation model and the target image generation model in the embodiments of this specification both use an improved version of the FLUX model with the Transformer structure as the core as the backbone network.

[0050] Specifically, the target word art generation model processes data as follows: The input word art description text is first encoded by the first text encoder 302's T5-XXL (Text-to-Text Transfer Transformer - XXL) encoder and CLIP (Contrastive Language-Image Pre-Training) text encoder, converting natural language into semantic vectors that the model can understand, thereby generating first text features. The image is then input into the first image encoder 304, the encoder portion of the VAE (Variational Autoencoder), which converts the image into latent features conforming to a Gaussian distribution, thereby generating the first image encoding. In the word art generation submodule 306, or the diffusion model, the first image encoding is combined with the semantic vector derived from the first text feature and the position information obtained from the rotational position encoding. This is processed through a 19-layer MM-DiT Block (Multi-Modal Diffusion Transformer Block) structure and a 38-layer Single-DiT Block (Single-stream Diffusion Transformer Block) structure to gradually denoise and optimize the latent features. During this process, the model guides image generation based on the first text feature, ensuring that the generated image matches the text description. The resulting latent features are input into the decoder portion of the VAE, where they are reconstructed into a pixel-level image, resulting in the final generated image.

[0051] The T5-XXL Encoder takes the raw text sequence corresponding to the text of the word art description as input. It first uses the T5 tokenizer to convert the raw text sequence into a token sequence, assigning each token a unique index. It also adds positional encoding to the text to capture the order of words in the text. The input token sequence is then processed through multiple self-attention layers and feed-forward layers. After the encoder's multiple layers of processing, it outputs a guide tensor of dimension (bs, 512, 4096), where bs represents the batch size, 512 is the number of tokens, and 4096 is the feature dimension of each token.

[0052] The CLIP ViT-L (Vision Transformer-Large) Text Encoder maps each word in the text describing the word art into a low-dimensional vector space through methods such as word embedding, obtaining the initial vector representation of the text. Positional encoding is also added. The Visual Transformer (ViT) structure in the CLIP model is then used to extract features from the embedded text vector. The ViT structure consists of multiple stacked Transformer blocks, each of which contains components such as self-attention layers, layer normalization, and feedforward neural networks. After a series of feature extraction and transformations, the CLIP ViT-L Text Encoder extracts the global semantic information of the text and outputs it as a vector.

[0053] The guide tensor output by T5 and the global semantic vector output by CLIP are used as the first text feature.

[0054] MM-DiT Block: Input: Prepare the first text feature: the guided tensor output by T5 (e.g., bs × 512 × 4096) or the global semantic vector output by CLIP. First image code: The latent variable generated by the VAE (e.g., bs × 64 × H × W, with 64 channels after Pack_Latents), which must first be flattened into a sequence (bs × (H × W) × 64). The first noise variable, Latent, undergoes a linear layer to adjust the feature dimension before being passed to the MM-DiT Block. The first noise variable, the first text feature, and the first image code are concatenated and passed to the attention layer for attention calculation. A feed-forward neural network (FFN) performs a nonlinear transformation on the attention output, further fusing the first noise variable, the first text feature, and the first image code. After residual connections and layer normalization, the image features are output to the next block, preserving the original information while incorporating text semantics to produce the first fused feature.

[0055] The Single-DiT Block receives the first fused feature output by the MM-DiT Block, fuses the first noise variable, the first text feature, and the first image code in the first fused feature into a single feature sequence through linear projection, and then refers to the processing process of the MM-DiT Block input attention layer and the feedforward neural network to obtain the second fused feature.

[0056] The modulated second fusion feature is decoded into the final image pixel value through the image decoding module VAE Decoder, and the target art word image that meets the text description is output.

[0057] In addition, Figure 3 The other structures are described as follows: Pooled: represents the pooling operation, which is used to compress and reduce the dimension of features, reducing the amount of data while retaining important features.

[0058] MLP (Multilayer Perceptron): Multilayer perceptron, used to perform nonlinear transformation and feature extraction on input features, converting them into a suitable feature space for subsequent fusion and other operations.

[0059] Sinusoidal Encoding: Used to encode timesteps. During the diffusion process, different timesteps require different representations. Sinusoidal encoding can convert timestep information into a continuous vector representation, providing the model with information about time and helping it make appropriate predictions at different diffusion steps.

[0060] Timestep: This represents the different stages of the diffusion process. It is a discrete variable that changes gradually as the diffusion process progresses. The model gradually removes noise based on the different timesteps, generating a latent representation that is closer to the true image.

[0061] Guidance: Related to text information, etc., used to guide the model to generate images that match the text description.

[0062] Ids: Identifiers (plural form), discrete variables, token id (token identification), used to distinguish different image or text tokens.

[0063] RoPE: Rotary Positional Encoding, a position encoding method.

[0064] Latent: noise variable, a random tensor, following a Gaussian distribution.

[0065] Linear layer: It is used to perform linear transformation on the input and map the input features to another feature space to meet the needs of subsequent calculations.

[0066] The complete processing flow of the target word art generation model is as follows: Input processing: 1. Text Encoding: The user-entered text (prompt) is fed into two text encoders, CLIP and T5. CLIP generates pooled features, which are processed by a multi-layer MLP layer. T5 outputs text features, which are then processed by a linear layer.

[0067] 2. Timestep encoding: The timestep information is encoded using sinusoidal encoding. Together with the guidance information, it is processed by MLP and then fused with the pooled feature processing results of CLIP. The fused text-related information is the first text feature.

[0068] The core module processes: The fused first text feature, the Linear-processed T5 output, together with the initial Latent (image potential representation, initially a Gaussian noise tensor), the first image code (for the artistic word generation model or image generation model in this application specification, it is also necessary to add the VAE Encoder encoding and the Linear-processed image features, that is, the first image code), and Ids (identifier) are input into a module composed of multiple MM-DiT-Block and Single-DiT-Block, and the second fused feature is output after processing.

[0069] Modulation and output: 1. Modulation: The second fused feature processed by the above modules passes through the Modulation module, which generates parameters such as scale, shift, and gate based on information such as text semantics to modulate the latent features. The modulated latent features are then processed by the Linear layer and fed back to the module for iterative processing (repeating the core module processing process T times).

[0070] 2. Image Generation: The iteratively processed Latent is input into the VAE Decoder (Variational Autoencoder Decoder), which converts the Latent features from (n, 4c, h, w) dimensions to (n, c, 2h, 2w). Finally, the decoder generates the target word art image (Image) with dimensions of (n, 3, 16h, 16w), where n is the batch size, c is the number of channels, h is the height, and w is the width.

[0071] In summary, the embodiments of this specification clearly demonstrate the complete process from input to output by elaborating on the functions and synergy of each module of the artistic word generation model, ensuring that the generated artistic word images are highly consistent with user needs in terms of style and details.

[0072] In the artistic word creation described in the embodiments of this specification, glyph design breaks through traditional constraints. By combining techniques such as generative adversarial networks and style transfer, elements such as calligraphic brushstrokes and graffiti textures can be incorporated into stroke forms, or structures can be rearranged through neural networks to create variations that combine readability with artistic expression. For example, dynamic glyphs generated by simulating fluid deformation using physical models, or personalized layouts optimized through reinforcement learning scenario feedback, elevate glyphs from information carriers to creative expression media, enabling widespread application in advertising, cultural and creative industries, and meeting diverse visual needs.

[0073] Step 106: Based on the target content text and the target word art image, generate a target image corresponding to the target content text.

[0074] Specifically, after generating the target wordart image, it is necessary to apply the target wordart image to the appropriate position of the target image according to the target content text, and determine the layout of other elements in the target image according to other descriptions in the target content text.

[0075] Specific instructions, such as Figure 4 As shown, Figure 4This is a schematic diagram of generating a first target image in an image generation method provided in an embodiment of this specification. The target content text input by the user is "This title is at the top. On the right side, there is a cartoon character of the Monkey King holding a staff. Below the title, there are two speech bubbles, one with stars around it and the other with a red outline. On the left side, there is a cartoon character dressed in red and yellow, sitting with hands in a prayer position. On the bottom right, there is a large blank white rectangle with afolded corner. The background is green with a blue border.", which means "The title is at the top. On the right, there is a cartoon image of Sun Wukong holding a golden hoop. Below the title, there are two speech bubbles, one surrounded by stars and the other with a red border. On the left, there is a cartoon character wearing red and yellow clothes, with his hands folded in a prayer position. On the bottom right, there is a large blank white rectangle with afolded corner. The background is green with a blue border." The word art title image "Journey to the West" can be the target word art image generated by the target word art generation model in the previous step. The target word art image is placed above the image according to the target content text. On the right, there is a cartoon image of the Monkey King holding a golden hoop. Below the title are two speech bubbles, one with stars around it and the other with a red border. On the left, a cartoon character dressed in red and yellow sits with his hands folded. In the lower right corner, there's a large, empty white rectangle with one corner folded. The background is green with a blue border.

[0076] Furthermore, in the embodiments of this specification, the target image can be generated by a target image generation model, and based on the target content text and the target art word image, a target image corresponding to the target content text is generated, including: inputting the target content text and the target art word image into the target image generation model; the target image generation model processes the target content text and the target art word image to obtain a target image corresponding to the target content text.

[0077] Specifically, the target image generation model, trained with a large amount of sample data, is able to accurately understand and generate images that align with the intended target content text. The target image generation model can simultaneously receive both the target content text and the target word art image. Pixel resizing is also required before the target word art is input. Since the target word art image will ultimately be placed within the target image, and typically only occupies a small portion of the target image, it needs to be scaled.

[0078] Based on this, the adjusted target art word image and the target content text are input into the target image generation model together. The target image generation model integrates multi-level feature extraction to generate an image that is highly consistent with the target content text, ensuring the harmony and unity of the art word and image elements.

[0079] Continuing with the previous example, after the user enters the target content text and the target WordArt model generates the target WordArt image, the pixel size of the target WordArt image is resized to 256x256, ensuring that both the width and height are integer multiples of 16 and smaller than the target image size. The resized WordArt image and the target content text are then fed into the target image generation model. The model processes them and produces a complete handwritten poster image that includes the target WordArt image, the Monkey King, the cartoon character in red and yellow clothing, and a dialog box.

[0080] In summary, the embodiments of this specification, by introducing a target image generation model, achieve a high degree of integration between artistic fonts and handwritten newspaper content, generating high-quality images with high accuracy. Users only need to input simple input to enjoy an efficient and easy-to-use creative experience that meets diverse style needs.

[0081] Furthermore, in order to accurately fuse the target content text and the target art word image, the target image generation model includes: a second text encoder, a second image encoder and an image generation submodule; accordingly, the target image generation model processes the target content text and the target art word image to obtain the target image, including: the second text encoder encodes the target content text to obtain a second text feature; the second image encoder encodes the target art word to obtain a second image code; obtains a second noise variable, splices the second noise variable, the second text feature and the second image encoder to obtain a second splicing feature; inputs the second splicing feature into the image generation submodule to generate the target image.

[0082] Specifically, the second text encoder is a module in the target image generation model used to encode text-type data and is responsible for extracting key features from the target content text. The second image encoder is a module in the target image generation model used to encode image-type data and is responsible for extracting key features from the target art word image. The second noise variable is used to introduce randomness and enhance image diversity. The image generation submodule is used to accurately capture the semantic association between text and image and generate personalized target images that meet user needs. The specific image generation submodule can be implemented using the transformer architecture.

[0083] Based on this, after obtaining the target content text and the target art word image, the target image generation model inputs the target content text into the second text encoder for encoding processing to obtain the second text feature, and inputs the target art word image into the second image encoder for encoding processing to obtain the second image code, randomly samples from the normal distribution to obtain the second noise variable, splices the second noise variable, the second text feature and the second image code to obtain the second splicing feature, and inputs it into the image generation sub-module, and finally generates a personalized target image that integrates the art word and the hand-written newspaper content.

[0084] Continuing with the previous example, the target image generation model takes the target content text and the target art image "Journey to the West". The target content text is fed into the target image generation model's text encoders T5 and CLIP for processing, yielding the second text feature A2. The target art image is fed into the target image generation model's image encoder VAE for processing, yielding the second image encoding B2. A second noise variable C2 is randomly sampled from a normal distribution. The three features C2, A2, and B2 are concatenated along the second dimension using torch.cat((txt, img), 1) , maintaining the batch size. The other dimensions are concatenated to yield the second concatenated feature D2. D2 is fed into the Transformer block for feature fusion and enhancement, outputting a personalized target image consistent with the target content text. This image not only retains the "Journey to the West" art style but also incorporates the details of the target content text.

[0085] The basic architecture adopted by the target image generation model in the embodiments of this specification is the same as the basic architecture adopted by the target art word generation model in the aforementioned embodiments. The more specific processing process of the target image generation model can refer to the processing process of the target art word generation model in the aforementioned embodiments, which will not be repeated here.

[0086] In summary, the embodiments of this specification achieve efficient fusion of text and images through the synergy of multiple modules, ensuring that the generated target image not only conforms to the style of artistic fonts but also accurately conveys the semantics of the text, thereby improving the practicality and personalized expressiveness of the image generation model.

[0087] The following is combined with Figure 5 , taking the application of the image generation method provided in this specification in the production of handwritten newspapers as an example, the image generation method is further explained. Figure 5 This is a flowchart of the processing process of an image generation method provided by an embodiment of this specification in a hand-written newspaper production scenario, which specifically includes the following steps.

[0088] Step 502: Input the artistic word description text and the standard font image.

[0089] Specifically, the standard font image, also known as the conditional image, is used to limit features such as strokes, structure, or font of the target artistic font to be generated.

[0090] Step 504: Generate an artistic word model.

[0091] Specifically, the word art generation model, also known as the target word art generation model, is a qualified word art generation model obtained after training with a large amount of sample data. It can process the input description text and standard glyph images and output target word art images that meet the requirements.

[0092] Step 506: Output the artistic title image.

[0093] Specifically, the word art title image is also the target word art image, and the word art title image can be used as input for the subsequent hand-written newspaper generation model.

[0094] Step 508: Input the handwritten newspaper description text and artistic title image.

[0095] Specifically, the hand-written newspaper description text is also the target content text. The hand-written newspaper description text needs to fully describe the overall content layout, theme style and required elements of the hand-written newspaper. The artistic title image, as the conditional image of the hand-written newspaper generation model, needs to be scaled to 256x256 pixels to ensure that the width and height are both integer multiples of 16.

[0096] Step 510: Generate a handwritten newspaper model.

[0097] Specifically, the hand-written newspaper generation model, also known as the target image generation model, is a qualified model obtained after training based on a large amount of sample data, and can process the input hand-written newspaper description text and artistic title images.

[0098] Step 512: Output the handwritten newspaper image.

[0099] Specifically, the hand-written newspaper picture, that is, the target image, combines the content elements of the artistic title and the hand-written newspaper description text, presenting a unique style and theme characteristics, meeting the user's personalized needs and improving creation efficiency.

[0100] It should be noted that the word art generation model and the handwritten newspaper generation model in the embodiments of this specification can operate independently or collaboratively. When operating independently, the word art generation model only requires the input of the word art description text and the standard font image, while the handwritten newspaper generation model requires the user to provide the word art title image while inputting the handwritten newspaper description text. When working collaboratively, the user can directly input the word art description text, the handwritten newspaper description text, and the standard font image. The word art generation model will generate the word art title image, which will serve as the input condition image for the handwritten newspaper generation model, thereby automatically completing the generation of the handwritten newspaper.

[0101] To sum up, the embodiment of this specification combines the artistic word generation model with the hand-written newspaper generation model. Users only need to simply input descriptive text and standard fonts to efficiently generate artistic word and hand-written newspaper. This method simplifies the creation process, lowers the design threshold, supports diversified style customization, meets the needs of different scenarios, changes the cumbersome steps of traditional hand-written newspaper production, realizes one-click automatic generation, simplifies the creation process, improves the generation effect, and provides users with an efficient and convenient personalized content generation solution.

[0102] The following combined Figure 6 , the training method of the artistic word generation model is further explained. Figure 6 This is a flowchart of a processing process of an artistic word generation model training method provided by an embodiment of this specification, which specifically includes the following steps.

[0103] Step 602: Acquire a plurality of first training samples, where the first training samples include word art description text samples, conditional image samples, and word art image samples.

[0104] Specifically, the first training sample is used to perform preliminary training on the initial word art generation model. The word art description text sample corresponds to the word art description text in the aforementioned embodiment, and the word art description text sample is used to describe the stylistic features of the word art image sample. The conditional image sample corresponds to the conditional image in the aforementioned embodiment, and the conditional image sample is used to constrain the strokes, structure, and / or glyph shape of the generated first word art image.

[0105] Furthermore, in order to efficiently obtain a large number of training samples during model training, the acquisition of multiple first training samples includes: obtaining a target image sample, segmenting the target image sample, and obtaining an art word image sample corresponding to the target image sample; inputting the art word image sample into a preset multimodal model to generate an art word description text sample corresponding to the art word image sample; rendering the text title corresponding to the art word image in a standard font to obtain a conditional image sample; and using the triple consisting of the art word image sample, the art word description text sample, and the conditional image sample as the first training sample.

[0106] Specifically, the target image sample refers to an image material containing artistic words, which can be obtained from online galleries, design portfolios or pictures uploaded by users. Segmentation processing can be implemented using a segmentation model, for example, segmentation models SAM, FastSAM, MobileSAM, etc. The preset multimodal model can realize the conversion between different modal data, such as converting image data into text descriptions. The preset multimodal model is, for example, GPT-4o, etc. The standard font refers to a specific font specified by the model, such as Dengxian, Songti, Microsoft Yahei and other fonts. The standard font is a font configuration file, that is, a ttf file, which is used to render the text in the text title of the input model into a standard glyph image.

[0107] Based on this, after obtaining the target image sample from the Internet, the segmentation model is used to extract the word art image sample from the target image sample, and then the preset multimodal model is used to generate the word art description text sample corresponding to the word art image sample; then the text title in the word art image is rendered with a standard font to obtain the conditional image sample, and finally a triplet data of <word art image sample, word art description text sample, conditional image sample> is formed. This triplet data is used as the first training sample to train the initial word art generation model. For example, a large number of hand-written poster images are collected from the Internet. For each hand-written poster image, the SAM model is used to extract the artistic word title from the hand-written poster image, and the extracted artistic word title is used as the artistic word image sample; the artistic word title is then input into the GPT-4o model to generate the corresponding artistic word description text sample; the text corresponding to the artistic word title is rendered in a standard font to obtain a conditional image sample; and finally, it is combined into a <artistic word image sample, artistic word description text sample, conditional image sample> triplet as the first training data.

[0108] In summary, the embodiments of this specification obtain target image samples from the network and combine them with the collaborative processing of the segmentation model and the multimodal model to efficiently generate a large number of high-quality artistic word training samples, thereby improving the training efficiency and generation effect of the initial artistic word generation model.

[0109] Step 604: Input the word art description text sample and the conditional image sample into the initial word art generation model to generate a first word art image corresponding to the word art description text sample, and adjust the parameters of the initial word art generation model according to the accuracy of the first word art image relative to the word art image sample until the first training stop condition is met, thereby obtaining an intermediate word art generation model.

[0110] Specifically, the initial word art generation model refers to an untrained initialization model, and the first word art image refers to a word art image generated by the initial model. The first training stopping condition refers to the initial word art generation model reaching a preset accuracy threshold or a training round limit after multiple iterations.

[0111] Based on this, the initial word art generation model is fed with a sample of the word art description text and a sample of the conditional image to generate a first word art image. This first word art image is then compared with the sample word art image to calculate the accuracy, and the model parameters are adjusted accordingly. The same training method is repeated multiple times until the first training stop condition is met, resulting in the initial word art generation model achieving high generation accuracy and stability. At this point, training is stopped and the initial word art generation model is converted into the intermediate word art generation model.

[0112] For example, before model training, it is necessary to construct an initial word art generation model. The initial word art generation model uses an improved version of the FLUX model with the Transformer structure as the core as the backbone network, and the learning rate is set to 1. In the initial training stage, a total of 761 groups of first training samples are constructed to train the initial word art generation model. According to the accuracy of generating the first word art image relative to the word art image sample, the weight parameters related to Q, K, and V of the attention calculation part corresponding to the conditional image features in the model are adjusted. After 5 rounds of iterative training, the model accuracy reached 92%, meeting the first training stopping condition, and the training was stopped to obtain the intermediate word art generation model.

[0113] In summary, the embodiments of this specification conduct preliminary training on the initial word art generation model through refined model training and parameter optimization, ensuring that it has high accuracy and stability when generating word art.

[0114] Furthermore, the initial word art generation model includes: a first initial text encoder, a first initial image encoder, and an initial word art generation submodule; accordingly, the word art description text sample and the conditional image sample are input into the initial word art generation model to generate a first word art image corresponding to the word art description text sample, including: inputting the word art description text sample into the first initial text encoder for encoding processing to obtain the corresponding first text sample feature; inputting the conditional image sample into the first initial image encoder for encoding processing to obtain the corresponding first image sample code; obtaining a first noise variable sample, splicing the first noise variable sample, the first text sample feature, and the first image sample code to obtain a first spliced sample feature; inputting the first spliced sample feature into the initial word art generation submodule to generate a first word art image.

[0115] Specifically, the overall structure of the initial artistic word model in this process is consistent with the structure of the target artistic word generation model in the aforementioned embodiment, and the model reasoning process of the initial artistic word generation model to generate the first artistic word image is similar to the reasoning process of the target artistic word generation model in the aforementioned embodiment. The specific implementation process can refer to the aforementioned embodiment, and the embodiments of this specification will not be repeated here.

[0116] Step 606: Filter a first target word art image from the plurality of first word art images and construct a first target training sample.

[0117] Specifically, after the first training is completed, the model performance needs to be further improved, and the intermediate art word generation model needs to be further trained. At this time, a new training sample can be constructed based on the first art word image generated by the initial art word generation model. However, since the accuracy or correctness of the first art word image may be biased, it needs to be screened and evaluated, and images with higher accuracy must be selected as new training samples to ensure the effectiveness of subsequent training and further improvement of model performance.

[0118] Furthermore, in order to further improve the performance of the artistic word generation model, after the initial artistic word generation model is preliminarily trained based on the first training sample, it is necessary to further construct a more accurate first target training sample and further train the intermediate artistic word generation model. The screening of the first target artistic word image from multiple first artistic word images and the construction of the first target training sample include: based on the accuracy of the first artistic word image relative to the artistic word image sample, screening the first target artistic word image with an accuracy higher than a first accuracy threshold from the first artistic word image; obtaining the target artistic word description text sample corresponding to the first target artistic word image; and using the triple consisting of the first target artistic word image, the target artistic word description text sample and the conditional image sample as the first target training sample.

[0119] Specifically, the first accuracy threshold is a standard for determining whether the first artistic word image is qualified.

[0120] Based on this, if the accuracy of the first artistic word image relative to the artistic word image sample is greater than the first accuracy threshold, it indicates that the first artistic word image has a high accuracy and can be used as a valid sample for the next training. That is, the first artistic word image is used as the first target artistic word image. Otherwise, it should be eliminated to ensure the quality of the training sample. After selecting the first target artistic word image, to improve the accuracy of the sample, the style of the first target artistic word image can be manually described to obtain a target artistic word description text sample. The triple data <first target artistic word image, the target artistic word description text sample, the conditional image sample> is used as the first target training sample.

[0121] For example, if the accuracy of a first art word image exceeds 85%, it will be selected as the first target art word image. The content of the first target art word image is "Thank you teacher", and the corresponding target art word description text sample can be described as "the font strokes are thick and even, the corners are smooth and rounded, and the color is warm". Combined with the conditional image sample, a triplet <"Thank you teacher" first target art word image, "the font strokes are thick and even, the corners are smooth and rounded, and the color is warm" target art word description text sample, standard font picture> is formed as the first target training sample for further training the intermediate art word generation model.

[0122] In summary, the embodiments of this specification ensure the high quality of the first target training sample by combining precise screening with manual description, thereby effectively improving the performance and generation effect of the intermediate artistic word generation model, and laying a solid foundation for the optimization of subsequent models.

[0123] Step 608: The intermediate artistic word generation model is trained based on the first target training sample until a second training stop condition is met, thereby obtaining a target artistic word generation model.

[0124] Specifically, during the training of the intermediate word art generation model, supervised training was performed for 40 epochs with a learning rate of 1 to ensure that the model fully learned the characteristics of the triplet data. During training, a weighted mean squared error loss function was used to continuously optimize the model parameters, improving the accuracy and aesthetics of the generated word art.

[0125] Furthermore, the process of further training the intermediate art word generation model based on the first target training sample is as follows: the intermediate art word generation model is trained based on the first target training sample until the second training stop condition is met to obtain the target art word generation model, including: inputting the target art word description text sample and the conditional image sample into the initial art word generation model, generating a second art word image corresponding to the target art word description text sample, and adjusting the parameters of the intermediate art word generation model according to the accuracy of the second art word image relative to the first target art word image, until the second training stop condition is met to obtain the target art word generation model.

[0126] The intermediate word art generation model includes: a first intermediate text encoder, a first intermediate image encoder and an intermediate word art generation submodule; accordingly, the target word art description text sample and the conditional image sample are input into the initial word art generation model to generate a second word art image corresponding to the target word art description text sample, including: inputting the target word art description text sample into the first intermediate text encoder for encoding processing to obtain the corresponding first target text sample feature; inputting the conditional image sample into the first intermediate image encoder for encoding processing to obtain the corresponding first target image sample code; obtaining a first target noise variable sample, splicing the first target noise variable sample, the first target text sample feature and the first target image sample code to obtain the first target splicing sample feature; inputting the first target splicing sample feature into the intermediate word art generation submodule to generate the second word art image.

[0127] Specifically, the overall structure of the intermediate artistic word model in this process is consistent with the structure of the target artistic word generation model in the aforementioned embodiment, and the model reasoning process of the intermediate artistic word generation model to generate the second artistic word image is similar to the reasoning process of the target artistic word generation model in the aforementioned embodiment. The specific implementation process can refer to the aforementioned embodiment, and the embodiments of this specification will not be repeated here.

[0128] To sum up, the embodiments of this specification have undergone two model trainings, and the first target training samples used in the second model training have been precisely screened and combined with manual descriptions, thereby ensuring the efficiency and accuracy of the model training, further improving the performance of the target art word generation model, and making the generated art word images more refined and in line with expectations, providing reliable guarantees for subsequent applications.

[0129] The following combined Figure 7 , further describes the image generation model training method. Figure 7 A flowchart of a processing process of an image generation model training method provided by an embodiment of this specification is shown, which specifically includes the following steps.

[0130] Step 702: Acquire a plurality of second training samples, where the second training samples include target content text samples, word art image samples, and target image samples.

[0131] Specifically, the second training sample is used to perform preliminary training on the initial image generation model, wherein the target content text sample corresponds to the target content text in the aforementioned embodiment, the target content text sample is used to describe information such as the position of the artistic characters and decorative elements contained in the target image sample, the overall style of the image, etc., the artistic character image sample corresponds to the artistic character image in the aforementioned embodiment, and the target image sample corresponds to the target image in the aforementioned embodiment.

[0132] Furthermore, in order to efficiently obtain a large number of training samples, the method of obtaining multiple second training samples includes: obtaining a target image sample, segmenting the target image sample, and obtaining an art word image sample corresponding to the target image sample; inputting the target image sample into a preset multimodal model to generate a target content text sample corresponding to the target image sample; and using a triple consisting of the target image sample, the art word image sample, and the target content text sample as a second training sample.

[0133] Specifically, the method for acquiring and processing target image samples, word art image samples, and target content text samples in this process is similar to the method for constructing the first training sample in the aforementioned embodiment. The difference is that the aforementioned embodiment is word art generation, and therefore, it is necessary to input word art image samples into a preset multimodal model to obtain word art description text samples corresponding to the word art image samples. In contrast, the present embodiment is image generation, and it is necessary to input target image samples into a preset multimodal model to obtain target content text samples corresponding to the target image samples. Other steps can refer to the aforementioned embodiment, and the embodiments of this specification will not be repeated here.

[0134] In summary, the embodiment of this specification preliminarily constructs a second training sample to facilitate subsequent preliminary adjustment of the initial image generation model and optimize the model parameters, thereby significantly improving the accuracy and stability of the model in the generation of artistic fonts and hand-written newspapers, ensuring the high quality and fit of the final output results.

[0135] Step 704: Input the target content text sample and the artistic word image sample into the initial image generation model to generate a first image corresponding to the target content text sample. According to the accuracy of the first image relative to the target image sample, adjust the parameters of the initial image generation model until the third training stop condition is met, thereby obtaining an intermediate image generation model.

[0136] Specifically, the initial image generation model is an untrained image generation model, and the first image is an image generated by the initial image generation model based on input training samples. The third training stopping condition is that the initial image generation model reaches a preset accuracy threshold or an upper limit of training rounds after multiple iterations.

[0137] The model reasoning process and model iteration process in the specific model training are similar to the reasoning process and iteration process of the initial artistic word generation model in the aforementioned embodiment. The specific process can refer to the aforementioned embodiment, and the embodiments of this specification will not be repeated here.

[0138] In summary, the embodiments of this specification continuously optimize the model parameters by comparing the prediction results of the initial model with the target image samples, ensuring that the generated image is closer to the real sample in terms of details and overall effect, thereby improving the expressiveness and reliability of the model in practical applications.

[0139] Furthermore, the initial image generation model includes: a second initial text encoder, a second initial image encoder and an initial image generation submodule; the target content text sample and the art word image sample are input into the initial image generation model to generate a first image corresponding to the target content text sample, including: inputting the target content text sample into the second initial text encoder for encoding processing to obtain the corresponding second text sample feature; inputting the art word image sample into the second initial image encoder for encoding processing to obtain the corresponding second image sample code; obtaining a second noise variable sample, splicing the second noise variable sample, the second text sample feature and the second image sample code to obtain a second spliced sample feature; inputting the second spliced sample feature into the initial image generation submodule to generate the first image.

[0140] Specifically, the model reasoning process in this process is consistent with the reasoning process of the target image generation model in the aforementioned embodiment. Please refer to the aforementioned embodiment for details, and this specification will not repeat them here.

[0141] Step 706: Filter a first target image from the plurality of first images and construct a second target training sample.

[0142] Specifically, after preliminary training of the initial image generation model, further training needs to be strengthened to improve model performance. Therefore, more accurate training samples need to be further constructed.

[0143] Furthermore, the method of screening the first target image from multiple first images and constructing the second target training sample includes: screening the first target image with an accuracy higher than a second accuracy threshold from the first image based on the accuracy of the first image relative to the target image sample; obtaining the target content description text sample corresponding to the first target image; and using the triple consisting of the first target image, the target content description text sample and the artistic word image sample as the second target training sample.

[0144] Specifically, the specific implementation process of screening the first target image and obtaining the target content description text sample in this process is similar to the process of screening the first target art word image and obtaining the target art word description text sample in the aforementioned embodiment. Please refer to the aforementioned embodiment for details, and the embodiments of this specification will not be repeated here.

[0145] In summary, the embodiments of this specification further optimize model performance by accurately screening and constructing high-quality training samples, ensuring that the details and overall effect of the generated images are closer to real samples, and improving the stability and accuracy of the model in practical applications.

[0146] Step 708: Train the intermediate image generation model based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model.

[0147] Specifically, the intermediate image generation model is trained based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model, including: inputting the target content description text sample and the art word image sample into the initial art word generation model to generate a second image corresponding to the target content text sample, and adjusting the parameters of the intermediate image generation model according to the accuracy of the second image relative to the first target image until the fourth training stop condition is met to obtain a target image generation model.

[0148] The intermediate image generation model includes: a second intermediate text encoder, a second intermediate image encoder and an intermediate image generation submodule; accordingly, the target content description text sample and the art word image sample are input into the initial art word generation model to generate a second image corresponding to the target content text sample, including: inputting the target content description text sample into the second intermediate text encoder for encoding processing to obtain the corresponding second target text sample feature; inputting the art word image sample into the second intermediate image encoder for encoding processing to obtain the corresponding second target image sample code; obtaining a second target noise variable sample, splicing the second target noise variable sample, the second target text sample feature and the second target image sample code to obtain the second target splicing sample feature; inputting the second target splicing sample feature into the intermediate image generation submodule to generate a second art word image.

[0149] It should be noted that the training process of the above-mentioned intermediate image generation model is similar to the training process of the intermediate image generation model in the aforementioned embodiment. Please refer to the aforementioned embodiment for details, and the embodiments of this specification will not be repeated here.

[0150] In summary, the embodiments of this specification ensure the stability and accuracy of the model when processing complex text and image features through preliminary training of the initial image generation model and further enhanced training.

[0151] The following is combined with Figure 8 Taking the application of the image generation method provided in this specification in the production of handwritten newspapers as an example, the training method of the artistic word generation model and the image generation model are further explained. Figure 8 This is a flowchart of the processing process of an artistic character generation model training method and an image generation model training method in a hand-written newspaper production scenario provided by an embodiment of this specification, which specifically includes the following steps.

[0152] Step 802: Collect handwritten newspaper pictures from the Internet.

[0153] Specifically, the hand-written newspaper picture is the target image sample.

[0154] Step 804: Cut out the artistic title from the handwritten newspaper image.

[0155] Specifically, the word art title is also a word art image sample, and the word art title can be extracted through a preset segmentation model.

[0156] Step 806: Use a standard font to render the text title corresponding to the artistic word.

[0157] Specifically, the corresponding conditional image can be obtained after rendering the text title with a standard font.

[0158] Step 808: Construct a data pair of <standard font image, artistic word description, artistic word image>.

[0159] The specific standard font image is also the conditional image, the word art description is the word art description text sample, the word art description is generated through a preset multimodal model, and the word art image (word art title) is the word art image sample.

[0160] Step 810: Pre-training of the artistic word generation model.

[0161] Specifically, the word art generation model in this step, that is, the initial word art generation model, is preliminarily trained based on the data generated in the previous step.

[0162] Step 812: Use the model to generate artistic characters with various styles, and filter and construct data pairs.

[0163] Specifically, the artistic word, that is, the first artistic word image, after the initial training of the artistic word generation model, multiple artistic words with different styles can be generated. Among these artistic words, the artistic words with higher accuracy are selected to construct a new data pair, that is, the first target training sample, which is used to further fine-tune the model.

[0164] Step 814: Train the model using the constructed data.

[0165] Specifically, after constructing the new training samples in the previous step, the artistic word generation model is trained again.

[0166] Step 816: Test the accuracy of the artistic word generated by the model.

[0167] The second word art image generated by the word art generation model is compared with the word art image sample to determine its accuracy.

[0168] Step 818: Does the accuracy of generating artistic words meet the requirements? Determine whether the accuracy of the artistic words generated by the model during training meets the requirements. If not, return to step 812. If so, execute step 820.

[0169] Step 820: Save the optimal model.

[0170] Save the currently trained word art generation model that meets the standards as the target word art generation model.

[0171] Step 822: Construct a data pair of <artistic word title picture, handwritten newspaper description, handwritten newspaper picture>.

[0172] Specifically, the word art title image is the word art image sample, the hand-written newspaper description and target content text sample, and the hand-written newspaper image is the target image sample.

[0173] Step 824: Pre-training the handwritten newspaper generation model.

[0174] Specifically, the hand-written newspaper generation model in this step, namely the initial hand-written newspaper generation model, is preliminarily trained based on the data constructed in step 822.

[0175] Step 826: Use the model to generate handwritten newspapers in various styles, and filter and construct data pairs.

[0176] Specifically, the hand-written newspaper is the first hand-written newspaper image generated by the hand-written newspaper generation model after preliminary training. The hand-written newspaper with higher accuracy is selected from them to construct a new data pair, namely the second target training sample, for further fine-tuning the model.

[0177] Step 828: Train the model using the constructed data.

[0178] Specifically, the second target training sample constructed in the previous step is used to train the handwritten newspaper generation model again.

[0179] Step 830: Test the accuracy of the handwritten newspaper generated by the model.

[0180] Specifically, the second hand-written newspaper image generated by the hand-written newspaper generation model is compared with the hand-written newspaper image sample to determine its accuracy.

[0181] Step 832: Does the accuracy of the generated handwritten newspaper meet the requirements? Specifically, determine whether the accuracy of the handwritten newspaper generated by the model during training meets the requirements. If not, return to step 826. If so, execute step 834.

[0182] Step 834: Save the optimal model.

[0183] The currently trained hand-written newspaper generation model that meets the standards is saved as the target hand-written newspaper generation model.

[0184] In summary, the embodiments of this specification utilize handwritten newspaper production data to train and fine-tune the initial handwritten newspaper generation model and the initial image generation model, ultimately achieving efficient generation of both artistic fonts and handwritten newspapers. Multiple iterations of model optimization have improved generation quality and accuracy, ensuring a seamless integration of artistic fonts and handwritten newspapers.

[0185] See also Figure 9 , Figure 9This is a schematic diagram illustrating generating a second target WordArt image in an image generation method provided in one embodiment of this specification. In this embodiment, the WordArt description text entered by the user is "This characters are presented in a playful, handwritten font with a bright green color. The font features whimsical, uneven strokes with a casual and lively appearance. The overall design is cheerful and inviting, suitable for a theme focused on spring outings." The input conditional image is "Spring Outing." Accordingly, the generated target WordArt image is "Spring Outing," displayed below. Its color is green, and its overall style closely matches the user's description.

[0186] See also Figure 10 , Figure 10This is a schematic diagram of generating a second target image in an image generation method provided in one embodiment of this specification. In this embodiment, the target content text entered by the user is "This title is at the left. On the top right, there is a large tree with green leaves and flowers around it. Below the tree, there is a rectangular blank writing frame bordered with small leaves and blossoms. In the middle right, there is a beautifully decorated cake with "Happy Mother's Day" written on it, surrounded by small stars. On the bottom right, there is a gift box with a ribbon, surrounded by smaller gift boxes and sparkles. The background is light pastel blue with scattered confetti and butterflies." The target word art image is the words "Mother's Day." Accordingly, the generated target image is the image shown below, which is brightly colored, well-organized, and perfectly presents all elements of the user's description.

[0187] Corresponding to the above method embodiment, this specification also provides an image generating device embodiment, Figure 11 FIG1 shows a schematic diagram of the structure of an image generating device provided by an embodiment of this specification. Figure 11 As shown, the device includes: The data acquisition module 1102 is configured to acquire a description text of a word art, a target content text, and a conditional image, wherein the conditional image is used to constrain the shape of the target word art; The word art generation module 1104 is configured to process the conditional image based on the word art description text to generate a target word art image corresponding to the conditional image; The image generation module 1106 is configured to generate a target image corresponding to the target content text based on the target content text and the target word art image.

[0188] Optionally, the word art generation module is further configured to input the word art description text and the conditional image into a target word art generation model; the target word art generation model processes the word art description text and the conditional image to obtain a target word art image corresponding to the conditional image.

[0189] Optionally, the target word art generation model includes: a first text encoder, a first image encoder, and a word art generation submodule; the word art generation module includes: A first text encoding submodule is configured to encode the word art description text using the first text encoder to obtain a corresponding first text feature; A first image encoding submodule is configured to perform encoding processing on the conditional image by the first image encoder to obtain a corresponding first image code; a first feature splicing submodule, which obtains a first noise variable, and splices the first noise variable, the first text feature, and the first image code to obtain a first splicing feature; The first art word generation submodule inputs the first splicing feature into the art word generation submodule to generate a target art word image.

[0190] Optionally, the image generation module is further configured to input the target content text and the target artistic word image into a target image generation model; the target image generation model processes the target content text and the target artistic word image to obtain a target image corresponding to the target content text.

[0191] Optionally, the target image generation model includes: a second text encoder, a second image encoder, and an image generation submodule; correspondingly, the image generation module includes: A second text encoding submodule is configured to perform encoding processing on the target content text by the second text encoder to obtain a second text feature; A second image encoding submodule is configured to perform encoding processing on the target artistic word by the second image encoder to obtain a second image code; a second feature splicing submodule configured to obtain a second early noise variable, and splice the second noise variable, the second text feature, and the second image encoder to obtain a second splicing feature; The image generation submodule is configured to input the second stitching feature into the image generation submodule to generate the target image.

[0192] The above is a schematic diagram of an image generation device according to this embodiment. It should be noted that the technical solution of this image generation device and the technical solution of the aforementioned image generation method are based on the same concept. For details not described in detail in the technical solution of the image generation device, please refer to the description of the technical solution of the aforementioned image generation method.

[0193] Corresponding to the above method embodiment, this specification also provides an embodiment of an artistic word generation model training device, Figure 12 FIG. 1 shows a schematic diagram of a structure of an artistic character generation model training device provided by an embodiment of this specification. Figure 12 As shown, the device includes: The first word art sample construction module 1202 is configured to obtain a plurality of first training samples, wherein the first training samples include word art description text samples, conditional image samples, and word art image samples; The first word art model training module 1204 is configured to input the word art description text sample and the conditional image sample into an initial word art generation model, generate a first word art image corresponding to the word art description text sample, and adjust the parameters of the initial word art generation model according to the accuracy of the first word art image relative to the word art image sample until a first training stop condition is met, thereby obtaining an intermediate word art generation model; The second word art sample construction module 1206 is configured to select a first target word art image from a plurality of first word art images and construct a first target training sample; The second word art model training module 1208 is configured to train the intermediate word art generation model based on the first target training sample until a second training stop condition is met to obtain a target word art generation model.

[0194] The above is a schematic diagram of a device for training a model for generating artistic characters according to this embodiment. It should be noted that the technical solution of this device for training an artistic character generation model is based on the same concept as the technical solution of the aforementioned method for training an artistic character generation model. For any details not described in detail in the technical solution of the device for training an artistic character generation model, please refer to the description of the technical solution of the aforementioned method for training an artistic character generation model.

[0195] Corresponding to the above method embodiment, this specification also provides an embodiment of an image generation model training device, Figure 13 FIG1 shows a schematic diagram of the structure of an image generation model training device provided by an embodiment of this specification. Figure 13 As shown, the device includes: The first image sample construction module 1302 is configured to obtain a plurality of second training samples, wherein the second training samples include target content text samples, word art image samples, and target image samples; The first image model training module 1304 is configured to input the target content text sample and the word art image sample into an initial image generation model, generate a first image corresponding to the target content text sample, and adjust parameters of the initial image generation model based on the accuracy of the first image relative to the target image sample until a third training stop condition is met, thereby obtaining an intermediate image generation model. The second image sample construction module 1306 is configured to filter a first target image from a plurality of first images and construct a second target training sample; The second image model training module 1308 is configured to train the intermediate image generation model based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model.

[0196] The above is a schematic diagram of an image generation model training device according to this embodiment. It should be noted that the technical solution of this image generation model training device and the technical solution of the aforementioned image generation model training method are based on the same concept. For details not described in detail in the technical solution of the image generation model training device, please refer to the description of the technical solution of the aforementioned image generation model training method.

[0197] Figure 14 14 shows a block diagram of a computing device 1400 according to one embodiment of the present disclosure. Components of the computing device 1400 include, but are not limited to, a memory 1410 and a processor 1420. The processor 1420 is connected to the memory 1410 via a bus 1430, and a database 1450 is used to store data.

[0198] Computing device 1400 also includes an access device 1440 that enables computing device 1400 to communicate via one or more networks 1460. Examples of such networks include a public switched telephone network (PSTN), a local area network (LAN), a wide area network (WAN), a personal area network (PAN), or a combination of communication networks such as the Internet. Access device 1440 may include one or more of any type of network interface (e.g., a network interface controller (NIC)) whether wired or wireless, such as an IEEE 802.14 wireless local area network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a universal serial bus (USB) interface, a cellular network interface, a Bluetooth interface, or a near field communication (NFC) interface.

[0199] In one embodiment of the present specification, the above components of the computing device 1400 and Figure 14 Other components not shown in the figure may also be connected to each other, for example, via a bus. Figure 14 The computing device structure block diagram shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art may add or replace other components as needed.

[0200] Computing device 1400 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, personal digital assistant, laptop computer, notebook computer, netbook computer, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or personal computer (PC). Computing device 1400 can also be a mobile or stationary server.

[0201] Among them, the processor 1420 is used to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned image generation method, artistic character generation model training method or image generation model training method.

[0202] The above is a schematic diagram of a computing device according to this embodiment. It should be noted that the technical solution of this computing device is based on the same concept as the technical solution of the aforementioned image generation method, artistic word generation model training method, or image generation model training method. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the aforementioned image generation method, artistic word generation model training method, or image generation model training method.

[0203] An embodiment of the present specification also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the steps of the above-mentioned image generation method, artistic character generation model training method, or image generation model training method.

[0204] The above is a schematic diagram of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium is based on the same concept as the technical solution of the aforementioned image generation method, word art generation model training method, or image generation model training method. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the aforementioned image generation method, word art generation model training method, or image generation model training method.

[0205] An embodiment of the present specification further provides a computer program, wherein when the computer program is executed in a computer, the computer is caused to execute the steps of the above-mentioned image generation method, artistic character generation model training method or image generation model training method.

[0206] The above is a schematic diagram of a computer program according to this embodiment. It should be noted that the technical solution of this computer program is based on the same concept as the technical solution of the aforementioned image generation method, word art generation model training method, or image generation model training method. For details not described in detail in the technical solution of the computer program, please refer to the description of the technical solution of the aforementioned image generation method, word art generation model training method, or image generation model training method.

[0207] The foregoing description of this specification describes specific embodiments. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0208] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, recording medium, USB flash drive, mobile hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electric carrier signal, telecommunication signal, and software distribution medium. It should be noted that the content of the computer-readable medium may be appropriately increased or decreased based on the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media does not include electric carrier signals and telecommunication signals.

[0209] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.

[0210] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0211] The preferred embodiments disclosed above are intended only to help illustrate this specification. The optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made based on the content of the embodiments of this specification. This specification selects and specifically describes these embodiments in order to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can better understand and utilize this specification. This specification is limited only by the claims and their full scope and equivalents.

Claims

1. An image generation method, characterized in that: include: Acquire a word art description text, a target content text, and a conditional image, wherein the conditional image is used to constrain the shape of the target word art; Based on the artistic word description text, the conditional image is processed to generate a target artistic word image corresponding to the conditional image; Based on the target content text and the target word art image, a target image corresponding to the target content text is generated.

2. The image generation method according to claim 1, characterized in that: The step of processing the conditional image based on the wordart description text to generate a target wordart image corresponding to the conditional image includes: Inputting the word art description text and the conditional image into a target word art generation model; The target wordart generation model processes the wordart description text and the conditional image to obtain a target wordart image corresponding to the conditional image.

3. The image generation method according to claim 2, characterized in that: The target artistic word generation model includes: a first text encoder, a first image encoder and an artistic word generation submodule; Accordingly, the target word art generation model processes the word art description text and the conditional image to obtain the target word art image, including: The first text encoder encodes the word art description text to obtain a corresponding first text feature; The first image encoder performs encoding processing on the conditional image to obtain a corresponding first image code; Obtaining a first noise variable, and performing splicing processing on the first noise variable, the first text feature, and the first image code to obtain a first splicing feature; The first splicing feature is input into the artistic word generation submodule to generate a target artistic word image.

4. The image generation method according to claim 1, wherein: Generating a target image corresponding to the target content text based on the target content text and the target artistic word image includes: Inputting the target content text and the target artistic word image into a target image generation model; The target image generation model processes the target content text and the target artistic word image to obtain a target image corresponding to the target content text.

5. The image generation method according to claim 4, characterized in that: The target image generation model includes: a second text encoder, a second image encoder and an image generation submodule; Accordingly, the target image generation model processes the target content text and the target word art image to obtain the target image, including: The second text encoder encodes the target content text to obtain a second text feature; The second image encoder performs encoding processing on the target artistic word to obtain a second image code; Obtaining a second noise variable, and performing splicing processing on the second noise variable, the second text feature, and the second image encoder to obtain a second splicing feature; The second stitching feature is input into the image generation submodule to generate the target image.

6. A method for training an artistic character generation model, characterized in that: include: Acquire a plurality of first training samples, wherein the first training samples include word art description text samples, conditional image samples, and word art image samples; Inputting the word art description text sample and the conditional image sample into an initial word art generation model, generating a first word art image corresponding to the word art description text sample, and adjusting parameters of the initial word art generation model according to the accuracy of the first word art image relative to the word art image sample until a first training stop condition is met, thereby obtaining an intermediate word art generation model; Selecting a first target word art image from a plurality of first word art images and constructing a first target training sample; The intermediate artistic word generation model is trained based on the first target training sample until a second training stop condition is met, thereby obtaining a target artistic word generation model.

7. The method for training an artistic character generation model according to claim 6, characterized in that: The obtaining of a plurality of first training samples comprises: Acquire a target image sample, perform segmentation processing on the target image sample, and obtain an artistic word image sample corresponding to the target image sample; Inputting the word art image sample into a preset multimodal model to generate a word art description text sample corresponding to the word art image sample; Rendering the text title corresponding to the word art image in a standard font to obtain a conditional image sample; A triplet consisting of the word art image sample, the word art description text sample, and the conditional image sample is used as a first training sample.

8. The method for training an artistic character generation model according to claim 6, wherein: The step of selecting a first target artistic word image from a plurality of first artistic word images and constructing a first target training sample includes: selecting, from the first artistic word images, a first target artistic word image having an accuracy rate higher than a first accuracy rate threshold, according to the accuracy rate of the first artistic word image relative to the artistic word image sample; Obtaining a target artistic word description text sample corresponding to the first target artistic word image; A triple consisting of the first target artistic word image, the target artistic word description text sample, and the conditional image sample is used as a first target training sample.

9. The method for training an artistic character generation model according to claim 6, wherein: The initial artistic word generation model includes a first initial text encoder, a first initial image encoder, and an initial artistic word generation submodule; Accordingly, the step of inputting the word art description text sample and the conditional image sample into an initial word art generation model to generate a first word art image corresponding to the word art description text sample includes: Inputting the artistic word description text sample into the first initial text encoder for encoding processing to obtain corresponding first text sample features; Inputting the conditional image sample into the first initial image encoder for encoding processing to obtain a corresponding first image sample code; Obtaining a first noise variable sample, and performing splicing processing on the first noise variable sample, the first text sample feature, and the first image sample code to obtain a first spliced sample feature; The first spliced sample feature is input into the initial artistic word generation submodule to generate a first artistic word image.

10. The method for training an artistic character generation model according to claim 8, wherein: The intermediate artistic word generation model is trained based on the first target training sample until a second training stop condition is satisfied to obtain a target artistic word generation model, including: The target artistic word description text sample and the conditional image sample are input into the initial artistic word generation model to generate a second artistic word image corresponding to the target artistic word description text sample, and the parameters of the intermediate artistic word generation model are adjusted according to the accuracy of the second artistic word image relative to the first target artistic word image until the second training stop condition is met to obtain the target artistic word generation model.

11. The method for training an artistic character generation model according to claim 10, wherein: The intermediate artistic word generation model includes: a first intermediate text encoder, a first intermediate image encoder and an intermediate artistic word generation submodule; Accordingly, the step of inputting the target word art description text sample and the conditional image sample into an initial word art generation model to generate a second word art image corresponding to the target word art description text sample includes: Inputting the target artistic word description text sample into the first intermediate text encoder for encoding processing to obtain corresponding first target text sample features; Inputting the conditional image sample into the first intermediate image encoder for encoding processing to obtain a corresponding first target image sample code; Obtaining a first target noise variable sample, and performing splicing processing on the first target noise variable sample, a first target text sample feature, and the first target image sample code to obtain a first target spliced sample feature; The first target splicing sample feature is input into the intermediate artistic word generation submodule to generate a second artistic word image.

12. A method for training an image generation model, characterized in that: include: Acquire a plurality of second training samples, wherein the second training samples include target content text samples, word art image samples, and target image samples; Inputting the target content text sample and the word art image sample into an initial image generation model to generate a first image corresponding to the target content text sample, and adjusting parameters of the initial image generation model based on the accuracy of the first image relative to the target image sample until a third training stop condition is met, thereby obtaining an intermediate image generation model; Screening a first target image from a plurality of first images and constructing a second target training sample; The intermediate image generation model is trained based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model.

13. The image generation model training method according to claim 12, characterized in that: The obtaining of a plurality of second training samples comprises: Acquire a target image sample, perform segmentation processing on the target image sample, and obtain an artistic word image sample corresponding to the target image sample; Inputting the target image sample into a preset multimodal model to generate a target content text sample corresponding to the target image sample; A triplet consisting of the target image sample, the artistic word image sample, and the target content text sample is used as a second training sample.

14. The image generation model training method according to claim 12, characterized in that: The step of screening a first target image from a plurality of first images and constructing a second target training sample comprises: screening, from the first images, first target images having an accuracy rate higher than a second accuracy rate threshold, according to the accuracy rate of the first images relative to the target image sample; Obtaining a target content description text sample corresponding to the first target image; A triplet consisting of the first target image, the target content description text sample, and the word art image sample is used as a second target training sample.

15. The image generation model training method according to claim 12, characterized in that: The initial image generation model includes: a second initial text encoder, a second initial image encoder and an initial image generation submodule; Inputting the target content text sample and the word art image sample into an initial image generation model to generate a first image corresponding to the target content text sample includes: Inputting the target content text sample into the second initial text encoder for encoding processing to obtain corresponding second text sample features; Inputting the artistic word image sample into the second initial image encoder for encoding processing to obtain a corresponding second image sample code; Obtaining a second noise variable sample, and performing splicing processing on the second noise variable sample, the second text sample feature, and the second image sample code to obtain a second spliced sample feature; The second spliced sample features are input into the initial image generation submodule to generate a first image.

16. The image generation model training method according to claim 14, characterized in that: The step of training the intermediate image generation model based on the second target training sample until a fourth training stop condition is satisfied to obtain a target image generation model includes: The target content description text sample and the artistic word image sample are input into the initial artistic word generation model to generate a second image corresponding to the target content text sample, and the parameters of the intermediate image generation model are adjusted according to the accuracy of the second image relative to the first target image until the fourth training stop condition is met, thereby obtaining the target image generation model.

17. The image generation model training method according to claim 16, characterized in that: The intermediate image generation model includes: a second intermediate text encoder, a second intermediate image encoder and an intermediate image generation submodule; Accordingly, the step of inputting the target content description text sample and the word art image sample into an initial word art generation model to generate a second image corresponding to the target content text sample includes: Inputting the target content description text sample into the second intermediate text encoder for encoding processing to obtain corresponding second target text sample features; Inputting the artistic word image sample into the second intermediate image encoder for encoding processing to obtain a corresponding second target image sample code; Obtaining a second target noise variable sample, and performing splicing processing on the second target noise variable sample, the second target text sample feature, and the second target image sample encoding to obtain a second target spliced sample feature; The second target splicing sample feature is input into the intermediate image generation submodule to generate a second artistic word image.

18. An image generating device, characterized in that: include: A data acquisition module is configured to acquire a description text of a word art, a target content text, and a conditional image, wherein the conditional image is used to constrain the shape of the target word art; An artistic word generation module is configured to process the conditional image based on the artistic word description text to generate a target artistic word image corresponding to the conditional image; The image generation module is configured to generate a target image corresponding to the target content text based on the target content text and the target artistic word image.

19. An artistic word generation model training device, characterized in that: include: A first word art sample construction module is configured to obtain a plurality of first training samples, wherein the first training samples include word art description text samples, conditional image samples and word art image samples; a first word art model training module configured to input the word art description text sample and the conditional image sample into an initial word art generation model, generate a first word art image corresponding to the word art description text sample, and adjust parameters of the initial word art generation model according to an accuracy rate of the first word art image relative to the word art image sample until a first training stop condition is met, thereby obtaining an intermediate word art generation model; A second artistic word sample construction module is configured to select a first target artistic word image from a plurality of first artistic word images and construct a first target training sample; The second artistic character model training module is configured to train the intermediate artistic character generation model based on the first target training sample until a second training stop condition is met to obtain a target artistic character generation model.

20. An image generation model training device, characterized in that: include: A first image sample construction module is configured to obtain a plurality of second training samples, wherein the second training samples include target content text samples, word art image samples and target image samples; a first image model training module configured to input the target content text sample and the word art image sample into an initial image generation model, generate a first image corresponding to the target content text sample, and adjust parameters of the initial image generation model based on an accuracy rate of the first image relative to the target image sample until a third training stop condition is met, thereby obtaining an intermediate image generation model; a second image sample construction module, configured to screen a first target image from a plurality of first images and construct a second target training sample; The second image model training module is configured to train the intermediate image generation model based on the second target training sample until a fourth training stop condition is met to obtain a target image generation model.

21. A computing device, characterized in that include: memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the method described in any one of claims 1-5, 6-11 or 12-17 are implemented.

22. A computer-readable storage medium, characterized in that It stores computer-executable instructions, which, when executed by a processor, implement the steps of the method described in any one of claims 1-5, 6-11 or 12-17.

23. A computer program product, characterized in that The method comprises computer instructions which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 5, 6 to 11 or 12 to 17.

Citation Information

Cited By

  • 3D wordart generation method based on structure-view angle double-stage diffusion model

    CN121414972A

  • Image restoration method and related equipment

    CN121563838A