Image-text design method, apparatus and system, and computer-readable storage medium
Through the combination of the stable diffusion model and the character design model, a vivid and interesting caption effect is achieved in the pictures, solving the problem of conventional text generation in the existing technology, and is suitable for a variety of commercial and creative applications.
Patent Information
- Application Number
- PCT/CN2023/141578
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-25
- Publication Date
- 2025-07-03
AI Technical Summary
When the prior art generates text information in pictures, the generated text is conventional and formal, and it is not interesting enough, making it difficult to achieve vivid and interesting supporting effects.
By receiving prompt information input from users, the stable diffusion model is used to generate the pictures to be texted, and combined with the character design model and the graphic generation model, the semantic text synthesis is realized, including the use of self-encoding modules, denoising modules and conditional encoders, as well as the iterative optimization of the differentiable rasterizer and semantic encoding module in the character design model.
The generated graphic and text combination effect is vivid and interesting, and can quickly provide texts that conform to the artistic conception and theme. It has commercial value and is suitable for promotional posters, product posters, ancient poetry and paintings, and picture book generation.
Smart Images

Figure CN2023141578_03072025_PF_FP_ABST
Abstract
Description
Graphic design method, device, system and computer-readable storage medium Technical Field
[0001] The embodiments of the present disclosure relate to, but are not limited to, the field of image design technology, and in particular to a graphic design method, device, system, and computer-readable storage medium. Background Art
[0002] Currently, the stable diffusion model can generate images, but it cannot generate useful text within them. One method for adding text to images involves using editing software (such as Microsoft Office) and specifying fonts from a font library to write text within a stable text box. However, the text generated by this method is conventional, rigid, and generally uninteresting.
[0003] Summary of the Invention
[0004] The following is a summary of the subject matter described in detail herein. This summary is not intended to limit the scope of the claims.
[0005] The present disclosure also provides a graphic design method, including:
[0006] receiving first prompt information input by a user, and obtaining a picture to be accompanied by a text according to the first prompt information;
[0007] Receiving second prompt information input by the user or obtaining third prompt information based on the picture to be accompanied by text, and generating a semantically-defined picture with text based on the second prompt information or the third prompt information;
[0008] The picture to be accompanied by text and the semantically modified picture with text are combined into a picture with text.
[0009] An embodiment of the present disclosure also provides a graphic design device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the graphic design method described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0010] An embodiment of the present disclosure further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, the graphic design method described in any embodiment of the present disclosure is implemented.
[0011] The present disclosure also provides a graphic design system, including a client and a server, wherein:
[0012] The client is configured to receive first prompt information input by a user and upload the first prompt information to the server, or receive first prompt information and second prompt information input by a user and upload the first prompt information and second prompt information to the server; receive and display a picture with text;
[0013] The server is configured to obtain a picture to be accompanied by a text based on the first prompt information, generate a semantically-matched picture to be accompanied by a text based on the second prompt information, or obtain third prompt information based on the picture to be accompanied by a text, generate a semantically-matched picture to be accompanied by a text based on the third prompt information, combine the picture to be accompanied by a text and the semantically-matched picture to form a picture with a text, and send the picture with a text to the client.
[0014] Other aspects will become apparent upon reading and understanding the drawings and detailed description.
[0015] Summary of the Figures
[0016] The accompanying drawings are intended to provide a further understanding of the technical solutions of the present disclosure and constitute a part of the specification. Together with the embodiments of the present disclosure, they are used to explain the technical solutions of the present disclosure and do not constitute a limitation of the technical solutions of the present disclosure. The shapes and sizes of the components in the drawings do not reflect the actual scale and are intended only to illustrate the contents of the present disclosure.
[0017] FIG1 is a flow chart of a graphic design method according to an exemplary embodiment of the present disclosure;
[0018] 2A and 2B are schematic diagrams of two client user interfaces provided by exemplary embodiments of the present disclosure;
[0019] FIG3 is a schematic structural diagram of a stable diffusion model provided by an exemplary embodiment of the present disclosure;
[0020] FIG4A is a schematic diagram of a Chinese font before deformation provided by an exemplary embodiment of the present disclosure;
[0021] FIG4B is a schematic diagram of the shape of the Chinese font in FIG4A after undergoing character design modeling;
[0022] FIG5A is a schematic structural diagram of a character design model provided by an exemplary embodiment of the present disclosure;
[0023] FIG5B is a schematic diagram of the structure of a graph-text generation model provided by an exemplary embodiment of the present disclosure;
[0024] 6A and 6B are schematic diagrams of two server-side processing flows provided by exemplary embodiments of the present disclosure;
[0025] FIG7 is a flow chart of another graphic design method provided by an exemplary embodiment of the present disclosure;
[0026] FIG8 is a schematic diagram of a picture with text provided by an exemplary embodiment of the present disclosure;
[0027] FIG9 is a schematic structural diagram of a graphic design device provided by an exemplary embodiment of the present disclosure;
[0028] FIG10 is a schematic structural diagram of a graphic design system provided by an exemplary embodiment of the present disclosure.
[0029] Details
[0030] To make the objectives, technical solutions and advantages of the present disclosure more clearly understood, the embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings. It should be noted that, unless there is a conflict, the embodiments and features in the embodiments of the present disclosure can be combined with each other in any manner.
[0031] Unless otherwise defined, the technical or scientific terms used in the embodiments of the present disclosure should have the ordinary meaning understood by people with ordinary skills in the field to which the present disclosure belongs. The words "first", "second" and similar words used in the embodiments of the present disclosure do not indicate any order, quantity or importance, but are only used to distinguish different components. The words "include" or "comprising" and similar words mean that the elements or objects preceding the word include the elements or objects listed after the word and their equivalents, without excluding other elements or objects.
[0032] As shown in FIG1 , an embodiment of the present disclosure provides a graphic design method, including:
[0033] Step 101: Receive first prompt information input by a user, and obtain a picture to be accompanied by a text according to the first prompt information;
[0034] Step 102: Receive second prompt information input by the user or obtain third prompt information based on the image to be accompanied by a text, and generate a semantically-defined accompanying text image based on the second prompt information or the third prompt information;
[0035] Step 103: Combine the image to be accompanied by text and the semantically modified accompanying text image into an image with accompanying text.
[0036] The graphic and text design method provided by the embodiment of the present disclosure obtains a picture to be accompanied by a text according to the first prompt information; generates a semantic accompanying text picture according to the second prompt information or the third prompt information; and combines the picture to be accompanied by a text and the semantic accompanying text picture into a picture with an accompanying text. This method can provide a painting with an accompanying text that is more in line with the artistic conception and theme more quickly, and the generated accompanying text is vivid and interesting, and has great practical value and commercial value.
[0037] In some exemplary embodiments, the picture to be accompanied by the text may be directly provided by the user. In this case, the user directly inputs the picture to be accompanied by the text without inputting the first prompt information.
[0038] In other exemplary embodiments, the picture to be accompanied by a text may be generated based on the first prompt information. In this case, the first prompt information is input into the picture generation model so that the picture generation model outputs the picture to be accompanied by a text.
[0039] The graphic and text design method of the embodiment of the present disclosure can be implemented on a server (that is, both the client interface and the server algorithm run on the server); it can also be implemented jointly by a server and a client, wherein the client can be used to provide a client interface, and the client interface receives the user's input (including the first prompt information, the second prompt information and the third prompt information in the following text, etc.), and sends the user's input to the server. The server is used to run the server algorithm to obtain the picture to be accompanied by the text according to the user's first prompt information, generate a semantic accompanying picture according to the user's second prompt information or according to the picture to be accompanied by the text, synthesize the picture to be accompanied by the text and the semantic accompanying picture into a picture with text, and send the picture with text to the client for display by the client interface.
[0040] Figures 2A and 2B are schematic diagrams of two exemplary client interfaces. For example, as shown in Figures 2A and 2B, the first prompt information may include at least one of the following: a positive prompt word, a negative prompt word, an input picture, etc.
[0041] In the embodiment of the present disclosure, the positive prompt words may include any one or more of the following categories: subject, medium, style, artist, resolution, additional details, color, lighting, etc.
[0042] Subject refers to the subject you want to see in the image. Describe it as detailed as possible to avoid under-description, such as a flying dog or a pancake advertisement. Style refers to the style of the generated image, including illustration, oil painting, or photography. Style refers to the artistic style of the subject image, such as Impressionism, Surrealism, Pop Art, etc. Painter refers to using a specific painter as a reference to generate images in their style. Of course, you can also use multiple painters to generate mixed styles. Resolution refers to the clarity and level of detail of the generated image. Additional details can be used to modify the image. Tone refers to controlling the overall color of the image by adding color keywords. You can apply colors to specific objects or the overall tone. Lighting refers to the description of the lighting in the image. Changing the lighting can have a huge impact on the image.
[0043] In the disclosed embodiment, negative prompt words can be used to remove objects, that is, to remove any content that you do not want to see in the picture. For example, if you enter "beard" in the negative prompt words, the generated character pictures will all be people without beards. Negative prompt words can also be used to modify pictures. For example, when you get a satisfactory character picture, you can use negative prompt words to make fine adjustments. For example, if the hair of the character in the character picture is blown by the wind, enter "wind" in the negative prompt words, and the hair of the character in the picture will no longer appear to be blown by the wind. At this time, nothing is removed, only some minor modifications are made to the subject. Negative prompt words can also be used to switch styles. For example, if you enter "cartoon" in the negative prompt words, the output picture will be more inclined to realism style rather than cartoon style.
[0044] In the disclosed embodiments, the input image can be used to specify the generated object subject or the features, background, etc. of the object subject. For example, by inputting one or more images of a specific object and then using positive or negative prompt words to describe the background, action, or expression you want to generate, you can make the specific object "flash" into the scene described by the positive or negative prompt words, and the action and expression can also be lifelike. For another example, by inputting one or more images of sunglasses and then inputting a teddy bear with positive prompt words, you can output a picture of a teddy bear wearing sunglasses. When no positive or negative prompt words are input and only one input image is input, the image to be captioned is the input image.
[0045] In some exemplary embodiments, obtaining a picture to be accompanied by a text according to the user's first prompt information includes:
[0046] The user's first prompt information is input into the stable diffusion model, so that the stable diffusion model outputs a picture corresponding to the user's first prompt information.
[0047] In the embodiment of the present disclosure, the image generation model may be a stable diffusion model (Stable Diffusion Model), however, the embodiment of the present disclosure is not limited to this.
[0048] To speed up the image generation process, the stable diffusion model does not run the diffusion process on the pixel-level image, but on a compressed version of the image. This compression (and subsequent decompression) is done by the AutoEncoder module.
[0049] The overall framework of the stable diffusion model is shown in Figure 3, which mainly consists of three modules, namely the autoencoder module, the denoising module and the conditional encoder, which are shown in the left, middle and right parts of Figure 3 respectively. Represents the input image, Represents the generated image; For the encoder, For the decoder; is the latent vector; is the potential vector after adding noise; τ θ It is a conditional encoder (which can be Transformer or CLIP) that implements semantic compression; ∈ θ It is the denoising module.
[0050] The autoencoder module contains paired encoders and decoder encoder Input picture Encoded as a latent vector Decoder Decoding latent vectors To best reconstruct the picture, that is, through the encoder Compress the input image into the latent space and pass it through the decoder Reconstruct these compressed data. Given an RGB image encoder Encode it into a continuous latent vector For example, H=W=512, h=w=64, c=4, so the latent vector Than input image It is 48 times smaller, thus significantly improving computational efficiency by performing the denoising process in a compressed, compact latent space.
[0051] Conditional parameters entered It can be text, semantic map, image, representation, etc. In order to preprocess the input under different modalities, a conditional encoder τ is used. θ Conditional parameters to be entered Mapped into an intermediate representation Then the intermediate representation is transformed into Mapped to the denoising module ∈ θ middle.
[0052] Denoising module∈ θ It is the core module of the diffusion image generation process and consists of a time-constrained U-Net structure. By using the cross-attention mechanism to enhance the U-Net structure, the internal diffusion model is transformed into a conditional image generator. The switch in the figure is used to control between different types of conditioning inputs: For text input, the conditional encoder τ is first used. θ These are converted into embeddings (vectors) by a network (e.g., Transformer, CLIP), which are then mapped to the U-Net layer via a (multi-head) Attention (Q, K, V). For other spatially aligned inputs (e.g., semantic maps, images, and inpainting), concatenation can be used for adjustment. The U-Net architecture introduces skip connections, which fuse shallow positional information with deep semantic information, facilitating network optimization.
[0053] When using it, you first need to train the self-encoding module, through the encoder Compress the input image, map the input image from pixel space to latent space, perform diffusion operation on the latent space, learn denoising in the latent space through the U-Net model, combine the input conditioning mechanism, denoise the noise result in the latent space, and finally pass the decoder The generated image can be obtained by restoring it to the original pixel space.
[0054] In some exemplary embodiments, the method further comprises:
[0055] Receive a fourth prompt message from the user, the fourth prompt message including M training pictures of the target subject in the pictures to be captioned and description data corresponding to each picture, where M is a natural number greater than 1;
[0056] The fourth prompt information is used to fine-tune the stable diffusion model to obtain an image generation model that can generate images of different scenes of the target subject.
[0057] In the disclosed embodiment, the DreamBooth method may be used to fine-tune the stable diffusion model.
[0058] In actual use, if you need to generate a scene described by a common prompt word, you can directly use the pre-trained stable diffusion model to generate an image. This can produce an image of average quality, but actual needs are generally more specific or require a beautiful scene. For example, if you want to generate a pancake poster, but the pre-trained stable diffusion model doesn't have the concept of pancakes, then if you input the English translation or Chinese pinyin of "pancake", the model will not output a correct image of a Chinese pancake. The disclosed embodiments can solve this problem by receiving a third prompt from the user before invoking the stable diffusion model and using the third prompt to fine-tune the stable diffusion model.
[0059] During fine-tuning, first, select a pseudo-word for the target subject in the picture to be captioned. This pseudo-word has never appeared in the tokenizer (tokenizer, which plays a very important role in natural language learning tasks. Its main task is to convert text input into input acceptable to the model. Usually the model can only accept numerical input. Therefore, the tokenizer will convert the text input into numerical input). It can be a real word or a non-existent word. For example, the word "Jianbing" does not exist in foreign tokenizers. The pseudo-word corresponding to pancakes can be set to "Jianbing"; then, save M training pictures of the target subject and the description data corresponding to each picture to the corresponding folder. If the target subject is a rigid body with stable shape, prepare no more than 10 training pictures. Because pancakes have different shapes, more training pictures are prepared. Prepare about 20 pictures of common pancakes and give a simple description of each training picture, such as one Jianbing in the plate, save it as a txt file with the corresponding name and place it in the same folder as the training images. Finally, use the Dreambooth method to fine-tune the stable diffusion model to obtain an image generation model that can generate images of different scenes for the target subject. This model can be used with positive prompt words containing "Jianbing" to generate beautiful images of Chinese pancakes.
[0060] Currently, the stable diffusion model only supports English input, but retraining a stable diffusion model that supports Chinese requires significant human resources for Chinese data collation, hardware setup, model training, and optimization. This disclosure supports Chinese text input on the user interface, translates the user-entered Chinese text into English words or pre-set pseudo-words, and then inputs them into the stable diffusion model, achieving this with minimal human resources.
[0061] In some exemplary embodiments, the user's first prompt information is Chinese text and includes the Chinese text of the fine-tuned target subject, and the method further includes:
[0062] The Chinese text other than the target subject in the user's first prompt information is translated into English words, and the Chinese text of the target subject is translated into the pseudo-words selected during fine-tuning.
[0063] In this way, when the user enters positive prompt words such as "pancake" and "advertisement", the image generation model can generate a poster with beautiful pictures of pancakes based on the pancake pictures during fine-tuning.
[0064] In some exemplary embodiments, the first prompt information further includes: the size of the generated image.
[0065] For example, the size of the generated image can be limited by the length and width. For example, the size of the generated image can be set to 300x300 pixels, that is, the image output by the image generation model includes 300 pixels in the length and width directions respectively.
[0066] Generative models are very effective in generating, modifying and smearing images. However, for Chinese, which has strict requirements on spatial structure and stroke order, it is very difficult to directly generate accompanying images using prompts.
[0067] This disclosure provides text, fonts, semantic information, etc. through a second prompt message, uses a character design model to achieve semantic deformation of the text, and then obtains a mask image through an image segmentation algorithm. The mask image is pasted on the image to be accompanied by the text according to the specified position and size (the image to be accompanied by the text can be an existing image or generated through the "text to image" and / or "image to image" method). The required text and image with text can be generated. The entire process only requires inputting the specified parameters to obtain vivid text and image end-to-end.
[0068] In the disclosed embodiment, the second prompt information input by the user includes a text and semantic information corresponding to the text. Semantic information can also be referred to as contextual information. For example, as shown in FIG4A and FIG4B , in this example, the semantic information corresponding to the word "Chinese" is set to "running." Of course, in other examples, the word "Chinese" can also be set with other semantic information, such as "holding a microphone and singing." Semantic information represents a state, or a presentation, rather than a single word.
[0069] In some exemplary embodiments, generating a semantic text-matching image based on the user's second prompt information includes:
[0070] Determine the font of the accompanying text;
[0071] Generate a text character image in the corresponding font according to the text and the font of the text;
[0072] The accompanying text character picture and semantic information corresponding to the accompanying text are input into the character design model so that the character design model outputs a semantic accompanying text picture.
[0073] FIG5A shows an exemplary character design model. As shown in FIG5A, the character design model includes a differentiable rasterizer (DiffVG), an image enhancement (Augment) module, a stable diffusion encoder (Encoder), a semantic encoding (CLIP) module and a noise prediction module (UNet), wherein l i To input characters, in FIG5A , character 1 is input. i Taking S as an example, enter the character l i There are a series of control points after vectorization The semantic word is assumed to be surfing. For deformed characters, For deformed characters The control points of the deformed characters are iteratively optimized. In each round of iteration, the deformed characters are first transformed by using a differentiable rasterizer (DiffVG). Control Points Rasterization converts the text from coordinate mode to image mode, and then enhances it through the image enhancement module. The text is then input into the pre-trained stable diffusion encoder. Based on the feature map of semantic information encoded by the semantic coding (CLIP) module, the feature map output by the encoder and the feature map of semantic information are input into the noise prediction module to control the transformation of the font to a semantic style.
[0074] During training, the overall loss function of the character design model is as follows:
[0075] Among them, the first loss function
[0076] The second loss function
[0077] The third loss function
[0078] The first loss function The difference between the Z-space representation of the character shape (generated by the stable diffusion encoder) and the Z-space representation of the semantic input (output by the Unet model, and the pixel space representation of the semantic input is output by the semantic coding (CLIP) module); the second loss function The third loss function represents the difference between the font style and local character shape of the deformed character and the font style and local character shape of the original character after the original character and the deformed character are respectively passed through the low-pass filter (LPF); Indicates that the coordinates of the original character and the deformed character are respectively passed through the triangulation operator After that, the difference between the shape skeleton of the deformed character and the shape skeleton of the original character is calculated.
[0079] In the formula, α and β t are weight factors, α is a constant between 0 and 1, and α = 0.5 for example. t Related to the number of steps t, in some examples, β t The calculation method is as follows:
[0080] Wherein, a, b and c are all constants, for example, a=100, b=300, c=30, and the number of steps t can be between 50 and 950.
[0081] Where, Represents control points The vector expression in z space, c represents the vector expression of the semantic prompt word (i.e., the text corresponding to the semantic prompt information) in z space. This formula is a concrete representation of the parameter update in the network model. θ represents the parameters of the stable diffusion model, x aug Represents the picture after data enhancement, z aug represents the z space after data enhancement, ∈ represents noise, represents the noise prediction model, y represents the semantic cue word, and w(t) is the t The related constant, α t and σ t Used to control the noise adding mechanism, is the expression after adding noise in z space, It's expectation.
[0082] Currently, the character design model in Figure 5A is developed based on English, which can solve the semantic deformation of single English characters. Through traversing character calls, the same semantic deformation of multiple characters can be achieved. At the same time, the loss function above maintains the game between the font style and the overall skeleton and appearance semantic deformation of the characters, and there is no essential difference for English and Chinese characters. Therefore, based on this character design model, a Chinese font library is provided, and through training, the semanticization of Chinese strokes can be achieved, that is, the strokes are deformed in the direction of similar semantics to achieve the final semanticization. For example, in Figure 4A, the text "中国人" is in the font "阿武魂体(AaWuHunTi, a Chinese font)", the semantic input is "跑(running)", and the generated semantic caption image is shown in Figure 4B.
[0083] The caption in the second prompt message is the word that will be finally shown on the output image, such as "中国人". The semantic information corresponding to the caption can be represented by a semantic prompt word, such as "跑(running)". The semantic information corresponding to the caption is used to guide the deformation direction of the handwriting. The semantic prompt word obtains the feature corresponding to the semantic information through the CLIP model (the upper right corner of Figure 5A), and then the feature corresponding to the semantic information and the feature output by the StableDiffusion encoder are input into the UNet model to iteratively predict noise and denoise, generating a font in semantic style.
[0084] CLIP (Contrastive Language-Image Pre-Training) is a deep learning model, which is a pre-training model that can process text and images simultaneously. Different from previous image classification models, CLIP does not use a large-scale labeled image dataset for training, but pre-trains from unlabeled image and text data through self-supervised learning, enabling the model to understand the semantic connection between images and text. The input of the CLIP model can be text, or an image, or both.
[0085] In some exemplary embodiments, the second prompt message of the user further includes a font, and the font in the second prompt message of the user is used to determine the font of the caption. Of course, the font of the caption can also be preset by the server side, and the embodiments of the present disclosure do not limit this.
[0086] Exemplarily, the font can be any type of font such as boldface, regular script, etc. The font determines the shape of the caption before deformation.
[0087] In some exemplary embodiments, the font is a Chinese font, the caption is Chinese text, and the semantic information corresponding to the caption is Chinese text or English text.
[0088] The input Chinese caption can be directly converted into a character image through the corresponding Chinese font library. When the semantic information corresponding to the caption is input through Chinese text, the Chinese text of the semantic information is translated into English text, and then the character image and the English text of the semantic information are input into the character design model to obtain a deformed caption image.
[0089] In the disclosed embodiments, when a caption image includes multiple characters, deformation may be performed on only one or more of the characters, or on all characters. In the disclosed embodiments, a character refers to a basic symbol representing data in the form of a graphic arranged in space using connected or adjacent strokes, such as letters, numbers, Chinese characters, operation symbols, punctuation marks, and the like.
[0090] In other exemplary embodiments, the third prompt information can also be obtained directly based on the picture to be accompanied by the text. The third prompt information includes the text and the semantic information corresponding to the text. In this case, the user does not need to input the second prompt information, and a semantic text picture is generated according to the third prompt information.
[0091] In some exemplary embodiments, obtaining the third prompt information according to the image to be accompanied by a text includes:
[0092] Inputting the image to be captioned into the image-text generation model so that the image-text generation model outputs first text information and second text information corresponding to the image to be captioned, wherein the first text information includes the content of the image to be captioned, and the second text information includes the content or features of the image to be captioned;
[0093] Segmenting the first text information to obtain one or more first phrases and a part of speech corresponding to each first phrase; segmenting the second text information to obtain one or more second phrases and a part of speech corresponding to each second phrase;
[0094] One or more first phrases are selected as accompanying text according to the part of speech; and one or more second phrases are selected as semantic information corresponding to the accompanying text according to the part of speech.
[0095] In an embodiment of the present disclosure, the image-text generation model can be a multimodal model. The image-text generation model can obtain the first text information and the second text information based on multiple pre-set prompt questions. For example, the pre-set prompt questions can be "What is the content of this picture?" or "What style is this picture?", etc. The first text information obtained can be "orange juice, eggs, sandwiches", and the second text information can be "cartoon". After word segmentation, one or more proper nouns "orange juice", "eggs", and "sandwiches" can be selected as captions based on part of speech, and one or more proper nouns "cartoon" can be selected as semantic information corresponding to the captions based on part of speech. In an embodiment of the present disclosure, the prompt questions corresponding to the first text information are mainly questions about the content of the picture, and the prompt questions corresponding to the second text information can be questions about the characteristics, color, shape of the main content of the picture, etc. In other examples, the pre-set prompt questions can also include "What color is this picture?", "What shape is xxx (the main body of the image) in this picture?", etc., and the semantic information corresponding to the caption is generated based on the answers to these prompt questions. After obtaining the caption and the semantic information corresponding to the caption, the caption image can be generated according to the method described above, and the caption image and the picture to be captioned can be combined into a captioned picture.
[0096] Figure 5B is a schematic diagram of the structure of an image-text generation model according to an exemplary embodiment of the present disclosure. As shown in Figure 5B, the image-text generation model according to an exemplary embodiment of the present disclosure may include four parts: a first encoder, a second encoder, a third encoder, and a first decoder. The first encoder is used to encode an input image. The first encoder decomposes the input image into multiple image blocks and encodes each image block into an embedding sequence to better process and extract local features.
[0097] The second encoder is used to encode the text. The second encoder can adopt a Transformer structure and includes multiple second encoding blocks. Each second encoding block includes a bidirectional self-attention layer and a feedforward network layer. During training, the second encoder is activated by calculating the image-text contrast loss. The image-text contrast loss is used to align the feature spaces of the first and second encoders by encouraging positive image-text pairs to have similar representations. Both the first and second encoders are single-mode encoders.
[0098] The third encoder is used to encode text based on the image. It can adopt a variant of the Transformer structure and include multiple third encoder blocks. Each third encoder block includes a bidirectional self-attention layer and a feedforward network layer. It also includes a cross-attention layer between the bidirectional self-attention layer and the feedforward network layer. Visual information is injected through the cross-attention layer. The second and third encoders construct a representation of the current input token through the bidirectional self-attention layer. During training, the third encoder is activated by calculating the image-text matching loss. The image-text matching loss is used to learn multimodal representations of image and text, capturing fine-grained alignment between vision and language. The image-text matching loss is a binary classification task that predicts whether an image-text pair is a positive or negative match based on multimodal features.
[0099] The first decoder is used to decode text based on the image (i.e., generate a text description of a given image). The first decoder includes multiple decoding blocks, each of which includes a causal self-attention layer and a feedforward network layer, as well as a cross-attention layer between the causal self-attention layer and the feedforward network layer. The first decoder predicts the next token by using the causal self-attention layer. During training, the first decoder is activated by calculating the language modeling loss, thereby training the model to maximize the likelihood of text in an autoregressive manner.
[0100] The first encoder, the second encoder, the third encoder, and the first decoder share all parameters of other layers except the self-attention layer. By sharing all parameters of these layers, the training efficiency can be improved while benefiting from multi-task learning.
[0101] When the accompanying text image includes multiple characters, the accompanying text image can be split into multiple accompanying text images of single characters. In this way, the position of the accompanying text image of each character can be designed separately, making the output image more exquisite.
[0102] In some exemplary embodiments, when the accompanying text character includes multiple characters, inputting the accompanying text character and semantic information corresponding to the accompanying text into the character design model includes:
[0103] Select at least some of the multiple characters (only the selected characters are deformed, and other characters are not deformed);
[0104] When at least part of the selected characters includes only a single character, inputting the selected single character into the character design model so that the character design model outputs a semantically-defined caption image corresponding to the single character;
[0105] When at least part of the selected characters include multiple characters, the selected multiple characters are input into the character design model one by one, so that the character design model outputs the semantic text picture corresponding to each character; or, all of the selected multiple characters are input into the character design model, so that the character design model outputs the semantic text pictures corresponding to all the characters, and the semantic text pictures corresponding to all the characters are cropped to obtain the semantic text pictures corresponding to a single character.
[0106] After obtaining a semantic text image, simply using a horizontal layout may not be enough to meet the needs. In the disclosed embodiment, by obtaining a semantic text image corresponding to a single character, it is convenient to design the position of the semantic text image corresponding to each character when posting the image later, avoiding a single, fixed layout method.
[0107] In some exemplary embodiments, combining the image to be accompanied with text and the semantically modified image with text into the image with text includes: pasting the semantically modified image with text onto the image to be accompanied with text to obtain the image with text.
[0108] In reality, graphic and text works have various layouts. To meet this requirement, the scaling ratio of the text and the x and y coordinates of the upper left corner of each text character can be passed in when the user inputs. After the semantic text image is segmented by the segmentation algorithm, the images can be pasted in order on the generated image to obtain the desired image effect.
[0109] In some exemplary embodiments, the user's second prompt information further includes pasting location information, and pasting the semantically-defined caption image onto the image to be captioned includes:
[0110] According to the paste position information, paste the semantic caption image onto the image to be captioned.
[0111] When the accompanying text includes multiple characters, the pasting position information may include (x, y) coordinate information of the upper left corner point of each accompanying text image.
[0112] In some exemplary embodiments, the user's second prompt information also includes a zoom ratio. Before pasting the semantic text image onto the image to be captioned, the method further includes: scaling the semantic text image according to the zoom ratio.
[0113] In some exemplary embodiments, as shown in FIG6A and FIG6B , before pasting the semantically-defined caption image onto the image to be captioned, the method further includes:
[0114] Perform image segmentation on the semantic text image to obtain the mask image corresponding to the semantic text image.
[0115] As shown in Figure 4B, the text picture is a picture with black text on a white background. If it is directly pasted onto the picture to be texted, the white background will cover up part of the information on the picture to be texted. Therefore, before pasting, the semantic text picture is first segmented to obtain a mask picture corresponding to the semantic text picture (the mask picture is a binary picture and can be directly obtained from the picture segmentation algorithm). Then, the mask picture is pasted to prevent the white background from covering up the information on the picture to be texted. For example, the semantic text picture can be segmented by image segmentation algorithms such as U-Net, SegNet, and HRNet, and this disclosure does not limit this.
[0116] As shown in Figure 7, the user enters the first prompt information in the client interface, such as positive and / or negative prompt words (prompt), the expected image size (size), the input preliminary image (input image), etc. These prompt information are used by the server algorithm to control the image generation model to generate the image to be matched with the text. The user also enters the second prompt information such as the matching text (word), font (font), semantic information (sementics), word scaling ratio (scale ratio), and the x, y coordinate sequences (coordinate sequences) of the characters in the matching text in the upper left corner of the picture in the client interface. These prompt information are used by the server algorithm to control the semantic generation and pasting effect of the matching text picture. When all these prompt information are transmitted to the server, the server algorithm will generate the corresponding matching text and transmit the generated matching text back to the client interface (click the Generate Image button in Figure 2) for display. The client interface can be an interactive web interface, however, the embodiment of the present disclosure is not limited to this. For example, still taking the aforementioned pancake as an example, assuming that the positive prompt words entered by the user are "pancake" and "advertisement", the accompanying text entered by the user is pancake, and the semantic information is "fire", the output picture with the accompanying text is shown in Figure 8.
[0117] The present disclosure proposes a method for designing images and texts. The method includes inputting a first prompt message into an image generation model to obtain an image to be accompanied by a text; without the aid of editing software, the method provides a text, a text font, and the semantic information that the text is intended to express, or directly generates a semantic text-based image based on the image to be accompanied by the text, thereby adding interest and liveliness; and the image to be accompanied by the text and the text image are pasted and synthesized to obtain an image with text and text. The images generated by this method are similar in shape, vivid, and interesting, and can achieve end-to-end image generation and text matching. This method has great practical value, such as promotional posters, product posters, ancient poetry paintings, picture book generation, etc. It can help solve many real-life painting and text matching problems and generate great commercial value.
[0118] An embodiment of the present disclosure also provides a graphic design device, comprising a memory; and a processor connected to the memory, wherein the memory is used to store instructions, and the processor is configured to execute the steps of the graphic design method described in any embodiment of the present disclosure based on the instructions stored in the memory.
[0119] As shown in FIG9 , in one example, a graphic design device may include: a processor 910, a memory 920, a bus system 930, and a transceiver 940, wherein the processor 910, the memory 920, and the transceiver 940 are connected via the bus system 930, the memory 920 is used to store instructions, and the processor 910 is used to execute the instructions stored in the memory 920 to control the transceiver 940 to send and receive signals. Specifically, under the control of the processor 910, the transceiver 940 may receive a first prompt message input by a user, or a first prompt message and a second prompt message input by a user, and the processor 910 obtains a picture to be accompanied by a text based on the first prompt message; receives a second prompt message input by a user or obtains a third prompt message based on the picture to be accompanied by a text, generates a semantically-defined picture to be accompanied by a text based on the second prompt message or the third prompt message, and combines the picture to be accompanied by a text and the semantically-defined picture to be accompanied by a text into a picture to be accompanied by a text.
[0120] It should be understood that the processor 910 may be a central processing unit (CPU), or may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc.
[0121] The memory 920 may include a read-only memory and a random access memory, and provides instructions and data to the processor 910. A portion of the memory 920 may also include a non-volatile random access memory. For example, the memory 920 may also store information about the device type.
[0122] In addition to the data bus, the bus system 930 may also include a power bus, a control bus, a status signal bus, etc. However, for the sake of clarity, various buses are labeled as the bus system 930 in FIG.
[0123] During implementation, the processing performed by the processing device can be completed by the hardware integrated logic circuit in the processor 910 or by instructions in the form of software. That is, the method steps of the embodiment of the present disclosure can be embodied as being executed by a hardware processor, or being executed by a combination of hardware and software modules in the processor. The software module can be located in a storage medium such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 920, and the processor 910 reads the information in the memory 920 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0124] As shown in FIG10 , the embodiment of the present disclosure further provides a graphic design system, including a client 1001 and a server 1002 , wherein:
[0125] The client 1001 is configured to receive a first prompt message input by a user and upload the first prompt message to the server 1002, or receive a first prompt message and a second prompt message input by a user and upload the first prompt message and the second prompt message to the server; receive and display a picture with text sent by the server 1002;
[0126] Server 1002 is configured to obtain a picture to be accompanied by a text based on the first prompt information, generate a semantically accompanied picture based on the second prompt information, or obtain a third prompt information based on the picture to be accompanied by a text, generate a semantically accompanied picture based on the third prompt information, combine the picture to be accompanied by a text and the semantically accompanied picture into a picture with a text, and send the picture with a text to client 1001.
[0127] In some exemplary embodiments, the client 1001 may provide an interactive web page interface, and receive the user's first prompt information, or the first prompt information and the second prompt information, through the web page interface.
[0128] Fig. 2 is a schematic diagram of an exemplary web page interface provided by the client 1001. Exemplarily, as shown in Fig. 2, the first prompt information may include at least one of the following: a positive prompt word, a negative prompt word, an input picture, and the like.
[0129] In some exemplary embodiments, the first prompt information may further include: the size of the generated image.
[0130] In some exemplary embodiments, obtaining a picture to be accompanied by a text according to the first prompt information includes:
[0131] The first prompt information is input into the stable diffusion model, so that the stable diffusion model outputs a picture corresponding to the first prompt information of the user.
[0132] In some exemplary embodiments, a stable diffusion model includes: an autoencoder module, a denoising module, and a conditional encoder;
[0133] The autoencoder module consists of a paired encoder and decoder. The encoder compresses the input image into a latent space, and the decoder reconstructs the denoised data.
[0134] The conditional encoder maps the input conditional parameters into an intermediate representation, and maps the intermediate representation to the denoising module through the cross-attention layer;
[0135] The denoising module performs a diffusion operation on the compressed data in the latent space and performs denoising in the latent space in combination with the conditional mechanism output by the conditional encoder.
[0136] In some exemplary embodiments, the client 1001 is further configured to:
[0137] Receive a third prompt message from the user, the third prompt message including M training pictures of the target subject in the pictures to be captioned and description data corresponding to each picture, where M is a natural number greater than 1, and upload the third prompt message to the server 1002;
[0138] The server 1002 is also configured to:
[0139] The third prompt information is used to fine-tune the stable diffusion model to obtain an image generation model that can generate images of different scenes of the target subject.
[0140] In some exemplary embodiments, the user's first prompt information is in Chinese text and includes the Chinese text of the fine-tuned target subject. The server 1002 is further configured to:
[0141] The Chinese text other than the target subject in the first prompt information is translated into English words, and the Chinese text of the target subject is translated into the pseudo-words selected during fine-tuning.
[0142] In some exemplary embodiments, the second prompt information includes a text and semantic information corresponding to the text, and generating a semantic text image based on the second prompt information includes:
[0143] Determine the font of the accompanying text;
[0144] Generate a text character image in the corresponding font according to the text and the font of the text;
[0145] The accompanying text character picture and semantic information corresponding to the accompanying text are input into the character design model so that the character design model outputs a semantic accompanying text picture.
[0146] In some exemplary embodiments, when the accompanying text includes multiple characters, inputting the accompanying text character image and semantic information corresponding to the accompanying text into the character design model includes:
[0147] selecting at least some of the plurality of characters;
[0148] When at least part of the selected characters includes only a single character, inputting the character image and semantic information corresponding to the accompanying text of the single character into the character design model so that the character design model outputs a semantic accompanying text image corresponding to the single character;
[0149] When at least part of the selected characters include multiple characters, the selected multiple characters and the semantic information corresponding to the text are input into the character design model one by one, so that the character design model outputs a semantic text picture corresponding to each character; or, all the characters in the selected multiple characters and the semantic information corresponding to the text are input into the character design model, so that the character design model outputs semantic text pictures corresponding to all the characters, and the semantic text pictures corresponding to all the characters are cropped to obtain a semantic text picture corresponding to a single character.
[0150] In some exemplary embodiments, the font of the accompanying text is Chinese font, the accompanying text is Chinese text, and the semantic information corresponding to the accompanying text is Chinese text or English text.
[0151] In some exemplary embodiments, combining the image to be accompanied with text and the semantically modified image with text into the image with text includes: pasting the semantically modified image with text onto the image to be accompanied with text to obtain the image with text.
[0152] In some exemplary embodiments, the user's second prompt information further includes pasting location information, and pasting the semantically-defined caption image onto the image to be captioned includes:
[0153] According to the paste position information, paste the semantic caption image onto the image to be captioned.
[0154] In some exemplary embodiments, the user's second prompt information also includes a zoom ratio. Before pasting the semantic text image onto the image to be captioned, the server 1002 is further configured to: zoom the semantic text image according to the zoom ratio.
[0155] In some exemplary embodiments, before pasting the semantic text-matching image onto the image to be matched with text, the server 1002 is further configured to:
[0156] Perform image segmentation on the semantic text image to obtain the mask image corresponding to the semantic text image.
[0157] The present disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the graphic design method described in any of the embodiments of the present disclosure. The method for driving interactive reading by executing executable instructions is substantially the same as the graphic design method provided in the above embodiments of the present disclosure and is not further described here.
[0158] In some possible implementations, various aspects of the graphic design method provided by the present disclosure may also be implemented in the form of a program product, which includes program code. When the program product is run on a computer device, the program code is used to enable the computer device to execute the steps of the graphic design method according to the various exemplary implementations of the present disclosure described above in this specification. For example, the computer device may execute the graphic design method recorded in the embodiments of the present disclosure.
[0159] The program product may employ any combination of one or more readable media. The readable medium may be a readable signal medium or a readable storage medium. The readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or component, or any combination thereof. More specific examples (a non-exhaustive list) of readable storage media include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0160] It will be appreciated by those skilled in the art that all or some of the steps, systems, and functional modules / units in the methods disclosed above may be implemented as software, firmware, hardware, and appropriate combinations thereof. In hardware implementations, the division between the functional modules / units mentioned in the above description does not necessarily correspond to the division of physical components; for example, a physical component may have multiple functions, or a function or step may be performed by several physical components in cooperation. Some or all components may be implemented as software executed by a processor, such as a digital signal processor or a microprocessor, or implemented as hardware, or implemented as an integrated circuit, such as an application-specific integrated circuit. Such software may be distributed on a computer-readable medium, which may include a computer storage medium (or non-transitory medium) and a communication medium (or temporary medium). As is well known to those skilled in the art, the term computer storage medium includes volatile and non-volatile, removable, and non-removable media implemented in any method or technology for storing information (such as computer-readable instructions, data structures, program modules, or other data). Computer storage media include, but are not limited to, RAM, ROM, EEPROM, flash memory or other memory technology, CD-ROM, digital versatile disks (DVD) or other optical disk storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store the desired information and can be accessed by a computer. In addition, it is well known to those skilled in the art that communication media generally embodies computer-readable instructions, data structures, program modules, or other data in a modulated data signal such as a carrier wave or other transport mechanism, and may include any information delivery media.
[0161] It should be noted that the above-described embodiments or implementations are merely illustrative and not restrictive. Therefore, the present disclosure is not limited to what is specifically shown and described herein. Various modifications, substitutions, or omissions may be made to the forms and details of the implementations without departing from the scope of the present disclosure.
Claims
1. A graphic design method, comprising: Receiving a first prompt message input by a user, and obtaining an image to be captioned according to the first prompt message; Receiving a second prompt message input by the user or obtaining a third prompt message according to the image to be captioned, and generating a semantic caption image according to the second prompt message or the third prompt message; Combining the image to be captioned with the semantic caption image into a captioned image.
2. The graphic design method according to claim 1, wherein The second prompt message includes a caption and semantic information corresponding to the caption. The generating a semantic caption image according to the second prompt message includes: Determining the font of the caption; Generating a caption character image in the corresponding font according to the caption and the font of the caption; Inputting the caption character image and the semantic information corresponding to the caption into a character design model, so that the character design model outputs the semantic caption image.
3. The graphic design method according to claim 2, wherein, The loss function of the character design model includes a first loss function, a second loss function, and a third loss function. The first loss function represents the difference between the latent space expression features of characters and the latent space expression features of semantic information. The second loss function represents the difference in font style and local character shape between the original character and the deformed character after passing through a low-pass filter respectively. The third loss function represents the difference in the outer shape skeleton between the coordinates of the original character and the coordinates of the deformed character after passing through a triangular operator respectively. The deformed character is the semantic caption image.
4. The graphic design method according to claim 1, wherein, The third prompt message includes a caption and semantic information corresponding to the caption. The obtaining a third prompt message according to the image to be captioned includes: Inputting the image to be captioned into a graphic-text generation model, so that the graphic-text generation model outputs first text information and second text information corresponding to the image to be captioned. The first text information includes the content of the image to be captioned, and the second text information includes the content or features of the image to be captioned; Performing word segmentation on the first text information to obtain one or more first word groups and the part of speech corresponding to each first word group; performing word segmentation on the second text information to obtain one or more second word groups and the part of speech corresponding to each second word group; Selecting one or more of the first word groups as the caption according to the part of speech; selecting one or more of the second word groups as the semantic information corresponding to the caption according to the part of speech.
5. The graphic design method according to claim 2 or 4, wherein, When the caption includes multiple characters, the inputting the caption character image and the semantic information corresponding to the caption into the character design model includes: Selecting at least some of the multiple characters; When the selected at least some characters include only a single character, inputting the character image of the single character and the semantic information corresponding to the caption into the character design model, so that the character design model outputs the semantic caption image corresponding to the single character; When the selected at least some characters include multiple characters, inputting the selected multiple characters one by one and the semantic information corresponding to the caption into the character design model, so that the character design model outputs the corresponding one for each character Semantic caption image; or, input all the characters among the selected multiple characters and the semantic information corresponding to the caption into a character design model, so that the character design model outputs a semantic caption image corresponding to all the characters, and crop the semantic caption image corresponding to all the characters to obtain a semantic caption image corresponding to a single character.
6. The graphic design method according to claim 2 or 4, wherein The font of the caption is a Chinese font, the caption is a Chinese text, and the semantic information corresponding to the caption is a Chinese text or an English text.
7. The graphic design method according to claim 1, wherein The step of synthesizing the image to be captioned with the semantic caption image into a captioned image includes: pasting the semantic caption image onto the image to be captioned to obtain the captioned image.
8. The graphic design method according to claim 7, wherein, The second prompt information further includes paste position information, and the step of pasting the semantic caption image onto the image to be captioned includes: Pasting the semantic caption image onto the image to be captioned according to the paste position information.
9. The graphic design method according to claim 7, wherein, The second prompt information further includes a scaling ratio. Before pasting the semantic caption image onto the image to be captioned, the method further includes: scaling the semantic caption image according to the scaling ratio.
10. The graphic design method according to claim 7, wherein, Before pasting the semantic caption image onto the image to be captioned, the method further includes: Performing image segmentation on the semantic caption image to obtain a mask image corresponding to the semantic caption image.
11. The graphic design method according to claim 1, wherein, The first prompt information includes at least one of the following: positive prompt words, negative prompt words, and input images.
12. The graphic design method according to claim 11, wherein, The step of obtaining the image to be captioned according to the first prompt information includes: Inputting the first prompt information into a StableDiffusion model, so that the StableDiffusion model outputs an image corresponding to the first prompt information.
13. The graphic design method according to claim 12, wherein, The StableDiffusion model includes: an autoencoder module, a denoising module, and a conditional encoder; The autoencoder module includes a paired encoder and decoder. The encoder compresses the input image into the latent space, and the decoder reconstructs the denoised data. The conditional encoder maps the input conditional parameters into an intermediate representation and maps the intermediate representation into the denoising module through a cross-attention layer. The denoising module performs a diffusion operation on the compressed data in the latent space and denoises in the latent space by combining the conditional mechanism output by the conditional encoder.
14. The graphic design method according to claim 12, the method further includes: Receiving a fourth prompt information from the user, the fourth prompt information includes training images of the target subject in M images to be captioned and description data corresponding to each image, where M is a natural number greater than 1; Fine-tuning the StableDiffusion model using the fourth prompt information to obtain an image generation model that can generate different scenes of the target subject's images.
15. The graphic design method according to claim 14, wherein, The step of fine-tuning the StableDiffusion model using the fourth prompt information includes: Selecting a pseudo-word for the target subject, and the pseudo-word does not pre-exist in the tokenizer; Placing the M training images of the target subject and the description data corresponding to each image in the same folder; Fine-tuning the StableDiffusion model using the pseudo-word and the training images.
16. The graphic design method according to claim 15, wherein The first prompt message of the user is in Chinese text and includes the Chinese text of the target subject to be fine-tuned. The method further includes: Translating the Chinese text other than the target subject in the first prompt message into English words, and translating the Chinese text of the target subject into the pseudo-word selected during fine-tuning.
17. A graphic design device, comprising a memory; and a processor connected to the memory, the memory being configured to store instructions, and the processor being configured to execute the steps of the graphic design method according to any one of claims 1 to 16 based on the instructions stored in the memory.
18. A computer-readable storage medium, having stored thereon a computer program, which when executed by a processor implements the graphic design method according to any one of claims 1 to 16.
19. A graphic design system, comprising a client and a server, wherein: The client is configured to receive the first prompt message input by the user and upload the first prompt message to the server, or receive the first prompt message and the second prompt message input by the user and upload the first prompt message and the second prompt message to the server; Receive the picture with captions and display it; The server is configured to obtain the picture to be captioned according to the first prompt message, generate a semantic caption picture according to the second prompt message or obtain a third prompt message according to the picture to be captioned, generate a semantic caption picture according to the third prompt message, synthesize the picture to be captioned and the semantic caption picture into a picture with captions, and send the picture with captions to the client.
Citation Information
Patent Citations
Automatic text matching method, device thereof and computer storage medium
CN109167939A
Image text matching method and device, terminal and computer readable storage medium
CN111383302A
Text image generation method and diffusion generation model training method
CN116797868A
Figure graph model training method and text graph method
CN116935169A
Image generation method and device, image model construction method and device, equipment and storage medium
CN117037179A
Cited By
Data chart illustration method, system and equipment based on AIGC
CN121074201A